Skip to content

Attest: product goal and approval plan ​

Status: approved on 24 September 2026; implementation in progress. This document plans work on the attest branches of CDX and Flutter ELOG. CDX is where the complete Attest product is built. ELOG and SSO are real products used to test Attest end to end. Workflow changes still require separate review.

Product goal ​

Attest is a product for testing Flutter applications across development, operations and release qualification. Its studio is the interactive surface; the CLI, API, execution agents and testing farm are equally important parts of the product. A new customer should be able to connect a Flutter application, define a meaningful API check or UI journey, run it in a safe environment, and understand the result and its evidence without learning the CDX or Caliber internals. Developers retain a versioned, reviewable configuration path for advanced cases. ELOG and SSO are the first reference products, not the defaults embedded in the platform.

Attest must work locally and be architected for a hosted control plane with remote execution agents. The same case definitions, execution protocol and evidence model should survive that transition. Scheduled runs, sharding, streaming results, private network access and a worker farm are product capabilities to build in stages, not UI-only promises.

The first complete customer journey is: install Attest, add a product, choose an existing environment or create an isolated local runtime, author one check, validate it, run it, diagnose a failure, and export the result. The studio must show what is configured, what is executable, what is only planned, and what evidence was actually retained.

End-to-end proof runs this journey twice: once with an independent sample Flutter product that has no Caliber code, and once against the actual ELOG/SSO adapters. A green unit suite alone is not product acceptance.

Current evidence and boundaries ​

  • CDX owns the CLI, contracts, dashboard, runners, orchestration, evidence and reports. ELOG owns ELOG/SSO catalogs, scenarios, datasets, seeds and runtime adapters. Preserve that boundary and existing plan/checkpoint IDs.
  • The CLI creates a generic starter workspace, but registry-only installation is still unproven because cdx_attest_cli does not resolve from the private registry. A starter with no real tests is correctly marked unavailable.
  • The dashboard already has Overview, Catalog, Runs, Trends and Configuration. Its recorded UX gaps include a single guided run path, addressable routes, exact configuration guidance, and direct links from failed gates to evidence.
  • The current local managed runtime is Supabase-specific. ELOG's adapter owns product-specific deployment and seed logic. Treat broader runtime support as a designed extension, not an existing promise.
  • UI execution uses Patrol; API/action checks use Dart; performance uses k6. Playwright is an external black-box test of the Attest web application in this plan. It does not silently replace the Flutter UI execution engine.

What current products teach us ​

ProductUseful pattern for AttestAttest decision
Playwright UI Mode and codegenPick an individual test, inspect each step, and see before/after evidence; recording can help author tests.Build a step-level result and evidence view. Prototype authoring from interaction only after the declarative step model is reliable.
Cypress StudioRecorded actions and assertions are reviewable and editable.Never hide executable behavior behind an opaque recording.
Postman request builder and API testsA request can be tried immediately, asserted, and grouped into a reusable suite.Provide a request/check editor with a response preview, then save a versioned case.
k6 Studio and thresholdsRecording lowers setup cost, while explicit thresholds decide performance success.Start with workload presets and visible thresholds; show measured values and units.
Sentry issue detailA failure is connected to logs, trace context, replay and attachments on one investigation page.Make each failed run an evidence-backed investigation, with the failing step and related artifacts in one place.
BrowserStack Flutter execution and debugging evidenceUploaded app/test builds run on selected devices in parallel, with session logs and video available afterward.Separate scheduling from workers; show device/runtime identity and replay artifacts for every shard.

These are interaction references, not a plan to copy their engines or claim their feature breadth.

Proposed product model ​

  1. Product: name, branding, source revision, app launch/build instructions, API base URLs, and capabilities. No ELOG-specific assumptions.
  2. Environment: an existing deployment, an attached local stack, or an isolated managed runtime. Its adapter declares prerequisites, health probes, data setup and teardown. The studio previews every mutating preparation step.
  3. Case: a versioned API, integration/action, Flutter UI, or performance definition with stable ID, owner, tags, preconditions, steps and assertions. Advanced code-backed cases remain first-class.
  4. Plan: selected cases, environment eligibility, data strategy and explicit acceptance gates. A passing case does not imply full qualification.
  5. Execution: immutable snapshot of source/config/data versions, selected scope, measurements, artifacts, failure reason and storage status. Retries link to an earlier execution instead of overwriting it.
  6. Shard: a bounded subset of an execution assigned to one compatible worker, with its own attempt, lease, result and evidence. Aggregation must preserve failed and missing shards rather than reporting a partial pass.
  7. Worker agent: a registered execution process with declared runners, device/runtime capacity and network reachability. It receives scoped work, streams events and uploads artifacts; it does not own acceptance rules.
  8. Authoring assistant: an optional AI or skill-driven helper that proposes cases and configuration. Its draft passes the same validation as manual edits and requires a human review before changing files or running tests.

Keep a typed, versioned schema as the source of truth. The studio edits that schema and shows a reviewable diff; it does not create a second hidden database of test definitions. Credentials are references to a secret provider or local session input, never committed in cases or embedded in report exports.

Product architecture and separation of concerns ​

ConcernOwner in CDXBoundary
AuthoringStudio editor, CLI and authoring skillProduces reviewable case/config drafts; cannot claim a run passed
Product contractTyped versioned schemas and validatorsIndependent of a specific UI, AI provider or execution location
Control planeAPI, scheduler, execution ledger, tenant permissionsPlans and assigns work; never embeds ELOG seeds or a runner implementation
Execution planeWorker agent and runner adaptersExecutes assigned shards in isolated runtimes and reports observed facts
Runtime providersLocal container, attached environment, later cloud/device providersOwn preparation, health, teardown and capability declaration
Evidence planeEvent stream, screenshots/video/logs, artifact storage and retentionImmutable link to execution/shard/step with redaction and access policy
PresentationDashboard, replay and reportsReads the same execution record; does not recalculate qualification differently
Product integrationELOG/SSO or a customer-owned adapterOwns app build, credentials, data, domain actions and acceptance meaning

Local mode can run the control plane and one worker in a single installation. Hosted mode separates them through the same authenticated API and durable work protocol. Worker registration needs capability matching, heartbeat, leased work, idempotent completion, cancellation and recovery from worker loss. The farm needs per-tenant quotas and concurrency limits so a k6 load run cannot starve UI devices or API checks. Cloud readiness includes tenant isolation, role-based access, secret storage, encrypted transport, artifact retention and private network connectivity; these are architecture and acceptance work, not claims that the current application already provides them.

Scheduled runs are a control-plane trigger for a versioned plan and environment. The eventual schedule mechanism must be reviewed separately before any workflow file or automatic trigger is written. Manual and API-triggered execution come first and exercise the same planner and ledger.

Authoring experience ​

First run. A guided setup shows five concrete checks: product launch, test environment, test data, first case, and a dry run. Each check has a working example and a validation action. Users can inspect a starter without claiming that it qualifies their product.

API case. Choose method and path, select an environment, add headers/body and secret references, send a sample request, then add assertions for status, headers, JSON fields, schema and timing. A case can extract a typed value for the next request. The editor previews the exact request and sanitized response.

Integration case. Compose a sequence across API calls or registered product actions with setup, observable assertions, and cleanup. Label this surface by behavior and dependencies so users understand how it differs from one API request. Avoid arbitrary shell execution in a customer-facing editor.

Flutter UI journey. Start with a named app state and test identity, then add visible steps such as launch, tap a semantic target, enter text, wait for a condition, assert text/state, and capture evidence. Show target validation and the executable step order before saving. A code-backed Patrol escape hatch handles gestures and native integrations the guided vocabulary cannot express. Recording is a later prototype, not a prerequisite for the first useful editor.

Performance case. Choose API workload or browser journey, load profile, duration and data source; set explicit k6 thresholds for latency, errors and optional domain metrics. A run shows the measured values beside thresholds and the exact tested environment.

Attest authoring skill and AI-assisted setup ​

Create a product-owned attest-authoring skill in CDX after the case schema is stable, with a distributable source under packages/attest/skills/attest-authoring/. Document how a customer installs or invokes it alongside the CLI. It should ask for the application entry points, environment boundaries, actors, data setup, intended behavior and acceptance evidence; inspect an allowed product repository; then generate versioned YAML/configuration and code-backed hooks only where required. Its output is a change set with reasons, validation results and exact run commands. The customer can review or edit it, then run attest validate and a dry run before executing the real plan. A fork/clone of a product repo should follow this same path, without copying Caliber-specific data into the generated files.

The authoring skill must be useful without a hosted model. A later provider interface can accept customer-supplied OpenAI, Gemini or other model credentials for suggestions. Keep model calls behind a provider adapter, allow opt-out, redact source and secrets according to policy, display what content would be sent, and never let generated YAML bypass the typed validator or execution permissions. Distinguish this authoring assistant from the worker agent that executes tests on the farm.

Product experience and information architecture ​

AreaUser's immediate questionMain action
HomeWhat is ready, blocked or failing?Continue setup or run
CasesWhat do we test and how do I add a case?Create/try/validate
PlansWhich cases count for this goal?Review scope
RunsWhat happened in a specific execution?Inspect/retry/export
ReplayWhat did the tester see before failure?Scrub synchronized video and steps
EnvironmentsWhat will Attest start or modify?Validate/prepare
FarmWhich workers/devices are available and where is work queued?Inspect capacity and shards
SchedulesWhich approved plans will run automatically?Review a schedule proposal
SettingsHow is this product connected?Edit and review config

The run flow is select scope → check readiness → preview preparation → run → inspect evidence. Navigation keeps product, environment and execution identity visible. Failed gates link directly to their cause and retained evidence. The first viewport favors the next useful action over summary cards. Mobile width remains readable for monitoring and triage; dense authoring may use a larger screen with an explicit explanation.

Use Sentry as a maturity reference for dense but legible investigation: a run detail should lead with outcome, failing step, environment and time, then offer logs, network evidence, screenshots, replay, traces and report artifacts in a coherent timeline. Use BrowserStack as a reference for worker/device identity and session evidence. Video replay is only shown when a recording exists; align video time with step events and logs, show gaps explicitly, and make recording retention, masking and access visible. A screenshot sequence remains useful when video is unavailable. No fabricated live charts or replay placeholders.

Before creating any new shared CDX UI primitive, search broadly across vyuh_cdx_ui, vyuh_table, vyuh_timelines, the entity system and existing editor surfaces. Record the candidates and why they fit or fail. Reuse shared Button, DialogShell, CdxSegmentedTabs, SectionCard, EmptyState, table and timeline primitives where appropriate. Add a shared component only when it has a reusable contract beyond Attest; keep test-specific step/replay widgets inside Attest. This protects the separation between CDX design system and product behavior.

The visual work is an inspectable loop: render the live Attest application, capture desktop and narrow snapshots, run external Playwright journeys against real ELOG/SSO states, record usability and accessibility findings, make a bounded batch of changes, then capture and verify again. Keep screenshots and the tested revision as evidence in the tracker. Review the result against the thirteen checks below and the existing CDX design language.

Thirteen design checks ​

This is Attest's proposed review checklist, synthesized from Nielsen Norman Group's usability heuristics and WCAG 2.2. It is not a claim that one published standard defines these exact thirteen items.

  1. Show current system state and truthful progress.
  2. Use the tester's language and explain API, integration, UI and performance.
  3. Make the next action obvious on each screen.
  4. Prefer recognition and examples over memorized schema fields.
  5. Reveal advanced controls when needed without hiding capability.
  6. Keep labels, actions, status and navigation consistent.
  7. Preview destructive preparation and prevent invalid runs.
  8. Allow cancellation, correction and safe retry.
  9. Explain errors with cause, impact and a next step.
  10. Support keyboard, focus, screen readers and WCAG 2.2 AA targets.
  11. Adapt layout and interaction to supported screen widths.
  12. Keep response time and loading feedback appropriate for long operations.
  13. Preserve provenance and distinguish results from qualification decisions.

Delivery sequence and proof ​

SliceCDX attest branchELOG attest branchAcceptance evidence
0. BaselineRender Attest and run external Playwright journeys; capture desktop/narrow snapshots, a11y and failure-state findings.Exercise ELOG and SSO setup, selection and current real runs without changing scenarios.Reproducible findings tied to revision and screenshots.
1. Product foundationFinalize typed case/environment/execution/shard contracts, provider boundaries, secrets policy and registry-only install; keep local and cloud paths compatible.Preserve ELOG/SSO adapter behavior, IDs and seed rules.A generated non-Caliber Flutter sample installs, validates and runs one real case; ELOG/SSO compatibility tests pass.
2. First-use and authoringGuided onboarding, API editor, UI step editor, integration composer, validation, config diff and an attest-authoring skill.Supply realistic ELOG/SSO examples through product-owned definitions.A new user can author and run supported API and UI cases; generated YAML passes the same validator as hand-written YAML.
3. Running and replayUnified run flow, preparation preview, streamed step state, truthful report, synchronized video/screenshot replay and failure recovery.Run clean ELOG/SSO cases and inspect real recordings/artifacts.Playwright checks Attest; Patrol checks Flutter journeys; failed preparation and assertions remain explainable.
4. Performance and UXk6 presets/threshold editor, measured scorecard, dashboard investigation design, responsive/keyboard/accessibility pass.Validate ELOG traffic examples and SSO eligibility.Thresholds control result; no fabricated metrics; snapshot review and regression suites pass.
5. Cloud control planeExpose authenticated API, tenant/project roles, durable work ledger, pluggable artifact storage and worker registration; prove local and hosted topologies use one contract.Connect ELOG/SSO as customers without moving product rules into CDX.An isolated deployment can submit, observe, cancel and retrieve an execution with no local filesystem dependency in the API.
6. Testing farmAdd capability-based dispatch, leases, bounded parallel shards, streamed events, worker loss recovery and quotas. Prepare schedule design for separate manual approval.Use ELOG/SSO to prove mixed API/UI/k6 assignment and aggregate outcomes.Multiple workers execute shards; missing/failed shards cannot yield a passing aggregate; video and logs link to the right shard.
7. AI assistanceAdd opt-in OpenAI/Gemini-compatible provider adapters for the authoring skill, with draft review, validation and secret controls.Try real ELOG/SSO journey requests and reject incorrect proposals.A natural-language request yields reviewable files and exact run commands; no model output can silently execute or change acceptance rules.
8. Customer acceptanceRun the entire first-customer path from a clean checkout and registry packages; repair gaps found through Playwright and real runner evidence.Use ELOG and SSO as separate acceptance fixtures, including failed preparation, failed assertion, replay and report.An independent Flutter sample and ELOG/SSO complete setup, authoring, execution, diagnosis and export with recorded evidence.

Each independently reviewable change gets a focused test, commit and push to its respective attest branch. The farm, AI providers and schedules are staged capabilities; they do not block early local improvements or justify claiming cloud readiness before the evidence exists. No merge, package publication, deployment, scheduled automation or workflow file is authorized by this plan. Before writing any workflow, document its trigger, permissions, cost, artifacts and rollback for separate manual review.

Decisions requested before implementation ​

  1. Primary customer: Is the first self-service user a QA analyst who should author cases entirely in the studio, or a Flutter developer who is comfortable keeping cases in the repository? Proposed default: support both, with the guided studio as the first path and code-backed cases as an escape hatch.
  2. Runtime scope: By “sub-OS instance,” do you mean an isolated Docker/ Supabase-like stack per run, or a full virtual machine? Proposed first release: pluggable isolated runtime providers, with existing deployments and local containers first; no claim of full VM provisioning until specified.
  3. Design principles: Do you have a particular named set of “13 universal design principles” in mind? The thirteen checks above are a proposed Attest checklist; replace it with your named source if you meant a specific one.
  4. First delivery milestone: You have confirmed cloud hosting and a worker farm as product direction. Should the first customer release include the hosted control plane, or should a complete local/self-hosted release ship first while cloud slices continue? Proposed sequence: prove the complete local customer path first, then release hosted service after isolation, permissions and worker-loss evidence is complete.

Approval of this plan authorizes slice 0 and staged implementation on the two attest branches. The four answers refine scope before behavior or UI changes. Schedules and workflow files retain a separate manual review gate.