Spec-driven testing
Upload the PRDs and user stories you already write. Is It Ready plans risk-based test scenarios from them, runs each one in a real browser, and tells you where the product disagrees with the spec — with no test scripts to write or maintain.
Your specs become the source of truth for QA. When a spec changes, the next run's strategy changes with it.
PRDs, user stories, feature docs, API contracts or design specs, in Markdown.
Risk-based scenarios: happy paths, negative cases, boundaries and edge cases.
Review the plan and pick scenarios, or let CI approve it automatically.
The agent drives your app step by step and records everything it sees.
Bugs, spec mismatches and UX scores, each with screenshots and steps.
Hand-written end-to-end suites encode what the product did on the day the test was written. They break when a button moves, they say nothing about flows nobody scripted, and they cannot tell you that the shipped behaviour quietly drifted away from what the product manager asked for. Most teams end up with a slow, flaky suite that covers the happy path and little else.
Is It Ready starts from intent instead. A spec is a Markdown document describing how a feature should work: which fields exist and how they are validated, what the user flows are, and what counts as success. When you create a mission— an objective such as “verify a new user can register and reach the dashboard” — you link the specs it should read. At the start of each run the AI generates a test strategy: a list of scenarios with the risk-based reasoning for why each one was chosen.
You can review that strategy and select scenarios before anything runs, or set autoApprove so planning and execution happen in one call, which is what a CI pipeline wants. Then an AI agent opens your application in a real browser and works through each scenario step by step, capturing screenshots, network traffic, console errors and accessibility violations as it goes. A judge compares what happened with what the spec says should have happened.
The output is not a wall of green ticks. Each scenario is passed, failed, or explicitly not scored — an inconclusive scenario never inflates your pass rate or fails your build. Every finding is typed as a bug, a UX issue, a spec mismatch or an observation, carries a severity and a confidence, and comes with steps to reproduce, the expected and actual result, and the screenshots that prove it. Findings can be filed to GitHub automatically.
Because the plan is derived rather than hand-coded, coverage grows with your documentation. The Test Planner reads every spec for an application, builds an inventory of its capabilities, maps that against your existing missions, and proposes new missions where coverage is missing, keeping the plan current as specs change.
Scenarios are planned across four risk categories, and every run is judged on more than pass or fail.
The intended flow works for the right user, end to end, exactly as the spec describes it.
Invalid input, unauthorized actions and missing permissions are refused with a clear message.
Empty fields, maximum lengths, special characters and edge values the spec's validation rules imply.
Back button, refresh mid-flow, duplicate submissions and dead ends a scripted test rarely covers.
Behaviour that works but contradicts the document — reported as its own finding type, not buried in a pass.
Every scenario scored 1–5 on nine usability lenses: clarity, findability, feedback, recovery, trust and more.
An illustrative, abridged run summary for a registration mission. The failed boundary scenario is a spec mismatch: the app works, but not the way the spec says it should.
Run 7f3c · Smoke test registration · completed in 6m 12s
Strategy: 6 scenarios from 2 specs (User Registration, Password Policy)
✓ happy_path New user registers and lands on the dashboard
✓ negative Duplicate email is refused with an inline error
✗ boundary 12-character password rejected
spec_mismatch · high · confirmed
expected: "Password: 12+ chars" accepts exactly 12
actual: "Password must be longer than 12 characters"
evidence: 3 screenshots, 1 console error
✓ boundary Name field accepts 100 characters
✓ edge_case Refresh on step 2 keeps entered values
~ edge_case Double-click on Submit creates one account (inconclusive)
UX clarity 4 · findability 5 · feedback 3 · error_prevention 3 · recovery 4
Result: 4 passed · 1 failed · 1 not scored → exit 1Create specs in the app, upload a zip of Markdown files, or push them from your docs system with the REST API. Upserts match on externalId first, then slug, so re-running a sync updates documents instead of duplicating them.
curl -X POST "$UCT_BASE_URL/api/specs" \
-H "Authorization: Bearer $UCT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"title": "User Registration",
"content": "# Registration\n- Email must be unique\n- Password: 12+ chars ...",
"applicationId": "<app-id>",
"category": "user-story",
"tags": ["auth", "registration"],
"externalId": "CONF-1234",
"externalSource": "confluence"
}'Start a run from any CI system, wait for it to finish, and fail the build on failed scenarios. Exit codes are 0 for pass, 1 for failures above your severity threshold, and 2 when the run could not complete.
name: Is It Ready
on:
pull_request:
branches: [main]
jobs:
qa:
runs-on: ubuntu-latest
steps:
- name: Run the registration mission
env:
UCT_BASE_URL: ${{ secrets.UCT_BASE_URL }}
UCT_API_KEY: ${{ secrets.UCT_API_KEY }}
run: |
RUN_ID=$(curl -sf -X POST "$UCT_BASE_URL/api/runs" \
-H "Authorization: Bearer $UCT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"missionId":"<mission-id>","autoApprove":true,"costProfile":"regression"}' \
| jq -r '.id')
until [[ "$(curl -sf "$UCT_BASE_URL/api/runs/$RUN_ID" \
-H "Authorization: Bearer $UCT_API_KEY" | jq -r '.status')" =~ ^(completed|failed|cancelled)$ ]]; do
sleep 15
done
FAILED=$(curl -sf "$UCT_BASE_URL/api/runs/$RUN_ID" \
-H "Authorization: Bearer $UCT_API_KEY" \
| jq '[.scenarioResults[] | select(.status == "failed")] | length')
test "$FAILED" -eq 0The same three REST calls work in GitHub Actions, GitLab CI, Jenkins, Azure Pipelines and CircleCI: POST /api/runs to start, GET /api/runs/:id to poll, and an evaluation step that decides the exit code. Pass a callbackUrl instead of polling if your pipeline prefers a webhook.
When the run finishes, export it with GET /api/runs/:id/export as Markdown for the job log, JUnit XML for your test reporter, or SARIF 2.1.0 for GitHub code scanning, where functional findings, accessibility violations and security findings appear as alerts on the pull request.
Run against a named environment such as staging per run, so the same mission tests every preview deployment without editing it.
Domain-verified security testing
Security scans and red-team missions only run against domains your organization has verified.
Credentials never stored in evidence
Sign-in values are redacted from run evidence; only the outcome and landing page are recorded.
Configurable retention
Run evidence is redacted after 90 days by default, adjustable per organization.
Isolated workspaces
Every app, spec, run and API key belongs to one organization and is invisible to every other.
CI-native output
Exit codes for pipelines, plus Markdown, JUnit and SARIF exports for GitHub code scanning.
Upload one PRD, create a mission, and see the strategy the AI plans from it in minutes.