OpenSubs Engineering
How we think about QA, and the seed that grows an AI testing team inside your codebase.
Picture a swarm of testers loose in a copy of your app. Each one plays a different user. They click, call the API, fire the webhooks, run the jobs, move the clock forward, and try things in every order a real person could. And every one of them sees everything at once: the screen, the request, the job, the database, the logs.
That’s Tester Swarm. Plant it in your repo and it maps your app, builds a sealed copy of it, sends the swarm through, proves every bug it finds, and hands you a list. You don’t write tests. You don’t write a spec. You read the list and decide.
Here’s how we think about QA, and why a swarm works.
npx tester-swarm init
Software has four parts
Strip any app down and you find the same four things.
- State: what’s stored. The database.
- Surfaces: where people and systems touch it. The UI and the API.
- Actions: things that change state. Save, delete, pay.
- Streams: what moves between them. Requests, background jobs, webhooks, events.
Three forces act on it from outside.
- Actors: who does the action. A user, an admin, the system, another user.
- The outside world: third-party services, payments, email.
- Time: scheduled jobs, expiry, retries.

What a bug is
Given those parts, a bug has a one-sentence definition:
After some sequence of actions, by some actor, at some point in time, the state or what a surface shows breaks a rule.

The worst bugs live at the seams. The screen says “saved” while the database never changed. One user reaches another’s data. A scheduled job runs on stale data. A payment answers twice.
Tests that only click screens can’t see the database. Tests that only call code never see what a user would do. Each sees one part, and the bugs sit between the parts.

A periodic table of QA
Every app is built from the same seven elements, in different amounts. A store, a SaaS dashboard and a mobile backend are the same table in different proportions.
| Element | What it is | Where it lives in code |
|---|---|---|
| State | What’s stored | Schema, migrations, models |
| Surface | Where it’s touched | Pages, routes, API handlers |
| Action | What changes state | Handlers and functions that write |
| Stream | What carries change | Queues, jobs, webhooks, events |
| Actor | Who does it | Roles, sessions, tokens |
| Outside | Services beyond the app | Payment, email and other third-party clients |
| Time | The clock | Cron, schedulers, expiries, retries |
RuleRule is not an element. It is what every reaction is measured against: what must stay true.
A bug is a reaction between elements. Each combination has its own known bug:
| When these meet | The bug it makes |
|---|---|
| Surface plus State | The screen and the database disagree. It says “saved” and nothing changed. |
| Actor plus Actor plus State | One user reaches another user’s data. |
| Time plus State | A scheduled job runs on stale data. |
| Outside plus Action | A payment answers twice, late, or never. |
| Stream plus Stream | Two webhooks arrive out of order. |
| Action plus Action | The same thing done twice, or two things at once. |
The table does three jobs. Mapping an app becomes identifying which elements it has. Every combination present becomes something to hunt. And coverage becomes visible: like Mendeleev’s table predicting elements before anyone found them, ours lists every combination your app contains, and the report marks the ones tested. An untested cell is a seam nobody tried, stated in advance instead of found by luck.
Act like a user. See like a god.
The whole method rests on one principle with two halves.

Act like a user.An agent may only do what its actor could do, through that actor’s surfaces. A user agent calls the user’s API. It never writes to the database directly. Otherwise it creates states no real user could reach and reports bugs no real user would hit.
See like a god. Observation is total. Every action is recorded across every part at once: the request, the background job, the database change, the outgoing call, the console, and what the app reported back.
So each finding arrives as one trace across all the seams: what was done, by whom, when, and where the parts stopped agreeing.

Simulations, not mocks
To act freely, the swarm needs a world it can’t damage. So it builds one.
A simulation is one running copy of the whole app in Docker. Everything that’s yours runs for real: your app, your real database. Everything that isn’t (payments, email, other third-party APIs) is simulated by a stand-in that behaves like the real service, stays consistent, and logs everything it was asked. A clock the swarm controls lets it jump ahead to next month’s scheduled jobs.

Synthetic users are written straight into the simulation’s database during setup. The simulation is the swarm’s own copy, so it holds every key and can mint any login. No real accounts. Nothing leaves your machine.
The swarm runs many simulations at once, one per agent, and pulls every trigger the map found: an API call, a webhook arriving, a scheduled job, the clock moving forward.

Review, don’t specify

Most testing tools make you write down what to test and what correct looks like. Every flow, every assertion, by hand. It’s slow, and it’s never finished.
Reading a long list of findings and saying “not this, not this, this one” is fast. So you never write a spec. The swarm over-finds, backs every finding with evidence, and you filter.

How does an agent know something is wrong in an app it has never seen? It reads intent from what your app already has, strongest first:
- The app’s own promises: tests, comments, error messages, UI text, docs, constraints in the schema.
- Sibling paths: one handler does it right and another doesn’t.
- Git history. A fix commit is the clearest promise there is: the diff shows wrong turning into right, in a part of the app that has broken before.
- Past bug lists, if you point to one.
- Universal rules that hold in any app: no server error on bad input, a refusal changes nothing, users only reach their own data.
- General knowledge of how this kind of software behaves.

Every finding must name the value it expected and where that expectation came from. “It looks wrong” is not a finding.

We ship a seed, not a test suite
The seed is small and general: the periodic table as data, one skill per stage, the scripts that run the team, and an install command. It holds no app code and no app-specific rules. It doesn’t need to know your app, only to recognize the elements in it.

Planted in a repo, it grows .qa/ through eight stages:
Interview. It drafts a map from your code, then asks you to fill the gaps, one question per element.
Map. Every screen, endpoint, table, job and outside service. Areas fall out of it.
Simulate. The app in Docker, outside services simulated, synthetic users, a controllable clock.
Brief. One brief per area, split when an area is too big, plus briefs for reactions that cross areas.
Hunt. One explorer per brief, all in parallel, each in its own simulation.
Verify. A fresh agent in a fresh simulation reproduces each finding from a clean start. Not reproduced means dropped.
Report. One card per root cause, sorted: bug, hardening, not a bug. Plus the cells nobody reached.
Decide. You mark each one. Your decisions are saved, and the next run reads them.



Plant the seed
npx tester-swarm initFree and open source, MIT licensed. OpenSubs is giving Tester Swarm away to everyone who builds software.