OpenSubs Engineering

How we think about QA, and the seed that grows an AI testing team inside your codebase.

Picture a swarm of testers loose in a copy of your app. Each one plays a different user. They click, call the API, fire the webhooks, run the jobs, move the clock forward, and try things in every order a real person could. And every one of them sees everything at once: the screen, the request, the job, the database, the logs.

That’s Tester Swarm. Plant it in your repo and it maps your app, builds a sealed copy of it, sends the swarm through, proves every bug it finds, and hands you a list. You don’t write tests. You don’t write a spec. You read the list and decide.

Here’s how we think about QA, and why a swarm works.

Tester Swarm by OpenSubsFree and open source, MIT licensed

npx tester-swarm init
View on GitHub
Three periwinkle testers in headphones work at laptops in front of a wall of log screens, faces lit by their screens, while a mint verifier with a clipboard reads over their shoulders.

Software has four parts

Strip any app down and you find the same four things.

  • State: what’s stored. The database.
  • Surfaces: where people and systems touch it. The UI and the API.
  • Actions: things that change state. Save, delete, pay.
  • Streams: what moves between them. Requests, background jobs, webhooks, events.

Three forces act on it from outside.

  • Actors: who does the action. A user, an admin, the system, another user.
  • The outside world: third-party services, payments, email.
  • Time: scheduled jobs, expiry, retries.
Fig. 1. Four parts, three forces. A square system split into four parts: State (what’s stored), Surface (UI, API), Action (save, delete, pay) and Stream (requests, jobs, webhooks). Three yellow arrows press on it from outside: Actor (user, admin, system), Outside (payments, email) and Time (scheduled jobs, expiry, retries).

What a bug is

Given those parts, a bug has a one-sentence definition:

After some sequence of actions, by some actor, at some point in time, the state or what a surface shows breaks a rule.

A gray beetle, the bug, caught mid-step.

The worst bugs live at the seams. The screen says “saved” while the database never changed. One user reaches another’s data. A scheduled job runs on stale data. A payment answers twice.

Tests that only click screens can’t see the database. Tests that only call code never see what a user would do. Each sees one part, and the bugs sit between the parts.

A tester kneels and aims its lantern like a debugger light. The beam sweeps across a wall of logs and pins a gray beetle inside a crosshair.

A periodic table of QA

Every app is built from the same seven elements, in different amounts. A store, a SaaS dashboard and a mobile backend are the same table in different proportions.

The seven elements: what each one is, and where it lives in code.
ElementWhat it isWhere it lives in code
StateWhat’s storedSchema, migrations, models
SurfaceWhere it’s touchedPages, routes, API handlers
ActionWhat changes stateHandlers and functions that write
StreamWhat carries changeQueues, jobs, webhooks, events
ActorWho does itRoles, sessions, tokens
OutsideServices beyond the appPayment, email and other third-party clients
TimeThe clockCron, schedulers, expiries, retries

RuleRule is not an element. It is what every reaction is measured against: what must stay true.

A bug is a reaction between elements. Each combination has its own known bug:

Which combinations of elements make which bug.
When these meetThe bug it makes
Surface plus StateThe screen and the database disagree. It says “saved” and nothing changed.
Actor plus Actor plus StateOne user reaches another user’s data.
Time plus StateA scheduled job runs on stale data.
Outside plus ActionA payment answers twice, late, or never.
Stream plus StreamTwo webhooks arrive out of order.
Action plus ActionThe same thing done twice, or two things at once.

The table does three jobs. Mapping an app becomes identifying which elements it has. Every combination present becomes something to hunt. And coverage becomes visible: like Mendeleev’s table predicting elements before anyone found them, ours lists every combination your app contains, and the report marks the ones tested. An untested cell is a seam nobody tried, stated in advance instead of found by luck.

Act like a user. See like a god.

The whole method rests on one principle with two halves.

A tester aiming a debugger light with a crosshair.

Act like a user.An agent may only do what its actor could do, through that actor’s surfaces. A user agent calls the user’s API. It never writes to the database directly. Otherwise it creates states no real user could reach and reports bugs no real user would hit.

See like a god. Observation is total. Every action is recorded across every part at once: the request, the background job, the database change, the outgoing call, the console, and what the app reported back.

So each finding arrives as one trace across all the seams: what was done, by whom, when, and where the parts stopped agreeing.

Fig. 2. Act like a user, see like a god. One action fans out into seven trace lanes: click, request, job, database change, outgoing call, console and what the screen showed. Each lane pulses a little later than the one above, and one yellow cursor line labeled “recorded at once” crosses all seven.

Simulations, not mocks

To act freely, the swarm needs a world it can’t damage. So it builds one.

A simulation is one running copy of the whole app in Docker. Everything that’s yours runs for real: your app, your real database. Everything that isn’t (payments, email, other third-party APIs) is simulated by a stand-in that behaves like the real service, stays consistent, and logs everything it was asked. A clock the swarm controls lets it jump ahead to next month’s scheduled jobs.

A tester in headphones typing on a laptop.

Synthetic users are written straight into the simulation’s database during setup. The simulation is the swarm’s own copy, so it holds every key and can mint any login. No real accounts. Nothing leaves your machine.

The swarm runs many simulations at once, one per agent, and pulls every trigger the map found: an API call, a webhook arriving, a scheduled job, the clock moving forward.

A grid of twelve identical rooms, each with one tester working at its own monitor. In one room, a lantern beam has caught a beetle.

Review, don’t specify

A tester standing with arms crossed, lantern at its side.

Most testing tools make you write down what to test and what correct looks like. Every flow, every assertion, by hand. It’s slow, and it’s never finished.

Reading a long list of findings and saying “not this, not this, this one” is fast. So you never write a spec. The swarm over-finds, backs every finding with evidence, and you filter.

A periwinkle explorer hands a card with a beetle sketched on it to a mint verifier, who studies it, stamp ready, in front of a wall of logs.

How does an agent know something is wrong in an app it has never seen? It reads intent from what your app already has, strongest first:

  1. The app’s own promises: tests, comments, error messages, UI text, docs, constraints in the schema.
  2. Sibling paths: one handler does it right and another doesn’t.
  3. Git history. A fix commit is the clearest promise there is: the diff shows wrong turning into right, in a part of the app that has broken before.
  4. Past bug lists, if you point to one.
  5. Universal rules that hold in any app: no server error on bad input, a refusal changes nothing, users only reach their own data.
  6. General knowledge of how this kind of software behaves.
A verifier reading a clipboard, one eyebrow raised.

Every finding must name the value it expected and where that expectation came from. “It looks wrong” is not a finding.

Fig. 4. Where agents get intent. Six sources ranked strongest first: 1 The app’s own promises (tests, comments, errors, UI text, docs), 2 Sibling paths, 3 Git fix history, 4 Past bug lists, 5 Universal rules, 6 General knowledge. They funnel into a human review card that reads “not this”, “not this” (struck through) and “this one” with a check.

We ship a seed, not a test suite

The seed is small and general: the periodic table as data, one skill per stage, the scripts that run the team, and an install command. It holds no app code and no app-specific rules. It doesn’t need to know your app, only to recognize the elements in it.

A mint verifier stamps a check mark onto a card under a desk lamp while a second verifier, arms crossed, watches a monitor end on a green check.

Planted in a repo, it grows .qa/ through eight stages:

  1. Interview. It drafts a map from your code, then asks you to fill the gaps, one question per element.

  2. Map. Every screen, endpoint, table, job and outside service. Areas fall out of it.

  3. Simulate. The app in Docker, outside services simulated, synthetic users, a controllable clock.

  4. Brief. One brief per area, split when an area is too big, plus briefs for reactions that cross areas.

  5. Hunt. One explorer per brief, all in parallel, each in its own simulation.

  6. Verify. A fresh agent in a fresh simulation reproduces each finding from a clean start. Not reproduced means dropped.

  7. Report. One card per root cause, sorted: bug, hardening, not a bug. Plus the cells nobody reached.

  8. Decide. You mark each one. Your decisions are saved, and the next run reads them.

Fig. 3. How the seed grows. A seed planted in the .qa/ row of a repo starts a pipeline of eight stages: 1 interview, 2 map, 3 simulate, 4 brief, 5 hunt, 6 verify, 7 report, 8 decide. Each stage drops what it leaves in .qa/: map (from interview and map together), simulation, briefs, findings, verified, report and decisions. A dashed arrow from decisions loops back: “next run reads it”.
A periwinkle tester in headphones types at a terminal full of log lines, the screen’s light falling across its face, a lantern and a coffee mug on the desk.
A verifier stamping a finding.

Plant the seed

npx tester-swarm init
View on GitHub

Free and open source, MIT licensed. OpenSubs is giving Tester Swarm away to everyone who builds software.