What we build
Eight checks, each one a question a team would ask anyway, each one answered by driving a real browser rather than by reading the code. Would a stranger understand this. Where do people give up. Is it fast, and usable by everyone. Did a bug you already fixed come back. What does the whole run add up to.
Whatever they find lands in one list. Every finding carries the steps that produced it and the frame from the moment it broke, so it can be opened rather than believed. You rule on it in one keystroke, and those rulings are what the accuracy figure is measured from.
What we believe
Four of them. Each is enforced somewhere in the product rather than stated here and hoped for, and each has a page that shows the mechanism.
Anything behavioural a model claims is re-driven in a real browser before it can block a release. What a browser cannot settle keeps the word unverified on it and is capped below ship-blocking, so nothing can hold up a release on a sentence alone. How a claim gets confirmed.
The tool reports its own accuracy from your verdicts, and refuses to print a rate it cannot support from data already on disk. Early on it tells you there is not enough evidence yet, which is the honest answer and the unimpressive one. How the figure is earned.
Confirm a finding and it becomes a standing check, minted while the bug is still there so it has to go red before anyone trusts it. A check that has never failed is reported as unproven rather than as passing. How a re-check is born.
What the detectors catch, what they miss, and the command that reproduces each number. A tool that will not publish its own misses is asking to be taken on faith, which is the thing it was built to stop you doing. What we have measured.
What we will not do
- nth Labs does not phone home. No telemetry and no analytics, in any mode. There is nothing to opt out of.
- No metering of your key. Runs on your own model key are never counted, capped or billed by us. On your own install the key never leaves your machine; on the hosted console it is stored for your account alone and spent only on runs you start.
- No score for anything unmeasured. Coverage names the surfaces a run never reached, and the score covers only the ones it did.
- No lock-in. Every check exports as a standalone Playwright spec that runs without us. The report is one HTML file. Your history is files on your own disk.
Where things stand
nth Labs is in beta and in daily use. Parts of it still change week to week, and anything that has not been measured is labelled unmeasured rather than rounded up. The eight checks, the console and the command line are all in the open build today. Create an account and point it at what you are building, or write first if you would rather ask something.
Or go straight to the console.