Until this week our audit was a fixed list of rules. Each rule looked for one kind of problem — a product missing from the sitemap, a page no link points to — and a language model wrote up what the rules had found.
It was good at what it had rules for. But every new site showed us something it had no rule for, and some of what it reported was wrong. On one small shop, five of the six problems it led with were false. Each fix was one more rule, and each new rule was the next site’s bug.
What we built instead
An AI auditor now reads your site the way a careful customer would. It opens pages in a real browser and looks at them — the screenshot, not only the HTML behind it. It still has the old rules, but as leads to check, not verdicts to print.
Every problem it reports goes straight to a second agent whose only job is to prove it wrong. That agent re-opens every address itself, and rejects the claim if the page does not say what the claim says, if the thing said to be missing is there after all, or if it is normal behaviour rather than a defect.
Then code has the last word. A finding may quote only text that is on a page we read. A status must be one a request actually returned. Every count — 49 of 148 products, 3 of 171 pages — is worked out by code from what was recorded. The model never produces a number.
How we judged it
We ran the old audit and the new one on four real websites, one at a time: a small online shop, an antiques shop, a documentation site of 430 pages, and a documentation site built entirely in the browser. For every finding either of them reported, we opened the live page and checked it by hand.
| Rules | Agent | |
|---|---|---|
| Findings reported | 14 | 14 |
| True | 8, and 1 partly | 13, and 1 true but overstated |
| False | 4 | 0 |
| Real problems found | 7 | 14 |
| Time per audit | 229–365 s | 158–442 s |
Across everything we tried on those four sites, 25 real problems turned up. The rules found 7 of them; the agent found 14.
What it finds that rules could not
- a shipping policy that names a different company;
- a banner promising free shipping from $35, and a policy that starts it at $50;
- a cream sold as “fragrance-free” whose ingredient list includes lavender and rose oil;
- two different prices for the same service on the same page.
Nobody would write any of these as a rule in advance. Each is something a careful reader notices, backed by a quotation that code can check against the page.
What we are careful about
Four sites, one run each, is thin evidence, and an agent does not look at exactly the same things twice. So we keep a set of test websites with problems planted in them on purpose, and two that are meant to have none, and we run every change against them. That test is being rebuilt for the new audit. Until it is, we say so — here.