The AI Lab
We build our own AI tools. We break them on our own work first.
Three systems run inside Phoenix Info: one for content and SEO production, one for engineering quality, one for client reporting. They are internal — not products we licence out. They exist because we would rather find the failure modes of production AI on our own operations than on a client's budget.
Content & SEO production
Handles the mechanical layer of content work at scale: source research and collation, brief generation from a topic map, structured-data validation, and QA passes across large content sets for internal contradictions or stale figures.
- Research collation with source links retained for verification
- Structured data and schema validation across every page
- Consistency checks over large sets — stale numbers, contradictions
- Named human editor on every published piece, accountable for it
Engineering quality
Wired into our CI. Reviews diffs for the classes of mistake humans reliably miss at the end of a long day, scaffolds test cases for new code paths, and runs pre-release checks against a standing checklist.
- Review assistance on every diff — a second pass, not a replacement
- Test scaffolding for new paths, then completed by an engineer
- Dependency and secret-handling checks before release
- Every merge still requires human approval. No exceptions.
Reporting & anomaly alerts
Consolidates analytics, ad platform and CRM data into client reporting, and — more usefully — flags the metric that moved before the monthly call, so a problem is being fixed rather than explained.
- Multi-source consolidation with one agreed definition per metric
- Anomaly detection against seasonal baselines, not last week
- Draft commentary, reviewed and signed off by the account lead
- Every figure traceable back to its source system
Why we are telling you this
Most agencies now claim AI capability. Very few have run it in production.
There is a large gap between using a chat assistant to draft copy and operating an AI system that other people depend on. The second one teaches you things the first never will.
We have had a retrieval system confidently cite a document that had been superseded. We have had an evaluation set that passed everything because it was accidentally testing the wrong output field. We have had a cost curve go non-linear on a Friday afternoon. Each of those is now a check in our build process, and none of them happened on a client's system.
When we scope an AI project for you, that experience is the substance of what you are buying. Not access to a model — you can get that yourself for twenty dollars a month.
- Grounding. Answers come from your approved sources, with citations, or the system says it does not know.
- Evaluation before launch. A measurable test set, so you know what it gets wrong before customers find out.
- Human review on consequence. Anything customer-facing or financially material gets a person in the loop.
- Cost ceilings. Per-feature budget caps and alerting, set on day one.
- Version control. Prompts and model versions tracked like code, because they behave like code.
- An honest no. If a rules engine does the job better, we say so.
Straight answers
The AI questions clients actually ask
Usually asked slightly nervously, because a lot of vendors are being vague about this right now.
Next step
Tell us what is actually broken.
A 30-minute call with an engineer, not a salesperson. You will get a straight read on whether we are the right team — including if the answer is no.