We build the system, not the deck
Our engineers and coding agents ship production AI with logging, evaluation and human review in place from the first commit.
Compliance bolted on afterwards costs more and works less
Logging nobody designed in cannot be reconstructed later. Human review added at the end becomes a rubber stamp. An evaluation suite written after launch measures whatever the system already does. Every control the Act asks for is cheaper and better as a design constraint, which is the argument for having the people who understand the obligations write the code.
What the work involves
- 01
Agents and automation
Systems that do work rather than answer questions: document processing, triage and routing, research and synthesis, back-office automation. Scoped to a workflow with a measurable before and after.
- 02
Retrieval and knowledge systems
Search and question-answering over your own material, with provenance on every answer. A system that cannot show where a claim came from cannot be relied on and cannot be audited.
- 03
Evaluation suites
Regression tests for model behaviour, so a prompt change or a model upgrade shows up as a failing test rather than as a customer complaint. This is what makes accuracy and robustness testable instead of aspirational.
- 04
Observability and audit trails
Structured logging of inputs, outputs, model versions and human interventions, with retention set on purpose rather than by default.
- 05
Human review by design
Review queues, confidence thresholds, escalation paths and override mechanisms the responsible people can operate under real workload.
- 06
Integration and handover
Into your stack, your CI, your cloud. We work in your repositories and hand over something your team can maintain without us.
How this usually runs
What moves the price is scope, and the scope is agreed before anyone writes code. A scoped build carries a fixed price against written acceptance criteria; an embedded team is monthly. We quote after a technical scoping call, not from a brief.
Scoped build, fixed price
A defined system with agreed acceptance criteria. Best where the problem is understood and the boundary is clear.
Embedded team, monthly
Our engineers working alongside yours on a continuing roadmap, with the governance artefacts produced during delivery rather than after it.
Rescue and hardening
For AI already in production without the evaluation, logging or review it needs. We assess first, then close the gaps against the live system.
What you get
- Working system in your repositories, on your infrastructure
- Evaluation suite with baseline results and CI wiring
- Structured logging and audit trail, with a stated retention policy
- Human review design documented against Article 14
- Technical documentation aligned to Annex IV where the system is high-risk
- Handover, runbook and a walkthrough with your team
A good fit if
- Teams with a workflow worth automating and no in-house AI engineering
- Companies whose AI feature shipped before anyone asked about governance
- Product teams needing evaluation and observability retrofitted
Before you ask
You use coding agents. Who is accountable for the code?
We are. Agents speed our engineers up; they do not replace review, testing or responsibility. Everything that ships is read by a human on our side and covered by tests, and we tell you where agents were used. It is the same transparency we would ask of a supplier.
Do we own the code?
Yes. Work is delivered into your repositories and the IP is yours, subject to the licences of any open-source components, which we list.
Can you build on top of an audit or a governance engagement?
That is the common path. The audit finds the gaps, and where a gap is best closed with software rather than policy, the same team builds it already knowing the obligations.
Start with a scoping call
Thirty minutes. Tell us what you run and where it is going wrong, and we will tell you what we would look at first.
