It works.
Normal conditions. Latency is flat, queues stay empty, every request is answered inside its budget.
p99 flat · queue depth 0–1 · errors 0%
We determine the operating limits, failure boundaries and technical constraints of complex systems — and prove them with evidence you can reproduce.
You are reading Board, for founders, owners and investors. Turn to Bench, the reading for engineers.
Most systems are tested while they work. We examine what happens when conditions stop being ideal.
Concurrent sessions
An illustrative examination of one agent workflow. Not a client result.
The question is not whether your system works. The question is where it stops working.
Every system has an operating range. Demos and happy-path tests only ever see the first part of it.
Normal conditions. Latency is flat, queues stay empty, every request is answered inside its budget.
p99 flat · queue depth 0–1 · errors 0%
Concurrency, resource pressure and latency start to interact. The tail moves long before the average does.
p99 rising · pools saturating · retries appear
The system crosses its operating boundary. Queues grow without bound, retries amplify load, and it does not recover on its own.
p99 unbounded · queue growing · errors climbing
Not your AI, your roadmap or your team. The running system — measured along five dimensions that meet at one boundary.
How the system actually executes, path by path.
call paths · critical sectionsWhat changes under simultaneous demand.
sessions · contention · locksWhere response time becomes unacceptable.
p50 · p99 · p99.9Where compute, memory and queues become the limit.
CPU · memory · queues · poolsHow reproducibly it behaves under equivalent conditions.
run-to-run variance · jitterOne critical workflow, the one whose failure would cost the most.
Controlled operating conditions, raised one step at a time.
Actual behaviour, measured at every step. Nothing assumed.
The boundary identified, together with the mechanism behind it.
Reproducible evidence: the harness, the conditions, the data.
A documented technical boundary: where the system’s behaviour changes, why, and under what conditions it happens again.
A generic benchmark scoreA consultant’s opinionPass or fail
The workflow holds its 2.0 s p99 budget up to 8 concurrent sessions. Above 8, tool-call retries saturate a shared connection pool and latency grows without bound.
Every report follows this structure: finding, operating range, boundary, mechanism, evidence, reproduction, action and confidence.
Each engagement answers one question, and leads to the next only if you need it. The first boundary tells you where to look.
Runtime evidence for a deal: the target’s system examined under load before capital depends on it. Runs alongside DY Research due diligence — one scope, one report — or on its own.
DY PROOF or DY Research?
Produces new evidence: the system examined under controlled load until its boundary is found and reproduced.
Ends in a documented boundaryReads the evidence that exists: the code, the tests and the proofs behind a technology’s claims.
Ends in a written verdict · dyresearch.github.io ↗Different questions, different prices, one standard of evidence. For a deal, the two run as one engagement.
A technical system can look impressive while carrying an undiscovered operating boundary.
DY PROOF provides an independent technical examination before capital, deployment or acquisition depends on the system.
Alongside DY Research due diligence, or on its own. The findings are technical; the investment decision stays yours.
No incentive to confirm the conclusion you were hoping for.
Every finding is tied to behaviour that was observed and recorded.
Conditions and method are documented, so any result can be re-run.
Your system, your data and your results stay private.
Engineering findings in engineering terms. No management theatre.
No examination issues or implies qualification under any safety or compliance standard.
The findings are technical. What to do with them is your decision.
You receive evidence and reasoning, set out so that they can be checked.
From the infrastructure that is built to the diagnostics that prove how systems behave. Each layer works to the same rule: nothing is claimed that cannot be checked.
AxonOS is open infrastructure, built and funded by the house and bound for independent foundation governance. Belonging to the ecosystem never shapes a finding. If a system under examination competes with AxonOS or builds on it, you are told at scoping — before you commit to anything.
Send one workflow. Get its boundary back in writing — before production finds it for you.