Engineering
Orchestrating agents without losing the plot
Multi-agent systems fail in the seams, not in the models. What we changed after watching a supervised network quietly drift for three weeks.
By Nils Aubert 2 min read
Taking two projects for Q2 Paris · Remote
Neuryn is an applied AI studio. We take teams from a first architecture sketch to agents, retrieval and evaluation running under real traffic — with the measurements to prove it holds.
Services · 01
We run one engagement at a time. You talk to the people who write the code, not to an account manager relaying it.
Two to three weeks to establish whether the thing is worth building at all. Problem framing, data audit, feasibility spikes, and a written recommendation you can act on — including the recommendation not to build.
Agents, retrieval pipelines and evaluation harnesses built to run unattended. Every system ships with the eval suite that proves it works and the traces that explain it when it does not.
The part most AI projects skip. Interfaces that expose model uncertainty honestly, make correction cheap, and keep a human in the loop where it actually matters.
Work · 02
A selection of recent engagements. Client names appear where we have permission to use them.
Three of thirty · 2024–2026 Ask about the rest
By the numbers
Studio · 03
“The hard part was never the model. It is knowing which problem deserves one.”
Most AI work fails long before the first token is generated — in a problem framed too loosely to have an answer, or a dataset that never contained one. We spend the first weeks on that question, and we are willing to lose the engagement over the answer.
Founded in 2021, Neuryn is a small team of engineers and designers who have shipped applied AI into legal, industrial and field-service contexts where being wrong has a cost. We work in the open: a weekly call, a live staging link, no reveal at the end.
Clients · 04
Neuryn told us in week two that two thirds of what we had scoped was not worth building. That conversation saved us a year. What they did build has been in production ever since, untouched.
The evaluation harness they left behind is the reason we can ship model changes on a Friday. It caught three regressions before our users ever saw them.
They are the rare team that treats the interface as part of the model problem, not as decoration bolted on at the end.
Journal · 05
What we learn shipping this work. One article a month, no editorial calendar forcing the pace.
Engineering
Multi-agent systems fail in the seams, not in the models. What we changed after watching a supervised network quietly drift for three weeks.
By Nils Aubert 2 min read
Evaluation
Ship the harness before the feature. Why we now write evaluation code first, and what it costs when you do not.
By Nils Aubert 2 min read
Design
Users do not experience p95. They experience whether the interface told them what was happening. What we do instead of chasing milliseconds.
By Mira Lindqvist 2 min read
Archive · categories · authors Read the journal
Questions · 06
The six questions that come up on every first call. If yours is not here, ask it directly.
A discovery engagement is €18,000 for two to three weeks and ends with a written recommendation. Build engagements run from €60,000 for a single production system to €250,000 for a multi-quarter platform. We quote after discovery, never before — a number given without context commits nobody.
Six to ten weeks from the end of discovery for a first system in front of real users. We ship to a staging environment in week two and keep it live for the duration, so there is never a reveal at the end.
Almost always. We pair with your team rather than working behind a wall, and we plan handover from the first week. If your engineers cannot maintain what we build, we have failed regardless of how well it performs.
Whichever the problem argues for. We benchmark candidates against your own data during discovery rather than committing to a provider in advance, and we build the abstraction that lets you switch later without a rewrite.
Yes. We work under NDA by default, run in your infrastructure when required, and have delivered under GDPR, HIPAA-adjacent and financial-services constraints. Data residency and retention are decided in week one, not retrofitted.
Thirty days of warranty on defects, then an optional retainer for model drift monitoring, eval maintenance and incremental work. No lock-in: the code, the evals and the documentation are yours, and they run without us.
Contact · 07
We answer within one business day, with an honest read — including when the answer is that you do not need us.