Research

Neo Lab

Where we train the models that make agents reliable at company scale.

Everything on this site rests on one assumption: that an agent can be trusted with a piece of a real business. That assumption is not free. Neo Lab is the group whose job is to keep making it true.

What we work on

Four problems between a demo and a company

Long-horizon reliability

A demo agent has to be right once. A company agent has to be right on Tuesday afternoon, on the four-hundredth run, on a task nobody is watching.

Tool use under real conditions

Most failures are not reasoning failures. They are a stale credential, a renamed field, a rate limit — and an agent that plows on regardless.

Cost per completed task

Tokens are the wrong unit. What a business buys is a finished piece of work, and that is the number we optimize against.

Judgment that transfers

The Company Brain is only useful if a model can act on it the way a senior person would. That transfer is a training problem, not a prompt.

Working with the lab

Model work at Neo Lab is driven by what breaks in production, so the most useful thing a company can give us is a hard workflow. If your firm has one — long-running, tool-heavy, expensive to get wrong — we would like to hear about it.

Neo Lab

Reliability is the product

The research shows up as fewer retries, fewer silent failures, and a lower bill for the same finished work.