Back to Conferences
Conferences

OpenRouter and AI Evaluation for Industrial Work

Teseo Data LabOctober 6, 20265 min readLeer en español
How AI Models Learn to Do Real Work, OpenRouter's event in San Francisco, September 30, 2026

On September 30, 2026, Teseo Data Lab attended How AI Models Learn to Do Real Work, an event hosted by OpenRouter in San Francisco. The invitation raised a question that matters to any company looking to bring artificial intelligence into its operations: how do you verify that a model can reliably do the work a business needs?

Poster for How AI Models Learn to Do Real Work, by OpenRouter, featuring Chris Clark and Osvald Nitski

The event announced a conversation between Osvald Nitski, Head of Product at Mercor, and Chris Clark, co-founder of OpenRouter, on how models are trained, evaluated and deployed. Topics included expert data, reinforcement learning environments and agents that can take on increasingly complex tasks.

Attendees at OpenRouter's How AI Models Learn to Do Real Work event in San Francisco
Attendees at the OpenRouter event in San Francisco.
A moment from the conversation on model training, evaluation and deployment.

For Teseo, this discussion connects with hands-on experience: working with fragmented data from sectors such as concrete, real estate and corporate procurement in Mexico and Latin America. It also raises a question about the next use for that information: could it become a reference for evaluating how well AI performs industrial tasks?

Evaluating AI against the work the company needs done

The reflection Teseo's leadership shared after the event focuses on the business value of evaluations. That value lies in checking whether a model solves a specific task with the company's own data, rules and constraints, beyond where it ranks on a general leaderboard.

A benchmark lets you compare capabilities under certain test conditions. To decide how to bring a model into an operation, those capabilities need to be tied to verifiable results: classifying a company correctly, extracting information from a source or recognizing when the evidence is insufficient.

AI evaluations, known as evals, are tests that compare a system's response or output against defined success criteria. For agents that use tools and carry out multiple steps, it also matters to check the final outcome of the task. This distinction is part of the evaluation methodology described by Anthropic.

In an industrial setting, a well-written answer must be backed by verifiable information. Mistaking a sales office for a production plant, for example, can skew a coverage analysis or a prospecting decision.

Industrial data as a reference for checking results

The data Teseo has structured could take on an additional use as ground truth: validated information that serves as a reference for checking answers. The opportunity its leadership pointed to is turning sector knowledge into criteria for evaluating how well models understand industrial work.

This requires distinguishing between having organized information and having a reliable reference. To use a record in an evaluation, it helps to specify what was verified, with what evidence, on what date and under what definition.

For example, a company profile may bring together location, activity, fleet and capacity. Each field sets a different bar: confirming that a facility exists, interpreting a category correctly or preserving the difference between an observed data point and an estimate.

Sector knowledge helps define what counts as a correct answer. It also makes it possible to recognize cases where the right answer is to state that there is not enough evidence.

What could be evaluated in the concrete industry

The concrete industry offers specific tasks for exploring this approach: locating facilities, classifying companies and extracting operational information. In their reflection, Teseo's leadership identified these examples as possible uses for the data structured around products such as ConcreteOS. The criteria below illustrate how they could become tests.

Industrial taskProposed evaluation criterion
Locate and validate concrete plantsIdentify the facility and back its location with evidence.
Classify companiesDistinguish activities such as concrete production, pumping and distribution.
Extract fleet and capacity dataKeep figures, units and sources intact without filling gaps with assumptions.
Identify commercial signalsLink each signal to evidence relevant to a business opportunity.
Separate facts from estimatesLabel what was observed, what was inferred and what remains unconfirmed.

One example would be asking several models to analyze the same information about a company. The review could check whether they distinguish a branch from a plant, whether they attribute equipment correctly and whether they recognize that data on production capacity is missing.

The comparison would then have operational meaning: it would show on which tasks each model delivers useful results and where it needs review.

ConcreteOS and the link to business processes

ConcreteOS provides an application context for this reflection. The platform developed by Teseo Data Lab offers solutions for ready-mix concrete companies, including quoting, sales management and cost analysis. That closeness to day-to-day operations helps frame specific questions about how to use AI.

The vision shared by leadership proposes a sequence: start from industrial data, build validated references, design evaluations, select models by task and explore integrating them into agents and business processes.

Under this approach, a future evaluation would have to consider both the quality of the answer and the conditions of use. A result could be accurate and still need adjustments because of its cost, its run time or the human review it requires.

Model selection would be tied to a defined job and a verifiable level of rigor. This is a development direction proposed by Teseo, and putting it into practice requires building and testing each component.

Teseo's opportunity in industrial intelligence

The opportunity Teseo sees is turning its experience with data and productive sectors into a foundation for evaluating how reliable AI is. That means connecting sources, validation criteria and business knowledge to decide which tasks can be built into a process, and with what controls.

This vision broadens the role of industrial market intelligence: the information that helps you understand a market could also be used to check whether an AI system interprets it correctly.

Teseo's attendance at the OpenRouter event left a concrete line of thinking for its development: bringing model evaluation closer to companies' everyday work, with sector data and results that can be reviewed.

Which task in your company would be worth evaluating before you automate it? Talk to Teseo Data Lab about your data, your processes and the criteria you need to explore an artificial intelligence application.

Want to analyze your project in Mexico?

Our team can generate a custom analysis with market intelligence specific to your area.

Request analysis

Frequently asked questions

What was How AI Models Learn to Do Real Work?
An event hosted by OpenRouter in San Francisco on September 30, 2026. It announced a conversation between Osvald Nitski, Head of Product at Mercor, and Chris Clark, co-founder of OpenRouter, on how models are trained, evaluated and deployed.
What is an AI evaluation, or eval?
A test that compares a system's response or output against defined success criteria. For agents that use tools and carry out multiple steps, it also matters to check the final outcome of the task.
What is ground truth in an evaluation?
Validated information that serves as a reference for checking answers. To use a record like this, it helps to specify what was verified, with what evidence, on what date and under what definition.
What could be evaluated in the concrete industry?
Tasks such as locating and validating plants, classifying companies, extracting fleet and capacity data, identifying commercial signals and separating facts from estimates. These are examples proposed by Teseo Data Lab; putting them into practice requires building and testing each component.