An AI agent that performs well in the demo and fails in production is the most repeated scenario in 2024 and 2025 LLM automation projects. The reason is almost always the same: nobody defined how to measure whether it works, and the system was never instrumented to show what it does when no one is watching. At JULDITEC we see this constantly when integrating agents into n8n flows or business processes on Liferay: the prototype is convincing, but it's missing the layer of evals (systematic evaluations) and observability needed to trust the system in the medium term.
Why "it works well" isn't a metric
With traditional software, a unit test has a binary outcome: pass or fail. With an LLM, the same input can produce different outputs, and "correct" depends on context, tone, whether it complies with a business policy, or whether it cites real sources. This forces a shift in mindset: instead of a single test, you need a set of evals that assess different dimensions (factual accuracy, format, safety, cost, latency) and that run continuously, not just before deployment.
The question an evals system must answer isn't "did it come out right?" but "does it still come out right after changing the prompt, the model, or the knowledge base?". That's the difference between testing an agent once and maintaining it in production with guarantees.
Types of evals that actually provide signal
Unit evals (golden datasets)
A set of representative cases is built (between 30 and 200 is usually enough to start) with the expected input and the success criterion. For example, for a support agent that looks up orders:
{
"input": "¿Cuál es el estado del pedido 10234?",
"expected_tool_call": "get_order_status",
"expected_contains": ["10234", "estado"],
"must_not_contain": ["no tengo acceso"]
}
Each case is run against the real agent and the output is compared against the defined rules. This catches regressions when the system prompt is changed or the base model is updated.
LLM-as-judge: one model evaluating another
For subjective criteria (tone, clarity, compliance with brand policy), a second LLM is used as a judge, with an explicit rubric. It's not perfect, but it scales much better than manual review:
from openai import OpenAI
client = OpenAI()
def eval_response(question, answer, rubric):
prompt = f"""Evalúa la siguiente respuesta según esta rúbrica: {rubric}
Pregunta: {question}
Respuesta: {answer}
Devuelve un JSON con: score (0-10), razon (texto breve)."""
result = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"}
)
return result.choices[0].message.content
It's worth calibrating the judge: run it against responses already scored by humans and check that it matches reasonably well before trusting it at scale.
Human-in-the-loop and production sampling
Automatic evals don't replace human review, they complement it. The usual approach is to sample a percentage of real conversations (between 2% and 10%, depending on volume) and have a human score them using the same rubric. This also helps surface new cases that need to be added to the golden dataset.
An eval that isn't updated with real production cases stops measuring what truly matters.
Observability: seeing what the agent does step by step
Evals answer "does it work well?" at specific moments. Observability answers "what's happening right now?". For an agent with tools, API calls, and reasoning steps, you need detailed traces: which prompt was sent, which model responded, which tool was invoked, how long it took, and how much each step cost.
Traceability tools: Langfuse and similar
Platforms like Langfuse or LangSmith let you instrument each agent call as a span within a trace. Integrating it into a custom agent is straightforward:
from langfuse.decorators import observe
from langfuse.openai import openai
@observe()
def consultar_pedido(order_id):
response = openai.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": f"Consulta el pedido {order_id}"}]
)
return response.choices[0].message.content
Each run is logged with latency, tokens consumed, estimated cost, and the full tree of calls, which makes it easier to diagnose why a response took eight seconds or why an agent called the same tool twice.
Observability in n8n flows
When the agent lives inside an n8n workflow, observability doesn't stop at the LLM: you also need to monitor the preceding and following nodes (database queries, webhooks, transformations). In these cases, we combine n8n's native Execution Log with sending events to Langfuse via an HTTP Request node, so that each workflow execution is linked to the model's trace:
// Nodo Function en n8n: envía metadata de la ejecución a Langfuse
const payload = {
name: "n8n_workflow_execution",
input: $json.input,
metadata: {
workflow_id: $workflow.id,
execution_id: $execution.id
}
};
return [{ json: payload }];
This is especially useful for agents that orchestrate multiple subprocesses: without this correlation, an intermittent failure can take days to track down.
Metrics that actually need watching
- •Success rate per task: the percentage of runs that meet the success criterion defined in the golden dataset, not a vague overall average.
- •Detected hallucinations: responses that state data not present in the consulted sources; detected by comparing the response against the retrieved context (especially critical in RAG systems).
- •Latency per step: in addition to the total, how long each tool or model call takes, to identify bottlenecks.
- •Cost per conversation: input and output tokens multiplied by the model's price, aggregated by task type.
- •Human escalation rate: how many conversations end up handed off to support, an indirect indicator of perceived quality.
These metrics should be visible in a dashboard accessible to the product team, not just in technical logs. Langfuse and LangSmith offer aggregated views; if the stack is more artisanal, exporting to Grafana on top of an events database works just as well.
How to build a continuous evaluation pipeline
The pattern we recommend in JULDITEC projects combines three layers that run at different times:
- •Pre-deployment: the golden dataset runs in CI/CD every time a prompt changes or the model is updated, blocking the merge if the success rate drops below a threshold.
- •Continuous production: every real interaction is traced; a periodic job runs LLM-as-judge on a sample and feeds the metrics dashboard.
- •Weekly human review: the team reviews cases flagged with low scores or high latency, and relevant new cases are added to the golden dataset.
This cycle turns evals into something alive, not a checklist run once before launch. An AI agent in production changes behavior with every model update from the provider, with every new document added to the knowledge base, with every tweaked prompt; without this pipeline, those changes are discovered through user complaints, not through metrics.
If you're about to deploy an agent right now, start with the minimum viable setup: 20 real test cases, a basic trace with Langfuse, and a weekly manual review of ten conversations. That's worth more than any sophisticated dashboard without real data behind it.
