Interview Kickstart is now Dexity.
Dexity
Register for Live Webinar
  1. Home
  2. /
  3. Intel
  4. /
  5. Upskilling Reality
Upskilling Reality

AI Evals in Production: The Error-Analysis-First Playbook (2026)

Evals — not model choice, not prompt cleverness — decide whether AI features work in production. In Dexity's analysis of live job descriptions, evals now appear in 56% of AI-engineer and 32% of product-manager postings, up from near-zero two years ago. Here's the error-analysis-first method teams use to ship AI they can measure instead of hope for.

Summarize with AIChatGPTClaude
  • 31 July 2026
  • 8 min read
Key facts
  • In Dexity's analysis of live job descriptions, evals now appear in 56% of AI-engineer postings.
  • In Dexity's analysis of live job descriptions, evals now appear in 32% of product-manager postings.
  • Evals went from a research topic to a hiring requirement in under two years, up from near-zero in 2024, according to Dexity's job-description analysis.
  • There are three canonical types of AI eval — human evals, code/assertion-based evals, and LLM-as-a-judge — and mature systems combine all three rather than picking one.
  • In the error-analysis-first method Dexity teaches, reviewing just a few dozen production traces is often enough to surface a system's recurring failure patterns.

Product managers evaluate AI and LLM features by starting with error analysis, not tooling: read a sample of real production traces, note what went wrong in plain language, group the failures into recurring categories, and count them to get your actual failure distribution. In Dexity's analysis of live job descriptions, evals now appear in 56% of AI-engineer and 32% of product-manager postings — up from near-zero in 2024. Only then do you build measurement — LLM-as-a-judge scorers, assertion checks, and monitoring — around the failures that matter. Evals, not model choice or prompt cleverness, are the difference between an AI feature that works in production and a demo that falls over. As models commoditize, the durable advantage is knowing whether your system actually works — and why.

Why did AI evals become a hiring requirement?

Two years ago "evals" was a research word. Today it's on the job description.

The reason is simple: AI systems are probabilistic. You can't ship them the way you ship deterministic code — you have to measure whether they produce correct, useful output, on your data, for your users.

Employers are now writing this into the job itself:

"Help customers develop evaluation frameworks to measure Claude's performance for their specific use cases." — Anthropic, Applied AI Architect job description (2026)

How do you run error analysis on your AI outputs?

The most common mistake is reaching for an eval tool first. The teams that ship reliable AI invert that — they start by looking at their own data:

  1. Read real traces. After any significant change, manually review a sample of your system's outputs (a few dozen is enough to see patterns).
  2. Note what's wrong, in open-ended language — don't force categories yet.
  3. Categorize the failures into recurring buckets (wrong retrieval, hallucinated field, tone, format, broken tool call).
  4. Count them. Now you know your actual failure distribution, not your imagined one.

Only then do you build measurement — because now you know what to measure. This error-analysis-first approach is the method practitioners like Hamel Husain and Shreya Shankar have made the standard for applied LLM evals (Hamel Husain, "Your AI Product Needs Evals"), and it's the same discipline Aman Khan lays out in Lenny's Newsletter's guide to evals for PMs.

What kinds of AI evals are there?

Once error analysis tells you what to measure, you choose how to measure it. There are three canonical approaches, and mature systems use all three in combination rather than picking one:

Eval type How it works Best used when
Human evals A person reviews outputs against a rubric and labels each one Early on; for subjective quality (tone, helpfulness, safety); and to produce the ground-truth labels the other methods are validated against
Code / assertion-based evals Deterministic checks assert on structure or content — valid JSON, required fields present, exact match, regex, no forbidden strings The correct answer is objectively checkable: output format, schema, tool-call arguments, presence of a required key
LLM-as-a-judge A model scores each output against a rubric, calibrated against your human labels Subjective criteria that need to run at scale, after you have human labels to validate the judge against

The practical ordering follows from that: humans first (to see failures and create labels), assertion-based checks wherever correctness is deterministic (they're cheap, fast, and never drift), and LLM-as-a-judge to scale the subjective judgments once a human-labeled set exists to keep the judge honest.

How do you scale from manual review to automated evals?

Once error analysis reveals the failure modes, you scale from manual review to automated measurement:

  • LLM-as-a-judge — encode each recurring failure as a judge prompt, validated against your human labels, so every new output gets scored.
  • Synthetic data — manufacture edge cases you can't yet collect in production.
  • Production monitoring — keep sampling live traces; the failure distribution drifts as usage grows.
  • A data flywheel — each round of analysis → judges → fixes → new traces compounds into a system that improves over time.

Open frameworks from the major AI labs make the plumbing straightforward (OpenAI Evals, Anthropic evaluation docs) — but the frameworks are the easy part. The judgment about what to measure comes from error analysis.

How do you design a good eval?

How do you write an eval rubric?

A rubric turns a vague "is this good?" into specific, independently checkable dimensions. Break "good" apart into the qualities that actually matter for your feature — for example correctness, retrieval relevance, format, tone, and safety — and score each one separately instead of collapsing everything into a single blurry number. For each dimension, write down what a pass and a fail look like, anchored to concrete examples pulled straight from your error analysis. Prefer a binary pass/fail or a short ordinal scale per dimension, and tie every criterion back to a real trace so two reviewers reading the same output land on the same label.

How big should your eval dataset be?

There's no universal number, and bigger isn't automatically better. Start small: as the error-analysis step above notes, a few dozen traces is often enough to surface your recurring failure patterns. Size the set for coverage of the failure modes you found — a dataset that contains real examples of each category beats a larger random sample that misses the cases you care about. Then grow it over time as production surfaces new failure modes, so the set keeps reflecting how the system actually behaves.

What are evals not?

  • Not public benchmarks. An MMLU score doesn't tell you whether your RAG bot answers your customers correctly. Evals are application-specific.
  • Not vibes. "It looks good in the demo" is not an eval — the point is a repeatable, counted failure distribution.
  • Not an infra project first. Tooling comes after error analysis; the reverse is the most common failure.
  • Not one-and-done. The failure distribution drifts; evals are a continuous loop, not a launch gate.

Why is the window to learn evals closing?

Evals went from research topic to hiring requirement in under two years — 56% of AI-engineer and 32% of PM JDs now call for them. The builders who learn error-analysis-first evals now are scarce; the ones who wait will be competing against teams whose products measurably improve every week.

Build a real eval system for your AI product

Reading about evals isn't the same as running them on your own product. Dexity's AI Evals for PMs course walks you through the full loop — data collection, error analysis, architecture-specific eval strategies, and regression detection — so you ship AI you can actually stand behind.

Frequently asked questions

What is an AI eval?

An application-specific test of whether your AI system produces correct, useful output — built from real failure analysis of your own traces, not public benchmarks.

How do you start doing evals?

Start with error analysis, not tooling: review a sample of real outputs, note and categorize failures, count them, then encode the recurring failures as LLM-as-a-judge prompts validated against your labels.

What are the main types of AI evals?

Three canonical approaches: human evals (a person labels outputs against a rubric), code/assertion-based evals (deterministic checks for format, schema, or exact content), and LLM-as-a-judge (a model scores outputs against a rubric, calibrated on human labels). Mature systems combine all three.

Are evals a PM skill or an engineering skill?

Both — evals appear in 56% of AI-engineer and 34% of product-manager job descriptions in our analysis.

Source: Dexity analysis of live AI-engineer and product-manager job descriptions across company career boards, US-inclusive, July 2026 (keyword-coded from full JD text; shares directional). Technical references: OpenAI Evals · Anthropic evaluation docs · Expert sources: Hamel Husain — Your AI Product Needs Evals · Aman Khan / Lenny's Newsletter — a PM's guide to evals · JD datasets & methodology · Dexity.com

More in Upskilling Reality

All Intel →
Upskilling Reality

How to Set Up Conversion Tracking in Google Tag Manager (GA4, Google Ads, Meta) — and Automate It With AI in 2026

There are two ways to set up conversion tracking in Google Tag Manager in 2026. The manual way: build a GA4 event tag and mark the event as a Key Event; add a Google Ads conversion tag, relying on a site-wide Google Tag to store the GCLID (a standalone Conversion Linker is only needed on legacy containers without one); and install the Meta Pixel base with standard events (plus the Conversions API for server-side, deduplicated by a shared Event ID). The faster way: connect Claude to your container through a real GTM MCP server — like Stape's hosted google-tag-manager-mcp-server, which runs via npx mcp-remote at https://gtm-mcp.stape.ai/mcp on Node.js v18+ with Google OAuth — then run an audit, describe the tags you want, stage them in a new container version, and review before you publish. This guide gives the exact manual steps per platform, a copy-paste MCP config, a comparison of real GTM MCP servers, and the verification checklist for both paths — with the human review-before-publish gate kept non-negotiable.

1 August 2026 · 12 min read
Upskilling Reality

AEO Is a Real Job Now — and the Fastest-Growing Skill in Marketing

Answer Engine Optimization (getting cited inside ChatGPT, Perplexity and Google AI answers) has stopped being a buzzword and become a hiring line item. AI referral traffic grew 357% year over year; 94% of CMOs are increasing AEO investment; and companies from Stripe to HubSpot to Anthropic are posting dedicated AEO roles at $75K–$210K. As Kaleigh Moore puts it, the function 'has separated from SEO the same way content marketing separated from copywriting a decade ago.' Here's the data, what the skill actually is, and how to build it in one sitting.

31 July 2026 · 8 min read
Upskilling Reality

What Even Is an 'AI Engineer'? — 425 JDs and One Reader's Comment Later

49% of Software Engineer JDs already mention AI/ML — half the market is hiring for what an “AI Engineer” does day-to-day. Across 425 fresh AI Engineer JDs the role splits: every posting wants LLM integration, but only 36% require agentic systems on top. The label is doing more work than the boundary deserves — and less than 5% of AI Engineer JDs even disclose salary.

31 July 2026 · 10 min read