Skip to main content
Remote Atlas
himalayasCurated job boardOnsiteMid5+ years listedFull Time

Data Scientist, Agent Evaluations & Quality

Clera

United States

Job description

About the Role

This role sits at the intersection of applied data science and AI product quality for a small, fast-moving AI productivity startup building autonomous agents that handle email, calendar, browser, and business software tasks. You will own the measurement of agent quality end-to-end: turning ambiguous product behavior into rigorous, actionable evaluation systems that directly guide engineering and product decisions.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, edge cases, ambiguous requests, and adversarial scenarios.

  • Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and measure grader agreement, false positives, and false negatives.

  • Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

  • 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks, grading systems, and success criteria for ML or AI systems in production.

  • Strong Python and SQL proficiency with the ability to build automated data pipelines and production-quality analysis code at scale.

  • Solid statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance.

  • Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes of language model systems.

  • Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, data, and application logic.

  • Experience communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

  • Comfort operating with high ownership in ambiguous, fast-moving environments, independently turning open-ended quality questions into evaluation systems.

  • Experience with LLM-as-a-judge systems, agentic or multi-step task evaluation, or benchmarking platforms for AI systems is a strong plus.

Location

On-site in Palo Alto, California, United States. Visa sponsorship is not available for this role.

Originally posted on Himalayas

Apply kit

Sign in to copy a field card for the employer’s ATS. We never submit applications for you.

arbeitnowCurated job boardRemoteSeniority not stated

AI Scientist

Mistral.ai

Zurich

Atlas fit 6

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tacklin…

  • ai/ml
  • devops
  • go
  • java
  • python

Sign in to track applications

Details
arbeitnowCurated job boardOnsiteSeniority not stated

Data Scientist, Product Analytics

Sonarsource

Geneva

Atlas fit 5

Who is Sonar? Sonar is driving the future of agent-centric software development. As the leader in AI code verification and governance, we solve a critical prob…

  • ai/ml
  • go
  • python
  • sql
  • strategy

Sign in to track applications

Details
arbeitnowCurated job boardUnknownSeniority not stated

Applied AI Scientist(m/f/d)

Nc Group

Berlin

Atlas fit 4

About us NET CHECK GmbH was founded in 1999 with the aim of improving the quality of communication networks. Since then, NET CHECK has developed into the leadi…

  • ai
  • ai domains
  • ai ethics
  • ai governance
  • ai model

Sign in to track applications

Details
arbeitnowCurated job boardOnsiteSeniority not stated

Bioengineer, Display & Library Engineering

Adaptyv

Lausanne

Atlas fit 3

Adaptyv is building an automated lab that lets AI agents run biology experiments. We're entering the era of agentic science where AI models can now design nove…

  • ai/ml
  • biology
  • python

Sign in to track applications

Details