Data catalog

Clinical datasets and interactive environments for frontier AI

Explore ready-to-license Kaelio datasets or commission a custom data program built around your capability, modality, specialty, or workflow. Kaelio supplies physician-generated and de-identified real-patient data with comprehensive clinician-authored ground truth for training, post-training, and evaluation.

Buyer diligence

Review the complete evidence package

For every ready-to-license dataset and custom program, Kaelio can provide the exact dataset specification, evaluation evidence, clinical ground-truth audit, and delivery plan your team needs to assess fit.

Review technical diligence

Clinical task families

Coverage across clinical tasks

Kaelio data programs cover the range of clinical work labs test models on, not only diagnosis. Each program can be delivered as clinical datasets or interactive environments with task-specific clinician-authored ground truth.

Diagnostic reasoning

Differential diagnosis and management under uncertainty, worked out over a sequential encounter.

AgentClinic · CRAFT-MD

Triage & escalation

When to reassure, work up, or send to the ED, graded for missed escalation and over-escalation alike.

HealthBench · emergency referrals

Longitudinal management

Multi-visit cases where the model tracks disease progression, treatment response, medication changes, adverse effects, and escalation decisions over time.

No established public benchmark

Documentation & summarization

Notes, summaries, and handoffs, graded for accuracy and hallucination by specialty-matched physicians.

HealthBench · health data tasks

Patient communication

Register-appropriate, health-literacy-aware answers for both a patient and a clinician.

HealthBench · expertise-tailored comms

Medication & pharmacy safety

Dosing, interactions, contraindications, and reconciliation, graded by pharmacists.

Pharmacist-graded

Agentic EHR & tool use

Retrieving records, placing orders, and tracking state across many tool calls over an EHR.

MedAgentBench · FHIR-AgentBench

Families span patient-facing encounters and clinician-facing workflows: history-taking, explanation, triage, medication questions, documentation, handoffs, and management planning. More on request: guideline adherence, multilingual and global health, mental-health crisis safety, and evidence synthesis.

By data track

How each family is delivered

The two forms, across modalities: interactive cases with trajectory grading, rubric-graded conversations, reasoning traces, diagnostic imaging, and safety data, each positioned against an established benchmark.

01Text · interactiveFeatured

Interactive clinical environments + trajectory grading

A physician specifies the hidden patient state, the available observations and artifacts, the action space, state transitions, escalation criteria, and stopping conditions. The model gathers information, requests artifacts, and proposes management, then is graded across the full trajectory.

Available from the Kaelio catalog

Research use

  • Sequential decision-making under partial observability
  • Clinical-reasoning-process evaluation (right answer, unsafe path)
  • Process-reward / verifier training and RLHF/RLAIF

Positioned against

AgentClinicCRAFT-MDSDBench / MAI-DxOMedAgentBench

Our edge

An OSCE-style encounter without the patient actor: physician-authored hidden state replaces the LM patient-simulator, and physician grading replaces the model judge, on novel cases in no training corpus at delivery. Those are the two fidelity gaps the benchmarks' own authors flag.

Environment specification

What an interactive environment contains

Each environment defines what the model can observe, what it can do, how the clinical state changes, and how clinicians evaluate the full trajectory.

01Hidden patient state
The underlying condition, history, and findings the model must uncover, without exposing the answer directly.
02Observations and artifacts
Dialogue turns, records, ECGs, images, lab reports, medication lists, and other evidence returned as the model acts.
03Action space
The clinically valid questions, examinations, investigations, treatments, referrals, and tools available to the model.
04State transitions
How the patient and available evidence change with treatment, time, and decisions across an encounter or multiple visits.
05Escalation criteria
The red flags and thresholds that require referral, urgent review, or emergency care.
06Stopping conditions
The conditions that end the trajectory, such as a committed diagnosis, an accepted plan, or escalation out of the environment.
07Trajectory evaluation
Clinician-authored criteria, severity labels, independent review, and adjudication applied to the model's complete path through the case.
Deliverable 1

Interactive cases and rubric

Physician-authored cases with their state, action space, and transitions; an adjudicated rubric with each criterion written and checked by multiple independent physicians; an error taxonomy separating reasoning-process failures from factual and safety ones; a held-out split; inter-rater reliability against targets fixed in the data specification; and optional baseline rollouts for current frontier models.

Deliverable 2

Matched process-supervision data

The same cases with step-by-step physician reasoning traces and physician grading of model trajectories (rubric scores and pairwise preferences), in a form you can use for supervised fine-tuning, preference and RLHF/RLAIF training, and process-reward or verifier training.

02Text

Physician rubric-graded conversations

Newly written, consented multi-turn clinical conversations graded by physicians against adjudicated rubrics for correctness, completeness, safety, and calibration.

Available from the Kaelio catalog

Research use

  • Open-ended dialogue evaluation
  • Preference / RLHF data
  • Grader calibration data

Positioned against

HealthBenchHealthBench HardMedHELM

Our edge

Human rubric grading is the ground truth model-graders are calibrated against. We deliver that human signal directly, on real physician-written conversations rather than synthetic ones.

03Text

Clinical reasoning + physician reasoning traces

Vignettes, differentials, and management plans with step-wise physician reasoning traces. This is supervised signal for medical verifiers and process-reward models.

Available from the Kaelio catalog

Research use

  • Supervised fine-tuning
  • Process supervision / PRM
  • Differential diagnosis under uncertainty

Positioned against

MedQA (USMLE)MedMCQAPubMedQAMedXpertQA

Our edge

The classics are saturated and contaminated, and they grade only the final multiple-choice answer. We produce uncontaminated, process-supervised reasoning traces: the physician-graded step labels PRM work currently lacks.

04Image + text

Multimodal diagnostic imaging

Physician diagnostic grounding and reasoning across radiology, dermatology, pathology, ECG, and clinical photographs.

Available from the Kaelio catalog

Research use

  • Multimodal diagnostic grounding
  • Image-grounded reasoning
  • Bias / robustness evaluation

Positioned against

VQA-RADPathVQAMIMIC-CXRGMAI-MMBench

Our edge

We take on all three documented failure modes at once: auto-label noise, demographic and skin-tone bias, and language shortcuts. Our data is physician-labeled, balanced against those biases, and built to resist shortcuts.

05Text / multimodal

Safety, red-teaming & hallucination

Physician-authored adversarial cases with a severity-graded harm/error taxonomy that separates reasoning-process failures from factual and safety ones.

Available from the Kaelio catalog

Research use

  • Harm detection
  • Hallucination evaluation
  • Clinical red-teaming

Positioned against

Med-HALTMedHELM (safety)HealthBench (safety)

Our edge

Most residual medical hallucinations are reasoning failures, yet benchmarks rarely measure real patient harm and red teams stay small and ad-hoc. Our physicians author adversarial cases and severity-grade the harm in each one.

And more, on request: procedural and multimodal video, decisions under resource limits, and other axes as they prove useful.

Talk through your clinical data requirements

Tell us the capability, clinical workflow, modality, and model-development stage you are working on. We will identify the relevant Kaelio data or scope a custom program with you.