NewMissing Records Detection: flags every visit, provider, and date missing from the file. See how →
Remote (US) Contract Evals & QA

AI Medical Records Evals Specialist

You are the human quality-control layer between our AI and the medical-legal professionals who rely on its output. You check what the system generated against the underlying record, line by line, and turn what you find into the eval sets that keep it honest.

About Medrecords AI

Medrecords AI helps attorneys, physicians, medical-legal professionals, insurers, and claims teams process large medical records faster and more accurately. The platform turns cases that run to thousands, sometimes tens of thousands, of pages into evidence-backed chronologies, summaries, timelines, and answers to questions about a patient's history.

Our users are IME and QME physicians, medical expert witnesses, personal injury and malpractice attorneys, legal nurse consultants, life care planners, claims adjusters, and TPAs. For them, a missed diagnosis, a wrong date, or an unsupported claim can change the outcome of a case, so accuracy, completeness, and traceability matter more than a fluent-sounding answer. The platform runs fast, around 100 pages a minute, so a 2,000-page case processes in about 20 minutes. Speed without accuracy just means confidently wrong answers, faster, which is exactly why this role exists.

Records are handled under HIPAA with signed BAAs, and the platform is SOC 2 compliant.

The role

You take a set of medical records and the chronology, summary, report, or Q&A answer our system generated from them, and you check the second against the first. Concretely, you decide:

  • What did the AI miss, misstate, or invent that isn't in the record?
  • Are the dates, diagnoses, medications, procedures, providers, and encounters right?
  • Does every citation actually point to the passage it claims to support?
  • Did the system conflate two encounters, two patients, or two events?
  • Did it state an interpretation as if it were a documented fact, or the reverse?
  • Would a medical-legal professional trust this output enough to put their name on it?

What you'll do

Evaluate outputs against source records. Review chronologies, summaries, reports, and Q&A responses against the underlying medical records. Verify individual facts. Flag omissions, hallucinations, unsupported conclusions, and citation mismatches.
Build the eval sets. Turn real and representative cases into gold-standard answers, timelines, and expected findings. Write edge cases around handwritten notes, OCR errors, duplicate records, conflicting information, and long multi-provider histories. Define what "pass," "needs review," and "fail" mean for each output type.
Run error analysis. Tell an extraction error apart from a reasoning error, a retrieval error, an OCR issue, and a workflow bug. Look for patterns across cases instead of treating each miss as a one-off. Escalate anything severe enough to affect a real report.
Close the loop with engineering. Turn what you find into reproducible test cases. Help engineers reproduce hard failures. Regression-test every new model, prompt, or workflow change against the existing eval set before it ships, so a fix in one place doesn't break another.

What we're looking for

  • You've reviewed medical records, clinical documentation, or medical-legal files closely enough to catch a wrong date or a misread lab value on your own.
  • You know medical terminology, diagnoses, procedures, medications, labs, and imaging well enough to read a chart without translation.
  • You can read a longitudinal record across multiple providers and specialties and hold the timeline in your head.
  • You can state, precisely, the difference between what a record documents and what someone inferred from it.
  • You write findings down clearly enough that an engineer who never saw the case can act on them.
  • You can hold a defined evaluation standard consistently across a large volume of repetitive review, without your judgment drifting case to case.
Strong plus if you have
  • Clinical background as an RN, LPN/LVN, medical coder, clinical documentation specialist, or nurse/medical reviewer.
  • Experience with IME/QME, personal injury, malpractice, workers' comp, disability, or insurance-claim records.
  • Experience building medical chronologies or record summaries for attorneys, physicians, insurers, or TPAs.
  • Exposure to LLM evaluation, prompt testing, data annotation, or AI QA, and familiarity with terms like hallucination, retrieval error, and RAG.
  • Comfort with spreadsheets, SQL, or Python for building and querying eval sets (not required).

What we value

Accuracy over speed. Evidence over assumption. The error nobody else noticed. It's the same standard the product is built on: if we can't cite it, we don't say it. If your instinct on reading a claim is "where exactly in the record does this come from," you'll fit here.

At a glance
Location
Remote, United States
Type
Contract
Team
Evals / Medical Records Quality Assurance
Compensation
Hourly rate, based on experience, clinical background, and level of expertise
To apply

Tell us about a specific record review, chart audit, or QA case where you caught something everyone else missed, and how you found it.

Email [email protected]
← All open roles