Martlet AI logoMartlet AI

Millions of charts in.
Audit-grade codes out.

Martlet AI runs retrospective review end-to-end: chase lists ranked by RAF impact, every HCC validated against the checks CMS applies at audit, 95% of cases closed automatically at 99% precision — and adds and deletes exported in the format your submission pipeline expects.

  • 95%

    of cases closed automatically, end-to-end

  • 99%

    precision on every code the model automates

  • 95%

    reduction in chart review time

  • Verify · Add · Delete

    confirm, capture, and clean up — in one pass

From your data to your submission file.

Charts, claims and prior coding go in. Verified codes and submission-ready files come out. One system, with no hand-offs to vendors or spreadsheets in between.

  1. Ingest

    Claims, EHR notes, scanned PDFs and bundled files — normalized and de-duplicated.

  2. Prioritize

    Charts ranked by expected RAF impact and documentation strength.

  3. Validate

    Every code checked against the record, with the evidence linked to its page.

  4. Review

    Exceptions routed to your coders, with the evidence already attached.

  5. Submit

    Adds and deletes exported in the format your submission pipeline expects.

Verify. Add. Delete.

Validation returns more than a pass or a fail. Every code comes back with its full evidence mapping — the supporting sentence, the encounter and date of service, the signing provider and their credentials, and the source page. That mapping sorts each code one of two ways.

Add — capture revenue

Documented but never coded. Returned ranked by RAF impact, with the evidence attached, so a coder confirms rather than goes hunting for the chart.

Delete — reduce risk

Problem-list carryovers, unsupported specialties, missing signatures. Flagged for removal with the reason, before submission rather than at reconciliation — and what can still be corrected after a window closes is set by CMS and changes over time (42 CFR 422.310).

On a 12,400-chart run: 11,780 closed automatically · 620 routed to reviewers · deltas exported — every one with page-level evidence attached.

The documentation rules, enforced on every chart.

Every chart runs through the same checks an audit would apply, before submission rather than years after it. CMS guidance changes; we update the checks to match, so validation follows the rules currently in force.

  • 01

    Clinical support

    Each diagnosis is checked for support in the encounter itself rather than accepted because it appeared before, using the MEAT framework most coding programs are held to, and the sentence carrying that support is linked to its page.

  • 02

    Encounter type

    Every encounter is identified and typed on its own, including telehealth, and assessed against the criteria that apply to that type of encounter. CMS guidance

  • 03

    Provider specialty

    Signing clinicians are resolved against the specialty list CMS publishes, which is reissued each year. Where a provider type does not resolve, the diagnosis is flagged with the reason rather than assumed.

  • 04

    Signature and attribution

    Each record is checked for a signature and for whether it can be attributed to the clinician who wrote the note. Anything missing, unattributable or unclear is flagged rather than assumed. Medicare signature requirements

  • 05

    The payment year being coded

    Risk adjustment works a payment year at a time, so we track which conditions have support inside the year being coded and which are carrying over from an earlier one — and score each against the model that applies to its own year.

progress_note_2025-03-14.pdf · page 4

Reviewed labs and medication adherence. A1c 8.1%. Type 2 diabetes with diabetic polyneuropathy — continue metformin, titrate gabapentin. Follow-up in 3 months. Discussed diet and foot care; monofilament exam performed.

ICD-10 E11.42 → HCC 37
Diabetes with chronic complications
v28
  • EncounterOffice visit · detected
  • Date of service03/14/2025
  • ProviderJ. Rivera, MD · credential verified
  • SignaturePresent · e-signed
  • MEATEvaluate · Treat
Closed automaticallyconfidence 0.97

All patient data shown is synthetically generated for illustration.

What manual chart review actually costs you.

A skilled coder reviews 40–50 charts a day at a 95% accuracy target. Outsourcing buys volume at $2–3 a chart, and either way the audit risk stays with you for ten years. These are the numbers a population-scale program has to work against.

  • 40–50

    charts per day for a manual HCC coder — the industry benchmark

  • $2–3

    per chart for outsourced review, before rework and QA

  • 95%

    the accuracy target most coding QA programs are held to

  • ~9.5%

    CMS's estimated Part C improper payment rate

Three ways to run retrospective coding, and where the cost sits in each.

The same work can be bought three different ways. What changes is who does it, what you are billed for, and who carries the audit risk once it is done.

Outsourced servicesAI-assisted toolsMartlet AI
Where it fits bestNo internal coding team, and no plan to build oneA coding team you want to make faster at its current sizePopulation-scale volume, with the standard and the audit trail kept in-house
PricingPer chart, or a percentage of captured RAFPer-seat licensing, scaled to the number of reviewersAnnual license. No per-chart fees, no success commission
Who does the workThe vendor's coding team, on the vendor's scheduleYour coders, reviewing each suggestion the tool makesThe engine closes 95%; your reviewers work the 5%
What the cost tracksCharts worked, or RAF capturedThe number of reviewer seats you staffA flat annual fee, independent of volume or captured RAF
Institutional knowledgeHeld in the vendor's tooling and processSplit between tool and teamStays in-house — your data, your rules, your audit trail
Audit postureYou carry the audit risk for coding done elsewhereDepends on each reviewer's judgmentEvidence packet exists the day the code is submitted

The submission calendar, and how the work is paced to it.

Most retrospective programs look only for what is missing, and run against whichever deadline is closest. Running the full population continuously means the adds and deletes are ready before a window opens rather than assembled while it closes.

Sweeps per payment year

CMS submission deadlines
Initial~First Friday of September, before the payment year
Mid-year~First Friday of March, during the payment year
Final~January 31 of the year after the payment year

CMS sets these dates and reissues them each payment year, so they move. Martlet AI works to the schedule in force, with adds and deletes ready ahead of each deadline.

The v28 model maps fewer codes, and the impact is not evenly spread.

The v28 model added HCC categories but removed 2,294 ICD-10 codes from risk adjustment, constrained diabetes coefficients to a flat ~0.166, and dropped codes like unspecified PVD (I73.9, formerly worth ~0.288 RAF) entirely. Wakely’s analysis shows plan-level impacts ranging from −20% to +10%.

  • 115

    HCC categories under v28, up from 86

  • 7,770

    ICD-10 codes that still map — down from 9,797

  • −3.12%

    CMS-estimated average risk score impact

  • −20% to +10%

    plan-level spread — your mix decides your number

Martlet AI maps every diagnosis under both v24 and v28, a payment year at a time: each code is validated against the model that applies to the year being coded, and codes that no longer map under the model in force are surfaced before you submit rather than at reconciliation. A run covering several years is never pushed through a single model.

Code it right once.
The audit is already answered.

Every chart closed here is verified against the same checks a RADV audit applies. If your contract is selected, the evidence packets already exist — nothing has to be reconstructed from records that are years old, by people who may no longer be there to ask.

See RADV readiness

What operators ask about retrospective coding.

How do you decide which charts to work?

Charts are ranked by expected RAF impact and by how strong the documentation behind them looks, so the work starts where the return is. The weighting is yours to set — by line of business, provider group, condition, or a rule your own program already uses — and you can also run the full population rather than a list, which is what most customers do once the throughput is there.

How do we know the codes it closes automatically are right?

Every automated code carries the sentence that supports it, the source page, the encounter and date of service, and the provider and signature status — so any decision can be opened and checked rather than taken on trust. You set the confidence threshold at which a code closes without review, and you can route any share of automatic closures into QA sampling. Confirmation and deletion rates are tracked by coder and reviewer, so drift shows up in your reporting rather than in an audit.

What happens to the cases it doesn't close?

They arrive in your reviewers' queues as exceptions, ranked so the highest-value and weakest-evidence ones surface first, with the chart already open to the page in question. A coder confirms a finding instead of going to look for it. Second-level review, QA sampling and sign-off run in the same system, and every action is recorded against the person who took it.

Does it remove codes as well as add them?

Yes, and the same validation pass produces both. Codes that are supported but were never submitted come back as adds, ranked by impact with the evidence attached. Codes already submitted that the record does not support come back as deletes, each with the reason it failed. Most retrospective programs only look for what is missing, which leaves the second half of the exposure untouched.

Does Martlet AI replace our coding team?

No. It changes what the team spends its day on. Instead of reading 40 to 50 charts each, your coders work the roughly 5% of cases routed for judgment, run second-level review and QA sampling, and own the standards the platform applies. The same team covers many times the volume, and the institutional knowledge stays in-house rather than at a vendor.

What does it ingest, and what comes back out?

In: clinical notes and PDFs, including scanned documents read with OCR, along with FHIR and HL7 feeds, CCDs, claims extracts and your prior coding, normalized and de-duplicated at ingestion. Out: adds and deletes in the format your submission pipeline expects, the evidence behind each code, and reporting on confirmation and deletion rates by coder, vendor and provider group.

Where does it run, and does PHI leave our network?

It runs inside your own environment — on-premises, in your private cloud, or air-gapped. No charts are shipped out, no outsourced coders touch them, and no PHI leaves your network. Updates ship as versioned releases your team applies on its own schedule, and every decision is recorded with the evidence and the model version behind it.

What happens when CMS guidance changes?

We update the checks to match, so validation follows the rules currently in force rather than the ones that applied when the platform was installed. Because each code is scored against the model and the guidance that apply to the year being coded, a run covering several payment years is never pushed through a single set of rules.

Bring 500 charts. We'll close them while you watch.

Chase prioritization, verification at 99% precision, exceptions queued, and submission deltas out — run on your charts, inside your environment.

  • 95%

    closed automatically

  • 99%

    precision on automated codes

  • 100,000+

    lives on the platform