AI in risk adjustment is judged on what it finds. The question that decides its value is whether a capture can be defended three years later at audit.
Every AI vendor in risk adjustment is answering a question you did not ask. They are answering how much the model finds. You already know finding was never the constraint. Your team has more suspects than it can work and has for years.
The question that decides whether AI in risk adjustment is worth anything is narrower and less flattering. Three years from now, when an auditor pulls a member and asks why a condition was captured, what does your organization say?
There are two possible answers. One of them is a sentence. The other is a project.
Two Kinds of AI, and Only One of Them Answers
The market talks about AI in risk adjustment as a single capability with degrees of sophistication. It is not. There are two designs, and they produce different liabilities.
The first generates diagnoses. It reads a chart, applies a model, and outputs codes. The vendor reports an accuracy rate against a gold standard set. This is the design that demos best, because the output arrives finished, and it is the design most of the AI marketing in this category is selling.
The second surfaces and explains. It reads the same chart, identifies the same conditions, and then does the part that matters: points to the specific line of documentation supporting each one and hands it to a certified coder to confirm, reject, or query. The AI does not assign anything. It removes the search.
| At each step | Generates diagnoses | Surfaces and explains |
|---|---|---|
| What it does with the chart | Applies a model and outputs codes | Points to the documentation |
| What the AI itself decides | Which codes to assign | Nothing. It removes the search |
| Who signs the capture | No one | A certified coder |
| What you hand the auditor | A model version and a confidence score | A chart line and a signature |
Under normal operations these look like variations on a theme, and the first one looks faster. Under RADV they are not the same product. A capture from the second design has a coder who signed it and a document line behind it. A capture from the first has a model version and a confidence score, and neither is a person, and neither is evidence.
Accuracy Is Not the Metric It Appears to Be
Every AI risk adjustment vendor publishes an accuracy figure, including Invent Health, which averages 95 percent coding accuracy in the Coder Workbench. Take all of them, including ours, less seriously than the demo suggests.
Accuracy is measured against a labeled test set, on charts the vendor selected, scored by a definition the vendor chose. It tells you the model performs on that set. It does not tell you the model performs on your population, your provider mix, your documentation habits, or the specialty notes that make up the messiest quarter of your charts.
There is only one accuracy test worth running, and it costs a vendor nothing to agree to. Give the tool charts you have already coded and compare. What did it catch that your process missed. What did it flag that your coders would reject. What did it claim it could prove and then fail to point to a line for. That number is about your organization rather than about the vendor, and any AI vendor unwilling to be measured that way has told you something.
Where the Real Exposure Sits
RADV audits sample members and extrapolate. This is the mechanic that changes the entire risk calculation on AI, and it is routinely underweighted in vendor evaluations.
A handful of indefensible captures in a sample does not cost you those captures. It produces an error rate applied across the contract. So the exposure created by a coding approach is not proportional to how often it is wrong. It is proportional to how often it is wrong in a way you cannot explain when someone asks.
That is what makes autonomous coding a structurally poor fit here regardless of its accuracy rate. A model that is right 97 percent of the time and unable to justify any individual decision has produced a population of captures where you cannot tell the good three percent from the bad. Meanwhile a model that is right slightly less often but attaches a chart line and a coder signature to every capture gives you something to hand the auditor for each one.
The industry has spent two years describing this as a trust problem with AI. It is not a trust problem. It is an evidence problem, and evidence is a design decision a vendor makes before you ever see the product.
Want to see this run against charts you have already coded?
The AI Work That Is Not Coding
Reading charts is the visible application of AI in risk adjustment and the one every vendor competes on. It is not where most organizations have the most to gain, because it addresses a step that is already staffed.
Deciding what to work is unstaffed. Every organization has more identified conditions than coder hours, and the sorting is usually done by RAF value, which is a number that knows nothing about whether the member has a visit scheduled, an active primary care relationship, or an opt-out on file. Modeling closure alongside value is a better use of prediction than reading another chart, and it changes the composition of the worklist rather than its length.
Predicting failure before submission is also unstaffed. Encounter files fail for reasons that repeat and are learnable: the provider whose files always miss a field, the claim type that carries a diagnosis it cannot support, the mapping that quietly drops a code. Catching those before a file leaves is worth more than any rejection report, because a rejection late in a cycle is a condition that may never be recovered.
Neither of these makes for a good demo. Both do more for a risk adjustment program in a year than a faster chart read.
How Invent Health Applies It
Invent Health was built as AI-native, which in practice means the AI runs across the whole loop rather than sitting on top of one step, and it starts with analytics, not coding.
Risk Analytics decides what is worth working. Suspect logic runs across historical, lab, pharmacy, comorbidity, and chart-derived signals, drawing on EMR data through CCD and FHIR. Conditions are ranked by how much they move the score and how likely they are to close, and internal capture is validated against CMS-recognized HCCs in the MOR with RAF progression tracked across payment cycles in the MMR.
The Coder Workbench confirms the condition is real. AI-assisted coding reads the clinical documentation, shows the supporting evidence, and recommends ICD-10-CM codes for a certified coder to confirm and sign. Invent Health’s NLP detects that supporting evidence with 85 percent or higher accuracy out of the box, and it is a two-pass, coder-in-the-loop model. The AI surfaces what matters, a person makes the call, and every code stays linked to the chart for audit-defensible lineage. Coders move faster because the search is gone, not because a review step was removed.
Encounter Submissions makes sure it counts. One engine generates and validates clean submissions, 837P, 837I, and DME for Medicare EDPS and Edge Server XML for ACA, flags missing diagnoses and non-risk-eligible claims carrying risk diagnoses before a file leaves, then reconciles MAO-002 and MAO-004 for Medicare and the Edge Server reports for ACA.
The design principle underneath all three is the same. The AI never invents a diagnosis and never assigns a code. It explains what it found, and a person decides. For organizations working with delegated groups and IPAs, that lineage rolls up to the group level, so the evidence trail holds wherever the documentation happened.
Why This Gets Decided in 2027
The tolerance for unexplainable captures has been higher than it should be because the consequences arrived slowly. That is ending on a published schedule.
The CY 2027 Rate Announcement ends unlinked chart reviews beginning with the 2027 payment year, so a chart review diagnosis has to tie back to an encounter CMS accepted, and the July HPMS memo on CY 2027 risk adjustment implementation sets out what has to be in place. V28 removed a large share of diagnosis codes from mapping and re-based the values, so fewer conditions carry payment and each one carries more weight. And RADV is moving toward every contract every year with findings extrapolated across the contract.
Put together, those three change what an AI capture is worth. Volume matters less, defensibility matters more, and the gap between a capture you can explain and one you cannot is now a number on a repayment notice. An organization choosing an AI approach this year is choosing what it will be able to say in an audit three years from now.
Test It on Charts You Have Already Coded
Point any AI risk adjustment tool at a period your team has already worked. Look at what it catches that you missed, what it flags that your coders would reject, and whether it can show you the line of documentation behind every condition it claims. The third one is the answer.
Frequently asked questions
How is AI used in risk adjustment?
AI reads clinical documentation to surface conditions that were never coded, ranks which gaps are worth working, and validates encounters before submission. The strongest applications recommend and explain rather than assign, leaving the coding decision with a certified coder who confirms each condition against documented evidence.
Is AI coding safe for RADV audits?
It is safe when every capture traces to a documented chart line and a human decision. Tools that surface supporting evidence and keep a coder in the loop produce defensible captures. Autonomous tools that assign diagnoses do not, and because RADV findings are extrapolated across a contract, indefensible captures get expensive quickly.
Can AI assign HCC codes on its own?
It can, but doing so creates audit exposure that usually outweighs the efficiency. A capture with no human decision and no cited evidence cannot be explained when an auditor asks why it was made. Coder-in-the-loop designs keep the speed while preserving the trail an audit requires.
What is explainable AI in risk adjustment?
Explainable AI shows the reasoning behind each recommendation, specifically the line of clinical documentation supporting a condition. In risk adjustment this matters more than in most fields, because a capture that cannot be traced to evidence and a coder is a capture that cannot be defended at audit.
How accurate is AI at HCC coding?
Vendor accuracy figures are measured against labeled test sets the vendor selects, so they indicate capability rather than performance on your population. The meaningful test is running the tool against charts your team has already coded and comparing what it caught, what it overcalled, and what it could actually prove.
Does AI replace medical coders in risk adjustment?
No. AI removes the manual chart search and surfaces conditions with their supporting evidence, but certified coders still confirm and sign every capture. The value is making skilled coders faster and more accurate, not removing human judgment from a process CMS audits on documentation.
How much faster is AI-assisted coding?
Speed gains come from eliminating the search for supporting documentation rather than from cutting review steps. When evidence arrives alongside each suspected condition, coders spend their time deciding instead of hunting, so throughput and accuracy improve together rather than trading against each other.
What should health plans ask AI risk adjustment vendors?
Whether the tool assigns codes or recommends them, whether it cites the specific documentation behind each condition, whether accuracy can be tested on the organization’s own previously coded charts, and whether the AI extends beyond chart reading into gap prioritization and encounter validation.
Can provider groups and health systems use AI for risk adjustment?
Yes. IPAs, medical groups, and health systems carrying risk run coding and documentation programs of their own, and many health systems own health plans outright. The audit standard is identical regardless of who holds the contract, so the evidence requirement on AI captures does not change.
Beyond coding, where else does AI help in risk adjustment?
In deciding which gaps to work and in predicting encounter failures before submission. Both are typically unstaffed. Modeling closure likelihood alongside score impact changes what a worklist contains, and catching predictable submission errors early prevents conditions from being lost after the coding work is done.
Does AI create RADV risk?
The technology does not, but the design choice can. Systems that produce captures no one can explain create exposure that extrapolation multiplies across a contract. Systems that attach evidence and a coder signature to every capture reduce audit work, because responding becomes a lookup rather than an investigation.
How does Invent Health use AI in risk adjustment?
AI runs across analytics, coding, and encounter submission on one platform. It ranks conditions by score impact and closure likelihood, surfaces supporting evidence for a certified coder to confirm and sign in the Coder Workbench, and validates encounters before submission. It never invents a diagnosis or assigns a code.