Home Services ApproachFor Partners Insights About Contact
Point of View  ·  Insurance & AI Governance

Insurance AI Has Entered Its Examination Era

July 2026 · 12 min read · NAIC  ·  Colorado  ·  NY DFS
Maria Chaloux
Dr. Ian McCulloh
Maria Chaloux & Dr. Ian McCulloh, PhD Octant Advisory  ·  Founder & Managing Partner  ·  Chief AI Strategy Officer

Artificial intelligence is no longer a future-state question for insurers. It is embedded in underwriting, pricing, claims, fraud detection, customer service, distribution, and operational triage. In Conning’s 2025 survey, 90% of insurers were in some stage of generative-AI evaluation and 55% in early or full adoption, while machine-learning adoption stood at 74%. The NAIC’s own line surveys put AI or machine-learning use, or planned use, at 88% of auto insurers and 84% of health insurers. The technology is already load-bearing.

What has changed is not simply that regulators noticed. It is that they are asking a harder and more consequential question.

Can the insurer prove, system by system and decision by decision, that its use of AI is governed, tested, accountable, and fair?

For several years, AI governance in insurance could be framed as responsible innovation: principles, policy statements, ethical commitments, model-risk practices, board updates. That period is ending. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, Colorado’s rules under SB21-169, and the New York DFS Circular Letter No. 7 are moving the industry into an examination era. The common demand is not that insurers stop using AI. It is that they produce evidence that AI is governed before, during, and after deployment.

This is an operating-model challenge more than a policy one. Most carriers did not build their AI capabilities with regulatory examination as the design center. Adoption grew through function-specific initiatives: actuarial teams modernized pricing and segmentation, claims teams deployed triage and severity tools, fraud teams added anomaly models, business units bought vendor platforms, and innovation groups experimented with generative AI. The result is not necessarily reckless. In many cases, it is fragmented. And fragmentation is precisely what the new supervisory posture exposes.

Insurance AI governance has become an evidentiary discipline. A carrier that cannot show who owns each AI system, what data it uses, how it was tested, how unfair discrimination was assessed, how vendors are controlled, and how the board gets independent visibility is no longer merely immature. It is exam-vulnerable.

Five implications for boards:

The regulatory center of gravity has moved

The most important development in insurance AI regulation is not any single rule. It is the convergence among them.

The NAIC Model Bulletin, adopted in December 2023, gives state insurance departments a common supervisory template. It expects insurers to maintain a written Artificial Intelligence Systems Program covering governance, risk management, internal controls, testing, monitoring, documentation, third-party oversight, and board or senior-management accountability. It is framed as guidance, but its practical force is greater: it describes what examiners may ask to see. As of mid-2026, 24 states and the District of Columbia had adopted the bulletin or substantially similar guidance, with the total passing half of all states.

Colorado has gone further. Under SB21-169, its governance-and-risk-management regulation (Regulation 10-1-1) is final for life insurers, effective November 2023, and an amended regulation extended that governance framework to private passenger auto and health, effective October 15, 2025. Colorado matters nationally because it turns abstract governance ideas into concrete obligations: inventory, accountability, consumer-harm prevention, risk management, and documentation. Its quantitative-testing requirements, which would operationalize disparate-impact testing and contemplate estimation methods such as BIFSG, remain in rulemaking rather than final, and are furthest along for life insurance.

New York’s Department of Financial Services took a serious but differently structured approach. Circular Letter No. 7, issued July 11, 2024, addresses external consumer data, AI systems, and other predictive models in underwriting and pricing. It emphasizes actuarial validity, unfair and unlawful discrimination, governance, transparency, documentation, and oversight of third-party tools. The message is direct: an insurer may use complex data and models, but it must be able to demonstrate that the resulting decisions are valid, controlled, and not discriminatory.

These regimes differ in legal form, scope, and detail. Read together, they establish a common supervisory architecture: a complete inventory of AI systems and predictive models that affect regulated decisions; senior-management and board accountability for AI risk; lifecycle governance from design through deployment, monitoring, and retirement; testing for unfair discrimination, including where models use facially neutral variables; documentation sufficient to support examination, validation, and remediation; oversight of third-party and vendor models as if the insurer built them; and evidence that business, actuarial, data science, compliance, legal, and risk functions have defined decision rights.

The center of gravity has moved from “have an AI policy” to “prove how AI is governed in practice.” In 2026 the NAIC began piloting an AI Systems Evaluation Tool to give market-conduct examiners a standardized way to review a carrier’s AI governance, the mechanism that turns the bulletin’s expectations into exam questions.

Exhibit: the regimes at a glance. Three regimes differ in legal form, scope, and detail, but read together they establish a common supervisory architecture.

RegimeScope, core requirement, and what an examiner may expect
NAIC Model BulletinScope: all lines; a model framework adopted state by state. Requirement: a written AI Systems Program covering governance, risk management, testing, monitoring, documentation, third-party oversight, and board accountability. Evidence: AI inventory, governance charter, testing records, vendor-oversight files, board reporting. Board implication: stand up enterprise AI governance an examiner can review.
Colorado SB21-169 / Reg. 10-1-1Scope: life (2023); governance framework extended to private passenger auto and health (Oct 15, 2025). Requirement: a binding governance and risk-management framework for external data, algorithms, and predictive models, with documentation and reporting to the Division. Evidence: ECDIS and model inventory, governance documentation, progress reports. Board implication: treat Colorado as the leading edge; quantitative-testing rules are still in development.
NYDFS Circular Letter No. 7Scope: New York-authorized insurers; underwriting and pricing. Requirement: demonstrate actuarial validity and the absence of unfair or unlawful discrimination, plus governance, transparency, and third-party oversight. Evidence: fairness and discrimination analysis, actuarial justification, data-source documentation, vendor controls. Board implication: be able to defend that AI-driven decisions are valid and non-discriminatory.
Colorado AI Act (SB24-205), contextualScope: cross-sector high-risk AI; developers and deployers; effective June 30, 2026 (non-insurance-specific). Requirement: reasonable care to avoid algorithmic discrimination; risk management, impact assessments, disclosures. Evidence: risk-management program, impact assessments, consumer notices. Board implication: a broader overlay to watch; the effective date has already slipped.

Why this is not ordinary model risk management

Many insurers will treat the new AI expectations as an extension of existing model risk management. That is a mistake. Traditional model governance is necessary and not sufficient. Insurance AI governance reaches beyond technical model validation. It covers the business decision the model is embedded in, the data supply chain feeding it, the vendor or platform that may operate it, the consumer impact of the decision, the fairness methodology used to test outcomes, the documentation retained for examiners, and the board’s visibility into aggregate AI risk. The object of governance is the decision the algorithm sits inside, not the algorithm itself.

That distinction has teeth. A pricing model may be statistically strong and still raise unfair-discrimination concerns through proxy variables. A claims-triage model may cut cycle time and still produce disparate review burdens if historical data reflect uneven treatment. A fraud model may be useful and poorly governed if it comes from a vendor whose data sources, feature engineering, and monitoring are opaque. A generative-AI tool may make no final coverage decision and still create regulated risk if it drafts denial rationales, summarizes claim files, or shapes adjuster judgment.

Can the carrier demonstrate that the AI-enabled decision is lawful, explainable enough for its use, actuarially or operationally justified, monitored over time, and controlled by accountable humans?

That is why AI governance cannot sit solely with data science, which understands the model but not the full conduct-risk picture, or solely with compliance, which understands the obligation but not the technical evidence. Actuarial, underwriting, claims, legal, compliance, enterprise risk, IT, data governance, procurement, internal audit, and business leadership each hold part of the answer. Regulators will increasingly expect one coherent answer.

The patchwork is the operating burden

State AI regulation is usually described as a patchwork. That is accurate but incomplete. The bigger problem is that the patchwork is an operating burden, not just a legal-tracking one. A national carrier writing across many states cannot govern AI separately for each jurisdiction without duplication, inconsistency, and control failure, and it cannot govern to the lowest common denominator. The practical answer is a single enterprise AI governance program strong enough to satisfy the strictest regimes the carrier operates in, with state-specific overlays where required.

That requires a different management posture. The useful question is not what each state requires. It is what enterprise control environment would let the carrier answer any examiner’s AI questions with confidence. That environment has to reconcile real tensions:

Carriers that resolve these centrally will scale AI with confidence. Carriers that resolve them locally, project by project, will accumulate hidden examination risk.

The gap most carriers need to close

The central weakness in many insurers is not a lack of AI activity. It is a lack of AI control coherence. AI use has grown faster than the governance architecture around it. One function keeps a model inventory, another a vendor inventory, a third data lineage, and compliance keeps policy attestations, individually useful, collectively insufficient. When an examiner asks where AI touches regulated decisions, the carrier often has no single authoritative view.

In the regulators’ own 2025 survey, 92% of health insurers said they had AI governance principles in place, and nearly one in three still did not regularly test their models for bias or discrimination.

The evidence sits in the regulators’ own data. In the NAIC’s 2025 health survey, 84% of insurers reported using AI or machine learning and 92% said they had governance principles modeled on the NAIC’s, yet nearly one-third still did not regularly test their models for bias or discrimination, more than a year after the NAIC recommended it. The same survey found 37% of health insurers using AI for prior authorization, 44% for claims adjudication, and 56% for utilization management: the consumer-facing decisions examiners care most about. The most common gaps are predictable:

These gaps decide whether the carrier can answer the examiner’s basic chain of questions: where do you use AI; which uses affect underwriting, pricing, claims, fraud, or other regulated decisions; who owns each system; what data does each system use; how did you validate the system before deployment; how did you assess unfair discrimination; how do you monitor drift, performance, and outcomes; what vendors are involved, and how do you oversee them; what happens when testing finds a problem; and what does the board see, and how does it know the reporting is reliable? An insurer that cannot answer those questions is not simply missing documentation. It is missing an operating model.

Unfair discrimination is the hardest test

The most sensitive issue in insurance AI governance is unfair discrimination, and it is where superficial governance is most likely to fail. Insurance depends on classification; insurers lawfully distinguish among risks every day. The difficulty is that AI and external data can create or amplify distinctions that are facially neutral but functionally problematic. Geography, credit attributes, purchasing behavior, online activity, device data, occupation, education, and household composition can correlate with protected characteristics or historical inequities, and machine learning can detect and operationalize those correlations even when protected-class variables are excluded.

That is why “we do not use race” is no longer enough. Regulators are focused on proxy discrimination, disparate impact, data provenance, actuarial justification, and the ability to explain why a variable or output is appropriate for the decision at issue. Colorado has drawn particular attention for moving toward quantitative testing; its proposed testing rule would use race and ethnicity estimation methods such as Bayesian Improved First Name Surname Geocoding (BIFSG) where direct protected-class data are unavailable. These proxy methods are tools with limitations, not settled answers; they carry measurement error and require careful methodological governance. Carriers need a defensible testing protocol, not a single statistical ritual. A credible program should define which systems require testing, which protected classes are assessed, which estimation methods are used, which outcome metrics and control variables are legitimate for the specific decision, which thresholds trigger escalation, how actuarial validity and fairness are reconciled, and how results are documented for management, the board, and examiners.

This is a governance question as much as a technical one, because statistical findings force business decisions. Someone must decide whether a disparity is explainable, whether a feature should be removed, whether a model should be recalibrated, whether a vendor must provide more evidence, whether a filing is affected, whether consumers need notice, and whether a regulator should be told. Those decisions require authority.

Vendor models are not a safe harbor

One of the fastest ways insurers have adopted AI is through vendors, and one of the fastest ways AI governance becomes fragile. Vendor tools are embedded across the value chain: data enrichment, risk scoring, underwriting workbenches, claims estimation, fraud detection, litigation analytics, segmentation, call-center automation, document ingestion, and generative-AI assistants. They may be valuable and they may be opaque; a carrier may not know all training data, feature engineering, model updates, monitoring results, or downstream uses.

Regulators have been clear on the principle: outsourcing does not outsource accountability. If a vendor model affects a regulated decision, the insurer remains responsible for governing it. Insurers report the exposure indirectly; in the NAIC’s 2025 health survey, 85% said their third-party contracts contain no terms that limit transparency to regulators, which means roughly one in seven does, precisely where the carrier is still on the hook. AI vendor management therefore has to go beyond ordinary procurement and cybersecurity review. It should include pre-contract diligence on model purpose, data sources, limitations, validation, bias testing, and monitoring; contractual rights to obtain documentation sufficient for examination; audit rights and cooperation for regulatory inquiries; change-management notice when the vendor modifies data, models, features, thresholds, or decision logic; restrictions on secondary data use, model training, and subcontractors; requirements for incident reporting and discrimination findings; and exit and contingency plans for high-risk AI systems.

The board-level point is plain. A vendor model that cannot be explained, tested, monitored, or documented to regulatory standards is an imported control weakness, not a shortcut.

What good looks like

A mature insurance AI governance program is a repeatable management system rather than a static binder. It begins with an authoritative AI inventory covering traditional predictive models, machine-learning systems, generative-AI tools, embedded vendor models, and automated decision systems. Each entry should identify the business owner, technical owner, vendor involvement, data sources, decision use, consumer impact, regulatory exposure, validation status, monitoring cadence, and documentation location. The inventory should feed a risk-tiering process: not every AI use needs the same controls; a back-office summarization tool is not a pricing model, and a marketing-analytics tool is not a claims-denial recommendation engine. High-risk uses should get enhanced review, testing, documentation, approval, monitoring, and board visibility.

The program should define decision rights. The business owns the use case and outcomes. Data science or actuarial owns technical development and validation. Compliance and legal own interpretation of regulatory obligations. Risk management owns independent challenge. Internal audit tests whether the program operates as designed. The board oversees the framework, receives meaningful reporting, and challenges management where risk accumulates. It should establish lifecycle controls: intake and classification before development or purchase; data review before model training or vendor onboarding; pre-deployment validation and unfair-discrimination testing; formal approval for high-risk systems; post-deployment monitoring for performance, drift, consumer outcomes, and complaints; periodic revalidation; incident and remediation protocols; and retirement or suspension criteria.

Finally, it should generate examination-ready evidence, artifacts the carrier can actually produce: the AI inventory, governance charter, and risk-tiering standard; model and data-source documentation; validation reports; unfair-discrimination testing protocols and results; vendor due-diligence files; approval records and monitoring dashboards; issue logs and remediation evidence; board and committee reporting; and internal audit or independent assessment findings. The standard is not perfection. It is disciplined, documented, risk-based governance that survives independent review.

The board’s role

Boards do not need to become model validators. They do need to become serious overseers of AI as a regulated business capability. The board should press management on seven questions:

Those questions should drive tangible reporting. A mature board pack should show AI inventory counts by business area and risk tier, high-risk systems by line of business and jurisdiction, validation and revalidation status, unfair-discrimination testing coverage and findings, vendor-model exposure, open remediation items, regulatory developments, and incidents, complaints, overrides, and control exceptions, alongside internal audit or independent assessment results. The board does not need every technical detail. It needs enough independent signal to know whether management’s AI program is real.

Why independent baselines matter

Self-assessment is useful and rarely enough, for three reasons. The first is credibility: regulators, boards, and senior executives discount governance assurances that come only from the teams deploying the technology. The second is integration: AI governance crosses functions that each see part of the picture, and an independent baseline can find where actuarial validation, data governance, compliance policy, vendor oversight, legal interpretation, model monitoring, and board reporting fail to connect. The third is timing: a carrier that waits for an examination to learn where its AI governance is weak will discover the gap under the least favorable conditions. A baseline done before scrutiny gives management time to remediate, sequence investment, clarify accountability, renegotiate vendor terms, and improve board reporting. A useful independent baseline answers five questions: where AI is actually used, which uses create the highest regulatory and consumer-conduct exposure, which controls exist and where they are fragmented, which documentation would fail under examination, and what should be fixed first.

The choice in front of carriers

Insurance AI has crossed a threshold. The question is no longer whether carriers can build, buy, or deploy AI. They can. The question is whether they can govern it at the standard regulators are now making visible. The next generation of AI advantage in insurance will not belong to the carriers with the most models, the most data, or the fastest pilots. It will belong to the carriers that can make AI examinable: inventoried, risk-tiered, tested, documented, monitored, vendor-controlled, and overseen by a board with an independent line of sight. That is the new standard.

About Octant Advisory

Octant Advisory helps organizations convert AI ambition into measurable performance. We work with leadership teams on the governance, workforce, and operating-model changes that make AI investment pay off. Ian McCulloh built and led a federal AI practice at national scale and directs AI executive education at Johns Hopkins. Maria Chaloux built the leadership team behind Accenture Federal Services’ growth over a decade, and spent two decades helping organizations identify and develop the leaders who drive transformation. Learn more at octantadvisory.com.

Notes & Sources

Conning, “2025 Survey on AI & Insurance Technology: The C-Suite Verdict,” June 25, 2025.

NAIC AI/ML market-conduct surveys by line (auto, homeowners, life, health), 2022-2025; auto ~88%, health 84% (2025 Health Insurance AI/ML Survey Report).

NAIC, 2025 Health Insurance AI/ML Survey Report (May 9, 2025): 84% AI/ML use; 92% with governance principles; nearly one-third not regularly testing for bias; prior authorization 37%, claims adjudication 44%, utilization management 56%; 85% of third-party contracts contain no terms limiting transparency to regulators.

NAIC, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers (adopted December 2023); state-adoption map: 24 states and DC as of mid-2026. Consult the NAIC live adoption map for the current count.

NAIC AI Systems Evaluation Tool, market-conduct examination pilot, 2026.

Colorado SB21-169 (2021); Regulation 10-1-1 for life insurers (effective Nov 14, 2023); amended Regulation 10-1-1 extending to private passenger auto and health (effective Oct 15, 2025).

Colorado Division of Insurance, draft Algorithm and Predictive Model Quantitative Testing Regulation (3 CCR 702-10), proposed 2023; not finalized; life-insurance focus; contemplates BIFSG estimation.

New York DFS, Insurance Circular Letter No. 7 (July 11, 2024): external consumer data, AI systems, and predictive models in underwriting and pricing.

Colorado SB24-205, Consumer Protections for Artificial Intelligence (effective June 30, 2026): cross-sector, not insurance-specific.

Supporting frameworks: NIST AI Risk Management Framework; ISO/IEC 42001; literature on proxy discrimination and protected-class estimation (BISG, BIFSG).

This document is for general information and does not constitute legal advice.

From Analysis to Action

Make Your AI
Examinable.

Octant Advisory helps insurers establish an independent baseline of where their AI governance actually stands against the expectations regulators are now applying. We do not build or implement the models, which is what lets us give boards and management a clear, independent view of whether AI governance is real, where it is weak, and what must be fixed first.

Back to Insights