Insurance's GenAI unstructured data challenge
The insurance industry is built on data: every policy written, claim filed, and risk assessed depends on it. Yet for most Property and Casualty, Life and Annuity, and even Specialty lines carriers — the vast majority of their data remains effectively hidden from the people now asked to transform the business, and from the AI Agents expected to deliver historic leaps in processing speed and productivity gains.
The hidden data is unstructured data or files, such as adjuster notes, broker submissions, litigation demand packages, medical records, inspection reports, and process documents, which account for more than 90% of all enterprise data in insurance, according to analysis by AXA XL's Head of Process Optimization, Data and AI. These documents sit in PDFs, emails, SharePoint, file servers, scanned forms, and handwritten notes — rich with underwriting intelligence and claims insight — but locked away from the analytical and AI systems that could put these insights to work.
The problem by insurance segment
The unstructured data problem is not uniform across all insurance segments. It cuts deepest in the places where the workflows are most document-intensive.
In Property and Casualty underwriting, 40% of the average underwriter's time is consumed by administrative and non-core tasks, according to Accenture research, much of it spent manually extracting, re-keying, and reconciling data from unstructured broker submissions. Commercial lines carriers often face surges in submissions they physically cannot process, leaving revenue on the table and straining broker relationships.
In Life and Annuity, the gap is arguably widest. Accenture's "Underwriting Rewritten" research found that only 12% of Life insurance underwriting executives say their organizations use unstructured data to a "large" or "very large" extent — compared to 63% in Commercial P&C and 69% in Personal/Retail P&C. Medical records, attending physician statements, and financial documents remain labor intensive to process, creating a structural ceiling on how fast L&A carriers can underwrite, onboard, and serve customers.
In Claims, the problem is even more consequential. Litigation demand packages, medical histories, adjuster narratives, and damage photographs all arrive as unstructured content. Insurance Thought Leadership, citing Accenture and Swiss Re, estimates that the industry's reliance on manual experts to handle this administrative work will result in $85–$160 billion in lost efficiency by 2027. Swiss Re further found that insurers using AI to extract insights from unstructured data can see a 12–25% improvement in their loss ratio compared with those that don't.
The AI readiness gap
Leaders in the insurance industry are aware of the opportunity. EY's 2025 GenAI Insurance Survey of 100 senior leaders across P&C and L&A found that 74% of insurers identify predictive analytics as a key investment area for underwriting and claims, with 77% having allocated budget specifically to GenAI initiatives. Moreover, Conning's 2025 C-Suite AI Survey found 90% of US insurers — equally split between P&C and L&A — in some stage of GenAI evaluation.
But ambition is running ahead of readiness. IRMI's analysis of insurance data management notes that the focus is now shifting from experimentation to scalable execution — and that data governance, integration, and trust are the primary obstacles standing between pilots and production.
The root cause is straightforward: AI models, whether generative or agentic, need clean, certified, accessible data to function reliably. Data structured in databases, spreadsheets, and reports were reliable with dependable data classification, definitions, and business context captured within ongoing data governance. But unstructured data, like a PDF form and the 10s or 100s of millions of them across an insurance company, have none of that information, just a filename, create/modify date, and size. Unused and unusable for any initiative, but now GenAI has put a huge spotlight on this hidden data. Unfortunately, files in their current raw form are like Crude oil. Not yet refined to fuel for the minivan or the AI rocket ship.
The agentic AI stakes
Agentic AI has huge promise — systems capable of reasoning across multiple steps, orchestrating workflows, and acting with minimal human intervention. McKinsey estimates that generative and agentic AI could unlock $50–$70 billion in additional revenue for the insurance industry, but only for carriers that can feed these systems the data require.
IRMI frames the dependency directly: agentic AI is data-hungry, and insurers who cannot deliver governed data in real time will struggle to operationalize it safely. An AI agent that continuously adjusts underwriting thresholds based on catastrophe alerts, or a claims agent that ingests litigation packages and medical records to drive faster resolution, is only as reliable as the data it can access and trust.
The results for early movers are already compelling. Accenture research shows carriers that have addressed their data foundations can now process 100% of broker submissions, double their submission-to-quote rates, and reduce premium leakage from missed underwriting controls. Deloitte's research on L&A specifically highlights how NLP tools extracting insights from unstructured medical records are significantly accelerating and improving the accuracy of risk assessments in life underwriting.
Fraud and claims verification: A high-stakes use case
The atrophy of unstructured data is highly visible — and very quantifiable in claims fraud. The FBI estimates that non-health insurance fraud costs the US more than $40 billion annually, a figure the Coalition Against Insurance Fraud puts even higher at over $300 billion when healthcare fraud is included. That cost is ultimately borne by policyholders in the form of higher premiums. The challenge is that the evidence of fraud — inconsistent repair invoices, conflicting medical narratives, suspicious adjuster notes, staged-loss documentation — lives almost entirely in unstructured content that legacy rules-based systems cannot read.
Generative AI changes the equation fundamentally. Unlike rules-based fraud detection that relies on predefined triggers, AI systems can analyze the textual content of claims submissions, claimant communications, damage photographs, and supporting documentation to identify narrative inconsistencies, linguistic anomalies, and contextual patterns that signal fabrication or exaggeration — signals invisible to structured data systems. Insurers deploying AI-enhanced fraud detection have reported 30–50% improvements in fraud identification rates, translating into hundreds of millions of dollars in savings annually for large carriers. Deloitte's research goes further, predicting that by deploying AI-powered multimodal technologies across the claims lifecycle, P&C insurers could reduce fraudulent claims and save between $80 billion and $160 billion by 2032. Critically, Insurance Thought Leadership notes that accelerating claims document handling with AI can reduce processing time by up to 90% — meaning legitimate claims are paid sooner, while fraudulent ones are caught earlier. Both outcomes depend on the same prerequisite: unstructured claims data and files that are classified, tagged, certified, accessible, and AI-ready.
Extending the data foundation to be fully AI-ready
The data exists. The AI capability exists. And the processes are in place in your data governance program. The gap between them is a mountain of unstructured data — one that can be scaled with a deliberate investment in classifying, tagging, governing, and preparing the files carriers have spent decades accumulating.
Most insurers have traditional tools to scan documents for sensitive information and ensure secure handling, but these fall far short of what GenAI needs. Today’s advanced scanning tools generate a metadata profile for each file, i.e., what kind of document it is (claims form, medical report, etc.), who the provider or adjuster is, who the policyholder is, as well as the dates, types, and values embedded in the document. Further, Graph and Vector databases used for AI require files to be provided in chunks with metadata for each chunk. The metadata profiles are also delivered to the semantic layer that makes Agentic AI scalable and, more importantly, accurate.
As IRMI puts it: if the last decade in insurance was about digital transformation, the next will be about data activation and AI enablement. Carriers that extend their data foundation to include unstructured data will capture the operational gains, loss ratio improvements, and competitive separation that AI promises. Those that don't will find themselves with sophisticated AI tools and no reliable data to run them on. Like a Ferrari with a can of muddy water for fuel.
Keep up with the latest from Collibra
I would like to get updates about the latest Collibra content, events and more.
Thanks for signing up
You'll begin receiving educational materials and invitations to network with our community soon.