Taxonomy-driven multipass extraction of structured data from unstructured documents

A taxonomy-driven multipass extraction system addresses inefficiencies in document analysis by segmenting and adaptively scheduling extraction passes, reducing computational and bandwidth usage while enhancing accuracy and enabling efficient transformation and comparative editing of legal and financial documents.

WO2026117649A1PCT designated stage Publication Date: 2026-06-04CENTARI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CENTARI INC
Filing Date
2025-11-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Current systems for extracting critical datapoints from legal and financial documents impose significant computational burdens due to inefficient document analysis workflows that treat each document as an isolated full-text problem, leading to excessive compute cycles, storage consumption, and bandwidth usage, and lack mechanisms to adapt based on input complexity.

Method used

A taxonomy-driven multipass extraction system that segments documents into snippets, generates semantic vector representations, and adaptively schedules extraction passes based on document complexity and confidence, using domain-specific taxonomies and large language models to enhance accuracy and reduce unnecessary computation.

Benefits of technology

The system reduces computational and bandwidth usage while improving accuracy by segmenting documents once, reusing vector embeddings, and dynamically scheduling extraction passes, enabling efficient transformation of unstructured documents into structured data and facilitating comparative editing across document sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025057222_04062026_PF_FP_ABST
    Figure US2025057222_04062026_PF_FP_ABST
Patent Text Reader

Abstract

An intelligent document analysis platform transforms unstructured documents into structured, searchable data by aligning them with a domain specific taxonomy. The system may segment each document into snippets, store vector embeddings and use a structure map to target key sections. For every datapoint defined in the taxonomy, the platform may automatically retrieve semantically similar snippets, construct prompt to a language model and extracts candidate values with supporting citations. A normalization phase may resolve conflicts and enforce categorical answer formats, while confidence scores may guide iterative refinement and fallback strategies. Users may receive normalized datapoints with highlighted citations via an interactive interface, and the platform can logs feedback to refine future extractions.
Need to check novelty before this filing date? Find Prior Art