Taxonomy-driven multipass extraction of structured data from unstructured documents
A taxonomy-driven multipass extraction system addresses inefficiencies in document analysis by segmenting and adaptively scheduling extraction passes, reducing computational and bandwidth usage while enhancing accuracy and enabling efficient transformation and comparative editing of legal and financial documents.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CENTARI INC
- Filing Date
- 2025-11-26
- Publication Date
- 2026-06-04
AI Technical Summary
Current systems for extracting critical datapoints from legal and financial documents impose significant computational burdens due to inefficient document analysis workflows that treat each document as an isolated full-text problem, leading to excessive compute cycles, storage consumption, and bandwidth usage, and lack mechanisms to adapt based on input complexity.
A taxonomy-driven multipass extraction system that segments documents into snippets, generates semantic vector representations, and adaptively schedules extraction passes based on document complexity and confidence, using domain-specific taxonomies and large language models to enhance accuracy and reduce unnecessary computation.
The system reduces computational and bandwidth usage while improving accuracy by segmenting documents once, reusing vector embeddings, and dynamically scheduling extraction passes, enabling efficient transformation of unstructured documents into structured data and facilitating comparative editing across document sets.
Smart Images

Figure US2025057222_04062026_PF_FP_ABST