Information Extraction Using Spatial Context Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information extraction technologies face challenges in effectively processing semi-structured text documents, which heavily rely on spatial context due to varying data locations and formats, leading to inefficiencies in extracting relevant information and ignoring non-relevant data.

Innovation Solution

A method and system for information extraction from semi-structured documents that utilize spatial context, involving a framework with entity and relational models, allowing for real-time learning and training with fewer samples, and generating structured data formats like information graphs, which can handle diverse document types such as loss run documents in the insurance industry.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If statistical machine learning techniques are used for information extraction from unstructured text documents, then extraction capability is improved, but the approach fails to handle semi-structured documents with varying spatial contexts

Engineering Contradiction:
Improvecapability to handle different document typesVSAvoidextraction accuracy for semi-structured documents
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system segments the information extraction process into distinct modules: entity model retrieval, relational model retrieval, spatial context analysis, and data element extraction. This segmentation allows each module to specialize in handling specific aspects of document processing, enabling reliable extraction from semi-structured documents with varying spatial contexts while maintaining versatility across document types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial context as an intermediary element that mediates between the textual content and the extraction process. By analyzing spatial relationships and positions of text elements, the system can accurately identify and extract information from semi-structured documents regardless of their specific layout variations, resolving the contradiction between handling diversity and maintaining extraction accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Extent of automation

If conventional machine learning methods are used for information extraction, then processing is automated, but manual data entry processes remain extensive and time-consuming

Engineering Contradiction:
Improveautomation of information extractionVSAvoidtime required for manual data entry
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by retrieving pre-defined entity models and relational models before processing the actual document extraction. These pre-fetched models contain the necessary structural information and spatial relationship definitions that enable the system to automatically extract and structure data without requiring extensive manual intervention or time-consuming processing, thus reducing the time loss while maintaining high automation levels.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If information extraction is performed on semi-structured documents with varying formats, then versatility is improved, but extraction complexity and difficulty increase

Engineering Contradiction:
Improvehandling of varying document formatsVSAvoidcomplexity of extraction framework
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal extraction framework that uses a single set of tools and processes to handle multiple document formats and semi-structured layouts. The entity models and relational models serve as universal templates that can be applied across different document types, eliminating the need for separate extraction mechanisms for each format and thereby reducing overall system complexity while maintaining high versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If key pieces of information are identified in semi-structured documents, then extraction precision is improved, but difficulty in distinguishing important from non-important information increases

Engineering Contradiction:
Improveprecision of information extractionVSAvoiddifficulty of identifying important information
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system replaces manual mechanical identification of important information with an automated model-based approach. Entity models and relational models provide structured definitions that automatically distinguish between important and non-important information based on pre-established criteria. This substitution eliminates the difficulty of manual identification while maintaining high precision in extracting only the relevant key pieces of information from semi-structured documents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11106906B2Systems and methods for information extraction from text documents with spatial context
Publication Date: 2021.08.31 AMERICAN INTERNATIONAL GROUP INC
  • US11106906B2 patent drawing
  • US11106906B2 patent drawing
  • US11106906B2 patent drawing

AI summary

Performing information extraction from an electronic document is disclosed. A method comprises: receiving a semi-structured input document; retrieving an entity model that provides one or more domain variable definitions for one or more domain variables, wherein the entity model and the input document correspond to a common domain; determining that the input document includes an entity that satisfies a first domain variable definition corresponding to a first domain variable; retrieving a relational model that provides, for the first domain variable, one or more relational definitions comprising spatial restrictions for one or more values corresponding to the first domain variable; extracting one or more data elements from the input document that satisfy the one or more relational definitions; and generating an information graph having a structured data format, wherein the one or more data elements extracted from the input document correspond to the first domain variable in the structured data format.