LLM Document Data Extraction and Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting key-value pairs from documents are labor-intensive, require extensive manual annotation or rule-based systems, and are sensitive to language and domain variations, limiting their adaptability and efficiency in information retrieval and knowledge management.

Innovation Solution

A method using large language models (LLMs) to extract key-value pairs from documents by indexing document content with taxonomies, conducting OCR, and transforming text content into searchable pairs, allowing for plain language search queries to be converted into structured database queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional hand-crafted rules or pattern matching techniques are used for information extraction, then extraction accuracy for specific domains can be maintained, but labor intensity and maintenance cost increase significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidlabor intensity
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual hand-crafted rule creation and pattern matching with machine learning-based automatic extraction systems. The system uses trained models to automatically identify and extract key-value pairs from documents, substituting the mechanical process of manual rule development with automated computational processes that reduce labor intensity while maintaining extraction accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from static hand-crafted rules to dynamic machine learning models that can adapt to different document types and domains. By training models on domain-specific data, the system adjusts its extraction parameters and patterns automatically, reducing the need for manual rule maintenance while preserving accuracy across varying document structures.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If machine learning-based techniques are used for information extraction, then adaptability to different document types improves, but requirement for large amounts of labeled training data increases

Engineering Contradiction:
Improveadaptability to document typesVSAvoidtraining data requirement
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent implements a universal information extraction system using large language models that can handle multiple document types and domains with a single model architecture. The LLM is trained on diverse, multi-domain data, enabling it to generalize across different document formats and extraction tasks without requiring separate models for each domain, thus reducing overall training data requirements while maintaining adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary training on comprehensive, multi-domain datasets to pre-establish the model's extraction capabilities across various document types. This preliminary action of training on diverse data beforehand enables the model to adapt to new document types with minimal additional training, reducing the immediate training data requirement for specific applications.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If machine learning models are used for information extraction, then performance sensitivity to language and domain variations decreases, but cost and time for creating training data increases

Engineering Contradiction:
Improveperformance stabilityVSAvoidtraining data creation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the manual process of creating and maintaining domain-specific training data with automated approaches using pre-trained large language models. The LLM's inherent understanding of multiple languages and domains, acquired during pre-training, reduces the need for extensive domain-specific data creation, thereby decreasing training time while maintaining performance stability across different languages and domains.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Ease of operation

If structured information is extracted from unstructured documents, then granular searching capability improves, but complexity of the extraction system increases

Engineering Contradiction:
Improvesearching capabilityVSAvoidextraction system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces complex rule-based extraction systems with machine learning-based automatic extraction that simplifies the overall system architecture. By using trained models to automatically identify and structure information, the system reduces the complexity of manual rule configuration and maintenance while enabling powerful granular searching capabilities through automated key-value pair extraction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240362398A1Document data extraction and searching
Publication Date: 2024.10.31 BASE64AI INC
  • US20240362398A1 patent drawing
  • US20240362398A1 patent drawing
  • US20240362398A1 patent drawing

AI summary

Systems and methods of the inventive subject matter are directed to the use of large language models to improve data extraction, storage, and searching. Specifically, platforms implementing embodiments of the inventive subject matter are configured to receive uploaded documents. Once received, the contents of the document can be extracted, and key-value pairs can be generated using that content by applying a taxonomy. For any text content that cannot be index using the applied taxonomy, the platform can apply OCR and then use an LLM to generate additional key-value pairs. Once key-value pairs are created and saved to a database, plan language user-generated search queries can be received. An LLM can once again be used to create database search queries, resulting in the ability to search though uploaded documents for specific content along with types of content.