Document Metadata Generation via Logical Segmentation and NLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional document analysis methods in specialized fields like finance and law are labor-intensive and inefficient due to lack of context awareness, limited scalability, and inability to handle unstructured data, leading to inaccuracies and increased costs.

Innovation Solution

A computer-implemented method using large language models to generate metadata for documents by segmenting text into logical segments, creating structured metadata summaries, and storing them in libraries, enabling efficient document analysis and generation of new documents based on analyzed metadata.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual human analysis is used for document management, then accuracy in understanding complex documents is improved, but productivity and efficiency deteriorate due to labor-intensive processes

Engineering Contradiction:
Improveaccuracy in document analysisVSAvoidproductivity in document management
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments documents into logical sections and processes each section independently through multiple NLP passes. This allows the system to maintain high accuracy by analyzing each segment with appropriate context while improving productivity through parallel processing and automated pipeline operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary NLP processing layer that acts as a mediator between raw documents and final analysis results. This intermediary layer performs context-aware disambiguation, entity recognition, and relationship extraction, enabling automated processing to achieve accuracy previously only attainable through human expertise.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If traditional NLP methods are used, then ease of operation is improved, but measurement precision deteriorates due to lack of context awareness

Engineering Contradiction:
Improveease of operationVSAvoidcontext awareness accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by first segmenting documents into logical sections and performing initial entity recognition before conducting detailed relationship extraction. This preliminary structuring enables subsequent processing to maintain high context awareness while keeping the overall system easy to operate through automated pipeline execution.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous useful action through multi-pass NLP processing where each pass builds upon previous results. The system continuously refines entity recognition, relationship extraction, and context disambiguation across multiple processing stages, maintaining high measurement precision while operating as an automated continuous pipeline.

Inventive Principle:
Principle #20Continuity of useful action

3Ease of manufacture

If rule-based techniques are used for metadata extraction, then ease of manufacture is improved, but adaptability deteriorates due to inability to handle variations in text formats

Engineering Contradiction:
Improveease of implementationVSAvoidadaptability to document variations
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic processing by using NLP models that can adapt to different document types, formats, and structures. The system dynamically adjusts its analysis approach based on the specific characteristics of each document while maintaining a consistent automated pipeline, achieving both ease of operation and high adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes processing parameters dynamically based on document characteristics. The NLP pipeline adjusts entity recognition thresholds, relationship extraction criteria, and segmentation strategies according to the specific document type and structure, enabling high adaptability while maintaining automated ease of operation.

Inventive Principle:
Principle #35Parameter changes

4Device complexity

If prior NLP methods are used, then device complexity is reduced, but reliability deteriorates due to inability to disambiguate phrases and words

Engineering Contradiction:
Improvesystem complexityVSAvoidreliability in disambiguation
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments documents into logical sections and processes entities within each segment with full contextual information. This segmentation approach improves reliability of disambiguation by maintaining local context while keeping system complexity manageable through modular processing stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where NLP processing results from one pass inform and refine subsequent processing passes. The system uses feedback from entity recognition to improve relationship extraction, and from relationship extraction to refine context disambiguation, enhancing reliability while maintaining automated operation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240346061A1System and method for document metadata analysis and generation
Publication Date: 2024.10.17 HONG EN
  • US20240346061A1 patent drawing
  • US20240346061A1 patent drawing
  • US20240346061A1 patent drawing

AI summary

Provided is a system and a computer implemented method for generating metadata for a document and documents therefrom. A document is received for analysis via a communication interface a document for analysis. Text is extracted from the document and using a structural analyser model a plurality of segment titles therein are identified and regular expressions derived therefrom. A segmented document is generated comprising extracted text in logical segments with corresponding segment titles by analysing the generated regular expressions. For at least some of the plurality of the logical segments of the segmented document structured metadata summaries of that segment are generated using a metadata creator model. Document metadata thereof is also generated using a metadata creator model. When a new document is requested, using a library of documents and corresponding metadata that document may be generated.