NLP System Categorizes Image Documents via ML Concept Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to efficiently process and analyze unsearchable or partially legible image documents, particularly in the oil and gas industry, due to variations in document quality, organization, and terminology.

Innovation Solution

A natural language processing system and method that uses machine learning to categorize and subcategorize text from image documents, generating a user interface with navigable document images and lists of concepts, allowing users to efficiently navigate and analyze documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If image documents are processed without post-processing, then processing time is reduced, but searchability and analysis capability deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoidsearchability
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary OCR text extraction and machine learning-based concept categorization on document images during the ingestion phase. This preliminary processing creates searchable concept tags and categories before the documents are viewed or analyzed, enabling rapid full-text search and concept-based navigation without requiring post-processing when documents are retrieved.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual review of numerous documents is performed, then analysis accuracy is improved, but time consumption increases

Engineering Contradiction:
Improveanalysis accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system replaces manual mechanical review with automated machine learning models that perform concept extraction, categorization, and document analysis. The ML models automatically identify key concepts, assign categories, and generate summaries, eliminating the need for manual review while maintaining high accuracy through trained algorithms that can process thousands of documents with consistent precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an intermediary layer of concept tags and categories between the raw document text and the user analysis process. This intermediary representation allows users to search and analyze documents by concept rather than manually reviewing text, significantly reducing review time while preserving analysis accuracy through the structured concept hierarchy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If documents are organized with unique terms and structures, then document specificity is improved, but system complexity increases

Engineering Contradiction:
Improvedocument varietyVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs a universal machine learning-based categorization framework that can handle diverse document types, unique terms, and varying structures through a single unified approach. The ML models are trained to recognize concepts across different document formats and terminologies, mapping them to a standardized category system, thereby managing document variety without increasing processing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transforms the variable parameters of unique document terms and structures into standardized concept categories through machine learning. By changing the representation parameters from original document-specific terminology to universal concept tags, the system maintains adaptability to various document types while simplifying the processing complexity through consistent categorical mapping.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If concept categorization is applied to all documents, then searchability is improved, but computational resources increase

Engineering Contradiction:
ImprovesearchabilityVSAvoidcomputational energy
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system applies concept categorization selectively rather than uniformly to all documents. It prioritizes categorization for frequently accessed documents, documents with high search probability, or those containing critical information, while using lighter processing for less important documents. This partial application of full categorization maintains searchability for critical searches while reducing overall computational energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250173044A1Natural language processing system and method for documents
Publication Date: 2025.05.29 THOMSON REUTERS ENTERPRISE CENTRE GMBH
  • US20250173044A1 patent drawing
  • US20250173044A1 patent drawing
  • US20250173044A1 patent drawing

AI summary

In various embodiments, the disclosed systems and methods may receive documents, analyze the documents, categorize portions of the analyzed documents, and present the images of the documents and at least a portion of the categories. The analysis may include identification of categories and the presentation may include indicia of the portion of the image of the document related to the category. The systems and methods disclosed may allow querying and/or reporting of a plurality of documents to facilitate processing.