NLP Document Categorization for Oil and Gas Legal Records

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The oil and gas industry faces challenges in processing and analyzing large volumes of legal documents, particularly those in image format, which are often of varying quality and legibility, making them unsearchable and difficult to navigate due to differences in terminology and organization.

Innovation Solution

A natural language processing system that uses machine learning to categorize and subcategorize document content, generating a user interface that allows navigation through identified concepts, enabling efficient retrieval and analysis of document content regardless of formatting or terminology variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If image format documents are used for storage and transmission, then document preservation and sharing are improved, but searchability and analysis capability deteriorate

Engineering Contradiction:
Improvedocument preservationVSAvoidsearchability
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary processing system that converts image format documents into structured text data through OCR and natural language processing. This intermediary layer enables the document to maintain its original image format for preservation while generating searchable text representations that enable full-text search and analysis capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual document review and keyword-based search with automated machine learning models that perform semantic analysis, entity recognition, and concept extraction. This substitution transforms the mechanical process of document analysis into an intelligent system that can understand and categorize document content automatically.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual review and processing of documents is performed, then accurate analysis can be achieved, but time consumption and labor resources increase significantly

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements a self-service system where machine learning models automatically perform document analysis, categorization, and information extraction without requiring manual intervention. The system serves itself by continuously learning from processed documents and improving its analysis capabilities, enabling rapid processing of large volumes of documents while maintaining high accuracy through automated quality control mechanisms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the fundamental parameters of document processing by transitioning from human-cognitive time scales to machine-processing time scales. The system processes documents in seconds or minutes rather than hours or days, while maintaining or improving accuracy through multiple processing stages including OCR, NLP, and machine learning validation.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If documents are organized with unique structures and terminology, then document-specific requirements are met, but cross-document comparison and analysis become difficult

Engineering Contradiction:
Improvedocument format flexibilityVSAvoidcross-document analysis
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a universal processing framework that can handle multiple document formats, structures, and terminology schemes simultaneously. The machine learning models are designed to recognize and adapt to various document types (leases, contracts, reports) while extracting standardized information categories, enabling cross-document comparison while preserving the unique characteristics of each document type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies local quality by allowing different document sections and types to be processed with specialized models tailored to their specific requirements, while maintaining a unified output structure. Each document type can have its own terminology and organization preserved locally, while the overall system enables global comparison through standardized categorization and entity recognition.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11954098B1Natural language processing system and method for documents
Publication Date: 2024.04.09 THOMSON REUTERS ENTERPRISE CENTRE GMBH
  • US11954098B1 patent drawing
  • US11954098B1 patent drawing
  • US11954098B1 patent drawing

AI summary

In various embodiments, the disclosed systems and methods may receive documents, analyze the documents, categorize portions of the analyzed documents, and present the images of the documents and at least a portion of the categories. The analysis may include identification of categories and the presentation may include indicia of the portion of the image of the document related to the category. The systems and methods disclosed may allow querying and/or reporting of a plurality of documents to facilitate processing.