Oilfield Document Scanner for Automated Metadata and Table Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The management of large volumes of unstructured oilfield documents in the oil and gas industry is inefficient, as users manually open and review each document to extract relevant data, making it difficult to search and classify the documents effectively.
Innovation Solution
A pluggable lightweight software utility that extracts metadata and tabular data from oilfield documents using term frequency-inverse document frequency (TF-IDF) and a document content classification model, storing the extracted information with the document in storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual document review is used to extract data, then data extraction accuracy is maintained, but productivity is significantly reduced
Solution Approach 1:
The patent replaces the manual mechanical review process with an automated computer-based system that uses optical character recognition (OCR), natural language processing (NLP), and machine learning algorithms to extract data from oilfield documents, thereby eliminating the need for manual intervention while maintaining high processing speeds
Solution Approach 2:
The system enables documents to be processed automatically without human intervention by using intelligent algorithms that can independently extract, classify, and store data from various document formats including scanned images and PDFs, making the extraction process self-serving
2Ease of operation
If documents are stored without classification, then storage simplicity is maintained, but ease of operation is reduced due to difficulty in searching
Solution Approach 1:
The system performs preliminary classification and metadata extraction during the document ingestion process itself, automatically categorizing documents by type, well location, and other relevant attributes before they are stored, so that when documents are later searched, they are already organized and easily retrievable
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the raw document content and the search functionality, where extracted attributes such as document type, well ID, and location serve as intermediate indices that enable efficient searching without requiring complex query processing of the full document content
Data Source
AI summary
A method involves extracting, from a file comprising an unstructured oilfield document, terms, calculating term frequency inverse document frequency (TF-IDF) of the terms to generate an input vector, execute a document content classification model on the input vector to generate a document content classification of unstructured oilfield document, and extract table information from a table in the unstructured oilfield document. The method further involves storing, with the file in storage, the document content classification and the table information.


