Automated Metadata Extraction for Wellbore Planning Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying subsurface exploration documents in a database are manual and time-consuming, leading to inefficiencies and errors in document identification for wellbore planning.

Innovation Solution

A method and system utilizing machine-learned models and natural language processing to automatically determine document categories, types, and metadata attributes, enabling efficient retrieval and planning of wellbores using earth property data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual labelling of exploration documents with metadata is used, then document identification capability is achieved, but time consumption and error rate increase

Engineering Contradiction:
Improvedocument identification accuracyVSAvoidtime for manual labelling
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables documents to automatically generate their own metadata through machine-learned models that process document content, titles, and earth property data. The NLP algorithm extracts metadata attributes autonomously without requiring manual human intervention, making the documentation process self-servicing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of labelling documents is replaced with an automated electronic system using machine-learned models and natural language processing algorithms. These computational systems automatically analyze document content and generate metadata, substituting human manual work with algorithmic processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If manual labelling of exploration documents is used, then document categorization is achieved, but error rate increases

Engineering Contradiction:
Improvedocument processing speedVSAvoidlabelling accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The machine-learned models are trained using feedback from labeled training data, continuously improving their accuracy. The system processes documents through multiple stages including preprocessing, model inference, and validation, with the ability to learn from correct and incorrect classifications to enhance future performance.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Human manual labelling is replaced with automated machine-learned models and NLP algorithms that consistently apply classification rules without human error. The electronic processing system maintains uniform standards across all documents, eliminating variability inherent in manual human work.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automated machine-learned models are used for metadata extraction, then processing efficiency is improved, but system complexity increases

Engineering Contradiction:
Improvemetadata extraction speedVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The automated system is divided into distinct modular components: a preprocessing module that prepares documents, machine-learned models that extract specific metadata attributes, and an NLP algorithm that processes text content. Each module performs a specific function and can be independently trained, maintained, and improved.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The machine-learned models and NLP algorithm serve multiple functions: they process different document types (seismic surveys, electromagnetic surveys, well log data), extract various metadata attributes (category, type, properties), and work with different data formats. This multi-functionality reduces the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250384065A1Method and system for metadata extraction for document identification
Publication Date: 2025.12.18 SAUDI ARABIAN OIL CO
  • US20250384065A1 patent drawing
  • US20250384065A1 patent drawing
  • US20250384065A1 patent drawing

AI summary

A method includes obtaining a document comprising earth property data regarding a geological region of interest, preprocessing the document to form at least one preprocessed document and determining, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document. The method further includes determining, using a natural language processing algorithm, metadata attributes of the document and a title of the document, and updating a database storing the document with the title, the category, the type and the metadata attributes. The method further includes identifying, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes and planning a wellbore path in the geological region of interest using the earth property data comprised in the document.