Lithology Training Set Balancing for Accurate Cuttings Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for lithology estimation from drill cuttings using digital images or video suffer from imbalanced representation of rock types, leading to reduced precision and accuracy in machine learning predictions due to overrepresentation or underrepresentation of certain rock types in the training dataset.

Innovation Solution

A method for generating a balanced training dataset by creating lithology vectors, excluding underrepresented rock types, and repeating image/vector pairs to achieve a more representative distribution of rock types, ensuring each type is accurately represented in the training set.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If uniform sampling is used to collect digital images of cuttings at regular depth intervals, then the data collection process is simple and systematic, but the representation of different rock types becomes imbalanced, leading to reduced prediction accuracy

Engineering Contradiction:
Improvedata collection simplicityVSAvoidlithology prediction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent changes the sampling parameter from uniform depth intervals to stratified sampling based on lithology frequency. The system calculates target sample sizes for different rock types based on their occurrence frequencies in the formation, then selectively samples images to achieve a balanced training set that reflects actual lithological distribution, thereby improving prediction accuracy while maintaining systematic data collection

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the drilling data into distinct lithological categories and applies different sampling strategies to each segment. By dividing the uniform sampling stream into lithology-specific subsets and applying targeted sampling weights, the system achieves balanced representation of rare and common rock types in the training dataset

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If rare rock types are included in the training dataset, then the model learns more comprehensive lithological patterns, but the dataset becomes imbalanced with rare types overrepresented relative to their actual occurrence

Engineering Contradiction:
Improvemodel lithological coverageVSAvoidprediction accuracy for common types
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies frequency-weighted sampling where the sampling probability of each rock type is adjusted according to its actual occurrence frequency in the formation. Rare rock types receive higher sampling weights to ensure adequate representation, while common rock types receive lower weights to prevent overrepresentation, achieving a balanced training set that maintains prediction accuracy across all lithologies

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If the training dataset is enlarged to include more rock type examples, then the model has more training data, but the imbalance between rock type representations increases, reducing model performance

Engineering Contradiction:
Improvetraining data volumeVSAvoidprediction accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent transforms the data collection parameter from simple count-based expansion to frequency-normalized sampling. By calculating the expected frequency of each lithology type in the formation and using this as a sampling target, the system can enlarge the training dataset while maintaining balanced representation proportions, ensuring that rare types are adequately represented without diluting the overall data quality

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250239055A1Systems and methods for preparing a lithologically balanced training set
Publication Date: 2025.07.24 SAUDI ARABIAN OIL CO
  • US20250239055A1 patent drawing
  • US20250239055A1 patent drawing
  • US20250239055A1 patent drawing

AI summary

A method for training a machine learning engine from a balanced training set is provided. The method includes receiving a plurality of images of cuttings from a geological formation, generating a lithology vector associated with each image of the plurality of images to form an image/vector set comprising a plurality of image/vector pairs, the lithology vector comprising a plurality of rock types and a percentage of each of the plurality of rock types identified in a respective image, and balancing the image/vector set based on an occurrence of the plurality of rock types across the plurality of image/vector pairs to generate the balanced training set.