Lithology Training Set Balancing for Accurate Cuttings Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for lithology estimation from drill cuttings using digital images or video suffer from imbalanced representation of rock types, leading to reduced precision and accuracy in machine learning predictions due to overrepresentation or underrepresentation of certain rock types in the training dataset.
Innovation Solution
A method for generating a balanced training dataset by creating lithology vectors, excluding underrepresented rock types, and repeating image/vector pairs to achieve a more representative distribution of rock types, ensuring each type is accurately represented in the training set.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If uniform sampling is used to collect digital images of cuttings at regular depth intervals, then the data collection process is simple and systematic, but the representation of different rock types becomes imbalanced, leading to reduced prediction accuracy
Solution Approach 1:
The patent changes the sampling parameter from uniform depth intervals to stratified sampling based on lithology frequency. The system calculates target sample sizes for different rock types based on their occurrence frequencies in the formation, then selectively samples images to achieve a balanced training set that reflects actual lithological distribution, thereby improving prediction accuracy while maintaining systematic data collection
Solution Approach 2:
The patent segments the drilling data into distinct lithological categories and applies different sampling strategies to each segment. By dividing the uniform sampling stream into lithology-specific subsets and applying targeted sampling weights, the system achieves balanced representation of rare and common rock types in the training dataset
2Adaptability or versatility
If rare rock types are included in the training dataset, then the model learns more comprehensive lithological patterns, but the dataset becomes imbalanced with rare types overrepresented relative to their actual occurrence
Solution Approach 1:
The patent applies frequency-weighted sampling where the sampling probability of each rock type is adjusted according to its actual occurrence frequency in the formation. Rare rock types receive higher sampling weights to ensure adequate representation, while common rock types receive lower weights to prevent overrepresentation, achieving a balanced training set that maintains prediction accuracy across all lithologies
3Quantity of substance
If the training dataset is enlarged to include more rock type examples, then the model has more training data, but the imbalance between rock type representations increases, reducing model performance
Solution Approach 1:
The patent transforms the data collection parameter from simple count-based expansion to frequency-normalized sampling. By calculating the expected frequency of each lithology type in the formation and using this as a sampling target, the system can enlarge the training dataset while maintaining balanced representation proportions, ensuring that rare types are adequately represented without diluting the overall data quality
Data Source
AI summary
A method for training a machine learning engine from a balanced training set is provided. The method includes receiving a plurality of images of cuttings from a geological formation, generating a lithology vector associated with each image of the plurality of images to form an image/vector set comprising a plurality of image/vector pairs, the lithology vector comprising a plurality of rock types and a percentage of each of the plurality of rock types identified in a respective image, and balancing the image/vector set based on an occurrence of the plurality of rock types across the plurality of image/vector pairs to generate the balanced training set.


