Class-Based Subsurface Data Processing for Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional subsurface data processing systems are subjective, inconsistent, and slow due to expert-centric workflows, and require large datasets and numerous measurements for effective machine learning applications, limiting their practicality and leading to knowledge loss when experts change roles or leave.
Innovation Solution
A class-based machine learning approach that reduces subsurface data into explainable classes using cross entropy clustering, Gaussian mixture models, and hidden Markov models, enabling efficient data processing and interpretation with fewer measurements, and automating workflows by storing updated parameters in a knowledgebase for machine learning processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If expert-centric workflows are used for subsurface data processing, then specialized knowledge and analysis quality are maintained, but processing time increases and results become subjective and inconsistent
Solution Approach 1:
The system performs preliminary clustering and classification of subsurface data into representative classes before actual processing. This pre-organization of data into standardized categories enables automated workflows to efficiently process new data without requiring expert intervention for each case, thus maintaining quality while reducing processing time.
Solution Approach 2:
The system creates standardized templates and models from expert-analyzed cases that can be replicated and applied to new data. These copied patterns of analysis preserve expert knowledge and ensure consistent, reliable results across different processing instances without requiring the original expert's continuous involvement.
2Reliability
If expert-defined workflows are used, then specialized knowledge is applied, but knowledge is lost when experts change jobs or leave
Solution Approach 1:
The system captures and stores processing parameters, classification rules, and analysis workflows directly from the data and expert interactions, creating a self-sustaining knowledge base. This automated knowledge capture ensures that expertise is preserved in the system itself rather than dependent on individual experts, preventing knowledge loss when personnel change.
Solution Approach 2:
The system replaces the mechanical dependency on human experts with an automated machine learning framework that encodes specialized knowledge into algorithms and models. This substitution transforms tacit expert knowledge into explicit, reproducible computational processes that remain within the organization regardless of personnel changes.
3Productivity
If machine learning is applied to subsurface data, then processing speed and automation improve, but large amounts of data and high number of measurements are required
Solution Approach 1:
The system extracts and separates the essential classification features from large volumes of subsurface data, identifying and retaining only the critical parameters that define different lithological classes. This extraction process reduces the effective data quantity needed for machine learning while preserving the information necessary for accurate processing and interpretation.
4Measurement precision
If traditional processing workflows are used, then detailed analysis is performed, but results are subjective and inconsistent depending on expert expertise
Solution Approach 1:
The system transforms subjective expert judgment into objective quantitative parameters through standardized classification schemes and statistical models. By converting qualitative expert assessments into measurable, reproducible parameters, the system maintains detailed analysis capability while ensuring consistent results across different users and applications.
Data Source
AI summary
A method and apparatus for subsurface data processing includes determining a set of clusters based at least in part on measurement vectors associated with different depths or times in subsurface data, defining clusters in a subsurface data by classes associated with a state mode, reducing a quantity of the subsurface data based at least in part on the classes, and storing the reduced quantity of the subsurface data and classes with the state model in a training database for a machine learning process.


