Ensemble ML Model for Well Production Prediction with Small Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Predicting well performance in unconventional reservoirs is challenging due to the need for massive hydraulic fracturing, and existing machine learning methods require significant data for optimal results, making them ineffective with small training datasets.
Innovation Solution
A method involving a machine learning model trained with historical well production data, geological, and petrophysical data, where multiple sets of initial model parameters are generated, and individually trained models are ranked based on validation data, with top-ranked models used to generate final predicted well production data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning methods are used for predicting well production, then prediction capability is improved, but the requirement for large training data sets increases
Solution Approach 1:
The patent segments the training process by dividing the limited training data into multiple subsets and training multiple individual ML models on different subsets. This segmentation allows the system to make effective use of limited data while generating multiple diverse models that can be combined for more reliable predictions.
Solution Approach 2:
The patent merges multiple individually trained ML models into an ensemble system. By combining the predictions from multiple models trained on different data subsets, the system achieves more robust and reliable predictions while still using a relatively small overall training data set.
2Reliability
If a single ML model is trained with limited data, then training time is reduced, but prediction accuracy and reliability deteriorate
Solution Approach 1:
The training process is segmented into multiple parallel training operations, where multiple models are trained simultaneously on different data subsets. This approach distributes the computational workload and allows for more comprehensive model training without proportionally increasing total training time.
Solution Approach 2:
The system performs preliminary actions by generating multiple sets of initial parameter guesses before training begins. This preparation enables the subsequent training process to proceed more efficiently with multiple models, as the initial parameter space has been pre-explored and organized.
3Reliability
If multiple individually trained ML models are generated, then prediction reliability is improved, but model complexity increases
Solution Approach 1:
The model system is segmented into multiple independent but parallel individual models, each trained on different data subsets. This segmentation creates a modular architecture where complexity is distributed across multiple simpler models rather than concentrated in one complex model.
Solution Approach 2:
The system employs parameter changes by generating multiple sets of initial parameter guesses and training models with different parameter configurations. This approach explores the parameter space more thoroughly and creates diverse models that, when combined, improve reliability without requiring any single model to be overly complex.
Data Source
AI summary
A method for predicting well production is disclosed. The method includes obtaining a training data set for a machine learning (ML) model that generates predicted well production data based on observed data of interest, generating multiple sets of initial guesses of model parameters of the ML model, using an ML algorithm applied to the training data set to generate multiple individually trained ML models based the multiple sets of initial model parameters, comparing a validation data set and respective predicted well production data of the individually trained ML models to generate a ranking, selecting top-ranked individually trained ML models based on the ranking, using the data of interest as input to the top-ranked individually trained ML models to generate a set of individual predicted well production data, and generating a final predicted well production data based on the set of individual predicted well production data.


