A method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking

By improving the Stacking integrated learning model, combining multiple machine learning algorithms and hyperparameter optimization, the problem of category imbalance in enhanced oil recovery technology screening is solved, the accuracy and efficiency of decision-making is improved, and efficient decision-making support is provided.

CN120046498BActive Publication Date: 2025-09-02BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510197443.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2025-01-23
Filing Date
2025-02-21
Publication Date
2025-09-02
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing method for screening of enhanced oil production technology has problems such as time-consuming, high cost and inaccurate decision-making, especially in the category imbalanced data sets, which are difficult to provide efficient and scientific decision-making support.

Method used

Build an intelligent screening decision model for enhanced oil production technology based on improved Stacking. By integrating multiple machine learning algorithm models, optimizing model structure, using random search and grid search to optimize hyperparameters, combining metamodels with simple structure and anti-overfit characteristics, improve Stacking integrated learning to overcome category imbalance problem.

Benefits of technology

It improves the accuracy of strengthening oil production technology screening and decision-making accuracy, provides more efficient and scientific decision-making support, solves the impact of category imbalance problem, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046498B_ABST
    Figure CN120046498B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of intelligent oil and gas field development research and development, specifically to a method for constructing an intelligent screening decision model for enhanced oil recovery (ERR) technologies based on improved Stacking. The method comprises the following steps: Step 1: Collecting data on successfully implemented EOR projects worldwide and forming a complete data set through data processing and analysis; Step 2: Selecting multiple conventional machine learning algorithms and constructing multiple intelligent EOR screening models based on conventional machine learning algorithms; Step 3: Using Stacking ensemble learning to fuse and construct an intelligent EOR screening model based on Stacking; Step 4: Improving Stacking by adding a new meta-model to construct an integrated intelligent screening model for EOR technologies based on improved Stacking. The present invention addresses the limitations of machine learning algorithms in EOR screening decisions and provides more accurate decision support for reservoir engineers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent development research of oil and gas fields, and particularly relates to a method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking. Background Art

[0002] With rapid economic development, the demand for energy is increasing worldwide. As one of the world's primary energy sources and a crucial raw material for modern industrial society, oil is crucial to contemporary society and industrial development. Enhanced oil recovery (ERR) technologies, as effective solutions for increasing oil production, have garnered widespread attention. They utilize advanced physical and chemical techniques and methods to improve the physical and chemical properties of reservoir formations and fluids, increase oil mobility within the reservoir, and enhance oil recovery. With technological advancements, over 20 EOR technologies are now widely used worldwide. However, due to limitations such as reservoir rock and fluid characteristics, reservoir engineers must comprehensively consider multiple factors to select the most appropriate EOR technology from a wide range of options to achieve optimal results. Choosing the optimal EOR technology based on target reservoir parameters to maximize oil production remains a major challenge for reservoir engineers. EOR technology selection utilizes a multi-criteria decision-making process to select the best feasible EOR technology for a target reservoir, taking into account reservoir rock and fluid properties. As the first step in implementing an enhanced oil recovery (ERR) project, efficient and accurate screening has a direct impact on improving the efficiency of the project and achieving the desired oil recovery results. Inappropriate ER technology can not only permanently damage the reservoir but also lead to significant economic losses. Therefore, accurate and reliable ER technology selection in the early stages of reservoir development is a prerequisite for ER project planning and design.

[0003] At present, there are three main methods for screening enhanced oil recovery technologies at home and abroad:

[0004] The first approach involves laboratory simulations / field pilots, which simulate the implementation of enhanced oil recovery (ERR) technologies based on reservoir rock and fluid properties to identify suitable EOR technologies. This approach requires simulating multiple EOR technologies, resulting in time-consuming and costly screening. The second approach involves conventional EOR technology selection and decision-making. This approach uses reservoir and fluid parameters from past successful EOR projects and the EOR technologies used to develop corresponding EOR technology selection criteria. These criteria are typically presented in a graphical format. The likelihood of successful implementation of each EOR technology is determined by considering several predefined selection parameters. However, searching for reservoir parameters using these EOR technology selection criteria may result in multiple matching EOR technologies, resulting in non-uniform selection results. The third approach involves advanced EOR technology selection and decision-making. This approach utilizes computational techniques, such as machine learning, to identify potential relationships between reservoir rock and fluid properties and EOR technology implementation success in past successful EOR projects, thereby establishing valuable EOR technology selection rules. While this decision-making process can provide specific solutions quickly, it is highly dependent on the dataset. There are significant disparities in the number of enhanced oil recovery technologies implemented globally. Most datasets used for advanced enhanced oil recovery technology selection and decision-making suffer from class imbalance. When a dataset contains a disproportionately large number of samples in one class, computer technology may tend to predict the majority class while overlooking the importance of the minority class. Therefore, due to class imbalance, its decision-making accuracy needs to be further improved. Summary of the Invention

[0005] To address the drawbacks of the aforementioned prior art, the present invention discloses a method for constructing an intelligent screening decision model for enhanced oil recovery technologies based on improved Stacking. The present invention collects statistics on successfully implemented enhanced oil recovery technologies worldwide, constructs and optimizes multiple intelligent screening models for enhanced oil recovery technologies based on conventional machine learning algorithms, and utilizes Stacking ensemble learning to fuse and construct an intelligent screening ensemble model for enhanced oil recovery technologies based on Stacking. This model is then improved to construct a novel intelligent screening model for enhanced oil recovery technologies based on improved Stacking, thereby overcoming the impact of category imbalance on model performance and addressing the shortcomings of the Stacking ensemble learning model. This improves the accuracy of enhanced oil recovery technology screening, thereby providing more efficient, scientific, and intelligent decision support for oil companies in selecting the optimal enhanced oil recovery technology.

[0006] The present invention specifically adopts the following technical solutions:

[0007] A method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking, such as Figure 1 As shown, the following steps are included:

[0008] First, data on successfully implemented enhanced oil recovery projects around the world are collected, and a complete data set is formed through data processing and analysis for use in enhanced oil recovery technology screening and decision-making.

[0009] Secondly, based on conventional machine learning algorithms, intelligent screening models for various enhanced oil recovery technologies are constructed and optimized.

[0010] Furthermore, based on the prediction performance of conventional machine learning algorithm models, Stacking ensemble learning is used to integrate and construct an intelligent screening model for enhanced oil recovery technology based on Stacking.

[0011] Finally, the Stacking ensemble model is improved to construct a new enhanced oil recovery technology intelligent screening ensemble model based on improved Stacking to overcome the impact of category imbalance on model performance, while making up for the defects of the traditional Stacking ensemble learning model, thereby improving the enhanced oil recovery technology screening accuracy.

[0012] According to the present invention, preferably, the sources of the enhanced oil recovery project data include:

[0013] Step 1: Collect data from successfully implemented enhanced oil recovery projects around the world, select reservoir rock and fluid properties that can effectively reflect the actual reservoir and fluid conditions and are closely related to the oil recovery process of enhanced oil recovery technology as characteristic parameters, and obtain a complete data set through data preprocessing and analysis.

[0014] Preferably, according to the present invention, the conventional machine learning algorithm includes:

[0015] Step 2. Based on the current research status, random forest, extreme gradient boosting, neural network, decision tree, support vector machine and logistic regression are selected as conventional machine learning algorithms. They are used to learn the data set obtained in step 1.3. With feature parameters as input and enhanced oil recovery technology as output, a variety of conventional enhanced oil recovery technology intelligent screening models are constructed. The performance of conventional machine learning algorithm models is evaluated using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators, providing a model basis for the subsequent construction of an integrated model.

[0016] Preferably, according to the present invention, the Stacking ensemble learning algorithm is:

[0017] Step 3. Based on the prediction performance of the conventional machine learning algorithm model in step 2, Stacking ensemble learning is used for fusion to build an intelligent screening model for enhanced oil recovery technology based on Stacking. The performance of the Stacking ensemble model is evaluated using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators, providing an improvement basis for the subsequent construction of an improved Stacking ensemble model.

[0018] Preferably, according to the present invention, the improved Stacking ensemble learning algorithm is:

[0019] Step 4. Improve Stacking by adding a new meta-model, and construct an intelligent screening integrated model for enhanced oil recovery technology based on improved Stacking. Evaluate the performance of the improved Stacking integrated model using accuracy, precision, recall, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators to verify the superiority of the improved Stacking integrated model. Finally, provide a high-precision intelligent screening integrated model for enhanced oil recovery technology based on improved Stacking to participate in the final enhanced oil recovery technology screening decision.

[0020] Preferably, according to the present invention, the specific process of step 1 is:

[0021] Based on previous enhanced oil recovery technology screening criteria and enhanced oil recovery technology displacement mechanism, feature selection and data collection are performed through steps 1.1 and 1.2 to generate the original data set. Data preprocessing and data analysis are performed through steps 1.3 and 1.4 to construct the complete data set for the final screening decision.

[0022] Step 1.1: Based on previous research on enhanced oil recovery technology screening and calculation formulas for reservoir volume, select reservoir rock and fluid properties that effectively reflect the actual reservoir and fluid conditions and are closely related to the enhanced oil recovery process as characteristic parameters. Specific parameters include: lithology, porosity, permeability, reservoir depth, crude oil density, crude oil temperature, crude oil viscosity, net thickness, and initial oil saturation.

[0023]

[0024] In formulas (1)-(3), OOIP is the original reservoir reserves, q is the volume flow rate of the fluid in the porous medium, A is the oil-bearing area, h is the net thickness, is the porosity, S0 is the initial crude oil saturation, B oiis the average original crude oil volume coefficient, k is the permeability, μ is the crude oil viscosity, Δp is the pressure difference across the two ends, API is the crude oil specific gravity, and SG is the relative density of crude oil at 60°F. OOIP and q are important factors determining oil production and recovery effectiveness. Equations (1) to (3) involve some of the specific parameters in step 1.1, indicating that the characteristic parameters selected in this invention are closely related to the oil recovery process. One of the main methods currently used to enhance oil recovery is to increase crude oil fluidity. Therefore, this formula indirectly demonstrates that the characteristic parameters selected in this invention have a direct impact on enhancing oil recovery.

[0025] Step 1.2: Using the results of the semiannual enhanced oil recovery project survey published by the Oil & Gas Journal as the main data source, the data were supplemented by reports from the U.S. Department of Energy, the American Association of Petroleum Geologists database, field reports, Chinese publications, and publications from the Society of Petroleum Engineers to form the original data set.

[0026] Step 1.3: Process the data set collected in step 1.2, including deleting duplicate data, replacing missing values ​​with feature means, and performing one-hot encoding on discrete parameters, to form a complete data set for enhancing oil production technology screening decisions.

[0027] Formula (4) is used to perform logarithmic processing on the continuous feature parameters with wide distribution in the data set to narrow the numerical range, thereby increasing the stability of the data and avoiding the negative impact of extreme data on model training.

[0028] x * =lg(x org ) (4)

[0029] In formula (4), x * represents the data after logarithmic processing; x org Represents the original sample data;

[0030] Step 1.4: Analyze the complete data set generated in step 1.3, including the quantitative distribution of each enhanced oil recovery technology and the distribution of characteristic parameters, explore the class imbalance problem in the data set, ensure that the data set can effectively reflect the universal geological characteristics of the global rocks in the data set, and use the Spearman correlation coefficient to analyze the correlation between characteristic parameters, providing a data basis for the subsequent construction of the enhanced oil recovery technology screening model. The formula for calculating the Spearman correlation coefficient is:

[0031]

[0032] In formula (5), r s is the Spearman correlation coefficient, d iis the rank difference of the two variables, that is, for the i-th observation, its rank difference on the two variables, and n is the number of observations.

[0033] Step 1.5: Use the holdout method to randomly divide the complete dataset formed in step 1.3 into training set, validation set, and test set in a ratio of 8:1:1 for the training, optimization, and evaluation of the intelligent screening model for enhanced oil production technology.

[0034] Preferably, according to the present invention, the specific process of step 2 is:

[0035] Step 2. Based on the current research status, random forest, extreme gradient boosting, neural network, decision tree, support vector machine and logistic regression are selected as conventional machine learning algorithms. They are used to learn the data set obtained in step 1.3. With feature parameters as input and enhanced oil recovery technology as output, a variety of conventional enhanced oil recovery technology intelligent screening models are constructed. The performance of conventional machine learning algorithm models is evaluated using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators, providing a model basis for the subsequent construction of an integrated model.

[0036] Step 2.1. Determine the evaluation indicators. Accuracy, precision, recall, and F1 score are selected as conventional evaluation indicators to evaluate the overall performance of the model. A confusion matrix is ​​drawn and the Kappa coefficient and MCC value are used to evaluate the model's ability to overcome the class imbalance problem. Among them, accuracy, precision, recall, and F1 score are affected by class imbalance and their values ​​are relatively high. The Kappa coefficient and MCC value can ignore the impact of class imbalance. Comparing the two can provide a more accurate evaluation of the model. The confusion matrix can be used to determine the model's classification effect on each enhanced oil recovery technology. The calculation formula is:

[0037]

[0038] In formulas (6)-(11), TP is the number of positive examples correctly predicted as positive by the model, TN is the number of negative examples correctly predicted as negative by the model, FP is the number of negative examples incorrectly predicted as positive by the model, FN is the number of positive examples incorrectly predicted as negative by the model, A is the accuracy, and A r is the accuracy of random classification. The values ​​of all quantitative evaluation indicators are between -1 and 1. The closer the value is to 1, the better the model classification effect.

[0039] Step 2.2: Determine the hyperparameter optimization method. Use a combination of random search and grid search for model optimization. Random search can quickly explore the breadth of the hyperparameter space, narrowing the search space and improving search efficiency. Grid search, which traverses the hyperparameter space, ensures search quality and determines the optimal hyperparameter combination for all models.

[0040] In step 2.3, the random forest algorithm is used to learn the dataset obtained in step 1.3. The hyperparameter optimization method used in step 2.2 is then used to optimize the model. This allows for the construction of an intelligent screening model for enhanced oil recovery technologies based on the random forest algorithm. The random forest algorithm is a decision tree-based ensemble learning algorithm. It uses bootstrap sampling to create multiple independent decision tree models, and the final decision is reached through voting or averaging. The random nature of its sample and feature selection makes it more suitable for high-dimensional and large-scale data, increasing the diversity of its decision-making process and preventing the model from focusing too much on certain data features and overfitting. Its parallel ensemble nature makes it highly robust to noise and outliers.

[0041] The specific process of the random forest algorithm is:

[0042] Step 1: Through self-service sampling, randomly form multiple sub-datasets for independent training of a single decision tree;

[0043] Step 2: Randomly select multiple features from the feature space to form a feature subset. When constructing each node of the decision tree, select the optimal feature from the feature subset for splitting;

[0044] Step 3: Repeat the above steps to generate multiple decision trees to form a random forest.

[0045] In step 2.4, the dataset obtained in step 1.3 is trained using the extreme gradient boosting algorithm. This algorithm is optimized using the hyperparameter optimization method from step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the extreme gradient boosting algorithm. The extreme gradient boosting algorithm is an ensemble learning algorithm based on gradient boosting decision trees. It constructs multiple gradient boosting decision trees in series, learns sequentially using a highly adaptive method, and ultimately obtains prediction results through weighting.

[0046] This algorithm accelerates convergence and avoids overfitting by introducing a second-order Taylor formula and regularization terms. It also leverages sparse perception to efficiently handle samples with missing eigenvalues. The extreme gradient boosting algorithm is widely used in various machine learning tasks due to its high efficiency, flexibility, and scalability in high-dimensional data processing, feature selection, and handling of missing values.

[0047] The specific process of the extreme gradient boosting algorithm is:

[0048] Step 1: Initialize a constant model as the initial prediction value;

[0049] Step 2: Iteratively construct multiple decision trees. In each iteration, determine a decision tree that minimizes the objective function;

[0050] Step 3: After each decision tree is constructed, its prediction result is added to the previous prediction value to obtain a new prediction value.

[0051] Step 4: Continue the iterative process until the preset number of iterations is reached or the change in the objective function is less than a certain threshold.

[0052] Step 2.5: Use a neural network algorithm to learn the dataset obtained in step 1.3. Optimize it using the hyperparameter optimization method from step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on a neural network algorithm. A neural network algorithm is a machine learning algorithm that simulates the interconnections and information transfer between neurons in the human brain to achieve data processing and pattern recognition. It learns and processes input data by adjusting the connection weights and activation functions between neurons. During model operation, the algorithm can gradually improve its performance by learning and adjusting its own connection weights and biases, thus exhibiting excellent adaptability. The nonlinear activation function gives neural networks powerful nonlinear modeling capabilities, enabling them to effectively handle complex nonlinear relationships and patterns.

[0053] The specific process of the neural network algorithm is as follows:

[0054] Step 1: Forward propagation, from the input layer to the output layer to calculate the output of the network;

[0055] Step 2: Calculate the loss and calculate the loss function based on the predicted results and the true label;

[0056] Step 3: Back propagation, calculate the gradient from the output layer to the input layer, and update the weights and biases;

[0057] Step 4: Iterative training, repeat the above process until the network converges.

[0058] Step 2.6: Use a decision tree algorithm to learn the dataset obtained in step 1.3. Optimize it using the hyperparameter optimization method from step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the decision tree algorithm. The decision tree algorithm is an important classification and regression method in data mining. It interrogates the features in a dataset and, based on the values ​​of different features, gradually segments the data, ultimately generating a tree-like structure to predict the category or value of new data. This algorithm dynamically selects the optimal splitting point and method based on the characteristics of the data and the nature of the features. Due to its flexibility and adaptability in feature selection and node splitting, decision trees can handle both numerical and categorical features, eliminating the need for complex data transformation or preprocessing.

[0059] The specific process of the decision tree algorithm is:

[0060] Step 1: Starting from the root node, select the feature that can minimize the impurity of the dataset as the split feature;

[0061] Step 2: Divide the dataset into multiple subsets based on the selected features;

[0062] Step 3: For the divided subsets, repeat the above process and recursively construct decision tree branches until the stopping condition is met.

[0063] In step 2.7, the support vector machine algorithm is used to learn the dataset obtained in step 1.3. The hyperparameter optimization method used in step 2.2 is then used to optimize the dataset. This allows for the construction of an intelligent screening model for enhanced oil recovery technology based on the support vector machine algorithm. The support vector machine algorithm, a generalized linear classifier for data classification, seeks a linear or nonlinear hyperplane that maximizes the distance between sample points on either side of the hyperplane. The final classification result is determined based on the distance between the sample and the hyperplane. The algorithm can flexibly adapt to different data distributions by selecting different kernel functions. Its final decision relies solely on the points in the training sample closest to the classification boundary, resulting in strong generalization and robustness. The principle of structural risk minimization also makes the support vector machine algorithm highly resistant to overfitting.

[0064] The calculation formula for the classification result obtained by the support vector machine algorithm is:

[0065]

[0066] In formula (12), f(x) represents the classification result for input data x, x represents the input data to be classified, i represents the sample number, which takes values ​​of 1, 2, ..., n, where n is the number of samples, and y i is the label of the i-th sample, K(x,x i ) is the kernel function, xi represents the feature vector of the i-th sample, α i is the Lagrange multiplier, b is the bias term, when f(x)>0, sample x belongs to the positive category; when f(x)<0, sample x belongs to the negative category.

[0067] Step 2.8: Use the logistic regression algorithm to learn the data set obtained in step 1.3, optimize it using the hyperparameter optimization method in step 2.2, and construct an intelligent screening model for enhanced oil recovery technology based on the logistic regression algorithm. As a generalized linear model, the logistic regression algorithm classifies data through a linear combination of a set of input features, and converts the output of the linear regression model into the probability of the category through a logistic function to obtain the final classification result. Due to its linear structure, the logistic regression algorithm has a relatively simple model and fast calculation speed. Following the principle of structural risk minimization, it has a strong anti-overfitting characteristic. The calculation formula for the classification result obtained by the logistic regression algorithm is:

[0068] z=w T x1+b(13)

[0069]

[0070] Formula (13) is a linear combination of input features. Where z represents the output value given based on the input x1, x1 represents the feature vector, which contains the eigenvalues ​​of the input data, and w represents the weight vector, which contains the weight of each feature in the model. T represents the transpose of the weight vector, b is the bias term, and σ(z) is the class probability. When σ(z) ≥ 0.5, the sample is classified as a positive sample; when σ(z) < 0.5, the sample is classified as a negative sample.

[0071] Step 2.9: Evaluate six common machine learning algorithm models using accuracy, precision, recall, F1 score, confusion matrix, Kappa coefficient, and MCC value, and rank them according to their performance to provide a model foundation for the subsequent construction of the integrated model.

[0072] Preferably, according to the present invention, the stacking ensemble learning is:

[0073] Step 3. Based on the prediction performance of the conventional machine learning algorithm model in step 2, Stacking ensemble learning is used for fusion to build an intelligent screening model for enhanced oil recovery technology based on Stacking. The performance of the Stacking ensemble model is evaluated using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators, providing an improvement basis for the subsequent construction of an improved Stacking ensemble model.

[0074] The specific process of step 3 is as follows:

[0075] Step 3.1: Stacking ensemble learning is an ensemble learning method used to improve prediction results. It integrates the prediction results of multiple different base models using a meta-model to obtain the final prediction result. Figure 2 It is the structure of Stacking integrated learning algorithm.

[0076] The specific implementation process of Stacking ensemble learning is as follows:

[0077] Step 1: Dataset division: Use the training set divided in step 1.4 as training data and use K-fold cross-validation to divide it into K subsets, where K-1 subsets are sub-training sets and the remaining subset is used as sub-validation set;

[0078] Step 2: Base model training, using different base models for independent training on the sub-training set;

[0079] Step 3: Generate new features, use the prediction results of all base models on the sub-validation set as new features, and combine them into a new training dataset. The prediction results of the base model on the test set divided in step 1.4 are combined into a new test dataset;

[0080] Step 4: Train the meta-model, input the prediction results of the base model into a meta-model as a new feature, and maximize the accuracy of the overall model by fitting the prediction results of the base model.

[0081] Stacking ensemble learning usually has a two-layer structure of base model and meta-model. The selection criteria of base model and meta-model are:

[0082] The selection of base models usually needs to follow accuracy and diversity, that is, the selected base model should have high predictive performance, be able to provide correct feature information to the meta-model, and be able to capture different aspects of the data when predicting the same target, so as to reduce dependence on a single model.

[0083] The metamodel selection should be simple in structure and robust against overfitting. Stacking ensemble learning, due to the integration of different base models, results in a complex model structure and is prone to overfitting. By selecting a metamodel with a simple structure and robust against overfitting, we can control model complexity while minimizing the burden on the overall model complexity.

[0084] Finally, based on the prediction performance of conventional machine learning algorithm models and combined with the selection criteria of base models and meta-models, the optimal model combination is determined to construct an intelligent screening model for enhanced oil recovery technology based on Stacking.

[0085] Step 3.2: Evaluate the stacking ensemble model using accuracy, precision, recall, F1 score, confusion matrix, Kappa coefficient, and MCC value, and use the learning curve to analyze the operation process of the stacking ensemble model to identify potential problems in stacking ensemble learning.

[0086] Preferably, according to the present invention, the improved Stacking ensemble learning is:

[0087] Step 4. Based on the prediction performance of the Stacking integrated model in step 3, Stacking is improved by adding a new meta-model, and an intelligent screening integrated model for enhanced oil recovery technology based on improved Stacking is constructed. The performance of the improved Stacking integrated model is evaluated using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value as evaluation indicators to verify the superiority of the improved Stacking integrated model. Finally, a high-precision intelligent screening model for enhanced oil recovery technology based on improved Stacking is provided to participate in the final enhanced oil recovery technology screening decision.

[0088] The specific process of step 4 is as follows:

[0089] Step 4.1: Based on the improved Stacking ensemble learning structure, a suitable model is selected as a new meta-model from the conventional machine learning algorithm model constructed in step 2. Based on the Stacking ensemble model constructed in step 3.1, a new enhanced oil recovery technology screening decision model based on improved Stacking is proposed.

[0090] The improved Stacking ensemble learning algorithm uses a new meta-model to fine-tune and optimize the initial meta-model's predictions. By learning the initial meta-model's prediction patterns and underlying patterns, it identifies and corrects any potential biases or errors, thereby improving the model's prediction accuracy and stability. Furthermore, the new meta-model further enhances feature representation, better capturing complex patterns and underlying patterns in the data. Figure 3 To improve the Stacking ensemble learning algorithm structure.

[0091] The selection of the new meta-model also follows the principle of simple structure and anti-overfitting characteristics to avoid further aggravation of the overfitting problem and improve the generalization ability of the model.

[0092] The specific implementation process of improving Stacking ensemble learning is as follows:

[0093] Step 1: Dataset division: Use the training set divided in step 1.4 as training data and use K-fold cross-validation to divide it into K subsets, where K-1 subsets are sub-training sets and the remaining subset is used as sub-validation set;

[0094] Step 2: Base model training, use different base models to perform independent training on the sub-training set.

[0095] Step 3: Generate new features, use the prediction results of all base models on the sub-validation set as new features, and combine them into a new training dataset. The prediction results of the base model on the test set divided in step 1.4 are combined into a new test dataset;

[0096] Step 4: Train the meta-model, input the prediction results of the base model into the initial meta-model as new features, and maximize the accuracy of the overall model by fitting the prediction results of the base model.

[0097] Step 5: Train the new meta-model: Use the prediction results of the initial meta-model on the new training dataset to train the new meta-model.

[0098] The base model is the base model selected in step 3, the initial metamodel is the metamodel selected in step 3, and the selection criteria for the new metamodel are: based on the predictive performance of the enhanced oil recovery technology intelligent screening model based on conventional machine learning algorithms, a model with a simple structure and anti-overfitting characteristics is selected as the new metamodel.

[0099] In step 4.2, evaluate the improved stacking ensemble model using accuracy, precision, recall, F1 score, confusion matrix, Kappa coefficient, and MCC value, and compare it with the models constructed in steps 2 and 3 to verify the effectiveness of the improved stacking ensemble model. Use learning curves to analyze the performance of the improved stacking ensemble model and identify the differences between the improved and unimproved stacking ensemble models to further verify the superiority of the model.

[0100] Step 4.3: By analyzing the accuracy, precision, recall, F1 score, confusion matrix, Kappa coefficient, and MCC value of the intelligent EOR screening model constructed in Steps 2, 3, and 4.1, we determined that the EOR screening decision-making process based on improved stacking can address the impact of class imbalance at the algorithmic level while avoiding overfitting, significantly improving the accuracy of EOR screening decisions. After obtaining well logging data, oilfield development engineers can use this model to screen EOR technologies, enabling them to more accurately and quickly identify suitable EOR technologies, providing a solution for oil companies to select the optimal EOR technology.

[0101] Compared with the prior art, the present invention has the following beneficial effects:

[0102] The present invention addresses the problem that the accuracy of screening decisions of existing enhanced oil recovery technology intelligent screening models constructed based on machine learning algorithms needs to be improved due to the influence of the class imbalance problem in the data set. An enhanced oil recovery technology intelligent screening decision model based on improved Stacking is proposed. By integrating the prediction results of multiple machine learning algorithm models, the impact of the class imbalance problem is solved at the algorithm level. The present invention constructs a conventional enhanced oil recovery technology intelligent screening model through six machine learning algorithms, and optimizes it by combining random search and grid search to provide a model foundation for constructing a Stacking integrated model. Based on the Stacking integrated learning algorithm structure and combined with the prediction performance of the conventional enhanced oil recovery technology intelligent screening model, a variety of models with higher accuracy are selected as base models, and a model with a simple structure and anti-overfitting characteristics is used as a meta-model to construct an enhanced oil recovery technology intelligent screening integrated model based on Stacking. By adding a meta-model with a simple structure and anti-overfitting properties, we improve Stacking ensemble learning and build an intelligent screening ensemble model for enhanced oil recovery technology based on improved Stacking. This model can further capture complex data information, improve the generalization ability of the model, and address the limitations of machine learning algorithms in enhanced oil recovery technology screening decisions. This provides more accurate decision support for reservoir engineers and also offers new solutions for the application of machine learning in other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] Figure 1 A flow chart is constructed for the intelligent screening decision model of enhanced oil recovery technology based on improved Stacking of the present invention;

[0104] Figure 2 Integrated learning structure diagram for Stacking;

[0105] Figure 3 To improve the Stacking ensemble learning structure diagram;

[0106] Figure 4 To collect the distribution of enhanced oil recovery technology projects in the data set;

[0107] Figure 5(a) and Figure 5(b) show the distribution of characteristic parameters;

[0108] Figure 6(a) to Figure 6(f) It is the confusion matrix of the conventional machine learning algorithm model on the test set;

[0109] Figure 7 The accuracy and Kappa coefficient difference performance of the conventional machine learning algorithm model;

[0110] Figure 8 The confusion matrix of the two Stacking ensemble models on the test set;

[0111] Figure 9 The learning curves for the two stacking integration models;

[0112] Figure 10 The accuracy and Kappa coefficient difference performance of the two Stacking ensemble models;

[0113] Figure 11 Confusion matrix of the two improved Stacking ensemble models on the test set;

[0114] Figure 12 The learning curves of two improved Stacking ensemble models;

[0115] Figure 13 The difference in accuracy and Kappa coefficient between the two improved Stacking ensemble models is shown. DETAILED DESCRIPTION

[0116] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings.

[0117] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0118] To verify the feasibility of the present invention, data was collected from successful enhanced oil recovery projects around the world. The collected data set was randomly divided into training, validation, and test sets. The performance of the proposed model was evaluated on the test set when faced with unknown data. The specific steps for EOR screening using the present invention are as follows:

[0119] Step 1: Data from successfully implemented enhanced oil recovery projects around the world were collected. Lithology, porosity, permeability, reservoir depth, crude oil density, crude oil temperature, crude oil viscosity, net thickness, and initial oil saturation were selected as characteristic parameters. The data were preprocessed and analyzed to form a complete dataset containing 956 data points for intelligent screening and decision-making of enhanced oil recovery technologies.

[0120] Figure 4 Figures 5(a) and 5(b) show the distribution of the number of enhanced oil recovery technologies and the distribution of their characteristic parameters in the dataset. These results indicate that the collected dataset suffers from a significant class imbalance, which can be used to validate the effectiveness of this method in addressing this class imbalance. The characteristic parameters in the dataset cover the numerical ranges involved in most oil exploration projects and, to a certain extent, effectively reflect the universal geological characteristics of the global rocks in the dataset, ensuring the representativeness and validity of the data.

[0121] To ensure that the dataset contains all necessary information and that the training, validation, and test sets are independent of each other to prevent data leakage, we used a holdout method to randomly split the dataset into a training set and an external validation set at an 8:2 ratio. Furthermore, the external validation set was randomly split into a validation set and a test set at a 1:1 ratio. The training set was used for model building, the validation set for model optimization, and the test set for verifying the model's effectiveness on new data. A random seed was set to ensure reproducibility of the results. Table 1 shows the dataset partitioning.

[0122] Table 1 Dataset division

[0123] Dataset Number of projects training set 764 Validation set 94 Test set 98 total 956

[0124] Step 2: The selected conventional machine learning algorithm was trained on the training set and optimized on the validation set. Random search was used to rapidly explore the breadth of the hyperparameter space, narrowing the search space and improving search efficiency. A grid search was used to traverse the hyperparameter space to ensure search quality. Based on the model's prediction results on the validation set, the optimal hyperparameter combination was determined, ultimately resulting in an intelligent screening model for six conventional enhanced oil recovery technologies. Finally, the model's generalization ability was evaluated on the test set.

[0125] Table 2 shows the performance of six common machine learning algorithms on the test set, including accuracy, precision, recall, F1 score, Kappa coefficient, and MCC value. Random forest and extreme gradient boosting, as ensemble learning algorithms, achieved the best performance, with prediction accuracy reaching up to 92.7%. Neural networks and decision trees, as representatives of traditional single classifiers, performed relatively poorly. Support vector machines and logistic regression, as simple linear models, struggle to solve complex multi-classification problems and, due to factors such as class imbalance, exhibit the worst performance.

[0126] Table 2 Performance of quantitative evaluation indicators of six conventional machine learning algorithm models

[0127]

[0128]

[0129] Figure 6(a) to Figure 6(f) The confusion matrix of six conventional machine learning algorithm models on the test set is shown in Figure 2. As can be seen from the figure, the main reason affecting the performance of the model is the poor classification of enhanced oil recovery technology with relatively small data sets.

[0130] Typically, when a model is trained on a class-imbalanced dataset, its accuracy, precision, recall, and F1 score tend to be biased towards the majority class, while neglecting the minority class. However, the Kappa coefficient and MCC value can mitigate the impact of class imbalance and thus provide a more realistic assessment of the model. When these two values ​​are relatively consistent, it indicates that the model has effectively overcome the impact of class imbalance. Figure 7 The accuracy and Kappa coefficient are used as indicators to evaluate the model's ability to overcome class imbalance. Among the six conventional machine learning algorithms, Random Forest and Extreme Gradient Boosting performed relatively well in overcoming class imbalance, but there is still significant room for improvement.

[0131] Step 3: Based on the predictive performance of conventional machine learning algorithm models in Step 2, and incorporating the selection criteria for stacking ensemble learning base models and meta-models, random forest, extreme gradient boosting, neural network, and decision tree models were selected as base models, and support vector machine and logistic regression models were selected as meta-models, respectively. Two stacking-based enhanced oil recovery technology intelligent screening ensemble models (Stacking-SVM and Stacking-LR) were constructed. These four different models ensured diversity in the base models. Their high prediction accuracy and different prediction effects on different enhanced oil recovery technologies also ensured the accuracy and complementarity of the base models. Although the support vector machine and logistic regression models had relatively low accuracy, their relatively simple algorithmic structures and support for the principle of structural risk minimization resulted in good interpretability and strong resistance to overfitting.

[0132] Table 3 shows the performance of the two stacking ensemble models on the test set in terms of accuracy, precision, recall, F1 score, Kappa coefficient, and MCC value. Although the two stacking ensemble models combine the prediction results of different models, their performance is actually lower than that of random forest and extreme gradient boosting, with the highest prediction accuracy reaching only 88.6%.

[0133] Table 3 Quantitative evaluation index performance of two Stacking ensemble models

[0134] Ensemble Model Accuracy Accuracy Recall F1 score Kappa MCC Stacking-SVM 0.865 0.864 0.866 0.859 0.806 0.809 Stacking-LR 0.886 0.888 0.885 0.881 0.835 0.837

[0135] Figure 8 The confusion matrix of the two stacking ensemble models on the test set is shown in Figure 2. As can be seen from the figure, the stacking ensemble model's prediction accuracy for enhanced oil recovery technology with less data is much lower than the overall accuracy of the model, which directly affects the overall performance of the model.

[0136] To determine the cause of the stacking ensemble model's performance degradation, we used learning curves to analyze the model's operation. The learning curve can be used to determine whether the model is overfitting or underfitting. Typically, the training score on the model's learning curve is slightly higher than the test score. However, if the training score significantly exceeds the test score, the model is overfitting. If both the training and test scores are low, the model is underfitting. Figure 9 The following are the learning curves for the two stacking ensemble models. The training scores and test scores of the two models at the final convergence point differ significantly, indicating a certain degree of overfitting, which leads to reduced model performance.

[0137] Figure 10 The accuracy and Kappa coefficient were used as indicators to evaluate the stacking ensemble model's ability to overcome class imbalance. The results show that while stacking ensemble learning can combine predictions from different models, it suffers from overfitting and is inferior to random forest and extreme gradient boosting algorithms in overcoming class imbalance.

[0138] Step 4: Based on the prediction performance of the Stacking ensemble model in Step 3, Stacking was improved by adding a new meta-model to construct an intelligent screening ensemble model for enhanced oil recovery technologies based on improved Stacking. The relatively well-performing Stacking-LR model was selected as the basis for improvement. Two improved Stacking ensemble models (Mlti-Stacking-LR-SVM and Mlti-Stacking-LR-LR) were constructed using the support vector machine and logistic regression models as new meta-models, respectively.

[0139] Table 4 shows the performance of the two improved Stacking ensemble models on the test set in terms of accuracy, precision, recall, F1 score, Kappa coefficient, and MCC value. Compared with the conventional enhanced oil recovery technology intelligent screening model and the two stacking-based enhanced oil recovery technology intelligent screening ensemble models, the performance of the two improved Stacking ensemble models was significantly improved, with the highest prediction accuracy reaching 96.9%.

[0140] Table 4 Quantitative evaluation index performance of two improved Stacking ensemble models

[0141]

[0142]

[0143] Figure 11The confusion matrix of the two improved Stacking ensemble models on the test set shows that both improved Stacking ensemble models can predict all enhanced oil recovery technologies well. The Mlti-Stacking-LR-LR model achieves 100% prediction accuracy for enhanced oil recovery technologies with smaller data sets.

[0144] In order to further verify the superiority of the improved Stacking ensemble model and explore the differences in Stacking ensemble learning before and after improvement, the operation process of the improved Stacking ensemble model was analyzed using the learning curve. Figure 12 The following are the learning curves of the two improved Stacking ensemble models. The gap between the training score and the test score of the Mlti-Stacking-SVM-LR and Mlti-Stacking-LR-LR models is significantly reduced at the final convergence, indicating that they can effectively alleviate the overfitting of the Stacking ensemble model.

[0145] Figure 13 The improved Stacking ensemble model's ability to overcome class imbalance was evaluated based on the difference between accuracy and Kappa coefficient. Compared with other models, the improved Stacking ensemble model exhibited a significantly lower difference between accuracy and Kappa coefficient. These results demonstrate that the improved Stacking-based intelligent screening ensemble model for enhanced oil recovery technologies can effectively predict all types of enhanced oil recovery technologies, regardless of data size. This effectively overcomes the impact of class imbalance, thereby improving the accuracy of enhanced oil recovery technology screening decisions. Ultimately, a high-precision intelligent screening decision model for enhanced oil recovery technologies based on improved Stacking is developed to participate in the final enhanced oil recovery technology screening decision. When presented with completely new reservoir data, its decision-making accuracy can reach up to 96.87%.

[0146] To address the problem that intelligent screening decisions for enhanced oil recovery technologies based on machine learning algorithms are highly dependent on datasets and susceptible to class imbalance in datasets, resulting in a need for improved decision-making accuracy, this paper provides a high-precision intelligent screening decision-making framework for enhanced oil recovery technologies based on improved stacking. By integrating the prediction results of multiple machine learning algorithm models and using a two-layer meta-model framework in tandem, this framework effectively overcomes the impact of class imbalance while avoiding overfitting. This decision-making method significantly improves the accuracy of enhanced oil recovery technology screening decisions while improving the efficiency and cost of screening.

Claims

1. A method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved stacking, characterized in that: The steps include: Step 1: Collect data from successfully implemented enhanced oil recovery projects around the world, select appropriate reservoir rock and fluid characteristics as characteristic parameters, and form a complete data set through data processing and analysis; Step 2: Select multiple conventional machine learning algorithms, use the characteristic parameters in the complete data set obtained in Step 1 as input, and the enhanced oil recovery technologies in the complete data set as output, and construct multiple intelligent screening models for enhanced oil recovery technologies based on conventional machine learning algorithms; Step 3: Based on the prediction performance of multiple enhanced oil recovery technology intelligent screening models based on conventional machine learning algorithms constructed in Step 2, Stacking ensemble learning is used to fuse and construct a Stacking-based enhanced oil recovery technology intelligent screening model; Step 4: Based on the prediction performance of the intelligent screening model for enhanced oil recovery technology based on Stacking constructed in Step 3, Stacking is improved by adding a new meta-model to construct an integrated model for intelligent screening of enhanced oil recovery technology based on improved Stacking; Step 3 is as follows: Step 3.1: Using the Stacking ensemble learning method, the prediction results of multiple different base models are integrated by using a meta-model, specifically: Step 3.1.

1. Dataset division: Use the training set divided in step 1.4 as the training data and use K-fold cross-validation to divide it into K subsets, where K-1 subsets are sub-training sets and the remaining subset is used as the sub-validation set; Step 3.1.2, base model training: Use different base models to perform independent training on the sub-training sets; Step 3.1.3: Generate new features: Use the prediction results of all base models on the sub-validation set as new features to form a new training dataset. Use the prediction results of the base models on the test set divided in step 1.4 to form a new test dataset. Step 3.1.4: Train the meta-model: Input the prediction results of the base model as new features into a meta-model. By fitting the prediction results of the base model, the accuracy of the overall model is maximized to obtain the intelligent screening model of enhanced oil recovery technology based on stacking. The base model and meta-model are selected from the multiple intelligent screening models for enhanced oil recovery technology based on conventional machine learning algorithms constructed in step 2. The selection criteria for the base model are: based on the prediction performance of the intelligent screening model for enhanced oil recovery technology based on conventional machine learning algorithms, different models that can provide correct feature information to the meta-model and capture different aspects of the data when predicting the same target are selected as the base model; The selection criteria for the meta-model are: based on the prediction performance of the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithms, a model with a simple structure and anti-overfitting characteristics is selected as the meta-model; Step 3.2: Evaluate the intelligent screening model of enhanced oil recovery technology based on Stacking through accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value, and use the learning curve to analyze the operation process of the intelligent screening model of enhanced oil recovery technology based on Stacking.

2. The construction method according to claim 1, characterized in that Step 1 includes the following sub-steps: Step 1.1: Collect data from successfully implemented enhanced oil recovery projects around the world and select reservoir rock and fluid properties that effectively reflect the actual reservoir and fluid conditions and are closely related to the enhanced oil recovery process as characteristic parameters; Step 1.2: Using the results of the semi-annual enhanced oil recovery project survey published by Oil & Gas Journal as the data source, we supplemented the data with other relevant documents from around the world to form the original data set. Step 1.3: preprocess the data in the original dataset, including deleting duplicate data, replacing missing values ​​with feature means, and performing one-hot encoding on discrete parameters to form a complete dataset, and logarithmize the continuous feature parameters in the dataset; Step 1.4: Use the holdout method to randomly divide the complete dataset obtained in step 1.3 into training set, validation set, and test set in a ratio of 8:1:

1.

3. The construction method according to claim 1 or 2, characterized in that The characteristic parameters include: lithology, porosity, permeability, reservoir depth, crude oil specific gravity, crude oil temperature, crude oil viscosity, net thickness and initial oil saturation.

4. The construction method according to claim 2, characterized in that In step 2, the conventional machine learning algorithms include random forest algorithm, extreme gradient boosting algorithm, neural network algorithm, decision tree algorithm, support vector machine algorithm and logistic regression algorithm.

5. The construction method according to claim 4, characterized in that Step 2 is as follows: Step 2.

1. Determine the evaluation indicators: Accuracy, precision, recall, and F1 score are selected as general evaluation indicators to evaluate the overall performance of the model. Confusion matrix, Kappa coefficient, and MCC value are used to evaluate the model's ability to overcome class imbalance. Step 2.2: Determine the hyperparameter optimization method: Use a combination of random search and grid search to optimize the model. Step 2.3: Use multiple conventional machine learning algorithms to learn the complete data set obtained in step 1, optimize it using the hyperparameter optimization method in step 2.2, and construct multiple intelligent screening models for enhanced oil recovery technologies based on conventional machine learning algorithms; Step 2.4: Use the evaluation indicators determined in step 2.1 to evaluate the predictive performance of the intelligent screening model for enhanced oil recovery technology based on conventional machine learning algorithms constructed in step 2.

3.

6. The construction method according to claim 5, characterized in that: Step 2.3 is as follows: (1) Constructing an intelligent screening model for enhanced oil recovery technology based on random forest algorithm: Step 1: Through self-service sampling, randomly form multiple sub-datasets for independent training of a single decision tree; Step 2: Randomly select multiple features from the feature space to form a feature subset. When constructing each node of the decision tree, select the optimal feature from the feature subset for splitting; Step 3: Repeat the above steps to generate multiple decision trees to form a random forest; (2) Constructing an intelligent screening model for enhanced oil recovery technology based on the extreme gradient boosting algorithm: Step 1: Initialize a constant model as the initial prediction value; Step 2: Iteratively construct multiple decision trees. In each iteration, determine a decision tree that minimizes the objective function. Step 3: After each decision tree is built, its prediction result is added to the previous prediction value to obtain a new prediction value; Step 4: Continue the iteration process until the preset number of iterations is reached or the change in the objective function is less than a certain threshold; (3) Constructing an intelligent screening model for enhanced oil recovery technology based on neural network algorithm: Step 1: Forward propagation, from the input layer to the output layer to calculate the output of the network; Step 2: Calculate the loss and calculate the loss function based on the predicted results and the true label; Step 3: Back propagation, calculate the gradient from the output layer to the input layer, and update the weights and biases; Step 4: Iterative training, repeat the above process until the network converges; (4) Constructing an intelligent screening model for enhanced oil recovery technology based on a decision tree algorithm: Step 1: Starting from the root node, select the feature that can minimize the impurity of the dataset as the split feature; Step 2: Divide the dataset into multiple subsets based on the selected features; Step 3: Repeat the above process for the divided subsets and recursively construct decision tree branches until the stopping condition is met; (5) Constructing an intelligent screening model for enhanced oil recovery technology based on support vector machine algorithm: The calculation formula for the classification result obtained by the support vector machine algorithm is: In formula (12), f(x) represents the classification result of input data x, x represents the input data to be classified, i represents the sample number, and its value is 1, 2, ..., n, where n is the number of samples, and α i is the Lagrange multiplier, y i is the label of the i-th sample, K(x,x i ) is the kernel function, x i represents the feature vector of the i-th sample, b is the bias term, when f(x)>0, sample x belongs to the positive category; when f(x)<0, sample x belongs to the negative category; (6) Constructing an intelligent screening model for enhanced oil recovery technology based on logistic regression algorithm: The calculation formula for obtaining the classification result of the logistic regression algorithm is: z=w T x1+b(13) In formula (13), z represents the output value given based on the input x1, w T represents the transpose of the weight vector, x1 represents the eigenvector, which contains the eigenvalues ​​of the input data, and b is the bias term; In formula (14), σ(z) is the class probability. When σ(z) ≥ 0.5, the sample is classified as a positive sample; when σ(z) < 0.5, the sample is classified as a negative sample.

7. The construction method according to claim 1, characterized in that Step 4 is as follows: Step 4.1: Based on the prediction performance of the intelligent screening model for enhanced oil recovery technology based on Stacking in step 3.2, a suitable model is selected as a new meta-model from the intelligent screening model for enhanced oil recovery technology based on conventional machine learning algorithms constructed in step 2. Based on the intelligent screening model for enhanced oil recovery technology based on Stacking constructed in step 3.1, an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking is proposed. Specifically, Step 4.1.

1. Dataset division: Use the training set divided in step 1.4 as the training data and use K-fold cross-validation to divide it into K subsets, where K-1 subsets are sub-training sets and the remaining subset is used as the sub-validation set; Step 4.1.2: Base model training: Use different base models to perform independent training on the sub-training sets; Step 4.1.3: Generate new features: Combine the prediction results of all base models on the sub-validation set as new features to form a new training dataset. Combine the prediction results of the base models on the test set divided in step 1.4 to form a new test dataset as new features for the test set. Step 4.1.4: Train the meta-model: Input the prediction results of the base model as new features into the initial meta-model, and maximize the accuracy of the overall model by fitting the prediction results of the base model; Step 4.1.5, training new meta-model: Use the prediction results of the initial meta-model on the new training data set to train the new meta-model to obtain the intelligent screening decision model of enhanced oil recovery technology based on improved stacking.

8. The construction method according to claim 7, characterized in that: In step 4.1, the base model is the base model selected in step 3, the initial metamodel is the metamodel selected in step 3, and the selection criteria for the new metamodel are: based on the predictive performance of the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithms, a model with a simple structure and anti-overfitting characteristics is selected as the new metamodel.

Citation Information

Patent Citations

  • Method for predicting ground synthetic electric field based on improved Stacking algorithm

    CN115544879A

  • Drilling and production cost prediction method based on feature selection and stacked heterogeneous integrated learning

    CN117648646A