Construction method for intelligent screening decision model of enhanced oil extraction technology based on improved Stacking
By constructing an intelligent screening decision model for strengthened oil production technology based on improved Stacking, the problem of difficult screening of strengthened oil production technology and unbalanced categories in the existing technology affecting decision-making accuracy, achieving more efficient and accurate screening of strengthened oil production and decision-making has been improved.
Patent Information
- Application Number
- CN202510197443.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-23
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing screening methods for strengthened oil production technology are limited by the characteristics of reservoir rocks and fluids, so it is difficult to accurately select the best strengthened oil production technology, resulting in insufficient improvement in oil production and the problem of category imbalance affecting decision-making accuracy.
Using an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking, we will collect data on enhanced oil recovery projects successfully implemented globally, build a variety of conventional machine learning algorithm models, and use Stacking integrated learning to integrate, improve the model to overcome category imbalance problems and the shortcomings of traditional Stacking models.
The accuracy and efficiency of enhanced oil production technology screening can be improved, and the best enhanced oil production technology can be selected more scientifically, which can enhance the improvement of oil production and reduce reservoir damage and economic losses.
Smart Images

Figure CN120046498A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent development research of oil and gas fields, and specifically relates to a method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking. Background Technique
[0002] With the rapid development of the economy, the demand for energy in various countries around the world is increasing day by day. As one of the main global energy sources, oil is also an important raw material for modern industrial society and is extremely important for the contemporary social and industrial development. Enhanced oil recovery technology, as an effective solution to increase oil production, improves the physical and chemical properties of reservoir formations and fluids by applying advanced technologies and means such as physics or chemistry, increases the fluidity of crude oil in the reservoir, and improves the oil recovery rate, thus attracting extensive attention. With the progress of technology, more than 20 enhanced oil recovery technologies have been widely applied around the world. However, restricted by conditions such as reservoir rock and fluid characteristics, reservoir engineers need to consider multiple factors comprehensively and select the most suitable technology from numerous enhanced oil recovery technologies to achieve the best benefits. How to select the best enhanced oil recovery technology according to the reservoir parameters of the target reservoir to maximize oil production has become the main challenge faced by reservoir engineers. The screening of enhanced oil recovery technology can select the best feasible enhanced oil recovery technology for the target reservoir by following the multi-criteria decision-making process according to the reservoir rock characteristics and fluid properties. As the first step in implementing an enhanced oil recovery project, the efficient and accurate screening results have a direct impact on improving the implementation efficiency of the enhanced oil recovery project and whether the expected oil production effect can be achieved. An inappropriate enhanced oil recovery technology will not only cause permanent damage to the reservoir but also result in significant economic losses. Therefore, accurate and reliable screening decisions for enhanced oil recovery technology are a prerequisite for the planning and design of enhanced oil recovery projects in the early stage of reservoir development.
[0003] Currently, the screening methods for enhanced oil recovery technology at home and abroad are mainly divided into three types:
[0004] The first one is laboratory simulation / field pilot test. According to the characteristics of reservoir rocks and fluids, the implementation process of enhanced oil recovery technology is simulated to determine the appropriate enhanced oil recovery technology. This method needs to simulate a variety of enhanced oil recovery technologies, which has the defects of long time consumption and high screening cost. The second one is the screening decision of conventional enhanced oil recovery technology. By obtaining reservoir and fluid parameters and the enhanced oil recovery technology used from past successfully implemented enhanced oil recovery projects, the corresponding screening criteria for enhanced oil recovery technology are formulated. This criterion is usually presented in the form of a chart. By considering several predefined screening parameters, the possibility of successfully implementing each enhanced oil recovery technology is determined. However, when searching for reservoir parameters through the screening criteria of enhanced oil recovery technology, there may be multiple matching enhanced oil recovery technologies at the same time, and the screening results are not unique. The third one is the screening decision of advanced enhanced oil recovery technology. Using computer technology represented by machine learning, the potential relationship between reservoir rock and fluid characteristics and the successful implementation of enhanced oil recovery technology is explored from past successful enhanced oil recovery projects, and valuable screening rules for enhanced oil recovery technology are established. Although this decision-making process can provide specific solutions in a short time, it has a strong dependence on the data set. There are huge differences in the implementation quantities of different enhanced oil recovery technologies globally. Most of the data sets used for the screening decision of advanced enhanced oil recovery technology have the problem of class imbalance. When the number of samples in a certain class in the data set is much larger than that of other classes, computer technology may tend to predict the majority class and ignore the importance of the minority class. Therefore, affected by the class imbalance problem, its decision-making accuracy needs to be further improved. Summary of the Invention
[0005] To solve the above-mentioned drawbacks of the prior art, the present invention discloses a construction method of an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking. The present invention conducts statistics on the enhanced oil recovery technologies successfully implemented globally, constructs and optimizes multiple intelligent screening models for enhanced oil recovery technology based on conventional machine learning algorithms, uses Stacking ensemble learning for fusion to construct a Stacking-based intelligent screening ensemble model for enhanced oil recovery technology, and improves it to construct a new intelligent screening model for enhanced oil recovery technology based on improved Stacking, so as to overcome the influence of the class imbalance problem on the model performance, make up for the defects of the Stacking ensemble learning model, and further improve the screening accuracy of enhanced oil recovery technology, thereby providing more efficient, scientific and intelligent decision-making support for oil companies to select the best enhanced oil recovery technology.
[0006] The present invention specifically adopts the following technical solutions:
[0007] A construction method of an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking, as Figure 1 shown, includes the following steps:
[0008] First, collect data on globally successful enhanced oil recovery (EOR) projects, and form a complete dataset through data processing and analysis for use in EOR technology screening decisions.
[0009] Secondly, construct and optimize multiple intelligent screening models for EOR technologies based on conventional machine learning algorithms.
[0010] Furthermore, based on the prediction performance of the conventional machine learning algorithm models, use Stacking ensemble learning for fusion to construct a Stacking-based intelligent screening model for EOR technologies.
[0011] Finally, improve the Stacking ensemble model to construct a new type of intelligent screening ensemble model for EOR technologies based on improved Stacking, to overcome the impact of the class imbalance problem on the model performance, and at the same time make up for the defects of the traditional Stacking ensemble learning model, thereby improving the screening accuracy of EOR technologies.
[0012] Preferably according to the present invention, the sources of the EOR project data include:
[0013] Step 1: Statistically analyze the data of globally successful EOR projects, select reservoir rock and fluid properties that can effectively reflect the true reservoir and fluid conditions of the reservoir and are closely related to the oil displacement process of EOR technologies as characteristic parameters, and obtain a complete dataset through data preprocessing and analysis.
[0014] Preferably according to the present invention, the conventional machine learning algorithms include:
[0015] Step 2: According to the current research status, select random forest, extreme gradient boosting, neural network, decision tree, support vector machine, and logistic regression as conventional machine learning algorithms, use them to learn the dataset obtained in Step 1.3, take the characteristic parameters as input and the EOR technology as the output result, construct multiple conventional intelligent screening models for EOR technologies, and evaluate the performance of the conventional machine learning algorithm models with accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators to provide a model basis for subsequent construction of the ensemble model.
[0016] Preferably according to the present invention, the Stacking ensemble learning algorithm is:
[0017] Step 3: Based on the prediction performance of the conventional machine learning algorithm model in Step 2, use Stacking ensemble learning for fusion to construct an intelligent screening model for enhanced oil recovery technology based on Stacking. Evaluate the performance of the Stacking ensemble model using accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators, providing an improvement basis for the subsequent construction of an improved Stacking ensemble model.
[0018] Preferably according to the present invention, the improved Stacking ensemble learning algorithm is:
[0019] Step 4: Improve Stacking by adding a new meta-model to construct an intelligent screening ensemble model for enhanced oil recovery technology based on the improved Stacking. Evaluate the performance of the improved Stacking ensemble model using accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators to verify the superiority of the improved Stacking ensemble model, and finally provide a high-precision intelligent screening ensemble model for enhanced oil recovery technology based on the improved Stacking to participate in the final screening decision of enhanced oil recovery technology.
[0020] Preferably according to the present invention, the specific process of Step 1 is:
[0021] According to the previous enhanced oil recovery technology screening criteria and the enhanced oil recovery technology displacement mechanism, through Steps 1.1 and 1.2, feature selection and data collection are carried out to generate an original data set, and through Steps 1.3 and 1.4, data preprocessing and data analysis are carried out on the data to construct a complete data set finally used for screening decision-making.
[0022] Step 1.1: According to the research experience of past scholars on enhanced oil recovery technology screening and calculation formulas such as reservoir volume, select reservoir rock and fluid characteristics that can effectively reflect the true reservoir and fluid conditions of the reservoir and are closely related to the enhanced oil recovery technology displacement process as feature parameters. The specific parameters include: lithology, porosity, permeability, reservoir depth, crude oil specific gravity, crude oil temperature, crude oil viscosity, net thickness, and initial oil saturation.
[0023]
[0024] In formulas (1)-(3), OOIP is the original oil reservoir reserve, q is the volume flow rate of the fluid in the porous medium, A is the oil-bearing area, h is the net thickness, is the porosity, S 0 is the initial crude oil saturation, B oiis the average original crude oil volume coefficient, k is the permeability, μ is the crude oil viscosity, Δp is the pressure difference at both ends, API is the crude oil specific gravity, and SG is the relative density of the crude oil at 60°F. OOIP and q are important factors determining oil production and recovery efficiency. Formulas (1) to (3) involve some specific parameters in Step 1.1, indicating that the characteristic parameters selected in the present invention are closely related to the oil production process. One of the main methods to improve oil recovery currently is to increase the fluidity of crude oil. Therefore, it can be indirectly shown through this formula that the characteristic parameters selected in the present invention have a direct impact on improving oil recovery.
[0025] Step 1.2: Using the semi-annual Enhanced Oil Recovery (EOR) project survey results in the Journal of Petroleum and Gas as the main data source, supplement the data through the US Department of Energy reports, the American Association of Petroleum Geologists database, field reports, Chinese publications, and Society of Petroleum Engineers publications, etc., to form an original dataset.
[0026] Step 1.3: Process the dataset collected in Step 1.2, including deleting duplicate data, replacing missing values with characteristic means, and performing One-Hot encoding on discrete parameters, etc., to form a complete dataset for EOR technology screening decisions.
[0027] Use formula (4) to perform logarithmic processing on the widely distributed continuous characteristic parameters in the dataset to narrow the numerical range, increase the stability of the data, and avoid the negative impact of extreme data on model training.
[0028] x * = lg(x org ) (4)
[0029] In formula (4), x * represents the data after logarithmic processing; x org represents the original sample data;
[0030] Step 1.4: Analyze the complete dataset formed in Step 1.3, including the quantity distribution of each EOR technology and the distribution of characteristic parameters, explore the class imbalance problem existing in the dataset, ensure that the dataset can effectively reflect the general geological characteristics of global rocks in the dataset, and analyze the correlation between characteristic parameters using the Spearman correlation coefficient to provide a data basis for the subsequent construction of the EOR technology screening model. The calculation formula of the Spearman correlation coefficient is:
[0031]
[0032] In formula (5), r s is the Spearman correlation coefficient, d iis the rank difference between two variables, that is, for the i-th observation, it is the difference in ranks on the two variables, and n is the number of observations.
[0033] Step 1.5: Use the holdout method to randomly divide the complete dataset formed in Step 1.3 into a training set, a validation set, and a test set according to a ratio of 8:1:1 for the training, optimization, and evaluation of the intelligent screening model for enhanced oil recovery technology.
[0034] Preferably according to the present invention, the specific process of Step 2 is as follows:
[0035] Step 2: According to the current research status, select random forest, extreme gradient boosting, neural network, decision tree, support vector machine, and logistic regression as conventional machine learning algorithms, use them to learn the dataset obtained in Step 1.3, take the feature parameters as the input, and the enhanced oil recovery technology as the output result, construct multiple conventional intelligent screening models for enhanced oil recovery technology, and evaluate the performance of the conventional machine learning algorithm models with accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators to provide a model basis for the subsequent construction of the integrated model.
[0036] Step 2.1: Determine the evaluation indicators. Select accuracy, precision, recall rate, and F1 score as conventional evaluation indicators to evaluate the overall performance of the model, draw a confusion matrix and use the Kappa coefficient and MCC value to evaluate the ability of the model to overcome the problem of class imbalance. Among them, accuracy, precision, recall rate, and F1 score are affected by class imbalance and have high values. The Kappa coefficient and MCC value can ignore the influence of class imbalance. The comparison between the two can provide a more accurate evaluation for the model. Through the confusion matrix, the classification effect of the model on each enhanced oil recovery technology can be determined. The calculation formulas are as follows:
[0037]
[0038] In formulas (6)-(11), TP is the number of positive examples correctly predicted as positive examples by the model, TN is the number of negative examples correctly predicted as negative examples by the model, FP is the number of negative examples wrongly predicted as positive examples by the model, FN is the number of positive examples wrongly predicted as negative examples by the model, A is the accuracy, and A r is the accuracy of random classification. The values of all quantitative evaluation indicators are between -1 and 1, and the closer the value is to 1, the better the classification effect of the model.
[0039] Step 2.2: Determine the hyperparameter optimization method. Select a combination of random search and grid search for model optimization. Random search can quickly explore the breadth of the hyperparameter space, narrow the search space, and improve search efficiency. Grid search is used to traverse the hyperparameter space to ensure search quality and determine the best hyperparameter combination for all models.
[0040] Step 2.3: Use the random forest algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to build an intelligent screening model for enhanced oil recovery technology based on the random forest algorithm. The random forest algorithm is an ensemble learning algorithm based on decision trees. Multiple independent decision tree models are created through bootstrap sampling, and the final decision result is obtained by voting or taking the average. The randomness of its sample and feature selection makes it more suitable for high-dimensional and large-scale data, increases the diversity of its decision-making process, and avoids overfitting due to the model's excessive focus on certain characteristics in the data. The parallel integration feature enables it to have strong robustness to noise and outliers.
[0041] The specific process of the random forest algorithm is as follows:
[0042] Step1: Randomly form multiple sub-datasets through bootstrap sampling for independent training of individual decision trees;
[0043] Step2: Randomly select multiple features from the feature space to form a feature subset. When constructing each node of the decision tree, select the optimal feature from the feature subset for splitting;
[0044] Step3: Repeat the above steps to generate multiple decision trees to form a random forest.
[0045] Step 2.4: Use the extreme gradient boosting algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to build an intelligent screening model for enhanced oil recovery technology based on the extreme gradient boosting algorithm. The extreme gradient boosting algorithm is an ensemble learning algorithm based on gradient boosting decision trees. Multiple gradient boosting decision trees are constructed in series and learned in an orderly manner in a highly adaptive method, and finally the prediction result is obtained by weighting.
[0046] By introducing the second-order Taylor formula and regularization term, this algorithm accelerates convergence, avoids overfitting, and efficiently processes samples with missing feature values using sparse perception. The extreme gradient boosting algorithm is widely used because of its advantages of high efficiency, flexibility, and scalability in high-dimensional data processing, feature selection, and handling of missing values, and performs excellently in various machine learning tasks.
[0047] The specific process of the extreme gradient boosting algorithm is as follows:
[0048] Step1: Initialize a constant model as the initial prediction value;
[0049] Step2: Iteratively construct multiple decision trees. In each iteration, determine a decision tree that can minimize the objective function;
[0050] Step3: After constructing each decision tree, add its prediction result to the previous prediction value to obtain a new prediction value.
[0051] Step4: Continuously perform the iterative process until the preset number of iterations is reached or the change in the objective function is less than a certain threshold.
[0052] Step 2.5: Use the neural network algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the neural network algorithm. The neural network algorithm is a machine learning algorithm used to simulate the interconnections and information transmission between neurons in the human brain to achieve data processing and pattern recognition. It realizes the learning and processing of input data by adjusting the connection weights and activation functions between neurons. During the operation of the model, this algorithm can adjust its own connection weights and biases through learning, gradually improving its performance, and thus showing good adaptability. The non-linear activation function enables the neural network to have a powerful non-linear modeling ability, capable of effectively processing complex non-linear relationships and patterns.
[0053] The specific process of the neural network algorithm is as follows:
[0054] Step1: Forward propagation, calculate the output of the network from the input layer to the output layer;
[0055] Step2: Calculate the loss, calculate the loss function according to the prediction result and the true label;
[0056] Step3: Backward propagation, calculate the gradient from the output layer to the input layer, and update the weights and biases;
[0057] Step4: Iterative training, repeat the above process until the network converges.
[0058] Step 2.6: Use the decision tree algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the decision tree algorithm. The decision tree algorithm is an important classification and regression method in data mining technology. By asking questions about the features in the dataset and gradually splitting the data according to the value results of different features, a tree-like structure is finally generated to predict the category or value of new data. This algorithm can dynamically select the best splitting point and splitting method according to the characteristics of the data and the nature of the features. Due to its flexibility and adaptability in feature selection and splitting nodes, the decision tree can handle both numerical and categorical features simultaneously without the need for complex data transformation or preprocessing.
[0059] The specific process of the decision tree algorithm is as follows:
[0060] Step1: Start from the root node and select the feature that can minimize the impurity of the dataset to the greatest extent as the splitting feature;
[0061] Step2: According to the selected feature, divide the dataset into multiple subsets;
[0062] Step3: For the divided subsets, repeat the above process and recursively construct the decision tree branches until the stopping condition is met.
[0063] Step 2.7: Use the support vector machine algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the support vector machine algorithm. As a generalized linear classifier that can classify data, the support vector machine algorithm finds a linear or non-linear hyperplane to maximize the distance between the sample points on both sides of the hyperplane, and determines the final classification result according to the distance from the sample to the hyperplane. This algorithm can flexibly adapt to different data distributions by selecting different kernel functions. Its final decision only depends on the points closest to the classification boundary in the training samples, so it has strong generalization ability and robustness. Based on the principle of structural risk minimization, the support vector machine algorithm has strong anti-overfitting characteristics.
[0064] The calculation formula for the support vector machine algorithm to obtain the classification result is:
[0065]
[0066] In formula (12), f(x) represents the classification result for the input data x, x represents the input data to be classified, i represents the sample number, taking values 1, 2, …, n, n is the number of samples, y i is the label of the i-th sample, K(x, x i ) is the kernel function, xi represents the feature vector of the i-th sample, α i is the Lagrange multiplier, b is the bias term. When f(x) > 0, the sample x belongs to the positive class; when f(x) < 0, the sample x belongs to the negative class.
[0067] Step 2.8: Use the logistic regression algorithm to learn the dataset obtained in Step 1.3, and optimize it using the hyperparameter optimization method in Step 2.2 to construct an intelligent screening model for enhanced oil recovery technology based on the logistic regression algorithm. As a generalized linear model, the logistic regression algorithm classifies data through a linear combination of a set of input features, and transforms the output of the linear regression model into the probability of a class through a logistic function to obtain the final classification result. Due to its linear structure, the logistic regression algorithm has a relatively simple model and fast calculation speed. Following the principle of minimizing structural risk, it has strong anti-overfitting characteristics. The calculation formula for the logistic regression algorithm to obtain the classification result is:
[0068] z = w T x 1 +b(13)
[0069]
[0070] Formula (13) is the linear combination of the input features. Among them, z represents the output value based on the input x 1 given, x 1 represents the feature vector, which contains each feature value of the input data, w represents the weight vector, which contains the weights of each feature in the model, w T represents the transpose of the weight vector, b is the bias term; σ(z) is the class probability. When σ(z) ≥ 0.5, the sample is classified as a positive sample; when σ(z) < 0.5, the sample is classified as a negative sample.
[0071] Step 2.9: Evaluate six conventional machine learning algorithm models through accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient, and MCC value, and sort them according to their performance to provide a model basis for constructing an ensemble model later.
[0072] Preferably according to the present invention, the Stacking ensemble learning is:
[0073] Step 3: Based on the prediction performance of the conventional machine learning algorithm models in Step 2, use Stacking ensemble learning for fusion to construct an intelligent screening model for enhanced oil recovery technology based on Stacking, and evaluate the performance of the Stacking ensemble model using accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators to provide an improvement basis for constructing an improved Stacking ensemble model later.
[0074] The specific process of step 3 is as follows:
[0075] Step 3.1: As an ensemble learning method for improving prediction results, Stacking ensemble learning integrates the prediction results of multiple different base models through a meta-model to obtain the final prediction result. Figure 2 is the algorithm structure of Stacking ensemble learning.
[0076] The specific implementation process of Stacking ensemble learning is as follows:
[0077] Step1: Dataset partitioning. Take the training set partitioned in step 1.4 as the training data, and use K-fold cross-validation to partition it into K subsets, where K - 1 subsets are sub-training sets and the remaining one subset is the sub-validation set;
[0078] Step2: Base model training. Use different base models to perform independent training on the sub-training sets;
[0079] Step3: Generate new features. Take the prediction results of all base models on the sub-validation set as new features and combine them into a new training dataset. The prediction results of the base models on the test set partitioned in step 1.4 form a new test dataset;
[0080] Step4: Train the meta-model. Take the prediction results of the base models as new features and input them into a meta-model. By fitting the prediction results of the base models, the accuracy of the overall model is maximized.
[0081] Stacking ensemble learning usually has a two-layer structure of a base model and a meta-model. The selection criteria for the base model and the meta-model are as follows:
[0082] The selection of the base model usually needs to follow accuracy and diversity, that is, the selected base model should have high prediction performance, be able to provide correct feature information to the meta-model, and be able to capture different aspects of the data when predicting the same target to reduce the dependence on a single model.
[0083] The selection of the meta-model follows simplicity of structure and anti-overfitting characteristics. Since Stacking ensemble learning fuses different types of base models, the model structure is complex and prone to the risk of overfitting. By selecting a meta-model with a simple structure and anti-overfitting characteristics, on the one hand, the complexity of the model can be controlled, and on the other hand, the burden on the complexity of the overall model can be minimized as much as possible.
[0084] Finally, based on the prediction performance of the conventional machine learning algorithm model, combined with the selection criteria of the base model and the meta-model, the best model combination is determined to construct an intelligent screening model for enhanced oil recovery technology based on Stacking.
[0085] Step 3.2: Evaluate the Stacking ensemble model through accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value, and analyze the running process of the Stacking ensemble model using learning curves to identify potential problems in Stacking ensemble learning.
[0086] Preferably according to the present invention, the improved Stacking ensemble learning is as follows:
[0087] Step 4: Based on the prediction performance of the Stacking ensemble model in Step 3, improve Stacking by adding a new meta-model to construct an intelligent screening ensemble model for enhanced oil recovery technology based on improved Stacking. Evaluate the performance of the improved Stacking ensemble model using accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value as evaluation indicators to verify the superiority of the improved Stacking ensemble model. Finally, provide a high-precision intelligent screening model for enhanced oil recovery technology based on improved Stacking to participate in the final decision-making of enhanced oil recovery technology screening.
[0088] The specific process of Step 4 is as follows:
[0089] Step 4.1: Based on the improved Stacking ensemble learning structure, select a suitable model from the conventional machine learning algorithm models constructed in Step 2 as the new meta-model, and propose a new type of screening decision model for enhanced oil recovery technology based on improved Stacking on the basis of the Stacking ensemble model constructed in Step 3.1.
[0090] The improved Stacking ensemble learning algorithm makes more refined adjustments and optimizations to the prediction results of the initial meta-model by adding a new meta-model. By learning the prediction patterns and potential laws of the initial meta-model, it discovers and corrects possible prediction biases or errors, thereby improving the prediction accuracy and stability of the model. In addition, the new meta-model can further enhance the feature representation ability and better capture complex patterns and deep-seated laws in the data. Figure 3 To improve the structure of the Stacking ensemble learning algorithm.
[0091] The selection of the new added meta-model also follows the principles of simple structure and anti-overfitting characteristics to avoid further aggravation of the overfitting problem and improve the generalization ability of the model.
[0092] The specific implementation process of the improved Stacking ensemble learning is as follows:
[0093] Step 1: Dataset division. Use the training set divided in Step 1.4 as the training data, and use K-fold cross-validation to divide it into K subsets, where K - 1 subsets are sub-training sets and the remaining one subset is used as the sub-validation set.
[0094] Step 2: Base model training. Use different base models to perform independent training on the sub-training sets.
[0095] Step 3: Generate new features. Use the prediction results of all base models on the sub-validation set as new features and combine them into a new training dataset. Use the prediction results of the base models on the test set divided in Step 1.4 to form a new test dataset.
[0096] Step 4: Train the meta-model. Use the prediction results of the base models as new features and input them into the initial meta-model. By fitting the prediction results of the base models, maximize the accuracy of the overall model.
[0097] Step 5: Train the new meta-model. Use the prediction results of the initial meta-model on the new training dataset to train the new meta-model.
[0098] The base model is the base model selected in Step 3, the initial meta-model is the meta-model selected in Step 3, and the selection criterion for the new added meta-model is: according to the prediction performance of the intelligent screening model for enhanced oil recovery technology based on conventional machine learning algorithms, select a model with a simple structure and anti-overfitting characteristics as the new added meta-model.
[0099] Step 4.2: Evaluate the improved Stacking ensemble model through accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value, and compare it with the models constructed in Step 2 and Step 3 to verify the effectiveness of the improved Stacking ensemble model. Use the learning curve to analyze the running process of the improved Stacking ensemble model to determine the differences between the improved and unimproved Stacking ensemble models, so as to further verify the superiority of the model.
[0100] Step 4.3: By analyzing the performance of the intelligent screening model for enhanced oil recovery technology constructed in Step 2, Step 3, and Step 4.1 in terms of accuracy, precision, recall, F1-score, confusion matrix, Kappa coefficient, and MCC value, determine that the screening decision for enhanced oil recovery technology based on the improved Stacking can solve the impact of the class imbalance problem at the algorithm level and avoid the overfitting problem, so as to significantly improve the accuracy of the screening decision for enhanced oil recovery technology. After obtaining the logging data, oilfield development engineers can use this model to screen enhanced oil recovery technologies and can find suitable enhanced oil recovery technologies more accurately and quickly, providing a solution for oil companies to select the best enhanced oil recovery technology.
[0101] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0102] In view of the problem that the intelligent screening model of enhanced oil recovery technology constructed based on machine learning algorithms in the prior art has room for improvement in the accuracy of screening decisions due to the influence of class imbalance problems in the dataset, the present invention proposes an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking. By integrating the prediction results of multiple machine learning algorithm models, the influence of class imbalance problems is solved at the algorithm level. The present invention constructs a conventional intelligent screening model for enhanced oil recovery technology through six machine learning algorithms, and uses a combination of random search and grid search for optimization to provide a model basis for constructing a Stacking integration model. Based on the Stacking ensemble learning algorithm structure, combined with the prediction performance of the conventional intelligent screening model for enhanced oil recovery technology, multiple models with higher accuracy are selected as base models, and a model with a simple structure and anti-overfitting characteristics is selected as a meta-model to construct an intelligent screening integration model for enhanced oil recovery technology based on Stacking. By adding a meta-model with a simple structure and anti-overfitting characteristics to improve Stacking ensemble learning, an intelligent screening integration model for enhanced oil recovery technology based on improved Stacking is constructed, further capturing complex data information, improving the generalization ability of the model, solving the limitations of machine learning algorithms in the screening decision of enhanced oil recovery technology, providing more accurate decision support for reservoir engineers, and at the same time providing new solutions for the application of machine learning in other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 is the flowchart for constructing the intelligent screening decision model for enhanced oil recovery technology based on improved Stacking of the present invention;
[0104] Figure 2 is the structure diagram of Stacking ensemble learning;
[0105] Figure 3 is the structure diagram of improved Stacking ensemble learning;
[0106] Figure 4 is the distribution of enhanced oil recovery technology projects in the dataset;
[0107] Figures 5(a) and 5(b) are the distribution of characteristic parameters;
[0108] Figure 6(a) to 6(f) is the confusion matrix of the conventional machine learning algorithm model on the test set;
[0109] Figure 7 is the performance of the difference between the accuracy and Kappa coefficient of the conventional machine learning algorithm model;
[0110] Figure 8 Confusion matrices of two Stacking integration models on the test set;
[0111] Figure 9 Learning curves of two Stacking integration models;
[0112] Figure 10 Performance of the difference in accuracy and Kappa coefficient of two Stacking integration models;
[0113] Figure 11 Confusion matrices of two improved Stacking integration models on the test set;
[0114] Figure 12 Learning curves of two improved Stacking integration models;
[0115] Figure 13 Performance of the difference in accuracy and Kappa coefficient of two improved Stacking integration models. Specific implementation manners
[0116] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings.
[0117] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners:
[0118] To verify the feasibility of the present invention, data was re - collected based on globally successful enhanced oil recovery technology projects, and the collected dataset was randomly divided into a training set, a validation set, and a test set. The performance of the model proposed by the present invention when facing unknown data was evaluated on the test set. The specific steps for EOR screening using the method of the present invention are as follows:
[0119] Step 1: Statistically analyze the data of globally successful enhanced oil recovery projects, select lithology, porosity, permeability, reservoir depth, crude oil specific gravity, crude oil temperature, crude oil viscosity, net thickness, and initial oil saturation as characteristic parameters, and pre - process and analyze the data. Finally, a complete dataset containing 956 pieces of data is formed for intelligent screening decision - making of enhanced oil recovery technology.
[0120] Figure 4 , and Figure 5 shows the quantity distribution and characteristic parameter distribution of enhanced oil recovery technologies in the dataset. The results show that there is a strong class imbalance problem in the collected dataset, which can be used to verify the effectiveness of the present invention in solving the class imbalance problem. And the characteristic parameters in the dataset cover the numerical ranges involved in most petroleum exploration projects, and can effectively reflect the general geological characteristics of global rocks in the dataset to a certain extent, ensuring the representativeness and effectiveness of the data.
[0121] To ensure that the dataset contains all necessary information and that the training set, validation set, and test set are independent of each other to avoid data leakage, the hold-out method is used to randomly divide the dataset into a training set and an external validation set in an 8:2 ratio. On this basis, the external validation set is randomly divided into a validation set and a test set in a 1:1 ratio. The training set is used for model construction, the validation set is used for model optimization, and the test set is used to verify the effectiveness of the model when facing new data. The repeatability of the results is ensured by setting a random seed. Table 1 shows the dataset division situation.
[0122] Table 1 Dataset division situation
[0123]
[0124]
[0125] Step 2: Learn the selected conventional machine learning algorithms on the training set and optimize them on the validation set. Use random search to quickly explore the breadth of the hyperparameter space, narrow the search space, and improve the search efficiency. Use grid search to traverse the hyperparameter space to ensure the search quality. Based on the prediction results of the model on the validation set, determine the best hyperparameter combination, and finally obtain six intelligent screening models for conventional enhanced oil recovery technologies. Finally, evaluate the generalization ability of the model on the test set.
[0126] Table 2 shows the accuracy, precision, recall, F1 score, Kappa coefficient, and MCC value performance of the six conventional machine learning algorithm models on the test set. As ensemble learning algorithms, random forest and extreme gradient boosting perform the best, and the highest prediction accuracy can reach 92.7%. As representatives of traditional single classifiers, neural network and decision tree perform relatively poorly. As simple linear models with simple structures, support vector machine and logistic regression are difficult to solve complex multi-class problems, and their performance is the worst under the influence of factors such as class imbalance.
[0127] Table 2 Performance of quantitative evaluation indicators of six conventional machine learning algorithm models
[0128]
[0129] Figure 6 is the confusion matrix of the six conventional machine learning algorithm models on the test set. As can be seen from the figure, the classification effect of enhanced oil recovery technologies with relatively small data volume is the main reason affecting the model performance.
[0130] Under normal circumstances, when a model faces a class imbalanced dataset, the accuracy, precision, recall, and F1 score will tend to be biased towards the prediction results of the majority class and ignore the prediction performance of the minority class, resulting in high values. However, the Kappa coefficient and MCC value can ignore the impact of class imbalance and thus provide a more realistic evaluation of the model. When the performance between the two is relatively consistent, it indicates that the model can effectively overcome the impact of class imbalance. Figure 7 The accuracy and Kappa coefficient are used as representatives to evaluate the model's ability to overcome the class imbalance problem based on their differences. Among the six conventional machine learning algorithm models, the random forest and extreme gradient boosting algorithms perform relatively well in overcoming the impact of class imbalance, but there is still much room for improvement.
[0131] Step 3: Based on the prediction performance of the conventional machine learning algorithm model in step 2, combined with the selection criteria of the Stacking ensemble learning base model and meta-model, the random forest, extreme gradient boosting, neural network and decision tree models are determined as the base models, and the support vector machine and logistic regression models are selected as meta-models respectively, and two Stacking-based enhanced oil recovery technology intelligent screening ensemble models (Stacking-SVM and Stacking-LR) are constructed. The four different models can ensure the diversity of the base model, and their high prediction accuracy and different prediction effects on different enhanced oil recovery technologies also ensure the accuracy and complementarity of the base model. Although the accuracy of the support vector machine and logistic regression models is relatively poor, they have good interpretability and strong anti-overfitting characteristics due to their relatively simple algorithm structure and support for the principle of structural risk minimization.
[0132] Table 3 shows the accuracy, precision, recall, F1 score, Kappa coefficient and MCC value of the two Stacking ensemble models on the test set. Although the two Stacking ensemble models combine the prediction results of different models, the model performance is lower than that of random forest and extreme gradient boosting, and the highest prediction accuracy can only reach 88.6%.
[0133] Table 3 Performance of quantitative evaluation indicators of two Stacking ensemble models
[0134] Integrated model Accuracy Precision Recall F1 score Kappa MCC Stacking-SVM 0.865 0.864 0.866 0.859 0.806 0.809 Stacking-LR 0.886 0.888 0.885 0.881 0.835 0.837
[0135] Figure 8 The confusion matrix of the two Stacking ensemble models on the test set. As can be seen from the figure, the prediction accuracy of the Stacking ensemble model for enhanced oil recovery technology with less data is much lower than the overall accuracy of the model, which directly affects the overall performance of the model.
[0136] To determine the reason for the performance degradation of the Stacking ensemble model, the learning curve was used to analyze the running process of the Stacking ensemble model. Whether the model is overfitting or underfitting can be judged through the learning curve. Usually, the training score of the model learning curve is slightly higher than the test score. However, when the training score is much higher than the test score, it indicates that the model has an overfitting problem. When both the training score and the test score are small, it indicates that the model has an underfitting problem. Figure 9 Figure 2 shows the learning curves of two Stacking ensemble models. There is a large difference between the training score and the test score when the two models finally converge, and there is a certain degree of overfitting problem, resulting in a decrease in the model performance.
[0137] Figure 10 Taking the accuracy rate and the Kappa coefficient as representatives respectively, the ability of the Stacking ensemble model to overcome the class imbalance problem was evaluated according to their differences. The results show that although Stacking ensemble learning can fuse the prediction results of different models, due to the influence of the overfitting problem, its ability to overcome the class imbalance problem is inferior to that of the random forest and extreme gradient boosting algorithms.
[0138] Step 4: Based on the prediction performance of the Stacking ensemble model in Step 3, the Stacking was improved by adding a new meta-model to construct an intelligent screening ensemble model for enhanced oil recovery technology based on the improved Stacking. The relatively better-performing Stacking-LR model was selected as the basis for improvement, and the support vector machine and the logistic regression model were used as the new meta-models respectively to construct two improved Stacking ensemble models (Mlti-Stacking-LR-SVM and Mlti-Stacking-LR-LR).
[0139] Table 4 shows the performance of the accuracy rate, precision, recall rate, F1 score, Kappa coefficient and MCC value of the two improved Stacking ensemble models on the test set. Compared with the conventional intelligent screening model for enhanced oil recovery technology and the two Stacking-based intelligent screening ensemble models for enhanced oil recovery technology, the performance of the two improved Stacking ensemble models has been significantly improved, and the highest prediction accuracy rate can reach 96.9%.
[0140] Table 4 Performance of the two improved Stacking ensemble models in quantitative evaluation indicators
[0141]
[0142] Figure 11The confusion matrices of the two improved Stacking ensemble models on the test set. As can be seen from the figure, the two improved Stacking ensemble models can make good predictions on the prediction effects of all enhanced oil recovery technologies. Among them, the prediction effect of the Mlti-Stacking-LR-LR model on enhanced oil recovery technologies with a small amount of data can reach 100%.
[0143] To further verify the superiority of the improved Stacking ensemble model and explore the differences between the Stacking ensemble learning before and after improvement, the learning curve is used to analyze the running process of the improved Stacking ensemble model. Figure 12 The learning curves of the two improved Stacking ensemble models. The gap between the training score and the test score of the learning curves of the Mlti-Stacking-SVM-LR and Mlti-Stacking-LR-LR models is significantly reduced at the final convergence, indicating that it can effectively alleviate the overfitting performance of the Stacking ensemble model.
[0144] Figure 13 Taking the accuracy rate and the Kappa coefficient as representatives respectively, the ability of the improved Stacking ensemble model to overcome the problem of class imbalance is evaluated according to their differences. Compared with other models, the difference between the accuracy rate and the Kappa coefficient of the improved Stacking ensemble model is significantly reduced. This result shows that the intelligent screening ensemble model of enhanced oil recovery technology based on the improved Stacking can make good predictions on all categories of enhanced oil recovery technologies regardless of the amount of data, can effectively overcome the influence of the problem of class imbalance, and thus improve the screening decision accuracy of enhanced oil recovery technology. Finally, a high-precision intelligent screening decision model of enhanced oil recovery technology based on the improved Stacking is provided to participate in the final screening decision of enhanced oil recovery technology. In the face of new reservoir data, its decision accuracy can reach up to 96.87%.
[0145] Aiming at the problem that the intelligent screening decision of enhanced oil recovery technology based on machine learning algorithms has a strong dependence on the data set, is vulnerable to the problem of class imbalance in the data set, and its decision accuracy needs to be improved, the present invention provides a high-precision intelligent screening decision framework of enhanced oil recovery technology based on the improved Stacking. By integrating the prediction results of multiple machine learning algorithm models and using the framework of cascading two-layer meta-models in series, it can effectively overcome the influence of the problem of class imbalance and avoid overfitting at the same time. While improving the screening efficiency of enhanced oil recovery technology and saving the screening cost, this decision-making method can greatly improve the accuracy rate of the screening decision of enhanced oil recovery technology.
Claims
1. A method for constructing an intelligent screening decision model for enhanced oil recovery technology based on improved Stacking, characterized in that: The steps include: Step 1: Collect data on successfully implemented enhanced oil recovery projects around the world, select appropriate reservoir rock and fluid characteristics as characteristic parameters, and form a complete data set through data processing and analysis; Step 2: Select a variety of conventional machine learning algorithms, use the characteristic parameters in the complete data set obtained in step 1 as input and the enhanced oil recovery technology in the complete data set as output, and build a variety of intelligent screening models for enhanced oil recovery technology based on conventional machine learning algorithms; Step 3: Based on the prediction performance of multiple intelligent screening models for enhanced oil recovery technologies based on conventional machine learning algorithms constructed in step 2, Stacking ensemble learning is used to fuse and construct an intelligent screening model for enhanced oil recovery technologies based on Stacking; Step 4: Based on the prediction performance of the intelligent screening model for enhanced oil recovery technology based on Stacking constructed in step 3, Stacking is improved by adding a new meta-model to construct an integrated model for intelligent screening of enhanced oil recovery technology based on improved Stacking.
2. The construction method according to claim 1, characterized in that: Step 1 includes the following sub-steps: Step 1.1: Collect statistics on the data of successfully implemented enhanced oil recovery projects around the world, and select reservoir rock and fluid properties that can effectively reflect the real reservoir and fluid conditions of the reservoir and are closely related to the oil recovery process of enhanced oil recovery technology as characteristic parameters; Step 1.2: Using the results of the semi-annual enhanced oil recovery project survey of the Oil & Gas Journal as the main data source, the data were supplemented by other documents with relevant data from around the world to form the original data set; Step 1.3, preprocessing the data in the original data set, including deleting duplicate data, replacing missing values with feature means, and performing One-Hot encoding on discrete parameters to form a complete data set, and logarithmizing the widely distributed continuous feature parameters in the data set; Step 1.4: Use the holdout method to randomly divide the complete data set obtained in step 1.3 into training set, validation set and test set in a ratio of 8:1:
1.
3. The construction method according to claim 1 or 2, characterized in that: The characteristic parameters include: lithology, porosity, permeability, reservoir depth, crude oil specific gravity, crude oil temperature, crude oil viscosity, net thickness and initial oil saturation.
4. The construction method according to claim 2, characterized in that: In step 2, the conventional machine learning algorithms include random forest algorithm, extreme gradient boosting algorithm, neural network algorithm, decision tree algorithm, support vector machine algorithm and logistic regression algorithm.
5. The construction method according to claim 4, characterized in that: Step 2 is as follows: Step 2.1, determine the evaluation indicators: select accuracy, precision, recall rate, and F1 score as conventional evaluation indicators to evaluate the overall performance of the model, and cite confusion matrix, Kappa coefficient and MCC value to evaluate the model's ability to overcome the class imbalance problem; Step 2.2, determine the hyperparameter optimization method: select a combination of random search and grid search to optimize the model; Step 2.3, using a variety of conventional machine learning algorithms to learn the complete data set obtained in step 1, optimizing it using the hyperparameter optimization method in step 2.1, and constructing a variety of intelligent screening models for enhanced oil recovery technologies based on conventional machine learning algorithms; The evaluation index determined in step 2.1 is used to evaluate the predictive performance of the intelligent screening model for enhanced oil recovery technology based on conventional machine learning algorithm constructed in step 2.
3.
6. The construction method according to claim 5, characterized in that: Step 2.3 is as follows: (1) Constructing an intelligent screening model for enhanced oil recovery technology based on random forest algorithm: Step 1: Through self-service sampling, randomly form multiple sub-datasets for independent training of a single decision tree; Step 2: Randomly select multiple features from the feature space to form a feature subset. When constructing each node of the decision tree, select the optimal feature from the feature subset for splitting; Step 3: Repeat the above steps to generate multiple decision trees to form a random forest; (2) Constructing an intelligent screening model for enhanced oil recovery technology based on extreme gradient boosting algorithm: Step 1: Initialize a constant model as the initial prediction value; Step 2: Iteratively construct multiple decision trees. In each iteration, determine a decision tree that can minimize the objective function. Step 3: After each decision tree is constructed, its prediction result is added to the previous prediction value to obtain a new prediction value; Step 4: Continue the iteration process until the preset number of iterations is reached or the change in the objective function is less than a certain threshold; (3) Constructing an intelligent screening model for enhanced oil recovery technology based on neural network algorithm: Step 1: Forward propagation, calculating the output of the network from the input layer to the output layer; Step 2: Calculate the loss and calculate the loss function based on the predicted results and the true labels; Step 3: Back propagation, calculate the gradient from the output layer to the input layer, and update the weights and biases; Step 4: Iterative training, repeat the above process until the network converges; (4) Constructing an intelligent screening model for enhanced oil recovery technology based on decision tree algorithm: Step 1: Starting from the root node, select the feature that can minimize the impurity of the data set as the split feature; Step 2: Divide the data set into multiple subsets based on the selected features; Step 3: Repeat the above process for the divided subsets and recursively construct decision tree branches until the stopping condition is met; (5) Constructing an intelligent screening model for enhanced oil recovery technology based on support vector machine algorithm: The calculation formula for obtaining the classification result by the support vector machine algorithm is: In formula (12), f(x) represents the classification result for the input data x, x represents the input data to be classified, i represents the sample number, and its value is 1, 2, …, n, where n is the number of samples, and α i is the Lagrange multiplier, y i is the label of the i-th sample, K(x,x i ) is the kernel function, x i represents the feature vector of the i-th sample, b is the bias term, when f(x)>0, sample x belongs to the positive category; when f(x)<0, sample x belongs to the negative category; (6) Constructing an intelligent screening model for enhanced oil recovery technology based on logistic regression algorithm: The calculation formula for obtaining the classification result by the logistic regression algorithm is: z=w T x1+n(13) In formula (13), z represents the output value given based on the input x1, and w T represents the transpose of the weight vector, x1 represents the eigenvector, which contains the eigenvalues of the input data, and b is the bias term; In formula (14), σ(z) is the class probability. When σ(z)≥0.5, the sample is classified as a positive sample; when σ(z)<0.5, the sample is classified as a negative sample.
7. The construction method according to claim 5, characterized in that: Step 3 is as follows: Step 3.1: Use the Stacking ensemble learning method to integrate the prediction results of multiple different base models by using a meta-model. Specifically: Step 3.1.1 Dataset division: Use the training set divided in step 1.4 as training data, and use K-fold cross-validation to divide it into K subsets, of which K-1 subsets are sub-training sets and the remaining subset is used as the sub-validation set; Step 3.1.2: Base model training: Use different base models to perform independent training on the sub-training sets; Step 3.1.3, generate new features: take the prediction results of all base models on the sub-validation set as new features, combine them into a new training data set, and use the prediction results of the base model on the test set divided in step 1.4 to form a new test data set; Step 3.1.4, training meta-model: input the prediction results of the base model into a meta-model as new features, and maximize the accuracy of the overall model by fitting the prediction results of the base model to obtain an intelligent screening model for enhanced oil recovery technology based on Stacking; Among them, the base model and meta-model are selected from a variety of intelligent screening models of enhanced oil recovery technology based on conventional machine learning algorithms constructed in step 2; Step 3.2: Evaluate the intelligent screening model of enhanced oil recovery technology based on Stacking through accuracy, precision, recall rate, F1 score, confusion matrix, Kappa coefficient and MCC value, and use the learning curve to analyze the operation process of the intelligent screening model of enhanced oil recovery technology based on Stacking.
8. The construction method according to claim 7, characterized in that: In step 3.1, The selection criteria of the base model are: according to the prediction performance of the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithms, different models with higher prediction performance, capable of providing correct feature information to the meta-model, and capable of capturing different aspects of the data when predicting the same target are selected as base models; The selection criteria of the meta-model are: according to the prediction performance of the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithms, a model with a simple structure and anti-overfitting characteristics is selected as the meta-model.
9. The construction method according to claim 7, characterized in that: Step 4 is as follows: Step 4.1: According to the prediction performance of the intelligent screening model of enhanced oil recovery technology based on Stacking in step 3.2, a suitable model is selected as a new meta-model from the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithm constructed in step 2. On the basis of the intelligent screening model of enhanced oil recovery technology based on Stacking constructed in step 3.1, an intelligent screening decision model of enhanced oil recovery technology based on improved Stacking is proposed, specifically: Step 4.1.1, data set division: take the training set divided in step 1.4 as training data, and use K-fold cross validation to divide it into K subsets, of which K-1 subsets are sub-training sets, and the remaining subset is used as the sub-validation set; Step 4.1.2: Base model training: Use different base models to perform independent training on the sub-training sets; Step 4.1.3, Generate new features: Take the prediction results of all base models on the sub-validation set as new features and combine them into a new training data set. Take the prediction results of the base model on the test set divided in step 1.4 to form a new test data set as the new features of the test set. Step 4.1.4: Train the meta-model: Input the prediction results of the base model as new features into the initial meta-model, and maximize the accuracy of the overall model by fitting the prediction results of the base model. Step 4.1.5, training new meta-model: using the prediction results of the initial meta-model on the new training data set to train the newly added meta-model, and obtain the intelligent screening decision model of enhanced oil recovery technology based on improved Stacking.
10. The construction method according to claim 9, characterized in that: In step 4.1, the base model is the base model selected in step 3, the initial metamodel is the metamodel selected in step 3, and the selection criteria for the new metamodel are: according to the predictive performance of the intelligent screening model of enhanced oil recovery technology based on conventional machine learning algorithms, a model with a simple structure and anti-overfitting characteristics is selected as the new metamodel.
Citation Information
Patent Citations
Method for predicting ground synthetic electric field based on improved Stacking algorithm
CN115544879A
Underground water harmful element intelligent early warning method based on integrated incremental learning
CN117010274A
Drilling and production cost prediction method based on feature selection and stacked heterogeneous integrated learning
CN117648646A
KASP primer intelligent typing evaluation method and system TAL-SRX based on multi-model fusion
CN119132403A
Cited By
Ocean pH data anomaly detection method fusing spatio-temporal feature stacking ensemble learning
CN120995319A
An anomaly detection method for ocean pH data by integrating spatiotemporal feature stacking ensemble learning
CN120995319B