Multi-element gas drive multi-component front edge prediction method and system based on optimized random forest algorithm

By constructing a multi-component leading edge prediction model of multi-component gas drive based on optimization random forest algorithm, the problem of inaccurate prediction of multi-component leading edge in the prior art is solved, and accurate prediction of multi-component leading edge changes of multi-component leading edges and optimization of gas injection schemes is achieved.

CN120197466APending Publication Date: 2025-06-24CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123270.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

There is a lack of effective method for predicting the leading edge of multi-gas driving in the prior art, especially under complex reservoir conditions, it is difficult to accurately predict the leading edge changes of multi-gas driving in multiple components.

Method used

Using the multi-component leading edge prediction method based on optimization random forest algorithm, the optimal prediction model is obtained for leading edge prediction by constructing a prediction model based on random forests and using different parameter combinations and corresponding leading edge simulation results.

Benefits of technology

Accurate prediction of the changes in the leading edge of multi-components of multi-gas drive is achieved, helping engineers design a reasonable gas injection plan to avoid gas waste and unnecessary intrusion, and improve oil and gas recovery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197466A_ABST
    Figure CN120197466A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-primitive gas-driven multi-component front edges, and discloses a multi-primitive gas-driven multi-component front edge prediction method and system based on an optimized random forest algorithm. According to the method, front edge simulation results under different parameter combinations are obtained based on a preset multi-element gas drive multi-component oil reservoir mechanism model, a prediction model based on the random forest is constructed, the prediction model is optimized by utilizing the different parameter combinations and the corresponding front edge simulation results, and an optimal prediction model is obtained. Finally, the optimal prediction model is used for conducting front edge prediction on the multi-element gas drive multiple components, the change of the front edges of the multi-element gas drive multiple components can be accurately predicted, engineers can be helped to reasonably design a gas injection scheme, and gas waste and unnecessary gas inflowing are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multi-component gas drive multi-component front technology, and particularly relates to a multi-component gas drive multi-component front prediction method and system based on an optimized random forest algorithm. Background Art

[0002] Gas drive front prediction is crucial in oilfield development, especially in aspects such as improving oil recovery, optimizing the production process, and reducing energy consumption. Accurate front prediction can effectively guide the operation strategy of gas drive injection, avoid improper gas breakthrough or waste, and thus maximize the oil and gas recovery efficiency. Whether it is production management or long-term oilfield development planning, front prediction is a very important link. Existing front prediction technologies include physical model prediction, empirical model prediction, etc.

[0003] Physical model prediction can simulate based on real reservoir conditions, with strong theoretical support and interpretability. However, it often requires a large amount of experimental data and physical parameters, has complex calculations, and is usually not accurate enough for gas drive front prediction, especially under complex reservoir conditions.

[0004] Empirical models are simple and easy to use, with small computational amounts, and can provide prediction results relatively quickly. However, due to relying on historical data, they are difficult to handle new and unseen reservoir situations and may fail under variable and complex reservoir conditions. In multi-component gas drive, the complexity of gas components increases significantly. Common multi-component gases (such as a mixture of CO2, N2, and CH4) not only have different densities, solubilities, and diffusion characteristics, but their behaviors in the reservoir are more complex. For multi-component gas drive, there is no effective front prediction method yet. Summary of the Invention

[0005] This application provides a multi-component gas drive multi-component front prediction method based on an optimized random forest algorithm to solve the problem that there is no effective front prediction method for multi-component gas drive in the prior art.

[0006] To solve the above technical problems, this application discloses a multi-component gas drive multi-component front prediction method based on an optimized random forest algorithm, and the method includes:

[0007] Obtaining front simulation results under different parameter combinations based on a preset multi-component gas drive multi-component reservoir mechanism model;

[0008] Constructing a prediction model based on a random forest;

[0009] Optimizing the prediction model using different parameter combinations and corresponding front simulation results to obtain an optimal prediction model;

[0010] Use the optimal prediction model to predict the front edge of multi-component gas flooding, and obtain the prediction results.

[0011] Preferably, optimize the prediction model by using different parameter combinations and the corresponding front edge simulation results to obtain the optimal prediction model, including:

[0012] Take different parameter combinations and the corresponding front edge simulation results as the input of the prediction model;

[0013] Optimize the prediction model by using a variety of preset different optimization methods to obtain multiple optimized models;

[0014] Verify each of the multiple optimized models to obtain the corresponding performance indicators;

[0015] Optimize the optimized model with the best performance indicators to determine the optimal prediction model.

[0016] Preferably, optimize the prediction model by using a variety of preset different optimization methods to obtain multiple optimized models, including:

[0017] Optimize the prediction model by using the method of hyperparameter optimization and feature selection to obtain the first optimized model;

[0018] Optimize the prediction model by using ensemble learning to obtain the second optimized model;

[0019] Optimize the prediction model by using the method of hyperparameter optimization and cross-validation to obtain the third optimized model;

[0020] Optimize the prediction model by using the method of learning curve and hyperparameter optimization to obtain the fourth optimized model.

[0021] Preferably, optimize the prediction model by using the method of hyperparameter optimization and feature selection to obtain the first optimized model, including:

[0022] Use grid search for hyperparameter tuning to determine the best hyperparameter combination;

[0023] Select the top 5 important features to train the prediction model corresponding to the best hyperparameter combination to obtain the first optimized model.

[0024] Preferably, optimize the prediction model by using ensemble learning to obtain the second optimized model, including:

[0025] Use voting regression to combine multiple regression models into a voting regression model;

[0026] Use the voting regression model to optimize the prediction model to obtain the second optimized model;

[0027] Among them, the multiple regression models include random forest, gradient boosting machine, and XGBoost.

[0028] Preferably, the method of using hyperparameter optimization and cross-validation is used to optimize the prediction model to obtain a third optimized model, including:

[0029] Use Bayesian optimization for hyperparameter search to determine the best hyperparameter combination;

[0030] Use 5-fold cross-validation to optimize the prediction model under the best hyperparameter combination to obtain a third optimized model.

[0031] Preferably, the method of using learning curves and hyperparameter optimization is used to optimize the prediction model to obtain a fourth optimized model, including:

[0032] Generate a training set and a validation set based on different parameter combinations and the corresponding frontal simulation results;

[0033] Under different training set sizes, determine the training scores and validation scores of the prediction model on the training set and the validation set, and draw learning curves;

[0034] Based on the training scores and validation scores, calculate and draw the errors of the training set and the validation set;

[0035] Perform hyperparameter optimization based on the errors to obtain a fourth optimized model.

[0036] Preferably, the parameter combinations include gas injection rate, bottom-hole flowing pressure, injection-production well spacing, well pattern form, formation dip angle, formation pressure, formation temperature, effective thickness, permeability, development of interbeds, permeability anisotropy, Lorenz coefficient, oil saturation, sedimentary rhythm, crude oil viscosity, multi-component gas composition, and multi-component gas composition.

[0037] Preferably, before constructing the prediction model based on random forest, the method further includes:

[0038] Preprocess different parameter combinations and the corresponding frontal simulation results;

[0039] The preprocessing method includes data cleaning, data standardization, and normalization.

[0040] This application also discloses a multi-component gas drive multi-component frontal prediction system based on an optimized random forest algorithm. The system includes:

[0041] A data acquisition module for obtaining frontal simulation results under different parameter combinations based on a preset multi-component gas drive multi-component reservoir mechanism model;

[0042] A model construction module for constructing a prediction model based on random forest;

[0043] A model optimization module, which is used to optimize a prediction model by using different parameter combinations and corresponding leading edge simulation results to obtain an optimal prediction model;

[0044] A prediction module, which is used to predict the leading edge of a multi-component gas drive by using the optimal prediction model to obtain a prediction result.

[0045] This application has at least the following beneficial effects:

[0046] 1. Based on a preset multi-component gas drive reservoir mechanism model, this application obtains leading edge simulation results under different parameter combinations, constructs a prediction model based on a random forest, and optimizes the prediction model by using different parameter combinations and corresponding leading edge simulation results to obtain an optimal prediction model. Finally, using the optimal prediction model to predict the leading edge of a multi-component gas drive can accurately predict the changes in the leading edge of a multi-component gas drive, which can help engineers reasonably design gas injection schemes and avoid gas waste and unnecessary gas breakthrough.

[0047] 2. This application uses an artificial intelligence method to automatically search for the optimal solution, greatly improving the efficiency of problem-solving; compared with traditional methods that require manual simulation, experiment, and analysis, the artificial intelligence method can obtain results more quickly.

[0048] 3. The multi-component gas drive leading edge prediction method based on an optimized random forest proposed in this application has strong global search ability, can effectively search for and update the optimal solution in the parameter space simultaneously, and improve the prediction performance of the model.

[0049] 4. The multi-component gas drive leading edge prediction method based on an ensemble algorithm proposed in this application combines the prediction results of multiple single models, improves the prediction accuracy and stability of the model, can handle more complex problems simultaneously, and provides better interpretability.

[0050] The additional aspects and advantages of this application will be given in the following description part, and these will become obvious from the following description or be understood through the practice of this application. Description of the Drawings

[0051] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0052] Figure 1 is a flowchart of the multi-component gas drive leading edge prediction method based on an optimized random forest algorithm provided by an embodiment of this application;

[0053] Figure 2 is a data table of simulation results under different parameter combinations provided by an embodiment of this application;

[0054] Figure 3Schematic diagram of the input page of the practical application system loaded with the prediction model provided by the embodiments of the present application;

[0055] Figure 4 Schematic diagram of the prediction result display of the practical application system loaded with the prediction model provided by the embodiments of the present application;

[0056] Figure 5 Comparison chart of the prediction result and the actual result of the component front obtained by optimizing the optimization model based on the hyperparameter and feature selection method provided by the embodiments of the present application;

[0057] Figure 6 Comparison chart of the prediction result and the actual result of the component front obtained by optimizing the optimization model based on the ensemble learning method provided by the embodiments of the present application;

[0058] Figure 7 Comparison chart of the prediction result and the actual result of the component front obtained by optimizing the optimization model based on the hyperparameter optimization and cross - validation method provided by the embodiments of the present application;

[0059] Figure 8 Comparison chart of the prediction result and the actual result of the component front obtained by optimizing the optimization model based on the learning curve and hyperparameter optimization method provided by the embodiments of the present application;

[0060] Figure 9 Comparison chart of the prediction result and the actual result of the rich C front obtained by optimizing the optimization model based on the hyperparameter and feature selection method provided by the embodiments of the present application;

[0061] Figure 10 Comparison chart of the prediction result and the actual result of the rich C front obtained by optimizing the optimization model based on the ensemble learning method provided by the embodiments of the present application;

[0062] Figure 11 Comparison chart of the prediction result and the actual result of the rich C front obtained by optimizing the optimization model based on the hyperparameter optimization and cross - validation method provided by the embodiments of the present application;

[0063] Figure 12 Comparison chart of the prediction result and the actual result of the rich C front obtained by optimizing the optimization model based on the learning curve and hyperparameter optimization method provided by the embodiments of the present application;

[0064] Figure 13 Comparison chart of the prediction result and the actual result of the effective component front obtained by optimizing the optimization model based on the hyperparameter and feature selection method provided by the embodiments of the present application;

[0065] Figure 14 Comparison chart of the prediction result and the actual result of the effective component front obtained by optimizing the optimization model based on the ensemble learning method provided by the embodiments of the present application;

[0066] Figure 15 This is a comparison graph of the prediction results and actual results of the effective component front obtained by optimizing the model based on the hyperparameter optimization and cross-validation method provided by the embodiments of this application;

[0067] Figure 16 This is a comparison graph of the prediction results and actual results of the effective component front obtained by optimizing the model based on the learning curve and hyperparameter optimization method provided by the embodiments of this application;

[0068] Figure 17 This is a comparison graph of the prediction results and actual results of the N2-rich front obtained by optimizing the model based on the hyperparameter and feature selection method provided by the embodiments of this application;

[0069] Figure 18 This is a comparison graph of the prediction results and actual results of the N2-rich front obtained by optimizing the model based on the ensemble learning method provided by the embodiments of this application;

[0070] Figure 19 This is a comparison graph of the prediction results and actual results of the N2-rich front obtained by optimizing the model based on the hyperparameter optimization and cross-validation method provided by the embodiments of this application;

[0071] Figure 20 This is a comparison graph of the prediction results and actual results of the N2-rich front obtained by optimizing the model based on the learning curve and hyperparameter optimization method provided by the embodiments of this application;

[0072] Figure 21 This is a comparison graph of the prediction results and actual results of the CO2-rich front obtained by optimizing the model based on the hyperparameter and feature selection method provided by the embodiments of this application;

[0073] Figure 22 This is a comparison graph of the prediction results and actual results of the CO2-rich front obtained by optimizing the model based on the ensemble learning method provided by the embodiments of this application;

[0074] Figure 23 This is a comparison graph of the prediction results and actual results of the CO2-rich front obtained by optimizing the model based on the hyperparameter optimization and cross-validation method provided by the embodiments of this application;

[0075] Figure 24 This is a comparison graph of the prediction results and actual results of the CO2-rich front obtained by optimizing the model based on the learning curve and hyperparameter optimization method provided by the embodiments of this application;

[0076] Figure 25 This is a comparison graph of the prediction results and actual results of the gas slug front obtained by optimizing the model based on different optimizations provided by the embodiments of this application;

[0077] Figure 26 Comparison diagram of the prediction result and the actual result of the front edge of the gas slug obtained by optimizing the optimization model based on the ensemble learning method provided by the embodiment of the present application;

[0078] Figure 27 Comparison diagram of the prediction result and the actual result of the front edge of the gas slug obtained by optimizing the optimization model based on the hyperparameter optimization and cross - validation method provided by the embodiment of the present application;

[0079] Figure 28 Comparison diagram of the prediction result and the actual result of the front edge of the gas slug obtained by optimizing the optimization model based on the learning curve and hyperparameter optimization method provided by the embodiment of the present application;

[0080] Figure 29 Structure diagram of the multi - component gas drive multi - front edge prediction system based on the optimized random forest algorithm provided by the embodiment of the present application. Detailed implementation manners

[0081] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0082] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0083] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0084] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions in this regard. For the technical problems existing in the prior art, the present application provides a multi-component leading edge prediction method and system for multi-component gas drive based on an optimized random forest algorithm, aiming to solve at least one of the technical problems in the prior art.

[0085] The following uses specific embodiments to detail the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0086] For the leading edge prediction of multi-component gas drive, new methods need to be explored, especially in combination with modern artificial intelligence algorithms, such as machine learning and deep learning. Machine learning and artificial intelligence algorithms have been widely used in many fields of oilfield development in recent years, especially in key links such as reservoir modeling, production optimization, oil and gas exploration, and leading edge prediction. With the progress of data acquisition technology, a large amount of real-time production data, geological data, and physical parameters have been accumulated during the oilfield development process, providing a rich foundation for the application of algorithms. Through these advanced algorithms, oilfield development can more accurately achieve decision support, improve development efficiency, and optimize resource allocation. Among them, the random forest algorithm, as a powerful ensemble learning method, is particularly prominent in the application of oilfield development, especially when dealing with high-dimensional data, complex non-linear relationships, and noisy data, showing its unique advantages.

[0087] Therefore, the embodiments of the present application provide a possible implementation method, such as Figure 1 shown, provides a flowchart of a multi-component leading edge prediction method for multi-component gas drive based on an optimized random forest algorithm. This solution can be executed by any electronic device. Optionally, it can be executed on the server side or the terminal device.

[0088] As Figure 1 shown, the method may include the following steps:

[0089] Step 101, obtaining leading edge simulation results under different parameter combinations based on a preset multi-component gas drive reservoir mechanism model.

[0090] In the embodiments of the present application, a multi-component reservoir mechanism model for multi-gas flooding is established based on the geological parameters of the research block, and numerical simulation of multi-gas flooding is carried out according to this model, and the multi-gas flooding multi-component front is divided, including component front, rich C front, effective component front, rich N2 front, rich CO2 front and gas slug front. Based on the multi-component reservoir mechanism model of multi-gas flooding, numerical simulation of the multi-gas flooding reservoir is carried out to obtain the numerical simulation results, that is, the front simulation results.

[0091] Among them, the collected geological parameters include data on front prediction and related influencing factors.

[0092] Step 102, construct a prediction model based on random forest.

[0093] The random forest algorithm shows its unique advantages when dealing with high-dimensional data, complex non-linear relationships and noisy data. Therefore, in the embodiments of the present application, a prediction model based on random forest is constructed to realize the prediction of multi-gas flooding multi-components.

[0094] Step 103, optimize the prediction model by using different parameter combinations and the corresponding front simulation results to obtain the optimal prediction model.

[0095] Step 104, use the optimal prediction model to predict the front of the multi-gas flooding multi-components to obtain the prediction results.

[0096] In the embodiments of the present application, reasonable planning and decision-making can be carried out according to the prediction results.

[0097] In the embodiments of the present application, based on the preset multi-component reservoir mechanism model of multi-gas flooding, the front simulation results under different parameter combinations are obtained, a prediction model based on random forest is constructed, and the prediction model is optimized by using different parameter combinations and the corresponding front simulation results to obtain the optimal prediction model. Finally, the optimal prediction model is used to predict the front of the multi-gas flooding multi-components, which can accurately predict the changes of the front of the multi-gas flooding multi-components, and can help engineers reasonably design the gas injection scheme to avoid gas waste and unnecessary gas breakthrough.

[0098] In an alternative embodiment, the parameter combinations include gas injection rate, bottom-hole flowing pressure, injection-production well spacing, well pattern, formation dip angle, formation pressure, formation temperature, effective thickness, permeability, development of interbeds, permeability anisotropy, Lorenz coefficient, oil saturation, sedimentary rhythm, crude oil viscosity, multi-gas components and multi-gas composition.

[0099] The table of the front simulation results under different parameter combinations is as Figure 2 shown, where the front simulation result is the multi-gas injection volume, that is, the gas breakthrough time.

[0100] In an optional embodiment, before constructing the prediction model based on the random forest, the method further includes:

[0101] Preprocess different parameter combinations and the corresponding leading edge simulation results;

[0102] The preprocessing method includes data cleaning, data standardization, and normalization.

[0103] Preprocess the data, check and filter invalid samples.

[0104] Optionally, data cleaning specifically includes: checking whether there are missing values in the data set, and if there are missing values, filling them with the mean value;

[0105] Data standardization specifically includes: using a standardization scaler to convert each column of features to a mean of 0 and a standard deviation of 1 to avoid the adverse effects of feature scale differences on the model;

[0106] Normalization specifically includes: using a min-max scaler to normalize the target variable y, converting the target variable to the range [0,1], and removing the dimensionality effect;

[0107] In an optional embodiment, using different parameter combinations and the corresponding leading edge simulation results to optimize the prediction model to obtain an optimal prediction model includes:

[0108] Taking different parameter combinations and the corresponding leading edge simulation results as the input of the prediction model;

[0109] Using a variety of preset different optimization methods to optimize the prediction model to obtain multiple optimized models;

[0110] Validating each of the multiple optimized models respectively to obtain the corresponding performance metrics;

[0111] Optimizing the optimized model with the best performance metrics to determine the optimal prediction model.

[0112] Optionally, embodiments of the present application make a sample set according to different parameter combinations and the corresponding leading edge simulation results, divide the sample set data, with 80% for training and 20% for testing; respectively try different optimized models for training, and compare the performance metrics of each optimized model on the validation set. The performance metrics include the coefficient of determination and the root mean square error.

[0113] In embodiments of the present application, a line chart between the true values and the predicted values of each optimized model can be drawn using visualization to intuitively display the difference between the true values and the predicted values.

[0114] Evaluate the accuracy and prediction ability of each optimized model through performance metrics. For the optimized models with excellent performance, further optimize them to obtain the optimal prediction model.

[0115] In an optional embodiment, a plurality of different preset optimization methods are used to optimize the prediction model, and a plurality of optimized models are obtained, including:

[0116] The prediction model is optimized using the method of hyperparameter optimization and feature selection to obtain the first optimized model;

[0117] The prediction model is optimized using ensemble learning to obtain the second optimized model;

[0118] The prediction model is optimized using the method of hyperparameter optimization and cross-validation to obtain the third optimized model;

[0119] The prediction model is optimized using the method of learning curve and hyperparameter optimization to obtain the fourth optimized model.

[0120] In an optional embodiment, the prediction model is optimized using the method of hyperparameter optimization and feature selection to obtain the first optimized model, including:

[0121] Grid search is used for hyperparameter tuning to determine the best hyperparameter combination;

[0122] The top 5 important features are selected to train the prediction model corresponding to the best hyperparameter combination to obtain the first optimized model.

[0123] In the embodiment of the present application, first, grid search is used for hyperparameter tuning. Different hyperparameter options are defined through a hyperparameter grid. Grid search traverses all combinations and selects the best hyperparameter combination. 5-fold cross-validation is used, and all CPU cores are used to accelerate the calculation. Selecting the minimization of the mean squared error (MSE) as the scoring criterion, and then feature selection is performed. Here, the top 5 important features are selected according to the feature importance of the model. The selected features are used to train the model to reduce the impact of unimportant features on the model performance.

[0124] In an optional embodiment, the prediction model is optimized using ensemble learning to obtain the second optimized model, including:

[0125] Voting regression is used to combine multiple regression models into a voting regression model;

[0126] The prediction model is optimized using the voting regression model to obtain the second optimized model;

[0127] Among them, the multiple regression models include random forest, gradient boosting machine, and XGBoost.

[0128] Combining multiple regression models can improve the prediction effect of the model.

[0129] In an alternative embodiment, the prediction model is optimized using the method of hyperparameter optimization and cross-validation to obtain a third optimized model, including:

[0130] Use Bayesian optimization for hyperparameter search to determine the optimal hyperparameter combination;

[0131] Use 5-fold cross-validation to optimize the prediction model under the optimal hyperparameter combination to obtain a third optimized model.

[0132] In the embodiment of the present application, first use Bayesian optimization for hyperparameter search to automatically select the optimal hyperparameter combination, and then use 5-fold cross-validation to ensure that each training will be performed on different training sets and test sets, thereby improving the stability of the model and reducing overfitting.

[0133] In an alternative embodiment, the prediction model is optimized using the method of learning curve and hyperparameter optimization to obtain a fourth optimized model, including:

[0134] Generate training sets and validation sets based on different parameter combinations and corresponding leading-edge simulation results;

[0135] Under different training set sizes, determine the training scores and validation scores of the prediction model on the training set and validation set, and plot the learning curve;

[0136] Based on the training scores and validation scores, calculate and plot the errors of the training set and validation set;

[0137] Perform hyperparameter optimization based on the errors to obtain a fourth optimized model.

[0138] In the embodiment of the present application, hyperparameter optimization can be performed according to the errors of the training set and cross-validation set, and the optimal hyperparameter combination can be printed to obtain the final fourth optimized model.

[0139] In an alternative embodiment, the optimal prediction model is used to predict the leading edge of multi-component gas flooding, and the implementation process is as Figure 3 shown:

[0140] Input data: injection rate, bottom-hole flowing pressure, injection-production well spacing, well pattern, formation dip angle, formation pressure, formation temperature, effective thickness, permeability, development of interlayers, permeability anisotropy, Lorenz coefficient, oil saturation, sedimentary rhythm, crude oil viscosity, multi-component gas components, and multi-component gas composition. In the embodiment of the present application, the input data is adopted by a preset leading-edge warning system, and the previously trained optimal prediction model is loaded using a function. The loaded optimal prediction model is used to predict the input data, the predicted results are output, and the predicted data is compared with the original data (i.e., the leading-edge simulation results obtained through the multi-component gas flooding reservoir mechanism model). The results are as Figure 4As shown

[0141] It can be seen that in the embodiments of the present application, the multi - gas injection amounts of the predicted component front, rich - C front, effective - component front, rich - N2 front, rich - CO2 front, and gas slug front are close to the original data.

[0142] In the embodiments of the present application, based on different parameters and the prediction simulation results of the multi - component front of multi - gas flooding, as the input data of the prediction model, and the data is normalized. By optimizing the front - edge prediction models based on different artificial intelligence algorithms, finally, a multi - gas - flooding multi - component front - edge prediction method based on the optimized random forest is selected. At the same time, an incremental learning method is adopted. When the model prediction deviation is detected, new data is used to perform incremental adjustment of the model, enabling the model to dynamically adapt to the changes in the data distribution. The method proposed by the present invention can make reasonable planning and decisions according to the prediction results, providing a basis and support for the effective control of gas channeling.

[0143] In summary, the embodiments of the present application adopt artificial intelligence methods to automatically search for the optimal solution, greatly improving the efficiency of problem - solving; compared with the traditional methods that require manual simulation, experiment, and analysis, the artificial intelligence methods can obtain results more quickly; the multi - gas - flooding multi - component front - edge prediction method based on the optimized random forest algorithm proposed in the embodiments of the present application has strong global search ability, can effectively search for and update the optimal solution in the parameter space simultaneously, and improve the prediction performance of the model; the intelligent prediction method for the gas - channeling time of the multi - gas - flooding multi - component front edge based on the ensemble algorithm proposed in the embodiments of the present application combines the prediction results of multiple single models, improves the prediction accuracy and stability of the model, can handle more complex problems simultaneously, and provides better interpretability; the incremental learning method proposed in the embodiments of the present application can use new data to perform incremental adjustment of the model when detecting the model prediction deviation, thereby improving the accuracy of the model.

[0144] Referring to Table 1, in order to verify the performance of the method in the embodiments of the present application, experimental comparison data between the front - edge prediction method in the embodiments of the present application and the prior art are provided based on the root - mean - square error (RMSE) and the coefficient of determination (R 2 )

[0145] Table 1 Performance evaluation of the multi - component front - edge prediction model based on the random forest optimization algorithm

[0146]

[0147]

[0148] The evaluation results of each optimized model obtained by further optimizing the random forest model in the embodiments of the present application are as Figures 5 - 28 shown. Among them, Figures 5 - 8The figure shows the comparison between the predicted results and the actual results of the component front obtained by optimized models based on different optimizations (hyperparameter and feature selection-based, ensemble learning-based, hyperparameter optimization and cross-validation-based, learning curve and hyperparameter optimization-based). Figures 9 - 12 The figure shows the comparison between the predicted results and the actual results of the rich C front obtained by optimized models based on different optimizations. Figures 13 - 16 The figure shows the comparison between the predicted results and the actual results of the effective component front obtained by optimized models based on different optimizations. Figures 17 - 20 The figure shows the comparison between the predicted results and the actual results of the rich N2 front obtained by optimized models based on different optimizations. Figures 21 - 24 The figure shows the comparison between the predicted results and the actual results of the rich CO2 front obtained by optimized models based on different optimizations. Figures 25 - 28 The figure shows the comparison between the predicted results and the actual results of the gas slug front obtained by optimized models based on different optimizations.

[0149] It can be seen that in the comparison charts of the front prediction results and the actual results of each optimized model, the error between the predicted values (predicted data) and the actual values (original data) obtained by the ensemble learning algorithm is the smallest, indicating that this method has higher accuracy and stronger generalization ability in dealing with the prediction of the evolution of the multi-component gas drive front. Especially by combining the advantages of multiple models such as random forest, gradient boosting machine, and XGBoost, ensemble learning can effectively reduce the overfitting or underfitting problems that may occur in a single model, thus providing more stable and reliable prediction results in the complex non-linear gas drive front evolution process. Therefore, the application of the ensemble learning framework in optimizing the multi-component gas drive process shows good potential and can provide more accurate data support for actual oilfield development.

[0150] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide a multi-component gas drive multi-component front prediction system based on an optimized random forest algorithm, as Figure 29 shown, the system includes:

[0151] A data acquisition module 2901, configured to obtain front simulation results under different parameter combinations based on a preset multi-component gas drive multi-component reservoir mechanism model;

[0152] A model construction module 2902, configured to construct a prediction model based on random forest;

[0153] A model optimization module 2903, configured to optimize the prediction model by using different parameter combinations and the corresponding front simulation results to obtain an optimal prediction model;

[0154] A prediction module 2904, configured to perform front prediction on the multi-component gas drive multi-component by using the optimal prediction model to obtain a prediction result.

[0155] In the embodiments of the present application, based on a preset multi-component reservoir mechanism model for multi-gas flooding, the frontal simulation results under different parameter combinations are obtained, a prediction model based on a random forest is constructed, and the prediction model is optimized using different parameter combinations and the corresponding frontal simulation results to obtain an optimal prediction model. Finally, the optimal prediction model is used to predict the front of the multi-gas flooding multi-component, which can accurately predict the changes in the front of the multi-gas flooding multi-component, and can help engineers reasonably design the gas injection plan to avoid gas waste and unnecessary gas breakthrough.

[0156] The multi-gas flooding multi-component front prediction system based on the optimized random forest algorithm provided in the embodiments of the present application can implement Figures 1 - 28 each process implemented in the method embodiments. To avoid repetition, it will not be described in detail here.

[0157] The multi-gas flooding multi-component front prediction system based on the optimized random forest algorithm in the embodiments of the present application can execute the multi-gas flooding multi-component front prediction method based on the optimized random forest algorithm provided in the embodiments of the present application. Their implementation principles are similar. The actions performed by each module and unit in the multi-gas flooding multi-component front prediction system based on the optimized random forest algorithm in the embodiments of the present application correspond to the steps in the multi-gas flooding multi-component front prediction method based on the optimized random forest algorithm in the embodiments of the present application. For the detailed functional description of each module of the multi-gas flooding multi-component front prediction system based on the optimized random forest algorithm, reference can specifically be made to the description in the corresponding multi-gas flooding multi-component front prediction method based on the optimized random forest algorithm shown above, and it will not be described in detail here.

[0158] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but also covers other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.

Claims

1. A multi-component gas flooding front prediction method based on an optimized random forest algorithm, characterized in that: The method comprises: Based on the preset multi-component gas flooding reservoir mechanism model, the front simulation results under different parameter combinations are obtained; Build a prediction model based on random forest; Optimizing the prediction model using different parameter combinations and corresponding leading edge simulation results to obtain an optimal prediction model; The optimal prediction model is used to perform front prediction on multi-component gas flooding to obtain prediction results.

2. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 1, characterized in that: The method of optimizing the prediction model by using different parameter combinations and corresponding leading edge simulation results to obtain an optimal prediction model includes: Using different parameter combinations and corresponding leading edge simulation results as inputs to the prediction model; The prediction model is optimized using a plurality of preset optimization methods to obtain a plurality of optimization models; Verifying the plurality of optimization models respectively to obtain corresponding performance indicators; The optimization model with the best performance index is optimized to determine the optimal prediction model.

3. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 2, characterized in that: The prediction model is optimized by using a plurality of preset optimization methods to obtain a plurality of optimization models, including: The prediction model is optimized using a hyperparameter optimization and feature selection method to obtain a first optimized model; Optimizing the prediction model using ensemble learning to obtain a second optimized model; The prediction model is optimized using a hyperparameter optimization and cross-validation method to obtain a third optimized model; The prediction model is optimized using a learning curve and hyperparameter optimization method to obtain a fourth optimized model.

4. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 3 is characterized in that: The method of using hyperparameter optimization and feature selection to optimize the prediction model to obtain a first optimization model includes: Use grid search to perform hyperparameter tuning and determine the best hyperparameter combination; The first five important features are selected to train the prediction model corresponding to the optimal hyperparameter combination to obtain the first optimization model.

5. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 3, characterized in that: The method of optimizing the prediction model by using ensemble learning to obtain a second optimization model includes: Use voting regression to combine multiple regression models into a voting regression model; Optimizing the prediction model using the voting regression model to obtain the second optimization model; Among them, multiple regression models include random forest, gradient boosting machine and XGBoost.

6. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 3, characterized in that: The method of using hyperparameter optimization and cross-validation to optimize the prediction model to obtain a third optimization model includes: Use Bayesian optimization to search for hyperparameters and determine the best hyperparameter combination; The prediction model under the best hyperparameter combination is optimized using 5-fold cross validation to obtain the third optimized model.

7. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 3, characterized in that: The method of using a learning curve and hyperparameter optimization to optimize the prediction model to obtain a fourth optimization model includes: Generate training sets and validation sets based on different parameter combinations and corresponding front edge simulation results; Under different training set sizes, determining the training score and the validation score of the prediction model on the training set and the validation set, and drawing a learning curve; Based on the training score and the validation score, calculate and plot the errors of the training set and the validation set; Hyperparameter optimization is performed based on the error to obtain the fourth optimization model.

8. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 1, characterized in that: The parameter combination includes gas injection rate, bottom hole flow pressure, injection-production well spacing, well pattern, formation inclination, formation pressure, formation temperature, effective thickness, permeability, interlayer development, permeability anisotropy, Lorentz coefficient, oil saturation, sedimentary rhythm, crude oil viscosity, multi-gas components and multi-gas composition.

9. The multi-component gas flooding front prediction method based on the optimized random forest algorithm according to claim 1, characterized in that: Before constructing the prediction model based on random forest, the method further includes: Preprocess different parameter combinations and corresponding leading edge simulation results; The preprocessing method includes data cleaning, data standardization and normalization.

10. A multi-component gas flooding front prediction system based on an optimized random forest algorithm, characterized in that: The system comprises: The data acquisition module is used to obtain the front simulation results under different parameter combinations based on the preset multi-component gas drive multi-component reservoir mechanism model; Model building module, used to build a random forest-based prediction model; A model optimization module, used to optimize the prediction model using different parameter combinations and corresponding leading edge simulation results to obtain an optimal prediction model; The prediction module is used to use the optimal prediction model to perform front prediction on multi-component gas flooding to obtain prediction results.