Multi-comprehensive intelligent polymer flooding reservoir yield prediction method and system based on polymer injection data and storage medium

By constructing a multi-integrated intelligent oil flood reservoir yield prediction method, using the infusion data to screen characteristics and optimize hyperparameters, efficient and accurate prediction of oil field output is achieved, and the problems of low computing efficiency and high artificial dependence in the existing technology are solved, and the real-time and accuracy of oil field development are improved.

CN120297481APending Publication Date: 2025-07-11NORTHEAST GASOLINEEUM UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510371084.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing oil field yield prediction methods rely on static geological data and ignore dynamic infusion data. They have low computational efficiency and high artificial dependence, making it difficult to meet the accuracy and real-time requirements for output prediction during oil field development.

Method used

The multi-integrated intelligent coal-driving reservoir yield prediction method based on injecting data is adopted. By constructing the XGBoost module, linear regression module, support vector machine module and random forest module, combined with the online learning framework and concept drift detection algorithm, the model is updated in real time, key features are screened and hyperparameters are optimized, and multi-model fusion is achieved.

Benefits of technology

It significantly improves the accuracy of reservoir production forecasts and the timeliness of production decisions, reduces calculation costs, and provides efficient oil field development support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297481A_ABST
    Figure CN120297481A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-comprehensive intelligent polymer flooding reservoir yield prediction method and system based on polymer injection data and a storage medium, relates to the field of oilfield development, and aims to solve the problems of dependence on static data, neglect of dynamic association, low calculation efficiency, high manual dependency and poor interpretability in the prior art. Comprising the steps of 1, collecting oil well data from an oil reservoir site, and extracting key features for predicting the oil reservoir yield to obtain feature optimization data; 2, an oil reservoir yield prediction model is constructed, the model comprises an XGBoost module, a linear regression module, a support vector machine module and a random forest module, and hyper-parameters of the XGBoost module are optimized; training the modules based on the training set, testing the performance of the modules by adopting the test set, and weighting the modules based on the fitting effect of the modules on the training set; 3, fusing the prediction results of all the modules to obtain an oil reservoir yield prediction result; and 4, receiving new polymer injection data in real time through an online learning framework, and triggering incremental updating of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oilfield development, and in particular, to a multi-comprehensive intelligent polymer flooding reservoir production prediction method, system and storage medium based on polymer injection data. Background Art

[0002] A polymer flooding reservoir refers to a reservoir that enhances oil recovery by injecting polymers. With the continuous deepening of oilfield development, traditional oilfield production prediction methods, such as numerical simulation methods and production decline methods, have gradually become difficult to meet the current oilfield production prediction requirements. These methods mainly rely on the static characteristics of reservoir geology for modeling and lack the effective utilization of dynamic polymer injection data, resulting in limited prediction accuracy. Especially in the polymer flooding oilfields in the Daqing Placanticline, the limitations of these traditional methods are more obvious.

[0003] In the prior art, oilfield production prediction mainly relies on static geological data while ignoring the importance of real-time data. These methods are often too complex, with high computational costs and insufficiently real-time parameter updates, making it difficult to capture the non-linear relationship of production changes. For example, although the numerical simulation method can provide relatively accurate reservoir simulation results, its calculation process is cumbersome, requiring a large amount of geological data and long calculation time, which is not conducive to quickly responding to real-time changes in the oilfield site. Similarly, due to its overly simplified mathematical model, the production decline method cannot accurately describe the complexity of reservoir production changes.

[0004] Therefore, the prior art lacks a more accurate and efficient multi-comprehensive intelligent polymer flooding reservoir production prediction method to meet the urgent requirements for production prediction accuracy and real-time performance in the process of oilfield development. Summary of the Invention

[0005] The technical problem to be solved by the present invention is:

[0006] The existing technologies have problems such as relying on static data, ignoring dynamic correlations, low computational efficiency, high manual dependency, and poor interpretability.

[0007] The technical solution adopted by the present invention to solve the above technical problems:

[0008] The present invention provides a multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data, including the following steps:

[0009] Step 1: Collect well data from the reservoir site, preprocess the well data, extract key features for reservoir production prediction from the preprocessed data to obtain feature-optimized data, and randomly divide the feature-optimized data into a training set and a test set;

[0010] Step 2: Construct a reservoir production prediction model. The reservoir production prediction model includes: an XGBoost module, a linear regression module, a support vector machine module, and a random forest module, and optimize the hyperparameters of the XGBoost module; train each module based on the training set, test the module performance using the test set, and weight each module based on the fitting effect on the training set.

[0011] Step 3: Fuse the prediction results of each module of the reservoir production prediction model to obtain the reservoir production prediction result.

[0012] Step 4: Real-time receive new polymer injection data through an online learning framework, and trigger model incremental updates in combination with the concept drift detection algorithm.

[0013] Further, in Step 1, the preprocessing of the well data is specifically as follows: fill in the missing values of the well data, delete or replace the outliers in the data, and normalize the data.

[0014] Further, in Step 1, identify the missing values in the data through statistical mean, median, or time series; identify the outliers in the data through box plot analysis or standard deviation analysis.

[0015] Further, in Step 1, use the Boruta algorithm to extract the key features for reservoir production prediction from the preprocessed data, and update the feature set in real time through a sliding window to obtain feature-optimized data.

[0016] Further, the optimization of the hyperparameters of the XGBoost module in Step 2 is specifically as follows: use the Optuna hyperparameter optimization algorithm to tune the hyperparameters of the XGBoost module, use the TPE algorithm for global hyperparameter iterative search, perform grid search within the optimal parameter region, and set an early stopping mechanism to prevent overfitting, and finally determine the optimal hyperparameter combination.

[0017] Further, the objective function of the XGBoost module in Step 2 is:

[0018]

[0019] The loss function in the objective function of the XGBoost module is represented by the true value y and the predicted value , n represents the number of samples, and Ω(f i ) is the regularization term;

[0020] The expression of the predicted value is:

[0021]

[0022] where denotes the sample prediction result after the t-th iteration, denotes the prediction result of the first t - 1 trees, f t (x i ) represents the t-th tree model.

[0023] Furthermore, in step two, weighting each module based on the fitting effect of each module on the training set is specifically to determine the weight through the R 2 value of each module on the training set:

[0024]

[0025] where

[0026]

[0027] where, y true,i denotes the actual production, y pred,i denotes the predicted production, denotes the average value of the actual production.

[0028] Furthermore, in step three, fusing the prediction results of each module of the reservoir production prediction model is specifically:

[0029]

[0030] where, Y is the predicted value of the multi - comprehensive intelligent model, W i is the weight of the i-th model, pred i is the predicted value of the i-th model.

[0031] The present invention also provides a multi - comprehensive intelligent polymer - flooding reservoir production prediction system based on polymer injection data. This system has program modules corresponding to the steps of the method described in any one of the above technical solutions, and when running, executes the steps in the above - mentioned multi - comprehensive intelligent polymer - flooding reservoir production prediction method based on polymer injection data.

[0032] The present invention also provides a computer - readable storage medium. The computer - readable storage medium stores a computer program, and the computer program is configured to implement the steps in the above - mentioned multi - comprehensive intelligent polymer - flooding reservoir production prediction method based on polymer injection data when called by a processor.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] The present invention relates to a multi-comprehensive intelligent polymer flooding reservoir production prediction method, system, and storage medium based on polymer injection data. Based on real-time polymer injection data from the oilfield site, features strongly correlated with production are screened out, and redundant and noisy features are removed. The model uses the XGBoost module, linear regression module, support vector machine module, and random forest module to establish independent prediction models. Global search and local grid optimization are performed on the XGBoost hyperparameters, and weights are dynamically allocated based on the fitting effects of each model to achieve accurate prediction of the polymer flooding reservoir production, effectively improving the timeliness and accuracy of production decisions and promoting the efficient progress of oilfield development.

[0035] The present invention significantly improves the model prediction accuracy and greatly reduces the manual calculation cost through dynamic feature screening, hyperparameter optimization, multi-model fusion, and distributed computing, providing efficient technical support for the production management and decision-making of polymer flooding oilfields. Brief Description of the Drawings

[0036] Figure 1 It is a flowchart of the multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data in the embodiment of the present invention;

[0037] Figure 2 It is a schematic diagram of the Optuna-XGBoost fusion model process in the embodiment of the present invention;

[0038] Figure 3 It is a schematic diagram of the model structure and function in the embodiment of the present invention. Detailed Embodiments

[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the drawings. Obviously, the described embodiments or examples are only a part of the embodiments or examples of the present invention, rather than all of them. All other embodiments or examples obtained by those of ordinary skill in the art without creative efforts based on the embodiments or examples of the present invention shall fall within the scope of protection of the present invention.

[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the drawings.

[0041] Specific Embodiment 1: Combining Figure 1 As shown, the present invention provides a multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data, which is characterized by including the following steps:

[0042] Step 1: Collect well data from the reservoir site, preprocess the well data, extract the key features for reservoir production prediction from the preprocessed data to obtain feature-optimized data, and randomly divide the feature-optimized data into a training set and a test set;

[0043] Step 2: Construct a reservoir production prediction model. The reservoir production prediction model includes: an XGBoost module, a linear regression module, a support vector machine module, and a random forest module, and optimize the hyperparameters of the XGBoost module; train each module based on the training set, test the performance of the module using the test set, and weight each module based on the fitting effect of each module on the training set;

[0044] Step 3: Integrate the prediction results of each module of the reservoir production prediction model to obtain the reservoir production prediction result;

[0045] Step 4: Receive new polymer injection data in real time through an online learning framework, and trigger model incremental updates in combination with a concept drift detection algorithm.

[0046] The reservoir production prediction model of this implementation plan is based on a machine learning algorithm with XGBoost as the main body, and combines linear regression, support vector machine, and random forest algorithms. The support vector machine (SVM) classifies or regresses by finding an optimal hyperplane. As an ensemble learning method, the random forest constructs multiple decision trees and integrates their prediction results to improve the accuracy and robustness of the overall model. Train and predict each module based on the screened data, and optimize the parameter settings through cross-validation to ensure that each module can be trained and tested on an independent data set, thereby evaluating and improving the generalization ability of the model.

[0047] Specific implementation plan 2: In step 1, collect the key data of well pressure, temperature, and flow rate through on-site monitoring equipment, including all polymer injection data and monthly oil production data from the start of tertiary oil recovery to the stop of polymer injection. In order to ensure the timeliness and adaptability of the data, set a flexible collection frequency according to the specific needs and on-site conditions of the oilfield, so as to balance the real-time nature of the data and the processing cost; the preprocessing of the well data is specifically: fill in the missing values of the well data, delete or replace the abnormal values in the data, and normalize the data through min-max normalization or Z-score normalization, converting all feature data to a unified scale, making the model training process more stable and also accelerating the convergence speed of the algorithm. Other parts of this implementation plan are the same as those of specific implementation plan 1.

[0048] Specific Embodiment 3: Identify missing values in the data through statistical mean, median, or time series; identify outliers in the data through box plot analysis or standard deviation analysis. This embodiment ensures the reliability and applicability of the data, providing a solid data foundation for establishing an accurate and efficient reservoir production prediction model. Other aspects of this embodiment are the same as those of Specific Embodiment 2.

[0049] Specific Embodiment 4: In Step 1, the Boruta algorithm is used to extract key features for reservoir production prediction from the preprocessed data, including: randomly generating "shadow features", comparing the original features with the randomly generated "shadow features" by inputting them into a random forest model for evaluation, comparing the contributions of the original features and the "shadow features" to the model's prediction ability, thereby determining the importance scores of each feature, evaluating the actual value of the features, retaining the features with importance scores significantly higher than those of the shadow features to obtain feature-optimized data; and updating the feature set in real time through a sliding window. This embodiment also eliminates those features that contribute little or nothing to the prediction, verifies data correlation, reduces the feature dimension, and optimizes the usage efficiency of computing resources; objectively measures the importance of features through the Boruta algorithm without being affected by prior knowledge or human bias. Other aspects of this embodiment are the same as those of Specific Embodiment 1.

[0050] Specific Embodiment 5: As Figure 2 shown, in Step 2, the hyperparameters of the XGBoost module are optimized as follows: The Optuna hyperparameter optimization algorithm is used to tune the hyperparameters of the XGBoost module. The TPE (Tree-structured Parzen Estimator) algorithm is used to perform global hyperparameter iterative search in the hyperparameter space. A grid search with a step size of 0.05 is initiated within the optimal parameter region, and an early stopping mechanism (terminate if the loss decrease is less than 1% for 5 consecutive iterations) is set to prevent overfitting, ultimately determining the optimal hyperparameter combination, such as the number of trees, depth, and learning rate. Other aspects of this embodiment are the same as those of Specific Embodiment 4.

[0051] The hyperparameters of the XGBoost module are tuned using a hybrid strategy of Bayesian optimization and local grid search, improving the prediction performance of the XGBoost module.

[0052] Specific Embodiment 6: The objective function of the XGBoost module is:

[0053]

[0054] The loss function in the objective function of the XGBoost module is represented by the true value y and the predicted value , n represents the number of samples, and Ω(F i) is the regularization term;

[0055] Adopt a combined loss function Its expression is:

[0056]

[0057] Among them, is the Huber loss, is the MSE dynamic weighted loss, λ is the intermediate value, which is dynamically adjusted according to the outlier ratio, and the calculation formula is:

[0058]

[0059] The expression of the predicted value is:

[0060]

[0061] Among them represents the sample prediction result after the t-th iteration, represents the prediction results of the first t - 1 trees, f t (x i ) represents the t-th tree model. Other parts of this implementation plan are the same as those of the fifth specific implementation plan.

[0062] The objective function of the XGBoost module in this implementation plan consists of two parts: the loss function and the regularization term. The loss function measures the deviation between the predicted value of the model and the actual value, while the regularization term prevents the model from being overly complex, thus avoiding overfitting. Introduce a combined loss function (Huber loss and MSE dynamic weighting) to suppress the interference of outliers; integrate the online learning mechanism, and detect data distribution drift and trigger retraining through the ADWIN algorithm.

[0063] Specific implementation plan seven: Weigh each module according to the fitting effect of each module on the training set. Specifically, determine the weight through the R 2 value on the training set:

[0064]

[0065] Among them

[0066]

[0067] Among them, y true,i represents the actual output, y pred,i represents the predicted output, represents the average value of the actual output.

[0068] Other parts of this implementation plan are the same as those of the sixth specific implementation plan.

[0069] Specific Embodiment 8: The fusion of the prediction results of each module of the reservoir production prediction model in Step 3 is specifically as follows:

[0070]

[0071] Among them, T is the predicted value of the multi-comprehensive intelligent model, and W i is the weight of the i-th model, and pred i is the predicted value of the i-th model, and the weight W i is calculated based on the evaluation indexes (such as R 2 value) of each model on the training set.

[0072] Other parts of this embodiment are the same as those of Specific Embodiment 7.

[0073] Specific Embodiment 9: In this embodiment, the mean absolute error (MAE) and the mean square error (MSE) are calculated to measure the deviation between the predicted value and the actual value of the model. The coefficient (R 2 ) is used to evaluate the fitting degree of the model to the data. Through these indexes, the accuracy and reliability of the model prediction are jointly ensured. Other parts of this embodiment are the same as those of Specific Embodiment 8.

[0074] As Figure 3 shown, a multi-comprehensive intelligent polymer flooding reservoir production prediction method (algorithm) based on polymer injection data proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on the algorithm. A multi-comprehensive intelligent polymer flooding reservoir production prediction system based on polymer injection data is developed by using a programming language based on the method proposed by the present invention. The system has program modules corresponding to the steps of the above technical solution, and executes the steps in the above multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data when running. The Dask framework is used to implement distributed parallel computing. The data preprocessing task is divided into 100 shards, and the GPU cluster is used to accelerate the training of the XGBoost module, and the parameter tree_method = 'gpu_hist' is set to achieve TB-level data parallel processing.

[0075] The computer program of the developed system (software) is stored on a computer-readable storage medium, and the computer program is configured to implement the steps of the above multi-comprehensive intelligent polymer flooding reservoir production prediction method when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.

[0076] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0077] The computational programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor and can implement these computational programs using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0078] Although the present disclosure is presented as above, the scope of protection of the present disclosure is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. A multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data, characterized in that, It includes the following steps: Step 1: Collect well data from the reservoir site, preprocess the well data, extract key features for reservoir production prediction from the preprocessed data to obtain feature-optimized data, and randomly divide the feature-optimized data into a training set and a test set; Step 2: Build a reservoir production prediction model, which includes an XGBoost module, a linear regression module, a support vector machine module, and a random forest module, and optimize the hyperparameters of the XGBoost module; Train each module based on the training set, test the module performance using the test set, and weight each module based on the fitting effect of each module on the training set; Step 3: Integrate the prediction results of each module of the reservoir production prediction model to obtain the reservoir production prediction result; Step 4: Receive new polymer injection data in real time through an online learning framework, and trigger model incremental update in combination with a concept drift detection algorithm.

2. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 1, wherein In Step 1, the preprocessing of the well data specifically includes: filling in missing values of the well data, deleting or replacing outliers in the data, and normalizing the data.

3. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 2, wherein In Step 1, missing values in the data are identified by statistical mean, median, or time series; outliers in the data are identified by box plot analysis or standard deviation analysis.

4. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 1, wherein In Step 1, the Boruta algorithm is used to extract key features for reservoir production prediction from the preprocessed data, and the feature set is updated in real time through a sliding window to obtain feature-optimized data.

5. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 4, wherein In Step 2, the optimization of the hyperparameters of the XGBoost module specifically includes: using the Optuna hyperparameter optimization algorithm to tune the hyperparameters of the XGBoost module, using the TPE algorithm for global hyperparameter iterative search, performing grid search within the optimal parameter region, and setting an early stopping mechanism to prevent overfitting, and finally determining the optimal hyperparameter combination.

6. The method for predicting the production of a multi-comprehensive intelligent polymer flooding reservoir based on polymer injection data according to claim 5, wherein The objective function of the XGBoost module in Step 2 is: Loss function in the objective function of the XGBoost module Represented by the true value y and the predicted value , where n represents the number of samples, and Ω(f i ) is the regularization term; The expression of the predicted value is: Among them represents the sample prediction result after the t-th iteration, represents the prediction result of the first t - 1 trees, F t (x i ) represents the t-th tree model.

7. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 6, wherein, In Step 2, the weighting of each module based on the fitting effect of each module on the training set specifically includes determining the weight through the R2 value of each module on the training set: where Among them, ytrue,i represents the actual output, and ypred,i represents the predicted output. represents the average value of the actual output.

8. The multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data according to claim 7, characterized in that, In Step 3, the integration of the prediction results of each module of the reservoir production prediction model specifically includes: Among them, Y is the predicted value of the multi-integrated intelligent model, and W i is the weight of the i-th model, and pred i is the predicted value of the i-th model.

9. A multi-comprehensive intelligent polymer flooding reservoir production prediction system based on polymer injection data, characterized in that, The system has program modules corresponding to the steps of the method described in any one of claims 1 to 8 above, and when running, executes the steps in the above-mentioned multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps in the multi-comprehensive intelligent polymer flooding reservoir production prediction method based on polymer injection data described in any one of claims 1 to 8 when called by a processor.

Citation Information

Cited By

  • Polymer flooding reservoir yield prediction method and system based on dual attention mechanism and storage medium

    CN121071408A

  • A method and system for predicting the production of polymer flooding reservoirs based on a double attention mechanism and a storage medium

    CN121071408B