4D intelligent prediction and diagnosis method and system for spatio-temporal evolution of digital twin oil and gas reservoir damage
By constructing a digital twin-based 4D intelligent prediction and diagnosis method for the spatiotemporal evolution of oil and gas reservoir damage, and utilizing multidimensional features and the LightGBM model, the uncertainty problem in predicting complex reservoir damage is solved, achieving high-precision prediction and dynamic updating of reservoir damage, and supporting the optimization of oil and gas field development.
Patent Information
- Application Number
- CN202511081568.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Existing technologies struggle to accurately capture reservoir damage characteristics in complex reservoirs, leading to increased uncertainty in damage prediction. Furthermore, the lack of specialized modeling systems for multi-physics and multi-temporal coupling processes makes it difficult to meet the accuracy requirements of oil and gas field development.
A digital twin-based intelligent 4D prediction and diagnosis method for the spatiotemporal evolution of oil and gas reservoir damage is constructed. By inputting multi-dimensional features into an intelligent model and combining a Light Gradient Boosting Machine (LightGBM) and a penalty coefficient to design a loss function, high-precision prediction and dynamic updating of reservoir damage rate are achieved.
It enables high-precision visual modeling and intelligent prediction of reservoir damage, supports real-time adjustment of development plans, and optimizes reservoir protection and development operations.
Smart Images

Figure CN120611633B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas engineering technology, specifically to a digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method, system, device, processor, and computer program product. Background Technology
[0002] Reservoir damage is a prevalent and critical problem in oil and gas field development, primarily manifested as physical or chemical alterations to the reservoir's pore structure, leading to decreased permeability and reduced production capacity. The causes of reservoir damage are complex, involving various physicochemical processes such as solid particle migration, pore blockage, chemical precipitation reactions, clay mineral swelling, and interfacial emulsification. These processes exhibit significant spatial and temporal dynamics. As oil and gas field development progresses towards complex well structures, deep formations, and low- to ultra-low-permeability reservoirs, reservoir damage problems are becoming increasingly prominent, posing a serious challenge to efficient oil and gas field development and management. Traditional techniques involve detailed analysis of reservoir cores in the laboratory, including the physical, chemical, and mineral properties of the rocks, to assess reservoir damage. However, laboratory core analysis is typically conducted under controlled conditions, which differ from the actual temperature, pressure, and fluid conditions in reservoirs. These differences lead to discrepancies between analytical results and actual conditions. Furthermore, for complex reservoirs, such as fractured or highly heterogeneous reservoirs, core analysis struggles to accurately capture all reservoir characteristics, increasing the uncertainty in damage prediction. Furthermore, the interpretation of core analysis results may be influenced by the analyst's experience and knowledge, and thus involves a degree of subjectivity.
[0003] In recent years, with the development of the Industrial Internet and edge computing technologies, "digital twin" technology has been gradually applied to the state modeling and evolution prediction of complex systems, demonstrating excellent real-time mapping and control capabilities in fields such as aerospace manufacturing and smart grids. In oil and gas engineering, digital twins can realize the virtual mirroring, evolutionary deduction, and dynamic feedback of complex geological bodies and development behaviors by constructing a multi-dimensional correspondence between the physical state of reservoirs and virtual models.
[0004] However, most existing studies on digital twins of oil and gas reservoirs focus on development process simulation (such as production prediction and pressure evolution) or equipment operation monitoring, lacking a specialized modeling system for the multi-physics, multi-temporal coupled process of reservoir "damage evolution." Especially in reservoir damage scenarios, micro-scale evolution processes such as pore-throat structure destruction, fluid property changes, and mineral interactions are difficult to accurately represent in macro-scale models, leading to a disconnect between "digital twins" and "physical mechanisms," making it difficult to meet the accuracy requirements of dynamic prediction and development control.
[0005] Therefore, there is an urgent need for a "full-process, full-scale, and full-state" digital twin system that integrates image reconstruction, parameter analysis, model learning, and feedback mechanisms to achieve high-precision visual modeling, intelligent prediction, and dynamic updating of oil and gas reservoir damage, providing scientific support for subsequent reservoir protection and development optimization. Summary of the Invention
[0006] The purpose of this invention is to provide a digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method, system, device, processor and computer program product to solve or at least partially solve the above-mentioned defects of the prior art.
[0007] To achieve the above objectives, a first aspect of the present invention provides a 4D intelligent prediction and diagnosis method for the spatiotemporal evolution of digital twin oil and gas reservoir damage, the method comprising:
[0008] The digital twins corresponding to the well section samples in the target reservoir are converted into corresponding multi-dimensional features, and the multi-dimensional features are input into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir in the future time period.
[0009] The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rates. The model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined in the following manner:
[0010] For each training data point in the multiple training data sets, the following processing steps are performed: inputting the multidimensional features corresponding to the training data point into a preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data point; and determining the deviation corresponding to the training data point based on the empirical value and observed value of the reservoir damage rate corresponding to the training data point.
[0011] The penalty coefficient is determined based on the deviation and hyperparameter corresponding to each training data point in the plurality of training data; and
[0012] The loss function is determined based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model.
[0013] Preferably, determining the deviation corresponding to each training data point includes:
[0014] Based on the empirical value of the reservoir damage rate corresponding to each training data point and observed values Determine the deviation corresponding to each training data point. :
[0015] .
[0016] Preferably, determining the penalty coefficient includes: based on the deviation corresponding to each training data point in the training data. and hyperparameters Determine the penalty coefficient :
[0017]
[0018] Where N is the total number of training data required for the loss function during the training process.
[0019] Preferably, determining the loss function includes: based on the penalty coefficient. and the deviation corresponding to each training data point The observed values and the model prediction value corresponding to the intelligent model. Determine the loss function :
[0020] .
[0021] Preferably, the preset empirical model is a multiple linear regression model, and the intelligent model is a Light Gradient Boosting Machine (LightGBM).
[0022] Preferably, the training data is established based on the SMOTE manual sampling technique.
[0023] Preferably, the modeling features corresponding to the intelligent model are determined according to the following method:
[0024] Multiple evaluation models are used to evaluate the importance of different feature combinations corresponding to the multidimensional features, and a feature importance ranking list corresponding to each evaluation model is generated.
[0025] Based on the feature importance ranking list corresponding to the multiple evaluation models, determine the comprehensive ranking position of feature importance for each feature combination; and
[0026] The feature combination corresponding to the minimum value of the comprehensive ranking of feature importance is used as the modeling feature.
[0027] Preferably, the multi-source input parameters include at least two of the following: lithology, burial depth, initial porosity, initial permeability, cementation type, montmorillonite content, illite content, chlorite content, kaolinite content, illite-montmorillonite mixed layer content, chlorite-montmorillonite mixed layer content, clay mineral content, quartz content, feldspar content, and formation water salinity.
[0028] A second aspect of this invention provides a 4D intelligent prediction and diagnosis system for the spatiotemporal evolution of digital twin oil and gas reservoir damage, the system comprising:
[0029] Feature extraction module: used to obtain digital twins corresponding to well section samples in the target reservoir, and convert the digital twins corresponding to well section samples in the target reservoir into corresponding multi-dimensional features;
[0030] Prediction module: used to input the multidimensional features into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir over a future time period;
[0031] The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rates. The model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined in the following manner:
[0032] For each training data point in the multiple training data sets, the following processing steps are performed: inputting the multidimensional feature corresponding to the training data point into a preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data point; and determining the deviation corresponding to the training data point based on the empirical value and observed value of the reservoir damage rate corresponding to the training data point.
[0033] The penalty coefficient is determined based on the deviation and hyperparameter corresponding to each training data point in the plurality of training data; and
[0034] The loss function is determined based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model.
[0035] A third aspect of the present invention provides a digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis device, comprising: a memory configured to store instructions; and a processor configured to call the instructions from the memory and to implement the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method when executing the instructions.
[0036] A fourth aspect of this invention provides a processor for running a program, wherein the program, when run, is used to execute: the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method.
[0037] The fifth aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method.
[0038] The digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method provided in this invention, on the one hand, designs the penalty coefficient and bias of the loss function based on the predicted values corresponding to the intelligent model and the empirical model, realizing the integration of knowledge in the field of reservoir damage into the intelligent model to form a knowledge-guided intelligent framework; on the other hand, by inputting multi-dimensional features into the intelligent model, it can simulate the damage evolution of different development stages such as water injection, fracturing, and acidizing, and optimize operating parameters. Through intelligent model prediction, the reservoir damage rate prediction results can also be obtained in real time, thereby supporting real-time adjustment of development plans.
[0039] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0040] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0041] Figure 1 This is a flowchart illustrating the 4D intelligent prediction and diagnosis method for spatiotemporal evolution of digital twin oil and gas reservoir damage provided in an embodiment of the present invention.
[0042] Figure 2 This is a schematic diagram of CT scan results provided in an embodiment of the present invention;
[0043] Figure 3 This is a schematic diagram of NMR scan results provided in an embodiment of the present invention;
[0044] Figure 4 This is a schematic diagram of the maximum ball algorithm provided in an embodiment of the present invention;
[0045] Figure 5 This is a schematic diagram of the pore radius provided in an embodiment of the present invention;
[0046] Figure 6 This is a schematic diagram of the throat length provided in an embodiment of the present invention. Detailed Implementation
[0047] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0048] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0049] To overcome the bottlenecks of traditional reservoir damage analysis methods, such as time response lag, insufficient spatial resolution, and lack of prediction dimensions, this invention proposes a digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method. It takes the four dimensions of "time, depth, space, and dynamic evolution" as a unified modeling object, constructs a three-in-one digital twin of "structure-parameter-behavior", and realizes reservoir damage rate prediction and evolution monitoring through 4D dynamic mapping and model feedback mechanism. Figure 1 This is a flowchart illustrating the 4D intelligent prediction and diagnosis method for the spatiotemporal evolution of digital twin oil and gas reservoir damage provided in this embodiment of the invention. Figure 1 As shown, the method may include: converting the digital twin of the well section sample in the target reservoir into the corresponding multi-dimensional features, and inputting the multi-dimensional features into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir in the future time period;
[0050] The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rate.
[0051] In practical applications, for any well section sample in the target reservoir S i Its digital twin is defined as:
[0052]
[0053] in, M i Representing three-dimensional structural features, including: pore radius Roar length and equivalent orifice-throat ratio Pore radius It can be constructed using CT scans and the maximum sphere algorithm, specifically through the following formula:
[0054] ;
[0055] In the above formula, Represents the entire coordinate space. x Indicates the scan coordinates. This represents the radius of the sphere corresponding to the scan coordinates. It represents the entire pore space.
[0056] In some embodiments, online computed tomography (CT) and nuclear magnetic resonance (NMR) techniques are first used to acquire the dynamic changes in reservoir pore structure and fluid distribution corresponding to the target reservoir in real time. Then, an equivalent pore and fracture network model is constructed based on the maximum sphere algorithm. Finally, the three-dimensional structural features are obtained by analyzing the constructed pore and fracture network model.
[0057] Specifically, each sample in the target reservoir was scanned 3600 times. The scanning pixel size was 200 nm, and the spatial resolution was 1 μm. The X-ray CT scanning steps were as follows: S1, the coal sample to be scanned was fixed on a precision sample stage, and the X-ray source was turned on; S2, a high-resolution detector detected the X-rays absorbed and attenuated by the scanned sample; S3, the computer software automatically recorded and stored the X-rays converted into electronic signals; S4, after a successful scan, the sample holder on the sample stage was rotated to a certain angle, and a new scan was repeated. When the sample holder had rotated 360°, the scanning of one coal sample was completed. The CT pore structure map and NMR fluid distribution map are shown below. Figure 2 and Figure 3 As shown.
[0058] The equivalent pore fracture network model is constructed using the maximum sphere algorithm: First, any point in the pore space is taken as the base point ( Figure 4 (The red part in the image) continuously searches for the largest inscribed sphere centered at that point and tangent to the skeleton boundary. Figure 4 (The red circle in the image); secondly, once all inscribed spheres have been found, other inscribed spheres contained within them will be removed, and the remaining inscribed spheres will form a sphere set; then, a clustering algorithm can be used to classify and merge the largest spheres and identify pores ( Figure 5 ) and throat ( Figure 6 Finally, pores can be represented by a larger sphere, and throats by a series of smaller spheres. Based on the maximum sphere algorithm for extracting pore and fissure network models, pores can be defined as the equivalent maximum inscribed sphere extracted from the pore and fissure network model. The radius of the inscribed sphere can be obtained using the equal-diameter expansion method of the equivalent sphere, which is the pore radius. Figure 5 The pore volume can be calculated based on the required pore radius.
[0059] P i This represents multiple input parameters, including reservoir properties (such as porosity). Penetration rate K ), mineral componentsC clay Formation water salinity σ w wait;
[0060] H i Indicates the historical evolution path, that is, different moments t ∈[ t 0, t n Below, the sequence of changes in reservoir damage rate:
[0061] ;
[0062] In the above formula, Indicates the sample in t k Reservoir damage rate at any given time.
[0063] In some embodiments, the multi-source input parameters include at least two of the following: lithology, burial depth, initial porosity, initial permeability, cementation type, montmorillonite content, illite content, chlorite content, kaolinite content, illite-montmorillonite mixed layer content, chlorite-montmorillonite mixed layer content, clay mineral content, quartz content, feldspar content, and formation water salinity.
[0064] In some embodiments, the model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined according to S101~S103:
[0065] S101. For each training data point among multiple training data points, perform processing steps S201~S202:
[0066] S201. Input the multi-dimensional features corresponding to the training data into the preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data.
[0067] In some embodiments, the preset empirical model is a multiple linear regression model. The multiple linear regression method finds the most suitable regression function with at least two independent variables from a large amount of data, which can be mathematically written as:
[0068]
[0069] In the above formula, X yes n × k The matrix of independent variables, Y It is the dependent variable. n dimensional vector, β yes k dimensional regression coefficient vector, ε It is a random error related to regression.
[0070] The above formula can be further transformed into:
[0071]
[0072] The least squares method can be used to estimate the regression coefficients. β The optimal regression equation is established by taking the value of and performing F-test and R-test.
[0073] when X When the rank is full, the regression coefficients can also be obtained using the following formula. β:
[0074]
[0075] In this embodiment of the invention, to make the multiple linear regression model applicable to a wider range of different oilfield reservoirs, the input parameters are expanded to the parameters described in Table 1, with the observed value being the reservoir damage rate. Modified β The values correspond one-to-one with F0~F14 in Table 1. To avoid introducing unnecessary errors and improve the usability of empirical formulas, this embodiment of the invention uses bias... ε A regression equation is introduced, and specific values are calculated. Finally, the regression coefficients of the multiple linear regression model are determined according to an embodiment of the present invention. β and bias ε as follows:
[0076] ; .
[0077] Table 1 Statistical Indicators of Datasets
[0078]
[0079] In some embodiments, the intelligent model is a Light Gradient Boosting Machine (LightGBM). LightGBM is a new implementation based on the gradient boosting machine algorithm, which is an optimization of the XGBoost algorithm. It uses the same gradient descent method as XGBoost, that is, at the current point, only a second-order Taylor expansion is performed on the loss function term in the objective function before solving for the optimum. This method retains more information about the objective function and improves the accuracy of prediction.
[0080] After performing a second-order Taylor expansion on the loss function, the objective function can be written as:
[0081] (1)
[0082] (2)
[0083] (3)
[0084] (4)
[0085] In the formula, x i These are input features. y i It is the actual value. It is the first t -1 step prediction value, For loss function, T The number of leaves w Let Ω be the weight vector. f t ) represents the regularization term, and γ and λ are user-defined hyperparameters.
[0086] Define a I j ={ i | q ( x i )= j Let} represent the set of input sample indices assigned to leaf node j. Then the objective function can be transformed into:
[0087] (5)
[0088] (6)
[0089] (7)
[0090] In the formula, G j and H j These are the first and second derivatives of each leaf node sample. Each term in ∑ is about... w j By solving the quadratic equation and setting its reciprocal to 0, the optimal weights can be obtained. .
[0091] (8)
[0092] Substituting the optimal weights into the objective function eliminates the weight variables. The objective function after eliminating the weight variables (Equation 13) can then guide the splitting of the t-th decision tree, much like the Gini index and squared error.
[0093] (9)
[0094] We aim to maximize the reduction in the objective function with each node split, expressed as follows:
[0095] (10)
[0096] The first, second, and third terms in the parentheses are the scores of the left child node, the right child node, and the parent node, respectively. Therefore, the division between left and right nodes is entirely determined by… L split Decide.
[0097] LightGBM abandons the level-wise decision tree growth strategy used by most gradient boosting decision tree (GBDT) tools, instead employing a leaf-wise algorithm with depth constraints. XGBoost uses a level-wise growth strategy, splitting leaves at the same level simultaneously during a single data traversal, making it easy to optimize with multi-threading, control model complexity, and reduce the risk of overfitting. However, the level-wise strategy is actually inefficient, treating leaves at the same level indiscriminately. Many leaves have low splitting gains, making searching and splitting unnecessary, thus increasing computational overhead. LightGBM's leaf-wise strategy finds the leaf with the highest splitting gain from all current leaves and splits it, repeating this process. Compared to the level-wise strategy, leaf-wise has the advantage of smaller errors and higher accuracy with the same number of splits. The disadvantage is that it may grow a deep decision tree, leading to overfitting. Therefore, LightGBM adds a maximum depth constraint on top of the leaf-wise strategy to prevent overfitting while maintaining high efficiency.
[0098] LightGBM's scalability is reflected in three aspects: using smaller sample sizes, less memory, and fewer features. This is achieved through three techniques: Gradient-based one-side sampling (GOSS), histogram algorithms, and exclusive feature bundling (EFB). GOSS is a sampling algorithm that aims to discard some samples that do not contribute to information gain, which greatly reduces computational and time costs.
[0099] Furthermore, compared to XGBoost's pre-sorting algorithm, LightGBM, based on the histogram algorithm, optimizes the time complexity from O(Data × features) to O(Bins × features). In terms of memory, the histogram algorithm saves approximately 7 times the memory consumption of the pre-sorting algorithm. The EFB algorithm is used to reduce the dimensionality of features; it can transform a large number of mutually exclusive features into low-dimensional dense features, effectively avoiding the computation of redundant zero-value features.
[0100] Therefore, LightGBM maintains high accuracy while being faster and having a smaller memory footprint. As oilfield datasets expand further in the future, LightGBM has potential applications in reservoir damage prediction and even the entire oil and jet fuel industry.
[0101] S202. Determine the deviation corresponding to the training data based on the empirical value and observed value of the reservoir damage rate corresponding to the training data;
[0102] Specifically, determining the deviation corresponding to each training data point includes: based on the empirical value of the reservoir damage rate corresponding to each training data point. and observed values Determine the deviation corresponding to each training data point. :
[0103] (11)
[0104] S102. Determine the penalty coefficient based on the deviation and hyperparameter corresponding to each training data in the plurality of training data;
[0105] Specifically, determining the penalty coefficient includes: based on the deviation corresponding to each training data point in the training data. and hyperparameters Determine the penalty coefficient :
[0106] (12)
[0107] Where N is the total number of training data required for the loss function during the training process.
[0108] It should also be noted that the penalty coefficient While maintaining sample weights, it facilitates further training of samples with large deviations from empirical correlations; hyperparameters Within the range of 0 to 1, a selection will be made using Bayesian optimization.
[0109] S103. Determine the loss function based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model.
[0110] Specifically, determining the loss function includes: based on the penalty coefficient. and the deviation corresponding to each training data point The observed values and the model prediction value corresponding to the intelligent model. Determine the loss function :
[0111] (13)
[0112] Furthermore, the loss function can be determined according to formulas (4) and (5). first gradient and second gradient as follows:
[0113] (14)
[0114] (15)
[0115] loss function first gradient and second gradient Integrate into LightGBM's custom loss function for intelligent model training.
[0116] It should be noted that the gradient of the loss function, that is, the derivative of the loss function with respect to the model parameters, is of great significance in machine learning and optimization. Its core role is to guide the direction of updating the model parameters in order to minimize the loss function.
[0117] For steps S101-S103, in order to integrate knowledge of reservoir damage into the intelligent model and form a knowledge-guided intelligent framework, the empirical model is used as a reference for model prediction in this embodiment of the invention. Traditional data-driven models often use mean squared error as the loss function, and adjust model parameters by observing changes in the loss function on the dataset to achieve the training objective. Therefore, a bias mechanism is designed... Penalty coefficient β DK The loss function is introduced as a weighting factor and a scaling factor.
[0118] In some embodiments, the following 4D mapping operator is constructed. Complete the process from twin input to reservoir damage rate. Predicted output:
[0119]
[0120] in, θ This represents the training parameters of the LightGBM intelligent model; M i Representing three-dimensional structural features, including: pore radius Roar length and equivalent orifice-throat ratio ; P i This represents multiple input parameters, including reservoir properties (such as porosity). Penetration rate K ), mineral componentsC clay Formation water salinity σ w wait; H i Indicates the historical evolution path, that is, different moments t ∈[ t 0, t n The sequence of changes in reservoir damage rate is shown below.
[0121] Time Dimension (T): The system can integrate historical data and real-time acquired parameters from multiple stages (drilling, completion, injection and production, acidizing, etc.) to construct a time series of damage evolution. Using the LightGBM model combined with historical conditions, it predicts future evolution trends. The prediction result at each time point is mapped to the current reservoir damage rate, forming an evolution path with memory characteristics. (Using time windows...) Construct a time series input by using the state difference sequence as the input variable. as follows:
[0122]
[0123] Depth Dimension (D): Based on the acquisition of reservoir CT scan images and mineral distribution characteristics, the system establishes a correspondence between the reservoir's vertical profile and geological strata, enabling the identification of mineral distribution, porosity and permeability structures, and changes in damage rates at different depths. This allows for the identification of different depth locations... Porosity Penetration rate Mineral components and formation water salinity Mapping to vector X D The definition is as follows:
[0124] ;
[0125] Spatial Dimension (S): Combining CT / NMR image processing with the maximum sphere algorithm to reconstruct a three-dimensional pore and fracture network, extracting equivalent pore throat structures, seepage channels, and damage-sensitive areas, achieving micron / millimeter-level three-dimensional spatial mapping. Simultaneously, it supports projecting the predicted damage rate data onto the 3D model according to wellbore location and grid coordinates to construct a reservoir damage rate distribution map. (The well section coordinates are then used for this purpose.) Mapped to a three-dimensional reservoir damage distribution function:
[0126]
[0127] In the above formula, Indicates the coordinate at time t. reservoir damage rate, This represents the cross-section of the reservoir damage rate at time t.
[0128] Dynamic Dimension (Δ): By training a LightGBM model with dynamic memory capabilities, an "empirical model bias-penalty factor" mechanism is introduced as a loss function constraint, enabling the model to not only fit the current input but also judge evolutionary trends. The system can dynamically update the model structure and parameters based on new input data, achieving continuous learning and iterative correction of damage expansion. A dynamic penalty loss function based on bias control is introduced:
[0129]
[0130] In the above formula, Indicates the penalty coefficient; This represents the deviation corresponding to each training data point; Represents the observed value; This represents the model prediction value corresponding to the intelligent model.
[0131] In summary, the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method proposed in this invention completes a closed loop from reservoir physical property acquisition, image modeling, feature extraction, intelligent prediction to visualization output through a combination of static and dynamic methods, image-structure fusion, and data-model collaboration. This system overcomes the limitations of existing technologies such as static assessment, two-dimensional prediction, and experience-driven approaches, providing a novel solution for damage monitoring and intelligent control of oil and gas reservoirs at different life cycles.
[0132] In some embodiments, the training data is constructed based on the synthetic minority class oversampling technique SMOTE.
[0133] Specifically, SMOTE is a synthetic artificial sampling data technique, an oversampling algorithm improved upon random sampling. It is used to overcome the problem of class imbalance in data. The SMOTE algorithm examines and simulates samples from several classes, and synthesizes data by oversampling samples from several classes.
[0134] The SMOTE algorithm generates data as follows: First, it applies the K-nearest neighbor method to each sample point. Then, it randomly selects N neighboring points from the K neighboring points. The difference between the sample and its nearest neighbor is multiplied by a random number between 0 and 1 to achieve data aggregation. Essentially, the algorithm samples in the feature space, not the data space. All neighboring points in the feature space have comparable features. Therefore, the SMOTE algorithm avoids the problem of undersampling leading to the loss of implicit information, and it also avoids the problem of overfitting caused by simple replication, as seen in oversampling methods.
[0135] Its mathematical expression can be written as:
[0136]
[0137] in, This represents the generated synthetic sample. This indicates that it represents the original minority class sample. Indicates from of K A sample randomly selected from the nearest neighbors, This represents a random number within the interval [0,1].
[0138] The inventors of this application discovered during their research that feature selection or data dimensionality reduction is one of the key steps in handling high-dimensional nonlinear problems. Reasonable feature selection can remove redundant and irrelevant features, which is beneficial for mining useful feature information and has a significant impact on classification and regression results. However, in previous reservoir damage prediction, feature selection was often arbitrary, easily leading to inaccurate model prediction results.
[0139] Based on this, embodiments of the present invention propose determining the modeling features corresponding to the intelligent model in the following manner:
[0140] S301. Use multiple evaluation models to evaluate the importance of different feature combinations corresponding to multidimensional features, and generate a feature importance ranking list for each evaluation model.
[0141] Specifically, the multiple evaluation models include: random forest, extreme gradient boosting tree, and light gradient boosting machine.
[0142] For each evaluation model: the feature importance ranking list corresponding to the evaluation model includes the feature importance ranking position and score corresponding to different feature combinations.
[0143] Random Forest assesses feature importance by calculating the decrease in out-of-bag (OOB) accuracy after adding noise to a feature, based on the out-of-bag error. XGBoost calculates the gain of each feature at the split node when constructing the decision tree; the greater the gain, the higher the feature importance. LightGBM, also based on the gradient boosting framework, assesses feature importance by calculating the gain at feature splits, and its histogram optimization technique makes feature evaluation more efficient.
[0144] S302. Based on the feature importance ranking list corresponding to the multiple evaluation models, determine the comprehensive ranking position of feature importance corresponding to any feature combination;
[0145] Specifically, for any given feature combination, the ranking of feature importance determined by the random forest corresponding to that feature combination is determined. Random forest determines the ranking of feature importance. , LightGBM Determined feature importance ranking and the number of features corresponding to this feature combination. Determine the overall ranking of the feature importance corresponding to this feature combination. ;
[0146]
[0147] S303. The feature combination corresponding to the minimum value of the comprehensive ranking of feature importance is used as the modeling feature.
[0148] For S303, the smaller the overall ranking of the feature importance corresponding to the feature combination, the higher the importance of the feature combination.
[0149] The method for determining the modeling features corresponding to the intelligent model proposed in this invention, on the one hand, evaluates the importance of feature combinations through three tree ensemble models (random forest, extreme gradient boosting tree, and light gradient boosting machine), which can ensure the distinguishability of features and the accuracy of evaluation results. Moreover, compared with unsupervised feature selection, it can more accurately reflect the relationship between features and targets. On the other hand, it uses ensemble sorting to provide a more comprehensive ranking of reservoir damage influencing factors.
[0150] This invention also provides a digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis system, the system comprising:
[0151] Feature extraction module: used to obtain digital twins corresponding to well section samples in the target reservoir, and convert the digital twins corresponding to well section samples in the target reservoir into corresponding multi-dimensional features;
[0152] Prediction module: used to input the multidimensional features into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir over a future time period;
[0153] The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rates. The model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined in the following manner:
[0154] For each training data point in the multiple training data sets, the following processing steps are performed: inputting the multidimensional feature corresponding to the training data point into a preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data point; and determining the deviation corresponding to the training data point based on the empirical value and observed value of the reservoir damage rate corresponding to the training data point.
[0155] The penalty coefficient is determined based on the deviation and hyperparameter corresponding to each training data point in the plurality of training data; and
[0156] The loss function is determined based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model.
[0157] This invention provides a 4D intelligent prediction and diagnosis device for the spatiotemporal evolution of digital twin oil and gas reservoir damage, comprising: a memory configured to store instructions; and a processor configured to retrieve the instructions from the memory and, when executing the instructions, to implement the 4D intelligent prediction and diagnosis method for the spatiotemporal evolution of digital twin oil and gas reservoir damage.
[0158] This invention provides a processor for running a program, wherein the program is executed to perform: the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method.
[0159] This invention provides a computer program product, including a computer program that, when executed by a processor, implements the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method as described above.
[0160] The digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method provided in this invention, on the one hand, designs the penalty coefficient and bias of the loss function based on the predicted values corresponding to the intelligent model and the empirical model, realizing the integration of knowledge in the field of reservoir damage into the intelligent model to form a knowledge-guided intelligent framework; on the other hand, by inputting multi-dimensional features into the intelligent model, it can simulate the damage evolution of different development stages such as water injection, fracturing, and acidizing, and optimize operating parameters. Through intelligent model prediction, the reservoir damage rate prediction results can be obtained in real time, thereby supporting real-time adjustment of development plans.
[0161] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0162] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0163] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0164] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0165] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0166] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0167] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0168] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0169] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A 4D intelligent prediction and diagnosis method for the spatiotemporal evolution of damage in digital twin oil and gas reservoirs, characterized in that, The method includes: The digital twins corresponding to the well section samples in the target reservoir are converted into corresponding multi-dimensional features, and the multi-dimensional features are input into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir in the future time period. The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rates. The three-dimensional structural features include pore radius, throat length, and equivalent pore-throat ratio. The multi-source input parameters include reservoir properties, mineral composition, and formation water salinity. The model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined in the following manner: For each training data point in the multiple training data sets, the following processing steps are performed: inputting the multidimensional features corresponding to the training data point into a preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data point; and determining the deviation corresponding to the training data point based on the empirical value and observed value of the reservoir damage rate corresponding to the training data point. The penalty coefficient is determined based on the deviation and hyperparameter corresponding to each training data point in the plurality of training data; and The loss function is determined based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model. The modeling features corresponding to the intelligent model are determined using the following method: Multiple evaluation models are used to evaluate the importance of different feature combinations corresponding to the multidimensional features, and a feature importance ranking list corresponding to each evaluation model is generated. Based on the feature importance ranking list corresponding to the multiple evaluation models, determine the comprehensive ranking position of feature importance for each feature combination; and The feature combination corresponding to the minimum value of the comprehensive ranking of feature importance is used as the modeling feature.
2. The method according to claim 1, characterized in that, Determining the bias corresponding to each training data point includes: Based on the empirical value of the reservoir damage rate corresponding to each training data point and observed values Determine the deviation corresponding to each training data point. : 。 3. The method according to claim 1, characterized in that, The determination of the penalty coefficient includes: Based on the deviation corresponding to each training data point in the training data and hyperparameters Determine the penalty coefficient : Where N is the total number of training data required for the loss function during the training process.
4. The method according to claim 1, characterized in that, The determined loss function includes: According to the penalty coefficient and the deviation corresponding to each training data point The observed values and the model prediction value corresponding to the intelligent model. Determine the loss function : 。 5. The method according to claim 1, characterized in that, The preset empirical model is a multiple linear regression model, and the intelligent model is a Light Gradient Boosting Machine (LightGBM).
6. The method according to claim 1, characterized in that, The training data was established based on the SMOTE manual sampling technique.
7. The method according to claim 1, characterized in that, The multi-source input parameters include at least two of the following: lithology, burial depth, initial porosity, initial permeability, cementation type, montmorillonite content, illite content, chlorite content, kaolinite content, illite-montmorillonite mixed layer content, chlorite-montmorillonite mixed layer content, clay mineral content, quartz content, feldspar content, and formation water salinity.
8. A digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis system, characterized in that, The system includes: Feature extraction module: used to obtain digital twins corresponding to well section samples in the target reservoir, and convert the digital twins corresponding to well section samples in the target reservoir into corresponding multi-dimensional features; Prediction module: used to input the multidimensional features into the intelligent model to determine the change sequence of the reservoir damage rate corresponding to the target reservoir over a future time period; The digital twin corresponding to the well section sample includes three-dimensional structural features, multi-source input parameters, and historical time-series data corresponding to reservoir damage rates. The three-dimensional structural features include pore radius, throat length, and equivalent pore-throat ratio. The multi-source input parameters include reservoir properties, mineral composition, and formation water salinity. The model parameters corresponding to the intelligent model are determined during training by minimizing a loss function, which is determined in the following manner: For each training data point in the multiple training data sets, the following processing steps are performed: inputting the multidimensional features corresponding to the training data point into a preset empirical model to determine the empirical value of the reservoir damage rate corresponding to the training data point; and determining the deviation corresponding to the training data point based on the empirical value and observed value of the reservoir damage rate corresponding to the training data point. The penalty coefficient is determined based on the deviation and hyperparameter corresponding to each training data point in the plurality of training data; and The loss function is determined based on the penalty coefficient, the bias corresponding to each training data point, the observed value, and the model prediction value corresponding to the intelligent model. The modeling features corresponding to the intelligent model are determined using the following method: Multiple evaluation models are used to evaluate the importance of different feature combinations corresponding to the multidimensional features, and a feature importance ranking list corresponding to each evaluation model is generated. Based on the feature importance ranking list corresponding to the multiple evaluation models, determine the comprehensive ranking position of feature importance for each feature combination; and The feature combination corresponding to the minimum value of the comprehensive ranking of feature importance is used as the modeling feature.
9. A digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis device, characterized in that, include: The memory is configured to store instructions; And a processor configured to retrieve the instructions from the memory and, when executing the instructions, to implement the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method according to any one of claims 1 to 7.
10. A machine-readable storage medium storing instructions thereon, characterized in that, This instruction is used to cause the machine to execute the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, The invention includes a computer program that, when executed by a processor, implements the digital twin oil and gas reservoir damage spatiotemporal evolution 4D intelligent prediction and diagnosis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Complex industrial system operation monitoring method and system based on digital twinning
CN115857447A