Intelligent prediction method and system for permeability of oil and gas reservoir
By acquiring multi-source geological data in real time and combining it with a dual-model parallel prediction of meta-learning and ensemble learning, the model fusion coefficients are dynamically generated. This solves the adaptation problem between small sample and big data stages in oil and gas reservoir permeability prediction, achieves adaptive improvement and continuity of prediction accuracy throughout the entire cycle, and reduces the cost of manual intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-06
- Publication Date
- 2026-04-07
AI Technical Summary
In the prediction of permeability of oil and gas reservoirs, existing technologies have limitations. Lightweight models designed with small samples quickly reach their performance bottleneck after data becomes abundant, while complex models designed with large datasets are difficult to train effectively in the early stages and lack dynamic evaluation and automatic switching mechanisms. This results in insufficient accuracy of prediction results and fails to guarantee the continuity and stability of engineering decisions.
By acquiring multi-source geological data in real time and combining parallel prediction with dual models based on meta-learning and ensemble learning, model fusion coefficients are dynamically generated to achieve linear weighted fusion. This ensures that the predicted values transition smoothly and continuously with changes in data conditions, reducing the cost of manual intervention and improving prediction accuracy and engineering practicality.
It achieves adaptive improvement in the prediction accuracy of oil and gas reservoir permeability throughout the entire cycle, avoids jumps in prediction results caused by changes in data conditions, ensures the continuous smoothness and generalization ability of prediction output, and reduces the cost of manual intervention.
Smart Images

Figure CN121808700A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas exploration technology, specifically to an intelligent prediction method and system for oil and gas reservoir permeability. Background Technology
[0002] With the deepening of the digital and intelligent transformation of the oil and gas industry, reservoir permeability prediction technology has evolved from traditional empirical formulas and statistical methods to intelligent prediction based on machine learning and deep learning. Existing technologies mostly adopt static prediction frameworks that are constructed once for specific data conditions.
[0003] However, oil and gas field exploration and development is a dynamic process where data gradually becomes abundant from extremely scarce. Existing static models cannot adapt to this continuously evolving data environment: lightweight models designed for small samples quickly reach their performance bottlenecks once data becomes plentiful, while complex models designed for large datasets are difficult to train effectively in the early stages due to insufficient data. Furthermore, current technologies lack mechanisms for dynamically evaluating and automatically switching models based on data conditions, relying on manual retraining and switching. This results in insufficient accuracy in predicting oil and gas reservoir permeability, failing to guarantee the continuity and stability of engineering decisions. Summary of the Invention
[0004] This invention provides an intelligent prediction method and system for oil and gas reservoir permeability, aiming to solve the technical problem of insufficient accuracy in the prediction results of oil and gas reservoir permeability in existing technologies.
[0005] In view of the above problems, the present invention provides an intelligent prediction method and system for oil and gas reservoir permeability.
[0006] In a first aspect, the present invention provides an intelligent prediction method for oil and gas reservoir permeability, comprising: Real-time acquisition of multi-source geological data of the target reservoir, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions; The multi-source geological data are simultaneously input into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. Combined with a preset period, multidimensional data features of the multi-source geological data are obtained. Combined with a switching decision unit, model fusion coefficients are dynamically generated. The model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task. Based on the model fusion coefficient, the first predicted value and the second predicted value are linearly weighted and fused to generate the final permeability prediction value of the target reservoir. The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
[0007] Secondly, the present invention provides an intelligent prediction system for oil and gas reservoir permeability, comprising: A multi-source data acquisition module is used to acquire multi-source geological data of the target reservoir in real time, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions; The dual-model parallel prediction module is used to simultaneously input the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. The fusion coefficient dynamic generation module is used to obtain multi-dimensional data features of the multi-source geological data in combination with a preset period, and dynamically generate model fusion coefficients in combination with the switching decision-maker. The model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task. The linear weighted fusion output module is used to perform linear weighted fusion of the first predicted value and the second predicted value according to the model fusion coefficient to generate the final permeability prediction value of the target reservoir. The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
[0008] One or more technical solutions provided in this invention have at least the following technical effects or advantages: This invention provides an intelligent prediction method and system for oil and gas reservoir permeability. It strengthens the prediction data foundation by collecting and accumulating multi-source geological data in real time, achieves accurate adaptation to different data scenarios by relying on parallel prediction of dual models, and completes intelligent matching of data conditions and prediction weights by dynamically generating model fusion coefficients through switching decision units. Finally, through linear weighted fusion and continuous adjustment of fusion coefficients, it not only achieves adaptive improvement of the prediction accuracy of reservoir permeability throughout the entire cycle, but also effectively avoids the jump in prediction results caused by changes in data conditions, ensuring the continuous smoothness of prediction output. At the same time, it reduces the cost of manual intervention and improves the engineering practicality and generalization ability of the prediction method. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating an intelligent prediction method for oil and gas reservoir permeability provided in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an intelligent prediction system for oil and gas reservoir permeability provided in an embodiment of the present invention; The components represented by each number in the attached diagram are explained below: The system includes a multi-source data acquisition module 11, a dual-model parallel prediction module 12, a dynamic generation module for fusion coefficients 13, and a linear weighted fusion output module 14. Detailed Implementation
[0011] This invention provides an intelligent prediction method and system for oil and gas reservoir permeability, which addresses the technical problem of insufficient accuracy in the prediction results of oil and gas reservoir permeability in existing technologies.
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0013] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0014] Example 1, as Figure 1 As shown, this invention provides an intelligent prediction method for oil and gas reservoir permeability, the method comprising: S100: Real-time acquisition of multi-source geological data of the target reservoir, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions.
[0015] In this embodiment of the invention, multi-source geological data of the target reservoir is acquired in real time, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions. Reservoir permeability is essentially the ability of a formation to allow fluids to pass through, and its magnitude is directly related to various geological factors such as formation lithology, pore structure, hydrocarbon potential, and structural characteristics. Single-type geological data can only reflect a single dimension of reservoir characteristics and cannot provide complete information support for subsequent permeability prediction. Furthermore, oil and gas field exploration and development is a dynamic process; initially, only a small amount of data can be acquired, and as drilling and logging operations progress, data is continuously supplemented. The subsequent dynamic switching and fusion of the first and second prediction models depends on the continuously accumulating data volume and data characteristics. Static or single-dimensional data will lead to poor model generalization ability and limited prediction accuracy. Therefore, it is necessary to acquire multi-type, cumulative geological data in real time to build a data foundation for full-cycle adaptive prediction.
[0016] Step S100 in the method provided in this embodiment of the invention includes: The multi-source geological data includes well logging curves, well logging curves, and seismic gathers.
[0017] First, the scope and types of multi-source geological data acquisition are determined. This multi-source geological data includes well logging curves, well logging data, and seismic gathers. Well logging curves are continuous curves formed by measuring formation physical parameters such as natural gamma (GR), acoustic transit time (AC), density (DEN), neutron porosity (CNL), and resistivity (RT) as the well is lowered into the well during drilling. These curves change with depth and are core data for identifying reservoir lithology and porosity. Well logging curves are curves formed by recording parameters such as cuttings characteristics, drilling fluid properties, and gas measurements in real time during drilling. These curves change with depth or drilling time and can quickly determine the hydrocarbon content of the reservoir. Seismic gathers are datasets formed in seismic exploration by artificially generating seismic waves and receiving reflected signals. Multiple seismic signal records corresponding to the same common midpoint (CMP) constitute a gather, reflecting the undulations of the formation structure and the lateral distribution characteristics of the reservoir.
[0018] For example, for Block A of a certain oil field, the sampling range is clearly defined as 10km of the block. 2 All exploration wells, development wells, and seismic exploration coverage areas are included. Core data types are: well logging curves: GR, AC, DEN, CNL, RT; well logging curves: cuttings oil saturation, gas logging total hydrocarbon value; seismic gathers: CMP gathers, ensuring coverage of key dimensions such as reservoir lithology, porosity, hydrocarbon content, and structure.
[0019] Secondly, real-time acquisition of multi-source geological data of the target reservoir is crucial. A real-time data acquisition link is established to simultaneously acquire multi-source geological data. This real-time data acquisition link is a closed-loop system consisting of downhole sensors, surface receiving equipment, data transmission modules, and terminal storage systems, enabling zero-delay data transmission from acquisition to storage.
[0020] For example, well logging curve acquisition: When drilling the newly added development well No. 4 in Block A of this oilfield to 1500m above the target reservoir top boundary, the logging-while-drilling tool is run into the drill pipe. The instrument measures parameters such as GR (range 50-120 API) and AC (range 220-350 μs / m) in the reservoir section from 1500 to 1800m in real time. The data is transmitted to the surface logging workstation in real time through the transmission cable inside the drill pipe. A data point is recorded every 10cm depth to form a continuous well logging curve. Well logging curve acquisition: During the drilling of well No. 4, the logging engineer collects cuttings samples every 30 minutes. The oil saturation (range 10%-45%) is measured using a cuttings analyzer, and the total hydrocarbon value (range 0.1%-5.2%) is recorded in real time using a gas meter. The data is automatically synchronized to the surface logging system, and well logging curves are generated according to drilling time (or depth). Seismic gather acquisition: During the seismic exploration operation in the block, 10 seismic source vehicles generate seismic waves on the surface, and 200 receivers and detectors receive the reflected signals from the strata. After analog-to-digital conversion, the signals are transmitted in real time to the seismic data processing center by a 5G transmission module. After processing, CMP seismic gathers in the 1500-3000m depth range of the block are formed, and each gather contains 200 seismic signal records.
[0021] Finally, a data accumulation and storage mechanism is established to dynamically update the data volume. This mechanism involves associating newly acquired multi-source geological data with historical data according to an index rule of block-well number-depth, forming a cumulative database to ensure that the data volume increases with the number of acquisitions.
[0022] For example, before drilling Well No. 4, the database already stored the logging curves, well logging data, and initial seismic gathers of exploration wells No. 1-3 in Block A of the oilfield, totaling 12,000 data points. After drilling Well No. 4, five new logging curves (GR, AC, etc.), one oil saturation logging curve, and 30 supplementary CMP seismic gathers were added, adding 8,000 data points, bringing the total database size to 20,000 data points. During subsequent data acquisition, new data was appended to the database according to the indexing rules after each acquisition, ensuring the data volume continuously increases with the number of acquisitions.
[0023] In this embodiment of the invention, by integrating three types of data—well logging, well logging, and seismic data—covering key characteristics such as reservoir lithology, porosity, hydrocarbon content, and structure, the problem of predictive bias caused by insufficient single data dimensions is solved. A real-time acquisition mechanism ensures timely data access, and an incremental storage mode enables dynamic expansion of data volume, adapting to the dynamic switching requirements of subsequent meta-learning models and ensemble learning models. A unified block-well-depth index rule and cumulative database are constructed, providing a structured and traceable data foundation for subsequent multi-dimensional data feature extraction and model training, reducing the complexity of data preprocessing.
[0024] S200: Simultaneously input the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value.
[0025] In this embodiment of the invention, the multi-source geological data is simultaneously input into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. Permeability prediction needs to adapt to the full-cycle data conditions of oil and gas field exploration, from the initial small-sample stage to the mature stage of big data: the meta-learning model excels at rapid adaptation to small samples and can provide robust predictions when data is scarce, but its performance easily reaches a bottleneck when data is abundant; the ensemble learning model relies on mining complex patterns from a large amount of data, resulting in higher prediction accuracy, but it is difficult to train effectively in the small-sample stage. Directly fixing a single model will lead to uneven prediction accuracy throughout the entire cycle; at the same time, outliers may exist in the multi-source geological data, such as invalid data caused by logging instrument malfunctions, which would affect the reliability of the prediction if directly input into the model. Therefore, it is necessary to first screen valid samples and dynamically update the dual models according to the sample size, and then obtain the dual prediction values through parallel input, providing high-quality, scenario-adaptive basic prediction results for subsequent dynamic fusion.
[0026] Step S200 in the method provided in this embodiment of the invention includes: Specifically, the process of simultaneously inputting the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain first and second predicted values includes: Analyze the multi-source geological data, extract valid sample data, and statistically obtain the number of valid samples; The effective sample size is iteratively determined by combining the preset trigger sample size; If the effective sample size is less than the trigger sample size, then the first prediction model is constructed and iteratively updated based on meta-learning according to the preset first sample increment trigger value. If the effective sample size is greater than or equal to the trigger sample size, then based on the preset second sample increment trigger value, the first prediction model is synchronously constructed and iteratively updated based on meta-learning, and the second prediction model is constructed and iteratively updated based on ensemble learning.
[0027] First, the multi-source geological data was analyzed to extract valid sample data and calculate the number of valid samples. Valid sample data refers to geological data that, after removing outliers and missing values exceeding the threshold, can completely reflect reservoir characteristics and can be correlated with permeability calibration values. Each sample contains a set of multi-source data and the corresponding true permeability value. The number of valid samples refers to the total number of valid sample data, which is a core indicator for judging whether the model training conditions are mature. A three-step method was used to screen valid samples: outlier detection, missing value handling, and correlation verification. Outliers exceeding the normal range in well logging curves were removed using the 3σ criterion, such as GR > 300 API and AC < 180 μs / m. For samples with a missing value ratio ≤ 10%, the K-nearest neighbor (K=5) algorithm was used to complete the missing values. Only samples that match the true permeability values of the core samples were retained, and the total number of remaining samples was calculated.
[0028] For example, the 20,000 original data points in Block A of the oilfield were processed as follows: 1,200 abnormal data points, such as GR=350API and AC=150μs / m, were removed; the CNL value of the sample at a depth of 1550m in Well No. 4 was missing, accounting for 5%, and was filled in by the average CNL value of the five adjacent valid samples; only samples that could be associated with the true permeability value of the core were retained, and finally 158 valid samples were obtained, that is, the number of valid samples = 158.
[0029] Secondly, the effective sample size is iteratively determined based on a preset trigger sample size. The trigger sample size is the critical sample size for deciding whether to activate the second prediction model, and it needs to be calibrated according to the complexity of the target reservoir. In the small sample stage, only the first prediction model is activated. A first sample increment trigger value is set. When the effective sample size is less than the trigger sample size, the threshold for the number of new samples added in the first prediction model iteration update is triggered. A second sample increment trigger value is set. When the effective sample size is greater than or equal to the trigger sample size, the threshold for the number of new samples added in the dual-model iteration update is triggered simultaneously.
[0030] For example, the lithological complexity of the sandstone reservoir in Block A of this oilfield is medium, and the preset trigger sample size N=80. Based on the data accumulation rate, the first sample increment trigger value ΔN1=20 is set, meaning the first model is updated every 20 new valid samples. The second sample increment trigger value ΔN2=50, meaning the dual model is updated synchronously every 50 new valid samples. The change in the valid sample size is monitored in real time, and a judgment is made every time a new batch of data is added. For example, if the initial valid sample size is 65, which is less than the second trigger sample size of 80, the iteration is based on the first sample increment trigger value. When data from well No. 4 is supplemented, the valid sample size increases to 158≥80, and the system switches to synchronously updating the dual model based on the second sample increment trigger value. Subsequently, every 50 new valid samples triggers synchronous iteration of the dual model.
[0031] Furthermore, if the effective sample size is less than the trigger sample size, the first prediction model is constructed and iteratively updated based on meta-learning according to a preset first sample increment trigger value. Iterative update refers to fine-tuning and optimizing the original model parameters based on newly added effective samples, avoiding full retraining and improving model update efficiency. If the effective sample size is less than the trigger sample size, the parameters of the first prediction model are fine-tuned only based on the newly added samples. For example, if the initial effective sample size is 65, and 20 new effective samples are added, only the specialized parameters of the first prediction model are fine-tuned.
[0032] Furthermore, if the effective sample size is greater than or equal to the trigger sample size, then based on a preset second sample increment trigger value, the first prediction model is synchronously constructed and iteratively updated using meta-learning, and the second prediction model is synchronously constructed and iteratively updated using ensemble learning. If the effective sample size is greater than or equal to the trigger sample size, based on the newly added samples, the parameters of the first model are simultaneously fine-tuned, and the feature matrix and weight decision layer of the second prediction model are updated. For example, when the cumulative effective sample size is 158 ≥ 80, 50 additional effective samples are added, and the meta-learning parameters of the first model are synchronously fine-tuned, and the base learner prediction results and ridge regression weights of the second model are updated.
[0033] The construction steps of the first prediction model include: Load a pre-built meta-learning base model, wherein the meta-learning base model is constructed through multiple diverse few-sample tasks based on geological constraints; Based on the multi-source geological data, a small amount of calibration data for the target reservoir is extracted and calibrated, wherein the small amount of calibration data is calibrated based on permeability; Based on the limited calibration data, the meta-learning basic model is rapidly and adaptively trained, and the specialized model parameters of the target reservoir are output according to the training results. The neural network model associated with the meta-learning base model is instantiated using the specialized model parameters to obtain the first prediction model.
[0034] First, a pre-built meta-learning foundation model is loaded. This model is constructed through multiple diverse small-sample tasks based on geological constraints. The meta-learning foundation model is a neural network model built on a Model-Independent Meta-Learning (MAML) framework. It obtains general initial parameters through multi-task training, enabling rapid adaptation to new tasks. Geological constraints are conditions set based on reservoir geology principles, such as sandstone reservoir porosity ≤40% and permeability ≤1000mD, used to divide the diverse small-sample tasks. Small-sample tasks treat reservoir data from different regions and strata as independent tasks, with each task containing only a small number of samples. Multi-source geological data from five sandstone reservoirs similar to the target reservoir are pre-collected. These are then divided into 10 small-sample tasks based on geological constraints. Each task corresponds to a reservoir prediction scenario for a sub-block. The MAML framework is trained on each small-sample task to optimize and obtain general initial parameters θ, forming the meta-learning foundation model, which is stored in a model library and loaded directly when needed. For example, a pre-trained MAML meta-learning basic model is loaded. This model is trained on 10 small sample tasks, such as shallow sandstone in Block B and deep sandstone in Block C. The initial parameters θ include the convolutional layer weights and fully connected layer biases of the neural network, which can adapt to the basic characteristics of sandstone reservoirs.
[0035] Secondly, based on the aforementioned multi-source geological data, a small amount of calibration data for the target reservoir is extracted and calibrated. This small amount of calibration data uses permeability as the calibration object. Calibration data refers to data calibrated using permeability as the calibration object; that is, each sample contains multi-source geological data features and the corresponding true permeability value from core experiments, used for adapting and training the meta-learning basic model. From the valid samples of the target reservoir, a small number of samples are randomly selected, and the true permeability value of each sample is measured through core experiments to complete the calibration. For example, from 158 valid samples in Block A of this oilfield, samples from 10 depth points, such as 1500m and 1520m, of Well No. 4 are selected. Each sample contains multi-source features such as GR (60-110 API) and AC (230-320 μs / m), and the corresponding true permeability values (50-800 mD) are measured through core experiments, forming 10 sets of calibration data.
[0036] Furthermore, based on the limited calibration data, the meta-learning base model is rapidly adaptively trained, and specialized model parameters for the target reservoir are output according to the training results. Rapid adaptive training refers to performing a small number of gradient descent iterations on the general initial parameters θ of the meta-learning base model based on a limited amount of calibration data to quickly adapt to the specific characteristics of the target reservoir. The specialized model parameters are specific model parameters obtained after adapting to the target reservoir, enabling the model to predict only for that target reservoir. Ten sets of calibration data are input into the meta-learning base model, with an inner loop learning rate of 0.01 and an outer loop learning rate of 0.001. Three-step gradient fine-tuning is performed to minimize the mean squared error (MSE) between the predicted permeability and the true value, outputting specialized model parameters θ1 adapted to the target reservoir. For example, ten sets of calibration data from block A are input into the MAML base model. After three steps of gradient fine-tuning, specialized parameters θ1 are obtained, where the convolutional layer weights are optimized for the correlation between GR and permeability in that block, and the fully connected layer biases are adapted to the porosity characteristics of that block.
[0037] Finally, the neural network model associated with the meta-learning base model is instantiated and solidified using the specialized model parameters to obtain the first prediction model. Instantiation refers to substituting the specialized model parameters θ1 into the neural network structure of the meta-learning base model, fixing the parameters without further adjustment, and forming a dedicated model that can be directly used for prediction. The specialized model parameters θ1 are written into the neural network of the meta-learning base model, which includes two convolutional layers and three fully connected layers. The parameters of each layer are solidified to generate a first prediction model that is only adapted to the target reservoir. For example, the specialized parameters θ1 are solidified into the MAML base network to obtain the first prediction model M1, which can directly input multi-source geological data from Block A and output predicted permeability values.
[0038] The construction steps of the second prediction model include: Construct a multi-scale feature fusion processor, wherein the multi-scale feature fusion processor includes a multi-scale feature extraction channel and an adaptive fusion channel based on a gated attention mechanism; Based on the multi-scale feature fusion builder, sample fusion features are generated, wherein the sample fusion features are associated with a calibrated penetration rate; Based on the sample fusion features and the calibrated penetration rate, several heterogeneous base learner models are constructed and trained, and the prediction outputs of the trained base learner models are obtained and concatenated into a feature matrix. Based on the feature matrix, an adaptive ensemble weight decision layer based on the ridge regression model is constructed and trained; By combining the multi-scale feature fusion processor, several base learner models, and the feature matrix, a second prediction model based on ensemble learning is constructed.
[0039] First, a multi-scale feature fusion unit is constructed, comprising a multi-scale feature extraction channel and an adaptive fusion channel based on a gated attention mechanism. The multi-scale feature fusion unit is a module used to extract microscopic and macroscopic features from multi-source geological data and adaptively fuse them. The multi-scale feature extraction channel includes microscopic and macroscopic feature channels. The gated attention mechanism automatically assigns fusion weights to features at different scales using a learnable weight vector, highlighting the contribution of key features. A dual-channel feature extraction structure is constructed: a 1DCNN channel with a 3×3 kernel size and a stride of 1, used for extracting local features; and a Bi-LSTM channel with a hidden layer dimension of 64 and 10 iteration steps, used for extracting global features. A gated fusion unit is designed, generating a weight vector through a fully connected layer, and then weighted summing the dual-channel features to obtain the fused features.
[0040] For example, a multi-scale feature fusion F1 is constructed for block A: the 1DCNN channel extracts the AC and DEN features of the local depth segment of well No. 4 in the range of 1500-1510m, and the Bi-LSTM channel extracts the global structural features of the seismic gathers of this block; the gated unit generates a weight vector [0.6, 0.4], which weights and fuses the local and global features to highlight the contribution of local lithological features.
[0041] Secondly, based on the multi-scale feature fusion builder, sample fusion features are generated, wherein the sample fusion features are associated with calibrated permeability. Sample fusion features refer to the feature vectors output by the multi-scale feature fusion builder that integrate information from different scales of multi-source data. Each feature vector is associated with a corresponding ground truth permeability value and is used to train the base learner. Valid samples from the target reservoir are input into the multi-scale feature fusion builder, which outputs a fusion feature vector for each sample and binds it to the ground truth permeability value from the core experiment for that sample, forming a training dataset. For example, multi-source data from 158 valid samples in Block A are input into the fusion builder F1, which outputs 158 128-dimensional fusion feature vectors, each bound to a corresponding ground truth permeability value, forming a training set T1.
[0042] Further, based on the sample fusion features and the calibrated penetration rate, several heterogeneous base learner models are constructed and trained, and the prediction outputs of the trained base learner models are obtained and concatenated into a feature matrix. Heterogeneous base learners refer to different types of machine learning models. To avoid the limitations of a single model, this embodiment selects Extremely Random Tree (ET), Lightweight Gradient Boosting Machine (LGBM), and Deep Residual Network (ResNet). The feature matrix is a matrix formed by concatenating the prediction outputs of multiple base learners according to the sample dimension, and serves as the input to the subsequent meta-learner. The three heterogeneous base learners, ET, LGB, and ResNet, are initialized separately and trained independently using the training set T1. After training, the fusion features of T1 are input into each base learner to obtain three predicted values for each sample, which are then concatenated into a feature matrix X_meta in the format of sample row × base learner column. For example, three base learners, ET, LGB, and ResNet, are trained using training set T1: the ET model outputs a predicted value of 65mD for sample 1; the LGB model outputs 72mD; and the ResNet model outputs 68mD. The three predicted values of 158 samples are concatenated to obtain the feature matrix X_meta (158 rows × 3 columns), where the data in the first row is [65, 72, 68].
[0043] Next, based on the aforementioned feature matrix, an adaptive ensemble weight decision layer based on the ridge regression model is constructed and trained. The adaptive ensemble weight decision layer refers to a meta-learner built based on the ridge regression model, used to learn the optimal combination weights of the prediction results from the base learners, achieving accurate ensemble. The ridge regression model is a linear regression model with L2 regularization, which can avoid overfitting and learn robust linear combination weights. Using the feature matrix X_meta as input and the true penetration rate of the training set T1 as the label, the ridge regression model is trained; the regularization parameter λ is optimized through K-fold cross-validation, outputting the optimal combination weight vector w. For example, inputting the feature matrix X_meta (158×3) and the true penetration rate into the ridge regression model, and optimizing λ=0.1 through K=5-fold cross-validation, the optimal weight vector w=[0.35,0.45,0.2] is obtained, meaning the LGB model has the highest prediction weights, and the ResNet model has the lowest.
[0044] Finally, the second prediction model based on ensemble learning is constructed by combining the multi-scale feature fusion engine, several base learner models, and the feature matrix. The multi-scale feature fusion engine, the three trained heterogeneous base learners, and the ridge regression weight decision layer are connected in a sequence of input, feature fusion, base learner prediction, and weight integration to form the complete second prediction model. For example, the structure of the second prediction model M2 is as follows: input multi-source geological data; fusion engine F1 extracts multi-scale features; ET, LGB, and ResNet output permeability predictions in parallel; the ridge regression model uses weights w=[0.35,0.45,0.2] to weight and sum to obtain the ensemble permeability prediction.
[0045] Based on this, the multi-source geological data is simultaneously input into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first predicted value and a second predicted value. Parallel input is used, with the same batch of multi-source geological data simultaneously input into the first prediction model M1 and the second prediction model M2. The two models operate independently, outputting the first predicted value P_meta and the second predicted value P_ensemble, respectively. For example, multi-source data from well No. 4 at a depth of 1550m is simultaneously input into M1 and M2: M1 outputs the first predicted value P_meta = 110mD; M2 outputs the second predicted value P_ensemble = 105mD, completing the acquisition of dual predicted values.
[0046] In this embodiment of the invention, abnormal data is eliminated through effective sample screening to ensure the reliability of the input model and avoid prediction bias caused by invalid data. Based on the dynamic updating of the dual models according to the sample size, the meta-learning model quickly adapts in the small sample stage, while the integrated model deeply mines patterns in the big data stage, overcoming the limitation of a single model adapting to full-cycle data. The meta-learning model is trained quickly with small samples, and the integrated model improves accuracy through multi-scale feature fusion and heterogeneous integration. The parallel input design synchronously outputs dual predicted values, providing efficient and high-quality basic data for subsequent dynamic fusion. The meta-learning base model is trained through multiple tasks; the first model can quickly adapt to the target reservoir, and the heterogeneous integration design of the second model reduces the risk of overfitting. Both models have strong generalization capabilities.
[0047] S300: Combined with a preset period, obtain the multi-dimensional data features of the multi-source geological data, and combined with the switching decision-maker, dynamically generate model fusion coefficients, wherein the model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task; The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
[0048] In this embodiment of the invention, multidimensional data features of the multi-source geological data are acquired in conjunction with a preset period. A switching decision-maker is then used to dynamically generate model fusion coefficients, which quantify the credible weight of the second prediction model in the current prediction task. The predicted values output by the dual models in parallel need to be dynamically assigned credible weights based on real-time data conditions: the meta-learning model is more robust in small sample stages, while the ensemble learning model has higher accuracy in large data stages. However, the sufficiency of data and the reliability of the model cannot be determined by a single indicator. Simultaneously, to avoid jumps in prediction results due to model switching, the fusion coefficients need to be continuously adjusted according to data conditions. Therefore, it is necessary to first comprehensively characterize the data state and model adaptability through multidimensional indicators, and then rely on a well-trained switching decision-maker to generate scientific fusion coefficients, achieving accurate mapping and continuous transition between data conditions and model weights.
[0049] Step S300 in the method provided in this embodiment of the invention includes: Among them, the multidimensional data features of the multi-source geological data are obtained by combining a preset period, including: Extract the real-time data volume of the multi-source geological data and calculate the data sufficiency index; Using the preset period as the acquisition time window, the multi-source geological data is extracted with a priority of near-time series data. By combining the aforementioned multi-source geological data with recent data extraction results, the entropy of new data value is calculated and obtained. Based on the recent data extraction results, the first prediction model, and the second prediction model, the prediction divergence and feature importance drift rate are analyzed and obtained. The data sufficiency index, the prediction divergence, the new data value entropy, and the feature importance drift rate are combined and output as the multidimensional data features.
[0050] The data sufficiency index is obtained by calculating the ratio of the effective sample size to a preset sample size threshold in multi-source geological data, and is used to quantify the sufficiency of multi-source geological data. The prediction divergence is obtained by calculating the variance of the prediction results of the first prediction model and the second prediction model on recent data, and is used to quantify the degree of consistency of the prediction results. The new data value entropy is obtained by calculating the difference between the feature distributions of recent data and historical data, and is used to quantify the novelty of information brought by the new data. The feature importance drift rate is obtained by calculating the rate of change of the feature importance of the first prediction model and the second prediction model between recent data and historical data, and is used to quantify the degree of change of key geological factors affecting the prediction results.
[0051] First, the real-time data volume of the multi-source geological data is extracted, and a data sufficiency index is calculated. The data sufficiency index is obtained by calculating the ratio of the effective sample size to a preset sample size threshold in the multi-source geological data, and is used to quantify the sufficiency of the multi-source geological data. The real-time data volume refers to the total number of effective samples that have been filtered at the current moment. The data sufficiency index I_d is the core indicator for quantifying data sufficiency. The data sufficiency index = real-time effective sample size / preset sample size threshold, with a value range of [0,1]. The closer to 1, the more sufficiency the data. The preset sample size threshold is a benchmark value set based on the minimum effective training requirements of the ensemble learning model, and must cover the minimum data volume for multi-scale feature mining and base learner training. The current effective sample size N_current and the preset sample size threshold N_threshold are counted in real time. The data sufficiency index is calculated using the formula I_d = N_current / N_threshold. When N_current ≥ N_threshold, the index is fixed at 1, representing completely sufficiency of data. For example, the current effective sample size of block A is N_current=158, and the preset sample size threshold is N_threshold=100. Substituting into the formula, we get I_d=158 / 100=1.58>1, so we take 1, which means that the data has fully met the training and prediction requirements of the ensemble learning model.
[0052] Secondly, using the preset period as the data collection time window, the multi-source geological data is extracted with a near-time-series priority. The preset period is used to define the time window for recent data and must fit the data collection rhythm of the oil and gas field, such as adding a batch of data weekly / monthly, to ensure that only the latest data reflects the current data status. Near-time-series priority means filtering by data collection time in reverse chronological order, prioritizing the retention of the most recent valid samples to avoid outdated data interfering with the judgment of current data conditions. The preset period T is set to 7 days, meaning that recent data consists of valid samples collected and filtered within the last 7 days. New valid samples added within the last 7 days are extracted from the database to form the recent dataset D_recent. If the number of new samples added within the last 7 days is insufficient, the period is extended to 14 days to ensure that the recent data has statistical representativeness.
[0053] For example, within the past 7 days, Block A completed the remaining data acquisition and screening of Well No. 4, adding 50 valid samples. These 50 samples were extracted as the recent dataset D_recent, covering the GR (75-115 API), AC (240-330 μs / m), and DEN (2.3-2.6 g / cm³) of the 1600-1800m reservoir section of Well No. 4. 3 The data characteristics, including multiple sources, are consistent with the logging / well logging data of well No. 4 in S200.
[0054] Furthermore, by combining the multi-source geological data with recent data extraction results, the new data value entropy is calculated. This new data value entropy is obtained by calculating the difference between the feature distributions of recent and historical data, and is used to quantify the novelty of the information brought by the new data. The new data value entropy E_n is an indicator that quantifies the novelty of the new data information. It is obtained by calculating the difference between the feature distributions of recent and historical data, and its value ranges from [0, +∞). A higher entropy value indicates more unknown information brought by the new data and a greater change in data distribution. The feature distribution difference measure uses KL divergence (relative entropy) to characterize the difference in probability distributions of the two datasets on core features such as GR, AC, and porosity. Historical data D_histor refers to valid samples collected and screened before a preset period. Extract core features (GR, AC, DEN, CNL, RT) from D_recent and D_history; calculate the probability distribution of each core feature in the two datasets respectively. For example, divide GR into intervals of 50-80API, 80-110API, and 110-140API, and count the sample proportion of each interval; calculate the distribution difference of each feature according to the KL divergence formula, and take the average value as the new data value entropy E_n.
[0055] For example, after extracting the core features of D_recent and D_history, the KL divergence of GR is calculated to be 0.32, the KL divergence of AC is 0.28, and the average KL divergence of other features is 0.40. Finally, E_n = (0.32 + 0.28 + 0.40 × 3) / 5 = 0.36, indicating that the new data brings a certain degree of novelty, but there is no drastic change in the distribution of historical data, which is consistent with the characteristic continuity of the same block of reservoir.
[0056] Then, based on the recent data extraction results, the first prediction model, and the second prediction model, the prediction divergence and feature importance drift rate are analyzed and obtained. The prediction divergence is obtained by calculating the variance of the prediction results of the first and second prediction models on recent data, used to quantify the consistency of the prediction results. The feature importance drift rate is obtained by calculating the rate of change of the feature importance of the first and second prediction models between recent and historical data, used to quantify the degree of change of key geological factors affecting the prediction results. The prediction divergence V_p is an indicator that quantifies the consistency of predictions between the first prediction model M1 and the second prediction model M2, obtained by calculating the variance of the prediction results of the two models on recent data, with a value range of [0, +∞). The smaller the variance, the more consistent the judgments of the two models, and the smaller the divergence. The feature importance drift rate ΔF is an indicator that quantifies the degree of change of key geological factors affecting the prediction results, analyzed by the SHAP value to measure the rate of change of the contribution of core features to the predictions of the two models, with a value range of [0, 1]. The closer to 0, the more stable the feature importance, and the smaller the drift. Calculate the prediction divergence: Input the recent dataset D_recent into M1 and M2 constructed by S200, respectively, to obtain two sets of predicted values P_meta_recent and P_ensemble_recent; calculate the overall variance of the two sets of predicted values according to the variance formula, i.e., V_p=Var(P_meta_recent-P_ensemble_recent). Calculate the feature importance drift rate: Analyze the contribution of core features in D_history to the prediction of M1 and M2 using SHAP values (e.g., porosity contributes 0.45 to M2). Similarly, analyze the contribution of core features in D_recent (e.g., porosity contributes 0.42 to M2). Calculate the drift rate of each core feature according to the formula ΔF=|current contribution-historical contribution| / historical contribution, and take the maximum value as the final ΔF.
[0057] For example, the predicted divergence: M1 predicted a value of 120 mD for the 1650 m depth sample of Well 4 in D_recent, while M2 predicted a value of 118 mD. The difference between the two sets of predicted values is relatively small, and the calculated variance V_p = 0.09, indicating that the two models have high consistency in recent data and little divergence. Feature importance drift rate: SHAP value analysis shows that porosity contributes 0.45 to M2 in D_history and 0.42 in D_recent. The drift rates of other core features such as GR and AC are all <0.1. Therefore, ΔF = |0.42 - 0.45| / 0.45 ≈ 0.07, indicating that the importance of key geological factors has not changed significantly, consistent with the geological patterns of the sandstone reservoir in Block A.
[0058] Finally, the data sufficiency index, the prediction divergence, the new data value entropy, and the feature importance drift rate are combined and output as the multidimensional data features. The calculated data sufficiency index I_d, prediction divergence V_p, new data value entropy E_n, and feature importance drift rate ΔF are integrated into a multidimensional feature vector in the format [I_d, V_p, E_n, ΔF], which serves as the input data for the switching decision-maker. For example, the current multidimensional data feature vector for block A is [1.0, 0.09, 0.36, 0.07], indicating that the current state data is completely sufficient, the model divergence is small, the new data novelty is moderate, and the feature importance is stable.
[0059] Among them, the model fusion coefficients are dynamically generated by combining the switching decision-maker, including: Collect historical data and calculate historical multidimensional data features based on the historical data; Based on the historical data, the true value of historical penetration rate is extracted. Combined with the first prediction model and the second prediction model, the fusion coefficients of multiple differentiated enumeration models are verified post-hocly. The historical model fusion coefficient is obtained by finding the one with the smallest error between the posterior penetration rate prediction value and the true value of historical penetration rate. Using the historical multidimensional data features as training input and the corresponding historical model fusion coefficients as training labels, a training sample set is constructed. The gating neural network is trained based on the training sample set to obtain the switching decision-maker; The multidimensional data features are input into the switching decision-maker, which outputs the current model fusion coefficients.
[0060] First, historical data was collected, and historical multidimensional data features were calculated using this data. Historical multidimensional data features refer to the multidimensional feature vectors for each stage calculated based on historical data using the methods described above. Historical data from five similar reservoirs were collected, with each reservoir divided into 10 data stages in chronological order, corresponding to different sample sizes and data distributions, covering small samples, large datasets, and the entire data lifecycle. For each data stage, historical multidimensional data features were calculated, resulting in a total of 50 sets of historical feature vectors.
[0061] For example, historical data were collected from Block B, which contains fluvial sandstone reservoirs similar to Block A. This block was accumulated from an initial 60 samples to 300 samples, divided into 10 stages. Stage 6 had 156 valid samples, and the calculated historical multidimensional feature vector for this stage was [1.0, 0.12, 0.40, 0.08]. Similarly, the historical multidimensional feature vectors for the remaining four similar reservoirs were obtained.
[0062] Secondly, based on the historical data, the true historical permeability values are extracted. Combining the first and second prediction models, the fusion coefficients of multiple differentiated enumeration models are post-validated. The historical model fusion coefficient is the one with the smallest error between the posterior permeability prediction value and the true historical permeability value. The true historical permeability value refers to the core experimental permeability value corresponding to the valid samples at each stage in the historical data. The enumeration model fusion coefficient refers to 11 candidate coefficients enumerated at 0.1 intervals within the interval [0,1], α_candidate=0,0.1,0.2,...,1.0, used to test the prediction effect under different weights. Post-validation refers to backtesting with historical data to find the fusion coefficient that minimizes the prediction error, which is taken as the optimal coefficient for that historical stage, i.e., the historical fusion coefficient. Extract the penetration rate true value vector Y_true for each historical data stage; input the historical data of that stage into M1 and M2 to obtain the historical predicted values P_meta_hist and P_ensemble_hist; for each enumerated coefficient α_candidate, calculate the predicted value according to the formula P_pred=α_candidate×P_ensemble_hist+(1-α_candidate)×P_meta_hist; calculate the mean squared error (MSE) of P_pred and Y_true, and select the α_candidate with the smallest MSE as the historical fusion coefficient α_hist for that stage. For example, for the data of stage 6 of block B, there are 156 valid samples. When enumerating α_candidate=0.7, the MSE of P_pred and Y_true is 8.2, which is the smallest among all candidate coefficients. Therefore, the historical fusion coefficient α_hist for this stage is 0.7.
[0063] Furthermore, a training sample set is constructed using the historical multidimensional data features as training input and the corresponding historical model fusion coefficients as training labels. The historical multidimensional data feature vector of each historical data stage is used as the training input X_train, and the corresponding historical fusion coefficient α_hist is used as the training label y_train. Data from all historical stages are integrated to construct the training sample set. For example, the feature vector [1.0, 0.12, 0.40, 0.08] of stage 6 in block B is used as a row of X_train, and the corresponding α_hist=0.7 is used as the corresponding element of y_train. Similarly, 50 sets of data from 10 stages across 5 reservoirs are integrated to form a complete training sample set.
[0064] Subsequently, the gated neural network is trained based on the training sample set to obtain the switching decision-maker. The gated neural network is a neural network model used to map multidimensional data features to fusion coefficients, consisting of an input layer (4 neurons, corresponding to 4 multidimensional features), hidden layers (2 layers, 32 neurons per layer, ReLU activation function), and an output layer (1 neuron, Sigmoid activation function, ensuring output α∈[0,1]). The switching decision-maker refers to the gated neural network after training, which has the decision-making ability to input multidimensional data features and output fusion coefficients. The training sample set is normalized, scaling each feature of X_train to the [0,1] interval; the training set and validation set are divided in a 7:3 ratio, with 50 training rounds and a learning rate of 0.001; the gated neural network is trained using MSE as the loss function, and training stops when the validation set loss converges, such as when the loss fluctuation is <0.001 for 5 consecutive rounds, resulting in the switching decision-maker D_switch. For example, after training, the MSE on the validation set stabilizes at 0.002, and switching the decision maker D_switch can accurately learn the mapping relationship between data conditions and model weights.
[0065] Finally, the multidimensional data features are input into the switching decision unit, which outputs the current model fusion coefficients. The current multidimensional data feature vector is normalized to match the normalization standard of the training sample set. The normalized feature vector is then input into the switching decision unit D_switch. The output layer outputs the current model fusion coefficient α_current through the Sigmoid function. This coefficient quantifies the confidence weight of the second prediction model in the current task. For example, inputting the feature vector [1.0, 0.09, 0.36, 0.07] into D_switch results in an output α_current = 0.75, representing that the confidence weight of the second prediction model is 75% and the confidence weight of the first prediction model is 25% in the current stage, consistent with the fusion coefficient pattern in stages with similar sample sizes in the past.
[0066] In this embodiment of the invention, the current data conditions are comprehensively evaluated from four dimensions of feature indicators, including data sufficiency, model consistency, information novelty, and pattern stability, to avoid decision bias caused by a single indicator. The switching decision generator is trained based on historical data of similar reservoirs, which can accurately learn the mapping pattern between data conditions and model weights. The fusion coefficient is dynamically adjusted with the continuous changes of multidimensional data features, rather than abruptly, laying a core foundation for the continuous and smooth transition of the final prediction value. Parameters such as preset period and enumeration coefficient interval can be adjusted according to the rhythm of on-site data collection. When new data is added, the multidimensional features are updated synchronously, and the fusion coefficient is automatically adjusted to ensure that the model weights always match the data conditions without manual intervention.
[0067] S400: Based on the model fusion coefficient, the first predicted value and the second predicted value are linearly weighted and fused to generate the final permeability prediction value of the target reservoir.
[0068] In this embodiment of the invention, the first and second predicted values are linearly weighted and fused according to the model fusion coefficients to generate the final permeability prediction value of the target reservoir. While weighting the first and second predicted values solely through the model fusion coefficients can dynamically adjust the model weights to adapt to data conditions, in extreme cases, it may lead to prediction results deviating from geological common sense. Furthermore, oil and gas reservoir permeability is constrained by rock physics laws, requiring the introduction of geological empirical formulas as a safety baseline. The constraint strength of these geological empirical formulas should be dynamically adjusted with the amount of data; strong constraints are needed to ensure rationality when data is scarce, while weak constraints are needed to unleash the potential of the data-driven model when data is abundant. Therefore, by quantifying the constraint strength through an exponential decay function and linearly weighting and fusing the physical constraint safety term with the dual-model predicted values, the geological rationality of the prediction results can be ensured, while maintaining the smooth transition characteristics brought by the fusion coefficients, ultimately generating an accurate and reliable final predicted value.
[0069] Step S400 in the method provided in this embodiment of the invention includes: Obtain prior geological empirical formulas and define a monotonically decreasing exponential decay function based on a preset decay constant; The cumulative data volume of the multi-source geological data is analyzed and input into the exponential decay function to obtain the constraint strength coefficient corresponding to the geological empirical formula; The product of the output of the geological empirical formula and the constraint strength coefficient is defined as the physical constraint safety term. Combined with the model fusion coefficient, the physical constraint safety term, the first predicted value, and the second predicted value are linearly weighted and fused to obtain the final permeability prediction value. The fusion coefficient corresponding to the physical constraint safety term is 1.
[0070] First, a priori geological empirical formula is obtained, and a monotonically decreasing exponential decay function is defined based on a preset decay constant. The geological empirical formula is an empirical formula for permeability calculation based on rock physics laws and numerous field experiments. It reflects the fundamental correlation between core reservoir parameters and permeability, providing a physically reasonable constraint for prediction. The decay constant is a parameter used to adjust the decay rate of the constraint strength. It needs to be calibrated according to the data accumulation rhythm and geological stability of the target reservoir, and its value is positive. The monotonically decreasing exponential decay function is a function that monotonically decreases as the accumulated data volume increases. It is used to map the data volume to a constraint strength coefficient, ensuring that the more abundant the data, the weaker the constraint of the geological empirical formula. Combining the lithological characteristics of the target reservoir, the regionally calibrated Coates model is selected as the geological empirical formula. Based on the data accumulation rate of the block, such as an average of 5-8 new valid samples per day, the preset decay constant N0=100 is used. The exponential decay function is defined as: β=exp(-N / N0), where N is the accumulated data volume of multi-source geological data, β is the constraint strength coefficient, and the value of β ranges from (0,1).
[0071] For example, the regionally modified Coates model for the sandstone reservoir in Block A is as follows: The values of C=1800, m=2.0, and n=2.9 were determined based on the core experimental data of the block. φ is the neutron porosity, which is calculated from the CNL logging curve. The exponential decay function is β=exp(-N / 180), where N is the cumulative effective sample size.
[0072] Secondly, the cumulative data volume of the acquired multi-source geological data is analyzed and input into the exponential decay function to obtain the constraint strength coefficient corresponding to the geological empirical formula. The cumulative data volume refers to the total number of valid samples accumulated from data acquisition in S100 to the final selection in S200, i.e., the current stable number of valid samples, reflecting the degree of data accumulation. The constraint strength coefficient β is a parameter that quantifies the constraint strength of the geological empirical formula; a larger β indicates a stronger constraint, and a smaller β indicates a weaker constraint. The cumulative valid sample volume N is extracted from the database; N is substituted into the preset exponential decay function to calculate the constraint strength coefficient β.
[0073] For example, the current cumulative effective sample size of block A is N=158. Substituting into the decay function β=exp(-158 / 100)=exp(-1.58)≈0.406, the constraint strength coefficient of the geological empirical formula is 0.406, indicating that the constraint is relatively weak when the data is sufficient, and only the basic correction capability is retained.
[0074] Finally, the product of the output of the geological empirical formula and the constraint strength coefficient is defined as the physical constraint safety term. Combined with the model fusion coefficient, the physical constraint safety term, the first predicted value, and the second predicted value are linearly weighted and fused to obtain the final permeability prediction value. The fusion coefficient corresponding to the physical constraint safety term is 1. The physical constraint safety term is the product of the output value of the geological empirical formula and the constraint strength coefficient. It reflects the physical laws of rocks and can dynamically adjust the constraint strength according to the amount of data, avoiding excessive intervention in model prediction. Linear weighted fusion refers to superimposing the meta-learning prediction contribution, the ensemble learning prediction contribution, and the physical constraint safety term according to a fixed weight ratio to generate the final prediction value, ensuring that the influence of each contribution is continuous and quantifiable. Based on the core parameters of the target reservoir, the reference value D is calculated by substituting them into the geological empirical formula; the physical constraint safety term is defined as β×D, and the fusion coefficient is fixed at 1, that is, it directly participates in the fusion in the form of β×D; the final permeability prediction value is calculated according to the formula P=(1-α)×P_meta+α×P_ensemble+β×D, where (1-α) is the weight of the meta-learning model, α is the weight of the ensemble learning model, and β is the weight of the physical constraint.
[0075] For example, the reference value D is calculated using the geological empirical formula: A sample from well No. 4 at a depth of 1650m is selected, corresponding to the same depth where P_meta = 110mD and P_ensemble = 105mD. The porosity φ is calculated as 18% using the CNL logging curve. Substituting this into the Coates model, D = 105mD is obtained. The physical constraint safety term is calculated: Given a constraint strength coefficient β ≈ 0.406, the physical constraint safety term = β × D = 0.406 × 105 ≈ 42.63mD. The final predicted value is calculated using linear weighted fusion: Substituting into the formula P = (1-α) × P_meta + α × P_ensemble + β × D, where α = 0.75, we get P = (1-0.75) × 110 + 0.75 × 105 + 42.63 ≈ 149mD. Therefore, the final predicted permeability value is 149mD.
[0076] In this embodiment of the invention, geological empirical formulas are used to provide a physical constraint baseline, preventing the model prediction results from deviating from the physical laws of reservoir rocks. Even if the data is abnormal, the constraint terms can be pulled back to a reasonable range. The exponential decay function allows the constraint strength to be dynamically adjusted with the amount of data, with strong constraints when data is scarce and weak constraints when data is abundant, balancing the needs of common sense protection and data mining. The continuous change of the fusion coefficient α and the gradual change of the constraint strength coefficient β together ensure that the final prediction value is continuously adjusted with data conditions, without abrupt changes, ensuring the continuity of engineering decisions. Parameters such as geological empirical formulas and decay constants can be flexibly calibrated according to the target reservoir lithology and data rhythm. Both the constraint strength coefficient and the fusion coefficient are automatically calculated without the need for manual weight adjustment, which can reduce operation and maintenance costs and improve the practicality and reliability of the prediction method.
[0077] Through the specific implementation methods described above, the embodiments of the present invention achieve the following technical effects: This invention provides an intelligent prediction method and system for oil and gas reservoir permeability. First, through real-time acquisition and data volume discrimination mechanisms, it enables on-demand construction and iterative updates of meta-learning and ensemble learning models, achieving full-cycle adaptive prediction capabilities from small sample startup to big data optimization. Then, through parallel prediction and feature analysis of dual models, it obtains complementary and robust first and second prediction values, achieving reliable output and uncertainty quantification of multi-perspective prediction results. Next, by dynamically generating fusion coefficients through switching decision units, it achieves intelligent weighting of the dual prediction results based on data characteristics and cycle stages, achieving online dynamic optimization and accuracy improvement of the prediction strategy. Finally, through linear weighted fusion output, it provides high-precision, highly robust continuous permeability prediction results for oilfield exploration and development, effectively unifying intelligent decision support and engineering application feasibility.
[0078] Example 2, as Figure 2 As shown, this invention provides an intelligent prediction system for oil and gas reservoir permeability, the system comprising: The multi-source data acquisition module 11 is used to acquire multi-source geological data of the target reservoir in real time, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions; The dual-model parallel prediction module 12 is used to simultaneously input the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. The fusion coefficient dynamic generation module 13 is used to obtain multi-dimensional data features of the multi-source geological data in combination with a preset period, and dynamically generate model fusion coefficients in combination with the switching decision-maker. The model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task. The linear weighted fusion output module 14 is used to perform linear weighted fusion of the first predicted value and the second predicted value according to the model fusion coefficient to generate the final permeability prediction value of the target reservoir. The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
[0079] In one embodiment, the multi-source data acquisition module 11 is further configured to: The multi-source geological data includes well logging curves, well logging curves, and seismic gathers.
[0080] In one embodiment, the dual-model parallel prediction module 12 is further configured to: Specifically, the process of simultaneously inputting the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain first and second predicted values includes: Analyze the multi-source geological data, extract valid sample data, and statistically obtain the number of valid samples; The effective sample size is iteratively determined by combining the preset trigger sample size; If the effective sample size is less than the trigger sample size, then the first prediction model is constructed and iteratively updated based on meta-learning according to the preset first sample increment trigger value. If the effective sample size is greater than or equal to the trigger sample size, then based on the preset second sample increment trigger value, the first prediction model is synchronously constructed and iteratively updated based on meta-learning, and the second prediction model is constructed and iteratively updated based on ensemble learning.
[0081] The construction steps of the first prediction model include: Load a pre-built meta-learning base model, wherein the meta-learning base model is constructed through multiple diverse few-sample tasks based on geological constraints; Based on the multi-source geological data, a small amount of calibration data for the target reservoir is extracted and calibrated, wherein the small amount of calibration data is calibrated based on permeability; Based on the limited calibration data, the meta-learning basic model is rapidly and adaptively trained, and the specialized model parameters of the target reservoir are output according to the training results. The neural network model associated with the meta-learning base model is instantiated using the specialized model parameters to obtain the first prediction model.
[0082] The construction steps of the second prediction model include: Construct a multi-scale feature fusion processor, wherein the multi-scale feature fusion processor includes a multi-scale feature extraction channel and an adaptive fusion channel based on a gated attention mechanism; Based on the multi-scale feature fusion builder, sample fusion features are generated, wherein the sample fusion features are associated with a calibrated penetration rate; Based on the sample fusion features and the calibrated penetration rate, several heterogeneous base learner models are constructed and trained, and the prediction outputs of the trained base learner models are obtained and concatenated into a feature matrix. Based on the feature matrix, an adaptive ensemble weight decision layer based on the ridge regression model is constructed and trained; By combining the multi-scale feature fusion processor, several base learner models, and the feature matrix, a second prediction model based on ensemble learning is constructed.
[0083] In one embodiment, the fusion coefficient dynamic generation module 13 is further configured to: Among them, the multidimensional data features of the multi-source geological data are obtained by combining a preset period, including: Extract the real-time data volume of the multi-source geological data and calculate the data sufficiency index; Using the preset period as the acquisition time window, the multi-source geological data is extracted with a priority of near-time series data. By combining the aforementioned multi-source geological data with recent data extraction results, the entropy of new data value is calculated and obtained. Based on the recent data extraction results, the first prediction model, and the second prediction model, the prediction divergence and feature importance drift rate are analyzed and obtained. The data sufficiency index, the prediction divergence, the new data value entropy, and the feature importance drift rate are combined and output as the multidimensional data features.
[0084] The data sufficiency index is obtained by calculating the ratio of the effective sample size to a preset sample size threshold in multi-source geological data, and is used to quantify the sufficiency of multi-source geological data. The prediction divergence is obtained by calculating the variance of the prediction results of the first prediction model and the second prediction model on recent data, and is used to quantify the degree of consistency of the prediction results. The new data value entropy is obtained by calculating the difference between the feature distributions of recent data and historical data, and is used to quantify the novelty of information brought by the new data. The feature importance drift rate is obtained by calculating the rate of change of the feature importance of the first prediction model and the second prediction model between recent data and historical data, and is used to quantify the degree of change of key geological factors affecting the prediction results.
[0085] Among them, the model fusion coefficients are dynamically generated by combining the switching decision-maker, including: Collect historical data and calculate historical multidimensional data features based on the historical data; Based on the historical data, the true value of historical penetration rate is extracted. Combined with the first prediction model and the second prediction model, the fusion coefficients of multiple differentiated enumeration models are verified post-hocly. The historical model fusion coefficient is obtained by finding the one with the smallest error between the posterior penetration rate prediction value and the true value of historical penetration rate. Using the historical multidimensional data features as training input and the corresponding historical model fusion coefficients as training labels, a training sample set is constructed. The gating neural network is trained based on the training sample set to obtain the switching decision-maker; The multidimensional data features are input into the switching decision-maker, which outputs the current model fusion coefficients.
[0086] In one embodiment, the linear weighted fusion output module 14 is further configured to: Obtain prior geological empirical formulas and define a monotonically decreasing exponential decay function based on a preset decay constant; The cumulative data volume of the multi-source geological data is analyzed and input into the exponential decay function to obtain the constraint strength coefficient corresponding to the geological empirical formula; The product of the output of the geological empirical formula and the constraint strength coefficient is defined as the physical constraint safety term. Combined with the model fusion coefficient, the physical constraint safety term, the first predicted value, and the second predicted value are linearly weighted and fused to obtain the final permeability prediction value. The fusion coefficient corresponding to the physical constraint safety term is 1.
[0087] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0089] This specification and accompanying drawings are merely illustrative examples of the invention and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Therefore, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is intended to include these modifications and modifications.
Claims
1. A smart prediction method for oil and gas reservoir permeability, characterized in that, include: Real-time acquisition of multi-source geological data of the target reservoir, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions; The multi-source geological data are simultaneously input into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. Combined with a preset period, multidimensional data features of the multi-source geological data are obtained. Combined with a switching decision unit, model fusion coefficients are dynamically generated. The model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task. Based on the model fusion coefficient, the first predicted value and the second predicted value are linearly weighted and fused to generate the final permeability prediction value of the target reservoir. The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
2. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, Real-time acquisition of multi-source geological data of the target reservoir, wherein the multi-source geological data includes well logging curves, well logging curves, and seismic gathers.
3. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, The steps for constructing the first prediction model include: Load a pre-built meta-learning base model, wherein the meta-learning base model is constructed through multiple diverse few-sample tasks based on geological constraints; Based on the multi-source geological data, a small amount of calibration data for the target reservoir is extracted and calibrated, wherein the small amount of calibration data is calibrated based on permeability; Based on the limited calibration data, the meta-learning basic model is rapidly and adaptively trained, and the specialized model parameters of the target reservoir are output according to the training results. The neural network model associated with the meta-learning base model is instantiated using the specialized model parameters to obtain the first prediction model.
4. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, The construction steps of the second prediction model include: Construct a multi-scale feature fusion processor, wherein the multi-scale feature fusion processor includes a multi-scale feature extraction channel and an adaptive fusion channel based on a gated attention mechanism; Based on the multi-scale feature fusion processor, sample fusion features are generated, wherein the sample fusion features are associated with a calibrated penetration rate; Based on the sample fusion features and the calibrated penetration rate, several heterogeneous base learner models are constructed and trained, and the prediction outputs of the trained base learner models are obtained and concatenated into a feature matrix. Based on the feature matrix, an adaptive ensemble weight decision layer based on the ridge regression model is constructed and trained; By combining the multi-scale feature fusion processor, several base learner models, and the feature matrix, a second prediction model based on ensemble learning is constructed.
5. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, Combined with a preset period, the multidimensional data features of the multi-source geological data are obtained, including: Extract the real-time data volume of the multi-source geological data and calculate the data sufficiency index; Using the preset period as the acquisition time window, the multi-source geological data is extracted with a priority of near-time series data. By combining the aforementioned multi-source geological data with recent data extraction results, the entropy of new data value is calculated and obtained. Based on the recent data extraction results, the first prediction model, and the second prediction model, the prediction divergence and feature importance drift rate are analyzed and obtained. The data sufficiency index, the prediction divergence, the new data value entropy, and the feature importance drift rate are combined and output as the multidimensional data features.
6. The intelligent prediction method for oil and gas reservoir permeability as described in claim 5, characterized in that, include: The data sufficiency index is obtained by calculating the ratio of the effective sample size to a preset sample size threshold in multi-source geological data, and is used to quantify the sufficiency of multi-source geological data. The prediction divergence is obtained by calculating the variance of the prediction results of the first prediction model and the second prediction model on recent data, and is used to quantify the degree of consistency of the prediction results. The new data value entropy is obtained by calculating the difference between the feature distributions of recent data and historical data, and is used to quantify the novelty of information brought by the new data. The feature importance drift rate is obtained by calculating the rate of change of the feature importance of the first prediction model and the second prediction model between recent data and historical data, and is used to quantify the degree of change of key geological factors affecting the prediction results.
7. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, By combining the switching decision-maker, the model fusion coefficients are dynamically generated, including: Collect historical data and calculate historical multidimensional data features based on the historical data; Based on the historical data, the true value of historical penetration rate is extracted. Combined with the first prediction model and the second prediction model, the fusion coefficients of multiple differentiated enumeration models are verified post-hocly. The historical model fusion coefficient is obtained by finding the one with the smallest error between the posterior penetration rate prediction value and the true value of historical penetration rate. Using the historical multidimensional data features as training input and the corresponding historical model fusion coefficients as training labels, a training sample set is constructed. The gating neural network is trained based on the training sample set to obtain the switching decision-maker; The multidimensional data features are input into the switching decision-maker, which outputs the current model fusion coefficients.
8. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, The multi-source geological data is simultaneously input into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain first and second predicted values. Prior to this, the process includes: Analyze the multi-source geological data, extract valid sample data, and statistically obtain the number of valid samples; The effective sample size is iteratively determined by combining the preset trigger sample size; If the effective sample size is less than the trigger sample size, then the first prediction model is constructed and iteratively updated based on meta-learning according to the preset first sample increment trigger value. If the effective sample size is greater than or equal to the trigger sample size, then based on the preset second sample increment trigger value, the first prediction model is synchronously constructed and iteratively updated based on meta-learning, and the second prediction model is constructed and iteratively updated based on ensemble learning.
9. The intelligent prediction method for oil and gas reservoir permeability as described in claim 1, characterized in that, Based on the model fusion coefficients, the first predicted value and the second predicted value are linearly weighted and fused to generate the final permeability prediction value of the target reservoir, and the method further includes: Obtain prior geological empirical formulas and define a monotonically decreasing exponential decay function based on a preset decay constant; The cumulative data volume of the multi-source geological data is analyzed and input into the exponential decay function to obtain the constraint strength coefficient corresponding to the geological empirical formula; The product of the output of the geological empirical formula and the constraint strength coefficient is defined as the physical constraint safety term. Combined with the model fusion coefficient, the physical constraint safety term, the first predicted value, and the second predicted value are linearly weighted and fused to obtain the final permeability prediction value. The fusion coefficient corresponding to the physical constraint safety term is 1.
10. An intelligent prediction system for oil and gas reservoir permeability, characterized in that, The system is used to implement the intelligent prediction method for oil and gas reservoir permeability according to any one of claims 1-9, the system comprising: A multi-source data acquisition module is used to acquire multi-source geological data of the target reservoir in real time, wherein the amount of multi-source geological data increases cumulatively with the number of data acquisitions; The dual-model parallel prediction module is used to simultaneously input the multi-source geological data into a first prediction model pre-built based on meta-learning and a second prediction model pre-built based on ensemble learning to obtain a first prediction value and a second prediction value. The fusion coefficient dynamic generation module is used to obtain multi-dimensional data features of the multi-source geological data in combination with a preset period, and dynamically generate model fusion coefficients in combination with the switching decision-maker. The model fusion coefficients are used to quantify the credibility weight of the second prediction model in the current prediction task. The linear weighted fusion output module is used to perform linear weighted fusion of the first predicted value and the second predicted value according to the model fusion coefficient to generate the final permeability prediction value of the target reservoir. The model fusion coefficient is continuously adjusted according to the data conditions within the preset period, so that the final penetration rate prediction value transitions smoothly and continuously as the data conditions change.
Citation Information
Patent Citations
Method for predicting permeability of shale oil reservoir based on artificial intelligence
CN118656622A
Intelligent reservoir logging evaluation method driven by double models
CN119807888A
Short-term photovoltaic power prediction method and system, computer equipment and medium
CN120893626A
Reservoir porosity and permeability prediction method based on dynamic committee integration model
CN120993519A
Intelligent prediction method for shale content logging curve of oil and gas reservoir
CN121072376A