Modeling method for reservoir prediction, and reservoir prediction method and system
By adopting a semi-supervised regression model with pseudo-label in petroleum reservoir prediction, using pseudo-label technology to expand the labeled samples, and combining clustering methods for variable selection, the problem of reduction in prediction accuracy caused by scarcity of labeled data is solved, and the accuracy and recognition accuracy of reservoir prediction are significantly improved.
Patent Information
- Application Number
- CN202510601575.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Due to the scarcity of labeled data, prior art has reduced prediction accuracy in petroleum reservoir prediction, especially in category identification with small sample sizes.
A semi-supervised regression model with pseudo-label is adopted, and a limited labeled sample is expanded into a fully labeled sample through pseudo-labeling technology. Variable selection is performed in combination with clustering methods, working data sets are constructed, and a semi-supervised regression model with pseudo-labeling is generated through a PI estimator.
It significantly improves the accuracy and robustness of reservoir prediction, especially in the case of scarce labeled data, improves the recognition accuracy of categories with small sample sizes, significantly better than conventional semi-supervised regression models and other classical regression methods.
Smart Images

Figure CN120214897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a modeling method for petroleum reservoir prediction, and also relates to a corresponding reservoir prediction method and a reservoir prediction system, belonging to the technical field of data processing. Background Art
[0002] Well logging data is a vertical depth profile obtained by well logging tools, which can reflect geological structure characteristics, such as lithology and reservoir type. These data are of great significance for improving geological modeling and enhancing the efficiency of petroleum exploration. However, due to limitations such as poor wellbore conditions, high labor and material costs, etc., the labeled data usually only accounts for a small part of the total data. This scarcity of labeled data significantly reduces the effectiveness of prediction models based on supervised learning, making reservoir prediction a challenging task.
[0003] To address this challenge, semi-supervised learning has received increasing attention in reservoir prediction in petroleum geological exploration. The core of semi-supervised learning lies in combining two types of data: (1) a relatively small or medium-sized "labeled" data set, which contains observations of the result variable and a set of covariates; (2) a much larger "unlabeled" data set, which only contains observations of the covariates. This characteristic makes semi-supervised learning particularly suitable for large data application scenarios where the result variable is difficult to obtain, but the covariates can be easily accessed.
[0004] Azriel et al. proposed a semi-supervised regression (SSR) model in their 2021 study, aiming to improve the least squares estimation (abbreviation: LSE) by transforming the regression problem into a mean estimation problem, and successfully demonstrated the potential of the semi-supervised regression model in petroleum reservoir prediction, providing an effective solution to the problem of scarce labeled data. The key to the semi-supervised regression model lies in utilizing the information in the unlabeled samples to provide additional optimization basis for the model, thereby achieving asymptotic uniform improvement in the least squares estimation. Specifically, by decomposing the regression problem into multiple independent mean estimation problems, the parameter estimation process is simplified, and at the same time, the unmodeled non-linear features that may exist in the data are utilized to further optimize the prediction performance of the semi-supervised regression model. Summary of the Invention
[0005] The primary technical problem to be solved by the present invention is to provide a modeling method for reservoir prediction to implement a semi-supervised regression model with pseudo-labels.
[0006] Another technical problem to be solved by the present invention is to provide a reservoir prediction method based on the semi-supervised regression model with pseudo-labels.
[0007] Another technical problem to be solved by the present invention is to provide a corresponding reservoir prediction system.
[0008] To achieve the above technical objectives, the present invention adopts the following technical solutions: According to the first aspect of the embodiments of the present invention, a modeling method for reservoir prediction is provided, including the following steps: Step 1: Prepare a data set, the data set contains Y observation samples, each sample consists of N seismic attributes, and the seismic attributes are numerical features characterizing logging events; Step 2: Use a clustering method to perform variable selection on the data set and screen out P key variables; Step 3: Based on the data set obtained in Step 1 and the P key variables screened in Step 2, construct a Y×P working data set; Step 4: On the working data set, use the pseudo-labeling technique to expand a limited number of labeled samples into fully labeled samples for training a semi-supervised regression model, and based on the training results of the PI estimator, generate a semi-supervised regression model with pseudo-labels; Wherein, Y, N, and P are all positive integers.
[0009] According to the second aspect of the embodiments of the present invention, a reservoir prediction method is provided, including the following steps: Input the data set to be predicted into the semi-supervised regression model with pseudo-labels constructed by the above modeling method; Use the semi-supervised regression model with pseudo-labels to perform reservoir prediction and output the prediction result.
[0010] According to the third aspect of the embodiments of the present invention, a reservoir prediction system is provided, including a processor and a memory, the processor and the memory are coupled; wherein, the memory is used to store a computer program; the processor runs the computer program stored in the memory to implement the above reservoir prediction method.
[0011] Compared with the prior art, the present invention effectively utilizes unlabeled data through the pseudo-labeling technique to expand labeled samples, significantly improving the accuracy of reservoir prediction. In the case of scarce labeled data, the present invention shows higher prediction effectiveness, especially improving the recognition accuracy of categories with fewer samples. Experimental results show that the present invention is superior to other methods in multiple evaluation metrics, including conventional semi-supervised regression models and other classical regression methods. Among them, the introduction of pseudo-labels significantly improves the prediction performance of the model in all cases, especially when the labeled data is insufficient, and the ability to identify rare geological features is particularly prominent. Therefore, the present invention effectively solves the problem of reduced prediction accuracy due to scarce labels, significantly improving the accuracy of reservoir modeling, oil and gas prediction, and logging data analysis, and is particularly suitable for actual scenarios in geological applications. Brief Description of the Drawings
[0012] Figure 1 Histogram of simulation data variables used for actual reservoir prediction; Figure 2 For Figure 1 Correlation matrix diagram of the shown simulation data; Figure 3A Schematic flow chart of the modeling method provided by the embodiment of the present invention; Figure 3B For Figure 3A Schematic flow chart of the pseudo-labeling technique used in; Figure 4 Schematic diagram of 7 variables obtained through variable selection in the first embodiment of the present invention; Figure 5 Schematic diagram of the original seismic horizons and well positions in the real data; Figure 6 Schematic diagram of the channels on the seismic section in the real data; Figure 7 Variable diagram of the working data set obtained through variable selection based on the real data in the first embodiment of the present invention; Figure 8 Scatter comparison diagram of the prediction results for well logging No. 1 in the real well logging data set using the modeling method provided by the embodiment of the present invention; Figure 9 Reservoir property heat map of the prediction results obtained using the modeling method provided by the embodiment of the present invention; Figure 10 Schematic diagram of 23 variables obtained through R-type clustering in the second embodiment of the present invention; Figure 11 In the second embodiment of the present invention, for Figure 10 Schematic diagram of 6 variables selected after principal component analysis clustering of the variables in. Detailed Description of the Invention
[0013] The technical content of the present invention will be described in detail below with reference to the drawings and specific embodiments.
[0014] The technical concept in the embodiments of the present invention is to expand the limited labeled samples into fully labeled samples through the pseudo-labeling technique, making full use of the information of unlabeled data, thereby improving the accuracy and robustness of reservoir prediction. The specific process includes: preparing an observation sample dataset containing seismic attributes, using a clustering method for variable selection to extract key variables, constructing a working dataset, expanding the labeled samples through the pseudo-labeling technique, training a semi-supervised regression model (SSR), and finally obtaining an optimized reservoir prediction model. This method significantly improves the accuracy of reservoir prediction in the case of scarce labeled data, is particularly suitable for categories with a small number of samples, and enhances the ability to identify rare geological features.
[0015] Here, the semi-supervised regression model itself will be briefly described first. The semi-supervised regression model is a method for performing regression analysis using limited labeled data and a large amount of unlabeled data. Considering only partial information, assuming that the number of labeled samples m is at least proportional to the number of unlabeled samples n, the semi-supervised regression model transforms the regression problem into a mean estimation problem for each parameter. Through a multiplicative step, the semi-supervised regression model provides asymptotically better parameter estimates than the standard least squares estimate (abbreviated as LSE), thus obtaining the best linear predictor. Specifically, the semi-supervised regression model utilizes the unmodeled non-linear information that may exist in the unlabeled data to optimize the model parameters by adjusting the regression coefficients. Taking the p-dimensional linear regression shown below as an example, the semi-supervised regression model simplifies the complex regression analysis process by decomposing the regression problem into p independent mean estimation problems and solving the estimation of each parameter separately, and provides more accurate prediction results in the asymptotic case.
[0016]
[0017] Among them, is the coefficient of the best linear predictor, defined as .
[0018] The remaining term satisfies .
[0019] Then, the p-dimensional estimation process is simplified into p independent simple regression problems, and p mean estimation problems are solved separately.
[0020] To define the PI estimator, let:
[0021] .
[0022] Among them, is the empirical mean based on the complete X sample of size . Now define and Here, represents the adjustment of the entire sample. Accordingly, the PI estimator is:
[0023] wherein, represents the least squares estimate of the regression model.
[0024] Herein, using to replace , we get: .
[0025] wherein,
[0026] On the other hand, the inventors used two sets of datasets in the public domain in the research. One set is the forward simulation data of the river channel, and the other set is the real data from a certain construction area. The structures of the simulation data and the real data are as follows: Simulation data: It contains 10,201 observation samples, and each sample consists of 31 seismic attributes. Each seismic attribute is a numerical feature representing the physical characteristics of seismic waves, such as amplitude, frequency, and arrival time, etc. These characteristics together characterize the seismic event. The structure of the dataset can be represented as a 10,201×31 matrix, where each row represents a sample and each column represents an attribute. The labeled and unlabeled samples are mixedly distributed in the dataset.
[0027] Real data: It also contains 10,201 observation samples, and each sample consists of 31 seismic attributes. Due to insufficient exploration in this area, only 50 samples are labeled, and the remaining 10,151 samples are unlabeled. The structure of the dataset is similar to that of the simulation data, being a 10,201×31 matrix, where the labeled and unlabeled samples are mixedly distributed.
[0028] Through these two sets of data, the inventors verified the effectiveness of the embodiments of the present invention in the case of scarce labeled data. The structure of the simulation data is shown in Table 1 below.
[0029] Table 1 Simulation data
[0030] Figure 1 Shows the distribution of covariates, indicating that the peaks of most variables are near the central value and are relatively symmetrically distributed around the mean. There is no variable presenting a uniform distribution in the figure.Figure 1 and Figure 2 The histogram and correlation matrix of the variables (X1-X31) are provided respectively, which comprehensively show the characteristics of the variables and their interrelationships. The histogram shows that the variables X1-X6, X9-X12, X17-X18, X22-X26 and X28-X30 show relative symmetry, while the variables X3, X7-X8, X13-X16, X19-X21 and X31 show skewed distribution. In particular, the variables X10-X11 show multimodal distribution. In addition, there are differences in the distribution range of the variables, for example, the distribution range of X1-X6 and related variables is narrow, while the distribution range of X3 and its corresponding variables is wide. It is worth noting that the variable X13 contains outliers. The correlation matrix reveals the correlation between the variables: for example, there is a strong positive correlation between X1 and X3 and their subsequent pairs; there is a strong negative correlation between X1 and X29 and other pairs; while there is a weak correlation or no correlation between X1 and X13. The diagonal elements are all 1, indicating the perfect correlation of the variables themselves. Based on the in-depth analysis of the characteristics of these variables, effective variable selection can be performed to retain variables that are strongly correlated with the target variable or that contribute significantly to the overall data pattern, while eliminating weakly correlated or meaningless variables. By additionally utilizing unlabeled data and correcting the estimation of linear parameters, the prediction accuracy of the model can be improved even if there are unmodeled nonlinear relationships in the real model. Based on the above comprehensive understanding of the characteristics of reservoir prediction data, an embodiment of the present invention provides a semi-supervised regression model with pseudo labels, which aims to significantly improve the accuracy of reservoir prediction.
[0031] First embodiment like Figure 3A and Figure 3B As shown, a modeling method for reservoir prediction provided by the first embodiment of the present invention includes at least the following steps.
[0032] Step 1: Prepare the dataset.
[0033] The data set contains Y observation samples, each of which consists of N seismic attributes. Each seismic attribute is a numerical feature that represents some physical properties of seismic waves, such as amplitude, frequency, and arrival time, which together characterize the logging event. The structure of the data set can be represented as a Y×N matrix, where each row of the matrix represents a sample and each column represents a seismic attribute.
[0034] Among these Y samples, M samples are labeled (hereinafter referred to as labeled samples), and the remaining Y-M samples are not labeled (hereinafter referred to as unlabeled samples). Labeled samples and unlabeled samples are mixedly distributed in the dataset.
[0035] Step 2: Use Q-type clustering based on target value correlation for variable selection.
[0036] In scenarios such as feature selection, dimensionality reduction, and variable grouping in machine learning, R-type clustering is usually more commonly used. It can help screen out irrelevant features, thereby effectively reducing the complexity of the model. In contrast, Q-type clustering mainly indirectly evaluates the importance of variables by analyzing the similarity between samples. Although R-type clustering is good at grouping variables, when dealing with strongly correlated features, it may group these related variables into the same class but cannot further eliminate the redundant variables among them. Given the variable selection requirements of the present invention, Q-type clustering is a more suitable choice because it can more directly reflect the importance of variables.
[0037] Based on the above selection, step two includes the following sub-steps: Step 2.1 Use the Pearson correlation calculation formula to calculate the correlation between each variable and the target value to identify those variables that exhibit a strong linear relationship. Among them, the target value is the labeled logging data, that is, the Y value.
[0038] In an embodiment of the present invention, the formula for Pearson correlation is:
[0039] Among them, x and y respectively represent the individual data points of the input variables (such as seismic attributes) and the target value (such as logging data) of M labeled samples. In the application scenario of reservoir prediction, is the seismic attribute of the i-th labeled sample, is the logging data of the i-th labeled sample, and are their respective means. Here, a series of operations are performed on M labeled X and Y for variable selection.
[0040] According to the calculation results, N1 (N1 < N) highly correlated variables are identified from them.
[0041] Step 2.2 Adopt Q-type clustering. According to the correlation patterns between the N1 variables (9 are selected in this embodiment) identified in the previous step and the target value, the variables are classified into different groups. Among them, the correlation pattern is the relationship structure between the target value and the N1 variables, mainly reflecting the similarity and difference between variables, and is classified through Q-type clustering.
[0042] Among them, the clustering objective can be expressed as minimizing the within-group variance , and the formula is as follows:
[0043] Among them, represents the result of the -th clustering, is the center of the clustering, They are the data points (variables) within each cluster.
[0044] According to the clustering results, further screening is carried out among the N1 variables, and finally P (P < N) key variables are selected (P = 7 in this embodiment), as Figure 4 shown.
[0045] Step 3: Use the data set in Step 1 and the P key variables selected in Step 2 to construct a working data set. Among them, the working data set is a Y×P matrix.
[0046] Step 4: On the working data set, use the pseudo-labeling technique to expand a limited number of labeled samples into fully labeled samples, generate pseudo-labels by screening unlabeled samples with high consistency, construct a fully labeled training set, and use it to train the semi-supervised regression model. Subsequently, based on the training results of the PI estimator, a semi-supervised regression model with pseudo-labels is generated.
[0047] In the following description, the definitions of each parameter or symbol can refer to the definitions in the semi-supervised regression model introduced above.
[0048] In an embodiment of the present invention, it specifically includes the following sub-steps: 4-1) Set the input and output as follows: Input: Observation value vector ; Matrix of the working data set (with p columns) ; Maximum number of iterations: ; Consistency threshold: ; Number of self-training models: ; Output: PI estimator for each k = 1, …… p ; 4.2) Initialize the semi-supervised regression model. Let the iteration counter and satisfy . Remove the k-th column from the matrix X and add an intercept column to X .
[0049] 4.3) Calculate the regression coefficients. Solve using the labeled samples: , Calculate the i-th component of the unlabeled sample .
[0050] 4.4) Adjust the response variable .
[0051] 4.5) Standardized regressors 。
[0052] 4.6) Adjusted regressors 。
[0053] 4.7) Calculate the PI estimator 。
[0054] 4.8) Calculate the consistency score.
[0055] 4.9) Determine the consistency metric.
[0056] 4.10) Update the data and retrain the model.
[0057] 4.11) Return the result 。
[0058] It should be noted that the labeled samples in Step 4 are obtained through the pseudo-labeling technique. Pseudo-labeling is a simple but efficient method aimed at improving the performance of the predictor by simultaneously utilizing labeled data and unlabeled data. In one embodiment of the present invention, the pseudo-labeling technique includes the following sub-steps: training an initial semi-supervised regression model based on the labeled samples, predicting the unlabeled samples to generate prediction results of multiple models; calculating the consistency score of each unlabeled sample, and screening out the unlabeled samples with a consistency score lower than a set consistency threshold as pseudo-labels; merging the pseudo-labels with the original labeled samples, updating the training set and iteratively optimizing the model until the convergence condition is met.
[0059] Specifically, first, the model is trained only on the labeled dataset. In each iteration, the model uses the prediction results of the previous iteration as the predicted values of the unlabeled samples and treats them as true labels. To improve the robustness of the model, this process is repeated n times (n is the number of models), and different random seeds are used each time during training to generate diverse prediction results. After obtaining the prediction results from multiple models, the next step is to evaluate the consistency of the pseudo-labels. The prediction results of all models are stored in a list, and by calculating the average of these prediction results, the prediction stability of each unlabeled sample can be evaluated. Specifically, the consistency evaluation is achieved by calculating the consistency score of each unlabeled sample in the predictions of multiple models, thereby screening out the samples with relatively stable prediction results. These pseudo-labels that have passed the consistency verification are used to expand the labeled dataset and further optimize the training effect of the model.
[0060]
[0061] Among them, is the prediction result of the k-th model for unlabeled samples. Specifically, the consistency score Ci of the i-th unlabeled sample (unlabeled sample i) is calculated, which is the sum of the absolute differences between the predicted values of each model and the average value.
[0062]
[0063] where is the prediction result of the k-th model for unlabeled sample i. This consistency check aims to ensure that only those prediction results that are consistent among different models are regarded as pseudo-labels.
[0064] Based on the set consistency threshold ( ) and the number of self-training models , samples with a consistency score lower than are selected to ensure the reliability of the selected pseudo-labels: .
[0065] These pseudo-labels are merged with the original labeled samples Y to form an updated labeled dataset, enabling the model to better utilize the unlabeled data. During the process of merging pseudo-labels, the feature matrix X of the working dataset is also extended to include data from the original labeled samples and the added pseudo-labels. In this way, the model can be trained using a more abundant dataset.
[0066] In an embodiment of the present invention, obtaining labeled samples using the pseudo-labeling technique includes the following sub-steps: First, input the observation value vector Y (original labeled samples), the feature matrix X (including p columns), the maximum number of iterations, the consistency threshold, and the number of self-training models. In each iteration, initialize an empty list (predictions_list) to store the prediction results. Then, train each model: set a random seed, train the model using the original labeled samples Y, and make predictions for the unlabeled samples, storing the prediction results in the list. Next, calculate the average predicted value of each unlabeled sample, and calculate the consistency score based on the prediction results of all models. The consistency score evaluates the stability of the prediction by calculating the sum of the absolute differences between the predicted values of each model and the average predicted value. Samples with high stability are selected as pseudo-labels according to the consistency score and the consistency threshold. Finally, these pseudo-labels are merged with the original labeled samples Y to form a new labeled dataset for subsequent model training. This process effectively utilizes the unlabeled data through multiple iterations and multi-model predictions, combined with consistency checks, enhancing the robustness and prediction performance of the model.
[0067] There are two stopping criteria for the above algorithm: First, when the maximum absolute change in the model parameters is less than the set consistency threshold, it indicates that the model parameters have stabilized and there are no significant changes, and the iterative process will terminate at this time; Second, if the number of iterations reaches the pre-set maximum allowable value, the iteration will also stop. This means that before reaching the maximum number of iterations, as long as the model converges, the iteration can be terminated in advance. The pseudo-labeling technique promotes the self-training process, which improves the model performance by training using both labeled data and unlabeled data. Specifically, for each unlabeled sample, first use multiple independently trained models to make predictions, and calculate the average value of its predicted values as the candidate pseudo-label; then calculate the consistency score by evaluating the dispersion degree of these prediction results. Finally, only keep the unlabeled samples with a consistency score lower than the consistency threshold as pseudo-labels and incorporate them into the training set. This mechanism ensures the reliability and accuracy of the pseudo-labels by strictly screening the samples with highly consistent prediction results.
[0068] To verify the actual effect of the present invention, we used a real logging dataset for experiments. This dataset contains 396,066 logging data points, which are distributed in the working grid and cover 9 different geological attributes of the formation.
[0069] Among these logging data, the data of 22 wells are with labeled values, while the data of the remaining 396,044 wells have no labeled values. Therefore, the composition of the entire dataset is as follows: the number of unlabeled samples is 396,044, the number of labeled samples is 22, and each sample contains 9 feature variables.
[0070] In the experiment, the data of 16 wells were selected from the data of these 22 known wells for model training, while the data of the remaining 6 wells were used as the blind well test set to evaluate the prediction performance of the model. The results of the blind well test are shown in Table 2.
[0071] By comparing the prediction results of the semi-supervised regression with pseudo-labels (abbreviated as PL-SSR) model and the conventional semi-supervised regression (SSR) model provided by the embodiments of the present invention, the performance and advantages of the PL-SSR model on the real dataset can be evaluated.
[0072] Table 2 Blind well test results
[0073] Figure 5 Shows the original seismic horizons and well locations of the real logging dataset, Figure 6 then shows the channel positions of this dataset on the seismic section. Variable selection is performed on the working dataset through the Q-type clustering method, Figure 7The selected variables are presented, specifically including: Peak Amplitude, spectral filtering (SF_amplitude_20, SF_amplitude_30, SF_amplitude_40, SF_amplitude_50, SF_amplitude_100) at different frequencies (20Hz, 30Hz, 40Hz, 50Hz, 100Hz), seismic impedance characteristics (standard Prg_CIP), a standardized property (standard perigram), standardized porosity calculation parameters (standard_Prg_CIP), and standardized reflection intensity (standardRSt).
[0074] The specific meanings of each variable are as follows: Peak Amplitude: It is related to the formation reflection intensity. High values may indicate hard or compact rock formations, while low values may be related to loose sedimentary layers.
[0075] SF_amplitude_20 - 100: The amplitudes at different frequencies reflect geological features at different depths or scales. High frequencies (such as 100Hz) reveal details or fractures, while low frequencies (such as 20Hz) reflect macroscopic sedimentary characteristics.
[0076] standard Prg_CIP: Standardized seismic impedance or key point data, used for analyzing the reflection contrast between geological layers and identifying subsurface fluids and lithologies.
[0077] standard RSt: Standardized reflection intensity, used for reservoir identification to reveal subsurface lithology or fluid distribution.
[0078] standard perigram: Standardized seismic profile data, used for identifying seismic reflection boundaries, faults, sedimentary environments, etc.
[0079] standard_Prg_CIP: Standardized porosity calculation method or parameters.
[0080] Figure 8 The prediction results of the PL - SSR model provided in the first embodiment of the present invention are compared with those of the conventional SSR model. It can be seen from Figure 8 that: The prediction results of the PL - SSR model are closer to the true values, indicating that its prediction effect is better than that of the conventional SSR model. In addition, Figure 9 The reservoir property heat map showing the prediction results of the PL - SSR model clearly presents the characteristics and trends of the river channels. The river channel contours are very distinct, further confirming the effectiveness of the pseudo - label technology and highlighting the significant advantages of the PL - SSR model compared with the conventional SSR model.
[0081] Second Embodiment Another modeling method for reservoir prediction is disclosed in the second embodiment of the present invention. Its main steps are similar to those of the first embodiment, but different technical combinations are adopted in the variable selection method. Specifically, in the second embodiment, R-type clustering and principal component analysis (PCA) are combined in step two for variable selection and dimensionality reduction processing.
[0082] In step 2.1, R-type clustering is first performed. R-type clustering is usually based on distance metrics, such as Euclidean distance, to evaluate the similarity between data points.
[0083] The definition of distance is:
[0084] where x represents the seismic attributes of the data points, y represents the well logging data of the data points, and n is the number of features or dimensions.
[0085] The definition of distance here involves the seismic attributes and well logging data of the data points, where the number of features or dimensions is a key parameter. Through R-type clustering, N' variables (N' is less than N) can be selected from the original N variables. In this embodiment, this number is determined to be 23 variables, as Figure 10 shown.
[0086] Next, in step 2.2, principal component analysis is performed. The goal of principal component analysis is to reduce the dimensionality of the data while retaining the maximum variance in the data. This is achieved by calculating the eigenvalues associated with each principal component, where the magnitude of the eigenvalue reflects the amount of information contained in the corresponding principal component.
[0087] Principal component analysis can be expressed as:
[0088] where represents the eigenvalue associated with the i-th principal component, and k represents the number of principal components required to retain at least 90% of the total variance.
[0089] In this embodiment, the number of principal components is selected based on the criterion of retaining at least 90% of the total variance. This method ensures that while reducing the dimensionality of the data, as much key information in the data as possible is retained. In this way, principal component analysis helps to further optimize the variable selection process, making the finally selected variables better represent the main features of the data.
[0090] After R-type clustering, N' variables (N' < N) are selected from the original N variables, and this number is determined to be 23 variables in this embodiment. Next, principal component analysis is performed on these 23 variables. Principal component analysis is a statistical method used to reduce the dimensionality of data while retaining the maximum variance in the data. By calculating the eigenvalues of each principal component, its contribution to the total variance of the data can be determined. In this embodiment, the number of principal components selected is based on the criterion of retaining at least 90% of the total variance. According to the results of principal component analysis, P key variables (P < N') are further selected from these 23 variables. These variables reduce the complexity of the data while retaining key information, thus optimizing the variable selection process and providing a more efficient dataset for subsequent modeling.
[0091] To comprehensively evaluate the performance of the PL-SSR model, the inventors systematically compared it with classical regression methods, including linear regression, ridge regression, Lasso regression, and random forest regression. The comparison covered the first embodiment (Method 1) and the second embodiment (Method 2), and the prediction results of each model were analyzed in detail under three different conditions. The experimental results revealed three key conclusions: (1) The effectiveness of reservoir prediction is highly sensitive to the number of available samples. Especially in the case of scarce labeled data, the performance differences of the models are significant.
[0092] (2) After introducing the pseudo-labeling technique, the overall performance of the semi-supervised regression model has been significantly improved, demonstrating the effectiveness of the pseudo-labeling technique in enhancing the model's prediction ability.
[0093] (3) Under the framework of the PL-SSR model, the model not only effectively addresses the challenge of insufficient labeled data but also particularly improves the recognition accuracy for categories with fewer samples, which is particularly important for the recognition of rare categories in geological applications.
[0094] Specifically, the PL-SSR model exhibits superior prediction performance under all test conditions, with lower root mean square error (RMSE) and mean absolute error (MAE), and a higher coefficient of determination (R²), indicating that the model can perform reservoir prediction more accurately. Especially in the case of scarce labeled data, the PL-SSR model has a significant improvement in the recognition accuracy for categories with fewer samples compared to other methods, which is of great significance for practical applications such as oil and gas prediction, reservoir modeling, and well logging data analysis.
[0095] Table 3 Data simulation results
[0096] Table 3 shows the prediction performance of different models using the first embodiment (Method 1) and the second embodiment (Method 2) on the original data and the data processed by two different methods. The evaluation metrics include root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²).
[0097] When comparing methods such as SSR, PL-SSR, Lasso, Ridge, RF, and PL-RF, Method 1 performs better in most cases, with lower RMSE and MAE values and higher R² values, indicating that Method 1 generally has higher prediction accuracy. However, under the PL-Lasso method, Method 2 shows lower RMSE and MAE values and slightly higher R² values, indicating that Method 2 has better prediction performance in specific situations.
[0098] It can be concluded that Method 1 is superior to Method 2 in most cases. It is worth noting that introducing pseudo-labels improves the model performance in all cases, highlighting the effectiveness of the pseudo-labeling technique in enhancing the model performance. Generally speaking, the semi-supervised regression model itself shows lower RMSE and MAE values and higher R² values, and after adding pseudo-labels, the RMSE is further reduced, and the model performance is further improved.
[0099] In summary, when the present invention is applied to logging data, through the overall performance comparison of different variable selection techniques and common regression methods, the following conclusions are obtained: Overall performance: Both PL-SSR and SSR show superior overall performance, with lower RMSE and MAE and higher R² values, indicating that their regression predictions are more accurate.
[0100] Effect of pseudo-labels: Adding pseudo-labels for training significantly improves the overall performance of the model. Especially for SSR, the regression prediction accuracy is most significantly improved after adding pseudo-labels.
[0101] Therefore, the present invention successfully solves the problem of reduced prediction accuracy caused by scarce labels, significantly improves the prediction accuracy, and is particularly suitable for geological applications such as oil and gas prediction, reservoir modeling, and logging data analysis.
[0102] Third Embodiment Based on the foregoing first embodiment or second embodiment, the third embodiment of the present invention further provides a reservoir prediction method, including the following steps: Input the dataset to be predicted into a semi-supervised regression model with pseudo-labels, where each seismic attribute in the dataset is a numerical feature representing the physical characteristics of seismic waves; Use the semi-supervised regression model with pseudo-labels to perform reservoir prediction and output the prediction result; Among them, the semi-supervised regression model with pseudo-labels is constructed by the modeling method for reservoir prediction described in the first embodiment (Method 1) or the second embodiment (Method 2).
[0103] Fourth Embodiment As Figure 11 shown, based on the above reservoir prediction method, the fourth embodiment of the present invention further provides a reservoir prediction system. The reservoir prediction system includes one or more processors and a memory. Among them, the memory is coupled to the processor and is used to store one or more programs. When the program is executed by the processor, the processor implements the reservoir prediction method in the above embodiments.
[0104] Among them, the processor is used to control the overall operation of the reservoir prediction system to complete all or part of the steps of the above reservoir prediction method. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. The memory is used to store various types of data to support the operation of the reservoir prediction system. These data can include, for example, instructions for any application program or method operating on the reservoir prediction system, as well as application program related data. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, etc.
[0105] In another exemplary embodiment, the present invention also provides a computer-readable storage medium including program instructions. When the program instructions are executed by the processor, the steps of the reservoir prediction method in any one of the above embodiments are implemented. For example, the computer-readable storage medium can be the above memory including program instructions, and the above program instructions can be executed by the processor to complete the above reservoir prediction method and achieve the same technical effects as the above method.
[0106] It should be noted that the above multiple embodiments are only examples. The technical solutions of each embodiment can be combined, and the order of each step can be changed, all within the protection scope of the present invention.
[0107] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0108] The modeling method, reservoir prediction method and system provided by the present invention are described in detail above. For those of ordinary skill in the art, any obvious changes made without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and shall bear corresponding legal liabilities.
Claims
1. A modeling method for reservoir prediction, characterized in that The steps include: Step 1: Prepare a data set, the data set includes Y observation samples, each sample consists of N seismic attributes, and the seismic attributes are numerical features that characterize logging events; Step 2: Use clustering method to select variables for the data set and screen out P key variables; Step 3: Based on the data set obtained in step 1 and the P key variables screened in step 2, construct a Y×P working data set; Step 4: On the working data set, a limited number of labeled samples are expanded into fully labeled samples by using pseudo-labeling technology for training the semi-supervised regression model, and a semi-supervised regression model with pseudo-labels is generated based on the training results of the PI estimator; Among them, Y, N, and P are all positive integers.
2. The modeling method for reservoir prediction according to claim 1, characterized in that The clustering method in step 2 is Q-type clustering, which specifically includes the following sub-steps: The Pearson correlation calculation formula is used to calculate the correlation between each variable and the target value, and N1 highly correlated variables are screened out, wherein the target value is the labeled logging data; The N1 variables are subjected to Q-type clustering, and based on the principle of minimizing intra-group variance, the variables are grouped and then P key variables are selected.
3. The modeling method for reservoir prediction according to claim 1, characterized in that The clustering method in step 2 is a combination of R-type clustering and principal component analysis, which specifically includes the following sub-steps: N' variables are selected from the original variables through R-type clustering; A principal component analysis was performed on the N' variables, and P key variables were selected based on the criterion of retaining at least 90% of the total variance.
4. The modeling method for reservoir prediction according to claim 1, characterized in that The pseudo-labeling technique in step 4 includes the following sub-steps: Based on the labeled samples, the initial semi-supervised regression model is trained, and predictions are made for each unlabeled sample to generate prediction results for multiple models. Calculate the consistency score of each unlabeled sample and select the unlabeled samples below the set consistency threshold as pseudo labels; Merge the pseudo labels with the original labeled samples, update the training set and iteratively optimize the model until the convergence condition is met.
5. The modeling method for reservoir prediction according to claim 4, characterized in that The following sub-steps are included: For each unlabeled sample, multiple independently trained models are first used to make predictions, and the average of the prediction results is calculated as the candidate pseudo-label; then, the consistency score is calculated by evaluating the degree of discreteness of the prediction results; finally, only the unlabeled samples with a consistency score lower than the consistency threshold are retained as pseudo-labels and included in the training set.
6. The modeling method for reservoir prediction according to claim 5, characterized in that The prediction stability of each unlabeled sample is evaluated using the following formula: ;in, is the prediction result of the kth model for unlabeled samples, and K is a positive integer.
7. The modeling method for reservoir prediction according to claim 6, characterized in that The consistency score Ci is calculated by the following formula: ;in, is the prediction result of the k-th model for the i-th unlabeled sample, where i is a positive integer.
8. A reservoir prediction method, characterized in that The following steps are involved: Inputting the data set to be predicted into the semi-supervised regression model with pseudo labels constructed by the modeling method described in any one of claims 1 to 7; A semi-supervised regression model with pseudo labels is used to predict reservoirs and output the prediction results.
9. A reservoir prediction system, characterized in that It comprises a processor and a memory, wherein the processor and the memory are coupled; wherein the memory is used to store a computer program; and the processor runs the computer program stored in the memory to implement the reservoir prediction method according to claim 8.
Citation Information
Patent Citations
Intelligent geoscience data pseudo label generation method based on machine learning
CN115859215A
Geological modeling method and system based on active domain adaptation learning, and medium
CN117173350A
Reservoir lithology prediction method based on multidisciplinary data fusion clustering algorithm
CN117452518A
Semi-supervised relation extraction method, system and equipment based on interaction consistency training and medium
CN117725917A
Multi-task classifier construction method combining pre-training and supervision fine tuning
CN118312839A