An intelligent prediction method for shale content logging curve of oil and gas reservoir

By combining the coefficient of determination and grid optimization methods to determine the parameters of the Gardner and Larionov formulas, generating a virtual dataset, and combining various machine learning algorithms, the problem of insufficient generalization ability of traditional methods is solved, and high-precision prediction of the logging curve of mud content in oil and gas reservoirs is achieved.

CN119647687BActive Publication Date: 2025-11-04XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411771668.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-11-04
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Traditional methods based on mathematical statistics and rock physics models have limited generalization ability and applicability in predicting the clay content of oil and gas reservoirs through logging curves. Furthermore, human adjustment of parameters can easily introduce errors, while machine learning methods struggle to explain rock physics relationships, resulting in insufficient prediction accuracy.

Method used

The optimal parameters of the Gardner and Larionov rock physics formulas were determined using the coefficient of determination and grid optimization methods. A virtual dataset was generated, and multiple machine learning algorithms were used for training. The prediction results were evaluated using cross-validation, correlation coefficient, and SHAP methods to form a fused dataset for predicting clay content.

Benefits of technology

It improved the accuracy of mud content prediction by 12%, significantly improved the accuracy and reliability of well logging curve datasets, and enhanced the interpretability and transparency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119647687B_ABST
    Figure CN119647687B_ABST
Patent Text Reader

Abstract

The application discloses an oil and gas reservoir shale content logging curve intelligent prediction method, and belongs to the technical field of rock physics and logging curve processing and interpretation, and comprises the following steps: in step 101, the Gardner formula parameters are determined by combining the determination coefficient and the grid optimization method, and then the virtual density data set capable of covering the measured density data set is generated by using the Gardner formula. By using the determination coefficient and the grid optimization method to determine the optimal parameters of the Gardner rock physics formula and the Larionov rock physics formula, the measured logging data abnormal values can be effectively corrected, the fusion data set for training various machine learning models can be generated, the cross-validation method, the correlation coefficient, the determination coefficient and the multi-dimensional SHAP method are comprehensively used to more comprehensively evaluate and explain the prediction results, compared with the traditional rock physics method, the prediction accuracy of the shale content is improved by 12%, the logging curve data set can be effectively corrected and expanded, and the accuracy and reliability of the oil and gas reservoir shale content prediction can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rock physics and well logging curve processing and interpretation technology, and particularly relates to an intelligent prediction method for well logging curves of mud content in oil and gas reservoirs. Background Technology

[0002] Oil and natural gas are not only non-renewable energy resources but also a strategic pillar for a nation's survival and development, playing a crucial role in ensuring energy security and promoting economic and social progress. However, due to limitations in my country's oil and gas exploration technology and costs, the exploration and development of deep, ultra-deep, thin, structurally complex oil and gas reservoirs with harsh surface conditions face significant technical challenges. Well logging data, as a more direct reflection of the physical properties of oil and gas formations, contains rich geological information and provides a reliable basis for accurately obtaining geological parameters. Through professional analysis of well logging curves, geologists can gain a detailed understanding of the regional geological structure and the physical characteristics of the formation rocks, thereby better grasping the underground geological conditions and establishing more accurate geological models. Therefore, well logging curves are an indispensable and important data source in oil and gas exploration and development.

[0003] Due to the complexity of the three-dimensional formation structure and the limitations of logging methods and equipment, logging curves typically only reflect formation parameters around the wellbore. When analyzing formations further away from the wellbore, parameters may vary significantly. Considering the enormous costs and resources involved in drilling, logging wells are usually located in specific study areas. This layout may result in gaps in formation information in un-well-drilled areas, or even overlook the existence of hidden reservoirs, leading to wasted resources. Furthermore, during logging, some data may be distorted or missing due to equipment limitations, mud invasion, or wellbore collapse. Therefore, utilizing existing logging data for data correction, completion, or prediction of logging curves in unknown areas has become an important means of reservoir prediction, effectively improving drilling reliability and safety while reducing exploration costs.

[0004] In the early stages of oil well logging, researchers mostly used mathematical and statistical methods to process and interpret logging data. For example, Archie (1942) could deduce porosity from measured resistivity and known formation water resistivity; Larionov (1969) proposed a method to estimate the clay content in reservoir rocks by using the relationship between shale volume and gamma-ray index; Gardner GHF, Gardner LW, and others (1974) described the relationship between acoustic velocity and density in rocks using their equations; Castagna, Batzle, and others (1985) derived an empirical formula for shear wave velocity based on P-wave velocity; Smith (2007) explored the relationship between resistivity logging and acoustic logging, with the main purpose of improving the interpretation of formation characteristics by using the correlation between acoustic velocity (P-wave or shear wave velocity) and resistivity.

[0005] This method, based on mathematical statistics and rock physics models, is the primary means of processing and interpreting well logging data. While this traditional method has laid a theoretical foundation and provided technical support for subsequent research, it still has limitations. This approach relies on rock physics models for mathematical and statistical analysis; however, it fails to fully consider the potential nonlinear correlations between data points, thus limiting the accuracy and applicability of the prediction results. In practical applications, this method based on mathematical statistics and rock physics models struggles to provide sufficiently accurate geological information. First, empirical models and statistical methods are often validated based on specific assumptions or limited data, resulting in insufficient generalization ability. Second, these methods rely on human selection and adjustment of model parameters, easily introducing subjective bias. Finally, when processing large-scale or high-dimensional data, traditional methods often struggle to effectively handle complex data structures.

[0006] With the rapid development of big data technology and the widespread application of deep learning in science and engineering, more and more researchers and production personnel are beginning to use data-driven methods such as artificial neural networks to process and interpret well logging curves. For example, Rolon, Mohaghegh et al. (2009) used a neural network model to synthesize resistivity curves from natural gamma, density, and neutron curve data, demonstrating higher prediction accuracy than traditional methods. Furthermore, researchers have also attempted to use advanced artificial intelligence algorithms such as deep learning, transfer learning, and few-shot learning to construct predictive models for well logging curves, which have performed well under specific conditions. However, due to the complex network structures of these models, it is difficult to accurately interpret and capture the rock physics relationship between the input and output curves, and they may still face challenges in dealing with nonlinear problems in new work areas and more complex formations.

[0007] In summary, traditional methods based on mathematical statistics and rock physics models suffer from limited generalization ability and applicability. Manually adjusting rock physics model parameters can lead to significant errors, while machine learning methods are black-box operations, unable to determine the rock physics relationship between output and input data. To address this, this invention proposes an intelligent prediction method for oil and gas reservoir clay content logging curves under rock physics constraints. This method utilizes the coefficient of determination and grid optimization to determine the optimal parameters of the Gardner and Larionov rock physics formulas. It can effectively correct outliers in logging data and generate virtual datasets for training various machine learning models. By comprehensively utilizing cross-validation, correlation coefficient, coefficient of determination, and multi-dimensional SHAP methods, the prediction results are evaluated and interpreted more fully. Compared with traditional rock physics methods, the prediction accuracy of clay content is improved by 12%. This method not only effectively corrects and expands the logging curve dataset but also significantly improves the accuracy and reliability of clay content prediction.

[0008] Based on this, the present invention designs an intelligent prediction method for logging curves of mud content in oil and gas reservoirs to solve the above problems. Summary of the Invention

[0009] The purpose of this invention is to propose an intelligent prediction method for logging curves of mud content in oil and gas reservoirs, in order to solve the problem that traditional methods based on mathematical statistics and rock physics models have limited generalization ability and applicability.

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] A method for intelligent prediction of clay content logging curves in oil and gas reservoirs, comprising:

[0012] In step 101, the parameters of the Gardner formula are determined by combining the coefficient of determination with the grid optimization method. Then, the Gardner formula is used to generate a virtual density dataset that can cover the measured density dataset. This virtual dataset is combined with the measured dataset to form a comprehensive density dataset.

[0013] In step 102, the parameters of the Larionov formula are determined using the coefficient of determination and grid optimization method, and a virtual gamma dataset that covers the actual measured gamma dataset is created based on this formula. This virtual dataset is then merged with the measured dataset to construct a complete gamma dataset.

[0014] In step 103, using the fused density dataset and gamma dataset, various machine learning algorithms, including Random Forest (RFs), Gradient Boosting Decision Tree (XGBoost), Multilayer Perceptron (MLP), and Support Vector Machine (SVM), are applied to train the model for regression analysis, in order to intelligently predict mud content. Cross-validation is used to evaluate the prediction accuracy of each algorithm.

[0015] In step 104, the correlation coefficient, determination coefficient, and multidimensional SHAP method are used to comprehensively interpret the prediction results of machine learning, and they are compared with traditional rock physics methods to analyze the improvement effect of the fusion dataset on the intelligent prediction performance.

[0016] As a further description of the above technical solution:

[0017] In step 101, the only variable involved in Gardner's empirical formula is the speed of sound, as shown in the following formula:

[0018] (1)

[0019] In the formula, ρ is the rock density (g / cc); Vt is the sound wave velocity (us / ft); a and b are constants that vary by region.

[0020] The Gardner empirical formula (1) fits the optimal parameters for the work area based on the coefficient of determination and grid optimization, with a=0.0665 and b=0.3923. The highest coefficient of determination for the fitting effect can reach 0.998, indicating the reliability and accuracy of the method. Then, a virtual velocity-density dataset is generated using the Gardner formula.

[0021] When generating the virtual velocity-density dataset, the following principles should be followed: the generated virtual dataset should cover the measured velocity-density dataset, while maintaining the characteristics of the measured dataset and avoiding excessive expansion of the measured dataset; optimal coverage is achieved by just covering the measured dataset. Virtual density data is generated by selecting a certain range of sound wave velocities and coefficients. The sound wave velocity Vt ranges from 115 to 150, with a step size of 0.01, representing the range of sound wave velocity variation. The coefficient 'a' in the density formula ranges from 0.0653 to 0.0675, with a step size of 0.00005. The parameter b = 0.3923 keeps the exponent constant in the density formula, representing the exponential relationship between sound wave velocity and density. Using these parameters, density curves for different 'a' values ​​are generated using code to simulate the virtual data relationship between sound wave velocity and density. A simple merging of the measured and virtual datasets yields the fused dataset.

[0022] As a further description of the above technical solution:

[0023] In step 102, based on the coefficient of determination and the Larionov correction formula for mesh optimization, the optimal parameter a=3.2 is fitted for this work area. There is a relationship between VSh and IRA as follows:

[0024] (2)

[0025] In the formula, IRA represents gamma rays or radioactivity index; VSh represents clay content.

[0026] (3)

[0027] In the formula, GR is the actual gamma logging curve reading; GRcs is the gamma reading of clean sandstone (without shale), with a value of 20; and GRSh is the gamma reading of shale, with a value of 140.

[0028] To generate a virtual gamma-mud content dataset using the Larionov formula, the following principles must be followed: the generated virtual dataset should cover the measured gamma-mud content dataset, while maintaining the characteristics of the measured dataset without excessively expanding it; optimal coverage is achieved by just covering the measured dataset. Virtual gamma curve (GR) data is generated by selecting a certain range of mud content (VSh) values ​​and parameters. We defined the mud content range as 0.01 to 0.3, with a step size of 0.0001. The parameter 'a' in the Larionov formula ranges from 1 to 5.5, with a step size of 0.03. Using this parameter, gamma ray curves at different 'a' values ​​were generated using code to simulate the virtual data relationship between mud content and gamma. A simple merging of the measured and virtual datasets yields the fused dataset.

[0029] As a further description of the above technical solution:

[0030] In step 103, algorithms such as Random Forest (RFs), Gradient Boosting Decision Tree (XGBoost), Multilayer Perceptron (MLP), and Support Vector Machine (SVM) are used to train machine learning based on the fused dataset to intelligently predict the mud content.

[0031] As a further description of the above technical solution:

[0032] In step 103, during model training, we selected three wells, J032, J061, and J034, as training wells. We used gamma ray, sonic waves, density, impedance, and velocity curves as input features, and mud content as the prediction target to train the model. Well J021 was used for model validation. The algorithm's prediction results were evaluated using the coefficient of determination (R²) and correlation coefficient (r). The results showed that the RFs algorithm performed best, with a coefficient of determination of 0.999 and a correlation coefficient of 0.998.

[0033] As a further description of the above technical solution:

[0034] In step 104, the accuracy of the intelligent prediction results of reservoir properties is evaluated using cross-validation, correlation coefficient, coefficient of determination (R2), and SHAP value method.

[0035] As a further description of the above technical solution:

[0036] Cross-validation is a statistical method widely used for evaluating the performance of machine learning models. Its core idea is to validate the model's stability and generalization ability by repeatedly partitioning the dataset. In practice, cross-validation divides the dataset into several equal-sized portions (called "folds"). Each time, one fold is selected as the test set, and the remaining folds are used as the training set. This training and testing process is repeated until each fold has been used as a test set at least once. The most common K-fold cross-validation divides the data into K subsets and performs K training and testing iterations. The final evaluation result is obtained by averaging the results of the multiple tests.

[0037] The correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables, and it is usually expressed by the formula:

[0038] (4)

[0039] In the formula, xi refers to the actual measured curve value (i.e., the true value), yi is the curve value predicted by the model (i.e., the predicted value), and and are the average values ​​of x and y, respectively. It is a statistical indicator used to measure the degree of correlation between two variables, usually expressed using the Pearson correlation coefficient. Its value ranges from -1 to 1: 1 indicates perfect positive correlation, -1 indicates perfect negative correlation, and 0 indicates no linear correlation. The calculation method is to divide the covariance of the two variables by the product of their standard deviations, reflecting the strength and direction of the linear relationship between the variables. Through the correlation coefficient, we can understand the linear consistency between the model's predicted values ​​and the actual values.

[0040] The coefficient of determination is defined as follows:

[0041] (5)

[0042] Where xi and yi are the original value and the machine learning model's predicted value, respectively, X is the average of xi, and n is the total number of samples. In R2, the sum of squares of the differences between the actual and predicted values ​​represents the total amount of information that our model has not captured, with the denominator being the amount of information contained in the true label. Therefore, both measure the proportion of 1 - the amount of information our model has not captured relative to the amount of information contained in the true label, so the closer both are to 1, the better. If the result is 0, it indicates that the model fits very poorly; if the result is 1, it indicates that the model has no errors.

[0043] SHAP, based on Shapley values ​​in game theory, provides a fair and consistent method for feature attribution. Within this framework, the contribution of each feature is calculated by considering all possible combinations of features and their average impact across those combinations, thus ensuring that the results are unaffected by the order of feature inputs.

[0044] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0045] This invention provides an intelligent prediction method for oil and gas reservoir clay content logging curves based on rock physics constraints. It comprehensively utilizes the coefficient of determination and grid optimization methods to determine the optimal parameters of the Gardner and Larionov rock physics formulas. This method can effectively correct outliers in measured logging data and generate a fusion dataset for training various machine learning models. It also comprehensively utilizes cross-validation, correlation coefficient, coefficient of determination, and multi-dimensional SHAP methods to provide a more comprehensive evaluation and interpretation of the prediction results. Compared with traditional rock physics methods, the prediction accuracy of clay content is improved by 12%. This method can not only effectively correct and expand the logging curve dataset, but also significantly improve the accuracy and reliability of oil and gas reservoir clay content prediction. Attached Figure Description

[0046] Figure 1 This is a flowchart of a specific embodiment of the intelligent prediction method for well logging curves of oil and gas reservoir clay content based on rock physics constraints according to the present invention;

[0047] Figure 2 The acoustic and density intersection plot is a fusion dataset composed of measured data and rock physical generated data of the present invention.

[0048] Figure 3 This invention provides a comparison between the generated density data based on the Gardner formula using the coefficient of determination and grid optimization, and actual well logging datasets.

[0049] Figure 4 The gamma and clay content of the fused dataset, consisting of the measured data of this invention and the data generated based on the Larionov rock physical relationship, are plotted.

[0050] Figure 5 This is a comparison chart of the results of intelligent prediction of clay content based on rock physics method and RFs algorithm of the present invention;

[0051] Figure 6 This diagram illustrates the process of using the multidimensional SHAP method to explain and evaluate the feature importance of the random forest algorithm in this invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] Please see the appendix Figure 1 -Appendix Figure 6 This invention provides a technical solution: an intelligent prediction method for logging curves of clay content in oil and gas reservoirs, comprising:

[0054] In step 101, the parameters of the Gardner formula are determined by combining the coefficient of determination with the grid optimization method. Then, the Gardner formula is used to generate a virtual density dataset that can cover the measured density dataset. This virtual dataset is combined with the measured dataset to form a comprehensive density dataset.

[0055] Scatter plots can display the scatter points between each pair of variables in a dataset, which helps to observe the relationships between different logging parameters. It can reveal whether there is a linear relationship or clustering trend between two parameters. By setting different shapes for different well categories, scatter plots can clearly show the distribution differences between different categories of data. This is particularly useful when studying the performance of different regions of formations. By observing scatter plots, researchers can quickly identify potential patterns, trends, and outliers in the data, which is invaluable for further geological analysis and decision support. The scatter plot shows anomalies in the density curve of well J021. Subsequently, the velocity-density logging scatter plot was corrected based on the Gardner formula using the coefficient of determination and grid optimization, and a density dataset was generated using sonic logging data, such as... Figure 3 As shown. The only variable involved in Gardner's empirical formula is the speed of sound, and the formula is as follows:

[0056] (1)

[0057] In the formula, ρ is the rock density (g / cc); Vt is the sound wave velocity (us / ft); a and b are constants that vary by region.

[0058] Gardner's empirical formula (1) uses the coefficient of determination and mesh optimization to fit the optimal parameters for the work area: a=0.0665; b=0.3923. The highest coefficient of determination for the fitting effect can reach 0.998, indicating the reliability and accuracy of the method.

[0059] Then, a virtual velocity-density dataset is generated using the Gardner formula. When generating the virtual velocity-density dataset, the following principles must be followed: the generated virtual dataset should cover the measured velocity-density dataset, while maintaining the characteristics of the measured dataset without excessively expanding it; optimal coverage is achieved by just covering the measured dataset. Virtual density data is generated by selecting a certain range of acoustic velocities and coefficients. The acoustic velocity Vt ranges from 115 to 150, with a step size of 0.01, representing the range of acoustic velocity variation. The coefficient 'a' in the density formula ranges from 0.0653 to 0.0675, with a step size of 0.00005. The parameter b = 0.3923 keeps the exponent constant in the density formula, representing the exponential relationship between acoustic velocity and density. Using these parameters, density curves for different 'a' values ​​are generated using code to simulate the virtual data relationship between acoustic velocity and density. A simple merging of the measured and virtual datasets yields a fused dataset. Compared to the measured dataset, the fused dataset has the following advantages: it not only maintains the integrity of the measured data but also compensates for its deficiencies. Effectively expanding the dataset under rock physics constraints can lay a solid data foundation for improving the generalization ability and prediction effect of subsequent intelligent predictions.

[0060] In step 102, the parameters of the Larionov formula are determined using the coefficient of determination and grid optimization method, and a virtual gamma dataset that covers the actual measured gamma dataset is created based on this formula. This virtual dataset is then merged with the measured dataset to construct a complete gamma dataset.

[0061] Based on the coefficient of determination and the Larionov modified formula for mesh optimization, the optimal parameter a=3.2 is fitted for this work area. There is a relationship between VSh and IRA as follows:

[0062] (2)

[0063] In the formula, IRA represents gamma rays or radioactivity index; VSh represents clay content.

[0064] (3)

[0065] In the formula, GR is the actual gamma logging curve reading; GRcs is the gamma reading of clean sandstone (without shale), with a value of 20; and GRSh is the gamma reading of shale, with a value of 140.

[0066] Then, a virtual gamma-mud content dataset was generated using the Larionov formula. The following principles were followed when generating the virtual gamma-mud content dataset: the generated virtual dataset needed to cover the measured gamma-mud content dataset, while maintaining the characteristics of the measured dataset without excessively expanding it; optimal coverage was achieved by just covering the measured dataset. Virtual gamma curve (GR) data was generated by selecting a certain range of mud content (VSh) values ​​and parameters. We defined the mud content range as 0.01 to 0.3, with a step size of 0.0001. The parameter 'a' in the Larionov formula ranged from 1 to 5.5, with a step size of 0.03. Using this parameter, gamma ray curves at different 'a' values ​​were generated using code to simulate the virtual data relationship between mud content and gamma. A simple merging of the measured and virtual datasets yielded a fused dataset. Compared to the measured dataset, the fused dataset has the following advantages: it not only maintains the integrity of the measured data but also compensates for its deficiencies. Effectively expanding the dataset under rock physics constraints can lay a solid data foundation for improving the generalization ability and prediction effect of subsequent intelligent predictions.

[0067] In step 103, using the fused density dataset and gamma dataset, various machine learning algorithms, including Random Forest (RFs), Gradient Boosting Decision Tree (XGBoost), Multilayer Perceptron (MLP), and Support Vector Machine (SVM), are applied to train the model for regression analysis, enabling intelligent prediction of mud content. Cross-validation is used to evaluate the prediction accuracy of each algorithm.

[0068] This study utilizes algorithms such as Random Forests (RFs), Gradient Boosting Decision Trees (XGBoost), Multilayer Perceptrons (MLPs), and Support Vector Machines (SVMs) to train machine learning on a fused dataset for intelligent prediction of mud content. RFs reduce the risk of overfitting to specific training data by constructing multiple decision trees and averaging or majority voting on the results. They also handle datasets with a large number of features well and do not require excessive feature selection or preprocessing. XGBoost improves prediction accuracy by minimizing error through the construction of multiple progressively optimized decision trees. Its advantages include built-in regularization to prevent overfitting, support for parallel processing to improve speed, and robust handling of missing values. MLPs utilize a multilayer neural architecture and learn through backpropagation to capture nonlinear relationships in the data. Their advantages include the ability to simulate complex functions and applicability to various tasks such as regression and classification. SVMs classify data by finding the hyperplane that maximizes the sample margin. Their advantages include good generalization performance, effective handling of high-dimensional data, and the ability to solve nonlinear problems using kernel functions.

[0069] During model training, we selected three wells, J032, J061, and J034, as training wells. We used gamma ray, sonic logging, density, impedance, and velocity curves as input features, and mud content as the prediction target to train the model. Well J021 was used for model validation. The algorithm's prediction results were evaluated using the coefficient of determination (R²) and correlation coefficient (r). The results showed that the RFs algorithm performed best, with a coefficient of determination of 0.999 and a correlation coefficient of 0.998.

[0070] In step 104, the correlation coefficient, determination coefficient, and multidimensional SHAP method are used to comprehensively interpret the prediction results of machine learning, and they are compared with traditional rock physics methods to analyze the improvement effect of the fusion dataset on the performance of intelligent prediction.

[0071] The accuracy of the intelligent prediction results of reservoir properties was evaluated using cross-validation, correlation coefficient, coefficient of determination (R2), and SHAP value method.

[0072] Cross-validation is a statistical method widely used for evaluating the performance of machine learning models. Its core idea is to validate the model's stability and generalization ability by repeatedly splitting the dataset. In practice, cross-validation divides the dataset into several equal-sized parts (called "folds"). Each time, one fold is selected as the test set, and the remaining folds are used as the training set. This training and testing process is repeated until each fold has been used as a test set once. The most common K-fold cross-validation divides the data into K subsets and performs K training and testing iterations. The final evaluation result is obtained by averaging the results of multiple tests. This method not only avoids the model evaluation bias that may result from a single data split but also effectively balances the data distribution between training and testing, improving the evaluation of the model's generalization ability. The advantage of this method is that it can fully utilize the dataset, especially when the amount of data is limited. Compared to a single-split training-test split, cross-validation provides a more robust performance evaluation and helps reduce the risk of overfitting, allowing for a more accurate measurement of the model's adaptability to new data.

[0073] The correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables, and it is usually expressed by the formula:

[0074] (4)

[0075] In the formula, xi refers to the actual measured curve value (i.e., the true value), yi is the curve value predicted by the model (i.e., the predicted value), and and are the average values ​​of x and y, respectively. It is a statistical indicator used to measure the degree of correlation between two variables, usually expressed using the Pearson correlation coefficient. Its value ranges from -1 to 1: 1 indicates perfect positive correlation, -1 indicates perfect negative correlation, and 0 indicates no linear correlation. The calculation method is to divide the covariance of the two variables by the product of their standard deviations, reflecting the strength and direction of the linear relationship between the variables. Through the correlation coefficient, we can understand the linear consistency between the model's predicted values ​​and the actual values.

[0076] The coefficient of determination is defined as follows:

[0077] (5)

[0078] Where xi and yi are the original value and the machine learning model's predicted value, respectively, X is the average of xi, and n is the total number of samples. In R2, the sum of squares of the differences between the actual and predicted values ​​represents the total amount of information our model did not capture, with the denominator being the amount of information contained in the true label. Therefore, both measures the proportion of 1 - the amount of information our model did not capture relative to the amount of information contained in the true label; hence, the closer both are to 1, the better. A result of 0 indicates a poor model fit; a result of 1 indicates no model errors. SHAP, based on Shapley values ​​in game theory, provides a fair and consistent method for feature attribution. Within this framework, the contribution of each feature is calculated by considering all possible combinations of features and their average impact across these combinations, thus ensuring that the results are unaffected by the order of feature inputs. The advantage of SHAP is that it not only improves the transparency of the model but also allows users to gain a deeper understanding of the logic behind the predictions. Furthermore, SHAP can be used to optimize feature engineering by quantitatively assessing the importance of each feature, helping to filter out the features that have the greatest impact on model performance, thereby simplifying model input. It also plays a crucial role in model error diagnosis, helping to identify potential causes of inaccurate predictions and guiding model adjustments and improvements. Through these capabilities, SHAP significantly enhances model interpretability and transparency, allowing users not only to see the prediction results but also to understand why the model reached those conclusions. This interpretability greatly increases the trustworthiness of the model, especially in practical applications, helping to make the decision-making process of machine learning models more transparent and reliable.

[0079] Evaluating predictive models simultaneously using multiple evaluation metrics allows for the analysis of the specific contributions of features to model predictions and provides a clear visual representation of the accuracy of the predicted values. This significantly enhances the transparency and credibility of the models, while also offering practical insights, leading to a better understanding and evaluation of these models. Model evaluation methods include... Figure 6 As shown, gamma plays a decisive role in the prediction of clay content. Analysis of the SHAP plot, R², and r indicates that the key to successful clay content prediction lies in the generation of a virtual training dataset with a certain rock physiological relationship to the prediction target, based on the correction of the Gardner and Larionov formulas using the coefficient of determination and grid optimization.

[0080] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent prediction of clay content logging curves in oil and gas reservoirs, characterized in that, include: In step 101, the parameters of the Gardner formula are determined by combining the coefficient of determination with the grid optimization method. Then, the Gardner formula is used to generate a virtual density dataset that can cover the measured density dataset. This virtual dataset is combined with the measured dataset to form a comprehensive density dataset. In step 102, the parameters of the Larionov formula are determined using the coefficient of determination and grid optimization method, and a virtual gamma dataset that can cover the actual measured gamma dataset is created based on this formula. This virtual dataset is then merged with the measured dataset to construct a complete gamma dataset. In step 103, using the fused density dataset and gamma dataset, various machine learning algorithms such as Random Forest (RFs), Gradient Boosting Decision Tree (XGBoost), Multilayer Perceptron (MLP), and Support Vector Machine (SVM) are applied to train the model for regression, so as to intelligently predict the mud content; the prediction accuracy of each algorithm is evaluated by cross-validation. In step 104, the correlation coefficient, determination coefficient, and multidimensional SHAP method are used to comprehensively interpret the prediction results of machine learning, and they are compared with traditional rock physics methods to analyze the effect of fusion dataset on improving intelligent prediction performance. In step 101, the only variable involved in Gardner's empirical formula is the speed of sound, as shown in the following formula: ρ=a(V t ) b (1); In the formula, ρ is the rock density (g / cc); Vt is the sound wave velocity (us / ft); a and b are constants that vary by region. Gardner's empirical formula (1) fits the optimal parameters a = 0.0665 and b = 0.3923 based on the coefficient of determination and mesh optimization. The highest coefficient of determination for fitting the optimal parameters is 0.998, which shows the reliability and accuracy of the method. Then, virtual velocity-density data are generated using the Gardner formula; When generating the virtual velocity-density dataset, the following principles should be followed: the generated virtual dataset should cover the measured velocity-density dataset, while maintaining the characteristics of the measured dataset and avoiding excessive expansion of the measured dataset; optimal coverage is achieved by just covering the measured dataset. Virtual density data is generated by selecting sound wave velocities and coefficients within a certain range. The sound wave velocity Vt ranges from 115 to 150, with a step size of 0.01, representing the range of sound wave velocity variation. The coefficient 'a' in the density formula ranges from 0.0653 to 0.0675, with a step size of 0.00005. The parameter b = 0.3923 keeps the exponent constant in the density formula, representing the exponential relationship between sound wave velocity and density. Using these parameters, density curves for different 'a' values ​​are generated using code to simulate the virtual data relationship between sound wave velocity and density. A simple merging of the measured and virtual datasets yields the fused dataset. In step 102, based on the coefficient of determination and the Larionov modified formula for mesh optimization, the optimal parameter a = 3.2 is fitted to this formula, and the following relationship exists between VSh and IRA: In the formula, IRA represents gamma rays or the radioactivity index; VSh represents the clay content. GR=I RA (GR Sh -GR cs )+GR cs (3); In the formula, GR is the actual gamma logging curve reading; GRcs is the gamma ray reading of clean sandstone, with a value of 20; GRSh is the gamma ray reading of shale, with a value of 140. To generate a virtual gamma-mud content dataset using the Larionov formula, the following principles must be followed: the generated virtual dataset should cover the measured gamma-mud content dataset, while maintaining the characteristics of the measured dataset without excessively expanding it; optimal coverage is achieved by just covering the measured dataset. Virtual gamma curve (GR) data is generated by selecting a certain range of mud content (VSh) values ​​and parameters. We defined the mud content range as 0.01 to 0.3, with a step size of 0.0001; the parameter 'a' in the Larionov formula ranges from 1 to 5.5, with a step size of 0.

03. Using this parameter, gamma ray curves at different 'a' values ​​were generated using code to simulate the virtual data relationship between mud content and gamma. A simple merging of the measured and virtual datasets yields the fused dataset.

2. The intelligent prediction method for logging curves of mud content in oil and gas reservoirs according to claim 1, characterized in that, In step 103, algorithms such as Random Forest (RFs), Gradient Boosting Decision Tree (XGBoost), Multilayer Perceptron (MLP), and Support Vector Machine (SVM) are used to train machine learning based on the fused dataset to intelligently predict the mud content.

3. The intelligent prediction method for logging curves of mud content in oil and gas reservoirs according to claim 2, characterized in that, In step 103, during model training, we selected three wells, J032, J061, and J034, as training wells. We used gamma ray, sonic waves, density, impedance, and velocity curves as input features, and clay content as the prediction target to train the model. Well J021 was used for model validation. The algorithm's prediction results were evaluated using the coefficient of determination (R²). 2 The results were evaluated using the coefficient of determination (C) and the correlation coefficient (r); the results showed that the RFs algorithm performed best, with a C coefficient of determination of 0.999 and a correlation coefficient of 0.

998.

4. The intelligent prediction method for logging curves of mud content in oil and gas reservoirs according to claim 3, characterized in that, Step 104: The accuracy of the intelligent prediction results of reservoir properties is evaluated using cross-validation, correlation coefficient, coefficient of determination (R2), and SHAP value method.

5. The intelligent prediction method for logging curves of clay content in oil and gas reservoirs according to claim 4, characterized in that, Cross-validation is a statistical method widely used for evaluating the performance of machine learning models. Its core idea is to validate the model's stability and generalization ability by repeatedly partitioning the dataset. In practice, cross-validation divides the dataset into several equal-sized subsets, selecting one subset each time as the test set and the remaining subsets as the training set. This process is repeated until each subset has been used as a test set. The most common method, K-fold cross-validation, divides the data into K subsets and performs K training and testing iterations. The final evaluation result is obtained by averaging the results of the multiple tests. The correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables, and it is usually expressed by the formula: In the formula, xi refers to the actual measured curve value, yi is the curve value predicted by the model, and and are the average values ​​of x and y, respectively; it is a statistical indicator used to measure the degree of correlation between two variables, usually expressed by the Pearson correlation coefficient. Its value is between -1 and 1: 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no linear correlation. The calculation method is to divide the covariance of the two variables by the product of their standard deviations. It reflects the strength and direction of the linear relationship between the variables. Through the correlation coefficient, we can understand the linear consistency between the model's predicted value and the actual value. The coefficient of determination is defined as follows: Where xi and yi are the original value and the machine learning model's predicted value, respectively, X is the average of xi, and n is the total number of samples. In R2, the sum of squares of the differences between the actual and predicted values ​​represents the total amount of information that our model has not captured, with the denominator being the amount of information contained in the true label. Therefore, both measure the proportion of 1 - the amount of information that our model has not captured to the amount of information contained in the true label, so the closer both are to 1, the better. If the result is 0, it indicates that the model fits very poorly; if the result is 1, it indicates that the model has no errors. SHAP, based on Shapley values ​​in game theory, provides a fair and consistent method for feature attribution. In this framework, the contribution of each feature is calculated by considering all possible combinations of features and their average impact under these combinations, thus ensuring that the results are not affected by the order of feature inputs.