Complex lithology identification generalization integration method based on game theory and machine learning optimization
Through the generalization integration method of complex lithology recognition based on game theory and machine learning optimization, the problems of insufficient data volume and high model cost in lithology recognition are solved, and efficient and accurate lithology recognition effect is achieved, which is suitable for lithology recognition in complex geological environments.
Patent Information
- Application Number
- CN202510481517.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology may lead to overfitting or insufficient generalization capabilities when there is a small amount of data or uneven sample in lithology recognition, and the stacked generalization model is costly and difficult to balance accuracy and efficiency.
The generalization integration method of complex lithologic recognition based on game theory and machine learning optimization is adopted. Through data preprocessing, feature selection, Bayesian hyperparameter optimization and game theory interpretive analysis, combined with SHAP interpretable method and random forest algorithm, the best input curve combination and basis model combination are selected to optimize model parameters to improve recognition accuracy and stability.
It significantly improves the accuracy and stability of lithology identification, is suitable for lithology identification in complex geological environments, and meets the actual needs of regional geological exploration and production.
Smart Images

Figure CN120354079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of oil and gas reservoir exploration and development, petroleum exploration, and artificial intelligence, and specifically relates to a generalization integration method for complex lithology identification optimized based on game theory and machine learning. Background Technique
[0002] Lithology identification occupies an important position in geophysical logging technology for oil and gas exploration and underground resource assessment tasks, directly affecting the construction of geological models and exploration decisions. Currently, lithology identification methods are mainly divided into two types: traditional evaluation and intelligent evaluation methods.
[0003] Traditional lithology evaluation mainly relies on the dual constraints of logging and core thin sections, and uses methods such as traditional experience, lithophysical models, and empirical cross-plot analysis to evaluate the lithology of the research horizon (CN119641332A, 2025; CN119616453A, 2025; CN118914257A, 2024). However, the lithology mapping relationship is usually implicit and non-linear, and traditional lithology methods rely on expert experience and simple statistical analysis methods, making it difficult to achieve ideal results in complex geological environments.
[0004] With the rapid development of artificial intelligence and machine learning technologies, intelligent lithology evaluation technologies have gradually emerged (CN117951476A, 2024; CN119152360A, 2024; CN119648828A, 2025). Many scholars have introduced machine learning methods into geophysical logging lithology identification. For example, Support Vector Machine (SVM) has been widely used in lithology classification and reservoir prediction (CN105388531B, 2015; CN119167227A, 2024). Random Forest has been successfully applied to lithology identification and geological modeling (CN119577589A, 2025).
[0005] However, it is often difficult for a single model to fully utilize the advantages of different algorithms. For this reason, the stacking method that integrates multiple models is introduced into lithology identification to further improve the generalization of the model, and good results have been achieved in tasks such as lithology classification and reservoir prediction (CN117932407A, 2024). Although the stacking method has improved the accuracy of lithology identification to a certain extent, there are still some challenges in practical applications: (1) Stacking often requires a large amount of training data to ensure that each base model can fully play its role. For some tasks with small data volume or unbalanced samples, it may lead to overfitting or insufficient generalization ability; (2) To build a good integrated model, it is crucial to reasonably select the base models and set relevant parameters. The relationship between the high-precision and diversity weights of the base models determines the success of the integrated model. (3) The stacking generalization model has a high computational cost, and how to balance accuracy and efficiency.
[0006] In view of the above defects, based on the existing well logging data of the oil reservoir, the present invention proposes a generalization integration strategy based on Stacking, which can handle the challenges faced by stacking generalization in geophysical well logging lithology identification through data preprocessing, feature selection, Bayesian hyperparameter optimization, and game theory interpretability analysis, significantly improving the accuracy and stability of lithology identification. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a generalization integration method for complex lithology identification based on game theory and machine learning optimization, which solves the problems existing in the prior art, such as overfitting or insufficient generalization ability for some tasks with small data volume or unbalanced samples, how to reasonably select base models and set relevant parameters, and the high computational cost of the stacking generalization model, and how to balance accuracy and efficiency.
[0008] The technical solution adopted by the present invention is as follows:
[0009] A generalization integration method for complex lithology identification based on game theory and machine learning optimization, comprising the following steps:
[0010] Step 1: Collect well logging data and core lithology data;
[0011] Step 2: Perform in-depth alignment on well logging curves and core lithology data;
[0012] Step 3: Use the SHAP interpretable method to perform feature selection on well logging curves, score the input curve combinations using the random forest algorithm, and select the best-performing input curve combination;
[0013] Step 4: Clean, standardize, and equalize the outliers in the logging data of the optimal input curve combination. Meanwhile, convert the logging data into a scale suitable for the machine learning model;
[0014] Step 5: Divide the logging data of the optimal input curve combination into a training set and a test set at a ratio of 7:3. The training set is used to train the model, and the test set is used to test the model performance; during the training and learning process, divide the training set into K - 1 folds and 1 fold, where the K - 1 folds are the training data and the 1 fold is the self - test data;
[0015] Step 6: Randomly select base model algorithms, which include UNet algorithm, TabNet algorithm, SVM algorithm, LSTM algorithm, KNN algorithm, and CatBoost algorithm;
[0016] Step 7: Use the Bayesian optimization method to optimize the parameters of the base model algorithms and embed the SHAP interpretable method at the same time;
[0017] Step 8: Use the SHAP interpretable method to select the combination of base model algorithms, score through the random forest algorithm, and select the combination of base model algorithms with the best performance;
[0018] Step 9: Take the prediction results of the combination of the best base model algorithms as new features and use the random forest meta - model algorithm to learn the new features;
[0019] Step 10: Compare and analyze the learning results of the random forest meta - model algorithm with the test set. Use the parameters Accuracy, f1 macro, and auc as the lithology identification accuracy indicators. If the identification effect of the random forest meta - model algorithm is better than that of the base model algorithm, promote and apply the improved method.
[0020] Preferably, the logging data in Step 1 includes compensated neutron, natural gamma, density, and well diameter.
[0021] Preferably, the SHAP interpretable method in Step 2 specifically measures the marginal contribution of each feature to the model output through the Shapley value.
[0022] Preferably, the calculation method of the Shapley value includes:
[0023] Step 2.1: For a given prediction f(χ), the Shapley value φ j The contribution for feature j is
[0024]
[0025] Among them, N is the set of all features, |N| is the total number of features, S is the feature subset that does not include feature j, f(S) is the prediction result of the model on the feature subset S, and f(S∪{j}) is the prediction result after adding feature j to the feature subset S;
[0026] Step 2.2: Accumulate the Shapley values of all features to obtain the final output of the model prediction. The formula is:
[0027]
[0028] Among them, f(x) is the prediction output of the model, φ0 is the global average prediction value, and φ j (x) is the contribution of feature j to the prediction result.
[0029] Preferably, in step 4, the cleaning of curve outliers is specifically as follows: By calculating the mean μ j and standard deviation σ j of each logging curve, eliminate the outliers that deviate more than 3 times the standard deviation from the mean. The calculation formula is:
[0030]
[0031] Among them, χ ij is the input logging feature;
[0032] The calculation formula for standardization in step 4 is:
[0033]
[0034] In the formula, χ i is the input logging feature, μ is the mean of the feature data, σ is the standard deviation of the feature data, and z i is the value after standardization;
[0035] The equalization process in step 4 is the SMOTE method.
[0036] Preferably, the SMOTE method includes the following steps:
[0037] Step 4.1: By selecting K minority-class lithology samples x i =(x i1 , x i2 ,..., x id ), the nearest neighbors x j1 , x j2 ,..., x jK in space;
[0038] Step 4.2: Generate a new synthetic sample x new =x i +λ·(x j-x i ), where λ ∈ [0, 1].
[0039] The beneficial effects of the complex lithology identification generalization integration method based on game theory and machine learning optimization of the present invention are as follows:
[0040] The present invention effectively improves the lithology identification accuracy in well logging data. When facing different lithology identification problems, the complex lithology identification generalization integration strategy based on game theory and machine learning optimization can meet the requirements of regional geological exploration and production practice, and has broad application and promotion prospects. Brief Description of the Drawings
[0041] Figure 1 is a flow chart of the present invention.
[0042] Figure 2 is a diagram showing the performance of the input feature combination of the present invention.
[0043] Figure 3 is a diagram for selecting the optimal input feature combination based on SHAP importance ranking of the present invention;
[0044] Figure 4 is a diagram comparing the data before and after balancing by the SMOTE method of the present invention.
[0045] Figure 5 is a diagram for optimizing the hyperparameters of the base model by Bayesian optimization of the present invention.
[0046] Figure 6 is a diagram for selecting the optimal base model combination before training the stacking generalization model of the present invention.
[0047] Figure 7 is a diagram comparing the lithology identification accuracy between the stacking generalization model and the single base model algorithm of the present invention.
[0048] Figure 8 is a diagram comparing the lithology identification accuracy of the stacking generalization model of the present invention with the average lithology identification accuracy of the base model algorithm. Detailed Embodiments
[0049] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0050] The following describes the specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0051] A generalization integration method for complex lithology identification optimized based on game theory and machine learning, comprising the following steps:
[0052] Step 1: Collect logging data and core lithology data;
[0053] Step 2: Perform depth alignment on logging curves and core lithology data;
[0054] Step 3: Use the SHAP explainable method to perform feature selection on logging curves, use the random forest algorithm to score the input curve combinations, and select the input curve combination with the best performance;
[0055] Step 4: Clean, standardize, and equalize the outliers of the logging data of the best input curve combination, and at the same time, convert the logging data into a scale suitable for the machine learning model;
[0056] Step 5: Divide the logging data of the best input curve combination into a training set and a test set according to 7:3. The training set is used to train the model, and the test set is used to test the model performance; during the training and learning process, the training set is divided into K - 1 folds and 1 fold, where the K - 1 folds are training data and the 1 fold is self - test data;
[0057] Step 6: Randomly select base model algorithms, and the base model algorithms include UNet algorithm, TabNet algorithm, SVM algorithm, LSTM algorithm, KNN algorithm, and CatBoost algorithm;
[0058] Step 7: Use the Bayesian optimization method to optimize the parameters of the base model algorithms, and at the same time, embed the SHAP explainable method;
[0059] Step 8: Use the SHAP explainable method to select the base model algorithm combination, score through the random forest algorithm, and select the base model algorithm combination with the best performance;
[0060] Step 9: Use the prediction results of the best base model algorithm combination as new features, and use the random forest meta - model algorithm to learn the new features;
[0061] Step 10: Compare and analyze the learning results of the random forest meta - model algorithm with the test set. Use the parameters Accuracy, f1 macro, and auc as lithology identification accuracy indicators. If the identification effect of the random forest meta - model algorithm is better than that of the base model algorithm, then promote and apply the improved method.
[0062] The logging data in Step 1 of this implementation plan includes compensated neutron, natural gamma ray, density, and well diameter.
[0063] The SHAP interpretable method in step 2 of this implementation scheme specifically measures the marginal contribution of each feature to the model output through the Shapley value.
[0064] The calculation method of the Shapley value of this embodiment includes:
[0065] Step 2.1: For a given prediction f(χ), the Shapley value φ j The contribution of feature j is
[0066]
[0067] Where N is the set of all features, |N| is the total number of features, S is the feature subset that does not contain feature j, f(S) is the prediction result of the model on feature subset S, and f(S∪{j}) is the prediction result after adding feature j to feature subset S;
[0068] Step 2.2: Add up the Shapley values of all features to get the final output predicted by the model. The formula is:
[0069]
[0070] Among them, f(x) is the predicted output of the model, φ0 is the global average predicted value, and φ j (x) is the contribution of feature j to the prediction result.
[0071] In step 4 of this implementation scheme, the outlier cleaning of the curve is specifically as follows: by calculating the mean μ of each logging curve j and standard deviation σ j , excluding outliers that deviate from the mean by more than 3 times the standard deviation, the calculation formula is:
[0072]
[0073] Among them, χ ij To input well logging characteristics;
[0074] The calculation formula for standardization in step 4 is:
[0075]
[0076] In the formula, χ i is the input logging feature, μ is the mean of the feature data, σ is the standard deviation of the feature data, z i is the standardized value;
[0077] The equalization process in step 4 is the SMOTE method.
[0078] The SMOTE method of this embodiment includes the following steps:
[0079] Step 4.1: By selecting K minority-class lithology samples \(x\) i =(x i1 , x i2 ,..., x id ), the nearest neighbors \(x\) j1 , x j2 ,..., x jK ;
[0080] Step 4.2: Generate a new synthetic sample \(x\) new =x i +λ·(x j -x i ), where λ ∈ [0, 1].
[0081] When implementing this implementation plan,
[0082] Step 1: Collect existing conventional logging data and core lithology data. The conventional data includes, but is not limited to, (AC), compensated neutron (CNL), natural gamma ray (GR), density (DEN), and caliper (CAL), etc.
[0083] Step 2: Perform depth alignment on the logging curves and core lithology data. The depth alignment method is commonly used in the field of geophysical logging, and the relevant formulas involved are common methods in the logging industry.
[0084] Step 3: Use the SHAP interpretable method to perform feature selection on the logging curves, score the input curve combinations using a random forest, and select the input curve combination with the best performance;
[0085] Among them, the marginal contribution of each feature to the model output is measured by the Shapley value. The Shapley value calculation process is as follows: (1) For a given prediction \(f(x)\), the contribution of the Shapley value \(\varphi\) j to feature j is:
[0086]
[0087] N is the set of all features, \(n\) is the total number of features, S is the feature subset that does not include feature j, \(f(S)\) is the prediction result of the model on the feature subset s, and \(f(S\cup\{j\})\) is the prediction result after adding feature j to the feature subset s;
[0088] Accumulate the Shapley values of all features to obtain the final output of the model prediction, and its formula is
[0089]
[0090] \(f(x)\) is the prediction output of the model, \(\varphi_0\) is the global average prediction value, and \(\varphi\) j(x) is the contribution of feature j to the prediction result. Figure 2 The representative uses random forest to score the input feature performance. When the number of input features is 14, the performance is the best. Figure 3 The SHAP interpretable method ranks the importance of each feature and finally selects the top 14 most important feature curves as input.
[0091] Step 4: Clean, standardize, and balance the best input curve combination logging data to remove curve outliers and convert the data into a scale suitable for the machine learning model. The purpose is to eliminate noise data, ensure the physical rationality of the input data, reduce the complexity of the input features, and reduce the risk of model overfitting.
[0092] Among them, outlier cleaning and data standardization are common methods in the well logging industry. The equalization process uses SMOTE technology to solve the problem of category imbalance and generate synthetic samples;
[0093] The curve normalization formula is:
[0094] Among them, x i is the input logging feature, μ is the mean of the feature data, σ is the standard deviation of the feature data, z i is the normalized value.
[0095] According to the 3σ criterion, the curve outliers are cleaned, and the formula is:
[0096]
[0097] The parameters in the formula have the same meaning as the standardized formula. Its purpose is to calculate the mean μ of each logging curve. j and standard deviation σ j , remove outliers that deviate from the mean by more than 3 times the standard deviation.
[0098] The SMOTE method is a method for processing unbalanced data sets by interpolating and synthesizing new samples based on the neighborhood information of minority lithology samples. i =(x i1 ,x i2 ,...,x id ) The spatial distance to the nearest neighbor x j1 ,x j2 ,...,x jK , generate a new synthetic sample x for each neighbor new =x i +λ·(x j -x i ), where λ∈[0,1]. Figure 4It shows that the SMOTE algorithm performs data balancing on the input feature curve after constant value cleaning and standardization. Based on the number of lithologies of the largest sample ( Figure 4 Diorite lithology in
[0099] Step 5: Divide the best combined logging data of input curves into a training set and a test set according to 7:3. The training set is used to train the model. During the training and learning process, the training set is further divided into K folds, where K - 1 folds are used as training data and the other 1 fold is used as self - verification data, while the test set is used to test the model performance;
[0100] Step 6: Randomly select base model algorithms. All existing algorithms can be selected as base model algorithms. In this case, 6 algorithms, namely UNet, TabNet, SVM, LSTM, KNN, and CatBoost, are randomly selected as base models;
[0101] Step 7: Use the Bayesian optimization method to optimize the parameters of each base model algorithm. The purpose is to scientifically and systematically screen algorithm hyperparameters to achieve the optimal performance of the algorithm. The data used in the optimization process is the training set data. In this step, the SHAP interpretable method is embedded to visualize the parameter selection, increasing the transparency and credibility of model parameter selection;
[0102] For example, Figure 5 in the case of the Unet base model algorithm, the best cross - validation score of the Unet base model algorithm is 0.8405, and the corresponding importance ranking of hyperparameters is: learning_rate>dropout>hidden_channels>num_layers. The optimal hyperparameter values are 0.0065, 0.1, 128, and 2 respectively.
[0103] Step 9: Use the SHAP interpretable method to select the combination of base model algorithms. Score through the random forest algorithm and select the combination of base model algorithms with the best performance, aiming to reduce the risk brought by the selection of base model algorithms and the impact on lithology recognition accuracy; Figure 6 It shows in
[0104] Step 9: Take the prediction results of the best combination of base model algorithms as new features and use the random forest meta - model algorithm to learn the new features;
[0105] Step 10: Compare and analyze the learning results of the random forest meta - model algorithm with the test set, using common parameters such as Accuracy, f1 macro, and auc as lithology recognition accuracy indicators.
Claims
1. A generalization integration method for complex lithology identification optimized based on game theory and machine learning, characterized in that, It includes the following steps: Step 1: Collect logging data and core lithology data; Step 2: Perform depth alignment on the logging curves and core lithology data; Step 3: Use the SHAP interpretable method to perform feature selection on the logging curves, use the random forest algorithm to score the input curve combinations, and select the input curve combination with the best performance; Step 4: Clean, standardize, and equalize the outliers of the logging data for the best input curve combination. At the same time, convert the logging data into a scale suitable for the machine learning model; Step 5: Divide the logging data of the best input curve combination into a training set and a test set according to 7:
3. The training set is used to train the model, and the test set is used to test the model performance; during the training and learning process, the training set is divided into K - 1 folds and 1 fold, where the K - 1 folds are training data and the 1 fold is self - test data; Step 6: Randomly select a base model algorithm, and the base model algorithms include UNet algorithm, TabNet algorithm, SVM algorithm, LSTM algorithm, KNN algorithm, and CatBoost algorithm; Step 7: Use the Bayesian optimization method to optimize the parameters of the base model algorithm, and at the same time, embed the SHAP interpretable method; Step 8: Use the SHAP interpretable method to select the base model algorithm combination, score through the random forest algorithm, and select the base model algorithm combination with the best performance; Step 9: Use the prediction result of the best base model algorithm combination as a new feature, and use the random forest meta - model algorithm to learn the new feature; Step 10: Compare and analyze the learning result of the random forest meta - model algorithm with the test set. Use the parameters Accuracy, f1 macro, and auc as the lithology identification accuracy indicators. If the identification effect of the random forest meta - model algorithm is better than that of the base model algorithm, then promote and apply the improved method.
2. The generalization integration method for complex lithology identification optimized based on game theory and machine learning according to claim 1, wherein The logging data in Step 1 includes compensated neutron, natural gamma, density, and well diameter.
3. The generalization integration method for complex lithology identification optimized based on game theory and machine learning according to claim 1, characterized in that The SHAP interpretable method in Step 2 specifically measures the marginal contribution of each feature to the model output through the Shapley value.
4. The generalization integration method for complex lithology identification optimized based on game theory and machine learning according to claim 3, wherein The calculation method of the Shapley value includes: Step 2.1: For the given prediction f(χ), the contribution of the Shapley value φ j for feature j is where N is the set of all features, |N| is the total number of features, S is a feature subset that does not include feature j, f(S) is the prediction result of the model on the feature subset S, and f(s∪{j}) is the prediction result after adding feature j to the feature subset S; Step 2.2: Accumulate the Shapley values of all features to obtain the final output of the model prediction, and its formula is: where f(x) is the predicted output of the model, φ0 is the global average prediction value, and φ j (x) is the contribution of feature j to the prediction result.
5. The generalization integration method for complex lithology identification optimized based on game theory and machine learning according to claim 1, characterized in that The cleaning of curve outliers in step 4 is specifically as follows: by calculating the mean μ of each logging curve j and the standard deviation σ j , removing the outlier points that deviate more than 3 times the standard deviation from the mean, and its calculation formula is: where χ ij is the input logging feature; The calculation formula for standardization in Step 4 is: where χ i is the input logging feature, μ is the mean of the feature data, σ is the standard deviation of the feature data, and z i is the value after standardization; The equalization process in Step 4 is the SMOTE method.
6. The generalization integration method for complex lithology identification optimized based on game theory and machine learning according to claim 5, and the SMOTE method includes the following steps: Step 4.1: By selecting K minority-class lithology samples x i =(x i1 , x i2 ,..., x id ), the nearest neighbors x j1 , x j2 ,..., x jK in terms of spatial distance; Step 4.2: Generate a new synthetic sample \(x\) for each neighbor new = \(x\) i + \(\lambda\cdot(x\) j - \(x\) i ), where \(\lambda\in[0, 1]\).
Citation Information
Patent Citations
Lithology identification method based on support vector regression machine and Kernel Fisher discriminant analysis
CN105388531A
Lithology identification and prediction method based on Stacking algorithm
CN117932407A
Construction method of lithology identification model of shale oil reservoir and lithology identification method
CN117951476A
Method for identifying lithology of metamorphic bed rock
CN118914257A
Multi-modal lithology identification method and system
CN119152360A