An intelligent identification method for the lithology of ore-bearing strata rocks in sandstone-type uranium deposits

By using random forest regression model and XGBoost model for logging curve missing value filling and lithology identification in sandstone-type uranium ore-bearing rocks, the problems of poor logging curve reconstruction and serious misjudgment in lithology identification in the existing technology are solved, and more efficient and accurate lithology identification is achieved.

CN119441895BActive Publication Date: 2025-06-17NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510036870.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-17
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art has problems of poor results and serious misjudgment in the logging curve reconstruction and lithology identification of sandstone-type uranium ore-bearing rocks, especially in the absence or distortion of deep lithology identification and logging data.

Method used

The random forest regression model and XGBoost model were used to fill missing values ​​of logging curves and lithologic recognition, and the best hyperparameters were found through Bayesian optimization algorithm, and model evaluation was carried out in combination with multiple evaluation indicators.

Benefits of technology

It improves the accuracy and efficiency of logging curve reconstruction and lithology identification, reduces lithology misjudgment caused by abnormal logging data, and is more in line with the actual geological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441895B_ABST
    Figure CN119441895B_ABST
Patent Text Reader

Abstract

This application belongs to the field of data processing and provides an intelligent identification method for the lithology of ore-bearing strata in sandstone-type uranium deposits, including: establishing a random forest regression model and an XGBoost model; obtaining a logging curve sample set and lithology data, dividing the logging curve sample set into original data and a validation set, and generating a first training set and a first test set; training to obtain a qualified random forest regression model; obtaining filled data based on the qualified random forest regression model; training to obtain a qualified second XGBoost model and a third XGBoost model based on the original data, filled data, and lithology data; and obtaining the lithology of the ore-bearing strata in the identified uranium deposit based on the qualified second XGBoost model and third XGBoost model. The lithology identified in this application is more in line with the actual geological situation, has higher accuracy, and reduces the misjudgment of lithology caused by abnormal GR logging data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to an intelligent recognition method for the lithology of ore-bearing strata rocks in sandstone-type uranium deposits. Background Art

[0002] Sandstone-type uranium deposits are a key performance metal and mineral resource. In existing mineral exploration, geophysical logging methods (hereinafter referred to as logging) are widely used. Logging is also an effective method for understanding and obtaining information on underground geological structures and reservoir characteristics. Logging data has important value in reservoir evaluation, and can be used for stratigraphic structure division, sedimentary facies identification, lithology identification, etc. Among them, lithology identification is of great significance in the study of uranium deposits. By identifying different lithologies under sedimentation, metamorphism or volcanic action, the main host rocks of uranium deposits (such as sandstone, black shale) and their associated enrichment rock types can be determined. At the same time, lithology identification helps to analyze the tectonic characteristics that control the formation of uranium deposits, such as fault zones, folds and magmatic activities, thereby revealing the influence of tectonic activities on the enrichment of uranium deposits and providing a basis for determining potential geological areas of uranium deposits.

[0003] Digital mapping is an important means of lithology identification. However, due to the depth limitation of surface geological mapping, it can only reflect the distribution of shallow lithology, and it is difficult to identify deep lithology. In the covered area, it is even more impossible to give the changes in the underlying lithology. To overcome the shortcomings of surface geological mapping, technologies such as geophysics and remote sensing have gradually been introduced into lithology identification and have made great progress. However, at present, there are still technical problems such as low lithology identification accuracy and serious lithology misjudgment.

[0004] At the same time, due to factors such as poor drilling conditions, instrument failures, logging condition differences, improper storage, etc., it is easy to cause data loss or distortion in some well sections of logging data, which affects the indication of lithology and fluids, brings difficulties to logging evaluation and interpretation, and even abandons some logging data due to exploration cost considerations. Given that re-logging is limited by cost and wellbore conditions, the need for low-cost reconstruction of missing logging curves is becoming increasingly urgent. Without increasing additional manpower and economic costs, in order to solve the problem of missing logging curves, those skilled in the art have tried various methods, such as traditional empirical models based on rock physics, multiple regression analysis and physical model inversion to generate artificial logging curves. However, these methods usually oversimplify the formation information, rely on subjective experience, and are difficult to effectively reflect the complex non-linear relationship between logging data. In the face of the challenges of strong formation heterogeneity and complex downhole conditions, traditional linear methods perform poorly.

[0005] Therefore, exploring and establishing a more efficient logging curve reconstruction method and improving the accuracy of lithology identification are crucial for improving the regional logging database and the accuracy of geophysical exploration interpretation. Summary of the Invention

[0006] In view of the problems existing in the prior art, the present application provides an intelligent recognition method for the lithology of ore-bearing layer rocks in sandstone-type uranium deposits, so as to solve the technical problems of poor reconstruction effect of logging curves and inaccurate lithology recognition in the prior art.

[0007] In a first aspect, an embodiment of the present application provides an intelligent recognition method for the lithology of ore-bearing layer rocks in sandstone-type uranium deposits, and the method includes:

[0008] Step S1: Establish a random forest regression model and an XGBoost model;

[0009] Step S2: Obtain a logging curve sample set and lithology data, divide the logging curve sample set into original data and a validation set, and generate a first training set and a first test set based on the original data;

[0010] Step S3: Train the random forest regression model based on the first training set and the first test set to obtain a qualified random forest regression model; based on the validation set and the qualified random forest regression model, obtain a predicted filled GR curve;

[0011] Step S4: Fill the predicted filled GR curve into the original data to obtain filled data;

[0012] Step S5: Generate a second training set, a second test set, a third training set, and a third test set based on the original data, the filled data, and the lithology data;

[0013] Step S6: Train the XGBoost model based on the second training set, the second test set, the third training set, and the third test set respectively to obtain a qualified second XGBoost model and a qualified third XGBoost model;

[0014] Step S7: Input the original data into the qualified second XGBoost model to obtain the lithology of the ore-bearing layer of the uranium deposit after original recognition; input the filled data into the qualified third XGBoost model to obtain the lithology of the ore-bearing layer of the uranium deposit after filled recognition.

[0015] In a possible implementation manner, the hyperparameters of the random forest regression model include the number of trees n_estimators, the maximum depth max_depth, the minimum number of samples for splitting min_samples_split, the minimum number of samples in leaf nodes min_samples_leaf, and the maximum number of features max_features;

[0016] The evaluation indexes of the random forest regression model include the mean absolute error MAE, the mean square error MSE, the root mean square error RMSE, and the R-squared value of the prediction model fitting accuracy 。

[0017] In a possible implementation, the hyperparameters of the XGBoost model include the number of trees n_estimators, the maximum depth max_depth, the learning rate learning_rate, the subsampling parameter subsample, the proportion of the features of the logging curve samples used when training each tree to all the features colsample_bytree, the minimum weight sum of a leaf node min_child_weight, and the minimum loss reduction gamma required to split a node;

[0018] The evaluation metrics of the XGBoost model include accuracy, precision, recall, the harmonic mean of precision and recall, and the confusion matrix.

[0019] In a possible implementation, the step S2 is specifically as follows:

[0020] Obtain a logging curve sample set and lithology data;

[0021] Classify the logging curve samples in the logging curve sample set according to the data integrity of the logging curve samples in the logging curve sample set, where the set of logging curve samples with complete data is the original data, and the set of logging curve samples with missing data is the validation set;

[0022] Take the apparent resistivity curve, well diameter curve, well temperature curve, spontaneous potential curve, and triple lateral resistivity curve of the logging curve samples in the original data as the first input data set, take the natural gamma logging curve as the first label data set, pair the data in the first input data set with the corresponding natural gamma logging curve in the first label data set to obtain the first logging data pair set; divide the first logging data pair set into a first training set and a first test set at a ratio of 8:2.

[0023] In a possible implementation, the step S3 is specifically as follows:

[0024] Define the hyperparameter range of the random forest regression model. Set the range of the number of trees n_estimators to (50, 500), the range of the maximum depth to max_depth (5, 30), the range of the minimum number of samples for splitting to min_samples_split (2, 10), the range of the minimum number of samples in a leaf node to min_samples_leaf (1, 5), and the range of the maximum number of features to max_features (0.5, 1.0), and set the random forest regression model as the default;

[0025] Initialize the Bayesian optimization algorithm. The evaluator selects the random forest default model, the search range selects the hyperparameter range of the random forest regression model, the number of search iterations selects 200, the cross-validation selects 5, the scoring selects the negative mean squared error, and the random number selects 66;

[0026] Use the Bayesian optimization algorithm to search the first training set. After obtaining the best hyperparameters, use the first test set to evaluate the performance of the random forest regression model to obtain a qualified random forest regression model;

[0027] Input the validation set into the qualified random forest regression model to obtain the predicted filled GR curve.

[0028] In a possible implementation, step S4 is specifically: Delete the abnormal part of the natural gamma log curve in the log curve sample of the original data, and then supplement the corresponding part of the predicted filled GR curve to the natural gamma log curve after deleting the abnormal part to obtain the filled data.

[0029] In a possible implementation, step S5 is specifically:

[0030] Use the apparent resistivity curve, natural gamma curve, well diameter curve, well temperature curve, spontaneous potential curve, and triple lateral resistivity curve of the original data and the filled data as the second input data set and the third input data set respectively, and use the lithology data as the label data set;

[0031] Pair the data in the second input data set with the corresponding lithology data in the label data set to obtain the second log data pair set; pair the data in the third input data set with the corresponding lithology data in the label data set to obtain the third log data pair set;

[0032] Divide the second log data pair set and the third log data pair set into the second training set, the second test set, the third training set, and the third test set according to the ratio of 7:3 respectively;

[0033] The lithology data includes mudstone, siltstone, fine sandstone, medium sandstone, coarse sandstone, and conglomerate.

[0034] In a possible implementation, step S6 is specifically:

[0035] Preprocess the second training set, the second test set, the third training set, and the third test set. Set the XGBoost model to not use the target variable for label encoding by default, and the evaluation metric of the XGBoost model is the log loss;

[0036] Define the hyperparameter range of the XGBoost model. Set the range of the number of trees n_estimators to (50, 500), the range of the maximum depth max_depth to (3, 10), the range of the learning rate learning_rate to (0.01, 0.3, 'log-uniform'), the range of the subsampling parameter subsample to (0.5, 1.0), the range of the proportion of features of the logging curve samples used when training each tree to the total features colsample_bytree to (0.5, 1.0), the range of the minimum sum of weights of a leaf node min_child_weight to (1, 10), and the range of the minimum loss reduction gamma required to split a node to (0, 5).

[0037] Initialize the XGBoost Bayesian algorithm. Select the XGBoost model for the estimator, the hyperparameter range of the XGBoost model for the search range, 50 for the number of search iterations, 3 for cross-validation, accuracy for scoring, and 45 for the random number.

[0038] Use the XGBoost Bayesian algorithm to search the second training set and the third training set respectively. After obtaining the optimal hyperparameters, perform performance evaluation on the second test set and the third test set respectively to obtain a qualified second XGBoost model and a qualified third XGBoost model.

[0039] In a possible implementation, the preprocessing is to use standardization to perform a unified standard normal distribution on the logging curve samples.

[0040] In a possible implementation, when evaluating the performance of the random forest regression model, the R-squared value indicating the accuracy of the prediction model fitting is

[0041] Based on the above invention content, compared with the prior art, the present application has achieved the following technical effects:

[0042] (1) In the present invention, by standardizing all input logging curves and fitting the missing values of GR logging data with a random forest regression model, a predicted filled GR curve is obtained. The predicted filled GR curve is used to fill the abnormal part of the natural gamma curve in the original data to obtain filled data. The lithology is identified using the filled data, and the identified lithology is more in line with the actual geological situation, with higher accuracy, and successfully reduces the misjudgment of lithology caused by abnormal GR logging data.

[0043] (2) By selecting reasonable sample data, the present invention reconstructs the data using a random forest regression model and identifies the lithology using an XGBoost model, thereby improving the computational efficiency of data reconstruction and lithology identification.

[0044] (3) During the model training process of the present invention, an advanced Bayesian optimization algorithm is used to find the optimal hyperparameters, which has obvious advantages over the low efficiency of traditional grid search. Moreover, during the parameter adjustment process, multiple evaluation metrics are combined for evaluation, further improving the accuracy of the hyperparameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic flowchart of an intelligent identification method for the lithology of rocks in an ore-bearing layer of a sandstone-type uranium deposit provided by an embodiment of the present application;

[0047] Figure 2 It is a schematic diagram of filling data provided by an embodiment of the present invention;

[0048] Figure 3 It is a comparison chart of the classification results of the original data and the filled data of Well Z1 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To better understand the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0051] Refer to Figure 1 , which is a schematic flowchart of an intelligent identification method for the lithology of rocks in an ore-bearing layer of a sandstone-type uranium deposit provided by an embodiment of the present invention. As Figure 1 shown, the method specifically includes:

[0052] Step S1: Establish a random forest regression model and an XGBoost model.

[0053] The hyperparameters of the random forest regression model include the number of trees n_estimators, the maximum depth max_depth, the minimum number of samples for splitting min_samples_split, the minimum number of samples in a leaf node min_samples_leaf, and the maximum number of features max_features.

[0054] The evaluation metrics of the random forest regression model include the mean absolute error MAE, the mean squared error MSE, the root mean squared error RMSE, and the R-squared value representing the accuracy of the prediction model fit. . The smaller the mean absolute error MAE, the smaller the average absolute value of the prediction error of the model, and the higher the accuracy of the model. The smaller the mean squared error MSE, the closer the predicted value is to the true value. The smaller the root mean squared error RMSE, the smaller the prediction error of the model, and the closer the predicted value is to the true value. The closer it is to 1, the better the fitting effect of the random forest regression model and the better the comprehensive performance.

[0055] The hyperparameters of the XGBoost model include the number of trees n_estimators, the maximum depth max_depth, the learning rate learning_rate, the subsampling parameter subsample, the proportion of the features of the well logging curve samples used when training each tree to all features colsample_bytree, the minimum weight sum of a leaf node min_child_weight, and the minimum loss reduction gamma required to split a node.

[0056] The evaluation metrics of the XGBoost model include accuracy, precision, recall, the harmonic mean of precision and recall, and the confusion matrix.

[0057] Accuracy: Represents the proportion of correctly predicted samples in the total samples.

[0058] Precision: Represents the proportion of samples predicted as positive classes that are actually positive classes.

[0059] Recall: Represents the proportion of samples that are actually positive classes and are correctly predicted as positive classes.

[0060] The harmonic mean of precision and recall (F1-score): Represents the harmonic mean of precision and recall, and is suitable for use when precision and recall need to be balanced.

[0061] Confusion Matrix: Represents the detailed distribution of the prediction results shown in matrix form.

[0062] Step S2: Obtain a logging curve sample set and lithology data, divide the logging curve sample set into original data and a validation set, and generate a first training set and a first test set based on the original data. Specifically as follows:

[0063] Step S21: Obtain a logging curve sample set and lithology data.

[0064] Step S22: Classify the logging curve samples in the logging curve sample set according to the data integrity of the logging curve samples in the logging curve sample set, where the set of logging curve samples with complete data is the original data, and the set of logging curve samples with missing data is the validation set.

[0065] Step S23: Use the apparent resistivity curve (RES), well diameter curve (WD), well temperature curve (WT), spontaneous potential curve (SP), and triple lateral resistivity curve (TAR) of the logging curve samples in the original data as the first input data set, and use the natural gamma logging curve (GR) as the first label data set. Pair the data in the first input data set with the corresponding natural gamma logging curve in the first label data set to obtain a first set of logging data pairs; divide the first set of logging data pairs into a first training set and a first test set at a ratio of 8:2.

[0066] In this embodiment, the data used is derived from the rock lithology and the corresponding logging curve sample set provided in the northern Songliao Basin. The logging curve sample set consists of 3 uranium mineralization industrial holes, with well numbers Z1, Z2, and Z3 respectively. The data of the mineralized section in the lower member of the Yaojia Formation is selected as the logging curve samples, totaling 3375 groups of logging curve samples. These logging curve samples cover various rock types, including mudstone, siltstone, fine sandstone, medium sandstone, coarse sandstone, and conglomerate, among which sandstone is a good uranium-bearing reservoir. Based on the above data, the distributions of the first training set, the first test set, and the validation set are shown in Table 1.

[0067] Table 1:

[0068]

[0069] Step S3: Train a random forest regression model based on the first training set and the first test set to obtain a qualified random forest regression model after verification; based on the validation set and the qualified random forest regression model after verification, obtain a predicted filled GR curve. Specifically include:

[0070] Step S31: Define the hyperparameter range of the random forest regression model. Set the range of the number of trees n_estimators to (50, 500), the range of the maximum depth to max_depth (5, 30), the range of the minimum number of samples for splitting to min_samples_split (2, 10), the range of the minimum number of samples in leaf nodes to min_samples_leaf (1, 5), and the range of the maximum number of features to max_features (0.5, 1.0). Set the random forest regression model as the default.

[0071] Step S32: Initialize the Bayesian optimization algorithm. Select the random forest default model as the estimator, select the hyperparameter range of the random forest regression model as the search_spaces, select 200 for the number of search iterations (n_iter), select 5 for cross-validation (cv), select negative mean squared error (neg_mean_squared_eeror) for scoring, and select 66 for the random number.

[0072] Step S33: Use the Bayesian optimization algorithm to search the first training set. After obtaining the optimal hyperparameters, use the first test set to evaluate the performance of the random forest regression model to obtain a qualified random forest regression model.

[0073] When evaluating the performance of the random forest regression model, use the R-squared value as the first evaluation index of the random forest regression model. In this embodiment, when adjusting the logging curve samples of the random forest regression model of Well Z1 to the optimal hyperparameters, the evaluation index values of the random forest regression model are shown in Table 2.

[0074] Table 2:

[0075]

[0076] Step S34: Input the validation set into the qualified random forest regression model to obtain the predicted filled GR curve.

[0077] Step S4: Fill the predicted filled GR curve into the original data to obtain the filled data. Specifically:

[0078] Delete the abnormal part of the natural gamma logging curve (GR) in the logging curve samples of the original data, and then supplement the part corresponding to the abnormal part of the natural gamma logging curve (GR) in the predicted filled GR curve to the natural gamma logging curve (GR) after deleting the abnormal part to obtain the filled data.

[0079] To illustrate the effect of the filled data, seeFigure 2 , which is a schematic diagram of the filling data of Well Z1 provided by an embodiment of the present invention. As Figure 2 shown, there are three color waveforms in the figure. The red dashed waveform is the abnormal part of the natural gamma log curve (GR) in the log curve sample of the original data. The black waveform is the natural gamma log curve (GR) after deleting the abnormal part. The green waveform on the right is the predicted filling GR curve, and the green waveform on the left is the natural gamma log curve (GR) of the filled data. From Figure 2 it can be seen that in Well Z1, there is only one obvious peak in the abnormal value of the natural gamma log curve, but there are different wave tips, and the thickness across is relatively large. The highest GR abnormal value is up to 4033.33 API (not shown in Figure 2 due to the too large abnormal value). The missing data is mainly in the first peak and a few local abnormal values, corresponding to the well depths of 545.75 - 546.4m, 551.1 - 551.6m, 556.65 - 563.2, 565.5 - 565.85, 566.2 - 580.2 to 591.15 - 593.25, with a total of 24.15m. The filled data within the peak shows completely different data volumes, with peaks and local minima. The predicted filled data conforms to the complex underground geological environment, and the fitting degree is nearly similar.

[0080] Step S5: Generate a second training set, a second test set, a third training set, and a third test set based on the original data, the filling data, and the lithology data. Specifically:

[0081] Take the apparent resistivity curve (RES), natural gamma curve (GR), well diameter curve (WD), well temperature curve (WT), spontaneous potential curve (SP), and triple lateral resistivity curve (TAR) of the original data and the filling data as the second input data set and the third input data set respectively, and take the lithology data as the label data set;

[0082] Pair the data in the second input data set with the corresponding lithology data in the label data set to obtain a second set of log data pairs; pair the data in the third input data set with the corresponding lithology data in the label data set to obtain a third set of log data pairs;

[0083] Divide the second set of log data pairs and the third set of log data pairs into a second training set, a second test set, a third training set, and a third test set according to the ratio of 7:3 respectively.

[0084] The lithology data includes mudstone, siltstone, fine sandstone, medium sandstone, coarse sandstone, and conglomerate.

[0085] Step S6: Train the XGBoost model based on the second training set, second test set, third training set, and third test set to obtain a qualified XGBoost model through testing.

[0086] Step S6: Respectively train the XGBoost model based on the second training set, second test set, third training set, and third test set to obtain a second qualified XGBoost model and a third qualified XGBoost model through testing. Specifically:

[0087] Step S61: Preprocess the second training set, second test set, third training set, and third test set. Set the default setting of the XGBoost model to not use the target variable for label encoding, and the evaluation metric of the XGBoost model is the log loss.

[0088] The preprocessing is to use standardization to perform a unified standard normal distribution on the logging curve samples.

[0089] Step S62: Define the hyperparameter range of the XGBoost model. Set the range of the number of trees n_estimators to (50, 500), the range of the maximum depth max_depth to (3, 10), the range of the learning rate learning_rate to (0.01, 0.3, 'log-uniform'), the range of the subsampling parameter subsample to (0.5, 1.0), the range of the proportion of the features of the logging curve samples used for training each tree to the total features colsample_bytree to (0.5, 1.0), the range of the minimum sum of weights of a leaf node min_child_weight to (1, 10), and the range of the minimum loss reduction gamma required to split a node to (0, 5).

[0090] Step S63: Initialize the XGBoost Bayesian algorithm. Select the XGBoost model for the estimator, select the hyperparameter range of the XGBoost model for the search_spaces, select 50 for the number of search iterations (n_iter), select 3 for the cross-validation (cv), select accuracy for the scoring, and select 45 for the random number.

[0091] Step S64: Use the XGBoost Bayesian algorithm to search the second training set and the third training set respectively. After obtaining the best hyperparameters, perform performance evaluations on the second test set and the third test set respectively to obtain a second qualified XGBoost model and a third qualified XGBoost model through testing.

[0092] Step S7: Input the original data into the second XGBoost model that has passed the test to obtain the lithology of the ore-bearing layer of the uranium deposit after original identification; input the filled data into the third XGBoost model that has passed the test to obtain the lithology of the ore-bearing layer of the uranium deposit after filled identification.

[0093] See Figure 3 , which is a comparison chart of the classification results of the original data and the filled data of Well Z1 provided by the embodiment of the present invention. Among them, Figure 3 The left figure in the middle is the classification result of the original data of Well Z1, Figure 3 The right figure in the middle is the classification result of the filled data of Well Z1. As Figure 3 shown in the left figure in the middle, the lithology identification accuracy of the original logging curves to be processed in Well Z1 is 97.3%. Among the input samples, there are 53 mudstones, 32 siltstones, 138 fine sandstones, 143 medium sandstones, and 16 coarse sandstones. The total number of samples is 382. The samples correctly predicted on the main diagonal are 51 mudstones, 30 siltstones, 135 fine sandstones, 1418 medium sandstones, and 15 coarse sandstones, totaling 372. From Figure 3 the left figure in the middle, it can be obtained that the mudstone and fine sandstone perform best in the model, and the other three lithologies perform the same, namely siltstone, medium sandstone, and coarse sandstone.

[0094] As Figure 3 shown in the right figure in the middle, the lithology identification accuracy of the filled data to be processed in Well Z1 is 98.6%. Among the input samples, there are 66 mudstones, 31 siltstones, 144 fine sandstones, 130 medium sandstones, and 11 coarse sandstones. The total number of samples is 382. The samples correctly predicted on the main diagonal are 65 mudstones, 31 siltstones, 143 fine sandstones, 127 medium sandstones, and 11 coarse sandstones, totaling 377. From Figure 3 the right figure in the middle, it can be obtained that the coarse sandstone performs best in the model, followed by the fine sandstone, the medium sandstone and the mudstone are worse, and the worst is the siltstone.

[0095] In the embodiment of the present application, by standardizing all input logging curves and fitting the missing values of GR logging data with a random forest regression model, a predicted filled GR curve is obtained. The predicted filled GR curve is used to fill the abnormal part of the natural gamma curve in the original data to obtain filled data. The filled data is used for lithology identification, and the identified lithology is more in line with the actual geological situation, with higher accuracy, and successfully reduces the misjudgment of lithology caused by abnormal GR logging data. By selecting reasonable sample data, using a random forest regression model for data reconstruction, and using an XGBoost model for lithology identification, the computational efficiency of data reconstruction and lithology identification is improved. During the model training process, an advanced Bayesian optimization algorithm is used to find the best hyperparameters, which has obvious advantages over the low efficiency of traditional grid search. Moreover, during the parameter adjustment process, multiple evaluation indicators are combined for evaluation, further improving the accuracy of hyperparameters.

[0096] The above are only specific embodiments of the present application. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An intelligent identification method for rock lithology of sandstone-type uranium ore-bearing strata, characterized in that: include: S1: Establish a random forest regression model and an XGBoost model; the hyperparameters of the random forest regression model include the number of trees n_estimators, the maximum depth max_depth, the minimum number of sample splits min_samples_split, the minimum number of leaf nodes min_samples_leaf and the maximum number of features max_features; the evaluation indicators of the random forest regression model include the mean absolute error MAE, the mean square error MSE, the root mean square error RMSE and the R-square value of the accuracy of the prediction model fitting ; When evaluating the performance of the random forest regression model, the R-square value of the prediction model fitting accuracy is used As the first evaluation metric for random forest regression models; S2: Obtain the well logging curve sample set and lithology data of the mineralized section of the lower section of the Yaojia Formation at a depth of 530 meters to 590 meters; classify the well logging curve samples in the well logging curve sample set according to the data completeness of the well logging curve samples in the well logging curve sample set, the set of well logging curve samples with complete data is the original data, and the set of well logging curve samples with missing data is the validation set; Taking the apparent resistivity curve, well diameter curve, well temperature curve, natural potential curve and three-lateral resistance curve of the logging curve samples in the original data as the first input data set, taking the natural gamma logging curve as the first label data set, pairing the data of the first input data set with the corresponding natural gamma logging curve in the first label data set to obtain a first logging data pair set; dividing the first logging data pair set into a first training set and a first test set in a ratio of 8:2; The lithology data include mudstone, siltstone, fine sandstone, medium sandstone, coarse sandstone, and conglomerate; S3: Define the hyperparameter range of the random forest regression model, set the number of trees n_estimators to (50, 500), the maximum depth to max_depth to (5, 30), the minimum number of sample splits to min_samples_split to (2, 10), the minimum number of leaf node samples to min_samples_leaf to (1, 5), and the maximum number of features to max_features to (0.5, 1.0), and set the random forest regression model to default; Initialize the Bayesian optimization algorithm, select the random forest default model for the evaluator, the hyperparameter range of the random forest regression model for the search range, 200 for the number of search iterations, 5 for cross validation, negative mean square error for scoring, and 66 for the random number; The first training set is searched using the Bayesian optimization algorithm, and after obtaining the optimal hyperparameters, the performance of the random forest regression model is evaluated using the first test set to obtain a verified and qualified random forest regression model; The validation set is input into a verified random forest regression model to obtain a predicted filling GR curve; S4: deleting the abnormal part of the natural gamma logging curve in the logging curve sample of the original data, and adding the part of the predicted filled GR curve corresponding to the abnormal part of the natural gamma logging curve to the natural gamma logging curve from which the abnormal part is deleted, to obtain filled data; S5: generating a second training set, a second test set, a third training set, and a third test set based on the original data, the filling data, and the lithology data; S6: training the XGBoost model based on the second training set, the second test set, the third training set, and the third test set, respectively, to obtain a second XGBoost model that passes the test and a third XGBoost model that passes the test; S7: Input the original data into the second XGBoost model that has passed the test to obtain the lithology of the ore-bearing layer of the uranium deposit after the original identification; input the filled data into the third XGBoost model that has passed the test to obtain the lithology of the ore-bearing layer of the uranium deposit after the filled identification.

2. The intelligent identification method of rock lithology of sandstone-type uranium ore-bearing strata according to claim 1 is characterized in that: The hyperparameters of the XGBoost model include the number of trees n_estimators, the maximum depth max_depth, the learning rate learning_rate, the subsampling parameter subsample, the proportion of the features of the logging curve samples used to train each tree to the total features colsample_bytree, the minimum weight of a leaf node and min_child_weight, and the minimum loss reduction gamma required to split a node; The evaluation indicators of the XGBoost model include accuracy, precision, recall, harmonic mean of precision and recall, and confusion matrix.

3. The intelligent identification method of rock lithology of sandstone-type uranium ore-bearing strata according to claim 1 is characterized in that: Step S5 is specifically as follows: The apparent resistivity curve, natural gamma curve, well diameter curve, well temperature curve, natural potential curve, and three-lateral resistance curve of the original data and the filled data are used as the second input data set and the third input data set respectively, and the lithology data is used as the label data set; Pairing the data in the second input data set with the corresponding lithology data in the label data set to obtain a second well logging data pair set; Pairing the data in the third input data set with the corresponding lithology data in the label data set to obtain a third well logging data pair set; The second logging data pair set and the third logging data pair set are divided into a second training set, a second test set and a third training set, a third test set in a ratio of 7:3 respectively.

4. The intelligent identification method of rock lithology of sandstone-type uranium ore-bearing strata according to claim 1 is characterized in that: Step S6 is specifically as follows: The second training set, the second test set, the third training set, and the third test set are preprocessed, and the XGBoost model is set by default to not use the target variable for label encoding, and the evaluation indicator of the XGBoost model is the logistic loss; Define the hyperparameter range of the XGBoost model, set the number of trees n_estimators to (50, 500), the maximum depth max_depth to (3, 10), the learning rate learning_rate to (0.01, 0.3, 'log-uniform'), the subsampling parameter subsample to (0.5, 1.0), the proportion of the features of the logging curve samples used to train each tree to the total features colsample_bytree to (0.5, 1.0), the minimum weight of a leaf node and min_child_weight to (1, 10), and the minimum loss reduction gamma required to split a node to (0, 5); Initialize the XGBoost Bayesian algorithm, select the XGBooost model as the evaluator, the hyperparameter range of the XGBoost model as the search range, 50 as the search iterations, 3 as the cross validation, accuracy as the score, and 45 as the random number; The XGBoost Bayesian algorithm was used to search the second training set and the third training set respectively. After obtaining the optimal hyperparameters, performance evaluation was performed on the second test set and the third test set respectively to obtain a qualified second XGBoost model and a qualified third XGBoost model.

5. The intelligent identification method of rock lithology of sandstone-type uranium ore-bearing strata according to claim 4 is characterized in that: The preprocessing is to use standardization to make the well logging curve samples uniformly distributed.