Geological feature identification system applying equilibrium evaluation system and preferential coupling intelligent algorithm and method thereof

By integrating multiple models through a balanced evaluation system and a best-fit coupled intelligent algorithm, along with sliding window technology, the problems of low accuracy and efficiency in geological feature identification have been solved, enabling high-precision and efficient underground resource exploration, reducing false positive errors, and adapting to complex geological conditions.

CN120995198APending Publication Date: 2025-11-21XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511057685.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies lack highly accurate, efficient, and adaptable intelligent identification methods for identifying geological features, especially natural fractures, lithology, and reservoir fluid types. Furthermore, intelligent algorithm models suffer from an imbalance in the attention given to positive and negative samples, leading to frequent false positive errors.

Method used

By employing a balanced evaluation system coupled with an optimal intelligent algorithm, features are extracted through multi-model collaborative integration and sliding window technology. The model is optimized by combining threat scoring and comprehensive score indicators, and conflicts are resolved using majority voting, thereby achieving efficient identification of geological features.

Benefits of technology

It improves the accuracy and efficiency of geological feature identification, reduces human influence, adapts to complex geological conditions, reduces false positive errors, and enhances the scientific nature and efficiency of underground resource exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995198A_ABST
    Figure CN120995198A_ABST
Patent Text Reader

Abstract

The invention discloses a geological feature identification system and method using an equilibrium evaluation system and a preferential coupling intelligent algorithm, and relates to the technical field of geological exploration, and the method comprises the following steps: data preprocessing; labeling the data; dividing data; training and constructing a model; performing model evaluation and optimization selection; carrying out model coupling identification; and blind well identification. Compared with a traditional identification method, the method has the advantages of sustainable optimization, high processing efficiency, high identification precision, small artificial influence and the like. According to the method, automatic analysis is achieved through intelligent models, the working efficiency is greatly improved, time and labor cost are saved, meanwhile, on one hand, the method can make full use of multiple types of data and can adapt to complex geological conditions, on the other hand, the advantages of multiple models are integrated, the complementary characteristics of the models are utilized, fault tolerance is improved, and the method is suitable for large-scale popularization and application. And the identification result is more stable and credible.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of geological exploration, and particularly relates to a geological feature identification system and method using a balanced evaluation system and a preferred coupling intelligent algorithm. BACKGROUND

[0002] In the fields of oil and gas exploration, mineral development and geological research, information such as natural fractures, lithology and reservoir fluid types is crucial for determining the distribution of resources, evaluating the developability of resources and formulating scientific and reasonable exploitation plans. Existing identification methods can be divided into direct observation, experimental analysis, rock mechanics, seismic exploration, well logging interpretation, intelligent algorithm and the like. The direct observation method can intuitively display and judge geological features, but is limited by observation conditions and human factors, and often can only obtain limited geological information and is greatly subjective. The experimental analysis method requires a series of physical and chemical experiments on rock samples to determine the properties and composition of the rock, thereby identifying the geological features, but the sample only represents local geological features and is complex to prepare, time-consuming, high in cost and low in reproducibility. The rock mechanics method mainly identifies geological features by studying the mechanical properties of rocks, but the relationship between the mechanical properties of rocks and geological features is complex, requiring a large number of experiments and numerical simulations, which is high in resource requirement, complex in calculation process and low in accuracy. Seismic exploration detects underground structures by artificially exciting seismic waves and measuring their propagation time and amplitude, but is limited by resolution, velocity variation and noise, and has limited ability to identify small fractures, thin reservoirs and other subtle features, and requires professional personnel and complex software for processing. Well logging interpretation identifies geological features by using well logging instruments to obtain data, but different types of well logging data have different application ranges and limitations. This method relies on expert experience to set thresholds and is highly subjective, cannot handle multi-parameter nonlinear coupling such as complex lithology identification, and well logging data is greatly affected by factors such as borehole conditions and instrument errors. Intelligent algorithms rely on algorithm models to mine fracture information and analyze implicit correlations in data to effectively identify fractures, but most of them rely on a single algorithm model, making it difficult for the model to fully adapt to the needs of geological feature identification, and the evaluation system in intelligent algorithms generally has the problem of imbalance in attention to positive and negative samples, with excessive emphasis on the identification efficiency of positive samples (such as precision / recall rate) and relative neglect of the discrimination performance of negative samples (such as specificity). This bias can lead to a large number of false positive errors induced by the model, bringing risks and costs that cannot be ignored in actual application. In summary, there is currently a lack of an intelligent identification method that can simultaneously meet the requirements of high precision, high efficiency and strong adaptability in identifying geological features such as natural fractures, lithology and reservoir fluids. SUMMARY

[0003] The present application intends to provide a geological feature identification system and method using a balanced evaluation system and a preferred coupling intelligent algorithm, which is based on the construction of a heterogeneous model coupled with an identification model based on the evaluation system, realizes efficient identification and analysis of key geological features such as natural fractures, lithology and reservoir fluid types, and aims to provide high-precision and automated technical means for oil and gas exploration, energy development and other fields, and to help improve the efficiency of underground resource exploration and the scientific nature of decision-making.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0005] The present application provides a geological feature identification method using a balanced evaluation system and a preferred coupling intelligent algorithm, comprising the following steps:

[0006] S1, data preprocessing: collecting various logging curve data and performing data preprocessing operations;

[0007] S2, data labeling: labeling the data information of the training set according to the determined division type;

[0008] S3, data division: dividing 70% of the processed data set into a training set and 30% into a validation set;

[0009] S4, model training and construction: according to the task requirements and selection principles, select multiple models, construct a multi-model collaborative integrated intelligent identification system, use sliding window technology to extract multi-dimensional features of the training set, select dominant features according to the recursive elimination method, then train the model, and optimize the hyperparameters during the training process to determine the optimal parameters of different models;

[0010] S5, model evaluation and optimization: based on all evaluation indicators of the confusion matrix, introduce threat score TS to evaluate the recognition ability of the model to positive class, self-defining NTS to evaluate the recognition ability of the model to negative class, self-defining NF1 and NBA, combining the proportion of F1 score (F1), balanced accuracy (BA) and global indicator (Matthews correlation coefficient, MCC) of the confusion matrix, designing comprehensive score indicator OS, equally allocating threat score TS, NTS and comprehensive score indicator OS, and comprehensively analyzing the model performance:

[0011] Positive class prediction performance = 0.5OS + 0.5TS

[0012] Negative class prediction performance = 0.5OS + 0.5NTS

[0013] The scores of all model positive / negative class prediction performance indicators are calculated and normalized, and then the models are sorted in descending order of the scores (positive class prediction performance ranking, negative class prediction performance ranking). According to the specific task, a suitable threshold is selected, and the model with better positive / negative class prediction performance is selected.

[0014] S6, model coupling identification: according to the task requirements of the required identification, the geological features are divided into positive and negative classes. After evaluating the trained models, the positive / negative class prediction performance scores are obtained. According to the selected threshold, the corresponding high-quality model is selected. After the to-be-identified data is predicted by the selected high-quality model, the prediction results of each model are obtained. These results are combined. If there is a conflict area, use all models to vote for the majority class in the conflict area to determine the final class of the conflict area, thereby obtaining the final output result.

[0015] S7, blind well identification: the coupled model is independently applied to a data set that has never been contacted, i.e., blind well data (test set), thereby obtaining the identification result of the blind well.

[0016] Further, in step S2, when the labeled data set has a class imbalance problem, resampling methods such as undersampling and oversampling or data augmentation methods are used to alleviate the class imbalance problem.

[0017] Further, in step S3, the multi-model collaborative integrated intelligent identification system includes RandomForest (RF), BaggingDecisionTrees (BDT), ExtraTrees (ET), GradientBoosting (GB), LightGBM (LGBM), CatBoost (CB), Adaboost (AB), and Stacking and Blending models with Logistic Regression (LR) as the meta-learner.

[0018] Further, the selection principles of the model include diversity, complexity, and interpretability.

[0019] The diversity is to select different types of models from multiple different algorithm types to reduce the correlation between models and increase the difference between models.

[0020] The complexity is that due to the difference in model principles and different conditions, the selected model itself should have certain discrimination ability when processing classification tasks, and preferably contains both simple and complex models.

[0021] The interpretability is to avoid selecting pure black box models as much as possible.

[0022] Further, in step S5, the self-defined NTS evaluation model is:

[0023]

[0024] Wherein, i, j represent the category, TN j,i represents the number of samples correctly predicted as category j when category i is the positive class;

[0025] The self-defined NF1 and NBA are:

[0026]

[0027] Wherein, NPV is the negative accuracy, Re is the recall rate, and Pr is the precision rate;

[0028] The comprehensive score index OS is:

[0029] OS = 0.5MCC + 0.125(F1 + BA + NF1 + NBA).

[0030] Further, in step S6, for the multi-classification task, that is, selecting category 1 to n as the positive class in turn, and the remaining categories as the negative class, the conflict judgment process is circulated respectively to obtain the prediction results when each category is the positive class, the prediction results are integrated, and then it is judged again whether there is a conflict area after integration, and finally the final full-class prediction result is obtained.

[0031] A geological feature identification system using a balanced evaluation system and a preferred intelligent algorithm coupling, comprising a data preprocessing unit, a data labeling unit, a model training and construction unit, a model evaluation and optimization unit, and a model coupling identification unit;

[0032] The data preprocessing unit is used to collect various logging curve data, increase geophysical data according to the task requirements of research depth and geological feature category, and perform preprocessing operations such as cleaning, normalization and standardization on the data;

[0033] The data labeling unit first determines the identified stratum feature category, then can select whether to subdivide the feature level or state according to the requirements, and finally labels the data materials of the training set according to the determined division type;

[0034] The model construction and training unit first divides the data into a training set and a validation set, constructs a multi-model collaborative integrated intelligent identification system, extracts multi-dimensional features of the training set using a sliding window technique, selects advantageous features according to a recursive elimination method, and then trains the model, and performs hyperparameter optimization in the training process to determine the optimal parameters of different models;

[0035] The model evaluation and the optimization unit are based on all evaluation indexes of the confusion matrix, the threat score TS is introduced to evaluate the recognition ability of the model to the positive class, and the NTS is self-defined to evaluate the recognition ability of the model to the negative class:

[0036]

[0037] The NTS is self-defined as NBA:

[0038]

[0039] The comprehensive score index OS is designed in combination with the F1 score (F1) of the confusion matrix, the balanced accuracy (BA) and the global index (Matthews correlation coefficient, MCC):

[0040] OS=0.5MCC+0.125(F1+BA+NF1+NBA)

[0041] The threat score TS and the NTS are equally allocated with the comprehensive score index OS, and the model performance is comprehensively analyzed:

[0042] The positive class prediction performance=0.5OS+0.5TS

[0043] The negative class prediction performance=0.5OS+0.5NTS

[0044] The model coupling recognition unit divides the geological features into the positive class and the negative class according to the task requirement of the required identification, after the to-be-identified data are predicted by the optimization model, the prediction results of the models are obtained, the results are combined, if there is a conflict area, the majority class voting is performed on the conflict area, the final class of the conflict area is determined, and finally the output result is obtained.

[0045] Further, for the multi-classification task, the model coupling recognition unit selects the class 1 to n as the positive class in turn, and the remaining classes as the negative class, respectively, the conflict judgment process is circulated, the prediction result when each class is taken as the positive class is obtained, then the prediction results are integrated, and then it is judged again whether there is a conflict area after integration, and finally the final full-class prediction result is obtained.

[0046] Compared with the prior art, the present application has the following beneficial effects:

[0047] 1. The application innovatively constructs a new evaluation system based on a confusion matrix. Specifically, the recognition ability of the positive class is evaluated by introducing a threat score TS evaluation model, and the threat score TS self-defined NTS evaluation model is imitated, the F1 score (F1), the balanced accuracy (BA), the self-defined NF1 and NBA evaluation models are imitated, and the global indicators (Matthews correlation coefficient, MCC) are combined. The F1, BA, NF1, NBA and MCC are comprehensively considered, and the comprehensive score index OS is constructed. Since the OS index can ensure the reliability of the model as a whole, it can avoid the decline of global performance caused by local optimization, and the TS and NTS indexes directly reflect the performance of the model in a specific class, which meets the actual task requirements. Therefore, equal weight allocation is implemented, that is, the OS weight is 0.5, and the TS or NTS weight is 0.5. Thus, the scores of all model positive / negative class prediction performance indicators are calculated and normalized, and then the models are sorted in descending order of score (positive class prediction performance sorting, negative class prediction performance sorting). Select 0.7 or a suitable value according to the specific task as the selection threshold, and select the model with better positive / negative class prediction performance. Comprehensive analysis of model performance can deeply analyze the model performance from multiple angles and ensure that the model with the best recognition performance in the crack identification task is selected.

[0048] 2. The application innovatively constructs a new evaluation system based on a confusion matrix. Specifically, according to the task requirements of the required identification, the geological features are divided into positive and negative classes. After evaluating the trained models, the positive / negative class prediction performance scores are obtained, and the corresponding high-quality models are selected according to the selection threshold. After the to-be-identified data is predicted by the selected high-quality models, the prediction results of each model are obtained, and these results are combined. If there is a conflict area, use all models to vote for the majority class in the conflict area to determine the final class of the conflict area, thereby obtaining the final output result. For multi-classification tasks, i.e., selecting classes 1 to n as positive classes and the remaining classes as negative classes, the conflict judgment process is repeated to obtain the prediction results when each class is taken as a positive class. Then, the prediction results are integrated, and then it is judged whether there is a conflict area after integration. Finally, the final full-class prediction result is obtained.

[0049] 3. Compared with the traditional identification method, the application has the advantages of sustainable optimization, high processing efficiency, high recognition accuracy, small artificial influence, etc. Moreover, the application uses intelligent models to realize automatic analysis, greatly improves work efficiency, saves time and labor cost, and at the same time, the method can fully utilize multi-type data and adapt to relatively complex geological conditions. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 It is a flow chart of a geological feature identification method using a balanced evaluation system and a high-quality coupled intelligent algorithm.

[0051] Figure 2 Fig. 1 is a schematic diagram for data labeling (taking fracture and non-fracture labeling of logging data based on imaging as an example);

[0052] Figure 3 Fig. 2 is a schematic diagram for sliding window technology;

[0053] Figure 4 Fig. 3 is a schematic diagram for evaluation index principle;

[0054] Figure 5 Fig. 4 is a ranking result of the model and a normalized ranking result; wherein Figure 5 (a) is a column chart of the model prediction performance index, Figure 5 (b) is a normalized histogram of the model prediction performance index;

[0055] Figure 6 Fig. 5 is a schematic diagram of model coupling principle;

[0056] Figure 7 Fig. 6 is a comparison chart of logging curve and advantage model identification result (taking fracture identification as an example), wherein Figure 7 (a) is a schematic diagram of fracture imaging near the depth of 2925m, Figure 7 (b) is a schematic diagram of fracture imaging near the depth of 2965m, Figure 7 (c) is a schematic diagram of fracture imaging near the depth of 3000m, Figure 7 (d) is a schematic diagram of fracture imaging near the depth of 3005m, Figure 7 (e) is a schematic diagram of fracture imaging near the depth of 3010m, Figure 7 (f) is a schematic diagram of fracture imaging near the depth of 3020m. DETAILED DESCRIPTION

[0057] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments and drawings. Herein, the illustrative embodiments of the present application and the descriptions thereof are used to explain the present application, but not as a limitation of the present application.

[0058] It should be further noted that, in order to avoid the present application being obscured by unnecessary details, only the processing steps closely related to the scheme according to the present application are shown in the drawings, and other details not closely related to the present application are omitted.

[0059] It should be emphasized that the term "comprise / comprising" is used herein to indicate the presence of a feature, element or step, but not to exclude the presence or addition of one or more other features, elements or steps.

[0060] It is emphasized here that the step markers mentioned in the following are not a limitation of the order of the steps, but it is to be understood that the steps can be carried out in the order mentioned in the examples, but also in a different order than in the examples, or several steps are carried out simultaneously.

[0061] Example I

[0062] Reference Figure 1 A geological feature identification method using a balanced evaluation system and a selection-coupling intelligent algorithm, comprising the following steps:

[0063] S1, data collection and preprocessing; collect various logging curve data, such as AC, AT10, AT90, CAL, CLL8, DEN, GR, SP, etc. to construct the original data set, further increase other types of geophysical data such as seismic data, rock mechanics data, etc. according to the continuous deepening of scientific research and the task requirements of the geological feature categories to be identified, and perform preprocessing operations such as cleaning, normalization, standardization, etc. on the data to improve data quality and make it suitable for model training.

[0064] S2, data labeling; first determine the identified formation feature categories, such as natural fractures, lithology, reservoir fluid, etc. Secondly, according to the needs, it can be selected whether to subdivide the feature level or state, such as stratigraphic fractures divided into vertical fractures, high-angle fractures, low-angle fractures, and horizontal fractures according to the angle; lithology is subdivided according to mineral composition and structure; reservoir fluid type is subdivided according to parameters such as saturation. Finally, label the data in the training set according to the determined division type, usually use Arabic numerals to represent different division types, reference Figure 2 which shows the process of labeling the logging curve data according to the fracture data in the imaging logging, the imaging logging data (micro-resistivity scan) on the left is divided into static and dynamic two parts, which shows the specific depth area corresponding to the rock fracture (blue line in the figure), that is, the position corresponding to the purple area in the figure, while the light yellow area corresponds to the non-fracture section. The data label labeling part on the right lists in detail the curve data (a, b, c in the figure represent the specific values of different logging curves) at each depth point and the corresponding data label, such as "0" representing the non-fracture area and "1" representing the fracture area. The logging curve data is labeled according to the imaging logging data by the above method.

[0065] If the labeled data set has the problem of class imbalance, resampling methods such as undersampling and oversampling or data enhancement methods can be used to alleviate the problem of class imbalance.

[0066] S3, data division: after selecting 3-5 well data as blind well data from the pretreated original data set, the remaining data set is randomly divided, usually 70% of the data set is divided into training set and 30% is divided into validation set.

[0067] S4, model construction and training; using machine learning or deep learning algorithm, selecting multiple different algorithm identification models such as RandomForest(RF), BaggingDecisionTrees(BDT), ExtraTrees(ET), GradientBoosting(GB), LightGBM(LGBM), CatBoost(CB), Adaboost(AB), Stacking and Blending with LogisticRegression(LR) as meta-learner, etc. It is worth noting that the model should be selected according to the task requirements and selection principles, which can be easily realized by those skilled in the art, and only part of the model is listed in the present application.

[0068] The selection principles of the model include:

[0069] 1. Diversity: from Bagging algorithm, Boosting algorithm, deep learning, reinforcement learning and other different algorithm types, select different types of models to reduce the correlation between models and increase the difference between models.

[0070] 2. Complexity: due to the difference of model principle and different adaptation conditions, when dealing with classification task, the selected model itself should have certain discrimination ability, and it is best to contain simple model and complex model at the same time.

[0071] 3. Explainability: according to the demand of classification task, geological decision needs causal support, etc., it is necessary to avoid selecting pure black box model as far as possible

[0072] Finally, based on the task demand and selection principle, an intelligent identification system of multi-model collaborative integration is constructed. Using sliding window technology, multi-dimensional(time domain, frequency domain, statistical characteristics, etc.) features of training set are extracted, and the window size is usually determined according to the characteristics of geological feature category trend, time series period, etc. For example, when identifying fracture, the mode value of vertical height of all known fractures is taken as the window size, and the step is usually 1 / 4 of the window size. Then the model is trained after selecting the dominant features according to the recursive elimination method, and the hyperparameter optimization is carried out in the training process to determine the optimal parameters of different models.

[0073] Reference Figure 3It is a schematic diagram of the sliding window technology, which is used to intuitively show the application of the sliding window technology in data analysis, and explain how to use the sliding window technology to extract features from the logging curve data. The left side of the figure shows the change of the logging curve (AC, GR, DEN curve in the figure) with depth. The middle part shows the size of the sliding window with a red dashed line box, and the arrow indicates the direction of the window sliding, that is, from the top of the curve to the bottom, and the window has a certain step, when the window starts to slide, it only slides a step distance at a time. The right side shows the process of extracting the curve segment from each window, each window corresponds to the depth area and specific value of the curve (AC, GR, DEN curve) segment. The sliding window technology can divide the long logging curve into multiple shorter windows, which is convenient for detailed analysis and feature extraction of the data in each window.

[0074] S5, model evaluation and optimization; in the process of model optimization, it is urgent to establish a unified and balanced evaluation system. The application establishes a unified identification evaluation system on the basis of considering the identification effect of positive and negative samples. The system takes all elements in the confusion matrix as the basis, explicitly includes the accuracy (Pr), recall (Re), specificity (Spe) and negative predictive value (NPV) and other indicators, and derives F1 score (F1), balanced accuracy (BA), self-defined NF1, NBA and other indicators to comprehensively measure the generalization performance of the classification model.

[0075] With reference to the specific embodiments Figure 4, which is a visual representation of the confusion matrix and model-related evaluation indicators, used to illustrate the performance of the classification model. The table on the left side of the figure builds a confusion matrix with the cross-classification of actual classes (positive and negative) and predicted classes (positive and negative), clearly defining the definitions of TP / FP / FN / TN: true positive (true positive - TP): the number of samples that are actually positive and predicted as positive by the model; false negative (false negative - FN): the number of samples that are actually positive but predicted as negative by the model; false positive (false positive - FP): the number of samples that are actually negative but predicted as positive by the model; true negative (true negative - TN): the number of samples that are actually negative and predicted as negative by the model. On the right side, evaluation indicators calculated based on the confusion matrix are listed, including precision, negative predictive value (NPV), recall, specificity, F1 score (F1), balanced accuracy (BA), and negative F1 score (NF1). Among them, precision measures the accuracy of predicting positive classes, NPV measures the accuracy of predicting negative classes; recall focuses on the proportion of actual positive classes that are correctly identified, and specificity focuses on the proportion of actual negative classes that are correctly identified; F1 score is the harmonic mean of precision and recall, and NF1 score is the harmonic mean of NPV and specificity. The confusion matrix intuitively reflects the classification of the model on positive and negative classes, while the evaluation indicators quantify the performance of the model from different angles, such as precision and recall focusing on the accuracy and integrity of positive class prediction, while specificity, NPV and NF1 focusing on the evaluation of negative classes. F1 and NF1 as the harmonic mean further integrate the performance of positive and negative classes, helping to comprehensively understand the performance differences of the model on different classes. Through these indicators, the performance of the model on different classes can be comprehensively evaluated, and the strengths and weaknesses of the model can be identified, guiding the optimization and improvement of the model.

[0076] In addition, the threat score (TS, also known as critical success index) is introduced to evaluate the model's ability to identify positive classes, and the NTS is self-defined to evaluate the model's ability to identify negative classes, and the MCC is a global indicator for model evaluation, and its calculation formula is shown in Table 1.

[0077] Table 1: Evaluation indicator calculation formula

[0078]

[0079]

[0080] It should be further noted that in a multi-classification task, TS and NTS for each class can be calculated by class-by-class (setting each class as positive class and the rest as negative class in turn), and then these values can be summarized (such as taking the average or other statistical quantities), thereby obtaining the positive class and negative class prediction performance evaluation of the model.

[0081]

[0082] where i, j represent the class, TN j,i denotes the number of samples correctly predicted as class j for class i as the positive class

[0083] To more accurately describe the overall classification performance of the model, a comprehensive score indicator Overall Score (OS) is designed according to the global indicator MCC and the classification performance indicators F1, BA, NF1, NBA, etc. Since MCC is a global indicator, it can comprehensively measure the prediction ability of the model for positive classes, negative classes, false positives, and false negatives. It takes into account all four classification results (positive class TP, negative class TN, false positive class FP, and false negative class FN), avoiding bias caused by uneven class distribution, so it is given a higher weight (0.5). The classification performance indicators F1, BA, NF1, NBA, etc. can complement the deficiencies of MCC from different angles, and each indicator has the same importance, so equal weight allocation is adopted, i.e. each indicator weight is 0.125:

[0084] OS = 0.5MCC + 0.125(F1 + BA + NF1 + NBA)

[0085] Considering that TS and NTS are evaluation indicators specifically for positive and negative classes, and do not consider the overall performance of the model, positive and negative class prediction performance indicators are designed according to OS and TS, NTS indicators. Since the OS indicator can ensure that the model is reliable overall, avoiding the decline in global performance caused by local optimization, and the TS, NTS indicators directly reflect the performance of the model in a specific class, which meets the actual task requirements. If the weight of TS / NTS is too high (such as 0.7), it may lead to over-optimization of a certain class, at the expense of overall performance; if the weight is too low (such as 0.3), it cannot effectively reflect the local demand, so equal weight allocation is implemented, i.e. the weight of OS is 0.5, and the weight of TS or NTS is 0.5:

[0086] Positive class prediction performance = 0.5OS + 0.5TS

[0087] Negative class prediction performance = 0.5OS + 0.5NTS

[0088] Through this comprehensive evaluation system, the performance of the model can be analyzed in depth from multiple angles, ensuring that the model with the best identification performance in the crack identification task is selected (the number of selected models can be determined by setting the threshold value of the positive and negative class prediction performance indicators, usually an odd number).

[0089] In particular, the embodiments show the performance of the models in predicting the positive class (fracture) and the negative class (non-fracture). Referring to Table 2, it shows the specific values of various evaluation indicators obtained when the trained models are evaluated using the evaluation system.

[0090] Table 2 Model evaluation table

[0091]

[0092] Referring to Figure 5 The figure shows the comparison of the performance of different models in fracture prediction. The data in the figure is calculated from the values of various evaluation indicators in Table 2. The horizontal bar chart (left side) shows the performance of the models (RandomForest (RF), BaggingDecisionTrees (BDT), ExtraTrees (ET), GradientBoosting (GB), LightGBM (LGBM), CatBoost (CB), Adaboost (AB), and Stacking (S-LR) and Blending (B-LR) models using Logistic Regression (LR) as the meta-learner) in predicting the positive class (fracture) and the negative class (non-fracture),

[0093] The horizontal axis represents the size of the performance indicator value, and the larger the value, the better the prediction performance. Each model shows its prediction performance for the positive class and the negative class. The positive class prediction performance reflects the ability of the model to identify fractures, and the negative class prediction performance reflects the ability of the model to identify non-fractures (e.g., the AB model has the strongest ability to identify fractures, and the LGBM has the weakest ability to identify fractures). The histogram (right side) normalizes the values. Through this display, the relative advantages and disadvantages of different models in predicting the positive and negative classes can be more intuitively compared, and the number of high-quality models to be selected can be determined. The number of models can be determined using a fixed threshold, or it can be analyzed according to the figure. For example, in the figure, the performance indicator decreases significantly after the 5th model, so the top five models can be selected as high-quality models.

[0094] S6, model coupling identification; according to the task requirements of the required identification, the geological features are divided into positive and negative classes. For multi-classification tasks, each class can be selected in turn as the positive class, and the remaining classes are taken as the negative class (consistent with the aforementioned calculation of TS and NTS for multiple classes). After the identification data is predicted by the selected models, the prediction results of each model are obtained, and these results are combined. If there is a conflict area, a majority vote is performed on the conflict area to determine the final class of the conflict area, thereby obtaining the final output result.

[0095] Referring toFigure 6 FIG. 2 is a workflow diagram of the multi-model fusion prediction, which describes the complete decision-making process from data input to the final prediction result. The main content includes: model selection and prediction, i.e., starting from the data to be identified, the optimal model is selected to divide the positive and negative categories. A certain category is set as the positive class, and the rest are the negative class, and then each model is used for prediction; preliminary result integration, i.e., the prediction results of the positive and negative classes are merged respectively for result integration; conflict processing, i.e., checking whether there is a conflict area, if there is, then full-model voting is performed on the conflict area to solve the inconsistency between model predictions; cycle verification, i.e., ensuring that each category has been divided and predicted as a positive class to ensure comprehensiveness. Finally, integrate all the prediction results of the categories to obtain the final prediction output.

[0096] For multi-classification tasks, the above process needs to be cycled to ensure that the prediction results of each category as a positive class can be obtained, and then these prediction results are integrated, and it is judged whether there is a conflict area, and finally the final full-category prediction result is obtained. At this point, the intelligent identification of the geological feature categories is completed.

[0097] S7, blind well identification; the coupling model is independently applied to a data set that has never been contacted (blind well), so as to obtain the identification result of the blind well. That is, through the environment of "zero information leakage", the identification performance is objectively verified, and deviation caused by repeated use of data or manual intervention is avoided.

[0098] Reference Figure 7 FIG. 3 is a comparison diagram of the logging curve and the model identification result. The logging curve part (left side of the figure) shows the change of various logging curves (such as CAL, GR, AC, etc.) with depth. The model identification result part (middle part) shows the identification results of the fractures by the high-quality models (ExtraTrees (ET), CatBoost (CB), Adaboost (AB), and Stacking (S-LR) and Blending (B-LR) models with LogisticRegression (LR) as the meta-learner) selected in the foregoing and the model after coupling, and are identified by different colors. The cumulative identification curve part (right side of the figure) shows the change curve of the cumulative number of models for fracture identification. The schematic diagram of the cumulative curve in the figure is a brief explanation of the identification curve, wherein "1" represents that the model successfully identifies the fracture area, and "0" represents that the model fails to identify the fracture area. The model cumulative identification curve can count the number of models for identifying the fracture at a certain depth, which helps to improve the accuracy and reliability of fracture identification. The fracture imaging data (a-f) at the bottom of the figure clearly shows the shape and distribution of the fractures at different depth sections for blind well testing.

[0099] Example Two

[0100] A geological feature identification system using a balanced evaluation system and a preferred coupling intelligent algorithm, comprising a data preprocessing unit, a data labeling unit, a model training and construction unit, a model evaluation and optimization unit, and a model coupling identification unit;

[0101] The data preprocessing unit is used to collect various logging curve data, increase geophysical data according to the task requirements of research depth and geological feature categories, and perform preprocessing operations such as cleaning, normalization and standardization on the data;

[0102] The data labeling unit first determines the identified formation feature category, then selects whether to subdivide the feature level or state according to the requirements, and finally labels the data materials of the training set according to the determined division type;

[0103] The model construction and training unit first divides the data into a training set and a validation set, constructs a multi-model collaborative integrated intelligent identification system, extracts multi-dimensional features of the training set using a sliding window technique, selects advantageous features according to a recursive elimination method, and then trains the model, and optimizes the hyperparameters during the training process to determine the optimal parameters of different models;

[0104] The model evaluation and optimization unit is based on all evaluation indicators of the confusion matrix, introduces threat score TS to evaluate the recognition ability of the model for positive classes, and customizes NTS to evaluate the recognition ability of the model for negative classes:

[0105]

[0106] Customize NF1, NBA:

[0107]

[0108] Combine the F1 score (F1), balanced accuracy (BA), and global indicators (Matthews correlation coefficient, MCC) of the confusion matrix to design a comprehensive score indicator OS:

[0109] OS = 0.5MCC + 0.125(F1 + BA + NF1 + NBA)

[0110] Assign equal weights to threat scores TS, NTS, and comprehensive score indicators OS, and comprehensively analyze the model performance:

[0111] Positive class prediction performance = 0.5OS + 0.5TS

[0112] Negative class prediction performance = 0.5OS + 0.5NTS

[0113] The model coupling identification unit divides the geological features into positive and negative classes according to the task requirements of the required identification. After the data to be identified is predicted by the selected model, the prediction results of each model are obtained. These results are combined. If there is a conflict area, the majority class voting is performed on the conflict area to determine the final class of the conflict area, so as to obtain the final output result. For a multi-classification task, class 1 to n are selected as positive classes in turn, and the remaining classes are selected as negative classes. The conflict judgment process is performed in cycles to obtain the prediction results when each class is selected as a positive class. Then, it is judged whether there is a conflict area after integration. Finally, the final full-class prediction result is obtained.

[0114] Those of ordinary skill in the art should understand that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether the implementation is in hardware or software depends on the specific application and design constraints imposed on the overall system. Skilled artisans can use various methods to implement the described functions using hardware, software, or a combination of both. Such implementation should not be considered beyond the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave in a transmission medium or communication link.

[0115] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.

[0116] In the present application, the features described and / or exemplified for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or in combination with or instead of the features of other embodiments.

[0117] The above description is only the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A geological feature identification method using a balanced evaluation system and a preferred coupling intelligent algorithm, characterized in that, Comprising the following steps: S1, data preprocessing: collecting various logging curve data and performing data preprocessing operations; S2, data labeling: labeling the data of the training set according to the determined classification type; S3, data division: dividing 70% of the processed data set into a training set and 30% into a validation set; S4, model training and construction: selecting multiple models according to task requirements and selection principles, constructing an intelligent recognition system with multiple model collaborative integration, extracting multi-dimensional features of the training set using sliding window technology, selecting dominant features according to the recursive elimination method, training the model, and optimizing the hyperparameters during training to determine the optimal parameters of different models; S5, model evaluation and optimization: based on all evaluation indicators of the confusion matrix, introduce threat score TS to evaluate the recognition ability of the model for positive class, self-defining NTS to evaluate the recognition ability of the model for negative class, self-defining NF1 and NBA, combining the proportion of F1 score (F1), balanced accuracy (BA) and global indicator (Matthews correlation coefficient, MCC) in the confusion matrix, designing comprehensive score indicator OS, equally allocating threat score TS, NTS and comprehensive score indicator OS, and comprehensively analyzing the model performance: Positive class prediction performance = 0.5OS + 0.5TS Negative class prediction performance = 0.5OS + 0.5NTS Calculate the scores of all model positive / negative class prediction performance indicators and normalize them, then sort the models according to the scores from large to small (positive class prediction performance sorting, negative class prediction performance sorting), select appropriate threshold according to specific task conditions, and select models with better positive / negative class prediction performance respectively; S6, model coupling recognition: according to the task requirements of the required identification, divide the geological features into positive and negative classes, obtain the positive / negative class prediction performance scores after evaluating the trained models, select the corresponding high-quality models according to the selected threshold, predict the data to be identified by the selected high-quality models, obtain the prediction results of each model, combine these results, if there are conflict areas, use all models to vote for the majority class in the conflict area to determine the final class of the conflict area, and thus obtain the final output result; S7, blind well identification; Applying the coupling model independently to a never-before-seen data set , i.e. blind well data (test set), thereby obtaining a result of identifying blind wells.

2. The geological feature identification method of claim 1, characterized in that, In step S2, when the labeled data set has the problem of class imbalance, resampling methods such as undersampling and oversampling or data augmentation methods are used to alleviate the class imbalance problem.

3. The geological feature identification method of claim 1, wherein, In step S3, the intelligent recognition system with multiple model collaborative integration includes RandomForest (RF), BaggingDecisionTrees (BDT), ExtraTrees (ET), GradientBoosting (GB), LightGBM (LGBM), CatBoost (CB), Adaboost (AB), and Stacking and Blending models with Logistic Regression (LR) as the meta-learner.

4. The geological feature identification method of claim 1, wherein, The selection principles of the models include diversity, complexity and interpretability; The diversity is to select different types of models from a plurality of different algorithm types to reduce the correlation between models and increase the difference between models; The complexity is that due to the difference in model principles and different conditions, the selected model itself should have certain discrimination ability when processing classification tasks, and preferably contains both simple and complex models; The interpretability is to avoid selecting pure black box models as much as possible.

5. The geological feature identification method of claim 1, wherein, In step S5, the self-defined NTS evaluation model is: where i, j represent the class, TN j,i denotes the number of samples correctly predicted as class j for class i as the positive class; The self-defined F1, NBA are: Wherein, NPV is the negative accuracy, Re is the recall rate, and Pr is the precision rate; The comprehensive score index OS is: OS=0.5MCC+0.125(F1+BA+NF1+NBA).

6. The geological feature identification method of claim 1, wherein, In step S6, for the multi-classification task, that is, class 1 to n are selected as positive classes in turn, and the remaining classes are negative classes, the conflict judgment process is circulated respectively to obtain the prediction results of each class as a positive class, the prediction results are integrated, and then it is judged again whether there is a conflict area after integration, and finally the final full-class prediction result is obtained.

7. A geological feature identification system using a balanced evaluation system and a preferred coupling intelligent algorithm, characterized in that, It comprises a data preprocessing unit, a data labeling unit, a model training and construction unit, a model evaluation and optimization unit, and a model coupling identification unit. The data preprocessing unit is used to collect a plurality of logging curve data, increase geophysical data according to the task demand of research depth and geological feature category, and perform preprocessing operations such as cleaning, normalization and standardization on the data; The data labeling unit first determines the identified formation feature category, then can select whether to subdivide the feature level or state according to the demand, and finally labels the data materials of the training set according to the determined division type; The model construction and training unit first divides the data into a training set and a validation set, constructs a multi-model collaborative integrated intelligent identification system, extracts multi-dimensional features of the training set by using a sliding window technology, selects advantageous features according to a recursive elimination method, and then trains the model, and performs hyperparameter optimization in the training process to determine the optimal parameters of different models; The model evaluation and optimization unit takes all evaluation indexes of the confusion matrix as the basis, introduces threat score TS evaluation model to evaluate the identification ability of the positive class, and self-defines NTS evaluation model to evaluate the identification ability of the negative class: The self-defined F1, NBA are: Combined with the F1 score (F1), the balanced accuracy (BA) and the global index (Matthews correlation coefficient, MCC) of the confusion matrix, a comprehensive score index OS is designed: OS=0.5MCC+0.125(F1+BA+NF1+NBA) The threat scores TS and NTS and the comprehensive score index OS are equally allocated, and the model performance is comprehensively analyzed: Positive class prediction performance=0.5OS+0.5TS Negative class prediction performance=0.5OS+0.5NTS The model coupling identification unit divides the geological features into positive and negative classes according to the required task demand, obtains the prediction results of each model after the selected optimal model predicts the data to be identified, merges these results, performs majority voting on the conflict area if there is a conflict area, determines the final class of the conflict area, and thus obtains the final output result.

8. The geological feature identification system of claim 7, wherein, The model coupling identification unit is for a multi-classification task, that is, class 1 to n are sequentially selected as positive classes, the remaining classes are negative classes, the conflict judgment process is circulated respectively, the prediction results when each class is a positive class are obtained, the prediction results are integrated, it is judged again whether there is a conflict region after integration, and finally the final full-class prediction result is obtained.