Method and system for constructing a geoscience binary visualization classification diagram based on machine learning
Through machine learning, the binary visual classification diagram of geology is constructed, and the SVM algorithm and SHAP algorithm are used to select features, and the feature combination and demarcation lines are automatically determined, which solves the problems of low accuracy and dependence on prior knowledge in the existing technology, achieving efficient and accurate classification effects.
Patent Information
- Application Number
- CN202410234441.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-03-01
AI Technical Summary
The accuracy of existing binary visualization diagrams of geology is low and relying on the prior knowledge of researchers, it is difficult to effectively distinguish the types of geological objects.
Using machine learning methods, the initial model is trained using the support vector machine (SVM) algorithm, combined with the SHAP algorithm to select the most influential features, and multiple binary visual classification diagrams are constructed, and feature combinations and dividing lines are automatically determined through machine learning to reduce dependence on prior knowledge.
It improves the accuracy and efficiency of geological object classification, reduces the error of human subjective factors, and can deal with almost all geological problems involving classification.
Smart Images

Figure CN118116510B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geology, and particularly relates to a method and system for constructing a geoscience binary visualization classification diagram based on machine learning. Background Art
[0002] Many important scientific problems in geology can be attributed to classification problems. For example, for a rock, whether it is formed in an arc environment or a collision environment, whether it is an ore-forming rock mass or a non-ore-forming rock mass, whether it is prone to copper mineralization or molybdenum mineralization, etc. To address these problems, people usually use certain geochemical components of rocks or minerals to construct visual diagrams that can be directly used based on prior knowledge. The most common visual geoscience diagrams are binary classification diagrams. For example, in order to better judge the emplacement tectonic environment of granitoids, previous researchers proposed the Y-Nb and (Y+Yb)-Rb granitoid emplacement tectonic environment discrimination diagrams based on whole-rock geochemical components. These binary visualization diagrams are widely used in geological research currently because they are easy to operate. Researchers only need to select the contents of a few elements to classify the research object.
[0003] However, more and more studies have shown that the accuracy of these previously proposed binary visualization diagrams is relatively low, and in many cases, they cannot effectively distinguish which type the research object belongs to. For example, on the whole-rock Y-Nb granitoid emplacement tectonic environment discrimination diagram, it is often encountered that sample points from the same granite body are plotted in both the volcanic arc granite area and the intraplate environment. In addition, the construction of binary visualization diagrams in traditional research strongly depends on the prior knowledge of researchers. When facing some scientific problems with less previous research, it becomes a difficult problem to select which characteristics of rocks or minerals to construct binary visualization diagrams. Summary of the Invention
[0004] An embodiment of the present invention provides a method and system for constructing a geoscience binary visualization classification diagram based on machine learning to solve the problems of inaccurate classification and dependence on the prior knowledge of researchers in the prior art.
[0005] To solve the above technical problems, the embodiment of the present invention discloses the following technical solutions:
[0006] One aspect of the present invention provides a method for constructing a geoscience binary visualization classification diagram based on machine learning for classifying geological body objects according to preset categories. The method includes:
[0007] Obtaining information data of multiple geological body objects, where the information data of the geological body objects includes characteristic data and category data of the geological body;
[0008] Construct a geological body dataset according to the information data of geological body objects. Each sample in the geological body dataset contains data of multiple preset features and the category of the corresponding geological body object, and the category of the geological body object belongs to the preset category;
[0009] Utilize the geological body dataset and adopt the support vector machine (SVM) algorithm to train an initial SVM model;
[0010] Based on the SHAP algorithm, analyze the initial SVM model, and select the first preset number of features that have the greatest impact on the output result of the initial SVM model from the preset features as training features;
[0011] Select the second preset number of features from the first preset number of training features in any combination manner to obtain multiple combination results of features, where the second preset number is less than the first preset number;
[0012] Utilize the geological body dataset to retrain the SVM model for the features in each combination result to obtain multiple new SVM models;
[0013] Construct a binary visualization classification diagram for each new SVM model. Among them, the boundary line between different categories in each binary visualization classification diagram is determined by the corresponding SVM model;
[0014] Obtain the feature data of the geological body object to be classified and project it onto all binary visualization classification diagrams, and determine the category of the geological body object to be classified according to the projection results in all binary visualization classification diagrams.
[0015] Optionally, the step of utilizing the geological body dataset and adopting the support vector machine (SVM) algorithm to train an initial SVM model includes:
[0016] Adopt a non-linear mapping method to map the data of the preset features and the category of the corresponding geological body object in each sample of the geological body dataset to a high-dimensional feature space;
[0017] Obtain an optimal classification hyperplane in the high-dimensional feature space;
[0018] Train an initial SVM model according to the optimal classification hyperplane.
[0019] Optionally, the method further includes:
[0020] Randomly divide the geological body dataset into a training dataset and a test dataset according to a preset ratio;
[0021] During the process of training the initial SVM model, the 5-fold cross-validation method is used to evaluate the performance of the initial SVM model based on the training data set, and the performance evaluation results are obtained; the performance evaluation results include precision, precision, recall, and F1 score;
[0022] The Bayesian optimization method and the performance evaluation results of the initial SVM model are used to optimize the parameters of the initial SVM model, and the optimized initial SVM model is obtained.
[0023] Optionally, the using the Bayesian optimization method and the performance evaluation results of the initial SVM model to optimize the parameters of the initial SVM model to obtain the optimized initial SVM model includes:
[0024] According to the performance evaluation results of the initial SVM model, a probability model between the parameters and performance of the initial SVM model is established;
[0025] The probability model is used to predict the next set of parameter combinations that make the performance evaluation results of the initial SVM model reach the optimal;
[0026] The initial SVM model is updated according to the new parameter combination and the performance evaluation results are obtained again;
[0027] Repeat the above steps until the performance evaluation results of the initial SVM model no longer increase, and the optimized initial SVM model is obtained using the current parameter combination.
[0028] Optionally, a binary visualization classification diagram is constructed for each new SVM model, where the dividing line between different classes in each binary visualization classification diagram is determined by the corresponding SVM model, including:
[0029] The preset decision function is used to determine the decision boundary, and the decision function is
[0030] f(x) = w·x + b
[0031] where w is the normal vector of the hyperplane in the SVM model, x is the feature in the SVM model, and b is the preset bias term;
[0032] The preset decision function and the decision boundary are used to determine the dividing line between different classes in the binary visualization classification diagram.
[0033] Optionally, the feature data of the geological body object to be classified is obtained and projected in all binary visualization classification diagrams, and the category of the geological body object to be classified is determined according to the projection results in all binary visualization classification diagrams, including:
[0034] For each binary visualization classification diagram, the classification results of the geological body object to be classified are obtained in the following manner:
[0035] Obtain the features corresponding to the binary visualization classification diagram as diagram features;
[0036] Project the data of the diagram features in the geological object to be classified onto the binary visualization classification diagram to obtain a classification result;
[0037] Count the number of each classification result, and take the classification result with the largest number as the category of the geological object to be classified.
[0038] Another aspect of the present invention provides a system for constructing a geoscience binary visualization classification diagram based on machine learning for classifying geological objects according to preset categories. The system includes:
[0039] An information data acquisition module configured to acquire information data of a plurality of geological objects, where the information data of the geological objects includes feature data and category data of the geological bodies;
[0040] A geological body data set construction module configured to construct a geological body data set according to the information data of the geological objects. Each sample in the geological body data set includes data of multiple preset features and the category of the corresponding geological object, and the category of the geological object belongs to the preset category;
[0041] An initial SVM model training module configured to use the geological body data set and train an initial SVM model using the support vector machine SVM algorithm;
[0042] A training feature selection module configured to analyze the initial SVM model based on the SHAP algorithm and select the first preset number of features that have the greatest impact on the output result of the initial SVM model from the preset features as training features;
[0043] A feature combination acquisition module configured to select the second preset number of features from the first preset number of training features in any combination manner to obtain multiple combination results of the features, where the second preset number is less than the first preset number;
[0044] A new SVM model training module configured to use the geological body data set to retrain the SVM model for the features in each combination result to obtain multiple new SVM models;
[0045] A binary visualization classification diagram module configured to construct a binary visualization classification diagram for each new SVM model. Among them, the dividing lines of different categories in each binary visualization classification diagram are determined by the corresponding SVM model;
[0046] A category determination module, configured to obtain feature data of a geological object to be classified and project it in all binary visualization classification diagrams, and determine the category of the geological object to be classified according to the projection results in all binary visualization classification diagrams.
[0047] A method and system for constructing a geoscience binary visualization classification diagram based on machine learning disclosed in the present invention automatically obtains the first preset number of training features that have the greatest impact on the classification result in a machine learning manner, and selects the second preset number of features from the training features in any combination to obtain various combination results of the features. For each combination result of features, an SVM model is trained respectively, and then a binary visualization classification diagram is constructed by using each SVM model. The category of the geological object to be classified is determined based on all binary visualization classification diagrams. The method adopted by the present invention greatly improves the classification efficiency and has a high accuracy. At the same time, different from the construction process of traditional geoscience binary visualization diagrams, the present invention hardly requires prior knowledge and can handle almost all geological problems involving classification in the geoscience field. Description of the Drawings
[0048] Figure 1 It is a schematic flow chart of a method for constructing a geoscience binary visualization classification diagram based on machine learning disclosed in an embodiment of the present invention;
[0049] Figure 2 It is a schematic flow chart of a method for constructing a geoscience binary visualization classification diagram based on machine learning disclosed in another embodiment of the present invention;
[0050] Figure 3 It is an implementation of an embodiment of the present invention Figure 1 The schematic flow chart of step S108;
[0051] Figure 4 It is a schematic structural diagram of a system for constructing a geoscience binary visualization classification diagram based on machine learning disclosed in an embodiment of the present invention. Detailed Embodiments
[0052] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0053] Figure 1 It is a schematic flow chart of a method for constructing a geoscience binary visualization classification diagram based on machine learning provided in an open embodiment of the present invention. As Figure 1 shown, the method includes the following steps:
[0054] Step S101: Obtain information data of multiple geological objects.
[0055] The information data of the geological body includes the characteristic data and category data of the geological body object.
[0056] In the embodiments disclosed in the present invention, the geological body object can be a mineral sample, a rock sample, a structure, an ore body, a deposit, etc. The characteristics of the geological body object can be the color and structure of the object, or the elements contained in the object, etc. The category of the geological body object is the specific category to which the geological body object belongs.
[0057] For the convenience of description and understanding, in the embodiments of the present invention, taking the zircon as the geological body object as an example, the implementation manner of the present invention is described. Among them, the geological body object is a zircon sample, the characteristics of the geological body object are the elements contained in the zircon sample, and the category of the geological body object is the source rock type of the zircon sample, that is, the ore-forming rock mass and non-ore-forming rock mass of the porphyry-skarn Cu polymetallic deposit.
[0058] In the embodiments disclosed in the present invention, the obtained zircon data of the ore-forming rock mass includes skarn Cu polymetallic deposits and porphyry Cu deposits; for the obtained zircon data of the non-ore-forming rock mass, the zircon data of the non-ore-forming rock mass at home and abroad is used. To make the results more representative, the data of each zircon sample needs to have preset element characteristics. For example, the zircon sample needs to include 15 rare earth element characteristics: La, Ce, Pr, Nd, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu and Y, and 6 trace element characteristics: Ti, Hf, Th, U, Eu / Eu* and Ce / Ce*, where the value of Eu* is calculated by the formula calculated, where Sm is the content value of the Sm element and Gd is the content value of the Gd element; the value of Ce* is calculated by the formula calculated, where La is the content value of the La element and Pr is the content value of the Pr element.
[0059] Step S102: Construct a geological body data set according to the information data of the geological body object.
[0060] Each sample in the geological body data set contains data of multiple preset characteristics and the category of the corresponding geological body object.
[0061] The category of the geological body object belongs to the preset category. In the specific embodiments of the present invention, the category of the geological body object is the source rock type of the zircon sample, that is, the ore-forming rock mass and non-ore-forming rock mass of the porphyry-skarn Cu polymetallic deposit.
[0062] Each sample in the geological body data set corresponds to a zircon sample, and different samples correspond to different zircon samples. For each sample, its data is obtained based on the information data of the corresponding zircon sample, and each sample contains data of a preset number of preset element characteristics and the geological body category corresponding to the corresponding zircon sample.
[0063] In a specific embodiment disclosed in the present invention, the preset element characteristics include 15 rare earth element characteristics: La, Ce, Pr, Nd, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu and Y, and 6 trace element characteristics: Ti, Hf, Th, U, Eu / Eu* and Ce / Ce*.
[0064] In one embodiment disclosed in the present invention, the geological volume data set can be constructed in the following manner:
[0065] (1) Determine whether there are two or more zircon samples with consistent information data.
[0066] If there are two or more zircon samples with consistent information data, that is, if there are two or more zircon samples with completely consistent information data, the zircon samples with consistent information data are taken as a group of repeated samples. In practical applications, there may be multiple groups of repeated samples. For each group of repeated samples, the information data of one sample is retained and the information data of the remaining samples is deleted so that there are no repeated samples.
[0067] If there are no two or more zircon samples with consistent information data, it is considered that there are no duplicate samples.
[0068] (2) Preselect 15 rare earth element characteristics and 6 trace element characteristics as preset element characteristics. In a specific embodiment disclosed in the present invention, the selected rare earth element characteristics are: La, Ce, Pr, Nd, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu and Y; the selected trace element characteristics are: Ti, Hf, Th, U, Eu / Eu* and Ce / Ce*. The 15 rare earth element characteristics and 6 trace element characteristics are used as preset element characteristics.
[0069] (3) After determining that there are no zircon samples with completely consistent information data, determine whether there are samples with all the pre-selected trace element characteristics based on the names and content values of the elements contained in the samples. For example, if a sample contains all the elements involved in Ti, Hf, Th, U, Eu / Eu* and Ce / Ce*, then the sample is considered to have all the pre-selected trace element characteristics.
[0070] If there is a sample having all the preselected trace element characteristics, such a sample is regarded as a sample to be retained.
[0071] Generally, there will not be a situation where there are no samples to be retained. If this happens, it means that all the zircon samples obtained are unqualified and step S101 needs to be executed again.
[0072] (4) For each sample to be retained, the following steps are performed:
[0073] A. Determine whether the sample to be retained has the characteristics of all preselected rare earth elements.
[0074] If the sample to be retained has the characteristics of all preselected rare earth elements, that is, the retained sample contains elements La, Ce, Pr, Nd, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu, and Y, then a sample is generated according to the information data of the sample to be retained and added to the geological body data set. This sample includes all the characteristics of preselected rare earth elements, all the characteristics of preselected trace elements, and the category of the corresponding geological body object.
[0075] If the sample to be retained does not have the characteristics of all preselected rare earth elements, determine the missing preselected rare earth element characteristics in the sample to be retained.
[0076] For example, determine that a sample to be retained lacks the Tm element characteristic.
[0077] B. For each missing preselected rare earth element in the sample to be retained, the following steps are performed:
[0078] Determine whether the sample to be retained has the characteristics of two rare earth elements adjacent to the missing preselected rare earth element characteristic in the periodic table of elements.
[0079] For example, the two rare earth elements adjacent to Tm in the periodic table of elements are Er and Yb respectively. Then determine whether the sample to be retained has the characteristics of these two rare earth elements Er and Yb, that is, determine whether the sample to be retained contains Er and Yb and has the content values of these two elements.
[0080] If so, use the data of the two adjacent rare earth element characteristics and calculate the data of the missing preselected rare earth element characteristic by interpolation method, that is, the content value of the missing preselected rare earth element in the sample to be retained. For example, the following formula can be used to calculate the content value of the missing Tm element in the sample to be retained:
[0081]
[0082] where, Tm N is the content value of the Tm element in the Nth sample to be retained calculated; Er N is the content value of the Er element in the Nth sample to be retained; Yb N is the content value of the Yb element in the Nth sample to be retained.
[0083] Supplement the data of the missing preselected rare earth element characteristic, that is, the name and content value of the missing preselected rare earth element, to the information data of the sample to be retained until the information data of the sample to be retained has the characteristics of all preselected rare earth elements.
[0084] C. Generate a sample based on the supplemented and complete information data of the samples to be retained and add it to the geological body data set.
[0085] Complete the construction of the geological body data set in the above manner, so that each sample in the geological body data set contains all preset element characteristics and the category of the corresponding geological body object.
[0086] Step S103: Use the geological body data set to train the initial SVM model using the support vector machine SVM algorithm.
[0087] In an embodiment disclosed by the present invention, before training the initial SVM model using the geological body data set, the samples in the geological body database need to be preprocessed in the following manner:
[0088] First, perform data cleaning on the samples in the zircon data set and delete the samples that meet the preset cleaning conditions.
[0089] Only the zircons formed during the magmatic crystallization process can be used to identify the source rock type and guide regional prospecting. Therefore, the constructed zircon data set only collects magmatic zircons and does not include zircons affected by hydrothermal alteration and metamorphic transformation. Also, since the data in the constructed zircon data set comes from different laboratories and researchers may not have paid attention to the genetic type of zircons when reporting, it is necessary to screen the samples in the zircon data set to exclude as much as possible the zircons affected by metamorphism and remove those zircons from the granites affected by metamorphism. The composition of zircons is extremely vulnerable to the pollution of accessory mineral inclusions, and such zircons usually have the characteristic of enriched light rare earths. In a specific embodiment disclosed by the present invention, the preset cleaning condition can be: only retain the samples where the zircon is La < 0.1 ppm.
[0090] The SVM support vector machine algorithm is a supervised learning algorithm used for classification and regression problems. Its main idea is to find an optimal hyperplane to separate data of different categories and maximize the distance from the closest data points to this hyperplane. The advantages of SVM include good generalization ability in high-dimensional spaces, the ability to handle small sample data, and sparsity on the support vectors of the decision boundary.
[0091] In an embodiment disclosed by the present invention, the following method can be used to execute step S103.
[0092] (1) Adopt a non-linear mapping method to map the data of preset features and the category of the corresponding geological body object in each sample of the geological body data set to a high-dimensional feature space.
[0093] (2) Obtain the optimal classification hyperplane in the high-dimensional feature space.
[0094] (3) An initial SVM model is trained based on the optimal classification hyperplane.
[0095] In another embodiment disclosed by the present invention, as Figure 2 shown, the method for constructing a geoscience binary visualization classification diagram further includes the following steps:
[0096] Step S201: Randomly divide the geological body dataset into a training dataset and a test dataset according to a preset ratio.
[0097] In a specific embodiment disclosed by the present invention, the training dataset and the test dataset can be divided according to a ratio of 8:2.
[0098] Step S202: During the process of training the initial SVM model, the performance of the initial SVM model is evaluated using a 5-fold cross-validation method based on the training dataset, and a performance evaluation result is obtained.
[0099] Among them, the performance evaluation result includes precision, accuracy, recall rate, and F1 score.
[0100] The following is a brief introduction to the 5-fold cross-validation method:
[0101] 1. Divide the dataset: Randomly divide the entire sphene training dataset into five mutually exclusive subsets, and each subset represents a part of the entire data.
[0102] 2. Model training and validation: Iterate five times. Each time, four of the subsets are used for training, and the remaining one subset is used to validate (evaluate) the performance of the model. This means that each subset will serve as the validation set once, and the model is trained on the other four subsets.
[0103] 3. Performance evaluation: In each iteration, the validation set is used to evaluate the performance of the model. In the present invention, performance metrics (accuracy rate, recall rate, and F1 score) are used to measure the performance of the model.
[0104] 4. Aggregate the results: Average the performance metric results of the five iterations to obtain the final performance evaluation result.
[0105] The advantage of five-fold cross-validation is that it makes full use of the data and reduces the risk of overfitting at the same time.
[0106] In the specific embodiment disclosed by the present invention, the following methods can be used to calculate the accuracy rate, precision, recall rate, and F1 score:
[0107] 1. Accuracy rate: It is the ratio of the number of all correctly predicted samples to the total number of samples.
[0108] The calculation formula is: Accuracy = (TP + TN) / (TP + FP + FN + TN)
[0109] 2. Precision: It is the ratio of the number of samples correctly predicted as positive by the model to the number of samples predicted as positive by the model.
[0110] The calculation formula is: Precision = TP / (TP + FP)
[0111] 3. Recall: It is the ratio of the number of samples correctly predicted as positive by the model to the number of all actual positive samples.
[0112] The calculation formula is: Recall = TP / (TP + FN)
[0113] 4. F1 Score: It is the harmonic mean of precision and recall.
[0114] The calculation formula is: F1 Score = 2 × Precision × Recall / (Precision + Recall)
[0115] Among them, TP is the number of positive class samples correctly predicted as positive by the model; TN is the number of negative class samples correctly predicted as negative by the model; FP: the number of negative class samples wrongly predicted as positive by the model; FN: the number of positive class samples wrongly predicted as negative by the model. The positive class is the target class of concern, and it is the class that the model is expected to accurately identify and predict. The negative class is other classes except the positive class, and it is the class that the model is expected to correctly exclude.
[0116] Step S203: Use the Bayesian optimization method and the performance evaluation results of the initial SVM model to optimize the parameters of the initial SVM model to obtain an optimized initial SVM model.
[0117] In an embodiment disclosed by the present invention, the following method can be used to obtain an optimized initial SVM model:
[0118] In machine learning, it is often necessary to adjust the hyperparameters of the model to achieve better performance, and Bayesian optimization can search for the best settings in the parameter space more efficiently.
[0119] The following are the general steps of Bayesian optimization:
[0120] 1. Define the objective function: Use the performance of the machine learning model (such as accuracy, F1 score, etc.) as the objective function, and maximize or minimize this function by adjusting the hyperparameters.
[0121] 2. Select a prior model: Select a surrogate model for the prior knowledge of the objective function. Usually, a Gaussian Process is selected. This surrogate model will help us estimate the unknown performance of the objective function in the parameter space.
[0122] 3. Select a sampling strategy: Select a sampling strategy, such as the "high uncertainty - high reward" strategy in Bayesian optimization, that is, sample in unknown places to explore, and at the same time sample in known high - reward places to exploit.
[0123] 4. Iterative optimization: Start iteratively training the model, evaluating the objective function, and sampling, and then update the prior model using the new observations. In this way, the model will gradually adjust the hyperparameters to find the optimal value of the objective function.
[0124] In an embodiment disclosed by the present invention, the following method can be used to optimize the model:
[0125] (1) According to the performance evaluation results of the SVM model, establish a probability model (i.e., surrogate model) between the SVM model parameters and the performance.
[0126] (2) Use the probability model to predict the next set of parameter combinations that will make the performance evaluation results of the SVM model reach the optimal.
[0127] In the embodiment disclosed by the present invention, the performance evaluation results reaching the optimal can be that all evaluation indicators (for example, precision, recall rate, and F1 - score) reach the maximum value at the current stage, or the sum of all evaluation indicators reaches the maximum value at the current stage, or it can also be other required ways.
[0128] (3) Update the SVM model according to the new parameter combination and re - obtain the performance evaluation results.
[0129] (4) Repeat the above steps, that is, repeatedly use the performance evaluation results of the SVM model to update the probability model, use the updated probability model to predict a set of parameter combinations that will make the performance evaluation results of the SVM model reach the optimal. Update the SVM model again according to the new parameter combination and re - obtain the performance evaluation results. Repeat the above steps multiple times until the performance evaluation results of the SVM model no longer increase (that is, the performance evaluation results are the same as those obtained in the previous iteration), and use the current parameter combination to obtain the optimized SVM model.
[0130] Step S104: Analyze the initial SVM model based on the SHAP algorithm, and select the first preset number of features that have the greatest impact on the output results of the initial SVM model from the preset features as training features.
[0131] In a specific embodiment disclosed by the present invention, the SHAP (SHapley Additive exPlanations) algorithm, an interpretability algorithm of a machine learning model, is used to perform SHAP analysis on the trained initial SVM model, and the 5 features that have the greatest impact on the model output result are determined according to the SHAP analysis result.
[0132] When all 21 preset element features are used as training input features, the SHAP value results of the support vector machine model are obtained. During the experiment, when using the support vector machine model to identify zircon from ore-forming rock masses and non-ore-forming rock masses, the 5 features that have the greatest impact on the output result are Gd, Dy, Yb, Y, and Lu, followed by Tm, Er, Tb, Ho, and Eu / Eu*. Compared with the above 10 elements, Eu, Nd, Sm, La, Th, and Pr have a lower impact on the output result, and the 5 elements that are the least important for ore-forming potential discrimination are Ti, Hf, U, Ce, and Ce / Ce*.
[0133] When the SHAP value of a certain feature is positive, it means that the feature has a positive impact on the predicted value of the model, that is, the more likely the feature variable is to cause the prediction to be an ore-forming rock mass; when the SHAP value of a certain feature is negative, it means that the feature has a negative impact on the predicted value of the model, that is, the more likely the feature variable is to cause the prediction to be a non-ore-forming rock mass. Therefore, in the specific embodiment disclosed by the present invention, Gd, Dy, Yb, Y, and Lu are selected as training features.
[0134] Step S105: Select the second preset number of features from the first preset number of training features in any combination manner to obtain multiple combination results of the features.
[0135] Wherein, the second preset number is less than the first preset number.
[0136] In a specific embodiment disclosed by the present invention, 2 features can be selected from 5 training features in any combination manner to obtain 10 combination results of the features.
[0137] Step S106: Use the geological body data set to retrain the SVM model for the features in each combination result to obtain multiple new SVM models.
[0138] Each of the 10 combination results contains 2 features. Using these 2 features, based on the geological body data set, the SVM model is retrained, and finally 10 new SVM models are obtained according to the 10 combination results.
[0139] Step S107: Construct a binary visualization classification diagram for each new SVM model, wherein the dividing line between different categories in each binary visualization classification diagram is determined by the corresponding SVM model.
[0140] For each new SVM model, a corresponding binary visualization classification diagram is established.
[0141] In an embodiment disclosed by the present invention, the following method can be adopted to determine the dividing line between different categories in each binary visualization classification diagram:
[0142] The decision boundary is determined by using a preset decision function, and the decision function is
[0143] f(x) = w·x + b
[0144] where w is the normal vector of the hyperplane of the SVM model, x is the feature in the SVM model, and b is a preset bias term; the sign (positive or negative sign) of the decision function determines the category of the sample point.
[0145] The dividing line between different categories in the binary visualization classification diagram is determined by using the preset decision function and the decision boundary.
[0146] The decision function is used to determine which category a sample point belongs to, and the decision boundary is the region where the output of the decision function is zero, that is, the dividing line between different categories.
[0147] Step S108: Obtain the feature data of the geological body object to be classified and project it in all binary visualization classification diagrams, and determine the category of the geological body object to be classified according to the projection results in all binary visualization classification diagrams.
[0148] In an embodiment disclosed by the present invention, as Figure 3 shown, the following method can be adopted to implement step S108.
[0149] For each binary visualization classification diagram, the classification result of the geological body object to be classified is obtained in the following manner:
[0150] Step S301: Obtain the features corresponding to the binary visualization classification diagram as the diagram features.
[0151] Step S302: Project the data of the diagram features in the geological body object to be classified in the binary visualization classification diagram to obtain the classification result.
[0152] Step S303: Count the number of each classification result, and take the classification result with the largest number as the category of the geological body object to be classified.
[0153] The method provided by the embodiments of the present invention is different from the conventional visualization diagram construction idea. The method disclosed in the embodiments of the present invention hardly requires prior knowledge and automatically selects the feature information of the X-axis and Y-axis of the binary visualization diagram through machine learning. In addition, in the previously constructed binary visualization diagrams, the boundaries of different types of objects are determined by researchers based on experience, so the boundaries are easily interfered by outliers. In the binary visualization diagram of the disclosed embodiments of the present invention, the boundaries of different types of objects are determined by a machine learning model, which not only eliminates the errors easily caused by subjective factors of people, but also can eliminate the interference of outliers to the greatest extent. The method disclosed in the embodiments of the present invention improves the analysis efficiency and can simultaneously construct 10 visualizations that can be compared with each other, improving the accuracy of the analysis.
[0154] Figure 4 The structure diagram of a system for constructing a geoscience binary visualization classification diagram based on machine learning disclosed in the embodiments of the present invention is used to classify geological bodies according to preset categories, such as Figure 4 As shown, the system includes the following modules:
[0155] The information data acquisition module 11 is configured to acquire the information data of multiple geological body objects. The information data of the geological body objects includes the feature data and category data of the geological body objects;
[0156] The geological body data set construction module 12 is configured to construct a geological body data set according to the information data of the geological body objects. Each sample in the geological body data set contains data of multiple preset features and the category of the corresponding geological body object, and the category of the geological body object belongs to the preset category;
[0157] The initial SVM model training module 13 is configured to use the geological body data set and train an initial SVM model by using the support vector machine SVM algorithm;
[0158] The training feature selection module 14 is configured to analyze the initial SVM model based on the SHAP algorithm and select the first preset number of features that have the greatest influence on the output result of the initial SVM model from the preset features as the training features;
[0159] The feature combination acquisition module 15 is configured to select the second preset number of features from the first preset number of training features in any combination manner to obtain multiple combination results of the features, and the second preset number is less than the first preset number;
[0160] The new SVM model training module 16 is configured to use the geological body data set to retrain the SVM model for the features in each combination result to obtain multiple new SVM models;
[0161] The binary visualization classification diagram module 17 is configured to construct a binary visualization classification diagram for each new SVM model, wherein the dividing line between different categories in each binary visualization classification diagram is determined by the corresponding SVM model;
[0162] The category determination module 18 is configured to obtain the feature data of the geological body object to be classified and project it in all binary visualization classification diagrams, and determine the category of the geological body object to be classified according to the projection results in all binary visualization classification diagrams.
[0163] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention. However, the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.
Claims
1. A method for constructing a geoscience binary visualization classification diagram based on machine learning, characterized in that, For classifying geological body objects according to preset categories, the method includes: Obtaining information data of multiple geological body objects, where the information data of the geological body objects includes characteristic data and category data of the geological body objects; among them, the geological body objects include mineral samples, rock samples, structures, ore bodies, and ore deposits, the characteristics of the geological body objects include the color, texture structure of the object, and the elements contained in the object, and the category of the geological body object is the specific category to which the geological body object belongs; Constructing a geological body data set according to the information data of the geological body objects, where each sample in the geological body data set contains data of multiple preset characteristics and the category of the corresponding geological body object, and the category of the geological body object belongs to the preset category; Using the geological body data set, training an initial SVM model with the support vector machine SVM algorithm; Analyzing the initial SVM model based on the SHAP algorithm, and automatically selecting the first preset number of characteristics that have the greatest impact on the output result of the initial SVM model from the preset characteristics as training characteristics; Selecting the second preset number of characteristics from the first preset number of training characteristics in any combination manner to obtain multiple combination results of the characteristics, where the second preset number is less than the first preset number; Using the geological body data set to retrain the SVM model for the characteristics in each combination result to obtain multiple new SVM models; Constructing a binary visualization classification diagram for each new SVM model, so as to simultaneously construct multiple mutually comparable binary visualization diagrams to improve the accuracy of analysis; among them, the dividing line between different categories in each binary visualization classification diagram is determined by the corresponding SVM model; Obtaining the characteristic data of the geological body object to be classified and projecting it in all binary visualization classification diagrams, and determining the category of the geological body object to be classified according to the projection results in all binary visualization classification diagrams, including: For each binary visualization classification diagram, obtaining the classification result of the geological body object to be classified in the following manner: Obtaining the characteristics corresponding to the binary visualization classification diagram as the diagram characteristics; Projecting the data of the diagram characteristics in the geological body object to be classified in the binary visualization classification diagram to obtain the classification result; Counting the number of each classification result, and taking the classification result with the largest number as the category of the geological body object to be classified.
2. The method according to claim 1, characterized in that, The step of using the geological body data set to train an initial SVM model with the support vector machine SVM algorithm includes: Adopting a non-linear mapping method to map the data of the preset characteristics and the category of the corresponding geological body object in each sample of the geological body data set to a high-dimensional feature space; Obtaining an optimal classification hyperplane in the high-dimensional feature space; Training an initial SVM model according to the optimal classification hyperplane.
3. The method according to claim 1, characterized in that The method further includes: Randomly dividing the geological body data set into a training data set and a test data set according to a preset ratio; During the process of training the initial SVM model, performing performance evaluation on the initial SVM model with a 5-fold cross-validation method based on the training data set to obtain a performance evaluation result; the performance evaluation result includes precision, recall rate, and F1 score; Using the Bayesian optimization method and the performance evaluation results of the initial SVM model, optimize the parameters of the initial SVM model to obtain an optimized initial SVM model.
4. The method according to claim 3, wherein The step of using the Bayesian optimization method and the performance evaluation results of the initial SVM model to optimize the parameters of the initial SVM model to obtain an optimized initial SVM model includes: According to the performance evaluation results of the initial SVM model, establish a probability model between the parameters of the initial SVM model and the performance. Use the probability model to predict the next set of parameter combinations that make the performance evaluation results of the initial SVM model reach the optimal. Update the initial SVM model according to the new parameter combination and obtain the performance evaluation results again. Repeat the above steps until the performance evaluation results of the initial SVM model no longer increase, and use the current parameter combination to obtain an optimized initial SVM model.
5. The method according to claim 1, wherein Construct a binary visualization classification diagram for each new SVM model. Among them, the dividing lines of different classes in each binary visualization classification diagram are determined by the corresponding SVM model, including: Use a preset decision function to determine the decision boundary. The decision function is: f(x) = w·x + b where w is the normal vector of the hyperplane in the SVM model, x is the feature in the SVM model, and b is a preset bias term. Use the preset decision function and the decision boundary to determine the dividing lines of different classes in the binary visualization classification diagram.
6. A system for constructing a geoscience binary visualization classification diagram based on machine learning, characterized in that, A system for classifying geological body objects according to a preset category, the system includes: An information data acquisition module configured to acquire information data of multiple geological body objects. The information data of the geological body objects includes feature data and category data of the geological body objects. Among them, the geological body objects include mineral samples, rock samples, structures, ore bodies, and ore deposits. The features of the geological body objects include the color, structure, and elements contained in the object. The category of the geological body object is the specific category to which the geological body object belongs. A geological body data set construction module configured to construct a geological body data set according to the information data of the geological body objects. Each sample in the geological body data set contains data of multiple preset features and the category of the corresponding geological body object. The category of the geological body object belongs to the preset category. An initial SVM model training module configured to use the geological body data set and train an initial SVM model using the support vector machine SVM algorithm. A training feature selection module configured to analyze the initial SVM model based on the SHAP algorithm and automatically select the first preset number of features that have the greatest impact on the output results of the initial SVM model from the preset features as training features. A feature combination acquisition module configured to select the second preset number of features from the first preset number of training features in any combination manner to obtain multiple combination results of the features. The second preset number is less than the first preset number. A new SVM model training module configured to use the geological body data set to retrain the SVM model for the features in each combination result to obtain multiple new SVM models. The binary visualization classification diagram module is configured to construct a binary visualization classification diagram for each new SVM model respectively, so as to construct multiple mutually comparable binary visualization diagrams simultaneously and improve the accuracy of analysis. Among them, the dividing lines of different classes in each binary visualization classification diagram are determined by the corresponding SVM model; The class determination module is configured to obtain the feature data of the geological body object to be classified and project it in all binary visualization classification diagrams, and determine the class of the geological body object to be classified according to the projection results in all binary visualization classification diagrams, including: For each binary visualization classification diagram, the classification result of the geological body object to be classified is obtained in the following manner: Obtain the features corresponding to the binary visualization classification diagram as diagram features; Project the data of the diagram features in the geological body object to be classified in the binary visualization classification diagram to obtain a classification result; Count the quantity of each classification result, and take the classification result with the largest quantity as the class of the geological body object to be classified.
Citation Information
Patent Citations
Andric rock structure background discrimination diagram method fused with machine learning
CN117113162A
Granite structure environment discrimination method based on machine learning
CN118520281A