A geochemical anomaly recognition and interpretation method based on sequential machine learning

By employing sequential machine learning methods, combined with deep autoencoder networks and XGBoost models, global and local interpretations of geochemical anomalies were achieved. This addresses the problem of insufficient information utilization in existing technologies, improves the interpretability of geochemical anomalies, and guides mineral exploration.

CN122494049APending Publication Date: 2026-07-31江西省地质局第十地质大队 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
江西省地质局第十地质大队
Filing Date
2026-06-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

While existing technologies make full use of all geochemical elements, they are unable to effectively identify and interpret comprehensive geochemical anomalies, resulting in an incomplete understanding of mineralization.

Method used

A geochemical anomaly identification and interpretation method based on sequential machine learning is adopted, including data collection and processing, preprocessing, deep autoencoder network identification, XGBoost regression model construction and Shapley method interpretation. By combining deep autoencoder networks and XGBoost models, global and local interpretations of geochemical anomalies can be achieved.

Benefits of technology

By making full use of all geochemical element information, the interpretability of comprehensive geochemical anomalies identified by machine learning models has been improved, guiding regional mineral exploration and providing a more comprehensive understanding of mineralization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494049A_ABST
    Figure CN122494049A_ABST
Patent Text Reader

Abstract

This invention relates to a method for geochemical anomaly identification and interpretation based on sequential machine learning, belonging to the field of geology and mineral resources. It involves collecting and organizing geochemical data based on mineral exploration needs; preprocessing the geochemical data to eliminate closure effects; identifying and interpreting comprehensive geochemical anomalies based on sequential machine learning using the preprocessed geochemical data; and delineating single-element geochemical anomaly zones based on the comprehensive geochemical anomaly identification results to determine the type and location of mineral deposits. This method can fully utilize all geochemical elements while providing global and local interpretations of comprehensive geochemical anomalies, improving the interpretability of comprehensive geochemical anomalies identified by the machine learning model and providing guidance for regional mineral exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of geology and mineral resources, specifically relating to a method for identifying and interpreting geochemical anomalies based on sequential machine learning. Background Technology

[0002] Geochemical elements constitute the basic material components of various geological bodies. As geological processes continue, elements in these bodies are constantly activated, migrated, and enriched. Influenced by the diversity of physicochemical conditions and their interactions, as well as the differences in the chemical behavior of elements, different combinations of geochemical elements are ultimately formed. Therefore, studying the patterns of element combination in different geological processes is of great significance for understanding the migration and enrichment of geochemical elements, assessing geological environments, and revealing the genesis of mineral deposits.

[0003] Because mineralization is a unique geological process, chemical elements originally dispersed in the Earth's crust, upper mantle, and hydrosphere accumulate in relatively confined geological environments, forming mineral deposits. This process also causes the distribution of certain chemical elements in the deposit area to deviate significantly from their normal ranges in the surrounding region, resulting in geochemical anomalies. Therefore, identifying geochemical element assemblages and effectively detecting geochemical anomalies are crucial tasks in geochemical exploration.

[0004] Geochemical element assemblages screening is the foundation for comprehensive geochemical anomaly interpretation. Traditional geochemical element assemblage screening methods include: (1) multivariate statistical methods, such as principal component analysis and factor analysis; and (2) geochemical element assemblage screening methods that integrate spatial structure, such as local spatial autocorrelation index, ROC curve, multifractal analysis, and geographic detectors. However, current traditional geochemical element assemblage screening methods only consider positive anomalies, resulting in a relatively one-sided understanding of mineralization provided by the obtained geochemical anomalies.

[0005] In the field of geochemical anomaly identification, artificial intelligence algorithms such as machine learning are widely used in geochemical anomaly identification because they can effectively handle large-volume, high-dimensional data with complex nonlinear relationships. These methods are mainly divided into: (1) pixel-based anomaly identification methods, such as restricted Boltzmann machines, artificial neural networks, a class of support vector machines, isolated forests, and deep autoencoder networks; (2) image-based anomaly identification methods, such as convolutional neural networks, convolutional autoencoder networks, and spatial-spectral bibranch neural networks; and (3) topological graph-based anomaly identification methods, such as figure neural networks.

[0006] While current mainstream geochemical element screening methods can interpret geochemical anomalies to some extent, they cannot fully utilize the mineralization information contained in all geochemical elements. Although using all geochemical elements and machine learning can fully learn the information contained in geochemical elements, it faces the problem of difficulty in interpreting the comprehensive geochemical anomalies identified. How to fully utilize all geochemical elements while improving the interpretability of comprehensive geochemical anomalies identified by machine learning models is a key technical problem that urgently needs to be solved. Summary of the Invention

[0007] To address the technical shortcomings of existing technologies, the purpose of this invention is to provide a geochemical anomaly identification and interpretation method based on sequential machine learning. This method considers all geochemical elements, and geochemical anomalies obtained using all geochemical elements can provide a more comprehensive understanding of mineralization. It can improve the interpretability of comprehensive geochemical anomalies identified by the machine learning model while fully utilizing all geochemical elements, thereby guiding regional mineral exploration.

[0008] To achieve the above objectives, the technical solution adopted by this invention is as follows: This invention discloses a method for geochemical anomaly identification and interpretation based on sequential machine learning, the method comprising the following steps:

[0009] S1. Collect and organize geochemical data based on mineral exploration needs;

[0010] S2. Preprocess geochemical data to eliminate the closure effect of geochemical data;

[0011] S3. Based on the preprocessed geochemical data, perform comprehensive geochemical anomaly identification and interpretation using sequential machine learning.

[0012] S4. Based on the comprehensive geochemical anomaly identification results, delineate single-element geochemical anomaly zones to determine the type and location of mineral deposits.

[0013] Furthermore, step S1 also includes collecting geological maps of the study area.

[0014] Furthermore, step S1 involves organizing the geochemical data, including unifying the data coordinate system, cleaning the collected geochemical data, and filling in missing values.

[0015] Furthermore, step S2 involves preprocessing the geochemical data by dividing the study area into a series of grid cells of equal size, performing inverse distance weighted interpolation on each geochemical element, extracting the obtained element content values ​​into the grid cells, performing compositional data analysis on the geochemical data, and using equidistant logarithmic ratio transformation to eliminate the closure effect of the geochemical data.

[0016] Furthermore, the formula for the isochronous logarithmic ratio transformation in step S2 is as follows:

[0017] (1)

[0018] in, These are primitive geochemical elements, and they are primitive variables. The new variable after the logarithmic ratio transformation is called. The number of geochemical elements.

[0019] Furthermore, step S3 includes the following sub-steps:

[0020] S31. Construct a comprehensive geochemical anomaly identification model based on deep autoencoder networks to identify comprehensive geochemical anomalies;

[0021] S32. Use XGBoost to construct a regression model between geochemical elements and integrated geochemical anomalies;

[0022] S33. The Shapley method is used to interpret the regression model constructed by XGBoost, and the global and local contributions of each geochemical element to the integrated geochemical anomaly are obtained, thereby realizing the global and local interpretation of the integrated geochemical anomaly.

[0023] Furthermore, in step S31, the learning rate, network depth, and number of neurons in each layer of the deep autoencoder network are continuously adjusted, and the reconstruction error obtained by selecting the set of parameters with the smallest loss function is taken as the comprehensive geochemical anomaly.

[0024] Furthermore, in step S32, the preprocessed geochemical elements are used as independent variables, and the comprehensive geochemical anomalies identified by the deep autoencoder network are used as dependent variables. Regression analysis is performed using a regression model constructed by XGBoost. Five-fold cross-validation is used to adjust the parameters of the regression model constructed by XGBoost, and the parameter combination with the smallest mean square error is selected as the optimal parameter combination.

[0025] Furthermore, in step S33, based on the optimal parameter combination analyzed by the regression model constructed by XGBoost, the Shapley method is used to interpret the integrated geochemical anomaly, obtaining the Shapley value of each element at each grid position, thereby achieving a local interpretation of the integrated geochemical anomaly. By calculating the mean of the absolute values ​​of the Shapley values ​​of each geochemical element, the contribution of each geochemical element to the integrated geochemical anomaly is obtained, thereby achieving a global interpretation of the integrated geochemical anomaly.

[0026] Furthermore, step S4 includes classifying the comprehensive geochemical anomalies identified by the deep autoencoder network using standard deviation, and screening high-value areas of comprehensive geochemical anomalies; converting the calculated local Shapley values ​​corresponding to each geochemical element into a raster map, screening out areas with Shapley values ​​greater than 0, and performing intersection analysis between the high-value areas of comprehensive geochemical anomalies and the areas with relevant geochemical element Shapley values ​​greater than 0 to obtain the corresponding geochemical element anomaly areas. Using the same method, the anomaly areas of each geochemical element can be obtained, thereby determining the type and location of the mineral deposit.

[0027] The beneficial technical effects of this invention are as follows: Compared with the existing geochemical anomaly identification methods, the geochemical anomaly identification and interpretation method based on sequential machine learning disclosed in this invention can fully utilize the information contained in all geochemical elements, and can also perform global and local interpretation of comprehensive geochemical anomalies, providing new ideas and methods for determining target mineral types and delineating single-element geochemical anomaly areas. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating a geochemical anomaly identification and interpretation method based on sequential machine learning, as shown in Embodiment 1 of the present invention.

[0029] Figure 2 This is a schematic diagram of a local interpretation of a comprehensive geochemical anomaly region obtained using the geochemical anomaly identification and interpretation method based on sequential machine learning described in Embodiment 1 of the present invention.

[0030] Figure 3 The curves showing the change of the loss function of the deep autoencoder network model for different learning rates;

[0031] Figure 4 This invention employs a geochemical anomaly identification and interpretation method based on sequential machine learning, as described in Embodiment 1, to identify comprehensive geochemical anomalies based on deep autoencoder networks.

[0032] Figure 5 The method for identifying and interpreting geochemical anomalies based on sequential machine learning, as described in Embodiment 1 of this invention, is based on the ROC curve of a comprehensive geochemical anomaly identified by a deep autoencoder network.

[0033] Figure 6 Correlation analysis between integrated geochemical anomalies identified by deep autoencoder networks and predicted values ​​from XGBoost regression models;

[0034] Figure 7 The overall contribution of different geochemical elements to the integrated geochemical anomaly reflects the global interpretation of the integrated geochemical anomaly;

[0035] Figure 8 The distribution map of Shapley values ​​for Cu reflects a local interpretation of the comprehensive geochemical anomalies.

[0036] Figure 9 This is a distribution map showing the abnormal distribution of Cu. Detailed Implementation

[0037] The present invention will now be further described with reference to the accompanying drawings and specific embodiments.

[0038] Example 1

[0039] In this embodiment, the geochemical anomaly identification and interpretation method based on sequential machine learning disclosed in this invention is illustrated by taking the identification and interpretation of Ag-Pb-Zn polymetallic deposits in the Jinxi-Lengshuikeng area as an example.

[0040] The Jinxi-Lengshui area is tectonically located on the western slope of the Wuyi Mountains, at the forefront of the collision zone between the Pacific and Eurasian plates, on the southern edge of the Jinning-Caledonian Cathaysia-Yangtze Plate amalgamation zone, and in the back-arc basin west of the Yanshanian Circum-Pacific Plate's Fuchong Belt. The tectonic unit belongs to the northern edge of the South China Orogenic Belt of the Cathaysia Plate, spanning two third-order tectonic units: the Wugongshan-Northern Wuyi (front-edge fold-thrust) Uplift Belt and the Central-Southern Wuyi Uplift of the Southeastern Jiangxi Uplift Belt. Since the Jinningian period, the study area has long been situated on an active continental margin, undergoing multiple phases of tectonic-magmatic-mineralization processes, resulting in the renowned Wuyishan copper-lead-zinc-gold-silver polymetallic metallogenic belt. The exposed strata in the area, from oldest to youngest, consist of the Neoproterozoic Zhoutanyan Formation, Wanyuanyan Formation, and Hongshan Formation; the Paleozoic Waiguankeng Formation, Zishan Formation, and Outangdi Formation; the Mesozoic Shuibei Formation, Ruyiting Formation, Daguding Formation, Ehuling Formation, Lengshuiwu Formation, and Maodian Formation; and the Cenozoic Quaternary Lianwei Formation. Among these, the Qingbaikou-Cambrian deep-to-shallow marine sedimentary clastic rocks form the main basement structure. The cover strata consist of Carboniferous-Cretaceous strata, representing shallow marine-terrestrial sediments. The area exhibits complex tectonic features, with well-developed NE and NNE-trending faults. North-south and northwest-trending faults form a supporting system for the NE-trending faults. The intersections of secondary structures in different directions provide pathways and ore-bearing sites for the formation of various minerals.

[0041] The region experienced intense and frequent volcanic-magmatic activity, with igneous rocks occurring from the Caledonian period to the Late Yanshanian period. Early Yanshanian intrusions are predominantly distributed in a northwest-trending pattern, with the early Yanshanian intrusions in Raoqiao, Wutaishan, Liandanping, Lengshuikeng, and Dagangxia area being subsurface volcanic rocks. Late Yanshanian intrusions are more scattered. Volcanic rocks are mainly distributed in Early Cretaceous and Late Jurassic strata, forming a set of continental intermediate-acidic volcanic rocks, creating several volcanic tectonic depressions such as Tiantaishan, Yuefengshan, and Meiyuancun. The intermediate-acidic igneous rocks of the Middle and Late Yanshanian periods are closely related to mineralization. The region experienced prolonged volcanic activity, resulting in complex lithology and lithofacies. The Qingbaikou and Nanhuai volcanic rocks, interbedded within metamorphic strata, are the oldest exposed volcanic rocks in the region, having undergone intense later deformation and metamorphism. Mesozoic volcanic rocks are mainly intermediate-acidic and intermediate-acidic to alkaline pyroclastic rocks and lava, primarily distributed in the Late Jurassic Ruyiting Formation and Daguding Formation, and the Early Cretaceous Ehuling Formation and Lengshuiwu Formation. Volcanic activity was primarily characterized by intense eruptions, with additional features of effusive and eruptive-depositional characteristics. Specifically, the Late Jurassic Daguding Formation consists of eruptive clastic flow deposits and eruptive sedimentary clastic rock deposits; the Early Cretaceous Ehuling Formation is mainly distributed within four primary volcanic structures: the Jinyuan eruptive basin, the Tiantaishan tectonic depression, the Liandanping uplift, and the eastern Guluoshan caldera. Volcanic eruptions were intense, primarily eruptive-effusive, and characterized by fissure-central eruptions. Volcanic activity generally migrated from northwest to southeast within the surveyed area. Both formations are the main ore-bearing strata of the Lengshuikeng super-large subvolcanic Ag-Pb-Zn polymetallic deposit within the area. In recent years, significant breakthroughs have been made in deep mineral exploration in this region, yielding a series of metallogenic insights that suggest the region possesses superior metallogenic geological conditions and enormous mineral exploration potential. Traditional geochemical anomaly identification methods are insufficient for effectively identifying and interpreting related geochemical anomalies. Therefore, we utilize a geochemical anomaly identification and interpretation method based on sequential machine learning to extract geochemical anomalies in this region, providing support for deep and peripheral mineral exploration predictions.

[0042] like Figure 1 As shown, this embodiment of the invention provides a method for geochemical anomaly identification and interpretation based on sequential machine learning, the method comprising the following steps:

[0043] S1. Collect and organize geochemical data based on mineral exploration needs.

[0044] Based on the required prediction accuracy, an appropriate scale is selected, and geochemical data is collected and organized according to identification needs. This also includes collecting geological maps of the study area. Organizing the geochemical data involves standardizing the data coordinate system, cleaning the collected geochemical data, and imputing missing values.

[0045] In this embodiment of the invention, 1:50,000 mineral geological maps and 1:50,000 stream sediment geochemical measurement data of the Jinxi-Lengshuikeng area were collected. Data units were standardized, and missing values ​​were supplemented.

[0046] The mineral geological map includes strata, structures, igneous rocks, and mineral deposits. The geochemical measurement data includes 13 geochemical elements such as Au, Ag, Cu, Pb, Zn, Mn, Sn, Mo, Cd, W, As, Sb, and Bi.

[0047] S2. Preprocess the geochemical data to eliminate the closure effect of the geochemical data.

[0048] Preprocessing of geochemical data includes dividing the study area into a series of equally sized grid cells, performing inverse distance weighting (IDW) interpolation on each geochemical element, and extracting the elemental abundance values ​​into the grid cells. Compositional data analysis is then performed on the geochemical data, using an isometric logarithmic ratio transformation to eliminate the "closing effect" in the geochemical data, resulting in transformed variables. The specific isometric logarithmic ratio transformation formula is as follows:

[0049] (1)

[0050] in, These are primitive geochemical elements, and they are primitive variables. The new variable after the logarithmic ratio transformation is called. The number of geochemical elements.

[0051] Continuing from the previous example, the study area was divided into a series of equally sized grid cells. Inverse distance weighted (IDW) interpolation was performed on the 13 geochemical anomaly data to obtain raster maps of the content of each element. The element content in each raster layer was extracted into the grid cells, and a logarithmic ratio (ilr) transformation was used to obtain new variables, which were then normalized.

[0052] S3. Based on the preprocessed geochemical data, perform comprehensive geochemical anomaly identification and interpretation using sequential machine learning.

[0053] like Figure 1As shown, firstly, an unsupervised deep autoencoder network is used to identify geochemical anomalies, resulting in a comprehensive geochemical anomaly. Then, XGBoost (eXtreme Gradient Boosting) is used to construct a regression model between geochemical elements and the comprehensive geochemical anomaly. Finally, the Shapley method is used to interpret the regression model constructed by XGBoost, obtaining the global and local contributions of each geochemical element to the comprehensive geochemical anomaly. The geochemical anomalies are then interpreted based on these global and local contributions.

[0054] Step S3 includes the following sub-steps:

[0055] S31. Construct a comprehensive geochemical anomaly identification model based on deep autoencoder networks to identify comprehensive geochemical anomalies.

[0056] A deep autoencoder network is used to reconstruct preprocessed geochemical element variables. The reconstruction error is calculated based on the original and reconstructed variables, and a comprehensive geochemical anomaly identification model is constructed accordingly. Figure 3 As shown, the learning rate, network depth, and number of neurons in each layer of the deep autoencoder network are continuously adjusted, such as... Figure 4 As shown, the reconstruction error obtained by selecting the set of parameters with the smallest loss function is taken as the comprehensive geochemical anomaly.

[0057] assumed This indicates the input of geochemical elements (input variables). As a hidden layer, This represents the reconstructed geochemical elements (i.e., the output variables). During the encoding process, the input data is mapped to the hidden layers through a non-linear activation function: During the decoding phase, the hidden layer is reconstructed into output data through an activation function: .in and As weight, and This is a bias term.

[0058] Deep autoencoders (DAEs) are an unsupervised variant of deep belief networks. The model consists of multiple restricted Boltzmann machines, employing a layer-by-layer greedy training algorithm to optimize network weights, thereby encoding and decoding input data. Deep autoencoders typically have a symmetrical structure, and the difference between the input and output is the reconstruction error.

[0059] (2)

[0060] Generally, the reconstruction error of geochemical anomalies is often greater than that of normal values. Therefore, the larger the reconstruction error at a given location, the higher the probability that it is an anomaly.

[0061] Continuing from the previous example, such as Figure 5 As shown, the geochemical anomalies identified by the deep autoencoder network were evaluated using ROC curves. The AUC value of the ROC curve was 0.8670, indicating that the geochemical anomalies identified by the deep autoencoder network have a high spatial correlation with known Ag-Pb-Zn polymetallic deposits, and the prediction results are good. The geochemical anomalies identified by the deep autoencoder network were classified according to their standard deviations to obtain the comprehensive geochemical anomaly zones.

[0062] S32. Construct an XGBoost regression model based on geochemical elements and integrated geochemical anomalies.

[0063] Using preprocessed geochemical elements as independent variables and comprehensive geochemical anomalies identified by deep autoencoder networks as dependent variables, regression analysis was performed using a regression model constructed with XGBoost. Five-fold cross-validation was used to adjust the parameters of the regression model constructed with XGBoost, and the parameter combination with the minimum mean square error was selected as the optimal parameter combination.

[0064] Extreme Gradient Boosting (XGBoost) is a scalable end-to-end tree boosting algorithm whose core objective is to learn an approximation function that minimizes the loss function with regularization.

[0065] (3)

[0066] (4)

[0067] In the formula, Let l be the representation in linear space, and l be the loss function used to measure the i-th sample. Predicted value Compared with the true value The differences between them; Indicates the complexity of the model; The number of leaf nodes in the tree; This represents the weight of each leaf node.

[0068] Assumption Let be the predicted value of the t-th tree for the i-th sample. In this case, the loss function can be expressed using an additive method:

[0069] (5)

[0070] In the formula, This represents the prediction result for the i-th sample after the t-th iteration.

[0071] Optimize the objective function using second-order Taylor expansion:

[0072] (6)

[0073] in, and These are the first and second gradients of the loss function, respectively.

[0074] For each tree structure, XGBoost uses split gain to control the optimal split of each leaf node:

[0075] (7)

[0076] in, Represents leaf nodes The sample set, Representing a tree structure, and These represent the left and right child nodes of the sample set after the split, respectively.

[0077] Continuing the previous example, using the new variables obtained by isochronous logarithmic ratio transformation and normalization of 13 geochemical elements (Au, Ag, Cu, Pb, Zn, Mn, Sn, Mo, Cd, W, As, Sb, and Bi) as independent variables, and the comprehensive geochemical anomaly identified by a deep autoencoder network as the dependent variable, regression analysis is performed using the XGBoost algorithm. Five-fold cross-validation is used to adjust the XGBoost model parameters, selecting the parameter combination with the minimum mean squared error (MSE). The corresponding dependent variable is then predicted using this parameter combination. Figure 6 As shown, a linear fit was performed between the anomalies identified by the deep autoencoder network and the prediction results of the XGBoost regression model. The results show that the slope of the fitted line is 0.9597, and the coefficient of determination of the goodness of fit (R²) is high. 2 The value of 0.9453 indicates that the model's prediction results are good.

[0078] S33. Identification and interpretation of integrated geochemical anomalies based on the Shapley method.

[0079] Based on the optimal parameter combination analyzed by the regression model constructed by XGBoost, the Shapley method is used to interpret the integrated geochemical anomaly, obtaining the Shapley value of each element at each grid position, thus achieving a local interpretation of the integrated geochemical anomaly. By calculating the mean of the absolute values ​​of the Shapley values ​​of each geochemical element, the contribution of each geochemical element to the integrated geochemical anomaly is obtained, thus achieving a global interpretation of the integrated geochemical anomaly.

[0080] Based on the overall contribution of geochemical elements to the integrated geochemical anomaly, a global interpretation of the integrated geochemical anomaly is carried out. The geochemical elements are ranked according to the magnitude of their contribution to the integrated geochemical anomaly. Based on this ranking, it can be determined which element combinations contribute significantly to the geochemical anomaly in the study area, and the potential mineralization types and target minerals can be inferred.

[0081] A local interpretation of the integrated geochemical anomaly is performed based on the Shapley value of each element at each grid location. This measures the contribution of each element to the integrated anomaly at each grid location. By combining the global contribution of the integrated geochemical anomaly with the local contribution at each grid location, it is possible to determine which elements are causing the anomalies at these locations, thereby identifying the elemental anomalies at certain local locations and potentially determining what mineral deposits might be found.

[0082] The Shapley method is a post-hoc additive feature attribution method based on cooperative game theory. It can explain machine learning and deep learning models at both global and local scales, thereby significantly improving the transparency of artificial intelligence models.

[0083] For a given sample instance, the Shapley value is calculated as follows:

[0084] (8)

[0085] in, The Shapley value represents the i-th geochemical element, reflecting its contribution to the overall geochemical anomaly. A positive Shapley value indicates that the element has a positive effect on the overall geochemical anomaly; conversely, a negative value indicates that the element has a negative effect on the overall geochemical anomaly. The larger the value, the greater the contribution of the i-th element to the overall geochemical anomaly. F represents a subset of all geochemical elements except element i; F represents the set of all geochemical elements. Indicates the XGBoost model on a subset The prediction results; Indicates in subset The predicted results after adding the i-th geochemical element; and is a weighting factor used to calculate the probability of occurrence of different combinations of geochemical elements.

[0086] Following the previous example, based on the optimal parameter combination obtained from XGBoost regression analysis, the Shapley method is used to interpret the integrated geochemical anomaly, obtaining the Shapley value for each element at each raster location, thus achieving a local interpretation of the integrated geochemical anomaly. For example... Figure 7As shown, by calculating the mean of the Shapley absolute values ​​of each geochemical element, the contribution of each geochemical element to the integrated geochemical anomaly can be obtained, thereby achieving a global interpretation of the integrated geochemical anomaly.

[0087] The order of contribution of different geochemical elements to the overall geochemical anomaly, from largest to smallest, is W > Cd > Sb > Pb > Mo > As > Sn > Zn > Cu > Au > Bi > Ag > Mn. Among these, W contributes the most to the overall geochemical anomaly; however, no large W deposits have been found in this region, while granitic magmatic rocks generally have high W content. Therefore, the W contribution may indicate the presence of concealed granitic intrusive bodies at depth, providing a heat source and material source for the enrichment of geochemical elements. The Cd–Sb–Pb element assemblage is an indicator element assemblage of epithermal Ag-Pb-Zn mineralization. These elements are typically widely distributed and anomalously stable, thus contributing significantly to the overall geochemical anomaly. Meanwhile, the high contribution of Pb suggests the possible presence of epithermal Pb mineralization in the region. Mo's contribution follows closely, indicating the possible existence of a porphyry-type mineralization system and the potential presence of porphyry-type Mo mineralization.

[0088] The As-Sn-Zn-Cu-Au-Bi-Ag sequence, ranking relatively low, likely directly indicates the main mineralization type in the region. Both Cu and Au contribute to the overall geochemical anomaly, suggesting the possible presence of porphyry-type to epithermal Cu mineralization at depth. The low Au contribution may be due to its inherently low content and limited distribution, resulting in a lower overall contribution to the geochemical anomaly. The lowest Ag contribution is likely due to its low enrichment and dispersed occurrence. Based on the ranking of geochemical element contributions to the overall geochemical anomaly, it can be inferred that the mineralization type in this region is likely dominated by a porphyry-epithelial-epithelial mineralization system, with Pb-Zn-Ag as the primary target mineral, and Cu-Au potentially present at depth. This inference is largely consistent with the deep exploration and verification results of the Lengshuikeng Ag-Pb-Zn polymetallic deposit in this region.

[0089] By projecting the Shapley values ​​of each element onto the graph, we obtain the distribution map of the Shapley values ​​of each geochemical element, such as... Figure 8 As shown, taking Cu as an example, the Shapley value of Cu is projected onto the figure to obtain the distribution map of Cu Shapley values, which is a local interpretation of the comprehensive geochemical anomaly.

[0090] S4. Based on the comprehensive geochemical anomaly identification results, delineate single-element geochemical anomaly zones to determine the type and location of mineral deposits.

[0091] The integrated geochemical anomalies identified by the deep autoencoder network were graded using standard deviation, and high-value areas of the integrated geochemical anomalies were screened. The local Shapley values ​​corresponding to each geochemical element were converted into raster maps. Relevant geochemical elements (such as Cu) were screened out, and areas with Shapley values ​​greater than 0 (i.e., those contributing positively to the integrated geochemical anomaly) were identified. Figure 2 As shown, by performing intersection analysis between areas with high values ​​of the overall geochemical anomaly and areas where the selected geochemical element's Shapley value is greater than 0, the corresponding geochemical element anomaly areas can be obtained. Using the same method, anomaly areas for each geochemical element can be obtained, thereby enabling the interpretation of the overall geochemical anomaly.

[0092] Continuing the previous example, areas with Shapley values ​​greater than 0 (i.e., positive contributions from geochemical elements) are selected. These are then overlaid with areas of abnormally high values ​​identified by a deep autoencoder network (e.g.,...). Figure 2 As shown), Figure 9 As shown, single-element geochemical anomalies were obtained, i.e., the corresponding geochemical element anomaly regions were identified. Based on the anomaly distribution of Cu, a relatively significant Cu anomaly exists in the Lengshuikeng area, suggesting the possible presence of Cu mineralization at depth. Simultaneously, deep boreholes ( Figure 9 The purple triangle area revealed a deep Cu ore body in the Cu anomaly zone, verifying the effectiveness of the anomaly identification and interpretation method.

[0093] As can be seen from the above embodiments, the geochemical anomaly identification and interpretation method based on sequential machine learning disclosed in this invention first organizes and preprocesses geochemical data; then, it uses a deep autoencoder network to identify comprehensive geochemical anomalies; secondly, it uses an XGBoost model to identify geochemical elements and comprehensive geochemical anomalies identified by the deep autoencoder network; finally, it uses the Shapley method to interpret the XGBoost regression model, delineates single-element geochemical anomaly areas, and determines the type and location of mineral deposits.

[0094] The method described in this invention is not limited to the embodiments described in the specific implementation. Other implementation methods derived by those skilled in the art based on the technical solution of this invention also fall within the scope of technical innovation of this invention.

Claims

1. A method for identifying and interpreting geochemical anomalies based on sequential machine learning, the method comprising the following steps: S1. Collect and organize geochemical data based on mineral exploration needs; S2. Preprocess geochemical data to eliminate the closure effect of geochemical data; S3. Based on the preprocessed geochemical data, perform comprehensive geochemical anomaly identification and interpretation using sequential machine learning. S4. Based on the comprehensive geochemical anomaly identification results, delineate single-element geochemical anomaly zones to determine the type and location of mineral deposits.

2. The sequential machine learning based geochemical anomaly identification and interpretation method according to claim 1, characterized in that: Step S1 also includes collecting geological maps of the study area.

3. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 1, characterized in that: Step S1 involves organizing the geochemical data, including unifying the data coordinate system, cleaning the collected geochemical data, and filling in missing values.

4. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 1, characterized in that: Step S2 involves preprocessing the geochemical data by dividing the study area into a series of equal-sized grid cells, performing inverse distance weighted interpolation on each geochemical element, extracting the content values ​​of each element into the grid cells, performing compositional data analysis on the geochemical data, and using equidistant logarithmic ratio transformation to eliminate the closure effect of the geochemical data.

5. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 4, characterized in that: The formula for the logarithmic ratio transformation in step S2 is as follows: (1) in, These are primitive geochemical elements, and they are primitive variables. The new variable after the logarithmic ratio transformation is called. The number of geochemical elements.

6. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 1, characterized in that, Step S3 includes the following sub-steps: S31. Construct a comprehensive geochemical anomaly identification model based on deep autoencoder networks to identify comprehensive geochemical anomalies; S32. Use XGBoost to construct a regression model between geochemical elements and integrated geochemical anomalies; S33. The Shapley method is used to interpret the regression model constructed by XGBoost, and the global and local contributions of each geochemical element to the integrated geochemical anomaly are obtained, thereby realizing the global and local interpretation of the integrated geochemical anomaly.

7. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 6, characterized in that: In step S31, the learning rate, network depth, and number of neurons in each layer of the deep autoencoder network are continuously adjusted, and the reconstruction error obtained by selecting the set of parameters with the smallest loss function is taken as the comprehensive geochemical anomaly.

8. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 7, characterized in that: In step S32, the preprocessed geochemical elements are used as independent variables, and the comprehensive geochemical anomalies identified by the deep autoencoder network are used as dependent variables. Regression analysis is performed using a regression model constructed with XGBoost. Five-fold cross-validation is used to adjust the parameters of the regression model constructed with XGBoost, and the parameter combination with the smallest mean square error is selected as the optimal parameter combination.

9. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 8, characterized in that: In step S33, the optimal parameter combination is analyzed based on the regression model constructed by XGBoost. The Shapley method is used to interpret the integrated geochemical anomaly, and the Shapley value of each element at each grid position is obtained, thereby realizing the local interpretation of the integrated geochemical anomaly. By calculating the mean of the absolute value of the Shapley value of each geochemical element, the contribution of each geochemical element to the integrated geochemical anomaly is obtained, thereby realizing the global interpretation of the integrated geochemical anomaly.

10. The method for geochemical anomaly identification and interpretation based on sequential machine learning according to claim 9, characterized in that: Step S4 includes classifying the comprehensive geochemical anomalies identified by the deep autoencoder network using standard deviation, and screening high-value areas of comprehensive geochemical anomalies; converting the calculated local Shapley values ​​corresponding to each geochemical element into a raster map, screening out areas with Shapley values ​​greater than 0, and performing intersection analysis between the high-value areas of comprehensive geochemical anomalies and the areas with Shapley values ​​greater than 0 of the relevant geochemical elements to obtain the corresponding geochemical element anomaly areas. Using the same method, the anomaly areas of each geochemical element are obtained, thereby determining the type and location of the mineral deposit.