Geochemical comprehensive anomaly identification method and system for complex lithologic region

By employing expectation-maximization clustering and the fast minimum covariance determinant method to decompose the mixed normal distribution in complex lithological regions and calculating robust Mahalanobis distance, the accuracy problem of geochemical anomaly identification in complex lithological regions is solved, and more reliable anomaly identification is achieved.

CN122045857APending Publication Date: 2026-05-15JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-04-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In complex lithological regions, existing technologies cannot effectively distinguish the mixed problem of multiple normal distributions, resulting in the Mahalanobis distance deviating from the chi-square distribution and failing to accurately identify geochemical anomalies.

Method used

The expected-maximization clustering algorithm is used to decompose the mixed normal distribution into multiple sub-distributions, and the robust Mahalanobis distance is calculated by the fast minimum covariance determinant method. The threshold for anomaly identification is determined by combining the chi-square distribution.

Benefits of technology

It restores the statistical validity of Mahalanobis distance, improves the accuracy and reliability of geochemical anomaly identification, and can identify weak but meaningful anomalies while suppressing false strong anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045857A_ABST
    Figure CN122045857A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of geochemical exploration and data processing thereof, and provides a geochemical comprehensive anomaly identification method and system for a complex lithologic region, and the method comprises the following steps: selecting a clustering index for clustering based on preprocessed geochemical data; clustering the preprocessed geochemical data by adopting an expectation maximization clustering algorithm based on a clustering index, and decomposing mixed normal distribution into a plurality of sub-bodies; for each daughter, calculating a robust mahalanobis distance of each daughter sample by adopting a fast minimum covariance determinant method; and according to the robust mahalanobis distance of each daughter sample, determining a threshold value for distinguishing the geochemical background and the anomaly based on chi-square distribution, and further identifying the geochemical comprehensive anomaly. According to the method, the statistical effectiveness of the mahalanobis distance of the geochemical sample in the complex lithologic region can be recovered to a certain extent, so that a geochemical comprehensive anomaly recognition result has statistical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of geochemical exploration and data processing technology, and particularly relates to a method and system for identifying comprehensive geochemical anomalies in complex lithological areas. Background Technology

[0002] Geochemical anomaly identification is a crucial technical step in mineral resource exploration, its core being the differentiation between diverse geochemical backgrounds and anomalies. Mahalanobis distance is a commonly used method for geochemical anomaly identification. Under conditions of a multivariate normal distribution, the Mahalanobis distance follows a chi-square distribution, which can be used to determine a statistically significant threshold for distinguishing between multivariate geochemical backgrounds and anomalies. However, in geochemical anomaly identification in lithologically complex areas, the problem of mixed normal distributions often exists. Under such conditions, directly calculating the Mahalanobis distance based on mixed data from the entire area will cause the Mahalanobis distance to deviate from the chi-square distribution, thus failing to accurately identify geochemical anomalies.

[0003] To address the aforementioned issues, existing technologies have proposed various solutions, primarily including robust estimation methods, data transformation methods, and local or adaptive improvement methods. Robust estimation methods improve the robustness of Mahalanobis distance calculations by mitigating the impact of extreme outliers on mean and covariance estimations. However, these methods are bound by the assumption of a "single global population," failing to distinguish between mixed multivariate normal distributions and struggling to address the issue of Mahalanobis distances deviating from the chi-square distribution within complex lithological regions. Data transformation methods only improve the overall distribution of the data at the numerical level, failing to eliminate the inherent structural variations caused by the mixing of multivariate normal distributions. While local adaptive methods or machine / deep learning methods can identify anomalies, their results lack clear statistical basis and geological interpretability.

[0004] In fact, the existing research has not recognized the fundamental reason why the Mahalanobis distance of regional geochemical samples deviates from the chi-square distribution, namely the problem of multivariate normal distribution mixing. Instead, it directly uses mixed data to calculate the Mahalanobis distance, which causes the Mahalanobis distance to deviate from the chi-square distribution. As a result, it is impossible to accurately identify geochemical anomalies. Therefore, it is often difficult to restore the statistical validity of the Mahalanobis distance and to guarantee the reliability of the comprehensive anomaly identification results. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for identifying geochemical anomalies in complex lithological regions, aiming to solve the aforementioned technical problems.

[0006] This invention is implemented as follows: a method for identifying geochemical anomalies in complex lithological regions, comprising the following steps:

[0007] Geochemical data of the study area were acquired and preprocessed to obtain preprocessed geochemical data.

[0008] Based on the preprocessed geochemical data, clustering indices are selected for clustering.

[0009] Based on the clustering index, the expectation-maximization clustering algorithm is used to cluster the preprocessed geochemical data, decomposing the mixed normal distribution into multiple sub-groups;

[0010] For each of the sub-entities, the robust Mahalanobis distance of each sub-entity sample is calculated using the fast minimum covariance determinant method;

[0011] Based on the robust Mahalanobis distance of each daughter sample, a threshold for distinguishing geochemical background from anomalies is determined using the chi-square distribution, thereby identifying integrated geochemical anomalies.

[0012] Furthermore, the preprocessing method is the central logarithmic ratio transformation method.

[0013] Furthermore, the method for selecting the clustering index is as follows: combining the geological background and lithological characteristics of the study area, non-mineralized lithological indicator elements that are sensitive to lithological changes and have weak correlations are selected as clustering indicators.

[0014] Furthermore, based on the clustering index, the preprocessed geochemical data is clustered using the expectation-maximization clustering algorithm to decompose the mixed normal distribution into multiple sub-groups. This step specifically includes:

[0015] The parameters of different normal distributions in the preprocessed geochemical data are estimated by maximum likelihood through iterative methods, and then the mixed normal distribution is divided into several categories; each category corresponds to a lithological subbody.

[0016] The optimal number of clusters is determined based on the Akaike information content criterion.

[0017] Furthermore, the steps for determining the optimal number of clusters based on the Akaike information content criterion specifically include:

[0018] The AIC values ​​of different numbers of clusters were compared using the Akaike Information Content Criterion, and the number of clusters with smaller AIC values ​​was selected.

[0019] The rationality of the clustering results is evaluated based on the distribution characteristics of the main lithological types in the study area.

[0020] Ensure that the number of samples in each sub-category is no less than the preset value of 30.

[0021] Furthermore, the robust Mahalanobis distance for each subsample using the fast minimum covariance determinant method includes the following steps:

[0022] By automatically selecting a subset of samples that can represent the distribution characteristics of the main body, and estimating the robust mean vector and covariance matrix based on the subset;

[0023] The robust Mahalanobis distance for each subsample is calculated based on the robust mean vector and covariance matrix.

[0024] Furthermore, based on the robust Mahalanobis distance of each daughter sample, the steps for determining the threshold distinguishing geochemical background from anomalies using the chi-square distribution specifically include:

[0025] The robust Mahalanobis distance of each subsample was statistically tested to confirm that it follows a chi-square distribution;

[0026] A preset quantile of the chi-square distribution was selected as the threshold for distinguishing geochemical background from anomalies.

[0027] Another objective of this invention is to provide a geochemical anomaly identification system for complex lithological areas, used to implement the aforementioned geochemical anomaly identification method, comprising:

[0028] The data acquisition and preprocessing module is used to acquire geochemical data of the study area and perform preprocessing to obtain preprocessed geochemical data.

[0029] The clustering index selection module is used to select clustering indexes for clustering based on the preprocessed geochemical data.

[0030] The mixed normal distribution decomposition module is used to cluster the preprocessed geochemical data based on the clustering index using the expectation-maximization clustering algorithm, thereby decomposing the mixed normal distribution into multiple sub-groups.

[0031] The Mahalanobis distance calculation module is used to calculate the robust Mahalanobis distance of each sub-sub sample using the fast minimum covariance determinant method for each sub-sub-sub.

[0032] The anomaly identification module is used to determine the threshold for distinguishing geochemical background from anomalies based on the robust Mahalanobis distance of each daughter sample and the chi-square distribution, thereby identifying comprehensive geochemical anomalies.

[0033] The geochemical anomaly identification method for complex lithological regions provided by this invention can, to a certain extent, restore the statistical validity of Mahalanobis distances of geochemical samples in complex lithological regions, making the geochemical anomaly identification results statistically significant. In addition, this invention can, to a certain extent, improve the reliability of geochemical anomaly identification results, mainly by not only identifying some weak but meaningful anomalies, but also suppressing some false strong anomalies. Attached Figure Description

[0034] Figure 1 A schematic flowchart illustrating the geochemical anomaly identification method for complex lithological regions provided in this embodiment of the invention.

[0035] Figure 2 This is a graph showing the relationship between the number of clusters and the AIC value of sediment samples from a river system.

[0036] Figure 3 This is a graph showing the clustering results of EM.

[0037] Figure 4 The robust Mahalanobis distance cumulative probability curves are for clusters 1(a), 2(b), 3(c), 4(d), 5(e), 6(f), and the entire region (g).

[0038] Figure 5 This is a schematic diagram of the Li-Be-Sn-F-Bi geochemical anomaly identified by the method (a) and the direct global FMCD method (b) of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0040] When using Mahalanobis distance to identify geochemical anomalies, a basic statistical requirement must be met: the data must follow a multivariate normal distribution. Only under this premise will the calculated Mahalanobis distance follow a chi-square distribution, allowing for the reasonable determination of the threshold for distinguishing background data from anomalies based on statistical criteria. Studies show that within a single lithological region, geochemical data typically follows a multivariate normal distribution; however, in complex lithological regions, data often originates from a mixture of multiple multivariate normal distributions. Ignoring this mixing phenomenon and directly using mixed data to calculate the Mahalanobis distance will cause the distance to deviate from the chi-square distribution, thus rendering geochemical anomaly identification statistically unfounded. The solution provided in this invention aims to distinguish mixed normal distributions in complex lithological regions, restore the statistical validity of the Mahalanobis distance, and thereby improve the accuracy and reliability of comprehensive geochemical anomaly identification. Specifically, this invention proposes a geochemical anomaly identification method that combines statistical rigor and geological interpretability: First, the expectation-maximization (EM) clustering algorithm is used to distinguish the mixed normal distributions in complex lithological regions. Then, the Mahalanobis distance is calculated for each normal distribution using the Fast Minimum Covariance Determinant (FMCD) method, providing technical support for the geochemical anomaly identification in complex lithological regions.

[0041] like Figure 1As shown, in one embodiment of the present invention, a method for identifying comprehensive geochemical anomalies in complex lithological regions is provided. This method employs a three-stage processing framework: hybrid normal distribution decomposition, robust Mahalanobis distance calculation, and comprehensive geochemical anomaly identification. It addresses the problem of identifying comprehensive geochemical anomalies in complex lithological regions and specifically includes the following steps:

[0042] Step 1: Decomposition of Mixed Normal Distribution: Obtain geochemical data of the study area and preprocess it to obtain preprocessed geochemical data; Based on the preprocessed geochemical data, select clustering indices for clustering; Based on the clustering indices, use the EM clustering algorithm to cluster the preprocessed geochemical data, decomposing the mixed normal distribution into multiple sub-groups;

[0043] Step 2, Robust Mahalanobis Distance Calculation: For each sub-sub, the robust Mahalanobis distance of each sub-sub sample is calculated using the FMCD method;

[0044] Step 3: Geochemical anomaly identification: Based on the robust Mahalanobis distance of each daughter body sample, a threshold for distinguishing between geochemical background and anomalies is determined based on the chi-square distribution, thereby identifying geochemical anomalies.

[0045] In a preferred embodiment of the present invention, in order to eliminate the influence of the closure effect, the embodiment of the present invention uses the central logarithmic ratio transformation method to process the geochemical data of the study area, which not only preserves all elemental information, but also has a certain robustness to outliers.

[0046] In a preferred embodiment of the present invention, the EM clustering algorithm is used to decompose the mixed normal distribution. First, a reasonable clustering index must be selected. In this embodiment of the present invention, the method for selecting the clustering index is as follows: combining the geological background and lithological characteristics of the study area, lithological indicator elements (such as SiO2, Al2O3, Y, Zr, etc.) that are non-mineralized, sensitive to lithological changes, and have weak correlations are preferentially selected as clustering indexes to improve the rationality and interpretability of the clustering results.

[0047] In a preferred embodiment of the present invention, the step of clustering the preprocessed geochemical data using the EM clustering algorithm based on clustering indices to decompose the mixed normal distribution into multiple sub-groups specifically includes:

[0048] The parameters of different normal distributions in the preprocessed geochemical data are estimated by maximum likelihood through iterative methods. Then, the mixed normal distribution is divided into several classes, resulting in several (approximate) normal distribution sub-body; each class corresponds to a lithological sub-body.

[0049] The optimal number of clusters is determined based on the Akaike Information Content Criterion (AIC).

[0050] In determining the optimal number of clusters, this embodiment of the invention comprehensively considers the following factors: First, the Akaike Information Content Criterion (AIC) is used to compare the AIC values ​​of different cluster numbers, selecting clusters with the smallest possible AIC values; second, the reasonableness of the clustering results is evaluated in conjunction with the distribution characteristics of the main lithological types in the study area; third, the number of samples contained in each subgroup is ensured to be sufficiently large, generally not less than a preset value of 30, to meet the robustness requirements of subsequent robust Mahalanobis distance calculations. Based on the above factors, the final number of clusters is determined.

[0051] In a preferred embodiment of the present invention, the method for calculating the robust Mahalanobis distance of each daughter sample using the FMCD method includes the following steps:

[0052] By automatically selecting a subset of samples that can represent the distribution characteristics of the main body, and estimating the robust mean vector and covariance matrix based on the subset;

[0053] Robust Mahalanobis distances for each subsample are calculated based on robust mean vectors and covariance matrices to restore the statistical validity of Mahalanobis distances.

[0054] Specifically, after completing the mixed normal distribution decomposition, the FMCD method is introduced within each sub-body to robustly estimate the mean vector and covariance matrix. The FMCD method automatically selects the majority sample subset representing the main distribution characteristics from the sub-body geochemical samples and estimates statistical parameters based on this subset, thereby effectively reducing the impact of individual outlier samples on the mean vector and covariance structure. After obtaining the robust mean and covariance matrix using the FMCD method, the robust Mahalanobis distance (MD) for each sub-body sample can be calculated. The specific calculation formula is as follows:

[0055] ;

[0056] In the formula, X is a variable normal distribution. The multidimensional observation vector; μ is the robust mean vector. A robust covariance matrix is ​​established. Finally, Mahalanobis distance values ​​at specific quantiles of the chi-square distribution are used to identify geochemical anomalies.

[0057] In a preferred embodiment of the present invention, the step of determining the threshold for distinguishing geochemical background from anomalies based on the robust Mahalanobis distance of each daughter sample and the chi-square distribution specifically includes:

[0058] The robust Mahalanobis distance of each subsample was statistically tested to confirm that it follows a chi-square distribution;

[0059] A preset quantile of the chi-square distribution was selected as the threshold for distinguishing geochemical background from anomalies.

[0060] In this embodiment of the invention, the robust Mahalanobis distance within each sub-subunit is statistically tested. Under the premise of meeting the statistical characteristics requirements, based on a specific quantile of the theoretical chi-square distribution (e.g., 0.975), the threshold for distinguishing between geochemical background and anomaly can be determined, thereby achieving accurate identification of geochemical anomalies.

[0061] Example: Taking the identification of Li, Be, Sn, F, and Bi geochemical anomalies in 1:200,000 stream sediments in northeastern Hunan as an example, the implementation process and application effects of the method provided in this embodiment of the invention in complex lithological areas are described in detail.

[0062] I. Data Preprocessing and Determination of Ore-forming Element Assemblage: Considering the closure effect that is common in geochemical data, a central logarithmic ratio transformation was performed on the 1:200,000 stream sediment geochemical data of northeastern Hunan to ensure that the data is suitable for subsequent analysis and processing.

[0063] Based on this, factor analysis was used to extract elemental combinations related to Li mineralization. The results showed that Li, Be, Sn, F, and Bi had high loadings on the same factor, indicating that these elements have a certain indicative role in lithium mineralization in the study area. Therefore, Li-Be-Bi-F-Sn was used as the combined variable for identifying comprehensive lithium anomalies.

[0064] II. Selection of Clustering Indicators: Considering that ore-forming elements and associated elements tend to accumulate significantly in local mineralization areas, and their distribution is mainly controlled by mineralization, they are excluded during the mixed normal distribution decomposition. Based on the characteristics of the study area, which is dominated by intermediate-acidic igneous rocks, sandstone, siltstone, metamorphic sandstone, and mudstone, elements that can reflect lithological characteristics, such as Ba, Ni, Sr, Th, Ti, U, V, Y, Zr, Al₂O₃, CaO, Fe₂O₃, K₂O, MgO, Na₂O, and SiO₂, are selected as clustering indicators for mixed normal distribution decomposition.

[0065] III. Mixed Normal Distribution Decomposition and Result Analysis Based on EM Clustering: After determining the clustering indices, the EM clustering algorithm was used to cluster geochemical samples of stream sediments in the study area. The trend of AIC value changing with the number of clusters (e.g., ...) was analyzed. Figure 2 As shown in the figure, considering the basic statistical requirements of lithological complexity and sample size in the study area, the optimal number of clusters was finally determined to be 6. The clustering results show good consistency with the main lithological units in the study area in terms of spatial distribution (e.g., ...). Figure 3 As shown in the figure, this indicates that the mixed normal distribution has been effectively decomposed.

[0066] Among them, cluster 1 (312 samples) spatially overlaps with clay minerals in the northwest of the study area; cluster 2 (199 samples) is mainly distributed in the mudstone and siltstone distribution areas of the study area; cluster 3 (865 samples) is widely distributed in the central and southern parts of the study area, corresponding to slurred sandstone; cluster 4 (302 samples) corresponds to sandstone in the study area; cluster 5 (554 samples) is the main igneous rock type in the study area, representing monzogranite plutons; and cluster 6 (215 samples) has a good correspondence with granite in the study area. Figure 3 The high correlation between the various types of sediment samples and regional lithology indicates that EM clustering has divided the stream sediment samples into samples with clear geological significance, laying a solid foundation for subsequent identification of geochemical anomalies.

[0067] IV. Robust Mahalanobis Distance Calculation for Sub-body Samples: After completing the mixed normal distribution decomposition, robust Mahalanobis distance calculations were performed for geochemical samples of each sub-body. First, the robust mean vector and covariance matrix of each sub-body sample were robustly estimated using the FMCD method; then, the robust Mahalanobis distance of each sub-body sample was calculated based on the two robust statistical parameters of the robust mean vector and covariance matrix.

[0068] V. Statistical Characteristic Tests and Outlier Identification: The cumulative probability curves of robust Mahalanobis distances for each subsample are very close to the cumulative probability curves of the theoretical chi-square distribution in terms of overall distribution shape (e.g., Figure 4 As shown in af), the cumulative probability curve of the robust Mahalanobis distance calculated directly from the whole area data without mixed normal distribution decomposition shows a significantly different situation, deviating significantly from the chi-square distribution (as shown in af). Figure 4 As shown in g). The above analysis shows that mixed normal distribution decomposition and robust parameter estimation can effectively restore the statistical validity of Mahalanobis distance. Based on this, using the chi-square distribution quantile as the anomaly threshold, and dividing the anomaly intensity into low, medium, and high levels at 1, 2, and 4 times its value, a comprehensive geochemical anomaly map of Li-Be-Sn-F-Bi based on the EM and FMCD combined method was drawn (as shown in g). Figure 5 (as shown in a).

[0069] VI. Evaluation of Method Effectiveness: To evaluate the application effectiveness of the method provided in this embodiment of the invention, robust Mahalanobis distances calculated directly using the FMCD method without mixed normal distribution decomposition were used to delineate the Li-Be-Sn-F-Bi geochemical anomaly (e.g., Figure 5 (As shown in b). The analysis results show that the method provided by the embodiments of the present invention reduces the anomaly area from 18.3% of the total area to 11.9%, while increasing the identification rate of known mineral deposits from 75% to 100%. This means that anomalies identified based on the method provided by the embodiments of the present invention can significantly reduce the exploration area and increase the probability of exploration success, thereby reducing exploration costs.

[0070] Furthermore, the anomalies identified by the method in this embodiment of the invention have more explicit geological significance and correspond better to the metallogenic geological background of known deposits in the study area, specifically in the following aspects:

[0071] (1) Eliminating spurious anomalies in high background areas: Due to geological processes such as magma differentiation, Li and associated elements are enriched in granite bodies, and these areas can be called high background areas. If the mixed normal distribution is not distinguished and the Mahalanobis distance is calculated directly using the FMCD method, large-area, high-intensity anomalies will be generated in high background areas, such as areas B1 and B2 (e.g., Figure 5 As shown in b). Among these, most anomalies have not yet shown any signs of mineralization. However, the method provided in this invention, through mixed normal distribution decomposition and robust Mahalanobis distance calculation of sub-body, effectively suppresses these high-intensity but meaningless spurious anomalies (such as...). Figure 5 As shown in a), it significantly improves the accuracy of anomaly detection.

[0072] (2) Enhancing weak anomalies in low-background areas: Compared with granite bodies, the content of Li and associated elements in other types of rocks in the study area is relatively low, which can be called low-background areas. If the mixed normal distribution is not distinguished and the Mahalanobis distance is calculated directly using the FMCD method, the anomaly intensity in low-background areas will be suppressed, or even made undetectable, such as areas A1, A2, A3, A4 and A5 (e.g. Figure 5 (As shown in a and b). Among them, regions A3 and A5 have known lithium deposits, but no anomalies are shown in the anomaly map delineated by the global FMCD method (e.g., Figure 5 As shown in b). The method provided in this embodiment of the invention effectively extracts such low-intensity but very important anomalies (e.g., through mixed normal distribution decomposition and robust Mahalanobis distance calculation of the sub-entities). Figure 5 (as shown in a). These anomalies are closely related to the NE-trending fault structures and the spatial relationship between the granite body and the surrounding rock contact zone, and have strong mineral exploration indication significance.

[0073] In summary, the value of the method provided in this invention lies not only in restoring the statistical validity of Mahalanobis distances in geochemical samples from complex lithological regions, but also in significantly improving the accuracy of geochemical anomaly identification. It transforms anomalies from statistical outliers into exploration targets with clear "geological constraints." The successful application of the method provided in this invention demonstrates that decomposing the mixed normal distribution in complex lithological regions and restoring the statistical validity of Mahalanobis distances is an important way to improve the accuracy of geochemical anomaly identification.

[0074] In another embodiment of the present invention, a geochemical anomaly identification system for complex lithological areas is also provided to implement the above-mentioned geochemical anomaly identification method, specifically including:

[0075] The data acquisition and preprocessing module is used to acquire geochemical data of the study area and perform preprocessing to obtain preprocessed geochemical data.

[0076] The clustering index selection module is used to select clustering indexes for clustering based on the preprocessed geochemical data.

[0077] The mixed normal distribution decomposition module is used to cluster preprocessed geochemical data based on clustering indices and employ the expectation-maximization clustering algorithm, decomposing the mixed normal distribution into multiple sub-groups.

[0078] The Mahalanobis distance calculation module is used to calculate the robust Mahalanobis distance of each sub-sub sample using the fast minimum covariance determinant method for each sub-sub-sub sample.

[0079] The anomaly identification module is used to determine the threshold for distinguishing geochemical background from anomalies based on the robust Mahalanobis distance of each daughter sample and the chi-square distribution, thereby identifying comprehensive geochemical anomalies.

[0080] It should be noted that each of the above modules can be implemented as a computer program, which can run on a computer device. The computer device's memory can store the computer program that makes up each module, enabling the processor to execute each step of the above method.

[0081] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0082] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.

[0083] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.

Claims

1. A method for identifying comprehensive geochemical anomalies in complex lithological regions, characterized in that, Includes the following steps: Geochemical data of the study area were acquired and preprocessed to obtain preprocessed geochemical data. Based on the preprocessed geochemical data, clustering indices are selected for clustering. Based on the clustering index, the expectation-maximization clustering algorithm is used to cluster the preprocessed geochemical data, decomposing the mixed normal distribution into multiple sub-groups; For each of the sub-entities, the robust Mahalanobis distance of each sub-entity sample is calculated using the fast minimum covariance determinant method; Based on the robust Mahalanobis distance of each daughter sample, a threshold for distinguishing geochemical background from anomalies is determined using the chi-square distribution, thereby identifying integrated geochemical anomalies.

2. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 1, characterized in that, The preprocessing method is the central logarithmic ratio transformation method.

3. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 1, characterized in that, The clustering index was selected by combining the geological background and lithological characteristics of the study area, and selecting non-mineralized lithological indicator elements that are sensitive to lithological changes and have weak correlations as clustering indexes.

4. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 1, characterized in that, Based on the clustering index, the preprocessed geochemical data is clustered using the expectation-maximization clustering algorithm to decompose the mixed normal distribution into multiple sub-groups. This process specifically includes: The parameters of different normal distributions in the preprocessed geochemical data are estimated by maximum likelihood through iterative methods, and then the mixed normal distribution is divided into several categories; each category corresponds to a lithological subbody. The optimal number of clusters is determined based on the Akaike information content criterion.

5. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 4, characterized in that, The steps for determining the optimal number of clusters based on the Akaike information content criterion specifically include: The AIC values ​​of different numbers of clusters were compared using the Akaike Information Content Criterion, and the number of clusters with smaller AIC values ​​was selected. The rationality of the clustering results is evaluated based on the distribution characteristics of the main lithological types in the study area. Ensure that the number of samples in each sub-category is no less than the preset value of 30.

6. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 1, characterized in that, The robust Mahalanobis distance for each subsample using the fast minimum covariance determinant method includes the following steps: By automatically selecting a subset of samples that can represent the distribution characteristics of the main body, and estimating the robust mean vector and covariance matrix based on the subset; The robust Mahalanobis distance for each subsample is calculated based on the robust mean vector and covariance matrix.

7. The method for identifying comprehensive geochemical anomalies in complex lithological areas according to claim 1, characterized in that, Based on the robust Mahalanobis distance of each daughter sample, the steps for determining the threshold for distinguishing geochemical background from anomalies using the chi-square distribution include: The robust Mahalanobis distance of each subsample was statistically tested to confirm that it follows a chi-square distribution; A preset quantile of the chi-square distribution was selected as the threshold for distinguishing geochemical background from anomalies.

8. A geochemical anomaly identification system for complex lithological areas, used to implement the geochemical anomaly identification method according to any one of claims 1-7, characterized in that, include: The data acquisition and preprocessing module is used to acquire geochemical data of the study area and perform preprocessing to obtain preprocessed geochemical data. The clustering index selection module is used to select clustering indexes for clustering based on the preprocessed geochemical data. The mixed normal distribution decomposition module is used to cluster the preprocessed geochemical data based on the clustering index using the expectation-maximization clustering algorithm, thereby decomposing the mixed normal distribution into multiple sub-groups. The Mahalanobis distance calculation module is used to calculate the robust Mahalanobis distance of each sub-sub sample using the fast minimum covariance determinant method for each sub-sub-sub. The anomaly identification module is used to determine the threshold for distinguishing geochemical background from anomalies based on the robust Mahalanobis distance of each daughter sample and the chi-square distribution, thereby identifying comprehensive geochemical anomalies.