Intelligent recognition method and system for geochemical anomaly based on machine learning

By constructing sparse activation weight vectors and autoencoder neural networks using machine learning methods and combining them with Gaussian mixture models, geochemical anomaly identification results are generated. This solves the problems of anomaly boundary drift and unclear demarcation in traditional methods, achieving higher identification accuracy and stability.

CN122224355APending Publication Date: 2026-06-16GEOLOGICAL SURVEY BUREAU OF YUNNAN PROVINCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEOLOGICAL SURVEY BUREAU OF YUNNAN PROVINCE
Filing Date
2026-04-30
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Traditional geochemical anomaly identification methods rely on fixed threshold division and manual interpretation, which makes it difficult to simultaneously characterize the correlation differences between sample points. This leads to anomaly boundary drift, unstable demarcation results, easy mixing of weak anomalies with background fluctuations, unclear regional classification, and a tendency for misjudgment and omission in mineral exploration.

Method used

A machine learning-based intelligent identification method for geochemical anomalies is adopted. By acquiring a set of multi-element concentration vectors, a sparse activation weight vector is constructed. An autoencoder neural network is used to extract the input signals of hidden nodes. Combined with a Gaussian mixture model and a probability energy distribution matrix, an anomaly category boundary structure is generated to evaluate the confidence of the anomaly state.

Benefits of technology

It improves the accuracy and clarity of anomaly detection, enhances the ability to identify weak anomaly transition areas, reduces the interference of background disturbances on the identification results, and improves the stability and accuracy of the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224355A_ABST
    Figure CN122224355A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of anomaly identification, in particular to a geochemical anomaly intelligent identification method and system based on machine learning, which comprises the following steps: obtaining a multi-element concentration vector and constructing an element combination index sequence and a sparse activation weight, extracting and compressing a representation through a self-encoding neural network, calculating a component response difference value based on a Gaussian mixture model and generating an anomaly category boundary structure, and combining the probability energy deviation of a neighborhood node and time compensation to output an identification result.In the present application, the feature expression is jointly constrained by multi-element combination sorting and occurrence frequency, the boundary candidate area is identified by combining compressed representation and component response comparison, and the spatial transition state is described by using category attribution weight and local reconstruction probability, thereby reducing background fluctuation interference, improving anomaly identification accuracy and category boundary clarity, enhancing result stability and weak anomaly identification ability, and facilitating geochemical anomaly result refinement interpretation and application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of anomaly recognition technology, and in particular to a method and system for intelligent recognition of geochemical anomalies based on machine learning. Background Technology

[0002] The field of anomaly identification technology involves the detection, judgment, and classification of abnormal phenomena in various types of data that deviate from normal distribution or patterns. It covers core matters such as statistical analysis, data acquisition, feature extraction, pattern discrimination, and result labeling. In geological and resource exploration scenarios, it typically includes the acquisition of geochemical measurement data, analysis of sample element content, study of spatial distribution patterns, and delineation of anomaly thresholds. It forms the basis for anomaly identification by systematically organizing and comparing the multi-element combination characteristics of different regions.

[0003] Among them, the traditional geochemical anomaly intelligent identification method refers to the process of identifying anomalous information in geochemical measurement data. It usually involves classifying the element content data of sampling points and setting fixed thresholds in combination with geological background values. Contour maps or profile maps are drawn by statistically sorting the ratios of single or multiple elements. Data points that are higher or lower than the threshold are marked according to empirical rules. At the same time, manual interpretation and classification are carried out in combination with regional distribution density and differences between neighboring sampling points, so as to complete the identification and classification of anomalous areas.

[0004] Traditional identification methods rely on fixed thresholds and manual interpretation. When faced with multi-element coupling relationships and spatial transition features, it is difficult to simultaneously characterize the correlation differences between sample points. Anomaly boundaries often drift due to empirical trade-offs, resulting in insufficient stability of the demarcation results. Weak anomalies and background fluctuations are easily mixed together, the regional categories are unclear, the strength levels of anomalies between different regions are difficult to compare uniformly, and the boundary transition zone lacks clear discrimination criteria. In subsequent mineral exploration judgments, misjudgments and omissions are prone to occur, affecting the consistency of results and the efficiency of verification. The depth of data utilization is also relatively limited. Summary of the Invention

[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a machine learning-based intelligent identification method for geochemical anomalies, comprising the following steps: To achieve the above objectives, the present invention adopts the following technical solution: a machine learning-based intelligent identification method for geochemical anomalies, comprising the following steps: S1: Obtain a set of multi-element concentration vectors, sort them according to the data size of each dimension to construct an element combination index sequence, detect the frequency of occurrence of the sample set associated with the element combination index sequence, and multiply it with the concentration gradient difference array associated with the multi-element concentration vector set to generate a sparse activation weight vector. S2: Input the set of multi-element concentration vectors into an autoencoder neural network to extract hidden node input signals, multiply them with sparse activation weight vectors to generate node weighted input signals, combine them with a preset truncation threshold to remove low-value data segments, and output a compressed representation parameter vector. S3: Call the Gaussian mixture model to receive the compressed representation parameter vector, calculate the component response value sequence associated with the global sampling nodes, sort them in descending order, extract the first two component values ​​and subtract them to obtain the component response difference sequence. S4: Based on the component response difference sequence, filter out global sampling nodes whose values ​​are lower than the preset sensitivity threshold, generate a set of boundary candidate nodes, extract the category belonging weight values ​​and assign them to the category, and generate an abnormal category boundary structure. S5: Extract the spatial set of neighboring nodes based on the anomaly category boundary structure, generate a probability energy distribution matrix by squaring the local reconstruction probability values, evaluate the confidence parameter of the anomaly state, and combine it with the time compensation coefficient to map the anomaly level and generate geochemical anomaly identification results.

[0006] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Obtain a set of multi-element concentration vectors, detect internal dimensional distribution attribute parameters based on basic mapping logic, perform scale alignment to extract a general feature matrix for the internal dimensional distribution attribute parameters, call the underlying sorting operator to perform spatial recombination transformation on the general feature matrix to construct a topological recombination parameter set, aggregate multi-dimensional feature nodes based on the topological recombination parameter set to establish a basic mapping combination, and generate an element combination index sequence. S102: Monitor the element combination index sequence, call the global scanning mechanism based on the state matching strategy to extract the sample distribution feature space, perform structural comparison on the sample distribution feature space to obtain the feature overlap indicator parameter, aggregate relevant response nodes to construct discrete similar clusters based on the feature overlap indicator parameter, perform numerical accumulation calculation on the discrete similar clusters to extract statistical benchmark numbers, and obtain the frequency of occurrence of the sample set. S103: Based on the frequency of occurrence of the sample set, call the numerical mapping model to detect the local spatial gradient difference representation term associated with the multi-element concentration vector set, perform dimension merging and transformation on the local spatial gradient difference representation term, construct the concentration gradient difference array, extract the multiplication calculation rules, perform product fusion operation on the frequency of occurrence of the sample set and the concentration gradient difference array, and obtain the sparse activation weight vector.

[0007] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: For a set of multi-element concentration vectors, call an autoencoder neural network to perform forward propagation mapping calculation to extract the hidden layer feature states associated with the input parameters. Perform linear mapping on the hidden layer feature states to analyze the node weight distribution pattern. Based on the node weight distribution pattern, aggregate the internal flow nodes of the network to generate hidden node input signals. S202: Call the sparse activation weight vector, perform dimension alignment to verify the compatibility of the spatial structure based on the hidden node input signal, perform element-wise multiplication on the two sets of sequences according to the compatibility to extract weighted response features, perform normalization processing on the weighted response features to construct a feature activation distribution array, and obtain the node weighted input signal; S203: Obtain a preset truncation threshold, compare the parameter values ​​of the weighted input signal of the node with the truncation threshold to extract amplitude difference identifiers, remove redundant data segments with values ​​lower than the truncation threshold based on the amplitude difference identifiers, retain key response features, perform dimension reorganization and compression calculation on the key response features, and obtain a compressed representation parameter vector.

[0008] As a further aspect of the present invention, the process of setting the truncation threshold specifically involves: collecting the globally weighted amplitude sequence of the sample set in the current data batch output by the hidden layer of the autoencoder neural network; calculating the mean and standard deviation parameters of the data distribution of the globally weighted amplitude sequence; and performing a summation calculation on the mean and standard deviation parameters of the data distribution to obtain the truncation threshold.

[0009] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Obtain the preset global sampling nodes, extract the feature mean matrix for the compressed representation parameter vector, configure the distribution kernel function included in the hybrid distribution model according to the feature mean matrix, calculate the probability density value of the global sampling nodes in the spatial dimension, perform normalization processing on the probability density value to extract the attribution probability array of individual components, aggregate the attribution probability array of all nodes, and establish the component response value sequence. S302: Based on the component response numerical sequence, detect the response probability values ​​recorded within the sequence, extract the swapping rules, perform cyclic comparison and address shifting processing on each response probability value corresponding to the same node within the component response numerical sequence, reconstruct the internal index mapping table according to the data decreasing characteristics, and perform spatial displacement rearrangement on the values ​​according to the internal index mapping table to generate a descending response numerical sequence. S303: For the descending response numerical sequence, extract the maximum probability value of the first data item and the second maximum probability value of the second data item, call the arithmetic rule to perform a subtraction operation on the maximum probability value and the second maximum probability value to obtain the boundary sensitive parameter, and merge the boundary sensitive parameters corresponding to each node in order according to the spatial distribution order of the nodes to obtain the component response difference sequence.

[0010] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Obtain a preset sensitivity threshold, perform item-by-item difference comparison between each value in the component response difference sequence and the sensitivity threshold, determine the amplitude attribute of each node, filter out global sampling nodes with difference amplitudes lower than the sensitivity threshold, perform coordinate parameter aggregation and recombination on global sampling nodes to establish a basic spatial index, and generate a set of boundary candidate nodes; S402: Based on the set of boundary candidate nodes, detect the local numerical space of the internal node mapping, perform segmentation processing on the local numerical space, construct a multi-dimensional numerical interval, count the frequency of data landing points of the set of boundary candidate nodes within the multi-dimensional numerical interval, call the operation rules to perform normalization transformation on the frequency of data landing points to extract the probability distribution representation term, and obtain the category belonging weight value. S403: Based on the category attribution weight value, extract the probability extreme value term of the boundary candidate node set, assign the associated node to the target label according to the probability extreme value term to determine the node's category, perform spatial optimization to extract the edge turning point coordinates for nodes with differentiated categories, perform connectivity calculation on the edge turning point coordinates to construct the spatial region contour, and establish an abnormal category boundary structure.

[0011] As a further aspect of the present invention, the process of setting the sensitive threshold specifically involves: extracting historical response difference items corresponding to the global sampling nodes; sorting the historical response difference items in ascending order to construct a benchmark difference sequence; calculating the slope change rate of adjacent numerical sequence segments in the benchmark difference sequence; extracting data nodes where the slope change rate exhibits extreme value abrupt changes; and setting the difference value corresponding to the data node as the sensitive threshold.

[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Based on the anomaly category boundary structure, delineate the associated target within the global sampling node, construct a neighborhood node space set, extract the local reconstruction probability value of each node in the neighborhood node space set, perform a square operation on the local reconstruction probability value to extract the energy metric, integrate each energy metric into a two-dimensional array according to the spatial arrangement, and generate a probability energy distribution matrix. S502: Based on the probability energy distribution matrix, calculate the arithmetic mean of all elements to determine the overall energy mean. Subtract the values ​​of the elements within the matrix from the overall energy mean to extract the absolute difference and construct an energy deviation parameter. Sort the energy deviation parameter in descending order according to the numerical index. Extract the extreme value ranking data and divide it with the overall energy mean to obtain the energy ratio coefficient. Read the preset equilibrium setting threshold and compare the two features to obtain the abnormal state confidence parameter. S503: Based on the anomaly confidence parameter, collect timestamp data of the sampling test process, call the attenuation model to process the timestamp data to extract time drift correlation terms, establish time compensation coefficients, perform item-by-item product calculations on the internal recorded values ​​of the anomaly confidence parameter and the time compensation coefficients to perform numerical fusion and recombination, map the anomaly level, and generate geochemical anomaly identification results.

[0013] As a further aspect of the present invention, the process of comparing two features by reading the preset equilibrium setting threshold specifically involves: extracting the energy deviation parameter of each of the top ten ranked items to establish extreme value ranking data; performing a division calculation using the extreme value ranking data as the dividend and the overall energy mean as the divisor, and extracting the division quotient to form the energy proportion coefficient; calculating the standard deviation index of all data elements in the probability energy distribution matrix; retrieving the tolerance multiplier set based on the numerical distribution law, and multiplying the standard deviation index by the tolerance multiplier to output the fluctuation assessment parameter; performing an addition operation between the fluctuation assessment parameter and the overall energy mean to set the equilibrium setting threshold; comparing each value in the energy proportion coefficient with the equilibrium setting threshold; outputting a high-order identifier when each value is greater than the equilibrium setting threshold, and outputting a low-order identifier when each value is not greater than the equilibrium setting threshold; and combining the high-order identifier and the low-order identifier in order to form the abnormal state confidence parameter.

[0014] A machine learning-based intelligent identification system for geochemical anomalies, comprising: The feature arrangement module obtains a set of multi-element concentration vectors, sorts them according to the data size of each dimension to construct an element combination index sequence, detects the frequency of occurrence of the sample set associated with the element combination index sequence, and calculates the sparse activation weight vector by multiplying it with the concentration gradient difference array associated with the set of multi-element concentration vectors. The compression mapping module inputs the multi-element concentration vector set into an autoencoder neural network to extract the hidden node input signal, multiplies it with the sparse activation weight vector to generate the node weighted input signal, combines it with a preset truncation threshold to remove low-value data segments, and outputs a compressed representation parameter vector. The response parsing module calls the Gaussian mixture model to receive the compressed representation parameter vector, calculates the component response value sequence associated with the global sampling nodes, extracts the first two component values ​​after sorting them in descending order, subtracts them, and obtains the component response difference sequence. The boundary discrimination module, based on the component response difference sequence, filters out global sampling nodes whose values ​​are lower than a preset sensitivity threshold, generates a set of boundary candidate nodes, extracts the category assignment weight values ​​and assigns them to categories, and generates an abnormal category boundary structure. The anomaly assessment module extracts a spatial set of neighboring nodes based on the anomaly category boundary structure, generates a probability energy distribution matrix by squaring the local reconstruction probability values, assesses the confidence parameter of the anomaly state, and maps the anomaly level in conjunction with the time compensation coefficient to generate geochemical anomaly identification results.

[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, weight constraints are constructed by combining multi-element concentration vectors according to their combination relationships and frequency of occurrence, and sparse activation intensity is formed by combining concentration gradient differences. This allows highly correlated features to be centrally represented. Boundary candidate information is extracted through compressed representation and mixed component response sorting. The spatial transition state and energy deviation are characterized by combining category assignment weights and local reconstruction probabilities. This continuously converges the abnormal boundary and category boundary, reduces the interference of background disturbances on the recognition results, improves the accuracy of anomaly judgment, the clarity of the boundary and the stability of the results, and enhances the ability to identify weak anomaly transition regions. This allows multi-element information and spatial neighborhood information to be used synergistically, supporting the refined interpretation of anomaly results. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0020] Please see Figure 1This invention provides a machine learning-based intelligent identification method for geochemical anomalies, comprising the following steps: S1: Obtain a set of multi-element concentration vectors, sort the multi-element concentration vector set according to the data size of each dimension in the multi-element concentration vector set to construct an element combination index sequence, detect the frequency of occurrence of the sample set associated with the element combination index sequence, and perform multiplication calculation on the frequency of occurrence of the sample set and the concentration gradient difference array associated with the multi-element concentration vector set to generate a sparse activation weight vector. S2: Input the multi-element concentration vector set into the autoencoder neural network to extract the hidden node input signal. The autoencoder neural network multiplies the hidden node input signal with the sparse activation weight vector to generate the node weighted input signal. The node weighted input signal is compared with the truncation set threshold. Data segments with values ​​lower than the truncation set threshold in the node weighted input signal are removed and the compressed representation parameter vector is output. S3: Call the Gaussian mixture model to receive the compressed representation parameter vector, calculate the component response numerical sequence associated with the global sampling nodes, sort the component response numerical sequence in descending order to generate a descending response numerical sequence, extract the first two component values ​​in the descending response numerical sequence and perform a subtraction operation to obtain the component response difference sequence. S4: Compare the component response difference sequence with the sensitivity threshold, filter the global sampling nodes with values ​​lower than the sensitivity threshold to generate a set of boundary candidate nodes, divide the numerical space associated with the set of boundary candidate nodes into numerical intervals to extract the category assignment weight values, assign the category to the set of boundary candidate nodes based on the category assignment weight values ​​and generate an abnormal category boundary structure. S5: Based on the anomaly category boundary structure, extract the spatial set of neighboring nodes from the global sampling nodes. Perform a square operation on the local reconstruction probability values ​​associated with the spatial set of neighboring nodes to generate a probability energy distribution matrix. Calculate the energy deviation parameter between the matrix element values ​​and the overall energy mean associated with the probability energy distribution matrix. Calculate the energy proportion coefficient based on the energy deviation parameter of the extreme value ranking. Compare the energy proportion coefficient with the equilibrium set threshold to generate an anomaly state confidence parameter. Obtain the time compensation coefficient. Perform a multiplication operation on the anomaly state confidence parameter and the time compensation coefficient to generate the geochemical anomaly identification result.

[0021] Please see Figure 2 The specific steps of S1 are as follows: S101: Obtain a set of multi-element concentration vectors, detect internal dimensional distribution attribute parameters based on basic mapping logic, perform scale alignment to extract a general feature matrix for the internal dimensional distribution attribute parameters, call the underlying sorting operator to perform spatial recombination transformation on the general feature matrix to construct a topological recombination parameter set, aggregate multi-dimensional feature nodes based on the topological recombination parameter set to establish a basic mapping combination, and generate an element combination index sequence. A multi-element concentration vector set was obtained by accessing a regional exploration database and extracting geochemical composition test records within a specified spatial range. The set of elemental concentration values ​​from these records constituted a multi-element concentration vector set, specifically covering five chemical components: copper, lead, zinc, gold, and silver. Sampling test results from 200 spatial data nodes were retrieved. All data were arranged in a two-dimensional data matrix, with rows corresponding to the 200 sampling point numbers and columns corresponding to the five element categories. A null value check was performed within this multi-element concentration vector set to identify locations containing missing values. For each missing value location, four neighboring sampling points were extracted. For the same element concentration values ​​at sample points, the arithmetic mean of the four values ​​is calculated. This average is then used to fill in the missing values, completing data imputation. This data acquisition and cleaning process directly provides a standardized input data foundation for subsequent machine learning applications. Based on the basic mapping logic, the internal dimensional distribution attribute parameters are detected. Specifically, for all 200 concentration values ​​of each element category, the arithmetic mean and standard deviation are calculated. For copper concentration, 200 values ​​are extracted and summed. The sum is divided by 200 to obtain the copper mean. The sum of the squares of the differences between each value and this mean is then calculated, divided by 200, and the square root is taken to obtain the copper standard deviation. The mean and standard deviation together constitute the internal dimension distribution attribute parameter. In the actual example, the copper concentration values ​​extracted from the 1st to the 5th sampling points were 12 mg / kg, 15 mg / kg, 14 mg / kg, 18 mg / kg, and 16 mg / kg, respectively. The cumulative result is calculated as 12 + 15 + 14 + 18 + 16 = 75. The arithmetic mean is calculated as 75 / 5 = 15, with a mean of 15 mg / kg. The sum of squares of deviations of each value is calculated as 9 + 0 + 1 + 9 + 1 = 20. The square of the standard deviation is calculated as 20 / 5 = 4, with a standard deviation of 2 mg / kg. By summing the mean and standard deviation of each element, the internal dimension distribution attribute parameter is detected. The internal dimensional distribution attribute parameter is scale-aligned to extract a general feature matrix. This scale alignment operation is performed on each initial concentration value in the multi-element concentration vector set. The value is subtracted from the mean of the corresponding element and then divided by the standard deviation of the corresponding element. After scale alignment, the concentration values ​​of all elements are distributed in the same numerical space with a mean of 0 and a standard deviation of 1, eliminating the difference in dimensional magnitude between different elements and forming a general feature matrix. When performing scale alignment on the copper element concentration of the first sampling point, the initial value of 12 is extracted, the mean of 15 is subtracted to obtain a difference of -3, and then divided by the standard deviation of 2. The calculation formula is 12-15=-3, -3 / 2=-1.5, and the aligned value is -1.5. Integrate all aligned values ​​to generate a general feature matrix. Call the underlying sorting operator to perform spatial recombination transformation on the general feature matrix to construct a topological recombination parameter set. Extract the full element alignment values ​​corresponding to each data node in the general feature matrix. Calculate the sum of the absolute values ​​of the five element alignment values. Use this sum of absolute values ​​as the sorting criterion and rearrange all data nodes according to descending order to generate the topological recombination parameter set. In the actual example, a data node contains five element alignment values ​​of -1.5, 0.8, 1.2, -0.5, and 1.0. The formula for calculating the sum of absolute values ​​is 1.5 + 0.8 + 1.2 + 0.5 + 1.0 = 5.0. The value 5.0 is used as the baseline for descending sorting of all nodes. Based on the topology reorganization parameter set, multi-dimensional feature nodes are aggregated to establish a basic mapping combination. The top 50 data nodes are extracted as multi-dimensional feature nodes based on their ranking. The aligned element concentration value arrays carried by each node are concatenated end-to-end according to spatial coordinate order. The resulting ultra-long one-dimensional array is the basic mapping combination. An element combination index sequence is generated. For each value within the basic mapping combination, a unique non-negative integer position identifier is assigned, incrementing by 1 sequentially until the end of the combination. All identifiers are arranged in order to form the element combination index sequence. Table 1 Initial Record of Element Concentrations at Sampling Points Sampling point number copper concentration Lead concentration Zinc concentration Gold concentration Silver concentration 1 12 45 80 0.5 2.1 2 15 42 75 0.6 2.0 3 14 48 82 0.4 2.3 4 18 50 85 0.8 2.5 5 16 46 78 0.5 2.2 As shown in Table 1, chemical composition concentration data of some spatial data nodes were extracted and used in the aforementioned calculation process of arithmetic mean and standard deviation. The data unit is milligrams per kilogram.

[0022] S102: Monitor the element combination index sequence, call the global scanning mechanism based on the state matching strategy to extract the sample distribution feature space, perform structural comparison on the sample distribution feature space to obtain the feature overlap indicator parameter, aggregate relevant response nodes to construct discrete similar clusters based on the feature overlap indicator parameter, perform numerical accumulation calculation on the discrete similar clusters to extract the statistical benchmark number, and obtain the frequency of occurrence of the sample set. The system monitors the element combination index sequence and uses a global scanning mechanism based on a state matching strategy to extract the sample distribution feature space. This global scanning mechanism has a fixed-length observation window, set to a width of 10 identifiers. The observation window slides sequentially along the element combination index sequence, starting from identifier 1, with each slide step being one identifier. At each stop point, a segment of value within the window is extracted to form a local data vector. All extracted local data vectors are combined to establish the sample distribution feature space. Structural comparison is performed on this feature space to obtain a feature overlap indicator parameter. Two local data vectors are randomly selected from the feature space, denoted as the base vector and the comparison vector, respectively. Their inner product is calculated, along with the geometric modulus of both the base and comparison vectors. These geometric moduli are then multiplied to obtain a modulus product term. Finally, the inner product is divided by the modulus product term to calculate the cosine similarity. This cosine similarity value is directly used as the feature overlap indicator parameter. In the actual example, the base vector contains numerical values... For vectors 1.0 and 2.0, containing the values ​​1.5 and 1.5 respectively, the inner product is calculated as 1.0*1.5 + 2.0*1.5 = 4.5. The square of the modulus of the base vector is calculated as 1.0*1.0 + 2.0*2.0 = 5.0. The square of the modulus of the comparison vector is calculated as 1.5*1.5 + 1.5*1.5 = 4.5. The square of the product of modulus terms is calculated as 5.0*4.5 = 22.5, which, after taking the square root, is approximately 4.74. The cosine similarity is calculated as 4.5 / 4.74 = 0.95, thus extracting the feature overlap indicator. With a parameter of 0.95, discrete similarity clusters are constructed by aggregating relevant response nodes based on the feature overlap indicator parameter. The similarity evaluation threshold is set to 0.85. For each local data vector in the sample distribution feature space, the cosine similarity value is compared with the similarity evaluation threshold. When the cosine similarity value is greater than or equal to 0.85, it is determined that the comparison vector and the base vector have a high degree of structural overlap. The comparison vector is then assigned to the subordinate set corresponding to the base vector. After completing the comparison and allocation of the global vectors, all vectors with a cosine similarity greater than or equal to 0 are integrated.For a vector group of 85, a discrete similarity cluster is established. For each discrete similarity cluster, a numerical accumulation calculation is performed to extract a statistical benchmark. Each vector group within the discrete similarity cluster is traversed, and the number of local data vectors within the group is counted one by one. The sum of these counts is recorded as the statistical benchmark, yielding the frequency of the sample set. This statistical benchmark is directly assigned to the total number of times the corresponding feature combination appears in the multi-element concentration vector set, thus determining the frequency of the sample set. This frequency provides a basic reference for the weight distribution in subsequent sparse activation. In a practical counting example, a discrete similarity cluster contains the first basic vector and 14 contrast vectors that meet the conditions. The accumulation calculation formula is 1 + 14 = 15, yielding a statistical benchmark of 15. This confirms that the frequency of the sample set with this specific element distribution structure is 15. S103: Based on the frequency of occurrence of the sample set, call the numerical mapping model to detect the local spatial gradient difference representation term associated with the multi-element concentration vector set, perform dimension merging and transformation on the local spatial gradient difference representation term, construct the concentration gradient difference array, extract the multiplication calculation rules, perform product fusion operation on the frequency of occurrence of the sample set and the concentration gradient difference array, and obtain the sparse activation weight vector. Based on the frequency of occurrence in the sample set, a numerical mapping model is invoked to detect the local spatial gradient difference representation term associated with the multi-element concentration vector set. The target sampling point at the center coordinate position is extracted from the multi-element concentration vector set, and the element concentration alignment values ​​of its four adjacent boundary sampling points in the four directions (north, south, east, and west) are read. The values ​​of the four boundary sampling points are subtracted from the value of the target sampling point to obtain four spatial gradient differences. These four spatial gradient differences constitute the local spatial gradient difference representation term. In the actual example, the copper element alignment value of the target sampling point is 1.2, and the alignment value of the adjacent point in the north is 1.8. The calculation formula is 1.8 - 1.2 = 0.6, resulting in a gradient difference of 0.6 in this direction. A dimension merging transformation is performed on the local spatial gradient difference representation term. The absolute values ​​of the above four spatial gradient differences are extracted and summed to construct a concentration gradient difference array. If the absolute values ​​of the differences in the four directions are 0.6, 0.4, 0.5, and 0.3 respectively... The summation calculation formula is 0.6 + 0.4 + 0.5 + 0.3 = 1.8. 1.8 is stored in the corresponding position of the concentration gradient difference array. The multiplication calculation rule is extracted, and a product fusion operation is performed on the frequency of occurrence of the sample set and the concentration gradient difference array. The frequency value of the specific sample set obtained in step S102, 15, is read, and the corresponding value 1.8 is extracted from the concentration gradient difference array. The two are directly multiplied to obtain the sparse activation weight vector. The multiplication fusion calculation formula is 15 * 1.8 = 27.0, resulting in a fusion result of 27.0. A nonlinear compression mapping is performed on this fusion result, and the scaling parameter 0.1 is called to multiply it, resulting in a calculation formula of 27.0 * 0.1 = 2.7. The scaled value 2.7 is filled into the corresponding data node position. This operation is repeated for all nodes, and all the calculated scaled values ​​are integrated to establish a one-dimensional data sequence, completing the acquisition of the sparse activation weight vector. This sparse activation weight vector has dual attributes: spatial variability and global occurrence frequency. Table 2. Spatial Neighborhood Gradient Difference Calculation Table Sampling point orientation Align initial values Target center point value Gradient calculation difference Absolute value record North neighboring point 1.8 1.2 0.6 0.6 South neighboring point 0.8 1.2 -0.4 0.4 East neighboring point 1.7 1.2 0.5 0.5 West neighboring point 0.9 1.2 -0.3 0.3 Table 2 lists the numerical calculation logic and absolute value extraction results of the target center point and its four directional neighboring points in the process of constructing the concentration gradient difference array.

[0023] Please see Figure 3 The specific steps of S2 are as follows: S201: For a set of multi-element concentration vectors, call an autoencoder neural network to perform forward propagation mapping calculation to extract the hidden layer feature states associated with the input parameters. Perform linear mapping on the hidden layer feature states to analyze the node weight distribution pattern. Based on the node weight distribution pattern, aggregate the internal flow nodes of the network to generate hidden node input signals. For a set of multi-element concentration vectors, an autoencoder neural network is invoked to perform forward propagation mapping to extract the hidden layer feature states associated with the input parameters. The autoencoder neural network consists of one input layer, two hidden layers, and one output layer. The input layer has 5 neurons, corresponding to the 5 element categories in the multi-element concentration vector set. The first hidden layer has 16 neurons, connected to the input layer using fully connected logic, and equipped with a linear rectified activation function to remove negative responses. The second hidden layer has 8 neurons, also connected to the first hidden layer using fully connected logic, and equipped with a tangent hyperbolic activation function to map values ​​to the range of -1 to +1. The output layer has 5 neurons for data reconstruction output. The forward propagation mapping calculation starts from the input layer receiving concentration data, multiplies it by connection weights and adds it to the bias term, then propagates it to the first hidden layer, and then from the first hidden layer to the second hidden layer to extract the second hidden layer's features. The output state of the hidden layer is used as the feature state of the hidden layer. A linear mapping is performed on the feature state of the hidden layer to analyze the node weight distribution law. The vector matrix composed of 8 feature values ​​of the output of the second hidden layer is extracted. The linear mapping matrix preset for this layer is called. The mapping matrix has a size of 8 rows and 1 column and contains 8 weight coefficients. The vector matrix composed of feature values ​​and the linear mapping matrix are multiplied to obtain a single-dimensional node influence factor sequence. The difference in the absolute value of the values ​​in the node influence factor sequence is analyzed as the node weight distribution law. According to the node weight distribution law, the internal flow nodes of the network are aggregated. The second hidden layer neuron nodes corresponding to the top 4 values ​​in the absolute value of the node influence factor sequence are extracted. These 4 nodes are extracted as key flow nodes. Their output feature values ​​are sequentially spliced ​​and merged to generate the hidden node input signal. The hidden node input signal contains a highly condensed combination of geochemical composition variation features. S202: Call the sparse activation weight vector, perform dimension alignment to verify the compatibility of the spatial structure based on the hidden node input signal, perform element-wise multiplication on the two sets of sequences according to the compatibility to extract weighted response features, perform normalization processing on the weighted response features to construct a feature activation distribution array, and obtain the node weighted input signal; The sparse activation weight vector is invoked, and the spatial structure compatibility is verified by performing dimension alignment based on the hidden node input signal. The data length parameters of the hidden node input signal and the sparse activation weight vector are extracted. A subtraction operation is performed between the two to check if the difference is 0. If the difference is 0, the spatial dimensions of the two sets of data are completely consistent, satisfying the dimension alignment requirement and establishing spatial structure compatibility. If a difference exists, zero-padding is performed at the end of the shorter sequence until the lengths of the two sequences are completely consistent. Based on compatibility, element-wise multiplication is performed on the two sequences to extract weighted response features. Following the corresponding sequence order, the first value of the hidden node input signal and the first value of the sparse activation weight vector are extracted and multiplied. This process is repeated until all corresponding values ​​are multiplied. In the actual example, the first value of the hidden node input signal is 0.5, and the first value of the sparse activation weight vector is 2.7, so the product formula is 0.5 * 2.7 = 1.3. 5. Record 1.35 as the first value of the weighted response feature. Perform normalization processing on the weighted response feature to construct a feature activation distribution array. Extract the maximum and minimum values ​​in the weighted response feature sequence. For each value in the sequence, subtract the minimum value to obtain the first difference, and subtract the minimum value from the maximum value to obtain the second difference. Divide the first difference by the second difference to calculate the normalized value. In the actual example, the maximum value is 4.5, the minimum value is 0.5, and the currently extracted value is 1.35. The formula for calculating the first difference is 1.35-0.5=0.85, the formula for calculating the second difference is 4.5-0.5=4.0, and the formula for normalization is 0.85 / 4.0=0.2125. Store the result 0.2125 in the new sequence. Integrate all the normalized values ​​to form a feature activation distribution array, obtain the node weighted input signal, and use the feature activation distribution array that has completed the normalization operation as the data support base for deep feature mining. Directly assign and record it as the node weighted input signal. S203: Obtain the preset truncation threshold, compare the parameter values ​​of the node weighted input signal with the truncation threshold to extract amplitude difference identifiers, remove redundant data segments with values ​​lower than the truncation threshold based on the amplitude difference identifiers, retain key response features, perform dimension reorganization and compression calculation on the key response features, and obtain the compressed representation parameter vector. The specific process of setting the truncation threshold is as follows: collect the global weighted amplitude sequence of the sample set in the current data batch output by the hidden layer of the autoencoder neural network, calculate the mean and standard deviation parameters of the data distribution of the global weighted amplitude sequence, and perform summation calculation on the mean and standard deviation parameters of the data distribution to obtain the truncation threshold. The process of setting a preset truncation threshold involves: collecting the globally weighted amplitude sequence output from the hidden layer of the autoencoder neural network within the current data batch; calculating the mean and standard deviation of the global weighted amplitude sequence; extracting 100 amplitude data points from the global weighted amplitude sequence; summing the sums of the values ​​and dividing by 100 to obtain the mean; summing the squares of the differences between the amplitude data points and the mean; dividing by 100 and taking the square root to obtain the standard deviation; and then calculating the sum of the squares of the differences between the amplitude data points and the mean. The standard deviation parameter is summed to obtain the truncation threshold. In the actual example, the calculated data distribution mean is 0.35, and the standard deviation parameter is 0.12. The summation formula is 0.35 + 0.12 = 0.47. 0.47 is set as the truncation threshold. For each node-weighted input signal, the parameter values ​​are compared with the truncation threshold to extract the amplitude difference identifier. Each parameter value in the node-weighted input signal is extracted and compared with 0.47. When the parameter value is greater than or equal to 0.47, the amplitude is recorded. The amplitude difference identifier is 1. When the parameter value is less than 0.47, the amplitude difference identifier is recorded as 0. Based on the amplitude difference identifier, redundant data segments with values ​​lower than the truncation threshold are removed to retain key response features. The node weighted input signal and the corresponding amplitude difference identifier are traversed. When the data position with the identifier 0 is scanned, the data at that position in the node weighted input signal is directly deleted and discarded. When the data position with the identifier 1 is scanned, the data at that position is extracted and saved. All saved data are integrated to construct key response features. Dimensional reorganization and compression calculation is performed on the key response features. The total amount of data contained in the key response features is set as the original dimension number. The target compression dimension number is set to half of the original dimension number and rounded down. Based on the target compression dimension number, the linear principal component mapping rule is called to calculate the eigenvalues ​​and eigenvectors of the key response feature covariance matrix. The eigenvectors associated with the largest number of eigenvalues ​​corresponding to the aforementioned target compression dimension number are extracted to construct a projection matrix. The key response features and the projection matrix are multiplied to obtain the compressed representation parameter vector. Please see Figure 4 The specific steps of S3 are as follows: S301: Obtain the preset global sampling nodes, extract the feature mean matrix for the compressed representation parameter vector, configure the distribution kernel function included in the hybrid distribution model according to the feature mean matrix, calculate the probability density value of the global sampling nodes in the spatial dimension, perform normalization processing on the probability density value to extract the attribution probability array of individual components, aggregate the attribution probability array of all nodes, and establish the component response value sequence. Pre-defined global sampling nodes are obtained. The two-dimensional coordinate parameters corresponding to all 200 sampling points within the exploration area are extracted as global sampling nodes. A feature mean matrix is ​​extracted from the compressed representation parameter vector. All values ​​in the compressed representation parameter vector are summed, and the calculation formula is: summing all values ​​and dividing by the total number of values ​​in the compressed representation parameter vector to obtain a single average value. This average value is multiplied by an identity matrix consisting entirely of 1s, where the number of rows and columns of the identity matrix is ​​equal to the total number of global sampling nodes (200). A scalar-matrix multiplication operation is performed to generate the feature mean matrix. Based on the feature mean matrix, the distribution kernel functions of the hybrid distribution model are configured. The hybrid distribution model adopts a Gaussian mixture structure composed of three Gaussian kernel functions. The numerical distribution in the feature mean matrix is ​​extracted as the expectation parameter benchmark of the Gaussian kernel function, and the covariance matrix benchmark parameter is configured as an identity matrix with all diagonal elements being 1. The initial mixing weights of the three Gaussian kernel functions are set to 0.3, 0.3, and 0.4, respectively. The probability of the global sampling nodes in the spatial dimension is calculated. The probability density values ​​are calculated by substituting the two-dimensional coordinate parameters of each global sampling node into the three pre-configured Gaussian kernel functions. Three local distribution probability values ​​are then calculated, each multiplied by its corresponding initial mixing weights of 0.3, 0.3, and 0.4, and summed. In the actual example, the three local distribution probability values ​​are 0.15, 0.20, and 0.10, respectively. The product and summation formula is: 0.15*0.3 + 0.20*0.3 + 0.10*0.4 = 0.045 + 0.06 + 0.04 = The overall probability density value is 0.145. Normalization is performed on the probability density value to extract the association probability array of individual components. The calculated probability density value of 0.145 is divided by the sum of the probability density values ​​of all sampled nodes in the entire domain. The calculation formula is 0.145 / 2.5=0.058. The association probability value of the node with a specific spatial cluster is obtained as 0.058. The calculation results of all 200 nodes are integrated to aggregate the association probability array of all nodes and establish the component response value sequence. S302: Based on the component response numerical sequence, detect the response probability value of the records within the sequence, extract the swapping rules, perform cyclic comparison and address shifting for each response probability value corresponding to the same node within the component response numerical sequence, reconstruct the internal index mapping table according to the data decreasing characteristics, perform spatial displacement rearrangement on the values ​​according to the internal index mapping table, and generate a descending response numerical sequence. Based on the component response numerical sequence, the response probability values ​​recorded within the sequence are detected. 200 response probability values ​​are extracted from the component response numerical sequence. A swapping rule is then applied to each response probability value corresponding to the same node within the component response numerical sequence, performing a loop comparison and address shifting process. Bubble alignment logic is invoked, starting from the first response probability value in the sequence, sequentially comparing the magnitudes of adjacent response probability values. If the preceding response probability value is less than the following response probability value, their address indices are extracted and their positions are swapped, causing the larger response probability value to move forward in the sequence. After completing the first round of full sequence alignment, the smallest response probability value is shifted to the end of the sequence. This loop comparison and address shifting process is repeated. The process continues until no adjacent values ​​in the sequence need to be swapped. The internal index mapping table is reconstructed based on the data decreasing characteristics. The full address index number after the position swap is completed is extracted. According to the order from the first to the last of the sequence, the original node number is bound to the current ranking number one by one to generate the internal index mapping table. The values ​​are spatially shifted and rearranged according to the internal index mapping table. The various auxiliary feature data corresponding to each original node are also synchronously extracted and repositioned according to the ranking number recorded in the internal index mapping table to generate a descending response value sequence. This descending response value sequence not only completes the strict sorting of the core response probability value from large to small, but also ensures the synchronous corresponding state of the node's accompanying attribute parameters. Table 3: Execution Ranking Changes of Bubble Shift Logic Initial position of the sequence Initial probability value The values ​​after the first round of comparison The values ​​after the second round of comparison Rank 1 0.058 0.082 0.095 Rank 2 0.082 0.095 0.082 Rank 3 0.095 0.058 0.060 Rank 4 0.060 0.060 0.058 As shown in Table 3, the numerical change trajectory of the four local response probability values ​​during the application of cyclic comparison and address shifting processes is presented, showing an arrangement in which the values ​​gradually shift from small to large.

[0024] S303: For the descending response numerical sequence, extract the maximum probability value of the first data item record and the second maximum probability value of the second data item record, call the arithmetic rule to perform a subtraction operation on the maximum probability value and the second maximum probability value to obtain the boundary sensitive parameter, and merge the boundary sensitive parameters corresponding to each node in order according to the spatial distribution order of the nodes to obtain the component response difference sequence. For a descending sequence of response values, the maximum probability value of the first data item and the second maximum probability value of the second data item are extracted. The response probability value at position 1 is directly read from the descending sequence and set as the maximum probability value. The response probability value at position 2 is read and set as the second maximum probability value. In the actual example, the maximum probability value is 0.25 and the second maximum probability value is 0.18. The arithmetic rule is applied to subtract the maximum and second maximum probability values ​​to obtain the boundary sensitive parameter. Subtracting the second maximum probability value 0.18 from the maximum probability value of 0.25 yields the formula 0.25 - 0.18. =0.07, the difference obtained by the operation is defined as the boundary sensitive parameter. This boundary sensitive parameter represents the degree of ambiguity in the classification of the corresponding spatial node. The smaller the difference, the more closely the probabilities of the node belonging to multiple categories are very close, and it is in the transitional edge of the spatial distribution. According to the spatial distribution order of the nodes, the boundary sensitive parameters corresponding to each node are merged in sequence. The above difference calculation process is repeated. Independent boundary sensitive parameters are calculated for each of the 200 sampling points in the whole domain. According to the original spatial latitude and longitude encoding order of the nodes, these 200 boundary sensitive parameters are arranged and combined in one dimension to obtain the component response difference sequence. Please see Figure 5 The specific steps of S4 are as follows: S401: Obtain the preset sensitivity threshold, compare each value in the component response difference sequence with the sensitivity threshold item by item, determine the amplitude attribute of each node, filter the global sampling nodes whose difference amplitude is lower than the sensitivity threshold, perform coordinate parameter aggregation and recombination on the global sampling nodes to establish a basic spatial index, and generate a set of boundary candidate nodes; The process of setting the sensitivity threshold is as follows: extract the historical response difference items corresponding to the sampling nodes of the whole domain, sort the historical response difference items in ascending order to construct the benchmark difference sequence, calculate the slope change rate of adjacent numerical sequence segments in the benchmark difference sequence, extract the data nodes where the slope change rate shows extreme value abrupt change, and set the difference value corresponding to the data node as the sensitivity threshold. The process of setting the preset sensitivity threshold involves extracting historical response difference items corresponding to all sampling nodes across the entire region. This is achieved by retrieving 1000 historical boundary sensitivity parameter values ​​obtained from five independent explorations of the same geological zone from the database. These historical response difference items are then sorted in ascending order to construct a baseline difference sequence. A low-level fast sorting algorithm is then called to sequentially rearrange these 1000 values ​​from minimum to maximum. The slope change rate of adjacent value segments in the baseline difference sequence is calculated. Finally, the baseline difference sequence is divided into groups of 10 data points. Divide a sequence segment into 100 segments. Calculate the slope by the ratio of the range to the number of data points within each segment. Extract the slope change rate as the difference between the slope of the second segment and the slope of the first segment. In the actual example, the slope of the first segment is 0.01, and the slope of the second segment is 0.03. The calculation formula is 0.03 - 0.01 = 0.02, resulting in a slope change rate of 0.02 for this segment. Extract data nodes where the slope change rate shows extreme abrupt changes. Scan all slope change rate values ​​and find the one with the largest value. The slope change rate between the 45th and 46th sequence segments is set to reach a maximum value of 0.15. The maximum difference data within the 45th sequence segment is extracted, which is 0.05. The difference value corresponding to the data node is set as the sensitivity threshold, i.e., the sensitivity threshold is set to 0.05. Each value in the component response difference sequence is compared with the sensitivity threshold item by item to extract each boundary sensitive parameter in the component response difference sequence. The magnitude of each node is determined by comparing it with the value 0.05. When the boundary sensitive parameter is greater than... When the value is equal to 0.05, the amplitude attribute of the node is determined to be an internally stable node. When the boundary sensitivity parameter is less than 0.05, the amplitude attribute of the node is determined to be an edge-blurred node. Global sampling nodes with difference amplitudes lower than the sensitivity threshold are filtered out. All spatial sampling points that are determined to be edge-blurred nodes are extracted. The coordinate parameters of the global sampling nodes are aggregated and reorganized to establish a basic spatial index. The x-coordinate and y-coordinate values ​​of all the above-filtered nodes are extracted and arranged in ascending order of x-coordinate first and y-coordinate second to establish an index list, generating a set of boundary candidate nodes. S402: Based on the set of boundary candidate nodes, detect the local numerical space of the internal node mapping, perform segmentation processing on the local numerical space, construct a multi-dimensional numerical interval, count the frequency of data landing points of the set of boundary candidate nodes within the multi-dimensional numerical interval, call the operation rules to perform normalization transformation on the frequency of data landing points to extract the probability distribution representation term, and obtain the category belonging weight value. Based on the boundary candidate node set, the local numerical space mapped by internal nodes is detected. The horizontal and vertical coordinate parameters of each node in the boundary candidate node set are extracted, and the coverage extreme value range of these coordinate parameters on the two-dimensional plane is determined. The minimum and maximum values ​​of the horizontal and vertical coordinates are obtained. A segmentation process is performed on the local numerical space to construct multi-dimensional numerical intervals. The difference between the maximum and minimum extreme values ​​of the horizontal and vertical coordinates is divided by 10, resulting in 100 regular two-dimensional grid regions as multi-dimensional numerical intervals. The frequency of data points within the multi-dimensional numerical intervals of the boundary candidate node set is counted. The coordinate parameters of all nodes in the boundary candidate node set are traversed to determine which specific two-dimensional grid region their spatial location belongs to. For each grid region, the number of nodes contained within it is accumulated. The sum is recorded as the frequency of data landing points. In the actual example, the 5th grid region contains 12 boundary candidate nodes, so the frequency of data landing points in this region is recorded as 12. The operation rule is called to perform a normalization transformation on the frequency of data landing points to extract the probability distribution representation. The frequency of data landing points in the specific grid region obtained above is divided by the total number of nodes in the boundary candidate node set. In the actual example, the boundary candidate node set contains a total of 60 nodes, and the normalization calculation formula is 12 / 60=0.20. The calculation result 0.20 is obtained as the probability distribution representation of this region. The category assignment weight value is obtained, and the probability distribution representation 0.20 is multiplied by the preset spatial penalty constant 1.5. The product calculation formula is 0.20*1.5=0.30. Finally, the category assignment weight value of the node in the target grid region is 0.30. S403: Based on the category attribution weight value, extract the probability extreme value term of the boundary candidate node set, assign the associated node to the target label according to the probability extreme value term to determine the node's category, perform spatial optimization to extract the edge turning point coordinates for nodes with differentiated categories, perform connectivity calculation on the edge turning point coordinates to construct the spatial region contour, and establish an anomaly category boundary structure. Based on the category assignment weights, extract the probability extrema of the boundary candidate node set. For each boundary candidate node, obtain the probability that it belongs to the abnormal region and the probability that it belongs to the background region, as determined by the mixed distribution model. Multiply these two probability values ​​by the category assignment weight of the grid region where the node is located, and record the larger of the product results as the probability extrema. Assign the associated nodes to the target labels to determine the node's category based on the probability extrema. If the extrema is obtained by multiplying the probability of the node belonging to the abnormal region by its weight, then the node is labeled with an abnormality-specific label. The item is obtained by multiplying the probability of a node belonging to the background region by its weight, and then a background-specific label is applied. For nodes with differentiated classification, spatial optimization is performed to extract the coordinates of edge turning points. In the two-dimensional spatial coordinate system, the Euclidean distance between each abnormal label node and its nearest background label node is calculated, and the midpoint coordinate parameter of the line connecting these distances is extracted and recorded as the edge turning point coordinates. Connectivity calculation is performed on the edge turning point coordinates to construct the spatial region contour. Using the Delaunay triangulation rule, all the extracted midpoint coordinates are used as vertices to connect line segments, forming a closed polygon network and establishing an abnormal category boundary structure. Please see Figure 6 The specific steps of S5 are as follows: S501: Based on the anomaly category boundary structure, delineate the associated target within the global sampling node, construct a neighborhood node spatial set, extract the local reconstruction probability value of each node in the neighborhood node spatial set, perform a square operation on the local reconstruction probability value to extract the energy metric, and integrate each energy metric into a two-dimensional array according to the spatial arrangement to generate a probability energy distribution matrix. Based on the anomaly category boundary structure, associated targets are delineated within the global sampling nodes. All sampling points within the polygonal network of the anomaly category boundary structure, as well as external sampling points less than 20 meters from the polygon boundary line, are extracted. These extracted points are collectively referred to as associated targets, and a neighborhood node space set is constructed. All associated targets are uniformly stored in a new data sequence as the neighborhood node space set. The local reconstruction probability value of each node within the neighborhood node space set is extracted. The output layer parameters of the autoencoder neural network are then called, and the original multi-element concentration input data entering the neighborhood node space set are compared with the model output layer. The reconstructed output data is used to calculate the difference. The absolute values ​​of the differences of the five elements are added together and the reciprocal is taken as the local reconstruction probability value. The local reconstruction probability value is squared to extract the energy metric. The local reconstruction probability value of a certain node is extracted as 0.8. The calculation formula is 0.8*0.8=0.64. The energy metric of the node is obtained as 0.64. Each energy metric is integrated into a two-dimensional array according to the spatial arrangement. The energy metric of all nodes is filled into the corresponding grid matrix points according to the distribution of the node's horizontal and vertical coordinates. The missing positions are filled with the value 0, generating a probability energy distribution matrix. S502: Based on the probability energy distribution matrix, calculate the arithmetic mean of all elements to determine the overall energy mean. Subtract the values ​​of the elements within the matrix from the overall energy mean to extract the absolute difference and construct an energy deviation parameter. Sort the energy deviation parameter in descending order according to the numerical index. Extract the extreme value ranking data and divide it with the overall energy mean to obtain the energy ratio coefficient. Read the preset equilibrium setting threshold and compare the two features to obtain the abnormal state confidence parameter. The process of comparing two features by reading the preset equilibrium threshold is as follows: First, extract the energy deviation parameter of each of the top ten ranked items to establish extreme value ranking data. Then, use the extreme value ranking data as the dividend and the overall energy mean as the divisor to perform a division calculation, extracting the quotient to form the energy proportion coefficient. Next, calculate the standard deviation index of all data elements in the probability energy distribution matrix. Then, retrieve the tolerance multiplier set based on the numerical distribution law, multiply the standard deviation index by the tolerance multiplier to output the fluctuation assessment parameter. Finally, perform an addition operation between the fluctuation assessment parameter and the overall energy mean to set the equilibrium threshold. Then, compare each value within the energy proportion coefficient with the equilibrium threshold. Output a high-order identifier when each value is greater than the equilibrium threshold, and output a low-order identifier when each value is not greater than the equilibrium threshold. Finally, combine the high-order and low-order identifiers in order to form the abnormal state confidence parameter. Based on the probability energy distribution matrix, the arithmetic mean of all elements is calculated to determine the overall energy mean. The values ​​of all energy metrics within the matrix are summed and then divided by the total number of non-zero elements. In the actual example, the sum is 120.0, and there are 40 non-zero elements. The mean is calculated as 120.0 / 40 = 3.0, yielding an overall energy mean of 3.0. The absolute difference between the values ​​of the elements within the matrix and the overall energy mean is then extracted to construct the energy deviation parameter. The energy metric value of a specific node in the matrix is ​​extracted as 5.5. The subtraction calculation formula is 5.5 - 3.0 = 2.5. Extracting the absolute value yields an energy deviation parameter of 2.5. Based on the numerical indicators, the energy deviation parameters are sorted in descending order. All energy deviation parameters are rearranged from largest to smallest. The process of comparing two features using a preset equilibrium threshold involves extracting the top ten energy deviation parameters to establish extreme value ranking data. This extreme value ranking data is used as the dividend, and the overall energy mean is used as the divisor to perform a division calculation. The quotient is extracted to form the energy proportion coefficient. In the actual calculation example... The energy deviation parameter ranking first is 4.5. The division calculation is 4.5 / 3.0 = 1.5, yielding an energy proportion coefficient of 1.5. The standard deviation of all data elements within the probability energy distribution matrix is ​​calculated, yielding a standard deviation of 1.2 for all non-zero elements. A tolerance multiplier based on the numerical distribution pattern is retrieved and set to 1.5. The standard deviation is multiplied by the tolerance multiplier to output the fluctuation assessment parameter, calculated as 1.2 * 1.5 = 1.8. This fluctuation assessment parameter is then multiplied by the overall energy distribution matrix. The average value is added to set the equilibrium threshold. The calculation formula is 1.8 + 3.0 = 4.8. The equilibrium threshold is set to 4.8. Each value in the energy ratio coefficient is compared with the equilibrium threshold. When each value is greater than the equilibrium threshold, a high-order identifier is output. When each value is not greater than the equilibrium threshold, a low-order identifier is output. Since 1.5 is less than 4.8, a low-order identifier is output. The high-order identifier and the low-order identifier are combined in order to form the abnormal state confidence parameter. The comparison results of all nodes are combined to obtain the abnormal state confidence parameter. S503: Based on the anomaly confidence parameter, collect timestamp data of the sampling test process, call the attenuation model to process the timestamp data to extract time drift correlation terms, establish time compensation coefficients, perform item-by-item multiplication calculations on the internal recorded values ​​of the anomaly confidence parameter and the time compensation coefficients to perform numerical fusion and recombination, map the anomaly level, and generate geochemical anomaly identification results. Based on the anomaly confidence parameter, timestamp data of the sampling and testing process is collected. The collection time (year, month, day) is read from each original geochemical data record. The attenuation model is used to process the timestamp data, extracting time drift correlation terms and establishing a time compensation coefficient. The number of days between the current analysis date and the sampling date is set. If the difference is 365 days, the attenuation model formula constant 0.001 is multiplied by the number of days, resulting in the calculation 365 * 0.001 = 0.365. Subtracting 0.365 from the value of 1 yields the final result. The calculation formula is 1-0.365=0.635, and the time compensation coefficient is 0.635. For the internal recorded values ​​of the anomaly confidence parameter, the numerical fusion and recombination are performed by multiplying the values ​​of each item with the time compensation coefficient. If a node is determined to be a high-level identifier with a base score of 2.0, the product calculation formula is 2.0*0.635=1.27, and the anomaly level is mapped. When the fusion score is greater than 1.0, it is mapped to a strong anomaly level; otherwise, it is mapped to a weak anomaly level, generating the geochemical anomaly identification result.

[0025] Please see Figure 7 A machine learning-based intelligent identification system for geochemical anomalies includes: The feature arrangement module obtains a set of multi-element concentration vectors, sorts them according to the data size of each dimension to construct an element combination index sequence, detects the frequency of occurrence of the sample set associated with the element combination index sequence, and calculates the sparse activation weight vector by multiplying it with the concentration gradient difference array associated with the set of multi-element concentration vectors. The compression mapping module inputs a set of multi-element concentration vectors into an autoencoder neural network to extract hidden node input signals. It then multiplies these signals with sparse activation weight vectors to generate node weighted input signals. Finally, it removes low-value data segments by combining a preset truncation threshold and outputs a compressed representation parameter vector. The response parsing module calls the Gaussian mixture model to receive the compressed representation parameter vector, calculates the component response value sequence associated with the global sampling nodes, extracts the first two component values ​​after sorting them in descending order, subtracts them, and obtains the component response difference sequence. The boundary discrimination module, based on the component response difference sequence, filters out global sampling nodes whose values ​​are lower than the preset sensitivity threshold, generates a set of boundary candidate nodes, extracts the category assignment weight values ​​and assigns them to categories, and generates an abnormal category boundary structure. The anomaly assessment module extracts the spatial set of neighboring nodes based on the anomaly category boundary structure, generates a probability energy distribution matrix by squaring the local reconstruction probability values, assesses the confidence parameter of the anomaly state, and maps the anomaly level in combination with the time compensation coefficient to generate geochemical anomaly identification results.

[0026] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection of the described technical solutions.

Claims

1. A machine learning-based intelligent identification method for geochemical anomalies, characterized in that, Includes the following steps: S1: Obtain a set of multi-element concentration vectors, sort them according to the data size of each dimension to construct an element combination index sequence, detect the frequency of occurrence of the sample set associated with the element combination index sequence, and multiply it with the concentration gradient difference array associated with the multi-element concentration vector set to generate a sparse activation weight vector. S2: Input the set of multi-element concentration vectors into an autoencoder neural network to extract hidden node input signals, multiply them with sparse activation weight vectors to generate node weighted input signals, combine them with a preset truncation threshold to remove low-value data segments, and output a compressed representation parameter vector. S3: Call the Gaussian mixture model to receive the compressed representation parameter vector, calculate the component response value sequence associated with the global sampling nodes, sort them in descending order, extract the first two component values ​​and subtract them to obtain the component response difference sequence. S4: Based on the component response difference sequence, filter out global sampling nodes whose values ​​are lower than the preset sensitivity threshold, generate a set of boundary candidate nodes, extract the category belonging weight values ​​and assign them to the category, and generate an abnormal category boundary structure. S5: Extract the spatial set of neighboring nodes based on the anomaly category boundary structure, generate a probability energy distribution matrix by squaring the local reconstruction probability values, evaluate the confidence parameter of the anomaly state, and combine it with the time compensation coefficient to map the anomaly level and generate geochemical anomaly identification results.

2. The intelligent identification method for geochemical anomalies based on machine learning according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Obtain a set of multi-element concentration vectors, detect internal dimensional distribution attribute parameters based on basic mapping logic, perform scale alignment to extract a general feature matrix for the internal dimensional distribution attribute parameters, call the underlying sorting operator to perform spatial recombination transformation on the general feature matrix to construct a topological recombination parameter set, aggregate multi-dimensional feature nodes based on the topological recombination parameter set to establish a basic mapping combination, and generate an element combination index sequence. S102: Monitor the element combination index sequence, call the global scanning mechanism based on the state matching strategy to extract the sample distribution feature space, perform structural comparison on the sample distribution feature space to obtain the feature overlap indicator parameter, aggregate relevant response nodes to construct discrete similar clusters based on the feature overlap indicator parameter, perform numerical accumulation calculation on the discrete similar clusters to extract statistical benchmark numbers, and obtain the frequency of occurrence of the sample set. S103: Based on the frequency of occurrence of the sample set, call the numerical mapping model to detect the local spatial gradient difference representation term associated with the multi-element concentration vector set, perform dimension merging and transformation on the local spatial gradient difference representation term, construct the concentration gradient difference array, extract the multiplication calculation rules, perform product fusion operation on the frequency of occurrence of the sample set and the concentration gradient difference array, and obtain the sparse activation weight vector.

3. The intelligent identification method for geochemical anomalies based on machine learning according to claim 1, characterized in that, The specific steps of S2 are as follows: S201: For a set of multi-element concentration vectors, call an autoencoder neural network to perform forward propagation mapping calculation to extract the hidden layer feature states associated with the input parameters. Perform linear mapping on the hidden layer feature states to analyze the node weight distribution pattern. Based on the node weight distribution pattern, aggregate the internal flow nodes of the network to generate hidden node input signals. S202: Call the sparse activation weight vector, perform dimension alignment to verify the compatibility of the spatial structure based on the hidden node input signal, perform element-wise multiplication on the two sets of sequences according to the compatibility to extract weighted response features, perform normalization processing on the weighted response features to construct a feature activation distribution array, and obtain the node weighted input signal; S203: Obtain a preset truncation threshold, compare the parameter values ​​of the weighted input signal of the node with the truncation threshold to extract amplitude difference identifiers, remove redundant data segments with values ​​lower than the truncation threshold based on the amplitude difference identifiers, retain key response features, perform dimension reorganization and compression calculation on the key response features, and obtain a compressed representation parameter vector.

4. The intelligent identification method for geochemical anomalies based on machine learning according to claim 3, characterized in that, The specific process of setting the truncation threshold is as follows: collect the globally weighted amplitude sequence of the sample set in the current data batch output by the hidden layer of the autoencoder neural network, calculate the mean and standard deviation parameters of the data distribution of the globally weighted amplitude sequence, and perform summation calculation on the mean and standard deviation parameters of the data distribution to obtain the truncation threshold.

5. The intelligent identification method for geochemical anomalies based on machine learning according to claim 1, characterized in that, The specific steps for S3 are as follows: S301: Obtain the preset global sampling nodes, extract the feature mean matrix for the compressed representation parameter vector, configure the distribution kernel function included in the hybrid distribution model according to the feature mean matrix, calculate the probability density value of the global sampling nodes in the spatial dimension, perform normalization processing on the probability density value to extract the attribution probability array of individual components, aggregate the attribution probability array of all nodes, and establish the component response value sequence. S302: Based on the component response numerical sequence, detect the response probability values ​​recorded within the sequence, extract the swapping rules, perform cyclic comparison and address shifting processing on each response probability value corresponding to the same node within the component response numerical sequence, reconstruct the internal index mapping table according to the data decreasing characteristics, and perform spatial displacement rearrangement on the values ​​according to the internal index mapping table to generate a descending response numerical sequence. S303: For the descending response numerical sequence, extract the maximum probability value of the first data item and the second maximum probability value of the second data item, call the arithmetic rule to perform a subtraction operation on the maximum probability value and the second maximum probability value to obtain the boundary sensitive parameter, and merge the boundary sensitive parameters corresponding to each node in order according to the spatial distribution order of the nodes to obtain the component response difference sequence.

6. The intelligent identification method for geochemical anomalies based on machine learning according to claim 1, characterized in that, The specific steps of S4 are as follows: S401: Obtain a preset sensitivity threshold, perform item-by-item difference comparison between each value in the component response difference sequence and the sensitivity threshold, determine the amplitude attribute of each node, filter out global sampling nodes with difference amplitudes lower than the sensitivity threshold, perform coordinate parameter aggregation and recombination on global sampling nodes to establish a basic spatial index, and generate a set of boundary candidate nodes; S402: Based on the set of boundary candidate nodes, detect the local numerical space of the internal node mapping, perform segmentation processing on the local numerical space, construct a multi-dimensional numerical interval, count the frequency of data landing points of the set of boundary candidate nodes within the multi-dimensional numerical interval, call the operation rules to perform normalization transformation on the frequency of data landing points to extract the probability distribution representation term, and obtain the category belonging weight value. S403: Based on the category attribution weight value, extract the probability extreme value term of the boundary candidate node set, assign the associated node to the target label according to the probability extreme value term to determine the node's category, perform spatial optimization to extract the edge turning point coordinates for nodes with differentiated categories, perform connectivity calculation on the edge turning point coordinates to construct the spatial region contour, and establish an abnormal category boundary structure.

7. The intelligent identification method for geochemical anomalies based on machine learning according to claim 6, characterized in that, The process of setting the sensitivity threshold is as follows: extract the historical response difference items corresponding to the global sampling nodes, sort the historical response difference items in ascending order to construct a benchmark difference sequence, calculate the slope change rate of adjacent numerical sequence segments in the benchmark difference sequence, extract the data nodes where the slope change rate shows extreme value abrupt change, and set the difference value corresponding to the data node as the sensitivity threshold.

8. The intelligent identification method for geochemical anomalies based on machine learning according to claim 1, characterized in that, The specific steps of S5 are as follows: S501: Based on the anomaly category boundary structure, delineate the associated target within the global sampling node, construct a neighborhood node space set, extract the local reconstruction probability value of each node in the neighborhood node space set, perform a square operation on the local reconstruction probability value to extract the energy metric, integrate each energy metric into a two-dimensional array according to the spatial arrangement, and generate a probability energy distribution matrix. S502: Based on the probability energy distribution matrix, calculate the arithmetic mean of all elements to determine the overall energy mean. Subtract the values ​​of the elements within the matrix from the overall energy mean to extract the absolute difference and construct an energy deviation parameter. Sort the energy deviation parameter in descending order according to the numerical index. Extract the extreme value ranking data and divide it with the overall energy mean to obtain the energy ratio coefficient. Read the preset equilibrium setting threshold and compare the two features to obtain the abnormal state confidence parameter. S503: Based on the anomaly confidence parameter, collect timestamp data of the sampling test process, call the attenuation model to process the timestamp data to extract time drift correlation terms, establish time compensation coefficients, perform item-by-item product calculations on the internal recorded values ​​of the anomaly confidence parameter and the time compensation coefficients to perform numerical fusion and recombination, map the anomaly level, and generate geochemical anomaly identification results.

9. The intelligent identification method for geochemical anomalies based on machine learning according to claim 8, characterized in that, The process of comparing two features by reading the preset equalization threshold is as follows: extract the energy deviation parameter of each of the top ten items to establish extreme value ranking data; use the extreme value ranking data as the dividend and the overall energy mean as the divisor to perform a division calculation, and extract the division quotient to form the energy ratio coefficient. Calculate the standard deviation of all data elements in the probability energy distribution matrix; retrieve the tolerance multiplier set based on the numerical distribution law, multiply the standard deviation by the tolerance multiplier to output the fluctuation assessment parameter; perform an addition operation on the fluctuation assessment parameter and the overall energy mean to set the equilibrium threshold. Compare each value within the energy ratio coefficient with the equilibrium setting threshold; When each value is greater than the set threshold for balance, a high-order identifier is output; when each value is not greater than the set threshold for balance, a low-order identifier is output. The high-order identifier and the low-order identifier are combined in order to form the confidence parameter of the abnormal state.

10. A machine learning-based intelligent identification system for geochemical anomalies, characterized in that, The system is used to implement the intelligent identification method for geochemical anomalies based on machine learning as described in any one of claims 1-9, and the system comprises: The feature arrangement module obtains a set of multi-element concentration vectors, sorts them according to the data size of each dimension to construct an element combination index sequence, detects the frequency of occurrence of the sample set associated with the element combination index sequence, and calculates the sparse activation weight vector by multiplying it with the concentration gradient difference array associated with the set of multi-element concentration vectors. The compression mapping module inputs the multi-element concentration vector set into an autoencoder neural network to extract the hidden node input signal, multiplies it with the sparse activation weight vector to generate the node weighted input signal, combines it with a preset truncation threshold to remove low-value data segments, and outputs a compressed representation parameter vector. The response parsing module calls the Gaussian mixture model to receive the compressed representation parameter vector, calculates the component response value sequence associated with the global sampling nodes, extracts the first two component values ​​after sorting them in descending order, subtracts them, and obtains the component response difference sequence. The boundary discrimination module, based on the component response difference sequence, filters out global sampling nodes whose values ​​are lower than a preset sensitivity threshold, generates a set of boundary candidate nodes, extracts the category assignment weight values ​​and assigns them to categories, and generates an abnormal category boundary structure. The anomaly assessment module extracts a spatial set of neighboring nodes based on the anomaly category boundary structure, generates a probability energy distribution matrix by squaring the local reconstruction probability values, assesses the confidence parameter of the anomaly state, and maps the anomaly level in conjunction with the time compensation coefficient to generate geochemical anomaly identification results.