Assay result data analysis method of water-based sediment sample
The three-dimensional clustering method addresses spatial fragmentation in K-means by weighting distances based on element similarity and spatial distribution, enhancing mineral detection accuracy in water system sediment samples.
Patent Information
- Application Number
- CN202510795810.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
When the prior art uses the K-means clustering algorithm to analyze water-based sediment samples, the geospatial continuity and proximity of the samples are ignored, resulting in a decrease in the accuracy of clustering results, and the abnormal concentration of individual samples and heavy metal pollution affect the classification results.
A three-dimensional clustering space is constructed, combining sample location and element concentration, and optimizing clustering distance through the first weighted correction and the second weighted correction is used to reduce spherical clustering, reduce the impact of heavy metal pollution, and improve clustering accuracy.
The clustering accuracy of ore body element distribution is improved, the probability of spherical clustering during high-dimensional clustering is reduced, and the accuracy of searching ore body distribution and analysis of test results is enhanced.
Smart Images

Figure CN120316539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for analyzing the test results of water system sediment samples. Background Art
[0002] The analysis of the test results of water system sediment samples is usually a key step in the fields of geochemistry, hydrogeology or environmental geology. By analyzing and processing the test data of water system sediment samples, characteristics such as the content, enrichment and dispersion, combination, and spatial distribution of the data can be understood. The purpose is to identify the element distribution from the collected sediment samples through data analysis and processing, which is beneficial to discovering favorable geological bodies, mineralized bodies, and ore bodies for mineralization and guiding geological prospecting. The collection of water system sediment samples is usually evenly collected on a 1:50,000 topographic map. The prior art usually uses the K-means clustering algorithm to cluster the concentration data of each element in the test results of water system sediment samples. Through clustering, the similarity structure between samples at different sampling points can be grasped, the geochemical type distribution of sediments can be revealed, possible abnormally enriched areas can be separated, and mineralized area detection can be assisted.
[0003] However, in the traditional method, only K-means clustering is performed based on the concentration data of each element, ignoring the continuity or proximity of samples in geographical space, which may cause spatially fragmented clustering and lead to a decrease in the accuracy of the clustering results in characterizing the ore distribution. At the same time, there are often individual samples with extra enrichment in the data of water system sediment samples, which will cause relatively high concentrations of individual elements or abnormal concentrations. For example, local heavy metal pollution causes a high concentration of a certain element, but the concentration of this element is not the real ore body element data. The K-means clustering will be biased by this abnormal extreme value, affecting the classification results.
[0004] Therefore, how to improve the accuracy of analyzing the test results of water system sediment samples using the K-means clustering algorithm has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a method for analyzing the test results of water system sediment samples to solve the problem of how to improve the accuracy of analyzing the test results of water system sediment samples using the K-means clustering algorithm.
[0006] An embodiment of the present invention provides a method for analyzing the test results of water system sediment samples, and the method includes the following steps: Obtain at least two water system sediment samples. For any element, extract the element concentration of the any element in each water system sediment sample, and combine the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample to construct a three-dimensional clustering space; When performing K-means clustering on data points in a three-dimensional clustering space, for any round of iteration, obtain the centroid of the any round of iteration, and perform a first weighted correction on the clustering distance between the any data point and the any centroid according to the element concentration similarity between the any data point and the any centroid to obtain a corrected distance; According to the element concentration difference between the any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similarity to the any data point, perform a second weighted correction on the corrected distance to obtain the final corrected distance between the any data point and the any centroid; Obtain the final corrected distance between each data point and each centroid in each round of iteration, and perform K-means clustering on all data points in the three-dimensional clustering space according to all the final corrected distances to obtain the clustering result of the any element. Analyze the test results of all water system sediment samples according to the clustering results of all elements.
[0007] Preferably, constructing the three-dimensional clustering space by combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample includes: Map the element concentration of the any test element in each water system sediment sample to the same three-dimensional clustering space according to the coordinates of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample. The x-axis of the three-dimensional clustering space represents the abscissa of the corresponding sampling point of each water system sediment sample in the two-dimensional topographic map, the y-axis represents the ordinate of the corresponding sampling point of each water system sediment sample in the two-dimensional topographic map, and the z-axis represents the element concentration of the any test element in each water system sediment sample.
[0008] Preferably, the performing a first weighted correction on the clustering distance between the any data point and the any centroid according to the element concentration similarity between the any data point and the any centroid to obtain a corrected distance includes: Calculate the reciprocal of the absolute value of the difference between the element concentration of the any data point and the element concentration of the any centroid to obtain a first element concentration similarity index; In the three-dimensional clustering space, calculate the standard deviation of the element concentrations of all data points in the neighborhood of the any data point, denoted as the first standard deviation, calculate the standard deviation of the element concentrations of all data points in the neighborhood of the any centroid, denoted as the second standard deviation, and calculate the reciprocal of the absolute value of the difference between the first standard deviation and the second standard deviation to obtain a second element concentration similarity index; Calculate the product between the first element concentration similarity index and the second element concentration similarity index to obtain the element concentration similarity degree between any data point and any centroid; Obtain the corrected distance between any data point and any centroid according to the element concentration similarity degree.
[0009] Preferably, the obtaining of the corrected distance between any data point and any centroid according to the element concentration similarity degree includes: Calculate the difference between the constant 1 and the element concentration similarity degree to obtain the position weight, and calculate the sum of the constant 1 and the element concentration similarity degree to obtain the concentration weight; In the three-dimensional clustering space, calculate the square of the distance between any data and any centroid on the x-axis to obtain the x-axis square term, calculate the square of the distance between any data and any centroid on the y-axis to obtain the y-axis square term, and calculate the square of the distance between any data and any centroid on the z-axis to obtain the z-axis square term; Use the position weight as the weight of the x-axis square term and the y-axis square term respectively, use the concentration weight as the weight of the z-axis square term, perform a weighted sum on the x-axis square term, the y-axis square term, and the z-axis square term to obtain the weighted sum result, and take the arithmetic square root of the weighted sum result as the corrected distance between any data point and any centroid.
[0010] Preferably, the second weighted correction of the corrected distance according to the element concentration difference between any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similar to that of any data point to obtain the final corrected distance between any data point and any centroid includes: Obtain the contamination degree of any data point according to the element concentration difference between any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similar to that of any data point; If the contamination degree is greater than or equal to the preset contamination threshold, perform a second weighted correction on the corrected distance between any data and any centroid according to the contamination degree of any data point to obtain the final corrected distance between any data point and any centroid; If the contamination degree is less than the preset contamination threshold, use the corrected distance between any data and any centroid as the final corrected distance.
[0011] Preferably, obtaining the pollution degree of any one of the data points according to the element concentration difference between any one of the data points and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentrations similar to that of any one of the data points includes: Obtaining similar data points with element concentrations similar to that of any one of the data points in a three-dimensional clustering space. If the number of the similar data points is 0, then set the pollution degree of any one of the data points to 1; If the number of the similar data points is not 0, then calculate the Euclidean distance between each of the similar data points and any one of the data points respectively, linearly normalize the average value of all the Euclidean distances, and obtain a first distribution characteristic value of any one of the data points; In the three-dimensional clustering space, denote all the data points within the neighborhood of any one of the data points as neighborhood data points, calculate the absolute value of the difference between the element concentration of each of the neighborhood data points and the element concentration of any one of the data points respectively, linearly normalize the average value of all the absolute values of the differences, and obtain a second distribution characteristic value of any one of the data points; Calculate the product between the first distribution characteristic value and the second distribution characteristic value of any one of the data points to obtain the pollution degree of any one of the data points.
[0012] Preferably, secondarily weighting and correcting the corrected distance between any one of the data and any one of the centroids according to the pollution degree of any one of the data points to obtain the final corrected distance between any one of the data points and any one of the centroids includes: Subtract the pollution degree of any one of the data points from the constant 1 to obtain a correction weight, and weight the corrected distance between any one of the data and any one of the centroids according to the correction weight to obtain the final corrected distance between any one of the data points and any one of the centroids.
[0013] Preferably, obtaining similar data points with element concentrations similar to that of any one of the data points in the three-dimensional clustering space includes: In the three-dimensional clustering space, for any other data point except any one of the data points, if the absolute value of the difference between the element concentration of any one of the data points and the element concentration of any other data point is less than or equal to a preset difference threshold, then denote any other data point as a similar data point with element concentrations similar to that of any one of the data points.
[0014] The beneficial effects of the embodiments of the present invention compared with the prior art are: The present invention obtains at least two water system sediment samples. For any element, the element concentration of the any element is extracted from each water system sediment sample. Combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample, a three-dimensional clustering space is constructed; when performing K-means clustering on the data points in the three-dimensional clustering space, for any round of iteration process, the centroid of the any round of iteration process is obtained. According to the element concentration similarity between any data point and any centroid, the clustering distance between the any data point and the any centroid is first weighted and corrected to obtain a corrected distance; according to the element concentration difference between the any data point and its neighborhood data points and the spatial distribution characteristics of the data points with element concentration similarity to the any data point, the corrected distance is second weighted and corrected to obtain the final corrected distance between the any data point and the any centroid; the final corrected distance between each data point and each centroid in each round of iteration process is obtained. According to all the final corrected distances, K-means clustering is performed on all the data points in the three-dimensional clustering space to obtain the clustering result of the any element. According to the clustering results of all elements, the test results of all water system sediment samples are analyzed. Among them, combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of any element in each water system sediment sample, a three-dimensional clustering space is constructed, and the K-means clustering algorithm is used to perform three-dimensional clustering on the data points in the three-dimensional clustering space, so that the clustering result synthesizes the characteristics of element concentration and element geographical spatial distribution, improves the clustering accuracy of the ore body element distribution, improves the accuracy of searching for the ore body distribution through the clustering result, and further improves the accuracy of analyzing the test results of water system sediment samples; when performing K-means clustering on the data points in the three-dimensional clustering space, according to the element concentration similarity and spatial distribution characteristics, the clustering distance between each data and the centroid is corrected, reducing the probability of spherical clustering when the K-means clustering algorithm performs high-dimensional clustering, and at the same time reducing the interference of the data points with abnormal element concentrations caused by heavy metal pollution on the clustering result of the K-means clustering, thereby improving the accuracy of the K-means clustering algorithm when performing high-dimensional clustering and improving the accuracy of analyzing the test results of water system sediment samples using the K-means clustering algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a flowchart of a method for analyzing the test results of water system sediment samples provided in the first embodiment of the present invention. Detailed implementation manners
[0017] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as a limitation to the present disclosure.
[0018] It should be noted that the terms "first", "second", etc. in the description of the present disclosure and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0019] In order to illustrate the technical solution of the present invention, it will be described below through specific embodiments.
[0020] Refer to Figure 1 , which is a flowchart of a method for analyzing the test results of water system sediment samples provided in the first embodiment of the present invention. As Figure 1 shown, the method may include: Step S101, obtain at least two water system sediment samples. For any element, extract the element concentration of the any element in each water system sediment sample, and combine the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample to construct a three-dimensional clustering space.
[0021] On a two-dimensional topographic map with a scale of 1:50000 according to Sampling points are evenly arranged in the basic sampling units to obtain at least two sampling points, and a stream sediment sample is collected at each sampling point. The sampling site is selected in the section where the water flow slows down and where various particle sizes of stream sediments are easy to converge, such as where the water flow changes from rapid to slow, where the water flow stagnates, the inner side of the river bend, behind large stones, etc., so that the proportion of each particle size in the stream sediment sample obtained by sampling at each sampling point is in a natural mixing state. Each stream sediment sample is subjected to steps such as drying, sieving, mixing, and packing, and then sent to the laboratory for chemical analysis. In the laboratory, methods such as four-acid digestion, aqua regia digestion, and nitric acid-hydrofluoric acid method are used to extract the concentration contents of different ore body elements such as Au, Ag, Mo, Sn, Co, Cr, Cu, Mn, Ni, Pb, Ti, V, W, Zn, As, Sb, Bi, and Hg in each stream sediment sample. Among them, the extraction of the concentration contents of each ore body element in each stream sediment sample is prior art and will not be elaborated here.
[0022] Under the traditional method, the K-means clustering algorithm is used to cluster the concentration contents of ore body elements in stream sediment samples, so as to grasp the similarity structure among stream sediment samples at different sampling points, reveal the geochemical type distribution of stream sediments, separate possible abnormally enriched areas, and assist in the detection of mineralized areas.
[0023] However, under the traditional method, only the concentration contents of each ore body element are used for K-means clustering, ignoring the continuity or proximity of the samples in the geographical space, which may cause spatially fragmented clustering and lead to a decrease in the accuracy of representing the ore body distribution. Therefore, in the embodiments of the present invention, in combination with the geographical space distribution of each stream sediment sample, the K-means clustering is performed on the concentration contents of each ore body element in each stream sediment sample, so that the clustering result synthesizes the concentration contents of each ore body element and the geographical space distribution characteristics of each stream sediment sample, improving the accuracy of representing the ore body distribution.
[0024] In an embodiment of the present invention, taking the element Au as an example, first, obtain the concentration content of the element Au in each stream sediment sample; then, perform linear normalization on all the concentration contents corresponding to the element Au to obtain the element concentration of the element Au in each stream sediment sample. At the same time, perform linear normalization on the longitude of each stream sediment sample in the two-dimensional topographic map to obtain the abscissa of each stream sediment sample, and perform linear normalization on the latitude of each stream sediment sample in the two-dimensional topographic map to obtain the ordinate of each stream sediment sample. Here, linear normalization is a prior art and will not be elaborated here; finally, map the element concentration of the element Au in each stream sediment sample into the same three-dimensional space, denoted as the three-dimensional clustering space. Among them, the x-axis of the three-dimensional clustering space represents the abscissa of the sampling point corresponding to each stream sediment sample, denoted as the abscissa axis, the y-axis represents the ordinate of the sampling point corresponding to each stream sediment sample, denoted as the ordinate axis, and the z-axis represents the element concentration of the element Au in each stream sediment sample, denoted as the concentration axis.
[0025] So far, by combining the position of each stream sediment sample in the two-dimensional topographic map with the element concentration of the element Au in each stream sediment sample, a three-dimensional clustering space is obtained, which is used to perform K-means clustering on the element concentration of the element Au in each stream sediment sample in the three-dimensional space.
[0026] Step S102, when performing K-means clustering on the data points in the three-dimensional clustering space, for any round of iterative process, obtain the centroid of the any round of iterative process, and perform a first weighted correction on the clustering distance between the any data point and the any centroid according to the element concentration similarity between the any data point and the any centroid to obtain the corrected distance.
[0027] Since the K-means clustering algorithm depends on the distance between the data points and the centroid, when the K-means clustering algorithm performs clustering on high-dimensional data, it is more likely to generate spherical clusters. However, in fact, the element concentration of the ore body elements in stream sediments is affected by geological origin or environmental factors (such as water flow velocity, mineralization zone influence), and the formed clusters are often non-spherical. Therefore, when the K-means clustering algorithm performs clustering on high-dimensional data, the clustering accuracy will decrease.
[0028] Therefore, in the embodiments of the present invention, when performing K-means clustering on data points in a three-dimensional clustering space, taking the i-th iteration process as an example, the centroid of the i-th iteration process is obtained. According to the similarity of element concentrations between the data points and the centroid, different weights are assigned to the distances between the data points and the centroid on each axis in the three-dimensional clustering space: the more similar the element concentrations between the data points and the centroid are, it indicates that the sampling points corresponding to the data points and the centroid are more likely to belong to the same ore body, and the data points should be more attributed to the cluster where the centroid is located. A larger weight should be assigned to the distance between the data points and the centroid on the concentration axis. On the contrary, a smaller weight should be assigned to the distance between the data points and the centroid on the concentration axis. On the other hand, the more similar the element concentration distribution in the neighborhood of the data points is to the element concentration distribution in the neighborhood of the centroid, it indicates that the sampling points corresponding to the data points and the centroid are more likely to belong to the same ore body. A larger weight should be assigned to the distance between the data points and the centroid on the concentration axis. On the contrary, a smaller weight should be assigned to the distance between the data points and the centroid on the concentration axis, so that data points with relatively far distances on the two-dimensional topographic map can also be more easily clustered into one cluster according to the element concentration, reducing the probability of the appearance of spherical clusters. Furthermore, according to the weights of the distances between the data points and the centroid on each axis in the three-dimensional clustering space, the corrected distance between the data points and the centroid is obtained to improve the clustering accuracy of the K-means clustering algorithm when clustering high-dimensional data.
[0029] Taking the q-th data point and the c-th centroid in the three-dimensional clustering space as an example, the steps to obtain the corrected distance between the q-th data point and the c-th centroid are as follows: (1) Obtain the similarity degree of element concentrations between the q-th data point and the c-th centroid.
[0030] Specifically, calculate the reciprocal of the absolute value of the difference between the element concentration of the q-th data point and the element concentration of the c-th centroid to obtain the first element concentration similarity index; In the three-dimensional clustering space, calculate the standard deviation of the element concentrations of all data points within the twenty-six neighborhood of the q-th data point, denoted as the first standard deviation, and calculate the standard deviation of the element concentrations of all data points within the twenty-six neighborhood of the c-th centroid, denoted as the second standard deviation. There is no limit here, and the implementer can set the neighborhood range according to the specific scenario. Calculate the reciprocal of the absolute value of the difference between the first standard deviation and the second standard deviation to obtain the second element concentration similarity index; Calculate the product of the first element concentration similarity index and the second element concentration similarity index to obtain the similarity degree of element concentrations between the q-th data point and the c-th centroid.
[0031] In an embodiment, the calculation formula for the similarity degree of element concentrations between the q-th data point and the c-th centroid is: Among them, represents the similarity degree of element concentration between the q-th data point and the c-th centroid, represents the element concentration of the c-th centroid, represents the element concentration of the q-th data point, represents the standard deviation of the element concentrations of all data points within the 26-neighborhood of the c-th centroid, represents the standard deviation of the element concentrations of all data points within the 26-neighborhood of the q-th data point, represents the absolute value symbol.
[0032] It should be noted that the smaller it is, the more similar the element concentration between the q-th data point and the c-th centroid is. Furthermore, the larger it is, the more likely the sampling points corresponding to the q-th data point and the c-th centroid belong to the same ore body, the more the q-th data point should belong to the cluster where the c-th centroid is located, and the greater the weight assigned to the distance between the q-th data point and the c-th centroid on the concentration axis; the standard deviation can characterize the uniformity of data distribution. the smaller it is, the more similar the element concentration distribution in the neighborhood of the q-th data point is to the element concentration distribution in the neighborhood of the c-th centroid. Furthermore, the larger it is, the more likely the sampling points corresponding to the q-th data point and the c-th centroid belong to the same ore body, the more the q-th data point should belong to the cluster where the c-th centroid is located, and the greater the weight assigned to the distance between the q-th data point and the c-th centroid on the concentration axis.
[0033] (2) According to the similarity degree of element concentration between the q-th data point and the c-th centroid, perform a first weighted correction on the clustering distance between the q-th data point and the c-th centroid to obtain the corrected distance between the q-th data point and the c-th centroid.
[0034] Specifically, calculate the difference between the constant 1 and the similarity degree of element concentration to obtain the position weight, and calculate the sum of the constant 1 and the similarity degree of element concentration to obtain the concentration weight; In the three-dimensional clustering space, calculate the square of the distance between the q-th data point and the c-th centroid on the x-axis (i.e., the horizontal coordinate axis) to obtain the x-axis square term, calculate the square of the distance between any data and any centroid on the y-axis (i.e., the vertical coordinate axis) to obtain the y-axis square term, and calculate the square of the distance between any data and any centroid on the z-axis (i.e., the concentration axis) to obtain the z-axis square term; Use the position weights as the weights of the x-axis squared term and the y-axis squared term respectively, and use the concentration weight as the weight of the z-axis squared term. Perform a weighted sum of the x-axis squared term, the y-axis squared term, and the z-axis squared term to obtain a weighted sum result. Take the arithmetic square root of the weighted sum result as the corrected distance between the q-th data point and the c-th centroid.
[0035] In one embodiment, the formula for calculating the corrected distance between the q-th data point and the c-th centroid is: Where, represents the corrected distance between the q-th data point and the c-th centroid, represents the similarity degree of element concentrations between the q-th data point and the c-th centroid, represents the x-axis coordinate of the c-th centroid in the three-dimensional clustering space, represents the x-axis coordinate of the q-th data point in the three-dimensional clustering space, represents the y-axis coordinate of the c-th centroid in the three-dimensional clustering space, represents the y-axis coordinate of the q-th data point in the three-dimensional clustering space, represents the z-axis coordinate of the c-th centroid in the three-dimensional clustering space, represents the z-axis coordinate of the q-th data point in the three-dimensional clustering space.
[0036] It should be noted that the larger, the more likely the sampling points corresponding to the q-th data point and the c-th centroid belong to the same ore body, the more the q-th data point should be assigned to the cluster where the c-th centroid is located, and the greater the weight assigned to the distance between the q-th data point and the c-th centroid on the concentration axis, that is the larger; in order to maintain weight balance, the weights of the distances between the q-th data point and the c-th centroid on the x-axis and y-axis should be reduced, that is .
[0037] So far, according to the similarity of element concentrations between the q-th data point and the c-th centroid, a first weighted correction is performed on the clustering distance between the q-th data point and the c-th centroid, and the corrected distance between the q-th data point and the c-th centroid is obtained, so that data points with relatively far distances on the two-dimensional topographic map can also be more easily clustered into one cluster according to element concentrations, reducing the occurrence probability of spherical clusters and improving the clustering accuracy of K-means clustering when clustering high-dimensional data.
[0038] Step S103: According to the element concentration difference between any one of the data points and its neighboring data points, and the spatial distribution characteristics of the data points with element concentrations similar to that of any one of the data points, perform a second weighted correction on the corrected distance to obtain the final corrected distance between any one of the data points and any one of the centroids.
[0039] Also, considering that in the sampled stream sediment samples, there may be individual stream sediment samples with additional enrichment, resulting in a relatively high element concentration of element Au. At the same time, if individual stream sediment samples are contaminated by heavy metals, it will also lead to a relatively high element concentration of element Au. However, the element concentration of element Au under heavy metal contamination is not the real ore body element data and belongs to abnormal extreme values. The K-means clustering algorithm will be biased by such abnormal extreme values, affecting the clustering result. Therefore, in the embodiments of the present invention, it is necessary to determine whether the q data points are data points with abnormal element concentrations caused by heavy metal contamination, and perform another correction on the corrected distance between the qth data point and the cth centroid, so as to reduce the possibility that the K-means clustering algorithm is biased by abnormal extreme values and improve the accuracy of analyzing the test results of stream sediment samples using the K-means clustering algorithm.
[0040] Although the element concentrations of ore body elements in different stream sediment samples may be different, the difference in the element concentrations of ore body elements in stream sediment samples at adjacent sampling points is relatively small. Heavy metal contamination belongs to artificial pollution, and the difference in the element concentrations of ore body elements in stream sediment samples at adjacent sampling points is relatively large. On the other hand, heavy metal contamination is artificially caused, and its scale is smaller than that of the ore body. At the same time, there may be heavy metal contamination caused by other industrial points in the two-dimensional topographic map, which is relatively discrete. Ore body elements are naturally formed and follow a gradually extending process during formation, and will form specific element concentrations in specific regions, that is, the sampling points of stream sediment corresponding to data points with similar element concentrations are relatively concentrated. Therefore, in the embodiments of the present invention, according to the element concentration difference between the qth data point and its neighboring data points, and the spatial distribution characteristics of the data points with element concentrations similar to that of the qth data point, the pollution degree of the qth data point is obtained to determine whether the qth data point is a data point with abnormal element concentrations caused by heavy metal contamination. Specifically: In the three-dimensional clustering space, for any other data point except the qth data point, if the absolute value of the difference between the element concentration of the qth data point and the element concentration of any other data point is less than or equal to 0.1 (not limited here, and the implementer can set it according to the specific scenario), then the any other data point is recorded as a similar data point with an element concentration similar to that of the qth data point; If the number of the similar data points is not zero, calculate the Euclidean distance between each of the similar data points and the q-th data point, linearly normalize the average value of all the Euclidean distances, and obtain the first distribution characteristic value of the q-th data point; In the three-dimensional clustering space, denote all the data points within the twenty-six neighborhood of the q-th data point as neighborhood data points. There is no limitation here, and the implementer can set the neighborhood range according to the specific scenario. Calculate the absolute value of the difference between the element concentration of each of the neighborhood data points and the element concentration of the q-th data point, linearly normalize the average value of all the absolute values of the differences, and obtain the second distribution characteristic value of the q-th data point; Calculate the product between the first distribution characteristic value and the second distribution characteristic value of the q-th data point to obtain the pollution degree of the q-th data point.
[0041] In one implementation manner, the calculation formula for the pollution degree of the q-th data point is: Wherein, represents the pollution degree of the q-th data point, represents the element concentration of the q-th data point, represents the element concentration of the j-th neighborhood data point of the q-th data point, represents the number of all the neighborhood data points of the q-th data point, represents the Euclidean distance between the q-th data point and the k-th similar data point of the q-th data point in the three-dimensional clustering space, represents the number of all the similar data points of the q-th data point, and norm() represents the linear normalization function.
[0042] It should be noted that, the larger, the greater the difference in the element concentration between the q-th data point and its neighborhood data points, and thus the larger, the greater the possibility that the q-th data point is a data point with abnormal element concentration caused by heavy metal pollution; the larger, the more discrete the distribution between the q-th data point and its similar data points, and the more in line with the spatial distribution characteristics of the data points with abnormal element concentration caused by heavy metal pollution. Thus the larger, the greater the possibility that the q-th data point is a data point with abnormal element concentration caused by heavy metal pollution.
[0043] Particularly, if the number of the similar data points of the q-th data point is zero, it indicates that the q-th data point is the heaviest metal pollution point with the most independent element concentration, rather than a vein point with a large coverage area of the mineralization area. Therefore, set the pollution degree of the q-th data point to 1.
[0044] Since The value range is [0, 1]. The data points with abnormal element concentrations caused by heavy metal pollution and the normal data points will be divided at both ends of the value range [0, 1]. Therefore, the midpoint 0.5 is used as the threshold, that is, the preset pollution threshold is set to 0.5. There is no limitation here, and the implementer can set it according to the specific scenario. If the pollution degree of the q-th data point is less than 0.5, it is determined that the q-th data point is a normal data point, and at this time, the corrected distance between the q-th data point and the c-th centroid is used as the final corrected distance; if the pollution degree of the q-th data point is greater than or equal to 0.5, it is determined that the q-th data point is a data point with abnormal element concentration caused by heavy metal pollution. At this time, according to the pollution degree of the q-th data point, the corrected distance between the q-th data point and the c-th centroid needs to be secondarily weighted and corrected to obtain the final corrected distance between the q-th data point and the c-th centroid, so as to reduce the possibility that the K-means clustering algorithm is pulled off the centroid by abnormal extreme values and improve the accuracy of analyzing the test results of water system sediment samples using the K-means clustering algorithm. Specifically: Subtract the pollution degree of the q-th data point from the constant 1 to obtain the correction weight, and according to the correction weight, weight the corrected distance between the q-th data point and the c-th centroid to obtain the final corrected distance between the q-th data point and the c-th centroid.
[0045] In one implementation, the calculation formula for the final corrected distance between the q-th data point and the c-th centroid is: Where, represents the final corrected distance between the q-th data point and the c-th centroid, represents the corrected distance between the q-th data point and the c-th centroid, represents the pollution degree of the q-th data point.
[0046] It should be noted that, the larger, the greater the possibility that the q-th data point is a data point with abnormal element concentration caused by heavy metal pollution, and a smaller weight should be given to the corrected distance between the q-th data point and the c-th centroid, that is the smaller, and thus the smaller, the smaller the final corrected distance between the q-th data point and the c-th centroid, that is, the smaller the clustering distance between the q-th data point and the c-th centroid for K-means clustering.
[0047] Step S104, obtain the final corrected distance between each data point and each centroid in each iteration process, perform K-means clustering on all data points in the three-dimensional clustering space according to all the final corrected distances to obtain the clustering result of any element, and analyze the test results of all water system sediment samples according to the clustering results of all elements.
[0048] According to step S103, when performing K-means clustering on the data points in the three-dimensional clustering space, the final corrected distance between each data point and each centroid in each iteration process is obtained. Based on all the final corrected distances, K-means clustering is performed on all the data points in the three-dimensional clustering space to obtain at least one clustering cluster of element Au. The concentration contents of element Au in the stream sediment samples corresponding to the data points belonging to the same clustering cluster are similar, and the geographical location distributions of the sampling points are aggregated.
[0049] Similarly, at least one clustering cluster of each other ore body element is obtained. Since the concentration contents of the ore body elements in the stream sediment samples corresponding to the data points belonging to the same clustering cluster are similar, and the geographical location distributions of the sampling points are aggregated, therefore, according to the regions where the sampling points are located in each clustering cluster corresponding to each ore body element, the geochemical type distribution of the stream sediments is revealed, possible anomalously enriched areas are separated, and mineralization area detection is assisted. For example, the more the number of members in the same clustering cluster, the more likely there is an ore body in the corresponding sampling point region within the clustering cluster, and the search for the ore body distribution is more accurate.
[0050] It should be noted that the focus of the present invention is to improve the accuracy of analyzing the assay results of stream sediment samples by using the K-means clustering algorithm by correcting the clustering distance between the data points and the centroids. Analyzing the assay results of stream sediment samples according to the clustering results is the prior art and will not be elaborated here.
[0051] In summary, the present invention obtains at least two water system sediment samples. For any element, the element concentration of the any element is extracted from each water system sediment sample. Combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample, a three-dimensional clustering space is constructed. When performing K-means clustering on the data points in the three-dimensional clustering space, for any round of iteration process, the centroid of the any round of iteration process is obtained. According to the element concentration similarity between any data point and any centroid, the clustering distance between the any data point and the any centroid is first weighted and corrected to obtain a corrected distance. According to the element concentration difference between the any data point and its neighborhood data points and the spatial distribution characteristics of the data points with element concentration similarity to the any data point, the corrected distance is second weighted and corrected to obtain the final corrected distance between the any data point and the any centroid. The final corrected distance between each data point and each centroid in each round of iteration process is obtained. According to all the final corrected distances, K-means clustering is performed on all the data points in the three-dimensional clustering space to obtain the clustering result of the any element. According to the clustering results of all elements, the test results of all water system sediment samples are analyzed. Among them, by combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of any element in each water system sediment sample, a three-dimensional clustering space is constructed, and the K-means clustering algorithm is used to perform three-dimensional clustering on the data points in the three-dimensional clustering space, so that the clustering result synthesizes the characteristics of element concentration and element geographical spatial distribution, improves the clustering accuracy of the ore body element distribution, improves the accuracy of finding the ore body distribution through the clustering result, and further improves the accuracy of analyzing the test results of water system sediment samples; when performing K-means clustering on the data points in the three-dimensional clustering space, according to the element concentration similarity and spatial distribution characteristics, the clustering distance between each data and the centroid is corrected, reducing the probability of spherical clustering when the K-means clustering algorithm performs high-dimensional clustering, and at the same time reducing the interference of the data points with abnormal element concentrations caused by heavy metal pollution on the clustering result of the K-means clustering, thereby improving the accuracy of the K-means clustering algorithm when performing high-dimensional clustering and improving the accuracy of analyzing the test results of water system sediment samples by using the K-means clustering algorithm.
[0052] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A data analysis method for the test results of water system sediment samples, characterized in that, The data analysis method for the test results of water system sediment samples includes: Obtain at least two water system sediment samples. For any element, extract the element concentration of the any element in each water system sediment sample, and combine the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample to construct a three-dimensional clustering space; When performing K-means clustering on the data points in the three-dimensional clustering space, for any round of iteration process, obtain the centroid of the any round of iteration process, and perform a first weighted correction on the clustering distance between the any data point and the any centroid according to the element concentration similarity between the any data point and the any centroid to obtain a corrected distance; Perform a second weighted correction on the corrected distance according to the element concentration difference between the any data point and its neighborhood data points and the spatial distribution characteristics of the data points with element concentration similarity to the any data point to obtain the final corrected distance between the any data point and the any centroid; Obtain the final corrected distance between each data point and each centroid in each round of iteration process, perform K-means clustering on all the data points in the three-dimensional clustering space according to all the final corrected distances to obtain the clustering result of the any element, and analyze the test results of all the water system sediment samples according to the clustering results of all the elements.
2. The data analysis method for the test results of a water system sediment sample according to claim 1, characterized in that The constructing a three-dimensional clustering space by combining the position of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample includes: According to the coordinates of the sampling point of each water system sediment sample in the two-dimensional topographic map and the element concentration of the any element in each water system sediment sample, map the element concentration of the any test element in each water system sediment sample into the same three-dimensional clustering space. The x-axis of the three-dimensional clustering space represents the abscissa of the corresponding sampling point of each water system sediment sample in the two-dimensional topographic map, the y-axis represents the ordinate of the corresponding sampling point of each water system sediment sample in the two-dimensional topographic map, and the z-axis represents the element concentration of the any test element in each water system sediment sample.
3. The data analysis method for the test results of a water system sediment sample according to claim 1, characterized in that, The performing a first weighted correction on the clustering distance between the any data point and the any centroid according to the element concentration similarity between the any data point and the any centroid to obtain a corrected distance includes: Calculate the reciprocal of the absolute value of the difference between the element concentration of the any data point and the element concentration of the any centroid to obtain a first element concentration similarity index; In the three-dimensional clustering space, calculate the standard deviation of the element concentrations of all the data points in the neighborhood of the any data point, denoted as the first standard deviation, calculate the standard deviation of the element concentrations of all the data points in the neighborhood of the any centroid, denoted as the second standard deviation, and calculate the reciprocal of the absolute value of the difference between the first standard deviation and the second standard deviation to obtain a second element concentration similarity index; Calculate the product between the first element concentration similarity index and the second element concentration similarity index to obtain the element concentration similarity degree between any data point and any centroid. Obtain the corrected distance between any data point and any centroid according to the element concentration similarity degree.
4. The data analysis method for the test results of a water system sediment sample according to claim 3, characterized in that, The obtaining the corrected distance between any data point and any centroid according to the element concentration similarity degree includes: Calculate the difference between 1 and the element concentration similarity degree to obtain the position weight, and calculate the sum of 1 and the element concentration similarity degree to obtain the concentration weight. In the three-dimensional clustering space, calculate the square of the distance between any data and any centroid on the x-axis to obtain the x-axis square term, calculate the square of the distance between any data and any centroid on the y-axis to obtain the y-axis square term, and calculate the square of the distance between any data and any centroid on the z-axis to obtain the z-axis square term. Use the position weight as the weight of the x-axis square term and the y-axis square term respectively, use the concentration weight as the weight of the z-axis square term, perform weighted summation on the x-axis square term, the y-axis square term, and the z-axis square term to obtain the weighted summation result, and take the arithmetic square root of the weighted summation result as the corrected distance between any data point and any centroid.
5. The data analysis method for assay results of a water system sediment sample according to claim 1, wherein, The second weighted correction of the corrected distance according to the element concentration difference between any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similar to that of any data point to obtain the final corrected distance between any data point and any centroid includes: Obtain the pollution degree of any data point according to the element concentration difference between any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similar to that of any data point. If the pollution degree is greater than or equal to the preset pollution threshold, perform a second weighted correction on the corrected distance between any data and any centroid according to the pollution degree of any data point to obtain the final corrected distance between any data point and any centroid. If the pollution degree is less than the preset pollution threshold, use the corrected distance between any data and any centroid as the final corrected distance.
6. The data analysis method for assay results of a water system sediment sample according to claim 5, wherein The obtaining the pollution degree of any data point according to the element concentration difference between any data point and its neighborhood data points, and the spatial distribution characteristics of the data points with element concentration similar to that of any data point includes: Obtain the similar data points with element concentration similar to that of any data point in the three-dimensional clustering space. If the number of similar data points is 0, set the pollution degree of any data point to 1. If the number of similar data points is not 0, calculate the Euclidean distance between each similar data point and any data point respectively, linearly normalize the average value of all Euclidean distances to obtain the first distribution characteristic value of any data point. In the three-dimensional clustering space, all the data points within the neighborhood of any one of the data points are denoted as neighborhood data points. The absolute value of the difference between the element concentration of each neighborhood data point and the element concentration of any one of the data points is calculated, and the average value of all the absolute values of the differences is linearly normalized to obtain the second distribution feature value of any one of the data points; Calculate the product between the first distribution feature value and the second distribution feature value of any one of the data points to obtain the pollution degree of any one of the data points.
7. A method for analyzing the test results data of a water system sediment sample according to claim 5, characterized in that, The second weighted correction of the corrected distance between any one of the data points and any one of the centroids according to the pollution degree of any one of the data points to obtain the final corrected distance between any one of the data points and any one of the centroids includes: Subtract the pollution degree of any one of the data points from the constant 1 to obtain a correction weight, and weight the corrected distance between any one of the data points and any one of the centroids according to the correction weight to obtain the final corrected distance between any one of the data points and any one of the centroids.
8. The data analysis method for the test results of a water system sediment sample according to claim 6, characterized in that, The obtaining of similar data points with element concentrations similar to those of any one of the data points in the three-dimensional clustering space includes: In the three-dimensional clustering space, for any other data point other than any one of the data points, if the absolute value of the difference between the element concentration of any one of the data points and the element concentration of any other data point is less than or equal to a preset difference threshold, then any other data point is denoted as a similar data point with an element concentration similar to that of any one of the data points.
Citation Information
Patent Citations
Method for delineating metallogenic prospective area in alpine mountainous area based on regional geochemistry
CN111983715A
Sedimentary facies boundary identification method and device
CN112347823A
Air quality detection and evaluation method based on multi-sensor data
CN116735807A
Underground resource detection method based on multi-source data fusion
CN117828379A
Method for identifying key risk source area of heavy metal pollution in mine water system sediment
CN118095837A