A method and system for accurately identifying the cause of heavy metal pollution in soil

By acquiring surface and deep features of soil sampling points, and combining river direction sorting and historical features of upstream abnormal sampling points to update weights, the SOM model is used to identify the causes of heavy metal pollution in soil. This solves the problem of identification bias in the dynamic diffusion process of the SOM algorithm and achieves higher identification accuracy.

CN120763552BActive Publication Date: 2026-01-02GUANGDONG INST OF ECO ENVIRONMENT & SOIL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511281787.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-01-02
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing self-organizing map (SOM) algorithms cannot effectively reflect the dynamic diffusion process of pollution in soil heavy metal pollution cause identification, leading to bias and errors in identification results and making it difficult to meet the needs of accurate identification.

Method used

By acquiring surface and deep features of soil sampling points and sorting them according to river direction, abnormal sampling points are identified. The initial weights are updated based on the historical features of upstream abnormal sampling points, and a pre-trained SOM model is used for identification.

Benefits of technology

It improves the accuracy of identifying the causes of heavy metal pollution in soil, and can comprehensively and accurately reflect the dynamic diffusion process of the pollution range and degree, avoiding the bias of static analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763552B_ABST
    Figure CN120763552B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of soil testing, in particular to a method and system for accurately identifying the causes of heavy metal pollution in soil, which comprises the following steps: obtaining the surface soil characteristics and deep soil characteristics of each sampling point in the soil to be detected, wherein the deep soil characteristics comprise heavy metal concentration data; sorting the sampling points according to the water flow direction of the river where the sampling points are located to obtain a sampling point sequence, and determining abnormal sampling points from the sampling points according to the surface soil characteristics and deep soil characteristics of each sampling point in the sampling point sequence and historical soil characteristics; updating the initial weight of the current sampling point according to the historical soil characteristics of the upstream abnormal sampling points upstream of the current sampling point in the sampling point sequence to obtain an updated weight; and inputting the surface soil characteristics, deep soil characteristics and updated weight of each sampling point into a pre-trained SOM model to identify the causes of heavy metal pollution in the soil to be detected by using the pre-trained SOM model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of soil testing, in particular to a soil heavy metal pollution cause accurate identification method and system. BACKGROUND

[0002] Soil heavy metal pollution refers to the accumulation of elements such as mercury and lead with significant biological toxicity in soil caused by human activities, which exceeds the natural background value and poses potential risks to the ecological environment. The causes of such pollution are complex and diverse, including wastewater, waste gas and waste residue discharged during industrial production, excessive use of fertilizers and pesticides in agricultural production, sewage irrigation, deposition of heavy metal particles in the atmosphere, and exhaust emissions from vehicles, all of which can cause abnormal increase of heavy metal content in soil. In the field of soil heavy metal pollution cause identification, traditional methods mainly rely on chemical analysis and statistical models, focusing on qualitative or semi-quantitative analysis of pollution characteristics. However, due to its reliance on a large amount of sampling data, it not only consumes a lot of manpower, material resources and time, but also performs poorly in terms of timeliness and accuracy.

[0003] In some scenarios, soil heavy metal pollution cause accurate identification methods based on neural networks have gradually attracted attention and been applied. Among them, the Self-Organizing Map (SOM) algorithm, as an unsupervised learning neural network, has been widely used in the field of soil heavy metal pollution cause accurate identification due to its advantages in processing multi-dimensional data and clustering analysis. This algorithm can cluster multi-dimensional data related to soil heavy metal pollution, thereby identifying the cause of pollution. However, soil heavy metal pollution has significant divergence, with its pollution range and degree dynamically spreading over time and environmental conditions. The static clustering characteristics of the SOM algorithm make it unable to effectively reflect this dynamic diffusion process. Therefore, when using the SOM algorithm to cluster and identify the cause of soil heavy metal pollution based on multi-dimensional data, due to the inability to identify the dynamic diffusion of pollution, the identification results often deviate or even error, making it difficult to meet the demand for accurate identification of soil heavy metal pollution causes. SUMMARY

[0004] In order to solve the technical problem of low accuracy of the identification result of the cause of soil heavy metal pollution, the purpose of the present application is to provide a soil heavy metal pollution cause accurate identification method and system.

[0005] To solve the above technical problems, the technical solutions adopted are as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for accurately identifying the cause of heavy metal pollution in soil, comprising: obtaining surface soil characteristics and deep soil characteristics of each sampling point in the soil to be detected, the deep soil characteristics including heavy metal concentration data; sorting each sampling point according to the water flow direction of the river where the sampling point is located to obtain a sampling point sequence, and determining an abnormal sampling point from each sampling point according to the surface soil characteristics and the deep soil characteristics of each sampling point in the sampling point sequence and historical soil characteristics; updating the initial weight of the current sampling point according to the historical soil characteristics of the upstream abnormal sampling point upstream of the current sampling point in the sampling point sequence to obtain an updated weight; inputting the surface soil characteristics, the deep soil characteristics and the updated weight of each sampling point into a pre-trained SOM model, and identifying the cause of heavy metal pollution in the soil to be detected by using the pre-trained SOM model.

[0007] Optionally, determining the abnormal sampling point from each sampling point according to the surface soil characteristics and the deep soil characteristics of each sampling point in the sampling point sequence and the historical soil characteristics comprises: determining the real-time soil similarity between each sampling point according to the surface soil characteristics and the deep soil characteristics of each sampling point in the sampling point sequence.

[0008] determining the historical soil similarity between each sampling point based on the surface soil characteristics and the deep soil characteristics in the historical soil characteristics of each sampling point;

[0009] determining the soil state consistency of each sampling point with the historical soil based on the historical soil similarity between each sampling point and the real-time soil similarity between each sampling point; determining the abnormal degree of the soil state consistency between each sampling point by using the soil state consistency of each sampling point; and determining the abnormal sampling point from each sampling point based on the abnormal degree.

[0010] Optionally, determining the real-time soil similarity between each sampling point according to the surface soil characteristics and the deep soil characteristics of each sampling point in the sampling point sequence comprises: calculating a first similarity of the surface soil characteristics between the current sampling point and other sampling points in the sampling point sequence, and a second similarity of the deep soil characteristics between the current sampling point and other sampling points in the sampling point sequence; and determining the real-time soil similarity between the current sampling point and other sampling points in the sampling point sequence based on the first similarity and the second similarity.

[0011] Optionally, determining the abnormal degree of the soil state consistency between each sampling point by using the soil state consistency of each sampling point comprises: determining a first number of sampling points whose soil state consistency is greater than a first threshold; and determining the abnormal degree of the soil state consistency between each sampling point according to the first number, the number of all sampling points in the sampling point sequence, and the soil state consistency of the sampling points greater than the first threshold.

[0012] Optionally, the initial weight of the current sampling point is updated according to the historical soil characteristics of the upstream abnormal sampling point upstream of the current sampling point in the sampling point sequence, to obtain an updated weight, including: determining the upstream abnormal sampling point upstream of the current sampling point from the abnormal sampling points in the order of the sampling point sequence; determining a soil state change degree value of the upstream abnormal sampling point relative to a historical sampling time according to a historical heavy metal concentration change curve in the historical soil characteristics of the upstream abnormal sampling point; determining a first influence degree of the current sampling point on the upstream abnormal sampling point according to the second number of the upstream abnormal sampling points, the third number of the abnormal sampling points, and the soil state change degree value; determining a second influence degree of the current sampling point on a downstream sampling point of the current sampling point according to the first influence degree; and updating the initial weight of each sampling point according to the second influence degree of each sampling point to obtain the updated weight.

[0013] Optionally, the soil state change degree value of the upstream abnormal sampling point relative to the historical sampling time is determined according to the historical heavy metal concentration change curve in the historical soil characteristics of the upstream abnormal sampling point, including: determining a fourth number of heavy metal elements in the historical soil characteristics of the upstream abnormal sampling point; determining a slope value of the historical heavy metal concentration change curve of the upstream abnormal sampling point at the historical sampling time, and a fluctuation amplitude of the historical heavy metal concentration change curve of the upstream abnormal sampling point; and determining the soil state change degree value according to the fourth number, the slope value, and the fluctuation amplitude.

[0014] Optionally, the first influence degree of the current sampling point on the upstream abnormal sampling point is determined according to the second number of the upstream abnormal sampling points, the third number of the abnormal sampling points, and the soil state change degree value, including: determining a first ratio of the second number to the third number; and determining a product between the first ratio and the soil state change degree value as the first influence degree.

[0015] Optionally, the initial weight of each sampling point is updated according to the second influence degree of each sampling point to obtain the updated weight, including: determining a second ratio between the second influence degree of each sampling point and a superposition value of the second influence degrees of all sampling points; and determining a product between the second ratio and the initial weight as the updated weight.

[0016] Optionally, the surface soil characteristics, the deep soil characteristics, and the updated weight of each sampling point are input into a pre-trained SOM model, and the pre-trained SOM model is used to identify the soil heavy metal pollution causes of the soil to be detected, including: combining the surface soil characteristics and the deep soil characteristics after normalization processing to obtain a pollution index vector of each sampling point; inputting the pollution index vector and the updated weight into the pre-trained SOM model for heavy metal pollution cause identification, and outputting a U matrix of the soil to be detected; and determining the soil heavy metal pollution causes of the soil to be detected based on a color scale change of the U matrix.

[0017] In a second aspect, an embodiment of the present application provides a soil heavy metal pollution cause accurate identification system, comprising: a processor and a memory; wherein the memory is used to store a computer program which can run on the processor; the processor is used to execute the program stored on the memory to realize the steps of the soil heavy metal pollution cause accurate identification method mentioned in the first aspect.

[0018] The present application has the following beneficial effects:

[0019] The embodiment of the present application determines the abnormal sampling points by the surface soil characteristics, deep soil characteristics and historical soil characteristics of each sampling point, and combines the ordering of the sampling points in the river flow direction, which not only considers the current state of the soil, but also takes into account the historical change rule, and also integrates the influence of the river as an important geographical element, can effectively reflect the dynamic diffusion process of the pollution range and degree, makes the identification of the abnormal sampling points more comprehensive and accurate, and avoids the problem that the deviation or even error of the identification result may be caused by the static analysis. And the initial weight of the current sampling point is updated according to the historical soil characteristics of the upstream abnormal sampling point of the current sampling point, fully considers the migration and dynamic diffusion of the soil pollution in the river basin, the pollution condition of the upstream often has an impact on the downstream, and the weight updating mechanism can make the SOM model focus on the key data which is greatly affected by the upstream pollution, so as to obtain more accurate clustering structure of the SOM model combined with the dynamic diffusion between different sampling points, and effectively improve the accuracy of the SOM model in identifying the cause of soil heavy metal pollution. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0021] Figure 1 A flow chart of a soil heavy metal pollution cause accurate identification method provided by an embodiment of the present application.

[0022] Figure 2 A schematic diagram of setting sampling points in farmland provided by an embodiment of the present application.

[0023] Figure 3 A structural schematic diagram of another soil heavy metal pollution cause accurate identification system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a method and system for accurate identification of the causes of heavy metal pollution in soil proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0026] The specific scheme of the method for accurately identifying the causes of heavy metal pollution in soil provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0027] Example 1:

[0028] Please see Figure 1 The flowchart illustrates a method for accurate identification of the causes of heavy metal pollution in soil according to an embodiment of the present invention, including:

[0029] Step S101: Obtain the surface soil characteristics and deep soil characteristics of each sampling point in the soil to be tested. The deep soil characteristics include heavy metal concentration data.

[0030] Specifically, this embodiment of the invention, based on the "Technical Guidelines for Heavy Metal Screening in Cultivated Land Soil," sets up sampling points in the soil to be tested. For example, such as... Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the setting of sampling points in farmland according to an embodiment of the present invention. Figure 2 In farmland areas, sampling points are set up using a 1km*1km grid method. In mining or industrial areas, a pollution source gradient distribution method is used. This method centers on known or suspected pollution sources and sets up sampling points along the possible directions of pollutant diffusion (such as river flow direction, prevailing wind direction, groundwater runoff direction, etc.), according to a gradient from near to far from the pollution source. Specifically, in high-risk areas surrounding the pollution source (such as within 50 meters of the pollution source boundary), the sampling point density can be increased to ensure the capture of data from areas with the highest pollutant concentration and most significant pollution characteristics. As the distance increases, sampling points are sequentially deployed at different gradient intervals such as 100 meters, 300 meters, and 500 meters, and the spacing can be flexibly adjusted according to the pollutant diffusion capacity (such as the mobility of heavy metals, soil permeability, etc.). Simultaneously, a control sampling point is set at the farthest end of the gradient sequence to reflect the background soil conditions unaffected by the pollution source. This allows for the acquisition of multiple sampling points for the soil to be tested.

[0031] Further, after setting the sampling points, the surface layer soil of 0cm-20cm of each sampling point is sampled layer by layer, and the color, texture and the like are recorded as the surface layer soil features. The deep layer soil samples of the depth of more than 20cm are recorded, and the concentration data of each heavy metal in each deep layer soil sample is detected by an inductively coupled plasma mass spectrometry (ICP-MS) detection technology, so as to obtain the deep layer soil features. In addition, the river distribution around the soil to be detected is obtained through the image shot by the unmanned aerial vehicle.

[0032] In step S102, the sampling points are sorted according to the water flow direction of the river where each sampling point is located, so as to obtain a sampling point sequence, and the abnormal sampling point is determined from each sampling point according to the surface layer soil features and the deep layer soil features of each sampling point in the sampling point sequence and the historical soil features.

[0033] Specifically, the embodiment of the present application has S sampling points. For the s-th sampling point, first, the other sampling points in the same irrigation river or flow-through river as the s-th sampling point are determined in the region of the soil to be sampled, which are referred to as the reference sampling points of the s-th sampling point. Then, the s-th sampling point and its reference sampling points are sorted according to the water flow direction of the irrigation river or flow-through river, so as to obtain the sampling point sequence of the s-th sampling point. The surface layer soil features and the deep layer soil features of the s-th sampling point collected in this detection are taken as the real-time soil features of the s-th sampling point, and the soil features collected in the historical detection process of the s-th sampling point are taken as the historical soil features.

[0034] Further, for the s-th sampling point, the surface layer soil features of the surface layer soil of the s-th sampling point constitute a surface layer soil feature set, such as {red, loam}. The concentration data of all heavy metals in the detection result of the deep layer soil of the s-th sampling point are taken as the deep layer soil features, and the concentration data of all heavy metals in the deep layer soil features are merged to constitute a deep layer soil feature set, such as {0.12, 0.05, 30, …}, wherein the concentration of cadmium, mercury, lead, … and other heavy metals is represented in sequence, and the unit is mg / kg.

[0035] Further, as an optional embodiment of the present application, the determining of the abnormal sampling point from the sampling points according to the surface soil features and deep soil features of each sampling point in the sampling point sequence and the historical soil features comprises: determining real-time soil similarities between the sampling points according to the surface soil features and deep soil features of each sampling point in the sampling point sequence; determining historical soil similarities between the sampling points based on the surface soil features and deep soil features in the historical soil features of each sampling point; determining soil state consistencies between the sampling points and the historical soil based on the historical soil similarities between the sampling points and the real-time soil similarities between the sampling points; determining abnormal degrees of the soil state consistencies between the sampling points by using the soil state consistencies of the sampling points; and determining the abnormal sampling point from the sampling points based on the abnormal degrees.

[0036] Specifically, the historical soil similarities between the sampling points in the embodiment of the present application are calculated by using the historical surface soil features and historical deep soil features in the historical detection process of each sampling point, and the calculation manner of the historical soil similarities is the same as that of the real-time soil similarities. The calculation process of the real-time soil similarities comprises: calculating first similarities of the surface soil features between a current sampling point in the sampling point sequence and other sampling points, and second similarities of the deep soil features between the current sampling point in the sampling point sequence and the other sampling points; and determining the real-time soil similarities between the current sampling point in the sampling point sequence and the other sampling points based on the first similarities and the second similarities.

[0037] More specifically, the first similarities are represented by Jaccard correlation coefficients, and the second similarities can be represented by Euclidean distances. The embodiment of the present application takes the s-th sampling point as an example, and calculates the real-time soil similarities between the s-th sampling point and other sampling points in the sampling point sequence of the s-th sampling point by using the following formula:

[0038]

[0039] In the above formula, represents the real-time soil similarity between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point. represents a Jaccard correlation coefficient, which is used to calculate the first similarity between and . represents a Euclidean distance, which is used to calculate the second similarity between and . is to avoid the denominator being 0. represents the surface soil features of the s-th sampling point. The surface soil feature of the i-th sampling point in the sampling point sequence of the s-th sampling point, excluding the s-th sampling point. The deep soil feature of the s-th sampling point. The deep soil feature of the i-th sampling point in the sampling point sequence of the s-th sampling point, excluding the s-th sampling point.

[0040] Wherein, The greater the value of the soil state similarity between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point. Since in the process of accurately identifying the cause of heavy metals in the soil, multiple soil feature data are usually collected for each sampling point. Therefore, the present embodiment of the application sets that R times of sampling data are statistically collected for each sampling point. Then, the historical soil similarity between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point in each statistical process is calculated using the above method, denoted as .

[0041] Further, for the r-th historical sampling data of the s-th sampling point, the present embodiment of the application calculates the difference between the real-time soil similarity and the historical soil similarity , and the normalized value of the absolute value of the difference (obtained by normalizing the absolute value of the difference using a normalization function) is denoted as , which represents the soil state consistency of the s-th sampling point and the i-th sampling point in the sampling point sequence with respect to the historical r-th historical soil feature data. If the calculated soil state consistency is greater, it indicates that the real-time soil state of the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point is less consistent with the soil state of the i-th sampling point in the sampling point sequence of the s-th sampling point in the r-th historical sampling process. It may indicate that the real-time soil state of the s-th sampling point has changed greatly compared to the historical r-th sampling, and the soil state of the i-th sampling point in the sequence s sampling point sequence.

[0042] Further, as an optional embodiment of the present application, determining the abnormality degree of the soil state consistency between the sampling points using the soil state consistency of each sampling point comprises: determining a first number of sampling points whose soil state consistency is greater than a first threshold value; and determining the abnormality degree of the soil state consistency between the sampling points according to the first number, the historical sampling times of each sampling point, and the soil state consistency of the sampling points greater than the first threshold value.

[0043] Specifically, the present embodiment calculates the soil state consistency of the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point with respect to the historical R times of sampling data in the above manner. Then, all soil state consistencies the number of sampling points greater than the first threshold value is denoted as When the value is greater and the corresponding soil state consistency is smaller, it indicates that the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point is more likely to be abnormal. The first threshold value can be determined according to actual conditions, and in the embodiment of the present application, the value is 0.3. The embodiment of the present application takes the s-th sampling point as an example, and calculates the abnormal degree of the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point by using the following formula:

[0044]

[0045] In the above formula, indicates the abnormal degree of the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point. indicates the number of all sampling points in the sampling point sequence of the s-th sampling point. indicates the first number of sampling points greater than the first threshold value in all indicates the proportion of the sampling points greater than the first threshold value in all indicates the proportion of the sampling points greater than the first threshold value in all the sampling point sequence of the s-th sampling point, and the greater the ratio, the greater the abnormal degree of the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point. indicates the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point in the historical j-th sampling data. indicates a normalization function, which is used for normalizing .

[0046] Further, after the abnormal degree is calculated, if the calculated is greater than a second threshold value, it indicates that the abnormal degree of the soil state consistency between the s-th sampling point and the i-th sampling point in the sampling point sequence of the s-th sampling point is greater, indicating that the soil state between the s-th sampling point and the i-th sampling point is abnormal, and the i-th sampling point is taken as an abnormal sampling point of the s-th sampling point. According to the above manner, all abnormal sampling points in the s-th sampling point sequence and the s-th sampling point whose soil state is abnormal are determined, and the number thereof is denoted as M. The second threshold value can be valued according to actual conditions, and in the embodiment of the present application, the value is 0.5.

[0047] In step S103, the initial weight of the current sampling point is updated according to the historical soil characteristics of the upstream abnormal sampling point upstream of the current sampling point in the sampling point sequence, to obtain an updated weight.

[0048] Specifically, after identifying the abnormal sampling points using the above method, this embodiment of the invention analyzes whether the abnormal sampling points in the sampling point sequence of the s-th sampling point and the soil condition of the s-th sampling point are caused by the flow of the same irrigation river or flowing river. First, this embodiment of the invention finds the upstream abnormal sampling points located upstream of the s-th sampling point among the above M abnormal sampling points, assuming a total of [number missing] upstream abnormal sampling points are found. One upstream abnormal sampling point. Then analyze. Whether the soil characteristic data and historical characteristic data collected from the upstream abnormal sampling points show any anomalies. Then, based on... The initial weights are updated by analyzing the degree of soil state change at each upstream abnormal sampling point relative to historical sampling times.

[0049] Furthermore, as an optional embodiment of the present invention, updating the initial weight of the current sampling point based on the historical soil characteristics of the upstream anomalous sampling points upstream of the current sampling point in the sampling point sequence to obtain the updated weight includes: determining the upstream anomalous sampling points upstream of the current sampling point from the anomalous sampling points according to the order in the sampling point sequence; determining the degree of soil state change of the upstream anomalous sampling point relative to the historical sampling time based on the historical heavy metal concentration change curve in the historical soil characteristics of the upstream anomalous sampling point; determining the first degree of influence of the current sampling point on the upstream anomalous sampling point based on the second number of upstream anomalous sampling points, the third number of anomalous sampling points, and the degree of soil state change; determining the second degree of influence of the current sampling point on the downstream sampling points of the current sampling point based on the first degree of influence; and updating the initial weight of each sampling point based on the second degree of influence of each sampling point to obtain the updated weight.

[0050] Specifically, in this embodiment of the invention, the sampling points in the sampling point sequence are arranged sequentially according to the river's flow direction. Therefore, upstream anomalous sampling points are determined from the anomalous sampling points based on the order of the sampling points in the sampling point sequence and the river's direction. The historical heavy metal concentration change curve refers to the curve showing the change in heavy metal concentrations of various types of heavy metals in the historical soil characteristics of the upstream anomalous sampling points during multiple historical sampling processes. For each metal, the historical heavy metal concentration change curve shows a slope value for the heavy metal concentration at each historical sampling time at the anomalous sampling point, as well as corresponding maximum and minimum metal concentration values. In this embodiment of the invention, the absolute value of the difference between the maximum and minimum metal concentration values ​​is used as the fluctuation amplitude. The degree of soil state change is calculated based on the slope value and the fluctuation amplitude.

[0051] Further, as an optional embodiment of the present application, the soil state change degree value of the upstream abnormal sampling point relative to the historical sampling time is determined according to the historical heavy metal concentration change curve in the historical soil feature of the upstream abnormal sampling point, comprising: determining a fourth number of heavy metal elements in the historical soil feature of the upstream abnormal sampling point; determining a slope value of the historical heavy metal concentration change curve of the upstream abnormal sampling point at the historical sampling time, and a fluctuation amplitude of the historical heavy metal concentration change curve of the upstream abnormal sampling point; and determining the soil state change degree value according to the fourth number, the slope value and the fluctuation amplitude.

[0052] Specifically, the embodiment of the present application takes the kth upstream abnormal sampling point of the st sampling point as an example, and calculates the soil state change degree value of the kth upstream abnormal sampling point relative to the historical sampling time by using the following formula:

[0053]

[0054] In the above formula, indicates the soil state change degree value of the kth upstream abnormal sampling point of the st sampling point relative to the historical sampling time. indicates the fourth number of heavy metal elements. indicates a normalization function, which is used to normalize . indicates the absolute value of the slope value of the historical heavy metal concentration change curve of the pth heavy metal element of the kth upstream abnormal sampling point of the st sampling point at the historical sampling time. indicates the fluctuation amplitude of the historical heavy metal concentration change curve of the pth heavy metal element of the kth upstream abnormal sampling point of the st sampling point. and The smaller the value of and is, the more stable the concentration fluctuation of the pth heavy metal element of the kth upstream abnormal sampling point of the st sampling point is. Therefore, The larger the value of is, the greater the change degree of the soil state of the kth upstream abnormal sampling point of the st sampling point relative to the historical sampling time is.

[0055] Further, as an optional embodiment of the present application, the first influence degree of the current sampling point on the upstream abnormal sampling point is determined according to the second number of upstream abnormal sampling points, the third number of abnormal sampling points and the soil state change degree value, comprising: determining a first ratio of the second number to the third number; and determining the product between the first ratio and the soil state change degree value as the first influence degree.

[0056] Specifically, the embodiment of the present application takes the kth upstream abnormal sampling point of the st sampling point as an example, and calculates the first influence degree of the kth upstream abnormal sampling point on the st sampling point by using the following formula:

[0057]

[0058] In the above formula, represents the first influence degree of the s-th sampling point by the k-th upstream abnormal sampling point. represents the third number of all abnormal sampling points in the sampling point sequence of the s-th sampling point and the soil state of the s-th sampling point is abnormal. represents the second number of upstream abnormal sampling points among all abnormal sampling points in which the soil state is abnormal and which are upstream of the s-th sampling point. represents the soil state change degree value of the k-th upstream abnormal sampling point of the s-th sampling point compared with the soil state change at the historical sampling time. The greater the value of the s-th sampling point is, the more sampling points upstream of the s-th sampling point in the sampling point sequence of the s-th sampling point are among all abnormal sampling points in which the soil state of the s-th sampling point is abnormal, and the s-th sampling point is more easily affected by the upstream sampling points. Therefore The greater the value of the s-th sampling point is, the greater the influence degree of the s-th sampling point by the k-th upstream abnormal sampling point is.

[0059] Further, in the embodiment of the present application, the first influence degree of the sampling point by other sampling points is calculated by the above method, and for the s-th sampling point, it is assumed that it affects Y downstream sampling points, and the number of the first influence degree of the Y downstream sampling points by the upstream s-th sampling point is also Y. Therefore, the second influence degree of the s-th sampling point on other downstream sampling points is the mean value of the first influence degree of all Y downstream sampling points by the upstream sampling point s, which is denoted as in the embodiment of the present application.

[0060] Further, since the influence degree of each sampling point by the upstream sampling point is different, and in the clustering process based on the SOM neural network algorithm, the winning neuron needs to be found through continuous iteration to achieve optimal clustering. Since the influence degree of each sampling point by the upstream sampling point is different, the sampling point with high influence degree needs to be given stronger influence in the iteration update process. Therefore, the initial weight of each sampling point is adjusted according to the second influence degree in the embodiment of the present application, and the sampling point with greater influence degree is given a larger weight value. In an optional embodiment of the present application, the initial weight of each sampling point is updated according to the second influence degree of each sampling point to obtain the updated weight, which includes: determining the second ratio between the second influence degree of each sampling point and the superposition value of the second influence degree of all sampling points; and determining the product between the second ratio and the initial weight as the updated weight.

[0061] Specifically, taking the s-th sampling point as an example, the updated weight of the s-th sampling point is calculated by the following formula in the embodiment of the present application:

[0062]

[0063] In the above formula, represents the update weight of the s-th sampling point. represents the second influence degree of the s-th sampling point on other sampling points. represents the number of all sampling points. is the initial weight of the s-th sampling point. represents the second influence degree of the q-th sampling point on other sampling points.

[0064] In step S104, the surface soil characteristics, deep soil characteristics and update weight of each sampling point are input into the pre-trained SOM model, and the pre-trained SOM model is used to identify the soil heavy metal pollution cause of the soil to be detected.

[0065] Specifically, the surface soil characteristics and deep soil characteristics of each sampling point are normalized in the embodiment of the application. Then, the normalized surface soil characteristics and deep soil characteristics are combined to form a pollution index vector of the sampling point. For example, the surface soil characteristics of the s-th sampling point have 3 characteristics, and the deep soil characteristics have 5 indexes, so the pollution index vector of the s-th sampling point is an 8-dimensional vector. Then, the pollution index vector and the update weight are input into the pre-trained SOM model, and the pre-trained SOM model is used to identify the soil heavy metal pollution cause of the soil to be detected.

[0066] Further, as an optional embodiment of the application, inputting the surface soil characteristics, deep soil characteristics and update weight of each sampling point into the pre-trained SOM model to identify the soil heavy metal pollution cause of the soil to be detected includes: combining the normalized surface soil characteristics and deep soil characteristics to obtain the pollution index vector of each sampling point; inputting the pollution index vector and the update weight into the pre-trained SOM model for heavy metal pollution cause identification, and outputting the U matrix of the soil to be detected; determining the soil heavy metal pollution cause of the soil to be detected based on the color scale change of the U matrix.

[0067] Specifically, the embodiment of the present application converts the surface soil characteristics of the s-th sampling point into a numerical form through encoding, and then normalizes the numerical form of the surface soil characteristics and the deep soil characteristics. Then, it is combined to form the pollution index vector of the s-th sampling point. For example, if the surface soil characteristics of the s-th sampling point have 3 characteristics and the deep soil characteristics have 5 indicators, then the pollution index vector of the s-th sampling point is an 8-dimensional vector. Then, the update weight of the pollution index vector of all sampling points is input into the SOM model, which can be used as the weight of the neighborhood neurons of the SOM model. The SOM model outputs the corresponding U matrix through competitive learning to reveal the pollution boundary and hotspot area. It is worth noting that the internal training and parameter adjustment process of the SOM model can be known in the art, and the embodiment of the present application will not be repeated here.

[0068] Further, after obtaining the U matrix, the embodiment of the present application determines the soil heavy metal pollution causes of the to-be-detected soil by the color scale change of the U matrix. Among them, the red color in the U matrix represents a topological distance range of greater than 85% quantile, a pollution level of high risk, and a pollution cause of multiple heavy metal element complex pollution. The yellow color in the U matrix represents a topological distance range of 60% to 85% quantile, a pollution level of medium risk, and a pollution cause of single heavy metal element pollution. The blue color in the U matrix represents a topological distance range of less than 60% quantile, a risk level of low risk, and a heavy metal element concentration within a normal range, which is not polluted.

[0069] The embodiment of the present application determines the abnormal sampling point by the surface soil characteristics, deep soil characteristics and historical soil characteristics of each sampling point, and combines the ordering of the sampling points in the river flow direction, which not only considers the current state of the soil, but also takes into account the historical change rule, and also incorporates the influence of the river as an important geographical element, can effectively reflect the dynamic diffusion process of the pollution range and degree, make the identification of the abnormal sampling point more comprehensive and accurate, and avoid the problem that static analysis can lead to deviation or even error of the identification result. And according to the historical soil characteristics of the upstream abnormal sampling point of the current sampling point, the initial weight of the current sampling point is updated, which fully considers the migration and dynamic diffusion of soil pollution in the river basin. The pollution of the upstream often has an impact on the downstream, and the weight updating mechanism can make the SOM model focus more on the key data affected by the upstream pollution, thereby effectively improving the accuracy of the SOM model in identifying the causes of soil heavy metal pollution.

[0070] Embodiment two:

[0071] Corresponding to the soil heavy metal pollution cause accurate identification method provided by the above embodiment, based on the same technical concept, the embodiment of the present application also provides a soil heavy metal pollution cause accurate identification system, which is used to execute the soil heavy metal pollution cause accurate identification method,Figure 3 This is a schematic diagram of another soil heavy metal pollution cause accurate identification system provided in one embodiment of the present invention, as shown below. Figure 3 As shown. Soil heavy metal pollution cause precise identification systems can vary significantly due to differences in configuration or performance. They may include one or more processors 301 and memory 302. Memory 302 stores computer programs that can run on processor 301. Processor 301 executes the programs stored in memory 302 to achieve the above... Figure 1 The various steps in the Chinese method embodiment. The memory 302 can be temporary or persistent storage. The application stored in the memory 302 may include one or more modules (not shown in the figures), each module may include a series of computer-executable instructions for the precise identification system of the causes of heavy metal pollution in soil.

[0072] Furthermore, the processor 301 can be configured to communicate with the memory 302 and execute a series of computer-executable instructions stored in the memory 302 on the soil heavy metal pollution cause precise identification system. The soil heavy metal pollution cause precise identification system may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more input / output interfaces 305, and one or more keyboards 306.

[0073] Specifically, in this embodiment, the soil heavy metal pollution cause accurate identification system includes a processor, a communication interface, a memory, and a communication bus; wherein, the processor, communication interface, and memory communicate with each other via the bus; the memory is used to store computer programs; the processor is used to execute the programs stored in the memory to achieve the above... Figure 1 The various steps in the method embodiments are the same as those in the above method embodiments, and have the same beneficial effects. To avoid repetition, the embodiments of the present invention will not be described again here.

[0074] It should be noted that the soil heavy metal pollution cause accurate identification system provided in this embodiment of the invention and the soil heavy metal pollution cause accurate identification method provided in this embodiment of the invention are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned soil heavy metal pollution cause accurate identification method, and has the same or similar beneficial effects. Repeated parts will not be repeated.

[0075] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0076] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the difference from other embodiments.

[0077] The embodiment of the present application also provides a computer readable storage medium, the computer readable medium stores one or more programs, when the one or more programs are executed by the electronic device comprising a plurality of application programs, the electronic device executes Figure 1 The method disclosed in the embodiment and the functions and advantages of the methods in the foregoing method embodiments are not repeated here.

[0078] The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disc or an optical disc, etc.

Claims

1. A method for accurately identifying the causes of heavy metal pollution in soil, characterized in that, include: The surface soil characteristics and deep soil characteristics of each sampling point in the soil to be tested are obtained, wherein the deep soil characteristics include heavy metal concentration data; The sampling points are sorted according to the direction of the river flow where each sampling point is located to obtain a sampling point sequence. Based on the surface soil characteristics, deep soil characteristics, and historical soil characteristics of each sampling point in the sampling point sequence, abnormal sampling points are identified from each sampling point. Based on the historical soil characteristics of the upstream abnormal sampling points in the sampling point sequence that are upstream of the current sampling point, the initial weight of the current sampling point is updated to obtain the updated weight; The surface soil characteristics, deep soil characteristics, and update weights of each sampling point are input into a pre-trained SOM model, and the causes of heavy metal pollution in the soil to be tested are identified using the pre-trained SOM model. The step of updating the initial weight of the current sampling point based on the historical soil characteristics of upstream abnormal sampling points in the sampling point sequence, which are upstream of the current sampling point, to obtain the updated weight includes: Determine the upstream abnormal sampling point that is upstream of the current sampling point from the abnormal sampling points according to the order in the sampling point sequence; Based on the historical heavy metal concentration variation curves in the historical soil characteristics of the upstream abnormal sampling points, the degree of soil state change of the upstream abnormal sampling points relative to the historical sampling time is determined. Based on the second number of upstream abnormal sampling points, the third number of abnormal sampling points, and the soil state change value, the first degree of influence of the current sampling point on the upstream abnormal sampling points is determined; Based on the first degree of influence, a second degree of influence of the current sampling point on the downstream sampling points of the current sampling point is determined; The initial weights of each sampling point are updated based on the second degree of influence of each sampling point to obtain the updated weights.

2. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 1, characterized in that, The step of determining abnormal sampling points from each of the sampling points based on the surface soil characteristics, deep soil characteristics, and historical soil characteristics of each sampling point in the sampling point sequence includes: The real-time soil similarity between each sampling point is determined based on the surface soil characteristics and deep soil characteristics of each sampling point in the sampling point sequence; The historical soil similarity between the sampling points is determined based on the surface soil characteristics and deep soil characteristics of the historical soil characteristics of each sampling point; Based on the historical soil similarity between each sampling point and the real-time soil similarity between each sampling point, the consistency of soil state between each sampling point and historical soil is determined. The degree of abnormality in soil condition consistency among the sampling points is determined by utilizing the soil condition consistency of each sampling point; Abnormal sampling points are determined from each of the sampling points based on the degree of abnormality.

3. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 2, characterized in that, The step of determining the real-time soil similarity between each sampling point based on the surface soil characteristics and the deep soil characteristics of each sampling point in the sampling point sequence includes: Calculate the first similarity of surface soil features between the current sampling point and other sampling points in the sampling point sequence, and the second similarity of deep soil features between the current sampling point and other sampling points in the sampling point sequence; The real-time soil similarity between the current sampling point and other sampling points in the sampling point sequence is determined based on the first similarity and the second similarity.

4. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 2, characterized in that, The determination of the degree of anomalousness in soil condition consistency among the sampling points by utilizing the soil condition consistency of each sampling point includes: Determine a first number of sampling points where the consistency of all soil conditions is greater than a first threshold. Based on the first quantity, the total number of all sampling points in the sampling point sequence, and the soil state consistency of sampling points greater than the first threshold, the degree of abnormality in soil state consistency among the sampling points is determined.

5. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 1, characterized in that, The step of determining the degree of soil state change of the upstream abnormal sampling point relative to the historical sampling time based on the historical heavy metal concentration change curve in the historical soil characteristics of the upstream abnormal sampling point includes: Determine the fourth quantity of heavy metal elements in the historical soil characteristics of the upstream abnormal sampling points; Determine the slope value of the historical heavy metal concentration change curve of the upstream abnormal sampling point at the historical sampling time, and the fluctuation amplitude of the historical heavy metal concentration change curve of the upstream abnormal sampling point; The degree of soil state change is determined based on the fourth quantity, the slope value, and the fluctuation amplitude.

6. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 1, characterized in that, The step of determining the degree of first influence of the upstream abnormal sampling points on the current sampling point based on the second number of upstream abnormal sampling points, the third number of abnormal sampling points, and the soil state change value includes: Determine a first ratio between the second quantity and the third quantity; The product of the first ratio and the soil state change value is determined as the first degree of influence.

7. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 1, characterized in that, The step of updating the initial weights of each sampling point based on the second degree of influence of each sampling point to obtain the updated weights includes: Determine a second ratio between the second influence degree of each sampling point and the sum of the second influence degrees of all sampling points; The product of the second ratio and the initial weight is determined as the updated weight.

8. The method for accurate identification of the causes of heavy metal pollution in soil according to claim 1, characterized in that, The step of inputting the surface soil characteristics, deep soil characteristics, and update weights of each sampling point into a pre-trained SOM model, and using the pre-trained SOM model to identify the causes of heavy metal pollution in the soil to be tested, includes: The surface soil characteristics and the deep soil characteristics are normalized and then combined to obtain the pollution index vector of each sampling point. The pollution index vector and the updated weights are input into a pre-trained SOM model to identify the causes of heavy metal pollution, and the U matrix of the soil to be tested is output. The cause of heavy metal pollution in the soil to be tested is determined based on the color level changes of the U matrix.

9. A precise identification system for the causes of heavy metal pollution in soil, characterized in that, include: Processor and memory; wherein the memory is used to store computer programs that can run on the processor; A processor is used to execute a program stored in memory to implement the steps of the method for accurate identification of the causes of heavy metal pollution in soil as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Water environment pollution analysis and management method and system based on big data

    CN117333772A

  • Watershed pollution prediction method based on multi-source data fusion

    CN120296576A