A soil pollutant type identification method and system

By acquiring multi-source soil data, quantifying uncertainty, and using Dempster-Shafer evidence theory for weighted fusion, the problem of pollutant type assessment bias in complex soil pollution scenarios is solved, improving identification efficiency and accuracy.

CN120781179BActive Publication Date: 2025-11-21ZHEJIANG HUIYU ENVIRONMENTAL ENG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511288686.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-21
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively address complex soil pollution scenarios, including compound pollution, pollutant speciation transformation, and data uncertainty, leading to biases in pollutant type assessment.

Method used

By acquiring multi-source soil data, quantifying the uncertainty of each data type, extracting key features related to soil pollutant types, speciation transformation, and soil physicochemical properties, and using Dempster-Shafer evidence theory for weighted fusion, soil pollutant types are identified.

Benefits of technology

It improves the efficiency of soil pollutant type identification, ensures the accuracy and reliability of identification results, effectively addresses compound pollution and pollutant form transformation, and reduces the impact of data uncertainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120781179B_ABST
    Figure CN120781179B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of soil pollutant identification, and discloses a soil pollutant type identification method and system, which comprises the following steps: acquiring multi-source soil data of a soil to be measured; determining the uncertainty corresponding to each data type according to the data type of the multi-source soil data, obtaining the confidence degree corresponding to each data type; extracting key features related to soil pollutant types, form transformation and soil physical and chemical properties from the multi-source soil data, obtaining multi-source key feature data; and identifying the pollutant type identification result of the soil to be measured based on the confidence degree and the multi-source key feature data. The pollutant type identification result of the soil to be measured is identified through the confidence degree corresponding to the uncertainty of each data type in the multi-source soil data and the multi-source key feature data related to soil pollutant types, form transformation and soil physical and chemical properties, thereby improving the identification efficiency of soil pollutant types.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of soil pollutant identification, and in particular, to a soil pollutant type identification method and system. BACKGROUND

[0002] Soil pollution is one of the serious environmental problems currently faced by the world, especially in areas where there has been heavy industrial activity or long-term pollution accumulation, there are often complex and diverse pollutants left in the soil. In ecological restoration or land use planning, it is crucial to accurately identify the type, form and spatial distribution of pollutants in the soil. However, existing technologies face many challenges in accurately obtaining this information.

[0003] Firstly, soil pollution often manifests as complex pollution, i.e. multiple pollutants coexist, and there may be complex interactions between these pollutants, such as synergistic or antagonistic effects. For example, the presence of certain heavy metals may affect the degradation process of organic pollutants, and vice versa, which makes the identification and risk assessment of single pollutants complex. Secondly, pollutants in the soil are not constant, they may undergo morphological transformation, for example, a highly toxic pollutant may be transformed into a less toxic form under certain conditions, or vice versa. Different data sources have different abilities to identify these morphological transformations, for example, high spectral remote sensing and ground sensor data may provide macroscopic information, while laboratory precise analysis data can provide detailed morphological information, but the real-time performance is poor. How to effectively fuse these heterogeneous data and accurately identify the morphological transformation of pollutants is a problem that needs to be solved urgently.

[0004] In addition, in actual soil data collection, various uncertainties exist. These uncertainties may come from data collection methods, instrument accuracy, sampling errors, limitations of analysis methods, and environmental noise interference, etc. If these uncertainties are not effectively quantified and processed, they will directly affect the accuracy and reliability of the pollutant identification results, and may lead to bias in the assessment of pollutant types, and thus affect the subsequent development of remediation strategies and decision-making.

[0005] Therefore, in order to solve the technical problem of bias in the assessment of pollutant types due to the inability to effectively deal with complex pollution, pollutant morphological transformation and data uncertainty in complex soil pollution scenarios, a soil pollutant type identification method and system are urgently needed. SUMMARY

[0006] The purpose of the present application is to provide a soil pollutant type identification method and system, which solves the problem of deviation in pollutant type evaluation caused by the inability to effectively respond to complex soil pollution scenarios, composite pollution, pollutant morphological transformation and data uncertainty, and can analyze and fuse the extracted multi-source key feature data under the premise of considering data reliability, thereby improving the identification efficiency of soil pollutant types.

[0007] In a first aspect, the present application provides a soil pollutant type identification method, comprising:

[0008] obtaining multi-source soil data of a soil to be tested;

[0009] determining the uncertainty corresponding to each data type according to the data type of the multi-source soil data, to obtain the confidence of each data type;

[0010] extracting key features related to soil pollutant types, morphological transformation and soil physicochemical properties from the multi-source soil data to obtain multi-source key feature data;

[0011] identifying the pollutant type identification result of the soil to be tested based on the confidence and the multi-source key feature data.

[0012] The soil pollutant type identification method provided by the present application can identify soil pollutant types, can analyze and fuse the extracted multi-source key feature data under the premise of considering data reliability, and improves the identification efficiency of soil pollutant types.

[0013] Optionally, obtaining multi-source soil data of a soil to be tested comprises:

[0014] obtaining initial multi-source soil data of a soil to be tested;

[0015] spatiotemporally aligning the initial multi-source soil data to obtain the multi-source soil data of the soil to be tested.

[0016] Optionally, the multi-source soil data comprises hyperspectral remote sensing data, ground real-time sensor data and laboratory precise analysis data.

[0017] Optionally, determining the uncertainty corresponding to each data type according to the data type of the multi-source soil data to obtain the confidence of each data type comprises:

[0018] quantifying the uncertainty of each data type in the multi-source soil data according to the collection method, instrument accuracy, sampling error and environmental interference of the multi-source soil data to obtain initial uncertainty parameters of each data type;

[0019] According to the correlation between each data type in the multi-source soil data, the initial uncertainty parameter is adjusted to obtain the uncertainty parameter of each data type.

[0020] Based on the preset confidence mapping rule, the confidence corresponding to the uncertainty parameter is determined.

[0021] Optionally, according to the correlation between each data type in the multi-source soil data, the initial uncertainty parameter is adjusted to obtain the uncertainty parameter of each data type, including:

[0022] The correlation between each data type in the multi-source soil data is identified;

[0023] Based on the correlation, the influence path and influence weight of the uncertainty of each data type are determined;

[0024] Based on the influence path and the influence weight, the uncertainty of each data type in the multi-source soil data is iteratively adjusted to obtain the uncertainty parameter of each data type.

[0025] Optionally, the key features related to soil pollutant types, form transformation and soil physicochemical properties are extracted from the multi-source soil data to obtain multi-source key feature data, including:

[0026] Multi-dimensional features related to soil pollutant types, form transformation and soil physicochemical properties are extracted from the multi-source soil data to obtain multi-dimensional feature data;

[0027] According to the multi-dimensional feature data, the correlation evidence reflecting the synergistic or antagonistic effect between pollutants is constructed, and the specific evidence indicating the form transformation of pollutants is constructed to obtain multi-source key feature data.

[0028] Optionally, based on the confidence and the multi-source key feature data, the pollutant type identification result of the soil to be tested is identified, including:

[0029] According to the confidence, a corresponding weight coefficient is set;

[0030] According to the multi-source key feature data and the weight coefficient, weighted fusion is performed to obtain the pollutant type identification result of the soil to be tested.

[0031] Optionally, according to the multi-source key feature data and the weight coefficient, weighted fusion is performed to obtain the pollutant type identification result of the soil to be tested, including:

[0032] The multi-source key feature data and the weight coefficient are fused by using a fusion framework based on Dempster-Shafer evidence theory to obtain preliminary pollutant type identification information.

[0033] It is determined whether there is identification type conflict information or type identification ambiguous information in the preliminary pollutant type identification information. If not, the preliminary pollutant type identification information is determined as the pollutant type identification result of the soil to be measured. If yes, the preliminary pollutant type identification information is corrected or prioritized according to a preset discrimination rule in combination with the identification type conflict information or type identification ambiguous information and corresponding multi-source key feature data to obtain the pollutant type identification result of the soil to be measured.

[0034] Optionally, the multi-source key feature data and the weight coefficient are fused by using a fusion framework based on Dempster-Shafer evidence theory to obtain preliminary pollutant type identification information, including:

[0035] An identification framework containing all potential soil pollutant types is defined;

[0036] The multi-source key feature data is converted into basic probability assignments for each pollutant type in the identification framework according to the weight coefficient;

[0037] The basic probability assignments are fused by using a Dempster combination rule to obtain combined evidence for each potential pollutant type;

[0038] Preliminary pollutant type identification information for each potential pollutant type is generated according to the combined evidence.

[0039] In a second aspect, the present application provides a soil pollutant type identification system, including:

[0040] An acquisition module is configured to acquire multi-source soil data of soil to be measured;

[0041] A determination module is configured to determine the uncertainty corresponding to each data type according to the data type of the multi-source soil data to obtain the confidence degree corresponding to each data type;

[0042] An extraction module is configured to extract key features related to soil pollutant types, form transformation and soil physicochemical properties from the multi-source soil data to obtain multi-source key feature data;

[0043] An identification module is configured to identify the pollutant type identification result of the soil to be measured based on the confidence degree and the multi-source key feature data.

[0044] The soil pollutant type identification system can analyze and fuse the extracted multi-source key feature data under the premise of considering data reliability, and improve the identification efficiency of the soil pollutant type.

[0045] Beneficial effects: The soil pollutant type identification method and system provided by the application can identify the pollutant type identification result of the to-be-tested soil by the confidence degree corresponding to the uncertainty of each data type in the multi-source soil data and the multi-source key feature data related to the soil pollutant type, the form transformation and the soil physical and chemical properties in the multi-source soil data, solve the problem that in a complex soil pollution scene, due to the inability to effectively deal with composite pollution, pollutant form transformation and data uncertainty, there is deviation in the pollutant type evaluation, can analyze and fuse the extracted multi-source key feature data under the premise of considering data reliability, and improve the identification efficiency of the soil pollutant type. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 The flowchart of the soil pollutant type identification method provided by the embodiments of the application.

[0047] Figure 2 The structural schematic diagram of the soil pollutant type identification system provided by the embodiments of the application.

[0048] Label explanation: 1, acquisition module; 2, determination module; 3, extraction module; 4, identification module. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. The components of the embodiments of the application described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the application provided in the drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the application. Based on the embodiments of the application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the application.

[0050] It should be noted that: similar labels and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the application, the terms "first", "second" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0051] Please refer to Figure 1 , Figure 1is a soil pollutant type identification method in some embodiments of the present application, used for identifying a soil pollutant type, comprising the steps of:

[0052] In step S101, multi-source soil data of the soil to be tested is obtained.

[0053] In step S102, according to the data types of the multi-source soil data, the uncertainty corresponding to each data type is determined to obtain the confidence of each data type.

[0054] In step S103, key features related to soil pollutant types, form transformation and soil physicochemical properties are extracted from the multi-source soil data to obtain multi-source key feature data.

[0055] In step S104, based on the confidence and the multi-source key feature data, the pollutant type identification result of the soil to be tested is obtained.

[0056] The soil pollutant type identification method solves the problem of deviation in pollutant type evaluation caused by the inability to effectively deal with complex soil pollution scenarios, such as composite pollution, pollutant form transformation and data uncertainty, and can analyze and fuse the extracted multi-source key feature data under the premise of considering data reliability, thereby improving the identification efficiency of soil pollutant types.

[0057] Specifically, in step S101, multi-source soil data of the soil to be tested is obtained. The multi-source soil data refers to a set of soil-related information obtained from different sources or by different means. The multi-source soil data includes hyperspectral remote sensing data, ground real-time sensor data and laboratory precise analysis data. The hyperspectral remote sensing data can be obtained by a UAV carrying a hyperspectral imager. The ground real-time sensor data can be obtained by real-time monitoring equipment such as soil pH sensors and heavy metal ion sensors laid on the site. The laboratory precise analysis data can be obtained by precise chemical analysis such as ICP-MS detection of heavy metals and GC-MS analysis of organic matter on collected soil samples.

[0058] Further, hyperspectral remote sensing data can provide large-scale, non-contact soil information. By analyzing the soil reflectance spectrum characteristics, the types and spatial distribution of pollutants in the soil can be preliminarily identified. This provides a macro background and an indication of key areas for subsequent detailed analysis, making up for the lack of spatial coverage of traditional point sampling, and is particularly suitable for initial screening and regional monitoring. At the same time, ground real-time sensor data can provide continuous and real-time monitoring information at specific points, capturing the dynamic changes and real-time response of pollutants, providing immediate evidence for the transformation and migration process of pollutants, making up for the lack of time resolution and depth information of remote sensing data, as well as the timeliness of laboratory analysis. Laboratory accurate analysis data can provide accurate and reliable quantitative information on the types, forms, and concentrations of pollutants through precise chemical analysis of soil samples.

[0059] Specifically, in step S101, multi-source soil data of the soil to be measured is obtained, including:

[0060] Obtaining initial multi-source soil data of the soil to be measured;

[0061] Performing spatio-temporal alignment on the initial multi-source soil data to obtain the multi-source soil data of the soil to be measured.

[0062] In step S101, initial multi-source soil data in the original state is collected. The initial multi-source soil data refers to original, uniformly processed soil-related data obtained from different sources, different collection devices, different time points, or different spatial locations, i.e., the original data of hyperspectral remote sensing data, ground real-time sensor data, and laboratory accurate analysis data. These original data may have inconsistent time stamps, geographic coordinate deviations, data format differences, and other problems. In order to ensure that these heterogeneous data can be effectively integrated and utilized, the initial multi-source soil data is subjected to spatio-temporal alignment processing to eliminate inconsistencies in the original data and obtain high-quality, consistent multi-source soil data.

[0063] Specifically, in step S102, according to the data types of the multi-source soil data, the uncertainty corresponding to each data type is determined to obtain the confidence corresponding to each data type, including:

[0064] According to the collection method, instrument accuracy, sampling error, and environmental interference of the multi-source soil data, the uncertainty of each data type in the multi-source soil data is quantified to obtain initial uncertainty parameters of each data type;

[0065] According to the correlation between each data type in the multi-source soil data, the initial uncertainty parameters are adjusted to obtain uncertainty parameters of each data type;

[0066] Based on the preset confidence mapping rule, the confidence corresponding to the uncertainty parameters is determined.

[0067] In step S102, based on factors such as the acquisition method of multi-source soil data, instrument accuracy, sampling error, and environmental interference, the inherent sources of uncertainty in the data acquisition process are systematically evaluated and quantified. This can be achieved using existing technologies such as statistical methods, error propagation models, fuzzy mathematics theory, or evaluation models based on expert experience.

[0068] Specifically, the quantification of uncertainties in hyperspectral remote sensing data includes:

[0069] (1) Regarding the data acquisition method: When planning the UAV flight mission, a ground sampling distance (GSD) threshold is set. If the actual flight altitude or sensor parameters cause the GSD to exceed the threshold, a spatial resolution uncertainty factor is introduced.

[0070] (2) Regarding instrument accuracy: Continuously monitor the signal-to-noise ratio (SNR) of the hyperspectral imager. If the SNR is lower than the preset threshold, increase the measurement noise uncertainty factor.

[0071] (3) Regarding environmental interference: When performing atmospheric correction, the atmospheric water vapor content and aerosol optical thickness are obtained. If the water vapor content or aerosol optical thickness exceeds the preset range, the atmospheric interference uncertainty factor is increased.

[0072] By weighted summation or multiplication of these factors, the initial uncertainty of the ability of hyperspectral remote sensing data to indicate specific pollutants is obtained.

[0073] Uncertainty quantification for real-time ground sensor data includes:

[0074] (1) Regarding the data acquisition method: When deploying the sensor array, record the sensor point density. If the actual density is lower than the preset spatial representativeness requirement, then introduce a spatial representativeness uncertainty factor.

[0075] (2) Regarding instrument accuracy: Record the sensor calibration frequency and calibration results. If the sensor calibration cycle exceeds the preset time or the calibration results show drift, increase the calibration drift uncertainty factor.

[0076] (3) Regarding environmental interference: Real-time monitoring of soil temperature and humidity at the sensor location. If the temperature or humidity exceeds the sensor's operating range, the uncertainty factor of environmental interference will increase.

[0077] By weighted summation or multiplication of these factors, the initial uncertainty of how ground-based real-time sensor data reflects soil physicochemical parameters in real time is obtained.

[0078] Quantification of uncertainty in precise laboratory analytical data, including:

[0079] (1) Regarding the collection method: Record the sampling points, depth, and number of soil samples. If the sampling points are sparsely distributed or do not cover key areas, introduce a sampling representativeness uncertainty factor.

[0080] (2) Regarding instrument accuracy: During the analysis, record the detection limit and recovery rate of the analytical method. If the detection limit is higher than the pollutant concentration or the recovery rate exceeds the acceptable range, the uncertainty factor of analytical accuracy will be increased.

[0081] (3) Regarding sampling error: Analyze the repeatedly collected samples and calculate the relative standard deviation of their measurement results. If the relative standard deviation exceeds the preset threshold, increase the sampling error factor.

[0082] By weighted summation or product of these factors, the initial uncertainty of the laboratory's precise analytical data for accurate identification of pollutant concentration and speciation is obtained.

[0083] Specifically, in step S102, the initial uncertainty parameters are adjusted based on the correlation between different data types in the multi-source soil data to obtain the uncertainty parameters for each data type, including:

[0084] Identify the relationships between different data types in multi-source soil data;

[0085] Based on the correlation, determine the impact path and impact weight of uncertainty for each data type;

[0086] Based on the influence path and influence weight, the uncertainty of each data type in multi-source soil data is iteratively adjusted to obtain the uncertainty parameters of each data type.

[0087] In step S102, identifying potential calibration dependencies, validation relationships, or complementarities among hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, and recognizing the correlations between different data types in multi-source soil data, lays the foundation for precise uncertainty adjustment. For example, it can be identified that real-time ground sensor data may be used to calibrate hyperspectral remote sensing data, while precise laboratory analysis data can serve as truth verification for both. Furthermore, hyperspectral remote sensing data and real-time ground sensor data may exhibit spatial correlations in certain pollutant characteristics. This explicit identification avoids blind or empirical adjustments to uncertainty parameters, ensuring the targeted nature of subsequent adjustments.

[0088] For example, hyperspectral remote sensing data can be used to extract reflectance characteristics in specific bands, real-time ground sensor data can be used to extract real-time parameters such as conductivity and pH value, and precise laboratory analysis data can be used to extract precise physicochemical properties such as heavy metal content and organic matter content. Subsequently, Pearson correlation coefficients can be used to analyze the linear correlations between these features, or mutual information methods can be used to analyze nonlinear correlations. For instance, it was found that certain spectral features in hyperspectral remote sensing data are strongly correlated with conductivity in real-time ground sensor data, while the conductivity in real-time ground sensor data is correlated with heavy metal content in precise laboratory analysis data.

[0089] Building upon this foundation, the impact paths and weights of uncertainties across different data types were further determined based on the identified correlations. This not only clarified the specific pathways through which uncertainties are transmitted and diffused between different data types—for example, how uncertainties in real-time ground sensor data affect uncertainties in hyperspectral remote sensing data—but also quantified the intensity of these impacts.

[0090] For example, the interactions between hyperspectral remote sensing data and real-time ground sensor data, as well as the interactions between hyperspectral remote sensing data and precise laboratory analysis data, can be analyzed to determine the impact paths and weights of uncertainty. This can be achieved through sensitivity analysis of historical data. Alternatively, an uncertainty propagation graph can be constructed, where nodes represent the uncertainty of each data type and edges represent impact paths. The corresponding impact paths are determined based on the linear or nonlinear correlations between hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, and the corresponding impact weights are determined based on the strength of the linear or nonlinear correlations. The stronger the correlation, the greater the impact weight.

[0091] Based on the established influence paths and weights, the uncertainties of each data type in multi-source soil data are iteratively adjusted to obtain uncertainty parameters for each data type. This is a dynamic optimization process. Using the quantified influence paths and weights, the uncertainty parameters of each data type are repeatedly corrected. In each iteration, the uncertainty of one data type is adjusted according to its influence from the uncertainties of other data types, and the adjusted uncertainty, in turn, affects other data types. This process continues until the changes in all uncertainty parameters are less than a preset threshold, reaching a convergence state. Therefore, the obtained uncertainty parameters more realistically reflect the complexity and intrinsic connections of multi-source data, providing a solid foundation for subsequent confidence calculations based on these uncertainty parameters, thus ensuring the accuracy and reliability of the final pollutant identification results.

[0092] For example, suppose the initial uncertainty parameters are: 0.15 for hyperspectral remote sensing data, 0.10 for real-time ground sensor data, and 0.05 for precise laboratory analysis data. In the first iteration, the uncertainty of the real-time ground sensor data is adjusted based on the uncertainty of the hyperspectral remote sensing data and its influence weight. For example, real-time ground sensor uncertainty = initial real-time ground sensor uncertainty + 0.6 * hyperspectral remote sensing uncertainty. Simultaneously, the uncertainty of the precise laboratory analysis data is also adjusted based on the uncertainty of the real-time ground sensor data and its influence weight. In subsequent iterations, these adjusted uncertainty parameters are used as new inputs to recalculate and correct the uncertainties of other data types until the change in all uncertainty parameters is less than a preset threshold (e.g., 0.001), reaching convergence. For example, after multiple iterations, the uncertainty of the hyperspectral remote sensing data might be adjusted to 0.12, the uncertainty of the real-time ground sensor data to 0.08, and the uncertainty of the precise laboratory analysis data to 0.06. These adjusted parameters more accurately reflect the mutual influence between data and the true propagation of uncertainty.

[0093] Specifically, in step S102, after obtaining the adjusted uncertainty parameter, it can be converted into a confidence level based on a preset confidence level mapping rule. For example, a sigmoid function can be defined to map the uncertainty parameter to a confidence level, such that the lower the uncertainty, the higher the corresponding confidence level, and the confidence level changes more gradually when the uncertainty is low, while it decreases more rapidly when the uncertainty is high. This ensures that the assessment of data quality is both sensitive and stable, providing a reliable basis for subsequent pollutant identification.

[0094] The preset confidence mapping rule refers to converting the quantified uncertainty parameters into an intuitive confidence metric that can be used for data fusion and decision-making. It can be implemented using existing techniques such as linear mapping functions, nonlinear mapping functions (such as the Sigmoid function), lookup tables, or piecewise functions.

[0095] Specifically, in step S103, key features related to soil pollutant types, speciation, and soil physicochemical properties are extracted from multi-source soil data to obtain multi-source key feature data, including:

[0096] Multidimensional feature data is obtained by extracting multidimensional features related to soil pollutant types, form transformations, and soil physicochemical properties from multi-source soil data.

[0097] Based on multi-dimensional feature data, we construct association evidence reflecting the synergistic or antagonistic effects between pollutants, and construct specific evidence indicating the transformation of pollutant forms, thus obtaining multi-source key feature data.

[0098] In step S103, multi-dimensional features related to soil pollutant types, speciation, and soil physicochemical properties are obtained from multi-source soil data, thus yielding multi-dimensional feature data. This initial step is fundamental and crucial, ensuring that information directly related to pollutant identification is comprehensively acquired from multiple sources, including hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data.

[0099] For example, reflectance, absorption characteristics, and vegetation indices in specific bands can be extracted from hyperspectral remote sensing data as multi-dimensional features, which can indicate the presence of heavy metals or organic matter in the soil. Soil parameters such as pH, redox potential, conductivity, and temperature can be obtained from real-time ground sensor data; these parameters are closely related to the migration and transformation of pollutants. Information such as the total content, valence state distribution, and organic matter content of specific elements in the soil can be obtained from precise laboratory analysis data. By acquiring these data, a multi-dimensional feature dataset can be constructed.

[0100] Based on multi-dimensional characteristic data, further association evidence reflecting synergistic or antagonistic effects between pollutants, as well as specific evidence indicating pollutant speciation, are constructed, ultimately yielding multi-source key characteristic data. This step further deepens and refines the initially acquired multi-dimensional characteristic data, transforming it into truly insightful key characteristic data. By constructing association evidence reflecting synergistic or antagonistic effects between pollutants, complex interactions that may arise when multiple pollutants coexist can be identified and quantified. For example, certain heavy metals may inhibit the degradation of organic pollutants, or organic matter may affect the migration of heavy metals. The construction of this association evidence transcends the limitations of single-pollutant analysis, providing a deeper understanding of the intrinsic mechanisms of compound pollution, thereby enabling more accurate assessment of pollution risks and identification of pollutant types. Simultaneously, by constructing specific evidence indicating pollutant speciation, the speciation changes of pollutants in the soil environment can be tracked and identified. For example, highly toxic hexavalent chromium may transform into less toxic trivalent chromium. The construction of this specific evidence distinguishes different speciations of pollutants, which is crucial for accurately assessing the bioavailability and toxicity of pollutants, as different speciations of pollutants exhibit drastically different environmental behaviors and ecological risks.

[0101] Among these, constructing evidence reflecting the synergistic or antagonistic effects between pollutants refers to identifying and quantifying evidence of the potential mutually reinforcing or mutually inhibiting effects that may occur when multiple pollutants coexist by analyzing multi-dimensional characteristic data. Constructing specific evidence indicating pollutant speciation refers to identifying and quantifying evidence of chemical or physical speciation changes of pollutants in the soil environment by analyzing multi-dimensional characteristic data.

[0102] Specifically, the correlation between heavy metal indicators and organic matter content in hyperspectral data can be analyzed to construct evidence reflecting synergistic or antagonistic effects between pollutants. For example, if a negative correlation is found between the content of a specific heavy metal and the degradation products of soil organic matter, this may indicate an inhibitory effect of heavy metals on organic matter degradation. This association can be quantified by establishing a multiple regression model or using methods based on graph neural networks.

[0103] Changes in the absorption peak shifts or intensities of specific elements in hyperspectral data can be used to preliminarily determine their valence state changes, constructing specific evidence indicating speciation transformation of pollutants. For example, hexavalent chromium and trivalent chromium exhibit different absorption characteristics in specific spectral regions. Combining this with precise determinations of chromium's valence state in laboratory data can further confirm and quantify this speciation transformation. For instance, if hyperspectral data shows a weakening of the characteristics of hexavalent chromium and an enhancement of the characteristics of trivalent chromium, and laboratory analysis also confirms this transformation, then this information constitutes specific evidence indicating speciation transformation.

[0104] Specifically, in step S104, based on confidence level and multi-source key feature data, the pollutant type identification results of the soil to be tested are obtained, including:

[0105] Set the corresponding weight coefficients based on the confidence level;

[0106] Based on multi-source key feature data and weighting coefficients, weighted fusion is performed to obtain the pollutant type identification results of the soil to be tested.

[0107] In step S104, key feature data with high confidence will have their corresponding weight coefficients set to higher values, while data with low confidence will be assigned lower weights. This weight coefficient setting enables intelligent differentiation and treatment of multi-source key feature data of different quality and reliability, thereby effectively avoiding the undue negative impact of low-quality or highly uncertain data on the final identification results and laying the foundation for subsequent accurate identification.

[0108] For example, if the confidence level is a value between 0 and 1, the weighting coefficient can be directly set to be proportional to that value, or it can be converted into an integer weight between 0 and 100 through a preset mapping function (e.g., a sigmoid function or a piecewise linear function). For example, for laboratory precision analysis data with an initial confidence level of 0.9, the weighting coefficient can be set to 90; for ground real-time sensor data with an initial confidence level of 0.6, the weighting coefficient can be set to 60.

[0109] Optionally, a conventional weighted average method can be used to weight and fuse multi-source key feature data using the weight coefficients converted from confidence levels. For example, suppose there are three types of pollutants to be identified: A, B, and C. Key feature data related to these three pollutants are extracted from hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, and their support for the presence of each pollutant is calculated. During weighted fusion, the support of each data source for a specific pollutant type can be multiplied by its corresponding weight coefficient. Then, all weighted support values ​​are summed and divided by the sum of all weight coefficients to obtain the final identification support for that pollutant type. The pollutant types with the highest support ranking or greater than a preset support threshold are identified as the pollutant type identification results for the soil to be tested.

[0110] Preferably, in step S104, a weighted fusion is performed based on multi-source key feature data and weighting coefficients to obtain the pollutant type identification result of the soil to be tested, including:

[0111] A fusion framework based on Dempster-Shafer evidence theory is adopted to perform weighted fusion of multi-source key feature data and weight coefficients to obtain preliminary pollutant type identification information;

[0112] Determine whether there is any conflicting or ambiguous information in the preliminary pollutant type identification information; if not, determine the preliminary pollutant type identification information as the pollutant type identification result of the soil to be tested; if so, based on the preset discrimination rules, combined with the conflicting or ambiguous information and the corresponding multi-source key feature data, correct or prioritize the preliminary pollutant type identification information to obtain the pollutant type identification result of the soil to be tested.

[0113] Specifically, in step S104, a fusion framework based on Dempster-Shafer evidence theory is used to perform weighted fusion of multi-source key feature data and weight coefficients to obtain preliminary pollutant type identification information, including:

[0114] Define an identification framework that includes all potential soil pollutant types;

[0115] Based on the weighting coefficients, the multi-source key feature data are converted into basic probability assignments for each pollutant type in the identification framework;

[0116] By employing the Dempster combinatorial rule and fusing the basic probability assignments, combined evidence is obtained for each potential pollutant type.

[0117] Based on the combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type.

[0118] It should be noted that the fusion framework based on Dempster-Shafer evidence theory is a mathematical tool for handling uncertainty and multi-source information fusion. It can be implemented by defining the identification framework, calculating the basic probability allocation, applying the combination rules, and generating preliminary identification information.

[0119] In step S104, an identification framework encompassing all potential soil pollutant types (e.g., heavy metals, organochlorine pesticides, petroleum hydrocarbons, etc.) is defined, clarifying the scope of pollutant identification and all possible targets. Based on this, multi-source key feature data is converted into basic probability assignments for each pollutant type within the identification framework, according to weighting coefficients. These weighting coefficients quantify the reliability of the data sources, ensuring that the support level of each data type for a specific pollutant type is accurately mapped to the basic probability assignment. This provides input with confidence information for subsequent evidence fusion, effectively handling data uncertainty and ensuring that the contributions of different data sources are reasonably weighted. Next, the Dempster combination rule is used to fuse the basic probability assignments, obtaining combined evidence for each potential pollutant type. The Dempster combination rule effectively aggregates information from multiple independent evidence sources. Even if these sources contain some degree of conflict or uncertainty, the rule yields a comprehensive and more persuasive combined evidence, enhancing the reliability and robustness of the identification results and providing a more solid foundation for subsequent decision-making. Finally, based on the combined evidence, preliminary pollutant type identification information for each potential pollutant type is generated. This step forms the basis for the final decision. It leverages the combined power of all available evidence, providing a solid foundation for subsequent conflict resolution or handling of ambiguous information. It ensures that the initial identification results are based on comprehensive and weighted multi-source data, improving the accuracy and credibility of the identification. This fusion approach, employing a framework based on Dempster-Shafer evidence theory, goes beyond simple weighted averaging, more comprehensively reflecting the accumulation of evidence and the propagation of uncertainty, thus laying the foundation for subsequent judgments.

[0120] For example, firstly, an identification framework is defined, which may include potential pollutant types such as "lead (Pb)," "cadmium (Cd)," "arsenic (As)," "hexavalent chromium (Cr(VI))," and "polychlorinated biphenyls (PCBs)." Then, based on multi-source key feature data extracted from hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, along with the corresponding weighting coefficients for these data types, this data is converted into basic probability assignments. For instance, if hyperspectral data indicates potential heavy metal pollution in the soil and its weighting coefficient is high, it can be converted into a basic probability assignment that assigns high confidence to heavy metal types such as "lead," "cadmium," and "arsenic," while assigning lower confidence to organic pollutants such as "PCBs." Simultaneously, if laboratory analysis data precisely determines the presence of hexavalent chromium and its weighting coefficient is extremely high, a basic probability assignment can be generated that assigns extremely high confidence to the "hexavalent chromium" type. Next, the Dempster combination rule is used to fuse these basic probability assignments from different data sources. For example, the basic probability assignments generated from hyperspectral data are combined with those generated from laboratory analysis data. If the hyperspectral data has a confidence level of 0.6 for lead, while the laboratory data has a confidence level of 0.3, Dempster's combination rule will consider both pieces of evidence to generate a combined confidence level for lead that is either higher or lower, while also addressing potential conflicts or uncertainties. Finally, based on the fused combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type. For example, if the combined evidence shows the highest confidence level for hexavalent chromium, significantly higher than other pollutant types, the preliminary pollutant type identification information could indicate that the primary pollutant type in the tested soil is hexavalent chromium. If the combined evidence shows high and similar confidence levels for both lead and cadmium, the preliminary identification information could indicate the presence of combined pollution and provide the relative confidence levels for the two pollutants.

[0121] Further, in step S104, it is determined whether there is conflicting or ambiguous information regarding pollutant type identification in the preliminary pollutant type identification information. This step introduces a self-diagnostic capability for the fusion results. In practical applications, even with fusion theory, the preliminary identification results may still be contradictory. For example, different pieces of evidence may strongly point to different pollutants, or the information may be ambiguous, such as insufficient evidence to clearly point to a single pollutant, but rather to a set of pollutants. By explicitly identifying these conflicts or ambiguities, unreliable results can be avoided, providing a clear basis for subsequent corrections. Conflicting pollutant type identification information refers to contradictory pollutant type judgments from different data sources or evidence during the preliminary pollutant type identification process. This can be identified by analyzing inconsistencies between evidence or calculating conflict coefficients. Ambiguous information regarding pollutant type identification refers to situations where, during the preliminary pollutant type identification process, evidence is insufficient to clearly point to a single pollutant type, but rather to a set of pollutant types, or the support for any single pollutant type is low or the uncertainty range is too large. This can be identified by assessing the uncertainty of the belief function or analyzing the diffusion of evidence.

[0122] For example, a conflict metric between pieces of evidence can be calculated, such as the conflict coefficient in the Dempster-Shafer theory. If this coefficient exceeds a preset threshold, it is determined that there is conflicting information regarding the identification type. Simultaneously, the uncertainty of the belief function can be assessed; for instance, if the belief level regarding any single pollutant type is below a certain threshold, or if the uncertainty range is too large, it is determined that there is ambiguous information regarding type identification.

[0123] If the preliminary pollutant type identification information does not contain conflicting or ambiguous information, then this preliminary pollutant type identification information is directly determined as the pollutant type identification result for the soil to be tested. This indicates that when the fusion results are clear and consistent, a final judgment can be given efficiently, avoiding unnecessary complex processing and improving identification efficiency. For example, if the preliminary identification information clearly indicates that the soil is mainly contaminated with lead, and the confidence level is high, then it is directly identified as lead contamination.

[0124] If conflicting or ambiguous information regarding pollutant type identification exists, the preliminary pollutant type identification information is corrected or prioritized based on pre-defined discrimination rules, combined with the conflicting or ambiguous information and corresponding multi-source key feature data, ultimately yielding the pollutant type identification result for the soil being tested. This step provides a mechanism for handling complex situations. The pre-defined discrimination rules can be based on expert knowledge, historical data, or the needs of specific application scenarios to guide how to resolve conflicts, such as prioritizing evidence from specific high-confidence data sources, using majority voting mechanisms, or handling ambiguity, such as prioritizing the most likely pollutants or requesting supplementary data. By combining the original multi-source key feature data, the correction process can trace back to the root cause of the conflict or ambiguity, allowing for targeted adjustments, such as re-evaluating conflicting evidence or refining the allocation of ambiguous beliefs. This correction or prioritization mechanism ensures that the final pollutant type identification result is well-considered, more robust, and reliable.

[0125] For example, when conflicts arise, pre-defined rules can prioritize laboratory analysis data, as it typically has higher reliability; alternatively, a majority voting mechanism can be used, selecting the pollutant type supported by the majority of evidence. When ambiguous information exists, pre-defined rules can guide the prioritization of the most likely pollutant types, for example, identifying the two pollutant types with the highest belief levels (such as cadmium and arsenic) as the primary suspects and recommending further supplementary sampling or analysis. During the correction process, original multi-source key characteristic data can be incorporated; for example, re-examining whether a particular sensor data point causing the conflict is anomaly, or refining ambiguous belief assignments, for example, by introducing additional auxiliary features (such as the impact of soil pH on pollutant mobility) to further differentiate similar pollutant types. Through these corrections or prioritizations, an accurate and reliable identification result of the soil pollutant type is ultimately output.

[0126] As can be seen from the above, this method for identifying soil pollutant types can analyze and fuse the extracted multi-source key feature data while taking into account the reliability of the data, thereby improving the efficiency of soil pollutant type identification.

[0127] refer to Figure 2 This application provides a soil pollutant type identification system for identifying soil pollutant types, including:

[0128] Module 1 is used to acquire multi-source soil data of the soil to be tested;

[0129] Module 2 is used to determine the uncertainty corresponding to each data type based on the data type of multi-source soil data, and to obtain the confidence level corresponding to each data type.

[0130] Extraction module 3 is used to extract key features related to soil pollutant types, form transformations and soil physicochemical properties from multi-source soil data to obtain multi-source key feature data;

[0131] The identification module 4 is used to identify the pollutant type of the soil to be tested based on confidence level and multi-source key feature data.

[0132] This soil pollutant type identification system solves the problem of biased pollutant type assessment caused by the inability to effectively deal with compound pollution, pollutant form transformation and data uncertainty in complex soil pollution scenarios. It can analyze and fuse extracted multi-source key feature data while taking into account data reliability, thereby improving the efficiency of soil pollutant type identification.

[0133] Specifically, when module 1 is executed, it acquires multi-source soil data of the soil to be tested. Multi-source soil data refers to a collection of soil-related information obtained from different sources or through different methods. Multi-source soil data includes hyperspectral remote sensing data, ground real-time sensor data, and laboratory precise analysis data. Among them, hyperspectral remote sensing data can be acquired by a drone equipped with a hyperspectral imager, ground real-time sensor data can be acquired by real-time monitoring equipment such as soil pH sensors and heavy metal ion sensors deployed on site, and laboratory precise analysis data can be obtained by performing precise chemical analysis on the collected soil samples, such as inductively coupled plasma mass spectrometry (ICP-MS) to detect heavy metals and gas chromatography-mass spectrometry (GC-MS) to analyze organic matter.

[0134] Furthermore, hyperspectral remote sensing data can provide large-scale, non-contact soil information. By analyzing soil reflectance spectral characteristics, the types and spatial distribution of pollutants in the soil can be preliminarily identified. This provides a macroscopic background and key area indications for subsequent refined analysis, compensating for the spatial coverage limitations of traditional point sampling, and is particularly suitable for initial screening and regional monitoring. Simultaneously, real-time ground sensor data can provide continuous, real-time monitoring information for specific locations, capturing the dynamic changes and real-time responses of pollutants, providing immediate evidence of pollutant transformation and migration processes, and compensating for the limitations of remote sensing data in temporal resolution and depth information, as well as the time lag in laboratory analysis. Precise laboratory analysis data, through precise chemical analysis of soil samples, can provide accurate and reliable quantitative information on pollutant types, forms, and concentrations.

[0135] Specifically, when acquiring multi-source soil data of the soil to be tested, module 1 executes the following:

[0136] Obtain initial multi-source soil data for the soil to be tested;

[0137] Spatiotemporal alignment of the initial multi-source soil data was performed to obtain multi-source soil data of the soil to be tested.

[0138] During execution, module 1 collects raw, initial multi-source soil data. This raw, unprocessed soil-related data refers to data acquired from different sources, using different acquisition devices, at different times, or in different spatial locations. This includes raw data from hyperspectral remote sensing, real-time ground sensors, and precise laboratory analysis. Such raw data may contain inconsistencies in timestamps, geographic coordinate deviations, and data format differences. To ensure effective integration and utilization of this heterogeneous data, spatiotemporal alignment processing is performed on this raw multi-source soil data to eliminate inconsistencies and obtain high-quality, consistent multi-source soil data.

[0139] Specifically, when determining the uncertainty corresponding to each data type based on the data types of multi-source soil data, and obtaining the confidence level corresponding to each data type, module 2 executes the following:

[0140] Based on the acquisition method, instrument accuracy, sampling error and environmental interference of multi-source soil data, the uncertainty of each data type in multi-source soil data is quantified to obtain the initial uncertainty parameters of each data type.

[0141] Based on the correlation between different data types in multi-source soil data, the initial uncertainty parameters are adjusted to obtain the uncertainty parameters for each data type.

[0142] Based on the preset confidence mapping rules, the confidence level corresponding to the uncertainty parameter is determined.

[0143] When module 2 is executed, it systematically evaluates and quantifies the sources of uncertainty inherent in the data acquisition process based on factors such as the acquisition method of multi-source soil data, instrument accuracy, sampling error, and environmental interference. This can be achieved using existing technologies such as statistical methods, error propagation models, fuzzy mathematics theory, or evaluation models based on expert experience.

[0144] Specifically, the quantification of uncertainties in hyperspectral remote sensing data includes:

[0145] (1) Regarding the data acquisition method: When planning the UAV flight mission, a ground sampling distance (GSD) threshold is set. If the actual flight altitude or sensor parameters cause the GSD to exceed the threshold, a spatial resolution uncertainty factor is introduced.

[0146] (2) Regarding instrument accuracy: Continuously monitor the signal-to-noise ratio (SNR) of the hyperspectral imager. If the SNR is lower than the preset threshold, increase the measurement noise uncertainty factor.

[0147] (3) Regarding environmental interference: When performing atmospheric correction, the atmospheric water vapor content and aerosol optical thickness are obtained. If the water vapor content or aerosol optical thickness exceeds the preset range, the atmospheric interference uncertainty factor is increased.

[0148] By weighted summation or multiplication of these factors, the initial uncertainty of the ability of hyperspectral remote sensing data to indicate specific pollutants is obtained.

[0149] Uncertainty quantification for real-time ground sensor data includes:

[0150] (1) Regarding the data acquisition method: When deploying the sensor array, record the sensor point density. If the actual density is lower than the preset spatial representativeness requirement, then introduce a spatial representativeness uncertainty factor.

[0151] (2) Regarding instrument accuracy: Record the sensor calibration frequency and calibration results. If the sensor calibration cycle exceeds the preset time or the calibration results show drift, increase the calibration drift uncertainty factor.

[0152] (3) Regarding environmental interference: Real-time monitoring of soil temperature and humidity at the sensor location. If the temperature or humidity exceeds the sensor's operating range, the uncertainty factor of environmental interference will increase.

[0153] By weighted summation or multiplication of these factors, the initial uncertainty of how ground-based real-time sensor data reflects soil physicochemical parameters in real time is obtained.

[0154] Quantification of uncertainty in precise laboratory analytical data, including:

[0155] (1) Regarding the collection method: Record the sampling points, depth, and number of soil samples. If the sampling points are sparsely distributed or do not cover key areas, introduce a sampling representativeness uncertainty factor.

[0156] (2) Regarding instrument accuracy: During the analysis, record the detection limit and recovery rate of the analytical method. If the detection limit is higher than the pollutant concentration or the recovery rate exceeds the acceptable range, the uncertainty factor of analytical accuracy will be increased.

[0157] (3) Regarding sampling error: Analyze the repeatedly collected samples and calculate the relative standard deviation of their measurement results. If the relative standard deviation exceeds the preset threshold, increase the sampling error factor.

[0158] By weighted summation or product of these factors, the initial uncertainty of the laboratory's precise analytical data for accurate identification of pollutant concentration and speciation is obtained.

[0159] Specifically, when determining the uncertainty parameters of each data type by adjusting the initial uncertainty parameters based on the correlation between different data types in the multi-source soil data, module 2 executes the following:

[0160] Identify the relationships between different data types in multi-source soil data;

[0161] Based on the correlation, determine the impact path and impact weight of uncertainty for each data type;

[0162] Based on the influence path and influence weight, the uncertainty of each data type in multi-source soil data is iteratively adjusted to obtain the uncertainty parameters of each data type.

[0163] When Module 2 is executed, it identifies potential calibration dependencies, validation relationships, or complementarities among hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data. This process also identifies the correlations between different data types within multi-source soil data, laying the foundation for precise uncertainty adjustment. For example, it can be identified that real-time ground sensor data may be used to calibrate hyperspectral remote sensing data, while precise laboratory analysis data can serve as truth verification for both. Furthermore, it can be noted that hyperspectral remote sensing data and real-time ground sensor data may exhibit spatial correlations in certain pollutant characteristics. This explicit identification avoids blind or empirical adjustments to uncertainty parameters, ensuring the targeted nature of subsequent adjustments.

[0164] For example, hyperspectral remote sensing data can be used to extract reflectance characteristics in specific bands, real-time ground sensor data can be used to extract real-time parameters such as conductivity and pH value, and precise laboratory analysis data can be used to extract precise physicochemical properties such as heavy metal content and organic matter content. Subsequently, Pearson correlation coefficients can be used to analyze the linear correlations between these features, or mutual information methods can be used to analyze nonlinear correlations. For instance, it was found that certain spectral features in hyperspectral remote sensing data are strongly correlated with conductivity in real-time ground sensor data, while the conductivity in real-time ground sensor data is correlated with heavy metal content in precise laboratory analysis data.

[0165] Building upon this foundation, the impact paths and weights of uncertainties across different data types were further determined based on the identified correlations. This not only clarified the specific pathways through which uncertainties are transmitted and diffused between different data types—for example, how uncertainties in real-time ground sensor data affect uncertainties in hyperspectral remote sensing data—but also quantified the intensity of these impacts.

[0166] For example, the interactions between hyperspectral remote sensing data and real-time ground sensor data, as well as the interactions between hyperspectral remote sensing data and precise laboratory analysis data, can be analyzed to determine the impact paths and weights of uncertainty. This can be achieved through sensitivity analysis of historical data. Alternatively, an uncertainty propagation graph can be constructed, where nodes represent the uncertainty of each data type and edges represent impact paths. The corresponding impact paths are determined based on the linear or nonlinear correlations between hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, and the corresponding impact weights are determined based on the strength of the linear or nonlinear correlations. The stronger the correlation, the greater the impact weight.

[0167] Based on the established influence paths and weights, the uncertainties of each data type in multi-source soil data are iteratively adjusted to obtain uncertainty parameters for each data type. This is a dynamic optimization process. Using the quantified influence paths and weights, the uncertainty parameters of each data type are repeatedly corrected. In each iteration, the uncertainty of one data type is adjusted according to its influence from the uncertainties of other data types, and the adjusted uncertainty, in turn, affects other data types. This process continues until the changes in all uncertainty parameters are less than a preset threshold, reaching a convergence state. Therefore, the obtained uncertainty parameters more realistically reflect the complexity and intrinsic connections of multi-source data, providing a solid foundation for subsequent confidence calculations based on these uncertainty parameters, thus ensuring the accuracy and reliability of the final pollutant identification results.

[0168] For example, suppose the initial uncertainty parameters are: 0.15 for hyperspectral remote sensing data, 0.10 for real-time ground sensor data, and 0.05 for precise laboratory analysis data. In the first iteration, the uncertainty of the real-time ground sensor data is adjusted based on the uncertainty of the hyperspectral remote sensing data and its influence weight. For example, real-time ground sensor uncertainty = initial real-time ground sensor uncertainty + 0.6 * hyperspectral remote sensing uncertainty. Simultaneously, the uncertainty of the precise laboratory analysis data is also adjusted based on the uncertainty of the real-time ground sensor data and its influence weight. In subsequent iterations, these adjusted uncertainty parameters are used as new inputs to recalculate and correct the uncertainties of other data types until the change in all uncertainty parameters is less than a preset threshold (e.g., 0.001), reaching convergence. For example, after multiple iterations, the uncertainty of the hyperspectral remote sensing data might be adjusted to 0.12, the uncertainty of the real-time ground sensor data to 0.08, and the uncertainty of the precise laboratory analysis data to 0.06. These adjusted parameters more accurately reflect the mutual influence between data and the true propagation of uncertainty.

[0169] Specifically, when module 2 is executed, after obtaining the adjusted uncertainty parameters, it can convert them into confidence levels based on a preset confidence level mapping rule. For example, a sigmoid function can be defined to map the uncertainty parameters to confidence levels, such that the lower the uncertainty, the higher the corresponding confidence level. Furthermore, the confidence level changes more gradually when the uncertainty is low, while it decreases more rapidly when the uncertainty is high. This ensures that the assessment of data quality is both sensitive and stable, providing a reliable basis for subsequent pollutant identification.

[0170] The preset confidence mapping rule refers to converting the quantified uncertainty parameters into an intuitive confidence metric that can be used for data fusion and decision-making. It can be implemented using existing techniques such as linear mapping functions, nonlinear mapping functions (such as the Sigmoid function), lookup tables, or piecewise functions.

[0171] Specifically, when extracting key features related to soil pollutant types, speciation, and soil physicochemical properties from multi-source soil data, and obtaining multi-source key feature data, module 3 executes the following:

[0172] Multidimensional feature data is obtained by extracting multidimensional features related to soil pollutant types, form transformations, and soil physicochemical properties from multi-source soil data.

[0173] Based on multi-dimensional feature data, we construct association evidence reflecting the synergistic or antagonistic effects between pollutants, and construct specific evidence indicating the transformation of pollutant forms, thus obtaining multi-source key feature data.

[0174] During execution, extraction module 3 retrieves multi-dimensional features related to soil pollutant types, speciation, and soil physicochemical properties from multi-source soil data, thereby obtaining multi-dimensional feature data. This initial step is fundamental and crucial, ensuring comprehensive acquisition of information directly related to pollutant identification from multiple sources, including hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data.

[0175] For example, reflectance, absorption characteristics, and vegetation indices in specific bands can be extracted from hyperspectral remote sensing data as multi-dimensional features, which can indicate the presence of heavy metals or organic matter in the soil. Soil parameters such as pH, redox potential, conductivity, and temperature can be obtained from real-time ground sensor data; these parameters are closely related to the migration and transformation of pollutants. Information such as the total content, valence state distribution, and organic matter content of specific elements in the soil can be obtained from precise laboratory analysis data. By acquiring these data, a multi-dimensional feature dataset can be constructed.

[0176] Based on multi-dimensional characteristic data, further association evidence reflecting synergistic or antagonistic effects between pollutants, as well as specific evidence indicating pollutant speciation, are constructed, ultimately yielding multi-source key characteristic data. This step further deepens and refines the initially acquired multi-dimensional characteristic data, transforming it into truly insightful key characteristic data. By constructing association evidence reflecting synergistic or antagonistic effects between pollutants, complex interactions that may arise when multiple pollutants coexist can be identified and quantified. For example, certain heavy metals may inhibit the degradation of organic pollutants, or organic matter may affect the migration of heavy metals. The construction of this association evidence transcends the limitations of single-pollutant analysis, providing a deeper understanding of the intrinsic mechanisms of compound pollution, thereby enabling more accurate assessment of pollution risks and identification of pollutant types. Simultaneously, by constructing specific evidence indicating pollutant speciation, the speciation changes of pollutants in the soil environment can be tracked and identified. For example, highly toxic hexavalent chromium may transform into less toxic trivalent chromium. The construction of this specific evidence distinguishes different speciations of pollutants, which is crucial for accurately assessing the bioavailability and toxicity of pollutants, as different speciations of pollutants exhibit drastically different environmental behaviors and ecological risks.

[0177] Among these, constructing evidence reflecting the synergistic or antagonistic effects between pollutants refers to identifying and quantifying evidence of the potential mutually reinforcing or mutually inhibiting effects that may occur when multiple pollutants coexist by analyzing multi-dimensional characteristic data. Constructing specific evidence indicating pollutant speciation refers to identifying and quantifying evidence of chemical or physical speciation changes of pollutants in the soil environment by analyzing multi-dimensional characteristic data.

[0178] Specifically, the correlation between heavy metal indicators and organic matter content in hyperspectral data can be analyzed to construct evidence reflecting synergistic or antagonistic effects between pollutants. For example, if a negative correlation is found between the content of a specific heavy metal and the degradation products of soil organic matter, this may indicate an inhibitory effect of heavy metals on organic matter degradation. This association can be quantified by establishing a multiple regression model or using methods based on graph neural networks.

[0179] Changes in the absorption peak shifts or intensities of specific elements in hyperspectral data can be used to preliminarily determine their valence state changes, constructing specific evidence indicating speciation transformation of pollutants. For example, hexavalent chromium and trivalent chromium exhibit different absorption characteristics in specific spectral regions. Combining this with precise determinations of chromium's valence state in laboratory data can further confirm and quantify this speciation transformation. For instance, if hyperspectral data shows a weakening of the characteristics of hexavalent chromium and an enhancement of the characteristics of trivalent chromium, and laboratory analysis also confirms this transformation, then this information constitutes specific evidence indicating speciation transformation.

[0180] Specifically, when identification module 4 identifies the pollutant type of the soil to be tested based on confidence level and multi-source key feature data, it performs the following:

[0181] Set the corresponding weight coefficients based on the confidence level;

[0182] Based on multi-source key feature data and weighting coefficients, weighted fusion is performed to obtain the pollutant type identification results of the soil to be tested.

[0183] When the identification module 4 is executed, the key feature data with high confidence will have its corresponding weight coefficient set to a higher value, while data with low confidence will be assigned a lower weight. This weight coefficient setting enables intelligent differentiation and treatment of multi-source key feature data of different quality and reliability, thereby effectively avoiding the undue negative impact of low-quality or highly uncertain data on the final identification result and laying the foundation for subsequent accurate identification.

[0184] For example, if the confidence level is a value between 0 and 1, the weighting coefficient can be directly set to be proportional to that value, or it can be converted into an integer weight between 0 and 100 through a preset mapping function (e.g., a sigmoid function or a piecewise linear function). For example, for laboratory precision analysis data with an initial confidence level of 0.9, the weighting coefficient can be set to 90; for ground real-time sensor data with an initial confidence level of 0.6, the weighting coefficient can be set to 60.

[0185] Optionally, a conventional weighted average method can be used to weight and fuse multi-source key feature data using the weight coefficients converted from confidence levels. For example, suppose there are three types of pollutants to be identified: A, B, and C. Key feature data related to these three pollutants are extracted from hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, and their support for the presence of each pollutant is calculated. During weighted fusion, the support of each data source for a specific pollutant type can be multiplied by its corresponding weight coefficient. Then, all weighted support values ​​are summed and divided by the sum of all weight coefficients to obtain the final identification support for that pollutant type. The pollutant types with the highest support ranking or greater than a preset support threshold are identified as the pollutant type identification results for the soil to be tested.

[0186] Preferably, when the identification module 4 performs weighted fusion based on multi-source key feature data and weight coefficients to obtain the pollutant type identification result of the soil to be tested, it performs the following:

[0187] A fusion framework based on Dempster-Shafer evidence theory is adopted to perform weighted fusion of multi-source key feature data and weight coefficients to obtain preliminary pollutant type identification information;

[0188] Determine whether there is any conflicting or ambiguous information in the preliminary pollutant type identification information; if not, determine the preliminary pollutant type identification information as the pollutant type identification result of the soil to be tested; if so, based on the preset discrimination rules, combined with the conflicting or ambiguous information and the corresponding multi-source key feature data, correct or prioritize the preliminary pollutant type identification information to obtain the pollutant type identification result of the soil to be tested.

[0189] Specifically, when identification module 4 uses a fusion framework based on Dempster-Shafer evidence theory to perform weighted fusion of multi-source key feature data and weight coefficients to obtain preliminary pollutant type identification information, it performs the following:

[0190] Define an identification framework that includes all potential soil pollutant types;

[0191] Based on the weighting coefficients, the multi-source key feature data are converted into basic probability assignments for each pollutant type in the identification framework;

[0192] By employing the Dempster combinatorial rule and fusing the basic probability assignments, combined evidence is obtained for each potential pollutant type.

[0193] Based on the combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type.

[0194] It should be noted that the fusion framework based on Dempster-Shafer evidence theory is a mathematical tool for handling uncertainty and multi-source information fusion. It can be implemented by defining the identification framework, calculating the basic probability allocation, applying the combination rules, and generating preliminary identification information.

[0195] During execution, identification module 4 defines an identification framework encompassing all potential soil pollutant types (e.g., heavy metals, organochlorine pesticides, petroleum hydrocarbons, etc.), clarifying the scope of pollutant identification and all possible targets. Based on this, multi-source key feature data is converted into basic probability assignments for each pollutant type within the identification framework, according to weighting coefficients. These weighting coefficients quantify the reliability of the data sources, ensuring that the support level of each data type for a specific pollutant type is accurately mapped to the basic probability assignment. This provides confidence-based input for subsequent evidence fusion, effectively handling data uncertainty and ensuring that the contributions of different data sources are reasonably weighted. Next, the Dempster combination rule is used to fuse the basic probability assignments, obtaining combined evidence for each potential pollutant type. The Dempster combination rule effectively aggregates information from multiple independent evidence sources. Even if these sources contain conflicting or uncertain information, the rule yields a comprehensive and more persuasive combined evidence, enhancing the reliability and robustness of the identification results and providing a more solid foundation for subsequent decision-making. Finally, based on the combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type. This step forms the basis for the final decision. It leverages the combined power of all available evidence, providing a solid foundation for subsequent conflict resolution or handling of ambiguous information. It ensures that the initial identification results are based on comprehensive and weighted multi-source data, improving the accuracy and credibility of the identification. This fusion approach, employing a framework based on Dempster-Shafer evidence theory, goes beyond simple weighted averaging, more comprehensively reflecting the accumulation of evidence and the propagation of uncertainty, thus laying the foundation for subsequent judgments.

[0196] For example, firstly, an identification framework is defined, which may include potential pollutant types such as "lead (Pb)," "cadmium (Cd)," "arsenic (As)," "hexavalent chromium (Cr(VI))," and "polychlorinated biphenyls (PCBs)." Then, based on multi-source key feature data extracted from hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data, along with the corresponding weighting coefficients for these data types, this data is converted into basic probability assignments. For instance, if hyperspectral data indicates potential heavy metal pollution in the soil and its weighting coefficient is high, it can be converted into a basic probability assignment that assigns high confidence to heavy metal types such as "lead," "cadmium," and "arsenic," while assigning lower confidence to organic pollutants such as "PCBs." Simultaneously, if laboratory analysis data precisely determines the presence of hexavalent chromium and its weighting coefficient is extremely high, a basic probability assignment can be generated that assigns extremely high confidence to the "hexavalent chromium" type. Next, the Dempster combination rule is used to fuse these basic probability assignments from different data sources. For example, the basic probability assignments generated from hyperspectral data are combined with those generated from laboratory analysis data. If the hyperspectral data has a confidence level of 0.6 for lead, while the laboratory data has a confidence level of 0.3, Dempster's combination rule will consider both pieces of evidence to generate a combined confidence level for lead that is either higher or lower, while also addressing potential conflicts or uncertainties. Finally, based on the fused combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type. For example, if the combined evidence shows the highest confidence level for hexavalent chromium, significantly higher than other pollutant types, the preliminary pollutant type identification information could indicate that the primary pollutant type in the tested soil is hexavalent chromium. If the combined evidence shows high and similar confidence levels for both lead and cadmium, the preliminary identification information could indicate the presence of combined pollution and provide the relative confidence levels for the two pollutants.

[0197] Furthermore, during execution, the identification module 4 determines whether there is conflicting or ambiguous information regarding pollutant type identification in the preliminary pollutant type identification information. This step introduces a self-diagnostic capability for the fusion results. In practical applications, even with fusion theory, the preliminary identification results may still be contradictory. For example, different pieces of evidence may strongly point to different pollutants, or the information may be ambiguous, such as insufficient evidence to clearly point to a single pollutant, instead pointing to a set of pollutants. By explicitly identifying these conflicts or ambiguities, unreliable results can be avoided, providing a clear basis for subsequent corrections. Conflicting pollutant type identification information refers to contradictory pollutant type judgments from different data sources or evidence during the preliminary pollutant type identification process. This can be identified by analyzing inconsistencies between evidence or calculating conflict coefficients. Ambiguous pollutant type identification information refers to situations where, during the preliminary pollutant type identification process, evidence is insufficient to clearly point to a single pollutant type, instead pointing to a set of pollutant types, or the support for any single pollutant type is low or the uncertainty range is too large. This can be identified by assessing the uncertainty of the belief function or analyzing the diffusion of evidence.

[0198] For example, a conflict metric between pieces of evidence can be calculated, such as the conflict coefficient in the Dempster-Shafer theory. If this coefficient exceeds a preset threshold, it is determined that there is conflicting information regarding the identification type. Simultaneously, the uncertainty of the belief function can be assessed; for instance, if the belief level regarding any single pollutant type is below a certain threshold, or if the uncertainty range is too large, it is determined that there is ambiguous information regarding type identification.

[0199] If the preliminary pollutant type identification information does not contain conflicting or ambiguous information, then this preliminary pollutant type identification information is directly determined as the pollutant type identification result for the soil to be tested. This indicates that when the fusion results are clear and consistent, a final judgment can be given efficiently, avoiding unnecessary complex processing and improving identification efficiency. For example, if the preliminary identification information clearly indicates that the soil is mainly contaminated with lead, and the confidence level is high, then it is directly identified as lead contamination.

[0200] If conflicting or ambiguous information regarding pollutant type identification exists, the preliminary pollutant type identification information is corrected or prioritized based on pre-defined discrimination rules, combined with the conflicting or ambiguous information and corresponding multi-source key feature data, ultimately yielding the pollutant type identification result for the soil being tested. This step provides a mechanism for handling complex situations. The pre-defined discrimination rules can be based on expert knowledge, historical data, or the needs of specific application scenarios to guide how to resolve conflicts, such as prioritizing evidence from specific high-confidence data sources, using majority voting mechanisms, or handling ambiguity, such as prioritizing the most likely pollutants or requesting supplementary data. By combining the original multi-source key feature data, the correction process can trace back to the root cause of the conflict or ambiguity, allowing for targeted adjustments, such as re-evaluating conflicting evidence or refining the allocation of ambiguous beliefs. This correction or prioritization mechanism ensures that the final pollutant type identification result is well-considered, more robust, and reliable.

[0201] For example, when conflicts arise, pre-defined rules can prioritize laboratory analysis data, as it typically has higher reliability; alternatively, a majority voting mechanism can be used, selecting the pollutant type supported by the majority of evidence. When ambiguous information exists, pre-defined rules can guide the prioritization of the most likely pollutant types, for example, identifying the two pollutant types with the highest belief levels (such as cadmium and arsenic) as the primary suspects and recommending further supplementary sampling or analysis. During the correction process, original multi-source key characteristic data can be incorporated; for example, re-examining whether a particular sensor data point causing the conflict is anomaly, or refining ambiguous belief assignments, for example, by introducing additional auxiliary features (such as the impact of soil pH on pollutant mobility) to further differentiate similar pollutant types. Through these corrections or prioritizations, an accurate and reliable identification result of the soil pollutant type is ultimately output.

[0202] As can be seen from the above, this soil pollutant type identification system can analyze and fuse the extracted multi-source key feature data while taking into account the reliability of the data, thereby improving the identification efficiency of soil pollutant types.

[0203] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0204] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0205] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0206] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0207] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for identifying soil pollutant types, characterized in that, Including the following steps: Acquire multi-source soil data for the soil to be tested; Based on the data types of the multi-source soil data, the uncertainty corresponding to each data type is determined, and the confidence level corresponding to each data type is obtained. Key features related to soil pollutant types, speciation, and soil physicochemical properties are extracted from the multi-source soil data to obtain multi-source key feature data. Based on the confidence level and the multi-source key feature data, the pollutant type identification result of the soil to be tested is obtained; Based on the data types of the multi-source soil data, the uncertainty corresponding to each data type is determined, and the confidence level corresponding to each data type is obtained, including: Based on the acquisition method, instrument accuracy, sampling error, and environmental interference of the multi-source soil data, the uncertainty of each data type in the multi-source soil data is quantified to obtain the initial uncertainty parameters of each data type. Based on the correlation between the data types in the multi-source soil data, the initial uncertainty parameters are adjusted to obtain the uncertainty parameters for each data type; Based on the preset confidence mapping rules, the confidence level corresponding to the uncertainty parameter is determined; Based on the correlation between different data types in the multi-source soil data, the initial uncertainty parameters are adjusted to obtain the uncertainty parameters for each data type, including: Identify the relationships between different data types in the multi-source soil data; Based on the aforementioned relationships, the impact paths and impact weights of uncertainty for each data type are determined; Based on the influence path and the influence weight, the uncertainty of each data type in the multi-source soil data is iteratively adjusted to obtain the uncertainty parameters of each data type.

2. The method for identifying soil pollutant types according to claim 1, characterized in that, Acquire multi-source soil data for the soil to be tested, including: Obtain initial multi-source soil data for the soil to be tested; The initial multi-source soil data is spatiotemporally aligned to obtain the multi-source soil data of the soil to be tested.

3. The method for identifying soil pollutant types according to claim 2, characterized in that, The multi-source soil data includes hyperspectral remote sensing data, real-time ground sensor data, and precise laboratory analysis data.

4. The method for identifying soil pollutant types according to claim 1, characterized in that, Key features related to soil pollutant types, speciation, and soil physicochemical properties are extracted from the multi-source soil data to obtain multi-source key feature data, including: Multidimensional features related to soil pollutant types, speciation, and soil physicochemical properties are extracted from the multi-source soil data to obtain multidimensional feature data. Based on the multi-dimensional feature data, we construct association evidence reflecting the synergistic or antagonistic effects between pollutants, and construct specific evidence indicating the transformation of pollutant forms, thus obtaining multi-source key feature data.

5. The method for identifying soil pollutant types according to claim 1, characterized in that, Based on the confidence level and the multi-source key feature data, the pollutant type identification results of the soil to be tested are obtained, including: Based on the confidence level, set the corresponding weight coefficients; Based on the multi-source key feature data and the weighting coefficients, a weighted fusion is performed to obtain the pollutant type identification result of the soil to be tested.

6. The method for identifying soil pollutant types according to claim 5, characterized in that, Based on the multi-source key feature data and the weighting coefficients, a weighted fusion is performed to obtain the pollutant type identification result of the soil to be tested, including: A fusion framework based on Dempster-Shafer evidence theory is used to perform weighted fusion of the multi-source key feature data and the weight coefficients to obtain preliminary pollutant type identification information; Determine whether there is any identification type conflict information or type identification ambiguity information in the preliminary pollutant type identification information; if not, determine that the preliminary pollutant type identification information is the pollutant type identification result of the soil to be tested; if so, according to the preset discrimination rules, combined with the identification type conflict information or type identification ambiguity information and the corresponding multi-source key feature data, correct or prioritize the preliminary pollutant type identification information to obtain the pollutant type identification result of the soil to be tested.

7. The method for identifying soil pollutant types according to claim 6, characterized in that, A fusion framework based on Dempster-Shafer evidence theory is used to perform weighted fusion of the multi-source key feature data and the weight coefficients to obtain preliminary pollutant type identification information, including: Define an identification framework that includes all potential soil pollutant types; Based on the weighting coefficients, the multi-source key feature data is converted into basic probability assignments for each pollutant type in the identification framework; By employing the Dempster combination rule and fusing the basic probability assignments, combined evidence for each potential pollutant type is obtained. Based on the combined evidence, preliminary pollutant type identification information is generated for each potential pollutant type.

8. A soil pollutant type identification system for identifying soil pollutant types, characterized in that, include The acquisition module is used to acquire multi-source soil data of the soil to be tested; The determination module is used to determine the uncertainty corresponding to each data type based on the data type of the multi-source soil data, and to obtain the confidence level corresponding to each data type. The extraction module is used to extract key features related to soil pollutant types, form transformations, and soil physicochemical properties from the multi-source soil data to obtain multi-source key feature data. The identification module is used to identify the pollutant type identification result of the soil to be tested based on the confidence level and the multi-source key feature data; The determination module is used to determine the uncertainty corresponding to each data type based on the data type of the multi-source soil data, and to obtain the confidence level corresponding to each data type, including: Based on the acquisition method, instrument accuracy, sampling error, and environmental interference of the multi-source soil data, the uncertainty of each data type in the multi-source soil data is quantified to obtain the initial uncertainty parameters of each data type. Based on the correlation between the data types in the multi-source soil data, the initial uncertainty parameters are adjusted to obtain the uncertainty parameters for each data type; Based on the preset confidence mapping rules, the confidence level corresponding to the uncertainty parameter is determined; Based on the correlation between different data types in the multi-source soil data, the initial uncertainty parameters are adjusted to obtain the uncertainty parameters for each data type, including: Identify the relationships between different data types in the multi-source soil data; Based on the aforementioned relationships, the impact paths and impact weights of uncertainty for each data type are determined; Based on the influence path and the influence weight, the uncertainty of each data type in the multi-source soil data is iteratively adjusted to obtain the uncertainty parameters of each data type.

Citation Information

Patent Citations

  • Accurate diagnosis method and device of pollution source and electronic device

    CN110376343A

  • Ship atmospheric pollutant monitoring method and system

    CN114037064A