Tracing method and system for new pollutants in various environmental media based on machine learning

By building a machine learning intelligent decision-making system, collecting and integrating diversified environmental data, the data integration and path tracking problems in the traceability of new pollutants have been solved, efficient and accurate pollution source positioning has been achieved, and traceability efficiency and accuracy have been improved.

CN120387598AActive Publication Date: 2025-07-29GUANGDONG INST OF ANALYSIS CHINA NAT ANALYTICAL CENT GUANGZHOU

Patent Information

Application Number
CN202510885466.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing technology has difficulty in dealing with changes in the chemical characteristics of organic pollutants in complex environments, insufficient data integration, strong experience relying on experts, and difficulty in cross-media pollutant transmission simulation and tracking, resulting in low traceability efficiency and inaccurate results.

Method used

By building an intelligent decision-making system based on machine learning, collecting multi-environment data to build a pollutant fingerprint library, performing spatiotemporal correlation analysis and transmission path reconstruction, combining graph convolution networks and multi-branch neural networks to fusion of multi-source data, localizing key pollution sources based on the pollution source contribution rate matrix, and using the C4.5 decision tree algorithm for traceability inference.

Benefits of technology

The automation, standardization and precise traceability of new pollutants has been achieved, the traceability efficiency and accuracy have been improved, the dependence on expert experience has been reduced, and the scientificity and transparency of the decision-making process has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387598A_ABST
    Figure CN120387598A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pollutant tracing, and discloses a method and a system for tracing new pollutants in various environmental media based on machine learning. The method comprises the following steps: collecting multivariate environment data to construct a pollutant fingerprint database; performing space-time correlation analysis on the fingerprint database and the environment data, and reconstructing a transmission path; converting the data, the fingerprint database and the path diagram into source node fusion to construct an analytical model, and forming a contribution rate matrix; and positioning the pollution source through the three-dimensional features based on the matrix. According to the method and the device, an intelligent decision-making system based on machine learning can be constructed by integrating multi-source heterogeneous data under a complex environment background, automatic, standardized and precise traceability of new pollutants is realized, and the traceability efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of pollutant source tracing, and particularly to a method and system for tracing new pollutants in multiple environmental media based on machine learning. Background Art

[0002] In the field of environmental pollution prevention and control, tracing the sources of new pollutants is a key task, which is of great significance for the water quality, air quality, and soil quality in multiple environmental media. Traditional pollutant source tracing methods mainly rely on manual experience judgment and simple mathematical models, such as the mass balance method, diffusion model, and statistical correlation analysis, etc. These methods usually work on conventional pollutants such as heavy metals, nitrogen, and phosphorus. The basic process includes on-site sampling and analysis, determination of pollution indicators, calculation of diffusion laws, and determination of pollution sources, etc. On this basis, with the development of analysis technologies, fingerprint recognition technology has begun to be applied to environmental pollution source tracing, improving the accuracy of source tracing by identifying the unique chemical characteristics and isotope compositions of pollutants. At the same time, the application of geographic information systems and remote sensing technologies has also provided new technical support for large-scale pollution source identification, making the spatial distribution characteristics of regional pollutants clearer.

[0003] However, there are still obvious deficiencies in the existing technologies for tracing new pollutants. First, traditional source tracing methods are difficult to deal with new pollutants in complex environments, especially those organic pollutants that degrade and transform in the environment, whose chemical characteristics often change, resulting in inaccurate source tracing judgments. Second, there is a general problem of insufficient data integration in existing source tracing technologies, and it is difficult to effectively integrate multi-source heterogeneous data, such as chemical analysis data, spatio-temporal distribution data, and human activity data, etc., resulting in a lack of systematicness in source tracing results. Third, traditional source tracing methods rely too much on expert experience and lack standardized and intelligent decision support tools, making the source tracing process subjective, inefficient, and the results have poor repeatability. In addition, with the continuous increase in the types of new pollutants in the environment, the existing pollutant fingerprint libraries are seriously lagging behind and cannot meet the needs of rapid identification and accurate source tracing of new pollutants. Finally, there are technical bottlenecks in the simulation and tracking of cross-media pollutant transport in existing source tracing technologies, and it is difficult to accurately reconstruct the complete transport path of pollutants from the source to the detection point. Summary of the Invention

[0004] This application provides a method and system for tracing new pollutants in multiple environmental media based on machine learning, which is used to build an intelligent decision-making system based on machine learning by integrating multi-source heterogeneous data in a complex environmental background, realizing automated, standardized, and precise source tracing of new pollutants, and improving the source tracing efficiency and accuracy.

[0005] In a first aspect, the present application provides a method for tracing new pollutants in multiple environmental media based on machine learning. The method for tracing new pollutants in multiple environmental media based on machine learning includes: collecting multivariate environmental data through a sensor network, and constructing a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data; performing spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and performing multi-source data fusion to construct a source analysis model and obtain a source contribution rate matrix; based on the source contribution rate matrix, performing key pollutant source location through pollutant characteristic dimension, spatial location dimension, and time characteristic dimension to obtain pollutant source location information.

[0006] In a second aspect, the present application provides a system for tracing new pollutants in multiple environmental media based on machine learning. The system for tracing new pollutants in multiple environmental media based on machine learning includes: a construction module for collecting multivariate environmental data through a sensor network and constructing a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data; a reconstruction module for performing spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; a fusion module for converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and performing multi-source data fusion to construct a source analysis model and obtain a source contribution rate matrix; a location module for performing key pollutant source location through pollutant characteristic dimension, spatial location dimension, and time characteristic dimension based on the source contribution rate matrix to obtain pollutant source location information.

[0007] In a third aspect, there is provided a device for tracing new pollutants in multiple environmental media based on machine learning, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the device for tracing new pollutants in multiple environmental media based on machine learning executes the above-mentioned method for tracing new pollutants in multiple environmental media based on machine learning.

[0008] In a fourth aspect, there is provided a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is made to execute the above-mentioned method for tracing new pollutants in multiple environmental media based on machine learning.

[0009] In the technical solution provided by this application, through the collection of multi-source environmental data and the construction of a pollutant fingerprint library, the accurate identification and feature extraction of new pollutants are achieved, thereby significantly improving the starting point accuracy of traceability; through the spatio-temporal correlation analysis of the fingerprint library and multi-source environmental data and the reconstruction of the transmission path, the migration and diffusion process of pollutants in the environment is visualized, solving the technical problem that it is difficult to track the pollutant transmission path in a complex environment by traditional methods; converting multi-source environmental data, fingerprint library and pollutant transmission path map into quantum states and performing multi-source data fusion, deeply mining the deep correlations between different types of data, where quantum computing exhibits parallel processing capabilities and dimensionality reduction efficiency that are incomparable to traditional computing methods in multi-dimensional feature processing. This innovative data fusion strategy solves the problem of insufficient data integration by traditional methods; based on the pollution source contribution rate matrix, key pollution sources are located through a three-dimensional decision space, realizing the gradual positioning from a large-scale area to an accurate geographical location, greatly improving the spatial resolution and accuracy of traceability. In particular, the application of the C4.5 decision tree algorithm in pollution source discrimination objectively quantifies the importance of each decision attribute through the information gain ratio compared with traditional empirical judgments, making the decision-making process more scientific and transparent, and eliminating the uncertainty brought by subjective human judgment in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as graph convolutional networks for processing transmission paths, multi-branch neural networks for feature fusion, and decision trees for constructing traceability inference chains, fully exerts the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion, and intelligent decision support, solving the problem of new pollutant traceability that cannot be handled by traditional methods, improving the traceability efficiency, and reducing the dependence on expert experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 FIG. is a schematic diagram of an embodiment of a method for tracing new pollutants in multiple environmental media based on machine learning in an embodiment of this application; Figure 2 FIG. is a schematic diagram of an embodiment of a system for tracing new pollutants in multiple environmental media based on machine learning in an embodiment of this application; Figure 3 FIG. is a schematic block diagram of the structure of a device for tracing new pollutants in multiple environmental media based on machine learning in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The embodiments of the present application provide a method and system for tracing new pollutants in multiple environmental media based on machine learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the term "including" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0013] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 One embodiment of the method for tracing new pollutants in multiple environmental media based on machine learning in the embodiments of the present application includes: Step S101: Collect multivariate environmental data through a sensor network, and construct a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data; Step S102: Perform spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; Step S103: Convert the multivariate environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and perform multi-source data fusion to construct a source analysis model and obtain a source contribution rate matrix; Step S104: Based on the source contribution rate matrix, perform key pollutant source location through the pollutant feature dimension, the spatial location dimension, and the time feature dimension to obtain the source location information.

[0014] It can be understood that the execution subject of the present application can be a system for tracing new pollutants in multiple environmental media based on machine learning, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution subject as an example.

[0015] Specifically, multi-source environmental data is collected through a sensor network and a pollutant fingerprint library is constructed. In this step, a multi-source sensor network including water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors, and biological response sensors is deployed, and physical and chemical parameters in the environment are collected in real time through low-power wide-area network technology. In the embodiments of the present application, the multi-environment media refer to water bodies, soil, and the atmosphere. Therefore, the solution of the present application is also for the source tracing of new pollutants in water bodies, soil, and the atmosphere. The original environmental data identifies outliers through the three-standard-deviation method, fills short-term missing data using linear interpolation, processes long-term missing data using multi-source interpolation methods, and then performs noise reduction through wavelet transform to finally obtain preprocessed environmental data. The preprocessed environmental data is standardized using the Z-score standardization method and is spatially and temporally matched and integrated with geographic information system data, meteorological data, human activity data, and historical monitoring records to form multi-source environmental data. At the same time, high-resolution mass spectrometry technology is used for non-targeted analysis of environmental samples, and a mass spectrometry feature table is obtained through peak identification, peak alignment, and peak area extraction. The mass spectrometry feature table is classified and clustered and database retrieval is performed. For unlisted compounds, structure analysis is carried out through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a candidate list of new pollutants. Finally, the mass spectrometry features, chromatographic features, physical and chemical properties, and environmental behavior features of each new pollutant are extracted, and feature optimization and screening are carried out through principal component analysis and recursive feature elimination algorithms to construct a pollutant fingerprint library.

[0016] Perform spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and multi-source environmental data. In this step, the multi-source environmental data is organized according to the time dimension, space dimension, and pollutant feature dimension to construct a three-dimensional data cube. The autoregressive integrated moving average model is applied to analyze the time series data of pollutant concentrations, which is decomposed into trend components, seasonal components, and random components to obtain time-varying characteristic data. Then, through geographically weighted regression and spatial autocorrelation analysis techniques, the global Moran index and local spatial autocorrelation indicators are calculated to identify the aggregation areas and abnormal hotspots of pollutants in space and generate spatial distribution characteristic data. Combining the time-varying characteristic data and the spatial distribution characteristic data, the variogram is used to analyze the spatial structure characteristics of pollutant concentrations, determine the spatial correlation distance and anisotropy characteristics, construct a spatial interpolation algorithm for pollutant concentrations, and generate a pollution distribution heat map of the study area. Based on the pollution distribution heat map, combined with the hydrodynamic equation, atmospheric diffusion equation, and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated, and their migration and diffusion processes are simulated to obtain multi-media transmission dynamic data. Finally, the backward trajectory analysis technology is applied, with the pollutant detection point as the end point, combined with the pollutant characteristic data in the fingerprint library, to reverse the transmission starting point and key transmission nodes of the pollutants and generate a pollutant transmission path map.

[0017] Convert multi - source environmental data, fingerprint libraries, and pollutant transmission path maps into source nodes and perform multi - source data fusion. Standardize the features of multi - source environmental data, use feature selection algorithms to screen significant features, perform dimensionality reduction through principal component analysis, and generate environmental data source nodes. Apply feature encoding methods to pollutant fingerprint identifiers in the fingerprint library, convert mass spectrometry features, chromatographic features, and physicochemical properties into numerical feature matrices, and generate fingerprint library source nodes through non - linear dimensionality reduction mapping. Process the pollutant transmission path map through a graph convolutional network, parametrically represent path nodes and connection relationships, extract path topological features and transmission kinetic features, and generate transmission path source nodes. Construct a multi - branch neural network for environmental data source nodes, fingerprint library source nodes, and transmission path source nodes, design an attention mechanism fusion layer to weight - integrate features, optimize the network structure through cross - validation, and obtain a set of fused source nodes. Based on the set of fused source nodes, construct a multi - layer perceptron network with three hidden layers, each layer containing 128, 64, and 32 neurons respectively, use the ReLU function as the activation function, the Softmax function for the output layer, select the cross - entropy loss as the loss function, and optimize the parameters through the Adam optimizer to obtain a source apportionment classifier. Use the source apportionment classifier to classify and predict the pollutant sources, calculate the probability distribution of each pollutant source category, combine the pollutant concentration data to calculate the contribution ratio of each source, and construct a three - dimensional data structure containing source type, contribution rate, and geographical coordinates to obtain a pollutant source contribution rate matrix.

[0018] Based on the pollution source contribution rate matrix, key pollution source location is carried out through the pollutant characteristic dimension, spatial location dimension and time characteristic dimension. In this step, the pollution source contribution rate matrix is split three-dimensionally, pollutant characteristic parameters, spatial coordinate information and time series data are extracted, a three-dimensional decision space is constructed, and the basic data for source tracing decision-making is obtained. Based on the pollutant characteristic parameters in the basic data for source tracing decision-making, chemical classification, molecular weight range division, functional group characteristic analysis and environmental degradation characteristic evaluation of pollutants are carried out, a decision-making path for the pollutant characteristic dimension is constructed, and the pollutant characteristic judgment result is obtained. Based on the spatial coordinate information, combined with the topographic and geomorphic features, water system distribution and administrative division, the spatial relationship analysis of the upstream area of the pollutant detection point, the distribution of surrounding potential pollution sources and land use types is carried out, a decision-making path for the spatial location dimension is constructed, and the spatial location judgment result is obtained. Based on the time series data, the differences between weekdays and weekends of pollutant detection, seasonal change characteristics and time correlation with specific events are analyzed, a decision-making path for the time characteristic dimension is constructed, and the time characteristic judgment result is obtained. The pollutant characteristic judgment result, spatial location judgment result and time characteristic judgment result are input into the decision tree model, the information gain ratio is calculated by the C4.5 algorithm for optimal splitting attribute selection, a source tracing inference chain is generated, and the candidate area of the pollution source is obtained. Finally, the grid refinement strategy is applied to the candidate area of the pollution source, temporary monitoring points are set, fingerprint feature comparison and analysis are carried out on the collected environmental samples and potential pollution source samples, the similarity score is calculated, the exact location of the pollution source is determined, and the pollution source location information is obtained.

[0019] In the embodiments of the present application, through the collection of multi-source environmental data and the construction of a pollutant fingerprint library, the accurate identification and feature extraction of new pollutants are realized, thus significantly improving the starting point accuracy of source tracing; through the spatio-temporal correlation analysis of the fingerprint library and multi-source environmental data and the reconstruction of the transmission path, the migration and diffusion process of pollutants in the environment is visualized, solving the technical problem that it is difficult to track the pollutant transmission path in a complex environment by traditional methods; converting multi-source environmental data, the fingerprint library and the pollutant transmission path map into quantum states and performing multi-source data fusion, deeply exploring the deep correlations between different types of data, where quantum computing exhibits parallel processing capabilities and dimensionality reduction efficiency that are unparalleled by traditional computing methods in multi-dimensional feature processing, and this innovative data fusion strategy solves the problem of insufficient data integration by traditional methods; based on the pollution source contribution rate matrix, through a three-dimensional decision space, key pollution sources are located, realizing the step-by-step positioning from a large-scale region to an accurate geographical location, greatly improving the spatial resolution and accuracy of source tracing. In particular, the application of the C4.5 decision tree algorithm in pollution source discrimination objectively quantifies the importance of each decision attribute through the information gain ratio compared with traditional empirical judgments, making the decision-making process more scientific and transparent and eliminating the uncertainty brought by subjective human judgments in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as the graph convolutional network for processing the transmission path, the multi-branch neural network for realizing feature fusion, and the decision tree for constructing the source tracing inference chain, gives full play to the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion and intelligent decision support, solving the problem of new pollutant source tracing that cannot be handled by traditional methods, improving the source tracing efficiency and reducing the dependence on expert experience.

[0020] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Deploy water quality monitoring sensors, air quality monitoring sensors, soil monitoring sensors and biological response sensors to form a sensor network, and collect environmental physical and chemical parameters through a low-power wide-area network to obtain raw environmental data; Apply the three-sigma method to the raw environmental data to identify outliers, use linear interpolation to fill short-term missing data, use multi-variate interpolation methods to fill long-term missing data, and perform noise reduction processing through wavelet transform to obtain preprocessed environmental data; Apply the Z-score standardization method to the preprocessed environmental data for standardization processing, and perform spatio-temporal matching and integration in combination with geographic information system data, meteorological data, human activity data and historical monitoring records to obtain multi-source environmental data; Use high-resolution mass spectrometry technology to perform non-targeted analysis on environmental samples, and perform peak identification, peak alignment and peak area extraction processing on the mass spectrometry data to obtain a mass spectrometry feature table; The mass spectral feature table is classified and clustered, and a database search is performed simultaneously. The structures of unlisted compounds are elucidated through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a candidate list of new pollutants. For each pollutant in the new pollutant candidate list, mass spectrometry characteristics, chromatographic characteristics, physicochemical properties and environmental behavior characteristics are extracted, and principal component analysis and recursive feature elimination algorithm are used to perform feature optimization screening to obtain a fingerprint library containing pollutant fingerprint identification.

[0021] Specifically, a multi-sensor network is deployed to collect environmental data. The sensor network includes water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors, and bio-response sensors. Water quality monitoring sensors are deployed at surface water bodies, groundwater sampling points, and sewage outlets to monitor water quality parameters such as pH, dissolved oxygen, total organic carbon, and heavy metal content in real time. Atmospheric monitoring sensors are deployed around industrial parks, urban transportation nodes, and sensitive areas to monitor atmospheric pollutants such as volatile organic compounds and particulate matter. Soil monitoring sensors are deployed in a grid-based manner across the study area, covering different land types, to monitor persistent organic pollutants in the soil. Bio-response sensors select indicator organisms sensitive to specific pollutants to capture pollutant accumulation within organisms. These sensors transmit data via low-power wide-area network technology, enabling real-time collection of environmental physical and chemical parameters and generating raw environmental data. It should be noted that in the embodiments of this application, water environment monitoring can also be achieved through on-site automatic online monitoring, drone remote sensing monitoring, unmanned boat mobile monitoring, and laboratory testing. Soil condition monitoring can also be achieved through on-site rapid testing, in-situ testing, drone remote sensing monitoring, and laboratory testing.

[0022] The triple standard deviation method is applied to the data to identify outliers. This method calculates the mean and standard deviation of the dataset and marks data points that deviate from the mean by more than three standard deviations as outliers. Outliers verified to be caused by equipment failure or environmental changes are corrected or removed. For missing data, different interpolation methods are used depending on the duration of the missing data. For short-term missing data (e.g., data missing within a few hours), linear interpolation is used to fill the missing value using a linear function calculated from valid data before and after the missing point. For long-term missing data (e.g., data missing for several days), multivariate interpolation methods based on historical data patterns are used, taking into account the historical variation patterns of multiple related variables to generate more realistic interpolated values. The data is then subjected to wavelet transform for noise reduction. By decomposing the signal into wavelet coefficients of different frequencies, high-frequency random noise is removed while retaining valid information. The data after these processes is referred to as preprocessed environmental data.

[0023] The pre-processed environmental data needs to be standardized. The Z-score standardization method is adopted, that is, for each data point, subtract the mean of the variable and then divide by the standard deviation to convert environmental parameters with different dimensions into a unified standard. The standardized data is spatially and temporally matched and integrated with auxiliary data sources, which include geographic information system data (topography, water system distribution, geological structure, and land use type), meteorological data (rainfall, wind direction and speed, temperature, and humidity), human activity data (industrial layout, agricultural activities, urbanization process, and traffic flow), and historical monitoring records. Optionally, the auxiliary data source can also be geographic information data, meteorological data, industrial and agricultural production status data, hydrological data (water bodies), land use information (soil), and human activity data (population, activity type) related to polluted areas.

[0024] These auxiliary data are collected and integrated through a data interface, spatially and temporally matched with the sensor network data to construct a complete multi-source environmental data. High-resolution mass spectrometry technology is used for non-targeted analysis of environmental samples, including liquid chromatography-quadrupole-time-of-flight mass spectrometry, gas chromatography-quadrupole-time-of-flight mass spectrometry, and ultra-high performance liquid chromatography-mass spectrometry. The mass spectrometry data is processed by professional analysis software for peak identification (identifying real peaks in the chromatogram and excluding noise), peak alignment (matching corresponding chromatographic peaks in different samples), and peak area extraction (calculating the area of each peak for quantitative analysis) to generate a mass spectrometry feature table. Optionally, in the embodiments of the present application, data analysis can also be performed on the characteristic spectral data and the corresponding signal intensity data.

[0025] Classify and cluster the mass spectrometry feature table and conduct database retrieval. The classification and clustering adopt molecular network analysis technology. Based on the similarity of the mass spectrometry features of compounds, they are grouped into several clusters, and each type of pollutant generates a unique mass spectrometry fingerprint. At the same time, the identified compounds are compared with the known environmental pollutant database and chemical registration database. Compounds not included in the existing database are marked as potential new pollutants. Conduct structure analysis on these potential new pollutants. Through isotope ratio analysis (inferring molecular composition by analyzing the isotope abundance distribution characteristics of elements), fragment ion analysis (inferring molecular structure according to the mass spectrometry fragmentation rules), and retention time prediction (predicting the retention behavior of compounds on the chromatographic column based on compound structure characteristics), deduce their molecular formulas and possible chemical structures to obtain a candidate list of new pollutants. Extract features for each pollutant in the candidate list of new pollutants and construct a fingerprint library. The extracted features include mass spectrometry features (exact mass, isotope pattern, characteristic fragment ions, ion abundance ratio), chromatographic features (retention time, peak shape), physicochemical properties (hydrophilic / hydrophobic, acid-base, volatility), and environmental behavior features (degradation rate, bioaccumulation). Optimize and screen these features through principal component analysis and recursive feature elimination algorithm. Principal component analysis realizes data dimensionality reduction by converting the original features into linearly independent new features (principal components) and retains the most informative feature combination; the recursive feature elimination algorithm screens out the most representative and discriminative feature set by repeatedly constructing models, evaluating feature importance, and removing the least important features. The screened features constitute the final pollutant fingerprint and are stored in the fingerprint library.

[0026] In a specific embodiment, the process of performing the step of classifying and clustering the mass spectrometry feature table may specifically include the following steps: Apply molecular network analysis technology to the mass spectrometry feature table to calculate the spectral similarity matrix between compounds. Compounds with a similarity score greater than 0.7 are grouped into the same cluster through the spectral similarity score to obtain a preliminary molecular network structure; Apply the hierarchical clustering algorithm to the preliminary molecular network structure for structure optimization, calculate the Euclidean distance between clusters, set the clustering threshold, and obtain an optimized molecular network map; Conduct two-way retrieval and matching of the mass spectrometry features in the molecular network map with the known environmental pollutant database and chemical registration database. Obtain the database matching results through comparison of molecular formulas, molecular weights, and characteristic fragment ions; Mark compounds with a matching degree lower than 80% in the database matching results, and conduct screening and filtering in combination with the retention time index to obtain a list of potential new pollutants; Conduct isotope ratio analysis on the compounds in the list of potential new pollutants, calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine elements, and deduce the molecular formula in combination with the element composition rules to obtain molecular composition data; Based on the molecular composition data, perform substructure analysis on characteristic fragment ions, apply the fragment tree algorithm to reconstruct the molecular skeleton, and optimize the molecular structure by combining electron density calculations to obtain a candidate list of new pollutants.

[0027] Specifically, apply molecular network analysis technology to the mass spectrometry feature table to calculate the spectral similarity matrix between compounds. Molecular network analysis technology is a method that classifies related compounds using spectral similarity. Its core is to calculate the cosine similarity between the mass spectra of different compounds. The specific operation is to regard the mass spectrum of each compound as a high-dimensional vector, with the mass-to-charge ratio of the mass spectrometry peaks as the dimension and the peak intensity as the value on this dimension, and then calculate the cosine similarity between the vectors of two compounds. The calculation result of the cosine similarity ranges from 0 to 1, and the value closer to 1 indicates that the two spectra are more similar. In this method, the threshold is set to 0.7, and compounds with similarity scores greater than this threshold are grouped into the same cluster to form a preliminary molecular network structure. This network structure intuitively shows the similarity relationship between compounds, and similar compounds cluster together in the network. Apply the hierarchical clustering algorithm to optimize the structure of the preliminary molecular network structure. The hierarchical clustering algorithm is a method that gradually merges or splits data points to form a hierarchical structure. In this method, the bottom-up agglomerative hierarchical clustering is adopted. First, calculate the Euclidean distance between clusters. The Euclidean distance is an index to measure the "straight-line distance" between two points in a multi-dimensional space, and the calculation formula is the square root of the sum of the squares of the coordinate differences of the two points. According to the Euclidean distance, start merging from the two closest clusters, gradually form larger clusters, and then set the clustering threshold to control the fineness of clustering to obtain an optimized molecular network map. This map more accurately reflects the relationship between compounds than the preliminary network structure, reducing the influence of noise and outliers.

[0028] Perform two-way retrieval and matching on the mass spectrometry features in the molecular network map with the known environmental pollutant database and the chemical registration database. Two-way retrieval means querying the matching items in the database starting from the mass spectrometry features and verifying the consistency with the mass spectrometry features starting from the database records. The matching process is carried out through three key indicators: molecular formula, molecular weight, and characteristic fragment ions. The molecular formula comparison checks whether the elemental composition is consistent, the molecular weight comparison checks whether the mass accuracy is within the allowable error range (usually 5 ppm), and the characteristic fragment ion comparison checks the presence of key structural fragments. Obtain the database matching results through comprehensive scoring, including the matching situation and matching degree score of each compound with the database records.

[0029] Compounds with a matching degree lower than 80% in the database matching results are marked. These compounds may be new pollutants or substances not included in the existing database. At the same time, screening and filtering are carried out in combination with the retention time index. The retention time index is a standardized representation of the retention time of a compound on the chromatographic column and has a strong indication of physicochemical properties. By comparing the actually measured retention time index with the theoretically predicted value or the retention time index of known compounds, the reliability of the matching results is further verified, false positives are excluded, and a list of potential new pollutants is formed. Isotope ratio analysis is carried out on the compounds in the list of potential new pollutants to calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine elements. Isotope ratio analysis uses the different isotopes existing in nature and their fixed abundance ratios of elements to infer the elements contained in the molecule. For example, carbon mainly exists in two isotopes, C-12 and C-13, with a natural abundance ratio of approximately 98.9:1.1, the abundance ratio of Cl-35 and Cl-37 of chlorine element is approximately 3:1, and the abundance ratio of Br-79 and Br-81 of bromine element is approximately 1:1. By analyzing the relative intensities of isotope peaks in the mass spectrum and combining with the constraints of element composition rules (such as the nitrogen rule: the molecular weight of a compound containing an even number of nitrogen atoms is even, and the molecular weight of a compound containing an odd number of nitrogen atoms is odd), the most likely molecular formula can be deduced to obtain molecular composition data.

[0030] Based on the molecular composition data, substructure analysis of characteristic fragment ions is carried out, and the fragment tree algorithm is applied to reconstruct the molecular skeleton. The fragment tree algorithm is a method to infer the molecular structure by constructing the hierarchical relationship of mass spectrometry fragments. All fragment ions are arranged in descending order of mass, the mass difference between adjacent fragments is calculated, and possible structural units (such as CH2, CO, NH, etc.) are matched. By connecting these structural units, the molecular skeleton is gradually reconstructed. Finally, the molecular structure is optimized by combining electron density calculation. Electron density calculation is based on quantum chemical theory, calculates the electron cloud distribution on each atom, and verifies the stability and rationality of the molecular structure, and finally obtains a candidate list of new pollutants.

[0031] For example, in the analysis of water body samples around a certain industrial area, a series of mass spectrometry peaks were detected by high-resolution mass spectrometry technology, and a mass spectrometry feature table containing 350 peaks was generated. The molecular network analysis technology was applied to calculate the similarity matrix between these peaks, and it was found that the similarity scores of three groups of peaks exceeded 0.7, respectively forming three preliminary clusters. Taking one of the clusters as an example, it contains 5 mass spectrometry peaks, and the m / z values are 315.0834, 317.0805, 319.0775, 315.0845, and 317.0816 respectively. The Euclidean distance between these peaks was calculated by the hierarchical clustering algorithm, and the distance threshold was set to 0.02. These 5 peaks were optimized into two sub-clusters: the first one contains three peaks with m / z of 315.0834, 317.0805, and 319.0775, and the second one contains two peaks with m / z of 315.0845 and 317.0816. The peak interval of the first sub-cluster is 2, and the intensity ratio is close to 3:1:0.1, which conforms to the isotope distribution characteristics of compounds containing two chlorine atoms. Comparing these characteristics with the environmental pollutant database, the record with the highest matching degree is the metabolite of a certain pesticide, but the matching degree is only 75%, lower than the threshold of 80%, so it is marked as a potential new pollutant. Further through isotope ratio analysis, it was confirmed that this compound does contain two chlorine atoms, and the possible molecular formula was deduced as C14H10Cl2O4. Applying the fragmentation tree algorithm to analyze its main fragment ions with m / z of 271.0968, 243.1019, and 215.0704, combined with electron density calculation, it was finally determined to be a derivative of a chlorophenoxy acid compound and added to the candidate list of new pollutants.

[0032] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Organize the multivariate environmental data according to the time dimension, space dimension, and pollutant characteristic dimension to construct a three-dimensional data matrix and obtain a spatio-temporal data cube; Apply the autoregressive integrated moving average algorithm to the pollutant concentration data in the spatio-temporal data cube for time series analysis, decompose the pollutant concentration change into trend components, seasonal components, and random components, and obtain time change characteristic data; Apply the geographically weighted regression and spatial autocorrelation analysis methods to the spatio-temporal data cube, calculate the global Moran index and local spatial autocorrelation indicators, identify the pollutant spatial aggregation areas and abnormal hotspots, and obtain spatial distribution characteristic data; Combine the spatial distribution characteristic data with the time change characteristic data, apply the variogram to analyze the spatial structure characteristics of the pollutant concentration, determine the spatial correlation distance and anisotropy characteristics, construct a spatial interpolation algorithm for the pollutant concentration, and obtain the pollution distribution heat map of the study area; Based on the pollution distribution heat map, combined with the hydrodynamic equation, the atmospheric diffusion equation, and the soil migration equation, calculate the transmission path parameters of pollutants in different environmental media, simulate the migration and diffusion process of pollutants in the environment, and obtain multi-media transmission dynamic data; Apply the backward trajectory analysis method to the multi-media transmission dynamic data. Taking the pollutant detection point as the end point, combined with the pollutant characteristic data in the fingerprint database, reverse infer the transmission starting point and key transmission nodes of the pollutants, and generate a pollutant transmission path map.

[0033] Specifically, organize the multivariate environmental data in three dimensions to construct a spatio-temporal data cube. The spatio-temporal data cube is a three-dimensional data structure that organizes environmental data according to the time dimension (including hourly, daily, weekly, monthly, and seasonal scales), the spatial dimension (including sampling point coordinates, administrative divisions, and geographical units), and the pollutant characteristic dimension (including concentration level, pollutant component structure, and fingerprint characteristics). The specific operation is to reconstruct the original environmental data table into a three-dimensional matrix, where each element of the matrix represents the concentration or characteristic value of a specific pollutant at a specific time point and a specific spatial position, forming a complete spatio-temporal data cube. Apply the autoregressive integrated moving average algorithm to the pollutant concentration data in the spatio-temporal data cube for time series analysis. The autoregressive integrated moving average algorithm, abbreviated as ARIMA, is a statistical model for processing non-stationary time series, which includes three parts: autoregressive term, differencing term, and moving average term. Specifically, when implementing, first conduct a stationarity test on the pollutant concentration time series. If it is not stationary, perform differencing processing to make it stationary; then determine the order of the model, including the autoregressive order, differencing order, and moving average order; then estimate the model parameters; finally, decompose the time series into a trend component (long-term change trend), a seasonal component (periodic change), and a random component (irregular fluctuation) through the model. This decomposition makes the time dynamic characteristics of the pollutant concentration clearer, forming time variation characteristic data, including the rising / falling trend of pollutant concentration, the intra-day / intra-week / seasonal fluctuation patterns, and the recognition results of sudden change events.

[0034] Apply the geographically weighted regression and spatial autocorrelation analysis methods to the spatio-temporal data cube. Geographically weighted regression is a regression analysis method that takes into account spatial non-stationarity. By assigning higher weights to observations that are closer in distance, local regression models are established at each geographical location to explore the spatial variation relationship between pollutant concentrations and environmental factors. Spatial autocorrelation analysis, on the other hand, studies the similarity between spatial units. Among them, the global Moran's I index is a statistic that quantifies the degree of spatial autocorrelation of the entire study area, with values ranging from -1 to 1. A positive value indicates the clustering of similar values, a negative value indicates the clustering of dissimilar values, and zero indicates a random distribution; local spatial autocorrelation indicators, such as the LISA statistic, identify local spatial clustering patterns, marking high-value clustering areas (hot spots), low-value clustering areas (cold spots), and spatial anomaly areas. Through these analyses, spatial distribution characteristic data are obtained to clarify the spatial distribution law of pollutants, especially the location and scope of the clustering areas and abnormal hot spots.

[0035] Combined with the spatial distribution characteristic data and the time-varying characteristic data, apply the variogram to analyze the spatial structure characteristics of pollutant concentrations. The variogram describes the relationship between the data difference between any two points in space and the distance. By calculating the semi-variance between sample point pairs at different distances and fitting the theoretical variogram model, the spatial correlation distance (range of influence) and anisotropic characteristics (differences in correlation in different directions) are obtained. Based on the results of the variogram analysis, a spatial interpolation algorithm for pollutant concentrations, such as Kriging interpolation, is constructed. This method is an optimal linear unbiased estimator that takes into account the spatial autocorrelation of sample points and generates a heat map of the pollution distribution in the study area to visually display the spatial distribution pattern of pollutants.

[0036] Based on the pollution distribution heat map, combined with the hydrodynamic equation, atmospheric diffusion equation and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated. The hydrodynamic equation describes the movement of pollutants in water bodies, mainly considering convection, diffusion and degradation processes; the atmospheric diffusion equation is based on the principle of conservation of mass, describing the diffusion behavior of pollutants in the atmosphere, considering factors such as wind direction, wind speed, and atmospheric stability; the soil migration equation considers factors such as soil properties and precipitation infiltration, and describes the vertical and horizontal migration of pollutants in the soil. By combining these equations, setting boundary conditions and initial conditions, and using numerical solution methods to calculate the migration and diffusion process of pollutants, multi-media transmission dynamic data are obtained, including the distribution of pollutant concentrations at each time point and each spatial position. The reverse trajectory analysis method is applied to the multi-media transmission dynamic data. Reverse trajectory analysis is a technology that reversely infers the transmission path of pollutants starting from the pollutant detection point. In specific implementation, starting with the pollutant detection point, the system combines transmission dynamics data with pollutant characteristic data (such as degradation rate and adsorption coefficient) and, based on concentration gradients and environmental conditions, gradually traces the pollutant's propagation path through various environmental media. Key transmission nodes (such as medium interfaces and flow direction change points) are identified, ultimately leading to the inference of possible pollution source areas. Through visualization, a pollutant transmission path map is generated, clearly showing the complete transmission path from the source to the detection point.

[0037] In a specific embodiment, the process of executing step S103 may specifically include the following steps: Perform feature standardization on multivariate environmental data, use feature selection algorithm to screen significant features, perform dimensionality reduction through principal component analysis, and generate environmental data source nodes; Apply feature encoding method to the pollutant fingerprint identification in the fingerprint library, convert mass spectrum characteristics, chromatographic characteristics and physical and chemical properties into numerical feature matrix, and generate fingerprint library source nodes through nonlinear dimensionality reduction mapping; The pollutant transmission path graph is processed through a graph convolutional network to parameterize the path nodes and connection relationships, extract the path topology features and transmission dynamics features, and generate the transmission path source nodes; A multi-branch neural network is constructed for the environmental data source nodes, fingerprint library source nodes, and transmission path source nodes. An attention mechanism fusion layer is designed to perform weighted integration of features. The network structure is optimized through cross-validation to obtain a fusion source node set. A multilayer perceptron network is constructed based on the fused source node set. The network consists of three hidden layers, each containing 128, 64, and 32 neurons, respectively. The activation function is the ReLU function, the output layer uses the Softmax function, and the loss function is the cross-entropy loss. The parameters are optimized using the Adam optimizer to obtain the source parsing classifier. Use a source parsing classifier to classify and predict the sources of pollutants, calculate the probability distribution of each pollution source category, calculate the contribution ratio of each source in combination with pollutant concentration data, construct a three-dimensional data structure containing source type, contribution rate, and geographical coordinates, and obtain a pollution source contribution rate matrix.

[0038] Specifically, perform feature standardization on the multivariate environmental data to make data with different dimensions comparable. Feature standardization includes methods such as min-max scaling (linearly transforming data to the [0,1] interval), Z-score standardization (subtracting the mean and dividing by the standard deviation), or percentile transformation. After standardization, use a feature selection algorithm to screen out significant features to reduce the data dimension and improve the model efficiency. Feature selection algorithms include methods based on statistical tests (such as chi-square test, F-test), model-based methods (such as feature importance scoring based on decision trees), and wrapper methods (using the performance of the target model as the evaluation criterion for feature subsets). The selected features are then reduced in dimension through principal component analysis. Principal component analysis transforms the original features into mutually orthogonal principal components through linear transformation, sorts them in descending order according to the variance ratio explained by the principal components, and selects the first few principal components as the new feature space, thus generating environmental data source nodes. Apply feature encoding methods to the pollutant fingerprint identifiers in the fingerprint library to convert non-numerical features into numerical forms that can be processed by machine learning algorithms. Specifically, convert mass spectrometry features (such as exact mass, isotope pattern, characteristic fragment ions, ion abundance ratio), chromatographic features (such as retention time, peak shape), and physical and chemical properties (such as hydrophilic / hydrophobic, acid-base, volatility) into a numerical feature matrix through one-hot encoding, label encoding, or embedding. Since the dimension of the converted features is usually very high, dimensionality reduction is required through non-linear dimensionality reduction mapping. Commonly used non-linear dimensionality reduction methods include t-SNE (t-distributed stochastic neighbor embedding) and UMAP (uniform manifold approximation and projection). These methods can preserve the local structure of high-dimensional data and better represent complex fingerprint feature relationships, and finally generate fingerprint library source nodes.

[0039] The pollutant transmission path map is processed through a graph convolutional network, which is a type of neural network specialized for processing graph-structured data. The transmission path map is represented as a mathematical graph structure, where nodes represent key points in space (such as monitoring stations, key path points), and edges represent the connection relationships between nodes (such as water flow direction, atmospheric transmission path). Then, the nodes and connection relationships in the graph are parametrically represented. Node features include position coordinates, pollutant concentration, etc., and edge features include direction, distance, transmission rate, etc. The graph convolutional network iteratively updates the node representations, aggregates the information of adjacent nodes, and extracts path topological features (such as connectivity, centrality) and transmission dynamics features (such as diffusion coefficient, decay rate), and finally generates a transmission path source node representing the entire transmission path structure. A multi-branch neural network is constructed for fusion of the environmental data source node, the fingerprint database source node, and the transmission path source node. The multi-branch neural network is a network structure that processes multiple different types of inputs in parallel. Each branch independently processes a type of data and then fuses at the backend of the network. In this method, the three source nodes are respectively input into three network branches for preliminary feature extraction, and then the three types of features are weighted and integrated through a designed attention mechanism fusion layer. The attention mechanism dynamically assigns weights to different features according to the importance of the current task, and pays more attention to the key features related to pollution source identification. The entire network structure is optimized through a cross-validation method. Cross-validation divides the dataset into a training set, a validation set, and a test set. Through multiple training-validation cycles, the network structure parameters with the best performance are selected, and finally a fused source node set is obtained.

[0040] A multi-layer perceptron network is constructed based on the fused source node set as a source apportionment classifier. The multi-layer perceptron is a feedforward neural network composed of an input layer, hidden layers, and an output layer. In this method, the network contains three hidden layers, with each layer containing 128, 64, and 32 neurons respectively, gradually decreasing in a pyramid structure, which helps to gradually extract high-level features. The ReLU activation function is used after each hidden layer. The ReLU function keeps positive inputs unchanged and sets negative inputs to zero, with advantages such as simple calculation and stable gradients. The Softmax function is used in the output layer to convert the outputs of multiple neurons into a probability distribution, where each value represents the probability that the sample belongs to a certain class. During the training process, the cross-entropy loss function is adopted to quantify the gap between the prediction result and the true label, and the learning rate is adaptively adjusted through the Adam optimizer to update the network parameters, and finally the source apportionment classifier is obtained.

[0041] Use a source analysis classifier to classify and predict the sources of pollutants, and calculate the probability distribution of each source category. Specifically, the pollutant sample data to be traced is processed through the aforementioned process, converted into a feature vector in the same format as the training data, and input into the trained source analysis classifier. The output layer gives the probability values of the sample belonging to each source category (such as industrial emissions, agricultural runoff, urban sewage, etc.). Combining with the pollutant concentration data, calculate the contribution ratio of each source, that is, the product of the probability value of each source and the corresponding pollutant concentration, and then perform normalization processing to obtain the contribution rate in percentage form. Finally, integrate the source type, contribution rate and geographical coordinate information (latitude and longitude, elevation, etc.) to construct a three-dimensional data structure containing these three dimensions, forming a source contribution rate matrix, providing basic data for subsequent key source location.

[0042] In a specific embodiment, the process of executing step S104 may specifically include the following steps: Perform three-dimensional splitting on the source contribution rate matrix, extract pollutant characteristic parameters, spatial coordinate information and time series data, construct a three-dimensional decision space, and obtain the basic data for source tracing decision-making; Based on the pollutant characteristic parameters in the basic data for source tracing decision-making, conduct chemical classification, molecular weight range division, functional group characteristic analysis and environmental degradation characteristic evaluation of the pollutants, construct a decision path in the pollutant characteristic dimension, and obtain the pollutant characteristic judgment result; Based on the spatial coordinate information in the basic data for source tracing decision-making, combine with the terrain and landform, water system distribution and administrative division, conduct spatial relationship analysis on the upstream area of the pollutant detection point, the distribution of surrounding potential pollution sources and the land use type, construct a decision path in the spatial position dimension, and obtain the spatial position judgment result; Based on the time series data in the basic data for source tracing decision-making, analyze the differences between weekdays and weekends of pollutant detection, seasonal change characteristics and time correlation with specific events, construct a decision path in the time characteristic dimension, and obtain the time characteristic judgment result; Input the pollutant characteristic judgment result, spatial position judgment result and time characteristic judgment result into the decision tree model, select the optimal splitting attribute by calculating the information gain ratio through the C4.5 algorithm, generate a source tracing inference chain, and obtain the candidate source area; Apply the grid refinement strategy to the candidate source area, set up temporary monitoring points, conduct fingerprint feature comparison and analysis on the collected environmental samples and potential pollution source samples, calculate the similarity score, determine the exact location of the pollution source, and obtain the pollution source location information.

[0043] Specifically, the pollution source contribution rate matrix is split into three dimensions. This matrix contains information on three dimensions: pollution source type, contribution rate, and geographical coordinates. The three-dimensional split means decomposing this complex data structure into three independent but interrelated data sets: pollutant characteristic parameters (including chemical properties, molecular structure characteristics, and environmental behavior data of pollutants), spatial coordinate information (including longitude, latitude, elevation, and regional boundary data), and time series data (including monitoring time points, seasonal cycles, and special event markers). These three types of data jointly construct a three-dimensional decision space, forming the basic data for source tracing decisions. The three-dimensional decision space can be regarded as a cube, where the x-axis represents pollutant characteristics, the y-axis represents spatial location, and the z-axis represents time. The position of each data point in space reflects the combined relationship of different dimensional characteristics. Based on the pollutant characteristic parameters in the source tracing decision basic data, multi-level pollutant property analysis is carried out. First is chemical classification, which classifies pollutants into categories such as organic compounds (such as polycyclic aromatic hydrocarbons, chlorinated hydrocarbons, esters) and inorganic compounds (such as heavy metals, nitrogen and phosphorus compounds) according to chemical structure and properties; then the molecular weight range is divided, and pollutants are divided into levels such as low molecular weight (<200), medium molecular weight (200 - 500), and high molecular weight (>500). Pollutants in different molecular weight ranges usually come from different types of emission activities; then the functional group characteristics are analyzed to identify the key functional groups (such as carboxyl groups, hydroxyl groups, amino groups, etc.) in the pollutant molecules. These functional groups are closely related to the source industries of the pollutants; finally, the environmental degradation characteristics are evaluated to analyze the persistence, bioaccumulation, and toxicity characteristics of pollutants in the environment. Pollutants from different sources usually have typical degradation characteristics. Combining these analysis results, a decision-making path for the pollutant characteristic dimension is constructed, forming a series of "if... then..." conditional judgment rules, such as "if the pollutant contains a specific combination of functional groups and the molecular weight is within a specific range, then it may come from a certain type of industrial activity", and finally the pollutant characteristic judgment result is obtained.

[0044] Analyze the topography and geomorphology, including elevation, slope, aspect, etc., which affect the diffusion path of pollutants; secondly, analyze the distribution of water systems, including river trends, lake distributions, groundwater flow directions, etc., as water systems are important carriers for pollutant migration; thirdly, consider administrative division factors, including the boundaries and jurisdiction scopes of different functional areas, which are directly related to the management of pollution sources. On this basis, focus on analyzing the upstream areas of pollutant detection points, and determine possible pollution source areas according to the water flow direction or air flow trajectory; at the same time, conduct a preliminary investigation on the distribution of surrounding potential pollution sources, including factories, farms, waste treatment facilities, etc.; in addition, consider the land use type, and there are obvious correlations between different land use natures (such as industrial land, agricultural land, residential land) and the emissions of specific pollutants. Through geographic information system methods such as spatial overlay analysis, buffer analysis, and shortest path analysis, construct a decision-making path in the dimension of spatial location to obtain a spatial location judgment result. Based on the time series data in the source tracing decision-making basic data, conduct multi-dimensional time pattern analysis. First, compare the differences in pollutant detection between weekdays and weekends. The concentration of industrial source pollutants is often higher on weekdays, and the concentration of domestic source pollutants may increase on weekends; secondly, analyze the seasonal change characteristics, such as the concentration of pollutants related to agricultural activities increases during specific farming seasons, and the pollutants related to heating increase in winter; thirdly, study the time correlation with specific events, such as the relationship between pollutant concentration changes and specific industrial production cycles, holidays, extreme weather events, etc. Through methods such as time series decomposition, pattern recognition, and association rule mining, construct a decision-making path in the dimension of time characteristics to obtain a time characteristic judgment result.

[0045] Take the judgment results of the previous three dimensions (pollutant characteristic judgment result, spatial position judgment result, and time characteristic judgment result) as inputs and feed them into a decision tree model for comprehensive analysis. A decision tree is a classification model with a tree structure that divides data into different categories through a series of questions. This method uses the C4.5 algorithm to construct a decision tree, which selects the optimal splitting attribute by calculating the information gain ratio. The information gain ratio measures the contribution of a certain attribute to classification and takes into account the splitting information of the attribute itself, avoiding bias towards attributes with more values. For each decision node, calculate the information gain ratio of all possible splitting attributes, select the attribute with the highest gain ratio as the splitting point, and recursively construct subtrees until the termination condition is met. Finally, form a complete decision tree, and each path from the root node to the leaf node forms a traceability inference chain. Different leaf nodes correspond to different possible pollution source areas. Combine the confidence scores of each inference chain to determine the candidate pollution source areas. Apply a grid refinement strategy to the determined candidate pollution source areas, divide the candidate areas into finer grid cells, and determine the grid size according to the regional characteristics and required positioning accuracy, usually gradually refining from several hundred meters to dozens of meters. Set up temporary monitoring points at key grid points to collect more targeted environmental samples, and at the same time collect direct samples of potential pollution sources as much as possible (such as sewage outlet water samples, flue gas samples, etc.). Conduct a detailed fingerprint feature analysis on the collected samples, extract molecular markers, elemental composition characteristics, etc., and compare and analyze them with the characteristics in the pollutant fingerprint library. Calculate the similarity score between samples. Commonly used similarity calculation methods include cosine similarity, Euclidean distance, Mahalanobis distance, etc. The higher the score, the greater the similarity between samples. By comparing the similarity between the samples at each point and the pollutants at the detection point, and combining the previous traceability decision results, finally determine the precise location of the pollution source of the pollutant and obtain the pollution source location information.

[0046] In a specific embodiment, the process of executing the step of inputting the pollutant characteristic judgment result, the spatial position judgment result, and the time characteristic judgment result into the decision tree model may specifically include the following steps: Merge the characteristics of the pollutant characteristic judgment result, the spatial position judgment result, and the time characteristic judgment result, construct a multi-dimensional decision data set containing all characteristics, and use a fully connected layer to combine the strongly correlated characteristics to obtain a decision attribute pool; Calculate the information entropy for each attribute in the decision attribute pool to quantify the uncertainty of the sample set. By calculating the probability of each type of sample appearing and performing a weighted sum, obtain an information entropy value set; Based on the information entropy value set, calculate the conditional entropy for each attribute in the decision attribute pool to quantify the classification ability of each attribute for the sample set. By calculating the entropy values of the sample subsets under each value of the attribute and performing a weighted sum, obtain a conditional entropy value set; According to the information entropy value set and the conditional entropy value set, calculate the information gain value of each attribute, and quantify the contribution degree of the attribute to classification by subtracting the conditional entropy from the original information entropy to obtain the information gain value set; Perform normalization processing on the information gain value set, calculate the information gain ratio, and balance the preference for multi-valued attributes by dividing the information gain by the split information value. Select the attribute with the largest information gain ratio as the split attribute of the current node to obtain the optimal split attribute sequence; According to the optimal split attribute sequence, construct a decision tree structure, recursively construct each branch node according to the attribute split rule until the termination condition is met, form a complete traceability decision tree, record the complete decision chain from the root node to the leaf node through the decision path tracking function, and obtain the candidate pollution source area.

[0047] Specifically, perform feature merging on the pollutant feature judgment result, the spatial position judgment result, and the time feature judgment result to form a complete decision basis. Feature merging integrates the judgment results of the three dimensions into the same data structure to construct a multi-dimensional decision data set containing all features. This process uses a fully connected layer to combine features with strong correlation. A fully connected layer is a neural network structure that can connect features from different sources into a unified vector representation. The specific operation is to represent the judgment results of the three dimensions as vectors. After being processed by the fully connected layer, features with strong correlation are given higher weights, while features with weak correlation have reduced weights. By this way, the feature combination is optimized, and finally a decision attribute pool is formed. The decision attribute pool contains multi-dimensional attributes such as "number of chlorine atoms in the pollutant", "water solubility level of the pollutant", "distance between the detection point and the industrial area", and "ratio of detection concentration on weekdays / weekends".

[0048] Calculate the information entropy for each attribute in the decision attribute pool to quantify the uncertainty of the sample set. Information entropy is a measure that describes the degree of chaos or uncertainty of a system. The larger the value, the more chaotic or uncertain the system is. When calculating specifically, first divide the samples in the decision attribute pool according to the pollution source category (such as industrial emissions, agricultural runoff, urban sewage, etc.), and count the proportion of each type of sample in the total samples as the probability of the occurrence of that category. Then calculate the logarithm of the probability of the occurrence of each type of sample multiplied by the probability itself, sum over all categories and take the negative to obtain the information entropy value of the sample set. Perform this calculation for each attribute in the decision attribute pool to form a set of information entropy values. This set reflects the difficulty of classifying the sample set without considering any attribute conditions. Based on the set of information entropy values, calculate the conditional entropy for each attribute in the decision attribute pool to quantify the classification ability of each attribute for the sample set. Conditional entropy describes the additional information required to classify the samples under the condition that the value of a certain attribute is known. The smaller the value, the greater the help of the attribute for classification. When calculating the conditional entropy, first divide the sample set into multiple subsets according to the different values of the attribute, calculate the information entropy of each subset, and then perform a weighted sum according to the proportion of the subset in the total samples to obtain the conditional entropy of the attribute. For example, for the attribute "whether the pollutant contains chlorine", divide the samples into two subsets of "containing chlorine" and "not containing chlorine", calculate the information entropy of these two subsets respectively, and then perform a weighted sum according to the sample proportion of the two subsets to obtain the conditional entropy of this attribute. Perform this process for each attribute in the decision attribute pool to form a set of conditional entropy values.

[0049] According to the set of information entropy values and the set of conditional entropy values, calculate the information gain value for each attribute. Information gain reflects the reduction in the uncertainty of the system after using a certain attribute to divide the samples. The larger the value, the greater the contribution of the attribute to classification. The method for calculating information gain is to subtract the conditional entropy from the original information entropy, and the obtained difference is the contribution of the attribute to reducing the system uncertainty. For example, if the information entropy of the original sample set is 0.9 and the conditional entropy after dividing by the attribute "whether the pollutant contains chlorine" is 0.3, then the information gain of this attribute is 0.6, indicating that after using this attribute to divide the samples, the uncertainty of the system is reduced by 0.6. Calculate the information gain for each attribute in the decision attribute pool to form a set of information gain values.

[0050] Normalize the set of information gain values and calculate the information gain ratio. The information gain ratio is a core concept of the C4.5 algorithm. It corrects the defect that the original information gain biases towards multi-valued attributes by dividing by the split information of the attribute itself. The split information refers to the degree of uniformity of the sample distribution after dividing the samples according to the attribute. The calculation method is similar to the information entropy, but the class probability is replaced by the sample proportion under each value. By dividing the information gain by the split information, the information gain ratio of each attribute is obtained. For example, if the information gain of "whether the pollutant contains chlorine" is 0.6 and the split information is 0.8, then its information gain ratio is 0.75; while "the type of functional group of the pollutant" has 7 values, the information gain is 0.65, and the split information is 1.8, then its information gain ratio is approximately 0.36. Compare the information gain ratios of all attributes and select the attribute with the largest gain ratio as the split attribute of the current node. Continue to recursively divide the sample set according to different attribute values to obtain the optimal split attribute sequence.

[0051] Construct a decision tree structure according to the optimal split attribute sequence. The decision tree consists of a root node, internal nodes, and leaf nodes. Each non-leaf node corresponds to a split attribute, each branch corresponds to a value of the attribute, and each leaf node corresponds to a classification result. The construction process starts from the root node, selects the attribute with the highest information gain ratio as the split point, divides the samples into each branch according to different values of this attribute, and then recursively executes the same split process for each branch until the termination condition is met. The termination conditions include: all samples belong to the same category, there are no more attributes available, and the number of samples in the branch node is less than the preset threshold. After the decision tree is constructed, record the complete decision chain from the root node to each leaf node through the decision path tracking function. Each decision chain represents a traceability inference path, corresponding to a possible pollution source type or area. Determine the final candidate pollution source area according to the distribution of samples in each leaf node and the class purity of the leaf node.

[0052] The above describes the method for tracing new pollutants in multiple environmental media based on machine learning in the embodiments of the present application. Next, the system for tracing new pollutants in multiple environmental media based on machine learning in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the system for tracing new pollutants in multiple environmental media based on machine learning in the embodiments of the present application includes: A construction module 201, configured to collect multivariate environmental data through a sensor network and construct a fingerprint database containing pollutant fingerprint identifiers; A reconstruction module 202, configured to perform spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint database and the multivariate environmental data to obtain a pollutant transmission path map; The fusion module 203 is used to convert the multi-source environmental data, the fingerprint database, and the pollutant transmission path map into source nodes, perform multi-source data fusion, construct a source analysis model, and obtain a pollution source contribution rate matrix. The positioning module 204 is used to perform key pollution source positioning based on the pollution source contribution rate matrix through the pollutant characteristic dimension, the spatial position dimension, and the time characteristic dimension, and obtain the pollution source location information.

[0053] Through the collaborative cooperation of the above-mentioned various components, by collecting multi-source environmental data and constructing a pollutant fingerprint database, the accurate identification and feature extraction of new pollutants are realized, thereby significantly improving the starting point accuracy of source tracing; by performing spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint database and multi-source environmental data, the migration and diffusion process of pollutants in the environment is visualized, solving the technical problem that it is difficult to trace the pollutant transmission path in a complex environment by traditional methods; converting the multi-source environmental data, the fingerprint database, and the pollutant transmission path map into quantum states and performing multi-source data fusion, fully exploring the deep associations between different types of data. Among them, quantum computing exhibits parallel processing capabilities and dimensionality reduction efficiency that are incomparable to traditional computing methods in multi-dimensional feature processing. This innovative data fusion strategy solves the problem of insufficient data integration by traditional methods; based on the pollution source contribution rate matrix, key pollution sources are located through a three-dimensional decision space, realizing step-by-step positioning from a large-scale area to an accurate geographical location, greatly improving the spatial resolution and accuracy of source tracing. In particular, the application of the C4.5 decision tree algorithm in pollution source discrimination objectively quantifies the importance of each decision attribute through the information gain ratio compared with traditional empirical judgments, making the decision-making process more scientific and transparent, and eliminating the uncertainty brought by subjective human judgment in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as graph convolutional networks for processing transmission paths, multi-branch neural networks for feature fusion, and decision trees for constructing source tracing inference chains, gives full play to the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion, and intelligent decision support, solving the problem of new pollutant source tracing that cannot be addressed by traditional methods, improving the source tracing efficiency, and reducing the dependence on expert experience.

[0054] Above Figure 2 The new pollutant source tracing system in multiple environmental media based on machine learning in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the new pollutant source tracing device in multiple environmental media based on machine learning in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0055] Figure 3FIG. 0 is a schematic structural diagram of a new pollutant tracing device in multiple environmental media based on machine learning provided by an embodiment of the present invention. The new pollutant tracing device 300 in multiple environmental media based on machine learning may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the new pollutant tracing device 300 in multiple environmental media based on machine learning. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the new pollutant tracing device 300 in multiple environmental media based on machine learning to implement the steps of the above-mentioned new pollutant tracing method in multiple environmental media based on machine learning.

[0056] The new pollutant tracing device 300 in multiple environmental media based on machine learning may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or, one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The shown structural diagram of the new pollutant tracing device in multiple environmental media based on machine learning does not limit the new pollutant tracing device in multiple environmental media provided by the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0057] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the new pollutant tracing method in multiple environmental media based on machine learning.

[0058] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0059] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a new pollutant source tracing device based on machine learning in a variety of environmental media (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0060] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for tracing new pollutants in multiple environmental media based on machine learning, characterized in that, The method includes: Collecting multivariate environmental data through a sensor network, and constructing a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data; Performing spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; Converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and performing multi-source data fusion to construct a source analysis model and obtain a source contribution rate matrix; Based on the source contribution rate matrix, performing key pollutant source location through pollutant characteristic dimension, spatial location dimension, and time characteristic dimension to obtain pollutant source location information.

2. The method for tracing new pollutants in multiple environmental media based on machine learning according to claim 1, wherein The step of collecting multivariate environmental data through a sensor network and constructing a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data includes: Deploying water quality monitoring sensors, air quality monitoring sensors, soil monitoring sensors, and biological response sensors to form a sensor network, and collecting environmental physical and chemical parameters through a low-power wide-area network to obtain raw environmental data; Identifying outliers in the raw environmental data using the three-sigma method, filling short-term missing data using linear interpolation, filling long-term missing data using multivariate interpolation, and performing noise reduction processing through wavelet transform to obtain preprocessed environmental data; Performing standardization processing on the preprocessed environmental data using the Z-score standardization method, and performing spatio-temporal matching and integration in combination with geographic information system data, meteorological data, human activity data, and historical monitoring records to obtain multivariate environmental data; Performing non-targeted analysis on environmental samples using high-resolution mass spectrometry technology, and performing peak identification, peak alignment, and peak area extraction on the mass spectrometry data to obtain a mass spectrometry feature table; Performing classification and clustering on the mass spectrometry feature table, and simultaneously performing database retrieval. For unrecorded compounds, performing structure analysis through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a list of new pollutant candidates; Extracting mass spectrometry features, chromatographic features, physical and chemical properties, and environmental behavior features for each pollutant in the list of new pollutant candidates, and performing feature optimization and screening using principal component analysis and recursive feature elimination algorithm to obtain a fingerprint library containing pollutant fingerprint identifiers.

3. The method for tracing new pollutants in various environmental media based on machine learning according to claim 2, characterized in that The step of performing classification and clustering on the mass spectrometry feature table, and simultaneously performing database retrieval. For unrecorded compounds, performing structure analysis through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a list of new pollutant candidates includes: Applying molecular network analysis technology to the mass spectrometry feature table to calculate the spectral similarity matrix between compounds, and classifying compounds with similarity scores greater than 0.7 into the same group according to the spectral similarity scores to obtain a preliminary molecular network structure; Applying the hierarchical clustering algorithm to the preliminary molecular network structure for structure optimization, calculating the Euclidean distance between groups, and setting a clustering threshold to obtain an optimized molecular network map; Performing two-way retrieval and matching on the mass spectrometry features in the molecular network map with a known environmental pollutant database and a chemical registration database, and obtaining a database matching result through comparison of molecular formula, molecular weight, and characteristic fragment ions; Mark the compounds with a matching degree lower than 80% in the database matching results, and perform screening and filtering in combination with the retention time index to obtain a list of potential new pollutants; Perform isotope ratio analysis on the compounds in the list of potential new pollutants, calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine elements, and deduce the molecular formula in combination with the elemental composition rules to obtain molecular composition data; Based on the molecular composition data, perform substructure analysis on the characteristic fragment ions, apply the fragment tree algorithm to reconstruct the molecular skeleton, and optimize the molecular structure in combination with electron density calculation to obtain a candidate list of new pollutants.

4. The method for tracing new pollutants in multiple environmental media based on machine learning according to claim 1, wherein, The spatio-temporal correlation analysis and transmission path reconstruction of the fingerprint library and the multi-source environmental data to obtain a pollutant transmission path map, including: Organize the multi-source environmental data according to the time dimension, space dimension, and pollutant characteristic dimension to construct a three-dimensional data matrix to obtain a spatio-temporal data cube; Apply the autoregressive integrated moving average algorithm to the pollutant concentration data in the spatio-temporal data cube for time series analysis, decompose the change of pollutant concentration into trend component, seasonal component, and random component to obtain time change characteristic data; Apply the geographically weighted regression and spatial autocorrelation analysis methods to the spatio-temporal data cube, calculate the global Moran index and local spatial autocorrelation index, identify the spatial aggregation area and abnormal hot spots of pollutants to obtain spatial distribution characteristic data; Combine the spatial distribution characteristic data with the time change characteristic data, apply the variogram to analyze the spatial structure characteristics of pollutant concentration, determine the spatial correlation distance and anisotropy characteristics, construct a spatial interpolation algorithm for pollutant concentration to obtain a pollution distribution heat map of the study area; Based on the pollution distribution heat map, combine the hydrodynamic equation, atmospheric diffusion equation, and soil migration equation to calculate the transmission path parameters of pollutants in different environmental media, simulate the migration and diffusion process of pollutants in the environment to obtain multi-media transmission dynamic data; Apply the backward trajectory analysis method to the multi-media transmission dynamic data, take the pollutant detection point as the end point, and combine the pollutant characteristic data in the fingerprint library to reverse the transmission starting point and key transmission nodes of pollutants to generate a pollutant transmission path map.

5. The method for tracing the source of new pollutants in various environmental media based on machine learning according to claim 1, wherein The conversion of the multi-source environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and the multi-source data fusion to construct a source analysis model to obtain a source contribution rate matrix, including: Perform feature standardization processing on the multi-source environmental data, use the feature selection algorithm to screen significant features, and perform dimensionality reduction processing through principal component analysis to generate an environmental data source node; Apply the feature encoding method to the pollutant fingerprint identification in the fingerprint library, convert the mass spectrometry features, chromatographic features, and physical and chemical properties into a numerical feature matrix, and generate a fingerprint library source node through non-linear dimensionality reduction mapping; Process the pollutant transmission path map through a graph convolutional network, parametrically represent the path nodes and connection relationships, extract the path topology features and transmission kinetic features to generate a transmission path source node; Construct a multi-branch neural network for the environmental data source node, the fingerprint database source node, and the transmission path source node. Design an attention mechanism fusion layer to weight and integrate features, and optimize the network structure through cross-validation to obtain a fused source node set. Construct a multi-layer perceptron network based on the fused source node set. The network contains three hidden layers, with each layer containing 128, 64, and 32 neurons respectively. The ReLU function is used as the activation function, the Softmax function is used in the output layer, and the cross-entropy loss is selected as the loss function. Parameter optimization is performed using the Adam optimizer to obtain a source parsing classifier. Use the source parsing classifier to classify and predict the pollutant sources, calculate the probability distribution of each pollutant source category, and calculate the contribution ratio of each source in combination with the pollutant concentration data. Construct a three-dimensional data structure containing source type, contribution rate, and geographical coordinates to obtain a pollutant source contribution rate matrix.

6. The method for tracing new pollutants in multiple environmental media based on machine learning according to claim 1, wherein, Based on the pollutant source contribution rate matrix, perform key pollutant source location through the pollutant feature dimension, spatial location dimension, and time feature dimension to obtain pollutant source location information, including: Perform three-dimensional splitting on the pollutant source contribution rate matrix, extract pollutant feature parameters, spatial coordinate information, and time series data, and construct a three-dimensional decision space to obtain the basic data for traceability decision-making. Based on the pollutant feature parameters in the basic data for traceability decision-making, perform chemical classification, molecular weight range division, functional group feature analysis, and environmental degradation characteristic evaluation on the pollutants, construct a decision-making path in the pollutant feature dimension, and obtain the pollutant feature judgment result. Based on the spatial coordinate information in the basic data for traceability decision-making, combine the topography, water system distribution, and administrative divisions to analyze the spatial relationship between the upstream area of the pollutant detection point, the distribution of surrounding potential pollutant sources, and the land use type, construct a decision-making path in the spatial location dimension, and obtain the spatial location judgment result. Based on the time series data in the basic data for traceability decision-making, analyze the differences between weekdays and weekends, seasonal change characteristics, and time correlation with specific events of pollutant detection, construct a decision-making path in the time feature dimension, and obtain the time feature judgment result. Input the pollutant feature judgment result, the spatial location judgment result, and the time feature judgment result into a decision tree model, select the optimal splitting attribute by calculating the information gain ratio through the C4.5 algorithm, generate a traceability inference chain, and obtain a candidate area for the pollutant source. Apply a grid refinement strategy to the candidate area for the pollutant source, set up temporary monitoring points, conduct fingerprint feature comparison and analysis on the collected environmental samples and potential pollutant source samples, calculate the similarity score, determine the precise location of the pollutant source, and obtain the pollutant source location information.

7. The method for tracing new pollutants in multiple environmental media based on machine learning according to claim 6, wherein Input the pollutant feature judgment result, the spatial location judgment result, and the time feature judgment result into a decision tree model, select the optimal splitting attribute by calculating the information gain ratio through the C4.5 algorithm, generate a traceability inference chain, and obtain a candidate area for the pollutant source, including: Perform feature merging on the pollutant feature judgment result, the spatial location judgment result, and the time feature judgment result to construct a multi-dimensional decision data set containing all features, and use a fully connected layer to combine features with strong correlation to obtain a decision attribute pool; Calculate the information entropy for each attribute in the decision attribute pool to quantify the uncertainty of the sample set, and obtain an information entropy value set by calculating the probability of each class of samples appearing and performing weighted summation; Based on the information entropy value set, calculate the conditional entropy for each attribute in the decision attribute pool to quantify the classification ability of each attribute for the sample set, and obtain a conditional entropy value set by calculating the entropy values of the sample subsets under each value of the attribute and performing weighted summation; According to the information entropy value set and the conditional entropy value set, calculate the information gain value for each attribute, and quantify the contribution degree of the attribute to classification by subtracting the conditional entropy from the original information entropy to obtain an information gain value set; Perform normalization processing on the information gain value set, calculate the information gain ratio, balance the preference for multi-valued attributes by dividing the information gain by the split information value, and select the attribute with the largest information gain ratio as the split attribute of the current node to obtain an optimal split attribute sequence; According to the optimal split attribute sequence, construct a decision tree structure, recursively construct each branch node according to the attribute split rule until the termination condition is met to form a complete traceability decision tree, and record the complete decision chain from the root node to the leaf node through the decision path tracking function to obtain the candidate source area of the pollution source.

8. A new pollutant tracing system in multiple environmental media based on machine learning, characterized in that, For implementing the method for tracing new pollutants in multiple environmental media based on machine learning as described in any one of claims 1-7, the system for tracing new pollutants in multiple environmental media based on machine learning includes: A construction module for collecting multivariate environmental data through a sensor network and constructing a fingerprint library containing pollutant fingerprint identifiers according to the multivariate environmental data; A reconstruction module for performing spatio-temporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; A fusion module for converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path map into source nodes and performing multi-source data fusion to construct a source analysis model to obtain a pollution source contribution rate matrix; A positioning module for performing key pollution source positioning based on the pollution source contribution rate matrix through the pollutant feature dimension, the spatial location dimension, and the time feature dimension to obtain pollution source location information.

9. A new pollutant tracing device in multiple environmental media based on machine learning, characterized in that, It includes a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements the method for tracing new pollutants in multiple environmental media based on machine learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, the processor is caused to execute the method for tracing new pollutants in multiple environmental media based on machine learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Polluted gas tracing method based on gas sensor array fingerprint identification

    CN112986497A

  • Atmospheric pollution tracing method and system based on neural network

    CN116881671A

  • Intelligent VOCs traceability and visualization method and system and storage medium

    CN118335229A

  • Water pollution intelligent traceability system based on three-dimensional fluorescence spectrum

    CN118883513A

  • Space-time big data fused drainage basin water quality pollution traceability analysis method and system

    CN119646471A

Cited By

  • Water and soil loss data analysis system based on isotopic tracing

    CN120779007A

  • New pollutant multi-medium migration path intelligent tracing method and system

    CN120869892A

  • Intelligent tracing method and system for multi-medium migration path of new pollutants

    CN120869892B

  • Mass spectrum gas source analysis method and system based on K-means clustering algorithm

    CN120929865A

  • Sea-land interaction zone groundwater pollution source tracing method based on multi-pollutant fingerprints

    CN121144894A