Method and system for tracing the source of new pollutants in various environmental media based on machine learning

By building a multi-source heterogeneous data fusion system for machine learning, the problems of insufficient data integration and transmission path simulation in the tracing of new pollutants have been solved, and the accurate identification of pollutants and improved tracing efficiency have been achieved, especially for tracing pollutants in complex environments.

CN120387598BActive Publication Date: 2025-09-19GUANGDONG INST OF ANALYSIS CHINA NAT ANALYTICAL CENT GUANGZHOU
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510885466.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-19
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies for tracing the source of new pollutants have problems such as difficulty in coping with changes in the chemical characteristics of organic pollutants in complex environments, insufficient data integration, lack of standardized intelligent decision-making support, and difficulty in simulating and tracking cross-media pollutant transmission, resulting in inaccurate and inefficient tracing results.

Method used

By building a multi-source heterogeneous data fusion system based on machine learning, using sensor networks to collect environmental data, building a pollutant fingerprint library, conducting spatiotemporal correlation analysis and transmission path reconstruction, combining graph convolutional networks and multi-branch neural networks for data fusion, locating key pollution sources based on the pollution source contribution rate matrix, and applying the C4.5 decision tree algorithm for source tracing reasoning.

Benefits of technology

It has achieved accurate identification and feature extraction of new pollutants, visualized the migration and diffusion process of pollutants, improved the spatial resolution and accuracy of tracing, reduced dependence on expert experience, and improved the efficiency and scientific nature of tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387598B_ABST
    Figure CN120387598B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of pollutant source tracing, and discloses a method and system for tracing the source of new pollutants in a variety of environmental media based on machine learning. The method includes: collecting multivariate environmental data to construct a pollutant fingerprint library; performing spatiotemporal correlation analysis on the fingerprint library and environmental data to reconstruct the transmission path; converting the data, fingerprint library, and path map into a source node fusion to construct an analytical model to form a contribution rate matrix; and locating the pollution source through three-dimensional features based on the matrix. In a complex environmental context, the present application can integrate multi-source heterogeneous data to construct an intelligent decision-making system based on machine learning, thereby realizing automated, standardized, and precise tracing of new pollutants and improving tracing efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pollutant source tracing, and in particular to a method and system for tracing the source of new pollutants in multiple environmental media based on machine learning. Background Art

[0002] In the field of environmental pollution prevention and control, tracing the source of new pollutants is a critical task, with important implications for the environmental quality of water, air, and soil in multiple environmental media. Traditional methods for tracing the source of pollutants rely primarily on manual empirical judgment and simple mathematical models, such as mass balance methods, diffusion models, and statistical correlation analysis. These methods typically target conventional pollutants such as heavy metals, nitrogen, and phosphorus. The basic process includes on-site sampling and analysis, pollution index determination, diffusion law calculation, and pollution source identification. On this basis, with the development of analytical technology, fingerprint recognition technology has begun to be applied to environmental pollution source tracing. By identifying the unique chemical characteristics and isotopic composition of pollutants, the accuracy of tracing has been improved. At the same time, the application of geographic information systems and remote sensing technology has also provided new technical support for large-scale pollution source identification, making the spatial distribution characteristics of regional pollutants more clear.

[0003] However, existing technologies still have obvious shortcomings in tracing the sources of new pollutants. First, traditional traceability methods are difficult to deal with new pollutants in complex environments, especially those organic pollutants that are degraded and transformed in the environment, whose chemical characteristics often change, resulting in inaccurate traceability judgments. Secondly, existing traceability technologies generally have the problem of insufficient data integration, and it is difficult to effectively integrate multi-source heterogeneous data, such as chemical analysis data, spatiotemporal distribution data, and human activity data, resulting in a lack of systematic traceability results. Furthermore, traditional traceability methods rely too much on expert experience and lack standardized and intelligent decision support tools, making the traceability process highly subjective, inefficient, and the results less reproducible. In addition, with the continuous increase in the types of new pollutants in the environment, the existing pollutant fingerprint library is seriously lagging behind and cannot meet the needs of rapid identification and accurate traceability of new pollutants. Finally, existing traceability technologies have technical bottlenecks in the simulation and tracking of cross-media pollutant transmission, making it difficult to accurately reconstruct the complete transmission path of pollutants from the source to the detection point. Summary of the Invention

[0004] This application provides a method and system for tracing the source of new pollutants in multiple environmental media based on machine learning. It is used to integrate multi-source heterogeneous data in a complex environmental context, build an intelligent decision-making system based on machine learning, realize the automated, standardized and precise tracing of new pollutants, and improve the efficiency and accuracy of tracing.

[0005] In the first aspect, the present application provides a method for tracing the source of new pollutants in multiple environmental media based on machine learning, and the method for tracing the source of new pollutants in multiple environmental media based on machine learning includes: collecting multivariate environmental data through a sensor network, and constructing a fingerprint library containing pollutant fingerprint identifiers based on the multivariate environmental data; performing spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; converting the multivariate environmental data, the fingerprint library and the pollutant transmission path map into source nodes and performing multi-source data fusion to construct a source parsing model to obtain a pollution source contribution rate matrix; based on the pollution source contribution rate matrix, key pollution sources are located through pollutant characteristic dimensions, spatial location dimensions and time characteristic dimensions to obtain pollution source location information.

[0006] In a second aspect, the present application provides a system for tracing the source of new pollutants in multiple environmental media based on machine learning, the system for tracing the source of new pollutants in multiple environmental media based on machine learning includes:

[0007] A construction module is used to collect multivariate environmental data through a sensor network and construct a fingerprint library containing pollutant fingerprint identification based on the multivariate environmental data;

[0008] A reconstruction module, configured to perform spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map;

[0009] A fusion module, configured to convert the multivariate environmental data, the fingerprint library, and the pollutant transmission path diagram into source nodes and perform multi-source data fusion, construct a source resolution model, and obtain a pollution source contribution rate matrix;

[0010] The positioning module is used to locate key pollution sources based on the pollution source contribution rate matrix through the pollutant characteristic dimension, spatial location dimension and time characteristic dimension to obtain pollution source location information.

[0011] In the third aspect, a device for tracing the source of new pollutants in multiple environmental media based on machine learning is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the device for tracing the source of new pollutants in multiple environmental media based on machine learning to execute the above-mentioned method for tracing the source of new pollutants in multiple environmental media based on machine learning.

[0012] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned method for tracing the source of new pollutants in multiple environmental media based on machine learning.

[0013] In the technical solution provided by this application, accurate identification and feature extraction of new pollutants are achieved through multivariate environmental data collection and construction of a pollutant fingerprint library, thereby significantly improving the starting point accuracy of traceability; spatiotemporal correlation analysis and transmission path reconstruction of the fingerprint library and multivariate environmental data are performed to visualize the migration and diffusion process of pollutants in the environment, solving the technical problem that traditional methods are difficult to track the transmission paths of pollutants in complex environments; multivariate environmental data, fingerprint library and pollutant transmission path diagram are converted into quantum states and multi-source data fusion is performed, fully exploring the deep correlation between different types of data, among which quantum computing demonstrates parallel processing capabilities and dimensionality reduction efficiency that traditional computing methods cannot match in multidimensional feature processing. This innovative data fusion strategy solves the problem of insufficient data integration of traditional methods; based on the pollution source contribution rate matrix, the three-dimensional decision space is used to make the decision. The key pollution source positioning has achieved gradual positioning from large-scale areas to precise geographical locations, greatly improving the spatial resolution and accuracy of tracing. In particular, the application of the C4.5 decision tree algorithm in pollution source identification, compared with traditional empirical judgment, objectively quantifies the importance of each decision attribute through the information gain ratio, making the decision-making process more scientific and transparent, and eliminating the uncertainty brought about by human subjective judgment in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as graph convolutional network processing transmission paths, multi-branch neural network realization of feature fusion, and decision tree construction of tracing reasoning chain, fully utilizes the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion and intelligent decision support, solves the problem of new pollutant tracing that traditional methods cannot cope with, improves tracing efficiency, and reduces dependence on expert experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0015] Figure 1 This is a schematic diagram of an embodiment of a method for tracing the source of new pollutants in multiple environmental media based on machine learning in an embodiment of the present application;

[0016] Figure 2 This is a schematic diagram of an embodiment of a system for tracing the source of new pollutants in multiple environmental media based on machine learning in an embodiment of the present application;

[0017] Figure 3 This is a schematic block diagram of the structure of a device for tracing the source of new pollutants in various environmental media based on machine learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The embodiments of the present application provide a method and system for tracing the source of new pollutants in a variety of environmental media based on machine learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.

[0019] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of a method for tracing the source of new pollutants in multiple environmental media based on machine learning includes:

[0020] Step S101: collecting multivariate environmental data through a sensor network, and constructing a fingerprint library containing pollutant fingerprint identifiers based on the multivariate environmental data;

[0021] Step S102: performing spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and multivariate environmental data to obtain a pollutant transmission path map;

[0022] Step S103: convert the multivariate environmental data, fingerprint database, and pollutant transmission path diagram into source nodes and perform multi-source data fusion to build a source resolution model and obtain a pollution source contribution rate matrix;

[0023] Step S104: Based on the pollution source contribution rate matrix, key pollution sources are located through the pollutant characteristic dimension, spatial location dimension, and time characteristic dimension to obtain pollution source location information.

[0024] It is understood that the execution subject of this application can be a new pollutant tracing system in various environmental media based on machine learning, or a terminal or server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.

[0025] Specifically, multivariate environmental data is collected through a sensor network and a pollutant fingerprint library is constructed. In this step, a multivariate sensor network including water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors and biological response sensors is deployed, and the physical and chemical parameters in the environment are collected in real time through low-power wide area network technology. In the embodiment of the present application, multi-environmental media refers to water bodies, soil, and atmosphere. Therefore, the present application scheme is also for tracing the source of new pollutants in water bodies, soil, and atmosphere. The original environmental data is identified by the triple standard deviation method for outliers, and the linear interpolation method is used to fill the short-term missing data. The multivariate interpolation method is used for long-term missing data, and then the noise is reduced by wavelet transform to finally obtain pre-processed environmental data. The pre-processed environmental data is standardized using the Z-score standardization method, and is spatially and temporally matched and integrated with geographic information system data, meteorological data, human activity data and historical monitoring records to form multivariate environmental data. At the same time, high-resolution mass spectrometry technology is used to perform non-targeted analysis on environmental samples, and a mass spectrum feature table is obtained through peak recognition, peak alignment and peak area extraction. The mass spectral signature table was classified and clustered, and database searches were performed. Structural elucidation of unlisted compounds was performed through isotope ratio analysis, fragment ion analysis, and retention time prediction to generate a candidate list of new pollutants. Finally, the mass spectral, chromatographic, physicochemical, and environmental behavior characteristics of each new pollutant were extracted. Principal component analysis and recursive feature elimination were used to optimize and screen the features and construct a pollutant fingerprint library.

[0026] Perform spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint database and multivariate environmental data. In this step, the multivariate environmental data is organized according to the time dimension, space dimension and pollutant characteristic dimension to construct a three-dimensional data cube. The autoregressive integral moving average model is applied to analyze the time series data of pollutant concentrations, decomposing them into trend components, seasonal components and random components to obtain time-varying characteristic data. Then, through geographically weighted regression and spatial autocorrelation analysis techniques, the global Moran index and local spatial autocorrelation index are calculated to identify the spatial aggregation areas and abnormal hotspots of pollutants and generate spatial distribution characteristic data. Combining the time-varying characteristic data and spatial distribution characteristic data, the variation function is applied to analyze the spatial structural characteristics of pollutant concentrations, determine the spatial correlation distance and anisotropy characteristics, construct a spatial interpolation algorithm for pollutant concentrations, and generate a pollution distribution heat map for the study area. Based on the pollution distribution heat map, combined with the hydrodynamic equation, atmospheric diffusion equation and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated, and their migration and diffusion process is simulated to obtain multi-media transmission dynamic data. Finally, the reverse trajectory analysis technology is applied, with the pollutant detection point as the end point, combined with the pollutant characteristic data in the fingerprint library, to reversely infer the pollutant transmission starting point and key transmission nodes, and generate a pollutant transmission path map.

[0027] Multivariate environmental data, fingerprint databases, and pollutant transmission pathways are converted into source nodes, and multi-source data fusion is performed. Feature normalization is performed on the multivariate environmental data, and a feature selection algorithm is used to screen significant features. Dimensionality reduction is performed through principal component analysis to generate environmental data source nodes. Feature encoding methods are applied to the pollutant fingerprints in the fingerprint database, converting mass spectral features, chromatographic features, and physicochemical properties into numerical feature matrices. Fingerprint database source nodes are generated through nonlinear dimensionality reduction. The pollutant transmission pathway diagram is processed using a graph convolutional network, parameterizing the pathway nodes and connectivity, extracting pathway topological features and transmission dynamics, and generating transmission pathway source nodes. A multi-branch neural network is constructed for the environmental data source nodes, fingerprint database source nodes, and transmission pathway source nodes. An attention mechanism fusion layer is designed to weightedly integrate features. The network structure is optimized through cross-validation to obtain a fused source node set. A multi-layer perceptron network is constructed based on the fused source node set, consisting of three hidden layers with 128, 64, and 32 neurons, respectively. ReLU functions are used as the activation function, Softmax functions are used in the output layer, and cross-entropy loss is selected as the loss function. Parameters are optimized using the Adam optimizer to obtain a source parsing classifier. The source apportionment classifier is used to classify and predict the sources of pollutants, calculate the probability distribution of each pollution source category, and calculate the contribution ratio of each source in combination with the pollutant concentration data. A three-dimensional data structure containing source type, contribution rate and geographic coordinates is constructed to obtain the pollution source contribution rate matrix.

[0028] Based on the pollution source contribution rate matrix, key pollution sources are located through the pollutant characteristic dimension, spatial location dimension, and time characteristic dimension. This step performs a three-dimensional split on the pollution source contribution rate matrix, extracts pollutant characteristic parameters, spatial coordinate information, and time series data, constructs a three-dimensional decision space, and obtains the basic data for tracing decision-making. Based on the pollutant characteristic parameters in the basic data for tracing decision-making, the pollutants are chemically classified, molecular weight range divided, functional group characteristics analyzed, and environmental degradation characteristics evaluated. A decision path for the pollutant characteristic dimension is constructed to obtain the pollutant characteristic judgment result. Based on spatial coordinate information, combined with topography, water system distribution, and administrative divisions, a spatial relationship analysis is conducted on the upstream area of ​​the pollutant detection point, the distribution of surrounding potential pollution sources, and land use types. A decision path for the spatial location dimension is constructed to obtain the spatial location judgment result. Based on time series data, the differences in pollutant detection between weekdays and weekends, seasonal variation characteristics, and temporal correlation with specific events are analyzed. A decision path for the time characteristic dimension is constructed to obtain the time characteristic judgment result. The results of pollutant feature, spatial location, and temporal feature judgments were input into a decision tree model. The C4.5 algorithm was used to calculate the information gain ratio for optimal splitting attribute selection, generating a traceability reasoning chain to identify candidate pollution source areas. Finally, a grid refinement strategy was applied to the candidate pollution source areas, temporary monitoring points were set up, and fingerprint features of collected environmental samples were compared with potential pollution source samples. Similarity scores were calculated to determine the precise location of the pollution source and obtain pollution source location information.

[0029] In the embodiments of the present application, accurate identification and feature extraction of new pollutants are achieved through multivariate environmental data collection and construction of a pollutant fingerprint library, thereby significantly improving the starting point accuracy of traceability; spatiotemporal correlation analysis and transmission path reconstruction of the fingerprint library and multivariate environmental data are performed to visualize the migration and diffusion process of pollutants in the environment, solving the technical problem that traditional methods are difficult to track the transmission paths of pollutants in complex environments; multivariate environmental data, fingerprint library and pollutant transmission path diagram are converted into quantum states and multi-source data fusion is performed, fully exploring the deep correlation between different types of data, among which quantum computing demonstrates parallel processing capabilities and dimensionality reduction efficiency that traditional computing methods cannot match in multidimensional feature processing. This innovative data fusion strategy solves the problem of insufficient data integration of traditional methods; based on the pollution source contribution rate matrix, key pollutants are identified through three-dimensional decision space. Pollution source positioning has achieved gradual positioning from large-scale areas to precise geographical locations, greatly improving the spatial resolution and accuracy of tracing. In particular, the application of the C4.5 decision tree algorithm in pollution source identification, compared with traditional empirical judgment, objectively quantifies the importance of each decision attribute through the information gain ratio, making the decision-making process more scientific and transparent, and eliminating the uncertainty brought about by human subjective judgment in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as graph convolutional network processing transmission paths, multi-branch neural network realization of feature fusion, and decision tree construction of tracing reasoning chain, fully utilizes the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion and intelligent decision support, solves the problem of new pollutant tracing that traditional methods cannot cope with, improves tracing efficiency, and reduces dependence on expert experience.

[0030] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0031] Deploy water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors, and biological response sensors to form a sensor network, collect environmental physical and chemical parameters through a low-power wide-area network, and obtain raw environmental data;

[0032] The triple standard deviation method was applied to the original environmental data to identify outliers, linear interpolation was used to fill short-term missing data, multivariate interpolation was used to fill long-term missing data, and wavelet transform was used to perform noise reduction to obtain preprocessed environmental data;

[0033] The pre-processed environmental data were standardized using the Z-score standardization method, and then spatial and temporal matching was performed to integrate the geographic information system data, meteorological data, human activity data, and historical monitoring records to obtain multivariate environmental data.

[0034] High-resolution mass spectrometry technology is used to perform non-targeted analysis of environmental samples. The mass spectrometry data are processed by peak identification, peak alignment and peak area extraction to obtain a mass spectrum feature table.

[0035] The mass spectral feature table is classified and clustered, and a database search is performed simultaneously. The structures of unlisted compounds are elucidated through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a candidate list of new pollutants.

[0036] For each pollutant in the new pollutant candidate list, mass spectrometry characteristics, chromatographic characteristics, physicochemical properties and environmental behavior characteristics are extracted, and principal component analysis and recursive feature elimination algorithm are used to perform feature optimization screening to obtain a fingerprint library containing pollutant fingerprint identification.

[0037] Specifically, a multi-sensor network is deployed to collect environmental data. The sensor network includes water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors, and bio-response sensors. Water quality monitoring sensors are deployed at surface water bodies, groundwater sampling points, and sewage outlets to monitor water quality parameters such as pH, dissolved oxygen, total organic carbon, and heavy metal content in real time. Atmospheric monitoring sensors are deployed around industrial parks, urban transportation nodes, and sensitive areas to monitor atmospheric pollutants such as volatile organic compounds and particulate matter. Soil monitoring sensors are deployed in a grid-based manner across the study area, covering different land types, to monitor persistent organic pollutants in the soil. Bio-response sensors select indicator organisms sensitive to specific pollutants to capture pollutant accumulation within organisms. These sensors transmit data via low-power wide-area network technology, enabling real-time collection of environmental physical and chemical parameters and generating raw environmental data. It should be noted that in the embodiments of this application, water environment monitoring can also be achieved through on-site automatic online monitoring, drone remote sensing monitoring, unmanned boat mobile monitoring, and laboratory testing. Soil condition monitoring can also be achieved through on-site rapid testing, in-situ testing, drone remote sensing monitoring, and laboratory testing.

[0038] The triple standard deviation method is applied to the data to identify outliers. This method calculates the mean and standard deviation of the dataset and marks data points that deviate from the mean by more than three standard deviations as outliers. Outliers verified to be caused by equipment failure or environmental changes are corrected or removed. For missing data, different interpolation methods are used depending on the duration of the missing data. For short-term missing data (e.g., data missing within a few hours), linear interpolation is used to fill the missing value using a linear function calculated from valid data before and after the missing point. For long-term missing data (e.g., data missing for several days), multivariate interpolation methods based on historical data patterns are used, taking into account the historical variation patterns of multiple related variables to generate more realistic interpolated values. The data is then subjected to wavelet transform for noise reduction. By decomposing the signal into wavelet coefficients of different frequencies, high-frequency random noise is removed while retaining valid information. The data after these processes is referred to as preprocessed environmental data.

[0039] Preprocessed environmental data requires standardization using the Z-score standardization method. This involves subtracting the mean of each variable from each data point and dividing by the standard deviation to convert environmental parameters of varying dimensions to a unified standard. The standardized data is then spatiotemporally aligned with auxiliary data sources, including geographic information system data (topography, water system distribution, geological structure, and land use type), meteorological data (precipitation, wind direction and speed, temperature, and humidity), human activity data (industrial layout, agricultural activity, urbanization, and traffic flow), and historical monitoring records. Optionally, auxiliary data sources can include geographic information data, meteorological data, industrial and agricultural production data, hydrological data (water bodies), land use information (soil), and human activity data (population and activity types) for pollution-related areas.

[0040] These auxiliary data are collected and integrated through the data interface, and are matched with the sensor network data in time and space to construct complete multivariate environmental data. High-resolution mass spectrometry technology is used to perform non-targeted analysis of environmental samples, including liquid chromatography-quadrupole-time of flight mass spectrometry, gas chromatography-quadrupole-time of flight mass spectrometry and ultra-high performance liquid chromatography-mass spectrometry. Mass spectrometry data is subjected to peak recognition (identifying the real peaks in the chromatogram and eliminating noise), peak alignment (matching the corresponding chromatographic peaks in different samples) and peak area extraction (calculating the area of ​​each peak for quantitative analysis) through professional analysis software to generate a mass spectrum feature table. Optionally, in an embodiment of the present application, data analysis can also be performed on the characteristic spectrum data and the corresponding signal intensity data.

[0041] The mass spectral signature table is classified and clustered, and database searches are performed. Molecular network analysis is used for clustering, categorizing compounds into clusters based on similarities in their mass spectral signatures. A unique mass spectral fingerprint is generated for each pollutant class. Identified compounds are then compared to databases of known environmental pollutants and chemical registration databases. Compounds not listed in existing databases are flagged as potential new pollutants. Structural elucidation of these potential new pollutants is performed, using isotope ratio analysis (inferring molecular composition based on the isotopic abundance distribution of the analyzed elements), fragment ion analysis (inferring molecular structure based on mass spectrometric fragmentation patterns), and retention time prediction (predicting chromatographic column retention based on compound structural features) to derive their molecular formulas and possible chemical structures, resulting in a candidate list of new pollutants. For each pollutant in the candidate list, features are extracted and a fingerprint library is constructed. The extracted features include mass spectral characteristics (accurate mass, isotope pattern, characteristic fragment ions, ion abundance ratio), chromatographic characteristics (retention time, peak shape), physicochemical properties (hydrophilicity / hydrophobicity, acidity / alkalinity, volatility), and environmental behavior characteristics (degradation rate, bioaccumulation). These features are optimized and screened using principal component analysis and recursive feature elimination. Principal component analysis reduces data dimensionality by converting the original features into linearly independent new features (principal components), retaining the most informative feature combinations. Recursive feature elimination, on the other hand, repeatedly builds models, assesses feature importance, and removes the least important features to select the most representative and discriminative feature set. These selected features constitute the final contaminant fingerprint, which is stored in a fingerprint library.

[0042] In a specific embodiment, the process of performing the step of classifying and clustering the mass spectrum feature table may specifically include the following steps:

[0043] Molecular network analysis technology was applied to the mass spectral feature table to calculate the spectral similarity matrix between compounds. Compounds with a similarity score greater than 0.7 were classified into the same group based on the spectral similarity score to obtain a preliminary molecular network structure.

[0044] Apply the hierarchical clustering algorithm to optimize the preliminary molecular network structure, calculate the Euclidean distance between clusters, set the clustering threshold, and obtain the optimized molecular network map;

[0045] The mass spectrometry features in the molecular network map were bidirectionally searched and matched with the known environmental pollutant database and the chemical registration database, and the database matching results were obtained by comparing the molecular formula, molecular weight and characteristic fragment ions;

[0046] Compounds with a matching degree of less than 80% in the database matching results were marked and filtered based on the retention time index to obtain a list of potential new pollutants;

[0047] Conduct isotope ratio analysis on compounds in the list of potential new pollutants, calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine, and deduce molecular formulas based on elemental composition constraints to obtain molecular composition data;

[0048] Based on the molecular composition data, the characteristic fragment ions are subjected to substructure analysis, the molecular skeleton is reconstructed using the fragment tree algorithm, and the molecular structure is optimized in combination with electron density calculation to obtain a candidate list of new pollutants.

[0049] Specifically, molecular network analysis technology is applied to the mass spectral feature table to calculate the spectral similarity matrix between compounds. Molecular network analysis technology is a method that uses spectral similarity to classify related compounds. Its core is to calculate the cosine similarity between mass spectra of different compounds. The specific operation is to regard the mass spectrum of each compound as a high-dimensional vector, with the mass-to-charge ratio of the mass spectral peak as the dimension and the peak intensity as the value on this dimension, and then calculate the cosine similarity of the vectors between each two compounds. The calculation result of cosine similarity ranges from 0 to 1. The closer the value is to 1, the more similar the two spectra are. In this method, the threshold is set to 0.7, and compounds with similarity scores greater than this threshold are classified into the same group to form a preliminary molecular network structure. This network structure intuitively shows the similarity relationship between compounds, and similar compounds are clustered in the network. A hierarchical clustering algorithm is applied to the preliminary molecular network structure for structural optimization. A hierarchical clustering algorithm is a method that gradually merges or splits data points to form a hierarchical structure. This method employs a bottom-up agglomerative hierarchical clustering approach. First, the Euclidean distance between clusters is calculated. Euclidean distance measures the "straight-line distance" between two points in multidimensional space and is calculated as the square root of the sum of the squares of their coordinate differences. Based on this Euclidean distance, clusters are merged, starting with the two closest ones, to gradually form larger clusters. A clustering threshold is then set to control the refinement of the clustering, resulting in an optimized molecular network map. This map more accurately reflects the relationships between compounds than the initial network structure, reducing the influence of noise and outliers.

[0050] The mass spectral features in the molecular network map are bidirectionally searched and matched against databases of known environmental pollutants and chemical registration databases. Bidirectional search involves both searching the database for matches based on the mass spectral features and verifying consistency with the database records. The matching process compares three key indicators: molecular formula, molecular weight, and characteristic fragment ions. Molecular formula comparison verifies elemental composition consistency, molecular weight comparison verifies mass accuracy within the tolerance (typically 5 ppm), and characteristic fragment ion comparison verifies the presence of key structural fragments. A comprehensive score is then generated to generate database matching results, including the match status and matching score for each compound against the database record.

[0051] Compounds with a database match score below 80% are flagged as potential new contaminants or substances not included in existing databases. Retention time index (RTI) is also used for screening and filtering. The RTI is a standardized representation of a compound's retention time on the chromatographic column and a strong indicator of its physicochemical properties. By comparing the measured RTI with theoretical predictions or those of known compounds, the reliability of the matching results is further verified, false positives are eliminated, and a list of potential new contaminants is compiled. Isotope ratio analysis is performed on the compounds on the potential new contaminant list to calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine. Isotope ratio analysis uses the different naturally occurring isotopes of an element and their fixed abundance ratios to infer the elements present in a molecule. For example, carbon primarily exists as two isotopes, C-12 and C-13, with a natural abundance ratio of approximately 98.9:1.1. Chlorine has an abundance ratio of Cl-35 to Cl-37 of approximately 3:1, and bromine has an abundance ratio of Br-79 to Br-81 of approximately 1:1. By analyzing the relative intensities of the isotope peaks in the mass spectrum and combining them with the constraints of elemental composition rules (such as the nitrogen rule: compounds containing an even number of nitrogen atoms have an even molecular weight, and compounds containing an odd number of nitrogen atoms have an odd molecular weight), the most likely molecular formula can be deduced and molecular composition data can be obtained.

[0052] Based on molecular composition data, characteristic fragment ions are subjected to substructure analysis, and the molecular backbone is reconstructed using a fragment tree algorithm. This algorithm infers molecular structure by constructing a hierarchical relationship among mass spectrometric fragments. All fragment ions are arranged from highest to lowest mass, the mass difference between adjacent fragments is calculated, and possible structural units (such as CH2, CO, and NH) are matched. By connecting these structural units, the molecular backbone is gradually reconstructed. Finally, the molecular structure is optimized using electron density calculations. Based on quantum chemistry theory, electron density calculations calculate the electron cloud distribution on each atom, verifying the stability and rationality of the molecular structure, and ultimately generating a candidate list of new pollutants.

[0053] For example, in an analysis of water samples from an industrial area, high-resolution mass spectrometry detected a series of mass spectral peaks, generating a mass spectral signature table containing 350 peaks. Molecular network analysis techniques were used to calculate a similarity matrix between these peaks, revealing that three groups of peaks had similarity scores exceeding 0.7, forming three preliminary clusters. For example, one of these clusters contained five mass spectral peaks with m / z values ​​of 315.0834, 317.0805, 319.0775, 315.0845, and 317.0816. Using a hierarchical clustering algorithm, the Euclidean distances between these peaks were calculated, setting a distance threshold of 0.02. These five peaks were then refined into two subclusters: the first containing three peaks at m / z 315.0834, 317.0805, and 319.0775, and the second containing two peaks at m / z 315.0845 and 317.0816. The peak interval of the first subgroup is 2, and the intensity ratio is close to 3:1:0.1, which is consistent with the isotope distribution characteristics of compounds containing two chlorine atoms. When these characteristics were compared with the environmental pollutant database, the record with the highest matching degree was a metabolite of a pesticide, but the matching degree was only 75%, which was lower than the threshold of 80%, so it was marked as a potential new pollutant. Further isotope ratio analysis confirmed that the compound does contain two chlorine atoms, and the possible molecular formula was deduced to be C14H10Cl2O4. The fragment tree algorithm was used to analyze its main fragment ions m / z as 271.0968, 243.1019 and 215.0704. Combined with electron density calculations, it was finally determined to be a derivative of a chlorinated phenoxy acid compound and added to the list of new pollutant candidates.

[0054] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0055] Organize multivariate environmental data according to time dimension, space dimension and pollutant characteristic dimension, construct a three-dimensional data matrix and obtain a spatiotemporal data cube;

[0056] The autoregressive integrated moving average algorithm is applied to the pollutant concentration data in the spatiotemporal data cube to perform time series analysis, decomposing the pollutant concentration changes into trend components, seasonal components, and random components to obtain time variation characteristic data;

[0057] Applying geographically weighted regression and spatial autocorrelation analysis methods to the spatiotemporal data cube, the global Moran index and local spatial autocorrelation index are calculated to identify spatial concentration areas and abnormal hot spots of pollutants and obtain spatial distribution characteristic data;

[0058] Combining spatial distribution characteristic data with temporal variation characteristic data, the variation function is applied to analyze the spatial structure characteristics of pollutant concentrations, determine the spatial correlation distance and anisotropy characteristics, and construct a spatial interpolation algorithm for pollutant concentrations to obtain a pollution distribution heat map of the study area;

[0059] Based on the pollution distribution heat map, combined with the hydrodynamic equation, atmospheric diffusion equation and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated, the migration and diffusion process of pollutants in the environment is simulated, and multi-media transmission dynamic data is obtained;

[0060] The reverse trajectory analysis method is applied to the dynamic data of multi-media transmission. Taking the pollutant detection point as the end point, combined with the pollutant characteristic data in the fingerprint library, the transmission starting point and key transmission nodes of the pollutants are reversed to generate a pollutant transmission path map.

[0061] Specifically, the multivariate environmental data is organized according to three dimensions to construct a spatiotemporal data cube. The spatiotemporal data cube is a three-dimensional data structure that structures environmental data according to the time dimension (including hourly, daily, weekly, monthly and seasonal scales), spatial dimension (including sampling point coordinates, administrative divisions and geographical units) and pollutant characteristic dimension (including concentration levels, pollutant component structure and fingerprint characteristics). The specific operation is to reconstruct the original environmental data table into a three-dimensional matrix. Each element of the matrix represents the concentration or characteristic value of a specific pollutant at a specific time point and a specific spatial location, forming a complete spatiotemporal data cube. The autoregressive integrated moving average algorithm is applied to the pollutant concentration data in the spatiotemporal data cube for time series analysis. The autoregressive integrated moving average algorithm, abbreviated as ARIMA, is a statistical model for processing non-stationary time series. It contains three parts: autoregressive term, difference term and moving average term. In practice, the pollutant concentration time series is first tested for stationarity. If it is not stationary, it is stabilized by performing a differencing process. The model order is then determined, including autoregressive, differencing, and moving average orders. The model parameters are then estimated. Finally, the model is used to decompose the time series into a trend component (long-term trend), a seasonal component (cyclical variation), and a random component (irregular fluctuations). This decomposition makes the temporal dynamics of pollutant concentrations clearer, generating temporal variation data, including rising / falling trends in pollutant concentrations, intraday / weekly / seasonal fluctuation patterns, and the identification of sudden change events.

[0062] Geographically weighted regression and spatial autocorrelation analysis methods were applied to the spatiotemporal data cube. Geographically weighted regression is a regression analysis method that accounts for spatial nonstationarity. By assigning higher weights to closer observations, a local regression model is established at each geographic location to explore the spatial variation between pollutant concentrations and environmental factors. Spatial autocorrelation analysis studies the similarity between spatial units. The global Moran index is a statistic that quantifies the degree of spatial autocorrelation across the entire study area. It ranges from -1 to 1, with positive values ​​indicating clustering of similar values, negative values ​​indicating clustering of dissimilar values, and zero indicating random distribution. Local spatial autocorrelation indicators, such as the LISA statistic, identify local spatial clustering patterns, marking high-value clusters (hot spots), low-value clusters (cold spots), and spatial anomalies. Through these analyses, spatial distribution characteristics are obtained, clarifying the spatial distribution patterns of pollutants, particularly the location and extent of clusters and anomalous hot spots.

[0063] Combining spatial distribution data with temporal variation data, we apply the variogram to analyze the spatial structure of pollutant concentrations. The variogram describes the relationship between data differences and distance between any two points in space. By calculating the semivariance between pairs of sample points at different distances and fitting a theoretical variogram model, we can determine the spatial correlation distance (range of influence) and anisotropy (differences in correlation across different directions). Based on the variogram analysis results, we construct a spatial interpolation algorithm for pollutant concentrations, such as the kriging interpolation method. This method is an optimal linear unbiased estimate that accounts for the spatial autocorrelation of sample points. This method generates a heat map of the pollution distribution in the study area, visually demonstrating the spatial distribution pattern of pollutants.

[0064] Based on the pollution distribution heat map, combined with the hydrodynamic equation, atmospheric diffusion equation and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated. The hydrodynamic equation describes the movement of pollutants in water bodies, mainly considering convection, diffusion and degradation processes; the atmospheric diffusion equation is based on the principle of conservation of mass, describing the diffusion behavior of pollutants in the atmosphere, considering factors such as wind direction, wind speed, and atmospheric stability; the soil migration equation considers factors such as soil properties and precipitation infiltration, and describes the vertical and horizontal migration of pollutants in the soil. By combining these equations, setting boundary conditions and initial conditions, and using numerical solution methods to calculate the migration and diffusion process of pollutants, multi-media transmission dynamic data are obtained, including the distribution of pollutant concentrations at each time point and each spatial position. The reverse trajectory analysis method is applied to the multi-media transmission dynamic data. Reverse trajectory analysis is a technology that reversely infers the transmission path of pollutants starting from the pollutant detection point. In specific implementation, starting with the pollutant detection point, the system combines transmission dynamics data with pollutant characteristic data (such as degradation rate and adsorption coefficient) and, based on concentration gradients and environmental conditions, gradually traces the pollutant's propagation path through various environmental media. Key transmission nodes (such as medium interfaces and flow direction change points) are identified, ultimately leading to the inference of possible pollution source areas. Through visualization, a pollutant transmission path map is generated, clearly showing the complete transmission path from the source to the detection point.

[0065] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0066] Perform feature standardization on multivariate environmental data, use feature selection algorithm to screen significant features, perform dimensionality reduction through principal component analysis, and generate environmental data source nodes;

[0067] Apply feature encoding method to the pollutant fingerprint identification in the fingerprint library, convert mass spectrum characteristics, chromatographic characteristics and physical and chemical properties into numerical feature matrix, and generate fingerprint library source nodes through nonlinear dimensionality reduction mapping;

[0068] The pollutant transmission path graph is processed through a graph convolutional network to parameterize the path nodes and connection relationships, extract the path topology features and transmission dynamics features, and generate the transmission path source nodes;

[0069] A multi-branch neural network is constructed for the environmental data source nodes, fingerprint library source nodes, and transmission path source nodes. An attention mechanism fusion layer is designed to perform weighted integration of features. The network structure is optimized through cross-validation to obtain a fusion source node set.

[0070] A multilayer perceptron network is constructed based on the fused source node set. The network consists of three hidden layers, each containing 128, 64, and 32 neurons, respectively. The activation function is the ReLU function, the output layer uses the Softmax function, and the loss function is the cross-entropy loss. The parameters are optimized using the Adam optimizer to obtain the source parsing classifier.

[0071] The source apportionment classifier is used to classify and predict the sources of pollutants, calculate the probability distribution of each pollution source category, and calculate the contribution ratio of each source in combination with the pollutant concentration data. A three-dimensional data structure containing source type, contribution rate and geographic coordinates is constructed to obtain the pollution source contribution rate matrix.

[0072] Specifically, feature standardization is performed on multivariate environmental data to make data of different dimensions comparable. Feature standardization includes methods such as min-max scaling (linearly transforming the data to the range [0,1]), Z-score standardization (subtracting the mean and dividing by the standard deviation), or percentile transformation. After standardization, feature selection algorithms are used to screen for significant features to reduce data dimensionality and improve model efficiency. Feature selection algorithms include statistical test-based methods (such as the chi-square test and the F test), model-based methods (such as feature importance scoring based on decision trees), and wrapper methods (using the performance of the target model as a criterion for evaluating feature subsets). The selected features are then subjected to dimensionality reduction using principal component analysis. Principal component analysis transforms the original features into mutually orthogonal principal components through linear transformation. The principal components are then sorted from largest to smallest according to the proportion of variance explained by the principal components, and the top few principal components are selected as the new feature space to generate environmental data source nodes. Feature encoding methods are applied to the pollutant fingerprint identifiers in the fingerprint library to convert non-numerical features into numerical forms that can be processed by machine learning algorithms. Specifically, mass spectrometric features (such as accurate mass, isotope pattern, characteristic fragment ions, ion abundance ratio), chromatographic features (such as retention time, peak shape), and physicochemical properties (such as hydrophilicity / hydrophobicity, acidity / alkalinity, volatility) are converted into numerical feature matrices through one-hot encoding, label encoding, or embedding. Because the dimensions of the converted features are usually very high, they need to be reduced through nonlinear dimensionality reduction mapping. Commonly used nonlinear dimensionality reduction methods include t-SNE (t-distributed stochastic neighbor embedding) and UMAP (uniform manifold approximation and projection). These methods can preserve the local structure of high-dimensional data, better represent complex fingerprint feature relationships, and ultimately generate fingerprint library source nodes.

[0073] The pollutant transmission pathway graph is processed using a graph convolutional network (GCN), a type of neural network specialized for processing graph-structured data. The transmission pathway graph is represented as a mathematical graph structure, where nodes represent key points in space (e.g., monitoring stations, key pathway points) and edges represent connections between nodes (e.g., water flow direction, atmospheric transmission pathways). The nodes and connections in the graph are then parameterized. Node features include location coordinates and pollutant concentration, while edge features include direction, distance, and transmission rate. The GCN iteratively updates node representations, aggregates information from adjacent nodes, and extracts path topological features (e.g., connectivity and centrality) and transmission dynamics (e.g., diffusion coefficient and decay rate). Ultimately, it generates a transmission pathway source node representing the entire transmission pathway structure. A multi-branch neural network is constructed for the environmental data source nodes, fingerprint library source nodes, and transmission pathway source nodes for fusion. A multi-branch neural network is a network structure that processes multiple different types of inputs in parallel, with each branch independently processing a specific type of data, which is then fused at the network backend. In this method, the three source nodes are fed into three network branches for preliminary feature extraction. The three types of features are then weighted and integrated through a designed attention mechanism fusion layer. The attention mechanism dynamically assigns weights to different features based on the importance of the current task, prioritizing key features relevant to pollution source identification. The entire network structure is optimized using cross-validation. Cross-validation divides the dataset into training, validation, and test sets. Through multiple training-validation cycles, the optimal network structure parameters are selected, ultimately resulting in a set of fused source nodes.

[0074] A multilayer perceptron network is constructed as a source parsing classifier based on the fused source node set. A multilayer perceptron is a feedforward neural network consisting of an input layer, a hidden layer, and an output layer. In this method, the network contains three hidden layers, each containing 128, 64, and 32 neurons, respectively. This pyramidal structure decreases the number of neurons layer by layer, facilitating the gradual extraction of high-level features. A ReLU activation function is used after each hidden layer. This function remains unchanged for positive inputs and sets negative inputs to zero, offering advantages such as computational simplicity and gradient stability. The output layer uses the Softmax function to convert the outputs of multiple neurons into a probability distribution, where each value represents the probability that a sample belongs to a particular class. During training, a cross-entropy loss function is used to quantify the gap between the predicted results and the true labels. The Adam optimizer is used to adaptively adjust the learning rate and update the network parameters, ultimately resulting in a source parsing classifier.

[0075] Use the source parsing classifier to classify and predict the sources of pollutants and calculate the probability distribution of each pollution source category. Specifically, the pollutant sample data to be traced is converted into a feature vector in the same format as the training data through the aforementioned processing flow, and input into the trained source parsing classifier. The output layer gives the probability value of the sample belonging to each pollution source category (such as industrial emissions, agricultural runoff, urban sewage, etc.). Combined with the pollutant concentration data, the contribution ratio of each source is calculated, that is, the product of the probability value of each source and the corresponding pollutant concentration, and then normalized to obtain the contribution rate in percentage form. Finally, the source type, contribution rate and geographic coordinate information (latitude and longitude, elevation, etc.) are integrated to construct a three-dimensional data structure containing these three dimensions, forming a pollution source contribution rate matrix, which provides basic data for the subsequent positioning of key pollution sources.

[0076] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0077] Perform three-dimensional splitting of the pollution source contribution rate matrix, extract pollutant characteristic parameters, spatial coordinate information and time series data, construct a three-dimensional decision space, and obtain basic data for tracing decision-making;

[0078] Based on the pollutant characteristic parameters in the basic data of traceability decision-making, the pollutants are chemically classified, their molecular weight ranges are divided, their functional group characteristics are analyzed, and their environmental degradation characteristics are evaluated. A decision path based on the pollutant characteristic dimension is constructed to obtain the pollutant characteristic judgment results.

[0079] Based on the spatial coordinate information in the basic data for tracing decision-making, combined with topography, water system distribution, and administrative divisions, a spatial relationship analysis is conducted on the upstream area of ​​the pollutant detection point, the distribution of surrounding potential pollution sources, and land use types. A decision path in the spatial location dimension is constructed to obtain the spatial location judgment result.

[0080] Based on the time series data in the basic data for tracing decision-making, the differences between weekdays and weekends in pollutant detection, seasonal variation characteristics, and temporal correlation with specific events are analyzed to construct a decision path in the time feature dimension and obtain the time feature judgment results.

[0081] The pollutant feature judgment results, spatial location judgment results, and temporal feature judgment results are input into the decision tree model. The information gain ratio is calculated using the C4.5 algorithm to select the optimal split attribute, generate a traceability reasoning chain, and obtain the candidate pollution source area.

[0082] A grid refinement strategy is applied to candidate pollution source areas, temporary monitoring points are set up, fingerprint feature comparison and analysis is performed on the collected environmental samples and potential pollution source samples, the similarity score is calculated, the precise location of the pollution source is determined, and the pollution source location information is obtained.

[0083] Specifically, the pollution source contribution rate matrix is ​​split into three dimensions. This matrix contains information in three dimensions: pollution source type, contribution rate, and geographic coordinates. Three-dimensional splitting refers to decomposing this complex data structure into three independent but interrelated data sets: pollutant characteristic parameters (including the chemical properties, molecular structure characteristics, and environmental behavior data of pollutants), spatial coordinate information (including latitude and longitude, elevation, and regional boundary data), and time series data (including monitoring time points, seasonal cycles, and special event markers). These three types of data together construct a three-dimensional decision space, forming the basic data for tracing decisions. The three-dimensional decision space can be viewed as a cube, where the x-axis represents the pollutant characteristics, the y-axis represents the spatial position, and the z-axis represents time. The position of each data point in space reflects the combined relationship of characteristics of different dimensions. Based on the pollutant characteristic parameters in the basic data for tracing decisions, a multi-level pollutant characteristic analysis is performed. First, chemical classification is performed, categorizing pollutants based on their chemical structure and properties into organic compounds (such as polycyclic aromatic hydrocarbons, chlorinated hydrocarbons, and esters) and inorganic compounds (such as heavy metals and nitrogen and phosphorus compounds). Molecular weight ranges are then stratified, with pollutants classified into low molecular weight (<200), medium molecular weight (200-500), and high molecular weight (>500). Pollutants in different molecular weight ranges typically originate from different emission activities. Functional group characterization is then performed to identify key functional groups (such as carboxyl, hydroxyl, and amino) within pollutant molecules. These functional groups are closely associated with the source industries. Finally, environmental degradation characteristics are evaluated to analyze the persistence, bioaccumulation, and toxicity of pollutants in the environment. Pollutants from different sources generally exhibit typical degradation characteristics. Combining these analytical results, a decision path is constructed based on the pollutant characteristic dimension, forming a series of "if...then..." conditional judgment rules, such as "If a pollutant contains a specific functional group combination and its molecular weight is within a specific range, it is likely derived from a certain type of industrial activity," ultimately resulting in a pollutant characteristic judgment result.

[0084] The topography is analyzed, including elevation, slope, and aspect, as these factors influence the diffusion paths of pollutants. Next, the distribution of water systems is analyzed, including the direction of rivers, lakes, and groundwater flow, as water systems are important vehicles for pollutant migration. Finally, administrative divisions are considered, including the boundaries and jurisdictions of different functional zones, which are directly related to pollution source management. Based on this, a focused analysis is conducted on the upstream areas of the pollutant detection points, identifying possible pollution sources based on water flow direction or airflow trajectory. A comprehensive survey of the distribution of potential pollution sources in the surrounding area is also conducted, including factories, farms, and waste treatment facilities. Land use types are also considered, as different land uses (such as industrial, agricultural, and residential) are clearly associated with the emission of specific pollutants. Using GIS methods such as spatial overlay analysis, buffer zone analysis, and shortest path analysis, a decision path is constructed for the spatial location dimension, resulting in spatial location judgments. Multidimensional temporal pattern analysis is conducted based on the time series data within the source tracing decision-making foundation. First, we compared pollutant detection differences between weekdays and weekends. Industrial pollutants tend to be higher on weekdays, while domestic pollutants may rise on weekends. We then analyzed seasonal variations, such as the increased concentrations of pollutants related to agricultural activities during specific farming seasons and the increase in heating-related pollutants during winter. Finally, we examined temporal correlations with specific events, such as the relationship between pollutant concentration changes and specific industrial production cycles, holidays, and extreme weather events. Using methods such as time series decomposition, pattern recognition, and association rule mining, we constructed a decision path based on the temporal feature dimension and generated temporal feature judgment results.

[0085] The results of the three aforementioned dimensions (pollutant characteristics, spatial location, and temporal characteristics) are fed into a decision tree model for comprehensive analysis. A decision tree is a tree-structured classification model that classifies data into different categories through a series of questions. This method uses the C4.5 algorithm to construct the decision tree. This algorithm selects the optimal splitting attribute by calculating the information gain ratio. The information gain ratio measures the contribution of an attribute to the classification and also considers the attribute's inherent splitting information to avoid bias toward attributes with a high number of values. For each decision node, the information gain ratios of all possible splitting attributes are calculated. The attribute with the highest information gain ratio is selected as the splitting point, and subtrees are recursively constructed until the termination criteria are met. This ultimately forms a complete decision tree. Each path from the root node to a leaf node forms a traceable inference chain. Different leaf nodes correspond to different possible pollution source areas. The confidence scores of each inference chain are combined to determine candidate pollution source areas. A grid refinement strategy is applied to the identified candidate pollution source areas, dividing the candidate areas into finer grid cells. The grid size is determined based on the regional characteristics and the required positioning accuracy, typically starting from several hundred meters to tens of meters. Temporary monitoring points are set up at key grid locations to collect more targeted environmental samples. Direct samples from potential pollution sources (such as sewage outlet water samples and flue gas samples) are collected whenever possible. The collected samples are subjected to detailed fingerprint feature analysis to extract molecular markers, elemental composition characteristics, and other features, which are then compared with the features in the pollutant fingerprint library. Similarity scores are calculated between samples. Commonly used similarity calculation methods include cosine similarity, Euclidean distance, and Mahalanobis distance. Higher scores indicate greater similarity between samples. By comparing the similarity between samples at each point and the pollutants at the detection point, combined with the previous tracing decision results, the precise location of the pollution source is ultimately determined, and the pollution source location information is obtained.

[0086] In a specific embodiment, the process of inputting the pollutant characteristic judgment result, the spatial position judgment result, and the time characteristic judgment result into a decision tree model may specifically include the following steps:

[0087] The pollutant feature judgment results, spatial location judgment results, and temporal feature judgment results are merged to construct a multidimensional decision dataset containing all features. The fully connected layer is used to combine the highly correlated features to obtain a decision attribute pool.

[0088] Calculate the information entropy for each attribute in the decision attribute pool to quantify the uncertainty of the sample set. By calculating the probability of each type of sample appearing and performing weighted summation, we can obtain a set of information entropy values.

[0089] Based on the information entropy value set, the conditional entropy is calculated for each attribute in the decision attribute pool to quantify the classification ability of each attribute on the sample set. The conditional entropy value set is obtained by calculating the entropy value of the sample subset under each attribute value and performing weighted summation.

[0090] According to the information entropy value set and the conditional entropy value set, the information gain value of each attribute is calculated, and the contribution of the attribute to the classification is quantified by subtracting the conditional entropy from the original information entropy to obtain the information gain value set;

[0091] Normalize the information gain value set and calculate the information gain ratio. Balance the preference for multi-valued attributes by dividing the information gain by the split information value. Select the attribute with the largest information gain ratio as the split attribute of the current node to obtain the optimal split attribute sequence.

[0092] According to the optimal splitting attribute sequence, a decision tree structure is constructed, and each branch node is recursively constructed according to the attribute splitting rule until the termination condition is met, forming a complete traceability decision tree. The complete decision chain from the root node to the leaf node is recorded through the decision path tracing function to obtain the candidate pollution source area.

[0093] Specifically, the results of pollutant characteristic, spatial location, and temporal feature judgments are merged to form a complete decision-making basis. Feature merging integrates the judgment results from these three dimensions into a single data structure to construct a multidimensional decision dataset encompassing all features. This process uses a fully connected layer to combine highly correlated features. A fully connected layer is a neural network structure that connects features from different sources into a unified vector representation. Specifically, the judgment results from the three dimensions are represented as vectors. After processing through the fully connected layer, highly correlated features are given higher weights, while weakly correlated features are given lower weights. This optimizes the feature combination and ultimately forms a decision attribute pool. This decision attribute pool includes multidimensional attributes such as "number of chlorine atoms in the pollutant," "water solubility level of the pollutant," "distance between the detection point and the industrial zone," and "weekday / weekend ratio of detection concentration."

[0094] Information entropy is calculated for each attribute in the decision attribute pool to quantify the uncertainty of the sample set. Information entropy is a measure of the degree of disorder or uncertainty in a system, with larger values ​​indicating greater disorder or uncertainty. The calculation begins by categorizing the samples in the decision attribute pool by pollution source category (such as industrial emissions, agricultural runoff, and urban sewage). The proportion of samples in each category in the total sample is calculated as the probability of occurrence of that category. The logarithm of the probability of occurrence of each category is then multiplied by the probability itself, and the sum is negated across all categories to obtain the information entropy value of the sample set. This calculation is performed for each attribute in the decision attribute pool, forming a set of information entropy values. This set reflects the difficulty of classifying the sample set without considering any other attributes. Based on this set of information entropy values, conditional entropy is calculated for each attribute in the decision attribute pool to quantify the classification ability of each attribute for the sample set. Conditional entropy describes the amount of additional information required to classify a sample given the value of a given attribute. A smaller value indicates that the attribute provides greater assistance in classification. To calculate conditional entropy, the sample set is first divided into multiple subsets based on the different values ​​of the attribute. The information entropy of each subset is calculated, and then a weighted sum is taken based on the subset's proportion in the total sample to obtain the conditional entropy of that attribute. For example, for the attribute "whether the pollutant contains chlorine," the samples are divided into two subsets: "containing chlorine" and "not containing chlorine." The information entropy of these two subsets is calculated separately, and then the weighted sum is taken based on the sample proportions of the two subsets to obtain the conditional entropy of this attribute. This process is repeated for each attribute in the decision attribute pool to form a set of conditional entropy values.

[0095] Based on the information entropy and conditional entropy sets, the information gain value for each attribute is calculated. Information gain reflects the reduction in system uncertainty after using a particular attribute to classify the sample. A larger value indicates a greater contribution of the attribute to the classification. Information gain is calculated by subtracting the conditional entropy from the original information entropy. The difference represents the attribute's contribution to reducing system uncertainty. For example, if the information entropy of the original sample set is 0.9, and the conditional entropy after classifying the sample using the "whether the pollutant contains chlorine" attribute is 0.3, the information gain of this attribute is 0.6, indicating that the system uncertainty is reduced by 0.6 after classifying the sample using this attribute. Information gain is calculated for each attribute in the decision attribute pool to form a set of information gain values.

[0096] The information gain value set is normalized to calculate the information gain ratio. The information gain ratio is a core concept of the C4.5 algorithm. By dividing the information gain by the attribute's own splitting information, it corrects the original information gain's bias toward multi-valued attributes. Splitting information refers to the uniformity of the sample distribution after the sample is partitioned by the attribute. Its calculation method is similar to information entropy, except that the class probability is replaced by the sample proportion under each value. The information gain ratio of each attribute is obtained by dividing the information gain by the splitting information. For example, if the information gain of "whether the pollutant contains chlorine" is 0.6 and the splitting information is 0.8, its information gain ratio is 0.75. Meanwhile, if "pollutant functional group type" has seven possible values, the information gain is 0.65 and the splitting information is 1.8, its information gain ratio is approximately 0.36. The information gain ratios of all attributes are compared, and the attribute with the largest gain ratio is selected as the splitting attribute for the current node. The sample set is then recursively partitioned according to different attribute values ​​to obtain the optimal splitting attribute sequence.

[0097] Construct a decision tree structure based on the optimal splitting attribute sequence. The decision tree consists of a root node, internal nodes, and leaf nodes. Each non-leaf node corresponds to a splitting attribute, each branch corresponds to a value of the attribute, and each leaf node corresponds to a classification result. The construction process starts from the root node, selects the attribute with the highest information gain ratio as the splitting point, and divides the samples into branches according to the different values ​​of the attribute. Then, the same splitting process is recursively performed on each branch until the termination condition is met. The termination conditions include: all samples belong to the same category, no more attributes are available, and the number of samples at the branch node is less than the preset threshold. After the decision tree is constructed, the decision path tracking function is used to record the complete decision chain from the root node to each leaf node. Each decision chain represents a traceability reasoning path, corresponding to a possible pollution source type or area. Based on the distribution of samples at each leaf node and the category purity of the leaf node, the final candidate pollution source area is determined.

[0098] The above describes the method for tracing the source of new pollutants in various environmental media based on machine learning in the embodiment of the present application. The following describes the system for tracing the source of new pollutants in various environmental media based on machine learning in the embodiment of the present application. Figure 2 In one embodiment of the present application, a system for tracing the source of new pollutants in multiple environmental media based on machine learning includes:

[0099] A construction module 201 is configured to collect multivariate environmental data through a sensor network and construct a fingerprint library containing pollutant fingerprint identifiers based on the multivariate environmental data;

[0100] A reconstruction module 202 is used to perform spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map;

[0101] Fusion module 203, for converting the multivariate environmental data, the fingerprint library and the pollutant transmission path diagram into source nodes and performing multi-source data fusion, building a source resolution model, and obtaining a pollution source contribution rate matrix;

[0102] The positioning module 204 is configured to locate key pollution sources based on the pollution source contribution rate matrix through the pollutant characteristic dimension, spatial location dimension, and time characteristic dimension to obtain pollution source location information.

[0103] Through the collaborative cooperation of the above components, through the collection of multi-dimensional environmental data and the construction of a pollutant fingerprint library, accurate identification and feature extraction of new pollutants are achieved, thereby significantly improving the starting point accuracy of traceability; the spatiotemporal correlation analysis and transmission path reconstruction of the fingerprint library and multi-dimensional environmental data are carried out to visualize the migration and diffusion process of pollutants in the environment, solving the technical problem that traditional methods are difficult to track the transmission path of pollutants in complex environments; the multi-dimensional environmental data, fingerprint library and pollutant transmission path map are converted into quantum states and multi-source data fusion is performed, fully exploring the deep correlation between different types of data. Among them, quantum computing shows parallel processing capabilities and dimensionality reduction efficiency that traditional computing methods cannot match in multi-dimensional feature processing. This innovative data fusion strategy solves the problem of insufficient data integration of traditional methods; based on the pollution source contribution rate matrix, through the three-dimensional decision space It locates key pollution sources and realizes gradual positioning from large-scale areas to precise geographical locations, greatly improving the spatial resolution and accuracy of tracing sources. In particular, the application of the C4.5 decision tree algorithm in pollution source identification, compared with traditional empirical judgment, objectively quantifies the importance of each decision attribute through the information gain ratio, making the decision-making process more scientific and transparent, and eliminating the uncertainty brought about by human subjective judgment in traditional methods; the deep integration of artificial intelligence algorithms and environmental science knowledge in the overall solution, especially in key links such as graph convolutional network processing transmission paths, multi-branch neural network realization of feature fusion, and decision tree construction of tracing reasoning chain, fully utilizes the unique advantages of machine learning algorithms in complex pattern recognition, multi-source data fusion and intelligent decision support, solves the problem of tracing the source of new pollutants that traditional methods cannot cope with, improves tracing efficiency, and reduces dependence on expert experience.

[0104] above Figure 2 From the perspective of modular functional entities, the system for tracing the source of new pollutants in multiple environmental media based on machine learning in an embodiment of the present invention is described in detail. Below, the device for tracing the source of new pollutants in multiple environmental media based on machine learning in an embodiment of the present invention is described in detail from the perspective of hardware processing.

[0105] Figure 3This is a schematic diagram of the structure of a device for tracing the source of new pollutants in multiple environmental media based on machine learning, provided by an embodiment of the present invention. This device 300, which can vary significantly depending on configuration or performance, may include one or more central processing units (CPUs) 310 (e.g., one or more processors), memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) storing applications 333 or data 332. The memory 320 and storage media 330 may be either transient or persistent storage. The program stored in the storage medium 330 may include one or more modules (not shown), each of which may include a series of instructions and operations within the device 300 for tracing the source of new pollutants in multiple environmental media based on machine learning. Furthermore, the processor 310 may be configured to communicate with the storage medium 330, executing the series of instructions and operations stored in the storage medium 330 on the device 300 for tracing the source of new pollutants in multiple environmental media based on machine learning, thereby implementing the steps of the aforementioned method for tracing the source of new pollutants in multiple environmental media based on machine learning.

[0106] The device 300 for tracing the source of new pollutants in various environmental media based on machine learning may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be understood by those skilled in the art that Figure 3 The structure of the device for tracing the source of new pollutants in multiple environmental media based on machine learning shown does not constitute a limitation on the device for tracing the source of new pollutants in multiple environmental media based on machine learning provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0107] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the method for tracing the source of new pollutants in multiple environmental media based on machine learning.

[0108] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0109] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a new pollutant source tracing device in multiple environmental media based on machine learning (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for tracing the source of new pollutants in multiple environmental media based on machine learning, characterized in that: The method comprises: Collect multivariate environmental data through sensor networks and build a fingerprint library containing pollutant fingerprint identification based on the multivariate environmental data; Performing spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; Converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path diagram into source nodes and performing multi-source data fusion to construct a source resolution model and obtain a pollution source contribution rate matrix; Based on the pollution source contribution rate matrix, key pollution sources are located through the pollutant characteristic dimension, spatial location dimension and time characteristic dimension to obtain pollution source location information, including: performing three-dimensional decomposition of the pollution source contribution rate matrix, extracting pollutant characteristic parameters, spatial coordinate information and time series data, constructing a three-dimensional decision space, and obtaining basic data for tracing decision-making; based on the pollutant characteristic parameters in the basic data for tracing decision-making, chemical classification, molecular weight range division, functional group characteristic analysis and environmental degradation characteristic evaluation are performed on pollutants, a decision path of pollutant characteristic dimension is constructed, and pollutant characteristic judgment results are obtained; based on the spatial coordinate information in the basic data for tracing decision-making, combined with topography, water system distribution and administrative divisions, the upstream area of ​​the pollutant detection point, the distribution of surrounding potential pollution sources and land use types are spatially related. Analyze and construct a decision path in the spatial location dimension to obtain a spatial location judgment result; based on the time series data in the said traceability decision basic data, analyze the differences between weekdays and weekends in pollutant detection, seasonal variation characteristics and time correlation with specific events, construct a decision path in the time feature dimension, and obtain a time feature judgment result; input the said pollutant feature judgment result, the said spatial location judgment result and the said time feature judgment result into the decision tree model, calculate the information gain ratio through the C4.5 algorithm to select the optimal split attribute, generate a traceability reasoning chain, and obtain the candidate area of ​​the pollution source; apply the grid refinement strategy to the said candidate area of ​​the pollution source, set up temporary monitoring points, perform fingerprint feature comparison analysis on the collected environmental samples and potential pollution source samples, calculate the similarity score, determine the precise location of the pollution source, and obtain the pollution source location information.

2. The method for tracing the source of new pollutants in multiple environmental media based on machine learning according to claim 1, characterized in that: The method of collecting multivariate environmental data through a sensor network and constructing a fingerprint library containing pollutant fingerprint identifiers based on the multivariate environmental data includes: Deploy water quality monitoring sensors, atmospheric monitoring sensors, soil monitoring sensors, and biological response sensors to form a sensor network, collect environmental physical and chemical parameters through a low-power wide-area network, and obtain raw environmental data; Applying the triple standard deviation method to the raw environmental data to identify outliers, using the linear interpolation method to fill in short-term missing data, using the multivariate interpolation method to fill in long-term missing data, and performing noise reduction processing through wavelet transform to obtain preprocessed environmental data; The pre-processed environmental data are standardized by applying the Z-score standardization method, and are combined with geographic information system data, meteorological data, human activity data and historical monitoring records to perform spatiotemporal matching integration to obtain multivariate environmental data; High-resolution mass spectrometry technology is used to perform non-targeted analysis of environmental samples. The mass spectrometry data are processed by peak identification, peak alignment and peak area extraction to obtain a mass spectrum feature table. Classify and cluster the mass spectral feature table, perform database search, and perform structural elucidation on unlisted compounds through isotope ratio analysis, fragment ion analysis, and retention time prediction to obtain a candidate list of new pollutants; The mass spectrum characteristics, chromatographic characteristics, physicochemical properties and environmental behavior characteristics of each pollutant in the new pollutant candidate list are extracted, and principal component analysis and recursive feature elimination algorithm are applied to perform feature optimization screening to obtain a fingerprint library containing pollutant fingerprint identification.

3. The method for tracing the source of new pollutants in multiple environmental media based on machine learning according to claim 2 is characterized in that: The mass spectrum feature table is classified and clustered, and a database search is performed at the same time. The structure of the unlisted compounds is analyzed by isotope ratio analysis, fragment ion analysis and retention time prediction to obtain a candidate list of new pollutants, including: Applying molecular network analysis technology to the mass spectral feature table to calculate the spectral similarity matrix between the compounds, and classifying compounds with a similarity score greater than 0.7 into the same group based on the spectral similarity score to obtain a preliminary molecular network structure; Applying a hierarchical clustering algorithm to optimize the structure of the preliminary molecular network structure, calculating the Euclidean distance between clusters, setting a clustering threshold, and obtaining an optimized molecular network map; Perform a bidirectional search and match between the mass spectrum features in the molecular network map and the known environmental pollutant database and the chemical registration database, and obtain database matching results by comparing molecular formula, molecular weight and characteristic fragment ions; Compounds with a matching degree of less than 80% in the database matching results are marked and filtered based on the retention time index to obtain a list of potential new pollutants; Conduct isotope ratio analysis on the compounds in the potential new pollutant list, calculate the isotope abundance ratios of carbon, nitrogen, chlorine, and bromine, and deduce molecular formulas based on elemental composition rules to obtain molecular composition data; Based on the molecular composition data, the characteristic fragment ions are subjected to substructure analysis, the molecular skeleton is reconstructed using the fragment tree algorithm, and the molecular structure is optimized in combination with electron density calculation to obtain a candidate list of new pollutants.

4. The method for tracing the source of new pollutants in multiple environmental media based on machine learning according to claim 1, characterized in that: The performing of spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map includes: Organizing the multivariate environmental data according to the time dimension, space dimension and pollutant characteristic dimension, constructing a three-dimensional data matrix, and obtaining a spatiotemporal data cube; Applying an autoregressive integrated moving average algorithm to perform time series analysis on the pollutant concentration data in the spatiotemporal data cube, decomposing the pollutant concentration changes into trend components, seasonal components, and random components, and obtaining time variation characteristic data; Applying geographically weighted regression and spatial autocorrelation analysis methods to the spatiotemporal data cube, calculating the global Moran index and local spatial autocorrelation index, identifying spatial concentration areas and abnormal hot spots of pollutants, and obtaining spatial distribution characteristic data; Combining the spatial distribution characteristic data with the temporal variation characteristic data, applying the variation function to analyze the spatial structural characteristics of pollutant concentrations, determining the spatial correlation distance and anisotropy characteristics, constructing a spatial interpolation algorithm for pollutant concentrations, and obtaining a pollution distribution heat map of the study area; Based on the pollution distribution thermodynamic map, combined with the hydrodynamic equation, atmospheric diffusion equation and soil migration equation, the transmission path parameters of pollutants in different environmental media are calculated, the migration and diffusion process of pollutants in the environment is simulated, and multi-media transmission dynamic data is obtained; The reverse trajectory analysis method is applied to the multi-media transmission dynamic data, with the pollutant detection point as the end point, combined with the pollutant characteristic data in the fingerprint library, to reversely infer the transmission starting point and key transmission nodes of the pollutants to generate a pollutant transmission path map.

5. The method for tracing the source of new pollutants in multiple environmental media based on machine learning according to claim 1, characterized in that: The step of converting the multivariate environmental data, the fingerprint library, and the pollutant transmission path diagram into source nodes and performing multi-source data fusion to construct a source resolution model and obtain a pollution source contribution rate matrix includes: Performing feature standardization on the multivariate environmental data, screening significant features using a feature selection algorithm, and performing dimensionality reduction processing through principal component analysis to generate environmental data source nodes; Applying a feature encoding method to the pollutant fingerprint identifiers in the fingerprint library, converting mass spectrum features, chromatographic features, and physicochemical properties into a numerical feature matrix, and generating fingerprint library source nodes through nonlinear dimensionality reduction mapping; The pollutant transmission path graph is processed through a graph convolutional network to parameterize the path nodes and connection relationships, extract the path topology characteristics and transmission dynamics characteristics, and generate the transmission path source nodes; Constructing a multi-branch neural network for the environment data source node, the fingerprint library source node, and the transmission path source node, designing an attention mechanism fusion layer to perform weighted integration of features, optimizing the network structure through cross-validation, and obtaining a fusion source node set; A multilayer perceptron network is constructed based on the fused source node set. The network includes three hidden layers, each layer includes 128, 64, and 32 neurons, respectively. The activation function uses the ReLU function, the output layer uses the Softmax function, and the loss function selects the cross entropy loss. The parameters are optimized using the Adam optimizer to obtain a source parsing classifier. The source analysis classifier is used to classify and predict the sources of pollutants, calculate the probability distribution of each pollution source category, calculate the contribution ratio of each source in combination with the pollutant concentration data, construct a three-dimensional data structure including source type, contribution rate and geographic coordinates, and obtain the pollution source contribution rate matrix.

6. The method for tracing the source of new pollutants in multiple environmental media based on machine learning according to claim 1, characterized in that: The pollutant feature judgment result, the spatial location judgment result, and the temporal feature judgment result are input into a decision tree model, and the information gain ratio is calculated by the C4.5 algorithm to select the optimal split attribute, generate a traceability reasoning chain, and obtain a candidate pollution source area, including: Merging the pollutant feature judgment results, the spatial location judgment results, and the temporal feature judgment results to construct a multidimensional decision dataset containing all features, and combining highly correlated features using a fully connected layer to obtain a decision attribute pool; Calculating information entropy for each attribute in the decision attribute pool to quantify the uncertainty of the sample set, and obtaining a set of information entropy values ​​by calculating the probability of occurrence of each type of sample and performing weighted summation; Based on the information entropy value set, conditional entropy is calculated for each attribute in the decision attribute pool to quantify the classification ability of each attribute on the sample set. The conditional entropy value set is obtained by calculating the entropy value of the sample subset under each attribute value and performing weighted summation. Calculating the information gain value of each attribute based on the information entropy value set and the conditional entropy value set, and quantifying the contribution of the attribute to the classification by subtracting the conditional entropy from the original information entropy to obtain the information gain value set; Normalizing the information gain value set, calculating the information gain ratio, balancing the preference for multi-valued attributes by dividing the information gain by the splitting information value, selecting the attribute with the largest information gain ratio as the splitting attribute of the current node, and obtaining the optimal splitting attribute sequence; According to the optimal splitting attribute sequence, a decision tree structure is constructed, and each branch node is recursively constructed according to the attribute splitting rule until the termination condition is met, forming a complete traceability decision tree. The complete decision chain from the root node to the leaf node is recorded through the decision path tracing function to obtain the candidate pollution source area.

7. A machine learning-based system for tracing the source of new pollutants in multiple environmental media, characterized by: A method for tracing the source of new pollutants in multiple environmental media based on machine learning according to any one of claims 1 to 6, wherein the system for tracing the source of new pollutants in multiple environmental media based on machine learning comprises: A construction module is used to collect multivariate environmental data through a sensor network and construct a fingerprint library containing pollutant fingerprint identification based on the multivariate environmental data; A reconstruction module, configured to perform spatiotemporal correlation analysis and transmission path reconstruction on the fingerprint library and the multivariate environmental data to obtain a pollutant transmission path map; A fusion module, configured to convert the multivariate environmental data, the fingerprint library, and the pollutant transmission path diagram into source nodes and perform multi-source data fusion, construct a source resolution model, and obtain a pollution source contribution rate matrix; A positioning module is used to locate key pollution sources based on the pollution source contribution rate matrix through the pollutant characteristic dimension, spatial location dimension and time characteristic dimension to obtain pollution source location information, including: performing three-dimensional decomposition of the pollution source contribution rate matrix, extracting pollutant characteristic parameters, spatial coordinate information and time series data, constructing a three-dimensional decision space, and obtaining basic data for tracing decision-making; based on the pollutant characteristic parameters in the basic data for tracing decision-making, chemical classification, molecular weight range division, functional group characteristic analysis and environmental degradation characteristic evaluation of pollutants are performed to construct a decision path for pollutant characteristic dimension and obtain pollutant characteristic judgment results; based on the spatial coordinate information in the basic data for tracing decision-making, combined with topography, water system distribution and administrative divisions, spatial analysis of the upstream area of ​​the pollutant detection point, the distribution of surrounding potential pollution sources and land use types is performed. The relationship between them is analyzed, a decision path of spatial location dimension is constructed, and a spatial location judgment result is obtained; based on the time series data in the said traceability decision basic data, the difference between weekdays and weekends detected pollutants, seasonal variation characteristics and time correlation with specific events are analyzed, a decision path of time feature dimension is constructed, and a time feature judgment result is obtained; the pollutant feature judgment result, the said spatial location judgment result and the said time feature judgment result are input into the decision tree model, the information gain ratio is calculated by the C4.5 algorithm to select the optimal splitting attribute, a traceability reasoning chain is generated, and a candidate pollution source area is obtained; a grid refinement strategy is applied to the said pollution source candidate area, temporary monitoring points are set, the fingerprint feature comparison analysis of the collected environmental samples and potential pollution source samples is performed, the similarity score is calculated, the precise location of the pollution source is determined, and the pollution source location information is obtained.

8. A device for tracing the source of new pollutants in multiple environmental media based on machine learning, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the method for tracing the source of new pollutants in multiple environmental media based on machine learning as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor executes the method for tracing the source of new pollutants in multiple environmental media based on machine learning as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Polluted gas tracing method based on gas sensor array fingerprint identification

    CN112986497A

  • Space-time big data fused drainage basin water quality pollution traceability analysis method and system

    CN119646471A

  • Site multi-medium pollution risk early warning and management and control system

    CN120013230A