A Source Tracing Method for High-Risk Pollutants in Industrial Parks Based on Mass Spectrometry Machine Learning
By constructing a fragmented tree ensemble dataset and a graph neural network model, structural analogs of high-risk pollutants in industrial parks are identified, solving the problem of missed detection in traditional methods and enabling effective source tracing in the absence of standards and spectral libraries.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2026-01-08
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to identify a large number of structural analogs in industrial park environments, leading to missed detections of high-risk pollutants. Furthermore, traditional methods that rely on standard spectral libraries and fragment matching cannot effectively trace the transformation products of high-risk pollutants.
A fragmented tree ensemble dataset with reaction relationship labels was constructed. A graph neural network model was used to train and screen fragmented trees in environmental water samples from industrial parks. Machine learning was used to identify the reaction relationships of high-risk pollutants and identify structural analogs.
It eliminates the reliance on complete standard spectral libraries, enhances the ability to identify structural analogs, and can identify unknown substances in the absence of standards and reference spectral libraries, making it suitable for tracing the source of high-risk pollutants in complex process scenarios.
Smart Images

Figure CN122135804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental analysis technology, and in particular to a method for tracing the source of high-risk pollutants and structural analogs in industrial parks based on mass spectrometry machine learning. Background Technology
[0002] The industrial park is home to a large number of enterprises engaged in fine chemicals, pharmaceuticals, pesticides, and materials. The high-risk pollutants emitted are complex in type and vary widely in concentration. These pollutants readily undergo a series of transformation processes in the environmental medium, including oxidation, reduction, hydrolysis, covalent bond breaking, and bonding, forming numerous unknown transformation products and structural analogs. Some of these substances possess high-risk characteristics such as persistence, toxicity, or endocrine disruption, posing a potential threat to the regional ecological environment and human health.
[0003] Currently, the identification of high-risk pollutants mainly relies on targeted detection based on standards or suspicious screening based on standard spectral libraries. These methods require prior acquisition of standards or high-quality mass spectrometry reference data for the target high-risk pollutant, relying on accurate mass matching of precursor ions and comparison with a limited number of characteristic fragments to confirm the substance. However, in real-world environments, many structural analogs (including precursors and downstream transformation products) lack available standards or standard spectra, and existing spectral libraries have limited coverage, leading to significant missed detections in the identification of high-risk spectral signals in environmental samples. Furthermore, traditional screening based on simple fragment matching or similarity struggles to utilize known reaction pathways and transformation rules to discover spectral signals that react with known high-risk pollutants, making it impossible to trace structural analogs backward from high-risk end products. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a method for tracing the source of high-risk pollutants and their structural analogs in industrial parks based on mass spectrometry machine learning. Starting from known high-risk pollutants or suspected end products, the method screens spectral signals that have a reaction relationship or similar reaction modes, identifies potential structural analogs, enriches the pollutant spectrum information, and realizes the spectral signal tracing of high-risk pollutants.
[0005] Technical Solution: To achieve the above objectives, the present invention provides a method for tracing the source of high-risk pollutants and their structural analogs in industrial parks based on mass spectrometry machine learning, comprising the following steps:
[0006] S1. Construct a fragmented tree combination dataset labeled "existing reaction relationship / no reaction relationship". The dataset includes several fragmented tree combinations, and each fragmented tree combination includes two fragmented trees representing different known high-risk pollutants.
[0007] S2. Train the graph neural network model using the fragmented tree combination dataset;
[0008] S3. Obtain fragmented tree sets of high-risk pollutants contained in environmental water samples from the industrial park;
[0009] S4. Select at least one fragmented tree with a known high-risk pollutant from the fragmented tree set as the target fragmented tree. Combine the target fragmented tree with other fragmented trees in the fragmented tree set to form a "target fragmented tree - candidate fragmented tree" combination. Input the combination into the trained graph neural network model to determine whether there is a reaction relationship and filter out the candidate fragmented trees with a reaction relationship.
[0010] S5. Determine whether the high-risk pollutant corresponding to the candidate fragmentation tree is a structural analog of the high-risk pollutant corresponding to the target fragmentation tree;
[0011] S6. Based on the judgment results of S5, conduct source tracing of structural analogues of high-risk pollutants within the industrial park.
[0012] Preferably, the method for constructing the fragmentation tree combined dataset in S1 is as follows: obtaining environmental reaction secondary mass spectrometry data of known high-risk pollutants and their transformation products from an environmental reaction database, and predicting the fragmentation tree of the known high-risk pollutants based on a fragmentation tree construction algorithm;
[0013] Based on the reaction relationships between high-risk pollutants recorded in the environmental reaction database and process reaction pathways, pairs of high-risk pollutants with direct upstream-downstream relationships or located in the same reaction network are selected, and the two fragmented tree combinations corresponding to the high-risk pollutant pairs are combined and marked as "reaction relationship exists". Pairs of high-risk pollutants with no known reaction relationship or belonging to different structural families are selected, and the fragmented tree combinations corresponding to the high-risk pollutant pairs are marked as "no reaction relationship exists". Thus, a fragmented tree combination dataset with "reaction relationship exists / no reaction relationship exists" labels is constructed.
[0014] Preferably, in the fragmented tree composite dataset, each fragmented tree is represented as a graph structure, wherein fragment ions are nodes, neutral loss is an edge, node features include at least fragment mass-to-charge ratio m / z, intensity, fragmentation score and mass deviation, and edge features include at least neutral loss score.
[0015] Preferably, in step S2, the graph neural network model is trained by dividing the dataset into a training set and a validation set. During the training phase, the parameters of the graph neural network model are optimized by minimizing the loss function between the predicted probability and the labels in the training set. During the evaluation phase, different probability thresholds are scanned on the validation set, and performance metrics including accuracy, recall, and F1 score are calculated. The probability threshold with the best performance is selected as the optimal discrimination threshold.
[0016] Preferably, the method for obtaining the fragmented tree set of high-risk pollutants contained in the environmental water samples of the industrial park as described in S3 is as follows: the water samples collected in the industrial park are filtered, high-risk pollutants in the water samples are extracted by solid phase extraction, the solid phase extraction column is eluted with solvent, and nitrogen blowing concentration is performed for pretreatment to obtain the test solution. The test solution is analyzed by liquid chromatography-high resolution mass spectrometry, and secondary mass spectrometry data of the characteristic peaks of each high-risk pollutant in the sample are obtained based on chromatographic separation. Based on the fragmented tree construction algorithm, the secondary mass spectrometry data are converted into fragmented trees containing fragment ions and neutral loss relationships to obtain the fragmented tree set.
[0017] Preferably, the graph neural network model in S4 calculates the probability of "existence of a reaction relationship" for each input "target fragmentation tree-candidate fragmentation tree" combination, and based on the optimal discrimination threshold, selects candidate fragmentation trees that have a reaction relationship with the high-risk pollutants corresponding to the selected target fragmentation tree.
[0018] Preferably, the probability of "existence of a reaction relationship" is calculated as follows: the two spectral pairs of "target fragmentation tree - candidate fragmentation tree" are input into the graph neural network model respectively to obtain the probability values in the two directions, and the maximum probability value is taken as the probability of "target fragmentation tree - candidate fragmentation tree" having a reaction relationship.
[0019] Preferably, in S5, for each candidate fragmentation tree, the precise mass difference between the precursor ion of the candidate high-risk pollutant corresponding to the candidate fragmentation tree and the precursor ion of the target high-risk pollutant with which there is a reaction relationship is calculated; the precise mass difference is matched with predefined environmental reaction and transformation rules; if the reaction rules are met, the candidate high-risk pollutant corresponding to the candidate fragmentation tree is determined to be a structural analog of the target high-risk pollutant; the target high-risk pollutant is the high-risk pollutant corresponding to the selected target fragmentation tree.
[0020] Preferably, based on the determination of belonging, the structural analogues are further subdivided into precursors or downstream conversion products of the target high-risk pollutant according to the magnitude and sign of the mass difference and the direction of the reaction in the reaction rules.
[0021] Preferably, in step S6, based on the determination results of step S5, the process information of each process unit in the industrial park, and the spatial distribution of sampling points, the location and evolution relationship of the target high-risk pollutant and its corresponding precursors and downstream conversion products in the park are analyzed, thereby tracing the source of the high-risk pollutant's structural analogues in the industrial park.
[0022] Beneficial Effects: This invention has the following advantages: 1. Eliminating dependence on complete standard spectral libraries: By learning reaction relationships at the fragmentation tree level, rather than relying solely on matching specific fragment peaks, unknown substances that react with known high-risk pollutants can be discovered even in the absence of standards and reference spectral libraries; 2. Enhancing the identification ability of structural analogs: By using known high-risk pollutants or suspected end products as secondary spectral signals of target pollutants and screening spectral signals with reaction relationships or structurally similar reaction patterns, potential structural analogs can be identified, enriching pollutant spectral information; 3. Machine learning discrimination integrating reaction mechanism information: A graph neural network model is constructed based on fragmentation tree diagrams, integrating fragment-neutral loss structures, molecular formula elemental composition, and other information for reaction relationship discrimination, which has stronger reaction identification and generalization capabilities compared to simple mass spectrometry similarity indicators; 4. Applicable to source tracing analysis in complex process scenarios in industrial parks: This invention can combine information on industrial park process flow, emission nodes, and spatial distribution to associate structural analogs in spectral signals, providing tool support for source tracing of high-risk pollutants in complex process systems. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0024] Figure 2 Flowchart for training and using graph neural network models;
[0025] Figure 3 In the industrial park sample used as an example, a schematic diagram of the spectral signals of candidate high-risk pollutants with reaction relationships and their molecular networks were screened after selecting the secondary spectral signals of the target high-risk pollutants;
[0026] Figure 4 This is a secondary mass spectrum of 4-nonylphenol;
[0027] Figure 5 The mass spectrum of 4-(9-carboxynonyl)phenol is shown.
[0028] Figure 6 The mass spectrum of 4-nonylphenol-polyoxyethylene ether is shown.
[0029] Figure 7 This is the secondary mass spectrum of 4-nonylphenoxyethoxyacetic acid;
[0030] Figure 8 This is a secondary mass spectrum of 4-nonylphenoxyacetic acid. Detailed Implementation
[0031] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0032] like Figure 1As shown, the method for tracing the source of high-risk pollutants and structural analogs in industrial parks based on mass spectrometry machine learning, as described in this invention, includes the following:
[0033] S1. Based on the industrial park's process reaction pathways, environmental reaction database, and the transformation relationships of known high-risk pollutants, a fragmented tree ensemble dataset is constructed, labeled "reaction relationship exists / no reaction relationship exists." This dataset includes several fragmented tree ensembles, each consisting of two fragmented trees representing different known high-risk pollutants. The construction process specifically includes:
[0034] S101. Obtain environmental reaction secondary mass spectrometry data of known high-risk pollutants and their transformation products from the environmental reaction database, and predict the fragmentation tree of each known high-risk pollutant based on the fragmentation tree construction algorithm.
[0035] The fragmentation tree construction algorithm adopts a fragmentation tree inference approach similar to that of the SIRIUS software: First, based on the precise mass and isotopic distribution of the parent ion and each fragment ion, candidate molecular formulas that meet the mass accuracy requirements are enumerated and scored for selection; on this basis, with the parent ion as the root node, possible neutral loss edges are constructed between all candidate fragment molecular formulas. Using an objective function that comprehensively considers the mass difference matching error, fragment intensity, isotopic fitting degree, and tree structure complexity, different candidate fragmentation trees are scored and compared, and the fragmentation tree with the highest score is selected as the optimal fragmentation tree representation for the high-risk pollutant.
[0036] S102. Based on the reaction relationships between high-risk pollutants recorded in the environmental reaction database and process reaction pathways, select pairs of high-risk pollutants that have a direct upstream-downstream relationship or are located in the same reaction network. Combine the two fragmentation trees corresponding to these pairs of high-risk pollutants and label them as "reaction relationship exists". Simultaneously, select pairs of high-risk pollutants that have no known reaction relationship or belong to different structural families, and label their corresponding fragmentation tree combinations as "no reaction relationship exists". This constructs a fragmentation tree combination dataset labeled "reaction relationship exists / no reaction relationship exists". In this training set, each fragmentation tree is represented as a graph structure, with fragment ions as nodes and neutral loss as edges. Node features include at least fragment mass-to-charge ratio m / z, intensity, fragmentation score, and mass deviation, while edge features include at least the neutral loss score.
[0037] S2. The graph neural network model is trained using a fragmented tree ensemble dataset. During training, the model parameters are optimized by minimizing the loss function between the predicted probability and the labels in the training set. During model evaluation, different probability thresholds are scanned on the independent validation set, and corresponding performance metrics such as accuracy, recall, and F1 score are calculated. The probability threshold with the best performance is selected as the optimal discrimination threshold. This results in a machine learning model that can input fragmented tree ensembles corresponding to any two high-risk pollutants and determine whether they have a reaction relationship.
[0038] like Figure 2 As shown, the graph neural network model includes:
[0039] Graph encoding branch: Taking the fragmented tree obtained by the fragmented tree construction algorithm from the secondary mass spectrometry data of the environmental reaction database as input, each fragmented tree is represented as a graph structure, with fragment ions as nodes and neutral loss as edges. The graph structure is processed by a multi-layer graph neural network for message passing and feature aggregation, and global pooling is used to map the entire fragmented tree into a fixed-length fragmented tree graph representation vector.
[0040] Molecular formula encoding branch: Taking the chemical structure annotation or molecular formula information corresponding to the above fragmented tree as input, the molecular formula is parsed into a vector obtained by counting elements, and a logarithmic transformation is performed on it to reduce the difference in the order of magnitude of different element counts. Then, it is projected and scaled through a fully connected layer and a normalization layer to obtain a molecular formula representation vector that matches the dimension of the fragmented tree graph representation.
[0041] Fusion and discrimination branch: For each substance, its fragmented dendrogram representation vector and molecular formula representation vector are concatenated to obtain the fusion representation of a single substance; for the substances at both ends of the spectral pair, their fusion representations are taken respectively, and the cosine similarity, difference vector and element-wise product between the two are calculated. The fusion representations at both ends and the above three relationship features are concatenated to form the spectral pair feature vector, which is input into a multilayer fully connected network for nonlinear mapping, and finally outputs the probability that the spectral pair "exists a reaction relationship".
[0042] S3. Collect and pre-treat environmental water samples from the industrial park to obtain an environmental sample fragmentation tree set. Specifically, surface water and wastewater samples are collected from different monitoring points within the industrial park, including the inlet and outlet of the wastewater treatment plant and the downstream receiving water body. The collected water samples are filtered to remove suspended particulate matter. Optionally, after pH adjustment, solid-phase extraction is used to enrich and purify high-risk pollutants in the water samples. The solid-phase extraction column is eluted with an appropriate solvent to obtain a sample solution, which is then concentrated by nitrogen blowing and redissolved in the corresponding solvent (e.g., methanol) to obtain a test solution enriched with high-risk pollutants. Subsequently, the test solution is analyzed using liquid chromatography-high resolution mass spectrometry (LC-MS / MS). Based on chromatographic separation, secondary mass spectrometry data of the characteristic peaks of each high-risk pollutant in the sample are obtained. Based on a fragmentation tree construction algorithm, the secondary mass spectrometry data is converted into a fragmentation tree containing fragment ions and neutral loss relationships, resulting in an environmental sample fragmentation tree set.
[0043] S4. Select at least one known high-risk pollutant (target high-risk pollutant) fragmentation tree from the environmental sample fragmentation tree set as the target fragmentation tree; combine the target fragmentation tree with other fragmentation trees in the environmental sample fragmentation tree set to form a "target fragmentation tree - candidate fragmentation tree" combination, and input it into the trained graph neural network model. The graph neural network model calculates the probability of "existence of reaction relationship" for each combination, and based on the optimal discrimination threshold, selects candidate fragmentation trees that have a reaction relationship with the secondary spectrum signal of the target high-risk pollutant, i.e., candidate high-risk pollutants with reaction relationship.
[0044] The probability of "existence of a reaction relationship" is calculated as follows: a two-way discrimination strategy is adopted, that is, two spectral pairs in the order of (secondary spectral signal of target high-risk pollutant, secondary spectral signal of candidate high-risk pollutant) and (secondary spectral signal of candidate high-risk pollutant, secondary spectral signal of target high-risk pollutant) are input into the graph neural network model respectively to obtain the probability values in the two directions, and the maximum probability value is taken as the probability that the spectral pair has a reaction relationship.
[0045] S5. For the obtained candidate fragmented trees, further combine the precise mass difference between the corresponding parent ion and the parent ion of the target high-risk pollutant, and refer to the predefined environmental reaction and transformation rules (such as side chain oxidation, carboxylation, dehalogenation, polyoxyethyleneization, etc.). If the reaction rules are met, determine whether the substance corresponding to the candidate fragmented tree belongs to the structural analogue of the target high-risk pollutant. On the basis of determining that there is a reaction relationship, according to the magnitude and sign of the mass difference and the reaction direction in the reaction rules, further subdivide the candidate pollutant into possible precursors or downstream transformation products of the target high-risk pollutant.
[0046] S6. Combine the judgment result with the process information of each process unit in the industrial park and the spatial distribution of different sampling points to analyze the location and evolution relationship of the target high-risk pollutants and their precursors and downstream conversion products in the park, so as to achieve the source tracing of the secondary spectrum signal of high-risk pollutants in the industrial park.
[0047] Taking the process of identifying the transformation products of nonylphenol in a water sample from an industrial park in Taizhou, Jiangsu Province, and tracing the source of the spectral signal as an example, it includes:
[0048] 1. Sample collection and mass spectrometry analysis
[0049] Surface water and wastewater samples were collected from multiple monitoring points around Lingfei Chemical Plant in Taizhou Industrial Park, including the park's main discharge outlet, the inlet and outlet of the wastewater treatment plant, and the downstream receiving water body. The collected water samples were first filtered through a membrane filter to remove suspended particulate matter. If necessary, the pH was adjusted to weakly acidic or neutral to improve the enrichment efficiency of high-risk pollutants. Subsequently, solid-phase extraction was used to enrich and purify the high-risk pollutants in the water samples. Methanol was used for elution, and the eluent was concentrated by nitrogen blowing and then redissolved in methanol to obtain the analyte solution for mass spectrometry analysis.
[0050] The above-mentioned analyte solution was analyzed using ultra-high performance liquid chromatography-high resolution mass spectrometry (UHPLC-MS / MS). Simultaneous acquisition of primary full scan (MS^1) and data-dependent secondary mass spectrometry (MS^2) data was performed in both positive and negative ion modes. Data preprocessing software (such as MSDIAL) was used to process the raw mass spectrometry data, including: baseline correction and denoising of the raw spectra; conversion of the profile spectrum to a centroid spectrum; peak detection and extraction for each scan spectrum; retention time correction and characteristic peak clustering for the same component in different scans along the time dimension; identification and labeling of isotope peak clusters; and removal or merging of isotope peaks. Finally, a retention time-aligned list of characteristic peaks and their corresponding MS^2 data were obtained. 2 Spectrum.
[0051] Based on this, the secondary mass spectrometry data corresponding to each characteristic peak are input into the fragmentation tree construction algorithm. The fragmentation tree prediction method of SIRIUS software is used: candidate molecular formulas are enumerated based on the precise mass and isotopic distribution of the parent ion and each fragment ion; possible neutral loss edges are constructed between candidate molecular formulas; and different fragmentation trees are scored using an objective function that comprehensively considers mass difference matching, fragment intensity, isotopic fitting degree, and tree structure complexity. The fragmentation tree with the highest score is selected as the optimal fragmentation tree representation for that characteristic peak. Through the above steps, a set of fragmentation trees corresponding to each characteristic peak in the water sample from Taizhou Industrial Park is obtained, which serves as the graph structure input data for the graph neural network model of this invention.
[0052] 2. Target Spectrum Signal Selection and Model Execution
[0053] Based on previous monitoring results and toxicological information, this embodiment selected the characteristic peak of 4-nonylphenol (4-NP) as the target spectral signal for high-risk contaminants. First, using the precise mass and retention time of the precursor ion, the characteristic peak corresponding to 4-nonylphenol was identified in the preprocessed data, and its MS signal was examined. 2 The fragmentation pattern is consistent with that of the literature or standard spectrum; then the secondary mass spectrometry data corresponding to the characteristic peak is reconstructed using the above fragmentation tree construction algorithm to obtain the fragmentation tree structure of 4-nonylphenol, which is used as the target fragmentation tree in the method of this invention.
[0054] During the graph neural network model's execution phase, the fragmentation tree of 4-nonylphenol is used as one input, and the fragmentation trees corresponding to all other characteristic peaks in the water sample from Taizhou Industrial Park are sequentially used as the other input. The graph neural network model of this invention constructs a set of target fragmentation trees and candidate fragmentation trees. Using the trained graph neural network model, a probability score for "reaction relationship" is calculated for each pair of high-risk pollutants, and candidate spectral signals showing a reaction relationship with 4-nonylphenol are selected based on the optimal discrimination threshold determined in the validation set.
[0055] like Figure 3 As shown, this embodiment further visualizes the results where the probability exceeds the discrimination threshold: using the chemical structure or presumed structure corresponding to each candidate feature as network nodes, and using the graph neural network model of this invention to determine "reaction relationship exists" as network edges, a reaction relationship molecular network with 4-nonylphenol as the core node is constructed for subsequent identification of transformation products and spectral signal source tracing analysis.
[0056] 3. Identification results of nonylphenol conversion products
[0057] Using the method of this invention, 4-nonylphenol was used as the target spectral signal in water samples from Taizhou Industrial Park. Figure 4 Multiple candidate spectral signals with reaction relationships were identified. Combining the precise mass of the parent ion, retention time difference, known transformation reaction types (including side-chain oxidation, carboxylation, and polyoxyethyleneization), and fragment characteristics, structural analysis was performed on the candidate spectral signals, ultimately confirming the following four nonylphenol transformation products or structural analogs:
[0058] (1) 4-(9-carboxynonyl)phenol (NP-TP235)
[0059] This high-risk pollutant undergoes stepwise oxidation at the nonyl side chain end of 4-nonylphenol, ultimately forming a carboxyl group, which is a typical ω-oxidation / carboxylation conversion product. The graph neural network model identifies the spectral pair corresponding to 4-nonylphenol and 4-(9-carboxynonyl)phenol as having a "reaction relationship" and forms an edge in the molecular network pointing from 4-nonylphenol to the terminal carboxylic acid. For example... Figure 5 The image shows the secondary mass spectrum of 4-(9-carboxynonyl)phenol. Its fragment peaks simultaneously retain the characteristic ions of the p-hydroxybenzene ring skeleton and the neutral loss information reflecting the carboxylation of the side chain, which is highly similar to the fragmentation mode of 4-nonylphenol in the aromatic ring region.
[0060] (2) 4-Nonylphenol polyoxyethylene ether (NP-TP263)
[0061] This high-risk contaminant is an ether-like derivative of 4-nonylphenol, with a polyoxyethylene chain attached to the molecule. It is one of the representative structures of nonylphenol polyoxyethylene ether surfactants. A graph neural network model determined that the spectra of 4-nonylphenol and 4-nonylphenol-polyoxyethylene ether show a reactive relationship; the parent ion mass increases by approximately one C2H4O repeating unit compared to 4-nonylphenol, and the retention time changes accordingly. Figure 6 The image shows a secondary mass spectrum of 4-nonylphenol-polyoxyethylene ether. The spectrum shows both aromatic ring fragments identical to those of 4-nonylphenol and characteristic fragments corresponding to the broken polyoxyethylene chains, proving that it is a polyoxyethyleneized product based on the 4-nonylphenol structure.
[0062] (3) 4-Nonylphenoxyacetic acid (NP-TP277)
[0063] This high-risk pollutant can be considered as 4-nonylphenol undergoing O-alkylation followed by esterification / condensation with acetic acid to form a phenoxyacetic acid structure. This is a type of transformation product of nonylphenol in the environment, where a carboxyl group is introduced into the side chain. A graph neural network model determines a reaction relationship between 4-nonylphenol and 4-nonylphenoxyacetic acid and connects them within the same cluster in the molecular network. Figure 8 The image shows the secondary mass spectrum of 4-nonylphenoxyacetic acid. Fragment peaks representing the benzene ring-nonyl skeleton and characteristic ions corresponding to the breakage of the acetic acid group can be observed. Its fragmentation mode has significant commonalities with the aromatic ring portion of 4-nonylphenol.
[0064] (4) 4-Nonylphenoxyethoxyacetic acid (NP-TP321)
[0065] This high-risk pollutant introduces an ethoxy unit into 4-nonylphenoxyacetic acid, structurally representing a product of 4-nonylphenol undergoing O-alkylation and ethoxy extension followed by binding with acetic acid. The graph neural network model also identifies a "reaction relationship" between this pollutant and 4-nonylphenol, demonstrating a possible stepwise transformation pathway through the tandem edge between them. Figure 7The image shows the secondary mass spectrum of 4-nonylphenoxyethoxyacetic acid. In addition to retaining the characteristic fragments of the 4-nonylphenol aromatic skeleton, the spectrum also shows a series of ion peaks reflecting the fragmentation of the ethoxy-acetic acid fragment, which shows a step-like difference from the spectrum of 4-nonylphenoxyacetic acid.
[0066] In addition to the four transformation products mentioned above, this embodiment can further display the secondary spectrum of 4-nonylphenol itself as a reference for the target spectral signal, such as... Figure 4 As shown. By comparison Figures 4-8 The five spectra clearly show that all transformation products exhibit a high degree of consistency with 4-nonylphenol in the aromatic ring fragment region, while the side chain and linking group fragment regions show their own unique neutral loss and fragment ions. The graph neural network model of this invention uses this fragmented tree pattern of "local conservatism + local variation" to determine whether there is a reaction relationship between the spectra.
[0067] 5. Source tracing analysis and effect description
[0068] Based on the analysis in this embodiment, a structure such as Figure 3 The molecular network of reaction relationships with 4-nonylphenol as the core node is shown. All nodes, including 4-nonylphenol, 4-(9-carboxynonyl)phenol, 4-nonylphenol-polyoxyethylene ether, 4-nonylphenoxyacetic acid, and 4-nonylphenoxyethoxyacetic acid, are connected by edges determined by the graph neural network model of this invention to indicate a "reaction relationship," forming a transformation spectrum extending from the parent nonylphenol towards carboxylation, polyoxyethyleneization, and acetylation. By combining the frequency and concentration distribution of each substance at different sampling points, the main transformation pathways and key intermediates of nonylphenol in wastewater treatment and downstream receiving water bodies in Taizhou Industrial Park can be inferred, providing direct evidence for the control and management of high-risk pollutants in the park.
[0069] This embodiment demonstrates that the method of the present invention can automatically screen and identify multiple transformation products and structural analogs of 4-nonylphenol, a known high-risk pollutant, in environmental samples from complex industrial parks, using the target spectral signal. Even if there is a lack of commercial standards for the relevant high-risk pollutants or they are not included in existing standard spectral libraries, effective source tracing can still be achieved through reaction relationships in secondary spectral fragmentation trees. This verifies the feasibility and superiority of the present invention in practical engineering scenarios.
Claims
1. A method for tracing the source of high-risk pollutants and their structural analogs in industrial parks based on mass spectrometry machine learning, characterized in that, Includes the following steps: S1. Construct a fragmented tree combination dataset labeled "reaction relationship exists / reaction relationship does not exist". The dataset includes several fragmented tree combinations, and each fragmented tree combination includes two fragmented trees representing different known high-risk pollutants. S2. Train the graph neural network model using the fragmented tree combination dataset; S3. Obtain fragmented tree sets of high-risk pollutants contained in environmental water samples from the industrial park; S4. Select at least one fragmentation tree with a known high-risk pollutant from the fragmentation tree set as the target fragmentation tree. Combine the target fragmentation tree with other fragmentation trees in the fragmentation tree set to form a "target fragmentation tree - candidate fragmentation tree" combination. Input the combination into the trained graph neural network model to determine whether there is a reaction relationship and filter out the candidate fragmentation trees with a reaction relationship. S5. Determine whether the high-risk pollutant corresponding to the candidate fragmentation tree is a structural analog of the high-risk pollutant corresponding to the target fragmentation tree; S6. Based on the judgment results of S5, conduct source tracing of structural analogues of high-risk pollutants within the industrial park.
2. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 1, characterized in that, The method for constructing the fragmented tree combined dataset in S1 is as follows: obtain environmental reaction secondary mass spectrometry data of known high-risk pollutants and their transformation products from the environmental reaction database, and predict the fragmented tree of the known high-risk pollutants based on the fragmented tree construction algorithm; Based on the reaction relationships between high-risk pollutants recorded in the environmental reaction database and process reaction pathway, high-risk pollutant pairs that have a direct upstream-downstream relationship or are located in the same reaction network are selected, and the two fragmented trees corresponding to the high-risk pollutant pairs are combined and marked as "having a reaction relationship". High-risk pollutant pairs with no known reaction relationship or belonging to different structural families are selected, and the fragmented tree assemblies corresponding to the high-risk pollutant pairs are marked as "no reaction relationship exists", thereby constructing a fragmented tree assemblies dataset labeled "reaction relationship exists / no reaction relationship exists".
3. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 2, characterized in that, In the fragmentation tree composite dataset, each fragmentation tree is represented as a graph structure, where fragment ions are nodes and neutral loss is an edge. Node features include at least fragment mass-to-charge ratio m / z, intensity, fragmentation score, and mass deviation, and edge features include at least the neutral loss score.
4. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 1, characterized in that, The graph neural network model described in S2 is trained by dividing the dataset into a training set and a validation set. During the training phase, the parameters of the graph neural network model are optimized by minimizing the loss function between the predicted probability and the labels in the training set. During the evaluation phase, different probability thresholds are scanned on the validation set, and performance metrics including accuracy, recall, and F1 score are calculated. The probability threshold with the best performance is selected as the optimal discrimination threshold.
5. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 1, characterized in that, The method described in S3 for obtaining the fragmented tree set of high-risk pollutants contained in environmental water samples from industrial parks is as follows: the water samples collected in the industrial park are filtered, high-risk pollutants in the water samples are extracted by solid phase extraction, the solid phase extraction column is eluted with solvent, and nitrogen blowing concentration is performed to obtain the test solution. The test solution is analyzed by liquid chromatography-high resolution mass spectrometry, and secondary mass spectrometry data of the characteristic peaks of each high-risk pollutant in the sample are obtained based on chromatographic separation. Based on the fragmentation tree construction algorithm, the secondary mass spectrometry data is converted into fragmentation trees containing fragment ions and neutral loss relationships, resulting in a fragmentation tree set.
6. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 1, characterized in that, The graph neural network model described in S4 calculates the probability of "existence of a reaction relationship" for each input "target fragmentation tree – candidate fragmentation tree" combination, and based on the optimal discrimination threshold, selects candidate fragmentation trees that have a reaction relationship with the high-risk pollutants corresponding to the selected target fragmentation tree.
7. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 6, characterized in that, The probability of "existence of a reaction relationship" is calculated as follows: the two spectral pairs of "target fragmentation tree - candidate fragmentation tree" are input into the graph neural network model respectively to obtain the probability values in the two directions, and the maximum probability value is taken as the probability of "target fragmentation tree - candidate fragmentation tree" having a reaction relationship.
8. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 1, characterized in that, In S5, for each candidate fragmentation tree, the precise mass difference between the parent ion of the candidate high-risk pollutant corresponding to the candidate fragmentation tree and the parent ion of the target high-risk pollutant with which there is a reaction relationship is calculated; the precise mass difference is matched with predefined environmental reaction and transformation rules; if the reaction rules are met, the candidate high-risk pollutant corresponding to the candidate fragmentation tree is determined to be a structural analog of the target high-risk pollutant; the target high-risk pollutant is the high-risk pollutant corresponding to the selected target fragmentation tree.
9. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 8, characterized in that, Based on the determination of their classification, and according to the magnitude and sign of the mass difference and the direction of the reaction in the reaction rules, the structural analogues are further subdivided into precursors or downstream conversion products of the target high-risk pollutants.
10. The method for tracing the source of high-risk pollutants and structural analogs in industrial parks according to claim 9, characterized in that, Based on the judgment results of S5, the process information of each process unit in the industrial park, and the spatial distribution of sampling points, S6 analyzes the location and evolution relationship of the target high-risk pollutant and its corresponding precursors and downstream conversion products in the park, thereby tracing the source of high-risk pollutants through structural analogues within the industrial park.