Method for screening and identifying high-risk pollutants in sewage based on fragmentation tree pre-training

By constructing a fragmented tree pre-training dataset and a graph neural network self-supervised pre-training, rapid screening and prioritization of high-risk pollutants in wastewater were achieved, and the target molecular structure was identified. This solved the problem of identifying pollutants in wastewater with many types and rapid changes in existing technologies, and improved the sensitivity and reliability of screening.

CN121765490BActive Publication Date: 2026-05-12NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-03-03
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for identifying high-risk pollutants in wastewater rely on standards and reference spectral libraries, which are insufficient to meet the needs of rapid screening for a wide variety of pollutants that are changing rapidly. Furthermore, complex matrix interference can lead to missed detections and misjudgments, resulting in insufficient robustness.

Method used

A fragmentation tree pre-training dataset is constructed, and a graph neural network is used for self-supervised pre-training to generate a fragmentation tree encoder. The fragmentation tree set is used to screen and prioritize high-risk pollutant signals and identify target molecular structures.

Benefits of technology

It enables rapid screening and prioritization of high-risk pollutants in wastewater, provides structural identification results, solves the dependence of existing methods on the completeness of standards and spectral libraries, and improves the sensitivity and reliability of screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765490B_ABST
    Figure CN121765490B_ABST
Patent Text Reader

Abstract

The application discloses a sewage high-risk pollutant screening and identification method based on fragmentation tree pre-training, comprising the following steps: constructing a fragmentation tree pre-training data set according to secondary mass spectrum data of known high-risk pollutants, pre-training a graph neural network encoder in a self-supervised manner to obtain a fragmentation tree editor, obtaining secondary mass spectrum data of suspected high-risk pollutant-related compounds in a sewage sample to be tested, constructing a fragmentation tree set, constructing a high-risk pollutant screening model based on the fragmentation tree editor, screening high-risk pollutant signals from the fragmentation tree set to obtain a high-risk candidate fragmentation tree set, generating a candidate molecule set of each high-risk candidate fragmentation tree, and identifying a target molecular structure and a matching score of the high-risk candidate fragmentation tree from the candidate molecule set. The application realizes rapid screening and priority sorting of high-risk pollutants in sewage, and further provides a candidate output of structure identification on the basis of screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring and analytical chemistry, and in particular to a method for screening and identifying high-risk pollutants in wastewater based on fragmented tree pre-training. Background Technology

[0002] Wastewater often contains a variety of high-risk pollutants or substances of high risk concern, such as industrial additives and intermediates, pharmaceutical and personal care product-related substances, pesticides and related substances, disinfection byproducts, surfactants and their derivatives, halogenated / sulfur-containing / phosphorus-containing functional compounds, and chemicals with risks of biotoxicity, carcinogenicity / mutagenicity / teratogenicity, or endocrine disruption. These pollutants in real wastewater samples have complex origins, diverse compositions, and wide concentration ranges, and are significantly affected by matrix effects, resulting in a large number of suspected high-risk related signals in the secondary mass spectrometry data acquired by LC-HRMS that require identification and priority treatment.

[0003] Current methods for identifying high-risk pollutants in wastewater include targeted detection based on standards and suspicious screening based on standard spectral libraries. Targeted detection typically uses standards as a reference, comparing the retention time and secondary spectra of the target signal for confirmation, offering high confirmation accuracy. However, this method relies on a pre-defined target list and the availability of standards, making it difficult to meet the application requirements of wastewater samples, which are characterized by a wide variety of high-risk pollutants, rapid changes, and the need for quick screening and dynamic expansion. Suspicious screening is usually based on a candidate list and a reference spectral library, performing precise mass matching of the target signal with candidate substances in the library, comparing characteristic fragments, and calculating spectral similarity to complete candidate screening. However, this method is highly dependent on the coverage and update speed of the reference spectral library. When the library is incomplete, different platforms / conditions lead to spectral differences, or the sample matrix has a significant impact, it is prone to missed detections, misjudgments, or the inability to consistently provide reliable screening results.

[0004] In addition, real wastewater samples generally have complex matrix interference, co-elution, ion suppression and multi-source cascading, which makes screening strategies that rely solely on local fragment matching or spectral similarity insufficiently robust. This makes it difficult to achieve highly sensitive screening and reliable prioritization of signals related to high-risk pollutants, thus affecting subsequent manual review, targeted confirmation and risk management decisions. Summary of the Invention

[0005] Purpose of the invention: The purpose of this invention is to provide a method for screening and identifying high-risk pollutants in wastewater based on fragmented tree pre-training, so as to achieve rapid screening and priority ranking of high-risk pollutants in wastewater, and further provide candidate outputs for structure identification based on the screening.

[0006] Technical Solution: To achieve the above objectives, the wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training described in this invention includes the following steps:

[0007] S1. Construct a fragmentation tree pre-training dataset based on the secondary mass spectrometry data of known high-risk pollutants;

[0008] S2. Use the fragmented tree pre-training dataset to perform self-supervised pre-training on the graph neural network encoder to obtain the fragmented tree editor.

[0009] S3. Obtain secondary mass spectrometry data of compounds related to suspected high-risk pollutants in the wastewater sample to be tested, and construct a fragmented tree set;

[0010] S4. Construct a high-risk pollutant screening model based on the fragmented tree encoder to screen the fragmented tree set for high-risk pollutant signals and obtain a high-risk candidate fragmented tree set.

[0011] S5. Generate a set of candidate molecules for each high-risk candidate fragmentation tree in the set of high-risk candidate fragmentation trees;

[0012] S6. Identify the target molecular structure and matching score of the high-risk candidate fragmentation tree from the candidate molecule set.

[0013] Preferably, the secondary mass spectrometry data includes precursor ion information of high-risk pollutants and fragment ion spectrum data corresponding to the precursor ion; the construction of the fragmented tree pre-training dataset also includes spectral preprocessing of the secondary mass spectrometry data, which includes: identifying discrete peaks in the fragment ion spectrum data of the secondary mass spectrometry data to obtain a fragment ion peak list; removing noise peaks and low-confidence peaks from the fragment ion peak list; labeling and merging or removing isotope peaks in the fragment ion peak list and / or the peaks near the precursor ion in the secondary mass spectrometry data; and identifying adducts based on the precursor ion information.

[0014] Preferably, the method for constructing the fragmented tree pre-training dataset is characterized by:

[0015] For each secondary mass spectrometry data, the parent ion corresponding to the parent ion peak is used as the root node, and each fragment ion peak in the fragment ion peak list is used as a candidate node to form a candidate node set.

[0016] For the root node and the set of candidate nodes, as well as among the candidate nodes, the candidate neutral loss is calculated based on the mass difference between the ion mass of the parent node and the ion mass of the child node. The parent node, child node, and neutral loss are then associated to generate a set of candidate edges. Each candidate edge carries a neutral loss mass value Δm.

[0017] Under the conditions of satisfying the constraints of mass conservation, the relationship between parent ions and fragments, and the rationality of elemental composition, a scoring function is used to score the credibility of candidate nodes and candidate edges. Through tree structure search or optimization, a set of edges is selected from the candidate nodes and candidate edges such that: a directed acyclic tree starting from the root node is formed, the total score is optimized, or a preset optimal criterion is met, thereby obtaining a fragmentation tree corresponding to the secondary mass spectrometry data.

[0018] Generate the graph structure representation of the fragmentation tree, and summarize the graph structure representations of the fragmentation trees generated from all secondary mass spectrometry data to form a fragmentation tree pre-training dataset;

[0019] The graph structure representation includes: a set of nodes with parent ions and fragment ions as nodes, a set of edges with neutral loss as directed edges, node features, and edge features; the node features include fragment ion mass-to-charge ratio m / z, peak intensity, mass deviation, and / or fragmentation score; the edge features include neutral loss mass, loss composition information, and loss score.

[0020] Preferably, the fragmented tree editor is used to map the graph structure of the fragmented tree into a fragmented tree graph-level representation vector, which is obtained by self-supervised pre-training of the graph neural network encoder. During the self-supervised pre-training process of the graph neural network encoder, one or more task heads are used to implement the output layer structure of the self-supervised pre-training task. The self-supervised pre-training task includes node mask reconstruction task, edge mask reconstruction task, contrastive learning task, and spectrogram and fragmented tree consistency constraint task. The parameters of the graph neural network encoder are updated by minimizing the total pre-training loss function.

[0021] Preferably, the total pre-training loss function is a weighted sum of the losses of each supervised pre-training task, expressed as:

[0022] ,

[0023] in, The total loss function for pre-training. For node mask reconstruction loss, For edge mask reconstruction loss, To compare learning loss, The loss represents the consistency between the spectrum and the fragmented tree. These are the weight coefficients for node mask reconstruction loss, edge mask reconstruction loss, contrastive learning loss, and consistency loss between the spectrogram and the fragmented tree, respectively.

[0024] Preferably, the screening model includes: a fragmented tree encoder and a screening head connected to the output of the fragmented tree encoder, wherein the screening head is a classification head used to map the fragmented tree graph-level representation vector output by the fragmented tree encoder into a screening probability score.

[0025] Preferably, threshold screening or Top-N screening is performed based on the screening probability score to obtain a set of high-risk candidate fragmented trees. Top-N screening involves sorting all fragmented trees from high to low according to their screening probability scores within the same batch, the same sampling point, or the same chromatographic time window, and then retaining the top N fragmented trees.

[0026] Preferably, the screening model is obtained through supervised fine-tuning; the supervised fine-tuning involves training the fragmented tree encoder and the screening head using a labeled fragmented tree dataset while keeping the fragmented tree encoder network structure unchanged, so that the screening probability score can be used to distinguish between "high-risk pollutants that are relevant or irrelevant".

[0027] Preferably, the method for generating the candidate molecule set of the high-risk candidate fragmentation tree is as follows: converting the parent ion mass of the high-risk candidate fragmentation tree to obtain the neutral mass M of the compound corresponding to the candidate fragmentation tree; searching in the structural candidate library with an error window of ±ppm of the neutral mass M, and screening to obtain a set of candidate molecules that meet the error window constraint, thereby forming a relationship between one fragmentation tree and multiple candidate molecules.

[0028] Preferably, the method for identifying the target molecular structure and matching score corresponding to the high-risk candidate fragmentation tree from the candidate molecule set is as follows:

[0029] A structure recognition model is constructed, including a fragmentation tree encoding branch and a molecular structure encoding branch. The fragmentation tree encoding branch maps the fragmentation tree graph structure to fragmentation tree embedding vectors. The molecular structure encoding branch maps candidate molecules to molecular embedding vectors. ;

[0030] For a given candidate fragmentation tree and all corresponding candidate molecules, the matching score between the fragmentation tree and all candidate molecules is calculated according to the matching function. The candidate molecules are sorted from high to low according to the matching score, and the top 10 candidate molecular structures are selected as the target molecular structures of the high-risk candidate fragmentation tree.

[0031] Beneficial effects: The present invention has the following advantages: 1. The present invention uses fragmented tree pre-training representation learning and graph neural network to mine deep features of mass spectrometry, so as to realize rapid screening and priority ranking of high-risk pollutants in wastewater, solve the dependence of existing targeted detection and suspicious screening on the completeness of standards and reference spectral libraries, and effectively cope with the screening needs of multiple types of pollutants and rapid changes; 2. On the basis of realizing rapid screening and priority ranking of high-risk pollutants in wastewater, it can further obtain the target molecular structure and matching score of high-risk pollutants, providing material structure basis for subsequent source tracing analysis. Attached Figure Description

[0032] Figure 1 This is a schematic flowchart of the method of the present invention;

[0033] Figure 2 This is a schematic diagram of the MS / MS data record obtained in Example 2 (positive ion mode);

[0034] Figure 3 This is a schematic diagram of the fragmentation tree constructed from MS / MS data records in Example 2;

[0035] Figure 4 This is a schematic diagram of the Top-K candidate molecular structures output in Example 2, showing the results of structure identification and sorting. Detailed Implementation

[0036] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.

[0037] Example 1

[0038] like Figure 1 As shown, the wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training provided by the present invention includes the following steps:

[0039] S1. Construct a fragmentation tree pre-training dataset based on the secondary mass spectrometry data of known high-risk pollutants.

[0040] S101. Obtain secondary mass spectrometry data of several known high-risk pollutants (including standards for high-risk pollutants and substances corresponding to regulatory or toxicity lists); wherein, the secondary mass spectrometry data are tandem mass spectrometry data (MS / MS, also known as secondary mass spectrometry data) records acquired by liquid chromatography-high resolution mass spectrometry, and each MS / MS data record includes at least:

[0041] Precursor ion information: mass-to-charge ratio m / z, charge number z, adduct form, retention time, collision energy and / or collection mode;

[0042] Fragment ion spectral data: Fragment ion spectra corresponding to the parent ion are stored in the form of data point sequences, which are several fragment ion mass-to-charge ratio m / z – peak intensity I data points; the fragment ion spectral data can be centroid data or profile data, where centroid data are centroided discrete peak data points, and profile data are contour-form continuous sampling data points.

[0043] S102. Perform spectral preprocessing on each secondary mass spectrometry data. The spectral preprocessing includes at least one or more of the following, and the processing targets are the fragment ion peak list and / or precursor ion information:

[0044] Peak extraction: Discrete peaks are identified in the fragment ion spectrum data (centroid or profile data point sequence) to obtain a fragment ion peak list; wherein, the fragment ion peak list is a set of several fragment ion peaks, each fragment ion peak at least contains the fragment ion mass-to-charge ratio m / z and peak intensity I, and may optionally include attributes such as peak width, signal-to-noise ratio and / or mass deviation; the output of the peak extraction (i.e., the fragment ion peak list) is used for subsequent fragmentation tree node candidate generation.

[0045] Denoising: The fragment ion peak list output by peak extraction is used as the processing object. The fragment ion peak list is a set of fragment ion peaks, each of which contains at least the fragment ion mass-to-charge ratio m / z and peak intensity I (optionally including attributes such as peak width, signal-to-noise ratio, and mass deviation). The fragment ion peak list is filtered by signal-to-noise ratio threshold, relative intensity threshold, Top-M retention and / or baseline / blank subtraction to remove noise peaks and low-confidence peaks. The denoised data is characterized by a reduced number of peaks, removal of low-intensity random peaks, and a more concentrated distribution of retained peak intensities.

[0046] Isotope identification: Based on isotope mass difference and relative abundance rules, isotope peaks are labeled and merged / removed from the fragment ion peak list and / or peaks near the parent ion to avoid isotope peaks being mistakenly identified as independent fragments.

[0047] Adduct identification: Based on the parent ion m / z, charge number z, and common adduct set, the adduct form of the MS / MS record is determined, and the mass of the parent ion is converted into a neutral mass M accordingly, providing a unified mass benchmark for subsequent mass conservation and candidate molecule retrieval.

[0048] After spectral preprocessing, standardized MS / MS records are obtained, which include: parent ion information (m / z, z, adduct form) and a list of fragment ion peaks after denoising and deisotopeing (m / z–I).

[0049] S103. The standardized MS / MS records are converted into fragmentation trees based on a fragmentation tree construction algorithm. The fragmentation tree is a directed acyclic tree structure with the parent ion of the MS / MS record as the root node, used to characterize the hierarchical fragmentation relationship from parent ion to fragment ion. Each fragmentation tree contains a set of nodes and a set of edges. Nodes represent the fragmentation products corresponding to the parent ion or fragment ion (characterized by ion m / z and molecular formula composition information); edges represent the fragmentation transfer relationship from parent node to child node, and the corresponding neutral loss is determined by the mass difference between parent and child nodes.

[0050] The process of constructing the fragmented tree pre-training dataset is as follows:

[0051] S1031. Root Node and Candidate Node Generation: For each standardized MS / MS record, the parent ion corresponding to the parent ion peak of the record is used as the root node; each fragment ion peak in the fragment ion peak list of the record is used as a candidate node to form a set of candidate nodes.

[0052] S1032, Candidate Edge Generation: For the root node and the set of candidate nodes, and between candidate nodes, calculate the candidate neutral loss based on the mass difference between the ion mass of the parent node and the ion mass of the child node, and establish the association between the parent node, child node, and neutral loss to generate a set of candidate edges; wherein, each candidate edge carries at least the neutral loss mass value Δm;

[0053] S1033. Constraints, Scoring, and Tree Structure Determination: Evaluate candidate nodes and edges and determine the tree structure under the following constraints:

[0054] Quality conservation constraint: For any candidate edge, the quality of the parent node must be greater than or equal to the quality of the child node, and the difference between the two must be consistent with the quality of the candidate neutral loss (within the allowable quality error range).

[0055] Constraints on the relationship between parent ion and fragments: Requires that the path formed by nodes and edges is reachable from the root node, and that the child nodes are reasonable successors generated by the fragmentation of the parent node;

[0056] Element composition rationality constraints: If a node or neutral loss is assigned an element composition / molecular formula candidate, the composition is required to satisfy chemical rationality rules (e.g., non-negative element count, valence / unsaturation constraints, common element set constraints, etc.).

[0057] Under the above constraints, a scoring function is used to score the credibility of candidate nodes and candidate edges. This scoring function considers at least: peak intensity, quality bias, loss rationality, prior knowledge of fragmentation patterns, or weights obtained through training. Furthermore, through a tree structure search or optimization process (including one or more of dynamic programming, branch and bound, maximum spanning tree / shortest path variants, or heuristic search), a set of edges is selected from the candidate nodes and candidate edges such that:

[0058] Form a directed acyclic tree starting from the root node;

[0059] To achieve the optimal total score or meet the preset optimal criteria;

[0060] This results in a fragmented tree corresponding to the MS / MS record.

[0061] S1034. Graph Structure Output and Dataset Formation: Output the graph structure representation of the fragmented tree (each MS / MS record corresponds to one fragmented tree), with the parent ion and fragment ions as nodes and the neutral loss as directed edges, and output the node features and edge features of the fragmented tree; summarize the fragmented tree sets corresponding to all MS / MS records to form a fragmented tree pre-training dataset; each fragmented tree is independent of the others as samples.

[0062] The node features include: fragment ion mass-to-charge ratio m / z, peak intensity, mass deviation, and / or fragmentation score; wherein, the fragment ion mass-to-charge ratio m / z and peak intensity are derived from the fragment ion peak list of secondary mass spectrometry data, the mass deviation is the difference between the observed m / z and the theoretical m / z (derived from molecular formula candidates or mass conservation), and the fragmentation score is the score calculated by the scoring function for node credibility or fragmentation rationality.

[0063] Edge features include: neutral loss quality, loss composition information, and loss score; where, neutral loss quality is obtained from the quality difference between parent and child nodes, loss composition information is the candidate or category label composed of elements matched by the neutral loss quality within the allowable error, and loss score is the score calculated by the scoring function for the reasonableness of the edge.

[0064] S2. Use the fragmented tree pre-training dataset to perform self-supervised pre-training on the graph neural network encoder to obtain the fragmented tree editor.

[0065] To enable subsequent screening and priority ranking of high-risk signals in the fragmented tree of wastewater samples, as well as candidate structure matching and Top-K output of the high-risk candidate signals obtained from the screening, the graph neural network encoder is self-supervised pre-trained based on the fragmented tree pre-training dataset constructed in step S1 to obtain a fragmented tree encoder with generalization ability.

[0066] The fragmented tree encoder maps a fragmented tree graph structure to a fragmented tree graph-level representation vector; the fragmented tree graph structure includes a set of nodes, a set of edges, and their node and edge features; the fragmented tree encoder performs message passing / attention aggregation on the nodes and edges, and outputs:

[0067] Node-level representation: Output the node embedding vector for each node;

[0068] Edge-level representation: Output the edge embedding vector for each edge;

[0069] Graph-level representation (fragmented tree embedding vector): The fragmented tree embedding vector is obtained by aggregating the node-level representations through the readout function, which is used to represent the overall structural information of the fragmented tree;

[0070] Optionally, the graph neural network encoder may be implemented using a graph convolutional network, a graph attention network, a graph transformer, or an equivalent variant thereof.

[0071] In the self-supervised pre-training process of a graph neural network encoder, one or more task heads are used in conjunction with the graph neural network encoder to implement the output layer structure of the self-supervised pre-training task, including at least one or more of the following:

[0072] Node decoding head: Connected to the node-level representation output, used to predict the node features of the masked node;

[0073] Edge decoding head: Connected to the output of edge-level representation or edge-related aggregated representation, used to predict the edge features of the masked edge;

[0074] Projection head: Connected to the graph-level representation output, used to map the fragmented tree embedding to the embedding space required for contrastive learning or consistency constraints;

[0075] The task head can be a multilayer perceptron, a linear layer, or a combination thereof. The task head provides learning signals during the self-supervised training phase. In subsequent screening or structure recognition applications, only the fragmented tree encoder can be retained as the backbone network, and the projection head can be retained as needed.

[0076] The graph neural network encoder is self-supervised pre-trained using the fragmented tree pre-training dataset. The self-supervised pre-training task includes at least one or more of the following, and the parameters of the graph neural network encoder are updated by minimizing the total pre-training loss function. The total pre-training loss function is a weighted sum of the losses of each supervision task.

[0077] 1. Node mask reconstruction task

[0078] Node masking reconstruction: For each fragmented tree sample, at least one node is selected to perform a node feature masking operation. The masking operation is one or a combination of the following methods: setting one or more feature dimensions of the node to zero, setting them to a preset mask mark, or replacing them with noise; inputting the masked fragmented tree into a graph neural network encoder to obtain a node-level representation; inputting the node-level representation into a node decoding head to output the feature prediction value of the masked node, and using the error between the prediction value and the true node features before masking as the node reconstruction loss; wherein, when multiple feature dimensions are masked, the node reconstruction loss is the weighted sum or mean of the errors of each masked dimension.

[0079] The node mask reconstruction task forces the graph neural network encoder to recover node attributes from the neighborhood structure and context, thereby learning the local structural patterns and semantic representations of nodes in the fragmented tree.

[0080] 2. Side mask reconstruction task

[0081] Edge masking reconstruction: For each fragmented tree sample, at least one edge is selected to perform an edge feature masking operation. The masking operation is one or a combination of the following methods: setting one or more feature dimensions of the edge to zero, setting them to a preset mask mark, or replacing them with noise; inputting the masked fragmented tree into a graph neural network encoder to obtain an edge-related representation (edge-level representation or edge representation constructed from node-level representation); inputting the edge-related representation into an edge decoding head to output the feature prediction value of the masked edge, and using the error between the prediction value and the true edge features before masking as the edge reconstruction loss.

[0082] The edge mask reconstruction task prompts the graph neural network encoder to learn edge semantics such as "neutral loss / fragmentation transition relationship", thereby improving its ability to characterize the relationship between fragmented paths and tree structure.

[0083] 3. Comparison of learning tasks

[0084] Contrastive learning: Each fragmented tree sample is mapped to a graph-level representation via a graph neural network encoder and then mapped to a contrastive learning embedding vector via a projection head; training sample pairs are constructed such that positive sample pairs are closer in the embedding space and negative sample pairs are farther apart; wherein, positive sample pairs include at least one of the following: fragmented tree pairs constructed from different MS / MS (tandem mass spectrometry) records of the same compound, or fragmented tree pairs related to the same parent ion; negative sample pairs include fragmented tree pairs from different compounds; the relative distance of the embedding vectors is constrained by a contrastive loss function, wherein the contrastive loss function includes at least one or more of supervised contrastive loss, information noise contrastive estimation loss (InfoNCE), or triplet loss.

[0085] The contrastive learning task enhances the discriminative power of graph-level representations in identifying "similar fragmented trees are more similar and dissimilar trees are more separable", enabling the graph neural network encoder output to be a global embedding suitable for sorting / retrieval and subsequent screening head learning.

[0086] 4. Consistency constraint task between spectral graph and fragmented tree

[0087] Spectrum-Fragmentation Tree Consistency Constraint: For fragmentation trees generated from the same secondary mass spectrometry record, the spectral representation and fragmentation tree representation are calculated separately, and their consistency in the embedding space is constrained by a consistency loss. The spectral representation can be obtained by encoding the fragment ion peak list by a spectral encoder, or by obtaining the peak intensity vector composed of the fragment ion peak list through a projection head. The fragmentation tree representation is obtained by the graph-level representation output of a graph neural network encoder. The consistency loss includes at least one or more of the following: cosine distance loss, mean square error loss, or contrastive consistency loss.

[0088] The consistency constraint task aligns the fragmented tree structure information with the spectral information, improving the stability and generalization ability of the graph neural network encoder under different acquisition conditions and noise disturbances.

[0089] 5. Pre-training total loss function

[0090] The pre-training total loss function is a weighted sum of the losses of each supervised task, expressed as:

[0091] ,

[0092] in, The total loss function for pre-training. For node mask reconstruction loss, For edge mask reconstruction loss, To compare learning loss, The loss represents the consistency between the spectrum and the fragmented tree. These are the corresponding weighting coefficients.

[0093] After self-supervised pre-training is completed, a fragmented tree encoder is obtained; the fragmented tree encoder serves as the backbone network for the subsequent screening model in step S4 and the structure recognition model in step S6, and can be used for downstream training and inference by freezing or fine-tuning.

[0094] S3. Obtain secondary mass spectrometry data of compounds related to suspected high-risk pollutants in the wastewater sample to be tested, and construct a fragmented tree set.

[0095] Water samples were collected at multiple sampling points along the wastewater generation and treatment process. These sampling points included at least one or more of the following: influent, pretreatment unit, secondary biological treatment unit, advanced treatment unit, and effluent. The wastewater samples were pretreated, and secondary mass spectrometry data of compounds suspected to be high-risk pollutants in the wastewater samples were acquired using liquid chromatography-high resolution mass spectrometry (LC-MS). Pretreatment included at least one or more of the following: filtration, solid-phase extraction, solvent elution, nitrogen blowing concentration, reconstitution and volume adjustment, and internal standard correction.

[0096] Perform the same spectral preprocessing and fragmentation tree construction as in S1 on the secondary mass spectrometry data to obtain a fragmentation tree set of the wastewater sample to be tested; where "consistent" means that at least the parent ion information field, fragment ion peak list field, mass error unit and peak filtering rules are consistent with the fragmentation tree construction requirements of S1, so that the constructed fragmentation tree can be correctly parsed by the fragmentation tree encoder.

[0097] S4. Construct a high-risk pollutant screening model based on the fragmented tree encoder to screen high-risk pollutant signals from the fragmented tree set and obtain a high-risk candidate fragmented tree set.

[0098] The screening model includes a fragmented tree encoder and a screening head connected to the output of the fragmented tree encoder. The fragmented tree encoder maps the graph structure of the fragmented tree to the graph-level representation vector of the fragmented tree, and the screening head is a classification head used to map the graph-level representation vector of the fragmented tree to a screening probability score.

[0099] The fragmented tree set is input into the screening model, and the fragmented tree encoder maps each fragmented tree to a fragmented tree graph-level representation vector. Then, the screening head outputs the screening probability score of the high-risk pollutant related signal corresponding to the fragmented tree. The high-risk pollutant related signal refers to the spectral signal related to the substance corresponding to the high-risk pollutant list.

[0100] Based on the screening probability scores, perform threshold filtering or Top-N filtering to obtain a set of high-risk candidate fragmentation trees:

[0101] Threshold filtering retains fragmented trees whose screening probability score is ≥ a preset threshold;

[0102] Top-N screening involves sorting all fragmented trees by screening probability score from highest to lowest within the same batch, the same sampling point, or the same chromatographic time window, and retaining the top N fragmented trees; where N is a preset integer and is less than the number of fragmented trees within that range.

[0103] The screening model is obtained through supervised fine-tuning. This supervised fine-tuning involves training the fragmentation tree encoder and screening head using a labeled fragmentation tree dataset while maintaining the fragmentation tree encoder network structure. This ensures that the screening probability score can distinguish between "related / unrelated high-risk pollutants." The labeled fragmentation tree dataset originates from at least one or more of the following: secondary mass spectrometry data of standard samples, publicly available spectral library data, spectral data of substances corresponding to regulatory or toxicity lists, and spectral data of wastewater samples that have been manually verified or confirmed by external evidence. Positive sample fragmentation trees are constructed from substances corresponding to high-risk pollutant lists, while negative sample fragmentation trees are constructed from non-high-risk substances, common background substances, or spectra unrelated to the target. Screening thresholds are scanned on the validation set, and the optimal screening discrimination threshold is determined based on the F1 score.

[0104] S5. Generate a set of candidate molecules for each high-risk candidate fragmentation tree in the set of high-risk candidate fragmentation trees.

[0105] For each candidate fragmentation tree in the high-risk candidate fragmentation tree set, the neutral mass M of the corresponding compound is calculated based on the parent ion m / z, charge number z, and adduct form recorded by its corresponding MS / MS. The neutral mass M is then used to search at least one structural candidate library with an error window of ±ppm (parts per million) of the neutral mass M to obtain a set of candidate molecules that meet the error window constraints.

[0106] The candidate molecule set refers to a set of candidate molecule records that satisfy the neutral quality M-window constraint and are retrieved from the structural candidate library for the same candidate fragmentation tree, thus forming a relationship between a fragmentation tree and multiple candidate molecules.

[0107] The structure candidate library is a dataset containing molecular structure records, which include at least: Simplified Linear Input Specifications for Molecular Structures (SMILES), molecular formula, precise mass, and optional toxicity / use annotation information; the structure candidate library includes at least one or more of the following: public chemical structure library, toxicity-related structure library, industrial park target list library, hospital key drug and metabolite library, or municipal wastewater concern material library.

[0108] S6. Identify the target molecular structure and matching score corresponding to the high-risk candidate fragmentation tree from the candidate molecule set.

[0109] S601. Construct a structure recognition model, including a fragmented tree coding branch and a molecular structure coding branch:

[0110] The fragmented tree encoding branch takes the fragmented tree graph structure and its node / edge features as input, and maps the fragmented tree into fragmented tree embedding vectors through the fragmented tree encoder. , Indicates a fragmented tree. The fragmented tree mapping function is represented; wherein the fragmented tree encoder can reuse the fragmented tree encoder obtained by pre-training in step S2 and fine-tune it as needed.

[0111] The molecular structure encoding branch takes the structural representation of candidate molecules as input (including but not limited to molecular graphs, SMILES, or molecular fingerprints), and maps each candidate molecule to a molecular embedding vector through a molecular encoder. , Indicates candidate molecules, This represents the candidate molecule mapping function.

[0112] Each high-risk candidate fragmentation tree and its corresponding set of candidate molecules are input into the structure recognition model to obtain the embedding vector of the fragmentation tree. and the embedding vector of each candidate molecule. .

[0113] S602. For each candidate molecule corresponding to the same candidate fragmentation tree, calculate the matching score between the fragmentation tree and the candidate molecule based on the matching function. The candidate molecules of the fragmentation tree are reordered according to the matching score, and the top-K candidate molecular structures are output as the target molecular structures of the high-risk candidate fragmentation tree, and the matching score between the target molecular structures and the fragmentation tree is used as the structure identification result; wherein, the matching function includes at least one or more of cosine similarity, dot product, bilinear mapping or learnable matching network.

[0114] S603, based on the high-risk candidate fragmentation tree obtained from S4 screening and the structure identification results output from S6, combined with the spatial distribution of sampling points and wastewater process flow information, analyzes the location, migration, and transformation trends of the high-risk pollutant-related substances in industrial wastewater, hospital wastewater, and municipal wastewater systems, thereby achieving source identification and tracing of high-risk pollutants. The method described in this invention achieves screening of high-risk pollutant-related signals and identification of material structures in wastewater without relying on a complete standard spectral library.

[0115] Example 2

[0116] Taking the process of screening endocrine disruptors for signals and identifying spectral signals in wastewater samples from a hospital in Suzhou, Jiangsu Province as an example, the process includes:

[0117] 1. Obtain the wastewater sample to be tested.

[0118] Take 1.0 L of the hospital wastewater sample and perform pretreatment on the sample. The pretreatment includes at least one or more of the following: filtration, solid phase extraction, solvent elution, nitrogen blowing concentration, reconstitution and volume adjustment and internal standard correction, to obtain the sample to be analyzed by liquid chromatography-high resolution mass spectrometry.

[0119] 2. Obtain secondary mass spectrometry data and construct a fragmentation tree set (taking positive ions as an example).

[0120] Secondary mass spectrometry data were acquired using liquid chromatography-high resolution mass spectrometry in positive ion mode, resulting in 10,757 MS / MS data records for the wastewater samples. Each MS / MS data record included at least the following: precursor ion information (m / z, charge number z, adduct form, retention time, collision energy, and / or acquisition mode) and fragment ion spectrum data (several m / z–I data points).

[0121] Perform the same spectral preprocessing and fragmentation tree construction as in step S1 on the MS / MS data records, including:

[0122] Peak extraction: Identify discrete peaks in fragment ion spectrum data to obtain a list of fragment ion peaks;

[0123] Denoising: The fragment ion peak list was initially filtered using a relative intensity threshold of 0.5% and a signal-to-noise ratio threshold of 10. Based on this, the top 120 fragment ion peaks were retained in descending order of relative intensity.

[0124] Isotope identification: Isotope peak labeling and merging / removal of peaks near the parent ion and / or fragment ion peaks;

[0125] Adduct identification: The adduct form is determined based on the parent ion m / z, charge number z, and common adduct set, and the parent ion mass is converted to neutral mass M accordingly;

[0126] Fragmentation tree construction: Standardized MS / MS data records are converted into fragmentation trees, with the parent ion as the root node, fragment ions as nodes, and neutral loss as directed edges.

[0127] After the above processing, this embodiment obtains a fragmented tree set of the wastewater samples to be tested. Each MS / MS data record corresponds to one fragmented tree, and the final number of fragmented trees (sample number) is 10757.

[0128] This embodiment selects a representative MS / MS data record as an example. Its precursor ion information is as follows: precursor ion mass-to-charge ratio m / z = 194.12, charge number z = 1, and adduct type is potassium adduct, i.e., a positive adduct formed by the analyte and potassium ions; its corresponding secondary mass spectrometry fragment ion spectrum contains a base peak fragment ion peak m / z = 135.04, as shown below. Figure 2 As shown. After performing spectral preprocessing and constructing a fragmentation tree for this MS / MS data record, a fragmentation tree graph structure with the parent ion as the root node, fragment ions as nodes, and neutral loss as directed edges is obtained, as shown. Figure 3 As shown, the neutral loss between the parent ion node and the base peak node can be represented as -C3H9N.

[0129] 3. Construct a high-risk pollutant screening model based on the fragmented tree encoder and conduct screening.

[0130] In this embodiment, the set of fragmented trees (a total of 10757 trees) is input into the screening model. The fragmented tree encoder outputs a fragmented tree graph-level representation vector for each fragmented tree, and the screening head outputs the corresponding high-risk pollutant related signal screening probability score. The high-risk pollutant related signal refers to the spectral signal related to the substance corresponding to the high-risk pollutant list. This embodiment uses a preset screening discrimination threshold of 0.512 to perform threshold screening, retaining fragmented trees with a screening probability score ≥ 0.512, thereby obtaining a high-risk candidate fragmented tree set. Optionally, within the same batch, the same sampling point, or the same chromatographic time window, fragmented trees can be sorted from high to low according to their screening probability scores, and the top N fragmented trees can be retained to achieve Top-N screening. The screening probability scores are as follows: Figure 4 The filtering and sorting results are shown.

[0131] 4. Generate a set of candidate molecules for each high-risk candidate fragmentation tree in the set of high-risk candidate fragmentation trees.

[0132] For each candidate fragmentation tree in the high-risk candidate fragmentation tree set, the mass-to-charge ratio (m / z), charge number (z), and adduct form of the parent ion are read from its corresponding MS / MS data record. The mass of the parent ion is then converted according to the adduct form to obtain the neutral mass M of the compound corresponding to the candidate fragmentation tree. A search is conducted in at least one structure candidate library using an error window of ±10 ppm for the neutral mass M, and a set of candidate molecules that meet the error window constraint is obtained, thus forming a relationship between one fragmentation tree and multiple candidate molecules. The structure candidate library is a dataset containing molecular structure records, which at least include: Simplified Linear Input Specifications (SMILES), molecular formula, and precise mass.

[0133] Taking the representative fragmented tree as an example, the neutral mass M is calculated based on its parent ion information (m / z=194.12, z=1, adduct form), and a set of candidate molecules is obtained by searching the structure candidate library with an error window of ±10 ppm. The set of candidate molecules contains several candidate molecule structure records, and each candidate molecule structure record contains at least SMILES and the exact mass.

[0134] 5. Identify the target molecular structure and matching score corresponding to the high-risk candidate fragmentation tree from the candidate molecule set, and output the Top-K candidate structures.

[0135] In this embodiment, each high-risk candidate fragmentation tree and its corresponding set of candidate molecules are input into the structure recognition model to obtain the embedding vector of the fragmentation tree. and the embedding vector of each candidate molecule in its candidate molecule set. For each candidate molecule corresponding to the same candidate fragmentation tree, the matching score between the fragmentation tree and the candidate molecule is calculated based on the matching function. The candidate molecule set is reordered according to the matching score, and the Top-K candidate molecule structures and their matching scores are output as the structure identification results. The matching function includes at least one or more of cosine similarity, dot product, bilinear mapping, or a learnable matching network. In this embodiment, K=10, and the Top-10 candidate molecule structure records (represented by SMILES) and their corresponding matching scores are output, as follows: Figure 4 As shown.

[0136] This embodiment constructs 10,757 fragmented trees, outputs screening probability scores from the screening model, and completes screening with a discrimination threshold of 0.512 to obtain a set of high-risk candidate fragmented trees. Further, a candidate molecule set is generated with an error window of ±10 ppm, matching scores are calculated and reordered, and Top-K candidate molecule structures are output, realizing the screening and structural identification of high-risk pollutant-related signals in the wastewater sample. The method described in this invention can still achieve the screening of high-risk pollutant-related signals and the output of candidate structures in wastewater samples without relying on a complete standard spectral library.

Claims

1. A wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training, characterized in that, Includes the following steps: S1. Construct a fragmentation tree pre-training dataset based on the secondary mass spectrometry data of known high-risk pollutants; S2. Use the fragmented tree pre-training dataset to perform self-supervised pre-training on the graph neural network encoder to obtain the fragmented tree editor. S3. Obtain secondary mass spectrometry data of compounds related to suspected high-risk pollutants in the wastewater sample to be tested, and construct a fragmented tree set; S4. Construct a screening model for high-risk pollutants based on the fragmented tree encoder to screen the fragmented tree set for high-risk pollutant signals and obtain a set of high-risk candidate fragmented trees. S5. Generate a set of candidate molecules for each high-risk candidate fragmentation tree in the set of high-risk candidate fragmentation trees; S6. Identify the target molecular structure and matching score of the high-risk candidate fragmentation tree from the candidate molecule set.

2. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 1, characterized in that, The secondary mass spectrometry data includes information on the precursor ions of high-risk pollutants and fragment ion spectrum data corresponding to the precursor ions; The construction of the fragmented tree pre-training dataset also includes spectral preprocessing of the secondary mass spectrometry data. The spectral preprocessing includes: identifying discrete peaks in the fragmented ion spectrum data of the secondary mass spectrometry data to obtain a fragmented ion peak list; removing noise peaks and low-confidence peaks from the fragmented ion peak list; annotating and merging or removing isotopic peaks from the fragmented ion peak list and / or the peaks near the parent ion of the secondary mass spectrometry data; and identifying adducts based on the parent ion information.

3. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 2, characterized in that, The method for constructing the fragmented tree pre-training dataset: For each secondary mass spectrometry data, the parent ion corresponding to the parent ion peak is used as the root node, and each fragment ion peak in the fragment ion peak list is used as a candidate node to form a candidate node set. For the root node and the set of candidate nodes, as well as among the candidate nodes, the candidate neutral loss is calculated based on the mass difference between the ion mass of the parent node and the ion mass of the child node. The parent node, child node, and neutral loss are then associated to generate a set of candidate edges. Each candidate edge carries a neutral loss mass value Δm. Under the conditions of satisfying the constraints of mass conservation, the relationship between parent ions and fragments, and the rationality of elemental composition, a scoring function is used to score the credibility of candidate nodes and candidate edges. Through tree structure search or optimization, a set of edges is selected from the candidate nodes and candidate edges such that: a directed acyclic tree starting from the root node is formed, the total score is optimized, or a preset optimal criterion is met, thereby obtaining a fragmentation tree corresponding to the secondary mass spectrometry data. Generate the graph structure representation of the fragmentation tree, and summarize the graph structure representations of the fragmentation trees generated from all secondary mass spectrometry data to form a fragmentation tree pre-training dataset; The graph structure representation includes: a set of nodes with parent ions and fragment ions as nodes, a set of edges with neutral loss as directed edges, node features, and edge features; the node features include fragment ion mass-to-charge ratio m / z, peak intensity, mass deviation, and / or fragmentation score; the edge features include neutral loss mass, loss composition information, and loss score.

4. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 1, characterized in that, The fragmented tree editor is used to map the graph structure of the fragmented tree into fragmented tree graph-level representation vectors, which are obtained by the graph neural network encoder through self-supervised pre-training. During the self-supervised pre-training process of the graph neural network encoder, one or more task heads are used to implement the output layer structure of the self-supervised pre-training task. The self-supervised pre-training task includes node mask reconstruction task, edge mask reconstruction task, contrastive learning task, and spectrogram and fragmented tree consistency constraint task. The parameters of the graph neural network encoder are updated by minimizing the total pre-training loss function.

5. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 4, characterized in that, The total pre-training loss function is a weighted sum of the losses of each supervised pre-training task, expressed as: , in, The total loss function for pre-training. For node mask reconstruction loss, For edge mask reconstruction loss, To compare learning loss, This represents the consistency loss between the spectrum and the fragmented tree. These are the weight coefficients for node mask reconstruction loss, edge mask reconstruction loss, contrastive learning loss, and consistency loss between the spectrogram and the fragmented tree, respectively.

6. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 1, characterized in that, The screening model includes a fragmented tree encoder and a screening head connected to the output of the fragmented tree encoder. The screening head is a classification head used to map the fragmented tree graph-level representation vector output by the fragmented tree encoder into a screening probability score.

7. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 6, characterized in that, Threshold screening or Top-N screening is performed based on the screening probability score to obtain a set of high-risk candidate fragmented trees. The Top-N screening involves sorting all fragmented trees from high to low according to their screening probability scores within the same batch, the same sampling point, or the same chromatographic time window, and then retaining the top N fragmented trees.

8. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 6, characterized in that, The screening model is obtained through supervised fine-tuning; the supervised fine-tuning is to train the fragmented tree encoder and the screening head using a labeled fragmented tree dataset while keeping the fragmented tree encoder network structure unchanged, so that the screening probability score can be used to distinguish between "high-risk pollutants that are relevant or irrelevant".

9. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 1, characterized in that, The method for generating the candidate molecule set of the high-risk candidate fragmentation tree is as follows: convert the mass of the parent ion of the high-risk candidate fragmentation tree to obtain the neutral mass M of the corresponding compound of the candidate fragmentation tree; The candidate molecules are searched in the structure candidate library using the ±ppm error window of the neutral mass M, and a set of candidate molecules that meet the error window constraints is obtained, thereby forming a relationship between a fragmentation tree and multiple candidate molecules.

10. The wastewater high-risk pollutant screening and identification method based on fragmented tree pre-training according to claim 1, characterized in that, The method for identifying the target molecular structure and matching score corresponding to the high-risk candidate fragmentation tree from the candidate molecule set is as follows: A structure recognition model is constructed, including a fragmentation tree encoding branch and a molecular structure encoding branch. The fragmentation tree encoding branch maps the fragmentation tree graph structure to fragmentation tree embedding vectors. The molecular structure encoding branch maps candidate molecules to molecular embedding vectors. ; For a given candidate fragmentation tree and all corresponding candidate molecules, the matching score between the fragmentation tree and all candidate molecules is calculated according to the matching function. The candidate molecules are sorted from high to low according to the matching score, and the top 10 candidate molecular structures are selected as the target molecular structures of the high-risk candidate fragmentation tree.