Method for analyzing water body environment pollutant conversion based on GC-HRMS data

By constructing a high-confidence structural similarity molecular network using GC-HRMS data, the problem of difficulty in revealing pathways and transformation patterns in the monitoring of water pollutants was solved, and the accurate identification and visualization of compound associations were achieved.

CN121275932APending Publication Date: 2026-01-06NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511441471.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods for monitoring water pollutants are insufficient to systematically reveal the migration pathways and transformation patterns of pollutants, have limited ability to identify unknown degradation products, involve complex data processing, and lack efficient visualization methods.

Method used

Based on GC-HRMS data, a molecular network with high confidence is constructed by using a structural similarity molecular network construction method, extracting mass spectrometry features using an automatic deconvolution algorithm, calculating the chemical structural similarity between compounds, and introducing chemical transformation rules for verification.

Benefits of technology

It achieves highly reliable identification of compound associations, systematically reveals environmental transformation pathways, provides objective macroscopic evaluation indicators, reduces false positive rates, and is suitable for the analysis of complex samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121275932A_ABST
    Figure CN121275932A_ABST
Patent Text Reader

Abstract

The invention discloses a method for analyzing water body environment pollutant conversion based on GC-HRMS (Gas Chromatography-High Resolution Mass Spectrometry) data, which comprises the following steps: acquiring GC-HRMS mass spectrum data of a water body environment, and extracting mass spectrum characteristics of each component; performing compound qualification based on the mass spectrum characteristics to obtain a candidate compound, and converting the candidate compound into a standardized structure definition formula; calculating the chemical structure similarity between the candidate compounds, and constructing a structure similarity molecular network; verifying connections in the structural similarity molecular network by applying a chemical conversion rule, and screening and outputting an environment conversion path with high reliability; according to the invention, structural group characteristics of compounds in complex samples can be systematically revealed; the false positive is reduced, the reliability of the incidence relation of the compounds in the molecular network is improved, and a more accurate and explainable result is provided for analysis of the compounds in a complex sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of environmental analytical chemistry and data mining technology, specifically to a method for analyzing the transformation of pollutants in aquatic environments based on GC-HRMS data. Background Technology

[0002] Aquatic pollutants include parent organic pollutants and their transformation products, which often undergo complex migration, degradation, and transformation processes in water bodies. Current methods for monitoring aquatic pollutants mainly rely on traditional target analysis or non-target analysis, which have the following shortcomings: difficulty in systematically revealing the migration pathways and transformation patterns of pollutants; limited ability to identify unknown degradation products and intermediates; complex data processing; and a lack of efficient visualization methods for pollutant trajectory analysis. Therefore, an efficient method is needed that can systematically track the migration and transformation trajectories of pollutants in the aquatic environment while simultaneously revealing unknown products and metabolic networks. Summary of the Invention

[0003] To address the aforementioned issues, this invention provides a method for constructing a structural similarity molecular network based on GC-HRMS non-targeted screening data.

[0004] A method for analyzing the transformation of pollutants in aquatic environments based on GC-HRMS data includes the following steps:

[0005] Acquire GC-HRMS mass spectrometry data of aquatic environment and extract mass spectrometry features from GC-HRMS mass spectrometry data;

[0006] Based on the mass spectrometry characteristics, compounds in the aquatic environment are identified, candidate compounds are obtained, and the candidate compounds are represented by standardized structural definitions.

[0007] The chemical structural similarity between candidate compounds is calculated based on the structural definition, and a structural similarity molecular network is constructed; wherein the nodes of the structural similarity molecular network are the structural definition and the edges are the chemical structural similarity.

[0008] The connections in the structurally similar molecular network are verified by applying chemical transformation rules, and environmental transformation pathways are screened out.

[0009] Explanation: The above method overcomes the limitations of existing mass spectrometry analysis techniques that rely on standards and preset rules. Leveraging the high-confidence qualitative capabilities of GC-HRMS, it achieves a fundamental transformation in molecular network nodes, moving from identifying unknown compounds to identifying known compounds. This allows the technique to move beyond fuzzy correlations based on mass spectrum similarity, directly calculating the chemical structural similarity between compounds to construct the network. This reduces false positives and improves the reliability of compound associations within the molecular network. Furthermore, based on chemical principles, transformation rules are introduced for post-verification and filtering. This method not only systematically discovers highly reliable specific environmental transformation pathways but also provides objective and comparable macroscopic evaluation indicators for the effectiveness of different environmental conditions or treatment processes by quantifying the overall network topology. Ultimately, it forms a complete, closed-loop analytical system from revealing microscopic mechanisms to assessing macroscopic activity, effectively avoiding the errors and uncertainties that may occur in traditional methods. It has strong versatility in fields such as environmental science and metabolic research.

[0010] Furthermore, the extraction of mass spectrometry features from GC-HRMS mass spectrometry data includes: extracting mass spectrometry features from GC-HRMS mass spectrometry data of the aquatic environment using an automatic deconvolution algorithm.

[0011] Note: An automatic deconvolution algorithm is used to extract mass spectrometry features from GC-HRMS mass spectrometry data of aquatic environment. It can efficiently and accurately process complex mass spectrometry data, automatically identify and separate the unique mass spectrometry features of each component, and avoid subjective errors and omissions that may be caused by manual extraction.

[0012] Further, the step of qualitatively identifying the compound based on the mass spectrometry characteristics to obtain candidate compounds, and representing the candidate compounds using standardized structural definitions, includes:

[0013] The mass spectrometry features are matched with a standard mass spectrometry library;

[0014] The matched candidate compounds are structurally converted to obtain standardized structural definitions.

[0015] For candidate compounds that fail to match, the elemental composition is deduced based on the mass number, and possible structural isomers are generated. The structural isomers are then converted into standardized structural definitions.

[0016] Note: The above method can quickly and accurately identify potential compounds by leveraging abundant existing mass spectrometry information, improving qualitative efficiency. Successfully matched candidate compound structures are converted into standardized structural definitions, ensuring uniformity and standardization of structural representation, facilitating subsequent unified processing and analysis. For unmatched candidate compounds, elemental composition is derived based on precise mass numbers, generating structural isomers, which are then converted into standardized structural definitions. This maximizes the value of the data, ensuring no potentially important compound information is overlooked, and comprehensively and systematically completing the compound qualitative and structural standardization work.

[0017] Furthermore, the standardized structural definition is in SMILES, InChI, or Mol format.

[0018] Note: The above formats are all widely recognized and universally accepted structural representation methods in the field of chemistry, which can ensure the standardization and normalization of compound structural information.

[0019] Furthermore, the successfully matched candidate compounds are assigned a high confidence level.

[0020] Note: Assigning the highest confidence level allows for the rapid and clear identification of these highly reliable compounds. This helps researchers prioritize these high-confidence substances in subsequent analyses, improving research efficiency.

[0021] Furthermore, the calculation of chemical structural similarity between candidate compounds based on the structural definition and the construction of a structural similarity molecular network include:

[0022] Extract the molecular fingerprint from the standardized structural definition;

[0023] The chemical structural similarity between each pair of candidate compounds is calculated based on the molecular fingerprint; the chemical structural similarity of the edges in the structural similarity molecular network is greater than the similarity threshold.

[0024] Note: The above method extracts molecular fingerprints using standardized structural definitions, transforming complex compound structures into quantifiable and comparable data. Based on these molecular fingerprints, the chemical structural similarity between candidate compounds is calculated, accurately and objectively measuring the degree of structural correlation between different compounds and avoiding errors from subjective judgment. Furthermore, setting a similarity threshold for edges in the structural similarity molecular network effectively filters out compounds with truly close structural relationships, constructing a high-quality molecular network with practical research value.

[0025] Furthermore, it also includes: calculating the mass spectrometry feature similarity between every two candidate compounds;

[0026] When constructing the structural similarity molecular network, mass spectrometry feature similarity is simultaneously introduced as an auxiliary connection edge, which together with the chemical structure similarity edge forms a multidimensional association network.

[0027] Explanation: The above method calculates the mass spectrometry feature similarity between candidate compounds and introduces it as an auxiliary connection edge when constructing the structural similarity molecular network. Together with the chemical structure similarity edge, it constructs a multidimensional association network, which can more comprehensively and accurately characterize the complex relationship between candidate compounds. It takes into account both the structural essence and the actual detection characteristics, which helps to screen out truly relevant compounds more accurately and avoid omissions or misjudgments that may be caused by a single similarity judgment.

[0028] Furthermore, the application of chemical transformation rules to verify the connections in the structurally similar molecular network and screen out environmental transformation pathways includes:

[0029] Chemical transformation rules are encoded into computer-executable reaction patterns; wherein the computer-executable reaction patterns are one of SMARTS, SMIRKS, and reaction SMILES, or implemented through a programmatic interface of a cheminformatics toolkit.

[0030] The standardized structural definitions corresponding to the nodes are substituted into computer-executable reaction modes for matching, and environmental transformation paths that conform to the rules of chemical transformation are selected.

[0031] Note: The above method encodes chemical transformation rules into computer-executable reaction patterns and provides a variety of flexible implementation methods. Whether using mature standard formats such as SMARTS, SMIRKS, and reaction SMILES, or leveraging the programmatic interface of cheminformatics toolkits, it greatly improves the convenience and versatility of rule application and can efficiently adapt to different research scenarios and needs.

[0032] Furthermore, when applying chemical transformation rules, environmental constraints are introduced for secondary screening; the environmental constraints include one or more of the following: redox potential, pH value, and light conditions of the aquatic environment.

[0033] Note: The above-mentioned secondary screening method can eliminate transformation pathways that are feasible under theoretical chemical transformation rules but cannot occur in actual aquatic environments due to the lack of corresponding conditions. This allows for the selection of highly reliable transformation pathways that truly conform to the actual conditions of aquatic environments, making the research results closer to the real environment.

[0034] Further, the topological parameters of the structurally similar molecular network are calculated, and the topological parameters include at least the average degree and weighted average degree of the structurally similar molecular network;

[0035] By comparing the topological parameters of molecular networks with structural similarity in samples from different aquatic environments, the transformation activity and structural association strength of compounds are quantitatively assessed; wherein, the transformation activity and structural association strength of compounds are used to characterize the pollution activity and risk of compounds in the aquatic environment.

[0036] Note: The above method, by calculating topological parameters such as the average degree and weighted average degree of structurally similar molecular networks, can quantify the tightness of the connections between nodes and the efficiency of information transmission in molecular networks from the perspective of the overall structural characteristics of the network, providing key indicators for analyzing the interactions between compounds.

[0037] The beneficial effects of this invention are:

[0038] This invention combines GC-HRMS non-targeted screening data with molecular structure fingerprint similarity calculation to achieve high-confidence molecular network construction based on structural information. Compared with traditional methods that rely on spectral similarity or prior rules, this invention can systematically reveal the structural family characteristics of compounds in complex samples without the need for standards or pre-set reaction rules. By introducing a confidence screening mechanism, false positives are reduced and the reliability of compound associations in the molecular network is improved. It can assist in the family annotation of unknown compounds and the identification of potential transformation relationships, thereby providing more accurate and interpretable results for the analysis of compounds in complex samples. The method of this invention has good versatility and scalability, and is applicable to different types of complex matrix samples, with broad application prospects in environmental chemistry, natural product research, and the analysis of new pollutants. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating Embodiment 1 of the present invention;

[0040] Figure 2 This is an example diagram of the structural similarity molecular network diagram and potential transformation relationship of the six wastewater treatment processes in the embodiments of the present invention. Detailed Implementation

[0041] To further illustrate the methods and effects of this invention, the technical solution of this invention will be clearly and completely described below in conjunction with experiments.

[0042] Organic pollutants in the aquatic environment include not only the original "parent" pollutants but also their "offspring" transformation products and intermediates produced through processes such as light and microbial activity. This is a dynamic and complex system. Traditional targeted analysis can only detect a few known, pre-defined pollutants, and is completely powerless against unknown transformation products. Non-targeted analysis, while capable of detecting thousands of compounds indiscriminately in a sample, generates massive amounts of data that are difficult to interpret. Existing methods (such as observing concentration changes and pre-defined reaction rules) that infer related transformation reactions are prone to guessing components and have a high false positive rate, much like blindly connecting scattered points without a map, easily leading to errors.

[0043] Molecular network technology is a visualization analysis method based on compound similarity relationships. It can group compounds with similar structures or spectral features into the same network cluster, thereby revealing potential correlations between compounds in a sample. This method can not only visually display the chemical composition distribution of complex samples, but also assist in the annotation of unknown compounds and support the comparison of differences between samples from different sources. Therefore, it has been widely used in fields such as natural product research, metabolomics, and environmental chemical analysis.

[0044] GC-HRMS-based non-targeted screening technology enables high-throughput detection and identification of organic compounds in complex samples, allowing for qualitative analysis even in the absence of reference standards. However, identifying potential compound transformation relationships from massive non-targeted screening data remains a significant technical challenge. Existing methods often rely on concentration trends or pre-defined reaction rules to infer compound transformation pathways, but these methods frequently suffer from high false positive rates and strong reliance on prior assumptions, making it difficult to accurately reflect the true molecular dynamics in complex systems.

[0045] Combining GC-HRMS with molecular network technology has become a significant trend in the in-depth mining of mass spectrometry data, enabling systematic and visual analysis of complex, non-targeted data. Molecular networks based on molecular structural similarity can directly utilize structural information to establish connections between compounds, aiding in the annotation of unknown compounds and revealing potential transformation pathways and structural group characteristics. This method can systematically reveal the relationships between compounds at the structural level without requiring pre-defined reaction rules, demonstrating strong versatility and scalability.

[0046] Based on the background technology and the above content, the existing LC-HRMS molecular network technology addresses the problems in current water pollution transformation research. LC-HRMS molecular network technology is a visualization data analysis strategy based on liquid chromatography-mass spectrometry (LC-MS) secondary mass spectrometry (MS / MS) data. Its core principle is to calculate the similarity between all spectra and connect nodes with high similarity to form a visualized molecular network diagram. In this network, each node represents a mass spectrometry signal whose chemical identity is not yet clearly defined, and the connections between nodes only indicate the degree of similarity between their original mass spectra.

[0047] However, this technology has significant limitations. First, the specific chemical structures corresponding to network nodes are unknown at the time of network construction, heavily relying on subsequent matching with standard spectral libraries for annotation, resulting in a large number of unknown nodes remaining unidentified. Second, connections established solely based on spectral similarity are unreliable; they cannot directly reflect the true chemical structural relationships between compounds, often generating numerous chemically illogical or impossible spurious transformation pathways, leading to a high false positive rate. Finally, this method struggles to effectively distinguish structural isomers and has limited ability to discover entirely new compounds not present in the pre-defined spectral library.

[0048] To overcome the problems of high false positive rates, strong reliance on prior assumptions, and insufficient structural association identification in existing methods for analyzing complex non-targeted screening data, this invention utilizes GC-HRMS non-targeted screening data and combines molecular structure fingerprint similarity calculation with a confidence screening mechanism to construct a molecular network. This improves the reliability and interpretability of the network, providing a more confident technical means for the analysis of compounds in complex samples and the identification of their structural associations. The specific solution is as follows:

[0049] like Figure 1 As shown, a method for analyzing the transformation of pollutants in aquatic environments based on GC-HRMS data includes the following steps:

[0050] S101. Obtain GC-HRMS mass spectrometry data of the aquatic environment and extract the mass spectrometry features from the GC-HRMS mass spectrometry data.

[0051] The extraction of mass spectrometry features of each component includes: extracting mass spectrometry features from GC-HRMS mass spectrometry data of the aquatic environment using an automatic deconvolution algorithm.

[0052] The GC-HRMS mass spectrometry data of the aquatic environment are obtained by monitoring the aquatic environment, and the content includes information such as peak area and retention time. Unlike LC-HRMS, the GC-HRMS data used in this invention is not parsed in common formats such as mzML. The specific content of obtaining the mass spectrum includes the following steps 1 to 3:

[0053] Step 1: Sample Collection; Water samples from different treatment units of a wastewater treatment plant were selected as the analysis objects. The influent of the plant mainly consists of approximately 80% domestic sewage and 20% industrial wastewater. Water samples were collected from the wastewater treatment plant influent, A2 / O reaction tank, secondary sedimentation tank, magnetic coagulation reactor, and sodium hypochlorite disinfection tank.

[0054] Step 2: Sample Pretreatment; Sample pretreatment employs liquid-liquid extraction to improve the extraction efficiency of organic compounds and ensure compatibility with subsequent GC×GC-QTOFMS analysis. The specific steps are as follows: Take 200 mL of homogenized water sample into a separatory funnel, add 10 g of sodium chloride to promote phase separation; then add 60 mL of dichloromethane, shake for 5-10 minutes, allow to stand for layering, and collect the lower organic phase. The above extraction process is carried out under neutral, acidic, and alkaline conditions, respectively. The three organic phases obtained are combined, concentrated to near dryness by rotary evaporation, and after adding an internal standard, diluted to 1 mL with n-hexane. Finally, filter through a 0.22 μm polytetrafluoroethylene (PTFE) membrane and transfer to a 2 mL amber sample vial for instrument analysis.

[0055] Step 3: GC-HRMS instrument analysis; Instrument analysis was performed using a GC×GC-QTOFMS system (Agilent 8890 gas chromatograph, Agilent 7250 quadrupole time-of-flight mass spectrometer, equipped with Snowfield Technology's SSM1800 solid-state thermal modulator), combined with non-targeted screening methods to obtain full scan data.

[0056] (1) Injection method: Split / splitless injection port was used, the temperature was set to 300℃, and the injection volume was 1μL in splitless injection mode. The carrier gas was high-purity helium, with a constant flow rate of 1.2mL / min.

[0057] (2) Chromatographic conditions: The one-dimensional column was a nonpolar DB-5 column (30m × 0.25mm × 0.25μm, Agilent Technologies), and the two-dimensional column was a moderately polar DB-17 column (1.3m × 0.18mm × 0.18μm, Agilent Technologies). The temperature program was as follows: initial temperature 50℃, increased to 300℃ at 5℃ / min, and held for 7min.

[0058] (3) Modulator parameters: The SV:C6-C40 modulation column of Xuejing Technology was used and set as follows: the main column inlet temperature was 50℃, which was increased to 300℃ at 5℃ / min and held for 7min; the initial temperature of the cold zone was 9℃, which was decreased to -51℃ at -50℃ / min and held for 18min, and then increased to 9℃ at 20℃ / min and held for 34min; the initial temperature of the secondary column outlet was 80℃, which was increased to 320℃ at 5℃ / min and held for 9min; the modulation period was set to 6s.

[0059] (4) Mass spectrometry conditions: an electron impact (EI) ion source was used, with an ion source temperature of 250℃, a quadrupole temperature of 150℃, and an ionization energy of 70eV; the scanning range of the time-of-flight analyzer was 50–800 Da, and the solvent delay time was 3 min.

[0060] S102. Based on the mass spectrometry characteristics, the compounds in the aquatic environment are qualitatively identified to obtain candidate compounds, and the candidate compounds are represented by standardized structural formulas, including:

[0061] The mass spectrometry features are matched with a standard mass spectrometry library;

[0062] The matched candidate compounds are structurally converted to obtain standardized structural definitions.

[0063] For candidate compounds that fail to match, the elemental composition is deduced based on the mass number, and possible structural isomers are generated. The structural isomers are then converted to obtain standardized structural definitions.

[0064] The candidate compounds that are successfully matched are assigned a high confidence level, and the candidate compounds that are not successfully matched are assigned a low confidence level; the standardized structural definition is in SMILES, InChI or Mol format.

[0065] This embodiment uses Canvas Panel V2.0 software from Snowscape Technology to perform qualitative analysis on the raw data acquired by GC×GC-QTOFMS. Compound identification is mainly completed through database comparison, including the NIST spectral library, the Agilent Technologies Precision Quality Pesticides and Environmental Pollutants Library, high-resolution databases, and relevant online databases. The data processing flow is as follows:

[0066] (1) Signal peak extraction and filtering: Extract effective signal peaks with a signal-to-noise ratio (S / N) greater than 3 from the raw data, remove interference peaks caused by column wear and solvent residue, and background peaks in blank samples; for strong peak tailing signals, merge them with adjacent peaks of the same type.

[0067] (2) Initial screening of candidate compounds: Based on the database comparison results, a matching score threshold of ≥700 is set (matching means successful matching, and non-matching means failure to match) to screen candidate compounds.

[0068] (3) Isotope verification: Under 70 eV electron bombardment ionization conditions, the presence of isotope ion clusters in candidate compounds is detected and compared with the consistency of standard compounds in the database to improve the accuracy of annotation.

[0069] (4) Retention index correction: A series of n-alkane standards were used, and the samples were injected under the same conditions as the samples to establish a GC×GC retention index template and calculate the experimental RI. Compounds with a deviation of less than 30 from the NIST retention index database were retained to further improve the reliability of the annotation.

[0070] (5) Results correction and integration: To reduce the impact of random errors, only compounds detected in at least two of the three parallel samples are retained; the final qualitative results are based on their average response values.

[0071] The above steps can effectively achieve high-confidence screening and identification of candidate compounds in complex samples, providing accurate input data for subsequent construction of molecular networks based on molecular structure similarity.

[0072] After obtaining the structural information of candidate compounds, they are first uniformly converted into standardized SMILES representations, and then molecular fingerprints are generated to encode the molecular skeleton and functional group features.

[0073] S103. Calculate the chemical structural similarity between each candidate compound based on the structural definition formula, and construct a structural similarity molecular network; wherein, the nodes of the structural similarity molecular network are the structural definition formulas, and the edges are the chemical structural similarity.

[0074] The step of calculating the chemical structural similarity between candidate compounds based on the defined structural formula and constructing a structurally similar molecular network includes:

[0075] Molecular fingerprints are extracted from the standardized structural definition; wherein, the molecular fingerprint is an extended fingerprint, and the molecular fingerprint includes a bit vector characterizing the molecular structure and a continuous vector characterizing the physicochemical properties;

[0076] The chemical structural similarity between each pair of candidate compounds is calculated based on the molecular fingerprint; the chemical structural similarity of the edges in the structural similarity molecular network is greater than the similarity threshold.

[0077] Also includes:

[0078] Calculate the mass spectrometric feature similarity between every two candidate compounds;

[0079] When constructing the structural similarity molecular network, mass spectrometry feature similarity is simultaneously introduced as an auxiliary connection edge, which together with the chemical structure similarity edge forms a multidimensional association network.

[0080] For example, Tanimoto similarity between compounds is calculated based on molecular fingerprints, and an initial structural similarity molecular network is constructed using compound pairs with a chemical structural similarity greater than 0.5 (i.e., the aforementioned similarity threshold is 0.5) as connecting edges. Here, nodes represent compounds, and edge weights represent structural similarity. Based on this initial network, a supervised screening strategy is introduced to eliminate compound pairs with high similarity but significant differences in their core structures, retaining only candidate pairs with higher structural consistency, thus forming a high-confidence structural similarity molecular network.

[0081] Average degree reflects the number of connections between each compound node and other nodes in a network, thus characterizing the degree of structural association between compounds. Weighted average degree further considers the weight of similarity and can reflect the strength of potential reaction connections. For example, when comparing different water treatment processes, if the network average degree obtained under a certain treatment process condition is significantly higher, it indicates that the transformation relationship between compounds under that condition is more active, meaning that potential reactions are more frequent; conversely, it indicates that the reaction is weaker and the products are dispersed.

[0082] In summary, existing technologies directly involve "peak detection and compound identification" plus calculating cosine similarity using MS / MS spectra. This means "similar spectra suggest similar structures," establishing a "mass spectrometry behavior correlation." Our proposed solution innovates in this step by embedding a crucial "chemical intelligent decoding layer." It utilizes GC-HRMS high-resolution mass spectrometry data to derive molecular formulas and achieve high-confidence identification, restoring the analyte from "mass spectrometry peaks" to "chemical structures," providing a true chemical identity for all subsequent analyses.

[0083] S104. Apply chemical transformation rules to verify the structurally similar molecular network, screen and output environmental transformation pathways;

[0084] The application of chemical transformation rules verifies the connections in the structurally similar molecular network, filters and outputs high-confidence environmental transformation pathways, including:

[0085] The chemical transformation rules (CTS chemical transformation rules) are encoded into computer-executable reaction modes; wherein the computer-executable reaction modes are one of SMARTS, SMIRKS, and reaction SMILES or implemented through a programmatic interface of a cheminformatics toolkit.

[0086] The standardized structural definitions corresponding to the nodes are substituted into computer-executable reaction modes for matching, and environmental transformation paths that conform to the rules of chemical transformation are selected.

[0087] When applying chemical transformation rules, environmental constraints are introduced for secondary screening; the environmental constraints include one or more of the following: redox potential, pH value, and light conditions of the aquatic environment.

[0088] For example, to further improve the reliability of transformation relationship identification, this method combines CTS chemical transformation rules to verify the chemical rationality of candidate compound pairs, and uses actual environmental conditions (such as reduction, oxidation, sterilization, etc.) for constraint screening to eliminate links that are not chemically feasible. The final network can not only reveal the group structure between compounds, but also identify potential structural transformation pathways.

[0089] This high-confidence structural similarity molecular network enables systematic analysis and screening of potential transformation relationships of compounds in complex samples without the need for standards or pre-defined reaction rules. The results can be visualized and analyzed using software such as Origin and Gephi, supporting environmental chemistry research and related applications.

[0090] S105. Calculate the topological parameters of the structurally similar molecular network, wherein the topological parameters include at least the average degree and weighted average degree of the structurally similar molecular network;

[0091] By comparing the topological parameters of molecular networks with structural similarity in samples from different aquatic environments, the transformation activity and structural association strength of compounds are quantitatively assessed; wherein, the transformation activity and structural association strength of compounds are used to characterize the pollution activity and risk of compounds in the aquatic environment.

[0092] like Figure 2 As shown, the high-confidence structural similarity molecular network constructed based on the above method can be further used to identify potential compound transformation relationships in complex samples. Based on candidate compound pairs, functional group change characteristics and structural evolution directions are introduced to compare and screen the pairing relationships between removed compounds and newly generated compounds.

[0093] It should be understood that existing literature describing "molecular networks for transformational studies" is based on spectral similarity and annotation propagation, resulting in a lack of verifiability. This invention, for the first time, proposes introducing a "chemically intelligent decoding layer" after network construction, using CTS chemical transformation rules plus environmental constraints for verification. This ensures that the obtained transformation relationships are not merely speculative inferences from spectra, but rather chemically plausible, environmentally feasible, and interpretable structural transformation pathways. Therefore, this invention solves the core problem of existing technologies "lacking high confidence and interpretability," demonstrating a level of inventiveness far exceeding simple network visualization. This approach can identify candidate transformation pairs with high structural similarity and chemical plausibility, and reveal possible transformation modes based on their functional group characteristics. Results show that in a reducing environment, candidate relationships such as acid reductive decarboxylation, ester hydrolysis, halogen dehalogenation, and carbon chain shortening can be screened; under anaerobic conditions, relationships such as carbon chain cleavage, redox transformation, and functional group hydrolysis and removal reactions can be identified; and under aerobic conditions, the transformation of alcohols to carbonyl compounds, the further oxidation of carbonyl groups to carboxylic acids, and the hydrolysis and deamination reactions of nitrogen-containing compounds can be identified. These results demonstrate that this method can reflect the evolutionary trends of compounds under different environmental conditions at the structural level and effectively identify potential transformation pathways related to changes in functional groups.

[0094] In summary, this method can not only identify high-confidence compound transformation relationships without the need for reference standards and pre-defined reaction rules, but also screen feasible transformation pathways by combining environmental constraints and chemical reaction mechanisms. It is suitable for molecular evolution analysis of complex samples and research on environmental process mechanisms.

[0095] Specifically, compared to existing LC-HRMS molecular networks that rely on MS / MS spectrum similarity, where node structures are usually non-qualitative results and cannot directly reflect chemical structures, this invention firstly identifies each node as a candidate compound with high confidence, rather than a spectral peak or predicted node; and secondly, it is based on structural fingerprint similarity, rather than just spectral similarity.

[0096] Secondly, this invention introduces a screening strategy that combines environmental constraints with structural evolution direction to ensure that potential transformation pathways are chemically rational.

[0097] The method of this invention does not rely on traditional spectral annotation or database-based inference of product structures. Instead, it directly identifies products and their transformation relationships through a high-confidence structure-guided analysis strategy, thus avoiding the limitations of conventional methods. Existing LC-HRMS molecular network technologies typically require standards (for library matching) or pre-defined reaction rules (for transformation interpretation). In contrast, this invention does not rely on standards or predetermined reaction rules. Instead, it directly constructs a compound structure-level network through high-confidence compound qualitative analysis, structural encoding, and fingerprint similarity calculation from GC-HRMS data.

Claims

1. A method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data, characterized in that, The method comprises the following steps: obtaining GC-HRMS mass spectrum data of the water environment, and extracting mass spectrum characteristics in the GC-HRMS mass spectrum data; qualifying compounds in the water environment according to the mass spectrum characteristics, obtaining candidate compounds, and representing the candidate compounds by standard structure definition formulae; calculating chemical structure similarity between the candidate compounds based on the structure definition formulae, and constructing a structure similarity molecular network; wherein nodes of the structure similarity molecular network are the structure definition formulae, and edges are the chemical structure similarity; checking the connection in the structure similarity molecular network by applying chemical transformation rules, and screening environmental transformation paths.

2. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 1, characterized in that, The method for extracting mass spectrum characteristics in the GC-HRMS mass spectrum data is to extract mass spectrum characteristics in the GC-HRMS mass spectrum data of the water environment by an automatic deconvolution algorithm.

3. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 1, wherein, The method for qualifying compounds in the water environment according to the mass spectrum characteristics, obtaining candidate compounds, and representing the candidate compounds by standard structure definition formulae comprises the following steps: matching the mass spectrum characteristics with a standard mass spectrum library; structure decoding of the candidate compounds that are successfully matched to obtain standard structure definition formulae; for the candidate compounds that are not successfully matched, deducing element composition based on mass number, generating possible structural isomers, and structure decoding of the structural isomers to obtain standard structure definition formulae.

4. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 3, characterized in that, The standard structure definition formulae are SMILES, InChI or Mol format.

5. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 1, wherein, The method for calculating chemical structure similarity between the candidate compounds based on the structure definition formulae, and constructing a structure similarity molecular network comprises the following steps: extracting molecular fingerprints from the standard structure definition formulae; calculating chemical structure similarity between every two candidate compounds based on the molecular fingerprints; the chemical structure similarity of edges in the structure similarity molecular network is greater than a similarity threshold.

6. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 5, wherein, The method further comprises the following steps: calculating mass spectrum characteristic similarity between every two candidate compounds; when constructing the structure similarity molecular network, synchronously introducing mass spectrum characteristic similarity as auxiliary connection edges to form a multi-dimensional correlation network together with the chemical structure similarity edges.

7. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 1, wherein, The method for checking the connection in the structure similarity molecular network by applying chemical transformation rules, and screening environmental transformation paths comprises the following steps: encoding chemical transformation rules into computer-executable reaction patterns; wherein the computer-executable reaction patterns are one of SMARTS, SMIRKS, reaction SMILES or realized through a programmatic interface of a chemical informatics toolkit; substituting the standard structure definition formulae corresponding to the nodes into the computer-executable reaction patterns to perform matching, and screening environmental transformation paths that conform to the chemical transformation rules.

8. The method for resolving transformation of environmental pollutants in water bodies based on GC-HRMS data according to claim 1, wherein, When applying the chemical transformation rules, environmental constraint conditions are introduced for secondary screening; the environmental constraint conditions include one or more of redox potential, pH value and light conditions of the water environment.

9. The method for analyzing water environment pollutant transformation based on GC-HRMS data according to claim 1, wherein topological parameters of the structure-similarity molecular network, the topological parameters at least comprising an average degree and a weighted average degree of the structure-similarity molecular network; comparing the topological parameters of the structure-similarity molecular networks of different water environmental samples, and quantitatively evaluating compound transformation activity and structure correlation strength; wherein the compound transformation activity and the structure correlation strength are used to characterize the compound pollution activity and risk in the water environment.