Automatic identification method for new pollutants and structural analogues thereof in environment complex sample
By constructing a molecular network through co-occurrence analysis and iterative annotation, and combining it with a database search strategy, the bottleneck problem of new pollutant identification in traditional methods is solved, and efficient identification and classification of new pollutants and their structural analogs in complex environmental samples are achieved.
Patent Information
- Application Number
- CN202511174611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to effectively identify and classify novel pollutants and their structural analogs in the environment, especially in complex environmental samples where identification bottlenecks are significant. Furthermore, traditional mass spectrometry annotation methods fail to effectively utilize covariation relationships between samples.
A molecular network was constructed using a method based on co-occurrence analysis and iterative annotation. Combined with a database search strategy, highly correlated substances were screened using the Pearson correlation coefficient, structural similarity was evaluated using the Tanimoto coefficient, and the spectral library was gradually expanded through iterative annotation to achieve automatic identification and classification of unknown pollutants.
It enables rapid identification and structural annotation of new pollutants and their structural analogs in non-target screening scenarios, improving the accuracy and efficiency of identifying unknown pollutants in complex environmental samples.
Smart Images

Figure CN120913655A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of environmental pollutant non-target screening, and in particular to an automatic identification method for new pollutants and their structural analogues in complex environmental samples based on co-occurrence analysis and iterative annotation, which is suitable for high-throughput screening and structural identification of new chemical pollutants. BACKGROUND
[0002] With the gradual elimination of traditional persistent pollutants and the widespread use of substitutes, new pollutants with structural diversity continue to enter the environment. These substitute compounds are usually structural analogues or transformation products, but due to the lack of toxicological data and standard spectral information, they bring great challenges to environmental monitoring and risk assessment. The current mainstream screening methods mainly rely on pre-set spectral library or high MS / MS similarity screening rules, which are difficult to effectively cover pollutants with novel structures or atypical spectral characteristics.
[0003] In recent years, high-resolution mass spectrometry (HRMS) has been widely used in environmental non-target screening, with the advantages of accurate mass measurement and high sensitivity, which can obtain tens of thousands of unknown characteristic peaks in complex samples. However, due to the lack of supporting structure analysis algorithm, these high-dimensional data still face the "recognition bottleneck". In addition, with the accumulation of large sample environmental monitoring data, correlation analysis methods based on the abundance co-variation relationship between samples have gradually emerged, which can assist in identifying compounds with environmental behavior correlation and improve the contextual accuracy of unknown pollutant identification. However, this kind of information has not been effectively utilized in traditional mass spectrometry annotation methods.
[0004] Molecular networking is a graph theory method based on the similarity connection of spectral characteristics, which has been widely used in pollutant annotation in recent years. However, traditional molecular networking relies on high similarity calculation of MS2 spectra, which is difficult to capture compounds with similar structures but significant differences in fragmentation patterns, especially in complex environmental samples. Therefore, there is an urgent need for a more flexible, data-driven annotation strategy to achieve efficient identification and classification of unknown pollutants. SUMMARY
[0005] The present application proposes an automatic identification method for new pollutants and their structural analogues in complex environmental samples based on co-occurrence analysis and iterative annotation, which integrates the co-occurrence relationship between samples and spectral similarity, constructs a reasoning-type molecular network with contextual information, and combines database search strategy to realize the gradual expansion and automatic annotation of structural analogues without relying on complete spectral library.
[0006] To achieve the above purpose, the present application provides an automatic identification method for new pollutants and their structural analogues in complex environmental samples, which comprises the following steps:
[0007] S1: inter-substance co-occurrence analysis: preprocessing the mass spectrum characteristic peak data of multiple environmental samples, calculating the co-occurrence relationship between the characteristic peaks by using the Pearson correlation coefficient, and retaining the substance peak intensity with a p value of the Pearson correlation coefficient less than a threshold value;
[0008] S2: high correlation substance screening: screening the mass spectrum characteristic peak data with high correlation relationship with the seed substance as the target annotated substance;
[0009] S3: constructing a molecular network: taking the above significantly correlated mass spectrum characteristic peak data as the basis, retaining the edges with high correlation scores, and removing the intrasource cracking substances;
[0010] S4: based on the substance quality of the to-be-annotated substance, searching the Pubchem database to obtain multiple candidate compounds for each to-be-annotated substance under the premise of meeting the set mass deviation threshold, and screening these candidate compounds based on the element composition of the molecular formula;
[0011] S5: using the Tanimoto coefficient to evaluate the structural similarity between each candidate compound and the seed substance, performing simple annotation according to the similarity score or advanced annotation according to the fragment matching result, and obtaining the final annotation result of the target substance;
[0012] S6: taking the substance with the annotation result as the seed substance in the next round of annotation, performing iterative annotation until no new annotation result is generated.
[0013] Further improvement of the application is that the preprocessing process of the mass spectrum characteristic peak data of multiple environmental samples in step S1 comprises:
[0014] Subtracting program blank interference: retaining the substance peak intensity data of the compound detected in at least three environmental samples, and when the compound concentration in the program blank exceeds the average value plus three times the standard deviation, subtracting the program blank value from the sample peak area;
[0015] Outlier processing: identifying outliers by using the interquartile range method, evaluating the significant influence of outliers on the Pearson correlation coefficient by Fisher Z transformation, if there is a significant difference, using the conservative correlation coefficient after removing outliers, and if not, using the original Pearson correlation coefficient.
[0016] Further improvement of the application is that the expression of the Pearson correlation coefficient in step S1 is:
[0017]
[0018] Wherein: x i ,y irespectively are the concentration data of substance X in sample i, the concentration data of substance Y in sample i; respectively are the average concentration of substance X in all samples, the average concentration of substance Y in all samples; n is the number of samples.
[0019] A further improvement of the present application is that the seed substance refers to a known compound which is well matched with the existing spectral library and has high confidence annotation information.
[0020] A further improvement of the present application is that in step S3, if the retention time deviation between two substances is less than 0.01 and has a correlation coefficient value of 0.90 or above, it is considered that the substance with smaller molecular weight is an in-source fragmentation substance of the substance with larger molecular weight.
[0021] A further improvement of the present application is that in step S4, the calculation formula of the mass deviation threshold is:
[0022]
[0023] A further improvement of the present application is that in step S5, the simple annotation is based on Morgan fingerprint and Tanimoto similarity, and the candidate compound with the highest similarity score is directly selected as the annotation result of the target substance; the advanced annotation is based on Tanimoto coefficient to screen Top-k candidate compounds with the most similar structure, use a fragmentation prediction tool to simulate the MS2 spectrum, match and score the predicted spectrum with the real MS2 spectrum of the target substance, and select the compound with the highest spectrum matching score as the annotation result of the target substance.
[0024] The beneficial effects of the present application include that the method of the present application can be widely applied to the automatic and rapid identification and structure annotation task of pollutants in non-target screening scenarios, and can be applied to the screening of new pollutants and their structural analogues in complex environmental samples. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 Flowchart of the automatic identification method of new pollutants and their structural analogues in complex environmental samples;
[0026] Figure 2 Schematic diagram of the co-occurrence analysis module between substances;
[0027] Figure 3 Working schematic diagram of the high correlation substance screening module;
[0028] Figure 4 Working schematic diagram of the molecular network editing module;
[0029] Figure 5 Working schematic diagram of the target annotation candidate structure acquisition module;
[0030] Figure 6 An iterative annotation work flow is shown. DETAILED DESCRIPTION
[0031] Other advantages and benefits of the present application will become apparent to those skilled in the art upon consideration of the disclosure herein. The application can be implemented or applied in other different embodiments and its details can be modified in various obvious ways without departing from the spirit of the application. It is to be understood that the following examples and features thereof can be combined with each other, if not incompatible.
[0032] It is to be noted that the drawings provided in the following examples only schematically illustrate the basic concept of the present application, and thus only show the components related to the present application in the drawings, but are not drawn according to the number, shape and size of the components in actual implementation. The shapes, number and proportions of the components in actual implementation can be changed arbitrarily, and the layout pattern of the components can be more complex.
[0033] Some exemplary embodiments of the present application are described for illustrative purposes, and it is to be understood that the present application can be implemented in other ways not specifically shown in the drawings.
[0034] As shown in Figure 1 The present embodiment provides a method for automatically identifying new pollutants and their structural analogues in complex environmental samples, as shown in Figure 1 The present embodiment includes an inter-substance co-occurrence analysis module, a high-correlation substance screening module, a molecular network editing module, a target annotation candidate structure acquisition module, and an iterative annotation module, which are combined into an automated algorithm based on Python. Among them:
[0035] As shown in Figure 2 The inter-substance co-occurrence analysis module uses Pearson correlation coefficient to evaluate the linear correlation between variables, and to efficiently calculate the correlation coefficients between substances in multiple samples. Multiple quality control measures are implemented during the correlation analysis to ensure the results are robust and reliable.
[0036] First, it is determined whether the substance is a blank pollution substance. When the concentration of a compound (substance) detected in a sample exceeds the average value of the concentration of the compound in the procedural blank sample plus three times the standard deviation, it is considered that the compound exists in the sample, and the peak area observed from the procedural blank sample is subtracted from the corresponding sample measurement value to further reduce background interference.
[0037] In addition, only the compounds detected in at least three independent samples were included in the correlation calculation. To further refine the analysis and address the influence of outliers, an outlier detection and evaluation procedure was added in this embodiment. In this embodiment, the interquartile range (IQR) method was applied to identify outliers (abnormal characteristic peaks) that deviate significantly from the distribution of measured intensities. Then the Pearson correlation coefficients before and after the removal of outliers were calculated, and the Fisher Z transformation was used to determine whether the removal of outliers significantly affected the calculated correlation. If a significant difference was observed, the lower Pearson correlation coefficient was retained to arrive at a more conservative estimate; if not, the original value was used. Finally, the characteristics that showed statistically significant correlation (p < 0.05) with the seed substance were selected for analysis in the subsequent module.
[0038] As shown in Figure 3 , the high correlation substance screening module screens substances with high correlation coefficients with the seed substance as candidate target annotation substances based on the correlation calculation results described above. For the first round of annotation, the seed substance is selected by selecting a substance from the sample that matches well in the spectral library as the starting seed substance.
[0039] As shown in Figure 4 , the molecular network editing module uses the above significantly correlated mass spectral characteristic peak data as the basis for molecular network editing, retaining edges with high correlation scores while removing intra-source cracking substances.
[0040] As shown in Figure 5 , the target annotation substance candidate structure acquisition module acquires multiple candidate structures (candidate compounds) for each target annotation substance based on the mass of the target annotation substance by searching the Pubchem database under the premise of meeting the set mass deviation threshold, and screens these candidate structures based on the elemental composition of the molecular formula. Finally, these candidate structures serve as the data basis for the subsequent iterative annotation module.
[0041] As shown in Figure 6 , the iterative annotation module uses the Tanimoto coefficient to evaluate the structural similarity between each candidate compound and its related seed substance, which is calculated based on the Morgan fingerprint. According to the similarity score (simple annotation) or fragment matching result (advanced annotation), the top-ranked candidate compound is selected as the final annotation result of the target annotation substance. Further, these substances with annotation results will serve as seed substances in the next round of annotation, so that the annotation is iteratively performed until no new annotation result is generated.
[0042] The Pearson correlation coefficient (r) is a statistical measure of the linear correlation between two variables. It ranges from -1 to 1, with the following interpretations: r = 1 indicates a perfect positive correlation, i.e., as one variable increases, the other also increases at a constant rate; r = -1 indicates a perfect negative correlation, i.e., as one variable increases, the other decreases at a constant rate; r = 0 indicates no linear correlation between the two variables. In the method of this embodiment, the Pearson correlation coefficient is used to analyze the co-occurrence relationship of peak areas of substances in samples. If two substances show similar trends in abundance changes in multiple samples (i.e., the Pearson correlation coefficient is high), it can be considered that they have similar sources, transformation pathways, or environmental behaviors, and thus may be structural analogs.
[0043] Let there be two substances X and Y, X = {x1, x2, x3, …, x n}, Y = {y1, y2, y3, …, y h}, where the data in {} are the peak area data, i.e., the concentration data, of the substances in different samples. The Pearson correlation coefficient is defined as follows. and are the means of variables X and Y, respectively; the numerator is the covariance, indicating the direction and strength of the common change of the two variables; the denominator is the product of the standard deviations, used for normalization, so that r ∈ [-1, 1].
[0044]
[0045] where: x i , y i are the concentration data of substance X in sample i and substance Y in sample i, respectively; are the mean concentration of substance X in all samples and the mean concentration of substance Y in all samples, respectively; n is the number of samples.
[0046] The interquartile range (IQR) method is a commonly used method for detecting outliers. The interquartile range (IQR) is defined as the difference between the third quartile (Q3) and the first quartile (Q1) of the data: IQR = Q3 - Q1. Here, Q1 (the first quartile) refers to the value at the 25th percentile of the data arranged in ascending order, and Q3 (the third quartile) refers to the value at the 75th percentile. The rules for using the IQR method to judge outliers are as follows: any data point outside this interval is considered an "outlier".
[0047] Low outlier lower limit: x < Q1 - 1.5 × IQR
[0048] High outlier upper limit: x > Q3 + 1.5 x IQR
[0049] Where: x is the concentration value of the substance in the sample;
[0050] The Fisher Z transformation is a statistical method for standardizing the correlation coefficient (r) to a normally distributed variable z, which facilitates the significance test or comparison of the difference between two correlation coefficients. The complete calculation formula for evaluating whether two correlation coefficients are significantly different is as follows. If P < 0.05, it is considered that the correlation coefficients before and after removing outliers have statistically significant difference.
[0051] Apply Fisher Z transformation:
[0052] Standard error of z value:
[0053] Construct the test statistic:
[0054] Calculate the two-tailed P value: p = 2(1-Φ(|Z stat |))
[0055] The seed substance refers to a known compound with high confidence annotation information that is well matched with the existing spectral library and serves as the starting point for iterative annotation in the unknown contaminant identification process. The target annotation substance refers to a characteristic node in the molecular network that is connected to the seed substance and has not been annotated but has potential structural similarity.
[0056] The molecular network is a graph analysis method based on MS / MS (tandem mass spectrometry) spectrum similarity, which is used for visualizing, clustering and annotating structurally similar compounds, and is widely used in non-target screening and metabolomics. The basic principle is that in high-resolution mass spectrometry analysis, different compounds produce characteristic fragment ion patterns in secondary mass spectrometry (MS / MS). If two compounds have similar structures, the fragment peaks they produce in MS / MS also tend to have similarities. The core idea is to measure the structural similarity between compounds using spectrum similarity (such as cosine similarity), and express this similarity as an edge in the graph structure. Each node in the network graph is a compound, and there is an edge between two nodes, representing the similarity of their spectra.
[0057] The in-source fragmentation products refer to the fragment ions generated by the fragmentation of the parent ions in the ion source before entering the tandem mass spectrometry analysis. They are not real compounds, but may be misidentified as independent substances in the mass spectrometry data. The judgment condition of the in-source fragmentation products is that if the retention time deviation between two substances is less than 0.01 and the correlation coefficient value is greater than 0.90, the substance with smaller molecular weight is considered to be the in-source fragmentation product of the substance with larger molecular weight.
[0058] The retention time (RT) and mass (Mass) of the substance are two key parameters for characterizing and identifying compounds. The retention time refers to the time required for a compound to be eluted from the chromatographic column and detected by the detector after being injected (usually in minutes). The mass generally refers to the mass-to-charge ratio (m / z) of the compound, i.e. the ratio of the mass of an ion to the number of charges.
[0059] The Pubchem database is an open access chemical information database maintained by the National Center for Biotechnology Information (NCBI) of the United States, and is one of the most comprehensive and widely used chemical databases in the world. It contains information on hundreds of millions of compounds and serves multiple fields such as drug discovery, toxicology, environmental science, and computational chemistry. The mass, structure, and molecular formula of the candidate substance can be obtained from the database.
[0060] The mass deviation threshold refers to the maximum error range allowed between the m / z value to be tested and the candidate physical theory m / z value in the database, usually expressed in ppm (parts per million), i.e. relative error of one millionth. The formula is as follows:
[0061]
[0062] The molecular formula (Molecular Formula) is a chemical formula that represents the types and number of atoms of each element in a compound, and is one of the most basic chemical representations.
[0063] The Tanimoto coefficient and Morgan fingerprint are tools for evaluating the structural similarity between compounds. The role of the Morgan fingerprint is to encode a molecular structure into a fixed-length binary vector (or integer vector), each bit indicating the presence or absence of a certain specific structural subfragment. The role of the Tanimoto coefficient is to evaluate the similarity between two molecular fingerprints, with a value range of [0, 1], and the higher the value, the more similar the structure. The calculation formula is as follows. Where A and B are the fingerprints (i.e. bit vectors) of two molecules; c = the number of bits that are 1 in common, a = the number of bits that are 1 in A, and b = the number of bits that are 1 in B. If two molecules are completely identical (fingerprints are the same), Tanimoto = 1; if there is no overlap at all, Tanimoto = 0.
[0064]
[0065] The simple annotation is a fast and automatic preliminary annotation method based on structural similarity score. For each target annotation substance, the Tanimoto similarity of its Morgan fingerprint with the seed substance is calculated in the candidate compound list, and the candidate structure with the highest score (Top-1) is directly selected as the annotation result of the target substance. This annotation method does not consider fragment information and only relies on structural fingerprint evaluation, and can be used for preliminary screening or large-scale preliminary annotation.
[0066] The advanced annotation considers both structural similarity and MS2 fragment matching information to improve the accuracy and reliability of the annotation. First, based on the Tanimoto coefficient, the Top-k (such as Top-5 or Top-10) most structurally similar candidate compounds are selected; for each candidate compound, a fragment prediction tool (such as: CFM-ID, MetFrag, FISh Scoring, etc.) is used to simulate its MS2 spectrum; the predicted spectrum is matched with the real MS2 spectrum of the target substance to score the spectrum matching; finally, the compound with the highest spectrum matching score (Top-1) is selected as the annotation of the target substance. This annotation method introduces MS2 hierarchical information, which can improve the annotation accuracy.
[0067] In one specific embodiment, the sample used by the embodiments of the present application is a surface sediment sample collected from a typical river basin water environment. The sample is freeze-dried and ground uniformly, and the extract is analyzed using a Vanquish UHPLC (Thermo Fisher) coupled with an Orbitrap Exploris 120 mass spectrometer. The H-ESI source is used, and the full scan (Full MS) + data-dependent acquisition (DDA) is performed in positive ion mode. The chromatographic column is an ACQUITY UPLC BEH C18 (2.1 mm x 100 mm, 1.7 μm), and the mobile phase is a water-acetonitrile system containing 0.1% formic acid, with a flow rate of 0.3 mL / min, and the total running time is 30 min. The MS1 acquisition range is 80-1000 m / z, the resolution is 60,000, the MS2 acquisition Top 4 ions are HCD energy 30%, and the resolution is 15,000. To improve the fragment coverage, a second round of DDA is performed on compounds that do not have high-quality MS2 data, and stepwise collision energy (30%, 60%, 100%) is used for acquisition.
[0068] All sample data are processed by Compound Discoverer 3.3 software to generate a complete data table containing m / z, retention time, peak intensity, and spectral information.
[0069] In the co-occurrence analysis stage, a peak area matrix is first constructed with characteristic peaks as variables and samples as observation values. To eliminate the interference of procedural blanks, features that only exist in procedural blanks are filtered out, and features that repeatedly appear in at least 3 environmental samples are retained for subsequent analysis. After logarithmic transformation and zero value processing of the peak area data, the Pearson correlation coefficient is used to evaluate the linear covariation relationship between each pair of features. To control background noise and outliers, the interquartile range method is introduced to identify observation values that deviate from the distribution, and the interval is Q1-1.5xIQR, ~Q3+1.5xIQR, and the r value is calculated for each pair of substances with and without removing outliers.
[0070] To evaluate the significant influence of outliers on co-occurrence evaluation, the Fisher r-to-z transformation method is introduced to convert two correlation coefficients r1, r2 into z1, z2 under normal distribution, then calculate the combined standard error SE and perform z-test. If a significant difference is observed, the lower correlation coefficient is retained to obtain a more conservative estimate; if not, the original value is used.
[0071] The seed substances as the starting point of iterative annotation are known compounds with high confidence annotation information that have good fragment matching scores (at least 70 points, with a full score of 100 points) with commercial spectral library mzCloud or self-built spectral library mzVault.
[0072] The features that appear in more than 20% of the samples and have a significant positive correlation (r>0.70, p<0.05) with at least one seed substance are screened in Compound Discoverer 3.3 software to generate a molecular network. After obtaining the molecular network file, based on the correlation calculation results, a co-occurrence enhanced molecular network is constructed. In this network, the nodes are substances, and the edges represent the significant co-occurrence relationship and fragment similarity between two substances. A source fragmentation filtering mechanism is introduced. If the retention time difference is <0.01 min and r>0.90, and the mass of the node is smaller, it is considered as a source fragmentation product, which is excluded from the network to avoid the spread of false positive annotations. At this time, the feature nodes connected to the seed substance in the molecular network and not yet annotated but with potential structural similarity are the target annotated substances in the first round of annotation.
[0073] For each substance to be annotated, search for candidate structures in the PubChem database within the set mass deviation threshold (5ppm) according to its m / z value. Exclude candidate substances with molecular formula element components outside the target range (CHONPSFClBrI), and only retain structures consistent with the target features.
[0074] Convert the candidate structures to Morgan fingerprints (radius 2, 1024 bits) using the RDKit toolkit, and calculate the structural similarity between them and the seed substance using the Tanimoto coefficient. Using the advanced annotation method, select the Top-10 candidate substances according to the similarity score, use the Fragment Prediction tool FISh Scoring in Compound Discoverer 3.3 software to predict the spectra of these candidate substances, match the predicted spectra with the real spectra of the target substance, and select the candidate compound with the best spectrum match as the final annotation result of the target substance.
[0075] The successfully annotated target substances will be used as seed substances in the next round of annotation, realizing the recursive expansion of the annotation path until no new annotated substances are produced or the maximum number of iterations is reached. The final output includes the co-occurrence analysis results and the substance annotation results of each round.
[0076] The above examples only exemplarily illustrate the principles and effects of the present application, and are not used to limit the present application. Any person skilled in the art can modify or change the above examples without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical idea disclosed by the present application should be covered by the claims of the present application.
Claims
1. A method for automatic identification of new pollutants and their structural analogues in environmentally complex samples, characterized by, The method comprises the following steps: S1: co-occurrence analysis between substances: pre-processing mass spectrum characteristic peak data of multiple environmental samples, calculating co-occurrence relationship between characteristic peaks by using Pearson correlation coefficient, and retaining substance peak intensity with p value of Pearson correlation coefficient less than a threshold value; S2: high correlation substance screening: screening mass spectrum characteristic peak data with high correlation relationship with the seed substance as the target annotated substance; S3: constructing a molecular network: taking the above significant correlation mass spectrum characteristic peak data as the basis, retaining edges with high correlation scores, and removing in-source cracking substances; S4: based on the mass of the substance to be annotated, a plurality of candidate compounds are obtained by searching the Pubchem database under the premise of meeting the set mass deviation threshold, and the candidate compounds are screened based on the element composition of the molecular formula; S5: using Tanimoto coefficient to evaluate the structural similarity between each candidate compound and the seed substance, and according to the similarity score, simple annotation or advanced annotation is performed according to the fragment matching result, and the final annotation result of the target substance is obtained; S6: taking the substance with the annotation result as the seed substance in the next round of annotation, and iteratively annotating until no new annotation result is generated.
2. The method according to claim 1, wherein, The pre-processing process of the mass spectrum characteristic peak data of the multiple environmental samples in step S1 comprises: Subtracting program blank interference: retaining the substance peak intensity data of the compounds detected in at least three environmental samples, and when the concentration of the compound in the program blank exceeds the average value plus three times the standard deviation, the program blank value is subtracted from the sample peak area; Outlier processing: using the interquartile range method to identify outliers, and using Fisher Z transformation to evaluate the significant influence of outliers on Pearson correlation coefficient, if there is a significant difference, the conservative correlation coefficient after removing outliers is used, and if not, the original Pearson correlation coefficient is used.
3. The method according to claim 1, wherein the method is characterized by, The expression of the Pearson correlation coefficient in step S1 is: where: x i y i are concentration data for substance X in sample i, concentration data for substance Y in sample i, respectively; are the mean concentration of substance X across all samples, the mean concentration of substance Y across all samples, respectively; n is the number of samples.
4. The method according to claim 1, wherein, The seed substance refers to a known compound with good matching with an existing spectrum library and high confidence annotation information.
5. The method according to claim 1, wherein the method is characterized by, In step S3, if the retention time deviation between two substances is less than 0.01 and the correlation coefficient value is above 0.90, it is considered that the substance with smaller molecular weight is the in-source cracking substance of the substance with larger molecular weight.
6. The method according to claim 1, wherein, In step S4, the calculation formula of the mass deviation threshold is:
7. The method according to claim 1, wherein, In step S5, the simple annotation is based on Morgan fingerprint and Tanimoto similarity, and the candidate compound with the highest similarity score is directly selected as the annotation result of the target substance; The advanced annotation is based on Tanimoto coefficient to select the top-k most similar candidate compounds in structure, using a fragment prediction tool to simulate the MS2 spectrum of the candidate compounds, matching and scoring the predicted spectrum with the real MS2 spectrum of the target substance, and selecting the compound with the highest spectrum matching score as the annotation result of the target substance.