A highly sensitive and high-throughput chemical substance annotation method and system

Through the combination of EISA technology and chemical exposure group database, high sensitivity and high throughput chemical substance detection are achieved, solving the limitations of low-concentration detection in the prior art, and improving detection efficiency and accuracy.

CN117110466BActive Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311031340.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2025-07-11
Estimated Expiration
2043-08-15

AI Technical Summary

Technical Problem

The prior art cannot effectively realize high sensitivity and high throughput chemical substance detection, especially under low concentration conditions, traditional mass spectrometry data analysis software cannot process EISA data, resulting in limitations in chemical substance annotation methods.

Method used

EISA technology was used to collect precursor ions and fragment ions, build a chemical exposure group database, set quality accuracy standards, perform precursor ion chromatogram correction and screening, and sort by calculating relevant characteristic scores and peak types to output the final screening results.

Benefits of technology

It significantly improves the number of chemical substances detected at low concentrations, improves detection sensitivity and flux, reduces false positive characteristics, and expands the scope of chemical substance identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117110466B_ABST
    Figure CN117110466B_ABST
Patent Text Reader

Abstract

The present invention discloses a highly sensitive and high-throughput chemical substance annotation method and system, which relates to the technical field of chemical substance identification, and includes collecting precursor ions and fragment ions of a sample to be detected by using the EISA technology; constructing a chemical exposome database, and sequentially extracting the precursor ion chromatogram of the sample to be detected as the target chemical substance; obtaining a target chromatogram after correction, screening all chromatographic peaks therein to obtain candidate chromatographic peaks and their peak types; traversing and matching each candidate chromatographic peak with the target chemical substance to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance; sorting each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and the relevant characteristic scores, and outputting the first several candidate chromatographic peaks in the sorting result as the final screening result of the chemical substances in the sample to be detected. The present invention can significantly increase the number of low-concentration chemical substances detected, and has the advantages of high sensitivity and high throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chemical substance identification, and more particularly, to a highly sensitive and high-throughput chemical substance annotation method and system. Background Art

[0002] Chemical substance annotation is an important part of exposomics research. The concept of exposure was originally used to explain the environmental drivers of health and disease. Now, it represents all exogenous and endogenous environmental exposures and related biological effects throughout the life cycle. The concentration ranges of xenobiochemicals in biological and environmental samples are very wide, reaching as low as ppb or nanomolar levels. Therefore, there is an urgent need for a highly sensitive chemical exposure localization technology with a wide chemical coverage to promote the development of exposure type research. Traditional chemical substance annotation requires obtaining the mass spectrometry (MS) data of precursor ions and the MS data of fragment ions separately through a tandem mass spectrometer. The precursor ion is also called the parent ion and is usually generated in the ion source. The generation of fragment ions usually requires the combination of an electrospray ionization (ESI) source and a tandem mass spectrometer. When the parent ion enters the collision cell, it reacts with inert gas molecules under the action of energy to generate fragment ions.

[0003] Targeted detection based on triple quadrupole mass spectrometry is generally considered the most sensitive technique in the monitoring of exposure biomarkers. However, this method can only monitor a limited number of chemicals, restricting its application in exposure research. The untargeted analysis method based on high resolution mass spectrum (HRMS) can detect a wide range of chemicals. Usually, the full scan acquisition method is adopted to obtain the m / z (mass-to-charge ratio) and retention time (RT) information of precursor ions, and then the data-dependent acquisition (DDA) mode or data-independent acquisition (DIA) mode is used to generate mass spectra (MS / MS). The data-dependent acquisition (DDA) mode is a main mode of data acquisition in tandem mass spectrometry. In the DDA mode, when the mass spectrometer performs an MS full scan, it will automatically perform MS / MS analysis on the list of precursor ions selected from the full scan spectrum. The DDA mode is usually used to collect fragment spectra of exposure features of interest. The selection of precursor ions depends on the ion intensity. Then, the mass spectral features of low-abundance exposures may never be selected to generate fragments in the collision cell. The advantage of the DDA mode is high selectivity, but the throughput is low and it is difficult to achieve batch acquisition of fragment ions of chemicals; the data-independent acquisition (DIA) mode is a mode of batch acquisition of fragment ions of chemicals in mass spectrometry analysis. In this mode, all ions within the selected m / z range will be fragmented and analyzed in the second stage of tandem mass spectrometry. This mode obtains tandem mass spectrometry data by fragmenting all ions entering the mass spectrometer at a given time or by sequentially separating and fragmenting the m / z range. Although the DIA mode can achieve batch acquisition of fragment ions of chemicals, its selectivity is low.

[0004] Informatics algorithms have certain limitations in the annotation of chemicals at a lower level. Currently, the data processing informatics algorithms used for autonomous chemical identification and annotation in exposure research mainly draw on algorithms in the field of metabolomics. However, the concentration of endogenous metabolites in biological samples is usually several orders of magnitude higher than that of exogenous chemicals in the same sample. Even in environmental samples, the concentration of chemical pollutants is low, usually around ppb. Early studies have shown that traditional metabolomics data preprocessing software such as XCMS will miss a large number of metabolic features, that is, low-abundance and / or poorly shaped features. Therefore, traditional peak extraction algorithms will lead to feature loss, affecting downstream substance identification.

[0005] Previously, the applicant proposed a highly sensitive analysis technique for chemical substances (EISA) based on the phenomenon of high-energy fragmentation in a mass spectrometry ion source. As a simple and effective method for generating in-source fragments in an electrospray ionization (ESI) source, EISA has changed the traditional strategy for generating fragments in tandem mass spectrometry. By optimizing the in-source fragmentation conditions, the in-source fragment generation using the full-scan mode of the EISA technique in HRMS can simulate the endogenous metabolites and peptides of the medium- and high-energy MS / MS fragment spectra through the collision-induced dissociation (CID) technique in the collision cell without affecting the intensity of the precursor ions, thus enabling the effective identification of low-concentration chemical substances. However, traditional mass spectrometry data analysis software such as XCMS cannot analyze EISA data. Therefore, there is an urgent need to propose a highly sensitive and high-throughput method for annotating chemical substances.

[0006] The prior art provides a method and application for predicting fragment ions of a compound, including: 1) determining the compound used for the experiment, the molecular formula of the parent ion of the compound, and the molecular formula of the fragment ion; 2) selecting stable isotopes according to the elemental composition of the parent ion and the fragment ion; 3) calculating the labeling situation of the parent ion and the fragment ion with the stable isotope according to the determined compound, the molecular formula of the parent ion and the fragment ion, and the number of stable isotope elements in the parent ion and the fragment ion; 4) calculating the number of all parent ions and all fragment ions labeled with the stable isotope; 5) calculating the exact mass numbers of the parent ion and the fragment ion in all the labeling situations according to the number of labeling situations obtained in step 4), and calculating the mass-to-charge ratio; 6) establishing a mass-to-charge ratio database of the fragment ions of the compound. This prior art needs to clarify the relationship between the parent ion and the fragment ion in the compound used for the experiment, while the EISA technique collects all the parent ions and fragment ions in one full scan, and the relationship between the parent ion and the fragment ion is not clear, so it is not applicable to EISA data. Summary of the Invention

[0007] In order to overcome the defect that the above prior art cannot accurately analyze EISA data, and thus cannot achieve high-sensitivity and high-throughput detection of chemical substances, the present invention provides a highly sensitive and high-throughput method and system for annotating chemical substances, which significantly increases the number of low-concentration chemical substances detected, with high sensitivity and high throughput.

[0008] To solve the above technical problems, the technical solution of the present invention is as follows:

[0009] The present invention provides a highly sensitive and high-throughput method for annotating chemical substances, including:

[0010] S1: Based on the EISA technique, collect the parent ions and fragment ions in the sample to be detected;

[0011] S2: Construct a chemical exposome database, including several chemical substances and their mass spectrometry data;

[0012] S3: Set the quality accuracy standard, and successively use the chemical substances in the chemical exposome database as target chemical substances, and correspondingly extract the parent ion chromatograms of the samples to be detected.

[0013] S4: Calibrate the parent ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes a number of chromatographic peaks.

[0014] S5: Screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types.

[0015] S6: For each candidate chromatographic peak, traverse and match the target chemical substance, and calculate the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance.

[0016] S7: Based on the peak types and relevant characteristic scores of the candidate chromatographic peaks, sort each candidate chromatographic peak corresponding to the target chemical substance to obtain a sorting result.

[0017] S8: Output the first several candidate chromatographic peaks in the sorting result as the final screening result of the chemical substances in the sample to be detected.

[0018] Preferably, in step S2, based on the mass spectrometry data of the chemical reference standard local database, open-source mass spectrometry database and public literature, construct a chemical exposome database; the chemical exposome database includes several chemical substances and their mass spectrometry data, and the mass spectrometry data of each chemical substance includes the chemical substance name, parent ion mass-to-charge ratio, fragment ion mass-to-charge ratio and intensity.

[0019] Preferably, in step S3, the specific method for setting the quality accuracy standard and successively using the chemical substances in the chemical exposome database as target chemical substances to correspondingly extract the parent ion chromatograms of the samples to be detected is as follows:

[0020] Successively use the chemical substances in the chemical exposome database as target chemical substances, and according to the parent ion mass-to-charge ratio of the target chemical substance, combined with the set quality accuracy standard, obtain the range of the parent ion mass-to-charge ratio to be extracted; in the sample to be detected, for the ions whose parent ion mass-to-charge ratio is within the range of the parent ion mass-to-charge ratio, perform the operation of extracting and constructing a chromatogram.

[0021] Preferably, in step S4, the specific method for calibrating the parent ion chromatogram of the sample to be detected to obtain a target chromatogram is as follows:

[0022] For the extracted chromatogram of the sample to be detected, delete the corresponding chromatogram parts before the start and after the end of the chromatographic elution, and use the remaining part of the chromatogram as the target chromatogram for subsequent analysis.

[0023] The chromatographic elution is set according to the actual situation. The chromatogram parts corresponding to before the elution starts and after the elution ends are deleted, and only the chromatogram part during the elution is retained, reducing the appearance of false positive features and making the detection result more accurate.

[0024] Preferably, the specific method of step S5 is as follows:

[0025] S5.1: Calculate the peak height and sawtooth index of all chromatographic peaks;

[0026] S5.2: Filter the chromatographic peaks with peak height lower than the preset peak height threshold, and regard the remaining chromatographic peaks as candidate chromatographic peaks;

[0027] S5.3: Use the existing mass spectrometry data analysis algorithm to detect the candidate chromatographic peaks;

[0028] S5.4: For the candidate chromatographic peaks that can be detected by the existing mass spectrometry data analysis algorithm, classify the peak types of the corresponding candidate chromatographic peaks as the first type; for the candidate chromatographic peaks that cannot be detected by the existing mass spectrometry data analysis algorithm, if its sawtooth index is less than the preset sawtooth index threshold, then classify the peak types of the corresponding candidate chromatographic peaks as the second type; otherwise, classify the peak types of the corresponding candidate chromatographic peaks as the third type.

[0029] Filtering the chromatographic peaks with peak height lower than the preset peak height threshold can reduce the number of candidate chromatographic peaks without affecting the accuracy of the detection result, improve the detection speed and reduce the appearance of false positive features. Secondly, when classifying the peak types of the candidate chromatographic peaks, not only consider whether they can be detected by the existing mass spectrometry data analysis algorithm, but also consider the sawtooth index, retain the features that cannot be detected by the existing mass spectrometry data analysis algorithm, and expand the scope of chemical substance identification. Regard the candidate chromatographic peaks of the third type as the reference type, and select whether to further filter according to the actual needs; if filtered, the peak types for subsequent analysis are two, namely the first type and the second type; if not filtered, the peak types for subsequent analysis are three.

[0030] Preferably, in step S5.1, if the retention time of the target chemical substance exists in the chemical exposure group database, only calculate the peak height and sawtooth index of the chromatographic peaks that appear in the retention time window.

[0031] By using the retention time window, the range and number of chromatographic peaks to be screened are reduced, and the subsequent matching speed and accuracy can be effectively improved.

[0032] Preferably, the specific method of step S6 is as follows:

[0033] S6.1: For each candidate chromatographic peak, sequentially match the fragment ions corresponding to the top of the candidate chromatographic peak with the fragment ions of the target chemical substance in the chemical exposure group database to obtain the matching fragment ions, their quantities, mass-to-charge ratios, and intensities at the top of each candidate chromatographic peak;

[0034] S6.2: Based on the matching fragment ions at the top of the candidate chromatographic peak and the fragment ions of the target chemical substance in the chemical exposure group database, calculate the first score and the second score;

[0035] S6.3: Based on the first score and the second score, calculate the relevant characteristic score of the candidate chromatographic peak corresponding to the target chemical substance;

[0036] S6.4: Repeat steps S6.1 - S6.3 to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance.

[0037] Preferably, in the step S6.2, the specific method for calculating the first score is:

[0038] Based on the quantity of the matching fragment ions at the top of the candidate chromatographic peak and the quantity of the fragment ions of the target chemical substance in the chemical exposure group database, calculate the first score, and the calculation formula is:

[0039]

[0040] In the formula, MFR i represents the first score of the matching fragment ions at the top of the i-th candidate chromatographic peak of the target chemical substance, N i represents the quantity of the matching fragment ions at the top of the i-th candidate chromatographic peak, and N T represents the quantity of the fragment ions of the target chemical substance in the chemical exposure group database.

[0041] Preferably, in the step S6.2, the specific method for calculating the second score is:

[0042] Based on the intensities of the matching fragment ions at the top of the candidate chromatographic peak and the intensities of the fragment ions of the target chemical substance in the chemical exposure group database, calculate the second score, and the calculation formula is:

[0043]

[0044] In the formula, SSM i represents the second score of the matching fragment ions at the top of the i-th candidate chromatographic peak of the target chemical substance, W Qi represents the intensity of the matching fragment ions at the top of the i-th candidate chromatographic peak, and W Ri represents the intensity of the i-th fragment ion of the target chemical substance in the chemical exposure group database.

[0045] Preferably, in the step S6.3, based on the first score and the second score, the specific method for calculating the relevant characteristic score of the candidate chromatographic peak corresponding to the target chemical substance is as follows:

[0046] Assign weight coefficients to the first score and the second score respectively, and calculate the relevant characteristic score of the candidate characteristic corresponding to the target chemical substance:

[0047] Score i =αMFR i +βSSM i

[0048] In the formula, Score i represents the relevant characteristic score of the candidate characteristic of the i-th candidate chromatographic peak corresponding to the target chemical substance, and α and β respectively represent the first and second weight coefficients.

[0049] Preferably, the specific method of the step S7 is as follows:

[0050] S7.1: Set the sorting priority of the peak types of the candidate chromatographic peaks, from high to low are the first type, the second type, and the third type;

[0051] S7.2: For all candidate chromatographic peaks belonging to the same peak type, arrange their relevant characteristic scores in descending order to obtain the candidate characteristics of the target chemical substance within each peak type;

[0052] S7.3: Concatenate the candidate characteristics of the target chemical substance within each peak type according to the sorting priority of the peak types to obtain the sorting result.

[0053] The present invention also provides a highly sensitive and high-throughput chemical substance annotation system for implementing the above annotation method, including:

[0054] A data acquisition module for collecting precursor ions and fragment ions in the sample to be detected based on the EISA technology;

[0055] A database construction module for constructing a chemical exposome database, including several chemical substances and their mass spectrometry data;

[0056] A chromatogram generation module for setting a mass accuracy standard, and successively taking the chemical substances in the chemical exposome database as target chemical substances to extract the precursor ion chromatograms of the sample to be detected;

[0057] A chromatogram correction module for correcting the precursor ion chromatograms of the sample to be detected to obtain a target chromatogram; the target chromatogram includes several chromatographic peaks;

[0058] A chromatographic peak screening module for screening all chromatographic peaks to obtain candidate chromatographic peaks and their peak types;

[0059] A data matching module for traversing and matching target chemical substances for each candidate chromatographic peak and calculating the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance;

[0060] A sorting module for sorting each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and relevant characteristic scores of the candidate chromatographic peak to obtain a sorting result;

[0061] A chemical substance detection module for outputting the candidate chromatographic peaks in the top several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected.

[0062] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0063] The present invention first uses the EISA technology to collect all precursor ions and fragment ions of the sample to be detected; then constructs a chemical exposome database including several chemical substances and their mass spectrometry data, and extracts the precursor ion chromatogram of the sample to be detected by taking the chemical substances as target chemical substances in turn; corrects the precursor ion chromatogram to obtain a target chromatogram containing several chromatographic peaks; then screens all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; for each candidate chromatographic peak, traverses and matches the target chemical substance to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance; finally, sorts each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and relevant characteristic scores of the candidate chromatographic peak, and outputs the candidate chromatographic peaks in the top several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected. The present invention can significantly increase the number of low-concentration chemical substances detected, with high sensitivity and high throughput. Description of the Drawings

[0064] Figure 1 It is a flowchart of a highly sensitive and high-throughput chemical substance annotation method described in Example 1.

[0065] Figure 2 It is a schematic diagram of the chemical exposome database described in Example 2.

[0066] Figure 3 It is a chromatogram extracted from the sample to be detected based on a preset mass accuracy standard described in Example 2.

[0067] Figure 4 It is the target chromatogram described in Example 2.

[0068] Figure 5 It is the chromatographic peak to be processed when there is a retention time described in Example 2.

[0069] Figure 6 When there is no retention time as described in Example 2, it is the chromatographic peak to be processed.

[0070] Figure 7 It is a schematic diagram comparing the detection results of the method provided in this example and the traditional peak extraction algorithm at different concentrations as described in Example 2.

[0071] Figure 8 It is a schematic diagram comparing the detection results of the method provided in this example and the traditional TMM acquisition method at different concentrations as described in Example 2.

[0072] Figure 9 It is a schematic diagram of the structure of a highly sensitive and high-throughput chemical substance annotation system as described in Example 3. Detailed implementation manner

[0073] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0074] To better illustrate this example, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;

[0075] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0076] The technical solutions of the present invention will be further described below with reference to the drawings and examples.

[0077] Example 1

[0078] This example provides a highly sensitive and high-throughput chemical substance annotation method, as Figure 1 shown, including:

[0079] S1: Based on the EISA technology, collect the precursor ions and fragment ions in the sample to be detected;

[0080] S2: Construct a chemical exposome database, including several chemical substances and their mass spectrometry data;

[0081] S3: Set the mass accuracy standard, and sequentially use the chemical substances in the chemical exposome database as target chemical substances to extract the precursor ion chromatograms of the sample to be detected;

[0082] S4: Calibrate the precursor ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes several chromatographic peaks;

[0083] S5: Screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types;

[0084] S6: For each candidate chromatographic peak, traverse and match the target chemical substances, and calculate the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substances;

[0085] S7: Sort each candidate chromatographic peak corresponding to the target chemical substances based on the peak type and relevant characteristic scores of the candidate chromatographic peaks to obtain a sorting result;

[0086] S8: Output the candidate chromatographic peaks in the top several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected.

[0087] In the specific implementation process, in this embodiment, first, use the EISA technology to collect all the precursor ions and fragment ions of the sample to be detected; then construct a chemical exposome database including several chemical substances and their mass spectrometry data, and extract the precursor ion chromatogram of the sample to be detected as the target chemical substances in turn; correct the precursor ion chromatogram to obtain a target chromatogram containing several chromatographic peaks; then screen all the chromatographic peaks to obtain candidate chromatographic peaks and their peak types; for each candidate chromatographic peak, traverse and match the target chemical substances to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substances; finally, sort each candidate chromatographic peak corresponding to the target chemical substances based on the peak type and relevant characteristic scores of the candidate chromatographic peaks, and output the candidate chromatographic peaks in the top several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected. This embodiment can significantly increase the number of low-concentration chemical substances detected, with high sensitivity and high throughput.

[0088] Embodiment 2

[0089] This embodiment provides a highly sensitive and high-throughput chemical substance annotation method, including:

[0090] S1 Based on the EISA technology, collect the precursor ions and fragment ions in the sample to be detected; the EISA technology collects all the precursor ions and fragment ions in the sample to be detected in a single full scan, and the corresponding relationship between the fragment ions and the precursor ions cannot be determined;

[0091] S2: Construct a chemical exposome database, including several chemical substances and their mass spectrometry data;

[0092] As Figure 2 shown, construct a chemical exposome database based on the mass spectrometry data of the chemical standard product local database, the open-source mass spectrometry database, and the public literature; the chemical exposome database includes several chemical substances and their mass spectrometry data, and the mass spectrometry data of each chemical substance includes the chemical substance name, precursor ion mass-to-charge ratio, fragment ion mass-to-charge ratio, and intensity;

[0093] S3: Set the quality accuracy standard, and successively use the chemical substances in the chemical exposome database as target chemical substances to correspondingly extract the parent ion chromatograms of the sample to be detected;

[0094] Specifically, successively use the chemical substances in the chemical exposome database as target chemical substances, and based on the parent ion mass-to-charge ratio of the target chemical substance and in combination with the set quality accuracy standard, obtain the range of the parent ion mass-to-charge ratio to be extracted; in the sample to be detected, for the ions whose parent ion mass-to-charge ratio is within the range of the parent ion mass-to-charge ratio, perform the operation of extracting and constructing a chromatogram. For example, if the parent ion mass-to-charge ratio of the target chemical substance in the chemical exposome database is 183.0991 Da and the preset quality accuracy standard is ±0.01 Da, then the range of the parent ion mass-to-charge ratio in the sample to be detected to be extracted is [183.0891, 183.1091]; as Figure 3 shown, in the sample to be detected, for all ions whose parent ion mass-to-charge ratio is within [183.0891, 183.1091], extract and construct chromatograms;

[0095] S4: For the chromatogram of the sample to be detected that is extracted, delete the corresponding chromatogram parts before the start of chromatographic elution and after the end of chromatographic elution, and use the remaining part of the chromatogram as the target chromatogram for subsequent analysis; each said target chromatogram includes a number of chromatographic peaks;

[0096] As Figure 4 shown, the chromatographic elution is set according to the actual situation. By deleting the corresponding chromatogram parts before the start of elution and after the end of elution and only retaining the chromatogram part during the elution process, the appearance of false positive features is reduced, making the detection result more accurate; in this embodiment, the chromatographic system starts elution at 90 s and ends elution at 900 s, then delete the parts before 90 s and after 900 s in the entire chromatogram, and retain the part between 90 s - 900 s to obtain the target chromatogram.

[0097] S5: Screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; specifically:

[0098] S5.1: If the target chemical substance in the chemical exposome database has a retention time, only calculate the peak height and sawtooth index of the chromatographic peaks that appear within the retention time window; as Figure 5 shown, if the target chemical substance has a retention time of 4.822 min in the chemical exposome database, which is converted to 289.32 s; then extract the chromatographic peaks in the chromatogram according to the preset time window for subsequent analysis. In this embodiment, the time window is ±90 s, and calculate the peak height and sawtooth index of the chromatographic peaks that appear within the time period of 289.32 ± 90 s;

[0099] As Figure 6As described above, if the retention time of the target chemical substance does not exist in the chemical exposure group database, calculate the peak height and sawtooth index of all the chromatographic peaks that appear;

[0100] S5.2: Filter the chromatographic peaks with peak height lower than the preset peak height threshold, and use the remaining chromatographic peaks as candidate chromatographic peaks;

[0101] S5.3: Detect the candidate chromatographic peaks using existing mass spectrometry data analysis algorithms;

[0102] S5.4: For the candidate chromatographic peaks that can be detected by the existing mass spectrometry data analysis algorithms, classify the peak types of the corresponding candidate chromatographic peaks as the first type; for the candidate chromatographic peaks that cannot be detected by the existing mass spectrometry data analysis algorithms, if the sawtooth index is less than the preset sawtooth index threshold, classify the peak types of the corresponding candidate chromatographic peaks as the second type; otherwise, classify the peak types of the corresponding candidate chromatographic peaks as the third type.

[0103] In this embodiment, the peak height threshold is 1000, the sawtooth index is 0.2, and the existing mass spectrometry data analysis algorithm is the XCMS algorithm.

[0104] If the retention time of the target chemical substance exists in the chemical exposure group database, the range and number of chromatographic peaks to be screened can be reduced, which can effectively improve the subsequent matching speed and accuracy.

[0105] Filtering the chromatographic peaks with peak height lower than the preset peak height threshold or sawtooth index greater than the preset sawtooth index threshold can reduce the number of candidate chromatographic peaks without affecting the accuracy of the detection results, improve the detection speed and reduce the false positive rate. Moreover, when classifying the peak types of the candidate chromatographic peaks, not only consider whether they can be detected by the existing mass spectrometry data analysis algorithms, but also consider the sawtooth index, retain the characteristics that cannot be detected by the existing mass spectrometry data analysis algorithms, and expand the scope of chemical substance identification. Take the candidate chromatographic peaks of the third type as the reference type, and select whether to further filter according to actual needs; if filtered, the peak types for subsequent analysis are two, namely the first type and the second type; if not filtered, the peak types for subsequent analysis are three.

[0106] S6: For each candidate chromatographic peak, traverse and match the target chemical substance, and calculate the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance; specifically:

[0107] S6.1: For each candidate chromatographic peak, sequentially match the fragment ions corresponding to the top of the candidate chromatographic peak with the fragment ions of the target chemical substance in the chemical exposure group database to obtain the fragment ions matched at the top of each candidate chromatographic peak and their quantity, mass-to-charge ratio, and intensity;

[0108] S6.2: Calculate the first score based on the number of fragmented ions matched at the top of the candidate chromatographic peak and the number of fragmented ions of the target chemical substance in the chemical exposome database. The calculation formula is as follows:

[0109]

[0110] In the formula, MFR i represents the first score of the fragmented ions matched at the top of the i-th candidate chromatographic peak of the target chemical substance, and N i represents the number of fragmented ions matched at the top of the i-th candidate chromatographic peak, and N T represents the number of fragmented ions of the target chemical substance in the chemical exposome database;

[0111] Calculate the second score based on the intensity of the fragmented ions matched at the top of the candidate chromatographic peak and the intensity of the fragmented ions of the target chemical substance in the chemical exposome database. The calculation formula is as follows:

[0112]

[0113] In the formula, SSM i represents the second score of the fragmented ions matched at the top of the i-th candidate chromatographic peak of the target chemical substance, and W Qi represents the intensity of the fragmented ions matched at the top of the i-th candidate chromatographic peak, and W Ri represents the intensity of the i-th fragmented ion of the target chemical substance in the chemical exposome database;

[0114] S6.3: The specific method for calculating the relevant characteristic score of the candidate chromatographic peak corresponding to the target chemical substance based on the first score and the second score is as follows:

[0115] Assign weight coefficients to the first score and the second score respectively, and calculate the relevant characteristic score of the candidate characteristic corresponding to the target chemical substance:

[0116] Score i = αMFR i + βSSM i

[0117] In the formula, Score i represents the relevant characteristic score of the candidate characteristic of the i-th candidate chromatographic peak corresponding to the target chemical substance, and α and β respectively represent the first and second weight coefficients; in this embodiment, α = 0.7 and β = 0.3;

[0118] S6.4: Repeat steps S6.1 - S6.3 to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance.

[0119] S7: Sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and related characteristic scores of the candidate chromatographic peak to obtain a sorting result; specifically:

[0120] S7.1: Set the sorting priorities of the peak types of the candidate chromatographic peaks, from high to low are the first type, the second type, and the third type;

[0121] S7.2: For all candidate chromatographic peaks belonging to the same peak type, arrange their related characteristic scores in descending order to obtain the candidate characteristics of the target chemical substance within each peak type;

[0122] S7.3: Concatenate the candidate characteristics of the target chemical substance within each peak type according to the sorting priorities of the peak types to obtain a sorting result.

[0123] S8: Output the candidate chromatographic peaks in the first several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected.

[0124] In the specific implementation process, the method provided in this embodiment is compared with traditional peak extraction algorithms, such as the XCMS algorithm, the MZmine3 algorithm, and the MSDIAL algorithm, for the characteristic detection of chemical substances at concentrations of 500 ppb and 0.8 ppb respectively, and is compared with the manual inspection results. The schematic diagram of the detection result comparison is as Figure 7 shown; it can be seen that the method in this embodiment detects the largest number of chemical substance characteristics at 500 ppb and 0.8 ppb, and has higher sensitivity.

[0125] The method provided in this embodiment is compared with the traditional TMM acquisition method. The object is a mixed standard containing 50 pollutants, and the number of pollutant characteristics at concentrations of 20 ppb, 4 ppb, and 0.8 ppb is obtained; as Figure 8 shown, it can be seen that the method provided in this embodiment can detect very little change in the chemical substance characteristics when the concentration is reduced, and far exceeds the number of chemical substance characteristics detected in the traditional TMM acquisition mode. The method provided in this embodiment can still accurately detect chemical substances at a low concentration of 0.8 ppb, with high sensitivity and high throughput.

[0126] Further set up an experiment for verification. Use the dust sample as the sample to be detected, and detect it by using the method provided in this embodiment and the traditional TMM acquisition method respectively. The detection results are shown in the following table:

[0127]

[0128]

[0129] In this experiment, the constructed chemical exposome database includes 200 pesticides. It can be seen that 25 pesticides were identified by the method of this embodiment; while the TTM collection method identified 13 pesticides through manual inspection, and chemical substances with MFR equal to 0 indicate that no parent ions were matched; moreover, the intensity of the parent ions identified by the method of this embodiment is much higher than that of the parent ions identified by the TTM collection method. That is, the method provided in this embodiment can collect the characteristics of fragment ions without affecting the abundance of parent ions, realizing high-sensitivity and high-throughput detection.

[0130] Example 3

[0131] This embodiment also provides a high-sensitivity and high-throughput chemical substance annotation system for implementing the annotation method described in Example 1 or 2, as Figure 9 shown, including:

[0132] A data collection module for collecting parent ions and fragment ions in a sample to be detected based on EISA technology;

[0133] A database construction module for constructing a chemical exposome database, including several chemical substances and their mass spectrometry data;

[0134] A chromatogram generation module for setting a mass accuracy standard, and sequentially using the chemical substances in the chemical exposome database as target chemical substances to extract the parent ion chromatograms of the sample to be detected correspondingly;

[0135] A chromatogram correction module for correcting the parent ion chromatograms of the sample to be detected to obtain target chromatograms; the target chromatograms include several chromatographic peaks;

[0136] A chromatographic peak screening module for screening all chromatographic peaks to obtain candidate chromatographic peaks and their peak types;

[0137] A data matching module for traversing and matching target chemical substances for each candidate chromatographic peak, and calculating the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance;

[0138] A sorting module for sorting each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and relevant characteristic scores of the candidate chromatographic peak to obtain a sorting result;

[0139] A chemical substance detection module for outputting the first several candidate chromatographic peaks in the sorting result as the final screening result of the chemical substances in the sample to be detected.

[0140] Identical or similar reference numerals correspond to identical or similar components;

[0141] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0142] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A highly sensitive and high-throughput method for chemical substance annotation, characterized in that Including: S1: Based on the EISA technology, collect precursor ions and fragment ions in the sample to be detected; S2: Construct a chemical exposome database, including several chemical substances and their mass spectrometry data; S3: Set the mass accuracy standard, and sequentially use the chemical substances in the chemical exposome database as target chemical substances, and correspondingly extract the precursor ion chromatogram of the sample to be detected; S4: Calibrate the precursor ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes several chromatographic peaks; S5: Screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; S6: For each candidate chromatographic peak, traverse and match the target chemical substances, and calculate the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance. The specific method is as follows: S6.1: For each candidate chromatographic peak, sequentially match the fragment ions corresponding to the top of the candidate chromatographic peak with the fragment ions of the target chemical substance in the chemical exposome database to obtain the fragment ions, their quantities, mass-to-charge ratios, and intensities matched at the top of each candidate chromatographic peak; S6.2: Calculate the first score and the second score based on the fragment ions matched at the top of the candidate chromatographic peak and the fragment ions of the target chemical substance in the chemical exposome database; The calculation formula for the first score is: In the formula, represents the first score of the fragment ions matched at the apex of the th candidate chromatographic peak of the target chemical substance, represents the number of fragment ions matched at the apex of the th candidate chromatographic peak, represents the number of fragment ions of the target chemical substance in the chemical exposome database; The calculation formula for the second score is: In the formula, represents the second score of the fragment ions matched at the apex of the th candidate chromatographic peak of the target chemical substance, represents the intensity of the fragment ions matched at the apex of the th candidate chromatographic peak, represents the intensity of the th fragment ion of the target chemical substance in the chemical exposome database; S6.3: Calculate the relevant characteristic score of the candidate chromatographic peak corresponding to the target chemical substance based on the first score and the second score; S6.4: Repeat steps S6.1 - S6.3 to obtain the relevant characteristic scores of each candidate chromatographic peak corresponding to the target chemical substance; S7: Sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and relevant characteristic score of the candidate chromatographic peak to obtain a sorting result; S8: Output the first several candidate chromatographic peaks in the sorting result as the final screening result of the chemical substances in the sample to be detected.

2. The highly sensitive and high-throughput chemical substance annotation method according to claim 1, characterized in that, In step S2, based on the mass spectrometry data of the chemical standard product local database, the open-source mass spectrometry database, and the public literature, construct a chemical exposome database; the chemical exposome database includes several chemical substances and their mass spectrometry data, and the mass spectrometry data of each chemical substance includes the chemical substance name, precursor ion mass-to-charge ratio, fragment ion mass-to-charge ratio, and intensity.

3. The highly sensitive and high-throughput chemical substance annotation method according to claim 1, characterized in that, In step S4, the specific method for calibrating the precursor ion chromatogram of the sample to be detected to obtain a target chromatogram is as follows: For the extracted chromatogram of the sample to be detected, delete the corresponding chromatogram parts before the start of chromatographic elution and after the end of chromatographic elution, and use the remaining chromatogram as the target chromatogram for subsequent analysis.

4. The highly sensitive and high-throughput chemical substance annotation method according to claim 2 or 3, characterized in that The specific method of step S5 is as follows: S5.1: Calculate the peak height and sawtooth index of all chromatographic peaks; S5.2: Filter out the chromatographic peaks with peak heights lower than the preset peak height threshold, and use the remaining chromatographic peaks as candidate chromatographic peaks; S5.3: Detect the candidate chromatographic peaks using existing mass spectrometry data analysis algorithms; S5.4: For candidate chromatographic peaks that can be detected by existing mass spectrometry data analysis algorithms, classify the peak types of the corresponding candidate chromatographic peaks as the first type; for candidate chromatographic peaks that cannot be detected by existing mass spectrometry data analysis algorithms, if their zigzag index is less than the preset zigzag index threshold, classify the peak types of the corresponding candidate chromatographic peaks as the second type; otherwise, classify the peak types of the corresponding candidate chromatographic peaks as the third type.

5. The highly sensitive and high-throughput chemical substance annotation method according to claim 1, wherein In the step S6.3, the specific method for calculating the relevant feature score of the candidate chromatographic peak corresponding to the target chemical substance based on the first score and the second score is as follows: Assign weight coefficients to the first score and the second score respectively, and calculate the relevant feature score of the candidate feature corresponding to the target chemical substance: In the formula, represents the relevant feature score of the candidate feature of the th candidate chromatographic peak corresponding to the target chemical substance, respectively represent the first and second weight coefficients.

6. The highly sensitive and high-throughput chemical substance annotation method according to claim 5, characterized in that, The specific method of the step S7 is as follows: S7.1: Set the sorting priority of the peak types of the candidate chromatographic peaks, from high to low as the first type, the second type, and the third type; S7.2: For all candidate chromatographic peaks belonging to the same peak type, arrange their relevant feature scores in descending order to obtain the candidate features of the target chemical substance within each peak type; S7.3: Concatenate the candidate features of the target chemical substance within each peak type according to the sorting priority of the peak types to obtain the sorting result.

7. A highly sensitive and high-throughput chemical substance annotation system, characterized in that, Including: A data acquisition module, configured to collect precursor ions and fragment ions in a sample to be detected based on the EISA technology; A database construction module, configured to construct a chemical exposome database, including several chemical substances and their mass spectrometry data; A chromatogram generation module, configured to set a mass accuracy standard, and sequentially use the chemical substances in the chemical exposome database as target chemical substances to extract the precursor ion chromatograms of the sample to be detected; A chromatogram correction module, configured to correct the precursor ion chromatogram of the sample to be detected to obtain a target chromatogram; the target chromatogram includes several chromatographic peaks; A chromatographic peak screening module, configured to screen all chromatographic peaks to obtain candidate chromatographic peaks and their peak types; A data matching module, configured to, for each candidate chromatographic peak, traverse and match the target chemical substance, and calculate the relevant feature score of each candidate chromatographic peak corresponding to the target chemical substance, including: For each candidate chromatographic peak, sequentially match the fragment ions corresponding to the top of the candidate chromatographic peak with the fragment ions of the target chemical substance in the chemical exposome database to obtain the fragment ions matched at the top of each candidate chromatographic peak and their quantity, mass-to-charge ratio, and intensity; Based on the fragment ions matched at the top of the candidate chromatographic peak and the fragment ions of the target chemical substance in the chemical exposome database, calculate the first score and the second score; The calculation formula of the first score is: In the formula, represents the first score of the fragment ions matched at the apex of the th candidate chromatographic peak of the target chemical substance, represents the number of fragment ions matched at the apex of the th candidate chromatographic peak, represents the number of fragment ions of the target chemical substance in the chemical exposome database; The calculation formula of the second score is: In the formula, represents the second score of the fragment ions matched at the top of the th candidate chromatographic peak of the target chemical substance, represents the intensity of the fragment ions matched at the top of the th candidate chromatographic peak, represents the intensity of the th fragment ion of the target chemical substance in the chemical exposome database; Based on the first score and the second score, calculate the relevant feature score of the candidate chromatographic peak corresponding to the target chemical substance; Repeat the above steps to obtain the relevant feature scores of each candidate chromatographic peak corresponding to the target chemical substance; A sorting module, configured to sort each candidate chromatographic peak corresponding to the target chemical substance based on the peak type and the relevant feature score of the candidate chromatographic peak to obtain a sorting result; A chemical substance detection module, which is used to output the candidate chromatographic peaks in the first several positions in the sorting result as the final screening result of the chemical substances in the sample to be detected.

Citation Information

Patent Citations

  • Chemical substance annotation method independent of retention time

    CN115453009A

  • Method for identifying medicinal materials using characteristic atlas

    CN1588050A