Method and device for mining novel chemical components of traditional Chinese medicine based on MS2 data

By constructing a hierarchical precursor ion list and metabolite molecular network, combining intelligent annotation and column chromatography, the problems of redundancy and low coverage of MS2 data were solved, efficient separation and identification of new chemical components of traditional Chinese medicine were achieved, research efficiency and purity were improved, and modernization of traditional Chinese medicine was promoted.

CN120340665APending Publication Date: 2025-07-18CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516589.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, under the MS2 data acquisition mode, the MS2 map data is redundant and inconsistent, resulting in low map coverage and low data quality of metabolite MS2, affecting the reliability of molecular network analysis, and the traditional separation and purification methods are inefficient, and the identification of new components is random and uncertain.

Method used

By constructing a hierarchical precursor ion list, dynamic window algorithm eliminates retention time overlap interference, combined with CAMERA algorithm, intelligent annotation of chromatographic peaks and adduct filtration, MS2 data were collected, metabolite molecular network was constructed, and candidate compounds were separated and purified by database comparison and column chromatography.

Benefits of technology

The efficiency and accuracy of traditional Chinese medicine ingredients research have been significantly improved. The efficiency of MS2 data collection has been increased by 40%, the spectrum coverage rate has been increased by 3-5 times, the separation efficiency has been increased by 8 times, the false positive rate has been reduced to below 5%, the purity of the compound has reached more than 95%, the development cycle of traditional Chinese medicine has been shortened by 60%, and the cost has been reduced by 45%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340665A_ABST
    Figure CN120340665A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for mining novel chemical components of traditional Chinese medicine based on MS2 data. The method comprises the following steps: acquiring precursor ion data from the traditional Chinese medicine; performing chromatographic peak identification through an XCMS R packet, and performing annotation through a CAMERAR packet to obtain an annotated chromatographic peak; grouping according to the annotated chromatographic peaks to obtain a plurality of ion groups; selecting representative precursor ions from each ion group according to ion strength and adduct stability, and generating an initial precursor ion list; sorting the precursor ions according to an RT ascending order, and dividing the precursor ions through a dynamic window algorithm to obtain a layered precursor ion list; mS2 data are collected in a segmented mode according to RT, and a metabolite molecular network is constructed according to the MS2 data; taking the unmatched node as a target node; obtaining a plurality of subfractions through column chromatography; and enriching and purifying the subfractions through semi-preparative liquid chromatography to obtain candidate compounds. According to the method, the map coverage rate of the metabolite MS2 can be increased, and the reliability of molecular network analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of analytical chemistry, and particularly to a method and device for mining new chemical components of traditional Chinese medicine based on MS2 data. Background Art

[0002] Traditional Chinese medicine, as a traditional medicine for treating various diseases in traditional Chinese medicine, can treat a variety of disease problems. Nowadays, with the continuous in-depth research on traditional Chinese medicine, many studies have separated and identified natural products from traditional Chinese medicine to clarify the action mechanism of traditional Chinese medicine, and significantly promoted the development of new therapeutic drugs. Therefore, discovering natural products from traditional Chinese medicine is one of the most direct and effective ways in drug research. However, traditional phytochemical separation and purification schemes are time-consuming and inefficient, and usually require multiple steps of separation and purification combined with mass spectrometry and spectroscopic analysis to determine the structure and novelty of new compounds. There is a problem of blind separation in traditional separation and purification methods, which makes the identification of new components have great randomness and uncertainty. Developing a method that can accurately and efficiently separate natural products is of great significance.

[0003] Currently, in the ordinary MS2 data acquisition mode (data-dependent acquisition, DDA), MS2 spectral data is obtained by non-targeted acquisition of precursor ions (i.e., MS1 data), which will result in redundant but inconsistent MS2 data, a large number of redundant nodes appear when building a molecular network, and similar compounds do not cluster. On the other hand, for precursor ions with low intensity, MS2 data may not be obtained, resulting in low spectral coverage and low data quality of metabolite MS2. Redundant nodes and low-quality MS2 data will complicate data analysis, affect the reliability of molecular network analysis, and hinder the research on natural components.

[0004] Therefore, how to invent a method for mining new chemical components of traditional Chinese medicine based on MS2 data, improve the spectral coverage of metabolite MS2, improve data quality, and improve the reliability of molecular network analysis has become an urgent problem to be solved. Summary of the Invention

[0005] To this end, the present invention provides a method and device for mining new chemical components of traditional Chinese medicine based on MS2 data. By collecting MS2 data based on a hierarchical precursor ion list, a concise metabolite molecular network MBMN is constructed, where each node represents a unique metabolite. By comparing and analyzing with a database, potential new compounds are discovered by combining MBMN and database annotation analysis, and the MS1 characteristics of these compounds are monitored based on ultra-high performance liquid chromatography-high resolution mass spectrometry to achieve the separation of new chemical components of traditional Chinese medicine.

[0006] To achieve the above object, the present invention provides the following technical solutions: A method for mining new chemical components of traditional Chinese medicine based on MS2 data, comprising:

[0007] Collecting precursor ion data from traditional Chinese medicine through a set collection system;

[0008] Identifying chromatographic peaks for the precursor ion data through the XCMS R package to obtain chromatographic peak data; Annotating the chromatographic peak data through the CAMERA R package to obtain annotated chromatographic peaks;

[0009] According to the annotated chromatographic peaks, removing the chromatographic peaks annotated as isotopes, and grouping the adduct ions of the same metabolite according to the neutral mass number to obtain a number of ion groups; Screening representative precursor ions from each of the ion groups according to the ion intensity and adduct stability to generate an initial precursor ion list;

[0010] Sorting the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; Dividing the sorted initial precursor ion list through a dynamic window algorithm to obtain a hierarchical precursor ion list;

[0011] Based on the hierarchical precursor ion list, collecting MS2 data in segments according to RT, so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; Importing the collected MS2 data into the online platform GNPS, and constructing a metabolite molecular network according to the structural similarity;

[0012] Comparing and matching the nodes in the metabolite molecular network with the database, and taking the unmatched nodes as target nodes; Separating the crude fractions containing the target nodes through column chromatography to obtain a number of sub-fractions; Enriching and purifying the sub-fractions corresponding to the target nodes through semi-preparative liquid chromatography to obtain candidate compounds.

[0013] As a preferred solution of the method for mining new chemical components of traditional Chinese medicine based on MS2 data, during the process of collecting the precursor ion data from traditional Chinese medicine through the set collection system, collecting through a UPLC-Q-TOF system from traditional Chinese medicine to obtain the precursor ion data.

[0014] As a preferred solution of the method for mining new chemical components of traditional Chinese medicine based on MS2 data, the chromatographic peak data includes: m / z value, retention time, range of m / z value, and range of chromatographic peak.

[0015] As a preferred solution of the method for mining new chemical components of traditional Chinese medicine based on MS2 data, during the process of annotating the chromatographic peak data through the CAMERA R package, the annotation includes: isotope, adduct, and PC group.

[0016] As an optimized solution for the method of mining new chemical components of traditional Chinese medicine based on MS2 data, in the process of dividing the sorted initial precursor ion list through the dynamic window algorithm to obtain the hierarchical precursor ion list, the dividing steps are as follows:

[0017] Set the RT tolerance window and group adjacent precursor ions step by step according to the RT overlap situation;

[0018] If the set RT range of the precursor ion has no overlap with the end RT of the previous SPL, classify the set precursor ion into this SPL;

[0019] Filter out precursor ions with intensities lower than the set threshold to generate a non-overlapping hierarchical precursor ion list.

[0020] The present invention also provides a device for mining new chemical components of traditional Chinese medicine based on MS2 data. Based on the above method for mining new chemical components of traditional Chinese medicine based on MS2 data, it includes:

[0021] A precursor ion data acquisition module for acquiring precursor ion data from traditional Chinese medicine through a set acquisition system;

[0022] A chromatographic peak identification and annotation module for identifying chromatographic peaks from the precursor ion data through the XCMS R package to obtain chromatographic peak data; and annotating the chromatographic peak data through the CAMERAR package to obtain annotated chromatographic peaks;

[0023] An initial precursor ion list generation module for removing chromatographic peaks annotated as isotopes according to the annotated chromatographic peaks, grouping adduct ions of the same metabolite by neutral mass number to obtain several ion groups; and screening representative precursor ions from each ion group according to ion intensity and adduct stability to generate an initial precursor ion list;

[0024] A hierarchical precursor ion list acquisition module for sorting the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; and dividing the sorted initial precursor ion list through the dynamic window algorithm to obtain a hierarchical precursor ion list;

[0025] A metabolite molecular network construction module for collecting MS2 data in segments according to RT based on the hierarchical precursor ion list to make each hierarchical precursor ion correspond to the fragment information of a single metabolite; importing the collected MS2 data into the online platform GNPS and constructing a metabolite molecular network according to structural similarity;

[0026] The target ion targeting and separation module is used to compare and match the nodes in the metabolite molecular network with a database, and regard the unmatched nodes as target nodes; separate the crude fraction containing the target nodes by column chromatography to obtain several sub-fractions; enrich and purify the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds.

[0027] As a preferred embodiment of the device for mining new chemical components of traditional Chinese medicine based on MS2 data, in the precursor ion data acquisition module, during the process of acquiring the precursor ion data from traditional Chinese medicine through the set acquisition system, the UPLC-Q-TOF system is used to acquire the precursor ion data from traditional Chinese medicine.

[0028] As a preferred embodiment of the device for mining new chemical components of traditional Chinese medicine based on MS2 data, in the chromatographic peak identification and annotation module, the chromatographic peak data includes: m / z value, retention time, range of m / z value, and range of chromatographic peak.

[0029] As a preferred embodiment of the device for mining new chemical components of traditional Chinese medicine based on MS2 data, in the chromatographic peak identification and annotation module, during the process of annotating the chromatographic peak data by the CAMERAR package, the annotation includes: isotope, adduct, and PC group.

[0030] As a preferred embodiment of the device for mining new chemical components of traditional Chinese medicine based on MS2 data, in the hierarchical precursor ion list acquisition module, during the process of dividing the sorted initial precursor ion list by the dynamic window algorithm to obtain the hierarchical precursor ion list, the sub-modules for division include:

[0031] The precursor ion grouping sub-module is used to set the RT tolerance window and group adjacent precursor ions step by step according to the RT overlap situation.

[0032] The precursor ion division sub-module is used to classify the set precursor ions into the SPL if the set RT range of the precursor ions has no overlap with the end RT of the previous SPL.

[0033] The precursor ion filtering sub-module is used to filter the precursor ions with intensities lower than the set threshold to generate a non-overlapping hierarchical precursor ion list.

[0034] The present invention has the following advantages: The present invention collects precursor ion data from traditional Chinese medicine through a set collection system; performs chromatographic peak identification on the precursor ion data through the XCMS R package to obtain chromatographic peak data; annotates the chromatographic peak data through the CAMERAR package to obtain annotated chromatographic peaks; according to the annotated chromatographic peaks, removes the chromatographic peaks annotated as isotopes, and groups the adduct ions of the same metabolite according to the neutral mass number to obtain a number of ion groups; screens representative precursor ions from each of the ion groups according to the ion intensity and adduct stability to generate an initial precursor ion list; sorts the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; divides the sorted initial precursor ion list through a dynamic window algorithm to obtain a hierarchical precursor ion list; based on the hierarchical precursor ion list, collects MS2 data in segments according to RT, so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; imports the collected MS2 data into the online platform GNPS, and constructs a metabolite molecular network according to the structural similarity; compares and matches the nodes in the metabolite molecular network with a database, and takes the nodes that cannot be matched as target nodes; separates the crude fraction containing the target nodes by column chromatography to obtain a number of sub-fractions; enriches and purifies the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds. The present invention significantly improves the efficiency and accuracy of the research on traditional Chinese medicine components through systematic technological innovation.First, the present invention constructs a hierarchical precursor ion list (SPLs), eliminates the interference of retention time overlap by using a dynamic window algorithm, combines the intelligent annotation and adduct filtering of chromatographic peaks by the CAMERA algorithm, improves the MS2 data acquisition efficiency by 40%, and increases the spectral coverage by 3-5 times. Secondly, through MBMN network topology analysis combined with database comparison, rapid elimination of known components and efficient screening of new compounds are achieved. Taking Euphorbia helioscopia L. as an example, 10 un-matched nodes among 1279 nodes are accurately locked as new components, and the false positive rate is controlled below 5%. Thirdly, a three-stage separation system of "network node → sub-fraction localization → semi-preparative liquid phase capture" is pioneered. The RT boundary of SPLs directly guides the sub-fraction cutting accuracy to reach ±3 seconds. Combined with DAD ultraviolet characteristic tracking, the purity of the target compound reaches more than 95%, and the separation efficiency is increased by 8 times compared with traditional column chromatography. In addition, the present invention breaks through the limitations of the traditional Chinese medicine field. Its SPLs algorithm can be migrated to the selection of precursor ions in proteomics, the MBMN topology analysis is applicable to the analysis of microbial metabolic networks, and the dynamic window RT management technology provides new ideas for suppressing interference peaks in clinical mass spectrometry. Finally, through the successful separation of 10 new compounds (such as C6 flavonoid glycosides and C8 triterpenoid saponins), its core value in analyzing the "multi-component - multi-target" action mechanism of traditional Chinese medicine is verified, which shortens the R & D cycle of traditional Chinese medicine by 60% and reduces the cost by 45%, provides a standardized technical tool for the modernization of traditional Chinese medicine, and helps to achieve the goal of scientific connotation analysis in the "Implementation Plan for Major Projects for the Revitalization and Development of Traditional Chinese Medicine". Brief Description of the Drawings

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.

[0036] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have substantial technical significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0037] Figure 1 It is a schematic flow chart of the method for mining new chemical components of traditional Chinese medicine based on MS2 data provided in Embodiment 1 of the present invention.

[0038] Figure 2Schematic diagram of the process for mining new chemical components of Euphorbia helioscopia in a possible embodiment provided in Embodiment 1 of the present invention;

[0039] Figure 3 Schematic diagram of the architecture of the device for mining new chemical components of traditional Chinese medicine based on MS2 data provided in Embodiment 2 of the present invention. Detailed implementation manners

[0040] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] Embodiment 1

[0042] Refer to Figure 1 , Embodiment 1 of the present invention provides a method for mining new chemical components of traditional Chinese medicine based on MS2 data, including the following steps:

[0043] S1. Collect precursor ion data from traditional Chinese medicine through a set acquisition system;

[0044] S2. Identify chromatographic peaks for the precursor ion data through the XCMS R package to obtain chromatographic peak data; annotate the chromatographic peak data through the CAMERAR package to obtain annotated chromatographic peaks;

[0045] S3. According to the annotated chromatographic peaks, remove the chromatographic peaks annotated as isotopes, and group the adduct ions of the same metabolite according to the neutral mass number to obtain several ion groups; screen representative precursor ions from each ion group according to the ion intensity and adduct stability to generate an initial precursor ion list;

[0046] S4. Sort the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; divide the sorted initial precursor ion list through the dynamic window algorithm to obtain a hierarchical precursor ion list;

[0047] S5. Based on the hierarchical precursor ion list, collect MS2 data in segments according to RT so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; import the collected MS2 data into the online platform GNPS and construct a metabolite molecular network according to the structural similarity;

[0048] S6. Compare and match the nodes in the metabolite molecular network with a database, and regard the nodes that cannot be matched as target nodes; separate the crude fractions containing the target nodes by column chromatography to obtain several sub-fractions; enrich and purify the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds.

[0049] In this embodiment, in step S1, precursor ion data is collected from traditional Chinese medicine by setting a collection system.

[0050] Specifically, the precursor ion data (LC-MS data) is collected from traditional Chinese medicine by a UPLC-Q-TOF system. The UPLC-Q-TOF system consists of an Agilent 1290 Infinity II (Santa Clara, CA) liquid phase system including a Diode-Array Detector (DAD) and an Impact II Q-TOF mass spectrometer (Bruker, Germany) equipped with an Electrospray Ionization (ESI) ion source.

[0051] Taking Euphorbia helioscopia L. as an example, the liquid phase conditions are as follows: detection is carried out at 40 °C using a Waters CORTECS UPLC C18 column (2.1×150 mm, 1.6 μm), 0.1% formic acid water is used as mobile phase A, acetonitrile is used as mobile phase B, the injection volume of each crude fraction sample is 5 μL, and the elution is carried out at a flow rate of 0.3 mL / min. The elution gradient is as follows: 0 - 10 min, 3 - 15% B; 10 - 11 min, 15% B; 11 - 14 min, 15 - 24% B; 14 - 15 min, 24 - 25% B; 15 - 25 min, 25 - 95% B; 25 - 27 min, 95% B; 27 - 27.5 min, 95 - 3% B; 27.5 - 30 min, 3% B. The mass spectrometry conditions are as follows: for the first-level mass spectrometry, data is collected in the positive ion mode of the ESI ion source (end plate offset potential -500 V, capillary voltage -4500 V), the atomization gas pressure is 2.0 bar, the flow rate of the drying gas (nitrogen) is 8.0 L / min, and the temperature is 220 °C. The scanning range is m / z 100 - 1300 Da, and the collection frequency is 4 Hz. The mass axis is calibrated using a sodium formate electrospray calibration solution.

[0052] In this embodiment, in step S2, chromatographic peak identification is performed on the precursor ion data by the XCMS R package to obtain chromatographic peak data; annotation is performed on the chromatographic peak data by the CAMERA R package to obtain annotated chromatographic peaks.

[0053] Specifically, the MS1 data is preprocessed and MS1 chromatographic peaks are identified using the XCMS R package. The main parameters for peak identification by XCMS (AMW - SMI) are shown in Table 1:

[0054]

[0055]

[0056] Table 1 Main parameters of the XCMS package

[0057] The detected chromatographic peak data includes the m / z (mass - to - charge ratio) value, retention time (rt), the range of m / z values (m / z min and m / z max), and the range of chromatographic peaks (rt min and rt max). Subsequently, the detected chromatographic peaks are annotated using the CAMERA R package, and the annotation rules are shown in Table 2. The resulting data matrix is exported as an Excel file (.xlsx) for further processing. The main parameters for peak annotation by CAMERA are shown in Table 2:

[0058]

[0059]

[0060] Table 2 Main parameters for peak annotation by CAMERA

[0061] In this embodiment, in step S3, according to the annotated chromatographic peaks, the chromatographic peaks annotated as isotopes are removed, and the adduct ions of the same metabolite are grouped by neutral mass number to obtain several ion groups; representative precursor ions are screened from each of the ion groups according to the ion intensity and adduct stability to generate an initial precursor ion list;

[0062] Specifically, first, the data matrix is processed. The R language is used to automatically remove the blank sample background solvent peaks from the sample dataset, and only the peaks of the sample compounds are retained. According to the CAMERA annotation (isotope, adduct, and PC group), the peaks annotated as isotopes (e.g., [M + 1], [M + 2], [M + 3]) are removed, and various adduct ions of the same metabolite (having the same neutral mass number, M) (e.g., [M + H] + , [M + Na] + , [M + NH4] +)Grouped based on neutral mass number, the precursor ions in a group originate from the same compound and have the same M and retention time. A representative precursor ion is selected from each group to exclude all other interfering ions and obtain a unique and characteristic MS2 spectrum for each putative metabolite (PM).

[0063] In this example, the preferred precursor ion adduct forms are [M+H] + and [M+NH4] + . In the same group, when the peaks of [M+H] + and [M+NH4] + both appear, the one with higher intensity is selected as the representative precursor ion; when only one of the peaks of [M+H] + and [M+NH4] + appears, regardless of the intensity of other adduct ions, only [M+H] + or [M+NH4] + is selected. When neither [M+H] + nor [M+NH4] + appears, the peak of the other adduct ion with higher intensity will be the representative precursor ion. All the remaining MS1 peaks that are completely unannotated and lack a clear neutral mass are considered their respective unique representative precursor ions.

[0064] Arrange these selected precursor ions in descending order of intensity. If an ion with lower intensity has more than 5 ions with intensities greater than it within the ±6-second window range of its retention time, it will be removed. This filtering ensures the simplicity and cleanliness of the precursor ion list while retaining the maximum number of high-intensity precursor ions for MS2 data acquisition to obtain high-quality MS2 spectra. The parameters for generating SPLs are shown in Table 3:

[0065]

[0066] Table 3 Generation parameters of the hierarchical precursor ion list

[0067] In this example, in step S4, the initial precursor ion list is sorted in ascending order of RT to obtain the sorted initial precursor ion list; the sorted initial precursor ion list is divided by the dynamic window algorithm to obtain the hierarchical precursor ion list;

[0068] Specifically, by selecting precursor ions, it is ensured that among multiple precursor ions from the same compound, only the most representative precursor ions within the same retention time range will be collected for MS2, and redundant precursor ions and those that produce incomplete spectra will not be collected, thereby reducing interference with the molecular network. Next, the retention times of the precursor ions are staggered to avoid retention time overlap, and the remaining precursor ions filtered and selected from each sample are further stratified into multiple sub-lists.

[0069] To generate a stratified precursor ion list, these precursor ions are first sorted in ascending order according to the start time (rtmin) of their retention time ranges. The specific steps for partitioning are as follows:

[0070] S41. Set the RT tolerance window and group adjacent precursor ions step by step according to the RT overlap situation;

[0071] Specifically, in the R script loop, the first SPL first puts in the first precursor ion, and the end time (rt max) of the retention time range of this precursor ion is used as the end point (spl end) of the first SPL.

[0072] S42. If the set RT range of the precursor ion has no overlap with the end RT of the previous SPL, then classify the set precursor ion into this SPL;

[0073] Specifically, subsequently, the ions with the smallest rt min are sequentially selected from the remaining ions for comparison. If rt min is greater than the spl end of the current SPL, it means that this ion does not overlap with the ions in the current SPL and can be put into the first SPL, and the spl end of this SPL is updated to the rt max of the newly added ion; if rt min is less than the current spl end, it means that this ion overlaps with the retention time ranges of other ions in the current SPL, and this ion will be put into a new SPL, and its rt max is used as the spl end of this SPL. This process is repeated, and the remaining precursor ions are compared with the spl end in sequence until all precursor ions are assigned to SPLs.

[0074] S43. Filter the precursor ions with intensities lower than the set threshold to generate a non-overlapping stratified precursor ion list.

[0075] Specifically, filter low-intensity ions (such as ions with intensities lower than the top 5% within the RT window), and finally generate a non-overlapping stratified precursor ion list.

[0076] In this embodiment, in step S5, based on the hierarchical precursor ion list, MS2 data is collected in segments according to RT, so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; the collected MS2 data is imported into the online platform GNPS, and a metabolite molecular network is constructed according to the structural similarity.

[0077] Specifically, according to the rt, rt min, and rt max of each precursor ion, the retention time error is automatically calculated, that is, rt plus or minus rt tolerance, which is used to confirm the sampling time range of each precursor ion when collecting MS2.

[0078] The collected MS2 data is imported into the online platform GNPS, and a metabolite molecular network (Metabolite-Based Molecular Network, MBMN) is constructed according to the structural similarity; each node represents a unique metabolite.

[0079] In this embodiment, in step S6, the nodes in the metabolite molecular network are compared and matched with the database, and the nodes that cannot be matched are used as target nodes; the crude fraction containing the target nodes is separated by column chromatography to obtain several sub-fractions; the sub-fractions corresponding to the target nodes are enriched and purified by semi-preparative liquid chromatography to obtain candidate compounds.

[0080] Specifically, the MBMN nodes are compared with the database in multiple dimensions (m / z, RT, fragment spectrum), and the unmatched nodes are screened as target nodes with potential new components. After determining the separated nodes, the step-by-step separation is started. The crude fraction containing the target compound is separated into a series of sub-fractions by column chromatography. A part of the sample is taken from each sub-fraction to prepare an LC-MS sample and measured under the same liquid phase conditions. Based on the MS1 characteristics in the network, that is, rt and m / z, the target compound is located and monitored, and the sub-fractions containing the target compound are combined. Repeatedly, the sub-fractions containing the target compound are further combined and separated. After the target compound is relatively pure, according to the DAD module in LC-MS, the maximum ultraviolet absorption value of the target compound is measured. On the semi-preparative liquid chromatography, using the same mobile phase in LC-MS and the measured maximum ultraviolet absorption, on the basis of ensuring the same elution order of the compound, the ultraviolet detection wavelength of the semi-preparative liquid is set to locate the target compound, and thus the pure product of the target compound is obtained using the semi-preparative liquid. Finally, after taking a part of the pure product to prepare an LC-MS sample, the MS2 data of the pure product is collected, and MS2 data is obtained while obtaining MS1 data. By comparing the MS2 spectrum of the final pure compound with the spectrum of the target node in the MBMN, and comparing the MS1 characteristics (i.e., m / z, rt) of the pure product with the MS1 characteristics of the node, it is confirmed that the final obtained pure product is the compound represented by the target node, realizing the targeted separation of the target component.

[0081] In a possible embodiment, a new chemical composition analysis example of Euphorbia helioscopia is provided as follows:

[0082] As Figure 2 shown, the steps to obtain the new chemical components of Euphorbia helioscopia are as follows:

[0083] T1. Sample preparation: After crushing the dried Euphorbia helioscopia sample (8.6 kg), ultrasonic extraction is carried out using 10 times the amount of 95% ethanol, extracting 2 times at 50 °C, 2 h each time. After filtering the two extraction solutions and combining them, rotary evaporation is used to concentrate to an extract paste, obtaining 1.1 kg of crude extract. The crude extract is suspended with 1 time the amount of pure water, and then successively extracted 5 times with petroleum ether and ethyl acetate respectively according to a volume ratio of 1:1. After combining the solutions of the same extractant and rotary evaporating and concentrating them to an extract paste respectively, finally 351 g of petroleum ether fraction sample, 284 g of ethyl acetate fraction sample and the remaining water fraction are obtained.

[0084] T2. LC-MS data acquisition: LC-MS data is acquired using a UPLC-Q-TOF system, which consists of an Agilent 1290 Infinity II (Santa Clara, CA) liquid phase system containing a Diode-Array Detector (DAD) and an Impact II Q-TOF mass spectrometer (Bruker, Germany) equipped with an Electrospray Ionization (ESI) ion source. Liquid phase conditions: Detection is carried out using a Waters CORTECS UPLC C18 column (2.1×150 mm, 1.6 μm) at 40 °C, 0.1% formic acid water as mobile phase A, acetonitrile as mobile phase B. The injection volume of each crude fraction sample is 5 μL, and elution is carried out at a flow rate of 0.3 mL / min. The elution gradient is as follows: 0 - 10 min, 3 - 15% B; 10 - 11 min, 15% B; 11 - 14 min, 15 - 24% B; 14 - 15 min, 24 - 25% B; 15 - 25 min, 25 - 95% B; 25 - 27 min, 95% B; 27 - 27.5 min, 95 - 3% B; 27.5 - 30 min, 3% B. Mass spectrometry conditions: For the first-stage mass spectrometry, acquisition is carried out in the positive ion mode of the ESI ion source (end plate offset potential -500 V, capillary voltage -4500 V), the nebulizing gas pressure is 2.0 bar, the flow rate of the drying gas (nitrogen) is 8.0 L / min, and the temperature is 220 °C. The scanning range is m / z 100 - 1300 Da, and the acquisition frequency is 4 Hz. The mass axis is calibrated using a sodium formate electrospray calibration solution.

[0085] T3. Selection of precursor ions after identification and annotation of 6 sample peaks: After peak identification of the original mzXML data of six samples and a blank solvent (methanol) control using the peak identification script of the R package XCMS (AMW-SMI algorithm), seven data matrices were processed. The R language was used to automatically remove the blank sample background solvent peaks from the six sample datasets, and only the peaks of sample compounds were retained. According to the CAMERA annotation (isotope, adduct, and PC group), the peaks annotated as isotopes (e.g., [M+1], [M+2], [M+3]) were removed, and various adduct ions of the same metabolite (with the same neutral mass number, M) (e.g., [M+H] + , [M+Na] + , [M+NH4] + ) were grouped based on the neutral mass number. The precursor ions in the group originated from the same compound with the same M and retention time. A representative precursor ion was selected from each group to exclude all other interfering ions and obtain a unique and characteristic MS2 spectrum for each putative metabolite (PM). These selected precursor ions were arranged in descending order of intensity. If the intensity of an ion with low intensity had more than 5 ions with intensities greater than that ion within the ±6-second window of its retention time, it was removed. In this way, all the precursor ion peaks in each sample were obtained. Among them, 1252, 1479, and 1443 precursor ion peaks were identified for the EA7-9 samples respectively; 1506, 1459, and 1340 precursor ion peaks were identified for the EA12-14 samples respectively. The CAMERA package was used to annotate these precursor ions. After peak identification and annotation of the original data of the blank control, the solvent peaks in the blank control were removed from the six samples using an R script. By comparing with the solvent peaks through the m / z tolerance (0.01 Da) and rt tolerance (±6 s), the precursor ion peaks in the samples that met both tolerances were removed to automatically deduct the solvent background and only retain the precursor ion peaks from the samples. The number of precursor ions remaining in the six samples of EA7-9 and EA12-14 was 1218, 1447, 1398, 1415, 1395, and 1279 precursor ions in sequence. Subsequently, these precursor ions were selected. By identifying the same annotation form ([M+num]), all isotope peaks were removed. Then, according to the precursor ion selection rule, the most representative precursor ion was selected from multiple precursor ions of the same compound with different adduct forms to represent the putative metabolite, and after intensity screening. Finally, the number of precursor ions remaining in the six samples of EA7-9 and EA12-14 was 130, 142, 125, 135, 155, and 142 precursor ions in sequence.

[0086] T4. Generate a hierarchical precursor ion list for LC-MS / MS data acquisition: Further hierarchicalize the representative precursor ions of each sample into multiple sub-lists. Through an R script, relying on the characteristics of retention time, judge the retention time overlap after sorting, so as to assign these precursor ions to different lists to separately acquire MS2, and layer by layer to achieve the "hierarchicalization" of precursor ions. The precursor ions in one sample are divided into several "Lists", which means that MS2 will be acquired several times for the same sample, and the precursor ions acquired each time are different. The specific method for generating hierarchical precursor ions is as follows: First, sort these precursor ions in ascending order according to the start time (rt min) of their retention time ranges. In the R script loop, the first SPL is filled with the first precursor ion, and the end time (rt max) of the retention time range of this precursor ion is used as the end point (spl end) of the first SPL. Subsequently, the ion with the smallest rt min is selected from the remaining ions in turn for comparison. If rt min is greater than the spl end of the current SPL, it means that this ion does not overlap with the ions in the current SPL and can be put into the first SPL, and update the splend of this SPL to the rt max of the newly added ion; if rt min is less than the current splend, it means that this ion overlaps with the retention time ranges of other ions in the current SPL, and this ion will be put into a new SPL, and its rtmax is used as the splend of this SPL. This process is repeated, and the remaining precursor ions are compared with splend in order until all precursor ions are assigned to SPLs. At the same time, the retention time error is automatically calculated according to the rt, rt min and rt max of each precursor ion, that is, rt plus or minus rt tolerance, which is used to confirm the sampling time range of each precursor ion when acquiring MS2. Finally, the precursor ions in the six samples EA7-9 and EA12-14 are sequentially assigned to 7, 9, 9, 9, 9 and 8 SPLs. For example, for sample EA14, sampling needs to be repeated 8 times, each corresponding to one SPL in its SPLs, and the retention times of the precursor ions in each SPL do not overlap. Some of the generated SPLs are shown in Table 4:

[0087]

[0088]

[0089]

[0090]

[0091] Table 4 Distribution of some precursor ions for MS2 analysis in SPLs

[0092] T5. LC-MS-guided targeted ion isolation: Nodes representing known compounds can be matched through comparison and matching with a database. When nodes cannot be matched, they are considered likely to represent potential new compounds. Ten of them were selected for targeted isolation, and the crude fraction containing the target compound was separated into a series of sub-fractions using column chromatography. A portion of the sample was taken from each sub-fraction to prepare an LC-MS sample and measured under the same liquid phase conditions. Based on the MS1 features in the network, namely rt and m / z, the target compound was located and monitored, and the sub-fractions containing the target compound were combined. The sub-fractions containing the target compound were repeatedly further combined and separated. After the target compound was relatively pure, according to the DAD module in LC-MS, the maximum ultraviolet absorption value of the target compound was measured. On semi-preparative liquid chromatography, using the same mobile phase as in LC-MS and the measured maximum ultraviolet absorption, and on the basis of ensuring the same elution order of the compounds, the ultraviolet detection wavelength of the semi-preparative liquid was set to locate the target compound, and thus the pure product of the target compound was obtained using semi-preparative liquid chromatography. Finally, after taking a portion of the pure product to prepare an LC-MS sample, the MS2 data of the pure product were collected, and MS2 data were obtained while obtaining MS1 data. By comparing the MS2 spectrum of the final pure compound with the spectrum of the target node in MBMN, and also comparing the MS1 features (i.e., m / z, rt) of the pure product with the MS1 features of the node, it was confirmed that the pure product finally obtained was the compound represented by the target node, and the targeted isolation of 10 target components was achieved. The target ions obtained from the targeted isolation are shown in Table 5:

[0093]

[0094] Table 5 Information on the isolated target nodes and their corresponding 10 components

[0095] In summary, the present invention collects precursor ion data from traditional Chinese medicine through a set collection system; identifies chromatographic peaks for the precursor ion data through the XCMSR package to obtain chromatographic peak data; annotates the chromatographic peak data through the CAMERA R package to obtain annotated chromatographic peaks; according to the annotated chromatographic peaks, removes the chromatographic peaks annotated as isotopes, and groups the adduct ions of the same metabolite according to the neutral mass number to obtain several ion groups; screens representative precursor ions from each of the ion groups according to the ion intensity and adduct stability to generate an initial precursor ion list; sorts the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; divides the sorted initial precursor ion list through a dynamic window algorithm to obtain a hierarchical precursor ion list; based on the hierarchical precursor ion list, collects MS2 data in segments according to RT, so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; imports the collected MS2 data into the online platform GNPS, and constructs a metabolite molecular network according to the structural similarity; compares and matches the nodes in the metabolite molecular network with a database, and takes the unmatched nodes as target nodes; separates the crude fraction containing the target nodes through column chromatography to obtain several sub-fractions; enriches and purifies the sub-fractions corresponding to the target nodes through semi-preparative liquid chromatography to obtain candidate compounds. The present invention significantly improves the efficiency and accuracy of the research on the components of traditional Chinese medicine through systematic technological innovation.First, the present invention constructs hierarchical precursor ion lists (SPLs), eliminates retention time overlap interference using the dynamic window algorithm, and combines the CAMERA algorithm for intelligent annotation and adduct filtering of chromatographic peaks, which improves the MS2 data acquisition efficiency by 40% and increases the spectral coverage rate by 3 - 5 times. Secondly, through MBMN network topology analysis combined with database comparison, rapid elimination of known components and efficient screening of new compounds are achieved. Taking Euphorbia helioscopia as an example, 10 un-matched nodes among 1279 nodes are accurately locked as new components, and the false positive rate is controlled below 5%. Furthermore, a three-level separation system of "network node → sub-fraction localization → semi-preparative liquid phase capture" is pioneered. The RT boundaries of SPLs directly guide the sub-fraction cutting accuracy to ±3 seconds. Combined with DAD ultraviolet feature tracking, the purity of the target compound reaches over 95%, and the separation efficiency is 8 times higher than that of traditional column chromatography. In addition, the present invention breaks through the limitations of the traditional Chinese medicine field. Its SPLs algorithm can be migrated to precursor ion selection in proteomics, MBMN topology analysis is applicable to the analysis of microbial metabolic networks, and the dynamic window RT management technology provides new ideas for suppressing interference peaks in clinical mass spectrometry. Finally, through the successful separation of 10 new compounds (such as C6 flavonoid glycosides and C8 triterpenoid saponins), its core value in analyzing the "multi-component - multi-target" action mechanism of traditional Chinese medicine is verified, which shortens the R & D cycle of traditional Chinese medicine by 60% and reduces the cost by 45%, provides a standardized technical tool for the modernization of traditional Chinese medicine, and helps to achieve the goal of scientific connotation analysis in the "Implementation Plan for Major Projects for the Revitalization and Development of Traditional Chinese Medicine".

[0096] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0097] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0098] Embodiment 2

[0099] See Figure 3 , Embodiment 2 of the present invention also provides a device for mining new chemical components of traditional Chinese medicine based on MS2 data, including:

[0100] Precursor ion data acquisition module 001, which is used to acquire precursor ion data from traditional Chinese medicine through a set acquisition system;

[0101] Chromatographic peak identification and annotation module 002, which is used to identify chromatographic peaks from the precursor ion data through the XCMS R package to obtain chromatographic peak data; and annotate the chromatographic peak data through the CAMERA R package to obtain annotated chromatographic peaks;

[0102] Initial precursor ion list generation module 003, which is used to remove chromatographic peaks annotated as isotopes according to the annotated chromatographic peaks, and group the adduct ions of the same metabolite by neutral mass number to obtain several ion groups; screen representative precursor ions from each ion group according to ion intensity and adduct stability, and generate an initial precursor ion list;

[0103] Stratified precursor ion list acquisition module 004, which is used to sort the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; divide the sorted initial precursor ion list through a dynamic window algorithm to obtain a stratified precursor ion list;

[0104] Metabolite molecular network construction module 005, which is used to collect MS2 data in segments according to RT based on the stratified precursor ion list, so that each stratified precursor ion corresponds to the fragment information of a single metabolite; import the collected MS2 data into the online platform GNPS, and construct a metabolite molecular network according to structural similarity;

[0105] Target ion targeted separation module 006, which is used to compare and match the nodes in the metabolite molecular network with a database, and regard the nodes that cannot be matched as target nodes; separate the crude fractions containing the target nodes by column chromatography to obtain several sub-fractions; enrich and purify the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds.

[0106] In this embodiment, in the precursor ion data acquisition module 001, during the process of acquiring the precursor ion data from traditional Chinese medicine through the set acquisition system, the precursor ion data is acquired from traditional Chinese medicine through a UPLC-Q-TOF system.

[0107] In this embodiment, in the chromatographic peak identification and annotation module 002, the chromatographic peak data includes: m / z value, retention time, range of m / z value, and range of chromatographic peak.

[0108] In this embodiment, in the chromatographic peak identification and annotation module 002, during the process of annotating the chromatographic peak data through the CAMERA R package, the annotation includes: isotope, adduct, and PC group.

[0109] In this embodiment, in the hierarchical precursor ion list acquisition module 004, during the process of dividing the sorted initial precursor ion list through the dynamic window algorithm to obtain the hierarchical precursor ion list, the sub-modules for division include:

[0110] The precursor ion grouping sub-module 041 is used to set an RT tolerance window and group adjacent precursor ions step by step according to the RT overlap situation;

[0111] The precursor ion division sub-module 042 is used to classify the set precursor ions into the SPL if the set RT range of the precursor ions has no overlap with the end RT of the previous SPL;

[0112] The precursor ion filtering sub-module 043 is used to filter out precursor ions with intensities lower than a set threshold to generate a non-overlapping hierarchical precursor ion list.

[0113] It should be noted that for the information interaction, execution process, etc. among the above system modules, since they are based on the same concept as the method embodiment in Embodiment 1 of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be elaborated here.

[0114] Embodiment 3

[0115] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which program codes for a method of mining new chemical components of traditional Chinese medicine based on MS2 data are stored, and the program codes include instructions for executing the method of mining new chemical components of traditional Chinese medicine based on MS2 data in Embodiment 1 or any possible implementation manner thereof.

[0116] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available media can be magnetic media (for example, floppy disks, hard disks, magnetic tapes), optical media (for example, DVDs), or semiconductor media (for example, solid state disks (SSDs)), etc.

[0117] Embodiment 4

[0118] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0119] The processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the method for mining new chemical components of traditional Chinese medicine based on MS2 data in Embodiment 1 or any possible implementation manner thereof by invoking the program instructions.

[0120] Specifically, the processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.

[0121] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).

[0122] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program code executable by the computing system, so that they can be stored in a storage system and executed by the computing system. And in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.

[0123] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it on the basis of the present invention, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A method for mining new chemical components of traditional Chinese medicine based on MS2 data, characterized in that, Comprising: Collecting precursor ion data from traditional Chinese medicine through a set collection system; Performing chromatographic peak identification on the precursor ion data through the XCMS R package to obtain chromatographic peak data; performing annotation on the chromatographic peak data through the CAMERAR package to obtain annotated chromatographic peaks; According to the annotated chromatographic peaks, removing the chromatographic peaks annotated as isotopes, and grouping the adduct ions of the same metabolite according to the neutral mass number to obtain a number of ion groups; screening representative precursor ions from each of the ion groups according to the ion intensity and adduct stability to generate an initial precursor ion list; Sorting the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; Dividing the sorted initial precursor ion list through a dynamic window algorithm to obtain a hierarchical precursor ion list; Based on the hierarchical precursor ion list, collecting MS2 data in segments according to RT so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; importing the collected MS2 data into the online platform GNPS and constructing a metabolite molecular network according to the structural similarity; Comparing and matching the nodes in the metabolite molecular network with a database, and taking the nodes that cannot be matched as target nodes; separating the crude fractions containing the target nodes by column chromatography to obtain a number of sub-fractions; enriching and purifying the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds.

2. The method for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 1, wherein During the process of collecting the precursor ion data from traditional Chinese medicine through the set collection system, collecting the precursor ion data from traditional Chinese medicine through a UPLC-Q-TOF system.

3. The method for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 2, wherein The chromatographic peak data includes: m / z value, retention time, m / z value range, and chromatographic peak range.

4. The method for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 3, wherein During the process of annotating the chromatographic peak data through the CAMERAR package, the annotation includes: isotope, adduct, and PCgroup.

5. The method for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 4, wherein During the process of dividing the sorted initial precursor ion list through the dynamic window algorithm to obtain the hierarchical precursor ion list, the dividing steps are: Setting an RT tolerance window and gradually grouping adjacent precursor ions according to the RT overlap situation; If the RT range of the set precursor ion has no overlap with the end RT of the previous SPL, then classifying the set precursor ion into the SPL; Filtering the precursor ions with intensities lower than the set threshold to generate a non-overlapping hierarchical precursor ion list.

6. An apparatus for mining new chemical components of traditional Chinese medicine based on MS2 data, using the method for mining new chemical components of traditional Chinese medicine based on MS2 data according to any one of claims 1-5, characterized in that, Comprising: A precursor ion data collection module for collecting precursor ion data from traditional Chinese medicine through a set collection system; A chromatographic peak identification and annotation module for performing chromatographic peak identification on the precursor ion data through the XCMS R package to obtain chromatographic peak data; performing annotation on the chromatographic peak data through the CAMERAR package to obtain annotated chromatographic peaks; An initial precursor ion list generation module, which is used to remove chromatographic peaks annotated as isotopes according to the annotated chromatographic peaks, group the adduct ions of the same metabolite by neutral mass number, and obtain a number of ion groups; screen representative precursor ions from each of the ion groups according to ion intensity and adduct stability, and generate an initial precursor ion list; A hierarchical precursor ion list acquisition module, which is used to sort the initial precursor ion list in ascending order of RT to obtain the sorted initial precursor ion list; Divide the sorted initial precursor ion list by a dynamic window algorithm to obtain a hierarchical precursor ion list; A metabolite molecular network construction module, which is used to collect MS2 data in segments according to RT based on the hierarchical precursor ion list, so that each hierarchical precursor ion corresponds to the fragment information of a single metabolite; import the collected MS2 data into the online platform GNPS, and construct a metabolite molecular network according to structural similarity; A target ion targeted separation module, which is used to compare and match the nodes in the metabolite molecular network with a database, and take the unmatched nodes as target nodes; separate the crude fractions containing the target nodes by column chromatography to obtain a number of sub-fractions; enrich and purify the sub-fractions corresponding to the target nodes by semi-preparative liquid chromatography to obtain candidate compounds.

7. The device for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 6, characterized in that, In the precursor ion data acquisition module, during the process of acquiring the precursor ion data from traditional Chinese medicine through the set acquisition system, the precursor ion data is acquired from traditional Chinese medicine through a UPLC-Q-TOF system.

8. The device for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 7, wherein, In the chromatographic peak identification and annotation module, the chromatographic peak data includes: m / z value, retention time, range of m / z value, and range of chromatographic peak.

9. The device for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 8, characterized in that, In the chromatographic peak identification and annotation module, during the process of annotating the chromatographic peak data by the CAMERAR package, the annotation includes: isotope, adduct, and PC group.

10. The device for mining new chemical components of traditional Chinese medicine based on MS2 data according to claim 9, characterized in that, In the hierarchical precursor ion list acquisition module, during the process of dividing the sorted initial precursor ion list by the dynamic window algorithm to obtain the hierarchical precursor ion list, the sub-modules of the division include: A precursor ion grouping sub-module, which is used to set an RT tolerance window and group adjacent precursor ions step by step according to the RT overlap situation; A precursor ion division sub-module, which is used to classify the set precursor ions into the SPL if the set RT range of the precursor ions has no overlap with the end RT of the previous SPL; A precursor ion filtering sub-module, which is used to filter precursor ions with intensities lower than the set threshold to generate a non-overlapping hierarchical precursor ion list.