Metabonomics MS / MS database based on liquid chromatography-mass spectrometry and construction method
By constructing a metabolomics MS/MS database based on LCMS, using AI prediction database to improve the accuracy and depth of metabolites identification, the problems of small quantities and low accuracy of metabolites in the existing technology have been solved, and the in-depth development of metabolites research has been achieved.
Patent Information
- Application Number
- CN202510195105.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The small number of metabolites and low accuracy in existing metabolomics technologies lead to an increased risk of deviation from research results and false conclusions.
The metabolomics MS/MS database construction method based on LCMS is adopted. By obtaining metabolites information from standard product databases, commercial databases and public databases, and using ESI-MS/MS machine learning model and random forest algorithm, the secondary mass spectrometry map and retention time are predicted, and the AI prediction library is constructed, and finally the metabolomics MS/MS database with extended capacity and function is constructed.
More than 95% of metabolites have been achieved, the accuracy and depth of metabolites identification have been improved, the database capacity has been expanded, cross-platform data exchange has been promoted, and the in-depth development of metabolomics research has been supported.
Smart Images

Figure CN120108577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bio-omics, and in particular to a metabolomics MS / MS database based on liquid chromatography-mass spectrometry and a construction method thereof. Background Art
[0002] Metabolomics is the newest member of the “omics” family, which aims to comprehensively characterize small molecules in biological samples. After decades of development, metabolomics has made significant progress, but its in-depth application is still subject to two core bottlenecks: one is the lack of depth in metabolite identification, which currently only reaches an identification rate of about 5-10%; the other is the limited accuracy of identification, which frequently results in false positive results due to complex situations such as isomers. These two dilemmas depend to a large extent on the capacity and functional performance of the metabolic database used. In a large number of published studies, researchers are prone to fall into the trap of pursuing the number of metabolites one-sidedly while ignoring the accuracy of identification. If there is a deviation in the qualitative link of metabolites, more identification numbers will not only fail to enhance the research, but may lead to deviations in research results and even wrong conclusions and insights. This issue has attracted widespread attention in the academic community. For example, Georgios Theodoridis et al. have clearly called for strengthening the rigor of metabolite identification and standardized reporting, which will effectively consolidate the intrinsic value and credibility of metabolomics as a science. Metabolomics databases play a key and decisive role.
[0003] In non-targeted metabolomics research, the process of metabolite qualitative analysis is not complicated, that is, by comparing the chromatographic mass spectrometry information of metabolites in the collected samples with the chromatographic mass spectrometry information of standard substances, qualitative analysis can be achieved if a match is achieved. The matching information mainly includes MS1 (primary mass spectrometry, which can obtain accurate molecular weight), MS2 (secondary mass spectrometry, which can obtain fragmentation characteristic information), RT (retention time, mainly used to distinguish isomers), and CCS (collision cross-section, a parameter collected by ion mobility mass spectrometry, whose main function is still to distinguish isomers). Therefore, the core information in the metabolomics database is MS1, MS2, RT and CCS.
[0004] Commonly used metabolomics databases can be roughly divided into the following types: 1. Self-built standard product library: Purchase or synthesize standard products by yourself, collect them on your own mass spectrometry platform, obtain MS1, MS2, RT and other information, and build a local standard product database; 2. Commercial database: The standard product information has been collected and integrated into a directly usable paid database. Most of them are constructed with information obtained from actual standard product collection. Commonly used ones include mzCloud, NIST, Metlin, etc. 3. Public database: Some units or laboratories make the standard sample spectra information collected by their own platforms or directly integrate the database public for free download and use. The MoNA library has integrated most of the commonly used public libraries. The disadvantage is that the information is relatively confusing and there are many problems in direct use. 4. Computer simulation database: This is a library constructed by generating predicted spectral information through computer simulation based on information such as the structural properties of compounds and metabolic reactions. There are many different prediction strategies and methods. With the continuous development of AI models, this type of library is expected to become a major trend.
[0005] From the perspective of qualitative accuracy, database 1>2>3>4, but the reality is that the standard substances available for purchase are very limited, probably only a few thousand or tens of thousands, and the cost is extremely high, which is simply a drop in the bucket for such a large number of metabolites. Therefore, computer simulation databases have become a solution with great potential, especially today when AI algorithms are advancing by leaps and bounds. This strategy can not only generate MS / MS spectra of compounds through AI models, but also predict RT and CCS values, further filter false positives, and improve the accuracy of identification. Summary of the invention
[0006] In view of the above-mentioned technical deficiencies, the purpose of the present invention is to provide a metabolomics MS / MS database based on liquid chromatography-mass spectrometry and a construction method to solve the problems of small number and low accuracy of metabolite qualitative analysis in the prior art.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for constructing a metabolomics MS / MS database based on liquid chromatography-mass spectrometry, the method comprising: The method comprises: Obtain mass spectrum information and chromatographic information from standard libraries and metabolite information from commercial and public databases; Based on the metabolite structure coding, an ESI-MS / MS machine learning model was constructed to predict and generate secondary mass spectra. Forest The algorithm predicts the retention time of metabolites and forms an AI prediction library; Construct a metabolomics MS / MS database based on standard library, commercial database, public database and AI prediction library; Data normalization and classification of metabolomics MS / MS database.
[0008] Preferably, in a possible implementation of the first aspect, the metabolite information acquisition method includes: Extracting secondary mass spectrometry data from commercial databases and public databases and formatting them into MSP files, wherein the content format of the MSP files is a unified preset format; Remove substances with molecular weights outside the preset range and remove information outside the scope of metabolomics research; De-redundancy processing is performed on repeated information.
[0009] Preferably, in a possible implementation manner of the first aspect, the commercial database content extraction method includes: Use mzVault to load the metabolite information file in the commercial database, which is a db file; Export the db file as a CSV file encoded in UTF-8; Extract the core information from the CSV file and format it into an MSP file.
[0010] Preferably, in a possible implementation manner of the first aspect, the core information includes CompoundName, ChemicalFormula, Ionization, ExtractedMass, Adduct, Polarity, ConfirmPrecursor, ConfirmEnergy and the corresponding secondary mass spectrometry fragment ion mass-to-charge ratio and response intensity value.
[0011] Preferably, in a possible implementation manner of the first aspect, the method for extracting content from a public database includes: Load the metabolite information file in the public database, which is a txt format file or an MSP file; Extract the metabolite information in the file and save it in a CSV file; Merge CSV files from public databases from different sources to build an MSP file.
[0012] Preferably, in a possible implementation of the first aspect, the AI prediction library construction method includes: Obtain the InChI and SMILES structural codes corresponding to metabolites based on the metabolite information in the standard library, commercial database and public database; Input the InChI or SMILES structure code into the ESI-MS / MS machine learning model to predict the secondary mass spectrum. Output three predicted MS / MS fragmentation information corresponding to the positive and negative energies of 10ev, 20ev, and 30ev for each metabolite, and organize the MS / MS fragmentation information into an MSP file. Random Forest ForestThe algorithm predicts the RT retention time of metabolites, builds a local model using SMILES and RT information based on the self-built library standard, inputs the SMILES structure encoding information of the metabolite to be predicted into the model, and outputs the predicted RT retention time of the metabolite.
[0013] Preferably, in a possible implementation of the first aspect, the key information for normalization and organization of metabolomics MS / MS database data includes basic metabolite information, database ID, structure code, metabolite classification, metabolite profile and metabolite source.
[0014] Preferably, in a possible implementation of the first aspect, the metabolomics MS / MS database classification includes a medical library, an animal library, a plant library, a microbial library, a gut flora library, a traditional Chinese medicine library, and an exposure group library.
[0015] In the second aspect, a metabolomics MS / MS database based on liquid chromatography-mass spectrometry, the database includes: Standard product library, including mass spectrometry information of self-purchased standards and standard product spectrum information integrated from mzCloud; Commercial databases, including the integrated NIST_2020_MS / MS_HR.db file; Public databases, including secondary mass spectrometry information from MassBank, HMDB, and MS-DIAL; AI prediction library, predicted secondary mass spectra and random forests generated based on metabolite structure encoding and ESI-MS / MS machine learning model algorithm Forest RT retention time predicted by the algorithm.
[0016] Preferably, in a possible implementation of the second aspect, the database content is stored by classification according to medical library, animal library, plant library, microbial library, intestinal flora library, traditional Chinese medicine library and exposure group library, and supports export in MSP file format.
[0017] The beneficial effects of the present invention are: integrating metabolite information from standard libraries, commercial databases, and public databases, and introducing AI prediction technology, thereby expanding the capacity and function of the metabolomics MS / MS database. By being compatible with MS / MS spectral libraries from multiple sources, the present invention maximizes the use of data resources, and the constructed database covers more than 620,000 metabolites and more than 21 million MS / MS spectra. In addition, the ESI-MS / MS machine learning model algorithm and random forest algorithm are used to predict the secondary mass spectra and retention times of metabolites, achieving more than 95% of the three-dimensional qualitative identification of metabolites (MS1, MS2, RT), improving the accuracy and depth of metabolite identification.
[0018] At the same time, according to the different sources and characteristics of metabolites, the database is subdivided into seven classification libraries: medical library, animal library, plant library, microbial library, intestinal flora library, traditional Chinese medicine library and exposure group library. At the same time, the present invention is also compatible with the original data analysis of almost all mainstream LCMS high-resolution mass spectrometry platforms, clearing the obstacles for data exchange and sharing between different platforms, and further promoting the in-depth development of metabolomics research. In summary, the present invention has shown excellent performance in improving identification accuracy, expanding database capacity and promoting cross-platform data exchange, providing strong support for metabolomics research. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 A flow chart of a method for constructing a metabolomics MS / MS database based on liquid chromatography-mass spectrometry is provided for this application.
[0021] Figure 2 An example diagram of the contents of an MSP file is provided for this application.
[0022] Figure 3 A qualitative result diagram of experimental map matching of an intestinal flora library, a medical library, and a plant library with an AI prediction library is provided for this application.
[0023] Figure 4 A RT prediction result diagram is provided for this application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] Embodiment 1: Figure 1 As shown, the present invention provides a method for constructing a metabolomics MS / MS database based on liquid chromatography-mass spectrometry, the method comprising: Obtain mass spectrometry and chromatographic information from standard libraries and metabolite information from commercial and public databases.
[0026] Specifically, the method for obtaining metabolite information includes: extracting secondary mass spectrometry data from commercial databases and public databases and formatting them into MSP files, wherein the content format of the MSP files is a unified preset format; removing substances with molecular weights outside a preset range and removing information outside the scope of metabolomics research; and performing redundancy processing on repeated information.
[0027] The content format of the MSP file is shown in Table 1. The content of the MSP file is as follows: Figure 2 The accompanying MSP format files facilitate the search software (such as Compound Discoverer 3.3, Progenesis QI, MS-DIAL, XCMS and other commercial and open source software) to identify and conduct qualitative analysis of metabolites.
[0028] Table 1 Content format of MSP file
[0029] The commercial database content extraction method includes: using mzVault to load the metabolite information file in the commercial database, which is a db file; exporting the db file as a UTF-8 encoded CSV file; extracting core information from the CSV file and formatting it into an MSP file. The core information includes CompoundName, ChemicalFormula, Ionization, ExtractedMass, Adduct, Polarity, Confirm Precursor, Confirm Energy, and the corresponding secondary mass spectrometry fragment ion mass-to-charge ratio and response intensity value.
[0030] The method for extracting content from a public database includes: loading a metabolite information file in a public database, which is a txt format file or an MSP file; extracting the metabolite information in the file and saving it in a CSV file; merging CSV files from public databases from different sources to construct an MSP file.
[0031] In this embodiment, the metabolomics MS / MS database is composed of four major types of databases, and the specific composition of each type of database is as follows: The standard product library is divided into two parts. One is the self-purchased standard products, and the mass spectrometry information of each standard product is collected. The other is the standard product spectrum information contained in the commercial database mzCloud. This part is mainly integrated by Thermo official from different laboratories based on Orbitrap high-resolution mass spectrometry. The commercial database, in addition to the mzCloud commercial database, also integrates the NIST_2020_MS / MS_HR.db database, which was developed by the National Institute of Standards and Technology of the United States; the public database integrates the MS / MS spectra of the most commonly used databases such as MassBank, HMDB, and MS-DIAL; the AI prediction library collects and organizes 46 database resources, and obtains the InChI, SMILES and other structural coding information corresponding to the metabolites. Based on the ESI-MS / MS machine learning model and the random forest prediction algorithm, it realizes the prediction from the metabolite SMILES coding to the MS / MS spectrum and RT.
[0032] The commercial databases included in the metabolomics MS / MS database are mainly mzCloud and NIST_2020_MS / MS. Both databases are files in db format, but the content in the file cannot be used directly. Information extraction and conversion are required. The specific method is as follows: Install the software Thermo mzVault in the Windows 10 system; click "Browse", "Open", and select the db database file to be opened; after loading, select all entries in the database, click "Built", "Export", and choose to export as a "CSV" file. The encoding must be UTF-8 to avoid garbled characters caused by special symbols in the metabolite name; the CSV file contains the core information in the db database: CompoundName, ChemicalFormula, Ionization, ExtractedMass, Adduct, Polarity, Confirm Precursor, Confirm Energy, and the corresponding secondary mass spectrometry fragment ion mass-to-charge ratio and response intensity value; write R language code to batch extract and organize the information in the CSV file according to the pre-prepared MSP file format to generate the final MSP file containing all metabolites and secondary mass spectrometry information.
[0033] Public databases used by metabolomics MS / MS databases, such as HMDB, MassBank, MoNA, and MS-DIAL, all contain a large amount of experimental metabolite secondary mass spectrometry information. The information provided in some databases is in txt format or MSP format. In order to standardize and unify, the information is extracted, filtered, and processed. According to the MSP file content format, the information is reorganized and the corresponding information tags are added so that the information can be traced and verified when the library is used. The specific operations are as follows: download the secondary mass spectrometry information from the database, txt or msp file; extract the metabolite information in the file, save it in a CSV file, and perform automatic code verification and manual correction based on information such as compound name and molecular formula to ensure the accuracy of the information; merge and remove redundancy from information from databases from different sources, and finally construct an MSP file.
[0034] Next, substances with molecular weights less than 50 and greater than 2000 were eliminated, and information outside the scope of metabolomics research, such as inorganic compounds and single elements, was eliminated; finally, missing information such as missing molecular formula, molecular weight, source, and ion mode were supplemented.
[0035] Based on the metabolite structure coding, the ESI-MS / MS machine learning model is used to predict and generate secondary mass spectra, and the random forest algorithm is used to predict the retention time of metabolites to form an AI prediction library.
[0036] Specifically, the AI prediction library construction method includes: obtaining the InChI and SMILES structural codes corresponding to the metabolites based on the metabolite information in the standard library, commercial database and public database; inputting the InChI or SMILES structural code into the ESI-MS / MS machine learning model to predict the secondary mass spectrum, and outputting three predicted MS / MS fragmentation information corresponding to the positive and negative energies of 10ev, 20ev and 30ev for each metabolite, and organizing the MS / MS fragmentation information into an MSP file; using the random forest algorithm to predict the RT retention time of the metabolites, using the SMILES and RT information based on the self-built library standard to build a local model, inputting the SMILES structural coding information of the metabolite to be predicted into the model, and outputting the predicted RT retention time of the metabolite.
[0037] In this embodiment, according to the metabolite information integrated in the metabolomics MS / MS database, the InChI and SMILES structural codes corresponding to the metabolites are obtained; the InChI or SMILES structural codes are input into the ESI-MS / MS machine learning model to predict the spectrum; for large-scale calculations, nearly 620,000 metabolite coding information are delivered to a large-scale cluster server in batches of 100 tasks per task for each 1,000 tasks for calculation; each metabolite will output three predicted MS / MS fragmentation information corresponding to the positive and negative energies of 10ev, 20ev, and 30ev respectively; the predicted MS / MS information is sorted according to the MSP file format designed above; the RT retention time of all metabolites is realized by using the random forest prediction algorithm, a local model is constructed based on the SMILES and RT information of the self-built library standard, and the SMILES structural coding information of the metabolite to be predicted is input into the model, thereby realizing the predicted RT information of the target metabolite.
[0038] A metabolomics MS / MS database was constructed based on standard library, commercial database, public database and AI prediction library.
[0039] Data normalization and classification of metabolomics MS / MS database.
[0040] Specifically, the key information of metabolomics MS / MS database data normalization includes metabolite basic information, database ID, structure code, metabolite classification, metabolite spectrum and metabolite source. Metabolomics MS / MS database classification includes medical library, animal library, plant library, microbial library, intestinal flora library, traditional Chinese medicine library and exposure group library.
[0041] In this embodiment, the metabolomics MS / MS database integrates libraries from 46 sources, and the metabolite information of different libraries or websites is not the same, and the information completeness varies greatly. In order to better manage and use the metabolites in the metabolomics MS / MS database, the information of the metabolites entered into the database is standardized, including 29 key information such as metabolite basic information, database ID, structure coding, compound classification, and source. The integration of this information can not only trace all relevant information of the metabolites, but also realize the ID conversion between multiple databases, breaking the information barriers between different compound libraries or websites.
[0042] The entries for metabolomics MS / MS database information are shown in Table 2.
[0043] Table 2 Metabolomics MS / MS database information entry items
[0044] Regarding the classification of metabolomics MS / MS databases, the biggest difference between metabolites and proteins and genes is that it is difficult for metabolites to have their species affiliation. There will be a large number of metabolites that exist in different species at the same time, rather than playing similar functions in different species in the form of sequence proximity or structural similarity. Only a small number of databases will provide information on the species origin and affiliation of metabolites. There are also a small number of databases that are specially constructed based on species, such as HMDB, which is a human metabolite database, YMDB, which is a yeast metabolome database, and Plantcyc, which is a plant metabolic database. Other comprehensive databases, such as NPASS, COCONUT, etc., provide specific corresponding species information. Based on this existing information, metabolites are divided into seven categories: medicine, animals, plants, microorganisms, intestinal flora, traditional Chinese medicine, and exposure groups. The specific steps are as follows: First, we organize the databases from a single source. For example, the metabolites in HMDB will be labeled with the medical library and animal library, the metabolites in YMDB will be labeled with the microbial library, and the metabolites in the Plantcyc library will be labeled with the plant library. Based on this rule, the metabolites in the single species library will be labeled with the corresponding species category.
[0045] The comprehensive species origin database is used to assign species attributes based on the species information provided. Taking the COCONUT library as an example, the database provides species Organisms information corresponding to each metabolite. For example, the metabolite L-homoserine provides 22 Organisms information in the library. The corresponding species attribution information is found based on these Organisms to determine which library the metabolite belongs to.
[0046] The metabolites of the intestinal flora library are collated. There are special libraries for intestinal flora, such as Fecal, gutMGene, etc. The metabolites in the library are labeled as intestinal flora libraries; the Chinese medicine library TCM is collated according to the information in 7 common Chinese medicine websites, and the metabolites in them are classified into the Chinese medicine library; the exposure group library, multiple metabolites contain this type of substance, such as HMDB, mzCloud and other libraries, and there will be corresponding classifications, such as pesticides, cosmetics, artificial products, etc. will be classified into the exposure group library.
[0047] In one embodiment, the method further comprises metabolomics MS / MS database integration, validation and testing.
[0048] Specifically, the metabolomics MS / MS database is a super-large metabolomics database with more than 620,000 metabolites and 21 million secondary spectra. The database contains the following: a master table containing all the information of 620,000 metabolites, each of which consists of 29 dimensions of information; MSP files of seven sub-libraries, each of which consists of MSP files in positive and negative modes, and contains a total of 21 million MS / MS spectra information. The number of metabolites in the seven sub-libraries of the metabolomics MS / MS database is shown in Table 3.
[0049] Table 3 Number of metabolites in the seven sub-databases of the metabolomics MS / MS database
[0050] like Figure 3 As shown in the figure, in order to verify the performance of the metabolomics MS / MS database, the project data analyzed by the intestinal flora library, medical library and plant library using the metabolomics MS / MS database were selected for statistical analysis, and the qualitative results of the experimental spectrum matching (level AB(i)) and the AI prediction library (B(ii)) were compared. Among the three major libraries, 62.2%, 60.8% and 61.2% of the identification results were verified in the qualitative results of the standard library, respectively, indicating that the AI library has a higher accuracy; at the same time, compared with the standard library, the application of the AI library can increase the number of identifications by about 30%-38%, which is a good supplement for the shortcomings of the standard library.
[0051] like Figure 4 As shown in the figure, in order to verify the reliability of the RT model, 100 standard samples were randomly selected from the self-built library for verification. The accuracy of RT prediction reached about 94%, and the average error was about 11.9s. With such performance, RT filtering will greatly reduce the probability of false positives and further improve qualitative reliability.
[0052] Embodiment 2: The present invention provides a metabolomics MS / MS database based on liquid chromatography-mass spectrometry, the database comprising: The standard library includes the mass spectrometry information of self-purchased standards and the spectral information of standards integrated from mzCloud; the commercial database includes the integrated NIST_2020_MS / MS_HR.db file; the public database includes the secondary mass spectrometry information of MassBank, HMDB, and MS-DIAL; the AI prediction library includes the predicted secondary mass spectra generated by the metabolite structure encoding and ESI-MS / MS machine learning model algorithm and the RT retention time predicted by the random forest algorithm.
[0053] The database content is stored by category: medical library, animal library, plant library, microbial library, intestinal flora library, traditional Chinese medicine library and exposure group library, and supports export in MSP file format.
[0054] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for constructing a metabolomics MS / MS database based on liquid chromatography-mass spectrometry, characterized in that: The method comprises: Obtain mass spectrum information and chromatographic information from standard libraries and metabolite information from commercial and public databases; Based on the metabolite structure coding, an ESI-MS / MS machine learning model was constructed to predict and generate secondary mass spectra. Forest The algorithm predicts the retention time of metabolites and forms an AI prediction library; Construct a metabolomics MS / MS database based on standard library, commercial database, public database and AI prediction library; Data normalization and classification of metabolomics MS / MS database.
2. The method for constructing a metabolomics MS / MS database according to claim 1, characterized in that: The metabolite information acquisition method comprises: Extracting secondary mass spectrometry data from commercial databases and public databases and formatting them into MSP files, wherein the content format of the MSP files is a unified preset format; Remove substances with molecular weights outside the preset range and remove information outside the scope of metabolomics research; De-redundancy processing is performed on repeated information.
3. The method for constructing a metabolomics MS / MS database according to claim 2, characterized in that: The commercial database content extraction method comprises: Use mzVault to load the metabolite information file in the commercial database, which is a db file; Export the db file as a CSV file encoded in UTF-8; Extract the core information from the CSV file and format it into an MSP file.
4. The method for constructing a metabolomics MS / MS database according to claim 3, characterized in that: The core information includes CompoundName, ChemicalFormula, Ionization, ExtractedMass, Adduct, Polarity, ConfirmPrecursor, ConfirmEnergy and the corresponding secondary mass spectrometry fragment ion mass-to-charge ratio and response intensity value.
5. The method for constructing a metabolomics MS / MS database according to claim 2, characterized in that: The public database content extraction method comprises: Load the metabolite information file in the public database, which is a txt format file or an MSP file; Extract the metabolite information in the file and save it in a CSV file; Merge CSV files from public databases from different sources to build an MSP file.
6. The method for constructing a metabolomics MS / MS database according to claim 5, characterized in that: The AI prediction library construction method includes: Obtain the InChI and SMILES structural codes corresponding to metabolites based on the metabolite information in the standard library, commercial database and public database; Input the InChI or SMILES structure code into the ESI-MS / MS machine learning model to predict the secondary mass spectrum. Output three predicted MS / MS fragmentation information corresponding to the positive and negative energies of 10ev, 20ev, and 30ev for each metabolite, and organize the MS / MS fragmentation information into an MSP file. Random Forest Forest The algorithm predicts the RT retention time of metabolites, builds a local model using SMILES and RT information based on the self-built library standard, inputs the SMILES structure encoding information of the metabolite to be predicted into the model, and outputs the predicted RT retention time of the metabolite.
7. The method for constructing a metabolomics MS / MS database according to claim 1, characterized in that: The metabolomics MS / MS database data normalization and collation key information includes metabolite basic information, database ID, structure code, metabolite classification, metabolite spectrum and metabolite source.
8. The method for constructing a metabolomics MS / MS database according to claim 1, characterized in that: The metabolomics MS / MS database classification includes a medical library, an animal library, a plant library, a microbial library, a gut flora library, a traditional Chinese medicine library, and an exposure group library.
9. A metabolomics MS / MS database based on LC-MS, characterized in that: The database includes: Standard product library, including mass spectrometry information of self-purchased standards and standard product spectrum information integrated from mzCloud; Commercial databases, including the integrated NIST_2020_MS / MS_HR.db file; Public databases, including secondary mass spectrometry information from MassBank, HMDB, and MS-DIAL; AI prediction library, predicted secondary mass spectra and random forests generated based on metabolite structure encoding and ESI-MS / MS machine learning model algorithm Forest RT retention time predicted by the algorithm.
10. The metabolomics MS / MS database according to claim 9, characterized in that: The database content is stored by classification according to medical library, animal library, plant library, microbial library, intestinal flora library, traditional Chinese medicine library and exposure group library, and supports export in MSP file format.
Citation Information
Patent Citations
Multi-species GC-MS endogenous metabolite database and establishment method thereof
CN111402961A
Method for establishing metabolite model and metabonomics database thereof
CN114283877A
Large-scale metabolome qualitative method based on molecular structure association network
CN114609318A
Construction method and application of intestinal microorganism related metabolite spectrum database
CN115083528A
High-throughput traditional Chinese medicine and metabolite database LutMet-TCM establishment method and application thereof
CN116153423A
Cited By
Construction method and device of non-target public spectrogram database and electronic equipment
CN120673860A