Standard compound library MetaTag database and application thereof

By establishing a MetaTag database, which provides characteristic information of compounds collected experimentally, the problems of incomplete coverage and insufficient information in existing databases are solved, enabling efficient identification of metabolites and discovery of unknown metabolites, reducing costs and improving the coverage and usability of the database.

CN121862218APending Publication Date: 2026-04-14INSTITUTE OF BASIC MEDICAL SCIENCES CHINESE ACADEMY OF MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing compound databases are incomplete, lack chemical and biomedical information, and are not suitable for the direct identification of metabolites, resulting in low search matching efficiency and false positive and false negative results.

Method used

A MetaTag database was established, containing 43,120 compounds, each with an 'Xm-Yn' type chemical structure, including virtual and physical parts, providing reliable characterization information for experimental acquisition, including chromatographic and mass spectrometric data. The physical parts are available for acquisition in a series of mixed standards.

Benefits of technology

It enables high-confidence annotation of known metabolites and rapid discovery of unknown metabolites, requiring no standards or only low-cost mixed standards, covering a wide range of metabolites and pathways, and the database is freely available.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862218A_ABST
    Figure CN121862218A_ABST
Patent Text Reader

Abstract

The invention discloses a standard compound library MetaTag database and application thereof.43120 compounds are recorded in the MetaTag database, the compounds have Xm-Yn type chemical structural characteristics, X represents an electrophilic substituent, Y represents a nucleophilic substituent, m and n represent the number of X and the number of Y respectively, the character '-' represents that X and Y form a covalent bond, and X and Y are both from endogenous metabolites; each compound is composed of a virtual part and an entity part; in the virtual part, each compound corresponds to a unique entry, and each entry comprises but is not limited to sub-entries such as a chemical information class, a biomedical information class, a related literature or an external database link; in the entity part, each compound can find a molecular entity in a mixed standard series of a MetaTag database, and the molecular entity is used for collecting LC-MS data to confirm the metabolite structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cheminformatics and bioinformatics, and more specifically to a standard compound library MetaTag database and its applications. Background Technology

[0002] Database technology stores the characteristic information of compounds in a structured manner, providing search and matching functions to help identify metabolites and elucidate their biomedical functions. Major databases related to metabolism include: Kyoto Encyclopedia of Genes and Genomes (KEGG), HMDB, GNPS, DeepMet, MetaboLights, MetaCyc, LipidMaps, PubChem, and ChEMBL.

[0003] Chromatography-mass spectrometry (LC-MS) is currently the mainstream technique for metabolite detection and analysis, offering advantages such as high sensitivity, high throughput, and applicability to complex samples. This technique can rapidly acquire multi-dimensional characteristic information of compounds (including but not limited to retention time, primary / secondary mass spectrometry, and ion mobility) for database construction.

[0004] Currently, the existing database still has the following problems: Problem 1: Existing databases do not fully cover all metabolites. On the one hand, there have been continuous reports of new metabolites not previously included in databases in recent years; on the other hand, a large number of non-redundant features still exist in the massive LC-MS data, making it impossible to find matching metabolites in existing databases.

[0005] Question 2: The database lacks necessary chemical and biomedical information. For example, KEGG and MetaCyc contain metabolite names, chemical structures, and metabolic pathways, but lack analytical chemical information such as chromatographic retention times and mass spectra, making them unusable for metabolite identification. Most metabolites in HMDB lack chromatographic characterization information, and a small number lack experimentally measured mass spectrometric characterization information, only providing predicted mass spectra.

[0006] Question 3: The databases are not dedicated to metabolism. The PubChem database contains over 100 million compound entries, and the ChEMBL database contains approximately 2.9 million compound entries. The vast majority of compound entries in these two databases are unrelated to metabolites, resulting in low efficiency in searching and matching metabolites.

[0007] Question 4: None of the databases directly provide the molecular entities of metabolites. In practical applications, the data acquisition conditions of the samples to be analyzed are difficult to maintain in consistency with the database's acquisition conditions, which may lead to false positives and false negatives in metabolite identification results. Although some databases provide commercial purchase links for metabolite standards, the cost of purchasing metabolite standards is high and not suitable for large-scale metabolite detection. For unknown metabolites, it is not even possible to purchase standards with known chemical structures. Summary of the Invention

[0008] In view of this, the present invention provides a standard compound library, the MetaTag database, and its applications. The MetaTag database provides free, reliable compound characteristic information collected experimentally. It can achieve high-confidence annotation of known metabolites and rapid discovery of unknown metabolites using conventional high-performance liquid chromatography-chromatography (HPLC) techniques, without the need for standards or only requiring a small amount of low-cost, readily available mixed standards.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: A standard compound library MetaTag database, wherein the MetaTag database contains 43,120 compounds, the compounds having the prefix "X". m -Y n The chemical structure features a “-” type, where X represents an electrophilic substituent, Y represents a nucleophilic substituent, m and n represent the number of X and Y respectively, and the hyphen “-” indicates that X and Y form a covalent bond. Both X and Y are derived from endogenous metabolites. Each compound consists of a virtual part and a physical part; In the virtual section, each compound corresponds to a unique entry, and each entry includes, but is not limited to, sub-entries such as chemical information, biomedical information, relevant literature, or external database links; In the entity section, each compound can be found as a molecular entity within a mixed series of standards in the MetaTag database, which is used to acquire LC-MS data to confirm the metabolite structure.

[0010] Preferred chemical information sub-items include: unique identifier, IUPAC name, common name, synonym, chemical formula, two-dimensional / three-dimensional chemical structure, SMILES string, InChI string, InChIKey string, precise molecular weight, mass-to-charge ratio, structure-related metabolites, information on the mixed standard to which the molecular entity belongs, chromatographic retention time Rt, chromatographic data acquisition parameters, primary MS / secondary MS2 mass spectra, collision cross section CCS, and mass spectrometry data acquisition parameters; Preferably, the biomedical information sub-entries include: biodistribution of compounds under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; biodistribution of structure-related metabolites under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; net release and net intake of compounds under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; net release and net intake of structure-related metabolites under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; flux of compounds under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; flux of structure-related metabolites under standard reference conditions, pathological conditions, physiological conditions, or artificial interference conditions; information on enzymatic and non-enzymatic reactions using compounds as products or substrates; and information on enzymatic and non-enzymatic reactions using structure-related metabolites as products or substrates. Biodistribution refers to the absolute concentration or relative abundance of compounds in various locations within a biological population or individual, including populations, individuals, organs, tissues, cells, and organelles. Net release and net intake are respectively: the ratio or difference between the outflow and inflow of metabolites in a single organ or tissue at a certain moment or within a certain period of time. A ratio greater than 1 or a difference greater than 0 indicates net release, a ratio less than 1 or a difference less than 0 indicates net intake, and a ratio equal to 1 or a difference equal to 0 indicates neither net release nor net intake. Flux is the total amount of metabolites that flow through an organ or tissue at a certain moment or within a certain period of time.

[0011] Preferably, X is the atomic group remaining after removing one hydroxyl group from the following metabolites: glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, histidine, tyrosine, tryptophan, glutamine, glutamic acid, asparagine, aspartic acid, serine, threonine, cysteine, lysine, arginine, methionine, cysteine, 5-hydroxytryptophan, γ-aminobutyric acid (GABA), dopa, para-aminobenzoic acid, 2-aminobutyric acid, citrulline, ornithine, creatine, 3-aminobutyric acid, theanine, sarcosine, kynurenine, β-alanine, lactic acid, pyruvic acid, citric acid, isocitrate, cis-aconitine, α-ketoglutarate, oxaloacetate. Acids, succinic acid, malic acid, fumaric acid, acetic acid, cholic acid, dehydrocholic acid, deoxycholic acid, folic acid, biotin, niacin, malonic acid, crotonic acid, itaconic acid, acetoacetic acid, α-ketoisovaleric acid, glycolic acid, glyceric acid, α-hydroxybutyric acid, β-hydroxybutyric acid, 3-hydroxyisobutyric acid, retinoic acid, quinoline acid, 4-pyridoxine, pantothenic acid, 4-hydroxyphenylacetic acid, kynurenic acid, 5-hydroxyindole-3-acetic acid, urocanic acid, pyroglutamic acid, propionic acid, butyric acid, valeric acid, hexanoic acid, decanoic acid, lauric acid, myristic acid, palmitic acid, palmitoleic acid, stearic acid, oleic acid, arachidonic acid, sialic acid, gluconic acid, glucuronic acid, bilirubin, carnitine, methanol, ethanol, propanol; The Y group refers to the atomic group remaining after removing one hydrogen atom from the heteroatom of the following metabolites: glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, histidine, tyrosine, tryptophan, glutamine, glutamic acid, asparagine, aspartic acid, serine, threonine, cysteine, lysine, arginine, methionine, cysteine, 5-hydroxytryptophan, GABA, dopa, para-aminobenzoic acid, 2-aminobutyric acid, citrulline, ornithine, creatine, 3-aminobutyric acid, theanine, sarcosine, kynurenine, β-alanine, guanosine, cytidine, adenosine, thymidine, uridine, inosine, guanine, cytosine, adenine, thymine, uracil, dimethylamine, dopamine, tyramine, histamine, and chromoside. Amines, 5-hydroxytryptamine, 3-methoxytyramine, norepinephrine, methylamine, ethanolamine, cysteine, taurine, taurine, putrescine, 1,3-propanediamine, guanidine, spermidine, spermine, glucosamine, galactosamine, pyridoxine, sphingosine, glucose, fructose, galactose, mannose, fucose, ribose, 2-deoxyribose, xylose, inositol, glycerol, galactitol, lactic acid, citric acid, isocitric acid, malic acid, cholic acid, folic acid, glycolic acid, glyceric acid, α-hydroxybutyric acid, β-hydroxybutyric acid, 3-hydroxyisobutyric acid, 4-pyridoxine, 4-hydroxyphenylacetic acid, kynurenic acid, 5-hydroxyindole-3-acetic acid, urocanic acid, pyroglutamic acid, sialic acid, gluconic acid, glucuronic acid, bilirubin, carnitine.

[0012] Preferably, the molecular entity of the compound included in the MetaTag database exists in at least one of the mixed standards in the mixed standard series; The mixed standard series consists of no fewer than 10 mixed standards, each containing one or more of the included compounds; The mixed standard is in the form of an acidic, alkaline, or neutral aqueous solution or an organic-aqueous mixture, and / or a solid in the form of powder, granules, or lumps without solvent, and / or a paste or suspension containing a small amount of solvent; For aqueous solutions or organic-aqueous mixtures, the concentration of each compound should be in the range of 10 nM to 100 μM. For pastes or suspensions, the apparent concentration of each compound should be in the range of 100 nM to 1 mM. These solutions should be diluted 10-1000 times with a diluent before loading.

[0013] Preferably, the structure-associated metabolite is a metabolite within the coverage of X and Y before removing a hydroxyl group (-OH) or a hydrogen atom (-H).

[0014] Another objective of this invention is to provide the aforementioned standard compound library, MetaTag database, for applications in annotating known metabolites, discovering new metabolites, identifying and characterizing drug metabolites, component analysis in nutrient mics and food science, and screening and validating disease biomarkers or physiological indicator metabolites.

[0015] As can be seen from the above technical solution, compared with the prior art, the present invention has the following technical effects: a. Reliability: The experimental data sub-entries in the MetaTag database are all derived from real experimental data collection, not computer predictions. All included compounds have corresponding molecular entities, enabling users to acquire customized analytical data. b. User-friendly: No standard products are required, or the cumbersome standard product synthesis process is eliminated, thereby reducing the cost of obtaining standard products and improving their availability; c. Diverse compound library: It covers a wide range of metabolites, metabolic pathways, metabolic organs or tissues, and includes compounds not included in other commonly used databases; d. Open source and open access: The use of the virtual part of the MetaTag database is completely free. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This is a diagram illustrating the overall framework of the MetaTag database of this invention. Figure 2 A heatmap showing the relative abundance of MetaTag database metabolites detected in rat organs or tissues; Figure 3 This is for the confirmation of propionyl taurine. The left image shows a chromatographic comparison of the sample (top) and the standard (bottom). The right image shows a comparison of the secondary mass spectra of the sample (top) and the standard (bottom). Detailed Implementation

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1 Taking Example 1 as an example, the standard compound library MetaTag database provided in this invention is established as follows: Enzymatic reaction data and serum metabolome concentration data from the KEGG database.

[0020] (1) Establishment of a virtual compound library A total of 4893 enzymatic reactions, primarily transferases and ligases, were extracted from the KEGG database. Reactions involving large molecules (proteins, nucleic acids, polysaccharides), non-human origins, or multiple steps were removed, leaving 1356 enzymatic reactions, defined as major metabolic reactions. Based on the type of chemical bonds formed, 1263 of these major metabolic reactions were further categorized into O-phosphorylation, N,O-acylation, N,O-alkylation, O-glycosylation, carbanion reactions, and sulfonation reactions. N-acylation and N-methylation reactions, which occurred more frequently, were prioritized for subsequent chemical reactions.

[0021] Based on the combined data of human and mouse serum metabolite concentrations, metabolites with serum concentrations not less than 10 μM were selected, and their suitability as N-acylation or N-methylation substrates was determined. The criteria were that metabolites containing carboxyl or amino groups could serve as N-acylation substrates, and those containing amino groups could serve as N-methylation substrates. Ultimately, 45 carboxyl-containing metabolites were identified and named Tag metabolites; 27 amino-containing metabolites were identified and named Meta metabolites. One metabolite from each of the Tag and Meta metabolites was randomly selected, and virtual N-acylation reactions were performed in pairs, generating 1215 possible products. Similarly, one metabolite from the Meta metabolites was randomly selected and virtual N-methylation reactions were performed, generating 27 possible products. The chemical structures of these 1242 products were searched in the Human Metabolomics Database (HMDB); 444 products were found, and 998 products were not found. A virtual compound library was established by compiling information on the chemical formula, precise molecular weight, chemical structure, and whether the product is included in HMDB for 1242 products.

[0022] (2) One-pot reaction array for efficient synthesis of virtual compound libraries For the 1215 potential N-acylation products in the virtual compound library, one or more of Schemes 1 to 6 were employed based on the chemical structural characteristics of the Meta and Tag metabolites. Schemes 1 to 4 involved carboxylic acid-amine condensation, Scheme 5 involved the anhydride method, and Scheme 6 involved the acyl chloride method. For the 27 potential N-methylation products in the library, Scheme 7 was employed.

[0023] Option 1: Dissolve Tag metabolite (0.1 mmol) in DMF (200 uL), add DIPEA (0.2 mmol) and HATU (0.12 mmol) in sequence, stir evenly at 0 °C for 5 min, add Meta metabolite (0.11 mmol), stir at 25 °C for 60 min and then stop the reaction.

[0024] Option 2: Dissolve the N-Boc protected Tag metabolite (0.1 mmol) in DMF (200 μL), then add DIPEA (0.2 mmol) and HATU (0.12 mmol) sequentially. Stir at 0 °C for 5 min, then add the Meta metabolite (0.11 mmol) and stir at 25 °C for 60 min. Reduce the reaction solution under reduced pressure, add 1 M HCl dioxane solution, and react at 0 °C for 15 min to remove the Boc protecting group, then terminate the reaction.

[0025] Option 3: Dissolve the Tag metabolite (0.1 mmol) in DMF (200 μL), then add DIPEA (0.2 mmol) and HATU (0.12 mmol) sequentially. Stir at 0 °C for 5 min, then add the O-Boc protected metabolite (0.11 mmol) and stir at 25 °C for 60 min. After evaporating the reaction solution under reduced pressure, add 1 M HCl dioxane solution and react at 0 °C for 15 min to remove the Boc protecting group.

[0026] Option 4: Dissolve the N-Boc protected Tag metabolite (0.1 mmol) in DMF (200 μL), then add DIPEA (0.2 mmol) and HATU (0.12 mmol) sequentially. Stir at 0 °C for 5 min, then add the O-Boc protected Metabolite (0.11 mmol) and stir at 25 °C for 60 min. After evaporating the reaction solution under reduced pressure, add 1 M HCl dioxane solution and react at 0 °C for 15 min to remove the Boc protecting group.

[0027] Option 5: Dissolve the acid anhydride (0.1 mmol) of Tag metabolite in DMF (200 uL), add DIPEA (0.2 mmol), stir evenly at 0 °C for 5 min, add Meta metabolite (0.11 mmol), stir at 25 °C for 60 min, and then terminate the reaction.

[0028] Option 6: Dissolve the acyl chloride (0.1 mmol) of Tag metabolite in acetonitrile (200 uL), add DIPEA (0.2 mmol), stir evenly at 0 °C for 5 min, add Meta metabolite (0.11 mmol), stir at 25 °C for 60 min, and then terminate the reaction.

[0029] Option 7: Dissolve the metabolite (0.2 mmol) in acetonitrile (200 uL), slowly add iodomethane (0.1 mmol) dropwise while stirring at 0 °C, and terminate the reaction after stirring at 25 °C for 60 min.

[0030] Virtual compounds that use the same synthesis scheme or produce the same reaction products are grouped together in the same reaction vessel. Each reaction vessel carries the synthesis reaction of 20-100 virtual compounds, and nearly 30 reaction vessels form a reaction array.

[0031] (3) Acquisition of chromatographic and mass spectrometric characteristics For each reaction vessel in the reaction array, the solvent inside the vessel was evaporated under reduced pressure to obtain the crude product. The crude product was redissolved in 1-5 mL of extraction solvent (acetonitrile-water or methanol-water solution), vortexed thoroughly, and allowed to stand for 10-60 min. 100 μL of the extract was centrifuged at high speed (16000 × g), and the supernatant was diluted 10-1000 times with the aforementioned extraction solvent and loaded onto a liquid chromatography-mass spectrometer (LC-MS). Using HILIC hydrophilic interaction chromatography and Orbitrap Exploris 480 mass spectrometry as examples, the instrument parameters were set as follows: Primary mass spectrometry scan range 70-1000, frequency 1 Hz, resolution 140000. Secondary mass spectrometry collision fragmentation energies were 10 eV, 20 eV, and 40 eV, with a resolution of 70000. Ionization modes were positive and negative ion modes. The chromatographic column used was an XBridge BEHAmide column (2.1 mm × 150 mm, particle size 2.5 μm, pore size 130 Å). Mobile phase A was an acetonitrile-water solution (95:5, pH 9.45) supplemented with ammonium acetate (20 mM) and ammonia (20 mM), and mobile phase B was acetonitrile. The flow rate was 0.15 ml / min. The mobile phase gradient was as follows: 0 min, 85% B; 2 min, 85% B; 3 min, 80% B; 5 min, 80% B; 6 min, 75% B; 7 min, 75% B; 8 min, 70% B; 9 min, 70% B; 10 min, 50% B; 12 min, 50% B; 13 min, 25% B; 16 min, 25% B; 18 min, 0% B; 23 min, 0% B; 24 min, 85% B; 30 min, 85% B. The autosampler temperature was 4 °C. The injection volume was 5 μL.

[0032] Primary mass spectrometry data for each sample is acquired. The raw files are converted to mzXML files using MSConvert, and the mzXML files are then read using the open-source El-MAVEN software. Chromatographic peaks are extracted manually or automatically based on the chemical formulas and molecular weight information of the virtual compound library. For detected virtual compounds, samples are loaded again, and secondary mass spectra are acquired. The secondary mass spectra are then manually or algorithmically determined to match the corresponding chemical structure. The criterion is that if fragment ions of metabolites or tag metabolites appear in the secondary mass spectrum, and the chromatographic peaks of these fragment ions have the same retention time and peak shape as the parent ion, then the secondary mass spectrum matches the chemical structure of the corresponding virtual compound, and the virtual compound is successfully synthesized. Furthermore, tertiary and even multi-level mass spectra of these virtual compounds can be acquired, along with drift time, collision cross-section, ion mobility, etc.

[0033] (4) Establishment of MetaTag v1.0, a standard compound library related to metabolites 1070 virtual compounds were confirmed by (3). The chromatograms, secondary or multi-stage mass spectra, retention times, drift times, collision cross sections, ion mobility and other chromatographic and mass spectrometric features of these compounds were entered into the database and associated with the chemical formulas, precise molecular weights and chemical structures that originally existed in the virtual compound library, thereby establishing a metabolite-related standard compound library MetaTag v1.0 containing both theoretical and actual data.

[0034] Example 2 Annotating and discovering metabolites in rat organs or tissues using the MetaTag database: Rat tissue sampling: Wild-type Sprague-Dawley rats (8-12 weeks old, male) were housed in a 25 ℃, 12h light-dark cycle environment, with free access to water and standard feed. Rats were fasted for 8 hours before the experiment. They were then anesthetized with high-concentration isoflurane in an induction chamber, and blood was collected from the tail tip. The rats were then euthanized quickly by injecting air directly into their hearts. Twenty-five organs and tissues were collected in sequence, including the skin, spleen, pancreas, kidney cortex, kidney medulla, liver, heart, lung, white adipose tissue (WAT), testis, white muscle (WM), red muscle (RM), bone marrow (BM), brown adipose tissue (BAT), eyes, cerebellum (Cb), brainstem (Bs), cerebral cortex (Cx), hippocampus (Hp), hypothalamus (Ht), duodenum, small intestine (SI), feces, large intestine (LI), and rectum. The extracted organs or tissues were rapidly transferred to liquid nitrogen for freezing, crushed, and weighed to approximately 20 mg. Metabolites were extracted from the tissues using 40 times their volume of a methanol-acetonitrile-water (4:4:2) mixture. The tissues were centrifuged at high speed (16000 × g, 4 °C), and the supernatant was collected and loaded onto LC-MS for chromatographic and primary mass spectrometry data acquisition.

[0035] Metabolite Annotation and Discovery: Chromatographic and primary mass spectrometry (PMS) data of compounds indexed in the MetaTag database were matched with the chromatographic and PMS data of biological samples. The matching criteria were as follows: the allowable error for retention time was ±0.1 min, and the allowable error for the mass-to-charge ratio of the precursor ion was no greater than 10 ppm. When both criteria were met, the chromatographic peak in the biological sample was labeled as a successfully matched compound. A total of 398 compounds were ultimately matched from the MetaTag database, such as... Figure 2As shown, their biological distribution in various organs and tissues exhibits clear regionality, with metabolite distributions being more similar in anatomically close tissues, such as the five regions of the brain and the two regions of the kidney. Among the successfully matched metabolites, 153 metabolites are included in HMDB and annotated as known metabolites; 245 metabolites are not included in HMDB and are annotated as newly discovered metabolites.

[0036] New metabolite confirmation: Taking propionyl taurine as an example among 245 metabolites. The WAT sample (400 μL) containing propionyl taurine was vacuum-dried, reconstituted with 40 μL of extraction solvent, centrifuged at high speed (16000 × g, 4 ℃), and the supernatant was loaded onto LC-MS for secondary mass spectra. Simultaneously, a 10 μM aqueous solution of pure propionyl taurine was prepared and loaded onto LC-MS with the same chromatographic and mass spectrometric parameters for secondary mass spectra. The chromatograms and secondary mass spectra of the WAT sample were compared with those of the standard. Figure 3 The retention times of the two are basically the same, and the cosine similarity of the secondary mass spectra is 0.98, which proves that the metabolite detected in the biological sample is propionyl taurine.

[0037] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0038] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A standard compound library MetaTag database, characterized in that, The MetaTag database contains 43,120 compounds, which have the prefix "X". m -Y n The chemical structure features a “-” type, where X represents an electrophilic substituent, Y represents a nucleophilic substituent, m and n represent the number of X and Y respectively, and the hyphen “-” indicates that X and Y form a covalent bond. Both X and Y are derived from endogenous metabolites. Each compound consists of a virtual part and a physical part; In the virtual section, each compound corresponds to a unique entry, and each entry includes, but is not limited to, chemical information, biomedical information, related literature, or external database link sub-entries; In the entity section, each compound can be found as a molecular entity within a mixed series of standards in the MetaTag database, which is used to acquire LC-MS data to confirm the metabolite structure.

2. The MetaTag database of a standard compound library according to claim 1, characterized in that, The chemical information sub-items include: unique identifier, IUPAC name, common name, synonym, chemical formula, two-dimensional / three-dimensional chemical structure, SMILES string, InChI string, InChIKey string, precise molecular weight, mass-to-charge ratio, structure-related metabolites, information on the mixed standards to which the molecular entity belongs, chromatographic retention time Rt, chromatographic data acquisition parameters, primary MS / secondary MS2 mass spectra, collision cross section CCS, and mass spectrometry data acquisition parameters; Here, a structure-related metabolite refers to a metabolite within the scope of (4) before the removal of a hydroxyl group (-OH) or a hydrogen atom (-H).

3. The MetaTag database of a standard compound library according to claim 1, characterized in that, The biomedical information category includes the following sub-entries: biodistribution of compounds under standard reference, pathological, physiological, or human-induced disturbance conditions; biodistribution of structure-related metabolites under standard reference, pathological, physiological, or human-induced disturbance conditions; net release and net intake of compounds under standard reference, pathological, physiological, or human-induced disturbance conditions; net release and net intake of structure-related metabolites under standard reference, pathological, physiological, or human-induced disturbance conditions; flux of compounds under standard reference, pathological, physiological, or human-induced disturbance conditions; flux of structure-related metabolites under standard reference, pathological, physiological, or human-induced disturbance conditions; information on enzymatic and non-enzymatic reactions using compounds as products or substrates; and information on enzymatic and non-enzymatic reactions using structure-related metabolites as products or substrates. Biodistribution refers to the absolute concentration or relative abundance of a compound in various locations within a biological population or individual, including populations, individuals, organs, tissues, cells, and organelles. Net release and net intake are respectively: the ratio or difference between the outflow and inflow of metabolites in a single organ or tissue at a certain moment or within a certain period of time. A ratio greater than 1 or a difference greater than 0 indicates net release, a ratio less than 1 or a difference less than 0 indicates net intake, and a ratio equal to 1 or a difference equal to 0 indicates neither net release nor net intake. Flux is the total amount of metabolites that flow through an organ or tissue at a certain moment or within a certain period of time.

4. The MetaTag database of a standard compound library according to claim 2, characterized in that, X is the radical remaining after removing one hydroxyl group from the following metabolites: glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, histidine, tyrosine, tryptophan, glutamine, glutamic acid, asparagine, aspartic acid, serine, threonine, cysteine, lysine, arginine, methionine, cysteine, 5-hydroxytryptophan, γ-aminobutyric acid, dopa, para-aminobenzoic acid, 2-aminobutyric acid, citrulline, ornithine, creatine, 3-aminobutyric acid, theanine, sarcosine, kynurenine, β-alanine, lactic acid, pyruvic acid, citric acid, isocitrate, cis-aconitine, α-ketoglutarate, oxaloacetic acid, succinic acid. Malic acid, fumaric acid, acetic acid, cholic acid, dehydrocholic acid, deoxycholic acid, folic acid, biotin, niacin, malonic acid, crotonic acid, itaconic acid, acetoacetic acid, α-ketoisovaleric acid, glycolic acid, glyceric acid, α-hydroxybutyric acid, β-hydroxybutyric acid, 3-hydroxyisobutyric acid, retinoic acid, quinoline acid, 4-pyridoxine, pantothenic acid, 4-hydroxyphenylacetic acid, kynurenic acid, 5-hydroxyindole-3-acetic acid, urocanic acid, pyroglutamic acid, propionic acid, butyric acid, valeric acid, hexanoic acid, decanoic acid, lauric acid, myristic acid, palmitic acid, palmitoleic acid, stearic acid, oleic acid, arachidonic acid, sialic acid, gluconic acid, glucuronic acid, bilirubin, carnitine, methanol, ethanol, propanol; The Y group refers to the atomic group remaining after removing one hydrogen atom from the heteroatom of the following metabolites: glycine, alanine, valine, leucine, isoleucine, proline, phenylalanine, histidine, tyrosine, tryptophan, glutamine, glutamic acid, asparagine, aspartic acid, serine, threonine, cysteine, lysine, arginine, methionine, cysteine, 5-hydroxytryptophan, γ-aminobutyric acid, dopa, para-aminobenzoic acid, 2-aminobutyric acid, citrulline, ornithine, creatine, 3-aminobutyric acid, theanine, sarcosine, kynurenine, β-alanine, guanosine, cytidine, adenosine, thymidine, uridine, inosine, guanine, cytosine, adenine, thymine, uracil, dimethylamine, dopamine, tyramine, histamine. Tryptamine, 5-hydroxytryptamine, 3-methoxytyramine, norepinephrine, methylamine, ethanolamine, cysteine, taurine, taurine, putrescine, 1,3-propanediamine, guanidine, spermidine, spermine, glucosamine, galactosamine, pyridoxine, sphingosine, glucose, fructose, galactose, mannose, fucose, ribose, 2-deoxyribose, xylose, inositol, glycerol, galactitol, lactic acid, citric acid, isocitric acid, malic acid, cholic acid, folic acid, glycolic acid, glyceric acid, α-hydroxybutyric acid, β-hydroxybutyric acid, 3-hydroxyisobutyric acid, 4-pyridoxine, 4-hydroxyphenylacetic acid, kynurenic acid, 5-hydroxyindole-3-acetic acid, urocanic acid, pyroglutamic acid, sialic acid, glucuronic acid, bilirubin, carnitine.

5. The MetaTag database of a standard compound library according to claim 1, characterized in that, The molecular entity of a compound included in the MetaTag database exists in at least one of the mixed standards in a series of mixed standards. The mixed standard series consists of no fewer than 10 mixed standards, each containing one or more of the included compounds; The mixed standard is in the form of an acidic, alkaline, or neutral aqueous solution or an organic-aqueous mixture, and / or a solid in the form of powder, granules, or lumps without solvent, and / or a paste or suspension containing a small amount of solvent; For aqueous solutions or organic-aqueous mixtures, the concentration of each compound should be in the range of 10 nM to 100 μM. For pastes or suspensions, the apparent concentration of each compound should be in the range of 100 nM to 1 mM. These solutions should be diluted 10-1000 times with a diluent before loading.

6. The MetaTag database of a standard compound library according to claim 4, characterized in that, The structure-associated metabolite is the metabolite within the coverage of X and Y before removing one hydroxyl group or one hydrogen atom.

7. The MetaTag database, a standard compound library as described in any one of claims 1-6, is used in annotating known metabolites, discovering new metabolites, identifying and characterizing drug metabolites, analyzing components in nutrient mics and food science, and screening and validating metabolites of disease biomarkers or physiological indicators.