Data processing device for metabolite analysis
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHIMADZU SEISAKUSHO LTD
- Filing Date
- 2023-10-02
- Publication Date
- 2026-07-17
AI Technical Summary
Current metabolite analysis methods using GC/MS and LC/MS are inefficient in analyzing thousands of metabolites simultaneously under the same conditions, leading to prolonged analysis times.
A metabolite analysis support system that includes an enzyme information storage unit, an analysis condition storage unit, a locus information input section, a metabolite specification section, and an analysis condition output section, which identifies metabolites affected by specific loci and outputs suitable analysis conditions for mass spectrometry.
This system enables rapid analysis of metabolites by targeting specific metabolites related to diseases or useful traits, allowing for mass analysis under optimized conditions, thereby speeding up the metabolite analysis process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a data processing device for metabolite analysis. [Background technology]
[0002] Metabolite analysis has been performed using gas chromatography mass spectrometry (GC / MS) and liquid chromatography mass spectrometry (LC / MS). Metabolism is a chemical reaction in the body of an organism. Among these, the metabolism related to the synthesis or conversion of polymeric compounds such as DNA, RNA, proteins, carbohydrates, and lipids, which are essential components for the maintenance of life, growth, and reproduction of organisms, and their constituent units (e.g., nucleic acids, amino acids, monosaccharides, sugar phosphates, fatty acids, organic acids, etc.) is called primary metabolism, and the other metabolisms are called secondary metabolism. Primary metabolism includes major metabolic pathways such as glycolysis, the TCA cycle (citric acid cycle), the electron transport system, and the pentose phosphate pathway, and the intermediate and final products in these metabolic pathways are called primary metabolites.
[0003] The method of comprehensively analyzing the types and concentrations of various metabolites contained in the body of an organism is called metabolomics, and it is expected to be applied to various fields. For example, in the medical field, it can be applied to the search for biomarkers for the early detection of diseases and the identification of causative substances of diseases, and in the food industry, it can be applied to quality evaluation and quality prediction, such as comparing products from different manufacturers or the origins of raw materials, or to the search for functional ingredients.
[0004] However, there are thousands of types of metabolites that are the subject of metabolomics, which is much more than the genes that are the subject of genomics or the proteins that are the subject of proomics. Although GC / MS and LC / MS can comprehensively analyze a large number of metabolites, it is not possible to analyze all metabolites simultaneously under the same conditions. In addition, there is a problem that the large number of metabolites makes it time-consuming to analyze the results of metabolite analysis. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2022-066655 Summary of the Invention [Problem to be solved by the invention]
[0006] The problem to be solved by the present invention is to expedite the analysis of metabolites contained in a sample using a mass spectrometer. [Means for solving the problem]
[0007] The metabolite analysis support system according to the present invention, which has been developed to solve the above problems, comprises: A system for supporting analysis of metabolites using a mass spectrometer, comprising: an enzyme information storage unit that stores information representing a relationship between a metabolite that is a substrate of one or more enzymes involved in a metabolic pathway and a metabolite that is a product of the one or more enzymes; an analysis condition storage unit that stores analysis conditions for analyzing metabolites by a mass spectrometer; a locus information input section for inputting information regarding a locus that affects expression of an enzyme of interest; a metabolite identification unit that identifies metabolites that are substrates and products of an enzyme affected by a gene locus based on information about the gene locus inputted into the gene locus input unit and information stored in the enzyme information storage unit; an analysis condition output unit that acquires from the analysis condition storage unit and outputs analysis conditions for analyzing the metabolites identified by the metabolite identification unit using a mass spectrometer; It is equipped with the following. Effect of the Invention
[0008] In the metabolite analysis support system according to the present invention, when information on a gene locus that affects the expression of an enzyme of interest to the user is input to the gene locus input unit, the substrate and product metabolites of the enzyme affected by the gene locus are identified based on the information and the information stored in the enzyme information storage unit, and analysis conditions for analyzing the metabolites using a mass spectrometer are output. Examples of enzymes of interest to the user include enzymes involved in the production of metabolites related to a specific disease or a specific constitution in humans, and enzymes involved in the production of metabolites related to useful traits in animals and plants. In other words, in the present invention, it is possible to narrow down the target metabolites that are the subject of metabolome-QTL analysis and perform mass analysis under analysis conditions suitable for the metabolites. Therefore, it is possible to speed up the work of analyzing metabolites contained in a sample collected from an organism using a mass spectrometer. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram of an embodiment in which the metabolite analysis support system according to the present invention is applied to a metabolite analysis system. [Diagram 2] FIG. 13 is a diagram showing an example of an analysis condition setting screen. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Many traits of various organisms, including humans, are determined by a combination of multiple genes (gene loci). Such traits are collectively called quantitative traits, and the gene loci that determine them are called quantitative trait loci (QTL). Some quantitative traits are influenced by multiple QTL. Therefore, if the QTL that affects an organism's useful quantitative traits is identified, it becomes possible to extract useful individuals using DNA markers that are highly related to the QTL. QTL analysis is the process of analyzing the QTLs that affect a certain quantitative trait, the effects of QTLs, and the number of QTLs from the DNA sequence of an organism.
[0011] In QTL analysis, next-generation sequencers, which can rapidly analyze genomic information of DNA fragments, have been used to identify QTLs that affect various traits. In analysis using next-generation sequencers, DNA is extracted from a sample taken from a target organism, and this DNA is fragmented into fragments with an appropriate narrow size distribution to prepare a library. This library is then subjected to the next-generation sequencer to sequence the base sequence of the original DNA.
[0012] The DNA base sequence (full sequence) determined by the next-generation sequencer can be analyzed using a homology search program called, for example, BLAST (Basic Local Alignment Search Tool). In an analysis using BLAST, the DNA base sequence of a target organism is compared with a database of base sequences of the same organism as the target organism to extract regions of homology, and the statistical significance (homology score) of the regions is calculated. Therefore, the user can identify regions of high and low homology from the size of the homology score.
[0013] A region with high homology is a region common to the biological species, and a region with low homology is a region that is heterogeneous (specific) in the target organism. Therefore, if the target organism has a characteristic trait not seen in the biological species, it can be presumed that the region with low homology is a region related to the characteristic trait. Therefore, by focusing on the base sequence of the low homology region and performing QTL analysis of the base sequence using an enzyme, which is a protein, as an indicator, it can be presumed that the QTL corresponding to the base sequence is a QTL that affects the expression of what kind of enzyme.
[0014] For example, the Reference Sequence (RefSeq) database, an open source database from the National Center for Biotechnology Information (NCBI), can be used for QTL analysis of a given region of a base sequence. The RefSeq database stores data on the genomic DNA, gene transcripts, and proteins produced from those transcripts of major organisms, from viruses to bacteria to eukaryotes. In other words, by using the RefSeq database to check the identification information (RefSeqID) of the QTL corresponding to the base sequence, information on the protein associated with the QTL can be obtained.
[0015] When the protein associated with a QTL is an enzyme in a metabolic pathway, it can be assumed that the QTL affects metabolites, which are the substrates and products of that enzyme. Therefore, in the present invention, metabolites contained in a sample collected from an organism are subjected to mass spectrometry, with the metabolites targeted.
[0016] Specifically, the present invention is a system for supporting the analysis of metabolites using a mass spectrometer, comprising an enzyme information memory unit that stores, for one or more enzymes involved in a metabolic pathway, information representing the relationship between the metabolites that are substrates and the metabolites that are products of the one or more enzymes; an analytical condition memory unit that stores analytical conditions for analyzing metabolites using a mass spectrometer; a gene locus information input unit for inputting information about a gene locus that affects the expression of an enzyme of interest; a metabolite identification unit that identifies metabolites that are substrates and products of an enzyme affected by the gene locus based on the information about the gene locus inputted in the gene locus input unit and the information stored in the enzyme information storage unit; and an analytical condition output unit that acquires from the analytical condition memory unit and outputs analytical conditions for analyzing the metabolites identified by the metabolite identification unit using a mass spectrometer.
[0017] Here, the "enzyme of interest" corresponds to an enzyme affected by a QTL related to a specific quantitative trait revealed by QTL analysis. In the metabolite analysis support system according to the present invention, when information on a gene locus that affects the expression of an enzyme of interest to the user is input to the gene locus input unit, the metabolites that are the substrate and product of the enzyme affected by the gene locus are identified based on the information and the information stored in the enzyme information storage unit, and analysis conditions for analyzing the metabolites by a mass spectrometer are output. Examples of enzymes of interest to the user include enzymes involved in the production of metabolites related to a specific human disease or specific constitution, and enzymes involved in the production of metabolites related to useful traits of animals and plants. In other words, in the present invention, it is possible to narrow down the target to a metabolite that is the subject of metabolome-QTL analysis and perform mass analysis under analysis conditions suitable for the metabolite. EXAMPLES
[0018] Hereinafter, a specific embodiment in which the present invention is applied to a metabolite analysis system will be described with reference to the drawings.
[0019] 1 is an overall schematic diagram of a metabolite analysis system according to the present embodiment. This metabolite analysis system includes a mass spectrometer, GC / MS 11 consisting of a GC section 111 and an MS section 112, an LC / MS 12 consisting of an LC section 121 and an MS section 122, a GC / MS analysis control section 13 for controlling the operation of each of the GC / MS 11 and the LC / MS 12, an LC / MS analysis control section 14, a base sequence determination section 15, and a control / processing personal computer (PC) 2. The base sequence determination section 15 is specifically a DNA sequencer, in particular a next-generation sequencer.
[0020] The control / processing PC2 includes a storage unit 20 and, as functional blocks, a data processing unit 211, a metabolite identification unit 212, a gene locus information input unit 213, an analysis condition setting unit 214, a method file creation unit 215, a method file output unit 216, a gene locus identification unit 217, a display processing unit 218, and an analysis result storage unit 219. The actual entity of the control / processing PC2 is a general personal computer, and each of the above functional blocks is realized by executing pre-installed software 21 for a metabolite analysis support system. In addition, an input unit 3 and a display unit 4 are connected to the control / processing PC2. In this embodiment, the control / processing PC2 corresponds to the metabolite analysis support system according to the present invention.
[0021] The storage unit 20 includes an enzyme information storage unit 201 , a method file storage unit 202 , and a metabolic map data storage unit 203 . The enzyme information storage unit 201 stores, for one or more enzymes, enzyme information that indicates the relationship between the enzyme involved in a metabolic pathway in a living body and the substrate and product metabolite of the enzyme. The enzyme information includes the metabolic pathway in which the enzyme participates and the order of metabolism in the metabolic pathway in addition to the enzyme, its substrate, and product metabolite.
[0022] The method file storage unit 202 stores information for creating a method file, such as analytical conditions (analysis name, target compound, etc.) used when analyzing a sample containing metabolites using GC / MS11 or LC / MS12, and an analysis method for the analysis results (retention index of known standard samples, mass spectrum, characteristic ion information (mass-to-charge ratio of ions, intensity ratio of multiple ions, etc.), chromatogram peak picking method, baseline correction method, etc.). Note that the number of analytical conditions when analyzing a certain metabolite using GC / MS11 or LC / MS12 is not limited to one, and there may be multiple conditions. The method file storage unit 202 corresponds to the analysis condition storage unit of the present invention.
[0023] The metabolic map data storage unit 203 stores data constituting one or more (usually more than one) metabolic maps. Here, a metabolic map refers to a chart showing metabolic pathways in the living body of a human or other organism. The metabolic map lists various compounds (metabolites) produced in the metabolic process, chemical reactions, enzymes involved in metabolism, and the like, allowing the flow of metabolism to be understood at a glance.
[0024] Next, a procedure for setting analysis conditions and creating a method file when analyzing metabolites contained in a sample using the above-mentioned metabolite analysis system with the GC / MS 11 or LC / MS 12 will be described.
[0025] First, the user (operator) operates the input unit 3 to instruct the display of an analysis condition setting screen. In response to this instruction, the display processing unit 218 reads out the analysis condition setting screen from the analysis condition setting unit 214 and displays it on the display unit 4. FIG. 2 shows an example of the analysis condition setting screen displayed on the display unit 4. On this screen, input fields for the Cas number, the locus name, the protein name, and the protein identification number (PubChemID) as input fields for gene locus information, and a display list displaying various conditions for mass spectrometry are displayed. The user operates the input unit 3 to input predetermined gene locus information into one or more of the input fields for gene locus information on the analysis condition setting screen. The gene locus information inputted into the analysis condition setting screen is inputted into the gene locus information input unit 213. The predetermined gene locus information is identification information of a gene locus (QTL) that affects the expression of an enzyme related to a metabolite whose presence or absence is to be confirmed in a sample, identification information of the enzyme, etc., and the enzyme related to a metabolite whose presence or absence is to be confirmed in a sample corresponds to the "enzyme of interest" in the present invention.
[0026] The input field for locus information includes a metabolic pathway selection field for selecting a metabolic pathway. When the input unit 3 is operated to select the metabolic pathway selection field, the names of a plurality of metabolic pathways are displayed as a pull-down list, and the metabolic pathway in which the "metabolite to be confirmed as being contained in the sample" is generated can be selected from this pull-down list. When a metabolic pathway is selected in the metabolic pathway selection field, the names of proteins, which are a plurality of enzymes involved in the metabolic pathway, are displayed as a pull-down list in the protein name input field. In this case, a predetermined protein name can be selected and input from the pull-down list. Of course, the user can also operate a keyboard or the like to input locus information in the form of characters or symbols into one or more of the input fields for locus information. When locus information is input into one of the input fields for locus information, other information (e.g., protein name, PubChemID) linked to that information (e.g., locus name) may be automatically displayed in the input field.
[0027] The locus information can also be obtained from a base sequence determined by extracting DNA contained in a sample and analyzing the genome information of the DNA using the base sequence determination unit 15 (next-generation sequencer). That is, when the user operates the input unit 3 to instruct input of base sequence information from the base sequence determination unit 15 to the locus identification unit 217 and input the enzyme of interest, the locus identification unit 217 acquires the base sequence information from the base sequence determination unit 15. Then, the locus identification unit 217 performs QTL analysis using the above-mentioned enzyme as an indicator for the acquired base sequence to identify a locus that affects the expression of the enzyme. The locus identified in the locus identification unit 217 is input to the locus information input unit 213.
[0028] The analysis condition setting screen in Figure 2 shows the displayed state after the following has been entered: "Kynurenine metabolism" as the metabolic pathway, "Vermillion" as the locus name, "TDO2" as the protein name, "161166" as the PubChemID, and "2922-83-0" as the Cas number.
[0029] When locus information is input in the input field for gene information, the metabolite identification unit 212 reads out metabolites, which are the substrate and product of the enzyme associated with the locus, from the enzyme information storage unit 201 based on the input locus information, and identifies these two metabolites as metabolites to be analyzed. In addition, in this embodiment, the metabolite identification unit 212 reads out metabolites, which are the substrate and product of the enzyme associated with the locus, as well as five metabolites located downstream of the substrate of the enzyme and five metabolites located upstream of the product, and identifies these metabolites as metabolites to be analyzed. Note that, when the number of metabolites located downstream or upstream is less than five, the number of metabolites identified as metabolites to be analyzed is less than five. That is, in this embodiment, when locus information is input, a maximum of 12 metabolites are identified as metabolites to be analyzed. In addition, the names of the identified maximum of 12 metabolites to be analyzed are displayed in order in the compound name list on the analysis condition setting screen.
[0030] In this embodiment, the maximum number of metabolites to be analyzed is set to 12, but the number of metabolites to be analyzed may be 2 (i.e., only metabolites that are substrates and products of the enzyme associated with the gene locus) or may be 3 or more. Also, the user may be able to operate the input unit 3 to specify the number of metabolites to be included in the metabolites to be analyzed other than the metabolites that are substrates and products of the enzyme associated with the gene locus.
[0031] When the metabolite identification unit 212 identifies the metabolites to be analyzed, the analysis condition setting unit 214 reads and sets information such as the analysis conditions used when analyzing a sample containing each metabolite to be analyzed using the GC / MS 11 and LC / MS 12 from the method file storage unit 202, and the analysis method of the analysis results. In addition, the method file creation unit 215 creates a method file that describes the procedure (schedule) for analyzing each metabolite to be analyzed, the analysis conditions, the analysis method of the analysis results, etc. The created method file is saved in the method file output unit 216 together with its identification information (method number, file name, etc.).
[0032] When a user selects and instructs identification information of a method file stored in the method file output unit 216, the method file output unit 216 reads out the selected method file and outputs it to the GC / MS analysis control unit 13 or the LC / MS analysis control unit 14. When the GC / MS analysis control unit 13 or the LC / MS analysis control unit 14 receives an instruction to start analysis, it controls the GC / MS 11 or LC / MS 12 in accordance with the method file and performs the analysis of the sample.
[0033] Since the method file contains analytical conditions suitable for each target metabolite, the target metabolite can be detected with high accuracy and high reliability by performing the analysis of the sample under these analytical conditions. In addition, the analysis is performed by focusing on a specific target metabolite among the many metabolites contained in the sample, thereby shortening the analysis time and enabling rapid analysis.
[0034] The analysis data collected by the GC / MS 11 and / or LC / MS 12 is input to the data processing unit 211, which processes the data based on the data analysis method described in the method file. In this method file, the retention index, mass spectrum, characteristic ion information, etc. are determined for each metabolite to be analyzed, so that the quantitative value of the metabolite to be analyzed can be reliably obtained by using these. When the quantitative values of all metabolites to be analyzed are obtained, the analysis results are stored in the analysis result storage unit 219.
[0035] After the quantitative values of the metabolites to be analyzed are obtained, when the user operates the input unit 3 to instruct the creation of a metabolic map, the display processing unit 218 reads out the metabolic map data from the metabolic map data storage unit 203, and also reads out the measurement results (quantitative values) of the metabolites listed on the metabolic map from the analysis result storage unit 219 and enters them in the corresponding locations on the metabolic map. From such a metabolic map, the user can easily recognize whether or not the sample contains the metabolites to be analyzed, and if so, how much of them are contained.
[0036] [Aspects] It will be apparent to those skilled in the art that the above-described exemplary embodiments are illustrative of the following aspects.
[0037] (Item 1) A metabolite analysis support system according to one aspect of the present invention comprises: A system for supporting analysis of metabolites using a mass spectrometer, comprising: an enzyme information storage unit that stores enzyme information representing a relationship between a metabolite that is a substrate of one or more enzymes involved in a metabolic pathway in the body of an organism and a metabolite that is a product of the one or more enzymes; an analysis condition storage unit that stores analysis conditions for analyzing metabolites by a mass spectrometer; a locus information input section for inputting information regarding a locus that affects expression of an enzyme of interest; a metabolite identification unit that identifies metabolites that are substrates and products of an enzyme affected by the gene locus as target metabolites based on information about the gene locus inputted into the gene locus input unit and information stored in the enzyme information storage unit; an analysis condition output unit that acquires from the analysis condition storage unit and outputs analysis conditions for analyzing the target metabolite identified by the metabolite identification unit using a mass spectrometer; It is equipped with the following.
[0038] According to the metabolite analysis support system of the first aspect, based on information on a gene locus that affects the expression of an enzyme of interest to the user and information stored in the enzyme information storage unit, it is possible to identify metabolites that are substrates and products of the enzyme affected by the gene locus, and obtain analytical conditions for analyzing those metabolites with a mass spectrometer. Therefore, by inputting enzymes involved in the production of metabolites related to a specific disease or specific constitution in humans, enzymes involved in the production of metabolites related to useful traits in animals and plants, etc. as enzymes of interest to the user, it is possible to narrow down the target metabolites to be subjected to metabolome-QTL analysis and perform mass spectrometry analysis under analytical conditions suitable for those metabolites.
[0039] (2) The metabolite analysis support system according to the second aspect of the present invention is the metabolite analysis support system according to the first aspect of the present invention, further comprising: The apparatus is equipped with a gene locus identification unit that performs QTL analysis of the base sequence of DNA contained in a sample collected from an organism using the enzyme of interest as an indicator to identify gene loci that affect the expression of the enzyme.
[0040] According to the metabolite analysis support system according to the second aspect, information regarding gene loci that affect the expression of an enzyme of interest can be easily obtained.
[0041] (3) The metabolite analysis support system according to the third aspect of the present invention is the metabolite analysis support system according to the second aspect of the present invention, further comprising: The system further comprises an input instruction unit for inputting information about the gene locus identified by the gene locus identification unit into the gene locus information input unit.
[0042] According to the metabolite analysis support system of paragraph 3, the task of inputting information about a gene locus that influences the expression of an enzyme of interest, identified by the gene locus identification unit, into the gene locus information input unit can be simplified.
[0043] (4) The metabolite analysis support system according to the fourth aspect of the present invention is the metabolite analysis support system according to the second or third aspect of the present invention, further comprising: It is equipped with a base sequence determination unit that determines the base sequence of DNA contained in a sample collected from an organism.
[0044] According to the metabolite analysis support system of the fourth aspect, a series of tasks can be easily performed, from determining the base sequence of DNA contained in a sample to identifying the gene locus that affects the expression of an enzyme of interest.
[0045] (Item 5) The metabolite analysis support system according to item 5 is the metabolite analysis support system according to item 4, in which the base sequence determination unit is a DNA sequencer. In this case, it is particularly preferable to use a DNA sequencer called a next-generation sequencer, since it is possible to rapidly analyze genomic information of DNA and determine its base sequence.
[0046] (Item 6) The metabolite analysis support system according to item 6 is a metabolite analysis support system according to any one of items 1 to 5, the enzyme information storage unit stores the enzyme information in the order of the metabolic pathway, The metabolite identifying section identifies a predetermined number of metabolites located downstream or upstream in a metabolic pathway relative to the reaction of the enzyme affected by the gene locus as metabolites to be analyzed.
[0047] (Item 7) The metabolite analysis support system according to item 7 is the metabolite analysis support system according to item 6, further comprising: The metabolite number designation unit is provided for designating the number of downstream or upstream metabolites to be included in the metabolites to be analyzed.
[0048] According to the metabolite analysis support system according to the sixth or seventh aspect, mass spectrometry can be easily performed by expanding the analysis target to the periphery of the metabolic pathway in which the enzyme of interest is involved. [Explanation of symbols]
[0049] 11…GC / MS 111…GC Department 112...MS Department 12…LC / MS 121...LC section 122…MS Department 13...GC / MS analysis control section 14...LC / MS analysis control section 15...Sequencer 2…Control and processing PC 20...Storage section 201...Enzyme information storage unit 202...Method file storage unit 203...metabolic map data storage unit 21...Software for metabolite analysis support system 211...Data processing unit 212...Metabolite identification department 213...Gene locus information input section 214…Analysis condition setting section 215...Method File Creation Department 216...Method file output section 217…Locus Identification Section 218...Display processing unit 219…Analysis result storage unit 3. Input section 4...Display section