Method and apparatus for constructing a database for microbial identification
By constructing a mass-to-charge ratio database using high-quality genomic data and weighting expressed proteins, the method addresses inaccuracies in microbial identification, enhancing the precision of microorganism classification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SHIMADZU SEISAKUSHO LTD
- Filing Date
- 2023-04-03
- Publication Date
- 2026-07-24
AI Technical Summary
Existing methods for microbial identification using mass spectrometry face challenges due to the variability of mass spectral patterns influenced by culture conditions and genetic diversity, leading to inaccuracies in predicted mass-to-charge ratio databases, which can result in misidentification of microorganisms.
A method and apparatus that constructs a mass-to-charge ratio database using high-quality genomic data from a genome database, predicting expressed proteins and weighting proteins likely to be expressed and detected, thereby reducing false peaks and improving identification accuracy.
The method enhances the quality and accuracy of microbial identification by ensuring the database includes only reliable mass-to-charge ratios, reducing the likelihood of misidentification and improving the precision of microorganism classification.
Smart Images

Figure 0007894610000002 
Figure 0007894610000003 
Figure 0007894610000004
Abstract
Description
Technical Field
[0001] The present invention relates to a method and an apparatus for constructing a database for microorganism discrimination.
Background Art
[0002] Non-Patent Document 1 discloses that there can be two approaches to the discrimination of microorganisms using mass spectrometry.
[0003] The first approach is a fingerprint method for discriminating an unknown microorganism by comparing the mass spectrum measured for the unknown microorganism with a database of mass spectra measured for each known microorganism. However, this method has problems such as the pattern of the mass spectrum of a microorganism being strongly affected by the culture medium conditions and the method of measurement reproducibility.
[0004] Regarding the problems of such a fingerprint method, as a second approach, a method based on bioinformatics using a genomic database has been attracting attention. In this method, an unknown microorganism is discriminated by comparing the mass spectrum measured for the unknown microorganism with a database of mass-to-charge ratios of proteins predicted from the genomic database. In this method, since the predicted mass-to-charge ratio is not affected by the culture medium conditions and the method of measurement reproducibility, the problems of the above fingerprint method can be solved.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In this second method, further improvement is desired in the quality of the predicted mass-to-charge ratio database. For example, it is possible that low-quality genomic data included in the genome database is affecting the quality of the predicted mass-to-charge ratio database.
[0007] This disclosure was made to address the aforementioned issues, and its purpose is to improve the quality of mass-to-charge ratio databases, which are constructed based on genome databases and used for identifying microorganisms using mass spectrometry. [Means for solving the problem]
[0008] A method for constructing a database for microbial identification relating to the first aspect of this disclosure comprises the steps of: obtaining microbial genomic data from a genome database; determining whether the obtained genomic data meets a criterion; predicting the expressed proteins for each genomic data determined to meet the criterion; and constructing a mass-to-charge ratio database which includes a list of mass-to-charge ratios for each genomic data predicted based on the predicted proteins.
[0009] The apparatus for constructing a database for microbial identification according to the second aspect of this disclosure constructs a database for microbial identification using microbial genomic data obtained from a genome database. The apparatus comprises a processor and a storage unit. The processor determines whether the acquired genomic data meets the criteria. The processor also predicts the expressed proteins for each genomic data determined to meet the criteria. The processor also constructs a mass-to-charge ratio database containing a list of predicted mass-to-charge ratios for each genomic data based on the predicted proteins. The processor also stores the mass-to-charge ratio database in the storage unit. [Effects of the Invention]
[0010] The method for constructing a database for microbial identification described herein allows for the construction of a mass-to-charge ratio database based only on genomic data that meets the criteria in the genome database. In other words, it is possible to improve the quality of mass-to-charge ratio databases constructed based on genome databases, which are used for microbial identification using mass spectrometry. [Brief explanation of the drawing]
[0011] [Figure 1] This is a schematic diagram showing the configuration of a microorganism identification system according to an embodiment of the present invention. [Figure 2] This is a flowchart outlining the processes performed by the device. [Figure 3] This is a functional block diagram of the device for constructing a mass-to-charge ratio database. [Figure 4] This is a functional block diagram of the device related to sample identification. [Figure 5] This flowchart shows the process for building a mass-to-charge ratio database. [Figure 6] This flowchart shows the subroutine for determining genome data. [Figure 7] This flowchart shows the process of adding new genome data. [Figure 8] This is a flowchart showing the process for identifying samples. [Figure 9] This flowchart shows another example of the process for identifying samples. [Figure 10] This figure shows the relationship between the total number of base sequences at a gene site in the genome and the estimated number of genes per genome. [Modes for carrying out the invention]
[0012] Embodiments of the present invention will be described in detail below with reference to the drawings. In the following description, the same or corresponding parts in the drawings will be denoted by the same reference numerals, and their descriptions will not be repeated in principle.
[0013] [Configuration of the Microorganism Discrimination System] FIG. 1 is a schematic diagram showing the configuration of a microorganism discrimination system 1000 according to an embodiment of the present invention.
[0014] Referring to FIG. 1, the microorganism discrimination system 1000 includes a public genome database 70, a public classification database 80, a network 90, and a device 100. In this specification, "database" is also referred to as "DB".
[0015] The public genome DB 70 is a database containing genomic data of organisms. A genome is genetic information on nucleic acids (deoxyribonucleic acid (DNA), ribonucleic acid (RNA)) possessed by an organism and includes the base sequence of the nucleic acid. In this specification, genomic data mainly refers to data of DNA sequences.
[0016] The public genome DB 70 is typically a DB containing a large number of genomic data of generally publicly available organisms, such as the genome DBs of NCBI (National Center for Biotechnology Information), DDBJ (DNA Data Bank of Japan), and EMBL (European Molecular Biology Laboratory). However, the examples of the public genome DB 70 are not limited to this, and for example, a genome DB that is not generally publicly available may also be included.
[0017] The public classification DB 80 is a database containing data related to the classification of organisms (hereinafter referred to as classification data). The classification of organisms is generally a classification based on the phylogenetic relationship between organisms indicated by classes such as family, genus, and species. In the classification of microorganisms, traditionally, classification has been made based on multiple indicators such as morphological observation, phenotype, chemotaxonomic index, protein analysis, and DNA analysis based on both phenotype and genome, but there is also a classification system based only on genomic information, and there are multiple classification systems.
[0018] The public classification DB80 is typically a DB containing publicly available biological classification data, such as DBs like GTDB (Genome Taxonomy Database), RDP (Ribosomal Database Project), Silva, etc. However, the examples of the public classification DB80 are not limited to this, and for example, it may include DBs that are not publicly available in general.
[0019] The network 90 is a network for the device 100 to communicate with the public genome DB70 and the public classification DB80. The network 90 is, for example, the Internet that interconnects a large number of government, corporate, public, and private networks on the earth.
[0020] The device 100 is a device for constructing a mass-to-charge ratio (m / z) DB for discriminating microorganisms using mass spectrometry. In this specification, discriminating microorganisms means taxonomically identifying microorganisms. That is, for example, it is to identify at least one of the genus, species, strain, and lineage of microorganisms. Therefore, the device 100 corresponds to an example of "a device for constructing a database for microorganism discrimination". Also, the device 100 is a device for discriminating microorganisms using mass spectrometry using the m / z DB. Therefore, the device 100 also corresponds to an example of "a microorganism discrimination device". Note that in this specification, the "type" of a microorganism or organism includes, for example, at least one of the "genotype, strain, or rank of a taxonomic group such as subspecies, species, genus, family, etc." of the microorganism or organism.
[0021] The device 100 includes a controller 101, a display 15, and an operation unit 14. The display 15 and the operation unit 14 are connected to the controller 101. The operation unit 14 is typically composed of a touch panel, a keyboard, a mouse, etc. The operation unit 14 receives the user's operation input to the processor 10. The display 15 is composed of, for example, a liquid crystal panel capable of displaying images. The display 15 displays an image related to receiving the user's operation input and displays the result of the processing by the processor 10.
[0022] The controller 101 has, as its main components, a processor 10, memory 11, a communication interface (I / F) 12, and an input / output I / F 13. Each of these parts is connected to each other via a bus so that they can communicate with one another.
[0023] The processor 10 is typically an arithmetic processing unit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). The processor 10 controls the operation of the device 100 by reading and executing programs stored in memory 11.
[0024] Memory 11 is implemented using storage devices such as ROM (Read Only Memory), RAM (Random Access Memory), and HDD (Hard Disk Drive). ROM can store programs executed by the processor 10. RAM can temporarily store data used during program execution by the processor 10 and can function as a temporary data memory used as a workspace. HDD is a non-volatile storage device. In addition to, or instead of, HDD, semiconductor storage devices such as flash memory may be used. The above programs and / or data may be stored in an external storage device accessible by the processor 10. Memory 11 corresponds to one embodiment of the "storage unit".
[0025] Communication I / F12 is a communication interface for exchanging various data with external devices, including the public genome DB70 and the public classification DB80, and is implemented by an adapter or connector. The communication method may be wireless communication such as Wi-Fi (Local Area Network) or wired communication such as USB (Universal Serial Bus).
[0026] The input / output interface 13 is an interface for exchanging various types of data between the processor 10 and external devices connected to the input / output interface 13. The external devices include the control unit 14 and the display 15. A mass spectrometer (MS) 16 may also be connected to the input / output interface 13. In this specification, the input / output interface 13 also includes devices that exchange data between a storage terminal such as a USB memory connected to the device 100 and the processor 10.
[0027] MS16 is an instrument for performing mass spectrometry of components contained in a sample, and is, for example, MALDI-TOF MS (Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry), MALDI-IT-TOF (Matrix-Assisted Laser Desorption / Ionization Ion Trap Time-of-Flight Mass Spectrometry), or scanning IT-MS, but is not limited to these. When MS16 is MALDI-TOF MS, ions generated by laser irradiation are extracted into a flight tube and allowed to fly, and then detected after being separated according to the flight time. The flight time correlates with the mass-to-charge ratio (m / z) of the components. As a result, a mass spectrum is obtained with m / z on the x-axis and the detected ion intensity on the y-axis.
[0028] In this specification, MS16 performs mass spectrometry of proteins in a sample. Therefore, in the mass spectrum, peaks are detected according to the m / z of the proteins in the sample. Thus, by referring to the pattern of the mass spectrum, or more specifically, the list of m / z values (referred to herein as the m / z list) for which peaks with a height above a predetermined threshold were obtained, the proteins contained in the sample can be identified. In this specification, the m / z values included in the m / z list represent the m / z values corresponding to the peaks in the mass spectrum.
[0029] Different types of organisms contain different proteins, resulting in different mass spectral patterns and m / z lists. Therefore, organisms can be identified based on their mass spectral patterns and m / z lists.
[0030] After performing mass spectrometry on the unknown microorganism sample, MS16 sends a sample list, which is a m / z list of the samples, to the instrument 100. The processor 10 identifies the samples based on the sample list.
[0031] Furthermore, the device 100 does not need to be composed of a single computer; it may be composed of multiple computers.
[0032] [2. Comparison with conventional devices] Traditionally, a fingerprinting method has been used to identify microorganisms using such mass spectrometers. This method involves constructing a database containing mass spectra actually measured for each microorganism, and then comparing the mass spectra of unknown microorganisms with those in the database.
[0033] However, building a practical database for the fingerprinting method requires measured mass spectral data from many types of microorganisms (e.g., over a thousand species). Furthermore, even within the same species of microorganism, the mass spectral pattern can vary due to genetic diversity, culture conditions, pretreatment for mass spectral measurement, and variability associated with repeated measurements. Therefore, considering these realities, a practical database requires a very large amount of mass spectral data; for example, dozens of data points for each type of microorganism, totaling tens of thousands of data points across all species. In other words, building a practical database requires actually culturing microorganisms and measuring mass spectra a very large number of times (e.g., tens of thousands of times), which is extremely costly.
[0034] Therefore, a new method for identifying microorganisms using mass spectrometry has attracted attention. This method involves using publicly available genome databases to predict expressed proteins, constructing an m / zDB (m / z list database) of m / z lists predicted from those proteins, and then utilizing this m / zDB. In this method, samples are identified by comparing the m / z lists contained in the m / zDB with a sample list, which is an m / z list corresponding to the peaks of the mass spectrum of an unknown microorganism. This method eliminates the need to actually culture the microorganism and measure its mass spectrum, and allows for the construction of a mass spectrum database more easily compared to the fingerprinting method described above.
[0035] However, even with this method, there was room for improvement in the quality of the predicted m / zDB and the accuracy of microbial identification using the m / zDB.
[0036] For example, this method would include low-quality genome data (e.g., genome data containing many undetermined bases) from publicly available genome databases in the m / zDB. As a result, there were concerns that the quality of the m / zDB would deteriorate, and the accuracy of microbial identification using the m / zDB would also decrease.
[0037] Therefore, in the apparatus 100 according to this embodiment, the m / zDB is constructed based only on high-quality genome data that meets predetermined criteria from the genome data obtained from the public genome DB 70. This improves the quality of the m / zDB. Furthermore, it improves the accuracy of microorganism identification using the m / zDB.
[0038] Furthermore, other problems were a concern with conventional methods of identifying microorganisms using predicted m / z databases. For example, there was concern that the predicted m / z list might contain false peaks that would not appear in the mass spectrum when measured. This is because even sequences that are predicted to express proteins based on genome data may not actually express proteins for some reason, or even if expressed, they may not be ionized, and therefore may not be detected as peaks in the measured mass spectrum. As a result, when comparing the sample list with the predicted m / z list, these false peaks could act as noise, potentially causing the sample list to match the m / z list of microorganisms unrelated to the sample. Therefore, there was a possibility that the sample might be misidentified as an unrelated microorganism. In other words, there was concern that the accuracy of microorganism identification would decrease.
[0039] Therefore, in the apparatus 100 according to this embodiment, microorganisms are identified by weighting proteins that are less likely to produce false peaks, that is, "proteins that are likely to be expressed in the living organism of microorganisms and are likely to be detected as peaks when mass spectrometry is measured." Thus, the possibility of misidentification of a microorganism due to the influence of false peaks is reduced. This improves the accuracy of microorganism identification.
[0040] [3. Overview of the device's processing] Figure 2 is a flowchart outlining the processing performed by the device 100. In step (hereinafter also referred to as ST) 101, the processor 10 of the device 100 constructs an m / z database from the genome data of the public genome DB 70. In ST 102, the processor 10 uses the m / z database to identify samples that are unknown microorganisms.
[0041] (3.1. Functional blocks related to the construction of m / zDB) Figure 3 is a functional block diagram of the device 100 for constructing the m / zDB, corresponding to ST101 in Figure 2. Referring to Figure 3, the device 100 includes a genome data acquisition unit 21, a genome data determination unit 22, a protein prediction unit 23, an m / zDB construction unit 24, and a storage unit 25.
[0042] The genome data collection unit 21 collects genome data from the public genome database 70. The genome data determination unit 22 determines whether the collected genome data meets predetermined criteria related to the quality of the genome data.
[0043] The protein prediction unit 23 predicts the expressed proteins for genome data that meets predetermined criteria. Specifically, it predicts the predicted gene region from the DNA sequence, and then predicts the amino acid sequence from the predicted gene region. Based on this amino acid sequence, it then predicts the expressed proteins.
[0044] The m / zDB construction unit 24 predicts an m / z list based on the predicted proteins, constructs an m / zDB, and stores it in the storage unit 25. The m / zDB includes, for example, two types of m / zDBs. One m / zDB is a general m / zDB containing m / z corresponding to all proteins predicted from the genome data. The other m / zDB is a specific m / zDB containing only m / z corresponding to proteins belonging to a particular group of proteins among the proteins predicted from the genome data. These two m / zDBs are used to identify samples that are unknown microorganisms, as explained in Figure 4.
[0045] The genome data acquisition unit 21, genome data determination unit 22, protein prediction unit 23, and m / zDB construction unit 24 correspond to the processor 10 in Figure 1. The storage unit 25 corresponds to the memory 11 in Figure 1.
[0046] (3-2. Functional blocks related to sample discrimination) Figure 4 is a functional block diagram of the sample discrimination apparatus 100, which corresponds to ST102 in Figure 2. Referring to Figure 4, the apparatus 100 includes an acquisition unit 31, a sample discrimination unit 32, an annotation unit 33, an output unit 34, and a storage unit 25.
[0047] The acquisition unit 31 acquires a sample list. The sample list is acquired, for example, from the MS16 connected to the device 100. The method of acquiring the sample list is not limited to this, and it may also be acquired, for example, from an external device that communicates with the device 100, or from a storage terminal connected to the device 100. The acquisition unit 31 further estimates and corrects the m / z measurement error included in the sample list as needed. The acquisition unit 31 corresponds to the processor 10 in Figure 1.
[0048] The sample discrimination unit 32 discriminates a sample by comparing the sample list with the m / z database stored in the storage unit 25, weighting the comparison based on the m / z values corresponding to proteins in a specific group. The sample discrimination unit 32 includes, for example, a primary screening unit 321 and a secondary screening unit 322. The primary screening unit 321 uses the m / z list contained in the specific m / z database to perform screening based on the m / z values corresponding to proteins in a specific group. The secondary screening unit 322 discriminates a sample by performing screening based on the m / z values corresponding to all proteins in the m / z list that corresponds to the genome data narrowed down in the primary screening, from the m / z list contained in the overall m / z database. The sample discrimination unit 32 corresponds to the processor 10 in Figure 1.
[0049] The annotation unit 33 links annotations, which are information about the predicted protein, to each m / z included in the sample list. For example, software is used to search for the name of the corresponding protein based on the protein's mass to link the annotations. The annotation unit 33 corresponds to the processor 10 in Figure 1.
[0050] The discrimination results and m / z annotations from the sample discrimination unit 32 are stored in the storage unit 25 and / or output by the output unit 34. The output unit 34 corresponds to the processor 10 and the display 15 or communication interface 12 in Figure 1. That is, the discrimination results and annotations are displayed on the display 15 and / or transmitted to an external device via the communication interface 12. This allows the user to recognize the discrimination results and annotations.
[0051] [4. Process flow for building m / zDB] (4-1. Building an m / z DB) Next, we will specifically explain the processing flow performed by device 100.
[0052] Figure 5 is a flowchart showing the process for building the m / zDB. The processes ST02 to ST28 shown in Figure 5 correspond to the process in ST101 in Figure 2.
[0053] Referring to Figure 5, in ST02, the processor 10 acquires microbial genome data from the public genome database 70. By acquiring genome data from multiple public genome databases 70, it is possible to comprehensively collect genome data for clinically or industrially important microbial species.
[0054] In ST04, the processor 10 integrates the acquired genome data to construct a collected genome database.
[0055] In ST06, the processor 10 determines whether the genomic data in the collected genome database meets predetermined criteria. These criteria are set so that only high-quality genomic data meets them. The specific details of the criteria are explained in Figure 6.
[0056] In ST08, processor 10 constructs a high-quality genome database containing genome data that has been determined to meet the criteria.
[0057] In ST10, processor 10 predicts the genes contained in the genome data included in the high-quality genome database. A gene refers to a specific region on DNA that is translated into a protein, or the information contained in that region. Gene prediction includes, for example, estimating the putative gene region on the genome data that is translated into a protein, using the translation start codon (ATG sequence) and stop codon (TGA sequence) as clues.
[0058] In ST12, processor 10 predicts the post-translational amino acid sequence from the predicted gene. Amino acid sequence prediction includes, for example, estimating the amino acids corresponding to each codon (three-base sequence) contained in the predicted gene region and concatenating them.
[0059] In ST14, processor 10 predicts post-translational modifications to a protein consisting of a predicted amino acid sequence. Post-translational modifications are modifications performed on a protein immediately after translation to transform it into a protein that actually functions in various parts of the body. Post-translational modifications include, for example, protein degradation including methionine removal and signal peptide removal, and specific chemical modifications including phosphorylation. Post-translational modifications are applied to most proteins and change their m / z. Therefore, by considering post-translational modifications, a more accurate m / z of a protein can be calculated.
[0060] In ST16, processor 10 predicts the protein with the predicted post-translational modifications.
[0061] In ST18, processor 10 predicts an m / z list for each genome data based on the protein. Specifically, the m / z corresponding to the protein is calculated based on the mass of the atoms contained in the protein. Preferably, the average mass of the element, which reflects the isotopic distribution of the element in nature, is used as the atomic mass. This allows for the calculation of a more accurate m / z.
[0062] In ST20A, processor 10 constructs a global m / zDB, which is a database of mass-to-charge ratios containing the m / z list. The global m / zDB contains all predicted m / z values for each genome data.
[0063] On the other hand, in ST22, processor 10 links annotations to the protein data predicted in ST16. Annotations generally contain information about the protein, including its name, function, etc. The linking of annotations is performed, for example, using general software that adds annotations according to m / z, but is not limited to this. For example, instrument 100 may create a table showing the relationship between m / z and annotations based on the public genome DB 70 and the public classification DB 80, and the linking may be performed using that table.
[0064] In this specification, footnotes are information about proteins, including information about the group to which a protein belongs. Information about a group of proteins includes at least one of the following: the protein's name, function, and family.
[0065] One advantage of linking annotations is that, based on the annotations, m / z values corresponding to proteins in a specific group can be selected and treated separately from m / z values corresponding to other proteins. Therefore, for example, it becomes possible to selectively weight m / z values corresponding to "groups of proteins that are likely to be expressed in vivo and are likely to be detected as peaks when mass spectrometry is measured" to identify microorganisms. This allows for sample selection by weighting "m / z values corresponding to proteins that are likely to be expressed in vivo and are likely to be detected as peaks when mass spectrometry is measured" compared to "m / z values corresponding to proteins that are not actually expressed as proteins in vivo or that do not appear in mass spectrometry even if expressed (false peaks)" in the m / z list predicted from genome data. Therefore, it is possible to suppress the reduction in sample selection accuracy caused by false peaks in the predicted m / z list acting as noise.
[0066] In order for a protein to be "highly likely to be expressed in vivo and detectable as a peak when a mass spectrum is measured," it is preferable that a group be selected based on at least one of the following conditions: the expression level is above a predetermined threshold; it has a function essential for maintaining life; a predetermined proportion or more of microorganisms classified as a predetermined type (e.g., microorganisms belonging to a predetermined family) have amino acid sequence similarity (homology) above a predetermined threshold; it is a basic protein; the mass-to-charge ratio can be analyzed within an error range of ±14 Da (more preferably within ±3 Da) when measured by MALDI-MS; the protein mass is within 4 to 30 kDa (more preferably 2 to 20 kDa); the number of protein types included in the group is above a predetermined number; and the number of microorganisms containing the genome data among microorganisms classified as a predetermined type (e.g., microorganisms belonging to a predetermined family) is above a predetermined proportion.
[0067] Furthermore, the above-mentioned functions essential for maintaining life include functions essential for at least one of the maintenance and proliferation of cells.
[0068] One example of a group determined by these conditions is ribosomal proteins. Other examples of this group include chaperones and DNA-binding proteins.
[0069] Furthermore, the group may not be limited to proteins that are significantly expressed in microorganisms in general, as exemplified above, but may also be proteins known to be significantly expressed in specific microorganisms. For example, by weighting specific proteins known to be significantly expressed in each genus, the likelihood of classifying a sample into the correct genus can be increased. In this specification, an example of a "significantly expressed protein" is a protein that exhibits an expression level above a predetermined threshold.
[0070] In ST24, processor 10 selects proteins that are predicted to belong to a specific group based on the group information contained in the annotations. In the subsequent ST26, processor 10 predicts a specific m / z list containing only the m / z predicted from the selected proteins. In ST20C, processor 10 constructs a specific m / z DB, which is an m / z database containing the specific m / z list.
[0071] Another advantage of linking annotations is that it makes it easier for users to understand which proteins each m / z in the m / z list corresponds to. From this perspective, in order to make annotations for m / z more readily available, in ST20B, processor 10 constructs an annotation database that aggregates annotations for m / z included in the overall m / z database.
[0072] Another advantage of linking annotations is that the validity of comparing the sample list with the m / z list included in the m / zDB can be examined by referring to the annotations. For example, annotations can be referenced in an m / z list in the m / zDB that has been determined to have a high degree of matching (precision) with the sample list. In this case, if the m / z list contains many m / z values corresponding to proteins that are presumed not to be expressed in that microorganism based on the annotations, the reliability of the m / z list itself is questionable, and therefore the validity of comparing it with the sample list is low, and the reliability of sample identification is also low. Similarly, if m / z values corresponding to functionally important and evolutionarily conserved proteins in the sample list coincide with noise m / z values in the m / z list, the validity of the comparison is low, and the reliability of sample identification is also low. Thus, when an m / z list with low validity for comparison with the sample list is found, the user can improve the reliability of identification by removing the m / z list.
[0073] Annotations in the annotation database are linked to m / z entries in the m / z database. For example, when an m / z entry in the m / z list in the m / z database is referenced, the m / z database and the annotation database are associated so that the corresponding annotation in the annotation database can also be referenced. Alternatively, the annotation database may be configured as part of the m / z database, with annotations corresponding to the m / z entries in the m / z database added to it.
[0074] In ST28, processor 10 acquires classification data from the public classification DB 80. In ST20D, processor 10 constructs a collected classification DB by integrating the collected classification data. At this time, by constructing the collected classification DB based on classification data from multiple public classification DBs 80, it is possible to incorporate a wide range of taxonomic systems. Therefore, by using the collected classification DB, it becomes possible to reflect various taxonomic systems in the identification results of microorganisms.
[0075] Furthermore, the collected classification database may include genome IDs, which are unique identifiers for each genome. These genome IDs are created, for example, based on the collected classification data.
[0076] The classification data in the collected classification database is associated with the data contained in the overall m / z database, specific m / z database, and annotation database, respectively. Therefore, a genome ID can be added to each genome data in the overall m / z database and specific m / z database. Furthermore, the contents of the collected classification database can be used to organize the overall m / z database and specific m / z database, or reflected in their contents. The collected classification database can also be used for other purposes in the device 100, such as when determining "specific proteins known to be significantly expressed only in specific species" as described above.
[0077] These four associated databases are collectively referred to as the microbial database. After constructing the microbial database using ST20A to ST20D, processor 10 terminates processing. This allows device 100 to use the microbial database to identify samples using mass spectrometry, as detailed in Figures 8 and 9.
[0078] The process shown in Figure 5 is performed, for example, once a year, in response to updates to the public genome database 70. This allows the updated content in the public genome database 70 to be reflected in the microbial database as needed, further improving the content of the microbial database.
[0079] (4-2. Interpretation of Genome Data) Figure 6 shows the genome data determination process. ST060 to ST069 shown in Figure 6 correspond to ST06 in Figure 2. The processes shown in Figure 6 are performed to remove low-quality genome data included in the collected genome database.
[0080] In ST060, processor 10 determines the quality of the genome data based on its completeness. Genome completeness is assessed using, for example, a group of single copy marker genes, which are known to exist as one copy each in the genome of a microbial organism. If the genome data is complete, all single copy marker genes should be present in the sample. However, if the genome data is incomplete, for example, if part of the genome data is missing or misread, the single copy marker genes included in the missing portion are lost. Therefore, the larger the missing or misread portion of the genome data, the fewer single copy marker genes there will be in the genome data. Thus, the number of single copy marker genes can be used as an indicator of genome data completeness. Specifically, completeness is calculated as a percentage proportional to the number of single copy marker genes present, with 100% representing the case where all single copy marker genes are present in the genome data.
[0081] Specifically, in ST060, processor 10 determines whether the completeness of the genome data is greater than the reference value T1. The reference value T1 is, for example, 50%. If the completeness is less than or equal to the reference value T1 (NO in ST060), in ST061, processor 10 removes the genome data. If the completeness is greater than the reference value T1 (YES in ST060), processor 10 proceeds to ST062.
[0082] In ST062, processor 10 determines the quality of the genome data based on the percentage of genome contamination. Contamination refers to the phenomenon where, for some reason, the DNA sequence of one genome data is mixed with the DNA sequence of another genome data. In other words, contamination typically means that the DNA sequences of multiple microorganisms are mixed together. If the percentage of single copy marker genes found when the genome data is not contaminated is set to 100%, then when contamination occurs, this percentage will be greater than 100%. Therefore, for example, the percentage of contamination is calculated based on the number of single copy marker genes found, with the case where no contamination occurs and all single copy marker genes are present in the genome data set to 100%. If the number of single copy marker genes found corresponds to (100+n)%, the percentage of contamination is n%, where n is a real number satisfying n>0. A high percentage of contamination suggests a high probability that the DNA sequences of multiple types of microorganisms are mixed together.
[0083] Specifically, in ST062, the processor 10 determines whether the contamination rate is less than the reference value T2. The reference value T2 is, for example, 20%. If the contamination rate is greater than or equal to the reference value T2 (NO in ST062), in ST063, the processor 10 removes the relevant genome data. If the contamination rate is less than the reference value T2 (YES in ST062), the processor 10 proceeds to ST064.
[0084] In ST064, processor 10 determines the quality of the genome data based on the number of contigs. A contig refers to a fragmented DNA sequence that has been divided into multiple DNA sequences. Therefore, the more contigs there are, the finer the DNA sequence is fragmented. If there are too many contigs, the gene regions that express proteins may also be fragmented, making accurate reading impossible. The number of contigs can be determined by counting how many fragments the DNA sequence contained in the genome data is divided into.
[0085] Specifically, in ST064, processor 10 determines whether the number of contigs is less than the threshold value T3. The threshold value T3 is, for example, 1000. If the number of contigs is greater than or equal to the threshold value T3 (NO in ST064), processor 10 removes the relevant genome data in ST065. If the number of contigs is less than the threshold value T3 (YES in ST064), processor 10 proceeds to ST066.
[0086] In ST066, processor 10 determines the quality of the genome data based on the number of undetermined bases. Undetermined bases refer to bases that could not be classified as either A, G, C, or T when the DNA sequence was decoded. DNA sequences containing many undetermined bases are highly likely to have difficulty properly identifying genes.
[0087] Specifically, in ST066, processor 10 determines whether the number of undecided bases is less than the reference value T4. The reference value T4 is, for example, 100,000. If the number of undecided bases is greater than or equal to the reference value T4 (NO in ST066), processor 10 removes the relevant genome data in ST067. If the number of contigs is less than the reference value T4 (YES in ST067), processor 10 proceeds to ST068.
[0088] In ST068, processor 10 determines the quality of the genome data based on whether the number of genes meets a standard value. This standard is used to determine whether the number of genes inferred from the genome data falls within a reasonable range. For example, if the number of genes inferred from the genome data is abnormally high, it is thought that for some reason, parts that are not actually genes are being inferred as genes. Such a cause could be, for example, an error in sequencing the DNA base sequence, where sequences that are not actually related to the start or end of transcription or translation are sequenced as sequences that are related to the start or end of transcription or translation. In this case, sequences that do not actually express proteins may be mistakenly identified as sequences that express proteins, and there is a concern that the predicted m / z list will contain many incorrect peaks. If such an m / z list is included in the m / zDB, the quality of the m / zDB will decrease, and the accuracy of sample identification will also decrease.
[0089] Specifically, in ST068, processor 10 determines whether the number obtained by dividing the number of genes in the genome data by the number of coding bases is less than the reference value T5. Gene coding bases generally refer to the bases included in the region related to protein expression on the DNA sequence. The reference value T5 is, for example, 0.00180. If the number obtained by division is greater than or equal to the reference value T5 (NO in ST068), processor 10 removes the genome data in ST069. If the number obtained by division is less than the reference value T5 (YES in ST068), processor 10 adds the genome data to the high-quality genome database.
[0090] Processor 10 performs ST060 to ST069 on all genomic data included in the collected genomics database.
[0091] The calculation methods for each of the criteria for completeness, contamination rate, number of contigs, number of undetermined bases, and validity of gene number are not limited to the examples above. For example, the validity of gene number may be determined by whether the number of genes contained in a single genome data is less than a predetermined threshold value.
[0092] The process shown in Figure 6 removes genomic data that does not meet the criteria from the collected genome database. In other words, low-quality genomic data from the public genome database 70 is removed, and only high-quality data is used to construct the m / z database. Therefore, the quality of the m / z database in the device 100 is improved.
[0093] (4-3. Addition of new genome data) The device 100 is also configured to allow the addition of new genome data to the m / zDB. This addition is performed, for example, when a user of the device 100 discovers a new microorganism and wishes to add its genome data.
[0094] Figure 7 shows the process of adding new genome data. In the flowchart of Figure 7, ST02 in the flowchart of Figure 5 has been changed to ST02A, and steps ST04 and ST08 in Figure 5 have been deleted. The processes from ST12 onwards in the flowchart of Figure 7 correspond to the processes from ST12 onwards in the flowchart of Figure 5.
[0095] In ST02A, the processor 10 acquires new genome data. Specifically, for example, the processor 10 acquires the genome data from an external device such as a DNA sequencer or storage device, or from a storage terminal such as a USB memory, via the input / output interface 13 or the communication interface 12.
[0096] In ST06, processor 10 determines whether the genome data meets the criteria. The criteria are set so that only high-quality genome data meets the criteria. If the new genome data meets the criteria, processor 10 proceeds to ST10. If the new genome data does not meet the criteria, processor 10 removes the new genome data.
[0097] In ST10, processor 10 predicts the genes included in the genome data and proceeds to ST12. The subsequent processing is the same as that shown in Figure 5, so the explanation will not be repeated. Therefore, processor 10 can add proteins whose expression is predicted from the new genome data to m / zDB if they meet the predetermined quality criteria.
[0098] This configuration allows for the addition of m / z lists predicted from newly acquired genome data to the m / zDB, thereby enriching its contents. As a result, the quality of the m / zDB is further improved, and the accuracy of sample identification using the m / zDB is also further enhanced.
[0099] [5. Processing flow for sample identification] (5-1.2 stage screening) The device 100 uses the m / zDB constructed as described above to identify the samples.
[0100] Figure 8 is a flowchart showing the process related to sample discrimination. The processes ST32 to ST54 shown in Figure 8 correspond to the process ST102 in Figure 2.
[0101] Referring to Figure 8, in ST32, the processor 10 obtains a sample list. The sample list is obtained, for example, from MS16. In ST34, the processor 10 determines whether or not to correct the m / z of the sample list. Whether or not to correct the sample list is, for example, set in advance by the user.
[0102] During analysis using mass spectrometers such as MALDI-TOF MS, the detected m / z value may be larger or smaller than the actual value depending on the mass of the protein contained in the sample and the instrument used. In other words, the sample list may contain some m / z shift as measurement error. On the other hand, the m / zDB included in instrument 100 is a theoretical value and therefore does not contain measurement error. Therefore, it is more accurate to shift the m / z in the sample list to cancel out the measurement error and then compare it with the m / zDB included in instrument 100 to identify the sample more accurately.
[0103] The estimation of measurement error is performed using the following procedure. First, the sample list containing the measurement error is compared directly with the "m / z list assumed to exclude the measurement error." Next, a predetermined value is searched for that, when the sample list is shifted by a predetermined value, yields a high degree of precision with the "m / z list assumed to exclude the measurement error." This predetermined value corresponds to the measurement error. Note that the predetermined value is searched for within the range of possible values for the measurement error.
[0104] The "m / z list assumed to be free of measurement errors" is, for example, an m / z list included in a specific m / z database that is considered unlikely to contain false peaks, but is not limited to this. It may also be an m / z list included in the overall m / z database, or another m / z list prepared for correcting measurement errors in the sample list.
[0105] If the m / z of the sample list is corrected (YES in ST34), in ST36, the processor 10 estimates the measurement error included in the sample list based on a specific m / z DB. In ST38, the processor 10 performs a correction by shifting the m / z of the sample list by the amount of the estimated measurement error.
[0106] When the m / z of the sample list is not corrected (NO in ST34), or following ST38, in ST40 to ST44, the processor 10 weights the sample list and the m / zDB by the m / z corresponding to the proteins included in a specific group and then compares them to discriminate the sample.
[0107] In ST40, as the primary screening, the processor 10 selects an m / z list with a matching rate with the sample list above a predetermined rank from among a specific m / zDB. More specifically, the “m / z list with a matching rate above a predetermined rank” is an m / z list in the m / z lists in the m / zDB used for screening, whose matching rate with the sample list is above a predetermined rank. For example, the m / z lists with the top N1 matching rates are selected as the m / z lists with a matching rate above a predetermined rank. N1 is an integer, for example, between 500 and 5000. Another example of the “m / z list with a matching rate above a predetermined rank” is an m / z list with a matching rate above a predetermined numerical value. The “m / z list with a matching rate above a predetermined numerical value” can be considered as the “m / z list with a rank above the rank corresponding to the number of m / z lists with a matching rate above a predetermined numerical value”.
[0108] In ST42, the processor 10 selects the m / z lists in the entire m / zDB corresponding to the top N1 m / z lists. In other words, it selects the m / z lists in the entire m / zDB of the genome corresponding to the top N1 m / z lists.
[0109] In ST44, as the secondary screening, the processor 10 discriminates the sample by selecting an m / z list with a high matching rate with the sample list from among the selected m / z lists in the entire m / zDB. For example, the top N2 m / z lists in the selected m / z lists in the entire m / zDB are selected as the m / z lists with a high matching rate. Note that N1 is an integer such that N2 < N1, for example, an integer between 1 and 100.
[0110] Once the sample discrimination is complete, in ST46, the processor 10 reflects the classification data in the discrimination results. For example, the processor 10 adds classification information (family, genus, species, lineage, etc.) for each of the N1 selected m / z lists of microorganisms.
[0111] Furthermore, the N1 m / z lists may be organized based on classification data. For example, a table may be created in which the N1 m / z lists are sorted in order of classification information. Alternatively, a diagram may be created showing the microorganisms corresponding to the N1 m / z lists on a phylogenetic tree. Alternatively, the number of microorganisms corresponding to a specific family, genus, species, or lineage included in the N1 m / z lists may be quantified. Specifically, a table may be created listing the family, genus, species, or lineage that appeared most frequently among the microorganisms corresponding to the N1 m / z lists. Additionally, the discrimination results may be further refined by reflecting classification data in the N1 m / z lists. Specifically, processing such as removing m / z lists that are taxonomically clear outliers may be applied. The processing exemplified above makes it possible to output discrimination results that reflect taxonomic perspectives. Furthermore, the processing exemplified above may be based on two or more taxonomic systems. This makes it possible to create discrimination results that reflect multiple taxonomic perspectives.
[0112] In ST48, processor 10 determines whether to predict the protein corresponding to the m / z included in the sample list, i.e., the protein that is expected to be expressed in the sample. Whether or not to predict a protein is set in advance by the user, for example.
[0113] If the protein is not predicted (NO at ST48), at ST50, the processor 10 outputs the discrimination result and terminates processing. The discrimination result is output, for example, by being displayed on the display 15.
[0114] If a protein is predicted (YES in ST48), in ST52, processor 10 links the annotations of the proteins corresponding to the m / z values included in the sample list. The annotations are information about the protein, as described above, and include information about the group in which the protein belongs. Specifically, for example, processor 10 adds entries for the name and function of the protein corresponding to the m / z value to the sample list. Alternatively, for example, processor 10 may create a table of the names and functions of the proteins corresponding to the m / z values included in the sample list, independent of the sample list.
[0115] When a protein annotation is linked, in ST54, the processor 10 outputs the discrimination result created in ST44 and ST46, and the annotation associated with the m / z included in the sample list in ST52, and then terminates processing. The discrimination result and annotation are output, for example, by being displayed on the display 15. By outputting information about the predicted expressed proteins from the sample list in this way, the user can easily recognize information about the proteins that are predicted to be expressed in the sample, thereby deepening their understanding of the sample. Furthermore, this information about the proteins can be referenced when reviewing the discrimination result of the sample, and can also be referenced when performing other analyses on the sample, thus enhancing user convenience.
[0116] The process shown in Figure 8 involves primary screening based on a specific m / z database and secondary screening based on the overall m / z database, which offers the following advantages. First, by focusing the primary screening on m / z values of functionally important and highly expressed proteins, it is possible to narrow down the list of m / z values with high precision while minimizing the influence of false peaks. Second, by performing secondary screening on all m / z values, it is possible to reflect the similarity of proteins other than those included in the specific group identified in the primary screening.
[0117] Furthermore, secondary screening may be performed by focusing on the m / z of proteins belonging to a specific group different from that of the primary screening. In this case, samples can be distinguished by focusing on two important proteins.
[0118] Furthermore, it is acceptable to combine three or more different screening methods in this manner. In summary, the device 100 can taxonomically classify samples through two or more screening steps, including m / z-based screening corresponding to proteins belonging to a specific group. By performing multiple different screenings in this way, the characteristics of each screening method can be utilized, thereby improving the overall accuracy of sample classification.
[0119] Furthermore, the device 100 can also be configured to identify samples based on classification data related to the classification of microorganisms.
[0120] For example, a specific m / z database (m / zDB) can be constructed to include only m / z values corresponding to groups of proteins commonly expressed within relatively high-level taxonomic groups (e.g., genera, which are higher-level taxonomic groups than species or strains). Conceptually, a specific m / zDB might include a group PA of proteins commonly expressed within a genus A. In this case, screening using this specific m / zDB can accurately determine whether a sample belongs to genus A or not. Similarly, by constructing a specific m / zDB to include each group of proteins commonly expressed within each genus, screening using this specific m / zDB can accurately determine the genus of a sample. Subsequently, a secondary screening can be performed to distinguish between species and strains within the determined genus. Therefore, the possibility of incorrect genus identification in the primary screening is reduced, and screening can be performed in a state suitable for species and strain identification in the secondary screening. Thus, the accuracy of sample identification can be improved.
[0121] (5-2. Weighting by score) In discriminating between samples, the method for weighting the m / z values corresponding to proteins in a specific group is not limited to the method using the specific m / z database described above. For example, when comparing a sample list with an m / z list included in the overall m / z database, the precision may be calculated such that when the m / z values corresponding to proteins in a specific group match, the precision is higher than when the m / z values corresponding to other proteins match.
[0122] Figure 9 is a flowchart showing another example of the process for sample discrimination. In the flowchart of Figure 9, steps ST40 to ST42 are changed to ST40A compared to the flowchart of Figure 8, and the other steps in Figure 9 are the same as in Figure 8.
[0123] In ST40A in Figure 9, the processor 10 selects m / z lists with high precision from the m / z lists in the overall m / zDB. For example, the top N3 m / z lists in the overall m / zDB are selected as m / z lists with high precision. Specifically, the processor 10 first calculates a score for each m / z list in the overall m / zDB by multiplying the number of matches between m / z in the m / z list and m / z in the sample list by a predetermined coefficient. Then, it selects m / z lists in the overall m / zDB whose scores are of a predetermined rank or higher. More specifically, "m / z lists whose scores are of a predetermined rank or higher" refers to m / z lists in the m / zDB used for screening whose scores are of a predetermined rank or higher. For example, the top N3 m / z lists with scores are selected as m / z lists whose scores are of a predetermined rank or higher. Another example of "m / z lists whose scores are of a predetermined rank or higher" is m / z lists whose scores are of a predetermined value or higher. A "m / z list with a score greater than or equal to a predetermined value" can be thought of as "m / z lists with a rank greater than or equal to the number of m / z lists with a score greater than or equal to a predetermined value."
[0124] In this case, the coefficient used to calculate the score is set so that when the m / z values for proteins belonging to a specific group match, the score is larger than when the m / z values for proteins not belonging to that group match. For example, the coefficient is set to 10 times. In other words, when the m / z values for proteins belonging to a specific group match, the score is more likely to be larger than when the m / z values for proteins not belonging to that group match, resulting in a higher precision calculation. Therefore, sample discrimination can be performed with weighting given to proteins belonging to a specific group. Thus, for example, samples can be distinguished with weighting given to functionally important and conserved proteins that are less likely to contain false peaks, thus improving the accuracy of sample discrimination.
[0125] Furthermore, in ST40A, the m / z values corresponding to proteins in a specific group may be selected from the m / z values included in the specific m / z database, or they may be selected by referring to the annotation database. Also, as mentioned above, it is not always necessary to use the specific m / z database when estimating the measurement error shown in ST36. Therefore, the construction of a specific m / z database is not necessarily required in the method of identifying microorganisms that is weighted by score coefficients.
[0126] Furthermore, a combination of methods for weighting—such as multi-stage screening using specific m / z databases and changing the score coefficients—may be used. For example, a primary screening could be performed using a specific m / z database of a certain group of proteins, and a secondary screening could be performed to distinguish samples by increasing the coefficient when the m / z matches that of a protein from another group.
[0127] [6. Experimental Examples] An example of an experiment conducted using the microbial identification system 1000 will be described.
[0128] (6-1. Database Construction) We obtained bacterial and archaeal genome sequences from the U.S. National Center for Biotechnology Information (NCBI) via an FTP server (using RefSeq v95, over 270,000 sequences). For all genome sequences, we performed gene estimation (using Prodigal) to predict gene loci and their products. The completeness and contamination of the results were estimated using checkM. We also measured the number of contigs, the number of undetermined bases (N), N50, and the number of genes relative to the genome base length using a computer. N50 is one of the indicators of the quality of the genome information (assembly), and represents the weighted average of the sequence lengths of contigs in the genome sequence assembly. N50 is the sequence length (base length) when the contigs are arranged in descending order and added up from top to bottom, reaching half the total length. For the predicted products (proteins) from the estimated genes, we performed methionine removal, signal protein prediction, and resulting cleavage fragment prediction (using SignalP) according to their amino acid composition, and calculated the mass of the predicted final protein. We digitized the theoretical protein mass information (overall m / z database), collected phylogenetic classification information for each genome (existing taxonomic information such as GTDB, Silva, and GreenGenes), and created a database where this information was linked under the same ID (classification database). Furthermore, for the predicted gene products (proteins), we estimated the function of each protein using the similarity with registration information in existing protein databases such as UniProKB and PFAM, and created a database (annotation database) where all theoretical protein masses and protein names are linked.
[0129] For each genome sequence, sequences that met any of the following criteria were considered to be of low quality and were excluded from the data: estimated completeness of 50% or less, contamination of 10% or more, number of contigs of 1,000 or more, N50 of 5kbp or less, or undetermined bases (N) of 100,000 or more. In addition, genome sequences with a ratio of (number of genes / total number of bases at a locus in the genome) of 0.00180 or more were deleted.
[0130] When a database was constructed using all genome entries without considering these criteria, and microbial identification was performed using protein measurement peak lists of known microbial strains obtained by MALDI-MS (AXIMA®, manufactured by Shimadzu Corporation), it was observed that genome entries with an extremely high number of genes per genome matched with a high probability, leading to incorrect results.
[0131] The specific experimental procedure for known microbial strains is as follows: Microbial groups obtained from the National Institute of Technology and Evaluation (NBRC), such as Escherichia coli NBRC 3301, Bacillus subtilis subsp. subtilis NBRC 13719, Microlunatus phosphovorus NBRC 101784, Bifidobacterium longum ATCC 15707, Clostridium acetobutylicum NBRC 13948, Arthrobacter globiformis NBRC 12137, Brachybacterium conglomeratum NBRC 15472, Streptomyces griseus subsp. griseus NBRC 12875, Tetrasphaera duodecadis NBRC 12959, Bacteroides fragilis ATCC 25285, and Sphingomonas yanoikuyae NBRC. Bacteria and archaea such as 15102, Xanthobacter autotrophicus NBRC 102463, Rhodobacter azotoformans NBRC 16436, Methanosarcina thermophila MST-A1, and Thauera linaloolentis NBRC 102519 were cultured in designated media. The culture medium was centrifuged (10000g, 2 min) to remove the media components, and the same amount of pure water was added to disperse the cells. The mixture was then centrifuged under the same conditions, and the supernatant was removed. 500 μL of pure water was added to the resulting cell precipitate to disperse the cells and obtain a cell dispersion. 500 μL of zirconia beads (φ0.5 mm) were added to a 1.5 mL screw-cap tube, and 500 μL of the aforementioned cell dispersion was added.The cells were crushed using a bead crusher (TOMY Seikou MS-100) at 4000 rpm for a total of 3 minutes. The crushed liquid was centrifuged (15000 g, 5 minutes), and 1 μL of the supernatant was mixed with 9 μL of 10 mg / ml CHCA (α-cyano-4-hydroxycinnamic acid) solution (50% acetonitrile aqueous solution containing 1% TFA (Trifluoroacetic acid)). 1 μL of this mixture was dropped onto a MALDI-MS sample plate and air-dried to prepare sample / matrix mixed crystals. MALDI mass spectra were obtained from these bacterial samples by measuring the m / z range of 2000-20000 in MALDI-MS linear mode. Peak picking was performed, and a peak list consisting of the m / z value and peak intensity (mV) of the detected peaks was created.
[0132] Subsequently, the database created above was used to verify the degree of agreement by comparing the peaks in the peak list with the theoretical m / z values in the database. As a result, in many cases, the measured peaks that matched within a certain range were theoretical peaks estimated from the genome information of the relevant bacterial species. However, among these, there was also genome information that showed a high degree of agreement with theoretical peaks estimated from genome information of species other than the relevant bacterial species.
[0133] Figure 10 shows the relationship between the total number of base pairs at a gene site in the genome and the estimated number of genes per genome. More specifically, Figure 10 shows the relationship between the base pair length of gene loci in each genome and the estimated number of genes, estimated from bacterial and archaeal genome information (over 270,000 entries) in RefSeq95. In Figure 10, when the relationship between the total number of base pairs and the estimated number of genes is shown for all genome entries (over 270,000 entries), many genomes were detected that were located above the dashed line in Figure 10 (the line where the total number of base pairs at a gene locus in the genome is 0.00180). It was hypothesized that this was due to false positives resulting from errors in reading the base pairs during genome sequencing, which led to the prediction of proteins that do not actually exist and the prediction of more theoretical peaks than are actually present.
[0134] On the other hand, by removing genome entries from the database with a data count (number of genes / total number of base pairs at a locus in the genome) of approximately 0.00180 or more, the results showed that the degree of agreement between the peak list from the above-mentioned microbial groups and the estimated theoretical peaks from the corresponding genome information was ranked higher. In other words, it was confirmed that appropriate evaluation can be performed by removing genome entries located above the dashed line in Figure 10 from the database. These findings indicate that constructing a database appropriately selected using the above method based on genome information registered in public databases is essential for constructing a database for microbial identification.
[0135] As a result of these selections, a total m / z database with 193,197 entries was created. In addition, a separate database representing the species level was created within it, based on GTDB, and a representative species-level database consisting of 31,760 entries was constructed.
[0136] (6-2. Algorithm Construction) Microorganisms obtained from NBRC, etc., such as Escherichia coli NBRC 3301, Bacillus subtilis subsp. globiformis NBRC 12137, Brachybacterium conglomeratum NBRC 15472, Streptomyces griseus subsp. griseus NBRC 12875, Tetrasphaera duodecadis NBRC12959, Bacteroides fragilis ATCC 25285, Sphingomonas yanoikuyae NBRC 15102, Xanthobacter autotrophicus NBRC Bacteria and archaea such as 102463, Rhodobacter azotoformans NBRC 16436, Methanosarcina thermophila MST-A1, and Thauera linaloolentis NBRC 102519 were cultured in designated media. The culture medium was centrifuged (10000g, 2 min) to remove the medium components, and the same amount of pure water was added to disperse the cells. The mixture was then centrifuged under the same conditions, and the supernatant was removed. These microbial groups include diverse lineages such as aerobic and anaerobic bacteria and methane-producing archaea, possessing diverse cell wall structures such as Gram-positive and Gram-negative, and also including actinomycetes. 500 μL of pure water was added to the precipitate of these cells to disperse the cells and obtain a cell dispersion. 500 μL of zirconia beads (φ0.5 mm) were added to a 1.5 mL screw-cap tube, and 500 μL of the aforementioned cell dispersion was added.The beads were crushed using a bead crusher (TOMY Seiko MS-100) at 4000 rpm for a total of 3 minutes. The crushed liquid was centrifuged (15000 g, 5 minutes), and 1 μL of the supernatant was mixed with 9 μL of 10 mg / ml CHCA solution (50% acetonitrile aqueous solution containing 1% TFA). 1 μL of this mixture was dropped onto a MALDI-MS sample plate and air-dried to prepare a sample / matrix mixed crystal. Next, measurements were performed using MALDI-MS (AXIMA®, Shimadzu Corporation) to obtain mass spectra for each sample strain.
[0137] The theoretical peak list in the database created above was compared with the experimentally measured peaks actually obtained from the cultured microorganisms. Peaks were considered to match if the experimentally measured peaks fell within a certain range from the theoretical peaks. Using a range of 200 ppm, the number of matching peaks was calculated for all entries in the overall m / z DB. However, it was observed that the genome entry with the highest degree of agreement was not necessarily the genome entry corresponding to the measured strain. Next, a database (specific m / z DB) was constructed, selecting proteins that are frequently detected by MALDI. Here, frequently detected ribosomal proteins were extracted based on the database (annotation DB), genome entries with high agreement to the experimentally measured peak list were selected from that database (for example, 500 to 5,000 entries were extracted), and a two-stage search algorithm was implemented to calculate the degree of agreement for these entries using the entire theoretical protein peak list. As a result, as shown in Table 1 below, an algorithm was constructed that could correctly estimate the phylogenetic group (genus, species) by selecting closely related genome entries from the experimentally measured peak list for all 15 strains mentioned above. Table 1 shows the results of classifying bacteria and archaea with various lineages, physiological characteristics, and cell wall characteristics using the algorithm.
[0138] [Table 1]
[0139] [Aspect] Those skilled in the art will understand that the above-described exemplary embodiments are specific examples of the following embodiments.
[0140] (Section 1) A method for constructing a database for identifying microorganisms according to one embodiment may include the steps of: obtaining microbial genome data from a genome database; determining whether the obtained genome data meets a criterion; predicting the expressed proteins for each genome data determined to meet the criterion; and constructing a mass-to-charge ratio database which includes a list of mass-to-charge ratios for each genome data predicted based on the predicted proteins.
[0141] According to the method for constructing a database for microbial identification described in paragraph 1, a mass-to-charge ratio database can be constructed based only on genomic data that meets the criteria in the genome database. In other words, the quality of the mass-to-charge ratio database constructed based on the genome database, which is used for microbial identification using mass spectrometry, can be improved.
[0142] (Paragraph 2) In the method for constructing a database for identifying microorganisms as described in Paragraph 1, the step of determining whether the criteria are met may include a step of determining whether the number of genes meets the criteria.
[0143] According to the method for constructing a database for microbial identification described in paragraph 2, genomic data in which the number of genes inferred from the genomic data does not fall within a reasonable range is removed and not reflected in the mass-to-charge ratio database. Therefore, the quality of the mass-to-charge ratio database is improved.
[0144] (3) In the method for constructing a database for identifying microorganisms as described in paragraph 1 or 2, the step of determining whether the criteria are met may include a step of determining based on genome integrity.
[0145] According to the method for constructing a database for microbial identification described in Section 3, if the genome data is incomplete, such as when some of the genome data is missing or misread, that genome data is removed and not reflected in the mass-to-charge ratio database. Therefore, the quality of the mass-to-charge ratio database is improved.
[0146] (Section 4) In a method for constructing a database for identifying microorganisms as described in any one of paragraphs 1 to 3, the step of determining whether the criteria are met may include a step of determining based on the rate of genome contamination.
[0147] The method for constructing a microbial identification database described in Section 4 allows for the removal of genomic data with a high rate of contamination. In other words, genomic data that is highly likely to contain mixed DNA sequences of multiple types of microorganisms will not be reflected in the mass-to-charge ratio database. Therefore, the quality of the mass-to-charge ratio database is improved.
[0148] (Clause 5) In a method for constructing a database for microbial identification as described in any one of paragraphs 1 to 4, the step of determining whether the criteria are met may include a step of determining based on the number of contigs.
[0149] The method for constructing a microbial identification database described in Section 5 allows for the removal of genomic data with a large number of contigs. Too many contigs can fragment gene regions expressing proteins, potentially preventing accurate reading. Therefore, by removing low-quality genomic data based on the number of contigs and preventing such data from being reflected in the mass-to-charge ratio database, the quality of the mass-to-charge ratio database can be improved.
[0150] (Section 6) In a method for constructing a database for identifying microorganisms as described in any one of paragraphs 1 to 5, the step of determining whether the criteria are met may include a step of determining based on the number of undetermined bases.
[0151] The method for constructing a database for microbial identification described in Section 6 allows for the removal of genomic data with a large number of undetermined bases. DNA sequences containing many undetermined bases are highly likely to fail to properly identify genes. Therefore, by removing low-quality genomic data based on the number of undetermined bases and preventing such data from being reflected in the mass-to-charge ratio database, the quality of the mass-to-charge ratio database can be improved.
[0152] (Section 7) In a method for constructing a database for microbial identification as described in any one of Sections 1 to 6, the step of constructing a mass-to-charge ratio database may include a step of linking a predicted protein or mass-to-charge ratio with information about the group in which the predicted protein is contained.
[0153] According to the method for constructing a database for microbial identification described in Section 7, it becomes possible to selectively process the mass-to-charge ratios corresponding to proteins in a specific group, based on information about the groups containing the proteins. A specific group is, for example, "a group of proteins that are likely to be expressed in the living organism of microorganisms and are likely to be detected as peaks when mass spectra are measured."
[0154] (Section 8) In the method for constructing a database for microbial identification as described in Section 7, the group information may include at least one of the protein name, protein function, and family.
[0155] The method for constructing a database for microbial identification described in Section 8 allows for the selective processing of mass-to-charge ratios corresponding to the same protein, proteins with the same function, or proteins of the same family. For example, it becomes possible to identify samples by weighting the mass-to-charge ratios corresponding to "groups of proteins that are likely to be expressed in the vivo environment of microorganisms and are likely to be detected as peaks when mass spectra are measured," based on information about the protein's name, function, or family.
[0156] (Section 9) A method for constructing a database for microbial identification as described in Section 7 or 8, the step of constructing a mass-to-charge ratio database may further include the step of constructing a specific mass-to-charge ratio database which includes a specific mass-to-charge ratio list having only mass-to-charge ratios predicted to be included in a particular group, based on information about the group.
[0157] According to the method for constructing a database for microbial identification described in Section 9, it becomes easy to selectively process the mass-charge ratios of proteins belonging to a specific group using a specific mass-charge ratio database. For example, screening based solely on the mass-charge ratios of proteins belonging to a specific group becomes possible.
[0158] (Section 10) In a method for constructing a database for identifying microorganisms as described in any one of Sections 7 to 9, a group is selected based on at least one of the following conditions: the expression level is above a predetermined threshold; the group has a function essential for maintaining life; a predetermined proportion or more of microorganisms have an amino acid sequence similarity above a predetermined threshold; the group is a basic protein; the mass-to-charge ratio can be analyzed within an error range of ±14 Da when measured by MALDI-MS; the protein mass is in the range of 4 to 30 kDa; and the group contains a predetermined number or more of protein types. The function essential for maintaining life may include a function essential for at least one of the maintenance and proliferation of cells.
[0159] According to the method for constructing a database for microbial identification described in Section 10, the mass-to-charge ratio corresponding to "proteins that are likely to be expressed in vivo and are detected as peaks when mass spectra are measured" that satisfy the above conditions can be selectively processed.
[0160] (Section 11) In a method for constructing a database for microbial identification as described in any one of paragraphs 7 to 10, the group may include at least one of a ribosomal protein, a chaperone, or a DNA-binding protein.
[0161] The method for constructing a database for microbial identification described in Section 11 allows for the selective processing of mass-to-charge ratios corresponding to proteins that are likely to be expressed in vivo and are detected as peaks when mass spectra are measured, such as ribosomal proteins, chaperones, and DNA-binding proteins. Therefore, it becomes possible to weight these proteins to identify samples.
[0162] (Section 12) In a method for constructing a database for microbial identification as described in any one of paragraphs 1 to 6, the step of constructing a mass-to-charge ratio database may include the step of constructing a total mass-to-charge ratio database that includes all predicted mass-to-charge ratios.
[0163] According to the method for constructing a database for microbial identification described in Section 12, it becomes possible to perform screening based on all mass-to-charge ratios using the overall mass-to-charge ratio database. Therefore, the similarity of proteins other than those included in a specific group can also be reflected in the selection of samples.
[0164] (Section 13) In a method for constructing a database for microbial identification as described in any one of paragraphs 7 to 11, the step of constructing a mass-to-charge ratio database may include the step of constructing a total mass-to-charge ratio database that includes all predicted mass-to-charge ratios.
[0165] According to the method for constructing a database for microbial identification described in Section 13, it becomes possible to perform screening based on all mass-to-charge ratios using the overall mass-to-charge ratio database. Therefore, the similarity of proteins other than those included in a specific group can also be reflected in the selection of samples.
[0166] (Section 14) A method for constructing a database for microbial identification as described in any one of paragraphs 1 to 13 may further include a step of obtaining classification data from a database containing classification data relating to the classification of microorganisms. The step of constructing a mass-to-charge ratio database may include a step of associating classification data with the mass-to-charge ratio database.
[0167] According to the method for constructing a microbial identification database described in Section 14, genome IDs created based on collected classification data can be associated with data contained in the overall mass-to-charge ratio database and the specific mass-to-charge ratio database, respectively. Furthermore, the collected classification data can be used to organize or incorporate into the content of the overall mass-to-charge ratio database and the specific mass-to-charge ratio database. The collected classification data can also be used for other applications in the instrument, such as determining "specific proteins known to be significantly expressed only in specific species."
[0168] (Section 15) In a method for constructing a database for microbial identification as described in any one of Sections 1 to 14, the prediction step may include the steps of predicting genes from genome data, predicting the post-translational amino acid sequence from the predicted genes, predicting post-translational modifications from the post-translational amino acid sequence, and predicting a protein with the predicted post-translational modifications.
[0169] According to the method for constructing a database for microbial identification described in Section 15, proteins that are actually expressed in vivo can be predicted from genome data. Therefore, the mass-to-charge ratio of proteins that are actually expressed in vivo can be reflected in the mass-to-charge ratio database, thereby improving the quality of the mass-to-charge ratio database.
[0170] (Section 16) A method for constructing a database for identifying microorganisms as described in any one of Sections 1 to 15 may include the steps of: acquiring new genome data; determining whether the new genome data meets the criteria; if the new genome data meets the criteria, predicting the proteins expressed from the new genome data, predicting the mass-to-charge ratios based on the prediction results, predicting a list of new mass-to-charge ratios; and adding the list of new mass-to-charge ratios to the mass-to-charge ratio database.
[0171] According to the method for constructing a microbial identification database described in Section 16, newly acquired genome data can be added to the m / zDB, thereby enriching its contents. As a result, the quality of the m / zDB is further improved, and the accuracy of sample identification using the m / zDB is also further enhanced.
[0172] (Clause 17) An apparatus for constructing a database for microbial identification according to one embodiment constructs a database for microbial identification using microbial genomic data obtained from a genome database. The apparatus comprises a processor and a storage unit. The processor determines whether the acquired genomic data meets the criteria. The processor also predicts the expressed proteins for each genomic data determined to meet the criteria. The processor also constructs a mass-to-charge ratio database containing a list of predicted mass-to-charge ratios for each genomic data based on the predicted proteins. The processor also stores the mass-to-charge ratio database in the storage unit.
[0173] The apparatus for constructing a database for microbial identification described in paragraph 17 allows for the construction of a mass-to-charge ratio database based only on genomic data that meets the criteria in the genome database. In other words, it is possible to improve the quality of mass-to-charge ratio databases constructed based on genome databases, which are used for microbial identification using mass spectrometry.
[0174] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims rather than the foregoing description, and all modifications within the meaning and scope of equivalents of the claims are intended. [Explanation of symbols]
[0175] 10 Processor, 11 Memory, 12 Communication I / F, 13 Input / Output I / F, 14 Operation Unit, 15 Display, 16 MS, 21 Genome Data Acquisition Unit, 22 Genome Data Determination Unit, 23 Protein Prediction Unit, 24 Construction Unit, 25 Storage Unit, 31 Acquisition Unit, 32 Sample Discrimination Unit, 33 Annotation Unit, 34 Output Unit, 70 Public Genome Database, 80 Public Classification Database, 90 Network, 100 Device, 101 Controller, 321 Primary Screening Unit, 322 Secondary Screening Unit, 1000 Microbial Discrimination System.
Claims
1. A method for constructing a database for microbial identification, which is executed by a computer, The steps involve obtaining microbial genome data from a genome database, The steps include determining whether the acquired genome data meets the criteria, For each genome data set that is determined to meet the criteria, the step is to predict the expressed proteins, A method for constructing a database for microbial identification, comprising the steps of constructing a mass-to-charge ratio database containing a list of mass-to-charge ratios for each genome data set, predicted based on the predicted protein.
2. A method for constructing a database for identifying microorganisms according to claim 1, wherein the step of determining whether the above criteria are met includes a step of determining whether the number of genes meets a standard value.
3. A method for constructing a microbial identification database according to claim 1 or 2, wherein the step of determining whether the above criteria are met includes a step of determining based on genome integrity.
4. A method for constructing a microbial identification database according to claim 1 or 2, wherein the step of determining whether the above criteria are met includes a step of determining based on the rate of genomic contamination.
5. A method for constructing a database for microbial identification according to claim 1 or 2, wherein the step of determining whether the above criteria are met includes a step of determining based on the number of contigs.
6. A method for constructing a database for identifying microorganisms according to claim 1 or 2, wherein the step of determining whether the above criteria are met includes a step of determining based on the number of undetermined bases.
7. A method for constructing a microbial identification database according to claim 1, wherein the step of constructing the mass-to-charge ratio database includes the step of linking predicted proteins or mass-to-charge ratios with information about the group containing the predicted proteins.
8. A method for constructing a microbial identification database according to claim 7, wherein the information relating to the group includes at least one of the protein name, protein function, and family.
9. A method for constructing a database for microbial identification according to claim 7 or 8, wherein the step of constructing the mass-to-charge ratio database further comprises the step of constructing a specific mass-to-charge ratio database which includes a specific mass-to-charge ratio list having only mass-to-charge ratios predicted to be included in a particular group, based on information about the group.
10. The group is selected based on at least one of the following conditions: the expression level is above a predetermined threshold; the group has a function essential for maintaining life; a predetermined proportion or more of the microorganisms have an amino acid sequence similarity above a predetermined threshold; the group is a basic protein; the mass-to-charge ratio can be analyzed within an error range of ±14 Da when measured by MALDI-MS; the protein mass is within 4 to 30 kDa; and the group contains a predetermined number or more of different protein types. A method for constructing a database for identifying microorganisms according to claim 7 or 8, wherein the functions essential for maintaining life include functions essential for at least one of cell maintenance and proliferation.
11. The group comprises at least one ribosomal protein, a chaperone, and a DNA-binding protein, and is a method for constructing a database for microbial identification according to claim 7 or 8.
12. The method for constructing a database for microbial identification according to claim 1 or 2, wherein the step of constructing the mass-to-charge ratio database includes the step of constructing an overall mass-to-charge ratio database that includes all predicted mass-to-charge ratios.
13. The method for constructing a database for microbial identification according to claim 7 or 8, wherein the step of constructing the mass-to-charge ratio database further comprises the step of constructing an overall mass-to-charge ratio database that includes all predicted mass-to-charge ratios.
14. The process further includes a step of obtaining classification data from a database containing classification data related to the classification of microorganisms. A method for constructing a database for identifying microorganisms according to claim 1 or 2, wherein the step of constructing the mass-to-charge ratio database includes the step of associating the classification data with the mass-to-charge ratio database.
15. The aforementioned prediction step is, Steps to predict genes from genome data, The steps include predicting the post-translational amino acid sequence from the predicted gene, A step to predict post-translational modifications from the post-translated amino acid sequence, A method for constructing a database for microbial identification according to claim 1 or 2, comprising the step of predicting a protein with predicted post-translational modifications.
16. The steps include acquiring new genome data for the aforementioned mass-to-charge ratio database, The steps include determining whether the new genome data meets the criteria, If the new genomic data meets the criteria, the steps include predicting the proteins expressed from the new genomic data, predicting the mass-to-charge ratio based on the prediction results, and predicting a new list of mass-to-charge ratios. A method for constructing a database for identifying microorganisms according to claim 1 or 2, comprising the step of adding the new list of mass-to-charge ratios to the mass-to-charge ratio database.
17. A device for constructing a database for microbial identification using microbial genome data obtained from a genome database, Processor and Equipped with a memory unit, The aforementioned processor, Determine whether the acquired genome data meets the criteria. For each genome data set that is determined to meet the criteria, the expressed proteins are predicted. We constructed a mass-to-charge ratio database containing a list of mass-to-charge ratios for each genome dataset, based on predicted proteins. A device for constructing a database for microbial identification, which stores the mass-to-charge ratio database in the memory unit.