Method for constructing database for discrimination of microorganisms, recording medium, device for constructing database for discrimination of microorganisms, program, method for discriminating microorganisms, and system for discriminating microorganisms

The method addresses the challenge of distinguishing bacterial species undifferentiable by MALDI-MS by predicting protein masses from genome data, reducing the operational burden through theoretical calculations.

WO2026004532A1PCT designated stage Publication Date: 2026-01-02SHIMADZU CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/020556
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-06
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing microbial identification methods using MALDI-MS struggle to distinguish bacterial species that are classified differently by ANI analysis, leading to a burden on operators due to the need for multiple mass spectrum analyses under varying conditions.

Method used

A method and system that utilize genome data to predict protein production and calculate theoretical mass-to-charge ratios, allowing identification of indistinguishable bacterial species by MALDI-MS without actual mass spectrometry measurements, reducing the operational burden.

Benefits of technology

Reduces the need for repeated mass spectrum analyses by identifying bacterial species based on theoretical protein mass values, independent of culture and sample preparation conditions, thereby easing the workload on operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025020556_02012026_PF_FP_ABST
    Figure JP2025020556_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for constructing a database for discrimination of microorganisms according to the present disclosure includes: a step (S12) for acquiring genome data of two kinds of microorganisms; a step (S14) for predicting a group of proteins produced by each of the two kinds of microorganisms; a step (S16) for producing a list of mass-charge ratios of each of the two kinds of microorganisms; a step (S20) for calculating the degree of similarity between the lists of the mass charge ratios; a step (S32) for generating information that includes the fact that the two kinds of microorganisms cannot be discriminated by MALDI-MS when the degree of similarity is equal to or larger than a predetermined value; and a step (S36) for outputting the information.
Need to check novelty before this filing date? Find Prior Art

Description

Method for constructing a database for microorganism discrimination, recording medium, device for constructing a database for microorganism discrimination, program, method for microorganism discrimination, and system for microorganism discrimination

[0001] The present invention relates to a method for constructing a database for microbial discrimination, a recording medium, an apparatus for constructing a database for microbial discrimination, a program, a method for microbial discrimination, and a system for microbial discrimination, and more particularly to the identification of bacterial species that cannot be distinguished by analysis using mass spectrometry.

[0002] The DNA sequence of each strain is used to classify bacterial strains into species, which are the basic unit of biological classification. For example, as disclosed in Goris J, Konstantinidis KT, Klappenbach JA, Coenye T, Vandamme P, Tiedje JM. DNA-DNA hybridization values ​​and their relationship to whole-genome sequence similarities. Int J Syst Evol Microbiol. 2007 Jan;57(Pt 1):81-91. (Non-Patent Document 1) and Jain C, Rodriguez-R LM, Phillippy AM, Konstantinidis KT, Aluru S. High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat Commun. 2018 Nov 30;9(1):5114. (Non-Patent Document 2), DNA-DNA hybridization (DDH) analysis and ANI (Average Nucleotide Identity) analysis are known as methods for determining whether two bacterial strains are of the same species.

[0003] In ANI analysis, the DNA sequence of the analyzed strain is fragmented into a fixed base length (e.g., 1020 bp) on a computer, and a similarity search is performed for each fragment against the DNA base sequence of the comparison strain. The ANI value between the DNA base sequences of the analyzed strain and the comparison strain is calculated based on the average similarity value calculated by the similarity search. If the ANI value is 95% or higher, the analyzed strain and the comparison strain are identified as being of the same species.

[0004] On the other hand, according to traditional microbial classification systems, groups of microorganisms that are considered to be the same species according to the above criteria may be listed as different species. To avoid such confusion, progress is being made in the construction of microbial classification systems that use genomic information, and their use is increasing.

[0005] In recent years, methods for identifying microorganisms contained in a sample using matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) have been developed. However, this method, which analyzes a sample using MALDI-MS and uses the mass spectrum obtained as an index to identify species, may not be able to distinguish two strains that are classified as different species.

[0006] Regarding bacterial species that cannot be identified using a microbial identification method using MALDI-MS, M58: Method for the Identification of Cultured Microorganisms Using Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry; 1st Edition, Clinical and Laboratory Standards Institute, 2017. (Non-Patent Document 3) discloses a group of bacterial species that cannot be identified using MALDI-MS.

[0007] Goris J, Konstantinidis KT, Klappenbach JA, Coenye T, Vandamme P, Tiedje JM. DNA-DNA hybridization values ​​and their relationship to whole-genome sequence similarities. Int J Syst Evol Microbiol. 2007 Jan;57(Pt 1):81-91.Jain C, Rodriguez-R LM, Phillippy AM, Konstantinidis KT, Aluru S. High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat Commun. 2018 Nov 30;9(1):5114.M58: Method for the Identification of Cultured Microorganisms Using Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry; 1st Edition, Clinical and Laboratory Standards Institute, 2017.

[0008] As disclosed in Non-Patent Document 3, by creating a list of bacterial species that cannot be distinguished by MALDI-MS in advance, when a bacterial species on the list is determined to be contained in a sample, the user can recognize that the sample may contain bacterial species other than the identified bacterial species. However, it was not clear whether the difficult-to-identify bacterial species on the list could not be distinguished because they were deemed to be nearly identical in ANI analysis due to deficiencies in the old classification system, or whether they were target bacterial species that could be determined to be different in ANI analysis but could not be distinguished by MALDI-MS. Therefore, analyzing all bacterial species by MALDI-MS and identifying the indistinguishable bacterial species may be a burden on the operator.

[0009] The present disclosure has been made in view of the above circumstances, and its purpose is to reduce the burden of the work involved in identifying bacterial species that cannot be identified by a microorganism identification method using MALDI-MS.

[0010] A method for constructing a database for microbial discrimination according to a first aspect of the present disclosure includes the steps of: (a) acquiring genome data for each of two types of microorganisms; (b) predicting a group of proteins produced by each of the two types of microorganisms based on the respective genome data; (c) calculating theoretical values ​​of observed masses measured when the group of proteins is analyzed by MALDI-MS, and creating a list of mass-to-charge ratios for each of the two types of microorganisms; (d) calculating the similarity between the lists of mass-to-charge ratios; (e) generating information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS if the similarity is equal to or greater than a predetermined value; and (f) outputting the information.

[0011] An apparatus for constructing a database for microbial discrimination according to a second aspect of the present disclosure includes a processor and a storage unit. The processor acquires genome data for each of two types of microorganisms and predicts a group of proteins produced by each of the two types of microorganisms based on the genome data. The processor calculates theoretical values ​​of observed masses measured when the protein groups are analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and creates a list of mass-to-charge ratios for each of the two types of microorganisms. The processor calculates the similarity between each list of mass-to-charge ratios. If the similarity is equal to or greater than a predetermined value, the processor generates information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS, and stores the information in the storage unit.

[0012] A program according to a third aspect of the present disclosure is executed by a processor mounted on a computer, causing the computer to (a) acquire genome data for each of two types of microorganisms; (b) predict, based on the genome data, a group of proteins produced by each of the two types of microorganisms; (c) calculate theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and create a list of mass-to-charge ratios for each of the two types of microorganisms; (d) calculate the similarity between the lists of mass-to-charge ratios; (e) generate information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS if the similarity is equal to or greater than a predetermined value; and (f) output the information.

[0013] According to the method for constructing a database for microorganism discrimination according to the present disclosure, it is possible to reduce the burden of the work of identifying bacterial species that cannot be discriminated by the microorganism discrimination method using MALDI-MS.

[0014] FIG. 1 is a schematic diagram showing the configuration of a microorganism discrimination system according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of information used by an analysis device when creating a discrimination database. FIG. 3 is a diagram showing an example of information stored in a discrimination database. FIG. 4 is a flowchart showing processing performed by a processor in an analysis device. FIG. 5 is a flowchart showing the subroutine of step S200 shown in FIG. 5. FIG. 6 is a flowchart showing the subroutine of step S400 shown in FIG. 4. FIG. 1 is a schematic diagram showing the configuration of a microorganism discrimination system according to a modified example. FIG. 2 is a flowchart showing the subroutine of step S200 according to a modified example. FIG. 3 is a flowchart showing the subroutine of step S400 according to a modified example. FIG. 4 is a diagram showing an example of calculation of similarity between lists of theoretical values ​​of observed mass. FIG. 5 is a diagram showing an example of a microorganism determined to be indistinguishable by MALDI-MS by the method for constructing a microorganism discrimination database according to the present embodiment.

[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail below with reference to the accompanying drawings. In the following description, the same or corresponding parts in the drawings are denoted by the same reference numerals, and their description will not be repeated in principle.

[0016] [Configuration of Microorganism Discrimination System] Figure 1 is a schematic diagram showing the configuration of a microorganism discrimination system 150 according to an embodiment of the present invention. The microorganism discrimination system 150 discriminates the type of microorganism contained in a sample based on a mass spectrum obtained by analyzing the sample by mass spectrometry. In this specification, microorganisms include, for example, bacteria, archaea, fungi, protozoa, algae, and viruses. In this specification, the "type" of a microorganism includes at least one of the genotype, strain, or rank of a phylogenetic taxonomic group such as subspecies, species, genus, or family of the microorganism.

[0017] 1, a microorganism discrimination system 150 includes a mass spectrometer 16, a public genome database 70, a network 90, and a processing device 100. In this specification, "database" is also referred to as "DB."

[0018] The public genome DB 70 is a database containing genome data of organisms. A genome is the genetic information on nucleic acids (deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)) possessed by an organism, and includes the base sequence of the nucleic acid. In this specification, genome data mainly refers to DNA sequences.

[0019] The public genome DB 70 is a DB containing a large amount of genome data of organisms that is publicly available. Examples of the public genome DB 70 include genome DBs of the National Center for Biotechnology Information (NCBI), the DNA Data Bank of Japan (DDBJ), and the European Molecular Biology Laboratory (EMBL). The public genome DB 70 is not limited to these, and may include, for example, genome DBs that are not publicly available.

[0020] The network 90 is a network through which the processing device 100 communicates with the public genome DB 70. The network 90 is, for example, the Internet.

[0021] The processing device 100 creates a discrimination DB 111, which is a database for microorganism discrimination used when discriminating the type of microorganism contained in a sample. In this specification, discriminating the type of microorganism refers to taxonomically identifying the microorganism. Discriminating a microorganism means, for example, identifying at least one of the genus, species, strain, and lineage of the microorganism. The processing device 100 corresponds to one example of an "apparatus for constructing a database for microorganism discrimination." The information contained in the discrimination DB 111 will be described later.

[0022] The processing device 100 also identifies the type of microorganism in the sample based on mass spectrum data obtained by mass spectrometry. Therefore, the processing device 100 is an example of a device that performs a microorganism discrimination method using a microorganism discrimination database constructed by a method for constructing a microorganism discrimination database.

[0023] The processing device 100 includes a controller 101, an input unit 14, and an output unit 15. The input unit 14 and the output unit 15 are connected to the controller 101. The processing device 100 is, for example, a computer. Note that the processing device 100 does not need to be configured by a single computer, and may be configured by multiple computers.

[0024] The controller 101 includes, as its main components, a processor 10, a memory 11, a communication interface (I / F) 12, and an input / output I / F 13. These components are connected to each other via a bus so as to be able to communicate with each other.

[0025] The processor 10 is an example of an electric circuit, and controls the operation of the processing device 100 by executing a given program. The program executed by the processor 10 may be stored in the memory 11, or may be stored in a storage device external to the processing device 100. The processor is, for example, a CPU (Central Processing Unit) or an MPU (Micro Processing Unit).

[0026] The memory 11 non-temporarily stores programs executed by the processor 10, mass spectrum data obtained by mass analysis, and databases. The databases and programs stored in the memory 11 include a discrimination DB 111 and an analysis program 112. The memory 11 includes volatile memory (e.g., RAM (Random Access Memory)) and non-volatile memory (e.g., ROM (Read Only Memory), a hard disk drive, and a solid state drive). The databases and / or programs may be stored in an external storage device accessible by the processor 10.

[0027] The communication I / F 12 is a communication interface for exchanging various data with external devices including the public genome DB 70 via the network 90. ​​The communication I / F 12 is realized by, for example, a network adapter. The communication method may be wireless communication such as Bluetooth (registered trademark) or wireless LAN, or wired communication using a USB (Universal Serial Bus) or the like.

[0028] The input / output I / F 13 is an interface for exchanging various types of data between the processor 10 and an external device connected to the input / output I / F 13. The external device includes an input unit 14 and an output unit 15. A mass spectrometer 16 is connected to the input / output I / F 13. In this specification, the input / output I / F 13 also includes a device that exchanges data between the processor 10 and a storage terminal connected to the processing device 100.

[0029] The mass spectrometer 16 is a device for performing mass analysis of components contained in a sample. Analysis in the mass spectrometer 16 generates mass spectrum data and includes measuring the mass-to-charge ratio (m / z) of substances contained in the sample. A mass spectrum is a graph in which the mass-to-charge ratio is plotted on the horizontal axis and the detected peak intensity is plotted on the vertical axis.

[0030] The mass analyzer 16 is a device for performing mass analysis of components contained in a sample, and may be, for example, a MALDI-TOF MS (Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry), a MALDI-IT-TOF MS (Matrix-Assisted Laser Desorption / Ionization Ion Trap Time-of-Flight Mass Spectrometry), or a scanning IT-MS, but is not limited to these. When the mass analyzer 16 is a MALDI-TOF MS, ions generated by laser irradiation are drawn into a flight tube, separated according to their time of flight, and then detected. The time of flight correlates with the mass-to-charge ratio of the components. As a result, a mass spectrum is obtained, with m / z on the horizontal axis and detected peak intensity on the vertical axis.

[0031] The mass spectrometer 16 performs mass analysis of proteins in a sample to generate mass spectrum data. In the mass spectrum data, peaks appear according to the mass-to-charge ratios of the proteins in the sample. Therefore, by referring to a list of mass-to-charge ratios that produce peaks with heights equal to or greater than a predetermined threshold in the mass spectrum data, the proteins contained in the sample can be identified.

[0032] After performing mass analysis of the sample, the mass spectrometer 16 transmits the mass spectral data of the sample and / or a list of mass-to-charge ratios in the mass spectral data of the sample to the processing device 100. The processor 10 identifies the microorganisms contained in the sample based on the mass spectral data of the sample and / or the list of mass-to-charge ratios in the mass spectral data of the sample.

[0033] The processing device 100 may control the mass spectrometer 16, or another control device (e.g., a computer) may be connected to the mass spectrometer 16, and the mass spectrometer 16 may be controlled by that control device.

[0034] [Comparative Example] When classifying bacterial strains into species, which are the basic unit of biological taxonomy, the DNA base sequence of each bacterial strain is used. As methods for classifying species using DNA base sequences, DDH analysis and ANI analysis are known, as shown in Non-Patent Documents 1 and 2.

[0035] In ANI analysis, the entire DNA sequence of the strain being analyzed is fragmented into a fixed base length (e.g., 1020 bp) on a computer, and a similarity search is performed for each fragment against the DNA sequence of the comparison strain.The ANI value between the genome sequences of the analysis strain and the comparison strain is calculated from the average of the similarity values ​​calculated from these.If the calculated ANI value is 95-97% or higher, the analysis strain and the comparison strain are identified as being of the same species.

[0036] In recent years, methods for distinguishing microorganisms in a sample using MALDI-MS have been developed. These methods can be performed more quickly and at lower cost than methods for distinguishing microorganisms using DNA sequences as an indicator. However, a method for distinguishing microorganisms using a mass spectrum obtained by analyzing a sample using MALDI-MS as an indicator may not be able to distinguish two strains that are classified as different species by ANI analysis. For example, consider a case where there are two strains that are classified as different species by ANI analysis. In this case, the DNA base sequences of the two strains are different. However, even if the DNA base sequences of the two strains are different, the mass numbers of the proteins produced by the two strains may be the same. In this case, because MALDI-MS cannot distinguish between different types of proteins, a method for distinguishing microorganisms using MALDI-MS cannot determine that the two strains are different species.

[0037] Regarding bacterial species that cannot be distinguished by a microbial discrimination method using MALDI-MS, Non-Patent Document 3 performs discrimination of microorganisms using MALDI-MS and shows a group of bacterial species that could not be distinguished. For example, Non-Patent Document 3 discloses that the bacterial species Bacillus cereus and Bacillus anthracis are bacterial species that cannot be distinguished by a discrimination method using MALDI-MS. Therefore, when a user obtains a result indicating that a sample contains Bacillus cereus using a discrimination method using MALDI-MS, the user needs to be aware that the sample may contain Bacillus anthracis rather than Bacillus cereus.

[0038] Until now, the task of identifying bacterial species that are classified as different species by ANI analysis but cannot be distinguished by a discrimination method using MALDI-MS has been carried out based on mass spectral data obtained by MALDI-MS. However, analyzing all bacterial species by MALDI-MS to identify the indistinguishable species can be a burden on the worker.

[0039] Furthermore, even if samples contain the same type of microorganism, different microbial culture conditions or sample preparation conditions can result in different mass spectra. Therefore, when identifying a bacterial species that cannot be distinguished by MALDI-MS based on mass spectrum data, it is necessary to obtain multiple mass spectrum data for a single type of microorganism by changing the microbial culture conditions or sample preparation conditions, which can place a heavy burden on the operator.

[0040] [Microorganism discrimination system according to the embodiment] Therefore, the microorganism discrimination system 150 according to the present embodiment identifies bacterial species that cannot be distinguished by MALDI-MS based on theoretical values ​​of observed protein masses predicted from DNA base sequences, rather than on measurement data. The microorganism discrimination system 150 can identify bacterial species that cannot be distinguished by MALDI-MS without analyzing each bacterial species by MALDI-MS. This reduces the burden on the operator of acquiring mass spectra for each bacterial species.

[0041] Furthermore, the theoretical value of the observed mass of a protein does not depend on the culture conditions of the microorganism or the preparation conditions of the sample. It is predicted that one group of proteins will be produced from one type of DNA base sequence. Therefore, the theoretical value of the observed mass in the mass spectrum obtained by analyzing a microorganism having that DNA base sequence by MALDI-MS can be uniquely determined. The operator can identify bacterial species that cannot be distinguished by MALDI-MS without considering the culture conditions of the microorganism and the preparation conditions of the sample. This reduces the burden on the operator of changing the culture conditions and the preparation conditions of the sample for one type of microorganism and performing MALDI-MS analysis.

[0042] [Processing Related to Database Construction] The processing of constructing a database by the processing device 100 will be described with reference to Fig. 2. Fig. 2 is a diagram for explaining the information used by the processing device 100 when constructing a database for microorganism discrimination. Fig. 2 shows information used in the process of determining whether or not three types of microorganisms, microorganism X, microorganism Y, and microorganism Z, can be distinguished by MALDI-MS, and the results of that process. Note that microorganism X, microorganism Y, and microorganism Z are microorganisms that are classified as different species in ANI analysis.

[0043] <1. Obtaining Genome Data> The processing device 100 obtains genome data of a microorganism from the public genome DB 70.

[0044] As shown in FIG. 2, the processing device 100 acquires genome data x of a microorganism X, genome data y of a microorganism Y, and genome data z of a microorganism Z from the public genome DB 70.

[0045] 2. Estimation of species-level taxa based on genome sequences. ANI values ​​are calculated based on the genome data of microorganisms, and species-level taxa are estimated based on the genome sequences. Specifically, ANI values ​​are calculated brute-force based on the collected genome sequences, and grouping is performed based on the 95% ANI value using software such as dRep. Alternatively, microbial species are organized using a classification system based on genome information, such as GENOME TAXONOMY DATABASE (GTDB) (https: / / gtdb.ecogenomic.org). The estimation of species-level taxa based on genome sequences is an example of identifying the species-level taxa of organisms corresponding to each genome data at the genome level based on each genome data.

[0046] 3. Prediction of Produced Proteins The processing device 100 predicts proteins produced by a microorganism based on the genome data of the microorganism. Specifically, it predicts gene coding regions from the genome data and predicts proteins produced from the genes. Generally, genome data includes coding regions for multiple genes. Therefore, it is possible to predict that multiple types of proteins will be produced from a single genome data set.

[0047] 2, the processing device 100 predicts, based on genome data x, that microorganism X will produce proteins A, B, and C. Similarly, the processing device 100 predicts, based on genome data y of microorganism Y, that microorganism Y will produce proteins D, E, and F, and, based on genome data z of microorganism Z, that microorganism Z will produce proteins G, H, and I.

[0048] 4. Creating a Mass-to-Charge Ratio List The processing device 100 creates a list of mass-to-charge ratios of peaks in mass spectrum data generated when a sample containing each microorganism is analyzed by MALDI-MS, based on the molecular weights of proteins predicted to be produced by each microorganism. The list of mass-to-charge ratios is a list of theoretical values ​​of observed masses when proteins produced by the corresponding microorganism are analyzed by MALDI-MS. The theoretical value of the observed mass of a protein is, for example, the molecular weight of the protein plus the masses of proton ions, sodium ions, and / or potassium ions added to the protein during mass spectrometry.

[0049] 2, the processing device 100 calculates the theoretical value of the observed mass corresponding to each mass-to-charge ratio obtained when proteins A to I are analyzed by MALDI-MS. For example, the theoretical observed mass of protein A is calculated to be 500, the theoretical observed mass of protein B is calculated to be 1000, and the theoretical observed mass of protein C is calculated to be 1500. In this case, the list of mass-to-charge ratios for microorganism X is "500, 1000, 1500." Furthermore, if the theoretical observed mass of protein D is calculated to be 400, the theoretical observed mass of protein E is calculated to be 1200, and the theoretical observed mass of protein F is calculated to be 1600, the list of mass-to-charge ratios for microorganism Y is calculated to be "400, 1200, 1600." Furthermore, if it is calculated that the theoretical observed mass of protein G is 500, the theoretical observed mass of protein H is 1000, and the theoretical observed mass of protein I is 1500, then the list of mass-to-charge ratios for microorganism Z would be "500, 1000, 1500".

[0050] 5. Grouping Based on Mass-to-Charge Ratio List The processing device 100 groups microorganisms based on the mass-to-charge ratio list obtained for each microorganism. First, the processing device 100 numbers the n target microorganisms registered in the discrimination DB 111 from 1 to n. Next, the processing device 100 extracts lists of mass-to-charge ratios for a first microorganism and a second microorganism. The processing device 100 calculates the similarity between the list of mass-to-charge ratios for the first microorganism and the list of mass-to-charge ratios for the second microorganism. If the similarity is equal to or greater than a predetermined threshold, the processing device 100 determines that the list of mass-to-charge ratios for the first microorganism and the list of mass-to-charge ratios for the second microorganism are similar, and assigns the first microorganism and the second microorganism to the same group. If the similarity is less than the predetermined threshold, the processing device 100 determines that the list of mass-to-charge ratios for the first microorganism and the list of mass-to-charge ratios for the second microorganism are dissimilar, and assigns the first microorganism and the second microorganism to different groups. The processing device 100 compares the lists of mass-to-charge ratios for all combinations of two microorganisms selected from the n microorganisms, and divides the n microorganisms into groups.

[0051] 2, the processing device 100 assigns microorganism X as number 1, microorganism Y as number 2, and microorganism Z as number 3. The processing device 100 then extracts a list of mass-to-charge ratios of microorganism X and a list of mass-to-charge ratios of microorganism Y. Since the list of mass-to-charge ratios of microorganism Y, "400, 1200, 1600," differs from the list of mass-to-charge ratios of microorganism X, "500, 1000, 1500," the processing device 100 assigns microorganism X and microorganism Y to separate groups, a first group and a second group, respectively.

[0052] Next, the processing device 100 extracts a list of mass-to-charge ratios of microorganism X and a list of mass-to-charge ratios of microorganism Z. The list of mass-to-charge ratios of microorganism Z, "500, 1000, 1500," is the same as the list of mass-to-charge ratios of microorganism X, "500, 1000, 1500." Therefore, the processing device 100 assigns microorganism Z to the first group, which is the same group as microorganism X.

[0053] The processing device 100 then extracts a list of mass-to-charge ratios of microorganism Y and a list of mass-to-charge ratios of microorganism Z. The list of mass-to-charge ratios of microorganism Y, "400, 1200, 1600," is different from the list of mass-to-charge ratios of microorganism Z, "500, 1000, 1500." Therefore, the processing device 100 separates microorganism Y and microorganism Z into separate groups.

[0054] 6. Determining Bacterial Species Cannot Be Identified by MALDI-MS By sorting n microorganisms, multiple groups are generated. The list of mass-to-charge ratios included in a group containing one assigned microorganism is different from the list of mass-to-charge ratios of other types of microorganisms. In other words, the mass spectral data obtained when analyzing the microorganism using MALDI-MS is different from the mass spectral data obtained when analyzing other types of microorganisms using MALDI-MS. Therefore, the processing device 100 determines that the microorganism is a bacterial species that can be identified by MALDI-MS. On the other hand, in a group containing two or more assigned microorganisms, each of the lists of mass-to-charge ratios included in the group is identical to the lists of mass-to-charge ratios of the other microorganisms included in the group. In other words, it is assumed that the mass spectral data obtained when analyzing two or more microorganisms included in the group using MALDI-MS is identical. Therefore, the processing device 100 determines that the two or more microorganisms are bacterial species that cannot be identified by MALDI-MS.

[0055] The processing device 100 outputs and stores group information including the grouping results and information on whether the classification was successful to a classification DB 111 in the memory 11. The group information and the classification DB 111 including the group information may be non-temporarily stored in a computer-readable, removable recording medium.

[0056] 2, since microorganism X and microorganism Z belong to the first group, the processing device 100 determines that microorganism X and microorganism Z are species that cannot be distinguished by MALDI-MS. Since the second group does not include microorganisms other than microorganism Y, the processing device 100 determines that microorganism Y is a species that can be distinguished by MALDI-MS.

[0057] By performing the processes described above in <1> to <6>, the processing device 100 can identify bacterial species that cannot be identified by MALDI-MS based on the theoretical values ​​of the observed masses of proteins estimated from genome data.

[0058] In this embodiment, the proteins included in the list of mass-to-charge ratios for each microorganism are all proteins predicted to be produced by each microorganism, but are not limited to this. The proteins included in the list of mass-to-charge ratios for each microorganism may be, for example, specific types of proteins. In MALDI-MS analysis, peaks corresponding to only a portion of the proteins produced by a microorganism may be observed. For example, ions derived from proteins produced in small amounts by a microorganism or proteins that are difficult to ionize may not be detected by MALDI-MS. Even if microorganisms V and W are separated into different groups based on the theoretical observed masses of proteins difficult to detect by MALDI-MS, the proteins that caused the separation into different groups cannot be detected by MALDI-MS. Therefore, the mass spectral data obtained from a sample containing microorganism V and the mass spectral data obtained from a sample containing microorganism X may be identical. In such cases, by identifying bacterial species that cannot be distinguished by MALDI-MS based only on specific types of proteins that are clearly detected by MALDI-MS, it is possible to prevent a bacterial species that cannot be distinguished by MALDI-MS from being determined to be distinguishable by MALDI-MS.

[0059] The specific types of proteins mentioned above may be proteins called biomarkers. Biomarkers are genes that are essential for the survival of an organism and are not prone to mutation. Examples of specific types of proteins include ribosomal proteins, chaperone proteins, and DNA-binding proteins.

[0060] Proteins included in the list of mass-to-charge ratios for each microorganism may be selected based on the magnitude of the theoretical value of the observed mass. For example, proteins whose theoretical observed mass falls within the m / z range of 1000-30000, which is a mass-to-charge ratio range with high measurement sensitivity in MALDI-MS analysis, may be selected as proteins to be included in the list of mass-to-charge ratios. Furthermore, for example, the mass-to-charge ratios of most ions derived from ribosomal proteins fall within the m / z range of 4000-20000. Furthermore, ribosomal proteins in the m / z range of 4000-10000 have high detection intensities in MALDI-MS. Therefore, proteins included in the list of mass-to-charge ratios for each microorganism may fall within the m / z range of 4000-20000 or the m / z range of 4000-10000.

[0061] In this embodiment, the similarity between two mass-to-charge ratio lists is calculated based on the proportion of the theoretical observed masses of proteins included in the list of mass-to-charge ratios of the target microorganism that match the theoretical observed masses of proteins included in the list of mass-to-charge ratios of the microorganism being compared.

[0062] For example, assume that the similarity between List Q, which is a list of mass-to-charge ratios of microorganism P, and List S, which is a list of mass-to-charge ratios of microorganism R, is calculated. List Q and List S each contain ten values ​​corresponding to the theoretical observed masses of proteins produced by the corresponding microorganisms. The processing device 100 extracts a first theoretical value, which is one value, from the ten theoretical values ​​contained in List Q and compares the first theoretical value with each of the ten theoretical values ​​contained in List S. If a theoretical value contained in List S matches the first theoretical value, the processing device 100 determines that the first theoretical value matches the theoretical observed mass of the protein contained in List S, which is the comparison target. Similarly, each of the other nine theoretical values ​​contained in List Q is compared with each of the ten theoretical values ​​contained in List S. If it is determined that nine of the ten theoretical values ​​contained in List Q match the theoretical values ​​in List S, the processing device 100 determines that the similarity between List Q and List S is 90%. If the similarity between the two lists is equal to or greater than a predetermined value, the processing device 100 determines that the two mass-to-charge ratio lists are similar and determines that the two corresponding types of microorganisms are bacterial species that cannot be distinguished by MALDI-MS.

[0063] If the number of theoretical values ​​contained in the two lists does not match, the processing device 100 may calculate the similarity between the two lists using the above-mentioned processing, or may determine that the two lists are not similar without calculating the similarity between the two lists.

[0064] In this embodiment, when comparing each theoretical observed mass of a protein included in the list of mass-to-charge ratios of a target microorganism with each theoretical observed mass of a protein included in the list of mass-to-charge ratios of a comparison microorganism, the two theoretical values ​​are deemed to match only if the two theoretical values ​​match. However, this is not limited to this. A predetermined tolerance range may be set for the theoretical observed mass of a protein included in the list of mass-to-charge ratios of a comparison microorganism, and the processing device 100 may determine that the two theoretical values ​​match if the theoretical observed mass of a protein included in the list of mass-to-charge ratios of a target microorganism falls within the predetermined tolerance range. The predetermined tolerance range can be determined by the user based on the mass accuracy and mass resolution of the mass spectrometer. The predetermined tolerance range is preferably 10 to 1000 ppm, and more preferably 10 to 800 ppm. For example, when comparing a protein with a theoretical observed mass of 500.00 and a protein with a theoretical observed mass of 499.95, if the tolerance range is 10 ppm, the theoretical observed masses are determined to not match. Here, if the tolerance is set to 1000 ppm, it is determined that the theoretical values ​​of the two observed masses match.

[0065] [Discrimination process using a microorganism discrimination database] The process in which the processing device 100 discriminates microorganisms contained in a sample using a microorganism discrimination database will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of information stored in the discrimination DB 111. As shown in Fig. 3, the discrimination DB 111 includes types of microorganisms, reference data, and group information.

[0066] The processing device 100 receives mass spectral data derived from a sample from the mass spectrometer 16. The processing device 100 executes the analysis program 112 and estimates the target microorganism, which is a microorganism contained in the sample, based on the received mass spectral data. Specifically, the processing device 100 calculates the similarity between the received mass spectral data and the reference data in the discrimination DB 111, and determines the type of microorganism corresponding to the reference data with the highest similarity as the target microorganism. The reference data may be, for example, mass spectral data obtained by analyzing a sample containing microorganisms using MALDI-MS and / or data on the theoretical observed mass of proteins predicted from genome data. If there is no reference data with a similarity to the received mass spectral data that is equal to or greater than a predetermined value, the processing device 100 determines that "the sample does not contain microorganisms."

[0067] Next, the processing device 100 refers to the group information of the target microorganism. If the target microorganism is a bacterial species that can be identified by MALDI-MS, the processing device 100 determines that "the target microorganism is contained in the sample." On the other hand, if the target microorganism is a bacterial species that cannot be identified by MALDI-MS, the processing device 100 determines that "the microorganism contained in the sample is a bacterial species that cannot be identified by MALDI-MS."

[0068] The processing device 100 displays on the output unit 15 one of the following judgment results: "The sample does not contain microorganisms," "The sample contains the target microorganisms," or "The microorganisms contained in the sample are a bacterial species that cannot be identified by MALDI-MS."

[0069] In FIG. 3 , the processing device 100 receives mass spectral data derived from a sample from the mass spectrometer 16 and calculates the degree of similarity with each of the reference data L, M, and N. If the received mass spectral data is most similar to the reference data L and this similarity is equal to or greater than a predetermined value, the processing device 100 determines that the microorganism X is the target microorganism. The processing device 100 references the group information for the microorganism X and confirms that the microorganism X is a bacterial species that cannot be distinguished by MALDI-MS. Therefore, the processing device 100 determines that "the microorganism contained in the sample is a bacterial species that cannot be distinguished by MALDI-MS." In this case, specifically, the sample may contain the microorganism Z rather than the microorganism X, and the processing device 100 cannot distinguish between the microorganism X and the microorganism Z based on the mass spectral data obtained by MALDI-MS.

[0070] As described above, the processing device 100 notifies the user that the microorganisms contained in the sample are microorganisms that cannot be identified by MALDI-MS, so that the user can easily understand this. The user can identify the microorganisms contained in the sample by a method other than MALDI-MS (for example, a method based on DNA sequences).

[0071] [Processing Flow] The flow of processing performed in the processing device 100 will be described below. Fig. 4 is a flowchart showing the microorganism analysis processing performed by the processor 10 in the processing device 100. In one implementation example, the processing in Fig. 4 is started by starting an analytical application program in the processing device 100. The microorganism analysis processing performed by the processor 10 in the processing device 100 includes a database construction process and a discrimination process.

[0072] 4, in step S100, the processing device 100 determines whether or not it has received an instruction to construct a discrimination DB 111, which is a database used for discriminating microorganisms. In one implementation example, when an analytical application program is started, a start-up screen is displayed on the output unit 15. The start-up screen includes one or more keys for selecting a menu. When a key for constructing a database is operated on the start-up screen of the output unit 15, an instruction to construct a database is input to the processing device 100.

[0073] If processor 100 determines that an instruction to construct a database has been received (YES in step S100), the control proceeds to step S200, and if not (NO in step S100), the control proceeds to step S300.

[0074] In step S200, processing device 100 executes a database construction process, and advances control to step S300.

[0075] 5 is a flowchart of the database construction processing subroutine in step S200. The database construction processing will be described with reference to FIG. 5. In step S200, the processing device 100 determines bacterial species that cannot be distinguished by MALDI-MS based on the theoretical values ​​of the observed masses of proteins predicted to be produced by the microorganisms, and constructs a database.

[0076] 5, in step S10, the processing device 100 selects target microorganisms contained in the discrimination DB 111. The target microorganisms may be all microorganisms, or a specific type of microorganism may be selected by the user. The specific type of microorganisms may be, for example, bacteria, archaea, fungi, and viruses.

[0077] In step S12, the processing device 100 acquires the genome data of the microorganism selected in step S10 from the public genome DB 70.

[0078] In step S14, the processing device 100 predicts a group of proteins produced by each microorganism based on the genome data of each microorganism acquired in step S12.

[0079] In step S16, the processing device 100 calculates the theoretical values ​​of the observed masses of the proteins predicted to be produced by each microorganism in step S14, and creates a mass-to-charge ratio list for each microorganism.

[0080] In step S18, the processing device 100 selects two types of microorganisms from the microorganisms selected in step S10, and extracts a list of mass-to-charge ratios corresponding to the two types of microorganisms.

[0081] In step S20, the processing device 100 calculates a first similarity, which is the similarity between the lists of mass-to-charge ratios extracted in step S 18. Fig. 6 shows the subroutine for calculating the first similarity in step S20.

[0082] Referring to FIG. 6, in step S50, the processing device 100 sets a tolerance range for each of the theoretical values ​​of the observed masses included in the first list of the two mass-to-charge ratio lists.

[0083] In step S52, the processing device 100 assigns numbers from 1 to n to the theoretical values ​​of n observed masses contained in a second list, which is different from the first list, of the two mass-to-charge ratio lists.

[0084] In step S54, the processing device 100 compares the kth (k is an integer in the range of 1≦k≦n) theoretical value with the tolerance range set in step S50 for the theoretical values ​​in the first list.

[0085] In step S56, the processing device 100 determines that the kth theoretical value matches one of the theoretical values ​​included in the first list if the kth theoretical value is within the allowable range of one of the theoretical values ​​in the first list.

[0086] In step S58, the processing device 100 determines whether the theoretical values ​​of the n observed masses included in the second list have been compared with the theoretical values ​​included in the first list. If the processing device 100 determines that all of the theoretical values ​​of the n observed masses included in the second list have been compared with the theoretical values ​​included in the first list (YES in step S58), it ends the first similarity calculation process subroutine in step S20 and returns to the process of Fig. 5. If the processing device 100 determines that all of the theoretical values ​​of the n observed masses included in the second list have not been compared with the theoretical values ​​included in the first list (NO in step S58), it returns to step S54.

[0087] 5 , in step S22, the processing device 100 determines whether the first similarity calculated in step S20 is equal to or greater than a first predetermined value. If the first similarity is equal to or greater than the first predetermined value (YES in step S22), the processing device 100 proceeds to step S24. If the first similarity is smaller than the first predetermined value (NO in step S22), the processing device 100 proceeds to step S26.

[0088] In step S24, the processing device 100 sorts the two types of microorganisms selected in step S18 into the same group.

[0089] In step S26, the processing device 100 sorts the two types of microorganisms selected in step S18 into different groups.

[0090] In step S28, the processing device 100 determines whether all combinations of two microorganisms selected from the microorganisms selected in step S10 have been processed. If it is determined that all combinations have been processed (YES in step S28), the processing device 100 proceeds to step S30. If it is determined that all combinations have not been processed (NO in step S28), the processing device 100 returns the process to step S18.

[0091] In step S30, it is determined whether each group contains two or more microorganisms. If the group contains two or more microorganisms (YES in step S30), the processing device 100 proceeds to step S32. If the group does not contain two or more microorganisms (NO in step S30), the processing device 100 proceeds to step S34.

[0092] In step S32, the processing device 100 determines that the bacterial species included in the group cannot be identified by MALDI-MS, and creates group information for each microorganism.

[0093] In step S34, the processing device 100 determines that the bacterial species included in the group can be identified by MALDI-MS, and creates group information for the microorganisms.

[0094] In step S36, the processing device 100 outputs and stores the group information created in steps S32 and S34 in the discrimination DB 111 of the memory 11. Thereafter, the processing device 100 ends the database construction processing subroutine in step S100 and returns the processing to FIG.

[0095] Returning to Fig. 4, in step S300, the processing device 100 determines whether or not it has received an instruction to identify microorganisms. In one implementation example, when a key for identifying the sample to be identified is operated on the startup screen of the output unit 15, an instruction to identify microorganisms is input to the processing device 100. If the processing device 100 determines that it has received the instruction for the analysis (YES in step S300), it proceeds to control to step S400. If the processing device 100 determines that it has not received the instruction for the analysis (NO in step S300), it returns control to step S100. When returning control to step S100, the processing device 100 causes the output unit 15 to display the startup screen.

[0096] Fig. 7 is a flowchart of the discrimination processing subroutine of step S400 in Fig. 4. The discrimination processing of microorganisms contained in a sample will be described with reference to Fig. 7. In step S400, the processing device 100 discriminates the microorganisms contained in the sample based on the database constructed in step S200 and mass spectrum data obtained by analyzing the sample.

[0097] Referring to FIG. 7, in step S70, processing device 100 receives mass spectrum data acquired by mass spectrometer 16 analyzing a sample.

[0098] In step S72, the processing device 100 compares the mass spectral data received in step S70 with each of the reference data stored in the discrimination DB 111, and calculates a second similarity that represents the similarity between the mass spectral data and each of the reference data.

[0099] In step S74, the processing device 100 determines whether the second similarity between the mass spectrum data received by the processing device 100 in step S70 and the most similar reference data is equal to or greater than a second predetermined value. If the second similarity between the most similar reference data is equal to or greater than the second predetermined value (YES in step S74), the processing device 100 proceeds to step S76. If the second similarity between the most similar reference data is less than the second predetermined value (NO in step S74), the processing device 100 proceeds to step S84.

[0100] In step S76, the processing device 100 identifies the microorganism corresponding to the reference data that is most similar to the mass spectrum data received in step S70 as the target microorganism contained in the sample.

[0101] In step S78, the processing device 100 refers to the discrimination DB 111 and determines whether the target microorganism identified in step S76 is of a bacterial species that cannot be identified by MALDI-MS. If the target microorganism is of a bacterial species that cannot be identified by MALDI-MS (YES in step S78), the processing device 100 proceeds to step S80. If the target microorganism is not of a bacterial species that cannot be identified by MALDI-MS (NO in step S78), the processing device 100 proceeds to step S82.

[0102] In step S80, the processing device 100 displays on the output unit 15 that the microorganisms contained in the sample are of a species that cannot be identified by MALDI-MS. Thereafter, the processing device 100 ends the identification processing subroutine and returns the process to FIG. 4.

[0103] In step S82, the processing device 100 displays on the output unit 15 that the sample contains the target microorganism. Thereafter, the processing device 100 ends the discrimination processing subroutine and returns the processing to FIG.

[0104] In step S84, the processing device 100 displays on the output unit 15 that the sample does not contain microorganisms. Thereafter, the processing device 100 ends the discrimination processing subroutine and returns the processing to FIG.

[0105] According to the method for constructing a database for microbial discrimination of this embodiment, bacterial species that cannot be identified by MALDI-MS can be determined based on the theoretical values ​​of observed masses of proteins predicted from genome data. Therefore, the operator does not need to analyze microorganisms by MALDI-MS in order to construct a database for microbial discrimination. According to the method for constructing a database for microbial discrimination of this embodiment, the burden on the operator who measures samples containing microorganisms can be reduced.

[0106] Furthermore, the theoretical values ​​of observed masses of proteins predicted from genome data are not affected by the culture conditions of microorganisms and the preparation conditions of samples. Therefore, according to the method for constructing a database for microorganism discrimination according to this embodiment, an operator can construct a database without considering the culture conditions of microorganisms and the preparation conditions of samples.

[0107] According to the microorganism discrimination system of this embodiment, when a microorganism contained in a sample is discriminated to be of a species that cannot be discriminated by MALDI-MS, the user is notified of this fact. Therefore, the user can recognize that the microorganism contained in the sample is of a species that is difficult to discriminate using a discrimination method based on MALDI-MS. In this case, the user can discriminate the type of microorganism contained in the sample using another discrimination method (for example, a microorganism discrimination method that compares DNA base sequences).

[0108] [Modification] In the above-described microorganism discrimination system 150, if the microorganism contained in the sample is a "bacterial species that cannot be distinguished by MALDI-MS," this fact is displayed on the output unit 15. In this case, the user cannot obtain any information about the microorganism contained in the sample other than the fact that the microorganism is a "bacterial species that cannot be distinguished by MALDI-MS." There are cases where it is preferable for the user to be able to recognize information about the type of microorganism contained in the sample.

[0109] In the microorganism discrimination system 150A according to the modified example, items common to the microorganisms included in a group are extracted based on the classification data of the microorganisms, and the items are saved as group information. When a microorganism included in a group is identified as a target microorganism, the items are displayed on the output unit 15. This allows the user to recognize that the microorganisms included in the sample are "bacterial species that cannot be identified by MALDI-MS" and the items common to these bacterial species in the classification data.

[0110] 8 is a schematic diagram showing the configuration of a microorganism discrimination system 150A which is a modified example of the microorganism discrimination system according to the present embodiment. The microorganism discrimination system 150A includes a mass spectrometer 16, a public genome database 70, a network 90, a processing device 100, and a public classification database 80. Note that in the modified example, the same components as those in the microorganism discrimination system 150 described in the embodiment are designated by the same reference numerals, and detailed description thereof will not be repeated. Furthermore, the content described in the embodiment can be combined with the modified example to the extent that it does not contradict the content.

[0111] The public classification DB 80 is a database containing data on the classification of organisms (hereinafter referred to as classification data). The classification of organisms is generally based on the relationships between organisms, as indicated by ranks (e.g., families, genuses, or species). Microorganisms are traditionally classified based on multiple indicators, including morphological observation, phenotypic traits, chemical taxonomic indicators, protein analysis, and DNA analysis, both of which are based on phenotypes and genomes. However, there are also classification systems based solely on genome information, and multiple classification systems exist. Examples of classification methods for classifying microorganisms based on genome information include DDH analysis and ANI analysis. Examples of the public classification DB 80 include the Genome Taxonomy Database (GTDB), the Ribosomal Database Project (RDP), and Silva's database. The public classification DB 80 is not limited to these and may include, for example, databases that are not publicly available.

[0112] The processing device 100 communicates with a public genome DB 70 and a public classification DB 80 via a network 90 .

[0113] After dividing microorganisms into groups, the processing device 100 acquires classification data for the microorganisms contained in each group containing two or more microorganisms from the public classification DB 80. Then, a common item in the classification data for the microorganisms contained in the group is set as a group name, which is the name of the group, and added to the group information. The common item is, for example, a genus name.

[0114] For example, if there is a group consisting of two bacterial species, Bacillus cereus and Bacillus mycoides, the two bacterial species are both in the genus "Bacillus" in the classification data, so the group name will be "Bacillus."

[0115] When the processing device 100 identifies the target microorganism as a bacterial species that cannot be identified by MALDI-MS, it causes the output unit 15 to display the name of the group to which the target microorganism belongs.

[0116] For example, if it is determined that the sample contains Bacillus cereus, the processing device 100 displays "Bacillus," which is the name of a group that includes that bacterial species, on the output unit 15. By displaying the group name, the user can recognize that the sample contains at least a bacterial species of the genus "Bacillus."

[0117] [Processing flow according to the modified example] Figure 9 is a flowchart showing the construction process of a database for microorganism discrimination performed by the processor 10 in the processing device 100 of the microorganism discrimination system 150A. In one implementation example, the subroutine for the analysis process in Figure 9 is executed when the processor 10 of the processing device 100 executes a given program. In Figure 9, the same components as those in the flowchart explained in Figure 5 are assigned the same reference numerals, and detailed explanations thereof will not be repeated.

[0118] In step S38, the processing device 100 acquires classification data for each of the bacterial species included in the group from the public classification DB 80.

[0119] In step S40, the processing device 100 adds to the group information a group name that is a common item in the classification data of each of the bacterial species included in the group acquired in step S38.

[0120] Fig. 10 is a flowchart showing the microorganism discrimination processing performed by the processor 10 in the processing device 100 of the microorganism discrimination system 150A. In one implementation example, the subroutine for the analysis processing in Fig. 10 is executed when the processor 10 of the processing device 100 executes a given program. In Fig. 10, the same components as those in the flowchart described in Fig. 7 are designated by the same reference numerals, and detailed description thereof will not be repeated.

[0121] In step S86, the processing device 100 causes the output unit 15 to display the group name of the group containing the target microorganism.

[0122] In a modified example, when a microorganism contained in a sample is determined to be a bacterial species that cannot be identified by MALDI-MS, part of the classification data for that microorganism is displayed on the output unit 15. This part of the classification data is an item that is common to the classification data for the microorganisms contained in the group. Therefore, the user can obtain information regarding the classification data for the microorganisms contained in the sample. For example, by knowing this information, the user may be able to determine whether or not it is necessary to implement a microorganism discrimination method other than the MALDI-MS discrimination method. Therefore, displaying the group name may reduce the burden on the user of implementing a microorganism discrimination method other than the MALDI-MS discrimination method.

[0123] [Example 1] Figure 11 shows an example of a method for calculating the similarity between two lists of mass-to-charge ratios created based on theoretical values ​​of observed masses of specific types of proteins. In Figure 11, the specific types of proteins are ribosomal proteins.

[0124] In FIG. 11 , the list is created for microorganisms that are taxonomically referred to as type strains. The processing device 100 creates a ribosomal protein mass-to-charge ratio list for each microorganism. The processing device 100 determines combinations of two lists from the created lists in a round-robin manner and calculates the similarity for all combinations. In this specification, this similarity is referred to as Ribosomal Protein Mass Identity (RPMI). The larger the RPMI value for two microorganisms, the more similar the theoretical values ​​of the observed masses of the ribosomal proteins of the two microorganisms are.

[0125] Figure 11 shows the number of combinations of microorganisms for each calculated RPMI value. Dark gray indicates the number of combinations of microorganisms belonging to different genera, and light gray indicates the number of combinations of microorganisms belonging to the same genus but different species.

[0126] As shown in Figure 11, the RPMI calculated for a combination of microorganisms belonging to the same genus but different species is generally larger than the RPMI calculated for a combination of microorganisms belonging to different genera. Microorganisms belonging to the same genus but different species often have similar mass-to-charge lists of ribosomal proteins.

[0127] The processing device 100 determines that a combination of microorganisms whose RPMI is equal to or greater than a predetermined value cannot be distinguished by MALDI-MS. The predetermined value is, for example, 95%.

[0128] [Example 2] Figure 12 shows an example of a microorganism determined to be a species that cannot be identified by MALDI-MS in a database constructed using publicly available genome data of organisms by the method for constructing a microorganism identification database according to the present disclosure.

[0129] FIG. 12 shows 24 groups each containing two or more microorganisms among the groups generated in the process of creating a database for microorganism identification, and the species names of the microorganisms contained in each group.

[0130] In Example 2, microorganisms belonging to the same group are of the same genus, and therefore, in Figure 12, the types of microorganisms are indicated by their species names.

[0131] For example, group number 3 includes four types of microorganisms: Acetobacter senegalensis, Acetobacter tropicalis, Acetobacter cerevisiae, and Acetobacter malorum. These four types of microorganisms are all microorganisms of the genus Acetobacter. Based on the genome data of each of these four types of microorganisms, the processing device 100 creates a list of mass-to-charge ratios of peaks in mass spectrum data generated when a sample containing each of these microorganisms is analyzed by MALDI-MS. The processing device 100 determines that each of the created lists is similar, and determines that these four types of microorganisms are bacterial species that cannot be distinguished by MALDI-MS.

[0132] Aspects It will be appreciated by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.

[0133] (Item 1) In one aspect, a method for constructing a database for microbial discrimination may include the steps of acquiring genome data for each of two types of microorganisms, predicting a group of proteins produced by each of the two types of microorganisms based on the genome data, calculating theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and creating a list of mass-to-charge ratios for each of the two types of microorganisms, calculating the similarity between the lists of mass-to-charge ratios, and, if the similarity is equal to or greater than a predetermined value, generating information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS, and outputting the information.

[0134] According to the method for constructing a database for microbial identification described in paragraph 1, it is possible to identify bacterial species that cannot be identified by MALDI-MS based on genome data, without using mass spectral data obtained by analyzing the microorganisms by MALDI-MS.

[0135] (Clause 2) The method for constructing a database for microbial discrimination described in paragraph 1 may further include a step of identifying, at the genome level, the species-level taxonomic group of the organism corresponding to each of the genome data based on each of the genome data.

[0136] According to the method for constructing a database for microbial identification described in paragraph 2, it is possible to construct a database that includes information on species-level taxonomic groups of bacterial species that cannot be identified by MALDI-MS, which are identified based on genome data.

[0137] (Item 3) In the method for constructing a database for microorganism discrimination according to item 1 or 2, the protein group may include only a predetermined type of protein.

[0138] According to the method for constructing a database for microbial identification described in item 3, bacterial species that cannot be identified by MALDI-MS are identified based on the theoretical values ​​of the observed masses of predetermined types of proteins.

[0139] (4) In the method for constructing a database for microorganism discrimination described in 3, the predetermined type of protein may include at least one of a ribosomal protein, a chaperone protein, and a DNA-binding protein.

[0140] According to the method for constructing a database for microbial discrimination described in Section 4, bacterial species that cannot be distinguished by MALDI-MS are identified based on the theoretical value of the observed mass of at least one of ribosomal proteins, chaperone proteins, and DNA-binding proteins.

[0141] (Item 5) In the method for constructing a database for microorganism discrimination according to item 3 or 4, the predetermined type of protein may be a protein whose theoretical value of observed mass is within a predetermined range.

[0142] According to the method for constructing a database for microbial identification described in item 5, bacterial species that cannot be identified by MALDI-MS are identified based on the theoretical values ​​of observed masses of proteins whose theoretical values ​​of observed masses are within a predetermined range.

[0143] (Item 6) In the method for constructing a database for microorganism discrimination according to Item 5, the predetermined range may be m / z 1000 to 30000.

[0144] According to the method for constructing a database for microbial discrimination described in Section 6, bacterial species that cannot be distinguished by MALDI-MS are identified based on the theoretical values ​​of observed masses of proteins whose theoretical values ​​of observed masses are in the range of m / z 1000 to 30000.

[0145] (Item 7) In the method for constructing a database for microorganism discrimination according to item 5 or 6, the predetermined range may be m / z 4000 to 20000.

[0146] According to the method for constructing a database for microbial discrimination described in item 7, bacterial species that cannot be distinguished by MALDI-MS are identified based on the theoretical values ​​of observed masses of proteins whose theoretical values ​​of observed masses are in the range of m / z 4000 to 20000.

[0147] (Item 8) In the method for constructing a database for microorganism discrimination according to any one of Items 5 to 7, the predetermined range may be m / z 4000 to 10000.

[0148] According to the method for constructing a database for microbial discrimination described in item 8, bacterial species that cannot be distinguished by MALDI-MS are identified based on the theoretical values ​​of observed masses of proteins whose theoretical values ​​of observed masses are in the range of m / z 4000 to 10000.

[0149] (Item 9) In the method for constructing a database for microorganism discrimination described in any one of Items 1 to 8, each of the two types of microorganisms may be selected from bacteria, archaea, fungi, and viruses.

[0150] According to the method for constructing a database for identifying microorganisms described in paragraph 9, it is possible to construct a database for identification of bacteria, archaea, fungi, and viruses.

[0151] (Clause 10) In the method for constructing a database for microbial discrimination described in any one of clauses 1 to 9, the calculating step may include the steps of: setting an allowable range for each theoretical value of the observed mass included in a first list among the lists of each mass-to-charge ratio; comparing a first theoretical value included in a second list among the lists of each mass-to-charge ratio with each theoretical value included in the first list; and determining that the first theoretical value matches one of the theoretical values ​​included in the first list if the first theoretical value is within the allowable range of one of the theoretical values ​​included in the first list.

[0152] According to the method for constructing a database for microbial discrimination described in paragraph 10, a tolerance range is set when comparing two lists of mass-to-charge ratios. As a result, even if two proteins have different theoretical observed mass values, if the difference falls within the tolerance range, the theoretical observed mass values ​​of the two proteins are determined to match. Measurement values ​​measured by a mass spectrometer may deviate from the theoretical observed mass values, and two proteins with different theoretical observed mass values ​​may not be distinguishable based on the measurements. By setting a tolerance range for two proteins that have different theoretical observed mass values ​​but cannot be distinguished based on measurements by a mass spectrometer, it is possible to determine that the theoretical observed mass values ​​of those proteins match.

[0153] (Item 11) In the method for constructing a database for microorganism determination according to item 10, the tolerance range may be 10 to 1000 ppm.

[0154] According to the method for constructing a database for microbial discrimination described in item 11, when comparing the theoretical values ​​of the observed masses of two proteins, if the theoretical value of one of the observed masses is included in a range of 10 to 1000 ppm of the theoretical value of the other observed mass, the theoretical values ​​of the observed masses of the two proteins are determined to match.

[0155] (Item 12) In the method for constructing a database for microorganism determination according to item 10 or 11, the tolerance range may be 10 to 800 ppm.

[0156] According to the method for constructing a database for microbial discrimination described in item 12, when comparing the theoretical values ​​of the observed masses of two proteins, if the theoretical value of one of the observed masses is included in a range of 10 to 800 ppm of the theoretical value of the other observed mass, the theoretical values ​​of the observed masses of the two proteins are determined to match.

[0157] (Item 13) In one aspect, the recording medium is a computer-readable recording medium that records a database for microbial identification constructed by the method for constructing a database for microbial identification described in any one of items 1 to 12.

[0158] The recording medium described in item 13 includes information specifying a bacterial species that cannot be identified by MALDI-MS based on genome data, without using mass spectral data obtained by analyzing the microorganism by MALDI-MS.

[0159] (Item 14) In one aspect, an apparatus for constructing a database for microbial discrimination comprises a processor and a storage unit. The processor acquires genome data for each of two types of microorganisms, predicts a group of proteins produced by each of the two types of microorganisms based on the genome data, calculates theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS), creates a list of mass-to-charge ratios for each of the two types of microorganisms, calculates a similarity between the lists of mass-to-charge ratios, and, if the similarity is equal to or greater than a predetermined value, generates information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS, and stores the information in the storage unit.

[0160] According to the device for constructing a database for microbial identification described in paragraph 14, it is possible to identify bacterial species that cannot be identified by MALDI-MS based on genome data, without using mass spectral data obtained by analyzing the microorganisms by MALDI-MS.

[0161] (Item 15) In one aspect, the program may be executed by a processor mounted on a computer to cause the computer to perform the following operations: acquire genome data for each of two types of microorganisms; predict a group of proteins produced by each of the two types of microorganisms based on the genome data; calculate theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS); create a list of mass-to-charge ratios for each of the two types of microorganisms; calculate a similarity between the lists of mass-to-charge ratios; and, if the similarity is equal to or greater than a predetermined value, generate information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS; and output the information.

[0162] According to the program described in item 15, it is possible to identify bacterial species that cannot be identified by MALDI-MS based on genome data, without using mass spectral data obtained by analyzing the microorganisms by MALDI-MS.

[0163] (Item 16) In one aspect, a method for distinguishing microorganisms may include the steps of: acquiring mass spectral data obtained by analyzing a sample with MALDI-MS; comparing the mass spectral data with a database for microorganism distinction constructed by the method for constructing a database for microorganism distinction described in any one of Items 1 to 11, and identifying a target microorganism contained in the sample; and, if the target microorganism is identified as at least one of the two types of microorganisms, notifying a user that the microorganism contained in the sample is a bacterial species that cannot be distinguished by MALDI-MS.

[0164] According to the microorganism discrimination method described in paragraph 16, if the microorganism contained in the sample is of a species that cannot be discriminated by MALDI-MS, the user is notified of this.

[0165] (Clause 17) The microorganism discrimination method described in Clause 16 may further include the steps of obtaining classification data for each of the two types of microorganisms, comparing the classification data and adding common items to the information, and notifying a user of the common items when the target microorganism is identified as at least one of the two types of microorganisms.

[0166] According to the microorganism discrimination method described in paragraph 17, when a microorganism contained in a sample is of a species that cannot be discriminated by MALDI-MS, part of the classification data regarding the microorganism is notified to the user.

[0167] (Item 18) A microorganism discrimination system in one aspect includes a mass spectrometer, an analyzer that receives data from the mass spectrometer, and a display device. The analyzer acquires mass spectrum data obtained by analyzing a sample with the mass spectrometer, compares the mass spectrum data with a microorganism discrimination database constructed by the method for constructing a microorganism discrimination database described in any one of Items 1 to 11, identifies a target microorganism contained in the sample, and, when the target microorganism is identified as at least one of the two types of microorganism, displays on the display device that the microorganism contained in the sample is a bacterial species that cannot be distinguished by MALDI-MS.

[0168] According to the microorganism discrimination system described in paragraph 18, if the microorganism contained in the sample is of a species that cannot be discriminated by MALDI-MS, the user is notified of this.

[0169] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not by the description of the above-described embodiments, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, it is intended that each technique in the embodiments can be implemented alone or, if necessary, in combination with other techniques in the embodiments to the extent possible.

[0170] 10 Processor, 11 Memory, 12 Communication I / F, 13 Input / Output I / F, 14 Input section, 15 Output section, 16 Mass spectrometer, 70 Public genome database, 80 Public classification database, 90 Network, 100 Processing device, 101 Controller, 112 Analysis program, 150, 150A Microorganism discrimination system.

Claims

1. A method for constructing a database for distinguishing microorganisms, comprising the steps of: acquiring genome data for each of two types of microorganisms; predicting a group of proteins produced by each of the two types of microorganisms based on the genome data; calculating theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and creating a list of mass-to-charge ratios for each of the two types of microorganisms; calculating the similarity of each of the lists of mass-to-charge ratios; generating information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS if the similarity is equal to or greater than a predetermined value; and outputting the information.

2. A method for constructing a database for microbial identification according to claim 1, further comprising a step of identifying, at the genome level, the species-level taxonomic group of the organism corresponding to each of the genome data based on each of the genome data.

3. A method for constructing a database for identifying microorganisms according to claim 1 or claim 2, wherein the protein group includes only proteins of a predetermined type.

4. The method for constructing a database for identifying microorganisms according to claim 3, wherein the predetermined types of proteins include at least one of ribosomal proteins, chaperone proteins, and DNA-binding proteins.

5. The method for constructing a database for identifying microorganisms according to claim 3, wherein the predetermined type of protein is a protein whose theoretical observed mass value is within a predetermined range.

6. The method for constructing a database for identifying microorganisms according to claim 5, wherein the predetermined range is m / z 1000 to 30000.

7. The method for constructing a database for identifying microorganisms according to claim 6, wherein the predetermined range is m / z 4000 to 20000.

8. The method for constructing a database for identifying microorganisms according to claim 7, wherein the predetermined range is m / z 4000 to 10000.

9. A method for constructing a database for identifying microorganisms according to claim 1 or claim 2, wherein each of the two types of microorganisms is selected from bacteria, archaea, fungi, and viruses.

10. A method for constructing a database for microbial identification as described in claim 1 or claim 2, wherein the calculating step comprises the steps of: setting an allowable range for each theoretical value of the observed mass included in a first list from among the lists of each mass-to-charge ratio; comparing a first theoretical value included in a second list from among the lists of each mass-to-charge ratio with each theoretical value included in the first list; and determining that the first theoretical value matches one of the theoretical values ​​included in the first list when the first theoretical value is within the allowable range of one of the theoretical values ​​included in the first list.

11. The method for constructing a database for identifying microorganisms according to claim 10, wherein the tolerance range is 10 to 1000 ppm.

12. The method for constructing a database for identifying microorganisms according to claim 11, wherein the tolerance range is 10 to 800 ppm.

13. A computer-readable recording medium for recording a database for microorganism identification constructed by the method for constructing a database for microorganism identification according to claim 1 or 2.

14. An apparatus for constructing a database for distinguishing microorganisms, comprising a processor and a memory unit, wherein the processor: acquires genome data for each of two types of microorganisms; predicts a group of proteins produced by each of the two types of microorganisms based on the genome data; calculates theoretical values ​​of observed masses measured when the group of proteins is analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS); creates a list of mass-to-charge ratios for each of the two types of microorganisms; calculates the similarity of each of the lists of mass-to-charge ratios; and, if the similarity is equal to or greater than a predetermined value, generates information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS; and stores the information in the memory unit.

15. A program that, when executed by a processor mounted on a computer, causes the computer to: obtain genome data for each of two types of microorganisms; predict proteins produced by each of the two types of microorganisms based on the genome data; calculate theoretical values ​​of observed masses measured when the proteins are analyzed by matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and create a list of mass-to-charge ratios for each of the two types of microorganisms; calculate the similarity of each of the lists of mass-to-charge ratios; and, if the similarity is equal to or greater than a predetermined value, generate information indicating that the two types of microorganisms cannot be distinguished by MALDI-MS; and output the information.

16. A method for distinguishing microorganisms, comprising the steps of: acquiring mass spectral data obtained by analyzing a sample using MALDI-MS; comparing the mass spectral data with a database for microorganism discrimination constructed by the method for constructing a database for microorganism discrimination described in claim 1 or claim 2, and identifying a target microorganism contained in the sample; and, if the target microorganism is identified as at least one of the two types of microorganisms, notifying a user that the microorganism contained in the sample is a bacterial species that cannot be distinguished by MALDI-MS.

17. A method for distinguishing microorganisms as described in claim 16, further comprising the steps of: acquiring classification data for each of the two types of microorganisms; comparing the classification data for each and adding common items to the information; and, when the target microorganism is identified as at least one of the two types of microorganisms, notifying a user of the common items.

18. A microbial discrimination system comprising: a mass spectrometry unit; an analysis device that receives data from the mass spectrometry unit; and a display device, wherein the analysis device acquires mass spectral data obtained by analyzing a sample with the mass spectrometry unit; compares the mass spectral data with a microbial discrimination database constructed by the method for constructing a microbial discrimination database described in claim 1 or claim 2; identifies a target microorganism contained in the sample; and, if the target microorganism is identified as at least one of the two types of microorganism, displays on the display device that the microorganism contained in the sample is a bacterial species that cannot be distinguished by MALDI-MS.

Citation Information

Patent Citations

  • Chromatograph mass analysis data processing device

    JP2014202582A

  • Microbe analysis method

    JP2021173613A

  • Microorganism identification method

    WO2017168740A1

  • Method and apparatus for constructing database for microbial identification

    WO2023204008A1