Microorganism analysis method and interpretation apparatus

The method and device leverage Asr fragments to address the limitations of conventional bacterial identification techniques, enabling efficient and accurate differentiation of Enterobacteriaceae bacteria by analyzing mass spectral patterns, thus simplifying the identification process.

WO2026018858A1PCT designated stage Publication Date: 2026-01-22SHIMADZU CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/025421
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-16
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Conventional methods for identifying bacteria of the order Enterobacteriaceae, such as MALDI and biochemical property tests, struggle to distinguish between similar species like Enterobacter asburiae and Enterobacter cloacae, as well as Escherichia albertii and Escherichia coli, due to reliance on ribosomal proteins and the need for extensive strain collection and analysis.

Method used

A method and device that utilize acid shock protein (Asr) fragments for bacterial identification by obtaining full-length sequences, identifying cleavage sites, and calculating discrimination ratios based on mass spectral patterns, allowing for rapid and efficient differentiation between bacterial groups without extensive strain collection.

Benefits of technology

Enables rapid and efficient identification of Enterobacteriaceae bacteria by simulating mass spectrometry results for numerous strains, reducing time and effort, and facilitating accurate differentiation of challenging bacterial species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025025421_22012026_PF_FP_ABST
    Figure JP2025025421_22012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention includes: a step (S0) for acquiring, for each of a first bacterial group and a second bacterial group belonging to the order Enterobacterales and being different from each other in at least one classification below the family level, a plurality of full-length sequences which are full-length amino acid sequences of a plurality of corresponding acid shock proteins; a step (S1) for identifying a cleavage sequence indicating a site in which the plurality of full-length sequences are fragmented; and steps (S2)-(S8) for obtaining a mass spectral pattern of a plurality of acid shock protein fragments obtained by fragmenting each of the plurality of full-length sequences, and calculating the identification ratios of the first bacterial group and the second bacterial group on the basis of the mass spectral pattern.
Need to check novelty before this filing date? Find Prior Art

Description

Microorganism analysis method and analysis device

[0001] The present invention relates to a method and an apparatus for analyzing microorganisms, and more particularly to a technique for classifying microorganisms using acid shock proteins.

[0002] Among the classifications of microorganisms, the order Enterobacteriale is known to include pathogenic bacteria such as enterohemorrhagic Escherichia coli, Salmonella, Shigella, and Yersinia pestis. The order Enterobacteriale also includes bacteria that are the subject of epidemiological investigations of food poisoning. Identifying the bacterial groups of the order Enterobacteriale, which include these important bacterial groups, is important in research on microorganisms and in medical settings.

[0003] Regarding bacteria belonging to the order Enterobacteriaceae, International Publication No. 2020 / 202861 (Patent Document 1) shows that when some strains belonging to the order Enterobacteriaceae are cultured under appropriate conditions, they produce acid shock protein (Acid shock protein, hereinafter also referred to as "Asr"). It has also been disclosed that Asr peaks showing different measured m / z were found in the mass spectra of different strains. In this way, it is possible to distinguish bacteria belonging to the order Enterobacteriaceae by observing Asr by mass spectrometry.

[0004] International Publication No. 2020 / 202861

[0005] Comparative study of clinical isolate identification by mass spectrometry (VITEK MS) and biochemical properties, Takuya Hattori et al., Medical Testing, 2014, Vol. 63, No. 5

[0006] However, in order to determine whether it is possible to actually identify Enterobacteriaceae by observing Asr by mass spectrometry, or to attempt to identify them, analysts must collect and culture a large number of strains of the target bacterial group and actually observe the Asr of these many strains by mass spectrometry, which is a time-consuming process.

[0007] Therefore, a method has been eagerly awaited that can easily obtain information on the identification of Enterobacteriaceae bacteria without actually collecting and culturing a large number of strains and performing mass spectrometry.

[0008] The present disclosure has been made to solve such problems, and its purpose is to obtain information relating to the identification of Enterobacteriaceae bacteria in a simple manner.

[0009] A method for analyzing microorganisms according to a first aspect of the present disclosure includes the steps of obtaining, for each of a first group of bacteria and a second group of bacteria that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family, a plurality of full-length sequences that are the full-length amino acid sequences of a plurality of corresponding acid shock proteins; identifying cleavage sequences that indicate sites at which the plurality of full-length sequences are fragmented; and determining mass spectral patterns of a plurality of acid shock protein fragments obtained by fragmenting each of the plurality of full-length sequences, and calculating a discrimination ratio between the first group of bacteria and the second group of bacteria based on the mass spectral patterns.

[0010] An analytical device according to a second aspect of the present disclosure includes a processor and a memory unit. The processor acquires multiple full-length sequences, which are the full-length amino acid sequences of multiple corresponding acid shock proteins, for each of a first bacterial group and a second bacterial group that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family. The processor identifies cleavage sequences that indicate sites at which the multiple full-length sequences are cleaved. The processor obtains mass spectral patterns of multiple acid shock protein fragments obtained by cleaving each of the multiple full-length sequences, and calculates a discrimination ratio between the first bacterial group and the second bacterial group based on the mass spectral patterns.

[0011] According to the method for analyzing microorganisms disclosed herein, information relating to the identification of bacterial groups of the order Enterobacteriaceae can be obtained in a simple manner.

[0012] 1 is a schematic diagram showing the configuration of an analysis system according to an embodiment. FIG. 1 is a flowchart showing a process relating to a method for evaluating the distinguishability of microorganisms according to embodiment 1. FIG. 1 is a flowchart showing a process for acquiring a plurality of full-length sequences. FIG. 2 is a diagram showing the acquired full-length sequence. FIG. 3 is a diagram showing the acquired full-length sequence. FIG. 4 is a diagram showing the acquired full-length sequence. FIG. 5 is a diagram showing the acquired full-length sequence. FIG. 6 is a diagram showing the acquired full-length sequence. FIG. 7 is a diagram showing the acquired full-length sequence. FIG. 8 is a flowchart showing a process for identifying cleavage sequences and acquiring a plurality of fragment sequence sets in one example. FIG. 9 is a diagram explaining fragment sequence sets. FIG. 10 is a diagram showing the relationship between theoretical molecular mass and actually measured m / z. FIG. 11 is a diagram showing the acquired fragment sequence set and molecular mass set. FIG. 12 is a diagram showing the acquired fragment sequence set and molecular mass set. FIG. 13 is a diagram showing the acquired fragment sequence set and molecular mass set. FIG. 14 is a diagram showing the acquired fragment sequence set and molecular mass set. FIG. 15 is a diagram showing the distinguishability of identical peak groups and bacterial groups. FIG. 16 is a diagram showing the distinguishability of identical peak groups and bacterial groups. FIG. 17 is a diagram showing the distinguishability of identical peak groups and bacterial groups. 1 is a diagram showing the distinguishability of identical peak groups and bacterial groups. 2 is a diagram showing the distinguishability of identical peak groups and bacterial groups. 3 is a diagram for explaining identical protein groups. 4 is a flowchart showing a process for calculating a second distinguishing ratio. 5 is a flowchart showing a process for determining molecular mass markers. 6 is a flowchart showing a process relating to a microorganism distinguishing method according to embodiment 1. 7 is a flowchart including a process for discriminating acid shock protein peaks. 8 is a flowchart showing a process for distinguishing between E. asburyae and E. cloacae. 9 is a flowchart showing a process relating to a method for evaluating the distinguishability of microorganisms according to embodiment 2. 10 is a flowchart showing a process relating to a microorganism distinguishing method according to embodiment 2.

[0013] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following, the same or corresponding parts in the drawings are denoted by the same reference numerals, and their description will not be repeated in principle.

[0014] 1, an analysis system 1000 includes a genome database 70, a network 90, and an analysis device 100. In this specification, the term "database" may also be referred to as a "DB."

[0015] In this specification, the genome DB 70 is a database containing information on the amino acid sequences of proteins contained in a certain organism. The genome DB 70 may be a DB containing information on the base sequences of nucleic acids (deoxyribonucleic acid (DNA), ribonucleic acid (RNA)) and containing information on the amino acid sequences of proteins mechanically deduced from the base sequence information.

[0016] The genome DB 70 is typically a publicly available DB, and in one example, is the National Center for Biotechnology Information (NCBI). In the protein database of NCBI, information about a given protein is managed in association with a given accession number.

[0017] The genome DB 70 may be, for example, a genome DB of the DNA Data Bank of Japan (DDBJ) or the European Molecular Biology Laboratory (EMBL). However, examples of the genome DB 70 are not limited thereto, and may include, for example, a genome DB that is not publicly available.

[0018] The network 90 is a network through which the analysis device 100 communicates with the genome DB 7. The network 90 is, for example, the Internet, which interconnects numerous government, corporate, public, and private networks around the world.

[0019] The analysis device 100 includes a controller 101, a display 15, and an operation unit 14. The display 15 and operation unit 14 are connected to the controller 101. The operation unit 14 is typically composed of a touch panel, a keyboard, a mouse, etc. The operation unit 14 accepts user operation input to the processor 10. The display 15 is composed of, for example, a liquid crystal panel capable of displaying images. The display 15 displays images related to the acceptance of the user operation input and displays the results of processing by the processor 10.

[0020] The controller 101 has, as its main components, a processor 10, a memory 11, a communication interface (I / F) 12, and an input / output I / F 13. These components are connected to each other via a bus so as to be able to communicate with each other.

[0021] The processor 10 is typically a calculation processing unit such as a central processing unit (CPU) or a micro processing unit (MPU). The processor 10 controls the operation of the analysis device 100 by reading and executing a program stored in the memory 11.

[0022] The memory 11 is realized by a storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), and a HDD (Hard Disk Drive). The ROM can store programs executed by the processor 10. The RAM can temporarily store data used during program execution by the processor 10 and can function as a temporary data memory used as a work area. The HDD is a non-volatile storage device. A semiconductor storage device such as a flash memory may be used in addition to or instead of the HDD. The programs and / or data may be stored in an external storage device accessible by the processor 10. The memory 11 corresponds to one embodiment of a "storage unit."

[0023] The communication I / F 12 is a communication interface for exchanging various data with an external device including the genome DB 70, and is realized by an adapter, a connector, etc. The communication method may be a wireless communication method using a wireless LAN (Local Area Network) or the like, or a wired communication method using a USB (Universal Serial Bus) or the like.

[0024] The input / output I / F 13 is an interface for exchanging various types of data between the processor 10 and external devices connected to the input / output I / F 13. The external devices include an operation unit 14 and a display 15. A mass spectrometer (MS) 16 may be connected to the input / output I / F 13. In this specification, the input / output I / F 13 also includes devices that exchange data between the processor 10 and a storage terminal such as a USB memory connected to the analysis device 100.

[0025] The MS 16 is a device for performing mass analysis of components contained in a sample. In one embodiment, the MS 16 is a MALDI-TOF MS (Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry). The MS 16 may be, for example, an IT-TOF (Matrix-Assisted Laser Desorption / Ionization Ion Trap Time-of-Flight Mass Spectrometry) or a scanning IT-MS, but is not limited to these. In the MS 16, ions generated by laser irradiation are drawn into a flight tube, separated according to their flight time, and then detected. The time of flight correlates with the m / z (mass-to-charge ratio) of the component, resulting in a mass spectrum with m / z on the horizontal axis and the intensity of the detected ions on the vertical axis.

[0026] The MS 16 performs mass analysis of proteins in a sample, for example. Thus, peaks are detected in the mass spectrum according to the m / z of the proteins in the sample. Therefore, by referring to the number (number) and / or positions (m / z) of peaks in the mass spectrum, the proteins contained in the sample can be determined. In this specification, the "number (number) and / or positions (m / z) of peaks in the mass spectrum" corresponds to one example of a "mass spectrum pattern."

[0027] Different types of organisms contain different proteins, which results in different mass spectral patterns, making it possible to identify organisms based on their mass spectral patterns.

[0028] In one embodiment, the MS 16 transmits the signal intensity obtained by mass spectrometry of the sample microorganism to the analysis device 100. The processor 10 creates a mass spectrum based on the signal intensity and analyzes the mass spectrum to identify the microorganism contained in the sample.

[0029] Analysis device 100 does not have to be configured by one computer, but may be configured by multiple computers.

[0030] [2. Conventional Identification of Microorganisms by Mass Spectrometry] Microorganism identification methods using MALDI, a type of mass spectrometry, are widely used and are gradually replacing conventional identification methods such as biochemical property tests because they are simple, rapid, and, above all, do not require the advanced skills required for prejudice in selecting a test method. Microorganism identification methods using MALDI are particularly common in the medical field. However, as shown in Non-Patent Document 1, there are cases in which some similar bacterial species cannot be distinguished from each other.

[0031] For example, it is difficult to distinguish between Enterobacter asburiae (E. asburiae) and Enterobacter cloacae (E. cloacae) using conventional MALDI identification methods. E. cloacae has been reported to have the highest mortality rate among infectious diseases caused by bacteria of the genus Enterobacter, and distinguishing it from E. asburiae is also clinically significant.

[0032] Furthermore, for example, Escherichia albertii (E. albertii) and Escherichia coli (E. coli) are also difficult to distinguish using conventional MALDI identification methods. For example, a problem has been pointed out in conventional MALDI identification methods that, even when E. albertii is the test bacterium, E. coli is the top hit. E. albertii is a gram-negative, facultative anaerobic bacillus enterobacterium that was recognized as a new species in 2003 and can cause diarrhea, abdominal pain, and fever in humans. Therefore, when E. coli or E. albertii is suspected as a causative bacterium, it must be correctly identified.

[0033] E. albertii lacks common biochemical properties, and depending on the strain, its biochemical properties may coincide with those of closely related species such as E. coli and Hafnia alvei. This has made it difficult to distinguish it from closely related species even using conventional biochemical property tests.

[0034] As described above, although the Enterobacteriaceae order includes clinically important microorganisms, there has been a problem in that it is difficult to identify them using conventional MALDI identification methods, conventional biochemical property tests, etc.

[0035] [3. Investigation by the Inventors into Identification Using Acid Shock Proteins] The inventors noticed that conventional microbial identification methods using MALDI (hereinafter also referred to simply as "conventional methods") mainly rely on peaks of ribosomal proteins, which account for the majority of peaks in the mass spectra of microorganisms, and explored the possibility of identifying microorganisms using other proteins. As a result, the inventors found that some microorganisms can be identified using Asr.

[0036] The gene encoding Asr is found in most Enterobacteriaceae species, except for the Morganellaceae family. The inventors demonstrated that Asr is produced by culturing numerous Enterobacteriaceae strains in a glucose-containing medium. They also demonstrated by mass spectrometry that Asr is fragmented at specific amino acid sequences. More specifically, the inventors demonstrated that peaks with molecular masses consistent with fragments predicted by gene sequences are observed in the measured mass spectra of numerous Enterobacteriaceae strains. Furthermore, they noted the diversity in Asr amino acid sequences and suggested that this diversity may be utilized to identify microorganisms.

[0037] As described above, by observing Asr using mass spectrometry for Enterobacteriaceae, it may be possible to distinguish between bacteria that are difficult to distinguish using conventional methods. However, to demonstrate this possibility, it is necessary to collect a large number of strains from each target bacterial group and actually observe the Asr of these many strains using mass spectrometry. Collecting a large number of strains belonging to a target bacterial group is not easy, and there are also difficulties such as carefully examining the attributes of each strain. For example, even if a large number of strains suspected to be E. cloacae are collected from human feces, biochemical property tests, etc., must be performed on each of the many strains to confirm that each of the many strains is actually E. cloacae.

[0038] [4. Microorganism Analysis Method According to an Embodiment] The microorganism analysis method according to an embodiment includes a method for evaluating the identifiability of microorganisms and a method for identifying microorganisms found to be identifiable by the evaluation method. More specifically, the method includes a method for evaluating the identifiability of a first group and a second group of bacteria that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family, and a method for identifying the first group and the second group of bacteria. The first group and the second group of bacteria are, for example, at least one of groups of bacteria belonging to different families in the same order, groups of bacteria belonging to different genera in the same family, groups of bacteria belonging to different species in the same genus, groups of bacteria belonging to different subspecies of the same species, and groups of bacteria that are different strains of the same species.

[0039] As used herein, a bacterial group refers to a group containing bacterial strains that are the same in at least one level of classification, such as family, genus, species, or strain. Furthermore, "identification" of a microorganism includes clarifying the classification of the microorganism at least at one level, such as family, genus, species, or strain. Identification of the microorganism at at least one level of classification, such as family, genus, species, or strain, is also referred to as "identification" by those skilled in the art. More specifically, "identification" of a microorganism includes determining whether the microorganism belongs to a first bacterial group or a second bacterial group.

[0040] Preferably, the first and second bacterial groups are bacterial groups that are indistinguishable from each other by conventional methods. In this case, if it can be determined that the first and second bacterial groups are distinguishable based on the Asr mass spectral pattern, the first and second bacterial groups can be rapidly and easily distinguished from each other using MALDI. However, the first and second bacterial groups may also be bacterial groups that are distinguishable from each other by conventional methods. In this case, if it can be determined that the first and second bacterial groups are distinguishable from each other by the Asr mass spectral pattern, the analyst can conveniently select whether to use the conventional method or discrimination based on the Asr mass spectral pattern to distinguish the first and second bacterial groups from each other.

[0041] In one example, the first and second bacterial groups are E. asburiae and E. cloacae, which are difficult to distinguish using conventional methods (see Embodiment 1, Examples 1 and 2). In another example, they are E. albertii and E. coli, which are difficult to distinguish using conventional methods (see Embodiment 2, Example 3). However, the analytical method according to this embodiment can be applied to any organism that has an Asr gene, and is not limited to the analysis of the above bacterial species.

[0042] The method for evaluating the distinguishability of microorganisms according to the embodiment includes the following three processes. In one example, the processes are performed by the processor 10. In another example, at least a part of the processes may be performed by an analyst operating the analytical device 100.

[0043] In the first process, the processor 10 obtains a plurality of corresponding full-length Asr amino acid sequences for each of the first and second bacterial groups. In one embodiment, the processor 10 obtains a plurality of corresponding full-length Asr amino acid sequences for each of the first and second bacterial groups from the genome DB 70. Hereinafter, the full-length Asr amino acid sequences are also referred to as "full-length sequences."

[0044] In the second process, the processor 10 identifies cleavage sequences that indicate sites at which multiple full-length sequences are fragmented. By analyzing a large number of Asr sequences from the order Enterobacteriaceae, the inventors have found that Asr cleavage sequences include a sequence in which the amino acids are arranged in the order glutamine-lysine-alanine-glutamine (QKAQ sequence) from the N-terminus to the C-terminus, and a sequence in which the amino acids are arranged in the order glutamine-asparagine-alanine-glutamine (QNAQ sequence).

[0045] In the third process, the processor 10 obtains mass spectral patterns of multiple ASr fragments obtained by fragmenting each of multiple full-length sequences, and calculates the discrimination ratio between the first and second bacterial groups based on the mass spectral patterns.

[0046] In one example, the mass spectrum pattern is the number and position (m / z) of Asr peaks (see Embodiment 1, Examples 1 and 2). In another example, the mass spectrum pattern is the number of Asr peaks (see Embodiment 2, Example 3).

[0047] By using the first to third processes described above to calculate the discrimination ratio based on the mass spectrum pattern using the amino acid sequence information of Asr for a large number of strains of each bacterial group, it is possible to evaluate discrimination feasibility without the need to actually collect a large number of bacterial strains and perform mass analysis of Asr. Therefore, the inventors can evaluate the discrimination feasibility of two bacterial groups without spending time and effort. As described above, the discrimination feasibility evaluation method according to the embodiment allows information on discrimination between bacterial groups of the order Enterobacteriaceae to be obtained in a simple manner.

[0048] Next, each step will be described in more detail with reference to Embodiments 1 and 2. [5. Embodiment 1] (5-1. Evaluation of Microorganism Identification Possibility) FIG. 2 is a flowchart showing a process relating to a method for evaluating microorganism identification possibility according to Embodiment 1. In one example, the process of FIG. 2 is performed by the processor 10. In another example, at least a part of the process of FIG. 2 may be performed by an analyst operating the analytical device 100. S0, S1, and S2 to S8 in FIG. 2 correspond to examples of a "first process," a "second process," and a "third process," respectively.

[0049] In S0, the processor 10 obtains a plurality of corresponding full-length sequences for each of the first and second bacterial groups.

[0050] In S1, the processor 10 identifies cleavage sequences that indicate the sites at which multiple full-length sequences are fragmented.

[0051] In S2 to S8, the processor 10 determines the number and positions (m / z) of Asr peaks, and calculates the discrimination ratio between the first and second bacterial groups based on the number and positions (m / z) of Asr peaks.

[0052] In S2, the processor 10 determines a set of amino acid sequences of multiple Asr fragments obtained by fragmenting each of the multiple full-length sequences. Hereinafter, the "amino acid sequences of the Asr fragments" are also referred to as "fragment sequences," and the "set of fragment sequences" is also referred to as "fragment sequence set."

[0053] In S3, the processor 10 determines a set of theoretical molecular masses of multiple Asr fragments corresponding to the set of multiple fragment sequences. Unless otherwise specified, "molecular mass" herein refers to a theoretical molecular mass calculated from an amino acid sequence. In one embodiment, the molecular mass is a theoretically calculated relative molecular mass (molecular weight). In another embodiment, the molecular mass may be a theoretically calculated m / z (theoretical m / z). In this specification, the m / z measured as a peak on an actually measured mass spectrum is referred to as the actually measured m / z. Hereinafter, a "set of molecular masses" will also be referred to as a "molecular mass set."

[0054] In S4, the processor 10 sets one or more identical peak groups by assigning, for a plurality of molecular mass sets, one or more sets that show the same peak in the measured mass spectrum to the same identical peak group. More specifically, if the first molecular mass set includes only the first molecular mass and the second molecular mass, the second molecular mass set includes only the third molecular mass and the fourth molecular mass, and the difference between the first molecular mass and the third molecular mass is within the above-mentioned error range, and the difference between the second molecular mass and the fourth molecular mass is also within the above-mentioned error range, the processor 10 assigns the first molecular mass set and the second molecular mass set to the same identical peak group.

[0055] As used herein, "identical peaks" refer to peaks that are indistinguishable or difficult to distinguish from one another in a mass spectrum due to a small difference in the molecular mass of the molecules being measured. "Identical peaks" are also referred to by those skilled in the art as "peaks within an error range." The error range corresponds, for example, to the resolution of the mass spectrometer. The error range is, for example, 200 to 1500 ppm (parts per million). In one embodiment, the error range is set to 500 ppm. For example, the difference between molecular weights 5149.8 and 5151.8 is (5151.8 - 5149.8) / 5151.8 = 0.0003882 ≈ 390 ppm. Therefore, if the error range (resolution of the mass spectrometer) is 500 ppm, the peak with molecular weight 5149.8 and the peak with molecular weight 5151.8 cannot be distinguished and are therefore treated as identical. In another embodiment, the error range is, for example, 800 ppm, and in this case, for example, the difference between m / z 5000 and m / z 5004 is (5004-5000) / 5000=0.008=800 ppm, so the peak at m / z 5000 and the peak at m / z 5004 are the "same peak."

[0056] In S5, the processor 10 calculates a first numerical value, which is the number of strains belonging to the first bacterial group, and a second numerical value, which is the number of strains belonging to the second bacterial group, among the strains containing full-length sequences corresponding to each of the same peak groups.

[0057] In S6, the processor 10 determines the smaller and larger of the first and second numerical values ​​for each of the same peak groups.

[0058] In S7, the processor 10 calculates a third numerical value which is the sum of the larger numerical values ​​of all the same peak groups, and a fourth numerical value which is the sum of the first numerical value and the second numerical value of all the same peak groups.

[0059] In S8, the processor 10 calculates a first discrimination ratio, which is the ratio of the third numerical value to the fourth numerical value, and ends the process.

[0060] 2, by calculating the corresponding molecular masses using the amino acid sequence information of Asr for a large number of strains of each bacterial group, it is possible to reproduce on a computer the results of mass spectrometry of Asr for a large number of actual bacterial strains collected, and to evaluate the feasibility of distinguishing between the two bacterial groups. Therefore, the inventors can evaluate the feasibility of distinguishing between the two bacterial groups without spending much time and effort.

[0061] Next, the above steps will be specifically described using an example relating to the discrimination between E. asburiae and E. cloacae.

[0062] (5-1-1. Acquisition of Full-Length Sequences) Figure 3 is a flowchart showing the process of acquiring multiple full-length sequences. Each step shown in Figure 3 corresponds to the subroutine of S0 in Figure 2, and is executed before S1.

[0063] In S01, the processor 10 determines a representative strain for each of the first and second bacterial groups. The representative strain is preferably a strain that is widely used in the first bacterial group (a so-called "standard strain"), but is not limited thereto and may be arbitrarily selected by the analyst from the first bacterial group.

[0064] In S02, the processor 10 acquires a plurality of full-length sequences having a sequence similarity of a predetermined rank or higher to the full-length sequence expressed in the representative strain.

[0065] The sequence similarity, also known as "homology" by those skilled in the art, indicates the percentage of matching between corresponding bases in two amino acid sequences. The processor 10 acquires multiple full-length sequences, for example, by performing a homology search in the genome DB 70 via the network 90. ​​In one embodiment, the genome DB 70 is NCBI. The homology search is preferably performed using BLAST (Basic Local Alignment Search Tool). In S12, the processor 10 may acquire multiple full-length sequences whose sequence similarity to the full-length sequence expressed in the representative strain is equal to or greater than a predetermined value.

[0066] The homology search may be performed using the accession number assigned to each amino acid sequence as a query, or the amino acid sequence itself as a query, or the DNA sequence of the Asr of a representative strain may be used as a query and then converted into an amino acid sequence.

[0067] Figures 4 to 12 show the full-length sequences obtained in the process of S12. Figures 4 to 8 show Table A (E. asburyae full-length sequence listing), which is the result of a search for E. asburyae by accession number, in five figures. Figures 9 to 12 show Table B (E. cloacae full-length sequence listing), which is the result of a search for E. cloacae by accession number, in four figures.

[0068] 4 to 12 show full-length sequences linked to accession numbers in descending order of similarity (highest homology) to the full-length sequence of the query type strain. More specifically, each row shows, from left to right, the "rank" of sequence similarity to the full-length sequence of the query type strain, the "accession (AC) number," the "full-length sequence," and the "SEQ ID NO." In this specification, the "SEQ ID NO" corresponds to the SEQ ID NO listed in the sequence listing.

[0069] "Accession numbers" are explained in detail below. "Accession numbers" indicate the accession numbers obtained as a result of a query search in the NCBI protein database. There are two main types of accession numbers used in the protein database: RefSeq number: NCBI Reference Sequence Database number ISNDC number: The International Nucleotide Sequence Database Collaboration number RefSeq numbers begin with "WP_" and are followed by numbers. ISNDC numbers begin with an alphabet, but there are no explicit rules for assigning alphanumeric characters.

[0070] Both the RefSeq number and the ISNDC are linked to the strain in which the amino acid sequence is expressed, along with the amino acid sequence. The same RefSeq number is assigned if the amino acid sequence is the same, even if the strains in which the amino acid sequence is expressed are different. Therefore, for example, even if the strains belong to different genera or species, the same RefSeq number is assigned if the amino acid sequence derived from genetic information is the same. In other words, the RefSeq number is a number that cannot identify classifications such as genus, species, or strain.

[0071] On the other hand, ISNDC numbers are assigned to each strain even if the amino acid sequence is the same. Therefore, even if the amino acid sequence is the same, different strains that express it will be assigned different ISNDC numbers.

[0072] As described above, the same amino acid sequence may be registered in NCBI with one RefSeq number and multiple ISNDC numbers. A collection of multiple accession numbers corresponding to the same amino acid sequence is called an identical protein group (IPG) in NCBI. Because an IPG contains multiple accession numbers, a representative number (RefSeq Selected Product) is set as a single accession number that represents the IPG. If an IPG contains a RefSeq number, the representative number is the RefSeq number. The accession number listed as a search result in NCBI is the representative number. Therefore, the "Accession Number" in Figures 4 to 12 indicates the representative number.

[0073] As described above, the process shown in FIG. 3 can obtain multiple full-length sequences whose sequence similarity to full-length sequences expressed in representative bacterial strains is at or above a predetermined rank. The predetermined rank is appropriately set by the analyst. The predetermined rank is preferably set to an appropriate value based on the amount of data (the number of full-length sequences to be processed) and the degree of homology between the full-length sequences. In one aspect of this embodiment, the predetermined rank is 10 to 20. In one example, the predetermined rank is 15. In this example, the top 15 or higher homologous items above the double ruled line L1 in FIG. 5 and the double ruled line L2 in FIG. 10 are used. The process shown in FIG. 3 can easily obtain multiple full-length amino acid sequences corresponding to each of the first and second bacterial groups. This allows for computer-based processing that simulates the actual cultivation of multiple bacterial strains belonging to each bacterial group, extraction of amino acids, and analysis.

[0074] The process shown in Fig. 3 may be replaced by another process that can obtain multiple full-length sequences corresponding to each bacterial group. For example, processor 10 may use all full-length sequences displayed as similar to full-length sequences expressed in representative bacterial strains (e.g., all full-length sequences in Figs. 4 to 12). In this case, the number of processes related to the subsequent steps from S4 onwards increases, but the accuracy of the evaluation of identifiability increases.

[0075] Furthermore, for example, the processor 10 may select multiple strains in each bacterial group and obtain the full-length sequences of each of the multiple strains. In this case, the "multiple strains in each bacterial group" may be strains that are distantly related to each other in each bacterial group. In this case, full-length sequences with theoretically low sequence similarity can be obtained in each bacterial group, thereby maximizing the variation in full-length sequences in each bacterial group. However, in general, two bacterial groups that cannot be distinguished by conventional methods are closely related and have high sequence similarity between their full-length sequences. Therefore, the process of Figure 3 is considered to adequately simulate the actual process of culturing multiple bacterial strains, extracting proteins, and analyzing them. Therefore, from the perspective of simplicity, it is preferable to use the process of Figure 3.

[0076] (5-1-2. Acquisition of fragment sequence sets and molecular mass sets) Figure 13 is a flowchart showing the process of identifying cleavage sequences and acquiring multiple fragment sequence sets in one embodiment. S1A and S2A shown in Figure 13 correspond to one embodiment of S1 and S2 in Figure 2, respectively.

[0077] In S1A, the processor 10 removes the signal peptide from the full-length sequence and identifies a cleavage sequence. Since the signal peptide of Asr extends from the N-terminus to the "AFA" sequence, the processor 10 searches for the "AFA" sequence and cleaves the signal peptide. The processor 10 then searches for and identifies a cleavage sequence (QKAQ sequence or QNAQ sequence) in the amino acid sequence from which the signal peptide has been removed.

[0078] In S2A, if there is a cleavage sequence in the full-length amino acid sequence from which the signal peptide has been removed, the processor 10 cleaves the cleavage sequence at the C-terminal end to create a set of fragment sequences.

[0079] Figure 14 is a diagram illustrating the fragment sequence set created in S22. Figure 14 shows the fragment sequence set corresponding to accession number AMA03210.1, which is the Asr gene of the E. asburiae type strain. The representative number of the IPG to which AMA03210.1 belongs in NCBI is WP_059346667.1 (listed in the row with the highest sequence similarity in Figure 4).

[0080] The Asr gene of AMA03210.1 encodes 135 amino acid residues, but in the process of S21, 21 residues are removed from the N-terminus as a signal peptide. The remaining 114 residues are the Asr body. In the process of S22, if the Asr body contains a QKAQ sequence or a QNAQ sequence, the Asr body is cleaved at the C-terminus of these sequences to fragment it. This fragmentation results in a set of fragment sequences including fragment sequences 1 to 4. As described above, processor 10 reproduces on a computer the process by which fragment sequences are generated in a cell that actually contains Asr.

[0081] The inventors have also revealed that some Asrs of some Enterobacteriaceae bacteria do not contain the QKAQ or QNAQ sequence and are not fragmented. For such Asrs, the processor 10 processes the fragment sequence as if it were the same sequence as the Asr itself, and the fragment sequence set includes one of the fragment sequences.

[0082] In S3, the processor 10 calculates the molecular mass for each fragment sequence in the fragment sequence set. This makes it possible to estimate the m / z position of the peak on the mass spectrum to which each fragment sequence corresponds. Therefore, the processor 10 can reproduce on a computer the process of extracting fragment sequences from cells that actually contain the fragment sequences, performing mass analysis to obtain a mass spectrum, and confirming the m / z corresponding to the position of the peak of the fragment sequence.

[0083] Fig. 15 shows the relationship between theoretical molecular mass and measured m / z. The lower part of Fig. 15 shows the theoretical m / z of Asr of the E. asburyae type strain shown in Fig. 14, and the upper part of Fig. 15 shows the measured mass spectrum of the E. asburyae type strain.

[0084] More specifically, the lower panel of Figure 15 shows a simulated mass spectrum created to have peaks at theoretical m / z positions corresponding to each molecular mass included in the molecular mass set. The horizontal axis in the lower panel of Figure 15 represents the theoretical m / z. Of fragment sequences 1 to 4 shown in Figure 14, fragment sequence 2 and fragment sequence 3 have the same amino acid sequence, and therefore have peaks at the same theoretical m / z positions in the mass spectrum. Therefore, three peaks are shown in the lower panel of Figure 15. Note that in the lower panel of Figure 15, in order to simulate a positive mode mass spectrum in which cations are selectively observed, peaks are displayed at positions obtained by adding the molecular weight of a proton, 1, to the molecular weight calculated from each sequence fragment.

[0085] The upper panel of Figure 15 shows the actual mass spectrum obtained by performing mass spectrometry in positive mode on an extract obtained by disrupting cells of the E. asburiae type strain. The horizontal axis of the upper panel of Figure 15 represents the measured m / z, and the vertical axis represents the % intensity (relative signal intensity with the highest peak designated as 100).

[0086] 15, the simulated mass spectrum in the lower row is similar to the measured mass spectrum in the upper row. More specifically, the theoretical m / z positions of the fragment sequences shown in the lower row are similar to the measured m / z positions of the measured peaks shown in the upper row. As described above, the molecular mass set calculated for a specific strain can reproduce the measured m / z of the Asr fragment peak obtained from the measured mass spectrum.

[0087] 16 to 21 show the third table (fragment sequence set table) showing the fragment sequence sets and molecular mass sets obtained in the processing of S2 to S3, divided into six figures.

[0088] Figures 16 to 21 show sets of fragment sequences obtained by removing the signal sequences from the full-length sequences (shown by accession numbers in Figures 16 to 21) ranked in the top 15 by sequence similarity in the tables of Figures 4 to 12 and fragmenting them at QKAQ or QNAQ sequences. The sequence number in the sequence listing is shown for each fragment sequence in the fragment sequence set. Also shown is the molecular mass corresponding to each fragment sequence, calculated based on each fragment sequence in the fragment sequence set.

[0089] (5-1-3. Creation of the Same Peak Group) In S4, the processor 10 assigns the same peak group when the molecular mass sets are the same or within the error range of each other. Below, a more detailed explanation will be given of the case where the full-length sequences are different but the molecular mass sets are the same.

[0090] Even when the full-length sequences are different, the molecular mass set is the same and they are assigned to the same peak group, as shown in the following first to third examples, for example.

[0091] The first example is when the signal peptide sequence is different but the Asr body sequence is the same, and therefore the molecular mass set of the Asr fragments is the same.

[0092] The second example is when the number of repeat sequences in the Asr body is different. Some Asr bodies may contain repeated fragments of the same amino acid sequence. Such amino acid sequences are called "repeated sequences." Even if the number of such repeated sequences is different, all of the repeated sequences have the same molecular mass. Therefore, the molecular mass set of the Asr fragments is the same.

[0093] The third example is when the amino acid composition in the Asr body is the same. The amino acid composition is the number of each amino acid residue in the amino acid sequence. If the amino acid composition is the same, the molecular mass will be the same even if the amino acid arrangement is different. Therefore, the molecular mass sets of the Asr fragments will be the same.

[0094] As described above, even if the full-length sequences are different, they may have the same molecular mass set. In other words, even if the full-length sequences correspond to different RefSeq numbers, they may have the same molecular mass set. In such cases, the processor 10 assigns different full-length sequences having different RefSeq numbers to the same peak group.

[0095] On the other hand, even if the full-length sequences have different RefSeq numbers and the molecular masses included in the molecular mass sets are different, if the difference in molecular mass is within the error range, the processor 10 still assigns them to the same peak group.

[0096] Next, a specific example will be given to explain the process of assigning molecular mass sets to the same peak group. Referring again to Figures 16 to 21, the molecular mass sets corresponding to the three accession numbers WP_045888858.1, WP_096150484.1, and WP_248195658.1 each contain molecular masses of 5149.8, 2069.4, 2099.5, and 2388.8, respectively. As described above, the molecular mass sets corresponding to the three accession numbers each contain the same multiple molecular mass values. In other words, the molecular mass sets corresponding to the three accession numbers each match each other. Furthermore, the molecular mass set for WP_193131548.1 contains molecular masses of 5151.8, 2069.4, 2099.5, and 2388.8. As described above, the molecular mass set of WP_193131548.1 and the molecular mass sets corresponding to the above three accession numbers have the same maximum molecular mass (2388.8), second-largest molecular mass (2099.5), and third-largest molecular mass (2069.4), but differ only in the minimum molecular mass. However, the difference between the minimum molecular mass of 5151.8 in the molecular mass set of WP_193131548.1 and the minimum molecular mass of 5149.8 in the molecular mass sets corresponding to the above three accession numbers is 390 ppm, within the error range (500 ppm), and therefore it is considered difficult to distinguish them in the measured mass spectrum. Therefore, the molecular mass sets corresponding to the above three accession numbers and the molecular mass set of WP_193131548.1 are assigned to the same identical peak group, identical peak group 1. Similarly, WP_013096840.1, WP_057072791.1, and WP_218664773.1 are grouped together in the same peak group 2, WP_163364963.1, AMX05469.1, and WP_032657610.1 are grouped together in the same peak group 3, WP_221346174.1 and WP_248725202.1 are grouped together in the same peak group 4, and WP_150057180.1 and WP_148791141.1 are grouped together in the same peak group 5.As described above, for each of the molecular mass sets in Figures 16 to 21, peaks predicted to have the same Asr mass spectrum peaks (in other words, Asr peaks with the same number and at the same positions) are assigned to the same peak group. As a result, 21 peak groups are set as shown in the tables in Figures 22 to 27.

[0097] Figures 22 to 27 are six diagrams showing Table D (Identification Possibility Table), which shows the identification possibility of identical peak groups and bacterial groups. Columns A to C, which relate to the processing of S4, list "Identity Peak Group," "Accession (AC) Number," and "Molecular Mass." "Molecular Mass" lists the molecular mass set corresponding to the fragment sequence set corresponding to a specific accession number. More specifically, one or more molecular masses corresponding to one or more fragments contained in the fragment sequence set are listed. The molecular masses in Figures 22 to 27 are expressed in Da.

[0098] (5-1-4. Evaluation of Identification Possibility of Bacterial Groups) In S5, the processor 10 confirms to which bacterial group the Asr corresponding to each identical peak group belongs. For example, the processor 10 uses information from the NCBI IPG to obtain a list of bacterial strains that have been reported to contain one or more full-length sequences corresponding to each identical peak group. Then, in S6, the processor 10 calculates, for each full-length sequence, the number of strains belonging to the first bacterial group and the number of strains belonging to the second bacterial group among the strains included in the list. Then, for each identical peak group, the processor 10 calculates, as a first numerical value, the sum of the strains belonging to the first bacterial group included in the list of full-length sequences corresponding to that identical peak group, and calculates, as a second numerical value, the sum of the strains belonging to the second bacterial group included in the list of full-length sequences corresponding to that identical peak group.

[0099] Figure 28 is a diagram illustrating IPGs. More specifically, Figure 28 is a screen displaying the IPG for the full-length sequence of the E. asburiae type strain, accession number WP_059346667.1, at NCBI. Referring to Figure 28, 10 search results, numbered 1 to 10, are shown for the IPG for WP_059346667.1. However, when the value corresponding to "Strain" is examined, search results for the same strain are included. When the value corresponding to "Source" for the search results for the same strain is examined, it is found that the same strain is provided by two different sources, Refseq and INSDC, and is displayed twice. When the strains displayed twice are reorganized as a single entity and the value of "Organism" is checked, it is found that there are five E. asburiae strains and one E. cloacae strain.

[0100] As described above, a list of bacterial strains containing one or more full-length sequences corresponding to each identical peak group was obtained, and the numbers of strains corresponding to the first bacterial group, the second bacterial group, and other bacterial groups were calculated. The results are shown in columns D to F of Figures 22 to 27. Columns D to F of Figures 22 to 27 show "first numerical value," "second numerical value," and "number of other bacterial strains." The first numerical value is the number of strains in the first bacterial group (E. asburiae in this example). The second numerical value is the number of strains in the second bacterial group (E. cloacae in this example). The "number of other bacterial strains" is the number of strains not included in either the first bacterial group or the second bacterial group. The "number of other bacterial strains" is counted when a strain not belonging to either the first bacterial group or the second bacterial group is displayed in "Organism" on the screen for explaining the IPG, as shown in Figure 28.

[0101] Next, the process of evaluating the identifiability of a bacterial group based on the first and second numerical values, which is described in S6 to S8, will be described.

[0102] The process of S6 makes it possible to determine the smaller or larger of the first and second numerical values ​​for each of the same peak groups. In other words, it is possible to determine whether the strains belonging to the first bacterial group or the second bacterial group are more prevalent for each of the same peak groups.

[0103] Therefore, processor 10 estimates that the bacterial strain analyzed by mass analysis belongs to the bacterial group corresponding to the larger numerical value, and performs the processes from S7 onward. In this specification, the "bacterial group corresponding to the larger numerical value" is referred to as the "estimated bacterial group," meaning the bacterial group that processor 10 estimates to correspond to a specific identical peak group. In S7 and S8, a first identification ratio is calculated, which is the ratio of the "total number of bacterial strains that match the estimated bacterial groups corresponding to all identical peak groups (third numerical value)" to the "total number of bacterial strains corresponding to all identical peak groups (fourth numerical value)." However, the "total number of bacterial groups corresponding to all identical peak groups" does not include "other bacterial strains" that are not included in either the first or second bacterial group. This is because the microbial analysis method according to this embodiment relates to the identification of a first bacterial group (e.g., E. asburiae) and a second bacterial group (e.g., E. cloacae), which cannot be distinguished by conventional methods, and therefore assumes that the bacterial groups to be identified are narrowed down to the first bacterial group and the second bacterial group.

[0104] Columns H and I in Figures 22 to 27 show the "number of strains matching the estimated bacterial group" and the "first and second numerical values" for each peak group, respectively. The bottom row of column H and the bottom row of column I show the third and fourth numerical values ​​calculated by S7 in Figure 2. In this example, the third numerical value is 556, and the fourth numerical value is 582. Therefore, when Asr fragments are observed by mass spectrometry of strains known to belong to group 1 or group 2, 556 of the 582 strains (96%) can be correctly identified. As described above, E. asburyae and E. cloacae can be distinguished by the molecular mass of the Asr fragment, and Asr can be used as a biomarker for E. asburyae and E. cloacae. Therefore, it was found that E. asburyae and E. cloacae can be distinguished by the molecular mass of the Asr fragment, based on the Asr gene information. The discriminability of C. cloacae could be evaluated.

[0105] In S63, the processor 10 may determine that the first and second bacterial groups are distinguishable by the molecular mass set of the Asr fragments if the first discrimination ratio is equal to or greater than a predetermined threshold. The predetermined threshold is a value greater than 0.5 and less than 1.0. The predetermined threshold may be, but is not limited to, 0.7, 0.8, 0.9, or 0.95, for example.

[0106] As described above, the method for evaluating identifiability according to this embodiment allows for a simple evaluation of the distinguishability between a first bacterial group and a second bacterial group based on information from the Asr gene. In particular, by evaluating distinguishability using multiple strains for each bacterial group, rather than just one, it is possible to replicate the process of actually culturing multiple strains and obtaining actual mass spectra. This point will be explained in more detail below.

[0107] When evaluating the distinguishability of two bacterial groups, even if two mass spectra obtained by analyzing one strain (e.g., a standard strain) for each of the two bacterial groups are distinguishable from each other, this does not necessarily mean that the two bacterial groups will be distinguishable even if the number of strains analyzed for each of the two bacterial groups is increased. For example, as shown in Figures 22 to 27, even within the same bacterial group, some strains are distinct from each other in their mass spectra (i.e., they form separate, identical peak groups). Therefore, it is possible that the bacterial groups are distinguishable for one strain but indistinguishable for many other strains. Conversely, even if two mass spectra obtained by analyzing one strain (e.g., a standard strain) for each of the two bacterial groups are indistinguishable, it is necessary to actually confirm whether the bacterial groups are indistinguishable even if the number of strains analyzed for each of the two bacterial groups is increased. Therefore, in the method for evaluating distinguishability according to this embodiment, by collecting and analyzing information on multiple (preferably many) bacterial groups, distinguishability can be evaluated with higher accuracy than when evaluating distinguishability for only one strain.

[0108] The step of evaluating the possibility of identifying the bacterial group may include a process of evaluating the possibility of identifying the bacterial group for each of the same peak groups, as shown in FIG. 29 below.

[0109] Fig. 29 is a flowchart showing the process of calculating the second classification ratio. The steps shown in Fig. 29 are executed after S6 in Fig. 2. Fig. 29 shows an example in which the steps are executed after S8 in Fig. 2.

[0110] In S9, the processor 10 calculates a second discrimination ratio, which is the ratio of the larger of the first and second numerical values ​​to the sum of the first and second numerical values.

[0111] Taking identical peak group 8 corresponding to accession number WP_059346667.1 of the full-length sequence of the E. asburyae type strain as an example again, the first numerical value (the number of E. asburyae strains) is 5, and the second numerical value (the number of E. cloacae strains) is 1. In this case, processor 10 determines the estimated bacterial group to be E. asburyae and calculates the second discrimination ratio to be 5 / 6. When strains known to belong to group 1 or group 2 are subjected to mass spectrometry and an Asr peak corresponding to identical peak group 8 is observed, it can be seen that 5 out of 6 strains (83%) can be correctly identified.

[0112] 29 allows for a simple evaluation of the possibility of distinguishing between the first and second bacterial groups for each of the identical peak groups. Therefore, for example, when a strain known to belong to the first or second bacterial group is subjected to mass spectrometry, if an actual measured peak corresponding to the identical peak group with a relatively low second discrimination ratio is obtained, the reliability of the discrimination can be considered to be relatively low, whereas if an actual measured peak corresponding to the identical peak group with a relatively high second discrimination ratio is obtained, the reliability of the discrimination can be considered to be relatively high.

[0113] (5-1-5. Determination of molecular mass marker) If it is determined through the above process that the first and second bacterial groups can be distinguished by Asr, in the following process of Figure 30, the processor 10 determines the molecular mass marker to be used for the distinction.

[0114] Fig. 30 is a flowchart showing the process of determining molecular mass markers. The steps shown in Fig. 30 are executed after S8 in Fig. 2. Fig. 30 shows an example in which the steps are executed after S8 in Fig. 2.

[0115] Referring to Figure 30, in step S10, if the processor 10 determines that the first and second bacterial groups are distinguishable by Asr, it determines that the set of molecular masses contained in each of the same peak groups is a molecular mass marker specific to the bacterial group (estimated bacterial group) corresponding to the larger numerical value.

[0116] 30 after evaluating the possibility of identifying a bacterial group for each of the identical peak groups in Fig. 29, the set of molecular masses contained in each identical peak group may be determined to be molecular mass markers specific to the putative bacterial group only if the second discrimination ratio is equal to or greater than a predetermined value. Furthermore, for molecular mass markers whose second discrimination ratio is equal to or less than a predetermined value, the value of the second discrimination ratio, or the fact that the value of the second discrimination ratio is equal to or less than the predetermined value, may be recorded as an annotation for reference.

[0117] By the process of Figure 30, a set of molecular masses of a predetermined identical peak group can be determined as molecular mass markers of the putative bacterial group of the predetermined identical peak group. Therefore, the molecular mass markers can be used to identify microorganisms. The details of this process are described below.

[0118] (5-2. Method for identifying microorganisms using molecular mass markers) (5-2-1. Method for identifying microorganisms) Fig. 31 is a flowchart showing a process for identifying microorganisms. Each step shown in Fig. 31 is executed after S10 in Fig. 30.

[0119] In S21, the processor 10 performs mass analysis on the specific bacterial strain to obtain a mass spectrum. The specific bacterial strain is preferably a strain belonging to the order Enterobacteriaceae, and more preferably a test bacterial strain that has been determined to belong to either the first or second bacterial group by a method other than the microbial analysis method of the present embodiment. Examples of methods other than the microbial analysis method of the present embodiment include conventional biochemical testing or conventional microbial identification methods using the MALDI method.

[0120] In one embodiment, if the specific strain is actually a type strain of E. asburiae, the processor 10 acquires the measured mass spectrum shown in the top row of FIG.

[0121] In S22, the processor 10 determines whether any peaks in the mass spectrum correspond to molecular mass markers for the first bacterial group. In one example, referring to the upper part of FIG. 15 , three peaks with m / z values ​​of 2100.8, 2390.4, and 5123.5 were detected in the mass spectrum. In a positive-mode mass spectrum, ions resulting from the addition of protons to molecules in the sample are observed. Therefore, the molecular weights corresponding to the three peaks with measured m / z values ​​of 2100.8, 2390.4, and 5123.5, after subtracting the mass of the protons, are 2099.8 Da, 2389.4 Da, and 5122.5 Da. Searching for molecular mass sets within an error range of 500 ppm from the molecular weights in FIGS. 22 to 27 reveals that the molecular mass sets in the same peak group 8 (5121.8 Da, 2099.5 Da, and 2388.8 Da) are relevant. The putative bacterial group of the same peak group 8 is E. asburiae. Therefore, the processor 10 determines that there is a peak in the mass spectrum that corresponds to the molecular mass marker of the first bacterial group.

[0122] If there is a peak corresponding to the molecular mass marker of the first group of bacteria (YES in S22), in S23, the processor 10 determines that the specific strain belongs to the first group of bacteria and ends the process. In one example, the processor 10 determines that the specific strain is E. asburyae.

[0123] If there is no peak corresponding to the molecular mass marker of the first bacterial group (NO in S22), in S24, the processor 10 determines whether or not there is a peak in the mass spectrum that corresponds to the molecular mass marker of the second bacterial group.

[0124] If there is a peak corresponding to the molecular mass marker of the second bacterial group (YES in S24), in S25 the processor 10 determines that the specific bacterial strain belongs to the second bacterial group, and ends the process.

[0125] If there is no peak corresponding to the molecular mass marker of the second bacterial group (NO in S22), the processor 10 determines that the specific bacterial strain does not belong to either the first or second bacterial group and terminates the process. However, if the specific bacterial strain is known to belong to either the first or second bacterial group by a method other than the microbial analysis method of this embodiment, the specific bacterial strain is determined to belong to either the first or second bacterial group in S23 or S25, and the process of S26 is generally not performed. The process of S26 is performed, for example, when an analyst analyzes a strain that does not belong to either the first or second bacterial group as a specific bacterial strain, but the specific bacterial strain actually does not belong to either the first or second bacterial group. Also, when an analyst believes a specific bacterial strain belongs to either the first or second bacterial group, but actually does not belong to either the first or second bacterial group. In such cases, the analyst can realize their mistake and redo the testing and / or analysis required to identify the specific bacterial strain.

[0126] As described above, according to the process of Figure 31, microorganisms can be easily identified using molecular mass markers of Asr fragments created by collecting pseudo-strains and performing mass analysis on a computer using the process of Figure 2.

[0127] (5-2-2. Method for identifying microorganisms, including discrimination of Asr peaks) Typically, a mass spectrum obtained by mass spectrometry of an extract of biological cells contains many peaks in addition to peaks corresponding to Asr fragments (hereinafter referred to as "Asr peaks"). Therefore, before comparing the mass spectrum peaks with molecular mass markers, a process may be performed to discriminate the Asr peaks from among the mass spectrum peaks. The Asr peak corresponds to one example of an "acid shock protein peak."

[0128] 32 is a flowchart including the process of distinguishing Asr peaks. In S41, a specific strain is cultured in a medium containing glucose and a medium not containing glucose.

[0129] In S42, mass spectra of the specific strain cultured in a medium containing glucose and a medium not containing glucose are obtained.

[0130] The specific bacterial strain is preferably a strain belonging to the order Enterobacteriaceae, and more preferably a strain that has been determined to belong to either the first or second bacterial group by a method other than the microbial analysis method according to this embodiment. In one example of S41 to S42, an analyst uses ordinary culture equipment and laboratory equipment to culture the specific bacterial strain, performs pretreatment for mass spectrometry on the cultured specific bacterial strain, and then operates MS 16 to perform mass spectrometry. In another example of S41 to S42, processor 10 controls automated culture equipment and laboratory equipment (not shown) to culture the specific bacterial strain, performs pretreatment for mass spectrometry on the cultured specific bacterial strain, and then performs mass spectrometry using MS 16.

[0131] In S43, the processor 10 compares the mass spectrum of the specific bacterial strain cultured in a glucose-free medium with the mass spectrum of the specific bacterial strain cultured in a glucose-containing medium, and identifies the Asr peak in the mass spectrum of the specific bacterial strain cultured in a glucose-containing medium. Research by the inventors has shown that bacteria belonging to the order Enterobacteriaceae produce Asr in a glucose-containing medium. Therefore, if the specific bacterial strain belongs to the order Enterobacteriaceae, and a peak corresponding to a predetermined m / z is detected in the mass spectrum of the specific bacterial strain cultured in a glucose-containing medium but not in the mass spectrum of the specific bacterial strain cultured in a glucose-free medium, the peak corresponding to the predetermined m / z is identified as an Asr peak.

[0132] In S44, the processor 10 determines whether or not there is a peak among the Asr peaks that corresponds to the molecular mass marker of the first bacterial group.

[0133] If there is a peak corresponding to the molecular mass marker of the first bacterial group (YES in S44), in S45 the processor 10 determines that the specific bacterial strain belongs to the first bacterial group, and ends the process.

[0134] If there is no peak corresponding to the molecular mass marker of the first bacterial group (NO in S44), in S46, the processor 10 determines whether there is a peak among the Asr peaks that corresponds to the molecular mass marker corresponding to the second bacterial group.

[0135] If there is a peak corresponding to the molecular mass marker of the second bacterial group (YES in S46), in S47 the processor 10 determines that the specific bacterial strain belongs to the second bacterial group, and ends the process.

[0136] If there is no peak corresponding to the molecular mass marker of the second bacterial group (NO in S46), in S48 it is determined that the strain subjected to mass analysis does not belong to either the first or second bacterial group, and the process is terminated.

[0137] According to the process of Figure 32, when a specific bacterial strain belongs to the order Enterobacteriaceae, culturing the strain in a glucose-containing medium can easily obtain a specific bacterial strain containing Asr intracellularly. Furthermore, culturing the strain in a glucose-free medium can easily obtain a specific bacterial strain not containing Asr intracellularly. Therefore, by comparing the mass spectrum of a specific bacterial strain cultured in a glucose-containing medium with that of a specific bacterial strain cultured in a glucose-free medium, the Asr peak can be easily distinguished from the numerous peaks in the mass spectrum. Therefore, by searching for a molecular mass marker corresponding to the Asr peak, it is possible to easily determine whether the specific bacterial strain belongs to the first or second bacterial group. Therefore, the specific bacterial strain can be more easily identified than when searching for a peak corresponding to a molecular mass marker from the numerous peaks in the mass spectrum, as in the process of Figure 31. Furthermore, although unlikely, this eliminates the risk of mistaking a peak of another protein having the same molecular weight as the Asr peak for an Asr peak.

[0138] (5-2-3. Method for distinguishing between E. asburyae and E. cloacae) Figure 33 shows a method for distinguishing between E. asburyae and E. cloacae, which is one example of the method shown in Figure 31. In the processing shown in Figure 33, E. asburyae and E. cloacae can be distinguished from each other using the molecular mass set (molecular mass marker) corresponding to the putative bacterial group obtained in the processing shown in Figure 2.

[0139] In S61, the processor 10 acquires a mass spectrum of a specific strain of bacteria. The specific strain is preferably a strain that is known to belong to at least one of E. asburiae and E. cloacae by a method other than the microbial analysis method according to this embodiment.

[0140] In S62, processor 10 determines whether any peaks in the mass spectrum correspond to molecular mass markers of E. asburyae. More specifically, processor 10 determines whether any mass spectrum peaks have measured m / z values ​​that fall within the error range for all of the multiple molecular masses included in the predetermined molecular mass markers of E. asburyae.

[0141] The molecular mass markers of E. asburyae are molecular mass sets corresponding to the same peak groups 1, 3, 8, 15, 16, 18, 20, and 21 in Figures 22 to 27, and are (5149.8, 2069.4, 2099.5, 2388.8), (5121.8, 2069.4, 2099.5, 2388.8), (5121.8, 2099.5, 2388.8), and (5121.8, 2076. The set of at least one molecular mass selected from the group consisting of (5121.8, 2069.4, 2388.8), (5121.8, 2069.4, 2364.8, 2388.8), (5121.8, 2069.4, 2099.5, 2113.5, 2388.8), and (5135.8, 2069.4, 2099.5, 2388.8). The notation (molecular mass 1, molecular mass 2) above indicates a set of molecular masses including two molecular masses, molecular mass 1 and molecular mass 2. For example, (5149.8, 2069.4, 2099.5, 2388.8) indicates a set of molecular masses containing the four m / z values ​​5149.8, 2069.4, 2099.5, and 2388.8. The above molecular mass markers for E. asburyae contain an error of 500 ppm.

[0142] Therefore, processor 10 determines whether any peaks in the mass spectrum have molecular masses within 500 ppm of the molecular mass marker for E. asburyae. More specifically, processor 10 calculates the corresponding theoretical m / z for each molecular mass included in the predetermined molecular mass marker and determines whether any peaks in the mass spectrum are within the error range of the theoretical m / z. Then, if peaks are detected in the mass spectrum within 500 ppm of the corresponding theoretical m / z for all molecular masses included in the predetermined molecular mass marker, processor 10 determines that any peaks in the mass spectrum correspond to the molecular mass marker for E. asburyae. Processor 10 may convert one or more measured m / z values ​​corresponding to one or more peaks in the mass spectrum into one or more molecular masses, and then determine whether any molecular mass markers are within the error range of the one or more molecular masses.

[0143] If there is a peak corresponding to the molecular mass marker of E. asburyae (YES in S62), in S63 the processor 10 determines that the specific strain belongs to E. asburyae, and ends the process.

[0144] If there is no peak corresponding to the molecular mass marker for E. asburiae (NO in S62), in S64, processor 10 determines whether or not there is a peak in the mass spectrum that corresponds to the molecular mass marker for E. cloacae. More specifically, processor 10 determines whether or not there is a mass spectrum peak for which the measured m / z falls within the error range for all of the multiple molecular masses included in the specified molecular mass marker for E. cloacae.

[0145] The molecular mass markers of E. cloacae are molecular mass sets corresponding to the same peak groups 2, 4 to 7, 9 to 14, and 17, and are (4152.7, 2611.1, 2099.5, 2388.8), (4122.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2099.5, 2021.4), (4180.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2085.4, 2388.8), (4168.7, 2611.1, 2099.5, 2388.8), The set of at least one molecular mass marker selected from the group consisting of (4152.7, 2611.1, 2125.5, 2388.8), (4152.7, 2611.1, 2099.5, 2416.9), (4152.7, 2611.1, 2069.4, 2388.8), (4223.8, 2611.1, 2099.5, 2388.8), (4010.5, 2611.1, 2099.5, 2388.8) and (5163.8, 2069.4, 2099.5, 2388.8). The molecular mass markers for E. cloacae include an error of 500 ppm. Therefore, the processor 10 determines whether the E. cloacae molecular mass markers are of the same or different molecular masses. In the same manner as in the case of the molecular mass marker of E. asburiae, it is determined whether or not there is a peak in the mass spectrum that exhibits a molecular mass within 500 ppm of the molecular mass marker of E. cloacae.

[0146] If there is a peak corresponding to the molecular mass marker of E. cloacae (YES in S64), in S65 the processor 10 determines that the specific strain belongs to E. cloacae, and ends the process.

[0147] If there is no peak corresponding to the molecular mass marker of E. cloacae (NO in S64), in S66, the strain subjected to mass analysis is determined to belong to neither E. asburyae nor E. cloacae, and the process is terminated. However, if the specific strain is a strain known to belong to either E. asburyae or E. cloacae by a method other than the microbial analysis method according to this embodiment, the specific strain is determined to belong to E. asburyae or E. cloacae in S63 or S65, and the process of S66 is not generally performed. The process of S66 is performed, for example, when a strain that does not belong to either E. asburyae or E. cloacae is analyzed as a specific strain, and the specific strain is actually determined to belong to neither E. asburyae nor E. cloacae. This may be the case when the specific strain that the analyst thought belonged to either E. asburyae or E. cloacae actually does not belong to either E. asburyae or E. cloacae. In such cases, the analyst can realize the mistake and redo the tests and / or analyses required to identify the specific strain.

[0148] As described above, according to the processing in Figure 33, E. asburiae and E. cloacae can be easily distinguished from each other by using the molecular mass marker of the Asr fragment, which was difficult to achieve using conventional microbial identification methods using MALDI that focused on ribosomal peaks.

[0149] [6. Embodiment 2] (6-1. Evaluation of Identification Potential of Microorganisms) As described above, the inventors were able to evaluate the identification potential of two bacterial groups based on the number and position (m / z) of Asr peaks by calculating the number and position (m / z) of Asr peaks. Through further investigation, the inventors found that the number of Asr peaks may differ between the two bacterial groups. This led the inventors to discover that it is also possible to evaluate the identification potential of bacterial groups based solely on the number of Asr peaks.

[0150] Specifically, the inventors compared the ATCC 11775 and NBRC 107761 strains, which are standard strains of closely related species E. coli and E. albertii, respectively, and found that the number of cleavage sequences differed. Table 1 shows the full-length sequences of the ATCC 11775 and NBRC 107761 strains.

[0151]

[0152] Referring to Table 1, the full-length sequence of the ATCC 11775 strain and the full-length sequence of the NBRC 107761 strain differ in the number of cleavage sequences QKAQ, two and one, respectively. That is, while there are three Asr fragments (three Asr peaks) in E. coli, there are two Asr fragments (two Asr peaks) in E. albertii. This demonstrates that the ATCC 11775 strain and the NBRC 107761 strain can be distinguished by the number of Asr peaks (i.e., the number of Asr fragments) or the number of cleavage sequences in the full-length sequence.

[0153] The inventors performed a process ( FIG. 34 ) relating to a method for evaluating the distinguishability of the first and second bacterial groups according to embodiment 2, and analyzed a large number of amino acid sequences for each of E. coli and E. albertii, demonstrating that E. coli and E. albertii can be distinguished by the number of Asr peaks (i.e., the number of Asr fragments) or the number of truncated sequences in the full-length sequence (see Example 3).

[0154] Fig. 34 is a flowchart showing a process relating to the method for evaluating the identifiability of microorganisms according to embodiment 2. In one example, the process in Fig. 34 is performed by processor 10. In another example, at least a part of the process in Fig. 34 may be performed by an analyst operating analytical device 100. S100, S101, and S102 to S110 in Fig. 34 correspond to examples of a "first process," a "second process," and a "third process," respectively.

[0155] 34 will be specifically described below using an example (Example 3) relating to the discrimination between E. coli (first bacterial group) and E. albertii (second bacterial group).

[0156] Steps S100 to S101 in Figure 34 correspond to steps S0 to S1 in Figure 2. In one embodiment of step S100, the processor 10 selects, from among a large number of full-length sequences obtained using NCBI, full-length sequences that include "Escherichia coli" in the genus / species name and full-length sequences that include "Escherichia albertii" in the 'Organism' column. The full-length sequence that includes "Escherichia coli" in the genus / species name will also be referred to as the full-length sequence of E. coli hereinafter. The full-length sequence that includes "Escherichia albertii" in the 'Organism' column will also be referred to as the full-length sequence of E. albertii hereinafter.

[0157] In one embodiment of S101, the processor 10 identifies cleavage sequences (QKAQ sequence and QNAQ sequence) of the full-length sequence of E. coli and the full-length sequence of E. albertii.

[0158] In S102 to S110, the processor 10 calculates the discrimination ratio between the first and second bacterial groups based on the number of Asr peaks in the mass spectrum.

[0159] As described above, the number of Asr peaks is equal to the number of Asr fragments. Furthermore, the number of Asr peaks corresponds to the number of cleavage sequences in the full-length sequence (the sum of the number of QKAQ sequences and the number of QNAQ sequences). More specifically, the number of Asr peaks is one greater than the number of cleavage sequences in the full-length sequence. Therefore, the process of S102 corresponds to a process of calculating the discrimination ratio between the first and second bacterial groups based on the number of Asr fragments or the number of cleavage sequences in the full-length sequence.

[0160] In S102, processor 10 calculates the number N1 of full-length sequences obtained for the first bacterial group and the number N2 of full-length sequences obtained for the second bacterial group from the full-length sequences obtained in S1. Hereinafter, the "full-length sequences obtained for the first bacterial group" and the "full-length sequences obtained for the second bacterial group" will also be simply referred to as the "full-length sequences of the first bacterial group" and the "full-length sequences of the second bacterial group."

[0161] In one example of S102, the processor 10 calculates the number N1 of full-length sequences of E. coli and the number N2 of full-length sequences of E. albertii obtained in S100. In this example, N1 and N2 are, for example, 10,673 and 8 (see Table 2).

[0162] In S103, the processor 10 counts the number of cleavage sequences in the full-length sequence of the first bacterial group and calculates the mode M1 of the number of cleavage sequences. The number of cleavage sequences is the sum of the number of QKAQ sequences and the number of QNAQ sequences. M1 is the number of cleavage sequences in the full-length sequence of the first bacterial group that corresponds to the largest number of full-length sequences.

[0163] In S104, the processor 10 tallies the number of cleavage sequences in the full-length sequences of the second bacterial group, and determines the mode M2 ​​of the number of cleavage sequences that corresponds to the largest number of full-length sequences among the cleavage sequences. The number of cleavage sequences is the sum of the number of QKAQ sequences and the number of QNAQ sequences. M2 is the number of cleavage sequences that corresponds to the largest number of full-length sequences among the cleavage sequences in the full-length sequences of the second bacterial group. Note that M2 is a different value from M1.

[0164] In one embodiment of S103, the processor 10 calculates M1 by counting the number of cleavage sequences in the full-length sequence of E. coli.

[0165] In one embodiment of S104, the processor 10 calculates M2 by tallying the number of cleavage sequences in the full-length sequence of E. albertii.

[0166] Table 2 shows the results of the count of the number of cleaved sequences in this example.

[0167]

[0168] Referring to Table 2, M1 is the "number of cleavage sequences" of 2, which corresponds to 10,638, the largest number in the "number of full-length sequences of E. coli." Also, M2 is the "number of cleavage sequences" of 1, which corresponds to 8, the largest number in the "number of full-length sequences of E. albertii."

[0169] In S105, the processor 10 calculates the number P1 of full-length sequences having the cleavage sequence of M2 in the full-length sequences of the first bacterial group.

[0170] In S106, the processor 10 calculates the number P2 of full-length sequences having the cleavage sequence of M1 in the full-length sequences of the second bacterial group.

[0171] In the example of Table 2, the value of P1 is 19, and P2 is 0. In S107, when P2 is 0, processor 10 sets the third discrimination ratio to 1 when the number of Asr peaks is (M1+1) (i.e., when the number of Asr peaks is 1 greater than M1). From the above, it can be seen that when the third discrimination ratio is 1, processor 10 can determine that a sample that is either the first bacterial group or the second bacterial group is 100% the first bacterial group when the number of Asr peaks is (M1+1).

[0172] In one example of S107, since the number of full-length sequences having two cleavage sequences in the full-length sequence of E. albertii is 0, the processor 10 sets the third discrimination ratio when the number of Asr peaks is 3 to 1. Therefore, the discrimination ratio (third discrimination ratio) between E. coli and E. albertii when the number of Asr peaks is 3 is calculated to be 100%. As a result, when a sample that is either E. coli or E. albertii has three Asr peaks, it is considered that the sample can be determined to be E. coli.

[0173] In S108, if the number of full-length sequences having the M2 cleavage sequence in the full-length sequence of the first bacterial group is 0, the processor 10 sets the fourth discrimination ratio when the number of Asr peaks is (M2+1) to 1. From the above, it can be seen that when the fourth discrimination ratio is 1, the processor 10 can determine that when the number of Asr peaks in a sample that is either the first bacterial group or the second bacterial group is (M2+1), the sample is 100% the second bacterial group.

[0174] In one embodiment of S108, the number of full-length sequences having one cleavage sequence in the full-length sequences of E. coli is not 0, so the processor 10 does not calculate the fourth discrimination ratio.

[0175] In S109, if P2 is not 0 and if the value obtained by subtracting P2 from N2 (i.e., N2-P2) is equal to or greater than a predetermined multiple of P2, the processor 10 sets the fifth discrimination ratio when the number of Asr peaks is (M1+1) to the value obtained by dividing (N2-P2) by N2 (i.e., (N2-P2) / N2). An example of the predetermined multiple is, but is not limited to, 100 times, and more specifically, 500 times.

[0176] In S110, if P1 is not 0 and (N1-P1) is equal to or greater than a predetermined multiple of P1, the processor 10 sets the fifth classification ratio when the number of Asr peaks is (M2+1) to (N1-P1) / N1. Examples of the predetermined multiple include, but are not limited to, 100, more specifically 300, and even more specifically 500. In one embodiment, the predetermined multiple is 500.

[0177] In S109 of this embodiment, the processor 10 calculates the fifth discrimination ratio when the number of Asr peaks is 2 as (10,673-19) / 10,673 because the number Q2 of full-length sequences having two cleavage sequences in the full-length sequence of E. albertii is not 0 and (10,673-19) is 500 times or more greater than 19. Since (10,673-19) / 10,673 = 10,654 / 10,673 = 0.9982, the processor 10 calculates that the fifth discrimination ratio between E. coli and E. albertii when the number of Asr peaks is 2 is 99.8% or greater. As a result, when a sample that is either E. coli or E. albertii has two Asr peaks, the sample is considered to be E. It is considered that it can be determined to be I. albertii (error is 0.2% or less).

[0178] That is, if (1) the full-length sequence of the first bacterial group contains a full-length sequence corresponding to the most frequent value M2 of the number of cleavage sequences in the second bacterial group, and (2) the number P1 of full-length sequences corresponding to M2 in the first bacterial group is extremely small compared to the total number N1 of full-length sequences in the first bacterial group (i.e., if P1 << (N1 - P1)), and an Asr peak (M2 + 1) is observed upon sample analysis, the sample can generally be determined to be the second bacterial group. More specifically, the sample can be determined to be the second bacterial group with a probability of (N1 - P1) / N1, with an error of P1 / N1 or less. As described above, the process of S109 also makes it possible to calculate the discrimination ratio in cases where (1) the full-length sequence of the first bacterial group contains a full-length sequence corresponding to the most frequent value M2 of the number of cleavage sequences in the second bacterial group, and (2) the number P1 of full-length sequences corresponding to M2 in the first bacterial group is extremely small compared to the total number N1 of full-length sequences in the first bacterial group.

[0179] In S110 of this embodiment, the processor 10 does not calculate the sixth discrimination ratio because the number of full-length sequences having two cleavage sequences in the full-length sequences of E. coli is 0.

[0180] As described above, E. albertii and E. coli can be distinguished by the number of Asr peaks, and Asr can be used as a biomarker for E. albertii and E. cloacae. Therefore, the distinguishability between E. albertii and E. coli could be evaluated based on information on the Asr gene (the number of cleavage sequences or the number of Asr fragments).

[0181] According to the process of FIG. 34, the discrimination ratio between the first and second bacterial groups according to the number of Asr peaks can be calculated by simple calculation.

[0182] As described above, according to the method for evaluating the distinguishability of microorganisms of embodiment 2, similar to embodiment 1, information regarding the distinction of Enterobacteriaceae bacteria can be obtained in a simple manner. In particular, according to embodiment 2, it is not necessary to calculate the m / z of fragment sequences from the full-length sequence, consider whether the m / z is distinguishable in the mass spectrum, and then calculate mass spectrum groups, as in embodiment 1; it is sufficient to simply count the number of truncated sequences in the full-length sequence. Therefore, information regarding the distinction of Enterobacteriaceae bacteria can be obtained even more easily than in embodiment 1. On the other hand, in embodiment 1, even for a first strain and a second strain that have the same number of Asr peaks, the distinction ratio between the first strain and the second strain can be calculated based on the positions of the Asr peaks.

[0183] Furthermore, the first embodiment may be further applied to a sample to which the second embodiment is applied, and the discrimination ratio based on the mass spectrum pattern may be calculated taking into account the position (m / z) of the Asr peak. This makes it possible to examine whether or not the discrimination ratio can be improved when the position of the Asr peak is also taken into account, even for portions that cannot be completely separated based on the number of Asr peaks alone, such as when there are two Asr peaks in E. coli and E. albertii.

[0184] (6-2. Method for identifying microorganisms based on the number of Asr peaks) Figure 35 is a flowchart showing processing relating to the method for identifying microorganisms according to embodiment 2. Each step shown in Figure 35 is executed after S104 in Figure 34. However, S102 in Figure 34 does not have to be executed before each step shown in Figure 35. In one example of Figure 35, the first bacterial group consists of E. coli, and the second bacterial group consists of E. albertii.

[0185] The process of S121 in Figure 35 corresponds to S21 in Figure 31. In S122, the processor 10 determines that the specific bacterial strain belongs to the first bacterial group if the number of Asr peaks in the mass spectrum is (M1 + 1). In one example, the processor 10 determines that the specific bacterial strain belongs to E. coli if the number of Asr peaks in the mass spectrum is 3.

[0186] In S123, if the number of Asr peaks in the mass spectrum is (M2 + 1), the processor 10 determines that the specific strain belongs to the second group of bacteria. In one embodiment, if the number of Asr peaks in the mass spectrum is 2, the processor 10 determines that the specific strain belongs to E. albertii.

[0187] As described above, according to the process of Fig. 35, the first and second bacterial groups can be easily distinguished from each other using the number of Asr peaks. In particular, E. coli and E. albertii, which have been difficult to distinguish between using conventional MALDI-based discrimination methods and conventional biochemical property tests, can be easily distinguished from each other based on the difference in the number of Asr peaks.

[0188] If the processing of Fig. 35 is performed after S107 of Fig. 34, S122 may be performed only when the third classification ratio is 1. Similarly, if the processing of Fig. 35 is performed after S108 of Fig. 34, S123 may be performed only when the fourth classification ratio is 1. With this configuration, it is possible to identify only strains that are considered to be identifiable without erroneous determination.

[0189] Furthermore, when the processing of Figure 35 is performed after S109 of Figure 34, S122 may be performed only when the fifth classification ratio is equal to or greater than a predetermined threshold. Similarly, when the processing of Figure 35 is performed after S110 of Figure 34, S123 may be performed only when the sixth classification ratio is equal to or greater than a predetermined threshold. The predetermined threshold is set appropriately. The predetermined threshold is, for example, 80%, more specifically 90%, or even more specifically 95%. With this configuration, it is possible to identify only strains that are considered to have a sufficiently low possibility of erroneous determination.

[0190] [7. Examples] [7-1. Example 1] The method for evaluating the distinguishability of microorganisms according to embodiment 1 was performed as follows, targeting E. asburyae and E. cloacae. (1) NCBI blastp was performed using ADF61781.1, the GenBank accession number for Asr of ATCC 35953, a type strain of E. asburyae, as a query, and the 52 accession numbers and full-length sequences obtained were listed in order of homology (Table A: Figures 4 to 8). (2) NCBI blastp was performed using ADF61781.1, the GenBank accession number for Asr of ATCC 13047, a type strain of E. cloacae, as a query, and the 46 accession numbers and full-length sequences obtained were listed in order of homology (Table B: Figures 9 to 12). (3) In Tables A and B, only the accession numbers and full-length sequences that were in the top 15 in terms of homology (the lines above the double ruled line L1 in Figure 5 and the lines above the double ruled line L2 in Figure 10) were retained, and the accession numbers and full-length sequences ranked 16th and below were excluded from further processing. (4) For the 30 full-length sequences processed in (3), signal peptides were excluded, and if any of the remaining amino acid sequences of the protein body contained a glutamine-lysine-alanine-glutamine (QKAQ) sequence or a glutamine-asparagine-alanine-glutamine (QNAQ) sequence, a set of fragment sequences was created by cleaving the C-terminal end of the sequence. A set of molecular masses corresponding to the set of fragment sequences was then calculated (Table C: Figures 16 to 21). (5) The 30 molecular mass sets obtained in (4) were compared with each other, and when all molecular masses in the molecular mass set matched within an error of 500 ppm, they were assigned to the same peak group (columns A to C in Table D (Figures 22 to 27)). (6) Using information on identical protein groups from NCBI, the strains belonging to each identical peak group obtained in (5) were listed, and the number of strains belonging to E. asburiae (first value) and the number of strains belonging to E. cloacae (second value) were calculated (columns D to E in Figures 22 to 27).Specifically, for each accession number corresponding to a given identical peak group, information on identical protein groups was displayed to confirm a list of strains containing the same full-length sequence as the accession number, and the number of E. asburiae strains and E. cloacae strains in the list was tallied. A script from FileMaker (Claris International Inc.) was used for this talliation. (7) Furthermore, the number of E. asburiae strains and the number of E. cloacae strains were tallied for each identical peak group, and the molecular mass set and accession number were assigned to the bacterial group with the largest number of corresponding strains for each identical peak group (presumed bacterial group: E. asburiae or E. cloacae in this example) (column G in Figures 22 to 27). In addition, the number of strains that matched the estimated bacterial population for each identical peak group was determined (column H in Figures 22 to 27). (8) A third value (bottom row of column H in Figure 27) was calculated, which is the "total number of strains that matched the estimated bacterial population corresponding to all identical peak groups." Then, a fourth value (bottom row of column I in Figure 27) was calculated, which is the "total number of strains that matched the estimated bacterial population corresponding to all identical peak groups." The third value was 562, and the fourth value was 582. Then, a first discrimination rate (96%) was calculated, which is the ratio of the third value to the fourth value. As described above, the differentiation possibility between E. asburyae and E. cloacae was evaluated based on the Asr gene information, and the result was that 556 strains (96%) of the 582 strains could be correctly differentiated. Therefore, the method for evaluating the microbial differentiation method according to embodiment 1 was able to differentiate E. asburyae and E. The differential diagnosis possibility of C. cloacae could be evaluated.

[0191] According to Example 1, the method for evaluating the microbial discrimination method according to Embodiment 1 can be said to be a method that can evaluate the discriminability of the first and second bacterial groups without actually collecting and analyzing the strains by reproducing on a computer the process of collecting and analyzing a large number of strains of the first and second bacterial groups. This makes it possible to simply and quickly evaluate the discriminability of the first and second bacterial groups that are expected to be discriminated.

[0192] [7-2. Example 2] The method for distinguishing microorganisms according to embodiment 1 was carried out as follows, with the aim of distinguishing between E. asburyae and E. cloacae.

[0193] The ATCC 35953 strain was used as a test bacterial strain, which was known to be either E. asburiae or E. cloacae, but the identity of the strain had not been determined. Mass spectra were obtained for the test bacterial strain using the following procedure. (1) The test bacterial strain was inoculated into IFO804 medium, a glucose-containing medium, and cultured at 37°C for 18 hours. The IFO804 medium contained 0.5% glucose, 0.5% hypopeptone, 0.5% yeast extract, and 0.1% magnesium sulfate heptahydrate. (2) The medium containing the grown bacterial cells was centrifuged, the supernatant was removed, and purified water was added to the remaining sediment to a turbidity of 1 and suspended. (3) 0.5 mL of the suspension of the sediment obtained in (2) was transferred to a separate centrifuge tube and centrifuged. The supernatant was discarded, and the sediment was collected again, after which it was resuspended in 50 μL of 1% TFA. (4) (3) was centrifuged again, and the supernatant was separated and diluted 100-fold with purified water. (5) 1 μL of a 5 mg / ml 4-hydroxycinnamic acid solution (solvent: a mixture of equal parts acetonitrile and purified water) was dropped onto a MALDI target plate and dried to a solid. (6) 1 μL of the diluted supernatant obtained in (4) was dropped onto the dried product obtained in (5), and the mixture was dried again to a solid, to prepare a sample for MALDI measurement. (7) A mass spectrometer (Shimadzu MALDI-8020) was set in positive linear mode with a measurement mass range of 1,500 to 10,000, and a mass spectrum was obtained (top panel of Figure 15). (8) For the three peaks at 2100.8, 2390.4, and 5123.5 in the mass spectrum acquired in (7), the atomic masses of the protons were subtracted to calculate the corresponding molecular masses of 2099.8 Da, 2389.4 Da, and 5122.5 Da. (9) A molecular mass set that falls within an error range of 500 ppm of the molecular masses was searched for in Figures 22 to 27. As a result, it was found that the molecular mass set in the same peak group 8 (molecular masses: 5121.8, 2099.5, 2388.8) corresponded. Since the estimated bacterial group in the same peak group 8 is E. asburyae, the test bacterial organism was determined to be E. asburyae. As described above, the microorganism identification method according to embodiment 1 was able to identify E. asburyae and E. The results were able to correctly identify strains that could not be identified as either P. cloacae or P. cloacae.

[0194] Example 2 shows that the method for identifying microorganisms according to embodiment 1 is a method that can accurately distinguish between group 1 (E. asburiae) and group 2 (E. cloacae), which could not be distinguished by conventional methods, using MALDI, which is the same measurement method as the conventional method. Therefore, it can be said that the method for identifying microorganisms according to embodiment 1 is a method that can distinguish important groups of bacteria that could not be distinguished before, taking advantage of the advantages of MALDI, namely, simplicity and speed.

[0195] [7-3. Example 3] In Example 3, the method for evaluating the identifiability of microorganisms according to embodiment 2 was performed using E. coli and E. albertii as the target organisms as follows. (1) The NCBI database was searched (May 2024) using the search term "acid shock protein," yielding 71,545 entries. Each of the NCBI entries typically includes information such as an amino acid sequence (including those predicted from gene sequences), an accession number, and a species name. (2) To exclude incomplete entries from the 71,545 entries obtained in (1), a filter was applied to the 71,545 entries, with the condition that the protein name contained "acid," "shock," and "partial," yielding 39,508 entries. As a result, 39,508 entries containing 39,508 full-length sequences were obtained. (3) From the 39,508 entries obtained in (2), those containing "Escherichia coli" in the genus / species name column were selected, yielding 10,673 entries. This resulted in 10,673 full-length E. coli sequences. (4) From the 39,468 entries obtained in (2), those containing "Escherichia albertii" in the "Organism" column were selected, yielding 8 entries. This resulted in 8 full-length E. albertii sequences. (5) From the 10,673 full-length sequences obtained in (3), the number of QKAQ sequences, which are cleavage sequences, was calculated and tabulated, yielding the results in Table 2. Note that when the 10,673 full-length sequences were searched for QNAQ sequences, which are other cleavage sequences, none of the 10,673 full-length sequences contained any QNAQ sequences. (6) The number of QKAQ sequences was calculated for the eight full-length sequences obtained in (4) and tabulated to obtain the results shown in Table 2. Furthermore, when the eight full-length sequences were searched for QNAQ sequences, which are other cleavage sequences, none of the eight full-length sequences contained any QNAQ sequences. (7) Based on Table 2, the possibility of distinguishing samples based on Asr peaks when they were known to be either E. coli or E. albertii was evaluated.

[0196] Referring to Table 2, all 8 full-length sequences of E. albertii had a cleavage sequence of 1. Furthermore, the majority (99.67% or more) of the full-length sequences of E. coli had a cleavage sequence of 2. Furthermore, of the full-length sequences of E. coli, the proportion of sequences with a cleavage sequence of 2 was 0.2% or less.

[0197] From the above, it was estimated that when a sample is known to be either E. coli or E. albertii, if the number of Asr peaks is 3, it can be determined to be Escherichia coli, and if it is 2, it can generally be determined to be Escherichia albertii (error rate of approximately 0.2%). Therefore, the method for evaluating the microbial identification method according to embodiment 2 was able to evaluate the possibility of distinguishing between E. coli and E. albertii.

[0198] According to Example 3, the method for evaluating the microbial discrimination method according to Embodiment 2 can be said to be a method that can evaluate the discriminability of the first and second bacterial groups without actually collecting and analyzing the strains by reproducing on a computer the process of collecting and analyzing a large number of strains of the first and second bacterial groups. This makes it possible to simply and quickly evaluate the discriminability of the first and second bacterial groups that are expected to be discriminated.

[0199] After carrying out Example 3, information on the items used in Example 3 (full-length sequences and QKAQ numbers) was recorded as SEQ ID NOs: 218 to 429 in the Sequence Listing and Tables 3 to 8. Details thereof are explained below in (8) and (9). The full-length sequences and QKAQ numbers of the 10,673 items corresponding to E. coli obtained in (8) and (3) were recorded as SEQ ID NOs: 218 to 424 in the Sequence Listing and Tables 3 to 5 and 7.

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206] The 10,673 entries obtained in (3) included 207 full-length sequences. These 207 full-length sequences are recorded as SEQ ID NOs: 218 to 424 in the sequence listing.

[0207] In addition, the number of QKAQs (QKAQ number) contained in each of SEQ ID NOs: 218 to 424 and the number of full-length sequences (sequence number) corresponding to each of SEQ ID NOs: 218 to 424 are recorded in Tables 3 to 5. For example, referring to Table 3, it can be seen that the full-length sequence of SEQ ID NO: 220 contains one QKAQ, and the number of full-length sequences having the full-length sequence of SEQ ID NO: 220 is one.

[0208] Table 7 shows the number of full-length sequences with QKAQ numbers of 0 to 3 for 207 full-length sequences, as well as the total number of full-length sequences. Referring to the first row of Table 7, it can be seen that there are 6 full-length sequences with a QKAQ number of 0, and a total of 10 full-length sequences with a QKAQ number of 0. (9) The full-length sequences and QKAQ numbers of the 8 full-length sequences corresponding to E. albertii obtained in (4) are recorded in SEQ ID NOs: 425 to 429 of the Sequence Listing and in Tables 6 and 8.

[0209] The eight items obtained in (4) contained five full-length sequences, which are recorded as SEQ ID NOs: 425 to 429 in the sequence listing.

[0210] In addition, the number of QKAQs contained in each of SEQ ID NOs: 425 to 429 (QKAQ number) and the number of full-length sequences corresponding to each of SEQ ID NOs: 425 to 429 (sequence number) are recorded in Table 6.

[0211] Table 8 shows the number of types of full-length sequences with each QKAQ number of 0 to 3 for eight types of full-length sequences, and the total number of full-length sequences.

[0212] In Example 3, Tables 3 to 8 are used to record information about the items obtained in Examples 3(3) and 3(4), and are not necessarily required for calculating the third to sixth discrimination ratios. However, in other examples of Embodiment 2, instead of creating Table 2 in Examples 3(5) to 3(7), Tables 7 and 8 may be created, and the third to sixth discrimination ratios may be calculated from these. As described above, the method for calculating the third to sixth discrimination ratios according to Embodiment 2 is not particularly limited, and any method may be used as long as it can correctly count the number of Asr peaks. Therefore, the user can appropriately set the calculation method depending on the amount of processing, the ease of viewing the processing steps and results, etc.

[0213] Aspects It will be understood by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.

[0214] (Item 1) A method for analyzing microorganisms according to one embodiment includes the steps of obtaining, for each of a first bacterial group and a second bacterial group that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family, a plurality of full-length sequences that are the full-length amino acid sequences of a plurality of corresponding acid shock proteins; identifying cleavage sequences that indicate the sites at which the plurality of full-length sequences are fragmented; and determining the mass spectral patterns of a plurality of acid shock protein fragments obtained by fragmenting each of the plurality of full-length sequences, and calculating the discrimination ratio between the first bacterial group and the second bacterial group based on the mass spectral patterns.

[0215] According to the method for analyzing microorganisms described in paragraph 1, by calculating the discrimination ratio based on the mass spectral pattern using the amino acid sequence information of Asr of a large number of strains of each bacterial group, it is possible to evaluate discrimination possibility without the need to actually collect a large number of bacterial strains and perform mass analysis of Asr. Therefore, the inventor can evaluate the discrimination possibility of two bacterial groups without spending time and effort. As described above, the method for evaluating discrimination possibility according to the embodiment allows information on discrimination of bacterial groups of the order Enterobacteriaceae to be obtained in a simple manner.

[0216] (2) In the method for analyzing microorganisms according to the first aspect, the step of calculating the discrimination ratio includes a step of obtaining a set of amino acid sequences of a plurality of acid shock protein fragments obtained by fragmenting each of a plurality of full-length acid sequences, a step of obtaining a set of theoretical molecular masses of a plurality of acid shock protein fragments corresponding to the set of amino acid sequences of the plurality of acid shock protein fragments, and a step of assigning one or more sets of molecular masses of the plurality of acid shock protein fragments that show the same peak in the measured mass spectrum to the same identical peak group, thereby obtaining one or more identical peaks. The method includes the steps of: setting groups; determining a first numerical value, which is the number of strains belonging to a first group, and a second numerical value, which is the number of strains belonging to a second group, among the strains containing full-length sequences corresponding to each of the same peak groups; distinguishing between the smaller and larger of the first and second numerical values ​​for each of the same peak groups; calculating a third numerical value, which is the sum of the larger numerical values ​​of all the same peak groups, and a fourth numerical value, which is the sum of the first and second numerical values ​​of all the same peak groups; and calculating a first discrimination ratio, which is the ratio of the third numerical value to the fourth numerical value.

[0217] According to the microbial analysis method described in paragraph 2, by calculating the corresponding molecular mass using the amino acid sequence information of Asr for a large number of strains of each bacterial group, it is possible to reproduce on a computer the results of actually collecting a large number of bacterial strains and performing Asr mass analysis, thereby evaluating the distinguishability. Therefore, the inventor can evaluate the distinguishability of the two bacterial groups without spending time or effort. As described above, information on the distinction between bacterial groups of the order Enterobacteriaceae can be obtained using a simple method. Furthermore, compared to paragraph 11, this method has the advantage that it is possible to calculate the discrimination ratio between the first and second bacterial strains based on the positions of the Asr peaks, even for first and second bacterial strains that have the same number of Asr peaks.

[0218] (Item 3) The method for analyzing microorganisms described in item 2 further includes a step of calculating a second discrimination ratio, which is the ratio of the larger value to the sum of the first and second values, for each of the same peak groups.

[0219] According to the method for analyzing microorganisms described in the third aspect, it is possible to easily evaluate the possibility of distinguishing between the first bacterial group and the second bacterial group for each of the same peak groups.

[0220] (Item 4) The method for analyzing microorganisms described in items 2 or 3 further includes a step of determining that the set of molecular masses contained in each of the same peak groups is the molecular mass marker of the bacterial group corresponding to the larger numerical value.

[0221] According to the method for analyzing microorganisms described in paragraph 4, a set of molecular masses of a predetermined identical peak group can be determined as molecular mass markers of the putative bacterial group of the predetermined identical peak group. Therefore, the molecular mass markers can be used to identify microorganisms.

[0222] (Item 5) In the method for analyzing microorganisms described in any one of Items 2 to 4, the step of obtaining multiple full-length sequences includes the steps of determining a representative strain for each of the first and second bacterial groups, and obtaining multiple full-length sequences whose sequence similarity to the full-length sequence expressed in the representative strain is at or above a predetermined rank.

[0223] According to the method for analyzing microorganisms described in paragraph 5, it is possible to perform processing on a computer that simulates the actual work of culturing multiple strains belonging to each bacterial group, extracting amino acids, and analyzing them.

[0224] (Item 6) In the method for analyzing a microorganism according to any one of Items 2 to 5, the step of determining the set of amino acid sequences of a plurality of acid shock protein fragments includes the steps of removing the signal peptide from each of the plurality of full-length sequences, and, if the amino acid sequence from which the signal peptide has been removed contains a glutamine-lysine-alanine-glutamine sequence or a glutamine-asparagine-alanine-glutamine sequence from the N-terminus, cleaving the sequence at the C-terminus to create a set of amino acid sequences of acid shock protein fragments.

[0225] According to the method for analyzing microorganisms described in item 6, the process by which fragment sequences are generated in cells that actually contain Asr can be reproduced on a computer.

[0226] (Item 7) The method for analyzing microorganisms described in Item 4 further comprises the steps of performing mass analysis on the specific bacterial strain to obtain a mass spectrum, determining that the specific bacterial strain belongs to the first bacterial group if any peak in the mass spectrum corresponds to a molecular mass marker of the first bacterial group, and determining that the specific bacterial strain belongs to the second bacterial group if any peak in the mass spectrum corresponds to a molecular mass marker of the second bacterial group.

[0227] According to the method for analyzing microorganisms described in item 7, microorganisms can be easily identified using molecular mass markers of Asr fragments created by collecting pseudo-strains on a computer and performing mass analysis.

[0228] (Item 8) The method for analyzing microorganisms according to item 7 further comprises culturing the specific strain in a medium containing glucose, and the step of acquiring a mass spectrum includes acquiring a mass spectrum of the specific strain cultured in the medium containing glucose.

[0229] According to the method for analyzing microorganisms described in item 8, a specific strain containing Asr in its cells can be easily obtained by culturing the microorganism in a medium containing glucose.

[0230] (Item 9) The method for analyzing microorganisms described in item 8 further comprises the steps of culturing a specific bacterial strain in a glucose-free medium, comparing the mass spectrum obtained by mass spectrometry with the mass spectrum of a specific bacterial strain cultured in a glucose-containing medium, and identifying acid shock protein peaks in the mass spectrum of the specific bacterial strain cultured in a glucose-containing medium. The step of determining that the specific bacterial strain belongs to the first bacterial group includes the step of determining that the specific bacterial strain belongs to the first bacterial group if any of the acid shock protein peaks corresponds to a molecular mass marker of the first bacterial group. The step of determining that the specific bacterial strain belongs to the second bacterial group includes the step of determining that the specific bacterial strain belongs to the second bacterial group if any of the acid shock protein peaks corresponds to a molecular mass marker of the second bacterial group.

[0231] According to the method for analyzing microorganisms described in item 9, the Asr peak can be easily distinguished from the many peaks in the mass spectrum.

[0232] (Item 10) In the method for analyzing microorganisms according to any one of Items 7 to 9, the first bacterial group consists of Enterobacter asburyae. The second bacterial group consists of Enterobacter cloacae. The molecular mass markers of Enterobacter asburyae are (5149.8, 2069.4, 2099.5, 2388.8), (5121.8, 2069.4, 2099.5, 2388.8), (5121.8, 2099.5, 2388.8), (5121.8, 2076.4, 2099.5, 2388.8), (5121.8, 20 and (5135.8, 2069.4, 2099.5, 2388.8). The molecular mass markers of cloacae are (4152.7, 2611.1, 2099.5, 2388.8), (4122.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2099.5, 2021.4), (4180.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2085.4, 2388.8), (4168.7, 2611.1, 2099.5, 2388.8), (4152.7 , 2611.1, 2125.5, 2388.8), (4152.7, 2611.1, 2099.5, 2416.9), (4152.7, 2611.1, 2069.4, 2388.8), (4223.8, 2611.1, 2099.5, 2388.8), (4010.5, 2611.1, 2099.5, 2388.8) and (5163.8, 2069.4, 2099.5, 2388.8).

[0233] According to the method for analyzing microorganisms described in paragraph 10, E. asburiae and E. cloacae, which have been difficult to distinguish using conventional MALDI-based microorganism discrimination methods centered on ribosomal peaks, can be easily distinguished by using the molecular mass marker of the Asr fragment.

[0234] (Item 11) In the method for analyzing microorganisms described in Item 1, the mass spectrum pattern includes the number of peaks of acid shock proteins in the mass spectrum, and the step of calculating the discrimination ratio includes the step of calculating the discrimination ratio based on the number of peaks of acid shock proteins.

[0235] According to the method for analyzing microorganisms described in paragraph 11, information regarding the identification of Enterobacteriale bacteria can be obtained in a simple manner. In particular, since it is only necessary to count the number of truncated sequences in the full-length sequence, information regarding the identification of Enterobacteriale bacteria can be obtained even more easily than with the method for analyzing microorganisms described in paragraph 2.

[0236] (Item 12) In the method for analyzing microorganisms according to item 11, the step of calculating the discrimination ratio based on the number of peaks of acid shock proteins includes the steps of: determining the number N1 of full-length sequences of the first bacterial group and the number N2 of full-length sequences of the second bacterial group among the plurality of full-length sequences obtained in the step of obtaining a plurality of full-length sequences; counting the number of cleavage sequences in the full-length sequences of the first bacterial group and determining the mode M1 of the number of cleavage sequences; counting the number of cleavage sequences in the full-length sequences of the second bacterial group and determining the mode M2 ​​of the number of cleavage sequences; determining the number P2 of full-length sequences having the cleavage sequence of M1 in the full-length sequence of the first bacterial group; determining the number P2 of full-length sequences having the cleavage sequence of M1 in the full-length sequence of the second bacterial group; setting a third discrimination ratio, which is the discrimination ratio between the first bacterial group and the second bacterial group when the number of acid shock protein peaks is (M1+1), to 1 if P2 is 0; and setting a fourth discrimination ratio, which is the discrimination ratio between the first bacterial group and the second bacterial group when the number of acid shock protein peaks is (M2+1), to 1 if P1 is 0.

[0237] According to the method for analyzing microorganisms described in item 12, the discrimination rate between the first bacterial group and the second bacterial group according to the number of Asr peaks can be calculated by simple calculation.

[0238] (Item 13) The method for analyzing microorganisms described in Item 12 further includes a step of determining the fifth discrimination ratio when the number of acid shock protein peaks is (M1+1) as ((N2-P2) / N2) when P2 is not 0 and (N2-P2) is a predetermined multiple of P2 or more.

[0239] According to the method for analyzing microorganisms described in paragraph 13, it is also possible to calculate the discrimination ratio in the case where (1) in the full-length sequence of the first bacterial group, there is a full-length sequence corresponding to the most frequent value M2 of the number of truncated sequences in the second bacterial group, and (2) the number P1 of full-length sequences corresponding to M2 in the first bacterial group is extremely small compared to the total number N1 of full-length sequences in the first bacterial group.

[0240] (Item 14) In the method for analyzing microorganisms described in Item 11, the step of calculating the discrimination ratio based on the number of acid shock protein peaks includes the steps of: counting the number of cleavage sequences in the full-length sequences of the first bacterial group and determining the most frequent value M1 of the number of cleavage sequences; counting the number of cleavage sequences in the full-length sequences of the second bacterial group and determining the most frequent value M2 of the number of cleavage sequences; performing mass analysis on the specific bacterial strain to obtain a mass spectrum; and determining that the specific bacterial strain belongs to the first bacterial group if the number of acid shock protein peaks in the mass spectrum is (M1+1), or determining that the specific bacterial strain belongs to the second bacterial group if the number of acid shock protein peaks in the mass spectrum is (M2+1).

[0241] According to the method for analyzing microorganisms described in item 14, the first and second bacterial groups can be easily distinguished from each other using the number of Asr peaks.

[0242] (Item 15) In the method for analyzing microorganisms according to item 14, the first bacterial group consists of Escherichia coli, and the second bacterial group consists of Escherichia albertii.

[0243] According to the method for analyzing microorganisms described in paragraph 15, E. coli and E. albertii, which have been difficult to distinguish between in conventional MALDI identification methods and conventional biochemical property tests, can be easily distinguished based on the difference in the number of Asr peaks.

[0244] (Item 16) An analytical device according to one aspect includes a processor and a memory unit. The processor acquires multiple full-length sequences, which are the full-length amino acid sequences of multiple corresponding acid shock proteins, for each of a first bacterial group and a second bacterial group that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family. The processor identifies cleavage sequences that indicate sites at which the multiple full-length sequences are cleaved. The processor obtains mass spectral patterns of multiple acid shock protein fragments obtained by cleaving each of the multiple full-length sequences, and calculates a discrimination ratio between the first bacterial group and the second bacterial group based on the mass spectral patterns.

[0245] According to the analysis device described in paragraph 16, by calculating the discrimination ratio based on the mass spectrum pattern using the amino acid sequence information of Asr of a large number of strains of each bacterial group, it is possible to evaluate discrimination possibility without the need to actually collect a large number of bacterial strains and perform mass analysis of Asr. Therefore, the inventor can evaluate the discrimination possibility of two bacterial groups without spending time and effort. As described above, according to the discrimination possibility evaluation method of the embodiment, information on discrimination of bacterial groups of the order Enterobacteriaceae can be obtained in a simple manner.

[0246] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims.

[0247] 10 processor, 11 memory, 14 operation unit, 15 display, 70 genome database, 90 network, 100 analysis device, 101 controller, 1000 analysis system, 12 communication I / F, 13 input / output I / F.

Claims

1. A method for analyzing microorganisms, comprising the steps of: obtaining a plurality of full-length sequences, which are the amino acid sequences of a plurality of corresponding full-length acid shock proteins, for each of a first bacterial group and a second bacterial group that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family; identifying cleavage sequences that indicate the sites at which the plurality of full-length sequences are fragmented; determining the mass spectral patterns of a plurality of acid shock protein fragments obtained by fragmenting each of the plurality of full-length sequences, and calculating the discrimination ratio between the first bacterial group and the second bacterial group based on the mass spectral patterns.

2. The step of calculating the discrimination ratio includes the steps of: determining a set of amino acid sequences of a plurality of acid shock protein fragments obtained by fragmenting each of the plurality of full-length sequences; determining a set of theoretical molecular masses of the plurality of acid shock protein fragments corresponding to the set of amino acid sequences of the plurality of acid shock protein fragments; setting one or more identical peak groups by assigning one or more sets of molecular masses of the plurality of acid shock protein fragments that show the same peak in an actually measured mass spectrum to the same identical peak group; determining a first numerical value which is the number of strains belonging to the first bacterial group and a second numerical value which is the number of strains belonging to the second bacterial group among strains containing the full-length sequence corresponding to each of the identical peak groups; distinguishing the smaller numerical value and the larger numerical value of the first numerical value and the second numerical value for each of the identical peak groups; and calculating a third numerical value which is the sum of the larger numerical value of all the identical peak groups and a fourth numerical value which is the sum of the first numerical value and the second numerical value of all the identical peak groups. and calculating a first discrimination ratio, which is a ratio of the third numerical value to the fourth numerical value.

3. The method for analyzing microorganisms according to claim 2, further comprising a step of calculating a second discrimination ratio, which is the ratio of the larger numerical value to the sum of the first numerical value and the second numerical value, for each of the identical peak groups.

4. The method for analyzing microorganisms according to claim 2 or 3, further comprising a step of determining the set of molecular masses contained in each of the identical peak groups as the molecular mass marker of the bacterial group corresponding to the larger numerical value.

5. A method for analyzing microorganisms as described in claim 2 or 3, wherein the step of obtaining the multiple full-length sequences includes the steps of: determining a representative strain for each of the first group of bacteria and the second group of bacteria; and obtaining the multiple full-length sequences whose sequence similarity to the full-length sequences expressed in the representative strains is at least a predetermined rank.

6. The method for analyzing microorganisms according to claim 2 or 3, wherein the step of determining the set of amino acid sequences of the plurality of acid shock protein fragments comprises the steps of: removing a signal peptide from each of the plurality of full-length sequences; and, if the amino acid sequence from which the signal peptide has been removed contains a glutamine-lysine-alanine-glutamine sequence or a glutamine-asparagine-alanine-glutamine sequence from the N-terminus, cleaving the sequence at the C-terminus to create the set of amino acid sequences of the acid shock protein fragments.

7. A method for analyzing microorganisms as described in claim 4, further comprising the steps of: performing mass analysis on a specific bacterial strain to obtain a mass spectrum; determining that the specific bacterial strain belongs to the first bacterial group if any peak in the mass spectrum corresponds to the molecular mass marker of the first bacterial group; and determining that the specific bacterial strain belongs to the second bacterial group if any peak in the mass spectrum corresponds to the molecular mass marker of the second bacterial group.

8. The method for analyzing microorganisms according to claim 7, further comprising a step of culturing the specific strain in a medium containing glucose, and the step of obtaining a mass spectrum includes a step of obtaining a mass spectrum of the specific strain cultured in a medium containing glucose.

9. The method for analyzing microorganisms according to claim 8, further comprising the steps of culturing the specific bacterial strain in a glucose-free medium, comparing the mass spectrum obtained by mass spectrometry with the mass spectrum of the specific bacterial strain cultured in the glucose-containing medium, and identifying acid shock protein peaks in the mass spectrum of the specific bacterial strain cultured in the glucose-containing medium, wherein the step of determining that the specific bacterial strain belongs to the first bacterial group comprises the step of determining that the specific bacterial strain belongs to the first bacterial group if any of the acid shock protein peaks corresponds to the molecular mass marker of the first bacterial group, and the step of determining that the specific bacterial strain belongs to the second bacterial group comprises the step of determining that the specific bacterial strain belongs to the second bacterial group if any of the acid shock protein peaks corresponds to the molecular mass marker of the second bacterial group.

10. The first bacterial group consists of Enterobacter asburyae, the second bacterial group consists of Enterobacter cloacae, and the molecular mass markers of the Enterobacter asburyae are (5149.8, 2069.4, 2099.5, 2388.8), (5121.8, 2069.4, 2099.5, 2388.8), (5121.8, 2099.5, 2388.8), (5121.8, 2076.4, 2099.5, 2388.8), (5121.8, 2 and (5135.8, 2069.4, 2099.5, 2388.8), (5121.8, 2069.4, 2364.8, 2388.8), (5121.8, 2069.4, 2099.5, 2113.5, 2388.8), and (5135.8, 2069.4, 2099.5, 2388.8), The molecular mass markers of cloacae are (4152.7, 2611.1, 2099.5, 2388.8), (4122.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2099.5, 2021.4), (4180.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2085.4, 2388.8), (4168.7, 2611.1, 2099.5, 2388.8), (4152.7, 2611.1, 2125.5, 2388.8), (4152.7, 2611.1, 2099.5, 2416.9), (4152.7, 2611.1, 2069.4, 2388.8), (4223.8, 2611.1, 2099.5, 2388.8), (4010.5, 2611.1, 2099.5, 2388.8) and (5163.8, 2069.4, 2099.5, 2388.8). The method for analyzing microorganisms according to claim 7, wherein the set of molecular masses is at least one selected from the group consisting of: (4152.7, 2611.1, 2099.5, 2416.9), (4152.7, 2611.1, 2069.4, 2388.8), (4223.8, 2611.1, 2099.5, 2388.8), (4010.5, 2611.1, 2099.5, 2388.8) and (5163.8, 2069.4, 2099.5, 2388.8).

11. The method for analyzing microorganisms according to claim 1, wherein the mass spectrum pattern includes the number of peaks of acid shock proteins in the mass spectrum, and the step of calculating the discrimination ratio includes the step of calculating the discrimination ratio based on the number of peaks of the acid shock proteins.

12. The step of calculating the discrimination ratio based on the number of peaks of the acid shock protein comprises the steps of: determining the number N1 of the full-length sequences of the first bacterial group and the number N2 of the full-length sequences of the second bacterial group from the plurality of full-length sequences obtained in the step of obtaining a plurality of full-length sequences; tabulating the numbers of truncated sequences in the full-length sequences of the first bacterial group and determining the mode M1 of the number of truncated sequences; tabulating the numbers of truncated sequences in the full-length sequences of the second bacterial group and determining the mode M2 ​​of the number of truncated sequences; determining the number P1 of full-length sequences having the M2 truncated sequence in the full-length sequences of the first bacterial group; and determining the number P2 of full-length sequences having the M1 truncated sequence in the full-length sequences of the second bacterial group; and if P2 is 0, setting a third discrimination ratio of 1, which is the discrimination ratio between the first bacterial group and the second bacterial group when the number of peaks of the acid shock protein is (M1 + 1). and when P1 is 0, setting a fourth discrimination ratio, which is a discrimination ratio between the first bacterial group and the second bacterial group when the number of peaks of the acid shock protein is (M2 + 1), to 1.

13. The method for analyzing microorganisms according to claim 12, further comprising the step of setting a fifth discrimination ratio when the number of acid shock protein peaks is (M1+1) to ((N2-P2) / N2) when P2 is not 0 and (N2-P2) is a predetermined multiple of P2 or more.

14. The method for analyzing microorganisms described in claim 11, wherein the step of calculating the identification rate based on the number of acid shock protein peaks comprises: a step of counting the number of cleavage sequences in the full-length sequence of the first bacterial group and determining the most frequent value M1 of the number of cleavage sequences; a step of counting the number of cleavage sequences in the full-length sequence of the second bacterial group and determining the most frequent value M2 of the number of cleavage sequences; a step of performing mass analysis on the specific bacterial strain to obtain a mass spectrum; a step of determining that the specific bacterial strain belongs to the first bacterial group if the number of acid shock protein peaks in the mass spectrum is (M1 + 1), or a step of determining that the specific bacterial strain belongs to the second bacterial group if the number of acid shock protein peaks in the mass spectrum is (M2 + 1).

15. The method for analyzing microorganisms according to claim 14, wherein the first bacterial group consists of Escherichia coli, and the second bacterial group consists of Escherichia albertii.

16. An analytical device comprising a processor and a memory unit, wherein the processor obtains a plurality of full-length sequences, which are the full-length amino acid sequences of a plurality of corresponding acid shock proteins, for each of a first bacterial group and a second bacterial group that belong to the order Enterobacteriaceae and differ from each other in at least one classification below family, identifies cleavage sequences that indicate sites at which the plurality of full-length sequences are cleaved, determines mass spectral patterns of a plurality of acid shock protein fragments obtained by cleaving each of the plurality of full-length sequences, and calculates a discrimination ratio between the first bacterial group and the second bacterial group based on the mass spectral patterns.

Citation Information

Patent Citations

  • AI big data real-time processing and analysis method

    CN118245680A

  • Analysis method, microbe identifying method and testing method

    WO2020202861A1

  • Microorganism classifying method, control device, and analysis device

    WO2024005120A1

  • Method for preparing analysis of microorganisms, and method for analyzing microorganisms

    WO2024005122A1