Methods for creating training data, machine learning models, methods for identifying microorganisms, analytical equipment, programs

By grouping microorganisms with similar genetic information and using amino acid sequence analysis, the method addresses misclassification issues in machine learning models, improving their accuracy in identifying microorganisms.

JP7865142B2Active Publication Date: 2026-05-26SHIMADZU SEISAKUSHO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SHIMADZU SEISAKUSHO LTD
Filing Date
2022-08-12
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing machine learning models struggle to accurately distinguish between microorganisms with similar genetic information using mass spectrometry, leading to misclassification, particularly among closely related species.

Method used

Create training data by grouping microorganisms with similar genetic information and assigning a common label to prevent discrimination within these groups, using methods such as amino acid sequence analysis to estimate closely related species.

Benefits of technology

Prevents misclassification among microbial species with similar genetic information by ensuring the machine learning model does not differentiate within these groups, enhancing the model's accuracy and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007865142000001
    Figure 0007865142000001
  • Figure 0007865142000002
    Figure 0007865142000002
  • Figure 0007865142000003
    Figure 0007865142000003
Patent Text Reader

Abstract

To provide a machine learning model that is not likely to erroneously discriminate the types of micro-organisms similar in genetic information to each other, and to generate learning data for creating the machine learning model.SOLUTION: Provided is a method for generating a learning data for a machine learning model for discriminating the type of a micro-organism by using mass spectrometry, the method including: a step S1 of acquiring genetic information for every type of the micro-organism; a step S2 of estimating the types of the micro-organisms similar in the genetic information to each other, and creating a group including the types of the micro-organisms estimated be similar in the genetic information to each other; and a step S3 of including information on the group in the learning data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for creating learning data, a machine learning model, a method for discriminating microorganisms, an analysis apparatus, and a program.

Background Art

[0002] Conventionally, a method for discriminating microorganisms based on a mass spectrum obtained by analyzing microorganisms by MALDI-MS (Matrix Assisted Laser Desorption / Ionization-Mass Spectrometry) has been known.

[0003] Particularly, in recent years, as a method for classifying microorganisms based on the mass spectrum, a method using a machine learning model has attracted attention. US Patent Application Publication No. 2020 / 0118805 (Patent Document 1) discloses a method for discriminating microorganisms using a machine learning model.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Non-Patent Document 2

[0006] However, when identifying microorganisms using mass spectrometry, it can be difficult to distinguish between species if their mass spectra are similar, even if they belong to different microbial classifications. For example, closely related species that are of the same genus but different species may be difficult to distinguish even by mass spectrum alone if their genetic information is similar. If such closely related species are included in the training data along with species that do not have closely related species, the machine learning model trained using that data may misclassify them among the closely related species. In other words, if the machine learning model receives the mass spectrum of a given species that has closely related species as input, it may output a classification result that identifies it as a closely related species rather than the given species.

[0007] This disclosure is made in light of these circumstances, and its purpose is to create a machine learning model that has a low probability of misclassifying microbial species with similar genetic information, and training data for creating such a machine learning model, in the field of machine learning that uses mass spectrometry to identify microorganisms. [Means for solving the problem]

[0008] A first aspect of this disclosure is a method for creating training data for a machine learning model that identifies types of microorganisms using mass spectrometry, comprising the steps of: obtaining genetic information for each type of microorganism; estimating types of microorganisms with similar genetic information; creating groups that include the types of microorganisms estimated to have similar genetic information; and including information about the groups in the training data.

[0009] A second aspect of this disclosure is an analysis device for creating training data for a machine learning model that identifies microorganisms using mass spectrometry. The analysis device comprises a memory and a processor. The memory stores the genetic information of microorganisms. The processor performs a method for creating training data using the genetic information stored in the memory. The processor obtains the genetic information of each microorganism, estimates microorganisms with similar genetic information, creates groups containing the microorganisms estimated to have similar genetic information, and includes information about the groups in the training data. [Effects of the Invention]

[0010] According to the method for creating training data described herein, by including information on groups containing microbial species that are estimated to have similar genetic information in the training data, it is possible to configure a machine learning model using said training data so that it does not distinguish between microbial species within a group (for example, closely related species). Therefore, misclassification among microbial species with similar genetic information can be prevented. In other words, a machine learning model with a low probability of misclassification among microbial species with similar genetic information, and training data for creating said machine learning model, can be provided. [Brief explanation of the drawing]

[0011] [Figure 1] This figure shows the configuration of the analysis device according to the embodiment. [Figure 2] This diagram illustrates the relationship between training data and a machine learning model according to the embodiment. [Figure 3] This is a diagram to explain neural networks. [Figure 4] This is a flowchart showing the process of creating training data according to the embodiment. [Figure 5] This flowchart shows an example of the process of adding group-related information to training data. [Figure 6] This is a flowchart showing the process for obtaining amino acid sequences. [Figure 7]It is a flowchart showing the estimation process of related species based on an amino acid sequence. [Figure 8] It is a diagram for explaining an array list and a protein list. [Figure 9] It is a flowchart showing an example of a calculation process for the similarity of an amino acid sequence. [Figure 10] It is a diagram for explaining a method of calculating similarity. [Figure 11] It is a flowchart showing an example of an estimation process of related species based on similarity. [Figure 12] It is a flowchart showing an example of an acquisition process of an amino acid sequence. [Figure 13] It is a diagram for explaining an estimation result of related species by a machine learning model according to an embodiment.

Best Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following, the same or corresponding parts in the drawings are denoted by the same reference numerals, and the description thereof will not be repeated in principle.

[0013] [1. Configuration of Analysis Device] FIG. 1 is a diagram showing the configuration of an analysis device 100 according to an embodiment. The analysis device 100 acquires a mass spectrum corresponding to the type of microorganism and creates learning data for a machine learning model that discriminates the type of microorganism using the mass spectrum. Referring to FIG. 1, the analysis device 100 includes a controller 101, a display 15, and an operation unit 14. The display 15 and the operation unit 14 are connected to the controller 101. The operation unit 14 is typically composed of a touch panel, a keyboard, a mouse, etc. The operation unit 14 receives a user's operation input to the processor 10. The display 15 is composed of, for example, a liquid crystal panel capable of displaying an image. The display 15 displays an image related to receiving a user's operation input and displays the result of processing by the processor 10.

[0014] The controller 101 mainly includes a processor 10, a memory 11, a communication interface (I / F) 12, and an input / output I / F 13. These components are communicably connected to each other via a bus.

[0015] The processor 10 is typically an arithmetic processing unit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). The processor 10 controls the operation of the analysis device 100 by reading and executing the programs stored in the memory 11. The programs include programs that, when executed by a computer, cause the computer to implement a method for creating learning data according to an embodiment.

[0016] The memory 11 is realized by a storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), and an HDD (Hard Disk Drive), for example. The ROM can store the programs executed by the processor 10. The RAM can temporarily store the data used during the execution of the programs in the processor 10 and can function as a temporary data memory used as a working area. The HDD is a non-volatile storage device. In addition to or instead of the HDD, a semiconductor storage device such as a flash memory may be employed. Note that the above programs and / or data may be stored in an external storage device accessible by the processor 10.

[0017] The communication I / F 12 is a communication interface for exchanging various data with an external device and is realized by an adapter or a connector, etc. Note that the communication method may be a wireless communication method such as a wireless LAN (Local Area Network) or a wired communication method using USB (Universal Serial Bus), etc.

[0018] The input / output interface 13 is an interface for exchanging various types of data between the processor 10 and external devices connected to the input / output interface 13. The external devices include the control unit 14 and the display 15. A mass spectrometer (MS) 16 may also be connected to the input / output interface 13.

[0019] MS16 is a device for performing mass spectrometry of components contained in microbial samples, such as MALDI-TOF MS (Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry). In MS16, ions generated by laser irradiation are extracted into a flight tube, allowed to fly, and then separated and detected according to their flight time. The flight time correlates with the mass-to-charge ratio (m / z) of the components. As a result, a mass spectrum is obtained with m / z on the x-axis and the detected ion intensity on the y-axis.

[0020] In this specification, MS16 performs mass spectrometry of proteins in a sample. Therefore, in the mass spectrum, peaks are detected according to the mass-to-charge ratio (m / z) of proteins in the sample. Thus, by referring to the pattern of the mass spectrum, or more specifically, the pattern of the peaks, it is possible to recognize the proteins contained in the sample.

[0021] Different types of microorganisms contain different proteins, and therefore their mass spectral patterns will also differ. In other words, the mass spectral pattern generally reflects the type of microorganism. In this specification, the "type" of a microorganism includes, for example, at least one of the "genotype, strain, or rank of a phylogenetic group such as subspecies, species, genus, or family."

[0022] MS16 performs mass spectrometry on a sample containing microorganisms and then transmits the sample's mass spectrum to the analysis device 100. The processor 10 creates training data for a machine learning model to identify microorganisms based on the mass spectrum. The processor 10 then creates the machine learning model using this training data. The processor 10 then uses the machine learning model to identify microorganisms based on the mass spectrum.

[0023] Furthermore, the analysis device 100 does not need to be composed of a single computer; it may be composed of multiple computers.

[0024] [2. Conventional methods for identifying microorganisms using mass spectrometry] As described above, the pattern of the mass spectrum reflects the type of microorganism. A method is known that utilizes this property to identify microorganisms based on the pattern of the mass spectrum obtained from MALDI-MS analysis.

[0025] In recent years, attempts have been made to use machine learning techniques to identify microorganisms. U.S. Patent Application Publication No. 2020 / 0118805 (Patent Document 1) discloses a configuration for identifying microorganisms using a machine learning model. Specifically, for example, the machine learning model is trained by providing it with mass spectra associated with known types of microorganisms as training data. As a result, when the mass spectrum of an unknown microorganism is input to the machine learning model, it can output the type (classification) of that microorganism.

[0026] However, in microbial identification methods based on such mass spectral patterns, it can be difficult to distinguish between different types of microorganisms if their mass spectra are similar. An example of this is explained below.

[0027] Generally, different species of microorganisms in the same genus have different genetic information, resulting in different mass spectra, making species identification possible through mass spectrometry. However, even among different species of microorganisms in the same genus, a very small number of species have similar genetic information. These different species of the same genus with similar genetic information are generally referred to as "closely related species." In this specification, a group of different species of the same genus with similar genetic information will be referred to as a "closely related species group." Furthermore, in this specification, the relationship between different species of the same genus that are closely related (i.e., have similar genetic information) will also be referred to as "being positioned as closely related species." In other words, a closely related species group consists of multiple species that are positioned as closely related species. Among different species of the same genus that are positioned as closely related species with similar genetic information, their mass spectra are also similar, making identification difficult. If species belonging to a particular group of closely related species are included in the training data as independent species, just like species that have no other closely related species (i.e., species that do not belong to either the group of closely related species or any other group of closely related species), there is a concern that the machine learning model trained on this data may misclassify some of the closely related species within that group. For example, if the mass spectrum of a specific species within a group of closely related species is input to the machine learning model, it may output another species within that group as the classification result.

[0028] If a machine learning model misclassifies closely related species, the following problems may arise: Some closely related species are already known, but if the user lacks knowledge of them, it may be difficult for the user to suspect a misclassification even if an incorrect classification result is output for a microorganism belonging to that group of closely related species. In this case, the user may believe the incorrect classification result. Furthermore, if an incorrect classification result is returned for an unknown closely related species, it is difficult for the user to know if that classification result is incorrect. As a result, there is a risk that the user will have difficulty determining whether the machine learning model's classification result is correct or not. In such a situation, even if a correct classification result is output for a species that has no closely related species, it may be difficult to determine whether the classification result is truly correct. In other words, the reliability of the machine learning model's classification results may decrease.

[0029] Therefore, in the method for creating training data according to this embodiment, types of microorganisms with similar genetic information (e.g., closely related species) are estimated in advance, groups containing these types of microorganisms with similar genetic information (e.g., closely related species groups) are created, and the types of microorganisms included in these groups are assigned a common label and used as training data. In a machine learning model trained with such training data, discrimination (classification) is performed without distinguishing between the microorganisms included in the group. Therefore, misclassification among microorganisms included in the group can be prevented. In other words, a machine learning model with a low possibility of misclassification among types of microorganisms with similar genetic information, and training data for creating such a machine learning model can be provided.

[0030] [3. Relationship between training data and machine learning model according to the embodiment] Figure 2 is a diagram illustrating the relationship between training data and a machine learning model according to the embodiment. The machine learning model, upon input of a mass spectrum of a microorganism, outputs the type of microorganism. This allows the machine learning model to identify microorganisms. In this specification, "identifying microorganisms" refers to taxonomically identifying the type of microorganism.

[0031] Such machine learning models include, for example, neural networks (Figure 3). In a neural network, each node is weighted to produce an appropriate output for a given input. This weighting is determined through learning from training data.

[0032] More specifically, in a neural network, when multiple inputs are given to the input layer, each input is multiplied by its weights, and the result of this multiplication is sent to the next layer. In the next layer, this result is multiplied by its weights, and the result of this multiplication is sent to yet another layer. Finally, the output is obtained from the output layer.

[0033] Machine learning models are trained as follows: First, a large number of mass spectra of known microorganisms are obtained. Next, the machine learning model is trained using sets of mass spectra (the input to the machine learning model) and the expected output (ground truth) of the microorganism (the species) as training data. This allows the machine learning model to learn to output the expected output for a given input. In the example in Figure 3, the weighting of each node is appropriately adjusted.

[0034] [4. Method for creating training data according to the embodiment] The following explains how to create training data for a machine learning model that identifies the type of microorganism using mass spectrometry, using Figures 4 and 5 as an example.

[0035] Figure 4 is a flowchart illustrating the process of creating training data according to an embodiment. Each step shown in Figure 4 is performed by the processor 10.

[0036] In step 1 (hereinafter, step 1 will be abbreviated as "S"), the processor 10 obtains genetic information for each microbial species.

[0037] In S2, the processor 10 estimates multiple microbial "species" (i.e., multiple microbial species that are considered closely related to each other) that have similar genetic information, and creates a closely related species group that includes these closely related species. In this specification, the closely related species group corresponds to one embodiment of the "group containing closely related species".

[0038] In S3, processor 10 includes information about closely related species groups in the training data and terminates processing.

[0039] Figure 5 is a flowchart illustrating in more detail the process of adding information about closely related species groups to the training data in S3.

[0040] S31 in Figure 5 corresponds to an example of S3 in Figure 4 and is performed after S2 in Figure 4. In S31, the processor 10 assigns the same label to closely related species included in the closely related species group and terminates processing. More specifically, it assigns the same label to microbial species (closely related species) that were estimated to belong to the same closely related species group in S2 of Figure 4.

[0041] In the machine learning model using the training data created by the processes shown in Figures 4 and 5, discrimination is not performed within the group of closely related species. Therefore, misclassification of closely related species can be prevented.

[0042] As mentioned above, Figures 4 onward illustrate a method for creating training data for a machine learning model that identifies microbial species. However, this method can also be adapted, as needed, for creating training data for machine learning models that identify microbial species of other ranks (e.g., genotype, strain, subspecies, genus, or family). In this case, as in Figure 4, the processor 10 estimates microbial species with similar genetic information and creates a group containing the microbial species estimated to have similar genetic information. The processor 10 then includes information about this group in the training data. For example, the processor 10 assigns the same label to the microbial species included in the group. As a result, the machine learning model using this training data does not perform discrimination among the microbial species included in the group. Such a machine learning model can prevent misclassification between microbial species with similar genetic information. In this specification, closely related species correspond to one example of "microbial species estimated to have similar genetic information."

[0043] (4-1. Method for estimating closely related species based on amino acid sequences) Figures 6 to 8 illustrate a more specific example of the method for creating training data according to this embodiment, as described above.

[0044] Figure 6 is a flowchart showing the process for obtaining the amino acid sequence. S11 in Figure 6 corresponds to an example of S1 in Figure 4.

[0045] In S11, the processor 10 obtains one or more amino acid sequences for each protein corresponding to various microorganisms.

[0046] Figure 7 is a flowchart showing the process for estimating closely related species based on amino acid sequences. Each step in Figure 7 corresponds to an example S2 in Figure 4 and is performed after S11 in Figure 6.

[0047] In S21, the processor 10 creates a sequence list for each type of protein, containing one or more amino acid sequences as elements.

[0048] In S22, the processor 10 creates a protein list for each type, with the sequence list as its element.

[0049] In S23, processor 10 calculates the similarity of protein lists between two different species.

[0050] In S24, processor 10 estimates closely related species based on similarity. In S25, processor 10 creates a related species group including closely related species and proceeds to S3.

[0051] The process shown in Figures 6 and 7 allows for the estimation of closely related species using information from the amino acid sequence of proteins. The amino acid sequences of microbial proteins, along with other genetic information such as DNA (Deoxyribonucleic Acid) sequences and RNA (Ribonucleic Acid) sequences, are widely available in public databases, making efficient collection possible.

[0052] Generally, species with similar genetic information tend to have similar mass spectra, making identification difficult. However, considering that RNA is transcribed from DNA, amino acids are translated from RNA to produce proteins, and that the m / z of proteins measured by mass spectrometry directly reflects the amino acid sequence, the similarity of amino acid sequences is considered to be the most correlated with the similarity of mass spectra.

[0053] Therefore, as shown in the processing in Figures 6 and 7, by obtaining amino acid sequences instead of other genetic information, estimating closely related species based on their similarity, and creating a group of closely related species that includes those species, it is possible to efficiently create a group that includes species that are difficult to distinguish by mass spectrometry.

[0054] However, the method for obtaining genetic information according to this embodiment is not limited to the above examples, and it is sufficient to obtain genetic information that reflects the classification of microorganisms. For example, the processor 10 may obtain DNA sequences and RNA sequences from a public database and calculate amino acid sequences in the processor 10. Also, for example, the genetic information may include information obtained from external storage devices other than public databases, or it may include genetic information obtained by the user through experiments.

[0055] Generally, amino acid sequence databases contain hierarchical information such as the classification of the microorganism (e.g., species), the name of the protein expressed in that microorganism, and the amino acid sequence corresponding to that protein name. In other words, the information for each amino acid sequence is stored linked to the classification of the corresponding microorganism and the corresponding protein name.

[0056] Furthermore, there are often multiple amino acid sequences corresponding to a given protein of a given species. This is because amino acid sequences corresponding to classifications below the species level (e.g., subspecies, strains) are often registered in databases. Since there are usually multiple subspecies or strains for a single species, there are usually multiple amino acid sequences corresponding to a single species. Therefore, in order to estimate closely related species using amino acid sequences, it is necessary to appropriately use one or more amino acid sequences corresponding to each of these proteins and determine their similarity.

[0057] Figure 8 illustrates the sequence list and protein list created in steps S21-S22 of Figure 7. Figure 8 shows the protein lists corresponding to species X and species Y. The protein list includes sequence lists for proteins a-c. Each sequence list contains one or more amino acid sequences.

[0058] As explained in Figures 6-8, by processing amino acid sequences, the similarity of amino acid sequences between two different species can be determined. Based on this similarity, closely related species can then be estimated.

[0059] (4-2. An example of a method for calculating similarity) Next, we will explain specific examples of methods for calculating amino acid sequence similarity using Figures 9 to 10.

[0060] Figure 9 is a flowchart showing an example of the process for calculating amino acid sequence similarity. The series of steps in Figure 9 corresponds to the example in S23 of Figure 7 and is performed after S22 of Figure 7.

[0061] In Figure 9, the similarity between species 1 and species 2, which are included in the "two different species" described in S23 of Figure 7, is calculated.

[0062] In S231, the processor 10 determines whether the specific differences between the amino acid sequences of each protein in the first protein list and each amino acid sequence in the second protein list satisfy the specific conditions, and calculates the specific number of proteins for which the specific differences may satisfy the specific conditions.

[0063] In S232, the processor 10 calculates the ratio of a specific number of proteins to the total number of proteins included in the first protein list as the similarity of the second type to the first type, and proceeds to S24.

[0064] Figure 10 is a diagram that provides a more detailed explanation of the similarity calculation method shown in Figure 9. Referring to Figure 10, the method for calculating the similarity between species Y and species X will be explained.

[0065] First, processor 10 compares the amino acid sequence aX1 of protein a from species X with the individual amino acid sequences contained in species Y. Then, it determines whether the specific differences in these amino acid sequences satisfy specific conditions.

[0066] Specific differences are numerical values ​​calculated based on, for example, the edit distance of amino acid sequences or the Average Amino-acid Identity method. The edit distance of amino acid sequences is a numerical value that indicates the difference between two amino acid sequences. More specifically, the edit distance is calculated based on the number of amino acid substitutions, deletions, and / or insertions required to make one amino acid sequence the same as the other. The Average Amino-acid Identity method is a method for easily comparing the similarity of amino acid sequences on a computer. More specifically, the Average Amino-acid Identity method is a method that fragments amino acid sequences on a computer and calculates the similarity of the entire amino acid sequence based on the similarity of those fragments.

[0067] The specific condition is determined based on a specific difference, and if that condition is met, it is considered that the two amino acid sequences can be determined to be similar. If the specific difference is the edit distance, the specific condition is, for example, that it is less than a predetermined natural number less than or equal to 10.

[0068] In the case of Figure 10, amino acid sequence aY2 was found to be an amino acid sequence contained in species Y that satisfies specific conditions, as it differs in specific amino acid sequence aX1 from protein a of species X. Therefore, it can be determined that an amino acid sequence similar to that of protein a of species X is contained in species Y. In other words, it can be determined that a similar protein to protein a of species X is contained in species Y.

[0069] Next, processor 10 compares the amino acid sequence bX1 of protein b of species X with each amino acid sequence contained in species Y. It then determines whether the specific difference between these amino acid sequences satisfies a specific condition. If no amino acid sequence of species Y is found that satisfies the specific difference between the amino acid sequence bX1 of protein b of species X and the specific condition, processor 10 searches for amino acid sequences of species Y that satisfy the specific difference for other amino acid sequences bX2 of protein b of species X. If no amino acid sequences of species Y that satisfy the specific difference for all amino acid sequences of protein b of species X are found, it can be determined that there are no amino acid sequences in species Y that are similar to the amino acid sequence of protein b of species X. In other words, it can be determined that there are no proteins in species Y that are similar to protein b of species X.

[0070] Processor 10 determines whether a similar amino acid sequence exists in species Y for all proteins of species X. In the example in Figure 10, no amino acid sequences similar to the amino acid combinations of proteins b and c were found in species Y.

[0071] Next, the processor 10 calculates the number of proteins in the protein list of species X that contain amino acid sequences that may satisfy specific conditions due to specific differences from each amino acid sequence of species Y. In other words, the number of proteins is the number of proteins in species Y that have amino acid sequences similar to that amino acid sequence. That is, the number of proteins in species X that contain proteins similar to those in species Y.

[0072] In the example in Figure 10, the number of specific proteins is 1. Here, since the total number of proteins in species X is 3, the similarity of species Y to species X is (number of specific proteins / total number) = (1 / 3).

[0073] The method described in Figures 9 and 10 allows for the accurate calculation of the differences between the amino acid sequences contained in two different species. Furthermore, based on these differences, the degree of similarity between the amino acid sequences of the two different species can be quantified.

[0074] (4-3. An example of a method for estimating closely related species based on similarity) Next, we will explain in more detail, using Figure 11, the method for estimating closely related species based on the similarity calculated above.

[0075] Figure 11 is a flowchart illustrating a specific example of the process for estimating closely related species based on similarity. Step S241 in Figure 11 corresponds to an example of step S24 in Figure 7 and is performed after step S23 in Figure 7.

[0076] In S241, processor 10 estimates closely related species by using a statistical method that employs at least one of the following for similarity: threshold, mean, standard deviation, and outlier test, and then proceeds to S25.

[0077] According to the process shown in Figure 11, statistically related species can be estimated by statistically processing the similarity. This allows for the estimation of related species in a simple manner, without involving the user's subjectivity.

[0078] (4-4. An example of how to obtain an amino acid sequence) The estimation of closely related species based on the similarity of amino acid sequences of proteins, as described above, may be performed using all proteins, or it may be limited to specific proteins.

[0079] Figure 12 is a flowchart showing a specific example of the amino acid sequence acquisition process. S111 in Figure 12 corresponds to an example of S11 in Figure 7.

[0080] In S111, the processor 10 retrieves only amino acid sequences from a database of amino acid sequences containing amino acid sequences corresponding to protein names, specifically those containing a particular string in the protein name. This particular string includes, for example, strings indicating housekeeping proteins and / or DNA-binding proteins. Housekeeping proteins refer to proteins necessary for maintaining basic cellular functions, such as ribosomal proteins. Strings indicating housekeeping proteins include, for example, at least one of "60kDa chaperonin," "Citrate synthase," "CTP synthase," and "RNA polymerase sigma factor RpoD." Strings indicating DNA-binding proteins include, for example, "DNA-binding."

[0081] According to the processing shown in Figure 12, related species can be estimated based on the similarity of specific proteins, rather than all of them. In particular, related species can be estimated based on the similarity of housekeeping proteins and / or DNA-binding proteins. That is, related species can be estimated based on the similarity of proteins whose function is important and whose amino acid sequences are considered to be conserved.

[0082] [5. Estimation results for closely related species] Next, we will explain the results of estimating closely related species using the method described above, with reference to Figure 13.

[0083] Figure 13 illustrates the estimation results of closely related species in the method for creating training data according to the embodiment. Figure 13 is a table arranged vertically, showing the similarity of 18 species of the genus Bacillus to a given species A and other species B. The number of proteins used for comparison is indicated in parentheses following the species name. The method for calculating similarity uses the specific condition that the edit distance is 3 or less. Furthermore, the study targets proteins whose names contain at least one of the following strings: "60kDa chaperonin", "Citrate synthase", "CTP synthase", "RNA polymerase sigma factor RpoD", or "DNA-binding".

[0084] Furthermore, as a method for estimating closely related species based on similarity, the Smirnov-Grubbs test, a type of outlier test, was performed on the similarity between a given species A and another species B, and the points where significance was found at the 5% level are indicated by diagonal lines.

[0085] As a result, the first group "amiloliquefaciens, atrophaeus, licheniformis, mojavensis, subtilis," the second group "megateirum, simplex," and the third group "mycoides, thuringiensis, weihenstephaensis," corresponding to the areas indicated by the diagonal lines, showed a high degree of similarity in the protein lists of the species within each group, and could be estimated to be closely related species. Of these, the first and third groups have been reported to be difficult to distinguish in practice (see Non-Patent Documents 1 and 2). In other words, the method for estimating closely related species in the method for creating training data according to this embodiment was able to actually estimate closely related species.

[0086] [Aspect] Those skilled in the art will understand that the above-described exemplary embodiments are specific examples of the following embodiments.

[0087] (Article 1) A method for creating training data according to one embodiment is a method for creating training data for a machine learning model that identifies types of microorganisms using mass spectrometry, comprising the steps of: obtaining genetic information for each type of microorganism; estimating types of microorganisms with similar genetic information; creating groups that include the types of microorganisms estimated to have similar genetic information; and including information about the groups in the training data.

[0088] The method for creating training data described in paragraph 1 can prevent misclassification among microbial species with similar genetic information. In other words, it is possible to provide a machine learning model with a low probability of misclassification among microbial species with similar genetic information, and training data for creating said machine learning model.

[0089] (Section 2) In the method for creating training data described in paragraph 2, the type of microorganism is a species of microorganism, and a microorganism includes one or more species, and the types of microorganisms with similar genetic information are closely related species.

[0090] The method for creating training data described in Section 2 can prevent misclassification among closely related species.

[0091] (Section 3) In the method for creating training data described in paragraph 2, the step of obtaining genetic information includes the step of obtaining one or more amino acid sequences for each protein corresponding to each species of microorganism. The step of creating groups includes, for each species, the step of creating a sequence list containing one or more amino acid sequences as elements for each protein; for each species, the step of creating a protein list using the sequence list as elements; the step of calculating the similarity of protein lists between two different species; the step of estimating closely related species based on the similarity; and the step of creating groups containing closely related species.

[0092] The method for creating training data described in Section 3 allows for the determination of the similarity of amino acid sequences between two different species. Based on this similarity, closely related species can then be estimated.

[0093] (Section 4) In the method for creating training data as described in paragraph 3, the two different species include species 1 and species 2, and the step of calculating similarity includes determining whether the specific differences between the amino acid sequences of each protein in the protein list of species 1 and each amino acid sequence in the protein list of species 2 satisfy specific conditions, calculating the specific number of proteins for which the specific differences may satisfy specific conditions, and calculating the ratio of the specific number to the total number of proteins in the protein list of species 1 as the similarity of species 2 from the perspective of species 1.

[0094] The method for creating training data described in Section 4 allows for the appropriate calculation of the differences between amino acid sequences contained in two different species. Furthermore, based on these differences, the degree of similarity between the amino acid sequences of the two different species can be quantified.

[0095] (Section 5) In the method for creating training data described in Section 4, the specific difference is a numerical value calculated based on the edit distance of the amino acid sequence or the Average Amino-acid Identity method.

[0096] The method for creating training data described in Section 5 allows for the appropriate calculation of the differences between amino acid sequences contained in two different species. Furthermore, based on these differences, the degree of similarity between the amino acid sequences of the two different species can be quantified.

[0097] (Section 6) In the method for creating training data described in paragraph 5, if the specific difference is the edit distance, the specific condition is that it is less than a predetermined natural number less than or equal to 10.

[0098] According to the method for creating training data described in Section 6, closely related species can be appropriately estimated based on the above-mentioned specific differences and specific conditions.

[0099] (Section 7) In the method for creating training data described in any one of sections 3 to 6, the step of obtaining amino acid sequences includes the step of obtaining only amino acid sequences from an amino acid sequence database containing amino acid sequences corresponding to protein names, where the protein name contains a specific string.

[0100] According to the method for creating training data described in Section 7, it is possible to estimate closely related species based on specific proteins, rather than all of them.

[0101] (Section 8) In the method for creating training data described in paragraph 7, the specific string includes a string that represents at least one of the housekeeping protein and the DNA-binding protein.

[0102] According to the method for creating training data described in Section 8, closely related species can be estimated based on the similarity of proteins that are considered to have important functions and conserved amino acid sequences, as indicated by the above strings.

[0103] (Section 9) In the method for creating training data described in any one of sections 3 to 8, the step of estimating related species includes the step of estimating related species by using a statistical method that uses at least one of the following for similarity: threshold, mean, standard deviation, or outlier test.

[0104] According to the method for creating training data described in Section 9, statistically related species can be estimated by statistically processing similarity. This allows for the estimation of related species in a simple manner, without involving user subjectivity.

[0105] (Section 10) A method for creating training data as described in any one of paragraphs 1 to 9, wherein the step to be included in the training data includes the step of assigning the same label to the types of microorganisms included in the group.

[0106] According to the method for creating training data described in Section 10, the machine learning model using the training data does not perform discrimination among the types of microorganisms included in the group. Such a machine learning model can prevent misclassification between types of microorganisms with similar genetic information.

[0107] (Section 11) A machine learning model created using training data created using the method for creating training data described in any one of items 1 to 10.

[0108] According to the machine learning model described in Section 11, misclassification between microbial species with similar genetic information can be prevented.

[0109] (Section 12) A method for identifying microorganisms, which uses the machine learning model described in Section 11 to identify microorganisms.

[0110] The method for identifying microorganisms described in paragraph 12 reduces the possibility of misidentification between microorganism species with similar genetic information.

[0111] (Section 13) In the method for identifying microorganisms described in paragraph 12, the machine learning model includes a neural network.

[0112] In the method for identifying microorganisms described in paragraph 13, the method for identifying microorganisms described in paragraph 12 can be performed using a neural network.

[0113] (Clause 14) An analytical device according to one embodiment is an analytical device that creates training data for a machine learning model that identifies microorganisms using mass spectrometry. The analytical device comprises a memory and a processor. The memory stores the genetic information of microorganisms. The processor executes a method for creating training data using the genetic information stored in the memory. The processor obtains the genetic information of each microorganism, estimates microorganisms with similar genetic information, creates groups containing the microorganisms estimated to have similar genetic information, and includes information about the groups in the training data.

[0114] The analysis device described in Section 14 can prevent misclassification among microbial species with similar genetic information. In other words, it can provide a machine learning model with a low probability of misclassification among microbial species with similar genetic information, and training data for creating said machine learning model.

[0115] (Section 15) A program that, when executed by a computer, causes the computer to perform the method of creating training data described in any one of items 1 through 10.

[0116] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the present invention is indicated by the claims rather than by the foregoing description, and all modifications within the meaning and scope equivalent to the claims are intended to be included. [Explanation of Symbols]

[0117] 10 Processor, 11 Memory, 12 Communication I / F, 13 Input / Output I / F, 14 Control Unit, 15 Display, 16 MS, 100 Analysis Device, 101 Controller.

Claims

1. A method for creating training data for a machine learning model that identifies types of microorganisms using mass spectrometry, The steps include obtaining genetic information for each type of microorganism, The steps include: estimating the types of microorganisms with similar genetic information, and creating a group that includes the types of microorganisms that are estimated to have similar genetic information; A method for creating training data for a machine learning model, comprising the step of creating training data for a machine learning model in which the mass spectrum of the microorganism is used as an explanatory variable, the labels commonly assigned to the group for the types of microorganisms included in the group are used as the dependent variable, and the labels assigned to each type of microorganism for the types of microorganisms not included in the group are used as the dependent variable.

2. The type of microorganism is the species of the microorganism, The aforementioned microorganisms include one or more species, The method for creating training data according to claim 1, wherein the types of microorganisms with similar genetic information are closely related species.

3. The step of obtaining the aforementioned genetic information is: The process includes obtaining one or more amino acid sequences for each protein corresponding to each of the aforementioned microorganisms. The step of creating the aforementioned group is: For each type, the step is to create a sequence list that contains one or more amino acid sequences as elements for each protein, For each type, the steps include creating a protein list using the sequence list as an element, A step of calculating the similarity of the protein lists between two different species, The steps include: estimating closely related species based on the aforementioned similarity; A method for creating training data according to claim 2, comprising the step of creating a group that includes the closely related species.

4. The two distinct species mentioned above include species 1 and species 2. The step of calculating the similarity is: The steps include: determining whether the amino acid sequence of each protein included in the first protein list satisfies specific conditions with respect to specific differences between it and each amino acid sequence included in the second protein list; and calculating the specific number of proteins for which such specific differences may satisfy specific conditions. A method for creating training data according to claim 3, comprising the step of calculating the ratio of the specified number to the total number of proteins included in the first type protein list as the similarity of the second type to the first type.

5. The method for creating training data according to claim 4, wherein the aforementioned specific difference is a numerical value calculated based on the edit distance of the amino acid sequence or the Average Amino-acid Identity method.

6. The method for creating training data according to claim 5, wherein, when the specified difference is the edit distance, the specified condition is that it is less than a predetermined natural number of 10 or less.

7. The step of obtaining the aforementioned amino acid sequence is: A method for creating training data according to claim 3, comprising the step of obtaining only amino acid sequences whose protein names contain a specific string from a database of amino acid sequences containing amino acid sequences corresponding to protein names.

8. The aforementioned specific string is a small amount of housekeeping protein and DNA-binding protein. A method for creating training data according to claim 7, which includes a string indicating at least one.

9. The step of estimating the closely related species is: A method for creating training data according to claim 3, comprising the step of estimating closely related species by using a statistical method that employs at least one of a threshold, mean, standard deviation, and outlier test for the similarity.

10. A machine learning model for determining the type of microorganism using a computer, The machine learning model is trained using training data created using the method for creating training data described in claim 1. A machine learning model that, when given the mass spectrum of the microorganism as an explanatory variable, outputs a label indicating the type or group of the microorganism corresponding to the mass spectrum as the target variable.

11. A method for identifying microorganisms, comprising inputting the mass spectrum of the microorganism as an explanatory variable into the machine learning model described in claim 10, and obtaining a label indicating the type or group of the microorganism corresponding to the mass spectrum as an objective variable.

12. The method for identifying microorganisms according to claim 11, wherein the machine learning model includes a neural network.

13. An analytical device that creates training data for a machine learning model that identifies microorganisms using mass spectrometry, A memory that stores the genetic information of microorganisms, The system comprises a processor that performs a method for creating training data using genetic information stored in the memory, The aforementioned processor, Obtain the aforementioned genetic information for each microorganism, We estimate microorganisms with similar genetic information and create a group containing the microorganisms estimated to have similar genetic information. An analysis device for creating training data for a machine learning model, in which the mass spectrum of the microorganisms is used as the explanatory variable, the label commonly assigned to the group for the types of microorganisms included in the group is used as the dependent variable, and the label assigned to each type of microorganism for the types of microorganisms not included in the group is used as the dependent variable.

14. A program that, when executed by a computer, causes the computer to perform the method for creating the learning data described in claim 1.