Microorganism estimation device and microorganism estimation program
The microorganism estimation device enhances the accuracy of microorganism identification by using autofluorescence data and multiple algorithms, addressing the limitations of existing methods and improving estimation accuracy to near 100%.
Patent Information
- Application Number
- JP2023188585
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-16
AI Technical Summary
Existing methods for identifying microorganisms, such as mass spectrometry and NMR, have limited correct answer rates, resulting in incorrect results for approximately 0.1% of cases, which can lead to delayed treatment and inappropriate drug use.
A microorganism estimation device that utilizes autofluorescence information and multiple algorithms to estimate microorganisms. The device calculates probabilities using these algorithms and references past correct answer rates to improve accuracy, especially when initial probabilities are below a predetermined threshold.
The device achieves more accurate and rapid microorganism identification, improving estimation accuracy from 99.9% to near 100% by leveraging multiple algorithms and past performance data.
Smart Images

Figure 2025076759000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus and a program for predicting the type of microorganism present in a sample to be predicted. [Background technology]
[0002] When detecting and removing microorganisms contained in foodstuffs, etc., or when determining a treatment plan for a patient suspected of having an infection caused by a microorganism, it is first necessary to identify the type of microorganism.
[0003] A method for identifying the type of microorganism has been used in which the microorganism in the suspected sample is cultured, isolated, and then examined for resistance to various drugs to identify the type of microorganism, but this conventional method requires at least about one week to identify the type of microorganism. Therefore, this identification method is not suitable for, for example, food supply sites where speed is required, and when identifying a microorganism to determine a patient's treatment plan, there is a risk that the patient's symptoms will worsen before the type of microorganism is identified, side effects due to the use of inappropriate drugs, and problems such as drug-resistant bacteria may occur.
[0004] Therefore, there is a demand for a method for identifying the type of microorganism in a shorter period of time, and one such method is a method for identifying the type of microorganism by examining the increase or decrease in chemical substances derived from the microorganism using mass spectrometry or NMR. In addition, as described in Patent Document 1, a method has been proposed in which the autofluorescence of known microorganisms is examined, and the microorganism is identified based on the autofluorescence of the microorganism in the sample to be estimated using an algorithm. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2022-140443 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, even when using high-precision analytical techniques such as mass spectrometry and NMR, the accuracy rate for identifying the type of microorganism is at most about 99.9%, and the remaining 0.1% The current situation is that the output of the test results incorrectly indicates that the microorganism is a different type from the actual microorganism.
[0007] The present invention has been made in consideration of such problems, and has an object to provide an apparatus etc. capable of estimating microorganisms in a short period of time with higher accuracy than conventional methods. [Means for solving the problem]
[0008] That is, the microbial estimation device according to the present invention includes an estimation unit that estimates a microbial species based on information about the microbial species' autofluorescence, The estimation unit estimates microorganisms using a plurality of algorithms, If the probability of the type of microorganism calculated by each of the algorithms is higher than a predetermined threshold, the type of microorganism is output as an estimated result; If the probability of the type of microorganism calculated by each of the algorithms is lower than a predetermined threshold, The method is characterized in that the microorganism is predicted based on the probability calculated by each of the algorithms and the past accuracy rates for each type of microorganism for each of the algorithms.
[0009] The present invention was completed only after the inventors investigated many algorithms suitable for estimating microorganisms and found that each algorithm has strengths and weaknesses in estimation for each type of microorganism.
[0010] When a microorganism is estimated by the microorganism estimation device configured as described above, in most cases (99.9%), an estimation result is output in which one or more algorithms estimate that the microorganism is highly likely to be present. However, even if estimation is performed using multiple algorithms, there are cases (the remaining 0.1% of cases) in which the probability of a certain microorganism is calculated to be low in all algorithms. According to the present invention, in such cases where a microorganism cannot be estimated with a high probability, the microorganism is estimated based on the probability calculated by each algorithm and the past accuracy rate for each type of microorganism for each algorithm. Therefore, even if the accuracy of the estimation result appears to be low at first glance, by adopting the estimation result of an algorithm that is known to be able to estimate the microorganism with high accuracy, the microorganism can be estimated as accurately as possible, and the estimation accuracy can be improved compared to the past. Effect of the Invention
[0011] According to the present invention, microorganisms can be estimated more accurately in a short period of time than ever before. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram showing a schematic diagram of the overall configuration of a microbial estimation device according to one embodiment of the present invention. [Diagram 2] FIG. 2 is a schematic diagram showing an example of a microbial estimation process using the microbial estimation device according to the present embodiment. [Diagram 3] 1 is a graph showing excitation wavelengths and fluorescence wavelengths for microbial autofluorescence, and an example of information regarding the microbial autofluorescence. [Figure 4] An image showing the correlation between the excitation wavelength used and the estimated results of microorganisms. [Diagram 5] An illustration showing the correlation between principal component analysis and microbial estimation results. [Figure 6] An illustration showing the variation in estimation accuracy for various microorganisms using multiple algorithms. [Figure 7]FIG. 4 is a diagram showing an example of a microbial estimation result obtained using the microbial estimation device according to the present embodiment. [Figure 8] FIG. 4 is a diagram showing an example of a microbial estimation result obtained using the microbial estimation device according to the present embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] An embodiment of a microbial estimation device of the present invention will be described below with reference to the drawings.
[0014] The microbial estimation device 100 of this embodiment can be used, for example, to ensure the quality of beverages and food by identifying the types of microorganisms contained in the beverages and food, or can function as a diagnostic support device that provides information that can be used as basis for doctors to make diagnoses and determine treatment plans by estimating the types of microorganisms contained in biological samples collected from a patient's body.
[0015] As shown in FIG. 1, the microorganism estimation device 100 includes a detection unit 1 that detects the autofluorescence of microorganisms, and an estimation device main body 2 that is connected to the detection unit 1 by wire or wirelessly.
[0016] The detection unit 1 includes, for example, a holding section for holding an estimation target sample containing a microorganism, a light source for emitting excitation light to the estimation target sample held in the holding section, a detector for detecting fluorescence emitted from the microorganism by the excitation light, and an output section for outputting information relating to the autofluorescence detected by the detector to the estimation device main body, either without processing or after processing.
[0017] Structurally, the estimation device main body 2 is a general-purpose computer equipped with a CPU, memory, an input / output interface, etc. Based on various application software (hereinafter referred to as programs) stored in the memory, the estimation device main body 2 is configured to function as, for example, an information acquisition section that acquires information on the autofluorescence of a microorganism output from a detection unit, an estimation section that estimates the type of the microorganism based on the information acquired by the information acquisition section, and a storage section that stores the estimation result estimated by the estimation section, as shown in Fig. 1, by the CPU and its peripheral devices working together. Hereinafter, the operation of the microorganism estimation device 100 according to this embodiment and each section will be described.
[0018] 2, first, excitation light is irradiated from a light source onto a suspected target sample set in a holder of the detection unit 1. The fluorescence emitted from the microorganisms excited by this excitation is detected by a detector. The excitation light can be changed as appropriate based on the type of sample to be estimated or the microbial species previously nominated by the doctor as candidates. Also, the light source may irradiate the same sample to be estimated with excitation light of various wavelengths, and the detector may detect the fluorescence for each excitation wavelength.
[0019] A set of excitation wavelengths and fluorescence intensities at the excitation wavelengths thus obtained, for example as shown in FIG. 3(a), are output from the detection unit 1 to the estimation device main body 2 as one piece of information on the autofluorescence of microorganisms. The information on the autofluorescence of microorganisms may include fluorescence intensities at a specific excitation wavelength as shown in FIG. 3(b), or may be a set of values including the fluorescence intensities at a plurality of excitation wavelengths obtained for one estimation target sample, or may be values obtained after removing noise from various detection values as described above by correction using a comparison target, etc. Although the preferred range of the excitation wavelengths described above varies somewhat depending on the species of bacteria contained in the estimation target sample, it is preferable to use one within the range of 200 nm to 400 nm, more preferably one within the range of 250 nm to 350 nm, and even more preferably one within the range of 280 nm to 320 nm, as shown in FIG. 4, for example. Similarly, although the preferred range of fluorescence wavelengths varies depending on the species of bacteria contained in the sample to be estimated, it is preferable to use a range of, for example, 400 nm or more and 600 nm or less, and it is more preferable to use a range of 450 nm or more and 550 nm or less. The abbreviations of the microorganisms shown in FIG. 4 indicate the following microorganisms. ACBA: Acinetobacter, ESCO: Escherichia coli, KLPN: Klebsiella pneumoniae, PSAE: Pseudomonas aeruginosa, SATY: Salmonella enterica, STAU: Staphylococcus aureus, STHA: Streptococcus pyogenes, STPN: Streptococcus pneumoniae, STPY: Streptococcus pyogenes.
[0020] The information acquisition unit has a function of acquiring information on the autofluorescence of the microorganism output from the detection unit 1. The information acquisition unit, for example, pre-processes a set of acquired excitation wavelengths and fluorescence intensities at the excitation wavelengths, and then outputs the pre-processed set to the estimation unit. As the pre-processing, for example, a fluorescence intensity detected using the same excitation wavelength for a reference sample adjusted to a composition as close as possible to the estimation target sample to be estimated may be subtracted from the fluorescence intensity acquired from the detection unit 1 as a reference value to calculate a difference in fluorescence intensity, and the difference may be supplied to the estimation unit. For example, as shown in FIG. 3(c), it may include values related to each principal component obtained by principal component analysis. When the above-mentioned principal component analysis or the like is performed as pre-processing, as shown in FIG. 5, the number of axes of the principal components to be extracted is preferably 3 or more, more preferably 6 or more, and even more preferably 10 or more. The more the number of axes of the principal components is, the higher the estimation accuracy is, but if the number is increased too much, the calculation processing becomes complicated, so the number of axes of the principal components is preferably 20 or less, and more preferably 15 or less. The information acquisition unit may output the acquired information to the estimation unit as is without performing preprocessing.
[0021] The estimation unit estimates the type of microorganism contained in the estimation target sample using a plurality of algorithms, based on the information acquired by the information acquisition unit and preprocessed if necessary.
[0022] The multiple algorithms are not particularly limited, but preferably include one or more of the following: QuadraticDiscriminantAnalysis, LogisticRegressionCV, ExtraTreesClassifier, LogisticRegression, Perceptron, LabelPropagation, LinearSVC, LabelSpreading, HistGradientBoostingClassifier, PassiveAggressiveClassifier, SGDClassifier, AdaBoostClassifier, CalibrateClassifierCV, RandomForestClassifier, GaussianNB, MLPClassifier, SVC, GradientBoostingClassifier, KNeighborsClassifier, DecisionTreeClassifier, BaggingClassifier, etc.
[0023] The memory unit may store the estimation result estimated by the estimation unit, for example, together with information regarding the autofluorescence of the microorganism that was the basis for obtaining the estimation result and the type of algorithm used for the estimation.
[0024] The estimation unit is configured to simultaneously calculate not only the type of microorganism but also the probability (first likelihood) that the microorganism contained in the estimation target sample is the estimated microbial species for multiple algorithms used to estimate the type of microorganism.
[0025] The storage unit may store a predetermined first threshold value for the first likelihood used to be calculated together with the estimation result. This first threshold is set individually for each algorithm for each microbial species, and is set in advance with reference to the accuracy rate of each algorithm for each microbial species. For example, the accuracy rate of Quadratic Discriminant Analysis (QDA), one of the above-mentioned algorithms, for KLPN, a microbial species, is very high at about 0.9 (i.e., 90%), so the first threshold for estimating QDA to be KLPN is set to a relatively low value of about 0.6. As such, the first threshold value can be changed as appropriate depending on the type of algorithm used, the required estimation accuracy, and the microbial species. However, it is preferable to set the first threshold value to, for example, 0.6 or more and 0.9 or less, where 1 is the likelihood that it is definitely the microorganism and 0 is the likelihood that it is completely unknown which microorganism it is.
[0026] The inventors have found that the accuracy rate of each algorithm for each microorganism varies, for example, as shown in Figure 6. Since each algorithm has strengths and weaknesses depending on the microbial species, even if the likelihood of a certain microorganism being identified appears low at first glance in all of the multiple algorithms as described above, it is possible that the algorithm used can accurately predict that a certain microorganism is identified. In fact, when a certain type of microorganism is suspended in LB medium or the like and the first likelihood is calculated by the method described above, as shown in Figure 7, depending on the algorithm, it may not be possible to estimate the correct microorganism even if the first likelihood is a very large number, while using another algorithm, it may be possible to estimate the correct microorganism even if the first likelihood seems low.
[0027] The accuracy rate of each algorithm for each microorganism can be obtained, for example, as follows. A calculation unit may be provided that allows the results of analysis using a different method such as a culture method or the actual microbial species obtained as a result of treatment to be inputted after the fact to the past estimation results stored in the storage unit, and calculates the accuracy rate for the estimation results estimated using each algorithm. Also, a user may input the accuracy rate calculated separately by hand and store it in the storage unit.
[0028] The estimation unit obtains the type of microorganism estimated by each algorithm and a first likelihood that the microorganism contained in the estimation target sample is that microorganism. After calculating the estimation result and the first likelihood for each algorithm, if the calculated first likelihood is the same as or exceeds the first threshold value in at least one of the multiple algorithms used for estimation, the estimation unit outputs the microorganism as it is as the estimation result. On the other hand, if the first likelihood is lower than the above-mentioned first threshold in all of the multiple algorithms used for the estimation, the estimation unit continues to perform the next step.
[0029] Specifically, for example, if a certain microorganism name is estimated for a certain sample by the above-mentioned algorithm QDA with a probability of 95% (i.e., the first probability is 0.95), the estimation unit will adopt the name of the microorganism as the final estimation result. Note that, if the QDA estimates the name of the microorganism with a probability of 95% (i.e., the first probability is 0.95), while another algorithm estimates a different type of microorganism with a first probability exceeding the first threshold, the estimation unit may adopt the estimation result by the algorithm with the largest difference between the first probability and the first threshold. Also, if multiple algorithms each output a different microorganism species as an estimation result, and there are multiple algorithms with the same difference between the first probability and the first threshold, the process will proceed to the next step in the same way as when the first probability is below the above-mentioned first threshold in all of the multiple algorithms used for estimation.
[0030] In this step, the estimation unit obtains a second likelihood using the accuracy rate when estimating each microorganism for each algorithm used for the estimation. Specifically, the memory unit stores the accuracy rate of each algorithm for each microorganism, and the estimation unit refers to this, and in addition to the microbial species estimated by a certain algorithm and the first certainty, uses the past accuracy rate of the algorithm for that microorganism, for example, multiplying the first certainty by the accuracy rate to obtain a second certainty, and compares this second certainty with a second threshold value predetermined for each algorithm, and if it exceeds the second threshold value, the microbial species may be output as the estimation result. On the other hand, if the second certainty is the same value as the second threshold value or is lower than the second threshold value, a result indicating that estimation is impossible may be output, or the estimation may be restarted from the beginning using another algorithm that has not yet been used, until the first certainty exceeds the first threshold value or the second certainty exceeds the second threshold value.
[0031] Specifically, for example, for the above-mentioned algorithm QDA (correct answer rate for KLPN: 90%), the second threshold may be set to, for example, 0.3 based on the accuracy rate similar to the first threshold, and for the other algorithm LRCV (correct answer rate for PSAE#2: 30%), the second threshold may be set to 0.4, etc., so that the second threshold for each microbial species of each algorithm may be predefined in advance. In this case, if KLPN is estimated by the above-mentioned algorithm QDA with a probability of 50% (first certainty is 0.5) and PSAE#2 is estimated by LRCV with a probability of 80% (first certainty is 0.8), the second certainty when it is estimated to be KLPN by QDA can be calculated as 0.45 from (correct answer rate, 0.9) × (first certainty, 0.3), which exceeds the second threshold. On the other hand, when the PSAE#2 is estimated by LRCV, the second likelihood can be calculated as 0.24 from (correct answer rate, 0.3) × (first likelihood, 0.8), which is below the second threshold. Therefore, in this case, the estimation unit adopts KLPN by QDA as the estimation result.
[0032] As in the case of the first certainty, if the KLPN is estimated by QDA with a second certainty exceeding the second threshold, while another algorithm estimates a different type of microorganism with a second certainty exceeding the second threshold, the estimation unit may adopt the estimation result by the algorithm with the largest difference between the second certainty and the second threshold. If multiple algorithms output different microorganism species as estimation results and there are multiple algorithms with the same difference between the second certainty and the second threshold, the estimation unit may output a result that estimation is impossible, or may continue the operation of starting the estimation over from the beginning using another algorithm that has not yet been used, until the first certainty exceeds the first threshold or the second certainty exceeds the second threshold.
[0033] <Effects of the microbial estimation device according to this embodiment> The thus configured microbial estimation device 100 of this embodiment can quickly and simply identify microbial species with a level of accuracy that was previously difficult to achieve even with precision equipment, as shown in Fig. 8. In the case of Fig. 8, the results of estimation of each isolated microorganism by the above-mentioned method are shown. According to the estimation device and estimation method described in this embodiment, the types of microorganisms such as streptococci such as Escherichia coli, Streptococcus pneumoniae, and Streptococcus pyogenes, Pseudomonas aeruginosa, enterobacteria such as Klebsiella and Proteus, Acinetobacter, and Salmonella can be estimated with high accuracy.
[0034] The present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present invention. For example, the functions of each part described in the above-described embodiment as being performed by a computer may be partially performed by a human user. In order to reduce the number of calculations and minimize the burden on the computer, it is preferable that the number of algorithms used by the estimation unit in one estimation process be between 2 and 50. Of the algorithms that the estimation unit can use, for example, about 20 types may be used to perform the first estimation, and if the estimation is difficult, the estimation may be redone using an algorithm other than the one used last time. [Explanation of symbols]
[0035] 100... Microorganism estimation device 1. Detection unit 2. Estimation device body
Claims
1. The present invention further includes an estimation unit that estimates a microorganism based on information about the autofluorescence of the microorganism, The estimation unit estimates microorganisms using a plurality of algorithms, If the probability of the type of microorganism calculated by each of the algorithms is higher than a predetermined threshold, the type of microorganism is output as an estimated result; If the probability of the type of microorganism calculated by each of the algorithms is lower than a predetermined threshold, A microbial estimation device which estimates a microorganism based on the probability calculated by each of the algorithms and the past accuracy rates for each type of microorganism for each of the algorithms.
2. The plurality of algorithms include QuadraticDiscriminantAnalysis, LogisticRegressionCV, ExtraTreesClassifier, LogisticRegression, Perceptron, LabelPropagation, LinearSVC, LabelSpreading, HistGradientBoostingClassifier, PassiveAggressiveClassifier, SGDClas 2. The microbial estimation device according to claim 1, wherein the classifiers are two or more selected from the group consisting of AdaBoostClassifier, AdaBoostClassifier, CalibrateClassifierCV, RandomForestClassifier, GaussianNB, MLPClassifier, SVC, GradientBoostingClassifier, KNeighborsClassifier, DecisionTreeClassifier and Bagging Classifier.
3. The microbial estimation device according to claim 1 or 2, wherein the threshold value is 60% or more and 90% or less.
4. A program for a microorganism estimation device having an estimation unit that estimates a microorganism based on information about the autofluorescence of the microorganism, It uses multiple algorithms to predict microorganisms, If the probability of the type of microorganism calculated by each of the algorithms is higher than a predetermined threshold, the type of microorganism is output as an estimated result; A microbial estimation program that causes a computer to function as an estimation unit that estimates a microorganism based on the probability calculated by each algorithm and the past accuracy rate for each type of microorganism for each algorithm when the probability of the type of microorganism calculated by each algorithm is lower than a predetermined threshold.
Citation Information
Patent Citations
Processing system, data processing device, display system and microscope system
JP2022140443A