Microorganism species identification method and device and computer readable storage medium

By collecting and analyzing single-cell Raman spectral data, and using an L1 regularization model to train a microbial species identification method, the problem of simultaneous characteristic peak identification and classification in existing technologies has been solved. This method enables rapid and accurate identification and characteristic peak extraction of microbial species, and is suitable for identifying target species in complex samples.

CN120801274APending Publication Date: 2025-10-17GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411379605.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing microbial species identification methods cannot achieve simultaneous recognition and classification of characteristic peaks of Raman spectral data. Conventional linear models are prone to overfitting, and nonlinear models and dimensionality reduction clustering methods cannot determine the characteristic peaks that contribute to species identification, resulting in inaccurate and unstable identification.

Method used

By employing a regularized model, the area under the Raman shift curve is determined by collecting single-cell Raman spectral data, and the model is trained using the L1 regularization algorithm to achieve rapid identification of microbial species and effective extraction of characteristic peaks.

Benefits of technology

It enables rapid, accurate, and interpretable identification of microbial species, is suitable for identifying target species in complex samples, reduces human intervention, and improves operational convenience and identification efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120801274A_ABST
    Figure CN120801274A_ABST
Patent Text Reader

Abstract

The invention provides a microbial species identification method and device and a computer readable storage medium, and belongs to the technical field of species identification. The method comprises the following steps: collecting a plurality of single-cell Raman spectrum data of a plurality of microorganisms; analyzing and processing the plurality of single-cell Raman spectrum data, and determining the area of each Raman shift under a curve in a preset range; and inputting the curve area of each Raman shift in the predetermined range into a trained regularization model, and determining the species type of each microbial single cell. According to the method, species identification of microorganisms can be realized in a real-time and lossless manner at the single cell level, the method is suitable for identifying cells of target species in a complex microbiome sample, and industrial transformation can be carried out in cooperation with a single cell sorting platform. In addition, according to the method, Raman spectrums of different microbial species are analyzed by utilizing a regularization algorithm, and synchronous integration of microbial characteristic peak identification and accurate classification is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of species identification, and in particular to a microorganism species identification method, device and computer readable storage medium. BACKGROUND

[0002] Raman spectroscopy has shown great potential in the identification of microorganism species (such as bacteria or fungi). The identification and classification of characteristic peaks based on Raman spectroscopy is usually divided into two independent steps, and the two steps cannot be achieved at one step at present.

[0003] The combination of Raman spectroscopy and machine learning algorithms provides new possibilities for microorganism species identification. By analyzing Raman spectroscopy data, a classification model of microorganisms can be established to achieve automated and high-throughput microorganism species identification. The machine learning models currently available for Raman spectroscopy analysis include conventional linear models, nonlinear models, and dimensionality reduction clustering methods.

[0004] Conventional linear models are prone to overfitting, unstable parameter estimation, and other problems due to the high dimensionality and multicollinearity of Raman spectroscopy data, and therefore a set of characteristic peaks must be determined in advance, but it is difficult to determine a small number of peak positions that contribute most to species identification; nonlinear models (such as deep neural networks, random forests, and support vector machines) classify based on the intensity values of all Raman shifts, although random forests and other models can output feature importance scores, but still only rank the importance of all Raman shifts, and such methods cannot determine a small number of characteristic peak combinations that contribute to species identification and balance overfitting and underfitting; dimensionality reduction clustering methods (such as principal component analysis, t-distributed stochastic neighbor embedding, and uniform manifold approximation and projection) cannot determine which characteristic peaks contribute to species identification, and it is also challenging to establish classification boundaries based on dimensionality reduction graphs.

[0005] Therefore, the existing microorganism species identification method still needs to be improved. SUMMARY

[0006] The present application aims to solve at least one of the existing problems. To this end, the present application proposes a method for simultaneous microorganism characteristic peak identification and accurate classification based on Raman spectroscopy, which realizes rapid identification of microorganism species and effective extraction of characteristic peaks through a regularization model, and improves the accuracy and interpretability of microorganism species identification.

[0007] Specifically, the present application provides the following technical solutions:

[0008] In a first aspect, the present application provides a method for identifying a microbial species. According to an embodiment of the present application, the method comprises: collecting single-cell Raman spectrum data of a plurality of microorganisms; processing the single-cell Raman spectrum data to determine an area under a curve of each Raman shift within a predetermined range; and inputting the area under the curve of each Raman shift within the predetermined range into a trained regularization model to determine a species type of each microbial single cell.

[0009] In some examples of the present application, the foregoing method can achieve real-time and non-destructive identification of microbial species at the single-cell level, and is suitable for identifying cells of a target species in a complex microbiome sample (e.g., feces), and is conducive to industrialization with a single-cell sorting platform. In addition, the foregoing method achieves simultaneous integration of identification of characteristic peaks and accurate classification of microorganisms by using a regularization algorithm to analyze Raman spectra of different microbial species.

[0010] In a second aspect, the present application provides a device for identifying a microbial species. According to an embodiment of the present application, the device comprises: a single-cell Raman spectrum data collection module configured to collect single-cell Raman spectrum data of a plurality of microorganisms; a Raman shift curve area determination module configured to process the single-cell Raman spectrum data to determine an area under a curve of each Raman shift within a predetermined range; and a discrimination module configured to input the area under the curve of each Raman shift within the predetermined range into a trained regularization model to determine a species type of each microbial single cell.

[0011] In some examples of the present application, the foregoing device can efficiently, accurately, and in real time and non-destructively achieve identification of microbial species at the single-cell level, and is suitable for identifying cells of a target species in a complex microbiome sample (e.g., feces). The device achieves simultaneous integration of identification of characteristic peaks and accurate classification of bacteria by using a regularization algorithm to analyze Raman spectra of different bacterial species. At the same time, the device reduces the need for manual intervention and improves operational convenience.

[0012] In a third aspect, the present application provides a computer program product. According to an embodiment of the present application, the computer program product comprises: computer instructions; and when part or all of the computer instructions are executed on a computer, the computer program product causes the method for identifying a microbial species as described in the first aspect to be performed.

[0013] In a fourth aspect, the present application provides a computing device. According to an embodiment of the present application, the computing device comprises: a processor and a memory; the memory is configured to store a computer program; and the processor is configured to execute the computer program to implement the method for identifying a microbial species as described in the first aspect.

[0014] In a fifth aspect, the present application provides a computer-readable storage medium. According to an embodiment of the present application, the computer-readable storage medium stores computer instructions or programs, when the computer instructions or programs are executed on a computer, the computer instructions or programs cause the microbial species identification method as described in the first aspect to be executed.

[0015] In some examples of the present application, the aforementioned computer program product, computing device and computer-readable storage medium implement the aforementioned microbial species identification method through automatic execution of computer instructions, achieving efficient automation, improving identification efficiency and accuracy. Secondly, based on the characteristics of the instructions, the above method has high consistency and reliability in different application scenarios.

[0016] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0018] Figure 1 The microbial species identification method flowchart provided for the specific embodiments of the present application;

[0019] Figure 2 The microbial species identification device schematic diagram provided for the specific embodiments of the present application;

[0020] Figure 3 The microbial species identification device schematic diagram provided for the specific embodiments of the present application (including model training and feature Raman shift extraction sub-module);

[0021] Figure 4 The electronic device schematic diagram provided for the specific embodiments of the present application;

[0022] Figure 5 The Raman spectrum result schematic diagram of Bacteroides intestinalis and Bacteroides xylanisolvens provided for the embodiments of the present application;

[0023] Figure 6 The classification performance result schematic diagram of the Bacteroides model provided for the embodiments of the present application;

[0024] Figure 7 The Raman spectrum result schematic diagram of oral streptococcus and salivary streptococcus provided for the embodiments of the present application;

[0025] Figure 8 A classification performance result diagram of a streptococcus model provided for an embodiment of the present application is shown in the figure;

[0026] Figure 9 A Raman spectrum result diagram of lactobacillus crispatus and lactobacillus jensenii provided for an embodiment of the present application is shown in the figure;

[0027] Figure 10 A classification performance result diagram of a lactobacillus model provided for an embodiment of the present application is shown in the figure;

[0028] Figure 11 A Raman spectrum result diagram of bifidobacterium longum and bifidobacterium adolescentis provided for an embodiment of the present application is shown in the figure;

[0029] Figure 12 A classification performance result diagram of a bifidobacterium model provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0031] It should be noted that the terms “first”, “second”, and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In the present application, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, “a plurality of” means two or more, unless otherwise specified.

[0032] In this paper, unless otherwise specified, the term “Raman spectrum” is an analytical technique based on Raman scattering effect. Raman scattering is the energy exchange between photons and molecules when light interacts with matter, resulting in a difference in frequency between scattered light and incident light. This frequency change is called Raman shift, which corresponds to the change of molecular vibration energy level, so it can be used to identify molecules and study their structure. In some examples of the present application, the inventors simultaneously integrate microbial species identification and characteristic peak identification based on Raman spectrum technology.

[0033] In this document, the term "single-cell Raman spectrum data" refers to spectral information obtained by analyzing a single cell through Raman spectroscopy technology, reflecting the vibration characteristics and chemical composition of the molecules in the cell. In some examples of the present application, the aforementioned single cell is selected from a single bacterium.

[0034] The present application is proposed to solve the technical problems existing in the prior art, such as the inability to simultaneously achieve microbial classification and characteristic peak identification; the overfitting and parameter instability of the machine learning method for identifying microbial species; the inability to determine the few characteristic peaks that contribute most to species identification; and the inability to determine which characteristic peaks contribute to classification.

[0035] To this end, the inventors have conducted in-depth research and proposed a microbial species identification method, a microbial species identification device, a computer program product, a computing device, and a computer storage medium. The following will be described in detail:

[0036] Bacterial identification method

[0037] In one aspect, the present application proposes a microbial species identification method, which refers to Figure 1 The method comprises the following steps:

[0038] S10, collecting single-cell Raman spectrum data of a plurality of microorganisms.

[0039] In some examples of the present application, the aforementioned microorganism types include fungi and / or bacteria. In some preferred examples of the present application, the aforementioned microorganism types are selected from bacteria.

[0040] In order to obtain more comprehensive, accurate and reliable microbial species identification results. In some examples of the present application, the aforementioned plurality of cells are selected from 100-2000 cells.

[0041] In this step, the single-cell Raman spectrum data of a plurality of microorganisms is obtained by culturing and sample preparation under suitable conditions. The prepared sample is subjected to spectral collection to determine the single-cell Raman spectrum data of the plurality of microorganisms. Those skilled in the art can understand that the aforementioned suitable conditions refer to conditions suitable for the growth and reproduction of microorganisms, including but not limited to suitable temperature, suitable oxygen concentration, suitable culture medium, suitable pH, etc.

[0042] For ease of understanding, the step of culturing and sample preparation under suitable conditions is described in detail with bacteria as an example:

[0043] (1) Take different species of bacterial strains, inoculate in broth or agar medium suitable for their growth, and culture for 12-48 h under conditions of temperature and oxygen concentration consistent with the growth characteristics of the bacteria. Centrifuge the bacterial solution at 3000-8000 rpm for 5 min, discard the supernatant, add sterile water to resuspend the precipitate, and resuspend or repeat 2-3 times by single centrifugation;

[0044] (2) Take 1-2 μL of the resuspended sample and spot it on the Raman chip. After natural air drying at room temperature, perform machine operation.

[0045] As known by those skilled in the art, in order to improve the signal-to-noise ratio, optimize the resolution, and enhance the identification ability of the characteristic peaks, the Raman spectral data acquisition parameters can be adjusted adaptively. In some examples of the present application, the aforementioned spectral acquisition parameters are selected from: excitation wavelength 532 nm, 600 g / mm grating, spectral range 400 cm -1 -3600 cm -1 , laser intensity 1 mW-20 mW, single spectrum acquisition time 1 s-10 s.

[0046] For ease of understanding, taking bacteria as an example, the acquisition of multiple single-cell Raman spectral data of multiple microorganisms is described in detail:

[0047] The HOOKE PRECIS CS-R300 visual single-cell Raman detection and sorting instrument is used for Raman spectral detection, and the acquisition parameters include: the laser wavelength of the incident light source is 532 nm, a 600 g / mm grating is used, the Raman shift in the range of 400 cm -1 -3600 cm -1 is scanned, the laser intensity is 3 mW-5 mW, the acquisition time is 1 s-5 s, the spectrum of each cell is collected once, and 100-2000 cells of each species are collected.

[0048] S20, analyzing and processing the multiple single-cell Raman spectral data to determine the area under the curve of each Raman shift within a predetermined range.

[0049] To improve the signal-to-noise ratio and the repeatability and stability of the measurement, the inventors select the area under the curve within a predetermined range for each Raman shift as the input of the trained regularization model. In some examples of the present application, the analysis of the plurality of single-cell Raman spectral data to determine the area under the curve within a predetermined range for each Raman shift comprises: A) quality control processing of the single-cell Raman spectral data; B) using the composite Simpson rule to analyze the single-cell Raman spectral data after the quality control processing to determine the area under the curve within a predetermined range for each Raman shift. By integrating the area under the peak, the ratio of the intensity of the signal to the noise can be improved, thereby improving the sensitivity and accuracy of the analysis. In addition, the peak area is less sensitive to changes in instrument parameters and other conditions, so more consistent results can be obtained under different experimental conditions, reducing the impact of subtle fluctuations in experimental conditions.

[0050] In step A), the quality and reliability of the data are ensured by quality control processing. In some examples of the present application, the foregoing quality control processing is selected from at least one of eliminating cosmic rays, Savitzky-Golay smoothing, adaptive iteratively reweighted penalized least squares baseline correction, min-max normalization, and signal-to-noise ratio filtering.

[0051] Cosmic rays are high-energy particle rays that can introduce false signals into Raman spectral data, often appearing as extremely high spikes. These spikes can interfere with the analysis of true signals, leading to false results. Therefore, in single-cell Raman spectral data processing, these abnormal spikes are identified and removed by algorithms, which can improve the reliability of the data and ensure that subsequent analysis is based on true signals, thereby improving the accuracy of microbial species identification.

[0052] Savitzky-Golay smoothing is a method of reducing random noise and preserving signal shape by fitting local data points with polynomials. When processing single-cell Raman spectral data, using Savitzky-Golay smoothing can effectively remove high-frequency noise, highlight true Raman peaks, and make the identification of characteristic peaks clearer, thereby facilitating subsequent analysis and identification.

[0053] The adaptive iteratively reweighted penalized least squares (airPLS) baseline correction method corrects baseline drift in spectral data through iterative analysis, ensuring that the true characteristics of the signal are preserved. When applied to single-cell Raman spectral data, airPLS can effectively remove the effects of background noise, making the calculation of characteristic peaks more accurate, thereby improving the reliability of species identification.

[0054] Min-max normalization scales the data to a fixed range (e.g., 0 to 1) to eliminate bias under different experimental conditions and ensure comparability between samples. In single-cell Raman spectroscopy data, this normalization can reduce data differences caused by different measurement conditions, ensuring consistency in the data used for model training, thereby improving the model's generalization ability.

[0055] Signal-to-noise ratio filtering removes signals with low signal-to-noise ratio by evaluating the ratio of signal intensity to background noise, retaining high-quality data. In single-cell Raman spectroscopy data analysis, this filtering method can ensure that only high-quality spectra are used for model training, reducing the interference of low-quality data on microbial species identification results, thereby improving overall accuracy.

[0056] Therefore, in some preferred examples of the present application, the cosmic ray elimination, Savitzky-Golay smoothing, adaptive iteratively reweighted penalized least squares baseline correction, min-max normalization and signal-to-noise ratio filtering are selected for single-cell Raman spectroscopy data processing to ensure the accuracy of subsequent analysis.

[0057] To improve the smoothness and readability of the data, effectively eliminate the influence of noise and enhance the continuity of the data. The inventors further added step A-1) after step A) and before step B), which includes interpolating the single-cell Raman spectroscopy data after quality control processing using a smoothing spline function. By adding this step, the position and intensity of the characteristic peaks can be more accurately identified, the calculation of the area under the curve can be optimized, and the accuracy of subsequent model prediction can be improved.

[0058] In step B), the inventors use the composite Simpson rule to determine the area under the curve for each Raman shift within a predetermined range. The composite Simpson rule is a numerical integration method used to estimate the value of a definite integral. It improves the accuracy of integration by fitting the integrand function to a polynomial, usually a quadratic polynomial, within small intervals. The basic idea of this method is to divide the integration interval into several small intervals, and then use the function values at the endpoints and midpoints of each small interval to estimate the area under the curve. By using the composite Simpson rule, the inventors can accurately quantify the area of the characteristic peaks in the spectral data, improve the ratio of signal intensity to noise, and enhance the sensitivity and accuracy of the analysis. In addition, it can provide consistent results under different experimental conditions, reduce the influence of minor fluctuations in experimental conditions, and improve the accuracy of microbial species identification.

[0059] In some examples of the present application, the aforementioned predetermined range refers to ±N cm -1 ) for each Raman shift (expressed in wave number, unit: cm -1wherein N is a positive integer, such as 8, 9, 10, 11, 12, 13, 14, or 15, etc. For example, the predetermined range of a certain Raman shift is 600 cm -1 ± 10 cm -1 The area under the curve of the Raman shift within the predetermined range refers to the area under the curve of the Raman shift ranging from 590 cm -1 - 610 cm -1 .

[0060] The aforementioned term “Raman shift” refers to the difference between the frequency of scattered light and the frequency of incident light, usually expressed in wavenumber (cm -1 ).

[0061] S30, inputting the area under the curve of each Raman shift within the predetermined range into the trained regularization model to determine the species type of each microbial single cell.

[0062] The aforementioned term “regularization model” refers to a mathematical model in machine learning and statistics that introduces a regularization term into the loss function to limit the complexity of the model, thereby improving its generalization ability and prediction performance. The regularization term can be L1 regularization or L2 regularization, wherein L2 regularization is achieved by adding the sum of squares of model weights to the loss function, thus tending to make the weight values close to zero but not completely zero; L1 regularization is achieved by adding the sum of absolute values of model weights to the loss function, thus tending to produce a sparse weight matrix and automatically setting the weights of unimportant features to zero. Therefore, L1 regularization can be used for feature selection to select a small number of features that have a significant impact on the prediction result. In some examples of the present application, the inventors innovatively apply L1 regularization to the Raman spectrum-based microbial classification model, achieving the simultaneous integration of microbial species identification and Raman feature peak recognition. In some examples of the present application, the aforementioned regularization model is selected from at least one of LASSO regression model, elastic network, neural network, and support vector machine. In some preferred examples of the present application, the aforementioned regularization model is selected from LASSO regression model. By introducing the L1 regularization term into the loss function, the complexity of the model is reduced, and the prediction accuracy is improved.

[0063] In some examples of the present application, the trained regularization model is obtained by dividing the training set and the test set using uniform random sampling, taking the area under the curve as the feature, and taking the species type as the target to train the regularization model. K-fold cross-validation is used for model parameter selection during training, wherein K is any integer greater than 2; based on the accuracy, recall rate, and ROC curve for the test set, the trained regularization model is determined.

[0064] In the training process of the above regularization model, in order to ensure the same number of samples for each species and prevent the model from being biased due to the excessive number of samples of some microorganisms, the inventors use a downsampling method to balance the number of samples in the training set from different species.

[0065] In some examples of the present application, the aforementioned division of the training set and the test set can be adaptively selected based on the number of samples, such as dividing the training set and the test set in a 7:3 manner, or dividing the training set, the validation set and the test set in a 7:2:1 manner, etc.

[0066] It can be understood that different model training methods can have some differences, and those skilled in the art can select the model training method based on the selected model type.

[0067] In some examples of the present application, the area under the curve of each Raman shift within the predetermined range is input into the above trained regularization model, and the species type of each microorganism single cell can be obtained. In addition, the L1 regularization model can also obtain the non-zero weight information of the model, and based on the non-zero weight information of the L1 regularization model, the characteristic Raman shift of each microorganism, that is, the characteristic peak of each microorganism and its combination mode, is determined.

[0068] In some examples of the present application, the penalty coefficient λ in the aforementioned regularization model is selected from the λ value corresponding to the minimum mean square error or the λ value within one standard error range of the minimum mean square error. The aforementioned λ value selection improves the generalization ability and robustness of the model while considering the prediction accuracy.

[0069] The above method uses a regularization algorithm to analyze Raman spectra from different bacterial species, realizes the simultaneous integration of feature peak identification and accurate classification, and the model has interpretability. Specifically,

[0070] 1. Feature peak identification: by compressing the feature weight to zero except for the most important features, the model is simplified, the high dimensionality and multicollinearity problem is solved, the instability of ordinary least squares estimation in the conventional linear model is avoided, and the features with non-zero weights are a small number of key feature peak combinations that are helpful for species identification.

[0071] 2. Accurate classification: the best regularization parameter is selected through cross-validation to improve the generalization ability of the model, which can handle the case where the number of features is much larger than the number of samples, while deep learning and other methods often require a large sample size to avoid overfitting.

[0072] 3. Interpretability: directly output the weight of a small number of key feature peaks, and intuitively combine the intensities of multiple feature peaks to directly show which Raman shifts are processed in which way by the model, and finally output the species identification result.

[0073] Microorganism species identification device

[0074] In another aspect, the present application provides a microorganism species identification device, referring to Figure 2 The device comprises a single-cell Raman spectrum data acquisition module 100, a Raman shift curve area determination module 200 and a discrimination module 300. Wherein,

[0075] 100 module, for acquiring a plurality of single-cell Raman spectrum data of a plurality of microorganisms.

[0076] In some examples of the present application, the aforementioned microorganisms include fungi and / or bacteria. In some preferred examples of the present application, the aforementioned microorganisms are selected from bacteria.

[0077] In some examples of the present application, the plurality of single-cell Raman spectrum data of the plurality of microorganisms is obtained by culturing the plurality of microorganisms under suitable conditions and sample preparation processing; performing spectrum acquisition on the prepared sample, and determining the single-cell Raman spectrum data of the plurality of microorganisms.

[0078] In some examples of the present application, the spectrum acquisition parameters are selected from: excitation wavelength 532nm, 600g / mm grating, spectral range 400cm -1 -3600cm -1 , laser intensity 1mW-20mW, single spectrum acquisition time 1s-10s.

[0079] 200 module, for analyzing and processing the plurality of single-cell Raman spectrum data, and determining the curve area under each Raman shift in a predetermined range.

[0080] In some examples of the present application, the plurality of single-cell Raman spectrum data is analyzed and processed to determine the curve area under each Raman shift in a predetermined range, comprising: A) quality control processing of the single-cell Raman spectrum data; B) using the composite Simpson rule to analyze the single-cell Raman spectrum data after quality control processing, and determining the curve area under each Raman shift in a predetermined range. Wherein, the quality control processing is selected from at least one of cosmic ray elimination, Savitzky-Golay smoothing, adaptive iteratively reweighted penalized least squares baseline correction, minimum-maximum normalization and signal-to-noise ratio filtering.

[0081] In some preferred examples of the present application, the aforementioned quality control processing is selected from cosmic ray elimination, Savitzky-Golay smoothing, adaptive iteratively reweighted penalized least squares baseline correction, minimum-maximum normalization and signal-to-noise ratio filtering.

[0082] In some examples of the present application, after step A) and before step B), further comprising:

[0083] A-1) interpolating the single-cell Raman spectrum data after quality control processing by using a smoothing spline function.

[0084] 300 module for inputting the area under the curve of each Raman shift within a predetermined range into a trained regularization model to determine the species type of each microbial single cell.

[0085] In some examples of the present application, the aforementioned regularization model is selected from an L1 regularization model. Wherein the L1 regularization model is selected from at least one of a LASSO regression model, an elastic network, a neural network, and a support vector machine.

[0086] In some preferred examples of the present application, the L1 regularization model is selected from a LASSO regression model.

[0087] In some examples of the present application, the identification module further comprises a model training and characteristic Raman shift extraction submodule 301( Figure 3 ) for training the regularization model while determining the characteristic Raman shift of each microorganism.

[0088] In some examples of the present application, the trained regularization model is obtained by training in the following manner: dividing into a training set and a test set by using uniform random sampling, taking the area under the curve as the feature and the species type as the target, training the regularization model, and using K-fold cross-validation for model parameter selection during training, where K is any integer greater than or equal to 2; based on the accuracy, recall rate and ROC curve for the test set, the trained regularization model is determined.

[0089] In some examples of the present application, the characteristic Raman shift of each microorganism is determined based on the non-zero weight information of the L1 regularization model.

[0090] In some examples of the present application, the aforementioned modules are connected, and the aforementioned connection includes physical connection or network connection.

[0091] Based on the above device, the mobile compensation value of the mobile platform can be efficiently and automatically determined.

[0092] The device can realize efficient and accurate real-time non-destructive microbial species identification, and is particularly suitable for target microbial identification of complex samples (such as feces). In addition, the device also provides flexible spectral acquisition parameter settings, and improves data quality and reduces noise influence through quality control processing and composite Simpson rule. The L1 regularization model is used for species classification, which can effectively solve the problem of multiple collinearity, reduce the risk of overfitting, and improve the robustness and accuracy of the model. Through uniform random sampling and K-fold cross-validation and other optimization training methods, the feasibility of the model in practical application is ensured, and good industrial transformation potential is possessed, especially in the fields of biotechnology and medical detection.

[0093] It should be understood that the device embodiments and the microbial species identification method embodiments described above can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, it will not be described here. Specifically, Figure 2 、 Figure 3 The device shown in the device can perform the embodiments of the microbial species identification method described above, and the operations and / or functions performed by each module in the device correspond to the method embodiments. For the sake of brevity, it will not be described here.

[0094] The system of the embodiments of the present application is described above in conjunction with the functional modules. It should be understood that the functional modules can be realized by hardware, or by instructions in the form of software, or by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit of hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware code processing performed by the processor, or executed by a combination of hardware and software modules in the code processing processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiments.

[0095] Computer program product, computing device and computer readable storage medium

[0096] In still another aspect, the present application proposes a computer program product, a computing device and a computer readable storage medium. Based on the foregoing computer program product, computing device or computer readable storage medium, so that the aforementioned microbial species identification method is executed.

[0097] The present embodiment takes an electronic device as an example for detailed description.

[0098] The term electronic device is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computing devices may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0099] like Figure 4 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0100] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0101] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the microbial species identification method. For example, in some embodiments, the microbial species identification method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the aforementioned microbial species identification method by other any appropriate means, such as by means of firmware.

[0102] In this application, logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions for implementing the logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a product of the manufacturing and / or processing. The computer-readable medium can include, but is not limited to, the following: an electronic connection (an electronic device with one or more wires), a portable computer diskette (a magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CD ROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the program can be printed, because the program can be electronically captured, for example, by optically scanning the paper or other medium, then in electronic form, and then compiled, interpreted, or processed in a suitable manner, if necessary, and stored in a computer storage. The various computer-readable storage media described in the present disclosure can represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" can include, but is not limited to, a wireless channel and various other media capable of storing, containing, and / or carrying the instructions and / or data.

[0103] It should be understood that each of the elements of the present application can be implemented in hardware, software, firmware, or combinations thereof. In the above embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, any one or combination of the following technologies known in the art can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0104] Those of ordinary skill in the art can understand that all or part of the steps carried out by the above-mentioned embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, one or a combination of the steps of the method embodiments is included.

[0105] In addition, each of the functional units in various embodiments of the present application can be integrated in one processing module, or each unit can exist physically separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware, or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0106] It should be understood that various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0107] Embodiments of the present application will be described in more detail below, examples of which are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0108] Example 1 Discrimination of Bacteroides intestinalis and Bacteroides xylanisolvens

[0109] 1. Single-cell Raman spectrum data collection

[0110] Pure strains of Bacteroides intestinalis and Bacteroides xylanisolvens were inoculated into modified GAM (Gifu Anaerobic Medium) broth, and after 24 h of activation at 37°C in an anaerobic glove box, 1 mL of bacterial solution was centrifuged at 3000 rpm for 5 min, and the supernatant was discarded. Sterile water was added to resuspend the precipitate by blowing. Both strains were obtained by isolation and purification from fecal samples of healthy volunteers. 1 μL of resuspended bacterial solution was spotted on an aluminized Raman chip, and after room temperature natural air drying, it was put on the machine.

[0111] HOOKE PRECIS CS-R300 visual single-cell Raman detection and sorting instrument was used for Raman spectrum detection, the laser wavelength of the incident light source was 532 nm, a grating of 600 g / mm was used, and the Raman shift in the range of 400 cm -1 -3600 cm -1 -1 was collected, and about 500 cells of each species were collected.

[0112] The measured Raman spectra were quality controlled, including cosmic ray elimination, Savitzky-Golay smoothing, airPLS baseline correction, min-max normalization, signal-to-noise ratio filtering, interpolation of data points of the spectrum by one-dimensional smoothing spline fitting, numerical integration using composite Simpson's rule, calculation of the area under the curve of each Raman shift ±10 cm -1 -600 cm -1 -1800 cm -1 -1800 cm

[0113] Among them, the Raman spectrum results of Bacteroides intestinalis and Bacteroides xylanisolvens are shown in Figure 5

[0114] 2. Obtain the trained LASSO regression model

[0115] 70% of the samples were included in the training set, and 30% of the samples were included in the test set. The LASSO algorithm was used to train the classification model to distinguish different species. The number of cross-validation folds was 10. When selecting the model, the value of the penalty coefficient λ was taken to be within one standard error range to the right of the λ value of the minimum mean square error.

[0116] Among them, the classification performance results of the Bacteroides model are shown in Figure 6 The classification accuracy of this model reached 96.37%.

[0117] 3. Results analysis

[0118] The single-cell Raman spectrum data obtained in step 1 was input into the trained LASSO regression model obtained in step 2, which could achieve accurate identification of Bacteroides intestinalis and Bacteroides xylanisolvens. Further, as shown in Table 1, based on the model weight information, the characteristic Raman shift / characteristic peak combination of Bacteroides intestinalis and Bacteroides xylanisolvens could also be obtained. Among them, the peak corresponding to the characteristic Raman shift is the characteristic peak.

[0119] Table 1

[0120]

[0121]

[0122] ​Example 2: Discrimination of Streptococcus oralis and Streptococcus salivarius

[0123] 1. Single-cell Raman spectrum data acquisition

[0124] Pure strains of Streptococcus oralis and Streptococcus salivarius were inoculated into tryptone soya broth medium, and cultured at 37°C in an aerobic environment for 24 h. After centrifugation at 3000 rpm for 5 min, the supernatant was discarded, and sterile water was added to resuspend the precipitate by blowing. Both strains were isolated and purified from saliva samples of healthy volunteers. 1 μL of the resuspended bacterial solution was spotted on an aluminized Raman chip, and then naturally air-dried at room temperature before being loaded into the machine.

[0125] The HOOKE PRECIS CS-R300 visual single-cell Raman detector and sorter were used for Raman spectrum detection. The laser wavelength of the incident light source was 532 nm, a 600 g / mm grating was used, and the Raman shift was scanned in the range of 400 cm -1 -3600 cm -1 -1. The laser intensity was 3 mW, the acquisition time was 3 s, the spectrum of each cell was collected once, and about 500 cells of each species were collected.

[0126] The measured Raman spectrum was quality controlled, including eliminating cosmic rays, Savitzky-Golay smoothing, airPLS baseline correction, minimum-maximum normalization, signal-to-noise ratio filtering, interpolating the data points of the spectrum by one-dimensional smoothing spline fitting, using the composite Simpson rule for numerical integration, and calculating the area under the curve of each Raman shift ±10 cm -1 -1. The laser intensity was 3 mW, the acquisition time was 3 s, the spectrum of each cell was collected once, and about 500 cells of each species were collected. -1 -1800 cm -1 -1. The laser intensity was 3 mW, the acquisition time was 3 s, the spectrum of each cell was collected once, and about 500 cells of each species were collected.

[0127] The Raman spectrum results of Streptococcus oralis and Streptococcus salivarius are shown in Figure 7 .

[0128] 2. Obtain a trained LASSO regression model

[0129] 70% of the samples were included in the training set, and 30% of the samples were included in the test set. The LASSO algorithm was used to train the classification model to distinguish different species. The number of cross-validation folds was 10. When selecting the model, the value of the penalty coefficient λ was taken to be within one standard error range to the right of the λ value of the minimum mean square error.

[0130] The classification performance results of the streptococcus model are shown in Table 1, and the classification accuracy of the model is 85.27%. Figure 8

[0131] 3. Result analysis

[0132] The single-cell Raman spectrum data obtained in step 1 is input into the trained LASSO regression model obtained in step 2, so that accurate identification of Streptococcus oralis and Streptococcus salivarius can be achieved. Further, based on the model weight information, the characteristic Raman shift / characteristic peak combination of Streptococcus oralis and Streptococcus salivarius can also be obtained, as shown in Table 2. The peak corresponding to the characteristic Raman shift is the characteristic peak.

[0133] Table 2

[0134]

[0135]

[0136] Example 3: Identification of Lactobacillus crispatus and Lactobacillus jensenii

[0137] 1. Single-cell Raman spectrum data acquisition

[0138] Pure strains of Lactobacillus crispatus and Lactobacillus jensenii were taken and inoculated into MRS (De Man, Rogosa and Sharpe) broth, and after 24 h of incubation at 37°C in an anaerobic glove box, 1 mL of bacterial solution was centrifuged at 3000 rpm for 5 min, and the supernatant was discarded. Sterile water was added to resuspend the precipitate. Both strains were isolated and purified from vaginal samples of healthy volunteers. 1 μL of the resuspended bacterial solution was taken and spotted on an aluminized Raman chip, which was then air-dried at room temperature and loaded into the machine.

[0139] The HOOKE PRECIS CS-R300 visual single-cell Raman detection and sorting instrument was used for Raman spectrum detection, the laser wavelength of the incident light source was 532 nm, a 600 g / mm grating was used, and the scanning range was 400 cm -1 -3600 cm -1 ​Raman shifts within the range of 600-1800 cm-1, laser intensity of 3 mW, collection time of 3 s, one spectrum collected per cell, approximately 500 cells collected per species.

[0140] The measured Raman spectra were quality controlled, including cosmic ray elimination, Savitzky-Golay smoothing, airPLS baseline correction, min-max normalization, signal-to-noise ratio filtering, interpolation of data points of the spectrum by one-dimensional smoothing spline fitting, numerical integration using composite Simpson's rule, calculation of the area under the curve of each Raman shift ±10 cm-1 -1 of 600 cm-1 -1 -1800 cm-1 -1 as model input features.

[0141] The Raman spectral results of L. crispatus and L. jensenii are shown in Figure 9 .

[0142] 2. Obtain the trained LASSO regression model

[0143] 70% of the samples were included in the training set, 30% of the samples were included in the test set, the LASSO algorithm was used to train the classification model to distinguish different species, the number of cross-validation folds was 10, and when selecting the model, the value of the penalty coefficient λ was taken to be within one standard error range to the right of the λ value of the minimum mean square error.

[0144] The classification performance results of the Lactobacillus model are shown in Figure 10 , and the classification accuracy of the model reached 99.72%.

[0145] 3. Results analysis

[0146] The single-cell Raman spectral data obtained in step 1 were input into the trained LASSO regression model obtained in step 2, which could achieve accurate identification of Lactobacillus crispatus and Lactobacillus jensenii. Further, as shown in Table 3, based on the model weight information, the characteristic Raman shift / characteristic peak combination of Lactobacillus crispatus and Lactobacillus jensenii could also be obtained. Among them, the peak corresponding to the characteristic Raman shift is the characteristic peak.

[0147] Table 3

[0148]

[0149]

[0150] Example 4: Discrimination of Bifidobacterium longum and Bifidobacterium adolescentis

[0151] 1. Single-cell Raman spectrum data acquisition

[0152] Pure strains of Bifidobacterium longum and Bifidobacterium adolescentis were inoculated into MRS (De Man, Rogosa and Sharpe) broth, and incubated at 37°C in an anaerobic glove box for 24 h. After centrifugation at 3000 rpm for 5 min, the supernatant was discarded, and the precipitate was resuspended by blowing with sterile water. Both strains were isolated and purified from fecal samples of healthy volunteers. 1 μL of the resuspended bacterial solution was spotted on an aluminized Raman chip, and then dried naturally at room temperature before being loaded into the machine.

[0153] The HOOKE PRECIS CS-R300 visual single-cell Raman detector and sorter were used for Raman spectrum detection. The laser wavelength of the incident light source was 532 nm, a 600 g / mm grating was used, and the Raman shift was scanned in the range of 400 cm -1 - 3600 cm -1 -1. The laser intensity was 3 mW, the acquisition time was 3 s, the spectrum of each cell was collected once, and about 500 cells of each species were collected.

[0154] The measured Raman spectrum was quality controlled, including eliminating cosmic rays, Savitzky-Golay smoothing, airPLS baseline correction, minimum-maximum normalization, signal-to-noise ratio filtering, interpolating the data points of the spectrum by one-dimensional smoothing spline fitting, using the composite Simpson rule for numerical integration, and calculating the area under the curve of each Raman shift ± 10 cm -1 - 1800 cm -1 -1. The area under the curve of the Raman shift of 600 cm -1 was used as the feature when the model was input.

[0155] The Raman spectrum results of Bifidobacterium longum and Bifidobacterium adolescentis are shown in Figure 11 .

[0156] 2. Obtain a trained LASSO regression model

[0157] 70% of the samples were included in the training set, and 30% of the samples were included in the test set, a classification model for distinguishing different species was trained using the LASSO algorithm, the number of cross-validation was 10, and when the model was selected, the value of the penalty coefficient λ was taken as the value of λ within one standard error range to the right of the value of λ of the minimum mean square error.

[0158] The classification performance results of the Bifidobacterium model are shown in Table 4, and the classification accuracy of the model is 99.51%. Figure 12

[0159] 3. Result analysis

[0160] The single-cell Raman spectrum data obtained in step 1 is input into the trained LASSO regression model obtained in step 2, and accurate identification of Bifidobacterium longum and Bifidobacterium adolescentis can be achieved. Further, as shown in Table 4, based on the model weight information, the characteristic Raman shift / characteristic peak combination of Bifidobacterium longum and Bifidobacterium adolescentis can also be obtained. Among them, the peak corresponding to the characteristic Raman shift is the characteristic peak.

[0161] Table 4

[0162]

[0163] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0164] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments without departing from the principles and spirit of the present application within the scope of the present application.​

Claims

1. A method for identifying microbial species, characterized in that: include: Collect multiple single-cell Raman spectroscopy data of various microorganisms; Analyzing and processing the plurality of single-cell Raman spectral data to determine the area under the curve of each Raman shift within a predetermined range; The area under the curve of each Raman shift within a predetermined range is input into a trained regularized model to determine the species type of each microbial single cell.

2. The method according to claim 1, characterized in that The regularization model is selected from the L1 regularization model; Optionally, the L1 regularization model is selected from at least one of a LASSO regression model, an elastic network, a neural network, and a support vector machine; Preferably, the L1 regularization model is selected from the LASSO regression model; Optionally, the trained regularized model is obtained by training in the following manner: The dataset was divided into training and test sets by uniform random sampling. The area under the curve was used as the feature and species type was used as the target to train the regularized model. K-fold cross validation was used for model parameter selection during training, where K is any integer greater than 2. The trained regularized model is determined based on the precision, recall and ROC curve for the test set.

3. The method according to claim 2, characterized in that Further including: Based on the non-zero weight information of the L1 regularized model, the characteristic Raman shift of each microorganism is determined.

4. The method according to claim 1, wherein Analyzing and processing the plurality of single-cell Raman spectral data to determine the area under the curve of each Raman shift within a predetermined range includes: A) performing quality control processing on the single-cell Raman spectroscopy data; B) using a composite Simpson's rule to analyze the single-cell Raman spectral data after the quality control process, and determining the area under the curve of each Raman shift within a predetermined range.

5. The method according to claim 4, characterized in that The quality control processing is selected from at least one of cosmic ray elimination, Savitzky-Golay smoothing, adaptive iterative reweighted penalty least squares baseline correction, minimum-maximum normalization and signal-to-noise ratio filtering; Preferably, the quality control processing is selected from cosmic ray removal, Savitzky-Golay smoothing, adaptive iterative reweighted penalized least squares baseline correction, minimum-maximum normalization and signal-to-noise ratio filtering.

6. The method according to claim 4, characterized in that After step A) and before step B), the method further comprises: A-1) Smoothing spline function is used to interpolate single-cell Raman spectroscopy data after quality control processing.

7. The method according to claim 1, characterized in that The penalty coefficient λ in the regularization model is selected from the λ value corresponding to the minimum mean square error or the λ value of the minimum mean square error plus the λ value within the range of one standard error.

8. The method according to any one of claims 1 to 7, characterized in that The microorganisms include: fungi and / or bacteria; Optionally, the plurality of single-cell Raman spectral data of the plurality of microorganisms are obtained by: Cultivating the plurality of microorganisms and performing sample preparation processing on an in-machine process under suitable conditions; Collecting spectra of the prepared samples to determine single-cell Raman spectral data of the multiple microorganisms; Optionally, the spectrum acquisition parameters are selected from: excitation wavelength 532nm, 600g / mm grating, spectral range 400cm -1 -3600cm -1 , laser intensity 1mW-20mW, single spectrum acquisition time 1s-10s.

9. A microbial species identification device, characterized in that: include: Single-cell Raman spectroscopy data acquisition module, used to collect multiple single-cell Raman spectroscopy data of various microorganisms; a Raman shift curve area determination module, configured to analyze and process the plurality of single-cell Raman spectral data to determine the area under the curve of each Raman shift within a predetermined range; an identification module, configured to input the area under the curve of each Raman shift within a predetermined range into a trained regularization model to determine the species type of each microbial single cell; Optionally, the identification module further comprises: a model training and characteristic Raman shift extraction submodule for training the regularization model and determining the characteristic Raman shift of each microorganism.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions or programs, and when the computer instructions or programs are executed on a computer, the microbial species identification method according to any one of claims 1 to 8 is executed.