A method of raman identification or sorting of unknown microbial strains
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO SINGLE CELL BIOTECH CO LTD
- Filing Date
- 2025-03-28
- Publication Date
- 2026-07-21
AI Technical Summary
Existing microbial screening technologies suffer from low screening efficiency, significant damage to cell viability, limited throughput, and difficulty in identifying novel strains. In particular, existing methods have limitations when isolating rare and novel strains in complex environments.
By combining Raman spectroscopy with machine learning algorithms, a Raman spectral library of strains is established. By real-time analysis of cell spectra acquired by flow Raman cytometry, unknown microbial species, especially low-abundance strains, are identified and sorted, simplifying the operation process and improving the effectiveness and accuracy of screening.
This technology enables the efficient and non-destructive separation of rare and novel bacterial species from complex environments, improves the accuracy of identification and sorting, expands the application range of Raman spectroscopy, reduces operating costs, and enhances cell viability.
Smart Images

Figure CN120319319B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microbial sorting technology, and in particular relates to a method for Raman identification or sorting of unknown microbial species. Background Technology
[0002] The natural world contains a vast array of microbial species, with humans having discovered only a tiny fraction of all microorganisms. This diversity means that many undiscovered microorganisms may possess unique biological characteristics and potential applications. In laboratory settings or production processes, long-term cultured microbial strains may degenerate, gradually weakening their survival capabilities and adaptability to growth conditions. In contrast, wild microorganisms in nature, through natural selection, possess stronger survival abilities and adaptability, making them more suitable for mutagenesis and cultivation under artificial conditions. Therefore, directly collecting samples from the target area for microbial isolation helps identify highly efficient strains best suited to the local environment. Furthermore, factors such as climate change influence ecosystem evolution; regularly re-isolating microbial strains from the natural environment helps researchers stay abreast of the latest developments and develop products adapted to current conditions.
[0003] Existing strain screening methods include dilution culture, selective culture, microfluidic screening, fluorescence flow cytometry sorting, and optical tweezers extraction. Conventional culture methods are generally limited by manual operation, resulting in low screening efficiency and the potential loss or missed screening of some strains during the screening process. Microfluidic screening requires the use of an oil phase and may be limited by throughput. Fluorescence flow cytometry sorting has requirements on cell size, and antibody dyes and excessive fluid pressure can damage cell viability. Optical tweezers technology is also limited by cell morphology and sorting efficiency during the separation process.
[0004] Raman spectroscopy, with its non-destructive and label-free characteristics, provides an efficient and precise analytical tool for scientific research and industrial applications. It can provide in-situ "molecular fingerprint" information of bacterial colonies, making it suitable for precise identification and classification of microorganisms and possessing great potential for screening new species in the environment. In microbiological research, using machine learning to process and analyze Raman spectral data can extract characteristic information for classifying and identifying microorganisms, thereby achieving rapid and accurate identification. However, in practical applications, machine learning still has certain limitations when processing diverse spectral signals, which can interfere with the accuracy of identification. Therefore, a new high-throughput Raman identification or sorting method for microbial species is proposed, which is of great significance for solving scientific problems and practical applications such as isolating rare and novel species from complex environments. Summary of the Invention
[0005] This invention provides a method for Raman identification or sorting of unknown microbial species. This method simplifies the process of strain identification or sorting, effectively improves the effectiveness and accuracy of screening low-abundance strains, expands the application scope of Raman spectroscopy, and is of great significance for solving scientific problems and practical applications such as isolating rare and novel strains from complex environments.
[0006] To achieve the above objectives, this invention provides a method for Raman identification or sorting of unknown microbial species. This method involves identifying the Raman spectral phenotypic characteristics of known species A and unknown non-A species, and distinguishing the spectral signals of species A and non-A species based on an algorithm model, thereby identifying the unknown non-A species; or
[0007] Non-A species are collected and sorted for culture to quickly identify unknown non-A species.
[0008] As a preferred method, Raman spectral phenotypic feature identification of known species A and unknown non-A species specifically involves:
[0009] For bacterial suspensions of N strains within the known species A, multiple Raman spectra of A1, A2...AN are collected to establish a Raman spectrum set {A1, A2...AN} for bacterial suspensions of various known strains;
[0010] Prepare simulated bacterial suspension samples containing strains within the range of species A and strains outside the range of species A. Wash the samples to obtain mixed bacterial cell suspensions. Then, establish a set of Raman spectra of the bacterial suspensions containing the strains in the mixed bacterial cell suspensions.
[0011] As a preferred method, multiple Raman spectra of A1, A2...AN are collected to establish a Raman spectrum set {A1, A2...AN} for various known strains. Specifically:
[0012] After washing the collected bacterial culture with sterile water, the cells were resuspended in buffer solution, with a final concentration preferably 1×10⁻⁶. 3 ~1×10 6 Within the cell / mL range, the detection carrier was then introduced separately, and its Raman spectrum was collected.
[0013] Preferably, the collected bacterial solution is selected from any one of frozen bacterial solution, fermented bacterial solution, or cell suspension extracted from the environment; it is understood that the present invention does not specifically limit the type of bacterial solution collected, and the above are only common examples.
[0014] As a preferred method, the acquisition conditions are: 532 laser energy 5-100mw, 0.5-2s; 50× lens.
[0015] As a preferred method, by distinguishing the spectral signals of species A and non-A, non-A species are collected and sorted and cultured based on an algorithm model, specifically as follows:
[0016] In the sorting mode with the algorithm model loaded, the spectral signal of the strain is determined in real time online to determine whether it is species A. If it is not species A, it is identified, and downstream docking counting is performed to identify non-species A in the environment; or
[0017] If the species is not A, it is sorted, followed by downstream docking culture and / or sequencing to sort it into non-A species in the environment.
[0018] As a preferred approach, the algorithm model is designed based on the prediction network for species A and non-A. Specifically, for the prediction of species A itself, an encoding-then-decoding method is used to train the network to understand A. Then, an external parameter factor is forcibly introduced to define some species A as non-A species. Based on the threshold of the algorithm model, the accuracy and actual recall of species A and non-A are calculated. By comparing the actual recall with the set threshold, the samples of species A in the collected spectrum that belong to the set {A1, A2...AN} and samples of non-A species that do not belong to the set are distinguished.
[0019] Preferably, external parameter factors include temperature control and reverse disturbance.
[0020] As a preferred option, the threshold is determined based on a preset recall rate setting.
[0021] Preferably, while determining that the strain belongs to species set A {A1, A2...AN}, the exact category of the strain belonging to said set is precisely distinguished.
[0022] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0023] This invention relies on high-throughput flow Raman cytometry combined with machine learning algorithms, making operation simpler. Compared to traditional strain screening methods, the identification or sorting method provided by this invention first establishes a Raman spectral model of the strain, then uses machine learning to identify the strain, and judges the cell spectra acquired by flow Raman cytometry in real time. Based on the identification results, non-A bacteria are identified or sorted from the mixed sample, especially for low-abundance strains, ensuring the effectiveness and accuracy of screening. This method does not require pre-labeling of samples, simplifying the operation process, saving costs, and improving cell viability compared to fluorescence flow cytometry. The flow Raman cytometry technology used in this method does not require oil phase encapsulation, increasing detection throughput compared to microfluidic technology. Furthermore, the established Raman spectral library can be reused in similar work, making it widely applicable.
[0024] More importantly, this method combines AI algorithms with Raman spectroscopy, avoiding interference from factors such as fluorescence background, noise, and baseline drift that often affect Raman spectral signals. Existing algorithms can only classify and identify specific types of spectral signals (referred to as spectral A), but they typically assign a known label to data the model has never seen before, failing to explore new species or variant bacteria, and thus failing to meet the needs of all practical applications. This invention can simultaneously and efficiently distinguish between spectral A and non-spectral A signals in Raman spectroscopy. This will help improve the accuracy and reliability of spectral data analysis and expand the application scope of Raman spectroscopy. Attached Figure Description
[0025] Figure 1 Neural network structure diagrams for species A and non-species A provided in embodiments of the present invention;
[0026] Figure 2 This provides specific data on different bacterial species identified for known and unknown samples in the embodiments of the present invention. Detailed Implementation
[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Example 1: Method for Identifying or Sorting Non-A Species
[0029] All known species are collectively referred to as species A. The method provided in this embodiment uses an algorithm to identify Raman spectral phenotypic features, efficiently distinguishing between spectral A and non-spectral A signals from complex samples, and sorting and culturing non-A species, especially low-abundance non-A species, thereby rapidly improving the efficiency of screening new species. The technical solution is as follows:
[0030] S1: Collect multiple Raman spectra of N strains of known species A, A1, A2...AN, and establish a Raman spectrum set {A1, A2...AN} for multiple known strains;
[0031] Specifically:
[0032] After washing the bacterial cultures of the N strains to be collected with sterile water, the final cell concentration was prepared to 1×10⁻⁶ using a buffer solution. 3 ~1×10 6 Within the cell / mL range, a quartz flow Raman chip was then introduced to collect multiple Raman spectra of N kinds of bacteria;
[0033] The collected bacterial suspensions were selected from any one of the following: frozen bacterial suspensions, fermented bacterial suspensions, and cell suspensions extracted from the environment.
[0034] The acquisition conditions were: 532 laser energy 5-100mw, 0.5-2s; 50× lens.
[0035] S2: Prepare a simulated bacterial suspension sample containing strains within the range of species A and strains outside the range of species A. Centrifuge and wash the sample to obtain a mixed bacterial cell suspension.
[0036] S3: Flow Raman acquisition. Mixed bacterial cell suspension is used as sample X for Raman spectroscopy acquisition. The algorithm model is used to distinguish the samples in the X-class spectrum that belong to the set {A1, A2...AN} from the open set that do not belong to the set. While determining whether a sample belongs to the set {A1, A2...AN}, the exact category of the set is also accurately distinguished.
[0037] The acquisition conditions were: 532 laser energy 5-100mw, 0.5-2s; 50× lens.
[0038] S4: Sorting. In the sorting mode with the algorithm model loaded, the system determines in real time whether the spectrum belongs to species A. If it is a species A strain, the A strain flows into the waste liquid outlet. If it is a non-species A strain, the A strain flows into the collection outlet for docking culture and functional verification to identify or sort non-species A in the environment.
[0039] Example 2: Algorithm Modeling Method
[0040] The algorithm model provided in this embodiment is based on the prediction network design for species A and non-species A, specifically as follows: Figure 1 As shown, the method of encoding and then decoding is used to train the network to understand A for the prediction of species A itself; then, external parameter factors are forcibly introduced to define some species A as non-A species, so that the feature set of species A can be defined relatively concentratedly.
[0041] It is understandable that by introducing external parameter factors, the original softmax score output by the network can be transformed from meaningless and unable to represent class confidence into a value that can reflect class characteristics and confidence. Thus, through uncertainty estimation and robustness enhancement, the model can more accurately distinguish between the training data distribution (In-Distribution, ID) and the non-training data distribution (OOD).
[0042] Temperature control and reverse disturbance are the introduced external parameter factors. The role of temperature control is reflected in the calculation of the softmax fraction, as shown in the following formula:
[0043]
[0044] Example: Suppose we are dealing with a three-class classification problem, and the model output is a 3-dimensional vector: [1, 2, 3]. We get the result: [0.09003057317038046, 0.24472847105479767, 0.6652409557748219].
[0045] Then we set t=2, which is essentially calculating the softmax using [1 / 2, 2 / 2, 3 / 2], and get the result: [0.1863237232258476, 0.30719588571849843, 0.506480391055654]. Next, we set t=0.5 and calculate the softmax using [1 / 0.5, 2 / 0.5, 3 / 0.5], and get the result: [0.015876239976466765, 0.11731042782619835, 0.8668133321973348]. When using this for OOD detection, we can obtain a suitable softmax score that can differentiate between OOD and ID.
[0046] The effect of the reverse perturbation is as follows:
[0047]
[0048] This perturbation is the gradient of the output with respect to x, obtained by adding a softmax operation to the output. ∈ is a hyperparameter used to adjust the degree of perturbation. Adding a small perturbation to the samples reduces the softmax score of the correct class label in the model output. The goal here is to increase the softmax score for any given input, regardless of the sample's class label. Adding the perturbation has a stronger effect on ID samples than on OOD samples, thus widening the gap in softmax scores between ID and OOD.
[0049] After adding the above two methods, for the input sample, the softmax score is calculated, the maximum value is taken as the confidence score, and then compared with the given threshold.
[0050]
[0051] When the confidence level is greater than the threshold, it can be considered an ID sample; otherwise, it is an OOD sample.
[0052] Example 3: Specific Case
[0053] Select Escherichia coli, Staphylococcus aureus, Listeria monocytogenes, Salmonella, Shigella, probiotic DH4, and probiotic DH8 bacterial suspensions stored at -80℃, streak them on LB agar plates, and incubate them overnight in an inverted 37℃ incubator.
[0054] Single clones were picked from the plate and inoculated into test tubes containing 2 mL of LB liquid medium. The culture was activated overnight at 37°C and 200 rpm.
[0055] Take 0.5 mL of activated bacterial solution, centrifuge at 3000 rpm and 4℃ for 5 min, discard the supernatant, wash twice with sterile water, and resuspend in 1 mL of sterile water; according to the cell concentration, take a certain amount of bacterial solution and add 1.5 mL of loading buffer, and collect Raman spectroscopy data under the following conditions: 532 laser energy 60 mw, 1 s; 50× lens.
[0056] Ultimately, under the same collection conditions, 1930 Escherichia coli, 2068 Staphylococcus aureus, 1646 Listeria monocytogenes, 1829 Salmonella, 1652 Shigella, 6409 DH4 probiotics, and 6546 DH8 probiotics were obtained using Raman spectra.
[0057] We know that the dataset contains *E. coli*, *Staphylococcus aureus*, *Listeria monocytogenes*, *Salmonella*, and *Shigella*, corresponding to categories 0, 1, 2, 3, and 4, respectively. We divide the dataset into training and test sets in a 9:1 ratio. We only use the training set for modeling and introduce two unknown classes, DH4 and DH8, into the test set for testing. The purpose is to identify the specific number of samples in the known and unknown classes. The threshold setting can be varied depending on the situation. Generally, the threshold can be understood as a preset recall rate for the training data; that is, for the modeling data, the threshold setting should result in the same recall rate for its own inference, thus applying this value to the data to be detected (inference). In this embodiment, when we set the threshold to 0.80 (i.e., the preset recall rate), the result is as follows... Figure 2 The table shows the precision and total data precision (i.e., recall) for each class. cls:0 precision: 0.9970; cls:1 precision: 0.9983; cls:2 precision: 0.9987; cls:3 precision: 0.9993; cls:4 precision: 0.9983.
[0058] Precision for known classes: 1.0, Recall for known classes: 0.7948; Precision for unknown classes: 0.8737, Recall for unknown classes: 1.0. It can be seen that the recall for known classes is close to the threshold; therefore, *E. coli*, *Staphylococcus aureus*, *Listeria monocytogenes*, *Salmonella*, and *Shigella* can be almost accurately identified with no false positives. Furthermore, the unknown classes DH4 and DH8 are identified relatively accurately, having almost no impact on the known classes.
Claims
1. A method for Raman identification or sorting of unknown microbial species, characterized in that, Raman spectral phenotypic features are used to identify known species A and unknown non-A species. Based on the algorithm model, the spectral signals of species A and non-A species are distinguished, thereby identifying unknown non-A species; or non-A species are collected and sorted for cultivation to quickly identify unknown non-A species. The specific steps for Raman spectral phenotypic feature identification of known species A and unknown non-A species are as follows: For bacterial suspensions of N strains within the known species A, multiple Raman spectra of A1, A2...AN are collected to establish a Raman spectrum set {A1, A2...AN} for bacterial suspensions of various known strains; Prepare simulated bacterial suspension samples containing strains within species A and strains outside species A. Wash the samples to obtain mixed bacterial cell suspensions. Then, establish a set of Raman spectra of the bacterial suspensions containing the strains in the mixed bacterial cell suspensions. The algorithm model is based on the prediction network design for species A and non-A. Specifically, for the prediction of species A itself, an encoding and decoding method is used to train the network to understand A. Then, an external parameter factor is forcibly introduced to define some species A as non-A species. The accuracy and actual recall of species A and non-A species are calculated based on the threshold of the algorithm model. By comparing the actual recall with the set threshold, the samples of species A belonging to the set {A1, A2...AN} in the collected spectrum and samples of non-A species that do not belong to the set are distinguished. External parameter factors include temperature control and reverse disturbance.
2. The method according to claim 1, characterized in that, Multiple Raman spectra of A1, A2...AN were collected to establish a Raman spectrum set {A1, A2...AN} for various known strains. After washing the collected bacterial culture with sterile water, the cells were resuspended in buffer solution, with a final concentration preferably 1×10⁻⁶. 3 ~1×10 6 Within the cell / mL range, the detection carrier was then introduced separately, and its Raman spectrum was collected.
3. The method according to claim 2, characterized in that, The collected bacterial solutions were selected from frozen bacterial solutions, fermentation bacterial solutions, and environmental sources. Any of the cell suspensions extracted from it.
4. The method according to claim 1, characterized in that, By distinguishing the spectral signals of species A and non-A, the non-A species are collected and sorted and cultured based on an algorithm model, specifically as follows: In the sorting mode with the algorithm model loaded, the spectral signal of the strain is determined in real time to determine whether it is species A. If it is not species A, it is identified and counted downstream to identify non-A species in the environment; or if it is not species A, it is sorted, cultured downstream and / or sequenced to sort non-A species in the environment.
5. The method according to claim 4, characterized in that, While determining that a strain belongs to species set A {A1, A2 ... AN}, the exact category of the strains belonging to said set is precisely distinguished.
6. The method according to claim 1, characterized in that, The threshold is determined based on the preset recall rate setting.