Method, device, medium and equipment for identifying single microbial cell species

By combining single-cell Raman spectral data with image data to construct multimodal features, the problems of poor data integrity and low identification accuracy in the prior art are solved, and rapid and accurate identification of microbial species are achieved.

CN114660040BActive Publication Date: 2025-05-27QINGDAO INST OF BIOENERGY & BIOPROCESS TECH CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210240203.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-05-27
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

Existing single-cell Raman detection technology cannot effectively combine single-cell images and spectral data, resulting in poor data integrity and low identification accuracy, especially in the identification of difficult-to-cultivate microorganisms such as Helicobacter pylori.

Method used

By comparing and analyzing the collected single-cell Raman spectral data with the data in the reference map database, the Raman spectral data that meets the conditions were screened out, and combined with the real-time collected single-cell image data, a single-cell phenolic database was constructed, and multimodal feature fusion was performed to achieve cell species identification.

Benefits of technology

It improves the data integrity and identification accuracy of single-cell species identification, and can more quickly and accurately identify microbial species, especially difficult-to-cultivate pathogens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114660040B_ABST
    Figure CN114660040B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device, medium and equipment for identifying single microbial cell species. The method includes comparing and analyzing the collected single-cell Raman spectroscopy data with the Raman spectroscopy data in the reference spectrum database to screen out the qualified Raman spectroscopy data; using the screened Raman spectroscopy data as samples, calculating based on specific spectral characteristic values in the spectral data samples to obtain the minimum sample spectral detection quantity of the samples; collecting the spectral data corresponding to the spectral detection quantity, and standardizing the spectral data through a calibration transfer model; storing the standardized spectral data and the single-cell image data collected in real time into an omics database; performing multi-modal feature fusion on the characteristic values of the images and spectra based on the cell images and spectral data in the single-cell phenomics database to realize the identification of the species of single-cell phenotype data, increasing the integrity of the data, thereby improving the accuracy of single-cell species identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microorganism detection, and particularly relates to a method, device, medium and equipment for identifying the types of single microorganism cells. Background Art

[0002] Clinically, the traditional identification of pathogenic bacteria mainly relies on the culture method. This method has the disadvantages of long detection time and the need to achieve a pure culture degree without other microorganisms before systematic identification can be carried out. Systematic identification is to detect the morphological structure, growth characteristics, antigenicity and pathogenicity of pathogenic bacteria, and use known standard immune sera to determine the genus, species and type of the isolated bacteria. The procedure for microorganism identification usually determines the species based on its morphology, growth, biochemical characteristics, etc., and finally determines the type based on the immunoserological examination of antigens. It generally takes 14 - 40 hours to obtain the identification result, and even longer for difficult-to-culture bacteria. In addition, pure culture strains can directly identify the types of microorganisms by mass spectrometry or by DNA amplification and sequencing methods. These methods that first culture and then identify usually take one to two days to obtain the identification result. Although the results are controllable, they have disadvantages such as long time consumption, high cost or high requirements for operators.

[0003] Currently, the existing "single-cell Raman" detection technology, that is, skipping cell culture and proliferation, directly characterizes the "growth" or "metabolism" phenotypes of the original single cells in the sample with single-cell precision, and achieves the goals of rapidity, phenotype-based, and wide applicability in principle. Raman spectroscopy is an efficient information recognition technology. By analyzing the inelastic scattering spectrum lines of specific incident light on compounds, Raman microspectroscopy can directly detect the vibrational or rotational energy levels of compound molecules. By analyzing the Raman characteristic spectrum lines, information on the molecular composition and structure of compounds can be obtained. However, for the identification of the types of pathogenic bacteria, especially difficult-to-culture pathogenic bacteria such as Helicobacter pylori, their culture time is long and the bacterial quantity is small. Therefore, rapid detection of types is required at the single-cell scale.

[0004] Raman spectroscopy is an efficient information recognition technology. By analyzing the inelastic scattering spectrum lines of specific incident light on compounds, Raman microspectroscopy can directly detect the vibrational or rotational energy levels of compound molecules. By analyzing the Raman characteristic spectrum lines, information on the molecular composition and structure of compounds can be obtained. However, the existing methods for detecting single-cell samples using Raman technology have problems such as the inability to combine single-cell images and spectral data, poor data integrity, and low identification accuracy. Summary of the Invention

[0005] Aiming at the above problems, the purpose of the present invention is to provide a method, device, medium and equipment for identifying the types of single microorganism cells, which can combine single-cell images and spectral data to form multimodal features, increase the data integrity, and thus improve the identification accuracy of single-cell types.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for identifying single microbial cell species, the method comprising:

[0008] Comparing and analyzing the collected single-cell Raman spectroscopy data with the Raman spectroscopy data in the reference spectrum database, and screening out the Raman spectroscopy data that meet the conditions;

[0009] Taking the screened Raman spectroscopy data as samples, calculating according to specific spectral characteristic values in the samples, and obtaining the minimum sample spectral detection quantity that is stable under the specific spectral characteristic values;

[0010] Collecting the spectral data corresponding to the spectral detection quantity, and standardizing the spectral data through a calibration transfer model to obtain standardized spectral data;

[0011] Constructing a single-cell phenomics database based on the standardized spectral data and the single-cell image data collected in real time;

[0012] Performing multi-modal feature fusion on the eigenvalues of the images and spectra based on the cell images and spectral data in the single-cell phenomics database;

[0013] Classifying the data after multi-modal feature fusion to obtain the cell species, and realizing the identification of the species of single-cell phenotype data.

[0014] Preferably, comparing and analyzing the collected single-cell Raman spectroscopy data with the Raman spectroscopy data in the reference spectrum database, and screening out the Raman spectroscopy data that meet the conditions, including:

[0015] Constructing a reference spectrum database, and storing the screened spectra in the reference database;

[0016] Using the CNN algorithm to compare and analyze the collected spectral data with the data in the reference spectrum database, and screening out the spectra with high similarity.

[0017] Preferably, using the CNN algorithm to compare and analyze the collected spectral data with the data in the reference spectrum database, and screening out the spectra with high similarity, including:

[0018] Taking the collected Raman spectroscopy data as test data, inputting it into the reference spectrum database, and outputting an N-dimensional output vector corresponding to N species through calculation, where N is a natural number;

[0019] Mapping a vector as input to the Softmax function, for a specific test data, the maximum probability value of the Softmax output is P, the mean of the maximum values of the Softmax function for all data of the same category in the test data is M, and the variance is S. If M - S / 2 ≤ P ≤ M + S / 2, then the Raman spectral data is the Raman spectral data screened to meet the conditions.

[0020] Preferably, collecting spectral data corresponding to the number of spectral detections, standardizing the spectral data through a calibration transfer model to obtain standardized spectral data, and adopting a piecewise direct standardization PDS algorithm, including:

[0021] Dividing the spectral data into a target set spectrum and an adjustment set spectrum;

[0022] Selecting a certain wavenumber as the center, expanding left and right according to a set range as a window, constructing a multiple regression model with the intensity value of the i-th wavenumber of the target set spectrum and the window matrix centered on i of the adjustment set spectrum, where i is a natural number; solving through partial least squares regression, placing the regression coefficients in the regression model on the main diagonal of the transformation matrix, and setting other elements to 0 to obtain a transformation matrix;

[0023] Transforming the collected spectral data into standardized spectral data through the transformation matrix.

[0024] Preferably, classifying the data after multi-modal feature fusion using a CNN classifier to obtain cell types.

[0025] Preferably, using the ReliefF algorithm to determine the weights of the eigenvalue of the image and the spectrum for fusion operation to form multi-modal features.

[0026] A device for identifying single-cell species of microorganisms, including:

[0027] A Raman spectral map screening module configured to compare and analyze the collected single-cell Raman spectral data with the Raman spectral data in the reference map database, and screen out the qualified Raman spectral data;

[0028] An analysis of the number of spectral detections module configured to use the screened Raman spectral data as a sample, calculate according to specific spectral eigenvalue in the spectral data sample, and obtain the minimum number of spectral detections for the sample to be stable under the specific spectral eigenvalue;

[0029] A spectral data standardization module configured to collect spectral data corresponding to the number of spectral detections, standardize the spectral data through a calibration transfer model to obtain standardized spectral data;

[0030] Build a general single-cell phenomics database module, which is configured to build a single-cell phenomics database from standardized spectral data and real-time collected single-cell image data;

[0031] A multimodal feature fusion module, which is configured to perform multimodal feature fusion on the eigenvalues of images and spectra based on the cell images and spectral data in the single-cell phenomics database;

[0032] A classification module, which is configured to classify the multimodal features through a CNN classifier to obtain cell types, and realize the identification of the types of single-cell phenotype data.

[0033] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for identifying single-cell species of microorganisms are implemented.

[0034] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method for identifying single-cell species of microorganisms as claimed are implemented.

[0035] Due to the above technical solutions adopted by the present invention, it has the following advantages:

[0036] The method of the present invention combines a single cell image and spectral data to form multimodal features, which increases the integrity of the data, thereby improving the accuracy of single-cell species identification. Description of the Drawings

[0037] Figure 1 It is a flowchart of the identification method provided by an embodiment of the present invention. Detailed Embodiments

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0040] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "assembly", "setting", and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0041] The method, device, medium and equipment for identifying single microbial cell species provided by the present invention combine a single cell image and spectral data to form multimodal features, improving the integrity and identification accuracy of the data.

[0042] Next, the method, device, medium and equipment for identifying single microbial cell species provided by the embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0043] As Figure 1 shown, the method for identifying single microbial cell species provided in this embodiment includes the following steps:

[0044] Step 101, screening single-cell Raman spectra.

[0045] The selected Raman spectra are stored in the reference spectrum database, and the Raman spectral data of the collected single microbial cells are compared and analyzed with the Raman spectral data in the reference spectrum database to screen out the Raman spectral data that meet the conditions;

[0046] Specifically, during the collection process of Raman spectra, a convolutional neural network CNN (Convolutional neural network) is used to intelligently screen the spectra. The steps include:

[0047] Construct a reference spectrum database and store the screened spectra in the reference database.

[0048] Use the CNN algorithm to compare the collected spectral data with the data in the reference spectrum database and analyze and screen out the spectra with high similarity.

[0049] Specifically, the raw and unprocessed spectral data collected initially are used as the network input. Through the calculations of the convolutional layer, pooling layer, and fully connected layer, an N-dimensional output vector corresponding to N species will be output at the fully connected layer. Mapping the vector as the input to the Softmax function, N probabilities that the spectrum belongs to N different categories can be obtained. The value with the maximum probability is the predicted category of the spectrum. The higher the spectrum quality, the better the classification effect, and the corresponding probability value P is larger. When setting a threshold T for the maximum probability, if the maximum value of the probability calculated by the Softmax for the spectral data is greater than the threshold, it is considered that the quality of this spectrum meets the quality control requirements.

[0050] Specifically, the Raman spectral data of the collected microbial single cells are used as test data and input into the reference spectrum database. After calculation, an N-dimensional output vector corresponding to N species is output, where N is a natural number.

[0051] Mapping the vector as the input to the Softmax function, the probability of the i-th category is output by the Softmax for a specific test data, where i is a natural number. And the maximum probability is selected from the output probabilities. The maximum probability value is P, and the mean of the maximum values of the Softmax functions of all data of the same category in the test data is M, and the variance is S. If M - S / 2 ≤ P ≤ M + S / 2, then the Raman spectral data are the qualified Raman spectral data screened out.

[0052] The Softmax function is shown in formula (1):

[0053]

[0054] where i is the i-th category, C represents the total number of categories, z i is the output of the fully connected layer for the i-th category, and Softmax(z i ) is the probability value of the test data in the i-th category, and e is the natural base.

[0055] Step 102: Analyze the spectral detection quantity.

[0056] Taking the screened Raman spectral data as samples, calculations are performed according to specific spectral feature values in the spectral data samples to obtain the minimum sample spectral detection quantity that is stable under the specific spectral feature values of the samples.

[0057] Specifically, the spectral detection quantity is calculated according to specific spectral feature values in the spectral data samples. In order to provide real-time feedback on the change of feature values with the sampling volume during the spectral acquisition process, a real-time sample quantity analysis method is designed and constructed. The sample quantity is calculated only for a specific spectral feature value, representing the minimum number of samples that are stable under a certain feature value of the sample. Real-time sample quantity analysis is to calculate and update the sample quantity in real time according to the existing spectra during the spectral acquisition process.

[0058] Specifically, taking the calculation of CDR (CD-Ratio) as an example, for the measurement of single-cell Raman spectroscopy, an initial data set Xn is randomly obtained from 30 measurement spectra out of 1000, and then the difference between the average CDR and CDRn is calculated. In 1000 statistics, the nth CDR is CDRn, where n is an integer with a value range from 1 to 1000, and the probability P that the relative error between CDRn and the CDR population is less than 5%. P is the probability that the sample eigenvalue tends to be stable when the sample size is 30. When P > 95%, it is reliable; when P is less than 95%, P is recalculated for the sample until P is greater than 95%.

[0059] Step 103, spectral data standardization.

[0060] Collect spectral data corresponding to the number of spectral detections, and standardize the spectral data through a calibration transfer model to obtain standardized spectral data, which is used to eliminate the signal difference changes caused by changes in detection conditions or sample environments;

[0061] Specifically, in the case of changes in detection conditions or sample environments, such as changes in sample detection environments, changes in sample morphologies, changes in detection parameters, and instrument replacements, Raman spectra usually exhibit intensity differences and wavelength shifts. If traditional quantitative or qualitative models directly predict these spectra, the prediction results will deviate.

[0062] In this embodiment, the problem of deviation in prediction results is solved by the Piecewise Direct Standardization (PDS) algorithm. Since the wavenumber and intensity changes of spectral data are regional and limited within a certain range. The spectral data is divided into a target set spectrum and an adjustment set spectrum. Thus, the value of a certain wavenumber in the target set spectrum is only significantly related to several points near the corresponding wavenumber in the adjustment set spectrum and has no relation to wavenumbers far apart. Select a certain wavenumber as the center and expand it left and right according to the set range as a window. A multiple regression model is constructed using the intensity value of the ith wavenumber in the target set spectrum and the window matrix centered on i in the adjustment set spectrum, as shown in formula (2), where i is a natural number:

[0063]

[0064] is the spectral matrix at the ith wavenumber of the target set spectrum, k is the piecewise half-window width, and the window width is 2k + 1; is the spectral matrix with a window width of 2k + 1 on both sides of the ith wavenumber of the adjustment set spectrum, and bi is the regression coefficient at the ith wavenumber;

[0065] $b_i$ is solved by partial least squares regression. The regression coefficient $b_i$ in the regression model is placed on the main diagonal of the transformation matrix $F$, and the other elements are set to 0, thus obtaining the transformation matrix $F$, as shown in formula (3).

[0066]

[0067] The new clinical sample XS ′ can be transformed into a standardized spectrum X that is the same as the pure culture space through the transformation matrix $F$, ′ s,std , as shown in formula (4):

[0068] X ′ s,std = X ′ S·F (4)

[0069] The spectrum used to establish the spectral transfer model is called the standard spectrum. The samples used to collect the standard spectrum are single-cell samples common under both clinical and pure culture conditions. When the sample size is 150 - 200, the spectral characteristics of the samples within the target set range can be covered.

[0070] Step 104, construct a general single-cell phenomics database.

[0071] The standardized spectral data after model transfer is combined with the single-cell image data collected in real-time during spectral real-time acquisition and stored in the omics database, thereby constructing a single-cell phenomics database;

[0072] Specifically, during the real-time acquisition of single-cell spectra, the obtained single-cell image data is stored. Combining the spectral data standardized by the calibrated transfer model, a general single-cell phenomics database is constructed to provide data support for subsequent identification and comparison.

[0073] Step 105, multi-modal feature fusion.

[0074] The cell images and spectral data in the constructed single-cell phenomics database are combined, and the ReliefF algorithm is used to determine the weights of the eigenvalues of the images and spectra for fusion operations, forming multi-modal features, and the multi-modal features are superimposed to form a longer vector as a description of the single-cell phenotype, used to characterize the single-cell object, achieving multi-modal feature fusion.

[0075] Step 106, construct a CNN classifier.

[0076] The single-cell phenotype data to be classified collected in real-time, after multi-modal feature fusion, is classified and compared using the eigenvalues to obtain the cell types, realizing the species identification of the single-cell phenotype data collected in real-time.

[0077] Specifically, the CNN classifier architecture consists of an initial convolutional layer, six residual layers, and a final fully connected layer. The residual layer contains shortcut connections between the input and output of each residual block, which makes the gradient propagation better and the training more stable. Each residual layer contains four convolutional layers, so the total depth of the network is twenty-six layers. The initial convolutional layer has sixty-four convolutional filters, and each convolutional layer has 100 filters. The architectural parameters of the initial convolutional layer, the six residual layers, and the final fully connected layer are selected by grid search and separated once training and validation on the species classification task.

[0078] Specifically, the activation function of the CNN classifier uses the Sigmoid function, and Φ(z) is the output after the nonlinearization of the activation function, which is used as the input of the next layer. The purpose is to introduce nonlinear factors to solve the problem of insufficient expression ability of the linear model, as shown in formula (5):

[0079]

[0080] Among them, z is the output of the input of this layer network multiplied by the weight and superimposed with the offset, which is used as the input of the activation function. The right side of formula (5) is the Sigmoid function.

[0081] The loss function of the CNN classifier is the cross entropy loss function. The loss function reflects the gap between the predicted data and the actual data. The function is shown in formula (6):

[0082]

[0083] Among them, i represents the i-th category, N is the total number of categories, and y (i) is the one-hot representation of the test data in the i-th category, is the probability distribution representation of the test data in the i-th category, and L is the loss value.

[0084] Due to the inherent characteristics of fusion feature data, there is When δ is 0, the loss function becomes Nan in a certain round of training, resulting in the function failing to converge. Therefore, the loss function is improved by truncating the parameter and giving a minimum non-zero value δ to ensure that the loss function is not Nan. The improvement is shown in formula (7):

[0085]

[0086] The Softmax function connected after the fully connected layer maps the outputs of multiple neurons to the interval (0,1), normalizes the output vector, highlights the largest value and suppresses other components far below the maximum value, thereby achieving multi-classification.

[0087] In this embodiment, Raman spectral data can be collected by a microbial single-cell species rapid identification instrument, which includes an excitation light module, a microscopic focusing module, a Raman main optical path and transmission module, a coaxial illumination module, an imaging module, an electric displacement platform, and a collection control module;

[0088] The excitation light module is used to emit laser light; the microscopic focusing module is used to focus the laser light on the sample to generate Raman signals;

[0089] The Raman main optical path and transmission module is used to obtain the Raman spectrum of cells on the sample and transmit the Raman spectrum information of the cells to the software automatic collection control module;

[0090] The coaxial illumination module is used to provide coaxial illumination light for the microscopic focusing module; the imaging module is used to photograph the cells, obtain the position information of each cell, form position correction information obtained by comparing each cell with a preset cell; and is also used to photograph and determine the collection position of the cells;

[0091] The collection control module is used to control the excitation light module, the Raman main optical path and transmission module, the microscopic focusing module, the coaxial illumination module, the imaging module, and the electric displacement platform to realize the collection of Raman spectral data.

[0092] The present invention also provides a microbial single-cell species identification device, including:

[0093] A Raman spectrum map screening module, configured to compare and analyze the collected single-cell Raman spectral data with the Raman spectrum data in the reference spectrum database, and screen out the qualified Raman spectral data;

[0094] An analysis of spectral detection quantity module, configured to use the screened Raman spectral data as a sample, calculate according to specific spectral characteristic values in the spectral data sample, and obtain the minimum sample spectral detection quantity that is stable under the specific spectral characteristic values;

[0095] A spectral data normalization module, configured to collect the spectral data corresponding to the spectral detection quantity, normalize the spectral data through a calibration transfer model, and obtain normalized spectral data;

[0096] A general single-cell phenomics database construction module, configured to store the normalized spectral data and the real-time collected single-cell image data into the omics database to construct a single-cell phenomics database;

[0097] A multi-modal feature fusion module, configured to perform multi-modal feature fusion on the feature values of the image and spectrum based on the cell image and spectral data in the single-cell phenomics database;

[0098] A classification module configured to classify the multi-modal feature fusion through a CNN classifier to obtain cell types, thereby realizing the identification of the types of single-cell phenotype data.

[0099] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of the method for identifying the types of microbial single cells.

[0100] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program, when executed by the processor, implements the steps of the method for identifying the types of microbial single cells.

[0101] The present invention is described in terms of the flowcharts and / or block diagrams of the method, apparatus (system), and computer program product according to the specific embodiments. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for identifying single - cell species of microorganisms, characterized in that, the method includes: Comparing and analyzing the collected single - cell Raman spectroscopy data with the Raman spectroscopy data in the reference spectrum database, and screening out the qualified Raman spectroscopy data; Taking the screened Raman spectroscopy data as samples, calculating according to the set spectral characteristic values in the samples, and obtaining the minimum sample spectral detection quantity that is stable under the set spectral characteristic values; Collecting the spectral data corresponding to the spectral detection quantity, and standardizing the spectral data through a calibration transfer model to obtain standardized spectral data; Constructing a single - cell phenomics database based on the standardized spectral data and the real - time collected single - cell image data; Performing multi - modal feature fusion on the characteristic values of the images and spectra based on the cell images and spectral data in the single - cell phenomics database; Classifying the data after multi - modal feature fusion to obtain the cell species, and realizing the identification of the species of single - cell phenotype data; Comparing and analyzing the collected single - cell Raman spectroscopy data with the Raman spectroscopy data in the reference spectrum database, and screening out the qualified Raman spectroscopy data, including: Constructing a reference spectrum database, and storing the screened spectra into the reference database; Using the CNN algorithm to compare and analyze the collected spectral data with the data in the reference spectrum database, and screening out the spectra with high similarity; Using the CNN algorithm to compare and analyze the collected spectral data with the data in the reference spectrum database, and screening out the spectra with high similarity, including: Taking the collected Raman spectroscopy data as test data, inputting it into the reference spectrum database, and outputting an N - dimensional output vector corresponding to N species after calculation, where N is a natural number; Mapping the vector as an input to the Softmax function. For a specific test data, the maximum probability value output by Softmax is P. The mean of the maximum values of the Softmax function for all data of the same category in the test data is M, and the variance is S. If M - S / 2 ≤ P ≤ M + S / 2, then the Raman spectroscopy data is the screened qualified Raman spectroscopy data.

2. The method for identifying single - cell species of microorganisms according to claim 1, characterized in that, Collecting the spectral data corresponding to the spectral detection quantity, and standardizing the spectral data through a calibration transfer model to obtain standardized spectral data. The piece - wise direct standardization (PDS) algorithm is adopted, including: Dividing the spectral data into a target set spectrum and an adjustment - required set spectrum; Selecting a certain wave number as the center, and expanding left and right according to the set range as a window. Constructing a multiple regression model with the intensity value of the i - th wave number of the target set spectrum and the window matrix centered on i of the adjustment - required set spectrum, where i is a natural number; Solving through partial least - squares regression, placing the regression coefficients in the regression model on the main diagonal of the transformation matrix, and setting other elements to 0 to obtain a transformation matrix; Transforming the collected spectral data into standardized spectral data through the transformation matrix.

3. The method for identifying single - cell species of microorganisms according to claim 1, characterized in that, The data after multi-modal feature fusion is classified using a CNN classifier to obtain the cell types.

4. The method for identifying microbial single-cell types as described in claim 1, characterized in that the ReliefF algorithm is used to determine the weights of the feature values of the image and the spectrum for fusion operations to form multi-modal features.

5. A device for identifying microbial single-cell types, characterized in that it includes: A Raman spectrum map screening module configured to compare and analyze the collected single-cell Raman spectrum data with the Raman spectrum data in the reference map database, and screen out the qualified Raman spectrum data; Comparing and analyzing the collected single-cell Raman spectrum data with the Raman spectrum data in the reference map database, and screening out the qualified Raman spectrum data, including: Constructing a reference map database and storing the screened maps in the reference database; Using the CNN algorithm to compare and analyze the collected spectrum data with the data in the reference map database, and screening out the maps with high similarity; Using the CNN algorithm to compare and analyze the collected spectrum data with the data in the reference map database, and screening out the maps with high similarity, including: Taking the collected Raman spectrum data as test data, inputting it into the reference map database, and outputting an N-dimensional output vector corresponding to N species after calculation, where N is a natural number; Mapping the vector as input to the Softmax function, the maximum probability value of the Softmax output for a specific test data is P, the mean value of the maximum values of the Softmax function for all data of the same category in the test data is M, and the variance is S. If M - S / 2 ≤ P ≤ M + S / 2, then the Raman spectrum data is the screened qualified Raman spectrum data; An analysis of spectral detection quantity module configured to use the screened Raman spectrum data as samples, calculate according to the set spectral feature values in the spectral data samples, and obtain the minimum sample spectral detection quantity that is stable under the set spectral feature values; A spectral data normalization module configured to collect the spectral data corresponding to the spectral detection quantity, and normalize the spectral data through a calibration transfer model to obtain normalized spectral data; A module for constructing a general single-cell phenomics database configured to construct a single-cell phenomics database with the normalized spectral data and the single-cell image data collected in real time; A multi-modal feature fusion module configured to perform multi-modal feature fusion on the feature values of the image and the spectrum based on the cell images and spectral data in the single-cell phenomics database; A classification module configured to classify the multi-modal feature fusion through a CNN classifier to obtain the cell types, and realize the identification of the types of single-cell phenotype data.

6. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the steps of the method for identifying microbial single-cell types according to any one of claims 1-4.

7. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the computer program, it implements the steps of the method for identifying single microbial cell species according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method for rapid identification of microalgae on single cell level

    CN103940801A

  • Single-cell phenotype database system and search engine

    CN104077307A