Cell Raman spectrum mass data modeling method
Through standardizing the biological sample culture and Raman spectroscopy acquisition process, combined with the network structure of the deep learning model, the problems of poor generalization capabilities and low inter-batch prediction accuracy in the prior art are solved, and efficient repeated prediction of biological samples of different batches are achieved.
Patent Information
- Application Number
- CN202311771122.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing Raman spectroscopy detection technology has challenges in model generalization capabilities and repeated prediction of data in different batches, which leads to the inability of the model to effectively predict data re-cultivated by the same sample, and the accuracy of repeated predictions between batches is seriously reduced.
By standardizing the culture and Raman spectral acquisition process of biological samples, including setting standard instrument parameters and acquisition conditions, spectral cleaning and real-time background subtraction, and finally building a network structure of a deep learning model to achieve high-precision repeat prediction of different batches of biological samples.
It effectively improves the generalization of the model, realizes efficient repeated prediction of biological samples in different batches, improves the prediction accuracy, and solves the problems of poor generalization ability of the model and low prediction accuracy among batches.
Smart Images

Figure CN120198906A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cell detection, and particularly relates to a method for modeling a large amount of cell Raman spectroscopy data. Background Art
[0002] Raman spectroscopy is a spectroscopic method for indirectly measuring the vibration state in a sample. It is the inelastic scattering of photons on a quantized system. Since the molecular vibration state is molecule-specific, Raman spectroscopy can be used as the "vibration fingerprint" of molecules. Due to the non-invasive, non-cultured, and label-free characteristics of Raman spectroscopy, it is widely used in biological samples and biomedical diagnosis, and can monitor the chemical composition and metabolism of single living microorganisms or tissues in real time.
[0003] Raman spectroscopy represents a collection of different molecular vibration modes and structures, including nucleic acids, proteins, lipids, and carbohydrates. The high-dimensional and complex Raman spectroscopy can provide rich information on the phenotype of single cells to distinguish different bacterial species or biological activity states. However, due to the inherently low quantum efficiency of the Raman effect, spectral variations between individuals of the same species, cell heterogeneity, and spectral overlap of different molecules, although various data preprocessing methods, chemometric analysis, machine learning, and deep learning methods are currently used to analyze complex biological sample Raman spectroscopy data, such as principal component analysis (PCA), support vector machine (SVM), random forest (RF), partial least squares (PLS), etc., it is extremely challenging to control the quality of biological sample Raman spectroscopy data and establish a reusable model. The most important problem is the poor generalization ability of the training model. Most of the current domestic and foreign scientific research achievements on Raman spectroscopy detection of biological samples have established models with quite good performance on the data of the modeling set, but they cannot predict the data of the same sample after re-cultivation. When changing batches, the model cannot be reused, and the correct rate of repeated prediction between batches is severely reduced. Finally, it is necessary to mix a part of the new batch data or new data with the original data and re-establish, migrate, or fine-tune the model again. As described in the literature "Deep Learning for Raman Spectroscopy: A Review" and "Rapid identification of pathogenic bacteria using Raman spectroscopy and deep learning", this seriously hinders and limits the practical application of Raman spectrometers in species discrimination and species detection in the biological field. Therefore, in view of the above problems, how to construct a complete modeling system, establish cell culture process specifications, and Raman data acquisition and preprocessing standards is particularly important for improving the generalization of the model and realizing repeated prediction of different batches by the model. Summary of the Invention
[0004] The present invention provides a method for modeling a large amount of cell Raman spectroscopy data, which can effectively improve the generalization of the model and achieve efficient repeated prediction of different batches of biological samples.
[0005] In order to achieve the above object, a method for modeling a large amount of cell Raman spectroscopy data includes the following steps:
[0006] Cultivate biological samples in a standardized manner;
[0007] Standardize the instrument parameters and acquisition conditions of the Raman spectroscopy acquisition end, and acquire Raman spectra of biological samples;
[0008] Clean the acquired spectra to remove invalid and abnormal spectra, ensuring that the number of valid spectra for each type > 500;
[0009] Subtract the real-time background spectrum from the valid spectrum data obtained after cleaning to obtain the spectrum after background subtraction, and perform data preprocessing on the spectrum after background subtraction;
[0010] Build a network structure of a deep learning model suitable for the data after the above preprocessing to achieve the purpose of high-precision repeated prediction of different batches of biological samples.
[0011] Preferably, the standardization of the instrument parameters and acquisition conditions of the Raman spectroscopy acquisition end is specifically:
[0012] Set the instrument parameters of the acquisition end as: integration time 10 ms - 10 min, laser power 1 mw - 300 mw;
[0013] After the instrument parameters are fixed, keep the peak positions and peak intensities of the background spectra collected from repeated experiments of different batches of biological samples the same;
[0014] Set the acquisition conditions as: the peak position of the quartz peak is in the range of 320 - 460 cm -1 and the maximum value of the peak intensity is in the range of A ± 10%, the peak position of the water peak is in the range of 3300 - 3500 cm -1 and the maximum value of the peak intensity is in the range of B ± 10%, and the values of A and B are determined by the setting of the instrument parameters.
[0015] Preferably, the dispensing concentration of the biological sample is not less than 1×10 2 cell / ml, and the injection flow rate is not less than 10 μL / min.
[0016] Preferably, the real-time background spectrum is a non-cellular spectrum recorded in chronological order during the Raman spectrum acquisition process. It can be understood that parameters such as the integration time of the instrument, the sample concentration, and the injection flow rate cooperate with each other to determine appropriate parameters, so that the cell spectrum and the background spectrum appear alternately at irregular intervals during the Raman spectrum acquisition process and are recorded according to time. According to the time information, each cell spectrum is subtracted from the background spectrum closest to that spectrum, and the purpose of real-time background subtraction can be achieved. Compared with the traditional background subtraction method (subtracting a fixed background spectrum collected at the beginning of the experiment), there is a substantial change and innovation in the method, which is more helpful to eliminate the influence of background information changes during the Raman acquisition process on the cell spectrum.
[0017] Preferably, for the signal-to-noise ratio calculation method, by calculating the peak area ratio of a certain peak in the Raman spectrum of a biological sample, the collected spectrum can be spectrally cleaned.
[0018] Preferably, the calculation method is as follows:
[0019]
[0020] Wherein, A Ram : The area size of the characteristic peak signal, A tot : The total area size of the characteristic peak signal, A Bg : The fluctuation size of the background peak.
[0021] Preferably, in the step of spectrally cleaning the collected spectrum, it further includes the step of eliminating randomly occurring cosmic rays in the spectrum.
[0022] Preferably, the second derivative and interpolation method are adopted. By taking the second derivative, the position of the cosmic ray in the spectrum is found, and multiple points near the cosmic ray are selected for interpolation. The result of the interpolation replaces the cosmic ray value to achieve the elimination of randomly occurring cosmic rays in the spectrum.
[0023] Preferably, it is characterized in that the interpolation method is selected from at least one of piecewise cubic Hermite interpolation, least squares interpolation, deep learning interpolation, and polynomial interpolation.
[0024] Preferably, the learning model is selected from traditional machine learning models in chemometrics and long short-term memory recurrent neural network models and convolutional residual neural network models in deep learning, including but not limited to the following: LDA, SVM, PLS, KNN, Random Forest, ResNet18, ResNet34, ResNet50, ResNet101, ResNet152, LeNet5, AlexNet, VGG16, GoogLeNet, LSTM, multi-layer LSTM, TCN, Transformer, LSTM combined with ResNet, TCN combined with ResNet, Transformer combined with ResNet.
[0025] Preferably, the preprocessing method is selected from smoothing, baseline removal, detrending, multiplicative scatter correction, derivation, data standardization, normalization, data augmentation transformation, or a combination of one or more of these methods.
[0026] Preferably, the biological sample is spherical, oval or rod-shaped microorganisms and animal and plant cells with a size of 0.5 - 50 μm. It can be understood that there are no special requirements for the starting cells of biological samples in repeated experiments of different batches. For example, they can be obtained from plate culture or from 16 - 48-hour overnight liquid culture.
[0027] Preferably, when the biological sample is yeast-like cells (including but not limited to yeast cells), the conditions for plate culture are as follows: Take the cryopreserved bacterial liquid from -80°C, streak it on a YPD plate, and place it in an incubator at 30°C for 36 - 48 hours. Then take out the plate. Randomly pick 1 - 3 monoclonal colonies on the plate and mix them into 100 - 200 μL of sterile water to prepare a cell mother liquor. Then store the plate in a 4°C refrigerator for use in other batch experiments. According to the requirements of high-throughput flow Raman acquisition, take 1 - 2 μL of the cell mother liquor, add 5 - 10 mL of the upper machine buffer to prepare a sample, and load the sample for Raman spectrum acquisition.
[0028] Preferably, when the biological sample is yeast-like cells (including but not limited to yeast cells), the conditions for overnight liquid culture are as follows:
[0029] Pick monoclonal colonies from the plate and inoculate them into a 15 mL centrifuge tube containing 2 mL of PYD liquid medium. Set the temperature at 30°C and the shaker oscillation frequency at 200 rpm, and culture overnight for 16 - 20 h. Take 100 μL of the overnight cultured bacterial liquid, set the temperature at 4°C, and centrifuge at 3000 rpm for 3 min. Resuspend it with sterile water, centrifuge and wash twice under the same conditions, discard the supernatant, and then add 500 - 900 μL of sterile water to resuspend it; Take 5 mL of the upper machine buffer and add 2 μL of the diluted bacterial liquid to prepare a sample suitable for Raman acquisition, and load the sample for Raman spectrum acquisition.
[0030] It is understandable that during the above-mentioned acquisition process, it is necessary to control the instrument parameters, cell concentration, and injection flow rate so that the cell spectrum and the background spectrum appear alternately at irregular intervals during the Raman spectrum acquisition process.
[0031] Preferably, when the biological sample is bacterial cells (including but not limited to yeast cells), its culture conditions are basically the same as those of yeast cells, with the differences being: bacteria are cultured by streaking on an LB plate, and after using the plate culture, 3 - 5 monoclonal colonies are randomly selected on the plate. For liquid culture, an LB liquid medium is used, and the culture temperature is set at 37°C.
[0032] Preferably, when the biological sample is mammalian cells (including but not limited to mammalian cells), for its culture conditions, the cells taken out from the -80°C refrigerator are cultured for 2 - 3 days after resuscitation. After adding trypsin to terminate digestion, 1 ml of cell suspension is taken into a 1.5 mL centrifuge tube. The remaining cells are passaged and then cultured for another 2 - 3 days before sampling again, and the culture passage is 5 - 10 generations. The taken-out cell suspension is centrifuged at 1000 rpm for 3 min at 23°C, resuspended with a cell buffer (components: glucose and sucrose), centrifuged and washed twice under the same conditions, the supernatant is discarded, and then 1 mL of buffer is added to resuspend. 50 μL of the diluted suspension is taken out, 5 mL of the upper machine buffer is added to prepare a sample suitable for Raman acquisition, and then the sample is loaded for Raman spectrum acquisition.
[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0034] The Raman spectrum large - batch data modeling method provided by the present invention standardizes the sample preparation process, Raman acquisition parameters and conditions, controls the flow rate and culture concentration during sampling, and subtracts the background in real - time. At the same time, it simplifies the spectrum cleaning method and the spectrum data pre - processing process, and builds a neural network structure suitable for the data format and a series of operations. It can effectively improve the generalization of the model and achieve efficient and repeated prediction of different batches of biological samples. Brief Description of the Drawings
[0035] Figure 1 It is a calculation diagram of the invalid and abnormal spectrum screening formula provided by the embodiment of the present invention;
[0036] Figure 2 It is a comparison diagram before and after spectrum processing provided by the embodiment of the present invention;
[0037] Figure 3 It is a flow chart taking the TCN model as an example provided by the embodiment of the present invention;
[0038] Figure 4 It is a comparison diagram of the extrapolated data accuracy rates of different learning models provided by the embodiment of the present invention, where the abscissa is the cell line number and the ordinate is the accuracy rate. Detailed implementation manners
[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Example 1
[0041] Cell samples: Three types of yeast cells are used: Saccharomyces cerevisiae, Schizosaccharomyces pombe, and Pichia pastoris, and four Bacterial cells Cells: Escherichia coli ATCC 35218, ATCC 25922, Staphylococcus aureus, and Staphylococcus epidermidis.
[0042] Cell culture:
[0043] For yeast cells , plate culture or overnight liquid culture can be adopted. Specifically:
[0044] The conditions for plate culture are specifically as follows:
[0045] Take the frozen bacterial liquid from -80°C, streak it on the YPD plate, place it in an incubator at 30°C for 36 - 48 hours, and then take out the plate. Randomly pick 1 - 3 monoclonal colonies on the plate and mix them into 100 - 200 μL of sterile water to prepare the cell mother liquor. Then, the plate is stored in a 4°C refrigerator for use in other batch experiments. According to the requirements of high-throughput flow Raman acquisition, take 1 - 2 μL of the cell mother liquor, add 5 - 10 mL of the upper machine buffer to prepare the sample, and load the sample for Raman spectrum acquisition.
[0046] The conditions for overnight liquid culture are specifically as follows:
[0047] Pick monoclonal colonies from the plate and inoculate them into a 15 mL centrifuge tube containing 2 mL of PYD liquid medium. Set the temperature at 30°C and the shaker oscillation frequency at 200 rpm, and culture overnight for 16 - 20 h. Take 100 μL of the overnight cultured bacterial liquid, set the temperature at 4°C, and centrifuge at 3000 rpm for 3 min. Resuspend it with sterile water, centrifuge and wash twice under the same conditions, discard the supernatant, and then add 500 - 900 μL of sterile water to resuspend; take 5 mL of the upper machine buffer and add 2 μL of the diluted bacterial liquid to prepare a sample suitable for Raman acquisition, and load the sample for Raman spectrum acquisition.
[0048] For bacterial cells , and its culture conditions are basically the same as those of yeast cells, except that: for bacterial culture, streak on the LB plate, randomly pick 3 - 5 monoclonal colonies on the plate after using plate culture, use LB liquid medium for liquid culture, and set the culture temperature at 37°C.
[0049] Raman spectrum acquisition:
[0050] Set the instrument parameters as follows: integration time 10 ms - 10 min, laser power 1 mW - 300 mW, select the 50x objective lens focusing ring at the 0.1 position, filter 100%, the software automatically adjusts the focus, and controls the background substance signal; after the instrument parameters are fixed, repeat the experiment in different batches to keep the peak positions and peak intensities of the collected background spectra the same; among them, the peak position of the quartz peak is at 320 - 460 cm -1 , the peak position of the water peak is at 3300 - 3500 cm -1 , and the peak intensity is affected by parameters such as the integration time and power.
[0051] The conditions set for this experiment are: integration time 0.1 s - 10 s, laser power 100 mW - 300 mW, continuous long - time acquisition, the quartz peak position (320 - 460 cm -1 ), the maximum value of the peak intensity fluctuates within the range of A ± 10%, the water peak position (3300 - 3500 cm -1 ), the maximum value of the peak intensity fluctuates within the range of B ± 10%, and the values of A and B are determined by the instrument parameters.
[0052] The sample concentration of the biological sample is not less than 1×10 2 cell / ml, and the injection flow rate is not less than 10 μL / min. Select appropriate sampling points to make the cell spectra and background spectra appear alternately at irregular intervals during the Raman spectrum acquisition process. For each type collected each time, the number of effective Raman spectra is greater than 500.
[0053] In this embodiment, all start from the flat plate. Among them, the yeast - like cells are biologically replicated nine batches, and the Escherichia coli are biologically replicated six batches. Among them, Escherichia coli ATCC 35218 and ATCC 25922 belong to the same species but different strains, and the others are all different species. Continuously sample the Raman spectra for 15 days.
[0054] Spectrum cleaning:
[0055] The original spectrum contains many invalid or abnormal spectra, mainly including:
[0056] (1) The saturated spectrum generated when turning on the white light to observe the cell morphology in the chip;
[0057] (2) The empty spectrum collected only for the buffer when the cell concentration is low or the cell flow is discontinuous, without collecting cells;
[0058] (3) The spectrum of broken cells generated by the broken cells remaining in the pre - treatment process of cell culture;
[0059] (4) The cell spectrum generated due to the laser focus of the cell flow being fixed at the cell edge;
[0060] (5) The spikes randomly appearing in the spectrum due to the influence of cosmic rays.
[0061] To meet the high-throughput requirements of the instrument and simplify the data cleaning process, by only calculating the peak area ratio of a certain characteristic peak of the biological sample and optimizing the calculation method of the signal-to-noise ratio, the elimination of invalid and abnormal spectra of types (1)-(4) can be achieved. The calculation method (1) is as follows:
[0062]
[0063] Among them, A Ram : The area size of a certain characteristic peak signal, A tot : The total area size of a certain characteristic peak signal, A Bg : The fluctuation size of the background peak. The above peak signals are as Figure 1 shown.
[0064] Furthermore, after achieving the elimination of invalid and abnormal spectra of types (1)-(4), it also includes the step of eliminating randomly appearing cosmic rays in the spectrum. Specifically: Through second-order derivative and piecewise cubic Hermite interpolation polynomial for interpolation, find the position where the cosmic rays in the spectrum are located. By selecting points adjacent to the cosmic rays and performing polynomial interpolation, the obtained result replaces the original cosmic rays to achieve the purpose of eliminating cosmic rays. The calculation method (2) is as follows:
[0065] Assume that the known function f(x) satisfies f(x i ) = f i and f'(x i ) = f i ' (i = 0, 1, 2,..., n) at n + 1 distinct nodes xi (i = 0, 1...) in the interpolation interval [p, q]. If the function G(x) exists and satisfies the following conditions:
[0066] ① The polynomial degree of G(x) on each subinterval is 3
[0067] ② G(x) ∈ C 1 [a, b]
[0068] ③ G(x i ) = f(x i ), G'(x i ) = f'(x i ), i = (0, 1,..., n)
[0069] then G(x) is called the Hermite interpolation polynomial of f(x) at n + 1 nodes x iThe piecewise cubic Hermite interpolation polynomial on it. Therefore, we have:
[0070]
[0071] Data processing:
[0072] Perform real-time background subtraction on the data of the effective spectrum obtained after cleaning, and normalize the data after background subtraction with respect to the maximum value of 65535 (the maximum resolution of the acquisition device is 2 16 , so the value range is 0 - 65535). Normalization is beneficial for the computer network to learn the data;
[0073] Combined with this embodiment, the experimental data is cleaned, and the effect is as Figure 2 shown. The left figure is the original spectrum containing invalid spectra, and the right figure is the spectrum after the data cleaning algorithm. By comparison, it can be seen that this method can better remove invalid and abnormal spectra. Only using functions (1)-(2) can meet the elimination of various types of abnormal data and at the same time meet the requirements of high-throughput of the instrument.
[0074] Build the network structure of the deep learning model:
[0075] After data cleaning and preprocessing, randomly select the Raman spectra collected in three days as the model extrapolation verification data, and the remaining data as the model building data. Randomly select 10% of the data in the model building data as the validation set, 10% as the test set, and 80% as the training set.
[0076] The learning model can be selected from at least one of LDA, SVM, PLS, KNN, Random Forest, ResNet18, ResNet34, ResNet50, ResNet101, ResNet152, LeNet5, AlexnNet, VGG16, GoogleNet, LSTM, multi-layer LSTM, TCN, Transformer, LSTM combined with ResNet, TCN combined with ResNet, Transformer combined with ResNet.
[0077] Combined with this embodiment, after a large number of data tests, different preprocessing methods are combined to explore the model. Taking the combination of TCN + ResNet as an example, the modeling process is as Figure 3 shown. We built a temporal convolutional network architecture. In Figure 3 , we built a network architecture that combines residual and temporal convolutional networks, considering residuals not only in the TCN layer but also in the fully connected layer.
[0078] As Figure 4As shown, on the premise of the same training dataset and test dataset, through the standard process proposed by the present invention, such as the cultivation method, Raman acquisition operation, parameter setting, data cleaning, data preprocessing method, real-time background reduction, model selection, etc., it is found that the accuracy and recall rate of the experimental results of a variety of different learning models (including but not limited to the models listed in the table) have been greatly improved. The optimal prediction accuracy between different batches of cell species and strains can reach over 95%, effectively solving the problem of low prediction accuracy or even inability to predict between different batches before.
[0079] Example 2
[0080] Cell samples: human B lymphoma cells, human T lymphoma cell leukemia cells, colon cancer cells, breast cancer cells, pancreatic cancer cells.
[0081] Cell culture:
[0082] After the cells taken out from the -80 refrigerator are resuscitated, they are cultured for 2 - 3 days. After adding trypsin to terminate digestion, 1 ml of cell suspension is taken into a 1.5 mL centrifuge tube. The remaining cells are passaged and then cultured for another 2 - 3 days before sampling again for use in different batches. The culture passage is 5 - 10 generations. The taken-out cell suspension is centrifuged at 1000 rpm for 3 min at 23°C, resuspended with cell buffer (components glucose and sucrose), centrifuged and washed twice under the same conditions, the supernatant is discarded, and then 1 mL of buffer is added to resuspend. 50 μL of the diluted suspension is taken out, 5 mL of the upper machine buffer is added to prepare a sample suitable for Raman acquisition, and the sample is loaded for Raman spectrum acquisition.
[0083] Raman spectrum acquisition:
[0084] The instrument parameters and acquisition process for collecting the Raman spectra of human cells strictly follow the process requirements of the case. The integration time is set to 0.5 s - 30 s, the laser power is 200 mw - 300 mw, the software automatically adjusts the focal length to control the background substance signal; after the instrument parameters are fixed, the peak positions and peak intensities of the background spectra collected in repeated experiments of different batches are kept the same; the flow rate and cell concentration are controlled, and appropriate sampling points are selected to make the cell spectra and background spectra appear alternately at irregular intervals during the Raman spectrum acquisition process. The number of Raman spectra collected for each type each time, the effective spectra are more than 500.
[0085] In this example, human B lymphoma cells, human T lymphoma cell leukemia cells, intestinal cancer cells, breast cancer cells, and pancreatic cancer cells are biologically replicated in six batches, and Raman spectra are continuously collected for about 20 days.
[0086] Data cleaning:
[0087] Similar to Example 1, an optimized signal-to-noise ratio algorithm is adopted, and an appropriate threshold is set to remove invalid and abnormal spectra, achieving the purpose of cleaning data.
[0088] Data preprocessing:
[0089] Similar to Example 1, the obtained effective spectral data after cleaning is subjected to real-time background subtraction, and the data ratio after background subtraction is normalized by the maximum value. Normalization is beneficial for the computer network to learn the data.
[0090] Construct the network structure of the deep learning model:
[0091] Select the Raman spectra collected in the last two days as the model extrapolation verification data, and the remaining data as the model modeling data. Through the standardized process and modeling method proposed by the present invention, as Figure 4 shown, it can be seen from the experimental results that various tested models can achieve repeated prediction of mammalian cells and cells of different culture passages, and the highest prediction accuracy rate reaches more than 90%.
[0092] In summary, the method provided by the present invention has a wide range of application scenarios.
Claims
1. A method for modeling a large amount of cell Raman spectroscopy data, characterized in that Including the following steps: Standardize the culture of biological samples; Standardize the instrument parameters and acquisition conditions of the Raman spectroscopy acquisition end, and acquire Raman spectra of biological samples; Clean the acquired spectra, remove invalid and abnormal spectra, and ensure that the number of valid spectra for each type > 500; Subtract the real-time background spectrum from the valid spectral data obtained after cleaning to obtain the spectrum after background subtraction, and perform data preprocessing on the spectrum after background subtraction; Build a network structure suitable for the deep learning model for the data after the above preprocessing to achieve the purpose of high-precision repeated prediction of biological samples in different batches.
2. The modeling method according to claim 1, characterized in that, The standardization of the instrument parameters and acquisition conditions of the Raman spectroscopy acquisition end specifically is: Set the instrument parameters of the acquisition end as: integration time 10 ms - 10 min, laser power 1 mW - 300 mW; After the instrument parameters are fixed, keep the peak positions and peak intensities of the background spectra collected by repeated experiments on biological samples in different batches the same; Set the acquisition conditions as follows: the peak position of the quartz peak is in the range of 320 - 460 cm -1 and the maximum peak intensity is in the range of A ± 10%, the peak position of the water peak is in the range of 3300 - 3500 cm -1 and the maximum peak intensity is in the range of B ± 10%. The values of A and B are determined by the settings of the instrument parameters.
3. The modeling method according to claim 1 or 2, characterized in that The dispensing concentration of the biological sample is not less than 1×10 2 cell / ml, and the injection flow rate is not less than 10 uL / min.
4. The modeling method according to claim 1, wherein The real-time background spectrum is the non-cell spectrum recorded in chronological order during the Raman spectrum acquisition process.
5. The modeling method according to claim 1, characterized in that, Optimize the signal-to-noise ratio calculation method. By calculating the peak area ratio of a certain peak of the Raman spectrum of a biological sample, the acquired spectra can be spectroscopically cleaned.
6. The modeling method according to claim 5, characterized in that The calculation method is as follows: Among them, A Ram : the area size of the characteristic peak signal, A tot : the total area size of the characteristic peak signal, A Bg : the fluctuation size of the background peak.
7. The modeling method according to claim 1, wherein In the step of spectroscopically cleaning the acquired spectra, it also includes the step of eliminating randomly appearing cosmic rays in the spectra.
8. The modeling method according to claim 1, wherein Adopt the second-order derivative and interpolation method. Through the second-order derivative, find the positions of cosmic rays in the spectra, select multiple points near the cosmic rays for interpolation, and the interpolation result replaces the cosmic ray values to achieve the elimination of randomly appearing cosmic rays in the spectra.
9. The modeling method according to claim 1, characterized in that The learning model is selected from traditional machine learning models in chemometrics and long short-term memory recurrent neural network models and convolutional residual neural network models in deep learning, including but not limited to the following: LDA, SVM, PLS, KNN, Random Forest, ResNet18, ResNet34, ResNet50, ResNet101, ResNet152, LeNet5, AlexnNet, VGG16, GoogleNet, LSTM, multi-layer LSTM, TCN, Transformer, LSTM combined with ResNet, TCN combined with ResNet, Transformer combined with ResNet.
10. The modeling method according to any one of claims 1-9, characterized in that, The biological samples are spherical, oval or rod-shaped microorganisms and animal and plant cells with a size of 0.5 - 50 μm.