Raman peak characteristic matching method, system, equipment and medium

By combining protein peak characteristics and important data characteristics of machine learning models, feature enhancement extraction in breast cancer spectrum analysis is achieved, and the problem of insufficient utilization of biometric information in the prior art is solved, and the accuracy and effectiveness of the analysis are improved.

CN119988996APending Publication Date: 2025-05-13XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510059167.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks effective utilization of biometric peak information in spectral analysis of breast cancer, which leads to difficulty in combining important data characteristics with biological background, affecting the accuracy and effectiveness of the analysis.

Method used

Raman peak feature matching method is used to combine protein peak features with important data features of machine learning models, and feature enhancement extraction of spectral analysis is achieved through matching and screening.

Benefits of technology

It improves the accuracy and effectiveness of spectral analysis, enhances support for cell classification, and improves the classification accuracy of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988996A_ABST
    Figure CN119988996A_ABST
Patent Text Reader

Abstract

The invention relates to a Raman peak characteristic matching method, system, equipment and medium, and the method comprises the following steps: firstly, collecting various receptor protein Raman spectrums and various single cell Raman spectrums related to cell typing, including normal cells and different types of cancer cells; secondly, biological peak value feature extraction is carried out on the collected receptor protein Raman spectrum, and important data feature extraction of a machine learning model is carried out on the single-cell Raman spectrum; then matching the extracted biological peak value features with important data features of a machine learning model, and simultaneously completing important data feature screening of the machine learning model; finally, performing feature enhancement on the cell Raman spectrum according to the matched and screened important data features of the machine learning model, and retraining the initial machine learning model; the classification accuracy of the machine learning model is effectively improved through the feature enhancement extraction of the Raman spectrum; the system, the equipment and the medium can perform cell Raman peak feature matching based on the Raman peak feature matching method. The method has the advantages of good matching precision and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of feature-enhanced spectral analysis, and in particular relates to a Raman peak feature matching method, system, equipment and medium. Background Art

[0002] Raman spectroscopy has become a key technology in contemporary scientific and technological research and industrial applications due to its advantages such as non-destructiveness, high resolution and rich molecular fingerprint information. Raman spectroscopy data is usually complex and has high-dimensional characteristics. The application of machine learning algorithms has increasingly promoted the analysis of Raman spectroscopy data. These algorithms have the potential to significantly improve the efficiency and accuracy of Raman spectroscopy analysis through automation and intelligent data processing. The core of Raman spectroscopy analysis by machine learning algorithms is the extraction of Raman spectral features, a step that directly affects the identification, classification and quantitative analysis of samples.

[0003] Feature enhancement involves improving the quality and expressiveness of features through various technical means. It is particularly effective in enhancing the contribution of important features while reducing the impact of irrelevant features. This process begins with the identification of important features, which can be achieved by dimensionality reduction and interpreting important data features from the model output. However, dimensionality reduction is time-consuming and prone to subjective evaluation errors. The interpretability of machine learning algorithms provides a valuable method for extracting important data features. This is particularly important in the field of spectral analysis, where it is crucial to identify potential biomarkers or biomolecules that are important for classification. However, machine learning algorithms often rely on data-driven pattern recognition and may lack the ability to integrate biological background knowledge. Although these models perform well in certain tasks, their effectiveness is limited if they are not deeply integrated with domain-specific knowledge. Current biological research often involves direct interpretation of important data features corresponding to biological feature peaks in model outputs without making full use of the effective information of biological feature peaks, which poses a huge challenge to combining important data features with specific background biological features to derive important features.

[0004] The concept of matching is traditionally used to analyze and compare spectra of different samples to identify similarities or differences, but this concept is rarely applied to feature extraction. Most studies ignore biometric factors when analyzing single-cell spectra, especially cell surface proteins, which is a common method for breast cancer typing.

[0005] A patent application for a non-destructive and label-free rapid breast cancer Raman spectroscopy pathological grading and staging method (CN114923893A) is filed. Breast tissue samples are obtained through clinical breast-conserving surgery, tissue pathological biopsy and tissue frozen sections. The tissue Raman spectrum is measured and pre-processed to obtain spectral feature difference information, clarify the spectral features of healthy breast tissue and invasive ductal carcinoma tissue, and propose the spectral feature differences of tryptophan, phenylalanine, β-carotene, lipids, nucleic acids, proteins and fatty acid biomarkers contained in invasive ductal carcinoma tissue under different stages and grades; through the spectral analysis methods of generalized discriminant analysis, principal component analysis-support vector machine and principal component analysis-linear discriminant analysis, it is applied to the classification and identification of Raman spectral data of different breast cancer lesions, and realizes the pathological staging and analysis of breast cancer based on Raman spectroscopy detection technology. This application first constructed a generalized discriminant analysis model for diagnosis in the staging and grading of breast tumors, and then proposed to use the principal component analysis method to reduce the dimension of the spectral data set and extract feature information, and then input the significant feature variables into the support vector machine algorithm and the linear discriminant analysis algorithm. This method uses principal component analysis for dimensionality reduction, which is time-consuming and prone to subjective evaluation errors. The input into the support vector machine algorithm and the linear discriminant analysis algorithm lacks the ability to integrate biological background knowledge. In addition, to obtain information on spectral feature differences, it is only a simple comparison of tissue Raman spectra of the healthy group and the cancerous group as well as breast cancer cases at different stages. The Raman shift corresponding to the change in peak intensity is mainly tryptophan, phenylalanine, β-carotene, lipids, nucleic acids, proteins and fatty acids. However, in spectral analysis, these biological feature information is not effectively used to achieve the effective fusion of important data features of biological features and machine learning models.

[0006] Therefore, it is imperative to develop a matching method for enhanced extraction of protein spectral peak features in combination with single-cell spectral data features for spectral analysis. Summary of the invention

[0007] In order to overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a Raman peak feature matching method, system, device and medium, which combine the protein peak features with the important data features of the machine learning model, realize the feature enhancement extraction of spectral analysis, and have universality to other machine learning algorithms that can output important data features, greatly improving the accuracy and effectiveness of spectral analysis in biological and medical applications.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A Raman peak feature matching method specifically comprises the following steps:

[0010] Step 1, collecting Raman spectra of multiple receptor proteins and multiple single-cell Raman spectra related to cell typing, including normal cells and different types of cancer cells;

[0011] Step 2, extracting biological peak features from the receptor protein Raman spectrum collected in step 1, and extracting important data features of the machine learning model from the single cell Raman spectrum;

[0012] Step 3, matching the biological peak features extracted in step 2 with the important data features of the machine learning model, and completing the screening of the important data features of the machine learning model;

[0013] Step 4: Enhance the features of the cell Raman spectrum according to the important data features of the machine learning model selected by matching in step 3, and retrain the initial machine learning model; through the feature enhancement extraction of Raman spectra, the classification accuracy of the machine learning model is effectively improved.

[0014] The step 2 extracts biological peak features from the collected Raman spectra of multiple receptor proteins respectively; extracts important data features of the machine learning model from the single cell Raman spectra; the specific method is:

[0015] Step 2.1, extract biological peak features from the collected Raman spectra of multiple receptor proteins:

[0016] Step 2.1.1, data preprocessing of the Raman spectra of each receptor protein including baseline correction, silica removal and multi-spectrum averaging;

[0017] Step 2.1.2, use drawing software to search for peaks in the Raman spectra of each receptor protein: first import the Raman spectrum of the receptor protein into the drawing software, the Raman spectrum of each receptor protein is composed of multiple continuous peaks, find the maximum peak point of a single peak in the Raman spectrum of the receptor protein, determine the baseline, set the baseline to the constant minimum value, and use the drawing software to set the fitting baseline of the fitting peak to ensure that the baseline of the subsequent fitting peak is consistent with the baseline of the imported original receptor protein Raman spectrum; then add the maximum peak point position of a single peak in the Raman spectrum of the receptor protein; and align the added maximum peak position to the spectrum to avoid the error caused by the maximum peak point position of a single peak not being on the Raman spectrum of the receptor protein; set the peak search to the positive direction, that is, the maximum peak point is above the x-axis in the drawing software; perform multiple rounds of iterative operations to automatically adjust the maximum peak point position of a single peak, so as to obtain the intensity corresponding to the maximum peak point position and the Raman shift of the maximum peak point; complete the peak search;

[0018] Step 2.1.3, use drawing software to perform peak fitting on the Raman spectra of each receptor protein separately; fit a single peak according to the maximum peak point position and corresponding intensity of the single peak found by peak search, that is, split the continuously superimposed protein Raman spectra into single peaks; according to the result of peak search, set the Raman shift of the single peak to be fitted as a fixed value in the fitting control; use Gaussian fitting algorithm to perform multiple rounds of iterations and converge; obtain the peak intensity, half-peak width (FWHM) of a single peak, that is, the peak width at half the height of the spectral peak, and peak area information;

[0019] Step 2.1.4, according to the half-peak width information of the single peak obtained by peak fitting in step 2.1.3, the matching interval of the receptor protein Raman spectrum is delineated, so as to perform Raman peak feature matching later, specifically: according to the peak search, the Raman shift of the maximum peak point of each receptor protein Raman spectrum is obtained, and the width of the half-peak width on the left and right of the Raman shift of the maximum peak point of each receptor protein Raman spectrum is a matching interval delineated for one receptor protein; the matching interval is delineated for the maximum peak point of each receptor protein Raman spectrum;

[0020] Step 2.2 Extract important data features of the machine learning model from the single-cell Raman spectra related to cell typing collected in step 1:

[0021] Step 2.2.1, firstly, preprocess the collected Raman spectra of each cell including baseline correction and silicon-based removal; divide the preprocessed cell Raman spectrum data into a training set and a test set;

[0022] Step 2.2.2, input the Raman spectrum data of each cell into the machine learning model for training; use grid search to train the hyperparameters gamma, kernel, C, class_weight and probability of the machine learning model, and obtain the optimal solution of the model training result by comparing different permutations and combinations of parameters; where gamma is the kernel function coefficient, and the default value is auto; kernel is the kernel function type, and the linear kernel function is selected; C is the penalty coefficient, and the optimal value of C is determined by grid search; class_weight is set to balanced, and the weight is automatically adjusted according to the frequency of the input sample, that is, the number of Raman spectra of different cells; probability is set to False, and the test set data is verified according to the optimal hyperparameters obtained from the training set. The accuracy of the test set, the cancer cell data is input into the machine learning model to complete the initial classification;

[0023] Step 2.2.3: Input the cell Raman spectral data of several Raman shift points and their corresponding intensities into the machine learning model, where each Raman shift point represents a feature; perform importance analysis on the input data features through the machine learning model, and output important data features of the machine learning model, specifically: train the machine learning model through step 2.2.2 to obtain the cell feature coefficient. The larger the cell feature coefficient, the greater the influence of the cell feature on the cell Raman spectral classification; obtain the importance ranking of the cell features by sorting the absolute values ​​of the cell feature coefficients in ascending order, and output a preset number of important data features of the machine learning model.

[0024] There are three principles for outputting a preset number of important data features of the machine learning model in step 2.2.3: first, the location of the aggregation of the output important data features of the model cannot conflict with the location of the obvious characteristic peak of the receptor protein, that is, the observable peak with high peak intensity and wide width; second, at the location of the obvious characteristic peak of the receptor protein, the number of important data features of the model no longer increases; finally, the matching intervals defined by the receptor protein all match the important data features of the model.

[0025] The specific method of step 3 is:

[0026] Step 3.1, matching and screening the matching intervals defined for each receptor protein obtained in step 2.1.4 with the important data features output by the machine learning model, specifically: matching each matching interval of each protein with the important data features output by the machine learning model within the corresponding Raman shift, retaining the important data features of the machine learning model within each matching interval of each protein, and eliminating the important data features of the machine learning model that are not within each matching interval of each protein; counting the number of important data features of the machine learning model matched by each matching interval of each protein, and at the same time counting the total number of important data features of the machine learning model matched by each protein;

[0027] Step 3.2, matching and screening the union matching regions of each receptor protein with the important data features output by the machine learning model, specifically: taking the union of the matching intervals of the receptor proteins obtained in step 2.1.4 to obtain the union region of each receptor protein; matching each matching region of the union region of each receptor protein with the important data features output by the machine learning model within the corresponding Raman shift, retaining the important data features of the machine learning model in each matching region of the union region of each receptor protein, and eliminating the important data features of the machine learning model that are not in each matching region of the union region of each receptor protein; calculating the percentage of the important data features of the machine learning model that are successfully matched, i.e. retained, to the total number of important data features of the machine learning model to obtain the matching success rate; in addition, calculating the percentage of the number of important data features of the machine learning model that are successfully matched for each receptor protein to the total number of important data features of the machine learning model that are successfully matched for the union region of each receptor protein; calculating the percentage of the important data features of the machine learning model that are successfully matched for each region of the union region of each receptor protein to the total number of important data features of the machine learning model.

[0028] The specific method of step 4 is:

[0029] Step 4.1, design a weighting formula based on the matching results obtained in step 3.2:

[0030] W(X)=[(1+T i K)x i ](i=1,2...,9) (1)

[0031] i represents the i-th region of the protein union region obtained after taking the union of the receptor protein intervals; T i represents the ratio of the number of important data features matched in the ith region of the protein union region to the total number of important data features successfully matched; the hyperparameter K is an integer multiple of the ratio of regional features to the total number of features; x i represents the initial eigenvalue of the ith region of the protein union region, and W(X) represents the weighted eigenvalue;

[0032] Step 4.2: Enhance the characteristics of the cell Raman spectrum according to the weighted formula in step 4.1:

[0033] Scheme I: The cell Raman spectra within the receptor protein union matching region are weighted according to the weighting formula, and the cell Raman spectra outside the receptor protein union matching region are not weighted; the weighted cell Raman spectra are re-input into the machine learning model and re-trained according to step 2.2.2 to obtain the accuracy of the training set and test set corresponding to the optimal value of the hyperparameter K;

[0034] Scheme II: The cell Raman spectra within the receptor protein union matching region are weighted according to the weighting formula, and the cell Raman spectra outside the receptor protein union matching region are reset to zero; the weighted cell Raman spectra are re-input into the machine learning model and re-trained according to step 2.2.2 to obtain the accuracy of the training set and test set corresponding to the optimal value of the hyperparameter K;

[0035] Step 4.3 compares the accuracy of the test set verified in step 2.2.2, the weighted accuracy of scheme I on the test set under the optimal hyperparameter K in step 4.2, and the weighted accuracy of scheme II on the test set under the optimal hyperparameter K to obtain the optimal weighted scheme with the highest accuracy; complete the cell Raman peak feature matching.

[0036] A Raman peak feature matching system, comprising:

[0037] The Raman data acquisition module is used in step 1 to measure multiple receptor proteins and single cells related to cell typing through a Raman spectrometer to obtain Raman spectra of receptor proteins and Raman spectra of single cells.

[0038] The Raman peak feature extraction module includes a protein peak feature extraction module and a machine learning model important data feature extraction module; the protein peak feature extraction module includes a data preprocessing module, a peak search module, a peak fitting module and a matching interval delineation module; wherein the data preprocessing module is used in step 2, and the protein Raman spectrum is subjected to baseline correction, silicon base removal and multi-spectrum averaging to obtain a standard protein Raman spectrum for subsequent data analysis; the peak search module is used in step 2, and the Raman spectrum of the receptor protein composed of multiple consecutive peaks is subjected to the maximum peak search technology to obtain the maximum peak point of a single peak; the peak fitting module is used in step 2, and the single peak is fitted by the Gaussian fitting algorithm according to the maximum peak point position and corresponding intensity of the single peak found by the peak search, so as to obtain the peak intensity, half-peak width (FWHM) and peak area information of the single peak; the matching interval delineation module is used in step 2, and the matching interval delineated for each receptor protein is obtained by taking the width of the half-peak width on the left and right of the Raman shift of the maximum peak point of the Raman spectrum of each receptor protein, that is, a matching interval delineated for a receptor protein, and obtaining the matching interval delineated for each receptor protein. The important data feature extraction module of the machine learning model includes a data preprocessing module, a machine learning model training module and an important data feature output module of the machine learning model; the data preprocessing module is used in step 2, and the cell Raman spectrum is baseline corrected and silicon-based removed to obtain a standard cell Raman spectrum for subsequent data analysis; the machine learning model training module is used in step 2, and the accuracy of the initial classification of the cell Raman spectrum is obtained by inputting the cell Raman spectrum into the machine learning model training. The important data feature output module of the machine learning model is used in step 2, and the absolute values ​​of the cell feature coefficients are sorted in ascending order to obtain a preset number of important data features of the machine learning model.

[0039] The Raman peak feature matching module includes an interval matching model important data feature module and a regional matching model important feature module; the interval matching model important data feature module is used in step 3, through the matching intervals respectively delineated by the receptor proteins, and the important data features of the machine learning model are matched to obtain the matching situation of each receptor protein with the important data features of the machine learning model; the regional matching model important feature module is used in step 3, through the matching and screening of the union area of ​​the receptor proteins with the important data features of the machine learning model, to obtain the matching and screening results of the union area of ​​the receptor proteins with the important data features of the machine learning model.

[0040] The feature enhancement module includes a weighted formula module, a feature enhancement module with different weighting schemes and a feature enhancement comparison module with different weighting schemes; the weighted formula module is used in step 4 to obtain the designed weighted formula by calculating the ratio of the number of important data features matched in the protein union region to the total number of important data features successfully matched; the feature enhancement module with different weighting schemes is used in step 4 to obtain the cell Raman spectra with different weighting scheme feature enhancements by designing different weighting scheme feature enhancements; the feature enhancement comparison module with different weighting schemes is used in step 4 to obtain the weighted accuracy of the cell Raman spectra with different weighting scheme feature enhancements on the test set by retraining the machine learning model.

[0041] An electronic device for Raman peak feature matching, comprising:

[0042] Memory for storing computer programs;

[0043] A processor is used to implement the Raman peak feature matching method described in steps 1 to 3 when executing the computer program.

[0044] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can perform cell Raman peak feature matching based on the Raman peak feature matching method described in steps 1 to 3.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] 1. Raman peak feature extraction: The present invention uses peak fitting technology to realize the extraction of protein biological feature peaks, and designs a method for dividing proteins into respective matching intervals. It uses a machine learning model to output important data features to realize the extraction of cell data features, and designs a principle for outputting important data features of a preset quantity model; it can enhance spectral features and improve matching accuracy.

[0047] 2. Raman peak feature matching: The present invention innovatively proposes the idea of ​​matching protein peak features with important data features of machine learning models, matches the matching intervals divided by each protein with the important data features of the machine learning model, takes the union of the matching intervals divided by each protein to obtain the union matching area of ​​the proteins, and matches the union matching area of ​​the proteins with the important data features of the machine learning model, which effectively makes up for the fact that machine learning algorithms usually rely on data-driven pattern recognition and lack the ability to integrate biological background knowledge.

[0048] 3. Feature enhancement: Based on the results of matching the union matching region of the protein with the important data features of the machine learning model, the present invention innovatively designs a weighted formula and proposes different feature enhancement schemes to enhance the cell spectral features. The feature-enhanced cell spectra are re-input into the machine learning training, and the accuracy of different weighted schemes on the test set is compared with the accuracy of the unweighted scheme on the test set. The ideal classification results can be obtained, which greatly improves the classification accuracy.

[0049] In summary, the present invention proposes a Raman peak feature matching method to integrate the biological feature peaks of proteins with important data features in machine learning models to enhance spectral features that are critical for cell classification. The RPFM method extracts protein biological feature peaks and important data feature information from the machine learning model. It matches these protein peak features with the important data features of the machine learning model and combines these features to screen out important data features. The cell spectrum is then feature enhanced and the machine learning model is retrained to enhance the classification effect of the cell. The present invention combines important features in the model output data features and biometric features to enhance the spectral features used for analysis. The matching method utilizes the data-driven model of machine learning while making up for the defect of not being able to consider specific professional background knowledge. The RPFM method can be successfully applied to different types of machine learning models and is universal. The RPFM method provides a new method for machine learning algorithms to enhance the analysis of Raman spectral features, thereby greatly improving the accuracy and effectiveness of spectral analysis in biological and medical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 The architecture and characteristics of the RPFM method of the present invention; wherein, Figure 1 (a) The structure of RPFM, Figure 1 (b) Detailed process of extracting protein peak features and important data features of the model, Figure 1 (c) Detailed process of matching protein peak features with important data features of the model.

[0051] Figure 2 It is the extraction of biological peak features of Raman spectra of three receptor proteins, among which, Figure 2(a) Biomacromolecule diagram of ER, HER-2 and PgR proteins. Figure 2 (b) Raman spectroscopy peak search for ER, HER-2 and PgR proteins. Figure 2 (c) Raman spectrum peak fitting for ER, HER-2 and PgR proteins. Figure 2 (d) Determine the matching interval of Raman spectra of ER, HER-2 and PgR proteins.

[0052] Figure 3 The three proteins were matched with the important data features output by the SVM model, among which, Figure 3 (a) The first 200 important data features output by the SVM model, Figure 3 (b) The number of successful matches between the three proteins and the first 200 important data features of the SVM model. Figure 3 (c) Matching diagram of important data features between ER protein and SVM model, 3(d) Matching diagram of important data features between HER-2 protein and SVM model, Figure 3 (e) PgR protein meets the important data features of the SVM model. Figure 3 (f) The number of successful matches between each of the 16 matching intervals of the ER protein and the important data features. Figure 3 (g) The number of successful matches between each of the 17 matching intervals of HER-2 protein and the number of important data features, Figure 3 (h) The number of successful matches between each of the 16 matching intervals of the PgR protein and important data features.

[0053] Figure 4 The union matching region of the three proteins is matched with the important data features output by the SVM model, where: Figure 4 (a) The protein union region matches the top 200 important data features of the SVM model. Figure 4 (b) The percentage of protein union regions that successfully matched the top 200 important data features in the SVM model. Figure 4 (c) The percentage of successful matches of the three proteins to the total number of successful matches, Figure 4 (d) The union of protein regions that match the top 200 important data features of the SVM model for each region, Figure 4 (e) The percentage of successful matches in each protein union region to the total successful matches.

[0054] Figure 5 To compare different weighting schemes for enhancing and classifying the Raman spectral features of breast cells, Figure 5 (a) Comparison of the original spectrum of SK-BR-3 and the weighted spectrum of Scheme I, Figure 5 (b) Model accuracy of hyperparameter K on training set and test set from 0 to 10, Figure 5 (c) Comparison of the original spectrum of SK-BR-3 and the weighted spectrum of Scheme II, Figure 5 (d) The accuracy of different schemes on the training set and test set, Figure 5 (e) Five classification accuracy and Kappa coefficient of RPFM applied to SVM model. Figure 5 (f) Classification confusion matrix of applying RPFM to retraining SVM model.

[0055] Figure 6 To demonstrate the universality of the RPFM method in machine learning algorithms that can output important data features of the model, Figure 6 (a) Model selection for RPFM application, Figure 6 (b) The accuracy and Kappa coefficient of the five-category RPFM applied to the logistic regression model. Figure 6 (c) Classification confusion matrix of retraining the logistic regression model using RPFM. Figure 6 (d) Accuracy and Kappa coefficient of five categories of RPFM applied to XGboost model, Figure 6 (e) Classification confusion matrix of applying RPFM to the retrained XGboost model. DETAILED DESCRIPTION

[0056] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0057] A Raman peak feature matching method specifically comprises the following steps:

[0058] Step 1, collecting Raman spectra of multiple receptor proteins and multiple single-cell Raman spectra related to breast cell typing, including normal breast cells and different types of breast cancer cells;

[0059] The receptor protein Raman spectra collected in step 1 include Raman spectra of ER protein, HER-2 protein and PgR protein; the single cell Raman spectra collected include Raman spectra of a normal breast cell MCF-10A cell and four breast cancer cells MDA-MB-231 cells, BT474 cells, SK-BR-3 cells and T47D cells.

[0060] The Raman spectroscopy data collection part collects receptor protein spectra and single cell spectra related to breast cell typing. The dataset contains 34 spectra of three proteins related to breast cell typing and 1040 spectra of five breast cell lines. The protein spectrum dataset includes 13 spectra of ER protein, 11 spectra of HER-2 protein and 10 spectra of PgR protein. The cell spectrum dataset includes 207 spectra of MCF-10A, 210 spectra of MDA-MB-231, 225 spectra of BT474, 188 spectra of SK-BR-3 and 210 spectra of T47D cells. The breast cell dataset is divided into a test set and a training set, of which 728 spectra are training datasets and 312 spectra are test datasets.

[0061] All cell lines were obtained from the Cell Resource Center of Peking Union Medical College. Polymerase chain reaction (PCR) and culture confirmed that the cell lines were free of mycoplasma contamination. Cells were cultured in Dulbecco's modified Eagle's medium (DMEM). DMEM contains 1% penicillin-streptomycin solution and 10% fetal bovine serum (FBS). The cell culture environment was maintained at 37°C, 5% carbon dioxide (CO2) and 95% air. Cells were washed three times with phosphate-buffered saline (PBS) before use. After each wash, the cells were spun at 1000 rpm for 3 minutes to ensure uniform cell distribution.

[0062] Protein spectra and cell spectra were tested under the same conditions. Single cells or proteins were placed on a silicon chip and their spectra were acquired using a Raman spectrometer (Finder One, Zhuoli, China). The spectrometer consisted of a multimode diode laser with an excitation wavelength of 532 nm and a back-illuminated, deep depletion, Peltier-cooled CCD detector. The numerical aperture (NA) of the objective lens was 0.65 and the working distance was 40. The integration time for each spectrum was set to 5 seconds.

[0063] Step 2: Extract biological peak features from the receptor protein Raman spectrum collected in step 1, and extract important data features of the machine learning model from the single cell Raman spectrum:

[0064] Step 2.1, extract biological peak features from the collected Raman spectrum of the receptor protein:

[0065] Step 2.1.1, data preprocessing of protein Raman spectra including baseline correction, silica removal and multi-spectrum averaging for subsequent data analysis;

[0066] Step 2.1.2, use the drawing software Origin 2022 to search for peaks of the three proteins ER protein, HER-2 protein and PgR protein, and find the maximum peak point of the protein Raman spectrum with a single peak; then search for peaks:

[0067] Import the Raman spectrum of the protein into the software and select peak analysis. First, you need to determine the baseline, set the baseline to the constant minimum value, and set the baseline to fit the peak at the same time to ensure that the baseline of the fitted peak is consistent with the baseline of the imported original protein Raman spectrum; manually double-click the mouse to add the maximum peak point position of a single peak in the protein Raman spectrum; and set it to align to the spectrum to avoid the error caused by the maximum peak point position of the manually added single peak not being on the protein spectrum. When searching for peaks, set the peak search method to local maximum, and find the maximum value point near the maximum peak point position of the added single peak. The maximum peak point of the manually added single peak is not necessarily the true maximum point. The peak search is set to positive, that is, the opening is upward; the local point uses the default setting of 2, and the peak screening method uses the default setting of 20% by height. The peak fitting limit uses the default value, and multiple rounds of iterations are performed according to the above settings to automatically adjust the maximum peak point position of a single peak of the protein, so as to obtain the intensity corresponding to the maximum peak point position; complete the peak search;

[0068] Step 2.1.3, use the drawing software Origin 2022 to perform peak fitting on the three proteins ER protein, HER-2 protein and PgR protein; fit a single peak according to the maximum peak point position and corresponding intensity of the single peak found by peak searching, that is, split the continuously superimposed protein Raman spectrum into single peaks; according to the results of peak searching, set the Raman shift of the single peak to be fitted to a fixed value in the fitting control; the core choice of peak fitting is to use the Gaussian fitting algorithm to perform multiple rounds of iterations while converging; the final fitting is performed in a convergent state, that is, the iteration of the algorithm is convergent. The above-mentioned peak fitting can provide useful information such as the peak intensity, half-peak width (FWHM) and peak area of ​​a single peak; the half-peak width refers to the peak width at half the height of the spectral peak;

[0069] Step 2.1.4, according to the half-peak width information of the single peak obtained by the peak fitting in 2.3 above, the matching intervals of the three proteins ER protein, HER-2 protein and PgR protein are delineated, so as to perform Raman peak feature matching later; the Raman shift of the maximum peak point of the Raman spectrum of each protein can be obtained by peak searching, and the width of the half-peak width on the left and right of the Raman shift of the maximum peak point of the Raman spectrum of each protein is a matching interval delineated for one protein; the matching interval is delineated for the maximum peak point of the Raman spectrum of each protein;

[0070] Step 2.2 Extract important data features of the machine learning model for single-cell Raman spectroscopy:

[0071] Step 2.2.1, firstly, the breast cell Raman spectrum is also preprocessed including baseline correction and silicon base removal; the preprocessed breast cell Raman spectrum data is divided into a training set and a test set in a ratio of 7:3;

[0072] Step 2.2.2, input the breast cell Raman spectroscopy data into the SVM model for training; use grid search to train the hyperparameters of the SVM model, gamma, kernel, C, class_weight and probability, and obtain the optimal solution for the model training results by comparing different permutations and combinations of parameters; where gamma is the kernel function coefficient, and the default value is auto; kernel is the kernel function type, and the linear kernel function is selected; C is the penalty coefficient, and the larger the C, the lower the tolerance of the model classifier to noise points within the boundary, the higher the classification accuracy, but it is prone to overfitting and poor generalization ability. The optimal value of C, 0.05, is determined by grid search; class_weight is set to balanced, and the weight is automatically adjusted according to the frequency of the input sample (i.e., the number of Raman spectra of different breast cells); probability is set to False, and the test set data is verified according to the optimal hyperparameters obtained from the training set, and the breast cancer cell data is input into the SVM model to complete the initial classification;

[0073] Step 2.2.3 outputs the important data features of the SVM model. The breast cell Raman spectroscopy data input to the SVM model is 2000 Raman shift points and their corresponding intensities, and each Raman shift point represents a feature; the importance analysis of the data features of the SVM model is used to understand the usefulness or value of each feature (variable or input) for classification. The purpose is to find the most important features that have the greatest impact on the model output. When the model achieves good results on the test set, the SVM machine learning model can output the important data features of the model during the training process. For linear SVM, specifically: through the training of the machine learning model, the breast cell feature coefficient is obtained, and the importance of the breast cell feature can be measured by the calculated breast cell feature coefficient. The larger the breast cell feature coefficient, the greater the impact of the breast cell feature on the breast cell Raman spectroscopy classification; by sorting the absolute values ​​of the breast cell feature coefficients in ascending order, the importance ranking of the breast cell features can be obtained, and the important data features of the preset number of SVM models can be output.

[0074] The step 2.2.3 outputs a preset number of important data features of the SVM model; there are three principles for outputting a preset number of important data features of the model. First, the location of the aggregation of the output important data features of the model cannot be inconsistent with the location of the obvious characteristic peaks of the three proteins ER protein, HER-2 protein and PgR protein. Secondly, at the location of the obvious characteristic peaks of the three proteins, the number of important data features of the model no longer increases. Finally, the matching intervals defined by the three proteins all match the important data features of the model.

[0075] Three proteins are associated with intrinsic subtypes of breast cancer. Intrinsic subtypes of breast cancer can be estimated based on the expression of estrogen receptor (ER), progesterone receptor (PgR), or human epidermal growth factor receptor 2 (HER-2) assessed by immunohistochemistry (IHC) or in situ hybridization (ISH). Based on the expression of these biomarkers and Ki-67 levels, breast cancer can be divided into four intrinsic subtypes: ductal A, ductal B, Erb-B2 overexpressing, and "basal-like". In order to visualize the biomacromolecular maps of the three proteins associated with breast cell typing, the amino acid sequences of the three proteins were consulted and found in the PDB protein database. Figure 2 a The three-dimensional structures of ER, HER-2 and PgR proteins are displayed using pymol software. After collecting protein Raman spectroscopy data, data preprocessing is required, including baseline correction, silica removal and multi-spectrum averaging. Raman spectra of the three proteins after data preprocessing. It can be seen that the peak characteristics of the three proteins are at 1005.1cm -1 、1453.3cm -1 and 1672.8cm -1 The most obvious place.

[0076] The three protein spectra were extracted through peak finding, peak fitting and matching interval demarcation. First, the Raman spectra of the three proteins need to be peak found. The protein Raman spectrum peak fitting program is performed using the Origin 2022 software program. Select 559.89cm -1 To 1881.46cm -1 The peaks in the region are fitted. The baseline mode is set to the constant minimum value, and the baseline is fitted while fitting the peak to ensure that the baseline of the later fitted peak is consistent with the baseline of the original spectrum imported by the protein. Manually double-click the mouse to add the maximum peak point position that looks like a single peak in the protein Raman spectrum. And set it to align to the spectrum to avoid the error caused by the maximum peak point position of the manually added single peak not being on the protein spectrum. When searching for peaks, the peak search method is set to local maximum, and the maximum value point near the maximum peak point position of the added single peak is found. The maximum peak point of the manually added single peak is not necessarily the true maximum point. The peak search is set to positive, that is, the opening is upward. The local point is set to 2 by default, and the peak screening method is set to 20% by height by default. The peak fitting limit uses the default value. Through the above settings, multiple rounds of iterations are performed to automatically adjust the maximum peak point position of a single peak of the protein, so as to obtain the intensity corresponding to the maximum peak point position. The peak search is completed. The peak search results of the three proteins are shown as follows Figure 2As shown in b. Fit a single peak based on the maximum peak point position and corresponding intensity of a single peak found by peak search, that is, split the continuously superimposed protein Raman spectrum into single peaks. According to the results of peak search, set the Raman shift of the single peak to be fitted in the fitting control. The core choice of peak fitting is to perform multiple rounds of iterations of the Gaussian fitting algorithm. The final fitting is performed in a convergent state, that is, the iteration of the algorithm is convergent. After the fitting is completed, a peak fitting report can be obtained. The peak fitting results of the three proteins are shown in Figure 2 c. The black spectrum represents the original spectrum, the green spectrum represents the result after fitting each peak according to the peak search results, and the red spectrum represents the cumulative fitting effect. The closer the red and black spectra are, the better the fitting effect. The evaluation index is the adjusted R 2 The adjusted R 2 The closer it is to 1, the better the fitting effect. Adjustment R of ER, HER-2 and PgR protein peak fitting 2 They are 0.89, 0.98 and 0.93 respectively. The overall fitting results are good, all reaching above 0.85. Peak fitting can provide useful information such as peak intensity, half-peak width (FWHM) and peak area of ​​a single peak. Half-peak width refers to the peak width at half the height of the spectral peak. The matching interval is defined by a half-peak width on both sides of the peak point. The results of dividing the matching interval for the three proteins are shown in the figure below. Figure 2 d. The green area represents the matching interval defined by the ER protein, with a total of 16 matching intervals. The blue area represents the matching interval defined by the HER-2 protein, with a total of 17 matching intervals. The purple area represents the matching interval defined by the PgR protein, with a total of 16 matching intervals. The characteristic information of the three proteins is fully utilized here, laying the foundation for matching.

[0077] Before extracting the important data features of the model, it is necessary to perform data preprocessing and model training on the mammary cell spectrum. The present invention collects Raman spectra of five mammary cells, including a normal mammary cell MCF-10A and four breast cancer cells. The four breast cancer cells are T47D, MDA-MB-231, BT474 and SK-BR-3. The five breast cancer cells have completed the data preprocessing of baseline correction and silicon-based removal. The data set is divided into a training set and a test set according to a ratio of 7:3. The training set is used for model training. The test set is used to test the final result of the model. It is not used for model training but as a simulated sample in practical applications to directly verify the classification effect of the model. The training set is divided into ten different validation sets. Nine groups of samples continue to be used as training sets, and the remaining group of samples is used as a validation set to test the model performance. The model will be trained ten times, and after ten-fold cross-validation, the average value of the ten model results will be used as the result value of the model. When the model is adjusted for hyperparameters, a parameter space will be given, that is, the value range of different parameters of the model. Grid search is a parameter exhaustive method. It will give all permutations and combinations of parameters in the user-specified parameter space and perform brute force tests. The model will be trained for each parameter and the optimal solution will be selected by comparing the model results obtained for each parameter. Although grid search will take a lot of time and cost, it can avoid significant result errors caused by artificially giving one or more parameter combinations. Classification accuracy is used as an evaluation indicator during training. The SVM model is a simple linear classification model. The SVM model is used to classify five types of breast cells into five categories. The main hyperparameters of the SVM model are gamma, kernel, C, class_weight, and probability. Gamma is the kernel function coefficient, and the default value is auto. Kernel is the kernel function type, and the linear kernel function is selected. C is the penalty coefficient. The larger C is, the lower the tolerance of the classifier to noise points within the boundary, the higher the classification accuracy, but it is easy to overfit and has poor generalization ability. The optimal value of 0.05 is determined by grid search. The class_weight is set to balanced, and then the weight is automatically adjusted according to the frequency of the input sample. Probability is set to False, which determines that each possible probability will be output according to the probability in the end. The accuracy of the training set is 89.74%. The model trained on the training set was applied to the test set and the accuracy on the test set was 88.78%.

[0078] Breast cancer cell data is input into the SVM model for initial classification, and the extraction of important data features is completed through the model. The SVM model can output the important data features of the model during the training process. For linear SVM, the importance of features can be measured by coefficients. The larger the coefficient, the greater the contribution of the feature to the classification. The importance ranking of the features can be obtained by sorting the absolute values ​​of the coefficients in ascending order. When the model achieves good results on the test set, it will output the ranking of the important data features of the model. This study outputs different numbers of important data features of the SVM model during the training process, and each red vertical line represents an important data feature of the model at the Raman shift point. It can be found that with the increase in the number of output model important data features, the aggregation band of the model important data features becomes more and more obvious at the characteristics of the three protein peaks. Most of the model important data features are consistent with the characteristics of the three protein peaks, but a few are inconsistent. In order to select an appropriate number of model important data features for matching, the changes in the number of important data features of the three Raman shift points were statistically analyzed. There are three principles for outputting an appropriate number of model important data features. First, the location of the aggregation of the output model important data features cannot be contrary to the location of the obvious characteristic peaks of the three proteins. Secondly, at the locations of the obvious characteristic peaks of the three proteins, does the number of important data features of the model no longer increase? Finally, do the matching intervals defined by the three proteins match the important data features of the model? The first 200 important data features output by the SVM model match the matching intervals determined by the three protein peak features. The first 200 important data features output by the SVM model are as follows: Figure 3 As shown in a.

[0079] Step 3: Match the biological peak features extracted in step 2 with the important data features of the machine learning model, and complete the screening of important data features of the machine learning model:

[0080] Step 3.1, match and screen the matching intervals of ER protein, HER-2 protein and PgR protein obtained in step 2.1.4 with the important data features output by the SVM model respectively:

[0081] Each matching interval of each protein is matched with the important data features output by the SVM model within the corresponding Raman shift, and the matching result is to retain the important data features of the SVM model within each matching interval of each protein, and eliminate the important data features of the SVM model that are not within each matching interval of each protein; the number of important data features of the SVM model matched by each matching interval of each protein is counted, and the total number of important data features of the SVM model matched by each protein is counted;

[0082] Step 3.2, matching and screening the union matching region of ER protein, HER-2 protein and PgR protein with the important data features output by the SVM model:

[0083] In reality, breast cells contain ER protein, HER-2 protein and PgR protein at the same time. This is the result of the combined action of these three proteins. The only difference is the content of these three proteins. Therefore, we should not only consider the matching intervals defined by the three proteins and match them with the important data features output by the SVM model. The matching intervals obtained in step 2.1.4 for ER protein, HER-2 protein and PgR protein are taken as the union, so as to obtain the union region of ER protein, HER-2 protein and PgR protein; each matching region of the union region of ER protein, HER-2 protein and PgR protein is matched with the important data features output by the SVM model within the corresponding Raman shift, the important data features of the SVM model within each matching region of the union region of ER protein, HER-2 protein and PgR protein are retained, and the important data features of the SVM model that are not within each matching region of the union region of ER protein, HER-2 protein and PgR protein are eliminated; while matching, the important data features of the SVM model output that are not within the region are eliminated, so that the screening of the important data features output by the SVM model is completed while matching the union region of ER protein, HER-2 protein and PgR protein with the important data features output by the SVM model. The percentage of the important data features of the SVM type that are successfully matched, i.e. retained, to the total number of important data features of the SVM model is calculated to obtain the matching success rate; in addition, the percentage of the number of important data features of the SVM model that are successfully matched by each of the ER protein, HER-2 protein, and PgR protein to the total number of important data features of the SVM model that are successfully matched by the union region of the ER protein, HER-2 protein, and PgR protein is calculated respectively. It is also necessary to calculate the percentage of the important data features of the SVM model that are successfully matched in each region of the union region of the ER protein, HER-2 protein, and PgR protein to the total number of important data features of the SVM model that are successfully matched by the union region of the ER protein, HER-2 protein, and PgR protein.

[0084] ER, HER-2, and PgR proteins were matched with the top 200 significant data features output by the SVM model, respectively. Figure 3 The radar plot in b shows the total number of successful matches of each of the three proteins to the top 200 significant data features output by the SVM model. Figure 3 c shows the matching results of 16 matching intervals delineated by ER proteins and the top 200 important data features output by the SVM model. Figure 3 d shows the matching results of 17 matching intervals delineated by HER-2 protein and the first 200 important data features output by the SVM model. Figure 3 e shows the matching results of the 16 matching intervals defined by PgR protein and the first 200 important data features output by the SVM model. Each matching interval of each protein is matched with the important data features output by the SVM model within the corresponding Raman shift. The result of the matching is that the important data features of the SVM model within each matching interval of each protein are retained, while the important data features of the SVM model outside each matching interval of each protein are eliminated. At the same time, the number of important data features matched by each matching interval of each protein is counted. The statistical results are shown in Figure 3 f to Figure 3 h.

[0085] Normally, breast cells contain three proteins, which are the result of the joint action of the three proteins. Therefore, it is necessary to take the union of the protein intervals to match the first 200 important data features output by the SVM model. The matching intervals of the three proteins are merged to obtain the union region of the proteins. Figure 4 a shows the matching results of the protein union region and the first 200 important data features output by the SVM model. It can be seen that the important data features that are concentrated on the peak features of the three proteins in the SVM model are basically retained, and only a few important data features of the SVM model that do not match the peak features of the three proteins are eliminated. In addition to proteins, cells also have lipids, nucleic acids, carbohydrates and other substances. Figure 4 b shows the percentage of the model important data features that were successfully matched in the first 200 important data features output by the SVM model. The matching success rate is 94.5%, which also shows that it is reasonable to select three proteins related to breast cell typing in the model important data feature screening. In addition, the percentage of the number of model important data features that were successfully matched by each of the three proteins and the total number of model important data features that were successfully matched by the union of the protein regions was calculated. Figure 4 c It can be seen that the percentages of ER protein and PgR protein are relatively high. Figure 4 d Statistics of the matching results between the union of protein regions and the top 200 important data features output by the SVM model. Figure 4 e shows the percentage of the number of matches between each protein region union and the important data features of the model to the total number of successful matches.

[0086] Step 4: Enhance the features of the breast cell Raman spectrum according to the important data features of the machine learning model selected by matching in step 3, and retrain the initial machine learning model; through the feature enhancement extraction of the Raman spectrum, the classification accuracy of the machine learning model is effectively improved:

[0087] Step 4.1, design a weighted formula based on the percentage of the total number of important data features of the SVM model that are successfully matched in each region of the important data feature union region of the matching result SVM model calculated in step 3.2, and enhance the features of the breast cell Raman spectrum according to the designed weighted formula. The following is a weighted formula designed based on the different matching success rates of each region in the union region of the three proteins:

[0088] W(X)=[(1+T i K)x i ](i=1,2...,9) (1)

[0089] i represents the i-th region of the protein union region obtained by taking the union of the ER protein, HER-2 protein and PgR protein intervals; T i represents the ratio of the number of important data features matched in the ith region of the protein union region to the total number of important data features successfully matched; an integer multiple of the hyperparameter K (ratio of regional features to total number of features); x i represents the initial eigenvalue of the ith region of the protein union region, and W(X) represents the weighted eigenvalue.

[0090] Step 4.2: Design different weighting schemes according to the weighting formula in step 4.1 to enhance the characteristics of the Raman spectrum of breast cells:

[0091] Scheme I: Weighted scheme, the breast cell Raman spectra within the protein union matching region are weighted according to the weighted formula, and the breast cell Raman spectra outside the protein union matching region are not weighted; the breast cell Raman spectra weighted according to Scheme I are re-input into the SVM model and re-trained according to step 2.2.2 to find the optimal value of the hyperparameter K and the corresponding accuracy of the training set and test set;

[0092] Scheme II: The breast cell Raman spectra within the protein union matching region are weighted according to the weighting formula, and the breast cell Raman spectra outside the protein union matching region are reset to zero; the breast cell Raman spectra weighted according to Scheme II are re-input into the SVM model and re-trained according to step 2.2.2 to obtain the accuracy of the training set and the test set corresponding to the optimal value of the hyperparameter K;

[0093] Step 4.3 compares the accuracy without weighting on the test set, the accuracy weighted on the test set by scheme I under the optimal hyperparameter K, and the accuracy weighted on the test set by scheme II under the optimal hyperparameter K to obtain the optimal weighting scheme (the one with the highest accuracy); calculate and display the confusion matrix and Kappa coefficient of breast cell Raman spectrum classification under the optimal weighting scheme to complete the breast cell Raman peak feature matching.

[0094] According to the designed weighting formula, different weighting schemes were tried to obtain ideal classification results after feature enhancement. For the weighting of scheme I, the spectra within the nine matching regions were weighted according to the designed weighting formula, and the spectra outside the matching regions were not weighted. Figure 5 Shown in a is the comparison spectrum of SK-BR-3 before and after weighting using Scheme I. Figure 5 b shows the optimal hyperparameter K for the weighting method of Scheme I. After enhancing the spectral features, the SVM model is retrained. When K = 7, the accuracy of the training set and the test set is the highest, and the accuracy of the test set is 97.12%. For the weighting of Scheme II, the spectra within the nine matching regions are weighted according to the designed weighting formula, and the spectra outside the matching region are zeroed. Figure 5 c shows the comparison spectra of SK-BR-3 before and after weighting in Scheme II weighting method. The accuracy of Scheme II weighting method on the test set is 94.55%. Figure 5 d Comparing the two weighting schemes, it can be found that the weighted method of Scheme I is better. The accuracy of the weighted method of Scheme I on the test set is 8.34% higher than that of the unweighted method. The accuracy of the weighted method of Scheme II is lower than that of the weighted method of Scheme I. This may be due to the loss of some fingerprint information by normalizing the spectra outside the matching area to zero. Figure 5 e shows the accuracy and Kappa coefficient of the five categories under the optimal weighting method. The accuracy of each category reached 90%, and the accuracy of four categories exceeded 98%. The Kappa coefficient of each category was between 0.81 and 1, indicating that the consistency test results were basically consistent. Figure 5 f shows the confusion matrix of the five categories under the optimal weighting method. After the RPFM method and feature enhancement, the classification accuracy has been greatly improved.

[0095] The RPFM method is universal in machine learning algorithms that can output important data features of the model. The RPFM method was first verified on the simplest linear SVM model. In addition to linear machine learning models, there are generalized linear machine learning models and tree machine learning models. The generalized linear classification model was selected for verification on the logistic regression model. The tree classification model was selected for verification on the XGboost model. These three different types of classification models are designed to improve the accuracy of machine learning model classification after the RPFM method and feature enhancement.

[0096] RPFM was applied to the generalized linear logistic regression model. The breast cell Raman spectroscopy data was input into the logistic regression model for training. An appropriate number of important data features of the logistic regression model were output. The matching intervals delineated by the three proteins were matched with the important data features output by the logistic regression model. The union matching region of the three proteins was matched with the important data features output by the logistic regression model. Weighted scheme I and weighted scheme II were tried to enhance the features of breast cell Raman spectra, and the breast cell Raman spectra with enhanced features were input into the logistic regression model for retraining. The accuracy of the unweighted test set, the accuracy of the weighted scheme I on the test set under the optimal hyperparameter K, and the accuracy of the weighted scheme II on the test set under the optimal hyperparameter K were obtained. The optimal weighting scheme was obtained. The confusion matrix and Kappa coefficient of breast cell Raman spectroscopy classification under the optimal weighting scheme were calculated and displayed. After the RPFM method and feature enhancement, the accuracy of the logistic regression model for breast cell classification was also improved.

[0097] RPFM was applied to the tree-type XGboost model. The breast cell Raman spectroscopy data was input into the XGboost model for training. An appropriate number of important data features of the XGboost model were output. The matching intervals delineated by the three proteins were matched with the important data features output by the XGboost model. The union matching region of the three proteins was matched with the important data features output by the XGboost model. Weighted scheme I and weighted scheme II were tried to enhance the features of breast cell Raman spectroscopy, and the breast cell Raman spectra after feature enhancement were input into the XGboost model for retraining. The accuracy of the unweighted test set, the accuracy of the weighted scheme I on the test set under the optimal hyperparameter K, and the accuracy of the weighted scheme II on the test set under the optimal hyperparameter K were obtained. The optimal weighted scheme was obtained. The confusion matrix and Kappa coefficient of breast cell Raman spectroscopy classification under the optimal weighted scheme were calculated and displayed. After the RPFM method and feature enhancement, the accuracy of the XGboost model for breast cell classification was also improved.

[0098] The applicability of the RPFM method to other machine learning algorithms was evaluated, for which important data features of the model can be output. Figure 6 As shown in a. The machine learning algorithm is ultimately used to complete the classification task. The RPFM method was first verified on the simplest linear SVM model. In addition to the linear machine learning model, there are also generalized linear machine learning models and tree machine learning models. The generalized linear classification model was selected to be verified on the logistic regression model. The tree classification model was selected to be verified on the XGboost model. These three different types of classification models are designed to improve the accuracy of model classification after the RPFM method and feature enhancement.

[0099] Extract important data features of the generalized linear logistic regression model. Apply the generalized linear logistic regression model to classify breast cells into five categories. The accuracy of the training set is 92.84%, and the accuracy of the test set is 90.06%. For generalized linear logistic regression, each feature corresponds to a model parameter. The larger the parameter, the greater the impact of the feature on the model prediction results, and the more important the feature is. Output the appropriate number of important data features of the logistic regression model. ER, HER-2, and PgR proteins are matched with the important data features output by the logistic regression model. Figure 6 b shows the accuracy and Kappa coefficient of the five categories under the optimal weighting method. The accuracy of each category is 89%, and the accuracy of four categories reaches 95% or above. The Kappa coefficient of each category is between 0.81-1, indicating that the consistency test results are almost consistent. Figure 6 c shows the confusion matrix of the five categories under the optimal weighting method. It can be found that after the RPFM method and feature enhancement, the accuracy of the logistic regression model for the five classifications of breast cells has also been greatly improved.

[0100] A similar processing step was used when using the tree-based XGboost model, just as the RPFM approach was used for the generalized linear logistic regression model. The tree-based XGboost model was used for the first five-class classification of breast cells. The accuracy was 94.86% on the training set and 88.14% on the test set. For the XGboost tree, after creating the boosted tree, it is relatively straightforward to obtain the importance score for each attribute using the gradient boosting algorithm. The importance score measures the value of the feature in building the boosted decision tree in the model. The more attributes in the model are used to build the decision tree, the higher its importance. Attribute importance is calculated by calculating and ranking each attribute in the data set. In a single decision tree, attribute importance is calculated by the amount of performance indicator improvement for each attribute split point. Nodes are responsible for weighting and recording the number of times, that is, for the split point improvement performance indicator, the larger the attribute (the closer to the root node), the larger the weight; the more boosted trees are selected, the more important the attribute is. The performance indicator can be purity (Gini index) for selecting split points, or other more specific error functions. Finally, the results of an attribute in all boosted trees are weighted and summed, and then averaged to obtain an importance score. Output an appropriate number of important data features of the XGboost model. ER, HER-2, and PgR proteins are matched with the important data features output by the XGboost model respectively. Figure 6 d shows the accuracy and Kappa coefficient of the five categories under the optimal weighting method. The accuracy of each category reached 85%, and the accuracy of three categories exceeded 92%. The Kappa coefficient of each category was between 0.81 and 1, indicating that the consistency test results were almost consistent. Figure 6e shows the confusion matrix of the five-class classification under the optimal weighting method. It can be found that after the RPFM method and feature enhancement, the accuracy of the XGboost model for the five-class classification of breast cells has also been greatly improved.

Claims

1. A Raman peak feature matching method, characterized in that: The specific steps include: Step 1, collecting Raman spectra of multiple receptor proteins and multiple single-cell Raman spectra related to cell typing, including normal cells and different types of cancer cells; Step 2, extracting biological peak features from the receptor protein Raman spectrum collected in step 1, and extracting important data features of the machine learning model from the single cell Raman spectrum; Step 3, matching the biological peak features extracted in step 2 with the important data features of the machine learning model, and completing the screening of the important data features of the machine learning model; Step 4: Enhance the features of the cell Raman spectrum according to the important data features of the machine learning model selected by matching in step 3, and retrain the initial machine learning model; through the feature enhancement extraction of Raman spectra, the classification accuracy of the machine learning model is effectively improved.

2. A Raman peak feature matching method according to claim 1, characterized in that: The step 2 extracts biological peak features from the collected Raman spectra of multiple receptor proteins respectively; extracts important data features of the machine learning model from the single cell Raman spectra; The specific method is: Step 2.1, extract biological peak features from the collected Raman spectra of multiple receptor proteins: Step 2.1.1, data preprocessing of the Raman spectra of each receptor protein including baseline correction, silica removal and multi-spectrum averaging; Step 2.1.2, use drawing software to search for peaks in the Raman spectra of each receptor protein: first import the Raman spectrum of the receptor protein into the drawing software. The Raman spectrum of each receptor protein is composed of multiple continuous peaks. Find the maximum peak point of a single peak of the receptor protein Raman spectrum, determine the baseline, set the baseline to a constant minimum value, and use the drawing software to set the fitting baseline of the fitting peak to ensure that the baseline of the subsequent fitting peak is consistent with the baseline of the imported original receptor protein Raman spectrum; Then, the maximum peak point position of a single peak in the Raman spectrum of the receptor protein is added; and the added maximum peak position is aligned to the spectrum to avoid the error caused by the maximum peak point position of a single peak not being on the Raman spectrum of the receptor protein; the peak search is set to the positive direction, that is, the maximum peak point is above the x-axis in the drawing software; multiple rounds of iterative operations are performed to automatically adjust the maximum peak point position of a single peak, so as to obtain the intensity corresponding to the maximum peak point position and the Raman shift of the maximum peak point; and the peak search is completed; Step 2.1.3, use drawing software to perform peak fitting on the Raman spectra of each receptor protein separately; fit a single peak according to the maximum peak point position and corresponding intensity of the single peak found by peak search, that is, split the continuously superimposed protein Raman spectra into single peaks; according to the result of peak search, set the Raman shift of the single peak to be fitted as a fixed value in the fitting control; use Gaussian fitting algorithm to perform multiple rounds of iterations and converge; obtain the peak intensity, half-peak width (FWHM) of a single peak, that is, the peak width at half the height of the spectral peak, and peak area information; Step 2.1.4, according to the half-peak width information of the single peak obtained by peak fitting in step 2.1.3, the matching interval of the receptor protein Raman spectrum is delineated, so as to perform Raman peak feature matching later, specifically: according to the peak search, the Raman shift of the maximum peak point of each receptor protein Raman spectrum is obtained, and the width of the half-peak width on the left and right of the Raman shift of the maximum peak point of each receptor protein Raman spectrum is a matching interval delineated for one receptor protein; the matching interval is delineated for the maximum peak point of each receptor protein Raman spectrum; Step 2.2 Extract important data features of the machine learning model from the single-cell Raman spectra related to cell typing collected in step 1: Step 2.2.1, firstly, preprocess the collected Raman spectra of each cell including baseline correction and silicon-based removal; divide the preprocessed cell Raman spectrum data into a training set and a test set; Step 2.2.2, input the Raman spectrum data of each cell into the machine learning model for training; use grid search to train the hyperparameters gamma, kernel, C, class_weight and probability of the machine learning model, and obtain the optimal solution of the model training result by comparing different permutations and combinations of parameters; where gamma is the kernel function coefficient, and the default value is auto; kernel is the kernel function type, and the linear kernel function is selected; C is the penalty coefficient, and the optimal value of C is determined by grid search; class_weight is set to balanced, and the weight is automatically adjusted according to the frequency of the input sample, that is, the number of Raman spectra of different cells; probability is set to False, and the test set data is verified according to the optimal hyperparameters obtained from the training set. The accuracy of the test set, the cancer cell data is input into the machine learning model to complete the initial classification; Step 2.2.3: Input the cell Raman spectral data of several Raman shift points and their corresponding intensities into the machine learning model, where each Raman shift point represents a feature; perform importance analysis on the input data features through the machine learning model, and output important data features of the machine learning model, specifically: train the machine learning model through step 2.2.2 to obtain the cell feature coefficient. The larger the cell feature coefficient, the greater the influence of the cell feature on the cell Raman spectral classification; obtain the importance ranking of the cell features by sorting the absolute values ​​of the cell feature coefficients in ascending order, and output a preset number of important data features of the machine learning model.

3. A Raman peak feature matching method according to claim 2, characterized in that: There are three principles for outputting a preset number of important data features of the machine learning model in step 2.2.3: first, the location of the aggregation of the output important data features of the model cannot conflict with the location of the obvious characteristic peak of the receptor protein, that is, the observable peak with high peak intensity and wide width; second, at the location of the obvious characteristic peak of the receptor protein, the number of important data features of the model no longer increases; finally, the matching intervals defined by the receptor protein all match the important data features of the model.

4. The Raman peak feature matching method according to claim 1, characterized in that: The specific method of step 3 is: Step 3.1, matching and screening the matching intervals defined for each receptor protein obtained in step 2.1.4 with the important data features output by the machine learning model, specifically: matching each matching interval of each protein with the important data features output by the machine learning model within the corresponding Raman shift, retaining the important data features of the machine learning model within each matching interval of each protein, and eliminating the important data features of the machine learning model that are not within each matching interval of each protein; Count the number of important data features of the machine learning model matched by each matching interval of each protein, and count the total number of important data features of the machine learning model matched by each protein; Step 3.2, matching and screening the union matching region of each receptor protein with the important data features output by the machine learning model, specifically: taking the union of the matching intervals of the receptor proteins obtained in step 2.1.4 to obtain the union region of each receptor protein; matching each matching region of the union region of each receptor protein with the important data features output by the machine learning model within the corresponding Raman shift, retaining the important data features of the machine learning model within each matching region of the union region of each receptor protein, and eliminating the important data features of the machine learning model that are not within each matching region of the union region of each receptor protein; The matching success rate is calculated by calculating the percentage of the important data features of the machine learning model that are successfully matched, i.e. retained, to the total number of important data features of the machine learning model. In addition, the percentage of the number of important data features of the machine learning model that are successfully matched to each receptor protein and the total number of important data features of the machine learning model that are successfully matched to the union region of each receptor protein is calculated respectively. The percentage of the important data features of the machine learning model that are successfully matched to each region of the union region of each receptor protein is calculated to the total number of important data features of the machine learning model.

5. The Raman peak feature matching method according to claim 1, characterized in that: The specific method of step 4 is: Step 4.1, design a weighting formula based on the matching results obtained in step 3.2: W(X)=[(1+T i K)x i ](i=1,2...,9) (1) i represents the i-th region of the protein union region obtained after taking the union of the receptor protein intervals; T i represents the ratio of the number of important data features matched in the ith region of the protein union region to the total number of important data features successfully matched; the hyperparameter K is an integer multiple of the ratio of regional features to the total number of features; x i represents the initial eigenvalue of the ith region of the protein union region, and W(X) represents the weighted eigenvalue; Step 4.2: Enhance the characteristics of the cell Raman spectrum according to the weighted formula in step 4.1: Scheme I: The cell Raman spectra within the receptor protein union matching region are weighted according to the weighting formula, and the cell Raman spectra outside the receptor protein union matching region are not weighted; the weighted cell Raman spectra are re-input into the machine learning model and re-trained according to step 2.2.2 to obtain the accuracy of the training set and test set corresponding to the optimal value of the hyperparameter K; Scheme II: The cell Raman spectra within the receptor protein union matching region are weighted according to the weighting formula, and the cell Raman spectra outside the receptor protein union matching region are reset to zero; the weighted cell Raman spectra are re-input into the machine learning model and re-trained according to step 2.2.2 to obtain the accuracy of the training set and test set corresponding to the optimal value of the hyperparameter K; Step 4.3 compares the accuracy of the test set verified in step 2.2.2, the weighted accuracy of scheme I on the test set under the optimal hyperparameter K in step 4.2, and the weighted accuracy of scheme II on the test set under the optimal hyperparameter K to obtain the optimal weighted scheme with the highest accuracy; complete the cell Raman peak feature matching.

6. A system according to any one of claims 1 to 5, characterized in that: include: A Raman data acquisition module is used in step 1 to measure multiple receptor proteins and single cells related to cell typing through a Raman spectrometer to obtain Raman spectra of receptor proteins and single cells; The Raman peak feature extraction module includes a protein peak feature extraction module and a machine learning model important data feature extraction module; the protein peak feature extraction module includes a first data preprocessing module, a peak search module, a peak fitting module and a matching interval delineation module; wherein the first data preprocessing module is used in step 2, and the protein Raman spectrum is subjected to baseline correction, silicon base removal and multi-spectrum averaging to obtain a standard receptor protein Raman spectrum for subsequent data analysis; the peak search module is used in step 2, and the Raman spectrum of the receptor protein composed of multiple consecutive peaks is subjected to maximum peak search technology to obtain the maximum peak point of a single peak; the peak fitting module is used in step 2, and a single peak is fitted by a Gaussian fitting algorithm according to the maximum peak point position and corresponding intensity of a single peak found by peak search to obtain the peak intensity, half-peak width (FWHM) and peak area information of a single peak; the matching interval delineation module is used in step 2, and the peak intensity, half-peak width (FWHM) and peak area information of a single peak are obtained by The width of the half-peak width on the left and right of the Raman shift of the maximum peak point of the Raman spectrum of each receptor protein is a matching interval defined for a receptor protein, and the matching interval defined for each receptor protein is obtained; the machine learning model important data feature extraction module includes a second data preprocessing module, a machine learning model training module and an output machine learning model important data feature module; the second data preprocessing module is used in step 2, and the cell Raman spectrum is baseline corrected and silicon-based removed to obtain a standard cell Raman spectrum for subsequent data analysis; the machine learning model training module is used in step 2, and the accuracy of the initial classification of the cell Raman spectrum is obtained by inputting the cell Raman spectrum into the machine learning model training; the output machine learning model important data feature module is used in step 2, and the absolute values ​​of the cell feature coefficients are sorted in ascending order to obtain a preset number of important data features of the machine learning model; The Raman peak feature matching module includes an interval matching model important data feature module and a region matching model important feature module; the interval matching model important data feature module is used in step 3 to match the important data features of the machine learning model through the matching intervals respectively delineated by the receptor proteins, so as to obtain the matching situation of each protein with the important data features of the machine learning model; the region matching model important feature module is used in step 3 to match and screen the union region of the receptor proteins with the important data features of the machine learning model, so as to obtain the matching and screening results of the union region of the receptor proteins with the important data features of the machine learning model; The feature enhancement module includes a weighted formula module, a feature enhancement module with different weighting schemes and a feature enhancement comparison module with different weighting schemes; the weighted formula module is used in step 4 to obtain the designed weighted formula by calculating the ratio of the number of important data features matched in the protein union region to the total number of important data features successfully matched; the feature enhancement module with different weighting schemes is used in step 4 to obtain the cell Raman spectra with different weighting scheme feature enhancements by designing different weighting scheme feature enhancements; the feature enhancement comparison module with different weighting schemes is used in step 4 to obtain the weighted accuracy of the cell Raman spectra with different weighting scheme feature enhancements on the test set by retraining the machine learning model.

7. An electronic device according to the Raman peak feature matching method according to any one of claims 1 to 6, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the Raman peak feature matching method according to any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it is capable of performing cell Raman peak feature matching based on the Raman peak feature matching method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Nondestructive label-free rapid breast cancer Raman spectrum pathology grading and staging method

    CN114923893A