Anomaly Detection Method for Serum SERS Spectral Data
The PCA dimensionality reduction and DBSCAN clustering algorithm remove outliers, which solves the problem that outliers in serum SERS spectral data affect data analysis, and improves the accuracy of the data and the performance of the classifier.
Patent Information
- Application Number
- CN202310831071.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-07-07
AI Technical Summary
Outliers exist in serum SERS spectral data, affecting the accuracy and robustness of data analysis and classifiers.
PCA technology is used to reduce the dimensionality of the spectral signal, and the DBSCAN clustering algorithm is used to remove abnormal spectral signals caused by poor experimental conditions, and outliers are identified and removed through clustering and parameter adjustment.
The outliers are effectively removed, which improves the accuracy and reliability of serum Raman spectroscopy data, thereby improving the accuracy of subsequent classification and analysis.
Smart Images

Figure CN117056840B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of biomedicine and spectral detection, and particularly relates to a method for detecting anomalies in serum SERS spectral data. Background Art
[0002] Raman spectroscopy can be used to study the structures of various biomolecules, such as peptide chains, carbohydrates, lipids, and nucleic acids. The advantage of this technique is that it can provide high-resolution molecular information without destruction, as it does not require any treatment or labeling of the sample. By analyzing serum SERS spectra, many biomolecules can be identified and characterized, including amino acids, nucleotides, lipid compounds, and polypeptides, etc. The changes in these molecules are closely related to the occurrence and development of diseases. Therefore, serum SERS spectral analysis can be used in applications such as early cancer diagnosis, treatment monitoring, and drug screening. At the same time, SERS technology can also be used to detect other types of biomolecules, such as bacteria, viruses, and cells. Due to its high sensitivity and high selectivity, SERS technology has broad application prospects in the fields of biomedicine, food safety, and environmental monitoring.
[0003] However, since serum samples contain a large number of biomolecules, such as proteins, metabolites, and nucleic acids, etc., the content and composition of these molecules are often affected by many factors, such as individual differences, environmental conditions, and experimental operations. These factors may lead to the existence of some outliers in the SERS spectral data set, thus affecting subsequent data analysis and modeling tasks, and reducing the accuracy and robustness of the classifier. Therefore, an effective method needs to be introduced to detect and remove outliers in SERS spectral samples to improve the performance and reliability of the classifier. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for detecting anomalies in serum SERS spectral data, which can effectively improve the reliability of the algorithm model.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for detecting anomalies in serum SERS spectral data, comprising the following steps:
[0007] Step 1, obtain multi-category raw serum SERS spectra.
[0008] Specifically, serum samples of multiple cancer test groups and healthy control groups are obtained according to the prevalence of different cancers in the population, and SERS characterization is performed to obtain the original spectral dataset. The proportion of the healthy control group and multiple cancer subjects is determined according to the actual prevalence of the cancer type in the population. For SERS characterization, silver nanocolloid is used to enhance the serum signal by SERS, and an OceanHood RMS1000 Raman spectrometer is used for measurement. The spectrometer is excited by a 785 nm laser, the laser power is set to 10 mW, and the integration time is set to 5000 mS.
[0009] Step 2: Perform spectral preprocessing on the obtained original serum SERS spectra to obtain standardized serum SERS spectra.
[0010] Specifically, spectral preprocessing of the original spectral data includes spectral ROI region cropping, spectral interpolation, fluorescence background signal removal, and spectral normalization. The ROI region for spectral cropping is limited to the fingerprint region of 400 - 1800 cm -1 . The spectral interpolation uses an interpolation algorithm, including one of the general algorithms such as linear interpolation, bilinear interpolation, and spline interpolation. The interpolation interval is set to at least one of 0.1 cm -1 , 0.5 cm -1 , 1 cm -1 , 2 cm -1 , etc. The fluorescence background signal is removed using the Vancouver algorithm. The normalization method uses at least one of the normalization methods based on the peak and the integral area under the spectrum.
[0011] Step 3: Use PCA technology to reduce the dimensionality of the preprocessed spectral signals to a low-dimensional space;
[0012] Specifically, using PCA technology to reduce the dimensionality of the preprocessed spectral signals to a low-dimensional space, the process is as follows: center all samples; calculate the covariance matrix XX T of the samples; perform eigenvalue decomposition on the matrix XX T ; extract the eigenvectors (W1, W2,..., W n ') corresponding to the largest n' eigenvalues; after normalizing all the eigenvectors, form an eigenvector matrix W; transform each sample X (i) in the sample set into a new sample Z (i) = W T X (i) ; finally, obtain the reduced-dimensional sample set D' = (Z1, Z2,..., Z m ).
[0013] Step 4: Use the DBSCAN clustering algorithm to remove abnormal spectral signals caused by too high excitation power, CCD saturation, and positioning deviation.
[0014] Specifically, by using the DBSCAN clustering algorithm, abnormal spectral signals caused by poor experimental conditions are identified and removed, such as too high excitation power, CCD saturation, and positioning deviation. Such processing can improve the accuracy and reliability of serum Raman spectral data, thus providing a more reliable basis for subsequent classification and analysis.
[0015] Step 5: Use the DBSCAN clustering algorithm and set the parameters epsilon and MinPts therein to cluster these sample sets.
[0016] Specifically, the parameter epsilon is obtained from the k-distance curve. Calculate the distances between each sample and all samples, select the distances of the k-th nearest neighbors and sort them from large to small to obtain the k-distance curve, and set the distance corresponding to the inflection point of the curve as epsilon.
[0017] The minimum value of the parameter MinPts can be obtained from the dimension Dim of the data set, that is, MinPts ≥ Dim + 1. If MinPts = 1, it means that all samples in the data set are core samples, that is, each sample is a cluster; if MinPts ≤ 2, the result is the same as that of single-linkage hierarchical clustering; therefore, MinPts must be greater than or equal to 3. Generally, it is considered that MinPts = Dim, and the larger the data set, the larger the value of MinPts selected.
[0018] Step 6: Obtain a SERS spectral data set without outliers.
[0019] Specifically, for each clustering cluster, calculate its center point by using the average value, median, etc. within the cluster, and then calculate the distance between each sample point and the center point. If the distance between a certain sample point and the center point is greater than a preset threshold, then this sample point is regarded as an outlier and removed from the clustering cluster, thus obtaining a SERS spectral data set without outliers.
[0020] Step 7: Based on the obtained SERS spectral data set without outliers, perform machine learning modeling to obtain a classification model for cancer screening.
[0021] Specifically, the machine learning classification model is built using the PyTorch framework. The model validation of the machine learning classification model adopts the 5-fold cross-validation method. The 5-fold cross-validation method evenly divides the original dataset into 5 equal parts. In turn, 1 part is taken as the test set, and the remaining 4 parts are used as the training set for model validation. It is executed 5 times until each fold is used as the test set to verify the model performance of the other 4 folds of training. The average of the five validations is taken to obtain the final validation effect of the model. The model parameters to be confirmed include epsilon, MinPts, etc. The model evaluation metrics of the machine learning classification model include accuracy, recall, precision, specificity, F1-score, receiver operating characteristic (ROC), confusion matrix, etc.
[0022] Step 8, evaluate the classification model using the 5-fold cross-validation method, and determine whether the current parameter combination reaches the best result; if so, execute Step 9; otherwise, update and optimize the parameters of the classification model, and execute Step 4;
[0023] Step 9, construct the best classification model using the optimized best parameter combination.
[0024] Step 10, input the serum SERS spectral data to be abnormally detected into the best classification model to obtain the abnormal detection result.
[0025] The present invention adopts the above technical solutions. The serum SERS spectral signals of different types of cancers are collected through SERS technology. First, the PCA technology is used to reduce the dimension of the preprocessed spectral signals to a low-dimensional space. Then, the DBSCAN algorithm is used to identify and remove the abnormal spectral signals caused by poor experimental conditions, and the sample points that are density-connected are clustered together. During clustering, the Raman spectral samples are clustered by setting the adjustment parameters epsilon and MinPts. The outliers are judged by calculating the center point and the distance between each sample point and the center point, and removed from the clustering cluster, thereby obtaining a SERS spectral dataset without outliers. Finally, a machine learning algorithm is used to classify the serum SERS spectra. This method can effectively resist the unsatisfactory classification results caused by outliers, thereby improving the accuracy of the classification algorithm model.
[0026] The beneficial effects of the present invention are as follows: In the spectral preprocessing process, the spectral cropping algorithm selects the fingerprint region rich in biochemical information; the fluorescence background removal algorithm removes the autofluorescence signal of biological tissues stacked with Raman signals; the normalization algorithm eliminates the signal intensity fluctuations caused by the change of laser excitation power of the Raman spectrometer. The above spectral preprocessing process standardizes the measured original spectra under the relative intensity of the same dimension for subsequent modeling. The outlier removal method based on PCA-DBSCAN can remove the outliers in the samples through clustering, which can effectively reduce the poor classification results caused by outliers. At the same time, introducing an efficient machine learning classification algorithm can significantly improve the classification accuracy. The five-fold cross-validation method can find the best parameter combination for modeling and output the best serum SERS classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments;
[0028] Figure 1 It is a flowchart of an outlier detection method for serum SERS spectral data of the present invention;
[0029] Figure 2 It is the sample data distribution in the present invention;
[0030] Figure 3 It is a flowchart of serum SERS data acquisition in the present invention;
[0031] Figure 4 It is the effect diagram of spectral preprocessing;
[0032] Figure 5 It is the removal of abnormal spectral signals caused by too high excitation power and CCD saturation;
[0033] Figure 6 It is the removal of abnormal spectral signals caused by measurement positioning deviation;
[0034] Figure 7 It is the clustering effect corresponding to different parameters within the same category (normal (NM), breast cancer (BC), lung cancer (LC));
[0035] Figure 8 It is the model evaluation index for different epsilon and MinPts modeling;
[0036] Figure 9 It is the model confusion matrix and test ROC for the preferred epsilon and MinPts modeling (epsilon = 0.0004, MinPts = 9). EMBODIMENTS
[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application.
[0038] As Figures 1 to 9 shown in one of the figures, the present invention discloses an abnormal detection method for serum SERS spectral data, which includes the following steps:
[0039] Step 1, obtain the original serum SERS spectra of multiple categories.
[0040] Specifically, according to the prevalence of different cancers in the population, serum samples of multiple cancer test groups and healthy control groups are obtained, and SERS characterization is performed to obtain the original spectral dataset. The proportion of the healthy control group and multiple cancer subjects is determined according to the actual prevalence of the cancer type in the population. For SERS characterization, silver nanocolloid is used to enhance the serum signal by SERS, and an OceanHood RMS1000 Raman spectrometer is used for measurement. The spectrometer is excited by a 785 nm laser, the laser power is set to 10 mW, and the integration time is set to 5000 mS.
[0041] Step 2, perform spectral preprocessing on the obtained original serum SERS spectra.
[0042] Specifically, performing spectral preprocessing on the original spectral data includes spectral ROI region cropping, spectral interpolation, fluorescence background signal removal, and spectral normalization. The ROI region for spectral cropping is limited to the fingerprint region of 400 - 1800 cm -1 . The spectral interpolation uses an interpolation algorithm including one of the general algorithms such as linear interpolation, bilinear interpolation, and spline interpolation. The interpolation interval is set to 0.1 cm -1 , 0.5 cm -1 , 1 cm -1 , 2 cm -1 or at least one of them. The Vancouver algorithm is used to remove the fluorescence background signal. The normalization method uses at least one of the normalization methods based on the peak and the integral area under the spectrum.
[0043] Step 3, use PCA technology to reduce the dimension of the preprocessed spectral signal to a low-dimensional space;
[0044] Specifically, using PCA technology to reduce the dimension of the preprocessed spectral signal to a low-dimensional space, the process is as follows: centralize all samples; calculate the covariance matrix XX T of the samples; perform eigenvalue decomposition on the matrix XX T ; extract the eigenvectors (W1, W2,..., W n'); After normalizing all the feature vectors, a feature vector matrix W is formed; for each sample X in the sample set (i) is transformed into a new sample Z (i) = W T X (i) ; Finally, the dimensionality-reduced sample set D' = (Z1, Z2, …, Z m ) is obtained.
[0045] Step 4: Use the DBSCAN clustering algorithm to remove abnormal spectral signals caused by too high excitation power, CCD saturation, and positioning deviation.
[0046] Specifically, by using the DBSCAN clustering algorithm to identify and remove abnormal spectral signals caused by poor experimental conditions, such as too high excitation power, CCD saturation, and positioning deviation. Such processing can improve the accuracy and reliability of serum Raman spectral data, thus providing a more reliable basis for subsequent classification and analysis.
[0047] Step 5: Use the DBSCAN clustering algorithm and set the parameters epsilon and MinPts therein to cluster these sample sets.
[0048] Specifically, the parameter epsilon is obtained from the k-distance curve. Calculate the distance between each sample and all samples, select the distance of the k-th nearest neighbor and sort it from large to small to obtain the k-distance curve, and set the distance corresponding to the inflection point of the curve as epsilon.
[0049] The minimum value of the parameter MinPts can be obtained from the dimensionality Dim of the data set, that is, MinPts ≥ Dim + 1. If MinPts = 1, it means that all samples in the data set are core samples, that is, each sample is a cluster; if MinPts ≤ 2, the result is the same as that of single-link hierarchical clustering; therefore, MinPts must be greater than or equal to 3, so generally MinPts = Dim, and if the data set is larger, the value of MinPts selected is also larger.
[0050] Step 6: Obtain a SERS spectral data set without outliers.
[0051] Specifically, for each clustering cluster, calculate its center point by using the average value, median, etc. inside the cluster, and then calculate the distance between each sample point and the center point. If the distance between a certain sample point and the center point is greater than a preset threshold, then regard this sample point as an outlier and remove it from the clustering cluster, so as to obtain a SERS spectral data set without outliers.
[0052] Step 7: Based on the obtained SERS spectral dataset without outliers, perform machine learning modeling to obtain a classification model for cancer screening.
[0053] Specifically, the machine learning classification model is built using the PyTorch framework. The model validation of the machine learning classification model adopts the 5-fold cross-validation method. The 5-fold cross-validation method evenly divides the original dataset into 5 equal parts. One part is taken as the test set in turn, and the remaining 4 parts are used as the training set for model validation. It is executed 5 times until each fold is used as the test set to verify the performance of the model trained by the other 4 folds. The average of the five validations is taken to obtain the final validation effect of the model. The model parameters to be confirmed include epsilon, MinPts, etc. The model evaluation metrics of the machine learning classification model include accuracy, recall, precision, specificity, F1-score, receiver operating characteristic curve (ROC), and confusion matrix, etc.
[0054] Step 8: Use the 5-fold cross-validation method to evaluate the classification model and determine whether the current parameter combination reaches the best result; if so, execute Step 9; otherwise, update and optimize the parameters of the classification model, and execute Step 4;
[0055] Step 9: Use the optimized best parameter combination to construct the best classification model.
[0056] Step 10: Input the serum SERS spectral data to be detected for outliers into the best classification model to obtain the outlier detection result.
[0057] Spectral data collection:
[0058] Determine the number of serum collections according to the prevalence rates of breast cancer (BC) and lung cancer (LC) in normal people (NM). Figure 2 Show the sample data distribution of different types of cancers. The sampling quantity ratio of NM:BC:LC in this embodiment is 300:271:161.
[0059] The SERS enhancement substrate uses silver nanocolloid. The specific synthesis steps are as follows: Take 4.5 ml of sodium hydroxide solution (0.1 mol / L) and 5 ml of hydroxylamine hydrochloride solution (0.06 mol / L) and mix them evenly. Then, pour the above mixture into 90 mL of silver nitrate aqueous solution (0.0011 mol / L) under vigorous stirring until a uniform milky gray mixture is obtained. The obtained silver nanocolloid solution is concentrated by a centrifuge at a speed of 10000 rpm for 10 min to obtain the concentrated silver nanocolloid.
[0060] Take 10 ul of the concentrated silver nanocolloid and 10 ul of the serum sample, mix and drop them on a pure aluminum foil sheet, incubate and air-dry for 120 min to obtain the serum SERS sample.
[0061] The RMS1000 Raman spectrometer was set to measure the serum SERS samples with a laser excitation power of 20 mW and an integration time of 5 - 10 s, and the original Raman spectral data was obtained. The specific implementation process is as Figure 3 shown.
[0062] Preprocessing of the spectrum:
[0063] For the measured original Raman spectrum, spectral cropping, spectral interpolation, fluorescence background subtraction, and normalization operations were performed in sequence to obtain a standardized serum SERS spectrum. Figure 4 It is the effect diagram of the spectral preprocessing process.
[0064] Method for removing outliers:
[0065] The PCA technique was used to reduce the dimensionality of the preprocessed spectral signal to a low - dimensional space. The process is as follows: centering all samples; calculating the covariance matrix of the samples; performing eigenvalue decomposition on the covariance matrix; taking out the eigenvectors corresponding to the largest n' eigenvalues; after standardizing all the eigenvectors, forming an eigenvector matrix; transforming each sample in the sample set into a new sample; and finally obtaining the reduced - dimensional sample set.
[0066] By using the DBSCAN clustering algorithm, abnormal spectral signals caused by too high excitation power, CCD saturation, and poor measurement position were identified and removed, as Figure 5 and Figure 6 shown.
[0067] The DBSCAN clustering algorithm was used and parameters such as Var - Epsilon Va - r and MinPts were set to cluster these sample sets, Figure 7 which are the clustering effects corresponding to different parameters.
[0068] For each clustering cluster, the center point was calculated using methods such as the average value and median within the cluster. Then, the distance between each sample point and the center point was calculated. If the distance between a certain sample point and the center point is greater than a preset threshold, then this sample point is regarded as an outlier and removed from the clustering cluster, thus obtaining a SERS spectral data set without outliers.
[0069] Machine learning modeling and evaluation:
[0070] The spectral data set after removing outliers was used for modeling evaluation to obtain the best model parameters and classification model.
[0071] In each round of training using the five - fold cross - validation method, 1 / 5 of the data was used as the test set, and the remaining 4 / 5 was used as the training set.
[0072] Evaluation uses sensitivity, F1-score, ROC (AUC) curve, and confusion matrix.
[0073] The model evaluation metrics of models built with different epsilon and MinPts parameters are as Figure 8 shown.
[0074] Preferably, the best classification model can be obtained when epsilon = 0.0004 and MinPts = 9. The model evaluation metrics of the model built with epsilon and MinPts (epsilon = 0.0004, MinPts = 9) are as Figure 9 shown, and the confusion matrix and test ROC curve are as Figure 9 shown.
[0075] The present invention adopts the above technical solutions to collect serum SERS spectral signals of different types of cancers through SERS technology. First, the PCA technology is used to reduce the dimension of the preprocessed spectral signals to a low-dimensional space. Then, the DBSCAN clustering algorithm is used to identify and remove abnormal spectral signals caused by too high excitation power, CCD saturation, and poor measurement positions. By using the DBSCAN algorithm and adjusting the parameters epsilon and MinPts therein, the Raman spectral samples are clustered. The outliers are judged by calculating the center point and the distance between each sample point and the center point, and they are removed from the clustering cluster, thereby obtaining a SERS spectral data set without outliers. Finally, a machine learning algorithm is used to classify the serum SERS spectra. This method can effectively resist the unsatisfactory classification results caused by outliers, thereby improving the accuracy of the classification algorithm model.
[0076] The present invention has the following beneficial effects: In the spectral preprocessing process, the spectral cropping algorithm selects the fingerprint region rich in biochemical information; the fluorescence background removal algorithm removes the autofluorescence signals of biological tissues stacked with Raman signals; the normalization algorithm eliminates the signal intensity fluctuations caused by the change of the laser excitation power of the Raman spectrometer. The above spectral preprocessing process standardizes the measured original spectra under the relative intensity of the same dimension for subsequent modeling.
[0077] The outlier removal method based on PCA-DBSCAN can remove the outliers existing in the samples through clustering, and can effectively reduce the poor classification results caused by outliers. At the same time, introducing an efficient machine learning classification algorithm can significantly improve the classification accuracy. The five-fold cross-validation method can find the best parameter combination for modeling and output the best serum SERS classification model.
[0078] Obviously, the described embodiments are some, but not all, of the embodiments of this application. Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Generally, the components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application is not intended to limit the scope of this application claimed, but merely represents selected embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.
Claims
1. An abnormal detection method for serum SERS spectral data, characterized in that: It includes the following steps: Step 1: Obtain the original SERS spectra of sera of multiple categories; Step 2: Perform spectral preprocessing on the obtained original SERS spectra of sera to obtain standardized SERS spectra of sera; Step 3: Use PCA technology to reduce the dimensionality of the spectral signals after preprocessing to a low-dimensional space to obtain a reduced-dimensional sample set; Step 4: Use the DBSCAN clustering algorithm to remove abnormal spectral signals caused by too high excitation power, CCD saturation, and positioning deviation; Step 5: Use the DBSCAN clustering algorithm and set the parameters epsilon and MinPts of the DBSCAN clustering algorithm to cluster the reduced-dimensional sample set; Step 6: Calculate the center points of each clustering cluster and calculate the distances between each sample point and the center points; when the distance between a sample point and the center point is greater than a preset threshold, the corresponding sample point is regarded as an outlier and removed from the clustering cluster, so as to obtain a SERS spectral data set without outliers; Step 7: Perform machine learning modeling based on the obtained SERS spectral data set without outliers to obtain a classification model for cancer screening; Step 8: Use the five-fold cross-validation method to evaluate the classification model and determine whether the current parameter combination reaches the best result; if so, execute Step 9; otherwise, update and optimize the parameters of the classification model and execute Step 4; Step 9: Construct the best classification model using the optimized best parameter combination; Step 10: Input the SERS spectral data of the sera to be abnormally detected into the best classification model to obtain the abnormal detection result.
2. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In Step 1, serum samples of multiple cancer test groups and healthy control groups are obtained according to the prevalence rates of different cancers in the population, and SERS characterization is performed to obtain an original spectral data set; the proportion of the healthy control group and multiple cancer subjects is determined according to the actual prevalence rate of this cancer type in the population; for SERS characterization, silver nanocolloid is used to enhance the serum signal, and an OceanHoodRMS1000 Raman spectrometer is used for measurement; the spectrometer is excited by a 785 nm laser, the laser power is set to 10 mW, and the integration time is set to 5000 mS.
3. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In step 2, the spectral preprocessing of the original spectral data includes spectral ROI region cropping, spectral interpolation, fluorescence background signal removal, and spectral normalization; the ROI region of spectral ROI region cropping is limited to the fingerprint region of 400 - 1800 cm -1 ; the spectral interpolation adopts an interpolation algorithm including one of linear interpolation, bilinear interpolation, and spline interpolation algorithms, and the interpolation interval is set to 0.1 cm -1 , 0.5 cm -1 , 1 cm -1 , 2 cm -1 ; at least one of them; the fluorescence background signal removal uses the Vancouver algorithm; The normalization method adopts at least one of the normalization methods based on peak value and based on the integral area under the spectrum.
4. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: The specific steps of Step 3 are as follows: Step 3-1: Centralize all samples; Step 3-2, calculate the covariance matrix XX of the samples T ; Step 3-3, perform eigenvalue decomposition on the matrix XX T and extract the eigenvectors (W1, W2, …, W n ’) corresponding to the largest n’ eigenvalues; Step 3-4: After standardizing all feature vectors, form a feature vector matrix W; Step 3-5, for each sample X in the sample set (i) transform it into a new sample Z (i) =W T X (i) ; finally, obtain the dimensionality-reduced sample set D’=(Z1, Z2, …, Z m ).
5. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In Step 4, the DBSCAN clustering algorithm is used to identify and remove abnormal spectral signals caused by poor experimental conditions.
6. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In Step 5, the parameter epsilon is obtained using the k-distance curve: calculate the distances between each sample and all samples, select the distance of the k-th nearest neighbor and sort it from large to small to obtain the k-distance curve, and the distance corresponding to the inflection point of the curve is set as epsilon; The minimum value of the parameter MinPts is obtained from the dimensionality Dim of the data set, that is, MinPts ≥ Dim + 1; and MinPts ≥ 3.
7. The abnormal detection method for serum SERS spectral data according to claim 6, characterized in that: In Step 5, the parameter MinPts takes the value of MinPts = Dim.
8. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In step 6, for each cluster, the center point is calculated using the average or median value within the cluster; then the distance between each sample point and the center point is calculated; when the distance between a certain sample point and the center point is greater than a preset threshold, the sample point is regarded as an outlier and removed from the cluster, so as to obtain a SERS spectral data set without outliers.
9. The abnormal detection method for serum SERS spectral data according to claim 1, characterized in that: In step 7, the machine learning classification model is built using the PyTorch framework, and the 5-fold cross-validation method is used for model validation of the machine learning classification model; the model parameters to be confirmed include epsilon and MinPts; the model evaluation metrics of the machine learning classification model include accuracy, recall, precision, specificity, F1 score, receiver operating characteristic curve, and confusion matrix.
10. The abnormal detection method of serum SERS spectral data according to claim 9, characterized in that: The 5-fold cross-validation method evenly divides the original data set into 5 equal parts. One of the 5 equal parts is taken as the test set in turn, and the remaining 4 parts are used as the training set for model validation; this is executed 5 times until each fold is used as the test set to verify the performance of the model trained by the other 4 folds; the average of the five validations is taken to obtain the final validation effect of the model.