Serum species identification model determination method, application method and related device
By constructing a PCA-LDA model based on Raman spectroscopy, serum samples are pretreated and dimensionalized, the problem of customs in identifying species of imported and exported serum samples is solved, rapid and non-destructive species identification is achieved, and regulatory efficiency is improved.
Patent Information
- Application Number
- CN202510185628.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
Customs ports need to quickly and non-destructive species identification of imported and exported serum samples to distinguish between human and bovine serum, especially in cases of suspected concealment or missed reporting.
Raman spectroscopy technology was used to combine principal component analysis (PCA) and linear discriminant analysis (LDA) to construct serum species identification model. The PCA-LDA model was constructed by pre-processing, dimensionality reduction and classification training of Raman spectral data of historical serum samples.
It realizes rapid and non-destructive species identification of serum samples without opening the sample outer packaging and performing any pre-treatment, which improves the efficiency and accuracy of customs supervision of serum samples.
Smart Images

Figure CN120123905A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of inspection technology for the import and export of special articles at customs ports, and particularly relates to a method for determining a serum species identification model, an application method, and related devices. Background Art
[0002] Raman spectroscopy technology is a promising rapid identification and analysis method that has developed rapidly in recent years. It has many advantages such as zero pollution, little or no pretreatment, non-contact, and small sample volume, and has been preliminarily explored and applied in various industries such as petroleum, food, and jewelry. At the same time, Raman spectroscopy technology has the advantages of easy operation, short measurement time, high sensitivity, and small sample volume required, and is suitable for quantitative research, database search, and qualitative research using differential analysis. Raman spectroscopy, as a vibrational spectroscopy, analyzes the scattered light with a frequency different from that of the incident light to obtain the vibration and rotation information of the substance molecules, thereby analyzing the composition of the substance and providing the possibility for non-destructive identification. The Raman spectral information of blood is very rich, and the structural information of hemoglobin contributes most significantly to the molecular vibration of the Raman spectrum. The structural differences of blood hemoglobin with different attributions lead to weak differences in their Raman spectra. Wael et al. first applied Raman spectroscopy to the identification of blood samples. Kelly Virkler first realized the species identification of blood from different species such as humans, dogs, and cats through Raman spectroscopy combined with advanced statistical analysis. McLaughlin et al. constructed a chemometric partial least squares discriminant analysis (PLS-LDA) model to extend the blood identification research to 11 species. However, the above Raman spectroscopy technology is mainly developed for whole blood. Among the blood products at the entry and exit, serum products account for the largest proportion. Bovine serum provides various essential nutrients and growth factors in cell growth and helps to maintain the biological characteristics of cells. Therefore, it plays an important role in scientific research activities and the production of biological products (such as vaccines) and has become one of the common biological materials in international trade. For this reason, the customs implements strict animal and plant quarantine and cargo certificate verification for the import and export of bovine serum to ensure the safety and compliance of biological materials. According to the "Regulations on the Health and Quarantine Management of Special Articles at the Entry and Exit" of the General Administration of Customs, the entry and exit human serum needs to be imported and exported under relatively strict approval and supervision conditions. The regulatory intensity of serum of animal origin is quite different from that of human serum. Since different species of blood products need to implement corresponding regulatory measures, the customs port supervision site needs to conduct species identification on biological materials such as serum from different sources (especially those suspected of concealment and misreporting). Therefore, there is an urgent need to develop a detection method that can identify the species sources of human and bovine serum samples. Summary of the Invention
[0003] The purpose of this application is to provide a method for determining a serum species identification model, an application method, and related devices, which can accurately identify the species source of serum samples. Without opening the outer packaging of the samples and without any pretreatment of the samples, non-destructive and rapid species identification of serum samples can be carried out.
[0004] To achieve the above object, this application provides the following solutions:
[0005] In the first aspect, this application provides a method for determining a serum species identification model. The method for determining the serum species identification model includes:
[0006] Obtain historical serum samples and corresponding Raman spectral data; the historical serum samples include: human serum samples and bovine serum samples.
[0007] Preprocess the Raman spectral data corresponding to the historical serum samples to obtain a data training set; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum rejection.
[0008] Use principal component analysis (PCA) to perform dimensionality reduction processing on the data training set to obtain a dimensionality-reduced data training set and a data validation set.
[0009] Based on the dimensionality-reduced data training set, perform internal cross-validation (CV) to train a linear discriminant classification model to obtain a number of classification results.
[0010] Construct a serum species identification model based on the principal components corresponding to the optimal classification results; the serum species identification model is a PCA-LDA model.
[0011] In the second aspect, this application provides an application method for a serum species identification model. The application method for the serum species identification model includes:
[0012] Obtain a serum sample to be tested and corresponding Raman spectral data; the serum sample to be tested includes: a human serum sample to be tested and a bovine serum sample to be tested.
[0013] Preprocess the Raman spectral data corresponding to the serum sample to be tested to obtain a data set to be tested; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum rejection.
[0014] Input the data set to be tested into the serum species identification model to obtain a prediction result to be obtained; the prediction result to be obtained is the species category; the serum species identification model is a model constructed based on the method for determining the serum species identification model described above.
[0015] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the method for determining the serum species discrimination model or the method for applying the serum species discrimination model described above.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for determining the serum species discrimination model or the method for applying the serum species discrimination model described above.
[0017] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0018] The present application provides a method for determining a serum species discrimination model, an application method, and related devices. The method for determining the serum species discrimination model includes: obtaining historical serum samples and corresponding Raman spectroscopy data; the historical serum samples include: human serum samples and bovine serum samples; preprocessing the Raman spectroscopy data corresponding to the historical serum samples to obtain a data training set; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum rejection; performing dimensionality reduction processing on the data training set by using principal component analysis to obtain a dimensionally reduced data training set and a data validation set; based on the dimensionally reduced data training set, performing internal cross-validation, training a linear discriminant classification model to obtain a plurality of classification results; constructing a serum species discrimination model based on the principal components corresponding to the optimal classification result; the serum species discrimination model is a PCA-LDA model. The present application can perform fast and non-destructive species discrimination on serum samples without opening the sample outer packaging and without preprocessing the samples. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0020] Figure 1 It is an application environment diagram of a method for determining a serum species discrimination model in an embodiment of the present application.
[0021] Figure 2 It is a flowchart of a method for determining a serum species discrimination model provided in an embodiment of the present application.
[0022] Figure 3Schematic diagram of the normalized spectrum in the serum spectrum preprocessing provided by an embodiment of the present application.
[0023] Figure 4 Schematic diagram of the preprocessing results of human and bovine sera provided by an embodiment of the present application.
[0024] Figure 5 Schematic diagram of the normal and abnormal human serum spectra after preprocessing with cosine distance provided by an embodiment of the present application.
[0025] Figure 6 Schematic diagram of the normal and abnormal bovine serum spectra after preprocessing with cosine distance provided by an embodiment of the present application.
[0026] Figure 7 Schematic diagram of the discrimination accuracy of the LDA model under different numbers of principal components provided by an embodiment of the present application.
[0027] Figure 8 Schematic diagram of the number of correctly classified cases in the cross-validation confusion matrix for the binary classification of human and bovine sera provided by an embodiment of the present application.
[0028] Figure 9 Schematic diagram of the correct classification rate in the cross-validation confusion matrix for the binary classification of human and bovine sera provided by an embodiment of the present application.
[0029] Figure 10 Schematic diagram of the LDA feature distribution of human and bovine sera provided by an embodiment of the present application.
[0030] Figure 11 Schematic diagram of the number of correctly classified cases in the external validation confusion matrix for the binary classification of human and bovine sera provided by an embodiment of the present application.
[0031] Figure 12 Schematic diagram of the correct classification rate in the external validation confusion matrix for the binary classification of human and bovine sera provided by an embodiment of the present application.
[0032] Figure 13 Schematic diagram of the process of an application method of a serum species discrimination model provided by an embodiment of the present application.
[0033] Figure 14 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0034] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0035] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Raman spectroscopy technology has significant advantages in identifying serum samples from different species, mainly due to the unique molecular vibration characteristics of the components of each biological serum. Serum is the main component of blood, referring to the pale yellow transparent liquid separated from plasma after removing fibrin after blood coagulation. The proteins, lipids, and other biomolecules in serum show significant structural and compositional differences among different species, and these differences appear as characteristic spectral peaks in Raman spectra. By accurately analyzing the intensity and position of the spectral peaks, the serum of a specific species can be effectively identified.
[0037] The method for determining the serum species identification model provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the obtained historical serum samples and the corresponding Raman spectroscopy data to the server 104. The historical serum samples include: human serum samples and bovine serum samples. After receiving the historical serum samples and the corresponding Raman spectroscopy data, for the historical serum samples and the corresponding Raman spectroscopy data, the server 104 preprocesses the Raman spectroscopy data corresponding to the historical serum samples to obtain a data training set. The preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum elimination. The principal component analysis is used to perform dimensionality reduction processing on the data training set to obtain a dimensionality-reduced data training set and a data validation set. Based on the dimensionality-reduced data training set, internal cross-validation is performed to train a linear discriminant classification model to obtain several classification results. A serum species discrimination model is constructed based on the principal components corresponding to the optimal classification results. The serum species discrimination model is a PCA-LDA model. The server 104 can feedback the obtained serum species discrimination model to the terminal 102. In addition, in some embodiments, the method for determining the serum species discrimination model can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly construct a serum species discrimination model for the historical serum samples and the corresponding Raman spectroscopy data, or the server 104 can obtain the historical serum samples and the corresponding Raman spectroscopy data from the data storage system and construct a serum species discrimination model for the historical serum samples and the corresponding Raman spectroscopy data.
[0038] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, and tablet computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0039] In an exemplary embodiment, as Figure 2 shown, a method for determining a serum species discrimination model is provided. This method is executed by a computer device, and can specifically be executed separately by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in
[0040] A1: Obtain historical serum samples and the corresponding Raman spectroscopy data; the historical serum samples include: human serum samples and bovine serum samples.
[0041] A2: Preprocess the Raman spectral data corresponding to historical serum samples to obtain a data training set; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum elimination.
[0042] A3: Perform dimensionality reduction on the data training set using principal component analysis to obtain a dimensionality-reduced data training set and a data validation set.
[0043] A4: Based on the dimensionality-reduced data training set, perform internal cross-validation, train a linear discriminant classification model, and obtain several classification results.
[0044] A5: Construct a serum species discrimination model based on the principal components corresponding to the optimal classification results; the serum species discrimination model is a PCA-LDA model.
[0045] As an alternative implementation, in step A1, it specifically includes:
[0046] A101: Calibrate the confocal laser Raman spectrometer using a silicon standard to obtain a calibrated confocal laser Raman spectrometer.
[0047] A102: Use the calibrated confocal laser Raman spectrometer to collect Raman spectral data of human serum samples and bovine serum samples to obtain Raman spectral data of human serum samples and Raman spectral data of bovine serum samples.
[0048] In this embodiment, a confocal laser Raman spectrometer is used to collect spectral signals. Before measurement, a dark room is selected as the environment, and the background signal is subtracted. The instrument is calibrated using a silicon standard. The laser light source is 785 nm, the laser power is selected as 200 mW, the scanning range is 264 - 2030 cm -1 , and the integration time is adjusted to 60 s. Start collecting Raman spectral signals of common species serum samples of humans and bovines.
[0049] As an alternative implementation, in step A2, it specifically includes:
[0050] A201: Perform spectral smoothing, fluorescence background subtraction, and outer packaging reagent bottle signal subtraction on the Raman spectral data of historical serum samples to obtain historical Raman spectral data with obvious peak information.
[0051] A202: Perform normalization on the historical Raman spectral data with obvious peak information to obtain historical Raman spectral data with obvious peak information after normalization processing.
[0052] A203: Based on the historical Raman spectral data with obvious peak information after normalization processing, identify and eliminate abnormal historical Raman spectral data to obtain a data training set.
[0053] In this embodiment, a series of preprocessing operations (spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction) are performed on the historical serum Raman spectra to obtain historical Raman spectral data with obvious peak information; abnormal spectra that deviate significantly from most of the spectral data are identified and removed to form a data training set.
[0054] The data training set is subjected to dimensionality reduction processing through principal component analysis. By performing internal cross-validation on sample data of different batches, a linear discriminant (PCA-LDA) classification model is trained; through comparison of classification effects, the optimal PCA principal components are selected to construct a PCA-LDA model; LDA is used to classify and discriminate the Raman spectral data after dimensionality reduction.
[0055] This application relates to a Raman detection method for non-destructive, rapid, and sensitive identification of human and bovine sera. The detection method principle of this application is to use a principal component discriminant analysis model (PCA-LDA) to analyze and identify the Raman spectra of sera of different species, collect spectral data of human and bovine serum samples with reagent tubes, and spectral data of the outer packaging reagent tubes of the serum samples; perform a series of preprocessing operations (spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction) on the serum Raman spectra to obtain Raman spectral data with obvious peak information, and perform normalization processing; identify and remove abnormal spectra that deviate significantly from most of the spectral data to form a data training set; perform dimensionality reduction processing on the data through principal component analysis, and select the optimal PCA principal components to construct a PCA-LDA model through internal cross-validation of sample data of different batches; use the trained model to classify sera from human and bovine sources; collect spectral data of human and bovine serum species intercepted at the port as an external validation set to evaluate the classification effect of the model. This application can perform rapid and non-destructive species identification on serum samples without opening the outer packaging of the samples and without performing pretreatment on the samples.
[0056] In addition, the method for determining the serum species identification model further includes external validation: using the spectral data of human and bovine sera intercepted at the port as an external validation set, and adopting cross-validation to evaluate the classification effect of the model on known species that did not participate in model building.
[0057] As an optional implementation manner, the method for determining the serum species identification model further includes:
[0058] A6: Input the data validation set into the serum species identification model to obtain a test result.
[0059] A7: According to the test result and the actual result, establish coordinate axes to draw an ROC curve; the horizontal axis of the ROC curve represents the predicted label, and the vertical axis represents the actual label.
[0060] A8: Evaluate the serum species discrimination model according to the ROC curve to obtain a first evaluation result.
[0061] The method for determining the serum species discrimination model further includes:
[0062] A9: Obtain an unknown species dataset; the unknown species dataset is Raman spectral data obtained based on other species except human serum samples and bovine serum samples.
[0063] A10: Input the unknown species dataset into the serum species discrimination model to obtain a prediction result.
[0064] A11: Draw a confusion matrix graph for binary classification and multi-class cross-validation of species in different batches according to the prediction result and the actual result; the abscissa of the confusion matrix graph represents the actual category of the sample, and the ordinate represents the predicted category.
[0065] A12: Evaluate the serum species discrimination model according to the confusion matrix graph to obtain a second evaluation result.
[0066] The following further elaborates on this application in conjunction with specific embodiments.
[0067] I. Implementation steps.
[0068] 1. Instruments.
[0069] Spectrometer: Confocal Raman spectroscopy system. Before measurement, the instrument was calibrated with a silicon standard.
[0070] Light source: Semiconductor laser, power 200 mW, excitation wavelength 785 nm.
[0071] 2. Serum samples.
[0072] To ensure sample diversity, 425 human serum samples were collected, including 381 provided by the Shanghai International Travel Health Center and 44 intercepted at the port. Bovine serum included 379 standard bovine serum samples purchased from different biotech companies and 5 intercepted port samples. The serum samples were all contained in 1.5 mL sterilized polypropylene plastic tubes and stored frozen at -80 °C. Before spectral collection, the serum samples were placed at room temperature and melted into a liquid state. Then the serum samples were placed on the sampling stage, and the laser focus point was adjusted to the serum position (about 1.5 mm from the outer side of the tube wall). The whole process was aseptic. The excitation light wavelength was 785 nm, the spectral scanning range was 264 - 2030 cm -1 , the laser power was 200 mW, the total integration time at a single collection point was 60 s, and spectral data at three different sites were collected for each sample.
[0073] 3. Data preprocessing.
[0074] All data preparation and construction of statistical models are based on the Scikit-learn machine learning library in the Python programming language. The spectrum is truncated to the data range of 200 - 2000 cm -1 , and the Savitzky-Golay (SG) smoothing algorithm is applied to each spectral data to eliminate random noise and outliers in the spectrum, improving the smoothness and signal-to-noise ratio of the spectrum. When collecting sample data, the spectral data of the polypropylene plastic tubes for the outer packaging of the samples was collected, and the airPLS algorithm was used to accurately remove the fluorescence background from the collected outer packaging spectra to ensure the accuracy of subsequent analysis. Since each sample data was collected 3 times, the data of the 3 samplings for each sample were averaged to form an average spectrum representing one sample. Based on the ratio of the peak heights of the characteristic peaks of the serum samples and the outer packaging spectra at 801 cm -1 , the coefficient of the outer packaging spectrum of the sample was calculated to deduct the signal of the outer packaging spectrum of the sample, obtaining the Raman spectrum of pure serum, and each average spectrum was normalized. Using the cosine distance of the spectrum as a metric, the abnormal spectra that deviated significantly from most of the spectral data were identified and removed, ensuring the accuracy and representativeness of the data used in subsequent analysis.
[0075] 4. Cross-validation and model construction.
[0076] Using a training dataset containing humans and cows, internal cross-validation was performed by K-Fold Cross Validation. The dataset was randomly and evenly divided into K mutually exclusive subsets. In each iteration, one of the K subsets was selected as the validation set, and the remaining K - 1 subsets were combined as the training set. After training, the current validation set was used to evaluate the performance of the model. By comparing the classification effects, the optimal PCA principal components were selected to construct the PCA-LDA model.
[0077] 5. Species classification.
[0078] LDA was used to classify and discriminate the Raman spectral data after dimensionality reduction.
[0079] 6. External validation.
[0080] To evaluate the classification effect of the model, the Raman spectral data of human and bovine sera intercepted at the port were used as the external validation set, and these samples were not involved in the model construction.
[0081] II. Experimental results.
[0082] a. Results of spectral preprocessing.
[0083] Each spectrum in the training data set has been preprocessed by fluorescence background subtraction and external packaging Raman signal subtraction. Taking human serum samples as an example, the normalized spectrum in the preprocessing process is shown in Figure 3 , where the solid line is the original average spectrum; the dotted line is the spectrum after deducting the fluorescence background; the combined line of the solid and dotted line is the spectrum after deducting the packaging bottle.
[0084] By averaging all human serum spectra in the data set, a spectrum representing human original serum is generated (solid line). After deducting background noise and external signals, a pure serum spectrum representing human serum signal is obtained (virtual and real combined spectrum line). For the serum sample signal of each species, the same preprocessing steps are used to generate a representative spectrum for each species. The spectral preprocessing results of the two species are compared, and the results are shown in Figure 2. Figure 4 shown.
[0085] From the pre-processed spectra, different species have relatively consistent Raman characteristic peaks, with the main peak appearing at 538cm -1 , 998cm -1 , 1238cm -1 , 1450cm -1 , 1650cm -1 However, there are variations in the relative intensities of the Raman bands between different species.
[0086] In order to ensure the accuracy and representativeness of the data used in subsequent analysis, the cosine distance of the spectrum was used as a metric to identify and eliminate abnormal spectra that deviated significantly from the majority of spectral data. If the calculated cosine distance was greater than the mean ±1 times the standard deviation, it was judged to be abnormal.
[0087] Figure 5 and Figure 6 The cosine distances of normal and abnormal human serum and bovine serum spectra after preprocessing are shown in Figure 2. After removing the abnormal spectra, 370 human sera remained, including 335 provided by Shanghai International Travel Health Care Center and 35 intercepted from ports; 332 bovine sera remained, including 327 standard bovine serum samples purchased from different biotechnology companies and 5 samples intercepted from ports.
[0088] b. Model construction and cross-validation results.
[0089] b1. Model construction.
[0090] A PCA-LDA model was constructed to classify species using 335 human serum samples provided by Shanghai International Travel Healthcare Center and 327 standard bovine serum samples purchased from different biotechnology companies as model training sets.
[0091] Considering that the components of samples of the same species and the same batch are relatively consistent, and there are certain compositional differences among serum samples of different batches and different biological companies, in order to ensure that the model can extract the characteristic information of different batches of each species, 10-fold cross-validation based on sample batches (K = 10) is used to verify the classification performance of serum spectra.
[0092] Figure 7 Table The discrimination accuracy of the LDA model under different numbers of principal components, from Figure 7 It can be seen that when the number of PCA factors is selected as 6, the classification effect of the model is the best. At this time, the cross-validation results of the model are as Figure 8 and Figure 9 shown.
[0093] Figure 8 and Figure 9 represent the cross-validation confusion matrix results of binary classification of human and bovine sera. In the figure, the rows represent the actual categories of the samples, and the columns represent the categories predicted by the model. The values on the diagonal represent the number of species correctly predicted. From Figure 8 it can be seen that all 335 human samples are correctly predicted as human, and all 327 animal samples are predicted as animals. The values outside the diagonal of the matrix are all 0, indicating that no human or bovine samples are mispredicted. From Figure 9 it can be intuitively seen that in the binary classification cross-validation of humans and cows, the recognition results are all distributed on the diagonal, and both human and bovine serum samples are correctly predicted, and the classification accuracy is 1.00.
[0094] b2. Classification results.
[0095] Figure 10 Shows the results after classifying human and bovine serum samples based on the first principal component extracted by linear discriminant analysis (LDA).
[0096] The values on the horizontal axis represent the sample numbers, which are marked according to the arrangement order of the samples. They are arranged from left to right in sequence to distinguish different samples. The vertical axis represents the first principal component extracted by LDA, and the values on the vertical axis represent the projection positions of the samples on the first principal component of LDA. Figure 10 Each point in Figure 10 represents a sample. From
[0097] b3. External verification results.
[0098] To test the classification effect of the model on samples outside the training dataset, the model was externally verified using the sample spectra intercepted at the port. The external verification dataset contained a total of 40 serum sample spectra, including 335 human and 5 bovine. The binary confusion matrix results of the external verification are as Figure 11 and Figure 12 shown.
[0099] From Figure 11 and Figure 12 it can be seen that when performing binary cross-validation on different species intercepted at the port, all samples were correctly classified and no false positive results occurred.
[0100] In an exemplary embodiment, as Figure 13 shown, a method for applying a serum species identification model is provided. The method for applying the serum species identification model includes:
[0101] B1: Obtain the serum sample to be tested and the corresponding Raman spectral data; the serum sample to be tested includes: the human serum sample to be tested and the bovine serum sample to be tested.
[0102] B2: Preprocess the Raman spectral data corresponding to the serum sample to be tested to obtain the dataset to be predicted; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum elimination.
[0103] B3: Input the dataset to be predicted into the serum species identification model to obtain the result to be predicted; the result to be predicted is the species category; the serum species identification model is a model constructed based on the determination method of the serum species identification model described in any one of the above.
[0104] As an alternative implementation, in step B1, it specifically includes:
[0105] B101: Calibrate the confocal laser Raman spectrometer using a silicon standard to obtain a calibrated confocal laser Raman spectrometer.
[0106] B102: Use the calibrated confocal laser Raman spectrometer to collect the Raman spectral data of the human serum sample to be tested and the bovine serum sample to be tested to obtain the Raman spectral data of the human serum sample to be tested and the Raman spectral data of the bovine serum sample to be tested.
[0107] As an alternative implementation, in step B2, it specifically includes:
[0108] B201: Perform spectral smoothing, fluorescence background subtraction, and outer packaging reagent bottle signal subtraction on the Raman spectral data of the serum sample to be tested to obtain the Raman spectral data to be tested with obvious peak information.
[0109] B202: Normalize the Raman spectral data to be measured with obvious peak information to obtain the Raman spectral data to be measured with obvious peak information after normalization processing.
[0110] B203: Based on the Raman spectral data to be measured with obvious peak information after normalization processing, identify and remove abnormal Raman spectral data to be measured to obtain a data set to be measured.
[0111] This application also provides an application scenario, which applies the application method of the above-mentioned serum species identification model. Specifically: The application method of the serum species identification model provided in this embodiment can be applied to the inspection scenario of the import and export of special items at customs ports. The inspection scenario of the import and export of special items at customs ports includes: a data acquisition link, a data preprocessing link, and a prediction link; obtaining a serum sample to be measured and the corresponding Raman spectral data; the serum sample to be measured includes: a human serum sample to be measured and a bovine serum sample to be measured; preprocess the Raman spectral data corresponding to the serum sample to be measured to obtain a data set to be measured; the preprocessing includes: spectral smoothing, fluorescence background subtraction, outer packaging reagent bottle signal subtraction, normalization, and abnormal spectrum removal; input the data set to be measured into the serum species identification model to obtain a prediction result to be obtained; the prediction result to be obtained is the species category.
[0112] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 14 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store historical serum samples and the corresponding Raman spectral data and serum samples to be measured and the corresponding Raman spectral data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it is used to implement a method for determining a serum species identification model or implement an application method of a serum species identification model.
[0113] Those skilled in the art can understand, Figure 14The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0114] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above method embodiments are implemented.
[0115] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the above method embodiments are implemented.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0117] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0118] The databases involved in the embodiments provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., and is not limited thereto. The processors involved in the embodiments provided in this application may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.
[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0120] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for determining a serum species identification model, characterized in that: The method for determining the serum species identification model comprises: Obtaining historical serum samples and corresponding Raman spectral data; the historical serum samples include: human serum samples and bovine serum samples; Preprocessing the Raman spectral data corresponding to the historical serum samples to obtain a data training set; the preprocessing includes: spectral smoothing, subtracting fluorescence background, subtracting the signal of the outer packaging reagent bottle, normalizing and eliminating abnormal spectra; The data training set is subjected to dimensionality reduction processing by principal component analysis to obtain a data training set and a data verification set after dimensionality reduction; Based on the dimension-reduced data training set, internal cross-validation is performed to train a linear discriminant classification model to obtain several classification results; A serum species identification model is constructed based on the principal components corresponding to the optimal classification results; the serum species identification model is a PCA-LDA model.
2. The method for determining the serum species identification model according to claim 1, characterized in that: Obtain historical serum samples and corresponding Raman spectral data, including: The confocal laser Raman spectrometer is calibrated by using silicon standard to obtain a calibrated confocal laser Raman spectrometer; The human serum samples and the bovine serum samples were collected using a calibrated confocal laser Raman spectrometer to obtain Raman spectrum data of the human serum samples and Raman spectrum data of the bovine serum samples.
3. The method for determining the serum species identification model according to claim 1, characterized in that: The Raman spectral data corresponding to the historical serum samples were preprocessed to obtain the data training set, which specifically includes: The Raman spectrum data of historical serum samples were processed by spectral smoothing, fluorescence background subtraction and external reagent bottle signal subtraction to obtain historical Raman spectrum data with obvious peak information; Normalizing the historical Raman spectrum data with obvious peak information to obtain normalized historical Raman spectrum data with obvious peak information; Based on the historical Raman spectral data with obvious peak information after the normalization process, abnormal historical Raman spectral data are identified and eliminated to obtain a data training set.
4. The method for determining a serum species identification model according to claim 1, characterized in that: The method for determining the serum species identification model also includes: Inputting the data validation set into the serum species identification model to obtain a test result; According to the test results and the actual results, a coordinate axis is established to draw a ROC curve; the horizontal axis of the ROC curve represents the predicted label, and the vertical axis represents the actual label; The serum species identification model is evaluated according to the ROC curve to obtain a first evaluation result.
5. The method for determining a serum species identification model according to claim 1, characterized in that: The method for determining the serum species identification model also includes: Acquire an unknown species data set; the unknown species data set is Raman spectral data obtained based on species other than human serum samples and bovine serum samples; Inputting the unknown species data set into the serum species identification model to obtain a prediction result; According to the predicted results and actual results, a confusion matrix diagram of binary classification and multi-classification cross-validation of different batches of species is drawn; the abscissa of the confusion matrix diagram represents the actual category of the sample, and the ordinate represents the predicted category; The serum species identification model is evaluated according to the confusion matrix diagram to obtain a second evaluation result.
6. An application method of a serum species identification model, characterized in that: The application method of the serum species identification model comprises: Acquire a serum sample to be tested and corresponding Raman spectrum data; the serum sample to be tested includes: a human serum sample to be tested and a bovine serum sample to be tested; Preprocessing the Raman spectral data corresponding to the serum sample to be tested to obtain a data set to be tested; the preprocessing includes: spectral smoothing, subtracting fluorescence background, subtracting the signal of the outer packaging reagent bottle, normalizing and eliminating abnormal spectra; The test data set is input into the serum species identification model to obtain the result to be predicted; the result to be predicted is the species category; the serum species identification model is a model constructed based on the determination method of the serum species identification model according to any one of claims 1 to 5.
7. The method for applying the serum species identification model according to claim 6, characterized in that: Obtain the serum sample to be tested and the corresponding Raman spectrum data, including: The confocal laser Raman spectrometer is calibrated by using silicon standard to obtain a calibrated confocal laser Raman spectrometer; The calibrated confocal laser Raman spectrometer is used to collect data from the human serum sample and the bovine serum sample to obtain Raman spectrum data of the human serum sample and the bovine serum sample to be tested.
8. The method for applying the serum species identification model according to claim 6, characterized in that: Preprocess the Raman spectrum data corresponding to the serum sample to be tested to obtain the test data set, specifically including: The Raman spectrum data of the serum sample to be tested are spectrally smoothed, the fluorescence background is subtracted, and the signal of the outer packaging reagent bottle is subtracted to obtain the Raman spectrum data to be tested with obvious peak information; Normalizing the Raman spectrum data to be measured with obvious peak information to obtain the normalized Raman spectrum data to be measured with obvious peak information; Based on the Raman spectrum data to be tested with obvious peak information after the normalization process, abnormal Raman spectrum data to be tested is identified and eliminated to obtain a data set to be tested.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for determining the serum species identification model described in any one of claims 1 to 5 or the method for applying the serum species identification model described in any one of claims 6 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for determining the serum species identification model described in any one of claims 1 to 5 or the method for applying the serum species identification model described in any one of claims 6 to 8 is implemented.