Application method based on serum Raman spectrum as biomarker in cognitive disorder diagnosis

Through serum Raman spectroscopy and machine learning technology, a cognitive impairment diagnosis system is established, which solves the problem of lack of cognitive impairment grading diagnosis and blood test methods suitable for general screening in the existing technology, and achieves low-cost and efficient grading diagnosis of cognitive impairment.

CN119915795APending Publication Date: 2025-05-02NANJING DRUM TOWER HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510147860.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing technology lacks a hierarchical diagnosis and blood test method for general screening of cognitive impairments suitable for general screening of the population, and domestic and foreign research focuses on the ATN framework, and the detection scheme is expensive and cannot be used for general screening.

Method used

Using serum Raman spectroscopy combined with machine learning technology, a cognitive impairment diagnosis system based on Raman spectroscopy technology is established, and cognitive hierarchical diagnosis of new samples is carried out through Raman spectroscopy detection and preprocessing, database module construction and optimization, and cognitive diagnosis module.

Benefits of technology

It provides a low-cost, high-receptivity, and easy-to-promote blood test method, which can perform hierarchical diagnosis of cognitive impairment, and combines machine learning algorithms to identify different cognitive hierarchical, which is suitable for general screening of populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119915795A_ABST
    Figure CN119915795A_ABST
Patent Text Reader

Abstract

The invention discloses an application method of a serum Raman spectrum as a biomarker in cognitive impairment diagnosis, a cognitive impairment diagnosis system is used for diagnosis, the cognitive impairment diagnosis system comprises a Raman spectrum detection and preprocessing module, a database module and a cognitive diagnosis module, and the application method comprises the following steps: step 1, collecting a sample; step 2, pre-treating a sample; 3, collecting and preprocessing a serum sample by using a Raman spectrum detection and preprocessing module; 4, performing model construction and optimization by utilizing a database module; step 5, performing efficiency evaluation; and 6, performing cognitive grading diagnosis on the new sample by using a cognitive diagnosis module. The method has remarkable technical progress, no cognitive impairment diagnosis blood test method suitable for the whole crowd exists clinically at present, cognitive impairment can be diagnosed through the Raman spectrum of serum, detailed cognitive function grading can be carried out, and the cost of a single sample is less than 100 yuan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cognitive impairment diagnosis, and in particular to an application method of serum Raman spectroscopy as a biomarker in cognitive impairment diagnosis. Background Art

[0002] Population screening is an important measure to identify cognitive impairment at an early stage and improve the brain health of the group. Cognitive function is divided into normal cognitive function, mild cognitive impairment and dementia. The purpose of cognitive impairment screening is to identify individuals with mild cognitive impairment and dementia in the population.

[0003] Raman spectroscopy refers to the interaction between the incident laser and the vibration of the chemical bond of the molecule, resulting in an inelastic collision, and the laser loses energy. This energy difference corresponds to the vibration frequency of the chemical bond, which is the fingerprint information of the molecular vibration, called Raman spectroscopy. Raman spectroscopy can reflect the energy of chemical bonds, determine the structure of compounds, and thus detect the chemical composition of the sample, the functional groups contained, and the changes caused by chemical reactions. Raman spectroscopy is applicable to a variety of biological samples, such as serum, plasma, cells, etc., and only 1.5ul of sample can be used to complete the detection. In addition, Raman spectroscopy does not require labeling, is low-cost, and has high accuracy. For biological samples, Raman spectroscopy can sensitively identify disease-related biochemical changes in samples and perform qualitative and quantitative analysis. For example, Raman spectroscopy can reflect the genetic background, metabolic dynamics, and microenvironmental changes of individual cells, and can perform qualitative and quantitative cell biochemical components in a multi-dimensional dynamic manner. Studies have found that Raman spectroscopy can identify chronic fatigue syndrome, early screening for tumors, reflect microbial metabolic activity, and assist in drug sensitivity testing. The rich individual biological information carried by Raman spectroscopy is highly compatible with machine learning technology. Using machine learning methods to build a library for Raman spectra of biological samples and establish algorithm models for disease identification and condition assessment are key steps in applying Raman spectroscopy to clinical practice. At present, the graded assessment of cognitive function needs to be completed by professionally trained psychiatrists through neuropsychological scales. Neuropsychological scale assessment requires a large amount of manpower from specialists, and each person takes more than 40 minutes. In addition, in screening work, the group's acceptance of neuropsychological scales is much lower than that of blood tests. Poor cooperation and refusal to test are common, and there is huge resistance to implementation in general screening. At present, the research on blood markers related to cognitive impairment at home and abroad focuses on the ATN framework. The proposed detection scheme is only applicable to AD-derived cognitive impairment, and the cost is high, so it cannot be used for general screening of the population. Therefore, the clinic urgently needs a new blood test method that is low-cost, highly acceptable, and easy to promote for graded diagnosis of cognitive impairment. Summary of the invention

[0004] The purpose of the present invention is to provide a method for using serum Raman spectroscopy as a biomarker in the diagnosis of cognitive impairment, which is in urgent need of clinical hematological methods for the diagnosis of cognitive impairment applicable to the entire population. It is intended to establish a cognitive impairment diagnosis system based on Raman spectroscopy technology through serum Raman spectroscopy and machine learning technology, and provide a new hematological detection method for the graded diagnosis of cognitive impairment. The specific technical scheme is as follows: A method for applying serum Raman spectroscopy as a biomarker in the diagnosis of cognitive impairment is disclosed. The method uses a cognitive impairment diagnosis system for diagnosis. The cognitive impairment diagnosis system includes a Raman spectroscopy detection and preprocessing module, a database module, and a cognitive diagnosis module. The method includes the following steps: Step 1: Sample collection: recruit participants to collect blood samples, and perform the Mini-Mental State Examination, Montreal Cognitive Assessment, and Daily Living Ability Scale tests; Step 2: Sample pretreatment: Collect blood samples from participants using an inert separation gel coagulation tube. After standing at room temperature for 10 minutes, centrifuge the blood samples at 1200g for 10 minutes to separate serum and obtain serum samples. Step 3: Use the Raman spectroscopy detection and preprocessing module to collect and preprocess serum samples; specifically, 1.5ul of serum sample was aspirated with a pipette and directly spotted on a low background noise chip. Raman spectra of serum samples were collected using a Raman spectrometer WITecalpha 300R. After collection, the Raman spectroscopy data was preprocessed, including filtering, peak removal, baseline correction, SNR screening, smoothing, and standardization.

[0005] Step 4: Use the database module to build and optimize the model; the details are as follows: first, observe the serum Raman spectrum data, and cluster the unnormalized original spectrum through an unsupervised clustering method according to its concentration characteristics. The number of clusters is determined by the elbow plot, and the characteristics of the spectrum are analyzed; then, according to the clustering results, the proportion of high intensity in the sample unsupervised clustering and the highest peak / reference peak of the original single spectrum (4800) are used as additional eigenvalues; then, the preprocessed and normalized Raman spectrum and additional eigenvalue information are combined to form the final eigenvalue column; finally, in non-image or sequence models, the final eigenvalue column data is reduced in dimension by PCA to ensure variable independence; the data set is divided into a training set and a test set by 85:15, and an external validation set is used to evaluate the generalization ability of the model in the cognitive diagnosis module; Step 5: Perform performance evaluation. After the performance evaluation, stack 3 to 4 models with good performance from traditional machine learning models and deep learning models to form an integrated learning model. During the stacking process, select the integration scheme with the smallest amount of calculation while ensuring the model performance. After stacking, determine the prediction result of each single spectrum by the maximum number of votes. Then, gather all the single spectra of the subject and use the majority mechanism to form the final prediction result of the subject. Step 6: Use the cognitive diagnosis module to perform cognitive grading diagnosis on new samples; specifically, collect Raman spectral data of new samples through the Raman spectral detection and preprocessing module, and use the model constructed in the database module to perform cognitive grading identification of samples; when external data verifies the effectiveness of the model, the cognitive diagnosis module outputs the recognition accuracy, sensitivity, and specificity of the external verification set; at the same time, the system tracks the condition of new samples input into the cognitive diagnosis module. After clinical diagnosis, the new sample data will enter the database module to further optimize the model.

[0006] Preferably, if the sample in step 2 cannot be centrifuged immediately, it needs to be stored in a 4 degree environment and centrifuged within 2 hours; if the serum cannot be collected for Raman spectrum immediately, the separated serum is collected in an EP tube and stored in a -80°C environment within 2 hours after sampling until the serum Raman spectrum is collected.

[0007] Preferably, if the serum sample is frozen in step 3, it needs to be rewarmed before spotting; the type of Raman spectrum collected is Raman fingerprint spectrum, i.e. spontaneous Raman, and its single spectrum collection parameters are: integration time 5s, laser power 20mw, spectral range 400~3400cm -1 ; Spectral data removed silent noise area of ​​1800~2700 cm -1 , is one-dimensional floating point data.

[0008] Preferably, the model building in step 4 is done in python 3.9; the libraries used include but are not limited to: scipy for statistical tools, sklearn for traditional machine learning models, tensorflow and pytorch for deep learning models, and matplotlib and seaborn for drawing; traditional machine learning models include support vector machines, linear discriminant analysis, random forests, gradient boosting, K-order nearest neighbors, artificial neural networks, etc., all of which are optimized for hyperparameters through GridSearchCV, CV=5, and the model is confirmed after selecting the best parameters and predicted on the test set, that is, the training set is divided into 5 parts through crossvalidation, and cross validation is tuned; deep learning models, i.e. deep neural network models, include feedforward neural networks FNN, recursive neural networks RNN, convolutional neural networks CNN, etc., all of which are fine-tuned based on commonly used architectures.

[0009] Preferably, the feedforward neural network FNN selects the multi-layer perceptron MLP architecture, the recurrent neural network RNN ​​selects the gated recurrent unit GRU architecture, and the convolutional neural network CNN selects the VGGNet architecture; the multi-layer perceptron MLP architecture consists of 5 hidden layers and 1 output layer, the gated recurrent unit GRU architecture consists of 2 GRU modules, 1 fully connected layer and 1 output layer, and the convolutional neural network CNN is fine-tuned on the VGGNet architecture by stacking multiple small convolution blocks to reduce parameters while ensuring the effectiveness of the model to optimize computing power requirements; in order to prevent overfitting, the adoption rate of the Dropout layer of the above architecture is 50%.

[0010] Compared with the closest prior art, the technical solution provided by the present invention has the following beneficial effects: The present invention uses serum Raman spectroscopy and machine learning technology to establish a cognitive impairment diagnosis system based on Raman spectroscopy technology, providing a new hematological detection method for the graded diagnosis of cognitive impairment; the present invention can not only diagnose cognitive impairment through serum Raman spectroscopy, but also perform detailed cognitive function grading, and combine with machine learning algorithms to build a serum Raman spectroscopy database to identify different cognitive grades; the present invention utilizes Raman spectroscopy, which is low in cost, high in accuracy, and requires a small sample size, and is highly adapted to the needs of general screening of the population, with a single sample cost of less than RMB 100. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a flow chart of a method for using serum Raman spectroscopy as a biomarker in the diagnosis of cognitive impairment according to the present invention; Figure 2 A schematic diagram of a multi-layer perceptron MLP architecture in the deep learning model of the present invention; Figure 3 This is a schematic diagram of the gated recurrent unit GRU architecture in the deep learning model of the present invention; Figure 4 A schematic diagram of the VGGNet architecture in the deep learning model of the present invention; Figure 5 This is a schematic diagram of the integrated learning model of the present invention; Figure 6 Schematic diagram of high-weight peaks of the cognitive grading diagnostic model of the present invention. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0013] See also Figure 1 : Example 1 A cognitive impairment diagnosis system based on serum Raman spectroscopy, the cognitive impairment diagnosis system includes a Raman spectroscopy detection and preprocessing module, a database module, and a cognitive diagnosis module. The Raman spectroscopy detection module is used to collect and preprocess the Raman spectra of samples, the database module is used for model construction and optimization, and the cognitive diagnosis module is used for cognitive graded diagnosis of new samples. This cognitive impairment diagnosis system based on Raman spectroscopy technology solves the problem of the lack of a blood test method for graded diagnosis of cognitive impairment suitable for population screening in the clinic. The cognitive impairment diagnosis system uses the Raman spectra of serum samples for graded diagnosis of cognitive impairment, provides a Raman spectroscopy method, collects Raman fingerprint spectrum data of serum samples of participants, and combines machine learning algorithms to build a serum Raman spectroscopy database of a certain scale to identify different cognitive grades (normal cognitive function, mild cognitive impairment, dementia).

[0014] A method for using serum Raman spectroscopy as a biomarker in the diagnosis of cognitive impairment comprises the following steps: Step 1: Sample collection; Participants were required to undergo Minimum Mental State Examination (MMSE), Montreal Cognitive Assessment (MoCA), and Daily Living Ability Scale tests, and blood samples were collected. MMSE and MoCA have characteristics of being dependent on education level. Therefore, the present invention uses the cutoff values ​​of MMSE and MoCA adjusted according to education level to grade the overall cognitive function of participants. The present invention defines a normal MMSE score as follows: illiterate participants score greater than 17 points, participants with an education level of 1-6 years score greater than 20 points, and participants with an education level of more than 6 years score greater than 24 points. The present invention defines a normal MoCA score as follows: illiterate participants score greater than 13 points, participants with an education level of 1-6 years score greater than 19 points, participants with an education level of 7-12 years score greater than 24 points, and participants with an education level of more than 12 years score greater than or equal to 26 points. In the present invention, normal cognitive function is defined as: 1) both MMSE and MoCA scores are normal; 2) the score of the Daily Living Ability Scale is 14 points. MCI is defined as: 1) cognitive impairment reported by patients, informants, or clinicians that cannot be explained by delirium or other psychiatric disorders; 2) normal MMSE score, but MoCA score does not reach the normal standard; 3) daily living ability scale score is 14 points. Dementia is defined as: 1) cognitive impairment reported by patients, informants, or clinicians that cannot be explained by delirium or other psychiatric disorders; 2) MMSE score does not reach the normal standard.

[0015] The inclusion criteria for participants are as follows: 1) meeting the definition of normal cognitive function, or MCI, or dementia as described above; 2) voluntarily participating in the present invention. The exclusion criteria for participants are as follows: 1) mental illness, such as depression, schizophrenia, etc.; 2) in certain acute states, such as acute cerebral infarction, acute cerebral hemorrhage, new craniocerebral trauma, acute migraine, vertigo attack, fever, etc.; 3) some conditions that cannot cooperate with this study, such as blindness, deafness, limb disability, etc.

[0016] Step 2: Sample pretreatment; Blood samples from participants were collected using an inert separation gel coagulation tube. After standing at room temperature for 10 minutes, the blood samples were centrifuged at 1200g for 10 minutes to separate the serum. If the sample cannot be centrifuged immediately, it must be stored in a 4-degree environment and centrifuged within 2 hours. If the serum cannot be collected for Raman spectroscopy immediately, the separated serum will be collected in an EP tube and stored at -80°C within 2 hours after sampling until the serum Raman spectrum is collected.

[0017] Step 3: Collect and pre-treat serum samples using Raman spectroscopy detection and pre-treatment modules; If the serum sample is frozen, it needs to be rewarmed before spotting. 1.5ul of serum sample is directly spotted on the low background noise chip using a pipette. Raman spectrometer WITec alpha 300R is used to collect Raman spectra of serum samples. The type of Raman spectrum collected in the present invention is Raman fingerprint spectrum (spontaneous Raman). Single spectrum collection parameters: integration time 5s, laser power 20mw, spectral range 400~3400 cm -1 Spectral data were removed to remove the silent noise area of ​​1800~2700 cm -1 , which is one-dimensional floating point data. After acquisition, the Raman spectrum data is preprocessed: filtering, peak removal, baseline correction, SNR screening, smoothing, and standardization.

[0018] Step 4: Use the database module to build and optimize the model; Serum Raman spectral data were observed and concentration characteristics were found. Therefore, the unnormalized original spectra were clustered by unsupervised clustering method, the number of clusters was determined by elbow plot, and the characteristics of the spectra were analyzed. It was found that the signal intensity showed three intervals of high, medium and low. According to the clustering results, the proportion of high intensity in the unsupervised clustering of samples and the highest peak / reference peak of the original single spectrum (4800) were used as additional eigenvalues. The pre-processed and normalized Raman spectra and additional eigenvalue information were combined to form the final eigenvalue column. In non-image or sequence models, the eigenvalue data was reduced in dimension (50 dimensions) by PCA and the independence of variables was ensured. In image or sequence models, the eigenvalue data used full spectrum data. The data set was divided into training set and test set by 85:15, and there was an external validation set to evaluate the generalization ability of the model in the cognitive diagnosis module.

[0019] The model is built in Python 3.9. The libraries used include but are not limited to: scipy for statistical tools, sklearn for traditional machine learning models, tensorflow and pytorch for deep learning models, and matplotlib and seaborn for drawing. Traditional machine learning models, including support vector machines, linear discriminant analysis, random forests, gradient boosting, K-order nearest neighbors, artificial neural networks (single layer), etc., are all optimized for hyperparameters through GridSearchCV (cv=5). After selecting the best parameters, the model is confirmed and predicted on the test set. That is, the training set is divided into 5 parts through cross validation, and cross validation is tuned. Deep learning models (deep neural network models), including feedforward neural networks FNN, recurrent neural networks RNN, convolutional neural networks CNN, etc., are all fine-tuned based on common architectures. The following is a detailed description of deep learning models.

[0020] (1) If Figure 2As shown, FNN selects the multi-layer perceptron MLP architecture. MLP is a widely used supervised learning model suitable for regression and classification tasks. MLP has strong expressive power and can approximate any continuous function, which is proved in the "Universal Approximation Theorem". MLP consists of multiple layers, each of which contains several neurons. In the present invention, MLP consists of 5 hidden layers and 1 output layer, the loss function adopts categorical_crossentropy, the optimizer is adam, and the output of the previous connection layer is converted into a probability output using a softmax optimizer, the learning rate is 0.001, the batch size is 128, and validation_split=0.2 (the training set is divided into 5 parts by cross validation). In order to prevent overfitting, the adoption rate of the Dropout layer is 50%, eliminating the contribution of 50% of the neurons to the next layer to reduce excessive reliance on certain neurons for classification.

[0021] (2) If Figure 3 As shown in the figure, RNN selects the gated recurrent unit GRU architecture. GRU is mainly used to process sequence data. Compared with LSTM, GRU has a simpler structure, fewer parameters, and relatively small computing power requirements. GRU can effectively capture long-range dependencies, and is particularly suitable for Raman spectroscopy data with a limited number of samples and a sequence relationship between wave numbers. The RNN model consists of 2 GRU modules, 1 fully connected layer, and 1 output layer. The loss function uses categorical_crossentropy, the optimizer is adam, and the softmax optimizer is used to convert the output of the previous connection layer into a probability output. The learning rate is 0.001, the batch size is 128, and validation_split=0.2. In order to prevent overfitting, the Dropout layer is adopted at a rate of 50%, eliminating the contribution of 50% of the neurons to the next layer to reduce excessive reliance on certain neurons for classification.

[0022] (3) If Figure 4As shown, CNN selects the VGGNet architecture. CNN is a common deep learning architecture for image processing and is suitable for Raman spectroscopy data. The CNN architecture of the present invention is fine-tuned on the VGGNet architecture. By stacking multiple small convolution blocks, the parameters are reduced while ensuring the effectiveness of the model to optimize the computing power requirements. The loss function uses categorical_crossentropy, the optimizer is adam, and the output of the previous connection layer is converted into a probability output using a softmax optimizer. The learning rate is 0.001, the batch size is 128, and validation_split=0.2. In order to prevent overfitting, the Dropout layer is adopted at a rate of 50%, eliminating the contribution of 50% of the neurons to the next layer to reduce excessive reliance on certain neurons for classification.

[0023] All the above machine learning models are evaluated on the test set, and the evaluation indicators include accuracy, sensitivity, and specificity. The indicators are defined as follows: (1) Accuracy: It is the ratio of the number of samples correctly predicted by the model to the total number of samples. It indicates the proportion of samples correctly classified by the model among all samples. Its calculation formula is: Accuracy = (TP+TN) / (TP+TN+FP+FN); where: TP (True Positive): the number of samples correctly classified as positive; TN (True Negative): the number of samples correctly classified as negative; FP (False Positive): the number of negative samples incorrectly classified as positive; FN (False Negative): the number of positive samples incorrectly classified as negative.

[0024] (2) Sensitivity: Also known as recall rate, it indicates the proportion of positive samples correctly identified by the model to all positive samples. It measures the model's ability to capture positive samples. Its calculation formula is: Sensitivity = TP / (TP+FN).

[0025] (3) Specificity: It indicates the proportion of negative samples correctly identified by the model to all negative samples. It measures the model's ability to identify negative samples. Its calculation formula is: Specificity = TN / (FP+TN).

[0026] Step 5: Perform performance evaluation. After the performance evaluation, stack 3 to 4 models with good performance among the above traditional machine learning models and deep learning models to form an integrated learning model. The integrated learning model can combine the advantages of different models and improve the accuracy and stability of prediction. In addition, the data features and patterns captured by each model may be different, and stacking helps to understand the data more comprehensively. Secondly, this method can also reduce the risk of overfitting of a single model, enhance the generalization ability of the model, and perform more robustly on new data. In addition, stacking allows different types of models to be flexibly combined, making full use of the characteristics of various algorithms, thereby providing more powerful solutions for complex problems. In the stacking process, while ensuring the performance of the model, select the integrated solution with the least computational effort. After stacking, the prediction result of each single spectrum is determined by majority vote. All the single spectra of the subject are then collected, and the majority mechanism is used to form the final prediction result of the subject.

[0027] Step 6: Use the cognitive diagnosis module to perform cognitive grading diagnosis on new samples; The Raman spectrum data of new samples are collected through the Raman spectrum detection and preprocessing module, and the model constructed in the database module is used to perform cognitive graded identification of samples. When the external data verifies the effectiveness of the model, the cognitive diagnosis module outputs the recognition accuracy, sensitivity, and specificity of the external verification set. At the same time, the system tracks the condition of new samples input into the cognitive diagnosis module. After clinical diagnosis, the new sample data will enter the database module to further optimize the model.

[0028] Example 2 The participants included in the present invention were recruited by neurologists from Nanjing Drum Tower Hospital, and a total of 260 participants were recruited from Nanjing Drum Tower Hospital and the community. In Example 2 of the present invention, an integrated learning model was constructed with the help of 220 serum samples (93 cases of normal cognitive function, 70 cases of mild cognitive impairment, and 57 cases of dementia), and a cognitive hierarchical recognition model was established. The cognitive hierarchical recognition model was verified in a cohort consisting of 14 participants with normal cognitive function, 16 participants with mild cognitive impairment, and 10 participants with dementia. The cognitive impairment grading diagnostic ability of the cognitive hierarchical recognition model was verified.

[0029] The statistical analysis of general data in the present invention was completed by SPSS (version 22). The count data were compared between groups using chi-square test or Fisher's exact test. All measurement data were tested for normality (PP chart) using descriptive statistical methods before inter-group comparison. The measurement data that met the normal distribution were compared between the three groups using one-way analysis of variance, and the data were described by mean ± standard deviation. The measurement data that did not meet the normal distribution were compared between the three groups using the Kruskal-Wallis H test, and the median (interquartile range) was used to describe the data. The statistical analysis of general data was statistically significant with P<0.05.

[0030] Furthermore, among the 220 participants recruited from Nanjing Drum Tower Hospital, the proportion of males in participants with normal cognitive function was 37.6%, with an average age of 66.753±9.185 years, an average education of 12.000 (7.000) years, an average MMSE score of 29.000 (2.000), and an average MoCA score of 26.000 (2.000); the proportion of males in participants with mild cognitive impairment was 47.1%, with an average age of 70.171±7.173 years, and an average education of 12.000 (7.000) years. The average years of education were 12.000 (6.000) years, the average MMSE score was 28.000 (3.000), and the average MoCA score was 22.000 (6.000); the proportion of males among participants with dementia was 42.1%, the average age was 73.702±9.575 years, the average years of education were 12.000 (9.000) years, the average MMSE score was 19.000 (10.000), and the average MoCA score was 14.500 (8.000). Among them, 5 patients with dementia did not complete the MoCA scale assessment. There were statistical differences in age, MMSE, and MoCA scores among participants with normal cognitive function, mild cognitive impairment, and dementia, but no statistical differences in gender and years of education. See Table 1 for details.

[0031] Table 1 Basic characteristics of database module participants

[0032] Note: One-way ANOVA was used to compare the differences between groups for the measurement data that met the normal distribution, and the values ​​of each group were described by mean ± standard deviation; Kruskal-Wallis H test was used to compare the differences between groups for the measurement data that did not meet the normal distribution, and the values ​​of each group were described by median (interquartile range). Abbreviations: MMSE: Mini-Mental State Examination; MoCA: Montreal Cognitive Assessment; *: 5 patients with dementia did not complete the MoCA assessment.

[0033] Furthermore, among the 40 participants recruited from the community, the proportion of males in participants with normal cognitive function was 35.7%, with an average age of 71.000±6.839 years, an average education of 10.000±5.129 years, an average MMSE score of 28.143±2.070, and an average MoCA score of 26.000 (3.250). The proportion of males in participants with mild cognitive impairment was 50.0%, with an average age of 66.688±9.329 years, and an average The average years of education were 10.625±4.829 years, the average MMSE score was 27.750±1.528, and the average MoCA score was 22.500 (6.000); the proportion of male participants with dementia was 50.0%, the average age was 67.800±8.29 years, the average years of education was 9.400±4.274 years, the average MMSE score was 18.100±6.173, and the average MoCA score was 14.000 (9.500). Among them, one patient with dementia did not complete the MoCA scale assessment. There were statistical differences in MMSE and MoCA scores among participants with normal cognitive function, mild cognitive impairment, and dementia, but no statistical differences in age, gender, and years of education. See Table 2 for details.

[0034] Table 2 Basic characteristics of participants in the cognitive diagnosis module

[0035] Note: One-way ANOVA was used to compare the differences between groups for the measurement data that met the normal distribution, and the values ​​of each group were described by mean ± standard deviation; Kruskal-Wallis H test was used to compare the differences between groups for the measurement data that did not meet the normal distribution, and the values ​​of each group were described by median (interquartile range). Abbreviations: MMSE: Mini-Mental State Examination; MoCA: Montreal Cognitive Assessment; *: One patient with dementia did not complete the MoCA assessment.

[0036] Conduct a systematic diagnostic performance test for the cognitive impairment grading system.

[0037] The database module of the present invention uses serum Raman spectra of 220 participants from Nanjing Drum Tower Hospital for database modeling. After preprocessing, there are 10,877 valid spectra in total. The 220 participants were divided into training set and test set at a ratio of 85:15. After division, the training set had a total of 186 participants, including 79 with normal cognitive function, 59 with mild cognitive impairment, and 48 with dementia; the test set had a total of 34 participants, including 14 with normal cognitive function, 11 with mild cognitive impairment, and 9 with dementia. In the process of constructing the integrated learning model, different machine learning models and deep learning models are combined. Under the premise of ensuring the effectiveness of the model, the integrated scheme with the least computational effort is selected; the model effectiveness is evaluated by accuracy, sensitivity, and specificity. After training, the integrated learning model is composed of linear discriminant analysis, artificial neural network (single layer), and multi-layer perceptron MLP stacked together. The test set accuracy is 0.85, the average sensitivity is 0.85, and the average specificity is 0.92. See for details. Figure 5 And as shown in Table 3.

[0038] Table 3 Database module test set performance table

[0039] The cognitive diagnosis module of the present invention uses serum Raman spectra of 40 participants from the community to verify the system performance. After preprocessing, there are 800 valid spectra. The serum Raman spectra of 40 participants from the community are input into the cognitive diagnosis module, and the accuracy of cognitive grading is 0.80, the average sensitivity is 0.80, and the average specificity is 0.89. As shown in Table 4.

[0040] Table 4 Cognitive diagnosis module system effectiveness verification table

[0041] Furthermore, the feature weights in the linear discriminant model were ranked, and the top three features were: Raman peak at 1602 cm -1 , Raman peak at 1002cm -1 , Raman peak at 1666cm -1 .like Figure 6 Shown (1) Raman peak at 1602 cm -1 : Related to the C=C plane bending mode of aromatic amino acids, such as phenylalanine and tyrosine, which play an important role in protein structure. 1602cm -1 It can reflect the level of mitochondrial activity, can be used to monitor cell biological activity, and has the ability to monitor brain oxygenation. It is called the "Raman spectral signature of life."

[0042] (2) Raman peak at 1002 cm -1 : Related to the stretching vibration of the CC aromatic ring of phenylalanine. 1002cm-1 It is one of the characteristic peaks of phenylalanine in Raman spectrum, usually called "benzene ring breathing mode", which is a direct manifestation of the vibration of the aromatic ring structure in the phenylalanine molecule and is often used to study the conformation and environment of biological molecules containing phenylalanine. -1 Changes in the intensity or position of this characteristic peak indicate that the local environment of phenylalanine in these aggregates has changed.

[0043] (3) Raman peak at 1666cm -1 :It is related to the vibration mode of Amide I of the protein and reflects the changes in the secondary structure of the protein (such as α-helix, β-fold and disordered structure). -1 Changes in the intensity or position of this characteristic peak can be used to assess conformational changes in the protein during the development and progression of cognitive impairment.

[0044] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, a person skilled in the art can still modify or make equivalent substitutions to the specific implementations of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention are within the scope of protection of the claims of the present invention to be approved.

Claims

1. A method for using serum Raman spectroscopy as a biomarker in the diagnosis of cognitive impairment, characterized in that: Diagnosis is performed using a cognitive impairment diagnosis system, the cognitive impairment diagnosis system comprising a Raman spectrum detection and preprocessing module, a database module, and a cognitive diagnosis module, and the application method thereof comprises the following steps: Step 1: Sample collection: recruit participants to collect blood samples, and perform the Mini-Mental State Examination, Montreal Cognitive Assessment, and Daily Living Ability Scale tests; Step 2: Sample pretreatment: Collect blood samples from participants using an inert separation gel coagulation tube. After standing at room temperature for 10 minutes, centrifuge the blood samples at 1200g for 10 minutes to separate serum and obtain serum samples. Step 3: Use Raman spectroscopy detection and preprocessing module to collect and preprocess serum samples; specifically, 1.5ul of serum sample is directly spotted on a low background noise chip using a pipette, and Raman spectrometer WITecalpha300R is used to collect Raman spectra of serum samples. After collection, Raman spectroscopy data is preprocessed, including filtering, peak removal, baseline correction, SNR screening, smoothing, and standardization; Step 4: Use the database module to build and optimize the model; the details are as follows: first, observe the serum Raman spectrum data, and cluster the unnormalized original spectrum through an unsupervised clustering method according to its concentration characteristics. The number of clusters is determined by the elbow plot, and the characteristics of the spectrum are analyzed; then, according to the clustering results, the proportion of high intensity in the sample unsupervised clustering and the highest peak / reference peak of the original single spectrum (4800) are used as additional eigenvalues; then, the preprocessed and normalized Raman spectrum and additional eigenvalue information are combined to form the final eigenvalue column; finally, in non-image or sequence models, the final eigenvalue column data is reduced in dimension by PCA to ensure variable independence; the data set is divided into a training set and a test set by 85:15, and an external validation set is used to evaluate the generalization ability of the model in the cognitive diagnosis module; Step 5: Perform performance evaluation. After the performance evaluation, stack 3 to 4 models with good performance from traditional machine learning models and deep learning models to form an integrated learning model. During the stacking process, select the integration scheme with the smallest amount of calculation while ensuring the model performance. After stacking, determine the prediction result of each single spectrum by the maximum number of votes. Then, collect all the single spectra of the subject and use the majority mechanism to form the final prediction result of the subject. In the above integrated learning model construction process, combine different machine learning models and deep learning models to evaluate the model performance by accuracy, sensitivity, and specificity. After training, the final integrated learning model is composed of linear discriminant analysis, artificial neural network (single layer), and multi-layer perceptron MLP stacked together. Step 6: Use the cognitive diagnosis module to perform cognitive grading diagnosis on new samples; specifically, collect Raman spectral data of new samples through the Raman spectral detection and preprocessing module, and use the model constructed in the database module to perform cognitive grading identification of samples; when external data verifies the effectiveness of the model, the cognitive diagnosis module outputs the recognition accuracy, sensitivity, and specificity of the external verification set; at the same time, the system tracks the condition of new samples input into the cognitive diagnosis module. After clinical diagnosis, the new sample data will enter the database module to further optimize the model.

2. The method for diagnosing cognitive impairment based on serum Raman spectroscopy according to claim 1, characterized in that: If the sample in step 2 cannot be centrifuged immediately, it must be stored in a 4 degree environment and centrifuged within 2 hours; if the serum cannot be collected for Raman spectrum immediately, the separated serum is collected in an EP tube and stored at -80°C within 2 hours after sampling until the serum Raman spectrum is collected.

3. The method for diagnosing cognitive impairment based on serum Raman spectroscopy according to claim 1, characterized in that: If the serum sample in step 3 is frozen, it needs to be rewarmed before spotting; the type of Raman spectrum collected is Raman fingerprint spectrum, i.e. spontaneous Raman, and its single spectrum collection parameters are: integration time 5s, laser power 20mw, spectral range 400~3400cm -1 ; Spectral data removed silent noise area of ​​1800~2700 cm -1 , is one-dimensional floating point data.

4. The method for diagnosing cognitive impairment based on serum Raman spectroscopy according to claim 1, characterized in that: In step 4, python is used for model building; the libraries used include but are not limited to: scipy is used for statistical tools, sklearn is used for traditional machine learning models, tensorflow and pytorch are used for deep learning models, and matplotlib and seaborn are used for drawing; traditional machine learning models include support vector machines, linear discriminant analysis, random forests, gradient boosting, K-order nearest neighbors, artificial neural networks, etc., all of which are optimized for hyperparameters through GridSearchCV, CV=5, and the model is confirmed after selecting the best parameters and predicted on the test set, that is, the training set is divided into 5 parts through crossvalidation, and cross validation is tuned; deep learning models, namely deep neural network models, include feedforward neural networks FNN, recursive neural networks RNN, convolutional neural networks CNN, etc., all of which are fine-tuned based on commonly used architectures.

5. The method for diagnosing cognitive impairment based on serum Raman spectroscopy according to claim 4, characterized in that: The feedforward neural network FNN selects a multilayer perceptron MLP architecture, the recurrent neural network RNN ​​selects a gated recurrent unit GRU architecture, and the convolutional neural network CNN selects a VGGNet architecture; the multilayer perceptron MLP architecture consists of 5 hidden layers and 1 output layer, the gated recurrent unit GRU architecture consists of 2 GRU modules, 1 fully connected layer and 1 output layer, and the convolutional neural network CNN is fine-tuned on the VGGNet architecture, and the adoption rate of the Dropout layer of the above architecture is 50%.

Citation Information

Cited By

  • GRU regression model and method for realizing intelligent detection of deltamethrin pesticide residues by using same

    CN120744869A

  • Serum Raman spectrum non-invasive pathological detection device

    CN121186010A