Biomarker for diagnosing Alzheimer disease

Through Olink technology, biomarkers such as NEFL, GIP, and TAFA5 were screened out, and combined with machine learning models, the problem of early diagnosis of Alzheimer's disease was solved, and the diagnostic effect with high sensitivity and high specificity was achieved, and the diagnostic accuracy was improved.

CN120490504APending Publication Date: 2025-08-15SHANGHAI MENTAL HEALTH CENT (SHANGHAI PSYCHOLOGICAL COUNSELLING TRAINING CENT) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510644166.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve early diagnosis of Alzheimer's disease. Common methods can only be detected after the patient has clinical symptoms, and lack non-invasive and efficient biomarkers.

Method used

3072 proteins were detected through Olink's ortho-extension analysis technology, and one or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21 or TNFRSF9 were screened as biomarkers. Combined with machine learning models, a method for diagnosing Alzheimer's disease was established.

Benefits of technology

High sensitivity and high specificity of early diagnosis of Alzheimer's disease have been achieved, with improved accuracy, and more biologically significant protein markers can be detected in plasma, providing an earlier diagnostic window.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120490504A_ABST
    Figure CN120490504A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological medicines, and particularly relates to a biomarker for diagnosing Alzheimer's disease. The biomarker comprises one or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21 or TNFRSF9 (tumor necrosis factor receptor like protein). The invention further relates to application of the biomarker in diagnosis of Alzheimer's disease, and a good diagnosis effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedicine technology, and specifically relates to a biomarker for diagnosing Alzheimer's disease. Background Art

[0002] Alzheimer's disease (AD), the most common type of dementia, is a chronic neurodegenerative disease characterized by memory impairment, personality changes, and behavioral degeneration.

[0003] The pathogenesis and causes of AD remain incompletely understood. The amyloid β hypothesis proposes that β-amyloid (Aβ) protein deposits, while the tau hypothesis proposes that tau neurofibrillary tangles impair brain function, ultimately leading to severe cognitive impairment, memory loss, and metabolic disturbances. While there is no definitive treatment for AD, appropriate care can delay disease progression and prevent further cognitive impairment. Research suggests that abnormalities in neurodegenerative biomarkers in AD may occur decades before clinical dementia is diagnosed, providing a valuable window for early diagnosis and intervention. Therefore, early diagnosis of AD is crucial for patients to mitigate or delay the onset of clinical symptoms. However, many potential AD patients go undiagnosed in a timely and accurate manner. In high-income countries, only 20-50% of dementia cases are recorded in primary care, and this percentage is even lower in low- and middle-income countries. This lack of formal diagnosis means that patients cannot access necessary treatment and care, and are unable to participate in critical support programs. Early diagnosis and intervention are crucial to narrowing this treatment gap.

[0004] As the prevalence of AD increases, the demand for more effective diagnostic methods is also growing. Currently, there are no ideal diagnostic markers for early detection. Commonly used AD diagnostic methods include head CT and MRI, amyloid protein PET imaging, genetic testing, cerebrospinal fluid amyloid protein, and Tau protein testing. These indicators are usually only detected or identified after Alzheimer's disease has progressed for many years and clinical symptoms have appeared, making early diagnosis difficult.

[0005] The use of serum / plasma biomarkers for the diagnosis of AD will achieve better compliance due to its non-invasive nature. Reported AD plasma protein markers include Aβ42 / 40, P-tau217, P-tau181, NfL, GFAP, etc. The detection of these biomarkers generally has high sensitivity and specificity. For example, p-Tau has been shown to have a sensitivity and specificity of over 80% in the diagnosis of AD. In addition, multiple proteins can be used in combination. In 2018, the National Institute of Aging and Alzheimer's Association (NIA-AA) proposed a biological definition of AD, namely the ATN (Aβ-Tau-Neurodegeneration, ATN) diagnostic criteria based on Aβ protein (A), pathological Tau protein (T) and neurodegeneration (N) biomarkers. This is an important breakthrough in the recent use of a combination of AD biomarkers to guide the standardization of early clinical intervention for AD. In 2019, Academician Ye Yuru from Hong Kong, China published an article titled "Large-scale plasma proteomic profiling identifies a high-performance biomarker panel for Alzheimer's disease screening and staging" in the journal AlzheimersDement. The article showed that the ATN combination marker had an AUC of 87.35% for AD diagnosis. At the same time, a panel of 19 proteins was found from 1,160 plasma proteins, which could accurately classify clinical AD (the AUC of the training set was 0.9816 and the AUC of the validation set was 0.9690).

[0006] The above studies demonstrate that plasma protein markers are not only highly accurate in diagnosing AD but can also achieve early diagnosis. However, the plasma protein markers screened in these studies were selected from a pool of over 1,000 candidate proteins, while the number of proteins in plasma far exceeds this number, potentially leading to the omission of valuable markers. Summary of the Invention

[0007] To address this issue, the present invention used Olink's proximity extension assay (PEA) to detect 3,072 proteins and screened for differentially expressed proteins in the plasma of Alzheimer's patients. Machine learning analysis of the differential protein data revealed seven proteins with significant differential expression between the two groups, demonstrating excellent diagnostic efficacy. The detailed information and differential expression of these seven proteins are shown in Table 4 below.

[0008] Therefore, in a first aspect, the present invention provides a biomarker for diagnosing Alzheimer's disease, wherein the biomarker comprises one or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21 or TNFRSF9.

[0009] In some embodiments, the biomarker may be only one of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21, or TNFRSF9, or a combination of two or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21, or TNFRSF9.

[0010] In some embodiments, the biomarkers comprise at least NEFL, GIP, and TAFA5.

[0011] In some embodiments, the biomarkers comprise at least NEFL, GIP, TAFA5, and TNFRSF9.

[0012] In some embodiments, the biomarkers comprise at least NEFL, GIP, TAFA5, FABP6, and TNFRSF9.

[0013] In some embodiments, the biomarkers comprise at least NEFL, GIP, TAFA5, FABP6, RBP2, and TNFRSF9.

[0014] In some embodiments, the biomarkers consist of NEFL, GIP, TAFA5, FABP6, RBP2, FGF21, and TNFRSF9.

[0015] As used herein, "diagnosing Alzheimer's disease" includes assessing whether a subject has Alzheimer's disease or is at risk of developing Alzheimer's disease (e.g., has an increased risk of developing Alzheimer's disease) or monitoring the severity or progression of Alzheimer's disease in a subject.

[0016] As used herein, the terms "biomarker" and "marker" are used interchangeably and can refer to a gene, mRNA, cDNA, antisense transcript, miRNA, polypeptide, protein, protein fragment, or any other nucleic acid sequence or polypeptide sequence that indicates the level of gene expression or protein production. When a biomarker indicates an abnormal process, disease or other condition in an individual, or a sign thereof, such biomarker is generally described as being overexpressed or underexpressed compared to the expression level or value of a biomarker that represents a normal process, the absence of a disease or other condition, or a sign thereof in the individual. "Upregulated," "overexpressed," and any variations thereof are used interchangeably to refer to a value or level of a biomarker in a biological sample that is greater than the value or level (or range of values or levels) of the biomarker typically detected in a similar biological sample from a healthy or normal individual. The term can also refer to a value or level of a biomarker in a biological sample that is greater than the value or level (or range of values or levels) of the biomarker that can be detected at different stages of a particular disease. "Downregulated," "underexpressed," and any variations thereof are used interchangeably to refer to a value or level of a biomarker in a biological sample that is less than the value or level (or range of values or levels) of that biomarker typically detected in a similar biological sample from a healthy or normal individual. The term may also refer to a value or level of a biomarker in a biological sample that is less than the value or level (or range of values or levels) of that biomarker that can be detected at different stages of a particular disease.

[0017] Furthermore, a biomarker that is overexpressed or underexpressed can also be referred to as being "differentially expressed," or having a "differential level," or "differential value," compared to a "normal" expression level or value of the biomarker, which indicates a normal process in an individual or the absence of a disease or other condition or a sign thereof. Thus, "differential expression" of a biomarker can also be referred to as a variation from the "normal" expression level of the biomarker.

[0018] The terms "differential biomarker expression" and "differential expression" are used interchangeably to refer to biomarkers in which the expression of the biomarker in subjects with a particular disease is activated to a higher or lower level relative to its expression in normal subjects, or in patients who respond differently to a particular therapy or have a different prognosis. The term also includes biomarkers whose expression is activated to a higher or lower level at different stages of the same disease. It should also be understood that differentially expressed biomarkers can be activated or inhibited at the nucleic acid level or protein level, or can undergo alternative splicing to produce different polypeptide products. Such differences can be evidenced by a variety of changes, including changes in the level of mRNA, miRNA, antisense transcripts, or protein surface expression, secretion, or other distribution. Differential biomarker expression can include comparing the expression of two or more genes or their gene products; or comparing the ratio of expression between two or more genes or their gene products; or even comparing two differently processed products of the same gene, wherein the processed products differ between normal subjects and subjects with the disease or between different stages of the same disease. Differential expression includes, for example, quantitative differences as well as qualitative differences in the spatial or cellular expression pattern of a biomarker between normal and diseased cells or between cells that have undergone different disease events or disease stages.

[0019] Biomarker expression level analysis methods include, but are not limited to, quantitative PCR, NGS, Northern blots, Southern blots, microarrays, SAGE, Simoa, immunoassays (ELISA, EIA, agglutination, nephelometry, turbidimetry, Western blotting, immunoprecipitation, immunocytochemistry, flow cytometry, Luminex assays), and mass spectrometry. The overall expression data for a given sample can be normalized / standardized using methods known to those skilled in the art to correct for different amounts of starting material, different efficiencies of extraction and amplification reactions.

[0020] In a second aspect, the present invention provides use of a reagent for detecting the expression level of a biomarker as described herein in the preparation of a product for diagnosing Alzheimer's disease.

[0021] In some embodiments, for example, in embodiments where the biomarker is one of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21, or TNFRSF9, the diagnosis comprises detecting the expression level of the biomarker from a test sample obtained from the subject and comparing it to the expression level of the biomarker detected from a control sample, and determining the subject as having or not having Alzheimer's disease, or having or not having a risk of having Alzheimer's disease (e.g., an increased risk of having Alzheimer's disease), or having or not having worsening or progressive Alzheimer's disease based on whether the expression level of the biomarker in the test sample is upregulated or the expression level of TAFA5, RBP2, FABP6, FGF21, or TNFRSF9 is downregulated in the test sample relative to the control sample, then the subject is diagnosed as having Alzheimer's disease, or having a risk of having Alzheimer's disease (e.g., an increased risk of having Alzheimer's disease), or having or not having worsening or progressive Alzheimer's disease. In some embodiments, the control sample includes a sample from a healthy or normal subject.

[0022] In some embodiments, for example, embodiments where the biomarker is a combination of two or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21, or TNFRSF9, the diagnosis is performed by:

[0023] (1) detecting the expression level of the biomarker from a test sample obtained from a subject;

[0024] (2) determining the score of the test sample based on the expression level of the biomarker;

[0025] (3) comparing the test sample's score to a threshold value; and

[0026] (4) determining the subject as having or not having Alzheimer's disease, or having or not having a risk of having Alzheimer's disease (e.g., an increased risk of having Alzheimer's disease), or having or not having worsening or progressive Alzheimer's disease based on whether the score of the test sample is higher or lower than a threshold value.

[0027] As used herein, the term "subject" includes human and non-human animals, preferably a human subject."Subject," "individual," and "patient" are used interchangeably herein.

[0028] As used herein, "test sample," "biological sample," and "sample" are used interchangeably herein to refer to any material, biological fluid, tissue, or cell obtained or derived from a subject. This includes blood (including whole blood, leukocytes, peripheral blood mononuclear cells, buffy coat, plasma, and serum), sputum, tears, mucus, nasal washes, nasal aspirates, breath, urine, semen, saliva, meningeal fluid, amniotic fluid, glandular fluid, lymph fluid, nipple aspirates, bronchial aspirates, synovial fluid, joint aspirates, ascites, cells, cell extracts, and cerebrospinal fluid. This also includes experimentally separated portions of all the aforementioned samples. For example, a blood sample can be separated into serum or into fractions containing specific types of blood cells, such as red blood cells or white blood cells (leukocytes). If desired, a sample can be a combination of samples from an individual, such as a combination of tissue and fluid samples. The term "biological sample" also includes materials containing homogenized solids (e.g., from stool samples, tissue samples, or tissue biopsy samples). The term "biological sample" also includes materials derived from tissue culture or cell culture. Any suitable method for obtaining a biological sample can be used; exemplary methods include, for example, phlebotomy, cotton swabs (e.g., cheek swabs), and fine needle aspiration biopsy. Samples can also be collected, for example, by microdissection (e.g., laser capture microdissection (LCM) or laser microdissection (LMD)), bladder washes, smears (e.g., PAP smears), or ductal lavage. A "biological sample" obtained or derived from a subject includes any such sample that has been processed in any suitable manner (e.g., fresh frozen or formalin-fixed and / or paraffin-embedded) after being obtained from the subject. The methods of the present invention as defined herein can begin with the sample obtained and, therefore, do not necessarily incorporate the step of obtaining a sample from a patient. These methods can be in vitro methods performed on isolated samples.

[0029] In preferred embodiments, the test sample is a serum or plasma or whole blood sample. In alternative cases, a whole blood sample can be used.

[0030] As used herein, reagents for detecting the expression level of a biomarker may include any reagent that detects the expression of a biomarker at the gene level and / or protein level, which may include, but are not limited to, reagents for one or more detection techniques selected from the group consisting of quantitative PCR, NGS, Northern blot, Southern blot, microarray, SAGE, Simoa, immunoassay (ELISA, EIA, agglutination, nephelometry, turbidimetry, Western blotting, immunoprecipitation, immunocytochemistry, flow cytometry, Luminex assay), and mass spectrometry.

[0031] In some embodiments, the score determined for the test sample in step (2) is obtained by calculating the sum of the expression levels of each biomarker multiplied by a corresponding coefficient.

[0032] In specific embodiments, the coefficients are determined by a machine learning model based on the selected biomarkers and their expression levels in a training set.

[0033] In some embodiments, the machine learning model may include a machine learning analysis of a lasso regression model.

[0034] In a more specific embodiment, the corresponding coefficient for each biomarker is as shown in Table 5 below.

[0035] In a more specific embodiment, the calculation formula for the score of the test sample is shown in Table 5 below.

[0036] As used herein, the threshold value is a score that can best distinguish Alzheimer's disease from normal subjects, obtained by comparing the scores of Alzheimer's disease samples with normal samples in a training set.

[0037] In a more specific embodiment, the threshold value is as shown in Table 6 below.

[0038] As used herein, a training set includes samples from subjects known to have Alzheimer's disease and samples from healthy subjects.

[0039] Thus, in some embodiments, the diagnosis is performed by:

[0040] (1) detecting the expression level of the biomarker from a training set including samples from subjects known to have Alzheimer's disease and samples from healthy or normal subjects, analyzing the expression level data using a machine learning model, and determining a corresponding coefficient for each biomarker and a diagnostic threshold for the biomarker;

[0041] (2) detecting the expression level of the biomarker from a test sample obtained from the subject to be tested;

[0042] (3) determining a score for the test sample based on the expression levels of the biomarkers by calculating the sum of the expression levels of each biomarker multiplied by a corresponding coefficient;

[0043] (4) comparing the test sample's score with a threshold value; and

[0044] (5) Determining the subject as having or not having Alzheimer's disease, or having or not having the risk of having Alzheimer's disease, or having or not having worsening or progressive Alzheimer's disease based on whether the score of the test sample is higher or lower than a threshold value.

[0045] In a third aspect, the present invention provides a product for diagnosing Alzheimer's disease, comprising a reagent for detecting the expression level of a biomarker as described herein.

[0046] In some embodiments, the article of manufacture comprises a kit.

[0047] In a fourth aspect, the present invention provides a system for diagnosing Alzheimer's disease, the system comprising:

[0048] (i) a data collection module: in which the expression levels of the biomarkers described herein are detected from a test sample obtained from a subject;

[0049] (ii) a data analysis module: in which a score of the test sample is determined based on the expression level of the biomarker; and

[0050] (iii) an output module: determining whether the subject has or does not have Alzheimer's disease, or has or does not have a risk of having Alzheimer's disease (e.g., an increased risk of having Alzheimer's disease), or has or does not have worsening or progressive Alzheimer's disease based on whether the score of the test sample is higher or lower than a threshold.

[0051] In some embodiments, in the data analysis module, the score of the test sample is determined by calculating the sum of the expression levels of each biomarker multiplied by a corresponding coefficient.

[0052] In specific embodiments, the coefficients are determined by a machine learning model based on the selected biomarkers and their expression levels in a training set.

[0053] In some embodiments, the machine learning model may include a machine learning analysis of a lasso regression model.

[0054] In a more specific embodiment, the corresponding coefficient for each biomarker is as shown in Table 5 below.

[0055] In a more specific embodiment, the calculation formula for the score of the test sample is shown in Table 5 below.

[0056] In a more specific embodiment, the threshold value is as shown in Table 6 below.

[0057] Therefore, in some embodiments, the system further comprises a training module for detecting the expression levels of the biomarkers from a training set comprising samples from subjects known to have Alzheimer's disease and samples from healthy or normal subjects, analyzing the expression level data through a machine learning model, and determining a corresponding coefficient for each biomarker and a diagnostic threshold for the biomarker.

[0058] Advantageous Effects of the Invention

[0059] The present invention aims to find specific protein markers for Alzheimer's disease. Compared with the existing technology, the present invention has the following two major advantages:

[0060] 1. Develop protein markers using Olink, a stable and reliable screening and validation technology platform

[0061] The present invention uses the 3072 panel of the Olink Explore platform to screen for differentially expressed proteins in Alzheimer's disease and normal samples. Previously, mass spectrometry-based proteomic detection methods were limited by their detection sensitivity, and it was difficult to detect more than 1000-2000 proteins in plasma / serum, and low-abundance protein markers with biological significance were often missed. Olink Explore's 3072 panel can detect more than 3000 protein markers in the same plasma or other body fluid sample, involving biological processes such as signal transduction, cellular metabolic processes, cell differentiation, and cell adhesion. This means that Olink plasma proteomic detection has achieved "unbiasedness" in the biological sense and is of epoch-making significance. Olink products can achieve good data reproducibility, have good linear effects, and can stably and reliably screen target markers in plasma. By comparing the protein expression of Alzheimer's disease and normal plasma samples, we obtained 995 significantly differentially expressed proteins (p<0.05). Further, through machine learning analysis of the lasso regression model, we obtained a marker panel that can distinguish Alzheimer's disease from normal samples with high specificity and high sensitivity.

[0062] 2. The marker combination obtained by the present invention has a higher accuracy rate

[0063] Among 995 significantly differentially expressed proteins, seven protein markers were identified through machine learning analysis using a lasso regression model. Compared with previously reported methods, these markers have a higher accuracy rate.

[0064] When present individually, these seven proteins demonstrated a moderate discriminatory effect for AD, with AUCs ranging from 0.84 to 0.90. A combination of just three proteins (NEFL, GIP, and TAFA5) achieved excellent discrimination for AD, achieving AUCs of 1 in the training set and 0.980 in the validation set. These three proteins can also be combined with other genes to form 4-7 gene panels, which achieve even better discrimination between Alzheimer's disease and normal controls. With 5-7 genes, the ability to distinguish between the two groups reached AUCs of 1 in both the training and validation sets. Table 2 below shows the AUC values for the discrimination of 3-7 protein panels for AD in the training and validation sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A volcano plot showing the differentially expressed proteins in Alzheimer's disease was shown.

[0066] Figure 2 The box plot of the expression differences of the seven screened protein markers in Alzheimer's disease and control samples is shown.

[0067] Figure 3 The ROC curves of the AD differentiation effects of the 7 screened protein markers are shown.

[0068] Figure 4 Receiver operating characteristic (ROC) curves showing the ability of combinations containing different numbers of proteins to distinguish Alzheimer's disease samples from normal controls. A: NEFL; GIP; TAFA5 combination. B: NEFL; GIP; TNFRSF9; TAFA5 combination. C: FABP6; NEFL; GIP; TNFRSF9; TAFA5 combination. D: RBP2; FABP6; NEFL; GIP; TNFRSF9; TAFA5 combination. E: RBP2; FABP6; NEFL; GIP; FGF21; TNFRSF9; TAFA5 combination. DETAILED DESCRIPTION

[0069] Olink's Proximity Extension Assay (PEA) technology is a high-throughput, high-specificity, high-sensitivity, and high-dynamic-range targeted proteome quantitative detection platform. It increases the specificity of target protein recognition through the specific binding of dual antibodies in the antigen-antibody reaction. During the detection phase, the immune reaction products are converted into nucleic acid products and detected by quantitative PCR or next-generation sequencing, ensuring the linear range of the results while increasing the sensitivity of the detection.

[0070] In this study, we aimed to find specific protein markers for Alzheimer's disease in plasma. We applied the 3072 panel product of OlinkPEA technology to screen differentially expressed proteins in Alzheimer's disease and normal samples. This product panel can simultaneously measure 3072 proteins, involving biological processes such as signal transduction, cellular metabolic processes, cell differentiation, and cell adhesion. By comparing the protein expression of Alzheimer's disease and normal plasma samples, we obtained 995 significantly differentially expressed proteins (p<0.05). The differential protein data were then subjected to machine learning analysis using the lasso regression model to screen the optimal protein combination panel as a marker for diagnosing Alzheimer's disease. Seven proteins were found to be significantly differentially expressed between the two groups, and had a good diagnostic effect.

[0071] The present invention is further described below with reference to specific examples, which, however, are not intended to limit the present invention in any way. Unless otherwise specified, the reagents, methods, and equipment used in the present invention are conventional reagents, methods, and equipment in the art.

[0072] Example: Screening and identification of specific protein markers for diagnosing Alzheimer's disease

[0073] Using Olink technology as a test to assess the risk of developing Alzheimer's disease or to monitor the severity or progression of Alzheimer's disease

[0074] 1. Olink detection and library preparation.

[0075] 1. Obtaining Blood Samples. Obtain blood samples from subjects for purposes of assessing the risk of developing Alzheimer's disease or monitoring the severity or progression of Alzheimer's disease. The same type of sample should be drawn from the control group (normal subjects who do not have Alzheimer's disease and are not at increased risk for Alzheimer's disease) and the test group (e.g., subjects being tested for possible Alzheimer's disease or for those at increased risk for Alzheimer's disease) using standard procedures typically used in hospitals or clinics.

[0076] 2. Prepare the sample source plate: Manually transfer the samples from the prepared sample plate to the sample source plate and add the controls to the sample source plate. Using a multichannel pipette, transfer 10 μL of each sample to the 384-well sample source plate.

[0077] 3. Sample dilution: Use The automated workstation dispenses sample diluent into the sample dilution plate and uses The automated workstation dilutes the sample.

[0078] 4. Use Dispense sample diluent into the sample dilution plate. The samples were diluted in four serial steps: from 1:1 (undiluted) to 1:10, 1:100, 1:1000 and approximately 1:100000.

[0079] 5. Incubation: In this step, four incubation mixtures are manually prepared for each panel, transferred to the reagent source plate, and mixed with the sample to form four incubation plates (1536), which are then incubated overnight.

[0080] 6. Use Transfer 0.2 μL of sample from the sample source plate and sample dilution plate to two incubation plates. The run takes approximately 15 minutes.

[0081] 7. From Remove the sample source plate, sample dilution plate, and reagent source plate 1 from the platform. If running 1-4 panels, process the plates as described in Table 1 below: Run 1 for 1-2 panels, and Run 1 and Run 2 for 3-4 panels.

[0082] Table 1:

[0083]

[0084] 8. Extension and pre-amplification (PCR1): Manually prepare PCR1 Mix and use The cells were distributed to incubation plates, which were renamed "PCR1 plates" and used for PCR reactions.

[0085] Table 2: PCR1 Mix Composition

[0086] Serial number Reagents Volume (μL) 1 Ultrapure water 27000.0 2 PCR1 Enhancer 3510.0 3 PCR1 solution 3510.0 4 PCR1 enzyme 351.0 Total 34371.0

[0087] 9. Prepare PCR1 plate and perform PCR1: In this step, use The prepared PCR1 mixture was distributed to four incubation plates, and the PCR reaction program for these plates was as follows: 50°C for 20 min, 95°C for 5 min (95°C for 30 s, 54°C for 1 min, 60°C for 1 min) x 25, 10°C hold.

[0088] 10. Mix PCR1 products and use Pool the PCR1 products into a PCR1 pooling plate.

[0089] 11. Amplification and sample indexing (PCR2): Manually prepare PCR2 Mix and then use The samples are mixed with index primers (unique for each sample) and then subjected to a second PCR reaction.

[0090] 12. Prepare PCR2 Mix: Manually prepare PCR2 Mix in a Falcon tube according to the order and volume indicated in Table 3 below:

[0091] Table 3:

[0092]

[0093]

[0094] 13. Prepare PCR2 Plate and perform PCR2: Use Mix the prepared 16 μL PCR2 Mix and 2 μL Index Primers. Then perform a second PCR reaction on the sample.

[0095] 14. Click Open and select the program Olink Index PCR2. Run the following PCR reaction program on these plates: 95°C for 3 minutes; (95°C for 30 seconds, 68°C for 1 minute) x 10; 10°C hold.

[0096] 15. Pool PCR2 product: use The Olink library was then manually transferred to a microcentrifuge tube per panel. Each tube contained amplicons from 96 samples, including controls.

[0097] 16. Library purification: Purify the Olink library using magnetic beads and transfer the eluate to a new microcentrifuge tube for each panel.

[0098] 17. Library quality control: Use a high-sensitivity DNA kit to perform quality control on the purified Olink library on a bioanalyzer.

[0099] 18. Perform second-generation sequencing on the library to detect the expression of the target protein.

[0100] 2. Data Analysis

[0101] 1. Analysis of differentially expressed protein data: The data output by the Olink project is the normalized protein expression value (Normalized Protein eXpression, or NPX value). NPX is a unique unit of the Olink project and uses the Log2 scale (scaling). In this project, the olinkR package was used to convert the bcl format file in the sequencing results into a counts file, which was then opened and analyzed using the NPX Explore Software software, and the NPX matrix was exported for subsequent differential analysis. Therefore, different analysis projects cannot be compared without bridging samples, and different proteins in the same project cannot be compared. The NPX calculation equation is as follows:

[0102] 1)ExtNPX_i,j=log2(counts(sample_jAssay_i) / counts(ExtCtrl_j))

[0103] a. Associate counts with internal references.

[0104] b. All assays and all samples, including negative controls, plate controls, and sample controls.

[0105] c.Log2 transformation obtains more normally distributed data.

[0106] 2)NPX_i,j=ExtNPX_i,j-median(ExtNPX(Plate Controls_i))

[0107] a. Perform plate standardization.

[0108] b. All experiments and samples per plate.

[0109] 3)NPX_Intnorm_i,j=(NPX_i,j–plate median(NPX_i))

[0110] a. Inter-board standardization (optional in multi-board projects).

[0111] b. For all slabs in the project.

[0112] NPX data can be used to discover expression changes of target proteins in a sample group, thereby using NPX data to establish protein features for subsequent analysis.

[0113] Based on the protein NPX Data, t-test analysis or variance analysis was performed on the differentially expressed proteins to screen for proteins with significant differences, and 995 proteins with significant differential expression were obtained (p<0.05). Figure 1 As shown, the distribution of differentially expressed proteins is displayed by a volcano plot, most of which are lowly expressed in Alzheimer's disease.

[0114] 2. Machine learning algorithm screening marker panel:

[0115] We used a Lasso regression model to analyze the differential protein data and screened for the optimal protein panel as a diagnostic marker for Alzheimer's disease. We identified seven proteins with significant differential expression between the two groups and demonstrated excellent diagnostic efficacy. These seven proteins are listed in Table 4.

[0116] Table 4:

[0117]

[0118]

[0119] The expression of 7 proteins in Alzheimer's disease and normal controls is as follows Figure 2 shown.

[0120] When the seven proteins exist alone, the diagnostic analysis is performed based on their respective up-regulation or down-regulation, which has a certain differentiation effect on AD, with an AUC range of 0.84 to 0.90. The ROC curve is shown in Figure 2. Figure 3 shown.

[0121] 3. Calculate the sample score based on the machine learning algorithm:

[0122] Next, based on machine learning analysis using the lasso regression model, the protein combination panel obtained was used to study different marker combination panels, and the scores of each sample between the two groups were used as the standard for diagnosing Alzheimer's disease. Specifically, all samples were randomly divided into training and test sets with a training set:test set ratio of 3:1, and 100 iterations were performed to screen the combination and algorithm formula with the highest sensitivity and specificity.

[0123] The seven protein markers screened were used to form five panel combinations, each containing 3-7 proteins, achieving optimal AD differentiation using different numbers of protein combinations. The most critical three-protein combination was NEFL + GIP + TAFA5. The score for each sample was calculated using the formula 3.78*NEFL + 4.97*GIP - 7.53*TAFA5. Using a threshold of -1.467 to separate Alzheimer's disease samples from healthy controls, samples above -1.467 were classified as Alzheimer's disease, and samples below -1.467 were classified as healthy controls. This achieved excellent AD differentiation: AUC = 1 in the training set and AUC = 0.980 in the validation set.

[0124] This 3-protein combination can also be combined with other highly specific protein markers in the 7-protein list to achieve even better differentiation. For combinations of 4-7 genes, the following formula can be used to calculate the score for each sample and further distinguish Alzheimer's disease samples from normal samples.

[0125] Table 5: Sample score calculation formulas for different marker combinations

[0126]

[0127] The thresholds of different marker combinations and the AUC values in the training set and validation set are shown in Table 6.

[0128] Table 6:

[0129]

[0130]

[0131] The ROC curves of the 3-7 protein combination in the training set and the validation set are as follows: Figure 4 As shown, it can be seen that it can achieve a good distinction effect on AD: AUC = 1 in the training set, and AUC = 0.982-1 in the validation set.

[0132] It should be noted that the preferred embodiments of the present invention are given in the specification and drawings of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. These embodiments are not intended to be additional limitations on the content of the present invention. The purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive. In addition, the above-mentioned technical features can be combined with each other to form various embodiments not listed above, which are all considered to be within the scope of the description of the present invention. Furthermore, it is obvious to those skilled in the art that improvements or changes can be made based on the above description, and all such improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A biomarker for diagnosing Alzheimer's disease, characterized in that: The biomarkers include one or more of NEFL, GIP, TAFA5, RBP2, FABP6, FGF21 or TNFRSF9.

2. The biomarker according to claim 1, characterized in that The biomarkers include at least NEFL, GIP and TAFA5.

3. The biomarker according to claim 1, characterized in that Diagnosing Alzheimer's disease includes determining whether a person has Alzheimer's disease or is at risk of developing Alzheimer's disease or monitoring the severity or progression of Alzheimer's disease.

4. Use of a reagent for detecting the expression level of a biomarker according to any one of claims 1 to 3 in the preparation of a product for diagnosing Alzheimer's disease.

5. The use according to claim 4, characterized in that The diagnosis is performed by the following steps: detecting the expression level of the biomarker from a test sample obtained from the subject, and comparing it with the expression level of the biomarker detected from a control sample, and determining whether the subject has or does not have Alzheimer's disease, or has or does not have a risk of developing Alzheimer's disease (e.g., an increased risk of developing Alzheimer's disease), or has or does not have worsening or progressive Alzheimer's disease based on the upregulation or downregulation of the expression level of the biomarker in the test sample; Further, if the expression level of NEFL or GIP is upregulated, or the expression level of TAFA5, RBP2, FABP6, FGF21 or TNFRSF9 is downregulated in the test sample relative to the control sample, the subject is diagnosed as having Alzheimer's disease, or being at risk of having Alzheimer's disease, or having or not having worsening or progressive Alzheimer's disease; Furthermore, the control sample includes a sample from a healthy or normal subject.

6. The use according to claim 4, characterized in that The diagnosis is performed by the following steps: (1) detecting the expression level of the biomarker from a test sample obtained from a subject; (2) determining the score of the test sample based on the expression level of the biomarker; (3) comparing the score of the test sample with the threshold; as well as (4) determining the subject as having or not having Alzheimer's disease, or having or not having a risk of having Alzheimer's disease, or having or not having worsening or progressive Alzheimer's disease based on whether the score of the test sample is higher or lower than a threshold value; Furthermore, the score of the test sample determined in step (2) is obtained by calculating the sum of the expression levels of each biomarker multiplied by the corresponding coefficient; Further, the coefficients are determined by a machine learning model based on the selected biomarkers and their expression levels in the training set; Furthermore, the corresponding coefficients of each biomarker are shown in Table 5.

7. A product for diagnosing Alzheimer's disease, characterized in that: The product comprises a reagent for detecting the expression level of the biomarker according to any one of claims 1 to 3.

8. A system for diagnosing Alzheimer's disease, characterized in that: The system comprises: (i) a data collection module: in which the expression level of the biomarker according to any one of claims 1 to 3 is detected from a test sample obtained from a subject; (ii) a data analysis module: in which a score of the test sample is determined based on the expression level of the biomarker; and (iii) Output module: Based on whether the score of the test sample is higher or lower than a threshold, determining whether the subject has or does not have Alzheimer's disease, or has or does not have the risk of developing Alzheimer's disease, or has or does not have worsening or progressive Alzheimer's disease.

9. The system according to claim 8, characterized in that In the data analysis module, the score of the test sample is determined by calculating the sum of the expression levels of each biomarker multiplied by a corresponding coefficient; Further, the coefficients are determined by a machine learning model based on the selected biomarkers and their expression levels in the training set; Furthermore, the corresponding coefficients of each biomarker are shown in Table 5 ; Furthermore, the calculation formula for the score of the test sample is shown in Table 5.

10. The system according to claim 8, wherein: The system also includes a training module: detecting the expression levels of the biomarkers from a training set including samples from subjects known to have Alzheimer's disease and samples from healthy or normal subjects, analyzing the expression level data through a machine learning model, and determining the corresponding coefficient of each biomarker and the diagnostic threshold of the biomarker.