Application of mRNA marker combination in preparation of detection product for screening or diagnosing Alzheimer disease
By selecting the combination of mRNA markers composed of specific genes and performing testing, and combining formulas to calculate the risk of Alzheimer's disease, the problem of limited accuracy in detection of blood markers for Alzheimer's disease in the prior art is solved, achieving higher screening and diagnostic accuracy and cost-effectiveness.
Patent Information
- Application Number
- CN202411937808.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has limited accuracy in the detection of blood marker for Alzheimer's disease, making it difficult to achieve the effectiveness of early screening and diagnosis.
The mRNA expression is detected by transcriptome or qPCR using a combination of mRNA markers selected from at least two or more combinations of NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK genes, and the risk of developing Alzheimer's disease is calculated using formula (1) or formula (2).
Improves the accuracy and specificity of screening or diagnosis of Alzheimer's disease, reduces the cost of testing, and provides a cost-effective early screening and diagnosis method.
Smart Images

Figure CN119932171A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine technology, and specifically relates to the application of mRNA marker combinations in the preparation of detection products for screening or diagnosing Alzheimer's disease. Background Art
[0002] Alzheimer's disease (AD) is a degenerative disease of the central nervous system that occurs in the elderly and pre-elderly, characterized by progressive cognitive dysfunction and behavioral impairment. Clinically, it manifests as memory impairment, aphasia, apraxia, agnosia, visual-spatial impairment, abstract thinking and calculation impairment, personality and behavioral changes, etc. AD is the most common type of dementia in the elderly, accounting for 50% to 70% of senile dementia. With the increase of age, the prevalence of AD gradually increases. After the age of 85, one in every three to four elderly people suffers from AD. With the continuous deepening of the understanding of AD, it is currently believed that there is an important pre-dementia stage and preclinical stage before the dementia stage of AD. In this stage, there are AD pathophysiological changes, but there are no or only mild clinical symptoms.
[0003] Risk factors for AD include low education, smoking, hypertension, hyperglycemia, and hypercholesterolemia. There are many theories about the pathogenesis of AD, among which the most influential is the β-amyloid (Aβ) waterfall theory, which believes that the imbalance between the generation and clearance of Aβ is the initiating event leading to neuronal degeneration and dementia. Mutations in the amyloid precursor protein gene, presenilin 1 gene, and presenilin 2 gene of familial AD can all lead to excessive generation of Aβ. Another important mechanism is the excessive phosphorylation of tau protein, which affects the stability of tubulin in the neuronal skeleton, leading to the formation of neurofibrillary tangles, and then destroying the normal function of neurons and synapses. In addition, there is the neurovascular hypothesis, which proposes that abnormal cerebrovascular function leads to neuronal cell dysfunction, and the ability to clear Aβ decreases, leading to cognitive impairment.
[0004] AD usually develops insidiously and progresses continuously, mainly manifested by cognitive impairment and psychiatric and behavioral symptoms. Currently, the most widely used AD diagnostic criteria were developed by the National Institute of Neurological Disorders, Language Disorders and Stroke and the Alzheimer's Disease and Related Disorders Association (NINCDS-ADRDA) in 1984, and later revised by the National Institute on Aging and the Alzheimer's Association (NIA-AA) in 2011, developing diagnostic criteria for different stages of AD and recommending diagnostic criteria for AD dementia and mild cognitive impairment for clinical use.
[0005] Auxiliary examinations for AD diagnosis mainly include biomarker examinations and imaging examinations of cerebrospinal fluid and blood. It is difficult and risky to obtain cerebrospinal fluid samples, which is painful for patients. Compared with cerebrospinal fluid, blood is easy to obtain and less invasive, making it an ideal specimen for disease screening and clinical testing. Currently, the biomarker examinations involving blood samples include Aβ42 / 40, plasma p-Tau181, plasma p-Tau217, plasma NfL, plasma GFAP, etc. These biomarkers are of certain significance for auxiliary diagnosis of AD, but their accuracy is limited. Therefore, there is an urgent need to develop an AD blood marker that has both cost advantages and can improve accuracy in clinical practice, so as to effectively realize the early screening and diagnosis of AD and benefit the general public. Summary of the invention
[0006] The main purpose of the present invention is to provide an mRNA marker combination, which can be used to prepare a detection product for screening or diagnosing Alzheimer's disease.
[0007] Another object of the present invention is to provide a reagent for detecting a combination of mRNA markers for use in preparing a detection product for screening or diagnosing Alzheimer's disease.
[0008] To achieve the above object, the present invention adopts the following technical solution:
[0009] In a first aspect of the present invention, an mRNA marker combination is provided, which is selected from at least two or more combinations of the following genes: NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK.
[0010] Preferably, the mRNA marker combination is derived from human peripheral blood.
[0011] Preferably, the mRNA marker combination is a combination of two genes: NPAP1, GRIN2A and HOXB13.
[0012] Preferably, the mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, KDR.
[0013] Preferably, the mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, SLC34A2, TERT, EGFR, ALK.
[0014] The second aspect of the present invention provides the use of the mRNA marker combination in the preparation of a detection product for screening or diagnosing Alzheimer's disease, wherein the mRNA marker combination is derived from human peripheral blood.
[0015] The third aspect of the present invention provides the use of a reagent for detecting the mRNA marker combination in preparing a detection product for screening or diagnosing Alzheimer's disease, wherein the detection sample of the mRNA marker combination is a human peripheral blood sample.
[0016] A fourth aspect of the present invention provides a method for assessing the risk of Alzheimer's disease, comprising:
[0017] (a) Set P AD The risk of developing Alzheimer's disease is calculated as follows:
[0018] P AD =Σ(P i ×B i ) (1)
[0019] Wherein, i is a candidate gene selected from the mRNA marker combination;
[0020] P i The risk of developing Alzheimer's disease when candidate gene i is positive;
[0021] B i is the detection result of candidate gene i expression, 1 is positive and 0 is negative;
[0022] (b) Set Y AD is the assessment value of suffering from Alzheimer's disease, and its formula is as follows:
[0023] Y AD =Σ(f i × i ) (2)
[0024] Among them, f i is the risk coefficient of candidate gene i;
[0025] x i is the detection value of the expression level of candidate gene i;
[0026] (c) extracting a combination of mRNA markers from the subject's peripheral blood, detecting the expression of candidate genes in the combination of mRNA markers by transcriptomics or qPCR, and then obtaining the risk of the subject suffering from Alzheimer's disease by formula (1) or formula (2); wherein:
[0027] P AD The value is between 0 and 100%, the smaller the value, the lower the risk; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level.
[0028] Y ADThe smaller the value, the lower the risk of AD for the individual or the lower the degree of AD in the peripheral blood mRNA expression level; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level.
[0029] Preferably, in step (a), the P values of all candidate genes i are i equal.
[0030] Preferably, in step (a), P i is the reciprocal of the number of candidate genes.
[0031] Preferably, in step (a), P i B is the basic data of healthy population i The ratio of the number of non-zero values to the number of non-zero values of all candidate genes.
[0032] Preferably, in step (b), the f of all candidate genes i i equal.
[0033] Preferably, in step (b), the f of all candidate genes i i are equal and both are 1.
[0034] Preferably, in step (b), f i It is the ratio of the mean expression of candidate gene i in the basic data of the AD group to the total mean expression of all candidate genes (i.e., the sum of the mean expression of all candidate genes).
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention provides a group of peripheral blood mRNA marker combinations, which are selected from at least two or more of NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK genes. The mRNA expression is detected by transcriptome or qPCR methods, and then the risk of the sample suffering from Alzheimer's disease can be obtained by formula (1) or formula (2), which can be used for screening or diagnosis of Alzheimer's disease.
[0037] 2. The detection cost of mRNA biomarkers in the present invention is lower than that of protein markers, especially the combination of mRNA biomarkers composed of multiple genes, which greatly improves the accuracy of Alzheimer's disease screening or diagnosis, with good specificity, high accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is the model ROC curve of the risk assessment value of AD for each individual in Example 1. DETAILED DESCRIPTION
[0039] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0040] Unless otherwise specified in the following examples, all reagents and materials used were commercially available.
[0041] The following examples propose an mRNA marker combination derived from human peripheral blood, which can be selected from at least two or a combination of more than two of the following genes: NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, ALK.
[0042] In some embodiments, the mRNA marker combination is a combination of three genes: NPAP1, GRIN2A and HOXB13; in some embodiments, the mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, KDR; in some embodiments, the mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, SLC34A2, TERT, EGFR, ALK.
[0043] The following examples propose the use of the above mRNA marker combination in the preparation of a detection product for screening or diagnosing Alzheimer's disease.
[0044] The following examples propose the use of the above-mentioned reagents for detecting mRNA marker combinations in the preparation of detection products for screening or diagnosing Alzheimer's disease.
[0045] The following embodiment proposes a method for assessing the risk of Alzheimer's disease, comprising:
[0046] (a) Set P AD The risk of developing Alzheimer's disease is calculated as follows:
[0047] P AD =Σ(P i ×B i ) (1)
[0048] i is a candidate gene, selected from a combination of mRNA markers;
[0049] P i The risk of developing Alzheimer's disease when candidate gene i is positive;
[0050] B i is the test result of candidate gene i, 1 is positive and 0 is negative;
[0051] (b) Set Y AD is the assessment value of suffering from Alzheimer's disease, and its formula is as follows:
[0052] Y AD =Σ(f i × i ) (2)
[0053] f i is the risk coefficient of candidate gene i;
[0054] x i is the detection value of the expression level of candidate gene i;
[0055] (c) extracting the mRNA marker combination from the subject's peripheral blood, detecting the expression of the candidate gene in the mRNA marker combination by transcriptomics or qPCR, and then calculating the risk of the subject suffering from Alzheimer's disease by formula (1) or formula (2); wherein: P AD The value is between 0 and 100%, the smaller the value, the lower the risk; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level.
[0056] Y AD The smaller the value, the lower the risk of AD for the individual or the lower the degree of AD in the peripheral blood mRNA expression level; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level.
[0057] In some embodiments, the P values of all candidate genes i are i equal; in some embodiments, P i is the reciprocal of the number of candidate genes; in some embodiments, P i B is the basic data of healthy population i The ratio of the number of non-zero values to the number of non-zero values of all candidate genes; in some embodiments, the f of all candidate genes i i equal; in some embodiments, the f of all candidate genes i i =1; in some embodiments, f i It is the ratio of the mean expression of candidate gene i in the basic data of the AD group to the total mean expression of all candidate genes (i.e., the sum of the mean expression of all candidate genes).
[0058] Example 1
[0059] 1. Detection of candidate gene mRNA expression
[0060] Seven healthy volunteers (HC group) and seven AD volunteers (AD group) were randomly selected as the screening group, and 2 mL of peripheral blood was collected and sent to a third-party company for eukaryotic transcriptome testing. The testing unit was Heyuan Biotechnology (Shanghai) Co., Ltd., and the testing platform was Illumina Novaseq 6000 sequencing platform.
[0061] Gene expression data were obtained through data analysis, and genes with data greater than 0 in the AD group were screened and arranged in ascending order based on the mean of the HC group. Then, genes with less data, significant expression differences, and possible associations with the nervous system or AD diseases were detected based on the HC, and candidate genes and their expression data were screened.
[0062] The protein encoded by the NPAP1 gene is nuclear pore-associated protein 1, which is closely related to the composition or function of the nuclear pore complex. The nuclear pore complex is an important channel for the exchange of substances and information between the cell nucleus and the cytoplasm, and is essential for the normal physiological function of cells. The nervous system is a highly complex and sophisticated system, including the brain, spinal cord, and nerve tissue. The cells in these tissues need to frequently exchange substances and information to maintain the normal function of the nervous system. As the main channel between the cell nucleus and the cytoplasm, the nuclear pore complex is essential for this communication need of nervous system cells. Therefore, it can be speculated that the NPAP1 gene may play a certain role in maintaining the normal physiological function of nervous system cells.
[0063] The GRIN2A gene encodes the NMDA (N-methyl-D-aspartate) receptor subunit NR2A, which is a type of glutamate receptor and belongs to the ionotropic glutamate receptor (iGluR). NMDA receptors play a vital role in the nervous system, especially in learning and memory, the balance of neuronal excitation and inhibition, and synaptic plasticity. The GRIN2A gene is located on human chromosome 16p13.2, and the NR2A subunit it expresses is an important component of the NMDA receptor.
[0064] The HNF1A gene is a transcription factor responsible for making a protein called hepatocyte nuclear factor-1α (HNF-1α). This protein is expressed in a variety of tissues, but is particularly critical to the pancreas, where it plays an important role in pancreatic development and insulin secretion. Recent studies have shown that insulin can cross the blood-brain barrier and that a variety of brain cells contain insulin receptors that are affected by abnormal insulin signaling and insulin resistance. Studies have found that insulin resistance is closely associated with the formation of plaques and tangles that lead to AD.
[0065] The ROS1 protein encoded by the ROS1 gene is a member of the insulin receptor family. It is a proto-oncogene that is highly expressed in a variety of tumor cell lines and can activate signaling pathways related to cell proliferation and differentiation, causing excessive cell proliferation. The ROS1 gene was first discovered in gliomas, and its expression was subsequently detected in a variety of nervous system tumors. These tumors include primary gliomas, meningiomas, etc. Although no direct relationship has been found between the ROS1 gene and AD, its role in cell growth, proliferation, and survival may be indirectly related to certain pathological processes of AD.
[0066] The HOXB13 gene belongs to the homeobox gene family, which is highly conserved in vertebrates and is essential for vertebrate embryonic development. The transcription factor encoded by the HOXB13 gene plays an important role in regulating cell growth, differentiation and development. The HOXB13 gene is involved in the formation and differentiation of the nervous system during embryonic development. It may affect the normal development of the nervous system by regulating the proliferation, migration and differentiation of nerve cells.
[0067] The SLC34A2 gene, also known as solute carrier family 34 member A2, encodes a multi-channel membrane protein. The protein it encodes is involved in the absorption and reabsorption of phosphate in the body, and has an important impact on bone health, cell signaling, energy metabolism, etc. The biological processes it participates in may affect the nervous system.
[0068] The protein encoded by the KDR gene is also named vascular endothelial growth factor receptor-2 (VEGFR-2), which is an important receptor for vascular growth factor signal transduction. It is involved in regulating the proliferation, chemotaxis and increased vascular permeability of vascular endothelial cells. The KDR gene plays a key role in multiple physiological processes such as the circulatory system, cardiovascular system, and central nervous system. Gene knockout experiments have shown that the KDR gene is necessary for angiogenesis in the central nervous system.
[0069] The TERT gene is one of the important genes encoding the telomerase complex, and the telomerase reverse transcriptase it encodes is the core catalytic subunit of telomerase. The main function of telomerase is to extend telomeres, which are the terminal structures of eukaryotic chromosomes and are essential for maintaining chromosome stability and genome integrity. Studies have shown that TERT can not only extend telomeres, but also act as a transcription factor to affect the expression of many genes that are closely related to neurogenesis, learning and memory, cell aging and inflammation. Restoring TERT levels can reduce cell aging and tissue inflammation, and this effect is particularly important in the nervous system, because aging and inflammation of the nervous system may lead to cognitive decline and the occurrence of neurodegenerative diseases.
[0070] The protein encoded by EGFR is a tyrosine kinase receptor that mediates cell proliferation and signal transduction of epidermal growth factor (EGF). EGFR receives growth factor signals, transmits them into cells, and promotes cell growth, proliferation, differentiation and migration through a series of signal transduction processes. EGFR plays an important role in the development of the nervous system and has an important influence on the survival and differentiation of neurons. In addition, overactivation of EGFR may be related to pathological processes such as neuronal apoptosis and inflammatory response.
[0071] The ALK gene encodes a receptor tyrosine kinase (RTK), which is a member of the receptor tyrosine kinase family. RTK is responsible for cell signal transduction, transmitting signals received on the cell surface to the cell. The ALK gene is usually expressed during the development of the nervous system (including the central and peripheral nervous systems) and is related to embryogenesis and cell proliferation and migration.
[0072] Based on the above formulas (1) and (2), x i The FPKM value of each gene in the transcriptome sequencing results, FPKM stands for fragments per thousand bases of transcription per million mapped reads, is used to measure gene expression levels. The FPKM value takes into account gene length and sequencing depth, and standardizes gene expression to make the expression levels of different genes and under different sequencing conditions comparable. The results are shown in Tables 1 and 2 (FPKM, accuracy is taken to four decimal places).
[0073] Table 1: Expression results of candidate genes in the screening group of HC volunteers in Example 1
[0074] symbol HC1 HC2 HC3 HC4 HC5 HC6 HC7 NPAP1 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 GRIN2A 0.0000 0.0000 0.0000 0.0040 0.0026 0.0000 0.0000 HNF1A 0.0051 0.0000 0.0000 0.0000 0.0000 0.0055 0.0066 ROS1 0.0029 0.0000 0.0000 0.0000 0.0055 0.0094 0.0037 HOXB13 0.0000 0.0000 0.0000 0.0000 0.0242 0.0000 0.0110 SLC34A2 0.0000 0.0172 0.0000 0.0075 0.0049 0.0110 0.0066 KDR 0.0000 0.0000 0.0000 0.0534 0.0077 0.0000 0.0262 TERT 0.0306 0.0439 0.0075 0.0000 0.0083 0.0187 0.0000 EGFR 0.0032 0.0054 0.0000 0.0259 0.0415 0.0435 0.0104 ALK 0.0456 0.0176 0.0362 0.0000 0.0066 0.0226 0.0045
[0075] Table 2: Expression results of candidate genes in the screening group of AD volunteers in Example 1
[0076] symbol AD1 AD2 AD3 AD4 AD5 AD6 AD7 NPAP1 0.0164 0.0113 0.0119 0.0078 0.0141 0.0303 0.0067 GRIN2A 0.0083 0.0092 0.0244 0.0127 0.0134 0.0336 0.0028 HNF1A 0.0124 0.0457 0.0845 0.0294 0.0640 0.0722 0.0102 ROS1 0.0350 0.0258 0.0307 0.0366 0.1449 0.1745 0.0405 HOXB13 0.0205 0.0378 0.0299 0.0292 0.0589 0.0977 0.0169 SLC34A2 0.0803 0.0569 0.1383 0.0469 0.0638 0.0262 0.0204 KDR 0.0098 0.0317 0.0526 0.0466 0.0621 0.0416 0.0324 TERT 0.0894 0.1165 0.1537 0.0849 0.0786 0.0613 0.0347 EGFR 0.0528 0.0739 0.0761 0.0519 0.1661 0.1842 0.0387 ALK 0.1308 0.0739 0.1479 0.1281 0.4604 0.4872 0.1532
[0077] 2. One-way ANOVA of candidate gene mRNA expression data
[0078] The gene expression data of 10 candidate genes were screened out, and the variance analysis function of Excel was used to perform one-way variance analysis to obtain the P value. P<0.05, the difference was significant, indicated by "*"; P<0.01, the difference was extremely significant, indicated by "**"; P≥0.05, the difference was not significant, indicated by "NS", as shown in Table 3.
[0079] Table 3: One-way ANOVA analysis of candidate genes
[0080]
[0081] As shown in Table 3, the expression data of candidate genes ROS1 and KDR were significantly different between the HC group and the AD group (P < 0.05), and the expression data of other candidate genes were extremely significantly different between the HC group and the AD group (P < 0.01).
[0082] 3. Calculate each volunteer’s risk of developing AD
[0083] (I) The peripheral blood mRNA marker combination of NPAP1, GRIN2A, HOXB13, HNF1A, and KDR was selected as the calculation index, and the risk of AD for each individual in the validation group was calculated using formula (1), which is as follows:
[0084] P AD =Σ(P i ×B i ) (1)
[0085] Case 1: Assuming that each candidate gene contributes equally to the risk of AD, the P value of each gene is i Equal, a total of 5 candidate genes, then P i =1 / 5=20%, convert the values of T-HC group data greater than 0 to 1, and calculate the P of each individual AD , the results are shown in Table 4.
[0086] Table 4: Formula (1) Case 1 P of each individual suffering from AD AD (PAD1)
[0087]
[0088]
[0089] Case 2: Assuming that each candidate gene contributes differently to the risk of AD, according to the characteristics of formula (1), the risk of AD is estimated using the undetected rate of each candidate gene, specifically:
[0090] P i = Number of individuals without detection of candidate gene i / Sum of the number of individuals without detection of all candidate genes, where:
[0091] The number of individuals with undetected candidate gene i is the number of individuals with a value of 0 in candidate gene i;
[0092] The sum of the number of individuals with undetected candidate genes is the number of all values 0;
[0093] The risk contribution of each candidate gene to AD was calculated based on the HC group data, and the results are shown in Table 5.
[0094] Table 5: P of HC group formula (1) case 2 i value
[0095] symbol Undetected quantity <![CDATA[P i ]]> NPAP1 7 28.00% GRIN2A 5 20.00% HNF1A 4 16.00% HOXB13 5 20.00% KDR 4 16.00% Total undetected quantity 25 100.00%
[0096] According to formula (1), using the data in Table 5, calculate the P of each individual AD , the results are shown in Table 6.
[0097] Table 6: PAD (PAD2) of each individual with AD in case 2 of formula (1)
[0098] symbol HC1 HC2 HC3 HC4 HC5 HC6 HC7 AD1 AD2 AD3 AD4 AD5 AD6 AD7 NPAP1 0 0 0 0 0 0 0 1 1 1 1 1 1 1 GRIN2A 0 0 0 1 1 0 0 1 1 1 1 1 1 1 HNF1A 1 0 0 0 0 1 1 1 1 1 1 1 1 1 HOXB13 0 0 0 0 1 0 1 1 1 1 1 1 1 1 KDR 0 0 0 1 1 0 1 1 1 1 1 1 1 1 PAD2 0.16 0 0 0.36 0.56 0.16 0.52 1 1 1 1 1 1 1
[0099] (II) The peripheral blood mRNA marker combination of NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK was selected as the calculation index, and the evaluation value of AD for each individual in the validation group was calculated using formula (2). The formula is as follows:
[0100] Y AD =Σ(f i × i ) (2)
[0101] Case 1: Assuming that each candidate gene contributes equally to the risk of AD, the fi of each gene is equal, and then Pi = 1, and the Y of each individual in Example 1 is calculated. AD (YAD1), the results are shown in Table 7.
[0102] Table 7: Formula (2) Case 1 Y of each volunteer suffering from AD AD (YAD1)
[0103] symbol YAD1 symbol YAD1 HC1 0.0875 AD1 0.4557 HC2 0.0842 AD2 0.4828 HC3 0.0437 AD3 0.7501 HC4 0.0908 AD4 0.4740 HC5 0.1013 AD5 1.1262 HC6 0.1107 AD6 1.2088 HC7 0.0691 AD7 0.3565
[0104] Case 2: Assuming that each candidate gene contributes differently to the risk of AD, the f of each gene i unequal, use statistical methods to estimate f i The value of the candidate gene i The f of each candidate gene was calculated by taking the ratio of the mean expression of the candidate gene in the AD group data to the total mean expression of all candidate genes in the peripheral blood mRNA marker combination. i , the results are shown in Table 8.
[0105] Table 8: f of each candidate gene in case 2 of formula (2) i
[0106] symbol AD-MEAN <![CDATA[f i ]]> NPAP1 0.0141 0.0203 GRIN2A 0.0149 0.0215 HNF1A 0.0455 0.0656 ROS1 0.0697 0.1005 HOXB13 0.0416 0.0599 SLC34A2 0.0618 0.0891 KDR 0.0396 0.0570 TERT 0.0884 0.1275 EGFR 0.0920 0.1326 ALK 0.2259 0.3258 SUM 0.6935 1.0000
[0107] According to formula (2), calculate the Y of each individual AD (YAD2), the results are shown in Table 9.
[0108] Table 9: Y of each individual in case 2 of formula (2) AD (YAD2)
[0109] symbol YAD2 symbol YAD2 HC1 0.0198 AD1 0.0748 HC2 0.0136 AD2 0.0639 HC3 0.0127 AD3 0.1044 HC4 0.0072 AD4 0.0741 HC5 0.0117 AD5 0.2141 HC6 0.0178 AD6 0.2252 HC7 0.0064 AD7 0.0691
[0110] 4. ROC Curve
[0111] By comparing the ROC curves and AUC values of different classification models, the advantages and disadvantages of the models can be intuitively evaluated. The larger the AUC value, the better the classification performance of the model. The ROC curve can also help understand the balance between the sensitivity and specificity of the classification model, thereby optimizing the performance of the model. The model for calculating the estimated value of the risk of AD for each individual using formula (1) case 1, formula (1) case 2, formula (2) case 1 and formula (2) case 2 in Example 1 was subjected to ROC curve analysis. The results are as follows: Figure 1 and as shown in Table 10.
[0112] It can be seen that the four ROC curves overlap, AUC=1, and the model classification performance is excellent. The best critical value of the model in case 1 of formula (1) in Example 1 is 0.8, the best critical value of the model in case 2 of formula (1) is 0.78, the best critical value of the model in case 1 of formula (2) is 0.2336, and the best critical value of the model in case 2 of formula (2) is 0.0419. At this time, the sensitivity and specificity of the corresponding model are both 1.
[0113] Table 10: Figure 1 The curvilinear coordinates
[0114]
[0115]
[0116] Example 2
[0117] Eight healthy volunteers (T-HC group) and three AD volunteers (T-AD group) were randomly selected as the validation group. 2 mL of peripheral blood was drawn and sent to a third-party company for eukaryotic transcriptome detection. The detection unit was Heyuan Biotechnology (Shanghai) Co., Ltd. and the detection platform was Illumina Novaseq 6000 sequencing platform. The gene expression data obtained are shown in Table 11 (FPKM, accuracy to four decimal places).
[0118] Table 11: Expression value results of candidate genes of volunteers in the validation group
[0119]
[0120]
[0121] The peripheral blood mRNA marker combination of NPAP1, GRIN2A, HOXB13, HNF1A, and KDR was selected as the calculation index, and the risk of AD for each individual in the validation group was calculated using formula (1). The formula is as follows:
[0122] P AD =Σ(P i ×B i ) (1)
[0123] Case 1: Assuming that each candidate gene contributes equally to the risk of AD, the P value of each gene is i Equal, a total of 5 candidate genes, then P i =1 / 5=20%, convert the values of the T-HC group data greater than 0 to 1, and calculate the P of each individual in the validation group AD , the results are shown in Table 12.
[0124] Table 12: Formula (1) Case 1 Calculation of P of each volunteer in the validation group suffering from AD AD
[0125]
[0126]
[0127] According to Example 1, the best critical value obtained by the ROC curve of the model for diagnosis is 0.8, less than 0.8 is the healthy group, greater than 0.8 is the AD group, the model can completely distinguish the HC group and the AD group, and the model performance is excellent. When used to assess the risk of healthy people suffering from AD, P AD The larger the value, the higher the risk.
[0128] Case 2: Assuming that each candidate gene contributes differently to the risk of AD, according to the characteristics of formula (1), the risk of AD is estimated using the undetected rate of each candidate gene, specifically:
[0129] P i = Number of individuals without detection of candidate gene i / Sum of the number of individuals without detection of all candidate genes, where:
[0130] The number of individuals with undetected candidate gene i is the number of individuals with a value of 0 in candidate gene i;
[0131] The sum of the number of individuals with undetected candidate genes is the number of all values 0;
[0132] The risk contribution of each candidate gene to AD was calculated based on the data in the screening group in Example 1 (see Table 5), and the risk P of AD for each individual in the validation group was calculated. AD , the results are shown in Table 13.
[0133] Table 13: Formula (1) Case 2 Calculation of P of each individual suffering from AD in the validation group AD
[0134]
[0135]
[0136] According to Example 1, the best critical value obtained by the ROC curve of the model for diagnosis is 0.78, less than 0.78 is the healthy group, greater than 0.78 is the AD group, the model can completely distinguish the HC group from the AD group, and the model performance is excellent. When used to assess the risk of AD in healthy people, P AD The larger the value, the higher the risk.
[0137] Example 3
[0138] The peripheral blood mRNA marker combination of NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK was selected as the calculation index, and the evaluation value of AD for each individual in the verification group in Example 2 was calculated using formula (2), and the formula is as follows:
[0139] Y AD =Σ(f i × i ) (2)
[0140] Case 1: Assuming that each candidate gene contributes equally to the risk of AD, the f of each gene i If they are equal, then let P i =1, calculate the Y of each individual in the validation group in Example 2 AD , the results are shown in Table 14.
[0141] Table 14: Formula (2) Case 1 Calculation of Y for each individual in the validation group suffering from AD AD
[0142] symbol T-HC1 T-HC2 T-HC3 T-HC4 T-HC5 T-HC6 T-HC7 T-HC8 T-AD1 T-AD2 T-AD3 <![CDATA[Y AD ]]> 0.1412 0.0836 0.0595 0.1028 0.1366 0.0751 0.1079 0.1409 0.4171 0.9455 0.3628
[0143] According to Example 1, the best critical value obtained by the ROC curve of the model for diagnosis is 0.2336. The group with a value less than 0.2336 is the healthy group, and the group with a value greater than 0.2336 is the AD group. The model can completely distinguish the HC group from the AD group, and the model performance is excellent. When used to assess the risk of AD in healthy people, Y AD The larger the value, the higher the risk; it is used to assess the severity of AD patients. AD The larger the value is, the greater the severity of the disease at the peripheral blood mRNA level.
[0144] Case 2: Assuming that each candidate gene contributes differently to the risk of AD, the f of each gene i unequal, use statistical methods to estimate f i The value of the candidate gene i It is: the ratio of the mean expression of the candidate gene in the AD group data to the total mean expression of all candidate genes in the peripheral blood mRNA marker combination.
[0145] According to Example 1, formula (2) Case 2: f of each candidate gene i The Y value (Table 8) of each individual in the validation group is calculated AD , the results are shown in Table 15.
[0146] Table 15: Formula (2) Case 2 Calculation of Y for each individual in the validation group suffering from AD AD
[0147] symbol T-HC1 T-HC2 T-HC3 T-HC4 T-HC5 T-HC6 T-HC7 T-HC8 T-AD1 T-AD2 T-AD3 <![CDATA[Y AD ]]> 0.0223 0.0112 0.0051 0.0169 0.0233 0.0102 0.0224 0.0276 0.0498 0.1480 0.0489
[0148] According to Example 1, the best critical value obtained by the ROC curve of the model for diagnosis is 0.0419, less than 0.0419 is the healthy group, greater than 0.0419 is the AD group, the model can completely distinguish the HC group and the AD group, and the model performance is excellent. When used to assess the risk of healthy people suffering from AD, Y AD The larger the value, the higher the risk; it is used to assess the severity of AD patients. AD The larger the value is, the greater the severity of the disease at the peripheral blood mRNA level.
[0149] In summary, the present invention proposes a peripheral blood mRNA marker combination selected from NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, and ALK. The mRNA expression of the candidate gene is first detected by a transcriptome or qPCR method, and then the risk of the individual suffering from Alzheimer's disease can be obtained by formula (1) or formula (2). The grading of the peripheral blood mRNA marker level of individuals suffering from Alzheimer's disease can also be evaluated. Combined with clinical practice, Alzheimer's disease can be better evaluated and used for screening, diagnosis and rating of Alzheimer's disease, with excellent performance.
[0150] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An mRNA marker combination, characterized in that: At least two or more combinations of the following genes are selected from: NPAP1, GRIN2A, HNF1A, ROS1, HOXB13, SLC34A2, KDR, TERT, EGFR, ALK; the mRNA marker combination is derived from human peripheral blood.
2. The mRNA marker combination according to claim 1, characterized in that The mRNA marker combination is a combination of three genes: NPAP1, GRIN2A and HOXB13.
3. The mRNA marker combination according to claim 1, characterized in that The mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, KDR.
4. The mRNA marker combination according to claim 1, characterized in that The mRNA marker combination is the following gene combination: NPAP1, GRIN2A, HNF1A, HOXB13, SLC34A2, TERT, EGFR, ALK.
5. Use of the mRNA marker combination according to any one of claims 1 to 4 in the preparation of a detection product for screening or diagnosing Alzheimer's disease, wherein the mRNA marker combination is derived from human peripheral blood.
6. Use of a reagent for detecting the mRNA marker combination according to any one of claims 1 to 4 in preparing a detection product for screening or diagnosing Alzheimer's disease, wherein the detection sample of the mRNA marker combination is a human peripheral blood sample.
7. A method for assessing the risk of Alzheimer's disease, characterized in that: The following steps are involved: (a) Set P AD The risk of developing Alzheimer's disease is calculated as follows: P AD =Σ(P i ×B i ) (1) Wherein, i is a candidate gene selected from the mRNA marker combination according to any one of claims 1 to 6; P i The risk of developing Alzheimer's disease when candidate gene i is positive; B i is the detection result of candidate gene i expression, 1 is positive and 0 is negative; (b) Set Y AD is the assessment value of suffering from Alzheimer's disease, and its formula is as follows: Y AD =Σ(f i ×x i ) (2) Among them, f i is the risk coefficient of candidate gene i; x i is the detection value of the expression level of candidate gene i; (c) extracting a combination of mRNA markers from the subject's peripheral blood, detecting the expression of candidate genes in the mRNA marker combination by transcriptomics or qPCR, and then obtaining the risk of the subject suffering from Alzheimer's disease by formula (1) or formula (2); wherein: P AD The value is between 0 and 100%, the smaller the value, the lower the risk; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level. Y AD The smaller the value, the lower the risk of AD for the individual or the lower the degree of AD in the peripheral blood mRNA expression level; below the critical value, it indicates that the individual has not yet suffered from AD; above the critical value, it indicates that the individual is very likely to suffer from AD or has already suffered from AD based on the peripheral blood mRNA expression level.
8. The method for assessing the risk of Alzheimer's disease according to claim 7, characterized in that: In step (a), the P of all candidate genes i i equal; and / or in step (a), P i is the reciprocal of the number of candidate genes; and / or in step (a), P i B is the basic data of healthy population i The ratio of the number of non-zero values to the number of non-zero values of all candidate genes.
9. The method for assessing the risk of Alzheimer's disease according to claim 7, characterized in that: In step (b), the f of all candidate genes i i equal; Or in step (b), the f of all candidate genes i i are equal and both are 1; and / or in step (b), f i It is the ratio of the mean expression of candidate gene i in the basic data of the AD group to the total mean expression of all candidate genes.