Alzheimer's disease risk prediction model and construction method and application thereof
By discovering six key Alzheimer's factors in plasma and using the GLM-Smooth algorithm, a highly accurate Alzheimer's early prediction model was constructed, solving the shortcomings of existing detection methods in terms of accuracy and interpretability, and improving the early prediction and prevention capabilities of the disease.
Patent Information
- Application Number
- CN202411936731.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
The existing Alzheimer's disease detection methods have insufficient accuracy and interpretability, which affects its commercial and clinical application value.
By discovering a combination of six key Alzheimer's factors in human plasma, an early prediction model that can be used to predict the probability of Alzheimer's disease was constructed using the novel GLM-Smooth machine learning algorithm.
It improves the early prediction ability of Alzheimer's disease, provides technical support for disease prevention and early treatment, and the model's AUC can reach 0.93, with excellent prediction effect and accuracy.
Smart Images

Figure SMS_1 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disease risk prediction model construction, and in particular to an Alzheimer's disease risk prediction model and a construction method and application thereof. Background Art
[0002] Alzheimer's Disease (AD) is a dementia and one of the most pressing medical and social challenges facing societies with aging populations.
[0003] The detection based on cell-free RNA (cfRNA) is a non-invasive detection method that can be used as a sample with simple body fluid sampling. It provides patients with a better experience, increases compliance, and reduces the complexity of sampling. cfRNA can provide dynamic gene expression information that reflects real-time physiological and pathological changes. Related studies have shown that the use of cfRNA for early disease detection and research on cfRNA omics have gradually become a new hot research direction in academia. The use of cfRNA omics for the detection of Alzheimer's disease began as early as 2020. With the deepening of research, more and more studies have found that cfRNA can be used as a biomarker for Alzheimer's disease. This non-invasive way to detect molecular changes in plasma has also strengthened people's understanding of the cause and progression of Alzheimer's disease, helped early screening and prevention of the disease, and provided support for the determination of effective treatment strategies, thereby delaying the progression of the disease. But overall, there are still relatively few relevant markers available, which limits its versatility and accuracy in the detection of Alzheimer's disease.
[0004] The trend of using artificial intelligence in the field of life sciences is increasing year by year, especially machine learning methods are increasingly being used in cfRNA-based disease prediction research, and have achieved remarkable results. However, in terms of Alzheimer's disease detection, Alzheimer's disease detection based on machine learning methods is still developing slowly, and there is still a lot of room for improvement in accuracy and interpretability, which affects its commercial and clinical use value and application value. Summary of the invention
[0005] The present invention aims to solve at least one of the above-mentioned technical problems existing in the prior art. To this end, the purpose of the present invention is to provide an Alzheimer's disease risk prediction model and its construction method and application. In the present invention, a combination of six key factors of Alzheimer's disease in human plasma was discovered for the first time. Based on their weights, a new GLM-Smooth machine learning algorithm was used to construct an early prediction model that can be used to predict the probability of onset of Alzheimer's disease, thereby improving the early prediction ability of Alzheimer's disease and providing strong technical support for the prevention and early treatment of Alzheimer's disease.
[0006] In a first aspect of the present invention, a group of diagnostic or predictive markers is provided, wherein the diagnostic or predictive markers include at least one of FZD6, GRIN2A, TUBA8, RYR3, NDUFS3, and TNF.
[0007] In some embodiments of the present invention, the diagnostic or predictive markers include at least 2, 3, 4, 5 or 6 of FZD6, GRIN2A, TUBA8, RYR3, NDUFS3 and TNF.
[0008] In some embodiments of the invention, the diagnostic or predictive markers include FZD6, GRIN2A, TUBA8, RYR3, NDUFS3 and TNF.
[0009] In some embodiments of the invention, the diagnostic or predictive marker is a combination of FZD6, GRIN2A, TUBA8, RYR3, NDUFS3 and TNF.
[0010] In some embodiments of the present invention, the diagnostic or predictive marker is an Alzheimer's disease diagnostic or predictive marker.
[0011] In some embodiments of the invention, the diagnosis comprises early diagnosis.
[0012] In the present invention, the term "early diagnosis" refers to diagnosis based on conventional technical means or clinical experience when no symptoms are found or appear, or when the initial symptoms have just appeared. In the present invention, "early diagnosis" refers to diagnosis during a period when conventional technical means are not yet available for the detection of Alzheimer's disease, or when the detection results of conventional technical means are not yet accurate.
[0013] In the present invention, the phrase "the detection results of conventional technical means are not yet accurate" means that the diagnostic accuracy using conventional technical means is equal to or lower than 60%, 50%, 40%, 30%, 20% or 10%; or there is a false positive rate greater than or equal to 30%, 40%, 50%, 60%, 70%, 80%, 90% or 99%.
[0014] In some embodiments of the present invention, the Alzheimer's disease diagnostic or predictive markers may further include: lifestyle, eating habits, environmental factors, laboratory test index information and basic clinical information of the subject, etc.
[0015] The second aspect of the present invention provides the use of a reagent for detecting the diagnostic or predictive markers described in the above aspects in the preparation of an Alzheimer's disease diagnosis and / or risk assessment prediction product.
[0016] In some embodiments of the present invention, the reagents for detecting the diagnostic or predictive markers described in the above aspects can be obtained based on conventional techniques in the art, including but not limited to: commercially available detection kits, synthetic primers, etc.
[0017] In the present invention, there is no limitation on the detection method of early diagnostic markers. According to actual needs, it can be at the gene level, protein level or biochemical level. The detection method includes but is not limited to: sequencing, chemiluminescence, enzyme-linked immunosorbent assay, radioimmunoassay, microparticle labeling and polymerase chain reaction (PCR), etc.
[0018] In some embodiments of the invention, the diagnosis comprises early diagnosis.
[0019] In some embodiments of the present invention, the Alzheimer's disease diagnosis and / or risk assessment prediction product includes a detection reagent, a detection kit and a detection chip.
[0020] The third aspect of the present invention provides an Alzheimer's disease diagnosis and / or risk assessment prediction product, which includes reagents for detecting the diagnostic or predictive markers described in the above aspects, and adjuvants.
[0021] In some embodiments of the present invention, the adjuvant comprises a pharmaceutically acceptable adjuvant.
[0022] In some embodiments of the present invention, the auxiliary agent includes at least one of a buffer, a solvent or a diluent, a preservative, a surfactant, a cleaning solution, and a sample pretreatment reagent.
[0023] In the present invention, the term "buffer" refers to a reagent used to adjust the pH value of the experimental environment or simulate a special environment (such as in vivo) to ensure the stability and accuracy of the experiment, including but not limited to phosphate buffered saline (PBS), Tris-HCl buffer, etc.
[0024] In the present invention, the term "solvent" refers to a liquid used to dissolve the corresponding test object or the sample to be tested, including but not limited to: organic solvents, such as ethanol, methanol, acetonitrile, ethyl acetate, dichloromethane, acetone, petroleum ether, benzene, toluene, ether, chloroform and isopropanol; inorganic solvents, such as water.
[0025] In the present invention, the term "diluent" refers to a reagent used to dilute a sample to be tested or an original reagent to reduce its concentration so that it can participate in a specific detection reaction, including but not limited to the above-mentioned buffer or solvent.
[0026] In the present invention, the term "preservative" is used to prevent the reagent from deteriorating, extend the shelf life of the reagent, and ensure the stability of the reagent during storage and transportation, including but not limited to: benzoic acid, sodium benzoate, parahydroxybenzoic acid ester compounds (such as its methyl, ethyl, propyl, and butyl esters), sorbic acid and its salts, o-phenylphenol, etc.
[0027] In the present invention, the term "surfactant" also includes dispersants, which refers to reagents used to increase the surface activity of a solution and help dissolve and disperse the components in a sample, including but not limited to: anionic surfactants (such as sodium dodecyl sulfate (SDS), sodium hexadecyl sulfate), cationic surfactants (such as benzalkonium chloride, benzalkonium bromide), nonionic surfactants (such as Tween 20, Tween 40, Tween 80) and amphoteric surfactants (such as lecithin).
[0028] In the present invention, the term "cleaning solution" refers to a reagent used to clean the reaction system, instrument pipelines and components in the detection reaction, including but not limited to: water, buffer solution.
[0029] In the present invention, the term "sample pretreatment reagents" refers to reagents used in sample separation, extraction, purification, enzymatic cutting and other processes before formal testing, including but not limited to: extracting solution (extractant), purification solution (purification material), dehydration reagent and digestive enzyme solution, etc.
[0030] In some embodiments of the present invention, in the Alzheimer's disease diagnosis and / or risk assessment prediction product, the reagents and adjuvants for detecting the markers described in the above aspects can be packaged independently or co-packaged.
[0031] In some embodiments of the present invention, the Alzheimer's disease diagnosis and / or risk assessment prediction product includes a detection kit or a detection set.
[0032] A fourth aspect of the present invention provides a method for constructing an Alzheimer's disease risk prediction model, comprising the following steps:
[0033] Collecting the expression of the diagnostic or predictive markers and the prevalence of Alzheimer's disease described in the above aspects, and constructing a feature matrix;
[0034] The feature matrix is fitted using a modeling algorithm to obtain the Alzheimer's disease risk prediction model.
[0035] In some embodiments of the present invention, the construction method further comprises: after the fitting is completed, optimizing the model parameters using a regularization algorithm.
[0036] In some embodiments of the present invention, the optimization includes: optimizing the coefficient β to obtain a higher prediction accuracy.
[0037] In some embodiments of the present invention, the modeling algorithm includes: GLM-Smooth, and random forest model.
[0038] In some embodiments of the present invention, the modeling algorithm is GLM-Smooth.
[0039] In some embodiments of the present invention, for a binary variable, once a positive result occurs, it is defined as positive.
[0040] In some embodiments of the present invention, the expression of the diagnostic or predictive marker can be obtained based on conventional methods in the art, including but not limited to: detection techniques at different detection levels (such as nucleic acid or protein levels).
[0041] In some embodiments of the present invention, the construction method specifically comprises:
[0042] Collecting the expression of the diagnostic or predictive markers and the prevalence of Alzheimer's disease described in the above aspects, converting the prevalence of Alzheimer's disease into a binary variable, and constructing a feature matrix;
[0043] The GLM-Smooth algorithm was used to fit the feature matrix, and then the regularization algorithm was used to optimize the coefficient β to obtain the Alzheimer's disease risk prediction model.
[0044] In some embodiments of the present invention, the collected diagnostic or predictive marker expression conditions are subjected to data processing.
[0045] In some embodiments of the present invention, the data processing includes: data quality control, data alignment, gene counting and normalization.
[0046] In some embodiments of the invention, the data quality control uses fastp v0.23.3.
[0047] In some embodiments of the invention, the data alignment uses STAR v2.7.10b.
[0048] In some embodiments of the invention, the gene counting uses Feature Counts v 2.0.6.
[0049] In some embodiments of the present invention, the normalization is performed using the following formula:
[0050]
[0051] in, represents the normalized gene expression; x i is the original expression of the corresponding gene; λ is the amplification constant, which is used to avoid the value being too small; n is the number of genes in the corresponding sample.
[0052] The fifth aspect of the present invention provides the use of the Alzheimer's disease risk prediction model constructed by the construction method described in the above aspects in the preparation of Alzheimer's disease diagnosis and / or risk assessment prediction products.
[0053] In some embodiments of the present invention, the Alzheimer's disease diagnosis and / or risk assessment prediction product includes an analysis module, and the analysis module obtains a prediction result by executing the Alzheimer's disease risk assessment prediction model constructed by the construction method described in the above aspects.
[0054] In some embodiments of the present invention, the Alzheimer's disease risk prediction model comprises:
[0055]
[0056] or
[0057]
[0058] Among them, P represents the predicted value of Alzheimer's disease risk, X 1-6 They represent the expression levels of FZD6, GRIN2A, RYR3, TNF, TUBA8 and NDUFS3 respectively.
[0059] In some embodiments of the present invention, the analysis module outputs a predicted value of the risk of Alzheimer's disease, and determines whether there is a risk of Alzheimer's disease based on the predicted value of the risk of Alzheimer's disease.
[0060] In some embodiments of the present invention, the Alzheimer's disease risk judgment standard based on the Alzheimer's disease risk prediction value is:
[0061] Comparison of the predicted risk of Alzheimer's disease and the risk threshold of Alzheimer's disease.
[0062] If the predicted value of the subject's Alzheimer's disease risk is less than the Alzheimer's disease risk threshold, it is judged that there is no risk of Alzheimer's disease or the risk of Alzheimer's disease is low;
[0063] If the predicted value of the subject's Alzheimer's disease risk is greater than or equal to the Alzheimer's disease risk threshold, it is determined that the subject has an Alzheimer's disease risk or a high Alzheimer's disease risk in some embodiments of the present invention.
[0064] In some embodiments of the present invention, the Alzheimer's disease diagnosis and / or risk assessment prediction product further includes: at least one of a detection module, a storage module and a transmission module.
[0065] In some embodiments of the present invention, the Alzheimer's disease diagnosis and / or risk assessment prediction product further includes: at least one input terminal and / or at least one output terminal.
[0066] In some embodiments of the present invention, the input end includes but is not limited to: keyboard, mouse, hard disk, magnetic disk, writable read-only compact disk (CD-ROM), photoelectric paper tape input device, card input device, optical character reader, magnetic tape input equipment, Chinese character input equipment, etc.
[0067] In some embodiments of the present invention, the output end includes but is not limited to: a mobile client (such as a mobile phone, a tablet computer, a laptop computer), a display, a printer, a light pen, a hard disk, a magnetic disk, a writable read-only compact disk (CD-ROM), etc.
[0068] A sixth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of calculating the Alzheimer's disease risk prediction model constructed by the construction method described in the above aspect are implemented.
[0069] A seventh aspect of the present invention provides a system, comprising an analysis module, wherein the analysis module contains the computer-readable storage medium described in the above aspects.
[0070] In some embodiments of the present invention, the system further comprises: an input end and an output end.
[0071] In some embodiments of the present invention, the input end and the output end are as defined in the above aspects.
[0072] In some embodiments of the present invention, the system further comprises: at least one of a detection module, a storage module and a transmission module.
[0073] In some embodiments of the present invention, the system includes: a detection module, a storage module and a transmission module.
[0074] An eighth aspect of the present invention provides a device in which the system described in the above aspect is installed, or the computer-readable storage medium described in the above aspect is contained.
[0075] The beneficial effects of the present invention are:
[0076] 1. The present invention proposes a set of key factor combinations for Alzheimer's disease for the first time, and finds that they are highly correlated with Alzheimer's disease. Based on this key factor combination, an original Alzheimer's disease risk prediction model is successfully constructed.
[0077] 2. The present invention is based on the key factor combination of Alzheimer's disease and the cfRNA algorithm GLM-Smooth of GLM (Generalized Linear Models), which realizes the construction of a model after cfRNA tuning, with an AUC of up to 0.93, and has excellent prediction effect and accuracy, so that it can predict high-risk populations with potential disease, avoid the inaccuracy of traditional simple bioinformatics analysis plus human subjective judgment verification, reduce interference from human factors to a certain extent, and improve the credibility and accuracy of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 Flow chart of the sequencing data steps in an embodiment of the present invention.
[0079] Figure 2 The figure is a flow chart of the modeling steps in the embodiment of the present invention.
[0080] Figure 3 This is a differential expression map of the key factors of Alzheimer's disease in the embodiments of the present invention between AD and healthy controls.
[0081] Figure 4 Violin plot based on the GLM-Smooth classifier scores.
[0082] Figure 5 This is a ROC graph of the Alzheimer's disease risk prediction model measured in an embodiment of the present invention. DETAILED DESCRIPTION
[0083] The present invention is further described in detail below by specific examples. Unless otherwise specified, the raw materials, reagents or devices used in the examples and comparative examples can be obtained from conventional commercial sources or can be obtained by prior art methods. Unless otherwise specified, the experiments or test methods are conventional methods in the art.
[0084] Example 1
[0085] In this embodiment, an Alzheimer's disease risk prediction model and a construction method thereof are provided, and the specific steps are as follows:
[0086] (1) Sample information used for model construction:
[0087] In this embodiment, the inventors collected 66 plasma samples from patients diagnosed with Alzheimer's disease and healthy controls, and the specific compositions are shown in the following table.
[0088] Table 1 Plasma sample information
[0089] type Healthy control group plasma Alzheimer's Plasma total Quantity (% of total) 25(37.88%) 41(62.12%) 66(100%)
[0090] (2) RNA extraction:
[0091] In this embodiment, RNA extraction was performed using a commercially available kit (Tiangen RNA extraction kit, DP501). Of course, those skilled in the art can also use other conventional methods in the art to extract RNA.
[0092] RNA extraction can be performed according to the kit instructions, or using the following steps:
[0093] Take 200 μL of plasma sample, add 900 μL of lysis buffer MZ (from the kit), shake and mix in a 2 mL centrifuge tube, and let stand at room temperature for 5 minutes. Centrifuge at 12000 rpm for 5 minutes, take the supernatant, and transfer it to a new 1.5 mL centrifuge tube. Add 0.2 mL of chloroform to the supernatant, shake vigorously for 15 seconds, and then let it stand at room temperature again for 5 minutes. Centrifuge at 10000 × g for 15 minutes at 2-8 ° C. At this time, the centrifuged solution will be divided into three layers, including a colorless aqueous phase layer (upper layer), an intermediate layer, and a pink organic phase layer (lower layer). RNA is mainly present in the aqueous phase. Take the aqueous phase layer and transfer it to a new centrifuge tube, add 1.5 times the volume of anhydrous ethanol, and precipitation may appear at this time. Gently invert to mix thoroughly. Add all the solution and precipitation to the adsorption column (miRspin, from the kit), let stand at room temperature for 2 minutes, centrifuge at 12000 rpm for 30 seconds at room temperature, and discard the effluent. Add 500μL of deproteinized solution MRD (from the kit), centrifuge at 12000rpm at room temperature for 30 seconds, and discard the effluent. Add 500μL of rinse solution RW (from the kit), centrifuge at 12000rpm at room temperature for 30 seconds, and discard the effluent. Repeat the rinse once. Then centrifuge at 12000rpm at room temperature for 1 minute to completely remove the residual ethanol. Place the adsorption column in a 1.5mL RNasc-free Tube, add 30μL of RNase-free Water to the center of the adsorption column, and place it at room temperature for 2 minutes. Centrifuge at 12000rpm at room temperature for 2 minutes. After elution, total RNA can be obtained and can be stored at -80℃.
[0094] (3) Fragmentation and removal of ribosomal rRNA:
[0095] Take 7.5 μL of the RNA sample eluted in step (2), and use a commercially available kit (Novozyme Ribo-off rRNA Depletion Kit (H / M / R), N406-01) to remove the rRNA (including cytoplasmic 28S, 18S, 5.8S, and 5S rRNA, and mitochondrial 16S, 12S rRNA), and retain the mRNA and other non-coding RNA. The operation steps are carried out according to the instructions for use of the kit, and the reaction procedures are: 94°C, 7min; 75°C, 2min; 60°C, 10min.
[0096] (4) Library construction:
[0097] Using Hieff NGS from Shanghai Yisheng Biotechnology Co., Ltd. TM Ultima Dual-Mode RNA Library Preparation Kit was used for library construction.
[0098] The specific steps are:
[0099] use RNA384CDIPrimer for (Applicable to The RNA samples after rRNA removal were barcoded using t-96×2T Set1, and then the samples were pooled to a concentration of 2 nM and sequenced using the PE150 strategy on Illumina's NovaSeq X Plus. At the same time, commercially available positive and negative control samples (nuclease-free water and human brain RNA (Takara)) were used to assess whether there was contamination during RNA extraction and library preparation.
[0100] (5) NGS sequencing:
[0101] NGS sequencing is performed according to conventional steps in the art, including: calculating the library concentration, diluting it according to the instructions for use of the kit, then denaturing and re-diluting the library, injecting the sample into the sample well of the kit, importing the sample information, cleaning the chip, inserting the kit and chip, confirming the machine settings, starting sequencing, and obtaining sequencing data.
[0102] (6) Processing of sequencing data:
[0103] The flowchart of this step is as follows Figure 1 shown.
[0104] A. Sequencing data cleaning:
[0105] Fastp v0.23.3 was used to perform data quality control on the sequencing data obtained in step (5), and low-quality data (including polyG and adapter sequences, sequences containing more than 5 N bases, and sequences with a Phred quality score of less than 15 for more than 40% of the bases) were cleaned up and low-quality sequences were removed. Then, barcode sequence detection and removal were performed to obtain high-quality sequencing data (FASTQ file).
[0106] B. Data comparison:
[0107] The high-quality sequencing data obtained in the above steps were compared with the human genome data using STAR v2.7.10b.
[0108] The specific steps are: use the human genome reference sequence and annotation files in the public database to build a genome index, and then compare the high-quality sequencing data (FASTQ file) obtained in the above steps with the constructed genome index to generate an alignment file (BAM format).
[0109] C. Gene count:
[0110] Gene counting was performed using Feature Counts v 2.0.6.
[0111] The specific steps are: align the generated alignment file (BAM format) with the gene annotation file (GTF format). Then, count the number of sequencing reads for each gene to generate a gene count matrix.
[0112] D. Normalization:
[0113] In order to enhance counting stability and reduce the deviation caused by library size, each gene was normalized using the following formula:
[0114]
[0115] in, represents the normalized gene expression; x i is the original expression of the corresponding gene; λ is the amplification constant, which is used to avoid the value being too small; n is the number of genes in the corresponding sample.
[0116] (7) Data analysis:
[0117] In this example, the GLM-Smooth algorithm based on GLM was used to analyze the processed sequencing data.
[0118] For the binary classification problem in data analysis (i.e., whether someone will have Alzheimer's disease), a feature matrix X (key feature) and a response variable y (whether someone has the disease) are set, where y can take the value of 0 or 1. The GLM-Smooth model is used to describe the relationship between the response variable and the feature:
[0119]
[0120] Among them, σ represents the logistic function and β represents the regression coefficient to be estimated.
[0121] Use the log-likelihood function as the loss function. Its negative log-likelihood function is:
[0122]
[0123] In order to solve the problems of multicollinearity and overfitting, a regularization term is further introduced:
[0124]
[0125] Among them, λ represents the regularization parameter, which is used to control the regularization strength.
[0126] Calculate the minimized loss function L(β). In this embodiment, the optimal solution can be obtained by calculating the gradient of the loss function with respect to β and setting it to zero. The formula for calculating the gradient of the loss function is as follows:
[0127]
[0128] Since the loss function is not quadratically differentiable, an analytical solution cannot usually be obtained directly. Therefore, it is necessary to further introduce the quasi-Newton method to obtain its numerical solution for approximation.
[0129] Based on the above method, by introducing regularization into the algorithm, the variance of the regression coefficient is effectively reduced, the robustness of the final model is improved, and finally the coefficient β that can be used to predict new samples is obtained.
[0130] The flowchart of the above steps is as follows Figure 2 shown.
[0131] Specifically, in this embodiment, the above 66 samples are randomly sampled and divided into a training set (42 samples, accounting for about 64%), a validation set (11 samples, accounting for about 16%), and a test set (13 samples, accounting for about 20%). Among them, the training set data is used to train the model. The model adjusts its own parameters by learning the relationship between the sample features and labels in the training set so that it can accurately predict new samples. The validation set is used to calculate the optimal hyperparameters, thereby effectively avoiding overfitting or underfitting and ensuring that the model performs better in practical applications.
[0132] Through preliminary screening, FZD6, GRIN2A, TUBA8, RYR3, NDUFS3, and TNF were selected as key characteristic genes (the expression differences of each gene in AD and healthy controls are shown in Figure 3 As shown, the gene information and its weight in the model construction of the embodiment of the present invention are shown in the following table).
[0133] Table 2 Key characteristic gene information and its weight
[0134]
[0135] In the above table, the weight indicates the contribution of the feature to the model prediction. The weight is related to the correlation between the factor itself and the disease. The higher the weight, the greater the correlation with the disease, and vice versa. The lower the weight, the smaller the correlation with the disease. Among the key factors used for prediction, it can be found that FZD6 is the most important and has the highest weight in the prediction results.
[0136] The expression data of the six key characteristic genes in the two data sets are used as input features, which can more specifically judge and predict the risk of Alzheimer's disease. This supervised learning can also improve the prediction accuracy of the model.
[0137] During the construction and subsequent validation set calibration process, five-fold cross validation was used to improve data utilization.
[0138] According to the steps in the above embodiment, the Alzheimer's disease risk prediction model finally obtained is:
[0139]
[0140] According to the above steps, the model formula can be converted to:
[0141]
[0142] Then, after obtaining the coefficient β of each key characteristic gene according to the above steps, the following final model can be obtained:
[0143]
[0144] or
[0145]
[0146] Among them, P represents the probability of Alzheimer's disease in a single sample (i.e., the predicted value), β0 represents the intercept (bias) of the model, and β 1-6 represents the weight of each gene expression value in plasma in the sample in the model (as shown in Table 2), X1-6 It represents the specific expression value (measured value) of each gene in the sample in plasma, and the order is FZD6, GRIN2A, RYR3, TNF, TUBA8, and NDUFS3.
[0147] The final prediction result, i.e. the judgment of the risk of Alzheimer's disease, is determined by the predicted value. When the predicted value reaches a significant value (risk threshold), after inputting new data into the trained model, the model outputs "AD" to indicate that there is a risk of Alzheimer's disease or a high risk of Alzheimer's disease; if it does not reach this value, the model outputs "none" to indicate that there is no risk of Alzheimer's disease or a low risk of Alzheimer's disease. The output statement is the final evaluation result.
[0148] In this embodiment, the risk threshold can be obtained based on the GLM-Smooth classifier score. Based on data analysis, the violin plot of the GLM-Smooth classifier score is as follows: Figure 4 shown.
[0149] based on Figure 4 The results shown show that the risk threshold for judging the risk of disease is 0.75924105, that is, when P (predicted value) ≥ 0.75924105, the sample is judged to have the risk of Alzheimer's disease or the risk of Alzheimer's disease is high; when P < 0.75924105, the sample is judged to have no risk of Alzheimer's disease or the risk of Alzheimer's disease is low.
[0150] Example 2
[0151] In order to verify the fairness and reliability of the above model, the test set split out in the above embodiment is used to evaluate the model performance.
[0152] In this embodiment, the performance of the Alzheimer's disease risk prediction model is evaluated from the indicators such as accuracy (ROC-AUC), sensitivity, and specificity.
[0153] The results are as follows Figure 5 As shown, the marked dots are the specificity and sensitivity points corresponding to the optimal threshold, with a sensitivity of 100% and a specificity of 80%.
[0154] After inputting the test set data into the Alzheimer's disease risk prediction model, the ROC curve was drawn based on the prediction results, and it was found that the AUC reached 93%. This shows that the Alzheimer's disease risk prediction model constructed in the embodiment of the present invention has high accuracy in decision-making and judgment of given data, and can be effectively used for predicting the risk of Alzheimer's disease.
[0155] In addition, the inventors further randomly selected 10 independent samples and used the Alzheimer's disease risk prediction model for prediction, and compared and analyzed them with the actual situation. The results are shown in the following table.
[0156] Table 3 Prediction of 10 independent samples
[0157] Sample No. Model prediction results Actual Results Is the prediction accurate? 1 0.25561355 0 yes 2 0.04808292 0 yes 3 0.97899813 1 yes 4 0.3603914 0 yes 5 0.97075241 1 yes 6 0.04808292 0 yes 7 0.85831537 1 yes 8 0.98979607 1 yes 9 0.75924105 1 yes 10 0.95625235 1 yes
[0158] In the table, 0 represents the healthy group and 1 represents the AD group.
[0159] It can be found from the above results that the Alzheimer's disease risk prediction model in the embodiment of the present invention can be used for the prediction of actual samples with high prediction accuracy.
[0160] Example 3
[0161] In this embodiment, an Alzheimer's disease risk prediction product is provided, which includes: an input end, an analysis module and an output end.
[0162] The input end is used for inputting data or information. For example, in this embodiment, the input end is used to receive the expression of FZD6, GRIN2A, RYR3, TNF, TUBA8 and NDUFS3 and related information, and then transmit the information to the analysis module. The input end can be optionally configured according to the needs of the usage scenario, including: a keyboard, a mouse, a hard disk, a magnetic disk, a writable read-only compact disk (CD-ROM), an optical paper tape input device, a card input device, an optical character reader, a magnetic tape input device, a Chinese character input device, etc.
[0163] The analysis module contains a storage medium on which the Alzheimer's disease risk prediction model in the above-mentioned embodiment is written, so as to execute the Alzheimer's disease risk prediction model in the analysis module. By receiving the information transmitted from the input end, the information is input into the Alzheimer's disease risk prediction model for calculation to obtain the prediction result. The prediction result will be exported from the analysis module to the output end. Among them, the analysis module can be optionally included according to the needs of the use scenario: a general-purpose computer, a cloud computing network, a single-chip microcomputer, an embedded system (such as a smart phone, home appliances, etc.), a dedicated processor (such as a graphics processor and a tensor processor), an algorithm processor (such as a neural computer and an array processor), etc.
[0164] The output end is used for outputting data or information, and optionally includes: a mobile client (such as a mobile phone, a tablet computer, a laptop computer), a display, a printer, a light pen, a hard disk, a magnetic disk, a writable read-only compact disk (CD-ROM), etc., depending on the requirements of the usage scenario.
[0165] According to the use requirements, the product can further include:
[0166] Detection module: used to receive samples for detection and output the detection results as input information.
[0167] Storage module: used to store the information or results of each module.
[0168] Transmission module: used to transmit information output by other modules, and can simultaneously transmit one or more information to one or more receiving ends.
[0169] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A set of diagnostic or prognostic markers, characterized in that The diagnostic or predictive markers include: at least one of FZD6, GRIN2A, TUBA8, RYR3, NDUFS3, and TNF; Preferably, the diagnostic or predictive markers include FZD6, GRIN2A, TUBA8, RYR3, NDUFS3 and TNF; Preferably, the diagnostic or predictive marker is an Alzheimer's disease diagnostic or predictive marker.
2. Use of a reagent for detecting the diagnostic or predictive marker of claim 1 in the preparation of a product for diagnosis and / or risk assessment of Alzheimer's disease; Preferably, the Alzheimer's disease diagnosis and / or risk assessment prediction product includes a detection reagent, a detection kit and a detection chip.
3. An Alzheimer's disease diagnosis and / or risk assessment prediction product, characterized in that: The Alzheimer's disease diagnosis and / or risk assessment prediction product comprises a reagent for detecting the diagnostic or predictive marker of claim 1, and an adjuvant; Preferably, the auxiliary agent includes at least one of a buffer, a solvent or a diluent, a preservative, a surfactant, a cleaning solution, and a sample pretreatment reagent.
4. A method for constructing an Alzheimer's disease risk prediction model, comprising the following steps: Collecting the expression of the diagnostic or predictive markers and the prevalence of Alzheimer's disease as described in claim 1, and constructing a feature matrix; The feature matrix is fitted using a modeling algorithm to obtain an Alzheimer's disease risk prediction model; Preferably, the construction method further comprises: after the fitting is completed, optimizing the model parameters using a regularization algorithm; Preferably, the modeling algorithms include: GLM-Smooth, random forest model.
5. Use of the Alzheimer's disease risk prediction model constructed by the construction method of claim 4 in the preparation of Alzheimer's disease diagnosis and / or risk assessment prediction products; Preferably, the Alzheimer's disease diagnosis and / or risk assessment prediction product includes an analysis module, and the analysis module obtains a prediction result by executing the Alzheimer's disease risk assessment prediction model constructed by the construction method described in claim 4.
6. The use according to claim 5, characterized in that The Alzheimer's disease risk prediction model includes: or Among them, P represents the predicted value of Alzheimer's disease risk, X 1-6 They represent the expression levels of FZD6, GRIN2A, RYR3, TNF, TUBA8 and NDUFS3 respectively.
7. The use according to claim 5, characterized in that The analysis module outputs a predicted value of the risk of Alzheimer's disease, and determines whether there is a risk of Alzheimer's disease based on the predicted value of the risk of Alzheimer's disease; Preferably, the Alzheimer's disease risk judgment criteria based on the Alzheimer's disease risk prediction value are: Comparison of the predicted risk of Alzheimer's disease and the risk threshold of Alzheimer's disease. If the predicted value of the subject's Alzheimer's disease risk is less than the Alzheimer's disease risk threshold, it is judged that there is no risk of Alzheimer's disease or the risk of Alzheimer's disease is low; If the predicted value of the subject's Alzheimer's disease risk is greater than or equal to the Alzheimer's disease risk threshold, the subject is judged to have an Alzheimer's disease risk or a high Alzheimer's disease risk.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of calculating the Alzheimer's disease risk prediction model constructed by the construction method described in claim 4 are implemented.
9. A system, characterized in that: The system includes an analysis module, wherein the analysis module contains the computer-readable storage medium of claim 8; Preferably, the system further comprises an input end and an output end.
10. A device, characterized in that: The device is installed with the system described in claim 9, or contains the computer-readable storage medium described in claim 8.