Method for screening biomarkers for screening early non-invasive Alzheimer's disease based on machine learning
By combining hematogenous cfRNA-seq data and brain-derived scRNA-seq data, key biomarker genes were screened out using machine learning methods, solving the problem of low accuracy of non-invasive screening in early Alzheimer's disease, and achieving high-precision non-invasive screening and personalized therapeutic support.
Patent Information
- Application Number
- CN202510158559.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to achieve high-precision non-invasive screening for early Alzheimer's disease, and traditional diagnostic methods have problems such as invasiveness, high cost and unstable diagnostic effects.
By obtaining blood-derived cfRNA-seq data and brain-derived scRNA-seq data from AD patients and controls, the genes with co-expression patterns were screened as biomarkers by machine learning method, and differential expression analysis was performed using pseudobulk analysis and DESeq2 software package to screen out key biomarker genes.
Non-invasive screening of early Alzheimer's disease has been achieved, which significantly improves the accuracy of diagnosis, can provide timely intervention treatment for high-risk groups, and distinguish AD patients with different disease progression, and support personalized treatment plans.
Smart Images

Figure CN120072059A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disease diagnosis, and particularly relates to a method and application for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning. Background Art
[0002] With the aggravation of the aging of the global population, the incidence rate of AD has been increasing year by year. Due to the lack of effective diagnostic methods for AD in clinical practice, most AD patients are diagnosed in the late stage of the disease, and there is currently no effective drug to reverse or prevent the progression of the disease. This makes the early screening of AD particularly urgent. However, these traditional diagnostic methods such as neuroimaging and cerebrospinal fluid (CSF) analysis rely on the detection of biomarkers, including amyloid-β (Aβ), total Tau protein (t-Tau), and phosphorylated Tau protein (p-Tau). However, these biomarkers are relatively specific for AD, and their levels may also increase in other neurodegenerative diseases, thus limiting the specific diagnosis of AD. Traditional diagnostic methods are invasive, costly, and have high risks. In addition, the methods, reagents, and reference ranges for CSF biomarker detection have not been standardized, resulting in uneven diagnostic effects.
[0003] Blood-based biomarkers (BBMs) study molecular changes through non-invasive procedures, avoiding the risks and discomforts of traditional surgeries. However, due to the existence of the blood-brain barrier, the development of blood-brain barrier drugs for neurodegenerative diseases such as AD is still more complex and challenging. Some studies have shown that Aβ and tau in the blood can replace traditional CSF markers for the diagnosis of AD, which is cost-effective. However, the measurement of the levels of Aβ and tau proteins is not sensitive and direct enough to accurately diagnose early AD. There is an urgent need for a new type of BBM that directly originates from brain lesions and can accurately reflect the potential mechanisms of the lesions. The cfRNA detected in the blood provides a non-invasive method to directly evaluate the status of multiple tissues. Studies have confirmed that the expression levels of cf-miRNAs in the blood can distinguish between people with normal cognitive ability, patients with mild cognitive impairment (MCI), and AD patients. More and more evidence supports the view that RNA molecules can cross the blood-brain barrier, and the brain-derived cfRNA detected in the blood can be used as a biomarker for non-invasive molecular analysis of nervous system diseases such as AD. However, there is currently no literature directly linking cfRNA and single-cell sequencing markers to AD biopsies for non-invasive monitoring of disease progression and treatment efficacy. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning in view of the deficiencies of the prior art.
[0005] The object of the present invention is achieved by the following technical solutions:
[0006] A method for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning, comprising:
[0007] Obtaining blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls and normalizing them;
[0008] Based on the normalized blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls, genes with common expression patterns in the two types of data are screened out as potential non-invasive Alzheimer's disease screening biomarkers.
[0009] Further, the pseudo-bulk analysis method is used to normalize the gene abundance levels in the blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls.
[0010] Further, the DESeq2 software package is used to perform differential expression analysis between the blood-derived cfRNA-seq data of AD patients and age-matched controls and between the brain-derived scRNA-seq data of AD patients and age-matched controls, and genes with a p-value < 0.05 in both types of data are screened out, and the threshold for scRNA-seq is an absolute log2 fold change >= 0.25, and the cfRNA-seq threshold is an absolute log2 fold change >= 0;
[0011] Then, the RFECV feature algorithm of Python is used to screen out the key biomarker genes with co-expression from the genes obtained above as biomarkers for early non-invasive Alzheimer's disease screening.
[0012] Further, the identified biomarker genes are specifically as follows:
[0013] RAB11FIP4, NEAT1, IKZF1, EZR, CREB5, MNDA, PTPN6, GGA2, YBEY, RUNX2, MTRNR2L12, SNX30, MTATP6P1, RBM47, MRPS23, STIP1, RPL6P27, FLOT2, HSPA8, CYTH1, ZBTB18, LBH, TSPYL1, RPS17, CCT5, RPL3P4, LAMTOR4, BCL6, GCA, BCL2, APLNR, ANPEP, ARL1, ACSL1
[0014] An early non-invasive Alzheimer's disease diagnosis method based on machine learning, comprising:
[0015] Obtain the blood-derived cfRNA-seq data of the patient to be tested and extract the gene expression data corresponding to the biomarkers obtained by the method of screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning;
[0016] Input the gene expression data into a trained AD diagnosis classifier to output a diagnosis result.
[0017] Furthermore, the AD diagnosis classifier is trained based on a training dataset with the goal of minimizing the error between the AD diagnosis classifier and the true value.
[0018] Furthermore, the AD diagnosis classifier is one of SVM, RF, and LR.
[0019] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the early non-invasive Alzheimer's disease diagnosis method based on machine learning is implemented.
[0020] A storage medium containing computer-executable instructions, wherein when the computer-executable instructions are executed by a computer processor, the early non-invasive Alzheimer's disease diagnosis method based on machine learning is implemented.
[0021] A computer program product, comprising a computer program / instructions, wherein when the computer program / instructions are executed by a processor, the steps of the early non-invasive Alzheimer's disease diagnosis method based on machine learning are implemented.
[0022] The beneficial effect of the present invention is that the biomarkers screened based on cfRNA-seq and scRNA-seq data of the present invention can be applied to non-invasive screening of AD, significantly improving the accuracy of AD diagnosis. Thus, it can provide timely intervention treatment for high-risk populations and may prevent disease progression; at the same time, it can distinguish AD patients with different disease progressions, helping to develop drug response monitoring indicators and providing support for personalized treatment plans for AD patients. Description of the Drawings
[0023] The present invention will be further described below with reference to the drawings and embodiments;
[0024] Figure 1 It is a diagram showing the intersection of differentially expressed genes of cfRNA in the AD patient group and the normal control group and AD lesion-related genes from different sources;
[0025] Figure 2Cell type map of single-cell transcriptome data from the brain tissues of AD patients;
[0026] Figure 3 Line graph of the expression level trends of characteristic genes of each cell type in cfRNA;
[0027] Figure 4 Bar graph of the quantity ratios of each cell type in AD patients and the control group;
[0028] Figure 5 Graph of the results of the association analysis between biomarker genes and known disease databases;
[0029] Figure 6 Graph of the results of the protein - protein interaction (PPI) analysis of biomarker genes;
[0030] Figure 7 Graph of the analysis of the expression changes of biomarker genes in different cell types under the influence of AD;
[0031] Figure 8 Graph of the tissue - specific expression results of biomarker genes in the GTEx database;
[0032] Figure 9 Graph of the results of the functional enrichment analysis of differentially expressed genes screened based on cfRNA data;
[0033] Figure 10 Graph of the prediction performance results of the AD diagnostic classifier constructed based on 34 biomarker genes;
[0034] Figure 11 Graph of the prediction performance results of the AD diagnostic classifier constructed based on biomarker genes screened only from cfRNA data;
[0035] Figure 12 Graph of the results of independently validating the prediction performance of the AD diagnostic classifier in different datasets;
[0036] Figure 13 Graph of the results of biomarker genes distinguishing AD samples at different progression stages in the MSBB dataset;
[0037] Figure 14 Graph of the expression of 34 biomarker genes in the MSBB and ROSMAP datasets. Detailed implementation methods
[0038] The present invention collected blood-derived cfRNA sequencing data of 242 samples, including 126 AD samples and 116 age-matched control samples. The obtained dysregulated genes were significantly enriched in pro-inflammatory biological processes and impaired nervous system functions, indicating that blood cfRNA can detect the pathological characteristics of AD. However, the classification model based on cfRNA dysregulated genes lacks robustness and accuracy and cannot effectively distinguish AD patients. To explore the potential of cfRNA in detecting brain lesions in AD, the present invention integrated cfRNA-seq data with brain-derived scRNA-seq data and screened out genes with common differential expression in the two types of data as potential biomarkers for early non-invasive Alzheimer's disease screening. Finally, a total of 34 signature genes were identified, which showed the highest importance scores in the feature selection algorithm and were differentially expressed in both the cfRNA and scRNA datasets. The machine learning model constructed based on the cfRNA expression patterns of these 34 genes can accurately predict AD patients and healthy individuals. In addition, these models can also effectively identify patients in the early stage of disease progression, which is crucial for timely initiation of treatment interventions. In addition, the classifier developed based on the expression levels of these 34 signature genes in the brain transcriptome data also showed strong predictive performance and can be used to evaluate the risk of an individual being likely to develop AD. These findings highlight the potential of these 34 genes as biomarkers for early non-invasive AD screening and pave the way for improving diagnostic accuracy and patient stratification in AD research and clinical practice. To better illustrate the purpose, technical solution and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments.
[0039] Example 1:
[0040] A method for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning according to the present invention includes:
[0041] Step 1: Obtain and standardize blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls;
[0042] In this embodiment, blood-derived cfRNA-seq data of 126 AD patients and 116 age-matched healthy controls and brain-derived scRNA-seq data of 46 AD patients and 42 healthy controls were collected and downloaded; due to the sparsity of single-cell sequencing data itself, the pseudo-bulk analysis method was used to standardize the gene abundance levels in scRNA data and cfRNA data.
[0043] Step 2: Based on the standardized blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls, genes with common differential expression in the two types of data are screened out. Using the RFECV feature algorithm in Python, key biomarker genes are screened out from the genes obtained above as biomarkers for early non-invasive Alzheimer's disease screening.
[0044] In this example, the DESeq2 software package is used to perform differential expression analysis between the blood-derived cfRNA-seq data of AD patients and age-matched controls and between the brain-derived scRNA-seq data of AD patients and age-matched controls, and genes with a p-value < 0.05 in both types of data are screened out. The threshold for scRNA-seq is an absolute log2 fold change ≥ 0.25, and the threshold for cfRNA-seq is an absolute log2 fold change ≥ 0. Using the RFECV feature algorithm in Python, key biomarker genes are screened out from the genes obtained above as biomarkers for early non-invasive Alzheimer's disease screening. A total of 34 biomarker genes are screened out in this example, as follows:
[0045] RAB11FIP4, NEAT1, IKZF1, EZR, CREB5, MNDA, PTPN6, GGA2, YBEY, RUNX2, MTRNR2L12, SNX30, MTATP6P1, RBM47, MRPS23, STIP1, RPL6P27, FLOT2, HSPA8, CYTH1, ZBTB18, LBH, TSPYL1, RPS17, CCT5, RPL3P4, LAMTOR4, BCL6, GCA, BCL2, APLNR, ANPEP, ARL1, ACSL1
[0046] The brain-derived RNA-seq data and cfRNA data of three AD-related datasets (ROSMAP, Mayo, and MSBB) are used for correlation analysis, and the differential expression patterns between AD patients and age-matched healthy control groups are compared. The analysis results show that, as Figure 1 shown, the genes showing differential expression detected in cfRNA are also altered in AD brain lesions. This indicates that cfRNA may originate from the brain and can cross the blood-brain barrier and enter the blood, thus providing a theoretical basis for the use of cfRNA in non-invasive diagnostic methods.
[0047] Furthermore, to investigate the molecular connections between blood cfRNA and brain-specific changes in AD patients, the Scanpy Python package was used to perform quality control and principal component analysis (PCA) on scRNA-seq data from four different brain regions (hippocampus, prefrontal cortex, frontal cortex, medial cortex). The bbknn Python package was used to remove batch effects from the data, and the data was then clustered and cell-annotated. Eight major cell types in the brain were annotated according to their respective canonical marker genes. The cell populations were classified as astrocytes, endothelial cells (Endo), microglia, mature oligodendrocytes (mOli), neurons, oligodendrocyte precursor cells (OPC), pericytes, and perivascular fibroblasts (PVF). The results are as Figure 2 shown, a trend of increased numbers of mature oligodendrocytes and microglia was observed in various brain regions of AD; in contrast, perivascular fibroblasts, pericytes, and endothelial cells were more prevalent in the control group, while oligodendrocyte precursor cells and neurons remained relatively stable in both groups. These trends in cell type changes are consistent with previous studies and conform to the pathological change characteristics of AD. As Figure 3 shown, the analysis indicated that most of the characteristic genes of the eight cell types could be detected in the cfRNA data, suggesting that brain-derived RNA can cross the blood-brain barrier and enter the blood. Quantitative analysis of the expression levels of these characteristic genes in the cfRNA data showed that microglia and mature oligodendrocyte markers were upregulated in AD. In contrast, the expression levels of markers for endothelial cells, neurons, oligodendrocyte precursor cells, pericytes, and perivascular fibroblasts were decreased, which is consistent with the changes in cell proportions observed in the single-cell data. This indicates that the blood cfRNA profile can be used for non-invasive detection of cell-specific characteristics in the brains of AD patients. Statistical analysis of various cell types was performed using the Ro / e method, as Figure 4As shown, it was found that neurons, mature oligodendrocytes, astrocytes, endothelial cells, microglia, and pericytes were the main factors leading to dysregulated gene expression, while oligodendrocyte precursor cells and perivascular fibroblasts had the least impact on gene expression. Gene ontology biological process pathway enrichment analysis of co-expressed differential genes using the "enrichR" function of clusterProfiler revealed that the downregulated pathways were involved in biological processes such as axon guidance, synaptic vesicle cycle, glutamatergic synapse, chemical synaptic transmission, neurotransmitter secretion, and nervous system development, which are typical features of neurodegenerative diseases; while the upregulated pathways, such as T cell receptor signaling pathway, B cell differentiation, and cytokine-mediated signaling pathway, were closely related to neuroinflammation, indicating enhanced neuroinflammation in AD patients. By integrating cfRNA-seq and scRNA-seq data, the present invention can comprehensively describe the progression of AD in the brain and blood, including neuronal death and neuroinflammation, which are the key pathological hallmarks of the disease. It shows that scRNA data and cfRNA data have the same gene expression changes and are highly correlated with the key pathological features of AD.
[0048] Based on the expression profiles of 34 key signature genes, the samples could be automatically clustered into AD patients and age-matched healthy controls. Through association analysis with known disease databases, such as Figure 5 As shown, it was found that BCL2 was the most important hub gene, associated with multiple diseases such as bipolar disorder, memory impairment, learning disability, AD, and major depressive disorder. In addition, PTPN6, STIP1, ANPEP, and TSPYL1 were also found to be associated with major depressive disorder. Through protein-protein interaction (PPI) analysis, as Figure 6 shown, it was further emphasized that BCL2, BCL6, HSPA8, and EZR were important hub genes. Studies have shown that inhibiting the expression of BCL6 could be used as a therapeutic target for central nervous system cancer, and HSPA8 is a molecular chaperone that can mediate autophagy and affect the hydrolysis of misfolded proteins.
[0049] As Figure 7 shown by single-cell expression profiling analysis, the expression patterns of these 34 biomarker genes were significantly different in different cell types affected by AD. Most genes had the greatest changes in microglia, followed by neurons, astrocytes, and oligodendrocytes. Tissue-specific expression pattern analysis based on RNA-seq data from different tissue sources in the GTEx database showed ( Figure 8) Compared with other tissues, the expression levels of most biomarker genes are significantly higher in the brain, including APLNR, MTATP6P1, MTRNR2L12, RAB11FIP4, SNX30, TSPYL1, and ZBTB18. These results indicate that these 34 key biomarker genes can be traced back to the origin of brain cell types and can also reflect the expression changes of different brain cell types during the onset of AD, thus describing the pathological characteristics of the disease.
[0050] Comparative Example 1:
[0051] This method is based on the blood-derived cfRNA-seq data of 126 AD patients and 116 age-matched healthy controls. The data is converted into FASTQ files using the fastq-dump pipeline. Then, quality control processing is performed on the data using fastq, and gene expression counting is carried out using featureCounts. Subsequently, the DESeq2 Bioconductor software package is used for count normalization and batch correction to adjust for the batch effects of sequencing depth and sample origin. The screening criteria for differentially expressed genes are set as p-value < 0.05 and absolute log2 fold change ≥ 0.25. Finally, a set of 2,658 downregulated genes and 431 upregulated genes are obtained.
[0052] Functional enrichment of the screened differentially expressed genes, as Figure 9 shown, reveals that the upregulated genes are enriched in pathways related to Parkinson's disease, AD, antigen processing and presentation, and cytokine production. This enrichment result is consistent with the pathological view that AD patients exhibit a pro-inflammatory state, which is prone to disease progression. The downregulated genes show significant enrichment in key pathways of nervous system development, including synapse organization, axonogenesis, synapse assembly, neuron development, and glutamatergic synapses. These pathways are crucial for the normal function of the nervous system, and their dysregulation reflects the loss of functional mechanisms in degenerative diseases. The results of gene functional enrichment indicate that blood cfRNA can detect the pathological characteristics of AD. These differential genes may all be biomarkers for AD screening, but the large number of genes poses a huge challenge for practical applications.
[0053] It can be seen that the method of the present invention for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning integrates cfRNA-seq data with brain-derived scRNA-seq data and can effectively screen out biomarkers that can most effectively identify Alzheimer's disease at an early stage.
[0054] Example 2
[0055] An early non-invasive Alzheimer's disease diagnosis method based on machine learning, comprising:
[0056] Obtain the blood-derived cfRNA-seq data of the patient to be tested and extract the gene expression data corresponding to the biomarkers obtained by the method for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning described in Example 1;
[0057] Input the gene expression data into a trained AD diagnosis classifier to output a diagnosis result.
[0058] The AD diagnosis classifier is trained based on a training data set with the goal of minimizing the error between the AD diagnosis classifier and the true value.
[0059] In this example, the collected data was divided into a training set (70%), a test set (20%), and a validation set (10%) for cross-validation. Finally, three different AD diagnosis classifiers were trained, including SVM, RF, and LR. Each classifier was cross-validated 10 times in both the training group and the validation group. And a receiver operating characteristic curve (ROC) was generated for each classifier and the corresponding area under the curve (AUC) value was determined. The results are as Figure 10 shown. The RF model constructed based on the cfRNA expression levels of 34 signature genes had the highest AUC, reaching 89%. The SVM and LR classifiers also showed high prediction performance (SVM: AUC = 0.80, LR: AUC = 0.81). The results indicate that combining brain scRNA data with cfRNA analysis in the blood can more precisely capture molecular changes in the brain. Compared with cfRNA-based biomarkers alone, these cfRNA- and scRNA-based biomarkers greatly improve the accuracy in diagnosing AD.
[0060] Comparative Example 2:
[0061] In this comparative example, among the 2658 downregulated genes and 431 upregulated genes screened based on cfRNA data alone in Comparative Example 1, 47 key genes were determined for the classification task using a feature selection algorithm. Three different AD diagnosis classifiers, including SVM, RF, and LR, were trained using the same method as in Example 2. Each classifier was cross-validated 10 times in both the training group and the validation group. And an ROC was generated for each classifier and the corresponding AUC value was determined. As Figure 11 shown, all three models achieved relatively good classification results (AUC >= 0.8). However, in the independent validation set, the highest AUC among the three models was only 0.66, which indicates that cfRNA-based biomarkers alone may lead to a high false positive rate when diagnosing AD. And the present invention uses brain single-cell scRNA data to establish a direct connection between cfRNA and genes related to AD brain lesions, thereby greatly improving the accuracy of diagnosing AD.
[0062] Example 3
[0063] To determine the applicability of these biomarker genes in classifiers constructed based on RNA-seq data from brain tissue sources, in this example, transcriptome data from brain sources were collected from three datasets, namely ROSMAP, Mayo, and MSBB, as the training dataset, which included brain tissue samples from AD patients and healthy control groups. Three different AD diagnostic classifiers, including SVM, RF, and LR, were trained using the same method as in Example 2. Each classifier was subjected to 10-fold cross-validation in both the training group and the validation group. As Figure 12 shown, the results indicate that in the Mayo dataset, the LS and SVM classifiers based on the expression of 34 biomarker genes had the highest AUC, reaching 94%. These genes also showed strong performance in the independent validation set, and the AUC values of the three classifiers were consistently higher than 86%. In addition, these biomarker genes also consistently maintained good predictive performance in the ROSMAP and MSBB datasets, indicating the predictive stability of the classifier based on biomarker genes in different datasets.
[0064] To evaluate the potential of the biomarker of the present invention in early screening, the RNA-seq data of brain tissues from AD patients across different disease stages were integrated. Unsupervised non-negative matrix factorization (NMF) clustering of AD samples in the MSBB dataset based on 34 key biomarker genes was able to divide the AD samples into two different groups. As Figure 13 shown, the results indicate that the first group was characterized by a higher average plaque load (plaqueMean) and an earlier death time. Statistical analysis further showed that the Braak stage and Clinical Dementia Rating (CDR) of the patients in this group were also higher; in contrast, the patients in Group 2 had a lower plaqueMean, a longer survival time, and lower Braak stage and CDR ratings. This indicates that the patients in the first group were in the late stage of disease progression, while the patients in the second group had mild symptoms and were classified as early AD. The results show that biomarker genes demonstrated a strong ability to distinguish AD patients at different stages. In addition, the expression levels of these biomarker genes in AD, normal aging, and MCI samples were also evaluated. As Figure 14 shown, it was found that most genes had the most obvious changes in AD samples. These results confirmed the role of the 34 biomarker genes in screening AD patients and emphasized their potential in the early detection of AD.
[0065] Corresponding to the foregoing embodiments of a method for early non-invasive diagnosis of Alzheimer's disease based on machine learning, the present invention further provides an electronic device. The electronic device includes one or more processors for implementing a method for early non-invasive diagnosis of Alzheimer's disease based on machine learning in the above embodiments.
[0066] The electronic device of the present invention can be on any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.
[0067] The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running. In terms of hardware, it includes a processor, a memory, a network interface, and a non-volatile memory. In addition, the any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.
[0068] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.
[0069] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0070] The embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for early non-invasive diagnosis of Alzheimer's disease based on machine learning in the above embodiments.
[0071] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0072] The above embodiments are used to explain the present invention rather than limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A method for screening biomarkers for early non-invasive Alzheimer's disease screening based on machine learning, characterized in that: include: Obtain and normalize blood-derived cfRNA-seq data and brain-derived scRNA-seq data from AD patients and age-matched controls; Based on the standardized blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls, genes with common expression patterns in the two types of data were screened as possible non-invasive Alzheimer's disease screening biomarkers.
2. The method according to claim 1, characterized in that Pseudo-bulk analysis was performed to normalize gene abundance levels in blood-derived cfRNA-seq data and brain-derived scRNA-seq data of AD patients and age-matched controls.
3. The method according to claim 1, characterized in that The DESeq2 software package was used to perform differential expression analysis between the blood-derived cfRNA-seq data of AD patients and age-matched controls, and between the brain-derived scRNA-seq data of AD patients and age-matched controls, and genes with a p-value < 0.05 in both data types, and the threshold of scRNA-seq was an absolute log2-fold change of ≥ 0.25, and the threshold of cfRNA-seq was an absolute log2-fold change of ≥ 0 were screened out; Then, the recursive feature elimination algorithm (RFECV) with cross-validation in Python was used to screen out the genes with common changes from the above genes as potential biomarkers for non-invasive Alzheimer's disease screening.
4. A method for early non-invasive diagnosis of Alzheimer's disease based on machine learning, characterized in that: include: Obtaining blood-derived cfRNA-seq data of the patient to be tested and extracting gene expression data corresponding to the biomarkers obtained by the method for screening biomarkers for non-invasive Alzheimer's disease screening based on a machine learning algorithm as described in claim 1; The gene expression data is input into a trained AD diagnosis classifier, and the diagnosis result is obtained as output.
5. The method according to claim 4, characterized in that The AD diagnosis classifier is obtained by training based on a training data set with the goal of minimizing the error between the AD diagnosis classifier and the true value.
6. The method according to claim 4, characterized in that The AD diagnosis classifier is one of support vector machine (SVM), random forest (RF) and logistic regression (LR).
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements an early non-invasive Alzheimer's disease diagnosis method based on machine learning as described in any one of claims 4-6.
8. A storage medium comprising computer executable instructions, which, when executed by a computer processor, implement an early non-invasive Alzheimer's disease diagnosis method based on machine learning as described in any one of claims 4 to 6.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of a method for diagnosing early-stage non-invasive Alzheimer's disease based on machine learning as described in any one of claims 4 to 6 are implemented.
Citation Information
Cited By
Immune metabolism crosstalk multi-omics anatomical Alzheimer disease prediction method and device
CN121148677A