Clinical decision support methods for predicting Moyamoya disease based on blood biomarkers
By analyzing the GEO database and using machine learning algorithms, the CFD and DKFZp434L192 genes were selected as biomarkers. A nomograph was constructed, which solved the shortcomings of Moyamoya disease prediction and enabled accurate identification of high-risk groups and clinical decision support.
Patent Information
- Application Number
- CN202411613938.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-11-13
AI Technical Summary
The lack of efficient biomarkers for predicting Moyamoya disease in existing technologies leads to a lack of effective decision support for the prediction and treatment of Moyamoya disease.
By analyzing datasets in the GEO database, differentially expressed genes were screened, enrichment analysis was performed using the Metascape database and gene ontology, and biomarkers were identified by combining LASSO and SVM-RFE machine learning algorithms. A nomograph was constructed to predict the risk of Moyamoya disease, and gene expression was verified using RT-qPCR. CFD and DKFZp434L192 genes were identified as biomarkers.
It enables effective prediction of high-risk groups for Moyamoya disease, provides rapid clinical decision support, provides a basis for the prevention and treatment of Moyamoya disease, and improves the accuracy and efficiency of prediction.
Smart Images

Figure CN119786008B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to a clinical decision support method for predicting Moyamoya disease based on the blood biomarkers CFD and DKFZp434L192. Background Technology
[0002] Moyamoya disease (MMD) is a chronic, rare, cerebrovascular occlusive disease characterized by progressive narrowing of the distal internal carotid artery (ICA) and an abnormal vascular network at the base of the brain. It has a poor prognosis and exhibits significant regional and ethnic variability. The occurrence and development of MMD are multifactorial, and its underlying causes and pathogenic mechanisms remain largely unclear.
[0003] The development of molecular biology and second-generation sequencing technologies has made it possible to explore the pathogenesis of diseases on a large scale at the gene and molecular level. Therefore, biomarker screening provides a new direction for disease prediction and treatment. However, currently, there is relatively little focus on using bioinformatics analysis of gene expression to explore differentially expressed genes (DEGs) related to disease occurrence and development, and screening for differentially expressed genes as a method for assessing the risk of Moyamoya disease. Summary of the Invention
[0004] To address the problems of existing technologies, the purpose of this invention is to provide a clinical decision support method for predicting Moyamoya disease based on the blood biomarkers CFD and DKFZp434L192, thus solving the current problem of the lack of efficient and identifiable biomarkers for Moyamoya disease. This invention uses machine learning to screen characteristic genes and analyzes the expression of these genes in the blood of patients to rapidly predict high-risk groups for Moyamoya disease, providing decision support for timely intervention and treatment.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A clinical decision support approach for predicting Moyamoya disease based on blood biomarkers includes the following steps:
[0007] Using CFD and DKFZp434L192 genes as biomarkers, a nomogram was constructed based on the expression of biomarkers in the dataset. Risk scores were calculated to predict the risk of developing Moyamoya disease in the general population. Specifically, differentially expressed genes between Moyamoya disease patients and healthy controls were first identified by analyzing datasets from the GEO database. Further analysis was conducted on the biological pathways and disease types enriched by these differentially expressed genes. Then, LASSO and SVM-RFE machine learning algorithms were used to identify Moyamoya disease-related biomarkers, and the predictive ability of these biomarkers was evaluated using ROC curves. Finally, CFD and DKFZp434L192 genes were selected, and a nomogram was constructed based on their expression in the dataset. Risk scores were calculated to predict the risk of developing Moyamoya disease in the general population, thus identifying high-risk individuals for Moyamoya disease in the general population. Clinical samples were further collected, and RT-qPCR was used to verify the expression of CFD and DKFZp434L192 genes in whole blood from Moyamoya disease patients.
[0008] Furthermore, the CFD and DKFZp434L192 genes were selected as candidate biomarkers through the following steps:
[0009] S1, to identify differentially expressed genes between patients with Moyamoya disease and healthy controls;
[0010] S2 further analyzed the biological pathways and disease types enriched by differentially expressed genes;
[0011] S3, screening for biomarkers to predict Moyamoya disease;
[0012] S4, to verify the differential expression of biomarkers in patients with Moyamoya disease and healthy controls.
[0013] Preferably, in step S1, differentially expressed genes between patients with moyamoya disease and healthy controls are identified by analyzing two moyamoya disease microarray datasets downloaded from the GEO database.
[0014] Preferably, in step S2, the enrichment analysis of differentially expressed genes is further performed using the Metascape database, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes to further analyze the biological pathways and disease types enriched by the differentially expressed genes.
[0015] Preferably, in step S3, biomarkers for patients with Moyamoya disease are screened using LASSO and SVM-RFE machine learning algorithms, and the predictive ability of the biomarkers is evaluated using ROC curves to obtain biomarkers with high predictive ability.
[0016] Preferably, in step S4, RT-qPCR is used to further validate the expression of candidate biomarkers after using an external dataset and extracting RNA from collected clinical blood samples.
[0017] Furthermore, the construction of the biomarker expression nomograph is based on CFD and the expression of DKFZp434L192, and is constructed using the "RMS" package.
[0018] Furthermore, the risk score is in the range of 0-1, and the closer the score is to 1, the higher the risk of developing Moyamoya disease.
[0019] Furthermore, the biomarker expression nomograph is used to predict high-risk groups for Moyamoya disease, providing rapid auxiliary decision support for the prevention and treatment of Moyamoya disease.
[0020] Compared with the prior art, the present invention has the following advantages:
[0021] (1) This invention analyzed two Moyamoya disease microarray datasets downloaded from the GEO database to identify differentially expressed genes between Moyamoya disease patients and healthy controls. It further analyzed the biological pathways and disease types enriched by these differentially expressed genes using the Metascape database, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes. It identified Moyamoya disease-related biomarkers using LASSO and SVM-RFE machine learning algorithms. The predictive ability of the biomarkers was evaluated using ROC curves, and the biomarkers were further validated using datasets and RT-qPCR. Finally, it was determined that the CFD and DKFZp434L192 genes can be used as biomarkers to predict high-risk populations for Moyamoya disease, and these biomarkers showed good predictive performance. A nomograph was constructed based on the expression of these two genes in the datasets to predict high-risk populations for Moyamoya disease in the general population, with good predictive results, thus solving the problem of the lack of efficient and identifiable biomarkers for Moyamoya disease.
[0022] (2) This invention identifies characteristic genes through machine learning, with the goal of clinical decision-making, and further assists clinicians in providing rapid decision support for the prevention and treatment of Moyamoya disease. Attached Figure Description
[0023] Figure 1 Heatmap (A) and volcano plot (B) of differentially expressed genes between the patient group and the control group in the relevant Moyamoya disease datasets GSE157628 and GSE141025.
[0024] Figure 2 The results of differential gene enrichment analysis are shown in the diagrams: A. GO enrichment analysis; B. DO enrichment analysis; C. KEGG enrichment analysis; D. Metascape analysis; E. PPI protein interaction network diagram.
[0025] Figure 3 The process of screening biomarkers in the Moyamoya disease dataset is illustrated by: AB. Lasso regression analysis; CD. SVM-RFE support vector machine recursive feature elimination analysis; E. Intersection of the two algorithms.
[0026] Figure 4 The diagram illustrates the process of validating biomarkers in the validation set. A. Expression levels of the intersecting genes in the merged training set; B. ROC of the intersecting genes in the merged training set; C. Expression levels of the intersecting genes in the validation set; D. ROC of the intersecting genes in the validation set.
[0027] Figure 5 Nonographs, decision curves, clinical impact curves, and calibration curves for candidate biomarkers (A: Nonograph; B: Decision curve; C: Clinical impact curve; D: Calibration curve).
[0028] Figure 6 This is a graph showing the relative expression levels of the target gene using qRT-PCR. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art. The technical terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the invention; unless otherwise specified, all raw materials, reagents, instruments, and equipment used in this invention are commercially available or can be prepared by existing methods.
[0031] Example 1
[0032] Download the GSE157628 and GSE141025 datasets from the GEO database. Use an online tool to screen for differentially expressed genes in Moyamoya disease; genes with p-values < 0.05 and |log2FC| > 1 were identified as differentially expressed genes.
[0033] 1.1 Filtering Moyamoya Disease-Related Datasets Using the GEO Database
[0034] The matrix files for datasets GSE157628 and GSE141025 were obtained from the GEO database (http: / / www.ncbi.nlm.nih.gov / geo / ) created and maintained by the National Center for Biotechnology Information (NCBI). Dataset GSE157628 includes 11 samples from patients with Moyamoya disease and 9 control samples. GSE141025 includes 4 samples from patients with Moyamoya disease and 4 control samples. Probes in each dataset were converted to gene symbols according to their probe annotation files. For multiple probes corresponding to the same gene symbol, the probe average was calculated as the final expression value of that gene. Datasets GSE157628 and GSE141025 were merged into a metadata queue for further ensemble analysis. The "SVA" package in R was used to eliminate batch effects. Background correction, inter-array normalization, and differential expression analysis between Moyamoya disease and control samples were performed using the limma package in R (http: / / www.bioconductor.org / ). Differentially expressed genes were defined as |logFC|>1 and p<0.05.
[0035] 1.2 Expression analysis of differentially expressed genes
[0036] First, hierarchical cluster analysis was performed on the 28 obtained samples to detect the correlation between samples. The limma package in R (http: / / www.bioconductor.org / ) was used for background correction, array normalization, and differential expression analysis between the samples and the control samples. The screening criteria were set as |logFC|>1 and P<0.05, indicating statistical significance. Heatmaps and volcano plots of the differentially expressed genes were then generated using R. (e.g., Figure 1 ).
[0037] 1.3 Gene functional annotation and enrichment analysis
[0038] To elucidate the biological function of Moyamoya disease, packages such as "clusterProfiler", "enrichplot", and "GSEABase" were used to perform functional and pathway enrichment analyses on differentially expressed genes (e.g., Figure 2 ).
[0039] Example 2
[0040] Screening and validation of marker genes
[0041] Candidate diagnostic biomarkers were further screened using the Least Absolute Shrinkage and Selection Operator (LASSO) model. The "glmnet" package in R was used for lasso regression, and the "e1071" package in R was used to construct a support vector machine (SVM) model for screening key genes. A Venn diagram was used to find the genes at the intersection of the two models, and the validation set GSE189993 was applied to validate the differentially expressed genes in the intersection. The "pROC" package in R was used to plot ROC curves to verify the specificity and sensitivity of the candidate differentially expressed genes. Two machine learning algorithms were used to screen diagnostic biomarkers from 76 key genes. LASSO identified 16 meaningful characteristic genes, such as... Figure 3 A and Figure 3 B. Using the SVM algorithm, a total of 8 genes with significant characteristics were identified, such as... Figure 3 C and Figure 3 D. The intersection of the two algorithms yields 6 overlapping genes, such as... Figure 3 E. These were CFD, DKFZp434L192, EN1, MIA, MYOT, and OGN, and their expression in the MMD group was significantly different from that in the control group (P<0.001). Figure 4 A. Their AUCs were 0.873, 0.873, 0.839, 0.895, 0.898, and 0.930, respectively, indicating that the above six genes have the potential to serve as diagnostic biomarkers for Moyamoya disease. Figure 4 B. Subsequently, we used the GSE189993 dataset as a validation set to validate the aforementioned six diagnostic biomarkers. The expression levels of these six genes showed that CFD and DKFZp434L192 were significantly upregulated in the MMD group (P<0.05), such as... Figure 4 C. ROC is an indicator for evaluating diagnostic accuracy and sensitivity. The diagnostic efficacy of the above six biomarkers was tested using ROC curves. The results showed that the AUCs of CFD and DKFZp434L192 were 0.831 and 0.732, respectively, indicating high predictive value. This suggests that these two biomarkers have good diagnostic value. Figure 4 D.
[0042] Example 3
[0043] Construction and evaluation of nomographs for predicting Moyamoya disease
[0044] A nomogram was constructed based on gene expression to predict the risk of Moyamoya disease patients. Figure 5 A). Decision curve analysis (DCA) shows that the "all influencing factors" curve is much higher than the gray line, indicating that the line graph model is more accurate. Figure 5B). To more intuitively evaluate the clinical efficacy of the nomogram model, clinical impact curves based on DCA curves were plotted. The thick solid line representing the high-risk curve is very close to the thick dashed line representing the true positive patient curve, indicating that the nomogram model has good predictive ability. Figure 5 C). The calibration curves indicate that the nomograph model has high accuracy in predicting the risk of moyamoya disease. Figure 5 D). The calibration curves show that the error between the risk of developing Moyamoya disease and the predicted risk is small and statistically significant, indicating that the nomograph model has high accuracy in predicting Moyamoya disease.
[0045] This invention collects blood samples from healthy individuals and patients with Moyamoya disease, and detects the mRNA expression levels of CFD and DKFZp434L192 using RT-qPCR. It was found that the expression levels of these two biomarkers can be detected at the mRNA level using RT-qPCR. Figure 6 ), which were then incorporated into the model as diagnostic biomarkers for Moyamoya disease.
[0046] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A clinical decision support method for predicting Moyamoya disease based on blood biomarkers, characterized in that, Includes the following steps: Using CFD and DKFZp434L192 genes as biomarkers, a nomogram based on biomarker expression was established; the risk score was calculated to predict the probability of developing Moyamoya disease in the general population.
2. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 1, characterized in that, The CFD and DKFZp434L192 genes were selected as candidate biomarkers through the following steps: S1, to identify differentially expressed genes between patients with Moyamoya disease and healthy controls; S2 further analyzed the biological pathways and disease types enriched by differentially expressed genes; S3, screening for biomarkers to predict Moyamoya disease; S4, to verify the differential expression of biomarkers in patients with Moyamoya disease and healthy controls.
3. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 2, characterized in that, In step S1, differentially expressed genes between patients with moyamoya disease and healthy controls are identified by analyzing two moyamoya disease microarray datasets downloaded from the GEO database.
4. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 2, characterized in that, In step S2, the enrichment analysis of differentially expressed genes is further performed using the Metascape database, Gene Ontology, and the Kyoto Encyclopedia of Genes and Genomes to further analyze the biological pathways and disease types enriched by the differentially expressed genes.
5. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 2, characterized in that, In step S3, biomarkers for patients with Moyamoya disease are screened using LASSO and SVM-RFE machine learning algorithms, and the predictive ability of the biomarkers is evaluated using ROC curves.
6. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 2, characterized in that, In step S4, RT-qPCR is used to further validate the expression of candidate biomarkers by using an external dataset and extracting RNA from collected clinical blood samples.
7. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 1, characterized in that, The aforementioned nomograph is based on the representation of CFD and DKFZp434L192 in the dataset and is constructed using the "RMS" package.
8. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 1, characterized in that, The risk score is in the range of 0-1, and the closer the score is to 1, the higher the risk of developing Moyamoya disease.
9. The clinical decision support method for predicting Moyamoya disease based on blood biomarkers according to claim 1, characterized in that, The nomograph is used to predict high-risk groups for Moyamoya disease, providing rapid auxiliary decision support for the prevention and treatment of Moyamoya disease.
Citation Information
Patent Citations
Immunoglobulin and / or toll-like receptor proteins associated with myelogenous haematological proliferative disorders and uses thereof
CN102046660A
Nanoparticle conjugates and uses thereof
US20180133343A1