Neurodegenerative disease classification system based on biological big data
By constructing a classification system for neurodegenerative diseases based on biological big data, and utilizing techniques such as Pearson correlation coefficient, t-test, and random forest classifier, we can identify lncRNAs associated with neurodegenerative diseases, thus solving the problem of insufficient accuracy of existing tools and achieving efficient lncRNA identification and disease diagnosis support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEAST FORESTRY UNIV
- Filing Date
- 2024-06-03
- Publication Date
- 2026-04-14
AI Technical Summary
Existing big data analysis tools lack accuracy and reliability in identifying long non-coding RNAs (lncRNAs) associated with neurodegenerative diseases. Traditional experimental methods are complex and costly, and cannot delve into the complex relationship between lncRNAs and programmed cell death (PCD).
A classification system for neurodegenerative diseases based on biological big data was constructed, including modules for data acquisition, scoring of programmed cell death lncRNAs, correlation assessment, gene optimization, and random forest classifier. Through Pearson correlation coefficient, t-test, GSEA enrichment analysis, and random forest classifier, lncRNAs and mRNAs associated with PCD were identified.
It improves the accuracy and efficiency of lncRNA recognition, providing support for the study of the pathogenesis and treatment strategies of neurodegenerative diseases, and realizing new ideas for early diagnosis and precision treatment.
Smart Images

Figure CN118711670B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics. Specifically, it relates to a classification system for neurodegenerative diseases based on biological big data. Background Technology
[0002] In the biomedical field, neurodegenerative diseases have always been a hot topic and a challenge in research. These diseases, such as Alzheimer's and Parkinson's, severely affect the neurological function of patients, leading to motor, cognitive, and behavioral disorders. Currently, although researchers have conducted extensive research on the molecular mechanisms of neurodegenerative diseases, the exact pathogenesis remains unclear. Programmed cell death (PCD), as a key process of neuronal loss in neurodegenerative diseases, is crucial for understanding disease progression and finding treatment strategies. Currently, programmed cell death includes apoptosis, necroptosis, pyroptosis, ferroptosis, entotic cell death, reticular cell death, lysosome-dependent cell death, autophagy-dependent cell death, intracellular alkalization death, oxygen free radical-induced death, copper death, and disulfide death.
[0003] However, existing technologies have many limitations in identifying long non-coding RNAs (lncRNAs) associated with the neurodegenerative disease PCD. Traditional experimental methods, such as gene chips and quantitative PCR, can detect lncRNA expression, but they are complex, costly, and struggle to comprehensively cover a large number of lncRNAs. Furthermore, these methods often provide only limited information and cannot delve into the complex relationship between lncRNAs and PCD.
[0004] In recent years, with the rapid development of biological big data technology, big data-based analysis methods have shown great potential in the biomedical field. However, existing big data analysis tools still have limitations in identifying lncRNAs related to the neurodegenerative disease PCD. These tools often lack disease-specific optimized algorithms, resulting in insufficient accuracy and reliability of the identification results. Summary of the Invention
[0005] The purpose of this invention is to address the problem of poor accuracy and reliability of existing big data analysis tools in classifying neurodegenerative diseases, and to propose a neurodegenerative disease classification system based on biological big data.
[0006] Classification systems for neurodegenerative diseases based on biological big data include:
[0007] The system includes a data acquisition module, a programmed cell death lncRNA scoring module, a lncRNA-programmed cell death correlation assessment module, a programmed cell death gene optimization module, a random forest classifier module, and a test module.
[0008] The data acquisition module is used to collect data on neurodegenerative diseases and preprocess the collected data to obtain preprocessed lncRNA and mRNA data of neurodegenerative diseases.
[0009] The programmed cell death lncRNA scoring module is used to calculate the Pearson correlation coefficient between each lncRNA and all corresponding mRNAs based on the preprocessed lncRNA and mRNA data output by the data acquisition module; calculate the correlation p-value between each lncRNA and all corresponding mRNAs using a t-test; convert the Pearson correlation coefficient and p-value between each lncRNA and all corresponding mRNAs into a risk assessment score between each lncRNA and all corresponding mRNAs, and sort the risk assessment scores from high to low to obtain a risk score ranking list between each lncRNA and all corresponding mRNAs;
[0010] The lncRNA-programmed cell death correlation assessment module is used to calculate the association strength (ES) between each lncRNA and all corresponding mRNAs based on the risk score ranking list between each lncRNA and all corresponding mRNAs obtained from the programmed cell death lncRNA scoring module. The GSEA functional enrichment analysis method is used to calculate the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs. Based on the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs, a set of lncRNAs significantly associated with programmed cell death is obtained. The set of lncRNAs significantly associated with programmed cell death is stored in BED format.
[0011] The programmed cell death gene optimization module is used to obtain mRNAs that are significantly associated with programmed cell death based on a set of lncRNAs that are significantly associated with programmed cell death, through GREAT cis-regulation and protein interaction network selection.
[0012] The Random Forest Classifier module is used to obtain a trained random forest classifier based on mRNAs that are significantly associated with programmed cell death obtained from the programmed cell death gene optimization module.
[0013] The module to be tested is used to classify mRNA data of the neurodegenerative disease to be tested based on a trained random forest classifier.
[0014] The beneficial effects of this invention are as follows:
[0015] This invention proposes a classification system for neurodegenerative diseases based on biological big data, which is of great significance for overcoming the shortcomings of existing technologies and improving the accuracy and efficiency of identification. This invention aims to rapidly and accurately screen lncRNAs closely related to the neurodegenerative disease PCD from massive gene expression data by utilizing advanced bioinformatics technologies and big data analysis methods, providing strong support for the research on disease pathogenesis and the development of treatment strategies. This invention's PCD-related lncRNA identification system based on biological big data improves the accuracy and efficiency of identification.
[0016] This invention utilizes biological big data to construct a scoring model for programmed cell death lncRNAs, which can accurately identify lncRNAs associated with programmed cell death. The construction method of this scoring model can provide empirical support and methodological reference for related research.
[0017] This invention utilizes lncRNA to identify and optimize mRNAs associated with programmed cell death and learns optimal classification features, thereby improving the accuracy of predicting the classification of neurodegenerative diseases.
[0018] The model of this invention can be applied to identify lncRNAs associated with different clinical features, providing preliminary support for further in-depth analysis of the specific functions of lncRNAs in the disease development process, and also providing new ideas and potential targets for the early diagnosis, precision treatment and drug development of neurodegenerative diseases. Attached Figure Description
[0019] Figure 1 A flowchart illustrating the method for identifying programmed death-related lncRNAs in neurodegenerative diseases based on biological big data;
[0020] Figure 2 A diagram illustrating the recognition of lncRNAs associated with programmed cell death;
[0021] Figure 3 P-value plot of mRNA extracted from Alzheimer's disease processes of cellular senescence, cell death and immune regulation based on homeostasis;
[0022] Figure 4 This is a frequency map of mRNAs extracted from the cellular senescence, cell death, and immune regulation processes in Alzheimer's disease based on homeopathic regulatory relationships;
[0023] Figure 5 P-value plot of mRNA extracted from the cellular senescence, cell death and immune regulation processes in Parkinson's disease based on homeostasis;
[0024] Figure 6This is a frequency map of mRNAs extracted from the cellular senescence, cell death, and immune regulation processes in Parkinson's disease based on homeopathic regulatory relationships.
[0025] Figure 7 This is a protein-protein interaction network diagram;
[0026] Figure 8 A layout diagram for degrees;
[0027] Figure 9 A high-resolution gene sequencing map;
[0028] Figure 10 The image shows the classification results of the trained random forest classifier on the mRNA data of Alzheimer's disease.
[0029] Figure 11 The image shows the classification results of the trained random forest classifier on the mRNA data of Parkinson's disease. Detailed Implementation
[0030] Specific Implementation Method 1: This implementation method's neurodegenerative disease classification system based on biological big data includes:
[0031] The system includes a data acquisition module, a programmed cell death lncRNA scoring module, a lncRNA-programmed cell death correlation assessment module, a programmed cell death gene optimization module, a random forest classifier module, and a test module.
[0032] The data acquisition module is used to collect data on neurodegenerative diseases and preprocess the collected data to obtain preprocessed lncRNA and mRNA data of neurodegenerative diseases, providing a clean and standardized data foundation for subsequent analysis.
[0033] The programmed cell death lncRNA scoring module is used to calculate the Pearson correlation coefficient between each lncRNA and all corresponding mRNAs based on the preprocessed lncRNA and mRNA data output by the data acquisition module; calculate the correlation p-value between each lncRNA and all corresponding mRNAs using a t-test; convert the Pearson correlation coefficient and p-value between each lncRNA and all corresponding mRNAs into a risk assessment score between each lncRNA and all corresponding mRNAs, and sort the risk assessment scores from high to low to obtain a risk score ranking list between each lncRNA and all corresponding mRNAs;
[0034] The lncRNA-programmed cell death correlation assessment module is used to calculate the association strength (ES) between each lncRNA and all corresponding mRNAs based on the risk score ranking list between each lncRNA and all corresponding mRNAs obtained from the programmed cell death lncRNA scoring module. The GSEA functional enrichment analysis method is used to calculate the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs. Based on the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs, a set of lncRNAs significantly associated with programmed cell death is obtained. The set of lncRNAs significantly associated with programmed cell death is stored in BED format.
[0035] The programmed cell death gene optimization module is used to obtain mRNAs that are significantly associated with programmed cell death based on a set of lncRNAs that are significantly associated with programmed cell death, through GREAT cis-regulation and protein interaction network selection.
[0036] The Random Forest Classifier module is used to obtain a trained random forest classifier based on mRNAs that are significantly associated with programmed cell death obtained from the programmed cell death gene optimization module.
[0037] The module to be tested is used to classify mRNA data of the neurodegenerative disease to be tested based on a trained random forest classifier.
[0038] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that the neurodegenerative diseases mentioned are Alzheimer's disease and Parkinson's disease;
[0039] The lncRNA is a long non-coding RNA; the mRNA is a messenger RNA.
[0040] The other steps and parameters are the same as in Specific Implementation Method 1.
[0041] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that: the data acquisition module is used to collect data on neurodegenerative diseases and preprocesses the collected data to obtain preprocessed lncRNA and mRNA data for neurodegenerative diseases, providing a clean and standardized data foundation for subsequent analysis; the specific process is as follows:
[0042] Download data on Alzheimer's disease and Parkinson's disease from the GEO (Gene Expression Omnibus) database;
[0043] The Alzheimer's disease data were preprocessed sequentially by gene name conversion, gene expression level regulation, and quantile standardization to obtain preprocessed lncRNA and mRNA data for Alzheimer's disease.
[0044] The Parkinson's disease data were preprocessed sequentially with gene name conversion, gene expression level regulation, and quantile standardization to obtain preprocessed lncRNA and mRNA data for Parkinson's disease.
[0045] Finally, high-quality and standardized lncRNA and mRNA expression datasets were obtained for subsequent analysis.
[0046] Other steps and parameters are the same as in specific implementation method one or two.
[0047] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that: the programmed cell death lncRNA scoring module is used to calculate the Pearson correlation coefficient between each lncRNA and all corresponding mRNAs based on the preprocessed lncRNA and mRNA data output by the data acquisition module; calculate the correlation p-value between each lncRNA and all corresponding mRNAs using a t-test; convert the Pearson correlation coefficient and p-value between each lncRNA and all corresponding mRNAs into a risk assessment score between each lncRNA and all corresponding mRNAs, and sort the risk assessment scores from high to low to obtain a risk score ranking list between each lncRNA and all corresponding mRNAs; the specific process is as follows:
[0048] The expression for the risk assessment score RS is:
[0049] RS(i,j)=-log 10 (P i,j )×sign(Corr i,j )
[0050] Where i and j represent lncRNA and mRNA, respectively;
[0051] Corr i,j Let i be the Pearson correlation coefficient between i and j;
[0052] sign represents the sign of the correlation coefficient;
[0053] Corr i,j A value greater than 0 indicates a positive correlation, sign(Corr) i,j The value is 1;
[0054] Corr i,j A value less than 0 indicates a negative correlation, sign(Corr) i,j The value is -1;
[0055] Corr i,j If the value is 0, convert 0 to 0.00001, sign(Corr) i,j The value is 1;
[0056] P i,j The correlation p-value between i and j;
[0057] RS(i,j) is the risk assessment score between i and j.
[0058] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0059] Specific Implementation Method 5: This implementation method differs from Specific Implementation Methods 1 to 4 in that: the correlation assessment module between lncRNA and programmed cell death is used to calculate the association strength ES between the mRNA of each lncRNA and the corresponding mRNA based on the risk score ranking list between each lncRNA and all corresponding mRNAs obtained by the programmed cell death lncRNA scoring module, and to obtain a set of lncRNAs that are significantly associated with programmed cell death based on the association strength ES between the mRNA of each lncRNA and the programmed cell death characteristic mRNA.
[0060] The set of lncRNAs significantly associated with programmed cell death was stored in BED format;
[0061] The specific process is as follows:
[0062] The set of mRNAs characteristic of programmed cell death was obtained through literature mining.
[0063] The mRNA sequence corresponding to each lncRNA obtained from the programmed cell death lncRNA scoring module was fed into the GSEA tool. The set of programmed cell death feature mRNAs was fed into the GSEA tool as the target gene set. The ES score of each lncRNA was calculated using the GSEA tool.
[0064] ES is a measure of the number of genes in a gene set that is ranked at the head or tail of the gene set. For each lncRNA, ES is determined by calculating the cumulative distribution of gene set members in a ranking list of risk scores.
[0065] Let's take a student's exam scores from elementary school to high school as an example. Imagine a student as a lncRNA, and the ranking of all their scores represents the ranking of their corresponding mRNAs. Let's assume all humanities subjects represent the set of programmed cell death mRNAs. If a student's humanities subjects rank highly, a large ES value indicates that the student excels in the humanities. This method is used to select students who are good at the humanities.
[0066] The statistical significance of the ES is determined by using a permutation test, which generates a null hypothesis distribution by randomly shuffling the labels of the gene set members and recalculating the ES.
[0067] Calculate the nominal P-value, which is the position of the observed ES in the null hypothesis distribution;
[0068] Adjust the P-value to control the false detection rate (FDR) to obtain the corrected P-value;
[0069] A high ES value indicates a positive correlation between lncRNA and the set of programmed cell death genes, while a low ES value indicates a negative correlation.
[0070] The statistical significance of ES was assessed using p-value and FDR;
[0071] For each lncRNA, output the ES, P-value and FDR-corrected P-value of the lncRNA and the PCD gene set; screen out statistically significant lncRNAs that play a role in regulating programmed cell death.
[0072] The other steps and parameters are the same as those in specific implementation methods one through four.
[0073] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that: the programmed cell death gene optimization module is used to obtain mRNAs significantly associated with programmed cell death based on a set of lncRNAs significantly associated with programmed cell death, through GREAT cistropic regulation and protein interaction network selection; the specific process is as follows:
[0074] 1) Input the set of lncRNAs in BED format that are significantly associated with programmed cell death into GREAT software for homeopathic regulatory function analysis to obtain the homeopathic regulatory function analysis results. In GREAT software, select the cell senescence, cell death and immune regulation processes in the homeopathic regulatory function analysis results and extract the mRNAs in the cell senescence, cell death and immune regulation processes.
[0075] 2) Frequency statistics analysis was performed on each extracted mRNA to determine the frequency distribution of each mRNA, and the top 10% of mRNAs were extracted;
[0076] 3) Differential expression analysis was performed on the top 10% of extracted mRNAs using the limma software package to screen out mRNAs with p ≤ 0.05 (significant changes in expression under pathological conditions);
[0077] 4) Input the mRNAs selected in 3) into the STRING database, extract the mRNA relationship pairs that are directly linked to the mRNAs selected in 3) from the STRING database, and use the relationship pairs to construct a concise and highly correlated protein interaction network.
[0078] The database itself provides the connection relationships between mRNAs. For example, if the input hypothesis is mRNA1, and mRNA2 and mRNA3 are connected to mRNA1, then I will extract the relationships between mRNA1 and mRNA2, and mRNA1 and mRNA3 to construct the network.
[0079] 5) The mRNA with the highest degree in the protein-protein interaction network was identified through degree centrality analysis, which is a mRNA that is significantly associated with programmed cell death.
[0080] Degree centrality analysis identified the most degree mRNAs in the protein-protein interaction network, namely hub mRNAs that are connected to multiple other nodes. Hub mRNAs play a key role in programmed cell death due to their central position in the network. They are mRNAs that are significantly associated with programmed cell death.
[0081] The other steps and parameters are the same as those in specific implementation methods one through five.
[0082] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that: the random forest classifier module is used to obtain a trained random forest classifier based on mRNAs significantly associated with programmed cell death obtained from the programmed cell death gene optimization module; the specific process is as follows:
[0083] The mRNAs significantly associated with programmed cell death obtained by the programmed cell death gene optimization module are used as input to the random forest classifier, and the prediction results of disease and normal samples (the samples collected by the acquisition module are labeled as disease or normal samples) are used as output to train the random forest classifier until convergence is obtained.
[0084] Each patient has the same number of mRNAs, but the expression levels of the mRNAs are different. I use each person's mRNAs to train a classifier, which determines whether a sample is diseased or normal.
[0085] I have 100 students, 50 from Primary School 1 and 50 from Primary School 2. I selected 10 subjects with significant differences between Primary School 1 and Primary School 2. Then, I trained a random forest model using these 10 subjects. The final result indicates whether each student belongs to Primary School 1 or Primary School 2. Next, I input the 10 subjects to be tested into the trained random forest model, and the final output is whether each of the 10 subjects belongs to Primary School 1 or Primary School 2. These 10 subjects are the final mRNA samples; Primary School 1 and Primary School 2 indicate disease or normality.
[0086] The neurodegenerative disease prediction model is based on U decision trees, which are used to improve the accuracy and robustness of classification. U is a positive integer. Each decision tree is trained on a random subset of the dataset, which reduces the risk of overfitting and increases the model's generalization ability.
[0087] The other steps and parameters are the same as those in specific implementation methods one through six.
[0088] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One through Seven in that: the module to be tested is used to classify the mRNA data of the neurodegenerative disease to be tested based on a trained random forest classifier; the specific process is as follows:
[0089] The mRNA data of the neurodegenerative disease to be tested is input into a trained random forest classifier. The trained random forest classifier outputs whether the mRNA data of the neurodegenerative disease to be tested is a disease sample or a normal sample.
[0090] We employed 10x cross-validation to evaluate the model's performance, ensuring our classifier maintained consistent prediction accuracy across different subsets of the training set. Finally, we evaluated the model's final performance using independent test sets. Additionally, we downloaded a neurodegenerative disease dataset from the GEO database—one not previously used in our analysis—as a test set, ensuring the validity of our evaluation results and the model's good generalization ability. The key performance metric was the area under the curve (AUC), which measures the model's ability to distinguish between diseased and healthy samples. Our goal was to achieve a high AUC value, indicating excellent classification performance.
[0091] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.
[0092] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A classification system for neurodegenerative diseases based on biological big data, characterized in that: The system includes: The system includes a data acquisition module, a programmed cell death lncRNA scoring module, a lncRNA-programmed cell death correlation assessment module, a programmed cell death gene optimization module, a random forest classifier module, and a test module. The data acquisition module is used to collect data on neurodegenerative diseases and preprocess the collected data to obtain preprocessed lncRNA and mRNA data of neurodegenerative diseases. The programmed cell death lncRNA scoring module is used to calculate the Pearson correlation coefficient between each lncRNA and all corresponding mRNAs based on the preprocessed lncRNA and mRNA data output by the data acquisition module; calculate the correlation p-value between each lncRNA and all corresponding mRNAs using a t-test; convert the Pearson correlation coefficient and p-value between each lncRNA and all corresponding mRNAs into a risk assessment score between each lncRNA and all corresponding mRNAs, and sort the risk assessment scores from high to low to obtain a risk score ranking list between each lncRNA and all corresponding mRNAs; The lncRNA-programmed cell death correlation assessment module is used to calculate the association strength (ES) between each lncRNA and all corresponding mRNAs based on the risk score ranking list between each lncRNA and all corresponding mRNAs obtained from the programmed cell death lncRNA scoring module. The GSEA functional enrichment analysis method is used to calculate the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs. Based on the association strength (ES) between each lncRNA and programmed cell death characteristic mRNAs, a set of lncRNAs significantly associated with programmed cell death is obtained. The set of lncRNAs significantly associated with programmed cell death is stored in BED format. The programmed cell death gene optimization module is used to obtain mRNAs that are significantly associated with programmed cell death based on a set of lncRNAs that are significantly associated with programmed cell death, through GREAT cis-regulation and protein interaction network selection. The Random Forest Classifier module is used to obtain a trained random forest classifier based on mRNAs that are significantly associated with programmed cell death obtained from the programmed cell death gene optimization module. The module to be tested is used to classify mRNA data of the neurodegenerative disease to be tested based on a trained random forest classifier.
2. The neurodegenerative disease classification system based on biological big data according to claim 1, characterized in that: The neurodegenerative diseases mentioned are Alzheimer's disease and Parkinson's disease; The lncRNA is a long non-coding RNA; the mRNA is a messenger RNA.
3. The classification system for neurodegenerative diseases based on biological big data according to claim 2, characterized in that: The data acquisition module is used to collect data on neurodegenerative diseases and preprocess the collected data to obtain preprocessed lncRNA and mRNA data for neurodegenerative diseases; the specific process is as follows: Download data on Alzheimer's disease and Parkinson's disease from the GEO database; The Alzheimer's disease data were preprocessed sequentially by gene name conversion, gene expression level regulation and quantile standardization to obtain the preprocessed lncRNA and mRNA data of Alzheimer's disease data; The Parkinson's disease data were preprocessed sequentially with gene name conversion, gene expression level regulation, and quantile standardization to obtain preprocessed lncRNA and mRNA data for Parkinson's disease.
4. The neurodegenerative disease classification system based on biological big data according to claim 3, characterized in that: The programmed cell death lncRNA scoring module is used to calculate the Pearson correlation coefficient between each lncRNA and all corresponding mRNAs based on the preprocessed lncRNA and mRNA data output by the data acquisition module. The correlation p-value between each lncRNA and all corresponding mRNAs was calculated using a t-test. The Pearson correlation coefficient and p-value between each lncRNA and all corresponding mRNAs were then converted into a risk assessment score for each lncRNA and all corresponding mRNAs. These risk assessment scores were then sorted from highest to lowest to obtain a risk score ranking list for each lncRNA and all corresponding mRNAs. The specific process is as follows: The expression for the risk assessment score RS is: RS(i,j)=-log 10 (P i,j )×sign(Corr i,j ) Where i and j represent lncRNA and mRNA, respectively; Corr i,j Let i be the Pearson correlation coefficient between i and j; sign represents the sign of the correlation coefficient; Corr i,j A value greater than 0 indicates a positive correlation, sign(Corr) i,j The value is 1; Corr i,j A value less than 0 indicates a negative correlation, sign(Corr) i,j The value is -1; Corr i,j If the value is 0, convert 0 to 0.00001, sign(Corr) i,j The value is 1; P i,j The correlation p-value between i and j; RS(i,j) is the risk assessment score between i and j.
5. The neurodegenerative disease classification system based on biological big data according to claim 4, characterized in that: The correlation assessment module between lncRNA and programmed cell death is used to calculate the association strength ES between each lncRNA and all corresponding mRNAs based on the risk score ranking list between each lncRNA and all corresponding mRNAs obtained by the programmed cell death lncRNA scoring module. The GSEA functional enrichment analysis method is used to calculate the association strength ES between each lncRNA and programmed cell death characteristic mRNAs. Based on the association strength ES between each lncRNA and programmed cell death characteristic mRNAs, a set of lncRNAs that are significantly associated with programmed cell death is obtained. The set of lncRNAs significantly associated with programmed cell death was stored in BED format; The specific process is as follows: Obtain a set of mRNAs that characterize programmed cell death; The mRNA sequence corresponding to each lncRNA obtained from the programmed cell death lncRNA scoring module was fed into the GSEA tool. The set of programmed cell death characteristic mRNAs was fed into the GSEA tool as the target gene set. The ES score of each lncRNA was calculated using the GSEA tool.
6. The neurodegenerative disease classification system based on biological big data according to claim 5, characterized in that: The programmed cell death gene optimization module is used to obtain mRNAs significantly associated with programmed cell death based on a set of lncRNAs significantly associated with programmed cell death, through GREAT cistropic regulation and protein-protein interaction network selection; the specific process is as follows: 1) Input the set of lncRNAs in BED format that are significantly associated with programmed cell death into GREAT software for homeopathic regulatory function analysis to obtain the homeopathic regulatory function analysis results. In GREAT software, select the cell senescence, cell death and immune regulation processes in the homeopathic regulatory function analysis results and extract the mRNAs in the cell senescence, cell death and immune regulation processes. 2) Frequency statistics analysis was performed on each extracted mRNA to determine the frequency distribution of each mRNA, and the top 10% of mRNAs were extracted; 3) Differential expression analysis was performed on the top 10% of extracted mRNAs using the limma software package, and mRNAs with p ≤ 0.05 were screened out; 4) Input the mRNAs selected in 3) into the STRING database, extract the mRNA relationship pairs that are directly linked to the mRNAs selected in 3) from the STRING database, and use the relationship pairs to construct a protein interaction network; 5) The mRNA with the highest degree in the protein-protein interaction network was identified through degree centrality analysis, which is a mRNA that is significantly associated with programmed cell death.
7. The classification system for neurodegenerative diseases based on biological big data according to claim 6, characterized in that: The random forest classifier module is used to obtain a trained random forest classifier based on mRNAs significantly associated with programmed cell death obtained from the programmed cell death gene optimization module; the specific process is as follows: The mRNAs significantly associated with programmed cell death obtained from the programmed cell death gene optimization module are used as input to the random forest classifier, and the prediction results of disease and normal samples are used as output to train the random forest classifier until convergence is obtained.
8. The classification system for neurodegenerative diseases based on biological big data according to claim 7, characterized in that: The module under test is used to classify mRNA data of the neurodegenerative disease under test based on a trained random forest classifier; the specific process is as follows: The mRNA data of the neurodegenerative disease to be tested is input into a trained random forest classifier. The trained random forest classifier outputs whether the mRNA data of the neurodegenerative disease to be tested is a disease sample or a normal sample.
Citation Information
Patent Citations
Colon adenocarcinoma pyroptosis related lncRNA prognosis model and construction method and application thereof
CN114627970A
Methods and Compositions for Treating Diseases Associated with Exhausted T Cells
US20210033595A1