Screening method of bipolar affective disorder and epilepsy co-pathogenic biomarkers

Through the combination of gene chip data analysis and machine learning algorithms, biomarkers of co-pathogenicity of bipolar disorder and epilepsy were screened out, solving the problem of single analysis methods and insufficient model construction in the existing technology, achieving more efficient biomarker screening and early diagnosis, and providing theoretical support for the treatment of bipolar disorder and epilepsy.

CN120544686APending Publication Date: 2025-08-26CAPITAL UNIVERSITY OF MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510628866.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing gene screening methods have problems with single analytical methods and insufficient model construction for the study of co-pathogenic mechanisms of bipolar disorder and epilepsy, and it is difficult to fully reveal complex genetic mechanisms, and traditional methods are difficult to capture the complex interaction relationship between genes.

Method used

Gene chip data analysis combined with machine learning algorithms are used to screen differential gene screening, weighted gene co-expression network analysis, protein interaction network construction and a combination of multiple machine learning algorithms to screen out bipolar disorder and epilepsy co-pathogenic biomarkers.

Benefits of technology

It improves the accuracy of biomarker screening, enhances the biological credibility of the results, provides a theoretical basis for the early diagnosis and treatment of bipolar disorder and epilepsy, and develops targets for comorbidity diagnosis kits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544686A_ABST
    Figure CN120544686A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics and medical diagnosis, in particular to a method for screening bipolar affective disorder and epilepsy co-pathogenic biomarkers, which comprises the following steps of: performing differential expression analysis on known gene chip expression profile data and finding out a significant correlation module by using a WGCNA algorithm so as to determine key differential genes; finding out hub genes in the core differential genes by using a PPI network, and screening out candidate core pathogenic genes of the bipolar affective disorder and the epilepsy by taking the result intersection of three machine learning methods of LASSO regression, support vector machine recursive feature elimination and random forest; the screened marker can be used for prediction and early prediction of bipolar affective disorder and epilepsy, early screening can be better completed, and a new perspective is provided for molecular mechanism research and early diagnosis of bipolar affective disorder and epilepsy co-pathopoiesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biometric identification technology, and specifically to a method for screening biomarkers for the co-causation of bipolar disorder and epilepsy. Furthermore, the marker screening results of the present invention provide a theoretical basis for conversion detection and also provide theoretical support for the clinical use of bipolar disorder and epilepsy medication and the evaluation of drug efficacy. Background Art

[0002] Bipolar disorder is a common mental disorder characterized by recurring mood swings (Reference [1]). Epilepsy is a disease characterized by spontaneous recurring seizures that affects people of all ages, races, social classes, and geographic regions (Reference [2]). These two diseases are common neurological diseases with some overlap in clinical manifestations and pathological mechanisms. Studies have suggested that there may be common pathogenic genes between the two. Traditional genetic screening methods usually rely on a single data type or simple statistical analysis, which makes it difficult to fully reveal the complex genetic mechanisms (Reference [3]).

[0003] In recent years, the development of gene chip technology has made it possible to obtain large-scale gene expression data, and machine learning algorithms have shown significant advantages in complex data analysis and pattern recognition.

[0004] However, the current diagnosis and treatment of these two diseases mainly rely on clinical symptoms and drug trials, but the research on the pathogenic mechanisms of these two diseases still has the following technical defects:

[0005] 1. Single analysis method:

[0006] Only using chip data analysis methods to screen candidate biomarkers without combining machine learning algorithms: In the study by Xueying Li et al., we found that they identified candidate biomarkers by performing differential analysis and weighted gene co-expression network analysis (WGCNA) on the obtained epilepsy dataset. The method used was simple and difficult to capture the complex interactions between genes. However, this study combined gene chip data analysis with machine learning algorithms to identify co-pathogenic biomarkers of bipolar disorder and epilepsy, achieving multi-dimensional data integration and reducing false positives; (Reference [4])

[0007] 2. Model construction:

[0008] Currently, there is no research on methods for combining gene chip data analysis with machine learning algorithms to screen for co-causative genes of bipolar disorder and epilepsy. Single disease models cannot capture the specific molecular characteristics of co-morbidity. This invention fills a gap in this field.

[0009] Technological breakthroughs:

[0010] To address the above-mentioned shortcomings, the present invention proposes for the first time a solution called "a method for screening biomarkers of co-pathogenicity for bipolar disorder and epilepsy". Based on gene chip data analysis and machine learning algorithms, it screens the co-pathogenic genes of bipolar disorder and epilepsy for the first time, providing a basis for the precise treatment of the two diseases and in-depth research on the pathogenic mechanisms of the two diseases. Summary of the Invention

[0011] In order to solve the above problems, the purpose of the present invention is to provide a method for screening biomarkers of co-pathogenesis of bipolar disorder and epilepsy. The markers can be used for the prediction and early prediction of bipolar disorder and epilepsy, which can better complete early screening and provide a new perspective for the study of the molecular mechanism of co-pathogenesis of bipolar disorder and epilepsy and early diagnosis.

[0012] Technical Solution

[0013] That is, the present invention includes the following aspects.

[0014] Step 1. Obtain the bipolar disorder gene chip dataset and the epilepsy gene chip dataset from the GEO database of the NCBI website (https: / / www.ncbi.nlm.nih.gov / ) and standardize them using the R packages “GEOquery” and “dplyr” in the R language (version 4.4.2), respectively.

[0015] Step 2. The bipolar disorder dataset and the epilepsy dataset were screened for differentially expressed genes using the R package "limma" in the R language. The screening conditions for bipolar disorder differentially expressed genes were |LogFC|>0 and p.value<0.05, and the screening conditions for epilepsy differentially expressed genes were |LogFC|>1 and p.value<0.05. Volcano plots were drawn for the bipolar disorder and epilepsy datasets respectively.

[0016] Step 3. For the bipolar disorder dataset and the epilepsy dataset, weighted gene co-expression network analysis (WGCNA) was performed using the R package "WGCNA" in the R language. The genes with the top 50% of the variance were selected to construct a scale-free co-expression network. The dynamic cutting tree algorithm was used to identify co-expression modules, and the correlation between the module eigengenes and the phenotype was calculated. A heat map of the correlation between the module and the phenotype was drawn, and significantly correlated modules (|r| ≥ 0.5, p < 0.01) were screened.

[0017] Step 4. Intersect the bipolar disorder differentially expressed genes and epilepsy differentially expressed genes obtained in step 2 with the bipolar disorder significantly associated module genes and epilepsy significantly associated module genes obtained in step 3 through a Venn diagram to preliminarily screen out key differentially expressed genes;

[0018] Step 5. The screened core differentially expressed genes were analyzed using the STRING database (https: / / string-db.org / ) to construct a protein-protein interaction network (PPI). A visual protein-protein interaction network diagram was constructed using Cytoscape 3.10.3 software and the MCODE plug-in (parameter settings K-Core = 2, Node Score Cutoff = 0.2, Degree Cutoff = 2, Max Depth = 100). The genes included in all modules calculated by MCODE were then selected as key genes.

[0019] Step 6. The R packages "glmnet," "e1071," and "randomForest" were used to perform 5-10-fold LASSO regression analysis on the key genes in the bipolar disorder dataset. 5-fold support vector machine recursive feature elimination (SVM-RFE) screening was performed. The random forest algorithm was used to set the number of decision trees to 500 to perform random forest analysis to select genes with gene importance greater than 0.7. Subsequently, a Venn diagram was drawn for the results of the three machine learning methods to obtain the core genes of bipolar disorder.

[0020] Step 7. Using the R packages "glmnet," "e1071," and "randomForest," we performed 5-10-fold LASSO regression analysis on the key genes in the epilepsy dataset, followed by 5-fold support vector machine recursive feature elimination (SVM-RFE) screening. We then used the random forest algorithm to set the number of decision trees to 500 and perform random forest analysis to select genes with a gene importance greater than 0.7. We then plotted a Venn diagram of the three machine learning results to obtain the core epilepsy genes.

[0021] Step 8. Take the intersection of the core genes for bipolar disorder and the core genes for epilepsy obtained in steps 6 and 7, draw a Venn diagram, and use the obtained intersection genes as candidate core pathogenic genes for bipolar disorder and epilepsy.

[0022] Beneficial effects

[0023] 1. Traditional differential gene analysis methods are usually based on statistical hypothesis testing and may not fully capture the complex relationships between genes. Machine learning algorithms (such as LASSO regression, support vector machine recursive feature elimination, random forest, etc.) can process high-dimensional data and improve screening accuracy through multi-algorithm fusion.

[0024] 2. Enhance the biological credibility of the results through PPI network.

[0025] 3. Provide targets for the development of diagnostic kits for comorbidities. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to make the above-mentioned and / or other purposes, features, advantages and examples of the present invention more obvious and easy to understand, the following is a brief introduction to the drawings required for use in the specific embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 Flowchart for the implementation of the present invention.

[0028] Figure 2 This is a volcano plot of the bipolar disorder dataset in an embodiment of the present invention.

[0029] Figure 3 This is a volcano plot of the epilepsy dataset in an embodiment of the present invention.

[0030] Figure 4 This is a heat map of the correlation between the bipolar disorder dataset module and phenotype in an embodiment of the present invention.

[0031] Figure 5 This is a heat map of the correlation between the epilepsy dataset module and phenotype in an embodiment of the present invention.

[0032] Figure 6 This is a Venn diagram of the differentially expressed genes between bipolar disorder and epilepsy and the modules significantly associated with bipolar disorder and epilepsy in the embodiment of the present invention.

[0033] Figure 7 A protein interaction network diagram was constructed for the core differentially expressed genes in the embodiments of the present invention.

[0034] Figure 8 This is a diagram of the protein interaction network module screened by the MCODE plug-in in an embodiment of the present invention.

[0035] Figure 9 3 is a cross-validation error diagram of the LASSO logistic regression coefficient of the bipolar disorder dataset in an embodiment of the present invention.

[0036] Figure 10 is a LASSO regression coefficient penalty graph of a bipolar disorder dataset in an embodiment of the present invention.

[0037] Figure 11 3 is a cross-validation root mean square accuracy graph of the bipolar disorder dataset in an embodiment of the present invention.

[0038] Figure 12 3 is a cross-validation root mean square error rate graph of the bipolar disorder dataset in an embodiment of the present invention.

[0039] Figure 13 is a random forest error curve graph of the bipolar disorder dataset in an embodiment of the present invention.

[0040] Figure 14 is a random forest feature importance graph of the bipolar disorder dataset in an embodiment of the present invention.

[0041] Figure 15 4 is a Venn diagram of the core genes screened by three machine learning algorithms for a bipolar disorder dataset in an embodiment of the present invention.

[0042] Figure 16 is a cross-validation error graph of the LASSO logistic regression coefficient of the epilepsy dataset in an embodiment of the present invention.

[0043] Figure 17 is a LASSO regression coefficient penalty graph of the epilepsy dataset in an embodiment of the present invention.

[0044] Figure 18 3 is a cross-validation root mean square accuracy graph of the epilepsy dataset in an embodiment of the present invention.

[0045] Figure 19 3 is a cross-validation root mean square error rate graph of the epilepsy dataset in an embodiment of the present invention.

[0046] Figure 20 is a random forest error curve diagram of the epilepsy dataset in an embodiment of the present invention.

[0047] Figure 21 is a random forest feature importance graph of the epilepsy dataset in an embodiment of the present invention.

[0048] Figure 22 1 is a Venn diagram of the epilepsy dataset in the embodiment of the present invention using three machine learning algorithms to screen core genes.

[0049] Figure 23 It is a Venn diagram for screening candidate core therapeutic genes for bipolar disorder and epilepsy core genes in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The preferred embodiments of the present invention are described above in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention. Any technical solution that meets the key features described in the claims of the present invention shall fall within the scope of protection of the present invention.

[0051] The present invention is described in detail below

[0052] like Figure 1 :Specific implementation process of the present invention

[0053] Example 1:

[0054] Step 1. Obtain the bipolar disorder gene chip dataset GSE5389 and the epilepsy gene chip dataset GSE28674 from the GEO database of the NCBI website (https: / / www.ncbi.nlm.nih.gov / ). Standardize them using the R packages “GEOquery” and “dplyr” in the R language (version 4.4.2), respectively, to obtain the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674.

[0055] Step 2. The bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674 were screened for differentially expressed genes using the R package "limma" in R language. The screening conditions for bipolar disorder differentially expressed genes were |LogFC|>0 and p.value<0.05, resulting in 2551 differentially expressed genes, including 1481 up-regulated genes and 1070 down-regulated genes. The screening conditions for epilepsy differentially expressed genes were |LogFC|>1 and p.value<0.05, resulting in 980 differentially expressed genes, including 750 up-regulated genes and 230 down-regulated genes. Volcano plots were drawn for the bipolar disorder and epilepsy datasets, as shown in Figure 2. Figure 2 Figure 3 .

[0056] Step 3. For the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674, weighted gene co-expression network analysis (WGCNA) was performed using the R package "WGCNA" in the R language. The genes with the top 50% variance were selected to construct a scale-free co-expression network. The dynamic cutting tree algorithm was used to identify the co-expression modules. The correlation between the module feature genes and the phenotype was calculated, and the module-phenotype correlation heat map was drawn as shown in Figure 3. Figure 4 Figure 5 , screened significantly correlated modules (|r| ≥ 0.5, p < 0.01), selected MEpink, MEbrown and MEred modules from the weighted gene co-expression network analysis (WGCNA) results of the bipolar disorder dataset GSE5389, which included a total of 1818 genes, and selected MEpink module from the weighted gene co-expression network analysis (WGCNA) results of the epilepsy dataset GSE28674, which included a total of 919 genes.

[0057] Step 4. The bipolar disorder differentially expressed genes and epilepsy differentially expressed genes obtained in step 2 were intersected with the bipolar disorder significantly associated module genes and epilepsy significantly associated module genes obtained in step 3 through a Venn diagram, and 113 key differentially expressed genes were initially screened out. Figure 6 .

[0058] Step 5. The screened core differentially expressed genes were analyzed using the STRING database (https: / / string-db.org / ) to construct a protein interaction network (PPI). Cytoscape 3.10.3 software and the MCODE plug-in (parameter settings K-Core = 2, Node Score Cutoff = 0.2, Degree Cutoff = 2, Max Depth = 100) were used to construct a visual protein interaction network diagram as shown in the figure. Figure 7 Figure 8 , then 14 genes included in all modules calculated by MCODE were selected as key genes.

[0059] Step 6. Use R packages "glmnet", "e1071" and "randomForest" to perform 5-fold LASSO regression analysis on the key genes in the bipolar disorder dataset. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 9 Figure 10 Four genes were obtained, and 5-fold support vector machine recursive feature elimination (SVM-RFE) was used to screen nonlinear relationships based on kernel functions to evaluate the contribution of feature classification boundaries. Figure 11 Figure 12 9 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 5 genes with gene importance greater than 0.7 were selected. Figure 13 Figure 14 , and then draw a Venn diagram for the results of the three machine learning methods, and get 6 core genes as follows Figure 15 .

[0060] Step 7. Perform 5-fold LASSO regression analysis on the key genes in the epilepsy dataset using R packages “glmnet”, “e1071”, and “randomForest”. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 16 Figure 17 6 genes were obtained, and 5-fold support vector machine recursive feature elimination was performed

[0061] (SVM-RFE) screening is based on kernel function to process nonlinear relationships and evaluate the contribution of feature classification boundaries. Figure 18 Figure 199 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 6 genes with gene importance greater than 0.7 were selected. Figure 20 Figure 21 , and then draw a Venn diagram for the results of the three machine learning, and get 8 core genes as follows Figure 22 .

[0062] Step 8. Take the intersection of the bipolar disorder core genes and epilepsy core genes obtained in steps 6 and 7, and draw a Venn diagram as shown in Figure 23 , and obtained two genes "RGS4" and "GABRA1" as candidate core pathogenic genes for bipolar disorder and epilepsy.

[0063] Lasso regression principle:

[0064] (1) By adding the L1 regularization term to the loss function, the model parameters are made as sparse as possible, thereby achieving feature selection and preventing overfitting.

[0065] (2) The objective function of Lasso regression can be expressed as:

[0066] L(w)=(y-Xw)2+λ||w||1

[0067] y: target variable;

[0068] X: feature matrix;

[0069] w: model parameters;

[0070] λ: regularization parameter;

[0071] ||w||1: 1-norm of w, that is, the sum of the absolute values ​​of all elements.

[0072] (3) Since the L1 regularization term is not differentiable at zero, Lasso regression is usually optimized using the coordinate descent method. This method only searches for the minimum value of the loss function on the current coordinate axis in each iteration without calculating the gradient of the function, which is more efficient.

[0073] (4) Application of Lasso regression:

[0074] Lasso regression has significant advantages in processing high-dimensional data, especially when the number of features is much larger than the number of samples. It can automatically select important features, simplify the model and improve interpretability.

[0075] The core principle of Support Vector Machine Recursive Feature Elimination (SVM-RFE):

[0076] (1) The principle of support vector machine recursive feature elimination (SVM-RFE) for screening genes is based on its classification ability. It maximizes the interval between samples of different categories by finding an optimal hyperplane, which is called the maximum margin hyperplane.

[0077] (2) During the SVM-RFE model training process, the contribution of each feature to the classification result will be evaluated. Features with higher importance have a greater impact on the classification result, while features with lower importance are considered to contribute less to the classification result.

[0078] (3) Application of SVM-RFE:

[0079] Suitable for high-dimensional small sample data; robust to noise and overfitting; can be combined with kernel functions to handle nonlinear relationships.

[0080] Random Forest Principle:

[0081] (1) The principle of random forest gene screening is to evaluate the influence of each gene on the classification results by constructing multiple decision trees, thereby determining which genes have a significant impact on the classification results. The random forest algorithm implements gene screening through the following steps:

[0082] Step 1. Build a decision tree: A random forest is composed of multiple decision trees. During the construction of each decision tree, a portion of samples and features are randomly selected from the original dataset for training.

[0083] Step 2. Feature selection: When splitting each decision tree node, a portion of features are randomly selected to select the best features. This process increases the diversity of the model.

[0084] Step 3. Voting or averaging: For classification tasks, random forests determine the final result by majority voting; for regression tasks, the final result is determined by calculating the average of the prediction results of all trees.

[0085] Step 4. Feature Importance Assessment: Random forests can assess the importance of each feature. This is typically done by calculating the contribution of each feature in each tree and then taking the average. Contribution metrics include the Gini index and the out-of-bag error rate.

[0086] (2) In gene screening tasks, random forests can be used to assess the importance of each gene. By training on a dataset, random forests can measure the contribution of each gene to the prediction model. These metrics can be used to rank the importance of genes, thereby selecting the most influential candidate genes as potential biomarkers or therapeutic targets.

[0087] (3) Application of Random Forest:

[0088] Applicable to high-dimensional nonlinear data; strong resistance to overfitting; can evaluate interactions (gene combinations).

[0089] Example 2:

[0090] Step 1. Obtain the bipolar disorder gene chip dataset GSE5389 and the epilepsy gene chip dataset GSE28674 from the GEO database of the NCBI website (https: / / www.ncbi.nlm.nih.gov / ). Standardize them using the R packages “GEOquery” and “dplyr” in the R language (version 4.4.2), respectively, to obtain the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674.

[0091] Step 2. The bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674 were screened for differentially expressed genes using the R package "limma" in R language. The screening conditions for bipolar disorder differentially expressed genes were |LogFC|>0 and p.value<0.05, resulting in 2551 differentially expressed genes, including 1481 up-regulated genes and 1070 down-regulated genes. The screening conditions for epilepsy differentially expressed genes were |LogFC|>1 and p.value<0.05, resulting in 980 differentially expressed genes, including 750 up-regulated genes and 230 down-regulated genes. Volcano plots were drawn for the bipolar disorder and epilepsy datasets, as shown in Figure 2. Figure 2 Figure 3 .

[0092] Step 3. For the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674, weighted gene co-expression network analysis (WGCNA) was performed using the R package "WGCNA" in the R language. The genes with the top 50% variance were selected to construct a scale-free co-expression network. The dynamic cutting tree algorithm was used to identify the co-expression modules. The correlation between the module feature genes and the phenotype was calculated, and the module-phenotype correlation heat map was drawn as shown in Figure 3. Figure 4 Figure 5 , screened significantly correlated modules (|r| ≥ 0.5, p < 0.01), selected MEpink, MEbrown and MEred modules from the weighted gene co-expression network analysis (WGCNA) results of the bipolar disorder dataset GSE5389, which included a total of 1818 genes, and selected MEpink module from the weighted gene co-expression network analysis (WGCNA) results of the epilepsy dataset GSE28674, which included a total of 919 genes.

[0093] Step 4. The bipolar disorder differentially expressed genes and epilepsy differentially expressed genes obtained in step 2 were intersected with the bipolar disorder significantly associated module genes and epilepsy significantly associated module genes obtained in step 3 through a Venn diagram, and 113 key differentially expressed genes were initially screened out. Figure 6 .

[0094] Step 5. The screened core differentially expressed genes were analyzed using the STRING database (https: / / string-db.org / ) to construct a protein interaction network (PPI). Cytoscape 3.10.3 software and the MCODE plug-in (parameter settings K-Core = 2, Node Score Cutoff = 0.2, Degree Cutoff = 2, Max Depth = 100) were used to construct a visual protein interaction network diagram as shown in the figure. Figure 7 Figure 8 , then 14 genes included in all modules calculated by MCODE were selected as key genes.

[0095] Step 6. Use R packages "glmnet", "e1071" and "randomForest" to perform 8-fold LASSO regression analysis on the key genes in the bipolar disorder dataset. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 9 Figure 10 Four genes were obtained, and 5-fold support vector machine recursive feature elimination (SVM-RFE) was used to screen nonlinear relationships based on kernel functions to evaluate the contribution of feature classification boundaries. Figure 11 Figure 12 9 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 5 genes with gene importance greater than 0.7 were selected. Figure 13 Figure 14 , and then draw a Venn diagram for the results of the three machine learning methods, and get 6 core genes as follows Figure 15 .

[0096] Step 7. Perform 8-fold LASSO regression analysis on the key genes in the epilepsy dataset using R packages “glmnet”, “e1071”, and “randomForest”. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 16 Figure 17 6 genes were obtained, and 5-fold support vector machine recursive feature elimination was performed

[0097] (SVM-RFE) screening is based on kernel function to process nonlinear relationships and evaluate the contribution of feature classification boundaries. Figure 18 Figure 199 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 6 genes with gene importance greater than 0.7 were selected. Figure 20 Figure 21 , and then draw a Venn diagram for the results of the three machine learning, and get 8 core genes as follows Figure 22 .

[0098] Step 8. Take the intersection of the bipolar disorder core genes and epilepsy core genes obtained in steps 6 and 7, and draw a Venn diagram as shown in Figure 23 , and obtained two genes "RGS4" and "GABRA1" as candidate core pathogenic genes for bipolar disorder and epilepsy.

[0099] Lasso regression principle:

[0100] (1) By adding the L1 regularization term to the loss function, the model parameters are made as sparse as possible, thereby achieving feature selection and preventing overfitting.

[0101] (2) The objective function of Lasso regression can be expressed as:

[0102] L(w)=(y-Xw)2+λ||w||1

[0103] y: target variable;

[0104] X: feature matrix;

[0105] w: model parameters;

[0106] λ: regularization parameter;

[0107] ||w||1: 1-norm of w, that is, the sum of the absolute values ​​of all elements.

[0108] (3) Since the L1 regularization term is not differentiable at zero, Lasso regression is usually optimized using the coordinate descent method. This method only searches for the minimum value of the loss function on the current coordinate axis in each iteration without calculating the gradient of the function, which is more efficient.

[0109] (4) Application of Lasso regression:

[0110] Lasso regression has significant advantages in processing high-dimensional data, especially when the number of features is much larger than the number of samples. It can automatically select important features, simplify the model and improve interpretability.

[0111] The core principle of Support Vector Machine Recursive Feature Elimination (SVM-RFE):

[0112] (1) The principle of support vector machine recursive feature elimination (SVM-RFE) for screening genes is based on its classification ability. It maximizes the interval between samples of different categories by finding an optimal hyperplane, which is called the maximum margin hyperplane.

[0113] (2) During the SVM-RFE model training process, the contribution of each feature to the classification result will be evaluated. Features with higher importance have a greater impact on the classification result, while features with lower importance are considered to contribute less to the classification result.

[0114] (3) Application of SVM-RFE:

[0115] Suitable for high-dimensional small sample data; robust to noise and overfitting; can be combined with kernel functions to handle nonlinear relationships.

[0116] Random Forest Principle:

[0117] (1) The principle of random forest gene screening is to evaluate the influence of each gene on the classification results by constructing multiple decision trees, thereby determining which genes have a significant impact on the classification results. The random forest algorithm implements gene screening through the following steps:

[0118] Step 1. Build a decision tree: A random forest is composed of multiple decision trees. During the construction of each decision tree, a portion of samples and features are randomly selected from the original dataset for training.

[0119] Step 2. Feature selection: When splitting each decision tree node, a portion of features are randomly selected to select the best features. This process increases the diversity of the model.

[0120] Step 3. Voting or averaging: For classification tasks, random forests determine the final result by majority voting; for regression tasks, the final result is determined by calculating the average of the prediction results of all trees.

[0121] Step 4. Feature Importance Assessment: Random forests can assess the importance of each feature. This is typically done by calculating the contribution of each feature in each tree and then taking the average. Contribution metrics include the Gini index and the out-of-bag error rate.

[0122] (2) In gene screening tasks, random forests can be used to assess the importance of each gene. By training on a dataset, random forests can measure the contribution of each gene to the prediction model. These metrics can be used to rank the importance of genes, thereby selecting the most influential candidate genes as potential biomarkers or therapeutic targets.

[0123] (3) Application of Random Forest:

[0124] Applicable to high-dimensional nonlinear data; strong resistance to overfitting; can evaluate interactions (gene combinations).

[0125] Example 3:

[0126] Step 1. Obtain the bipolar disorder gene chip dataset GSE5389 and the epilepsy gene chip dataset GSE28674 from the GEO database of the NCBI website (https: / / www.ncbi.nlm.nih.gov / ). Standardize them using the R packages “GEOquery” and “dplyr” in the R language (version 4.4.2), respectively, to obtain the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674.

[0127] Step 2. The bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674 were screened for differentially expressed genes using the R package "limma" in R language. The screening conditions for bipolar disorder differentially expressed genes were |LogFC|>0 and p.value<0.05, resulting in 2551 differentially expressed genes, including 1481 up-regulated genes and 1070 down-regulated genes. The screening conditions for epilepsy differentially expressed genes were |LogFC|>1 and p.value<0.05, resulting in 980 differentially expressed genes, including 750 up-regulated genes and 230 down-regulated genes. Volcano plots were drawn for the bipolar disorder and epilepsy datasets, as shown in Figure 2. Figure 2 Figure 3 .

[0128] Step 3. For the bipolar disorder dataset GSE5389 and the epilepsy dataset GSE28674, weighted gene co-expression network analysis (WGCNA) was performed using the R package "WGCNA" in the R language. The genes with the top 50% variance were selected to construct a scale-free co-expression network. The dynamic cutting tree algorithm was used to identify the co-expression modules. The correlation between the module feature genes and the phenotype was calculated, and the module-phenotype correlation heat map was drawn as shown in Figure 3. Figure 4 Figure 5 , screened significantly correlated modules (|r| ≥ 0.5, p < 0.01), selected MEpink, MEbrown and MEred modules from the weighted gene co-expression network analysis (WGCNA) results of the bipolar disorder dataset GSE5389, which included a total of 1818 genes, and selected MEpink module from the weighted gene co-expression network analysis (WGCNA) results of the epilepsy dataset GSE28674, which included a total of 919 genes.

[0129] Step 4. The bipolar disorder differentially expressed genes and epilepsy differentially expressed genes obtained in step 2 were intersected with the bipolar disorder significantly associated module genes and epilepsy significantly associated module genes obtained in step 3 through a Venn diagram, and 113 key differentially expressed genes were initially screened out. Figure 6 .

[0130] Step 5. The screened core differentially expressed genes were analyzed using the STRING database (https: / / string-db.org / ) to construct a protein interaction network (PPI). Cytoscape 3.10.3 software and the MCODE plug-in (parameter settings K-Core = 2, Node Score Cutoff = 0.2, Degree Cutoff = 2, Max Depth = 100) were used to construct a visual protein interaction network diagram as shown in the figure. Figure 7 Figure 8 , then 14 genes included in all modules calculated by MCODE were selected as key genes.

[0131] Step 6. Use R packages "glmnet", "e1071" and "randomForest" to perform 10-fold LASSO regression analysis on the key genes in the bipolar disorder dataset. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 9 Figure 10 Four genes were obtained, and 5-fold support vector machine recursive feature elimination (SVM-RFE) was used to screen nonlinear relationships based on kernel functions to evaluate the contribution of feature classification boundaries. Figure 11 Figure 12 9 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 5 genes with gene importance greater than 0.7 were selected. Figure 13 Figure 14 , and then draw a Venn diagram for the results of the three machine learning methods, and get 6 core genes as follows Figure 15 .

[0132] Step 7. Perform 10-fold LASSO regression analysis on the key genes in the epilepsy dataset using R packages “glmnet”, “e1071”, and “randomForest”. Use L1 regularization to achieve feature sparsification and preliminarily screen key variables such as Figure 16 Figure 17 Six genes were obtained, and 5-fold support vector machine recursive feature elimination (SVM-RFE) was used to screen nonlinear relationships based on kernel functions to evaluate the contribution of feature classification boundaries. Figure 18 Figure 19 9 genes were obtained, and the random forest algorithm set the number of decision trees to 500 for random forest analysis. 6 genes with gene importance greater than 0.7 were selected. Figure 20 Figure 21 , and then draw a Venn diagram for the results of the three machine learning, and get 8 core genes as follows Figure 22 .

[0133] Step 8. Take the intersection of the bipolar disorder core genes and epilepsy core genes obtained in steps 6 and 7, and draw a Venn diagram as shown in Figure 23 , and obtained two genes "RGS4" and "GABRA1" as candidate core pathogenic genes for bipolar disorder and epilepsy.

[0134] Lasso regression principle:

[0135] (1) By adding the L1 regularization term to the loss function, the model parameters are made as sparse as possible, thereby achieving feature selection and preventing overfitting.

[0136] (2) The objective function of Lasso regression can be expressed as:

[0137] L(w)=(y-Xw)2+λ||w||1

[0138] y: target variable;

[0139] X: feature matrix;

[0140] w: model parameters;

[0141] λ: regularization parameter;

[0142] ||w||1: 1-norm of w, that is, the sum of the absolute values ​​of all elements.

[0143] (3) Since the L1 regularization term is not differentiable at zero, Lasso regression is usually optimized using the coordinate descent method. This method only searches for the minimum value of the loss function on the current coordinate axis in each iteration without calculating the gradient of the function, which is more efficient.

[0144] (4) Application of Lasso regression:

[0145] Lasso regression has significant advantages in processing high-dimensional data, especially when the number of features is much larger than the number of samples. It can automatically select important features, simplify the model and improve interpretability.

[0146] The core principle of Support Vector Machine Recursive Feature Elimination (SVM-RFE):

[0147] (1) The principle of support vector machine recursive feature elimination (SVM-RFE) for screening genes is based on its classification ability. It maximizes the interval between samples of different categories by finding an optimal hyperplane, which is called the maximum margin hyperplane.

[0148] (2) During the SVM-RFE model training process, the contribution of each feature to the classification result will be evaluated. Features with higher importance have a greater impact on the classification result, while features with lower importance are considered to contribute less to the classification result.

[0149] (3) Application of SVM-RFE:

[0150] Suitable for high-dimensional small sample data; robust to noise and overfitting; can be combined with kernel functions to handle nonlinear relationships.

[0151] Random Forest Principle:

[0152] (1) The principle of random forest gene screening is to evaluate the influence of each gene on the classification results by constructing multiple decision trees, thereby determining which genes have a significant impact on the classification results. The random forest algorithm implements gene screening through the following steps:

[0153] Step 1. Build a decision tree: A random forest is composed of multiple decision trees. During the construction of each decision tree, a portion of samples and features are randomly selected from the original dataset for training.

[0154] Step 2. Feature selection: When splitting each decision tree node, a portion of features are randomly selected to select the best features. This process increases the diversity of the model.

[0155] Step 3. Voting or averaging: For classification tasks, random forests determine the final result by majority voting; for regression tasks, the final result is determined by calculating the average of the prediction results of all trees.

[0156] Step 4. Feature Importance Assessment: Random forests can assess the importance of each feature. This is typically done by calculating the contribution of each feature in each tree and then taking the average. Contribution metrics include the Gini index and the out-of-bag error rate.

[0157] (2) In gene screening tasks, random forests can be used to assess the importance of each gene. By training on a dataset, random forests can measure the contribution of each gene to the prediction model. These metrics can be used to rank the importance of genes, thereby selecting the most influential candidate genes as potential biomarkers or therapeutic targets.

[0158] (3) Application of Random Forest:

[0159] Applicable to high-dimensional nonlinear data; strong resistance to overfitting; can evaluate interactions (gene combinations).

[0160] The conventional techniques in the above embodiments are prior arts known to those skilled in the art, and thus will not be described in detail here.

[0161] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.

[0162] Although the present invention has been described in detail and some specific examples have been cited, it is obvious to those skilled in the art that various changes or modifications can be made without departing from the spirit and scope of the present invention. Although the above specific embodiments have shown, described and pointed out the novel features applied to various embodiments, it should be understood that various omissions, replacements and changes can be made to the form and details of the described methods without departing from the spirit of the present disclosure. In addition, the various features and methods described above can be used independently of each other or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Many of the above embodiments include similar components, and therefore, these similar components are interchangeable in different embodiments. Although the present invention has been disclosed in the context of certain embodiments and examples, it should be understood by those skilled in the art that the present invention can extend beyond the specifically disclosed embodiments to other alternative embodiments and / or applications and their obvious modifications and equivalents. Therefore, the present invention is not intended to be limited by the specific disclosure of the preferred embodiments herein.

[0163] Matters not covered in the present invention are all known technologies.

[0164] References:

[0165] [1]Luigi F Saccaro,Jasper Crokaert et al.Structural and functionalMRI correlates of inflammation in bipolar disorder:A systematic review.Journal of Affective Disorders,2023,15:325:83-92

[0166] [2]Roland D Thijs,Rainer Surges et al.Adult epilepsy.Lancet,2023,402(10399):414-424

[0167] [3]Jimmy Li,Lawrence Ledoux-Hutchinson et al.Prevalence ofBipolarSymptoms or Disorder in Epilepsy:A Systematic Review and Meta-analysisNeurology,2022,10,98(19):e1913-e1922

[0168] [4]XueyingLi,Leiwu et al.Ferroptosis-Related Gene Signatures inEpilepsy:Diagnostic and Immune Insights.MolecularNeurobiology,2025,62(2):1998-2011。

Claims

1. A method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy, characterized in that: The following steps are involved: Step 1. Obtain a bipolar disorder gene expression dataset from the brain, including a bipolar disorder patient group and a control group, and an epilepsy gene expression dataset, including an epilepsy patient group and a control group, from the GEO database; Step 2. Using differential analysis, identify differentially expressed genes with significant expression changes between the bipolar disorder patient group and the control group data set, and differentially expressed genes with significant expression changes between the epilepsy patient group and the control group data set; Step 3. Use weighted gene co-expression network analysis (WGCNA) to calculate the correlation between module characteristic genes and each disease phenotype, and screen significantly correlated module genes; Step 4. Intersect the bipolar disorder differentially expressed genes and epilepsy differentially expressed genes obtained in step 2 with the bipolar disorder significantly associated module genes and epilepsy significantly associated module genes obtained in step 3 through a Venn diagram to preliminarily screen out key differentially expressed genes; Step 5. Use the key differentially expressed genes from step 4 to construct a protein-protein interaction network (PPI) and screen out key genes; Step 6. Key genes in the bipolar disorder gene expression dataset were screened using LASSO regression, support vector machine recursive feature elimination (SVM-RFE), and random forest. A Venn diagram was then drawn for the results of the three machine learning methods to identify core bipolar disorder genes. Step 7. Key genes in the epilepsy gene expression dataset were screened using LASSO regression, support vector machine recursive feature elimination (SVM-RFE), and random forest. A Venn diagram was then drawn for the results of the three machine learning methods to identify core epilepsy genes. Step 8. Take the intersection of the core genes for bipolar disorder and the core genes for epilepsy obtained in steps 6 and 7, draw a Venn diagram, and use the obtained intersection genes as candidate core pathogenic genes for bipolar disorder and epilepsy.

2. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that The differential analysis in step 2 was performed using the "limma" package in R language (4.4.2). The screening of differentially expressed genes in bipolar disorder met the criteria of |LogFC|>0, p.value<0.05, and the screening of differentially expressed genes in epilepsy met the criteria of |LogFC|>1, p.value<0.

05.

3. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that The weighted gene co-expression network analysis in step 3 used the "WGCNA" package in the R language (4.4.2). The genes with the top 50% variance were selected to construct a scale-free co-expression network. The soft threshold power selection was based on the scale-free topology criterion. The dynamic cutting tree algorithm was used to identify co-expression modules, calculate the correlation between module eigengenes and phenotypes, and screen for significantly correlated modules (|r| ≥ 0.5, p < 0.05).

4. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that The protein interaction network in step 5 was drawn using the STRING database and the MCODE plug-in in Cytoscape software (3.10.3) to draw a protein interaction network (PPI) to screen hub genes from key differentially expressed genes, with the parameter settings of K-Core = 2, Node Score Cutoff = 0.2, Degree Cutoff = 2, and Max Depth = 100.

5. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that The LASSO regression in step 6 includes 5-10 fold cross-validation LASSO regression, the support vector machine recursive feature elimination uses 5-fold cross-validation support vector machine recursive feature elimination, and the random forest algorithm sets the number of decision trees to 500 and selects genes with importance greater than 0.

7.

6. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that The LASSO regression in step 7 includes 5-10 fold cross-validation LASSO regression, the support vector machine recursive feature elimination uses 5-fold cross-validation support vector machine recursive feature elimination, and the random forest algorithm sets the number of decision trees to 500 and selects genes with importance greater than 0.

7.

7. The method for screening biomarkers of co-pathogenicity of bipolar disorder and epilepsy according to claim 1, characterized in that In step 8, the intersection of the three machine learning results is taken to obtain an intersection list containing two candidate genes, including RGS4 and GABRA1.