Method for screening copathogenic core gene of psoriasis and schizophrenia

Through gene chip expression profile analysis and multiple machine learning methods, the hub genes co-expressed by psoriasis and schizophrenia were screened out, which solved the problem of insufficient mining of co-expressed genes in the existing technology, achieved more accurate diagnostic and treatment strategies, and improved the reliability of research on co-disease mechanisms.

CN120526850APending Publication Date: 2025-08-22CAPITAL UNIVERSITY OF MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510625044.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Existing research lacks systematic mining of genes co-expressed by psoriasis and schizophrenia, insufficient specificity of traditional markers, resulting in limited development of diagnostic and therapeutic targets, and low reliability in screening of a single machine learning algorithm, making it difficult to meet clinical needs.

Method used

Gene chip expression profile analysis combined with LASSO/SVM-RFE/RF triple machine learning method was used to identify the hub genes co-expressed by psoriasis and schizophrenia. Through PPI network construction and functional enrichment analysis, the co-pathogenic core genes were screened out, and their diagnostic performance was verified through ROC curves.

Benefits of technology

It significantly improves the accuracy and biological interpretation of co-pathogenic core gene screening, improves the specificity of diagnosis, provides a theoretical basis for personalized medical care for psoriasis and schizophrenia, and improves patients' quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526850A_ABST
    Figure CN120526850A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological feature recognition, and particularly relates to a method for screening a copathogenic core gene of psoriasis and schizophrenia, which comprises the following steps: performing differential expression analysis on known gene chip expression profile data to obtain a differential expression gene, taking an intersection of the differential expression genes of two diseases to obtain a coexpression differential gene, and screening the copathogenic core gene of psoriasis and schizophrenia. According to the method, a PPI network is used for finding hub genes and key proteins in co-expression differential genes, then results of LASSO, SVM-RFE and RF machine learning methods are taken to be intersected to screen out co-pathogenic core genes, and finally an ROC curve is used for detecting the disease diagnosis capability of the co-pathogenic core genes. The average accuracy of the copathogenic core gene for diagnosing psoriasis and schizophrenia can reach the AUC value of 0.84, which is higher than that of single gene chip expression profile data analysis or single machine learning recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biometric identification technology, and specifically relates to a method for screening core genes that are co-causative of psoriasis and schizophrenia, further providing a theoretical basis for the study of the co-morbidity mechanism and clinical diagnosis of the two diseases. Background Art

[0002] Psoriasis (PsO) is a common chronic inflammatory skin disease associated with physical and psychological burdens. It is characterized by damage to multiple systems such as the skin, nails, and joints, accompanied by symptoms such as significant itching and scaling. [1] Approximately 2-3% of the world's population is affected [2,3] The incidence of psoriasis has been reported to range from 30.3 to 321.0 per 100,000 person-years and is associated with age, sex, geographic location, race, and other genetic and environmental factors. [4,5] .

[0003] Schizophrenia (SCZ) is a serious mental illness with a profound impact on public health, with a global lifetime prevalence of approximately 1% across diverse geographic, cultural, and socioeconomic categories. [6] China's schizophrenia burden rate is 276.6, which is much higher than the global schizophrenia burden rate of 177.2. [7] Studies have shown that this disease may be related to abnormal neurodevelopment, immune disorders, etc. [8] , which usually begins in late adolescence or early adulthood [9] Most patients experience long-term adverse effects characterized by a range of symptoms including hallucinations, delusions or disruption of thought processes, disorganized thinking, and social withdrawal.

[10] .

[0004] Our team has previously published two papers. Zhu Peixin et al. conducted a meta-analysis on the comorbidity of psoriasis and schizophrenia, published in Dermatologic Therapy.

[11] The results showed that although psoriasis and schizophrenia differ in clinical manifestations and pathological mechanisms, there is a significant comorbidity between the two. Ruina et al. identified schizophrenia pathogenic genes and related pathways by analyzing the changes in gene transcription levels in the amygdala of mice with social isolation as a model of schizophrenia, and published their findings in the Chinese Journal of Neurology.

[12] Other studies have suggested that psychological stress, genetics, and immune system disorders are important factors in the comorbidity between psoriasis and schizophrenia, particularly suggesting that immune system abnormalities may play a significant role in the pathogenesis of both diseases. Elevated levels of inflammatory factors are considered a common feature of psoriasis and schizophrenia. Psoriasis causes skin lesions by overactivating the immune response, while the immunological hypothesis of schizophrenia also emphasizes the role of the immune system in neuroinflammation and psychiatric symptoms.

[0005] However, existing studies lack systematic exploration of co-expressed genes of the two diseases, and traditional markers are not specific enough, which limits the development of diagnostic and therapeutic targets. Gene chips and RNA sequencing technologies provide high-throughput data support for the study of molecular mechanisms of diseases, but how to screen clinically significant genes from massive data remains a challenge. Compared with the references

[13] For example, the ChIP-seq method relies on manual functional annotation, which is more time-consuming. The results of eQTL analysis also have the risk of false positives. This method uses gene chip expression profiling for the first time to identify hub genes and signaling pathways that are co-expressed in psoriasis and schizophrenia. It also combines LASSO / SVM-RFE / RF triple machine learning to quickly and automatically identify key candidate genes from thousands of genes. This method is suitable for large-scale data mining and integrates multi-algorithm machine learning cross-validation (such as taking the LASSO+SVM-RFE+RF intersection of key co-expressed DEGs) to effectively reduce single technical errors.

[0006] Machine learning algorithms (such as LASSO, SVM-RFE, etc.) can optimize gene screening through feature selection. Compared with reference

[14] Support vector machine (SVM), random forest (RF), naive Bayes (NB), and gradient boosting (XGBoost) were used to model global attributes or node attributes separately, but a single algorithm was easily affected by data noise: its prediction model performance was low, the random forest model had an AUC of only 0.68, an accuracy of 68.6%, and limited diagnostic efficacy. Other models (such as SVM AUC = 0.63) performed even worse and could not meet clinical needs. The ROC curves of experimental datasets (GSE14905 and GSE53987) and external datasets (GSE6710 and GSE21138) were used for validation. The results showed that the average AUC of this method for co-pathogenic core genes (such as PRSS3 and GADD45B) could reach 0.84, and it had higher reliability across platforms and cohorts.

[0007] The above two articles

[13]

[14] Focusing only on the abnormalities of brain structural networks in schizophrenia makes it difficult to explain the comorbid association between schizophrenia and other immune diseases (such as psoriasis). Summary of the Invention

[0008] To overcome the shortcomings of existing technologies, such as insufficient specificity of disease markers, low reliability of single-algorithm screening, and weak research on cross-disease comorbidity mechanisms, the present invention proposes a method for screening co-pathogenic core genes for psoriasis and schizophrenia. This method innovatively utilizes gene chip expression profiling to identify co-expressed hub genes and signaling pathways in both diseases. LASSO / SVM-RFE / RF triple machine learning is then used to screen hub genes in the gene chip expression profile data of psoriasis and schizophrenia, respectively. The three machine learning results for psoriasis and the three machine learning results for schizophrenia are intersected one-to-one to screen key co-expressed DEGs. The key co-expressed DEGs that overlap among the three machine learning methods are then selected as co-pathogenic core genes, achieving precise localization of co-pathogenic core genes. The machine learning method uses triple feature selection of LASSO regression, SVM-RFE, and random forest. The complementary advantages of the algorithms effectively overcome the overfitting and feature bias problems of a single model, significantly improving the accuracy and biological interpretability of screening for co-pathogenic core genes. The specific technical solution is as follows:

[0009] A method for screening core genes co-causative for psoriasis and schizophrenia, comprising:

[0010] 1. Data acquisition and preprocessing: Gene chip data of PsO and its healthy control group, and gene chip data of SCZ and its healthy control group were downloaded from public databases and standardized respectively;

[0011] 2. Differential gene screening: screen the DEGs of the two diseases separately, and obtain the co-expressed DEGs by taking the intersection of the DEGs;

[0012] 3. PPI network construction and hub gene screening: Construct a PPI network of differentially expressed genes, perform density analysis on the PPI network, and screen hub genes;

[0013] 4. Enrichment analysis: GO and KEGG enrichment analysis were performed to reveal the hub genes in biological processes (BP), cellular components (CC), and molecular functions (MF);

[0014] 5. Three machine learning methods to screen co-pathogenic core genes: LASSO / SVM-RFE / RF triple machine learning was used to screen hub genes in the gene chip expression profile data of psoriasis and schizophrenia, respectively. The three machine learning results of psoriasis and schizophrenia were matched one by one and the intersection was taken to screen key co-expressed DEGs. The key co-expressed DEGs that overlapped with the three machine learning methods were then selected as co-pathogenic core genes.

[0015] 6. Diagnostic efficacy verification: The AUC values ​​of the co-pathogenic core genes in PsO and SCZ were displayed using ROC curves, and the reliability of the co-pathogenic core genes was further confirmed using external datasets.

[0016] As a further preferred solution, the data acquisition and preprocessing steps described in the embodiments of the present application include:

[0017] The GSE14905 gene chip expression profile data of PsO, which includes 21 normal skin biopsy samples from 21 healthy donors and 56 skin biopsy samples from 28 psoriasis patients with matched lesional and non-lesional tissues, as well as 5 samples from psoriasis patients who only provided lesional skin biopsies, for a total of 82 samples, and the GSE53987 gene chip data of SCZ, which includes the prefrontal cortex (Brodmann Area 46), striatum, and hippocampus of cadaveric brains of schizophrenia patients and controls (n = 19 in each group), were downloaded from the GEO database.

[0018] As a further preferred solution, the differential gene screening step described in the examples of the present application includes:

[0019] The R package "Limma" in R language was used to screen the DEGs of the two diseases respectively. For GSE14905, P value <0.05 and |log2 FC|>1 were used to identify differentially expressed genes. For GSE53987, P value <0.05 and |log2 FC|>0.3 were used to identify differentially expressed genes. 841 DEGs were found in the psoriasis group, and upregulated genes included RIMS3, MSL3, etc., and downregulated genes included JUN, NR4A1, etc.; 506 DEGs were found in the schizophrenia group, and upregulated genes included CASP7, BAG3, etc., and downregulated genes included NCOA2, TRAK2, etc.; Venn diagrams were drawn for the DEGs to screen out 64 co-expressed DEGs.

[0020] As a further preferred solution, the PPI network construction and hub gene screening steps described in the embodiments of the present application include:

[0021] The co-expressed DEGs obtained after differential analysis were entered into the STRING database (https: / / string-db.org / ), and the confidence level was set to 0.4 to screen out hub genes. 23 hub genes were screened out, and a protein-protein interaction network of differentially expressed genes was constructed. A visualized PPI network was constructed using Cytoscape 3.7.2 software. Among them, ZFP36, EGR1, CDKN1A, and CEBPD proteins were closely connected and had a high influence weight on other proteins, and were the key proteins in the co-expressed DEGs.

[0022] As a further preferred solution, the enrichment analysis step described in the embodiment of the present application includes:

[0023] GO and KEGG enrichment analysis were performed using the R package "clusterProfier" in R language. The BPs of the co-expressed DEGs involved positive regulation of the degradation process of nuclear transcript mRNAs that depend on deadenylation, regulation of epithelial cell differentiation, regulation of the degradation process of nuclear transcript mRNAs that depend on deadenylation, response to glucocorticoids, kidney development, response to corticosteroids, renal system development, and cellular response to glucocorticoid stimulation; CC involved the CCR4-NOT complex; MF involved 14-3-3 protein binding, NAD+ protein poly ADP ribosyltransferase activity, NAD+ protein ADP ribosyltransferase activity, mRNA 3'-UTR AU-rich region binding, and pentosyltransferase activity; pathways involved the FoxO signaling pathway, transcriptional dysregulation in cancer, thyroid cancer, cellular senescence, endometrial cancer, basal cell carcinoma, human T-cell leukemia virus type 1 infection, and melanoma.

[0024] As a further preferred solution, the three machine learning analysis steps described in the embodiment of the present application include:

[0025] In the psoriasis group, the LASSO cross-validation coefficient curve yielded a λ of 0.02302865, resulting in 11 non-zero coefficient genes. In the schizophrenia group, the LASSO cross-validation coefficient curve yielded a λ of 0.05758736, resulting in 7 non-zero coefficient genes.

[0026] Nine genes with the lowest 5xCV error and best 5xCV accuracy were identified using SVM-REF in the psoriasis group, and 11 genes with the lowest 5xCV error and best 5xCV accuracy were identified using SVM-RFE in the schizophrenia group;

[0027] The hub genes were input into the RF classifier, and the top 15 genes were displayed on the importance scale. In the psoriasis group, the number of optimal trees was selected as 1, the importance was >1, and a total of 9 genes were identified. In the schizophrenia group, the number of optimal trees was selected as 1, the importance was >1, and a total of 5 genes were identified.

[0028] After the machine learning results of the psoriasis group and the schizophrenia group were intersected, there were 5 key co-expressed DEGs in SVM-RFE, namely MAFF, CDKN1A, EPRS, and PRSS3; there were 3 co-expressed DEGs in LASSO, namely EGR1, PRSS3, and GADD45B; there were 3 co-expressed DEGs in RF, namely TOB1, GADD45B, and ZFP36L1;

[0029] SVM-RFE and LASSO screened out one overlapping gene (PRSS3), and RF and LASSO screened out one overlapping gene (GADD45B). PRSS3 and GADD45B are the core pathogenic genes of psoriasis and schizophrenia.

[0030] As a further preferred solution, the diagnostic efficacy verification step described in the embodiment of the present application includes:

[0031] The ROC curve was used to display the AUC value to evaluate the diagnostic efficacy of co-pathogenic core genes. An AUC value > 0.6 indicated good diagnostic performance. In the psoriasis group (GSE14905), PRSS3 (AUC = 0.909) and GADD45B (AUC = 0.924) showed high diagnostic value; in the schizophrenia group (GSE53987), PRSS3 (AUC = 0.786) and GADD45B (AUC = 0.853) also performed well. Validated by external datasets (GSE6710 and GSE21138), the diagnostic AUCs of PRSS3 and GADD45B were both > 0.7, confirming their reliability across datasets.

[0032] In the above technical solution, the present invention provides a method for screening core genes that are co-causative of psoriasis and schizophrenia, which has the following beneficial effects:

[0033] 1. The method for screening co-pathogenic core genes of psoriasis and schizophrenia described in the present invention can comprehensively evaluate the expression patterns and functions of genes from multiple dimensions by comprehensively applying multiple analysis methods (differential expression analysis, functional correlation analysis, support vector machine-recursive feature elimination and other machine learning feature selection, etc.), thereby more accurately identifying co-pathogenic core genes in the co-expressed DEGs related to the pathogenicity of psoriasis and schizophrenia.

[0034] 2. In the method for screening the core co-pathogenic genes of psoriasis and schizophrenia described in the present invention, a PPI network is used to determine the hub genes, and the intersection of the results of three machine learning algorithms, LASSO, SVM-RFE, and RF, is taken to determine the co-pathogenic core genes, which helps to screen out potential key co-pathogenic genes of psoriasis and schizophrenia and improve the specificity of diagnosis.

[0035] 3. The method for screening co-pathogenic core genes of psoriasis and schizophrenia described in the present invention can identify the co-pathogenic core genes of co-expressed DEGs of psoriasis and schizophrenia, providing a theoretical basis for developing more accurate diagnostic tools and treatment strategies, realizing personalized medical treatment for psoriasis and schizophrenia, and improving the quality of life of patients.

[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the process of the present invention;

[0038] Figure 2 Volcano plot of differentially expressed genes between psoriasis and schizophrenia; A: Volcano plot of DEGs in psoriasis, green indicates downregulated genes, purple indicates upregulated genes, and gray indicates genes with unchanged expression; B: Volcano plot of DEGs in schizophrenia, green indicates downregulated genes, purple indicates upregulated genes, and gray indicates genes with unchanged expression;

[0039] Figure 3 Venn diagram of co-expressed DEGs in psoriasis and schizophrenia;

[0040] Figure 4 PPI network and hub genes; each circle represents a co-expressed DEGs. The darker the color and the larger the area of ​​the circle, the higher the influence weight on other co-expressed DEGs.

[0041] Figure 5 Results of GO and KEGG enrichment analysis; A: GO enrichment results, the horizontal axis represents the number of genes, the vertical axis represents the functions involved in the genes, and the redder the color, the more significant the enrichment; B: KEGG enrichment results, the horizontal axis represents the number of genes, the vertical axis represents the pathways involved in the genes, and the redder the color, the more significant the enrichment;

[0042] Figure 6The results of LASSO regression analysis; A: Cross-validation curve diagram for schizophrenia. The upper horizontal axis shows the number of variables required for the corresponding model; the lower horizontal axis is the logarithm of the penalty coefficient λ, and the vertical axis represents the mean squared error (MSE). The lambda.min dotted line on the left side of the figure indicates the horizontal axis corresponding to the minimum MSE, and the lambda.1se dotted line on the right side indicates the position one standard error away from lambda.min; B: Coefficient path diagram for schizophrenia. The horizontal axis has the same meaning as in A, and the vertical axis represents the coefficient of each variable in the model. Lines of different colors represent different variables; C: Cross-validation curve diagram for psoriasis. The horizontal and vertical axes have the same meanings as in A; D: Coefficient path diagram for psoriasis. The horizontal axis has the same meaning as in A and C, and the vertical axis represents the coefficient of each variable in the model. Lines of different colors represent different variables. Figure 7 The results of SVM-RFE machine learning are shown below. A: Error rate graph for schizophrenia, with the error rate on the ordinate and the number of models on the abscissa. Nine genes were obtained at an error rate of 0.0492. B: Accuracy graph for schizophrenia, with the accuracy on the ordinate and the number of models on the abscissa. Nine genes were obtained at an accuracy of 0.951. C: Error rate graph for the psoriasis group, with the error rate on the ordinate and the number of models on the abscissa. 11 genes were obtained at an error rate of 0.114. D: Accuracy graph for the psoriasis group, with the accuracy on the ordinate and the number of models on the abscissa. 11 genes were obtained at an error rate of 0.886.

[0043] Figure 8 RF machine learning results; A: Random forest plot of psoriasis. The number of optimal trees is 1. The black dashed line is the error, the green dashed line is the control group, and the red dashed line is the disease group. B: Display of the top 15 key genes in psoriasis, sorted by importance.

[0044] C: Random forest plot of psoriasis. The number of optimal trees is 1. The black dashed line is the error, the green dashed line is the control group, and the red dashed line is the disease group. D: Display of the top 15 key genes in the schizophrenia group, sorted by importance.

[0045] Figure 9 The intersection results of machine learning; A: intersection of the results of the SVM schizophrenia group and the psoriasis group; B: intersection of the results of the LASSO schizophrenia group and the psoriasis group; C: intersection of the results of the RF schizophrenia group and the psoriasis group; Figure 10 The expression of core genes; A: PRSS3 expression in the psoriasis group; B: PRSS3 expression in the schizophrenia group; C: GADD45B expression in the psoriasis group; D: GADD45B expression in the schizophrenia group;

[0046] Figure 11ROC curves of the experimental groups; A: diagnostic effect of GADD45B in the psoriasis group; B: diagnostic effect of PRSS3 in the psoriasis group; C: diagnostic effect of GADD45B in the schizophrenia group; D: diagnostic effect of PRSS3 in the schizophrenia group;

[0047] Figure 12 Figure 2 is the ROC curve diagram of the validation group; A: diagnostic effect of GADD45B in the psoriasis group; B: diagnostic effect of PRSS3 in the psoriasis group; C: diagnostic effect of GADD45B in the schizophrenia group; D: diagnostic effect of PRSS3 in the schizophrenia group. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0049] Example 1: A method for screening core genes co-causative for psoriasis and schizophrenia

[0050] Step 1: Data acquisition and preprocessing

[0051] From the Gene Expression Omnibus (GEO),

[0052] The psoriasis gene chip dataset (GSE14905) and the schizophrenia gene chip dataset (GSE53987) were downloaded from https: / / www.ncbi.nlm.nih.gov / geo / . Both datasets are based on the GPL570 platform (HG-U133_Plus_2). GSE14905 contains 82 samples, including skin biopsy samples from 21 healthy controls and samples from 61 psoriasis patients (including lesion and non-lesion tissues); GSE53987 contains gene expression data of the prefrontal cortex, synaptic striatum, and hippocampus of 38 samples (19 schizophrenia patients and 19 healthy controls) from cadaver brains.

[0053] The “GEOquery” package of R language (version 4.4.2) was used to read the raw data files, and the “dplyr” package was used for data normalization. Probes without gene annotations were removed, duplicate genes were merged, and the maximum expression value was retained. Finally, a standardized expression matrix was generated and saved as expression matrix files for the psoriasis group (including disease and control samples) and the schizophrenia group (including disease and control samples), respectively.

[0054] Step 2: Differentially expressed genes (DEGs) analysis

[0055] The R package "limma" was used for differential analysis. The screening criteria for the psoriasis group were |log2 FC|>1 and adjusted p value (adj.p)<0.05, and 841 DEGs were obtained (upregulated genes such as RIMS3 and MSL3, downregulated genes such as JUN and NR4A1); the screening criteria for the schizophrenia group were |log2 FC|>0.3 and adj.p<0.05, and 506 DEGs were obtained (upregulated genes such as CASP7 and BAG3, downregulated genes such as NCOA2 and TRAK2). The volcano plot ( Figure 2 ) to show the distribution of differentially expressed genes, and a Venn diagram ( Figure 3 ) and the intersection was taken to obtain 64 co-expressed DEGs.

[0056] Step 3: Protein-protein interaction (PPI) network construction and hub gene screening

[0057] The 64 co-expressed DEGs were input into the STRING database (https: / / string-db.org / ) to construct a PPI network. The confidence level was set to 0.4, and 23 hub genes (such as ZFP36, EGR1, and CDKN1A) were screened out. The network was visualized using Cytoscape 3.7.2 software ( Figure 4 ).

[0058] Step 4: Functional enrichment analysis

[0059] The GO and KEGG enrichment analysis of the 23 hub genes was performed using the R package "clusterProfiler". GO analysis showed that the hub genes were significantly enriched in items such as "epithelial cell differentiation regulation" (BP), "CCR4-NOT complex" (CC), and "14-3-3 protein binding" (MF); KEGG analysis showed that the hub genes were mainly enriched in pathways such as "FoxO signaling pathway" and "transcriptional dysregulation in cancer" ( Figure 5 ).

[0060] Step 5: Machine Learning Feature Selection

[0061] 5.1 LASSO regression screening

[0062] The LASSO model was constructed, and the optimal regularization parameter λ was determined by 10-fold cross validation (λ = 0.023 for the psoriasis group and λ = 0.058 for the schizophrenia group). Eleven genes with non-zero coefficients in the psoriasis group (such as PRSS3) and seven genes with non-zero coefficients in the schizophrenia group (such as GADD45B) were screened ( Figure 6 ).

[0063] LASSO regression algorithm:

[0064] enter:

[0065] Training sample X=[X1,X2,...,X n ] T (Standardization)

[0066] Continuous label y = [y1, y2, ..., y n ]

[0067] Regularization parameter λ (or selected by cross-validation)

[0068] initialization:

[0069] Characteristic coefficient vector β = [β1, β2, ..., β p ]

[0070] Algorithm flow:

[0071] Optimization objective function:

[0072] LASSO minimizes the loss function with L1 regularization:

[0073]

[0074] Use Coordinate Descent to iteratively update each coefficient β j , until convergence. Feature screening:

[0075] L1 regularization will compress some coefficients to 0 (β j =0), the features corresponding to the non-zero coefficients are retained.

[0076] Feature importance is sorted by the absolute value of the coefficient:

[0077] c j =|β j |

[0078] Generate a sorted list of features:

[0079] Press c j Sort descending:

[0080] r=argsort([c1,c2,...,c p ],reverse=True)

[0081] Output:

[0082] Non-zero coefficient feature subset and ranking r.

[0083] 5.2 SVM-RFE screening

[0084] The support vector machine recursive feature elimination (SVM-RFE) algorithm was used to screen out 9 key genes (such as MAFF) in the psoriasis group and 11 key genes (such as EPRS) in the schizophrenia group with the lowest error rate as the criterion ( Figure 7 ).

[0085] SVM-RFE algorithm:

[0086] enter:

[0087] Training sample X0 = [X1, X2, ..., X l ] T (l samples, n features)

[0088] Classification label y = [y1, y2, ..., y l ] (usually a binary label, such as disease / normal)

[0089] initialization:

[0090] Surviving feature subset S = [1, 2, ..., n] (initially contains all features)

[0091] Feature sorting list r = [] (used to store the reverse ranking of the removed features)

[0092] Algorithm flow:

[0093] Loop until all features are eliminated (that is, S is an empty set):

[0094] a. Limit training samples:

[0095] Only retain the feature columns corresponding to the current surviving feature subset S:

[0096] X=X0(:,S)

[0097] b. Train the SVM classifier:

[0098] Use (X, y) to train a linear SVM and obtain the support vector coefficient α and bias term b:

[0099] α=SVM-train(X,y)

[0100] c. Calculate the weight vector:

[0101] The feature weight w is calculated by weighting the coefficient of the support vector and the sample label:

[0102]

[0103] (In practice, w is the normal vector of the SVM hyperplane, reflecting the contribution of each feature to the classification)

[0104] d. Generate feature ranking criteria:

[0105] For each feature i∈S, calculate its importance index c i =(w i ) 2 (Squaring avoids sign interference).

[0106] e. Eliminate the least important features:

[0107] Find the least important feature f = arg min(ci), remove it from S, and insert it at the head of the sorted list r:

[0108] r=[S(f),r],S=S\{f}

[0109] Termination conditions:

[0110] The loop ends when all features are eliminated (S=[]).

[0111] Output:

[0112] A ranked list of features r (arranged in reverse order of elimination, with the most important features at the end).

[0113] 5.3 Random Forest (RF) Screening

[0114] A random forest model was constructed, and with a Gini coefficient > 1 as the threshold, 9 important genes (such as TOB1) in the psoriasis group and 5 important genes (such as ZFP36L1) in the schizophrenia group were screened ( Figure 8 ).

[0115] RF algorithm:

[0116] enter:

[0117] Training sample X = [X1, X2, ..., X n ]T(n samples, p features)

[0118] Classification / regression labels y = [y1,y2,...,y n ]

[0119] Parameters: number of decision trees T, number of feature samples m when each tree splits (usually )

[0120] initialization:

[0121] Feature importance list Importance = [0,0,...,0] (length p)

[0122] Algorithm flow:

[0123] Train a random forest model:

[0124] For each decision tree t (t=1, 2, ..., T):

[0125] a. Bootstrap the original data to generate a sub-dataset D t ,

[0126] b. When each node of the tree splits, M candidate features are randomly selected from p features,

[0127] c. Select the best splitting feature based on Gini impurity (classification) or mean squared error (regression),

[0128] d. Record the reduction in impurity ΔGini brought by feature J at each split t,j ,

[0129] Calculate feature importance:

[0130] For each feature j:

[0131]

[0132] Or based on OOB error:

[0133] a. For each tree t, calculate the prediction error using OOB samples that did not participate in training

[0134] b. Randomly shuffle the OOB sample values ​​of feature \(j\) and recalculate the error

[0135] c. Importance score:

[0136]

[0137] Generate a sorted list of features:

[0138] Arrange the features in descending order of importance score:

[0139] r=argsort(Importance,reverse=True)

[0140] Output:

[0141] List of feature importance scores Importance and ranking r.

[0142] Step 6: Core gene screening and verification

[0143] The intersection genes of three machine learning methods (LASSO, SVM-RFE, RF) were used to obtain the core genes PRSS3 and GADD45B co-expressed in psoriasis and schizophrenia ( Figure 9 ), through the box plot ( Figure 10) showed differences in their expression between the disease group and the control group: PRSS3 was significantly upregulated in psoriasis (p<0.001) and slightly downregulated in schizophrenia (p=0.03); GADD45B was significantly downregulated in psoriasis (p<0.001) and slightly upregulated in schizophrenia (p=0.02).

[0144] Step 7: Diagnostic Capability Assessment

[0145] The ROC curve was used to evaluate the diagnostic efficacy of the core genes. In the psoriasis group, PRSS3 (AUC = 0.909) and GADD45B (AUC = 0.924) showed high diagnostic value; in the schizophrenia group, PRSS3 (AUC = 0.786) and GADD45B (AUC = 0.853) also performed well ( Figure 11 ), verified by external datasets (GSE6710 and GSE21138), the diagnostic AUCs of PRSS3 and GADD45B were both > 0.7 ( Figure 12 ), confirming its reliability across datasets.

[0146] Example 2: A method for screening core genes co-causative for psoriasis and schizophrenia

[0147] Step 1: Data acquisition and preprocessing

[0148] From the Gene Expression Omnibus (GEO),

[0149] The psoriasis gene chip dataset (GSE14905) and the schizophrenia gene chip dataset (GSE53987) were downloaded from https: / / www.ncbi.nlm.nih.gov / geo / . Both datasets are based on the GPL570 platform (HG-U133_Plus_2). GSE14905 contains 82 samples, including skin biopsy samples from 21 healthy controls and samples from 61 psoriasis patients (including lesion and non-lesion tissues); GSE53987 contains gene expression data of the prefrontal cortex, synaptic striatum, and hippocampus of 38 samples (19 schizophrenia patients and 19 healthy controls) from cadaver brains.

[0150] The “GEOquery” package of R language (version 4.4.2) was used to read the raw data files, and the “dplyr” package was used for data normalization. Probes without gene annotations were removed, duplicate genes were merged, and the maximum expression value was retained. Finally, a standardized expression matrix was generated and saved as expression matrix files for the psoriasis group (including disease and control samples) and the schizophrenia group (including disease and control samples), respectively.

[0151] Step 2: Differentially expressed genes (DEGs) analysis

[0152] The R package "limma" was used for differential analysis. The screening criteria for the psoriasis group were |log2 FC|>1 and adjusted p value (adj.p)<0.05, and 841 DEGs were obtained (upregulated genes such as RIMS3 and MSL3, downregulated genes such as JUN and NR4A1); the screening criteria for the schizophrenia group were |log2 FC|>0.3 and adj.p<0.05, and 506 DEGs were obtained (upregulated genes such as CASP7 and BAG3, downregulated genes such as NCOA2 and TRAK2). The volcano plot ( Figure 2 ) to show the distribution of differentially expressed genes, and a Venn diagram ( Figure 3 ) and the intersection was taken to obtain 64 co-expressed DEGs.

[0153] Step 3: Protein-protein interaction (PPI) network construction and hub gene screening

[0154] The 64 co-expressed DEGs were input into the STRING database (https: / / string-db.org / ) to construct a PPI network. The confidence level was set to 0.4, and 23 hub genes (such as ZFP36, EGR1, and CDKN1A) were screened out. The network was visualized using Cytoscape 3.7.2 software ( Figure 4 ).

[0155] Step 4: Functional enrichment analysis

[0156] The GO and KEGG enrichment analysis of the 23 hub genes was performed using the R package "clusterProfiler". GO analysis showed that the hub genes were significantly enriched in items such as "epithelial cell differentiation regulation" (BP), "CCR4-NOT complex" (CC), and "14-3-3 protein binding" (MF); KEGG analysis showed that the hub genes were mainly enriched in pathways such as "FoxO signaling pathway" and "transcriptional dysregulation in cancer" ( Figure 5 ).

[0157] Step 5: Machine Learning Feature Selection

[0158] 5.1 LASSO regression screening

[0159] The LASSO model was constructed, and the optimal regularization parameter λ was determined by 5-fold cross validation (λ = 0.023 for the psoriasis group and λ = 0.058 for the schizophrenia group). Eleven genes with non-zero coefficients in the psoriasis group (such as PRSS3) and seven genes with non-zero coefficients in the schizophrenia group (such as GADD45B) were screened ( Figure 6 ).

[0160] LASSO regression algorithm:

[0161] enter:

[0162] Training sample X=[X1,X2,...,X n ] T (Standardization)

[0163] Continuous label y = [y1, y2, ..., y n ]

[0164] Regularization parameter λ (or selected by cross-validation)

[0165] initialization:

[0166] Characteristic coefficient vector β = [β1, β2, ..., β p ]

[0167] Algorithm flow:

[0168] Optimization objective function:

[0169] LASSO minimizes the loss function with L1 regularization:

[0170]

[0171] Use Coordinate Descent to iteratively update each coefficient β j , until convergence. Feature screening:

[0172] L1 regularization will compress some coefficients to 0 (β j =0), the features corresponding to the non-zero coefficients are retained.

[0173] Feature importance is sorted by the absolute value of the coefficient:

[0174] c j =|β j |

[0175] Generate a sorted list of features:

[0176] Press c j Sort descending:

[0177] r=argsort([c1,c2,...,c p ],reverse=True)

[0178] Output:

[0179] Non-zero coefficient feature subset and ranking r.

[0180] 5.2 SVM-RFE screening

[0181] The support vector machine recursive feature elimination (SVM-RFE) algorithm was used to screen out 9 key genes (such as MAFF) in the psoriasis group and 11 key genes (such as EPRS) in the schizophrenia group with the lowest error rate as the criterion ( Figure 7 ).

[0182] SVM-RFE algorithm:

[0183] enter:

[0184] Training sample X0 = [X1, X2, ..., X l ]T(l samples, n features)

[0185] Classification label y = [y1, y2, ..., y l ] (usually a binary label, such as disease / normal)

[0186] initialization:

[0187] Surviving feature subset S = [1, 2, ..., n] (initially contains all features)

[0188] Feature sorting list r = [] (used to store the reverse ranking of the removed features)

[0189] Algorithm flow:

[0190] Loop until all features are eliminated (that is, S is an empty set):

[0191] a. Limit training samples:

[0192] Only retain the feature columns corresponding to the current surviving feature subset S:

[0193] X=X0(:,S)

[0194] b. Train the SVM classifier:

[0195] Use (X, y) to train a linear SVM and obtain the support vector coefficient α and bias term b:

[0196] α=SVM-train(X,y)

[0197] c. Calculate the weight vector:

[0198] The feature weight w is calculated by weighting the coefficient of the support vector and the sample label:

[0199]

[0200] (In practice, w is the normal vector of the SVM hyperplane, reflecting the contribution of each feature to the classification)

[0201] d. Generate feature ranking criteria:

[0202] For each feature i∈S, calculate its importance index c i =(w i ) 2 (Squaring avoids sign interference).

[0203] e. Eliminate the least important features:

[0204] Find the least important feature f = arg min(ci), remove it from S, and insert it at the head of the sorted list r:

[0205] r=[S((f),r],S=S\{f}

[0206] Termination conditions:

[0207] The loop ends when all features are eliminated (S=[]).

[0208] Output:

[0209] A ranked list of features r (arranged in reverse order of elimination, with the most important features at the end).

[0210] 5.3 Random Forest (RF) Screening

[0211] A random forest model was constructed, and with a Gini coefficient > 1 as the threshold, 9 important genes (such as TOB1) in the psoriasis group and 5 important genes (such as ZFP36L1) in the schizophrenia group were screened ( Figure 8 ).

[0212] RF algorithm:

[0213] enter:

[0214] Training sample X = [X1, X2, ..., X n ]T(n samples, p features)

[0215] Classification / regression labels y = [y1,y2,...,y n ]

[0216] Parameters: number of decision trees T, number of feature samples m when each tree splits (usually )

[0217] initialization:

[0218] Feature importance list Importance = [0,0,...,0] (length p)

[0219] Algorithm flow:

[0220] Train a random forest model:

[0221] For each decision tree t (t=1, 2, ..., T):

[0222] a. Bootstrap the original data to generate a sub-dataset D t .

[0223] b. At each node split in the tree, m candidate features are randomly selected from the p features.

[0224] c. Select the best splitting feature based on Gini impurity (classification) or mean squared error (regression).

[0225] d. Record the reduction in impurity ΔGini brought by feature j at each split t,j .

[0226] Calculate feature importance:

[0227] For each feature j:

[0228]

[0229] Or based on OOB error:

[0230] a. For each tree t, calculate the prediction error using OOB samples that did not participate in training

[0231] b. Randomly shuffle the OOB sample values ​​of feature \(j\) and recalculate the error

[0232] c. Importance score:

[0233]

[0234] Generate a sorted list of features:

[0235] Arrange the features in descending order of importance score:

[0236] r=argsort(Importance,reverse=True)

[0237] Output:

[0238] List of feature importance scores Importance and ranking r.

[0239] Step 6: Core gene screening and verification

[0240] The intersection genes of three machine learning methods (LASSO, SVM-RFE, RF) were used to obtain the core genes PRSS3 and GADD45B co-expressed in psoriasis and schizophrenia ( Figure 9 ), through the box plot ( Figure 10) showed differences in their expression between the disease group and the control group: PRSS3 was significantly upregulated in psoriasis (p<0.001) and slightly downregulated in schizophrenia (p=0.03); GADD45B was significantly downregulated in psoriasis (p<0.001) and slightly upregulated in schizophrenia (p=0.02).

[0241] Step 7: Diagnostic Capability Assessment

[0242] The ROC curve was used to evaluate the diagnostic efficacy of the core genes. In the psoriasis group, PRSS3 (AUC = 0.909) and GADD45B (AUC = 0.924) showed high diagnostic value; in the schizophrenia group, PRSS3 (AUC = 0.786) and GADD45B (AUC = 0.853) also performed well ( Figure 11 ), verified by external datasets (GSE6710 and GSE21138), the diagnostic AUCs of PRSS3 and GADD45B were both > 0.7 ( Figure 12 ), confirming its reliability across datasets.

[0243] References

[0244] [1] Korman, NJ Management of psoriasis as a systemic disease: what is the evidence? Br J Dermatol 182,840–848(2020).

[0245] [2]Guo,J.et al.Signaling pathways and targeted therapies forpsoriasis.Signal Transduct Target Ther 8,437(2023).

[0246] [3] Parisi, R. et al. National, regional, and worldwide epidemiology of psoriasis: systematic analysis and modeling study. BMJ 369, (2020).

[0247] [4]Damiani,G.et al.The Global,Regional,and National Burden ofPsoriasis:Results and Insights From the Global Burden of Disease2019Study.Front Med(Lausanne)8,743180(2021).

[0248] [5]Parisi,R.et al.National,regional,and worldwide epidemiology ofpsoriasis:systematic analysis andmodelling study.BMJ 369,m1590(2020).

[0249] [6]Marder,S.R.&Cannon,T.D.Schizophrenia.N Engl JMed 381,1753–1761(2019).

[0250] [7]Charlson,F.J.et al.Global Epidemiology and Burden ofSchizophrenia:Findings From the Global Burden ofDisease Study2016.SchizophrBull 44,1195–1203(2018).

[0251] [8]Huo,C.et al.Abnormalities in behaviour,histology and prefrontalcortical gene expression profiles relevant to schizophrenia in embryonic day17MAM-Exposed C57BL / 6mice.Neuropharmacology 140,287–301(2018).

[0252] [9]Qi,M.,Zhu,P.,Wang,H.,He,Q.&Huo,C.Abnormalities in behaviorrelevant to schizophrenia in embryonic day 17 MAM-exposed rodent models:Asystematic review and meta-analysis.Pharmacology Biochemistry andBehavior245,173888(2024).

[0253]

[10] Kahn, RS et al. Schizophrenia. Nat Rev Dis Primers 1, 15067 (2015).

[0254]

[11] Peixin Zhu, Qi He, Xiyan Liu, Chunyue Huo. Risk of Schizophrenia in patients with Psoriasis. Dermatologic Therapy. 2023(5)1:10.

[0255]

[12] Luina, Gao Ao, Zhao Qi, et al. Study on the transcriptional level of amygdala in social isolation model mice with schizophrenia[J]. Chinese Journal of Neurology, 2023, 22(07): 649-656. DOI: 10.3760 / cma.j.cn115354-20230504-00260

[0256]

[13] Ma L, Shcherbina A, Chetty S.Variations and expression features ofCYP2D6 contribute to schizophrenia risk.Mol Psychiatry.2021 Jun;26(6):2605-2615.doi:10.1038 / s41380-020-0675-y.Epub 2020 Feb 11.PMID:32047265; PMCID:PMC8440189.

[0257]

[14] Jo YT,Joo SW,Shon SH,Kim H,Kim Y,Lee J.Diagnosing schizophreniawith network analysis and amachine learning method.Int JMethods PsychiatrRes.2020 Mar;29(1):e1818.doi:10.1002 / mpr.1818.Epub 2020Feb 5.PMID:32022360;PMCID:PMC7051840.

Claims

1. A method for screening core genes that are co-causative for psoriasis and schizophrenia, characterized in that: The following steps are involved: Step 1. Obtain gene chip expression profile data of several cell types in psoriasis and schizophrenia from public databases and perform normalization; Step 2. Differentially expressed genes (DEGs) between the two diseases were screened by differential analysis, and the intersection was used to obtain co-expressed DEGs; Step 3. Construct a protein interaction network (PPI) of co-expressed DEGs and screen hub genes; Step 4. Perform GO and KEGG pathway enrichment analysis on hub genes; Step 5. Three machine learning methods, Least Absolute Shrinkage and Selection Operator (LASSO), Support Vector Machine Recursive Feature Elimination (SVM-RFE), and Random Forest (RF), were used on the hub genes in the gene chip expression profile data of psoriasis and schizophrenia, respectively. The three machine learning results of psoriasis and schizophrenia were matched one by one and the intersection was taken to screen the key co-expressed DEGs. The key co-expressed DEGs that overlapped among the three machine learning results were then selected as the co-pathogenic core genes. Step 6. Use ROC curve to verify the disease diagnostic efficacy of these co-pathogenic core genes.

2. The method for screening core genes co-causative for psoriasis and schizophrenia according to claim 1, characterized in that: The psoriasis cell type datasets in step 1 include gene expression profile data of normal skin cell biopsy samples from normal healthy donors and skin cell biopsy samples from psoriasis patients, and psoriasis patients have matched gene expression profile data of lesion and non-lesion tissue cells; the schizophrenia cell type datasets include brain cell gene expression profile data of the prefrontal cortex, associative striatum, and hippocampus of cadaver brains from schizophrenia patients and controls.

3. The method for screening core genes co-causative for psoriasis and schizophrenia according to claim 1, characterized in that: DEGs screening in step 2 was completed using the R language "Limma" package. The threshold for psoriasis was |log2 FC|>1 and the adjusted P value <0.05, and the threshold for schizophrenia was |log2FC|>0.3 and the adjusted P value <0.

05.

4. The method for screening core genes co-causative for psoriasis and schizophrenia according to claim 1, characterized in that: The specific steps of screening hub genes described in step 3 are to input the co-expressed DEGs obtained after differential analysis into the STRING database (https: / / string-db.org / ), the confidence level was set to 0.4 to screen out hub genes, construct the protein-protein interaction network of differentially expressed genes, and use Cytoscape 3.7.2 software to construct a visualized PPI network.

5. The method for screening core genes co-causative for psoriasis and schizophrenia according to claim 1, characterized in that: The intersection analysis of the three machine learning algorithms described in steps 5 and 6 was used to improve the screening specificity. The lambda value of LASSO regression was determined by 5-10 fold cross validation. SVM-RFE was used to screen the optimal gene combination by recursive feature elimination. Random forest was used to select the best tree with a number of 500 and an importance threshold of 0.8-1.

6. The method for screening core genes co-causative for psoriasis and schizophrenia according to claim 1, characterized in that: In step 7, the ROC curve is used to obtain the AUC value to verify the efficacy of the co-pathogenic core gene in diagnosing the disease.