A comorbidity mechanism identification method and system based on complement-coagulation axis and ceRNA network

By combining transcriptome data and ceRNA network construction, the comorbidity mechanism of endometriosis and subacute cutaneous lupus erythematosus was identified, which solved the problem of difficulty in accurately identifying the comorbidity mechanism in the existing technology and realized reproducible and generalizable molecular mechanism identification and analysis.

CN122157775APending Publication Date: 2026-06-05FOSHAN MATERNAL & CHILD HEALTH CARE HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOSHAN MATERNAL & CHILD HEALTH CARE HOSPITAL
Filing Date
2026-03-02
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies lack systematic identification and analysis methods for the comorbidity mechanisms between endometriosis and subacute lupus erythematosus, making it difficult to achieve accurate identification and analysis of comorbidity mechanisms.

Method used

By combining transcriptome data, differential expression analysis, immune cell infiltration modeling, complement-coagulation pathway enrichment screening, machine learning feature extraction, and ceRNA network construction are performed to form a reproducible and scalable molecular mechanism identification process to identify the comorbidity mechanism between the two diseases.

Benefits of technology

This study has enabled the accurate identification of the comorbidity mechanism between endometriosis and subacute lupus erythematosus, providing a reproducible and scalable molecular mechanism identification and analysis method, and offering technical support for the study of comorbidity mechanisms in autoimmune diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157775A_ABST
    Figure CN122157775A_ABST
Patent Text Reader

Abstract

The application relates to a comorbidity mechanism identification method and system based on a complement-coagulation axis and a ceRNA network, and relates to the technical field of biomedical testing. The method comprises the following steps: integrating transcriptome expression characteristics of different diseases and performing immune pathway analysis, taking the complement-coagulation axis as a breakthrough point, identifying an inflammatory signal module shared by endometriosis and subacute cutaneous type red lupus, and further screening candidate target points with cross expression characteristics and immune regulation functions. At the regulation mechanism level, a circRNA-miRNA-mRNA ternary regulation network containing a circRNA node is constructed to obtain more complete non-coding regulation axis information; meanwhile, a plurality of algorithms are introduced to cooperatively screen and verify an independent data set, a repeatable, verifiable and expandable molecular mechanism mining process is formed, and thus the comorbidity mechanism identification analysis between diseases is realized, and technical support and path reference are provided for comorbidity mechanism research of autoimmune diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical testing technology, and in particular to a method and system for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network. Background Technology

[0002] Current research on the potential association between endometriosis (EMs) and lupus-related diseases is largely based on retrospective epidemiological studies, suggesting that EMs patients may have a higher risk of systemic autoimmune diseases. Existing comorbidity mechanisms research mainly focuses on systemic lupus erythematosus (SLE), while systematic studies on the potential association between its cutaneous subtype—subacute cutaneous lupus erythematosus (SCLE)—and EMs are relatively insufficient. SCLE shares some mechanistic overlap with SLE in terms of immune abnormalities, inflammatory signal activation, and complement system involvement, and its tissue origin is more clearly defined, making it a feasible model for mechanistic research. However, current techniques still have significant limitations.

[0003] Specifically, existing technologies mostly focus on differential analysis of single diseases or screening using single algorithms, lacking cross-disease joint modeling and mechanism mining based on systematic transcriptome data, as well as systematic integration and external validation of key pathways and regulatory networks. This results in limited reproducibility and generalizability of screening results, making it difficult to fully reflect the complexity of common disease mechanisms.

[0004] Therefore, existing technologies lack systematic identification and analysis methods for the comorbidity mechanisms between EMs and SCLE, making it difficult to achieve accurate identification and analysis of comorbidity mechanisms. Summary of the Invention

[0005] This application provides a method and system for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network. By combining transcriptome data from two diseases, and through differential expression analysis, immune cell infiltration modeling, complement-coagulation pathway enrichment screening, machine learning feature extraction, and ceRNA network construction, a reproducible and scalable molecular mechanism identification process is formed, enabling the identification and analysis of comorbidity mechanisms and solving the technical problem that existing technologies cannot accurately identify and analyze comorbidity mechanisms.

[0006] Firstly, this application provides a method for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network, which can be applied to the identification of comorbidity mechanisms in at least two diseases, such as EMs and SCLE, including:

[0007] The raw microarray data in the transcriptome dataset were standardized and batch effect corrected to establish a standardized expression matrix. The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease.

[0008] Based on the expression matrix, pathway activity scores and common disease pathway features were extracted from the first and second training sets to screen for differential pathways.

[0009] Based on the expression matrix, a weighted co-expression network was constructed using the expression profiles of the first validation set and skin tissue to identify key expression modules between the two diseases.

[0010] GO enrichment analysis and KEGG pathway enrichment analysis were performed on key expression modules to screen significant entries. In addition, differential expression analysis was performed on the first and second training sets to screen differentially expressed genes shared by the two diseases.

[0011] Pathway-related genes were extracted from the obtained complement and coagulation pathways, and differentially expressed genes were combined to screen for comorbidity hub genes. Expression trend analysis and diagnostic efficacy assessment were performed using the first and second validation sets to identify key genes.

[0012] Immune cell lineage characteristics and interrelationships were analyzed on the first and second training sets to obtain immune cell infiltration results, which included immune cell lineage characteristics and interaction patterns of the two diseases.

[0013] Based on the results of immune cell infiltration, the expression patterns of key genes in two diseases were identified as immune cell-related, revealing the expression correlation of immune features.

[0014] Based on expression correlation, immune pathway activity analysis and key gene regulation assessment were conducted according to key genes and immune cell lineage characteristics to obtain the downstream participation mode of key genes.

[0015] Based on the established ceRNA ternary network, the upstream regulatory network of key genes was identified, and a comorbidity identification mechanism was established by combining downstream participation patterns.

[0016] Secondly, this application proposes a comorbidity mechanism identification system based on the complement-coagulation axis and ceRNA network, including:

[0017] The data acquisition module is used to standardize and correct batch effects on the raw microarray data in the transcriptome dataset, and to establish a standardized expression matrix. The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease.

[0018] The differential pathway analysis module is used to perform pathway activity scoring and extract common disease pathway features from the first and second training sets based on the expression matrix, and to screen differential pathways.

[0019] The co-expression analysis module is used to construct a weighted co-expression network based on the expression matrix of the first validation set and the expression profile of skin tissue, and to identify the key expression modules between the two diseases.

[0020] The differential expression analysis module is used to perform GO enrichment analysis and KEGG pathway enrichment analysis on key expression modules to screen significant entries, and to perform differential expression analysis on the first and second training sets to screen differentially expressed genes shared by the two diseases.

[0021] The key gene identification module is used to extract pathway-related genes from the obtained complement and coagulation pathways, screen comorbidity hub genes by combining differentially expressed genes, and use the first and second validation sets to perform expression trend analysis and diagnostic efficacy assessment to identify key genes.

[0022] The immune infiltration module is used to analyze the immune cell lineage characteristics and interactions of the first and second training sets to obtain immune cell infiltration results, which include the immune cell lineage characteristics and interaction patterns of the two diseases.

[0023] The expression correlation analysis module is used to identify the immune cell-related expression patterns of key genes in two diseases based on the results of immune cell infiltration, and to obtain expression correlations that reveal immune characteristics.

[0024] The downstream participation pattern analysis module is used to analyze the activity of immune pathways and evaluate the regulatory role of key genes based on expression correlation and the characteristics of key genes and immune cell lineages, so as to obtain the downstream participation pattern of key genes.

[0025] The comorbidity identification module is used to identify the upstream regulatory network of key genes based on the established ceRNA ternary network, and to establish a comorbidity identification mechanism by combining downstream participation modes.

[0026] In summary, this application addresses the technical problem of the lack of accurate identification of comorbidity mechanisms between EMs and SCLE in existing technologies. First, it utilizes transcriptomic data for pathway and differential expression matrix analysis to identify enriched pathways and differentially expressed genes between the two diseases. Then, a weighted co-expression network is constructed, and expression clustering is used to identify functional aggregation features and pathway involvement, laying the foundation for core gene extraction and network construction. Next, key characteristic genes are screened, and a dual-algorithm screening of complement-coagulation comorbidity hub genes is performed using a model set. Expression trends and efficacy of comorbidity hub genes are evaluated to identify potential cross-disease diagnostic biomarkers, thereby determining target key genes. Finally, transcriptomic data analysis is used to analyze the immune cell lineage characteristics and correlations of each disease, exploring the compositional characteristics of the immune microenvironment and intercellular interaction patterns between the two diseases, revealing the expression correlation of immune features. Furthermore, by comparing the activity differences of immune signaling pathways using expression correlations and target key genes, disease-specific activation pathways are identified, immune pathway activity analysis and key gene regulatory roles are evaluated, downstream participation patterns are established, and a regulatory network identification of circRNA–miRNA–mRNA ceRNA networks is established to trace upstream regulatory mechanisms and identify potential comorbidity regulatory axes. Finally, using downstream participation patterns and upstream regulatory networks, a comorbidity identification mechanism is established. Based on a reproducible, verifiable, and scalable molecular mechanism mining process, the identification and analysis of comorbidity mechanisms among diseases are realized, providing technical support and pathway reference for the study of comorbidity mechanisms in autoimmune diseases, and solving the technical problem that existing technologies cannot accurately identify and analyze comorbidity mechanisms. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0028] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 A flowchart illustrating a method for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network, provided for embodiments of this application;

[0030] Figure 2 This is a graph showing the GSVA pathway enrichment analysis results for the GSE7305 (EMs) and GSE95474 (SCLE) datasets, provided as an example in this application.

[0031] Figure 3This is an example illustration provided in this application of gene modules identified based on WGCNA analysis that are significantly associated with EMs and SCLE;

[0032] Figure 4 This is a functional enrichment analysis plot of GO and KEGG for a key co-expression module provided as an example in this application;

[0033] Figure 5 This is an example of a diagram illustrating the identification and intersection analysis of differentially expressed genes between EMs and SCLE provided in this application;

[0034] Figure 6 This is a functional enrichment analysis diagram of differentially expressed genes in EMs and SCLEs provided as an example in this application;

[0035] Figure 7 This application provides an example of a machine learning algorithm-based screening map of key genes co-expressed in the coagulation cascade.

[0036] Figure 8 This is an example of an application providing a validation diagram of the co-expression of coagulation and complement-related genes (COL3A1, C1S, C1R);

[0037] Figure 9 This is an example of an association analysis diagram of immune cell infiltration provided in this application;

[0038] Figure 10 This is a correlation analysis diagram of immune cells and the expression levels of COL3A1 and C1S genes provided in an example of this application;

[0039] Figure 11 This is a diagram illustrating the differentially expressed immune-related signaling pathways in EMs (GSE7305) and SCLEs (GSE95474) and their association with key genes COL3A1 and C1S, provided as an example in this application.

[0040] Figure 12 This is a diagram of the circRNA–miRNA–mRNA (ternary ceRNA) regulatory network provided as an example in this application;

[0041] Figure 13 This is a schematic diagram of a comorbidity mechanism identification system based on the complement-coagulation axis and ceRNA network, provided in an optional embodiment of this application. The modules in system 1300 are used to implement... Figure 1 The corresponding processing functions in steps 110 to 190 of the method shown. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] To facilitate understanding of the embodiments of this application, further explanations and descriptions will be provided below in conjunction with the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of this application.

[0044] The following section, based on the establishment of the comorbidity mechanism between EMs and SCLE, details the specific implementation process.

[0045] Figure 1 This document provides a flowchart illustrating a method for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network, applicable to the identification of comorbidity mechanisms in at least two diseases. This embodiment uses the identification of comorbidity mechanisms in EMs and SCLE as an example for detailed explanation. Specifically, the method may include the following steps:

[0046] Step 110: Standardize and batch effect correct the raw microarray data in the transcriptome dataset to establish a standardized expression matrix.

[0047] The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease.

[0048] In this embodiment, a transcriptome dataset can be obtained from a preset data source. This transcriptome dataset mainly consists of publicly available transcriptome data related to at least two diseases. For example, taking EMs and SCLE as examples, five high-quality publicly available transcriptome datasets related to EMs and SCLE can be obtained from the GEO database. Taking EMs as the first disease and SCLE as the second disease as an example (the following examples focus on identifying the comorbidity mechanism between EMs and SCLE), the dataset may include:

[0049] GSE7305 (EMs original training set, i.e., EMs training set): contains endometrial samples from 10 EMs patients and normal endometrial samples from 10 non-ectopic individuals; GSE95474 (SCLE training set): includes skin tissue from 5 SCLE patients and skin tissue from 5 healthy individuals; GSE120103 (EMS validation set): contains endometrial samples from 18 healthy individuals and endometrial tissue from 18 patients with ovarian EMs (rASRMIV stage); GSE184989 (SCLE skin tissue expression profile): includes skin tissue expression profiles from 43 SCLE patients and 13 healthy individuals; GSE81071 (SCLE validation set): independently collected skin samples from 13 healthy individuals and 43 SCLE patients, which can be used as an external validation set.

[0050] The collected transcriptome datasets can be preprocessed as raw microarray data, including standardization and batch effect correction, to establish standardized expression matrices, serving as a unified basis for subsequent comorbidity mechanism identification and analysis. Specifically, each type of transcriptome dataset can have a corresponding expression matrix constructed; for example, the five high-quality datasets mentioned above can generate at least five corresponding standardized expression matrices.

[0051] Optionally, the raw microarray data in the transcriptome dataset can be standardized and batch effect corrected to establish a standardized expression matrix. Specifically, this may include: obtaining the transcriptome dataset from the data source; performing background correction and normalization on the transcriptome dataset using the RMA algorithm, and performing unified probe mapping in conjunction with annotation information to update the transcriptome dataset; performing correction processing on the transcriptome dataset using the sva package, and removing low-expression probes and missing samples to construct a standardized expression matrix; wherein each expression matrix corresponds to each data point in the transcriptome dataset.

[0052] In related technologies, existing technologies for identifying comorbidity mechanisms among different diseases have the following significant limitations:

[0053] ① Existing technologies still fall short in providing high-throughput support at the mechanistic level: Although some existing technologies, such as epidemiological studies, have suggested a possible comorbidity trend between EMs and SCLE, most rely on questionnaires or case-control designs and lack in-depth mechanistic exploration based on systematic transcriptome data, making it difficult to provide stable target identification evidence at the molecular level. ② Existing analytical methods are relatively limited, and cross-disease integration is insufficient: Most existing technologies focus on differential analysis of single diseases, and the application of cross-disease joint modeling and machine learning-assisted screening is still not widespread, which may affect the generalization ability and stability of the results. ③ The coverage of regulatory dimensions is somewhat limited: Current research on the cross-mechanism between EMs and SCLE has not systematically integrated multiple key pathways such as immune microenvironment characteristics, complement-coagulation signaling, and ceRNA regulatory networks, making it difficult to fully reflect the complexity of the common mechanisms of the diseases. ④ There is still room for improvement in model construction and validation strategies: Some regulatory network studies are still based on a single dataset and lack external validation or cross-platform support; at the same time, most non-coding RNA regulatory axes are concentrated on miRNA and lncRNA, and the inclusion of elements such as circRNA is insufficient, which may limit the complete analysis of regulatory networks.

[0054] To address the aforementioned technical problems, this application primarily implements a systematic method for identifying key genes and their regulatory pathways in the comorbidity mechanisms among diseases, thereby forming a reproducible and scalable molecular mechanism identification process. This application will provide a detailed description of each step and embodiment.

[0055] In the specific implementation, the raw microarray data in the transcriptome dataset undergoes preprocessing. For example, the Robust Multi-array Average (RMA) algorithm can be used to standardize and correct the background of the raw microarray data, thus completing background correction and normalization. Furthermore, probe IDs can be uniformly mapped using GPL annotation information provided by the platform. Then, batch effect correction can be performed using the sva package. During correction, samples with low expression probes and missing values ​​can be removed, establishing a standardized expression matrix to ensure consistency across platforms. For example, the aforementioned five transcriptome datasets can be used to construct five standardized expression matrices, with a total sample size of over 170 cases, covering endothelial endometrial tissue of EMs in situ and normal endometrium, as well as SCLE-related skin lesions and healthy control tissues, providing sufficient data support and a stable foundation for subsequent analysis.

[0056] Step 120: Based on the expression matrix, perform pathway activity scoring and extract common disease pathway features from the first and second training sets to screen differential pathways.

[0057] This embodiment uses the expression matrix as input to score functional pathway activity and extract common disease pathway features. Based on the expression matrix obtained in the preceding steps, Gene Set Variation Analysis (GSVA) ​​is used to score pathway activity and extract common disease pathway features from the datasets of the two diseases. Differences between the disease group and the control group in different disease datasets are compared to screen for differentially expressed pathways. These differentially expressed pathways can reveal a common immune mechanism between the two diseases in a comorbid context and provide a mechanistic precursor for subsequent co-expression module extraction and focused analysis of immune-related core pathways.

[0058] In the specific implementation, based on the expression matrix, pathway activity scores and common disease pathway features are extracted from the first and second training sets to screen differential pathways. Specifically, this may include: performing gene set variation analysis on the first and second training sets based on the expression matrix; using the obtained datasets as a benchmark, processing with pathway gene set templates to construct a pathway activity score matrix; and analyzing the differences between the disease group and the control group based on the pathway activity score matrix to determine pathway expression patterns and screen differential pathways.

[0059] First, GSVA was used to score pathway activity in the GSE7305 and GSE95474 datasets. Then, Reactome and KEGG (Kyoto Encyclopedia of Genes and Genomes) databases were used as pathway gene set templates (also known as reference gene sets) to construct a pathway activity scoring matrix for all samples, thus achieving pathway activity scoring for the samples. Reactome and KEGG databases can be used as pathway gene set templates.

[0060] Next, differential comparisons are performed based on the pathway activity score matrix, including comparing differential pathway expression patterns between the disease group and the control group. Differential pathways are screened according to preset screening thresholds or criteria to extract common pathway features of the disease. Preferably, the screening threshold is: |logFC| > 0.5 and P < 0.05.

[0061] For example, taking EMs and SCLE as examples, refer to Figure 2 As shown, Figure 2 The results of GSVA pathway enrichment analysis of the GSE7305 (EMs) and GSE95474 (SCLE) datasets are presented, including the main pathway activity scores of the two datasets. Figure 2 (A) and Figure 2 (B) Heatmaps showing the significantly different pathways in samples of endometriosis and subacute cutaneous lupus erythematosus, respectively; Figure 2 (C) and Figure 2(D) is a bar chart showing statistically significant pathway enrichment in the two types of diseases mentioned above.

[0062] As can be seen from the actual experiments, the differential pathway screening results indicate that 92 differential pathways were enriched in the GSVA analysis of EMs samples, and 49 pathways were significantly enriched in the SCLE samples. These pathways mainly involve metabolic regulation, chronic inflammatory response, and immune activation. Ranked by logFC value, the top 9 most significantly enriched pathways in endometriosis are: Systemic Lupus Erythematosus pathway, Other glycan degradation, Asthma, Aminoacyl tRNA biosynthesis, Proteasome, DNA replication, Protein export, Biosynthesis of valine, leucine, and isoleucine, and Homologous recombination (see...). Figure 2 (As shown in (A) and (C)). In SCLE samples, the top nine most significantly enriched pathways were: ribosome, primary immunodeficiency, graft-versus-host disease, homologous recombination, allograft rejection, cytosolic DNA sensing pathway, systemic lupus erythematosus pathway, apoptosis, and type I diabetes mellitus (see...). Figure 2 (As shown in (B) and (D)). This suggests that SLE-related immune abnormalities may constitute a common mechanistic basis for the two diseases. Furthermore, typical immune pathways such as graft-versus-host disease, primary immunodeficiency, IL-17, and type 1 diabetes are also synchronously active in both diseases, reinforcing the rationale for "systemic immune activation" as a potential comorbid link between endometriosis and SCLE. Therefore, these results reveal consistent activation in immune-related pathways such as the systemic lupus erythematosus pathway, homologous recombination, and graft-versus-host disease, suggesting a possible shared immune mechanism in their comorbid context. This provides a mechanistic precursor for subsequent co-expression module extraction and focused analysis of core immune-related pathways, thus providing data and a theoretical basis for subsequent steps.

[0063] It should be noted that the expression matrix corresponds to the transcriptome data. Therefore, when performing various processing on the transcriptome data based on the expression matrix in this application, and when processing in combination with the expression matrix, the corresponding expression matrix and transcriptome data are usually used for processing. For example, EMs transcriptome data are usually processed in association with their corresponding expression matrix.

[0064] Step 130: Based on the expression matrix, a weighted co-expression network is constructed for the first validation set and the skin tissue expression profile to identify the key expression modules between the two diseases.

[0065] In practical implementation, based on each expression matrix, WGCNA (Weighted Correlation Network Analysis) can be used to process samples in the first validation set and the skin tissue expression profile to construct a WGCNA weighted co-expression network. Then, co-expression modules are identified in the weighted co-expression network, and key functional modules highly correlated with the disease phenotype are selected as key expression modules. These key expression modules are closely related to the pathological phenotype.

[0066] In the specific implementation, the above-mentioned construction of a weighted co-expression network based on the expression matrix of the first validation set and the expression profile of skin tissue to identify key expression modules between the two diseases can specifically include: constructing weighted gene co-expression networks for the first validation set and the expression profile of skin tissue based on the expression matrix and using a weighted co-expression network construction method; identifying co-expression modules for the weighted gene co-expression network using a dynamic pruning algorithm, calculating the Pearson correlation coefficient between the co-expression modules and the disease state, and screening key expression modules that are highly correlated with the disease phenotype.

[0067] In this embodiment, using the expression matrices corresponding to GSE120103 and GSE184989 as a basis, WGCNA weighted gene co-expression networks are constructed for the samples of GSE120103 and GSE184989 respectively, to obtain the weighted gene co-expression networks of EMs and SCLE. In the dynamic pruning algorithm identification module, the dynamic pruning algorithm is used to identify the co-expression modules in the weighted co-expression network.

[0068] For each co-expression module, the Pearson correlation coefficient between the module and the disease state is calculated. Based on the correlation coefficient, key functional modules that are highly correlated with the disease phenotype are selected from all co-expression modules as key expression modules. A preset screening threshold can be used when selecting modules highly correlated with the disease phenotype.

[0069] For example, refer to Figure 3 As shown, Figure 3This study demonstrates gene modules that are significantly associated with EMs and SCLEs based on WGCNA analysis, including co-expression modules constructed by weighted gene co-expression network analysis in the GSE120103 (EMs) and GSE184989 (SCLE) datasets and their correlation with phenotypes. Figure 3 (A) and Figure 3 (D) represents the network topology analysis results under different soft thresholds, used to determine the optimal parameters for constructing scale-free networks; Figure 3 (B) and Figure 3 (E) is a gene clustering tree diagram, with each module represented by a different color; Figure 3 (C) and Figure 3 (F) is a heatmap showing the correlation between modules and phenotypes, displaying the correlation coefficients and P-values ​​between each module and the disease state.

[0070] As can be seen from the actual experimental analysis, nine gene modules (i.e., co-expression modules) can be identified in EMs. Among them, the pink, red, and brown gene modules are significantly positively or strongly correlated with disease status (where the threshold for screening positive correlation can be set as: |correlation coefficient| > 0.75, P < 0.01) (see [link to relevant documentation]). Figure 3 (A)-(C) are shown), involving processes such as hormone responses and extracellular matrix remodeling; eight gene modules can be identified in SCLE (see Figure 3 (D)-(F) are shown. Among them, the modules such as black, red and yellow are involved in interferon signaling and skin barrier function, suggesting that the two diseases may have common immune regulation modules. Analysis shows that these modules are closely related to the lesion phenotype, and functional enrichment analysis suggests that they include immune response, interferon signaling and skin barrier-related pathways.

[0071] Step 140: Perform GO enrichment analysis and KEGG pathway enrichment analysis on key expression modules to screen significant entries; and perform differential expression analysis on the first and second training sets to screen differentially expressed genes shared by the two diseases.

[0072] In this embodiment, based on the key expression modules, pathway enrichment analysis is further performed on the module genes of the key expression modules to screen significant entries, which can serve as the basis for subsequent core gene extraction and network construction.

[0073] Specifically, the key expression modules were subjected to enrichment analysis of the three major categories of GO (Gene Ontology) (such as BP, MF, CC) and KEGG pathway using the clusterProfiler (v4.2.2) of the R package. Significant entries could be selected according to preset thresholds (such as P < 0.05) to clarify the functional clustering characteristics of the modules. Significant entries can be used to identify the functional clustering characteristics and pathway participation of each module.

[0074] For example, refer to Figure 4 The diagram shows the GO and KEGG functional enrichment analysis of key co-expressed modules. It primarily displays the functional annotation results of modules related to disease state selected by WGCNA. Among them, Figure 4 (A) shows the GO and KEGG enrichment analysis results of the MEpink, MEred, and MEbrown modules in the EMs dataset; Figure 4 (B) shows the GO and KEGG enrichment analysis results of the MEblack, MEred, and MEyellow modules in the SCLE dataset.

[0075] As can be seen, the enrichment analysis results of the three major categories and the KEGG pathway in the actual experiment show that:

[0076] In the EMs samples, the corresponding MEpink module (47 genes) was significantly enriched in low-density lipoprotein receptor particle synthesis and metabolism, foam cell differentiation inhibition, sodium ion transmembrane transport, receptor ligand activity, viral receptor recognition, and hijacking molecule function (BP: regulation of synaptic transmission, MF: virus receptor activity, anion: sodium symporter activity), suggesting that it may be involved in metabolic-immune cross-linking and virus recognition pathways. Thus, the MEpink module focuses on immune-metabolic cross-linking pathways such as lipoprotein metabolism, foam cell differentiation, and virus recognition.

[0077] The MEred module (98 genes) was significantly enriched in photoreceptor activity, gap junction channels, non-membrane spanning protein tyrosine kinase activity, and hydrolase activity (MF: hydrolase activity on 12C-N bonds, CC: cytoplasmic side of plasma membrane), suggesting its association with transmembrane signal transduction and local intercellular communication. Therefore, the MEred module is related to photoreceptor activity, gap junctions, and non-membrane tyrosine kinase signal transduction processes.

[0078] The MEbrown module (181 genes) is clustered in extracellular matrix remodeling, glycosaminoglycan / heparin binding, female pregnancy regulation, response to steroid hormones, platelet dense granules, and vesicle lumen structure (BP: response to steroid hormone, CC: platelet dense granule lumen, MF: extracellular matrix structural constituent), suggesting its potential involvement in endometrial remodeling and hormone regulation pathways (see [link to relevant documentation]). Figure 4 (As shown in (A)). It can be seen that the MEbrown module is significantly enriched in extracellular matrix remodeling, hormone response and platelet granule function.

[0079] In SCLE samples, the MEblack module was significantly enriched in keratinocyte differentiation, epidermal cell development, skin barrier formation (cornified envelope), Toll-like receptor binding, IL-17 signaling pathway, and pathogen infection (such as Staphylococcus aureus infection and amebiasis) (BP: epidermal cell differentiation, KEGG: Staphylococcus aureus infection, IL-17 signaling pathway), suggesting a close relationship with innate skin immunity and susceptibility to infection.

[0080] The MEred module is strongly enriched in antiviral defense responses, type I interferon signaling pathways, viral infections (influenza A, hepatitis C), double-stranded RNA binding, and cytokine / chemokine activity (BP: response to virus, MF: cytokine activity, KEGG: NOD-like receptor signaling pathway), representing typical systemic inflammation and antiviral immune features in SCLE. Therefore, the MEred module is associated with type I interferon signaling and antiviral responses.

[0081] The MEyellow module is primarily enriched in T cell activation and regulation, lymphocyte activity regulation, membrane microdomain structure, and G protein-coupled receptor signaling pathways (BP: T cell activation, CC: membrane microdomain, KEGG: inflammatory bowel disease, malaria, leishmaniasis) (see [link]). Figure 4 (As shown in (B)), this suggests that it may be involved in adaptive immune activation and membrane receptor-mediated signal regulation. It is evident that the MEyellow module primarily participates in T cell activation and membrane receptor signal regulation.

[0082] In summary, the GO / KEGG enrichment results of the key expression modules indicate that they are closely related to immune processes such as interferon signaling, IL-17-related pathways, and T cell-mediated responses. This supports the possibility that EMs and SCLE may share the "immune activation-tissue damage" related mechanism, laying a biological foundation for the subsequent construction of core genes and networks.

[0083] When screening for significant entries, differential gene identification and co-expressed gene intersection extraction from dual-disease datasets can be performed simultaneously. Specifically, expression matrices can be used to perform differential expression analysis on disease-related datasets (including GSE7305 and GSE95474) in the transcriptome dataset to screen for differentially expressed genes shared by the diseases, such as those shared by EMs and SCLE. Enrichment analysis can reveal immune-related signaling pathways where these differentially expressed genes significantly cluster, which can serve as potential intersection points in the comorbidity mechanisms of EMs and SCLE. In specific implementation, the aforementioned differential expression analysis on the first and second training sets to screen for differentially expressed genes shared by the two diseases can specifically include: using the limma package to perform differential expression analysis on the first and second training sets based on the expression matrix to screen for differentially expressed data; and using Venn diagrams to identify differentially expressed genes shared by the two diseases for the differentially expressed data.

[0084] In this embodiment, differential expression analysis was performed on the GSE7305 and GSE95474 datasets using the limma package based on the expression matrix. Similarly, a pre-set screening threshold (e.g., |logFC| > 0.5 and P < 0.05) was used to filter differentially expressed data. Based on the filtered differentially expressed data, Venn diagrams were used to identify differentially expressed genes (DEGs) shared by the two diseases.

[0085] For example, see EMs and SCLE. Figure 5 As shown, Figure 5 The results of differentially expressed gene identification and intersection analysis between EMs and SCLE are presented, including the identification results of differentially expressed genes (DEGs) in the EMs dataset (GSE7305) and the SCLE dataset (GSE95474) and their intersection. Figure 5 (A) and (B) are the differential expression volcano plot and heatmap of the two sets of data, respectively, showing the distribution of significantly upregulated and downregulated genes; Figure 5 (C) is a Venn diagram of differentially expressed genes in the two diseases, showing the overlap between common and specifically expressed genes; Figure 5 (D) shows the intersection of co-expressed differentially expressed genes and coagulation cascade-related genes, used to screen key candidate genes that may be involved in comorbidity mechanisms.

[0086] Through actual experiments, 1118 DEGs were identified in GSE7305, of which 590 were upregulated and 528 were downregulated (see [link to experiment]). Figure 5 (A)); 1214 DEGs were identified in GSE95474, of which 783 were upregulated and 431 were downregulated (see [link]). Figure 5 (as shown in (B)). Through intersection analysis of the identified differentially expressed genes, a total of 388 DEGs were determined in the dataset (see [reference]). Figure 5 (as shown in (C)).

[0087] Reference Figure 6 As shown, Figure 6 The functional enrichment analysis of differentially expressed genes co-expressed by EMs and SCLE is presented, showing the enrichment results of the selected differentially expressed genes (388 in total) under different functional dimensions. Figure 6 (A) represents biological processes (BP). Figure 6 (B) represents cellular components (CC). Figure 6 (C) represents the three GeneOntology (GO) functional classifications of molecular function (MF), and, Figure 6 (D) shows the KEGG pathway enrichment results, involving key mechanisms such as inflammatory response, immune regulation, and cell signaling.

[0088] As can be seen, KEGG (Kyoto Encyclopedia of Genes and Genomes) enrichment analysis showed that among the 388 DEGs, they significantly clustered in multiple immune-related signaling pathways, including TNF signaling, NF-κB signaling, vitamin B6 metabolism, tyrosine metabolism, and the relaxin pathway (see details). Figure 6 (As shown in (A)-(D)). The complement and coagulation cascade pathway (KEGG: hsa04610, hsa04611) was also significantly enriched, suggesting that this pathway may serve as a potential intersection point in the comorbidity mechanism of endometriosis and SCLE, possessing potential mechanistic and interventional value. Furthermore, this pathway has also been widely reported to be closely related to inflammatory responses, immune escape, and tissue damage in immune diseases such as systemic lupus erythematosus. The results of this application further support its possible bridging role in the two diseases, providing a theoretical basis for subsequent exploration of targeted mechanisms and development of intervention strategies.

[0089] Step 150: Extract pathway-related genes from the obtained complement and coagulation pathways, screen comorbidity hub genes by combining differentially expressed genes, and use the first and second validation sets to perform expression trend analysis and diagnostic efficacy assessment to identify key genes.

[0090] In this embodiment, 205 pathway-related genes can be extracted from hsa04610 (complement pathway) and hsa04611 (coagulation pathway) in the KEGG database. The intersection of the pathway-related genes and the obtained differentially expressed genes is taken to extract the comorbidity hub genes shared by the two diseases, thereby achieving the screening of complement-coagulation comorbidity hub genes.

[0091] To further identify core candidates with potential functional importance in both diseases, expression analysis and diagnostic efficacy assessment were performed on the extracted comorbidity hub genes using a validation set. This validated the potential diagnostic value and discriminative performance of the comorbidity hub genes in both diseases, and identified key genes that can serve as cross-disease diagnostic biomarkers.

[0092] In one optional embodiment, this embodiment extracts pathway-related genes from the obtained complement and coagulation pathways, screens comorbidity hub genes by combining differentially expressed genes, and uses a first validation set and a second validation set to perform expression trend analysis and diagnostic efficacy assessment to identify key genes. Specifically, this may include: obtaining complement and coagulation pathways from the KEGG database and extracting pathway-related genes; taking the intersection of pathway-related genes and differentially expressed genes to obtain a cross-gene set including cross-genes; using a random forest model to evaluate the classification contribution of each cross-gene and selecting a first candidate gene set; and using a support vector machine model to select a second candidate gene set from each cross-gene using cross-validation; performing intersection analysis based on the first and second candidate gene sets to obtain comorbidity hub genes; using the first and second validation sets to perform expression trend analysis on comorbidity hub genes and calculate receiver operating characteristic (ROC) curves to assess the diagnostic efficacy of comorbidity hub genes in the two diseases and determine key genes.

[0093] This embodiment employs a dual-algorithm approach for screening complement-coagulation comorbidity hub genes (Random Forest and SVMRFE):

[0094] First, the intersection of 205 pathway-related genes and differentially expressed genes (mainly including 388 differentially expressed genes DEGs shared by endometriosis and SCLE) was used to obtain 11 cross genes (see [link to relevant documentation]). Figure 5 (D) forms a cross-gene set (also known as a candidate gene).

[0095] To further identify core candidates with potential functional importance in both diseases, a Random Forest and Support Vector Machine (SVM-RFE) algorithm was introduced to screen feature genes using cross-genes as input.

[0096] In the random forest model, the random forest algorithm (with the number of decisions ntree=500 and mrty set to automatic optimization) is used for analysis. The analysis primarily focuses on the frequency of occurrence of cross-genes in the model and the Mean Decrease Gini index to assess their classification contribution. Based on the assessment results, the feature importance of each cross-gene is analyzed and ranked. Then, the genes with the highest feature importance (e.g., the top 5) are selected as candidates, resulting in the first candidate gene set.

[0097] In SVM-RFE, the main method used is 10-fold cross-validation to perform dual screening of cross genes. Based on the radial basis function kernel (RBF kernel), cross genes with small impact on model prediction can be recursively eliminated through cross-validation to determine the contribution of cross genes to the improvement of classification performance. The top few genes with the greatest contribution to the improvement of classification performance can be selected as the second candidate gene set.

[0098] Subsequently, the model configuration with the smallest classification error in the two candidate gene sets is selected as the optimal output. The intersection of the results is taken to further determine the stable and recurring key candidate genes as comorbidity hub genes. Comorbidity hub genes can serve as clear representations of comorbidity pathway hubs between diseases.

[0099] It should be noted that when screening for comorbidity hub genes from cross-genes, each disease can be screened separately to obtain candidate genes for each disease.

[0100] For example, refer to Figure 7 , Figure 7 The results of screening co-expressed genes related to the coagulation pathway using two algorithms, random forest and support vector machine recursive feature elimination, are presented in EMs and SCLE. Figure 7 (A) shows the feature genes selected by the two models in EMs and their intersection. Figure 7 (B) shows the corresponding results in SCLE. Important candidate genes common to both diseases were screened and used for subsequent network construction and functional verification.

[0101] As can be seen from the actual experiment, the screening results of the dual algorithm are:

[0102] In the endometriosis (GSE7305) dataset, the random forest model identified 5 key genes (mainly C3AR1, COL3A1, PIK3R1, C1S, and C1R), while the SVM-RFE model screened out 8 genes (mainly COL3A1, C1S, C1R, C3AR1, SERPING1, C1QC, C1QA, and PIK3R1). The intersection of the candidate genes output by the two models revealed 4 highly important genes: C3AR1, COL3A1, C1S, and C1R (see [link to relevant documentation]). Figure 7 (As shown in (A)).

[0103] In the SCLE (GSE95474) dataset, the Random Forest model identified 10 feature genes, including COL3A1, C1QC, P2RY1, C1S, SERPING1, MAPK13, PIK3R1, C1R, C1QB, and C1QA. The SVM-RFE model identified another group of 10 genes with significant overlap. It is evident that both models jointly identified 9 genes sharing the same key genes: COL3A1, C1QC, P2RY1, C1S, SERPING1, PIK3R1, C1R, C1QB, and C1QA (see [link to relevant documentation]). Figure 7 (As shown in (B)).

[0104] Based on the gene screening results of the two diseases, it can be seen that COL3A1, C1S and C1R were repeatedly screened in all analyses, suggesting that they may serve as the hub of the comorbid pathway between endometriosis and SCLE in the complement-coagulation cascade, and can be considered as comorbid hub genes.

[0105] The results above indicate that, in the context of comorbid endometriosis and SCLE, the complement-coagulation cascade pathway, as one of the most significantly enriched pathways for differentially expressed genes, may constitute a key pathological hub between the two diseases. Therefore, this invention focuses on this pathway to identify key genes and model regulatory mechanisms, demonstrating clear mechanistic target guidance.

[0106] After screening out the comorbidity hub genes in this embodiment, these comorbidity hub genes can be further validated and evaluated as key genes. Specifically, expression trends and diagnostic efficacy can be evaluated in the validation set (i.e. validation data), and the expression level of key genes can be detected.

[0107] Specifically, the expression trends and diagnostic efficacy of comorbidity hub genes (COL3A1, C1S, C1R) in the validation set were assessed.

[0108] Expression trend analysis of comorbidity hub genes was performed using validation sets (including GSE120103 and GSE81071) to determine the expression level of each gene. ROC curves were plotted to evaluate the identification efficacy of the ROC curve and the area under the receiver operating characteristic curve (AUC) was calculated to assess its potential diagnostic value in the two diseases, thereby identifying key genes that can serve as potential cross-disease diagnostic biomarkers.

[0109] For example, refer to Figure 8 , Figure 8The results of expression validation of co-expressed coagulation and complement-related genes (COL3A1, C1S, C1R) were presented, including the expression of the three selected co-expressed key genes in the sub-EMs (GSE120103) and SCLE (GSE81071) datasets and their discriminative ability. Figure 8 (A) and Figure 8 (B) shows the expression differences of the three genes in the two diseases and the results of ROC curve analysis, which are used to evaluate their ability to distinguish potential biomarkers.

[0110] As can be seen, actual experiments show that:

[0111] In the GSE120103 dataset, compared with normal endometrial tissue, COL3A1 expression was significantly downregulated in patients with endometrial emphysema (AUC: COL3A1 AUC = 0.719), while C1S expression was upregulated (AUC: C1S AUC = 0.707) (see [link to dataset]). Figure 8 (As shown in (A)), while C1R expression showed no significant change.

[0112] In the GSE81071 dataset, compared with skin samples from healthy individuals, SCLE patients showed higher levels of COL3A1 (AUC = 0.923), C1S (AUC = 0.967), and C1R (AUC = 1.000) (see [link to dataset]). Figure 8 (B) The expression levels were significantly increased, demonstrating extremely high discrimination power.

[0113] The above results indicate that COL3A1 and C1S have good discriminative performance in both diseases and have the potential to identify diseases across diseases, supporting their role as potential cross-disease diagnostic biomarkers, i.e., as key genes.

[0114] Step 160: Analyze the immune cell lineage characteristics and interactions of the first and second training sets to obtain immune cell infiltration results, which include the immune cell lineage characteristics and interaction patterns of the two diseases.

[0115] In this embodiment, the types of immune cells are determined from the data samples of the two diseases. By analyzing the immune cell lineage characteristics (also known as immune features) and interrelationships of the immune cells, the compositional characteristics of the immune microenvironment of the two diseases and the interaction patterns between cells are systematically explored, and the results of immune cell infiltration are obtained.

[0116] Optionally, performing immune cell lineage characteristics and interrelationship analysis on the first and second training sets to obtain immune cell infiltration results may include: estimating the abundance of immune cell types in the first and second training sets using the CIBERSORTx algorithm to obtain abundance estimation information; and based on the abundance estimation information, analyzing the compositional characteristics of the immune microenvironment and the interaction patterns between cells among various diseases using statistical testing algorithms and Pearson correlation analysis to obtain immune cell infiltration results.

[0117] In the specific implementation, immune cell lineage characteristics and interrelationships were analyzed for EMs and SCLE: First, an immune infiltration analysis algorithm (such as the CIBERSORTx algorithm) was used to estimate the abundance of 22 immune cell types in the samples (GSE7305 and GSE95474), obtaining abundance estimates to determine immune cell lineage characteristics. This abundance estimate indicates the changes in immune cells in different diseases (such as an increase or decrease in cell proportion). Then, through statistical testing algorithms (Wilcoxon test) and Pearson correlation analysis, the compositional characteristics of the immune microenvironment and the interaction patterns between cells in each disease were systematically explored, determining the synergistic / antagonistic relationships between different immune cells and obtaining the infiltration analysis results of immune cells.

[0118] For example, refer to Figure 9 , Figure 9 The correlation analysis diagram for immune cell infiltration shows the characteristics of immune cell infiltration in EMs (GSE7305) and SCLE (GSE95474) and their correlation with the expression levels of COL3A1 and C1S. Figure 9 (A) and Figure 9 (D) is a stacked bar chart showing the distribution of different types of immune cells in the two disease samples; Figure 9 (B) and Figure 9 (E) represents the types of immune cells that showed significant differential expression in the disease group; Figure 9 (C) and Figure 9 (F) is a heatmap showing the correlations between various types of immune cells, used to analyze the compositional characteristics of the disease immune microenvironment and its potential regulatory mechanisms.

[0119] As can be seen, the actual experimental results show that:

[0120] In EMs, activated NK cells, regulatory T cells (Tregs), and follicular helper T cells were significantly downregulated, while the abundance of plasma cells, M2 macrophages, and resting memory CD4+ T cells was significantly increased (P<0.05). Correlation analysis showed that plasma cells were moderately negatively correlated with follicular helper T cells (r=–0.55) and activated NK cells (r=–0.60); plasma cells were positively correlated with M2 macrophages (r=0.50); ​​resting memory CD4+ T cells were strongly negatively correlated with follicular helper T cells (r=–0.74) and activated NK cells (r=–0.55), but positively correlated with M2 macrophages (r=0.58); activated NK cells were also significantly negatively correlated with M2 macrophages (r=–0.67) and follicular helper T cells (r=–0.76) (see [link to relevant documentation]). Figure 9 (As shown in (A)-(C)).

[0121] In SCLE, CD4+ memory activated T cells were significantly elevated, while M2 macrophages, resting dendritic cells, resting mast cells, memory B cells, and resting memory CD4+ T cells were significantly downregulated (P<0.05). The correlation network among immune cells was as follows: memory B cells were negatively correlated with CD4+ memory activated T cells (r=–0.74), but positively correlated with resting memory CD4+ T cells (r=0.84), M2 macrophages (r=0.55), resting dendritic cells (r=0.50), and resting mast cells (r=0.55); resting memory CD4+ T cells were strongly negatively correlated with CD4+ memory activated T cells (r=–0.90), and also negatively correlated with M2 macrophages (r=–0.90). =0.60), resting dendritic cells (r=0.71), and resting mast cells (r=0.68) showed positive correlations; CD4+ memory activated T cells showed significant negative correlations with M2 macrophages (r=-0.76), resting dendritic cells (r=-0.70), and resting mast cells (r=-0.90), respectively; in addition, there were also positive correlations between M2 macrophages and resting mast cells (r=0.75), and between resting dendritic cells and resting mast cells (r=0.52).

[0122] The above results indicate that both endometriosis and SCLE exhibit significant immune cell lineage remodeling and infiltration imbalance, but their compositional patterns differ. Endometriosis is more characterized by an increase in immunosuppressive macrophages (M2), plasma cells, and resting T cells, while SCLE shows an increase in immune-activating T cells and widespread depletion of resting immune components (see [link to study]). Figure 9 (D)-(F) shown), suggesting that the two diseases have different inflammatory ecosystem backgrounds in terms of immune imbalance mechanisms.

[0123] Step 170: Based on the results of immune cell infiltration, identify the immune cell-related expression patterns of key genes in the two diseases to obtain the expression correlation that reveals immune characteristics.

[0124] Based on the obtained immune cell infiltration results, we analyzed and identified the immune cell-related expression patterns of key genes in EMs and SCLEs, evaluated the expression correlation between the screened key genes and the abundance of various immune cells in EMs and SCLEs, and determined the positive / negative correlation between the expression level of key genes and immune cells.

[0125] Optionally, identifying the expression patterns of key genes in the two diseases based on immune cell infiltration results to obtain expression correlations that reveal immune characteristics may include: assessing the expression correlation between key genes and the abundance of various immune cells in the two diseases based on immune cell infiltration results.

[0126] In practice, based on the infiltration analysis results, the expression correlation between the key genes C1S and COL3A1 and the abundance of various immune cells in EMs and SCLEs is evaluated.

[0127] Reference Figure 10 , Figure 10 This is a graph showing the correlation between immune cells and the expression levels of COL3A1 and C1S genes. Figure 10 This demonstrates the use of EMs (corresponding to) Figure 10 (A)) and SCLE (corresponding) Figure 10 (B) shows a heatmap of the correlation between the expression levels of COL3A1 and C1S genes and the degree of infiltration of various immune cells.

[0128] As can be seen, the actual results show that:

[0129] In endometriosis samples, the expression level of the C1S gene was significantly positively correlated with the following three types of immune cells. Specifically, for the C1S gene, it was positively correlated with immunosuppressive cells such as M2 macrophages, plasma cells (r=0.53), and resting CD4+ memory T cells (r=0.72). Simultaneously, the C1S gene was significantly negatively correlated with the following three types of immune cells: follicular helper T cells (r=–0.85), activated NK16 cells (r=–0.59), and regulatory T cells (r=–0.46). For the COL3A1 gene, it showed the opposite trend in endometriosis samples, generally exhibiting a negative correlation with the aforementioned cell types, specifically: negatively correlated with plasma cells (r=–0.48), resting memory CD4+ T cells (r=–0.52), and M2 macrophages (r=–0.73) (see [link to relevant documentation]). Figure 10 (As shown in (A)).

[0130] In SCLE, the C1S gene was significantly positively correlated with CD4+ memory-activated T cells (r=0.69), but negatively correlated with resting mast cells (r=–0.80); COL3A1 expression was positively correlated with only two cell types in SCLE: resting dendritic cells (r=0.67) and resting mast cells (r=0.72) (see [link to relevant documentation]). Figure 10 (As shown in (B)).

[0131] It is evident that the C1S and COL3A1 genes exhibit expression specificity at the immune cell subset level in different diseases, potentially participating in disease pathogenesis by regulating immune cell infiltration. In other words, the specific immune cell-related expression patterns of these two genes under different disease backgrounds suggest that they may participate in disease progression by regulating immune cell infiltration.

[0132] Step 180: Based on expression correlation, immune pathway activity analysis and key gene regulatory role assessment are performed according to key genes and immune cell lineage characteristics to obtain the downstream participation mode of key genes.

[0133] Based on the identified key genes and immune characteristics, immune pathway activity analysis and key gene regulation assessment are performed. This includes comparing the activity differences of typical immune signaling pathways among diseases and assessing the localization and expression patterns of key genes in immune signaling pathways. In other words, the downstream participation patterns of target key genes in immune pathway activity are obtained, thereby realizing immune pathway activity analysis and key gene regulation assessment.

[0134] In its specific implementation, the above-mentioned approach uses expression correlation as a benchmark to analyze immune pathway activity and assess the regulatory role of key genes based on key genes and immune cell lineage characteristics, thereby obtaining the downstream participation patterns of key genes. Specifically, this may include: using ssGSEA to score the activity of immune-related pathways based on the obtained pathway set, comparing the differences in the overall activity composition of immune-related pathways between the two diseases, and identifying disease-specific activation pathways by comparing the differences in component pathway scores between the two diseases; and using specific activation pathways as the criterion, analyzing the correlation coefficient between key gene expression and pathway activity based on key genes and immune cell lineage characteristics to obtain the downstream participation patterns of key genes.

[0135] In this embodiment, ssGSEA was used to compare the overall activity differences between EMs and SCLE in 10 immune-related pathways, and to evaluate the localization and expression patterns of C1S and COL3A1 in the pathways.

[0136] For example, the activity differences of typical immune signaling pathways in EMs and SCLE can be compared as information on the activity of immune signaling pathways in each disease. It can also assess the localization and expression patterns of C1S and COL3A1 in the pathway, evaluate the regulatory role of key genes, and determine the downstream involvement patterns.

[0137] Specifically, immune pathway activity was analyzed: using key genes as samples, ssGSEA was used to score the activity of 10 immune-related pathways in each sample. The set of immune-related pathways was derived from the aforementioned KEGG and Reactome databases. The Wilcoxon rank-sum test was used to compare the pathway scores between the EMs and SCLE groups to identify disease-specific activation pathways. Targeted analysis was then conducted on each disease-specific activation pathway.

[0138] Further calculations were made of the Pearson correlation coefficients between the expression of C1S and COL3A1 and the activity of each pathway. Using a significance threshold (e.g., P < 0.05) as the standard, the potential regulatory role of key genes in the immune pathway network was analyzed to clarify the downstream participation pattern. This downstream participation pattern can be understood as the localization and expression pattern of key genes in the pathway.

[0139] For example, refer to Figure 11 The differentially expressed immune-related signaling pathways in EMs (GSE7305) and SCLEs (GSE95474) and their association with key genes COL3A1 and C1S are shown. Figure 11 (A) and (B) are diagrams showing the significant differences in immune pathway expression between the two diseases; Figure 11 (C) and Figure 11 (D) presents the results of correlation analysis between the expression levels of COL3A1 and C1S genes and the activity of various major immune pathways, revealing their potential role in the immune regulatory network.

[0140] As can be seen from the actual experiment:

[0141] For 10 typical immune pathways (including but not limited to Complement, VEGF, EGFR, JAK / STAT, IL-6, WNT, Toll-like receptor, etc.): In EMs, the expression of complement, VEGF, EGFR, JAK / STAT, and IL-6 pathways was relatively upregulated, while the activity of WNT, Toll-like receptor, and other pathways was relatively low (see...). Figure 11 (A) shows that in SCLE, complement and JAK / STAT are significantly activated, while the WNT pathway is significantly downregulated (see [reference]). Figure 11 (As shown in (B)).

[0142] Further analysis revealed that the key gene C1S was positively correlated with complement (r = 0.56), VEGF (r = 0.47), and JAK / STAT (r = 0.64) pathways in EMs (see [link to relevant documentation]). Figure 11 (C) shows that in SCLE, it was significantly correlated with complement (r = 0.93), JAK / STAT (r = 0.77), and the Toll-like receptor pathway (r = 0.65), while it was negatively correlated with ERBB2 (r = –0.75) and ERBB4 (r = –0.70) (see [link to relevant documentation]). Figure 11 (as shown in (D)).

[0143] Meanwhile, the key gene COL3A1 was positively correlated with the WNT pathway in both diseases (r = 0.78), but negatively correlated with the complement pathway (r = –0.66). This suggests that the key gene COL3A1 may play a potential role in the cross-regulation of structure and immunity.

[0144] Based on the above experimental results, it can be seen that key genes such as C1S and COL3A1 can be located at specific pathway hubs.

[0145] Step 190: Based on the established ceRNA ternary network, determine the upstream regulatory network of key genes, and combine it with the downstream participation mode to establish a comorbidity identification mechanism.

[0146] In practical implementation, after determining the downstream involvement mode, the upstream regulatory mechanisms of key genes in immune pathway activity are further traced back to identify the upstream ceRNA regulatory network of key genes (COL3A1 and C1S) based on miRNA-circRNA-interaction information. Specifically, a circRNA-miRNA-mRNA ternary network is established using ceRNA theory as the upstream regulatory network. In this ternary network, the miRNAs and their corresponding circRNAs that are co-targeted with the key genes are predicted by algorithms, thereby determining the upstream regulatory relationships of key genes, identifying potential co-pathogenic regulatory axes, clarifying the target key genes regulated by common ceRNA mechanisms, and using the circRNA / miRNA / mRNA axis to construct the upstream network regulating the common pathological pathways of EMs and SCLE, ultimately establishing a complete upstream ceRNA regulatory network identification mechanism.

[0147] Finally, by combining downstream participation patterns and upstream regulatory networks, a comorbidity identification mechanism can be established to identify and analyze the comorbidity between various diseases, especially between EMs and SCLE.

[0148] Optionally, the above-mentioned determination of the upstream regulatory network of key genes based on the established ceRNA ternary network, combined with downstream participation patterns, to establish a comorbidity identification mechanism may include the following sub-steps:

[0149] Sub-step 2901: Based on the ceRNA theory, a ternary network of miRNA, mRNA, and circRNA is established, and key genes are input into the ternary network.

[0150] Sub-step 2902: In the ternary network, an algorithm is used to predict the miRNAs that are commonly targeted by key genes to obtain candidate miRNAs.

[0151] Sub-step 2903 involves retrieving circRNA binding information based on candidate miRNAs, constructing an interaction network between miRNAs and circRNAs, and screening for circRNAs with target abundance.

[0152] Sub-step 2904 involves regulatory analysis based on candidate miRNAs and circRNAs to identify the upstream regulatory network that constitutes the common pathological pathways of various diseases, and to establish a comorbidity identification mechanism by combining downstream participation patterns.

[0153] A unified explanation is provided for sub-steps 2901-2904:

[0154] In practical implementation, firstly, algorithms such as miRWalk and TargetScan are used to predict miRNAs that co-target the target key genes, i.e., to predict miRNAs that are co-regulated by COL3A1 and C1S, and their intersection is taken to obtain candidate miRNAs (26 candidate miRNAs can be selected from the miRNAs corresponding to COL3A1 and C1S). Then, based on the candidate miRNAs, the binding information of the circRNAs regulated by these miRNAs is retrieved from a pre-defined database (such as StarBase v2.0), thereby constructing a miRNA–circRNA interaction network. Through cross-prediction, circRNAs with high binding abundance are screened to obtain the target abundance circRNAs. Finally, regulatory analysis is performed using the candidate miRNAs and the target abundance circRNAs to determine the upstream regulatory network that constitutes the common pathological pathways of the diseases.

[0155] For example, refer to Figure 12 The diagram shows the construction of a ternary ceRNA regulatory network. Figure 12 The process and results of constructing a ceRNA regulatory network with COL3A1 and C1S as the core target genes are presented. Figure 12 (A) Results of using miRWalk and TargetScan tools to predict miRNAs that simultaneously target COL3A1 and C1S; Figure 12(B) to (H) show the upstream regulatory circRNA prediction maps for the seven key miRNAs (including hsa-miR-129-5p, hsa-miR-487a-3p, hsa-miR-1276, hsa-miR-3622b-5p, hsa-miR-4458, hsa-let-7c-5p and hsa-let-7b-5p); Figure 12 (I) is a Sankey diagram of the integrated circRNA–miRNA–mRNA ternary regulatory network, which clearly shows the correspondence between each regulatory level and reveals the potential co-pathogenic regulatory axis.

[0156] As can be seen, in actual experiments, consensus predictions show that: 26 candidate miRNAs (refer to...) Figure 12 (A) shows 7 miRNAs that can form a regulatory relationship with circRNAs. These 7 candidate miRNAs are: hsa-miR-129-5p, hsa-miR-487a-3p, hsa-miR-1276, hsa-miR-3622b-5p, hsa-miR-4458, hsa-let-7c-5p, and hsa-let-7b-5p.

[0157] Among them, hsa-miR-129-5p binds to 25 candidate circRNAs (see reference). Figure 12 (As shown in (B)), hsa-miR-487a-3p binds to 5 candidate circRNAs (refer to...) Figure 12 (as shown in (C)), hsa-miR-1276 binds to 12 candidate circRNAs (see reference). Figure 12 (D) shows that hsa-miR-3622b-5p binds to 10 candidate circRNAs (see reference). Figure 12 (As shown in (E)), hsa-miR-4458 binds to 19 candidate circRNAs (see reference). Figure 12 (As shown in (E)), hsa-let-7c-5p binds to 18 candidate circRNAs (see reference). Figure 12 (as shown in (G)), hsa-let-7b-5p binds to 22 candidate circRNAs (refer to...) Figure 12 (H) is shown.

[0158] Further analysis revealed that hsa-let-7b-5p, hsa-let-7c-5p, and hsa-miR-4458 can be co-regulated by multiple circRNAs, including but not limited to: BOD1L, CHD4, CPSF4, E2F4, GNG5, IARS, KPNA2, LOC653513, NME4, PABPC4, PDE4DIP, PSAP, SIN3A, SYVN1, XKR8, etc. (see reference). Figure 12 (As shown in (I)), this embodiment clarifies that the aforementioned circRNA / miRNA / mRNA axis can constitute an upstream network regulating the common pathological pathways of EMs and SCLE. A circRNA–miRNA–mRNA ternary regulatory network was constructed as the upstream regulatory network, and it was determined that COL3A1 and C1S can be regulated by a common ceRNA mechanism. By combining the upstream regulatory network with the downstream participation mode, a comorbidity recognition mechanism was established. This comorbidity recognition mechanism constitutes a robust and cross-disease-transferable molecular mechanism mining process, providing technical support and pathway reference for the study of comorbidity mechanisms in autoimmune diseases.

[0159] As can be seen, this application's embodiments, based on systematic integration of transcriptome expression features and immune pathway analysis, starting from the complement-coagulation axis, identify highly overlapping inflammatory signaling modules between endometriosis and subacute cutaneous lupus erythematosus, and explicitly propose key genes as dual targets with cross-expression characteristics and immune regulatory functions, filling the gap in the study of the comorbidity mechanism of these two diseases. At the regulatory mechanism level, a ceRNA regulatory network including circRNA nodes is constructed, overcoming the limitations of traditional miRNA / lncRNA analysis and enhancing the complete identification of non-coding regulatory axes. Furthermore, this embodiment simultaneously introduces multiple algorithms for collaborative screening and combines them with an external validation set to form a robust and transferable molecular mechanism mining process across diseases, providing technical support and pathway reference for the study of comorbidity mechanisms in autoimmune diseases, and solving the technical problem that existing technologies cannot accurately identify and analyze comorbidity mechanisms.

[0160] The comorbidity mechanism identification process established in this application, which integrates multi-omics integration, immune mechanism modeling, and upstream regulation prediction, is applicable not only to EMs and SCLE, but can also be extended to molecular mechanism research and target discovery for other chronic autoimmune diseases with overlapping immune pathways.

[0161] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should know that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps may be performed in other orders or simultaneously.

[0162] like Figure 13As shown in the embodiments of this application, a comorbidity mechanism identification system 1300 based on the complement-coagulation axis and ceRNA network is also provided, comprising:

[0163] The data acquisition module 1310 is used to standardize and correct batch effects on the raw microarray data in the transcriptome dataset, and to establish a standardized expression matrix. The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease.

[0164] The differential pathway analysis module 1320 is used to perform pathway activity scoring and disease common pathway feature extraction on the first and second training sets based on the expression matrix, and to screen differential pathways.

[0165] The co-expression analysis module 1330 is used to construct a weighted co-expression network based on the expression matrix of the first validation set and the expression profile of skin tissue, and to identify the key expression modules between the two diseases.

[0166] The differential expression analysis module 1340 is used to perform GO enrichment analysis and KEGG pathway enrichment analysis on key expression modules to screen significant entries; and to perform differential expression analysis on the first and second training sets to screen differentially expressed genes shared by the two diseases.

[0167] The key gene identification module 1350 is used to extract pathway-related genes from the obtained complement pathway and coagulation pathway, screen comorbidity hub genes by combining differentially expressed genes, and use the first validation set and the second validation set to perform expression trend analysis and diagnostic efficacy assessment to identify key genes.

[0168] The immune infiltration module 1360 is used to analyze the immune cell lineage characteristics and interactions of the first and second training sets to obtain immune cell infiltration results, which include the immune cell lineage characteristics and interaction patterns of the two diseases.

[0169] The expression correlation analysis module 1370 is used to identify the immune cell-related expression patterns of key genes in two diseases based on the results of immune cell infiltration, and to obtain the expression correlation that reveals immune characteristics.

[0170] The downstream participation pattern analysis module 1380 is used to analyze the activity of immune pathways and evaluate the regulatory role of key genes based on expression correlation and the characteristics of key genes and immune cell lineages, so as to obtain the downstream participation pattern of key genes.

[0171] The comorbidity identification module 1390 is used to identify the upstream regulatory network of key genes based on the established ceRNA ternary network, and to establish a comorbidity identification mechanism by combining the downstream participation mode.

[0172] It should be noted that the comorbidity mechanism identification system based on complement-coagulation axis and ceRNA network provided in this application embodiment can execute the comorbidity mechanism identification method based on complement-coagulation axis and ceRNA network provided in any embodiment of this application, and has the corresponding functions and beneficial effects of the execution method.

[0173] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0174] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for identifying comorbidity mechanisms based on the complement-coagulation axis and ceRNA network, characterized in that, Applied to the identification of comorbidity mechanisms in at least two diseases, including: The raw microarray data in the transcriptome dataset were standardized and batch effect corrected to establish a standardized expression matrix. The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease. Based on the expression matrix, pathway activity scores and common disease pathway features were extracted from the first and second training sets to screen for differential pathways. Based on the expression matrix, a weighted co-expression network was constructed using the expression profiles of the first validation set and skin tissue to identify key expression modules between the two diseases. GO enrichment analysis and KEGG pathway enrichment analysis were performed on key expression modules to screen significant entries. In addition, differential expression analysis was performed on the first and second training sets to screen differentially expressed genes shared by the two diseases. Pathway-related genes were extracted from the obtained complement and coagulation pathways, and differentially expressed genes were combined to screen for comorbidity hub genes. Expression trend analysis and diagnostic efficacy assessment were performed using the first and second validation sets to identify key genes. Immune cell lineage characteristics and interrelationships were analyzed on the first and second training sets to obtain immune cell infiltration results, which included immune cell lineage characteristics and interaction patterns of the two diseases. Based on the results of immune cell infiltration, the expression patterns of key genes in two diseases were identified as immune cell-related, revealing the expression correlation of immune features. Based on expression correlation, immune pathway activity analysis and key gene regulation assessment were conducted according to key genes and immune cell lineage characteristics to obtain the downstream participation mode of key genes. Based on the established ceRNA ternary network, the upstream regulatory network of key genes was identified, and a comorbidity identification mechanism was established by combining downstream participation patterns.

2. The method according to claim 1, characterized in that, The raw microarray data in the transcriptome dataset were standardized and batch effect corrected to establish a standardized expression matrix, including: Obtain the transcriptome dataset from the data source; The RMA algorithm was used to perform background correction and normalization on the transcriptome dataset, and the probe was uniformly mapped in combination with annotation information to update the transcriptome dataset. The transcriptome dataset was corrected using the sva package, and low-expression probes and missing samples were removed to construct a standardized expression matrix. Each expression matrix corresponds to a data point in the transcriptome dataset.

3. The method according to claim 1, characterized in that, Based on the expression matrix, pathway activity scores and common disease pathway features were extracted from the first and second training sets to screen for differentially expressed pathways, including: Gene set variation analysis was performed on the first and second training sets based on the expression matrix. Using the obtained datasets as a benchmark, pathway gene set templates were used to construct pathway activity scoring matrices. Based on the pathway activity score matrix, the differences between the disease group and the control group were analyzed to determine the pathway expression pattern and screen differentially expressed pathways.

4. The method according to claim 1, characterized in that, Based on the expression matrix, a weighted co-expression network was constructed using the expression profiles of the first validation set and skin tissue to identify key expression modules between the two diseases, including: Based on the expression matrix, a weighted co-expression network construction method was used to construct weighted gene co-expression networks for the first validation set and the skin tissue expression profile, respectively. For weighted gene co-expression networks, a dynamic pruning algorithm is used to identify co-expression modules and calculate the Pearson correlation coefficient between co-expression modules and disease states to screen key expression modules that are highly correlated with disease phenotypes.

5. The method according to claim 1, characterized in that, Differential expression analysis was performed on the first and second training sets to screen for differentially expressed genes shared by the two diseases, including: Based on the expression matrix, differential expression analysis was performed on the first and second training sets using the limma package to filter differential data. For the differentially expressed data, Venn diagrams were used to identify differentially expressed genes shared by the two diseases.

6. The method according to claim 1, characterized in that, Pathway-related genes were extracted from the obtained complement and coagulation pathways. Comorbidity pivot genes were screened using differentially expressed genes, and expression trend analysis and diagnostic efficacy assessment were performed using a first and second validation set to identify key genes, including: The complement pathway and coagulation pathway were obtained from the KEGG database, and pathway-related genes were extracted. The intersection of pathway-related genes and differentially expressed genes is used to obtain a cross-gene set that includes the cross-genes; The first candidate gene set is selected by classifying and evaluating the contribution of each cross-gene using a random forest model, and the second candidate gene set is selected from each cross-gene using a support vector machine model with cross-validation. Based on the intersection analysis of the first and second candidate gene sets, the comorbidity hub gene was obtained. Expression trend analysis of comorbidity hub genes was performed using the first and second validation sets, and receiver operating characteristic (ROC) curves were calculated to evaluate the diagnostic efficacy of comorbidity hub genes in the two diseases and identify key genes.

7. The method according to claim 1, characterized in that, Immune cell lineage characteristics and interrelationships were analyzed on the first and second training sets to obtain results on immune cell infiltration, including: The CIBERSORTx algorithm was used to estimate the abundance of immune cell types in the first and second training sets to obtain abundance estimation information. Based on the abundance estimation information, the compositional characteristics of the immune microenvironment and the interaction patterns between cells among various diseases are analyzed by statistical testing algorithms and Pearson correlation analysis to obtain the results of immune cell infiltration. Among them, the expression patterns of key genes in the two diseases were identified based on the results of immune cell infiltration, and the expression correlation of immune characteristics was revealed. This included: assessing the expression correlation between key genes and the abundance of various immune cells in the two diseases based on the results of immune cell infiltration.

8. The method according to claim 1, characterized in that, Based on expression correlation, immune pathway activity analysis and key gene regulatory roles were assessed according to key genes and immune cell lineage characteristics to obtain downstream participation patterns of key genes, including: Based on the acquired pathway set, ssGSEA or GSVA is used to score the activity of immune-related pathways. The differences in the overall activity composition of immune-related pathways between the two diseases are compared. By comparing the differences in the component pathway scores between the two diseases, disease-specific activation pathways are identified. Based on specific activation pathways, the correlation coefficients between key gene expression and pathway activity were analyzed according to key genes and immune cell lineage characteristics to obtain the downstream participation patterns of key genes.

9. The method according to claim 1, characterized in that, Based on the established ceRNA ternary network, the upstream regulatory network of key genes was identified. Combined with downstream participation patterns, a comorbidity identification mechanism was established, including: Based on the ceRNA theory, a ternary network of miRNA, mRNA, and circRNA was established, and key genes were input into the ternary network. In the ternary network, an algorithm is used to predict the miRNAs that are commonly targeted by key genes to obtain candidate miRNAs; Based on candidate miRNAs, information on the binding of regulated circRNAs is retrieved, an interaction network between miRNAs and circRNAs is constructed, and circRNAs with target abundance are screened. Regulatory analysis based on candidate miRNAs and circRNAs identifies the upstream regulatory network that constitutes the common pathological pathways of various diseases, and establishes a comorbidity identification mechanism by combining downstream participation patterns.

10. A comorbidity mechanism identification system based on the complement-coagulation axis and ceRNA network, characterized in that, include: The data acquisition module is used to standardize and correct batch effects on the raw microarray data in the transcriptome dataset, and to establish a standardized expression matrix. The transcriptome dataset includes the first training set and the first validation set for the first disease, the second training set and the second validation set for the second disease, and the skin tissue expression profile for the second disease. The differential pathway analysis module is used to perform pathway activity scoring and extract common disease pathway features from the first and second training sets based on the expression matrix, and to screen differential pathways. The co-expression analysis module is used to construct a weighted co-expression network based on the expression matrix of the first validation set and the expression profile of skin tissue, and to identify the key expression modules between the two diseases. The differential expression analysis module is used to perform GO enrichment analysis and KEGG pathway enrichment analysis on key expression modules to screen significant entries, and to perform differential expression analysis on the first and second training sets to screen differentially expressed genes shared by the two diseases. The key gene identification module is used to extract pathway-related genes from the obtained complement and coagulation pathways, screen comorbidity hub genes by combining differentially expressed genes, and use the first and second validation sets to perform expression trend analysis and diagnostic efficacy assessment to identify key genes. The immune infiltration module is used to analyze the immune cell lineage characteristics and interactions of the first and second training sets to obtain immune cell infiltration results, which include the immune cell lineage characteristics and interaction patterns of the two diseases. The expression correlation analysis module is used to identify the immune cell-related expression patterns of key genes in two diseases based on the results of immune cell infiltration, and to obtain expression correlations that reveal immune characteristics. The downstream participation pattern analysis module is used to analyze the activity of immune pathways and evaluate the regulatory role of key genes based on expression correlation and the characteristics of key genes and immune cell lineages, so as to obtain the downstream participation pattern of key genes. The comorbidity identification module is used to identify the upstream regulatory network of key genes based on the established ceRNA ternary network, and to establish a comorbidity identification mechanism by combining downstream participation modes.