SLE-PAH typing method based on transcriptomics, prediction model construction method and application
Through the transcriptomic-based SLE-PAH classification method, SLE-PAH is divided into subtypes of different inflammation degrees, and a predictive model of moderate inflammation subtype is constructed, which solves the problem of unclear pathogenesis of SLE-PAH in the prior art, and achieves the formulation of targeted treatment plans for different diseases and the accurate identification of patients with moderate inflammation subtypes.
Patent Information
- Application Number
- CN202510118898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The pathogenesis of SLE-PAH in the prior art is not fully understood, and it is difficult to formulate targeted and effective treatment plans for different degrees of diseases.
Using the transcriptomic-based SLE-PAH classification method, SLE-PAH was divided into three subtypes with different PAH levels, different inflammation degrees, and different inflammation-related signaling pathway activation levels, and named high inflammatory subtypes, moderate inflammatory subtypes and low inflammatory subtypes, and predictive models of moderate inflammatory subtypes were constructed.
The pathogenesis of SLE-PAH was elucidated through transcriptomic typing method, and potential SLE-PAH treatment targets were screened. The constructed prediction model has strong predictive ability and can accurately identify patients with moderate inflammatory subtypes.
Smart Images

Figure CN120048336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics, and specifically relates to a method for typing SLE-PAH based on transcriptomics, a method for constructing a prediction model, and applications thereof. Background Art
[0002] Systemic lupus erythematosus (SLE) is a systemic autoimmune disease that can exhibit severe multi-organ damage. Pulmonary arterial hypertension (PAH) is a severe complication of SLE with a poor prognosis and often becomes the main cause of death in SLE patients. Approximately 3% of SLE patients can develop PAH, and the 3-year survival rate of SLE-PAH is 81.3% - 89.4%.
[0003] PAH is a subtype of pulmonary arterial hypertension and is often referred to as type I PH9. The underlying mechanism of SLE-PAH has not been fully elucidated. Genome studies of idiopathic PAH (IPAH) have shown that mutations in bone morphogenetic protein receptor type 2 (BMPR2) play a role in the pathogenesis of PAH10. The absence of the BMPR2 signaling pathway has also been found in SLE-PAH, indicating that SLE-PAH is involved in part of the pathogenesis of IPAH11. However, a unique pathogenesis related to abnormal immune system function has also been reported as the pathogenesis of SLE-PAH. Some studies have found a correlation between immune complexes and SLE-PAH. It has been reported that positive anti-ribonucleoprotein antibodies can increase the risk of PAH in SLE patients. Some studies have reported the deposition of antinuclear antibodies, immunoglobulins, and complement in the pulmonary vessels, similar to the situation in lupus nephritis. At the same time, abnormal immune cell function also promotes the development of SLE-PAH. The mutation of HLA-DQA1*03:02 in the major histocompatibility complex (MHC) region has been found to be associated with a poor prognosis of SLE-PAH17. In addition, it has been demonstrated that overactivation of T cells is an important factor in the inflammation and vascular remodeling of SLE-PAH18. In addition, previous studies have shown that immunosuppressive therapy can bring better outcomes for patients with SLE-PAH19. These studies together indicate that in addition to the mechanisms specific to PAH, inflammation plays a crucial role in the development of SLE-PAH. Summary of the Invention
[0004] To this end, the technical problem to be solved by the present invention is to provide a method for typing SLE-PAH based on transcriptomics, a method for constructing a prediction model, and applications thereof. It solves the technical problem that the pathogenesis of SLE-PAH in the prior art is not fully understood, making it difficult to formulate targeted and effective treatment plans for different degrees of diseases.
[0005] The present invention provides a technical solution: a method for typing SLE-PAH based on transcriptomics, the method comprising the following steps:
[0006] S1. Collect research samples, including SLE patients and SLE-PAH patients;
[0007] S2. Collect and process data to obtain transcriptome samples of SLE patients and SLE-PAH patients;
[0008] S3. Analysis based on transcriptome samples;
[0009] Classify SLE-PAH into three subtypes with different degrees of PAH, different degrees of inflammation, and different activation levels of inflammation-related signaling pathways, and name them high-inflammation subtype, moderate-inflammation subtype, and low-inflammation subtype respectively.
[0010] Preferably, the S2 includes:
[0011] S2.1. Data collection and risk assessment. The data collection includes collecting blood samples from patients, and the risk assessment includes evaluating PAH using a four-layer risk assessment system;
[0012] S2.2. Sample processing and sequencing. Collect peripheral blood samples from patients, extract total RNA, construct a library and sequence it;
[0013] S2.3. Data processing and quality control.
[0014] Preferably, the S3 includes:
[0015] S3.1. Perform differential gene expression analysis on SLE patients and SLE-PAH patients;
[0016] S3.2. Multi-pathway enrichment analysis, including Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment, and Gene Set Enrichment Analysis (GSEA);
[0017] S3.3. Consensus non-negative matrix factorization, using a clustering method based on non-negative matrix factorization (cNMF).
[0018] Preferably, the S3 further includes:
[0019] S3.4. Perform principal component analysis (PCA) according to the detection results of S3.1;
[0020] S3.5. Immune infiltration analysis (CIBERSORT);
[0021] S3.6. Perform protein-protein interaction analysis and hub gene identification according to the detection results of S3.1;
[0022] S3.7. Predict common transcription factors of hub genes;
[0023] S3.8. Disease enrichment analysis.
[0024] Preferably, the high - inflammation subtype is characterized by the mildest manifestation of PAH.
[0025] Preferably, the high - inflammation subtype is characterized by a significant increase in IL6.
[0026] Preferably, the moderate - inflammation subtype is characterized by the severest manifestation of PAH and IPAH.
[0027] Preferably, the moderate - inflammation subtype has 12 hub genes: SLC25A39, CDC34, FKBP8, SLC2A1, ALAS2, FOXO3, UBXN6, DMTN, DCAF12, BCL2L1, SLC4A1, KLF1; and the characteristic transcription factors are GATA1 and KLF1.
[0028] The present invention also provides a technical solution: a method for constructing a prediction model of the SLE - PAH moderate - inflammation subtype based on the above - mentioned typing method, screening 4 characteristic genes, and constructing a prediction model using multiple logistic regression. The 4 characteristic genes are ALAS2, FKBP8, SLC2A1, and UBXN6.
[0029] The present invention also provides a technical solution: a detection kit for the SLE - PAH moderate - inflammation subtype based on the above - mentioned typing method. The kit includes reagents for detecting characteristic genes, and the characteristic genes are ALAS2, FKBP8, SLC2A1, and UBXN6.
[0030] Beneficial effects:
[0031] The transcriptomics - based SLE - PAH typing method provided by the present invention analyzes the transcriptomes of SLE - PAH and SLE patients, determines three SLE - PAH subtypes showing distinguishable different inflammation levels through cNMF, clarifies the pathogenesis of SLE - PAH, and screens potential treatment targets for SLE - PAH. For the first time, a cohort study on classifying and characterizing SLE - PAH based on transcriptome is disclosed, aiming to pave the way for the precise management of SLE - PAH.
[0032] The method for constructing a prediction model of SLE - PAH based on the transcriptomics - based SLE - PAH typing method provided by the present invention determines 4 specific genes as risk factors. A calibration plot is used for internal validation, and the results show that the model fits very well. An ROC curve is plotted to evaluate the prediction ability of the model. The C - index of the model is 0.954, indicating that the model has strong prediction ability. A Nomogram is drawn for convenient clinical operation, and patients with a score above 98 are recommended to be classified into the moderate - inflammation subtype. The accuracy of the model in identifying patients with the moderate - inflammation subtype reaches 86%.
[0033] This invention is a study that for the first time compares the transcriptomes of SLE patients and SLE-PAH patients and describes the characteristics of different subtypes of SLE-PAH. This study is based on the largest and high-quality SLE cohort in China, and includes patients with comprehensive clinical data and clear clinical diagnoses. The sample size of this invention's study is relatively large, which makes the results more reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to make the content of this invention easier to be clearly understood, the following further details this invention according to specific embodiments of this invention and in combination with the drawings.
[0035] Figure 1 Differential expression analysis of patients with systemic lupus erythematosus and systemic lupus erythematosus-pulmonary arterial hypertension according to this invention; A shows a volcano plot of DEGs between SLE and SLE-PAH, the threshold of log2 fold change is set to 0.5, up-regulated genes are marked in red, and down-regulated genes are marked in blue; B shows a heat map between systemic lupus erythematosus and systemic lupus erythematosus-pulmonary arterial hypertension; C shows GSEA between systemic lupus erythematosus and systemic lupus erythematosus-pulmonary arterial hypertension groups, this analysis is based on gene sets marked by MsigDB, and the length of the gray line represents -log10 p value; D shows GO analysis between SLE and SLE-PAH groups, each circle represents an enriched pathway, and the result is visualized by Cluego; E shows GSEA based on GO gene sets; F shows the expression of 21 ISGs in SLE and SLE-PAH patients, the SLE-PAH group is marked in red, and the SLE group is marked in blue (*P value < 0.05); G shows a schematic diagram of batch effect;
[0036] Figure 2Schematic diagram of cNMF clustering analysis results for SLE-PAH patients and SLE patients in the present invention; A-C show cNMF clustering of 1324 DEGs in SLE-PAH patients; A is cNMF ranking screening, based on high latent coefficient and silhouette coefficient, the optimal 3 levels are selected; B is the consensus matrix of SLE-PAH patients; C is PCA analysis, P1-P3 represent the 1st-3rd classes of SLE-PAH patients identified by cNMF; D-F show CNMF clustering of 1324 DEGs in SLE-PAH patients and SLE patients; D is cNMF two-dimensional ranking screening, based on higher latent coefficient and silhouette coefficient, the optimal ranking is selected as 4; E is the consensus matrix of SLE-PAH patients and SLE patients; F is PCA analysis, P1-P3 represent the 1st-3rd classes of SLE-PAH patients identified by cNMF; G shows the comparison of SLE-PAH patient classification by two cNMF analyses, and the number of SLE-PAH patients classified into the same group is highlighted in gray; H shows the three-group PCA analysis of SLE-PAH patients and SLE patients;
[0037] Figure 3 GSEA of different subtypes of SLE-PAH in the present invention. Signature gene sets from MSigDB were used by GSEA to identify subgroup-specific features. The dots represent the regulatory direction (orange for upregulation; blue for downregulation), and the length of the gray line represents the significance of the statistical test. Only significant enrichment sets (at least in one subgroup) are shown in the figure;
[0038] Figure 4 Analysis of SLE-PAH subtypes and high-inflammation groups in the present invention; A shows a Venn diagram of DEGs between each subtype and SLE, with clusters 1-3 representing high, medium, and low inflammation subtypes respectively; B shows GO enrichment analysis based on the uniquely upregulated genes in the high-inflammation subgroup, showing the 30 gene sets with the smallest p-values; C shows GO enrichment analysis based on the uniquely upregulated genes in the high-inflammation subgroup; D shows the ISG scores of each subtype of SLE-PAH patients and SLE patients. The scores of 28 genes are shown on the left, and the scores of 21 genes are shown on the right; E-H show GSEA based on GO gene sets, and the description of each gene group is marked at the top; I shows the prediction of immune cell proportions using CIBERSORT;
[0039] Figure 5Analysis of the SLE-PAH subtype and the high-inflammation group in the present invention; A shows the GO enrichment analysis of unique differential genes based on the intermediate inflammation subgroup, and the results are visualized by Cluego; B shows the KEGG enrichment analysis of unique differential genes based on the intermediate inflammation subgroup, and the results are visualized by Cluego; C shows the disease enrichment analysis of unique differential genes based on the intermediate inflammation subgroup, and each gray block represents the enrichment degree of the genes on the x-axis in the diseases on the y-axis; D shows the prediction of TFs on the chEA3 website. The bar chart represents the contribution of evidence from different database sources to the prediction ranking of TFs. GATA1 and KLF1 are the only two common transcription factors among the 12 hub genes; E shows the protein-protein interaction network between the 13 hub genes and TFs; F shows the expression levels of the 13 hub genes and TFs in different subtypes of SLE-PAH and SLE.
[0040] Figure 6 Construction of a transcriptome-based prediction model in the present invention; A shows the coefficient distribution map and LASSO regression path of each gene; B shows the 10-fold cross-validation of LASSO regression. The black vertical line with red dots represents the cross-validation curve and the standard error of the partial likelihood deviation; C shows the model calibration curve, comparing the predicted probability with the actual probability to verify the effectiveness of the model; D shows the ROC curve of the model. The red line represents the specificity and sensitivity of the prediction model. The cut-off point is selected according to the maximum Youden index, and the AUC value of the model is marked as 0.954 in the figure; E shows the prediction model Nomogram. Detailed implementation mode
[0041] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. The principles and features of the present invention are described below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The embodiments given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0042] Example 1
[0043] This example provides a method for typing SLE-PAH based on transcriptomics, including the following steps:
[0044] S1. Collect research samples, including SLE patients and SLE-PAH patients
[0045] The Chinese Systemic Lupus Erythematosus Treatment and Research Group (CSTAE) is the largest follow-up cohort of patients with systemic lupus erythematosus in China. This study was based on the patient cohort of the Chinese Systemic Lupus Erythematosus Research Collaboration (CSTAR). All enrolled patients met the 2012 Systemic Lupus International Collaborating Clinics (SLICC) classification criteria or the 2019 European League Against Rheumatism / American College of Rheumatology (EULAR / ACR) SLE classification criteria. In this study, SLE-PAH patients were diagnosed by right heart catheterization (RHC) and met the hemodynamic criteria: 1) mean pulmonary artery pressure (mPAP) ≥ 21 mmHg; 2) pulmonary vascular resistance (PVR) 2 - 3 WU; 3) and PAWP ≤ 15 mmHg. All patients gave written informed consent. A total of 100 SLE-PAH patients and 95 SLE patients were included in this study.
[0046] S2. Collection of data and processing to obtain transcriptome samples of SLE patients and SLE-PAH patients
[0047] S2.1. Data collection and risk assessment
[0048] The follow-up of collecting patients' blood samples was used as the baseline. Demographic characteristics, clinical assessments, laboratory tests, and medical management were recorded. The COMPERA 2.0 four-layer risk assessment system was used to evaluate PAH, and patients in the intermediate or high-risk strata were defined as high-risk, and the rest were low-risk.
[0049] S2.2. Sample processing and sequencing
[0050] Peripheral blood samples of patients were collected. Red blood cells were lysed with ammonium chloride-potassium lysis buffer (Gibco). Through The blood RNA Kit (PreAnalytiX) extracts total RNA from peripheral blood samples. VAHTM mRNA Capture Beads (YEASEN) are used to isolate mRNA from the total RNA. When the mRNA quantity is in the range of 10 - 40 nanograms (ng), standard library construction is carried out. The Hieff NGS Ultima dual-mode mRNA Library Prep Kit for Illumina (YEASEN) is used for library construction. The mRNA is fragmented, and the first-strand cDNA is synthesized by reverse transcription using random primers, and then the second-strand cDNA is synthesized. The double-stranded cDNA is subjected to end repair and an A base is added to the 3' end. Then the sequencing adapter is ligated to the double-stranded DNA, and the ligation product is purified. PCR amplification is performed using a Thermal cycler S1000 (Bio-Rad) to generate the library. The library is subjected to quality control (QC) by qPCR (StepOne Plus (ABI)) to ensure the concentration is greater than 3 nM. The library passing the QC is arranged for subsequent sequencing. The sequencing is performed using a Novaseq X-plus (Illumina).
[0051] S2.3, Data processing and quality control
[0052] The RNA-seq data is output in fastq format and quality-controlled by FastQC (v 0.11.9) with default parameters, and all samples pass the QC. The sequencing data is aligned to the GRCh38 reference sequence using STAR (v 2.7.10b) 23. Then the sam file is sorted using samtools (v 1.16.1), and the bam file output by samtools is converted to read counts by featureCounts (a part of subread). All the above processes are executed in Linux.
[0053] Then the read count data is processed in R (v 4.3.1). Low-expression genes (the number of samples with gene counts ≥ 10 is less than the sample size of the smallest group [n = 95]) are regarded as noise and filtered out. Principal component analysis (PCA) does not find batch effects. Counts per million (CPM) is calculated from gene counts, and the calculation formula is as follows: CPM = Count * 10 6 / sum(Count). Transcripts per million (TPM) values are calculated from gene counts, and the formula is as follows:
[0054] S3. Analysis based on transcriptome samples
[0055] SLE-PAH is divided into three subtypes with different degrees of PAH, namely the high-inflammation subtype, the moderate-inflammation subtype, and the low-inflammation subtype.
[0056] S3.1. Differential gene expression analysis
[0057] To detect differentially expressed genes (DEGs), the R package DESeq2 was used with the read counts of 18,630 genes as input. Differential gene expression analysis was performed on SLE patients and SLE-PAH patients to demonstrate the pathogenesis of SLE-PAH. Analyses were also conducted between subtypes of SLE and SLE-PAH to identify the characteristics of each subtype.
[0058] S3.2. Pathway enrichment analysis
[0059] To elucidate the pathways involved in the pathogenesis of SLE-PAH, various pathway enrichment analyses were performed in the present invention. Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment, and Gene Set Enrichment Analysis (GSEA) were conducted in the clusterProfiler R package (v4.10.0). For the pathway datasets, the MsigDB hallmark gene sets (50 annotations) were used. The enrichment of GO and KEGG was observed in Cytoscape using Cluego. The DEGs used for GSEA were sorted according to the value calculated by the following formula: value = -log10(P value of differential gene expression analysis) / (sign(log2(Fold Change)).
[0060] S3.3. Consensus non-negative matrix factorization
[0061] To investigate the ideal number of clusters and classification subgroups of SLE-PAH, a clustering method based on non-negative matrix factorization (cNMF) was adopted. Clustering based on cNMF was performed using the NMF R package (v0.26). The TPM of 1324 DEGs detected between SLE-PAH and SLE was used for clustering based on cNMF. NMF was iterated 50 times, assuming a rank between 2 and 7, to determine the optimal rank and thus the number of subgroups. The optimal rank of the clustering was determined based on high silhouette and contour values. The seed 111111 was set manually to ensure the reproducibility of the analysis.
[0062] S3.4. Principal component analysis (PCA)
[0063] Principal component analysis was performed based on the TPM values of the previously detected DEGs. PCA was based on the FactoMineR R package (v2.9).
[0064] S3.5, Immune infiltration analysis (Cell-type Identification By Estimating Relative Subsets Of RNA Transcripts, CIBERSORT)
[0065] The transcriptome data and the TPM values of the LM22 marker genes were used as inputs for cell type identification by estimating relative subsets of RNA transcripts. CIBERSORT (v0.1.0) was run in absolute mode (sig.score as the absolute method) without quantile normalization.
[0066] S3.6, Protein-protein interaction analysis and hub gene identification
[0067] Protein-protein interaction (PPI) analysis was performed on the previously identified DEGs. The network was generated on STRING (website: https: / / cn.string-db.org / ). Then, PPI analysis was performed using the "MCODE" algorithm in Cytoscape software (v3.10.2).
[0068] Hub genes were identified among the DEGs using the cytoHubba plugin of Cytoscape (v0.1) software. Thirteen algorithms provided by cytoHubba were used to screen for hub genes. Genes designated as hub genes were finally determined by more than eight algorithms.
[0069] S3.7, Transcription factor prediction
[0070] Common transcription factors of the hub genes were predicted on ChEA3 (website: https: / / maayanlab.cloud / chea3 / ) 27).
[0071] S3.8, Disease enrichment analysis
[0072] The purpose of performing disease enrichment analysis was to determine the correspondence between genes and diseases. The CPM values of the previously identified DEGs were used for disease enrichment analysis. The DisGeNET enrichment function in the DOSE R package (v3.28.2) was used for disease enrichment analysis (using default parameters).
[0073] In this example, it was clarified that the pathogenesis of SLE-PAH is related to inflammation.
[0074] Specifically, differential gene expression analysis was performed on 100 SLE-PAH patients and 95 SLE patients. A total of 882 upregulated DEGs and 442 downregulated DEGs were identified, see Figure 1A and B in []. GSEA was performed based on DEGs. The inflammatory level in the SLE-PAH group was significantly higher than that in the SLE group. The upregulated DEGs were significantly enriched in TNFα signaling through the NFκB, interferon γ response, interferon α response, and inflammatory response gene sets, as shown in Figure 1 C in []. To better understand the pathogenic mechanism, GO analysis was performed, as shown in Figure 1 D in []. This analysis showed that the gene sets related to immune regulation and defense response were significantly enriched, which are the main functions of known TNFα and IFN. Multiple pathways involved in inflammatory activation, such as leukocyte activation, positive regulation of inflammatory response, and activation of immune response, were also enriched. This indicates that the inflammatory level in SLE-PAH patients is higher than that in SLE patients. Additional GSEA further confirmed the activation of the TNFα signaling pathway, as shown in Figure 1 D and E in []. Further, the expression of 21 interferon-stimulated genes (ISGs) between SLE-PAH and SLE was detected to clarify the activation of IFN-related pathways. These genes can indirectly indicate the activation of type I IFN. The expression of most of these genes was significantly increased in SLE-PAH, as shown in Figure 1 F in []. Notably, the gene sets related to the adaptive system, especially the function of B lymphocytes, were also identified, as shown in Figure 1 D in []. The enrichment of B cell activation, B cell receptor signaling pathway, and B cell regulation of immunity indicates that B cell activation is involved in SLE-PAH. In summary, the results indicate that inflammation is involved in the pathogenesis of SLE-PAH. At the same time, the activation of adaptive immunity, especially the adaptive immunity related to B cells, may also be involved in the occurrence of SLE-PAH.
[0075] In this example, three different subtypes of SLE-PAH were determined by CNMF clustering.
[0076] Specifically, cNMF clustering analysis was initially performed in 100 SLE-PAH patients based on the previously determined 1324 DEGs. See Figure 1 , where A, B, C, and G show that through rank screening, three different clusters were determined. A further 95 patients with systemic lupus erythematosus were included, and clustering analysis was performed again to verify the robustness of this clustering analysis. See Figure 1 , where D and E show the analysis results. Four clusters were determined, one of which consisted only of SLE patients, and two clusters consisted mainly of SLE-PAH patients. The classification of SLE-PAH in the two clustering analyses was compared, and it was found to be relatively stable with no significant differences ( Figure 1F). The ability to distinguish different subtypes of SLE-PAH from SLE was evaluated by principal component analysis. The above results indicate that although certain common features exist between cluster groups 1 and 2 of SLE-PAH and SLE, they can still be distinguished. However, cluster group 3 shows no distinguishable differences from SLE.
[0077] In this example, it was confirmed that different subtypes of SLE-PAH patients exhibit different degrees of inflammation.
[0078] Specifically, the differences between SLE-PAH subtypes were further elucidated by GSEA, and three clusters were found to exhibit different degrees of inflammation. See Figure 3 . According to the GSEA results, these three clusters were designated as high inflammation, moderate inflammation, and low inflammation. The activation of inflammation-related signaling pathways in the high-inflammation and moderate-inflammation SLE-PAH subtypes was significantly increased. In contrast, the low-inflammation group showed similar levels of these pathways compared to SLE patients. The three subtypes exhibited distinct clinical characteristics, as shown in Table 1. There was a statistically significant difference in the severity of PAH among the subtypes of SLE-PAH patients (P = 0.019). Patients in the high-inflammation group showed the mildest PAH. On the contrary, patients in the moderate-inflammation group showed the most severe PAH, and almost all patients at high risk of PAH belonged to this group. At the same time, there were no significant differences in demographic characteristics such as gender and age, as well as immunosuppressive therapy and PAH-specific therapy, among the three subtypes of SLE-PAH patients; there were significant differences in clinical manifestations and inflammation levels. See Figure 2 、 Figure 3 and Table 2.
[0079] Table 1 Clinical and demographic characteristics of SLE-PAH patients
[0080]
[0081]
[0082] Table 2 Clinical characteristics of SLE patients and SLE-PAH patients
[0083]
[0084]
[0085] In this example, it was confirmed that IL-6 plays a role in the high-inflammation SLE-PAH subtype.
[0086] Specifically, for the differential gene expression analysis between SLE and each subtype of SLE-PAH, see Figure 4Among them, A shows that the highly inflammatory subtype of SLE-PAH has the most significant differences compared with the SLE control group. A total of 4,433 DEGs were identified, of which 3,298 were highly inflammatory subtype-specific DEGs. As mentioned before, the low-inflammatory subtype has the smallest difference from SLE, with only 1,031 different genes. GO analysis and KEGG enrichment analysis were performed on the uniquely upregulated genes (1,707 out of 3,298) in the highly inflammatory subtype to further clarify its characteristics. See Figure 4 B and C in Figure 4 D in Figure 3 shows the calculation of the ISG score for each subtype, reflecting the IFN level. The highly inflammatory group has the highest ISG score. The TNFα-related pathway is also significantly upregulated in the highly inflammatory subtype. The GSEA results show that the IL6 JAK STAT3 signaling pathway is significantly upregulated in the highly inflammatory subtype and not observed in other subtypes. See Figure 4 The PI3K AKT mTOR signaling pathway, usually activated by IL-6, is also upregulated in the highly inflammatory subtype. KEGG enrichment analysis shows that IL6-related pathways such as Th17 cell differentiation, in addition to non-specific inflammatory pathways. Further, GSEA based on the GO gene set was used to study the regulation of the IL6 pathway in the highly inflammatory subtype. Figure 4 The results in G and H in
[0087] show a significant increase in the production of IL-6. By analyzing the proportion of immune cells through CIBERSORT,
[0088] I in Figure 5 shows that the highly inflammatory subtype has the lowest proportion of Treg cells, which may be caused by IL-6. The highly inflammatory subtype also has the highest proportion of neutrophils, which is consistent with its inflammatory level. Figure 5as shown in B of Figure 5 C. Vasoconstriction and remodeling are the main pathogenic mechanisms of PAH, and abnormal activation of the sympathetic nervous system in PAH patients has also been reported. The results showed that the SLE-PAH subtype with moderate inflammation exhibited the characteristics of IPAH. Then, the hub genes of the SLE-PAH moderate inflammation subtype were identified by cytoHubba, and the common TFs of the hub genes in chEA3 were predicted. A total of 12 hub genes were identified, and they shared two common TFs: GATA1 and KLF1. Notably, GATA itself is one of the 12 hub genes, see Figure 5 D and E in Figure 5 F. Overall, the SLE-PAH subtype with moderate inflammation showed the characteristics of IPAH in addition to elevated inflammation levels. These two TFs, GATA1 and KLF1, could be potential therapeutic targets.
[0089] In summary, in this example, by analyzing the transcriptomes of peripheral samples of SLE-PAH patients and comparing them with the transcriptomes of SLE patients, the pathogenesis of SLE-PAH was elucidated. Using transcriptome data, cNMF was used to classify SLE-PAH patients into three subtypes according to different PAH severities. In addition, the mechanistic characteristics of each category were clarified. For the most severe subgroup of PAH, hub genes and their common transcription factors were discovered, which could be used as prospective therapeutic targets.
[0090] In this example, the pathogenesis of SLE-PAH was elucidated by differential expression gene analysis and pathway enrichment analysis of SLE and SLE-PAH. The research results showed that inflammation plays an important role in SLE-PAH, and key cytokines involved in the pathogenesis were identified.
[0091] TNFα is a cytokine produced by various immune cells, including macrophages, lymphocytes, natural killer cells, etc. TNFα is generally considered an immunomodulator that affects the development of immune cells and mediates inflammatory processes. At the same time, TNFα is an important regulator of apoptosis and can play pro-apoptotic and anti-apoptotic roles simultaneously depending on different circumstances. In this example, TNFα-related pathways and gene sets were enriched in SLE-PAH by GSEA, indicating that the function of TNFα is activated in SLE-PAH compared with SLE. The infection-related gene sets may also be caused by the activation of TNFα signaling. In addition to inducing endothelial cell apoptosis, the abnormal activation of immune cells induced by TNFα may also be involved in the occurrence of SLE-PAH.
[0092] The IFN family is another group of cytokines that play roles in defending against viral infections and regulating inflammatory processes. In this example, it was found that both the IFNα and IFNγ signaling pathways were upregulated in SLE-PAH patients compared with SLE patients. The IFN signature was also analyzed, and the enrichment analysis was verified.
[0093] Overall, the inflammatory level in SLE-PAH patients is higher than that in general SLE patients. Through drug treatment targeting certain pathways, including TNFα and IFN that play important roles, there may be new treatment methods for SLE-PAH. In this example, it was also suggested that IL-6 and its downstream pathways, especially JAK / STAT3, may be therapeutic targets for SLE-PAH patients with high inflammation.
[0094] In this example, SLE-PAH patients with confirmed moderate inflammation showed the most severe PAH and had the characteristics of IPAH at the same time. It is worth noting that almost all high-risk SLE-PAH patients were classified into the moderate inflammation group. The results of gene set enrichment and PPI analysis in this example showed that in addition to activating inflammation, the intermediate subgroup had different potential mechanisms. Disease enrichment analysis identified diseases related to vascular stability, nerve impulse transmission, and smooth muscle stimulation, which have a common pathogenic process with IPAH. In addition, among the 12 hub genes, only FKBP8 and BCL2L1 mainly play roles in the immune process. SLC2A1 has been found in a PAH rat model and induces metabolic disorders. GATA1 is a transcription factor that is expected to regulate all 12 hub genes including itself. Generally speaking, the central genes of most moderate inflammation subtypes are involved in the pathogenesis of PAH. The moderate inflammation subtype shows both PAH-specific pathological factors related to vascular smooth muscle homeostasis and excessive inflammation activation related to SLE, causing a double blow and leading to more severe PAH in this group. That is, patients who simultaneously show inflammation activation and abnormal PAH-specific pathogenic characteristics have a much higher severity of PAH.
[0095] Example 2
[0096] This example provides a method for constructing a prediction model for the moderately inflammatory subtype of SLE-PAH. Based on the typing method described in Example 1, 4 characteristic genes were screened, and a prediction model was constructed using multiple logistic regression. The 4 characteristic genes are ALAS2, FKBP8, SLC2A1, and UBXN6.
[0097] Specifically, the development and validation of the prediction model include:
[0098] Using the Least Absolute Shrinkage and Selection Operator (LASSO) model to classify 12 hub genes and 1 transcription factor as the best prediction genes. The lambda value was determined by 10-fold cross-validation. 4 genes were selected to establish the prediction model. A prediction model was constructed using multiple logistic regression. The bootstrap method was used for internal validation, and the performance of the model was evaluated using the Receiver Operating Characteristic (ROC) and calibration curves.
[0099] Statistical analysis was performed, including:
[0100] When presenting baseline data, categorical data were presented as percentages, and continuous data were presented as the mean and standard error (SE) of the normal distribution. To compare continuous variables following a common distribution, the Student's t-test was applied. For categorical variables, the Pearson chi-square test or Fisher's exact test was used. A p-value less than 0.05 was considered statistically significant. Statistical tests were performed using R (4.3.1) and Graphpad Prism (10.1.2) software.
[0101] In this example, a prediction model for identifying SLE-PAH patients with the moderately inflammatory subtype was established. Specifically, considering that patients with the intermediate inflammatory subtype of SLE-PAH have more severe clinical symptoms, a prediction model was constructed to identify such patients. The log2-transformed CPM values of 13 hub genes and TFs from 100 SLE-PAH patients were used to construct the model. Through LASSO regression, 4 genes (ALAS2, FKBP8, SLC2A1, UBXN6) were finally included in the final model, see Figure 6 A and B in. Then, multiple-factor logistic regression was performed to construct the model, as shown in Table 3 below. All 4 genes were identified as risk factors. A calibration plot was used for internal validation, and the results showed that the model fit very well, see Figure 6 C in. An ROC curve was plotted to evaluate the predictive ability of the model. The C-index of this model was 0.954, indicating that the model had strong predictive ability, see Figure 6 D in. A Nomogram was drawn for easy clinical operation, see Figure 6E in it. Patients with a score above 98 are recommended to be classified into the moderate inflammation subtype. In summary, 86% of patients with the moderate inflammation subtype can be identified by the model.
[0102] Table 3 Results of multivariate logistic regression
[0103]
[0104]
[0105] In this embodiment, a prediction model based on the expression levels of four genes is established to identify SLE-PAH patients with the severe subtype, laying a foundation for clinical practice.
[0106] Example 3
[0107] This embodiment provides a detection kit for the moderate inflammation subtype of SLE-PAH. Based on the typing method described in Example 1, the kit includes reagents for detecting characteristic genes, and the characteristic genes are ALAS2, FKBP8, SLC2A1, and UBXN6.
[0108] Example 4
[0109] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the transcriptomics-based SLE-PAH typing method described in Example 1 or the method for constructing a prediction model for the moderate inflammation subtype of SLE-PAH described in Example 2 are completed.
[0110] Example 5
[0111] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the transcriptomics-based SLE-PAH typing method described in Example 1 or the method for constructing a prediction model for the moderate inflammation subtype of SLE-PAH described in Example 2 are completed.
[0112] The electronic device can be a mobile terminal and a non-mobile terminal. The non-mobile terminal includes a desktop computer, and the mobile terminal includes a smart phone (such as an Android phone, an IOS phone, etc.), smart glasses, a smart watch, a smart bracelet, a tablet computer, a laptop computer, a personal digital assistant, and other mobile Internet devices that can perform wireless communication.
[0113] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0114] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0115] In the implementation process, the steps of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with this embodiment may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0116] Obviously, the above embodiments are only examples given for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A transcriptomics-based SLE-PAH typing method, characterized in that: The method comprises the following steps: S1. Collect research samples, including SLE patients and SLE-PAH patients; S2, data collection and processing to obtain transcriptome samples of SLE patients and SLE-PAH patients; S3, analysis based on transcriptome samples; SLE-PAH was divided into three subtypes with different PAH degrees, different inflammation degrees, and different activation levels of inflammation-related signaling pathways, which were named high inflammation subtype, moderate inflammation subtype, and low inflammation subtype, respectively.
2. The typing method according to claim 1, characterized in that: The S2 includes: S2.1, data collection and risk assessment, the data collection includes collecting blood samples from patients, and the risk assessment includes evaluating PAH using a four-tier risk assessment system; S2.2, sample processing and sequencing: collect peripheral blood samples from patients, extract total RNA, establish libraries and sequence; S2.
3. Data processing and quality control.
3. The typing method according to claim 1, characterized in that: The S3 includes: S3.1, differential gene expression analysis was performed in SLE patients and SLE-PAH patients; S3.2, multiple pathway enrichment analysis, including gene ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment and gene set enrichment analysis (GSEA); S3.3, consensus non-negative matrix factorization, adopts a clustering method based on non-negative matrix factorization (cNMF).
4. The typing method according to claim 3, characterized in that: The S3 further includes: S3.4, perform principal component analysis (PCA) based on the test results of S3.1; S3.5, immune infiltration analysis (CIBERSORT); S3.
6. Perform protein-protein interaction analysis and hub gene identification based on the test results of S3.1; S3.7, prediction of common transcription factors of hub genes; S3.
8. Disease enrichment analysis.
5. The typing method according to claim 1, characterized in that: The hyperinflammatory subtype has the mildest features of PAH.
6. The typing method according to claim 5, characterized in that: The hyperinflammatory subtype is characterized by a significant increase in IL6.
7. The typing method according to claim 1, characterized in that: The moderately inflammatory subtype has features of the most severe PAH manifestations and IPAH.
8. The typing method according to claim 7, characterized in that: The moderate inflammatory subtype has 12 hub genes: SLC25A39, CDC34, FKBP8, SLC2A1, ALAS2, FOXO3, UBXN6, DMTN, DCAF12, BCL2L1, SLC4A1, and KLF1; and the characteristic transcription factors are GATA1 and KLF1.
9. A method for constructing a prediction model for moderate inflammatory subtype of SLE-PAH based on the typing method according to any one of claims 1 to 8, characterized in that: Four characteristic genes were screened and a prediction model was constructed using multivariate logistic regression. The four characteristic genes were ALAS2, FKBP8, SLC2A1, and UBXN6.
10. A detection kit for moderate inflammatory subtype of SLE-PAH based on the typing method according to any one of claims 1 to 8, characterized in that: The kit comprises reagents for detecting characteristic genes, and the characteristic genes are ALAS2, FKBP8, SLC2A1, and UBXN6.