Systemic lupus erythematosus activity detection marker composition, model construction and application thereof
By using immunomics and transcriptomics technologies, specific markers of CD4+ T cells, CD8+ T cells, and B cells were analyzed, and a multi-parameter model was constructed. This solved the problem of predicting the active phase of SLE and enabled precise diagnosis and personalized treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
- Filing Date
- 2026-01-28
- Publication Date
- 2026-06-02
AI Technical Summary
Current technologies lack effective methods to predict the active phase of systemic lupus erythematosus (SLE), and traditional treatment strategies are difficult to meet individualized needs, lacking precise biomarkers and molecular subtyping guidance.
Using immunomics and transcriptomics technologies, a multi-parameter diagnostic model was constructed by analyzing the V and J gene combinations of the TCR chains of CD4+ T cells and CD8+ T cells, the BCR IgH chain of B cells, and the expression of specific genes. Combined with machine learning, this model can distinguish between healthy individuals, those in the stable phase of SLE, and those in the active phase.
It enables accurate diagnosis of active SLE, improves early diagnosis rate and dynamic monitoring of disease activity, supports adjustments to personalized treatment plans, and enhances the sensitivity and specificity of diagnostic tools.
Smart Images

Figure CN122128419A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of disease medical detection technology, specifically relating to a systemic lupus erythematosus activity marker composition and its model construction and application. Background Technology
[0002] Systemic lupus erythematosus (SLE) is a typical disease caused by an adaptive immune system disorder, characterized by the loss of autoimmune tolerance, leading to the abnormal activation of a large number of autoreactive T cells and B cells. These abnormal cells can produce a large number of autoantibodies (such as anti-dsDNA and anti-Sm antibodies), further inducing immune complex deposition and tissue damage. In-depth research into the function and molecular mechanisms of these autoreactive lymphocytes is crucial for understanding the pathological mechanisms of SLE.
[0003] However, the immunopathological mechanisms of SLE are highly heterogeneous, and the pathogenic mechanisms can vary significantly from patient to patient. For example, in some patients, T-cell helper abnormalities (such as increased Tfh cells and Th1 / Th17 cell imbalance) are the main driving force, while in others, abnormal B-cell activation and plasma cell proliferation are the main factors. Therefore, elucidating this heterogeneity is crucial for identifying novel molecular targets for SLE and guiding personalized treatment, but current research still has many limitations.
[0004] The clinical manifestations and treatment responses of SLE vary from person to person, and traditional "one-size-fits-all" treatment strategies are insufficient to meet the individualized needs of patients. The core of precision medicine lies in guiding personalized treatment through biomarkers and molecular subtyping. Immune multi-omics (including immunomic sequencing and transcriptomics) is a cutting-edge technology for studying the complexity of the immune system in SLE. However, current research mostly applies these technologies independently, lacking integrated analysis of multi-omics data.
[0005] On the other hand, systemic lupus erythematosus (SLE) is divided into stable and active phases. Although existing methods are relatively mature for diagnosing stable SLE, there is still a lack of diagnostic tools to predict the active phase of SLE.
[0006] Therefore, how to utilize existing technologies to discover biomarkers for the detection and diagnosis of active SLE and obtain more precise and personalized treatment methods is a new direction for the prevention and treatment of this disease. However, this problem has not yet been studied in sufficient depth and urgently needs to be solved. Summary of the Invention
[0007] The present invention aims to solve the aforementioned technical problems by providing a biomarker composition for detecting systemic lupus erythematosus (SLE) activity, its model construction, and its application. The technical objective of this invention is to provide a method for the auxiliary diagnosis of active SLE based on immunomic sequencing and transcriptome sequencing. This method aims to differentiate between healthy individuals, those in the stable phase of SLE, and those in the active phase of SLE through multi-parameter analysis, as well as individuals with severe symptoms within the active SLE population, thus providing a new strategy for the accurate diagnosis of SLE.
[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: The present invention first provides a systemic lupus erythematosus (SLE) activity marker composition, which includes: a combination of V and J genes of the TCR alpha chain and TCR beta chain of CD4+ T cells and CD8+ T cells, a combination of V and J genes of the BCR IgH chain of B cells, and CD38, S100A11, IFITM3, GPX1 and NCF1 genes.
[0009] Furthermore, the combinations of V and J genes in the TCR alpha chain include: TRAV20_TRAJ6, TRAV4_TRAJ9, TRAV10_TRAJ18, TRAV13−1_TRAJ20, TRAV38−2 / DV8_TRAJ27, TRAV38−2 / DV8_TRAJ39, TRAV9−2_TRAJ44, TRAV38−2 / DV8_TRAJ45, TRAV17_TRAJ47, TRAV12−3_TRAJ48, and TRAV38−2 / DV8_TRAJ49.
[0010] Furthermore, the combinations of V and J genes in the B cell BCR IgH chain include: IGHV1−2_IGHJ1, IGHV3−15_IGHJ1, IGHV3−23_IGHJ1, IGHV3−48_IGHJ2, IGHV4−59_IGHJ2, IGHV3−64_IGHJ3, IGHV3−66_IGHJ3, IGHV3−7_IGHJ3, IGHV1−58_IGHJ4, IGHV4−34_IGHJ4, IGHV1−18_IGHJ5, IGHV2−26_IGHJ5, IGHV1−18_IGHJ6, IGHV1−24_IGHJ6, and IGHV1−3_IGHJ6.
[0011] Furthermore, the B cell BCR IgH chain includes the following subtypes: IgM, IgD, IgA, and IgG.
[0012] A second objective of this invention is to provide a systemic lupus erythematosus activity detection model comprising the activity detection marker composition described above.
[0013] The third objective of this invention is to provide a method for constructing a detection model for systemic lupus erythematosus (SLE), comprising the following steps: (1) Sample collection and sequencing: Peripheral blood samples were collected from healthy individuals, individuals in the stable phase of SLE, and individuals in the active phase. Immunome sequencing was performed on the TCR alpha chain and beta chain of CD4+ T cells and CD8+ T cells, respectively. IgH sequencing of BCR heavy chain of B cells and transcriptome sequencing of CD8+ T cells were also performed. (2) Immunome analysis: By calculating the clonal diversity, VJ combination and inter-sample similarity of BCR heavy chain of CD4+ T cells, CD8+ T cells and B cells respectively, the immunoome characteristics of healthy people, stable SLE and active SLE were revealed; (3) Transcriptome analysis: By analyzing differential gene expression and signaling pathways in CD8+ T cells, we identified genes that were specifically highly expressed in patients with active SLE. (4) Diagnostic model construction: Based on the key parameters of the immunomic and transcriptomic structures of CD4+ T cells, CD8+ T cells, and B cells, a diagnostic model was constructed to distinguish between the stable phase of SLE and the active phase of SLE, as well as the population with severe symptoms among the active SLE population.
[0014] Furthermore, the diagnostic model is constructed by including the inverse Simpson diversity of the TCR beta chain, the DE50 index, and the normalized Shannon diversity index for the SLE population.
[0015] Furthermore, a diagnostic classification model was constructed using ROC curves.
[0016] Furthermore, the construction of the diagnostic model includes a summary analysis of specific diversity and clonal parameters and parameter combinations of TCR and BCR, loss of VJ combinations, and differentially expressed genes in the transcriptome. A multi-parameter model is constructed by combining machine learning, and the AUC value is obtained by predicting using the model and analyzing using ROC curves.
[0017] The fourth objective of this invention is to provide a systemic lupus erythematosus (SLE) activity detection model for the development of reagents for detecting active SLE.
[0018] The beneficial effects of this invention are as follows: (1) This invention utilizes immunomics and transcriptome blood sequencing technologies to reveal the status, function, and dynamic change patterns of immune cell subsets in SLE patients, elucidating the multidimensional characteristics of immune imbalance and laying a theoretical foundation for precision medicine and stratified treatment. This invention also differentiates between the stable and active phases of SLE patients, developing a method for detecting SLE activity. This method will promote the development of precision medicine for SLE and provide reference for research on other autoimmune diseases (such as rheumatoid arthritis and multiple sclerosis).
[0019] (2) The method of this invention utilizes multi-omics data and machine learning technology to discover a series of biomarkers that reflect the characteristics of a patient's immune system, thereby achieving precise stratification of patients, such as the stable and active phases of SLE. Developing simple and effective diagnostic tools (such as ELISA kits or qPCR detection) can not only improve the early diagnosis rate of SLE but also dynamically monitor the patient's disease activity and predict the risk of relapse. The clinical application of these tools will help physicians adjust treatment plans in a timely manner, avoiding overtreatment or undertreatment. Attached Figure Description
[0020] Figure 1 Analysis of clonal and diversity parameters in infected and non-infected groups.
[0021] Figure 2 Clinical relevance analysis of TCR beta chain clonalness and diversity.
[0022] Figure 3 Combinations of TCR beta chain clonalness and diversity parameters can predict infection.
[0023] Figure 4 Clinical relevance analysis of BCR IgM chain clonalness and diversity.
[0024] Figure 5 Combinations of BCR IgM chain clonalness and diversity parameters can predict infection.
[0025] Figure 6 Analysis of differences in BCR IgM chain VJ combination frequencies between infected and non-infected groups.
[0026] Figure 7 Specific BCR IgM chain VJ combinations showed highly significant differences between the infected and non-infected groups.
[0027] Figure 8 Similarity analysis between TCR beta chain samples.
[0028] Figure 9 Analysis of the similarity difference of TCR beta chains in infected and non-infected groups.
[0029] Figure 10 Similarity analysis between BCR IgM chain samples.
[0030] Figure 11 Analysis of the similarity difference of BCR IgM chains between infected and non-infected groups.
[0031] Figure 12 Clinical relevance analysis of differentially expressed genes in infected and non-infected groups. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is described in detail below with reference to embodiments. It should be noted that the following embodiments are for explanation and illustration only and are not intended to limit the invention. Non-essential improvements and adjustments made by those skilled in the art based on the above description are still within the scope of protection of this invention.
[0033] Example 1
[0034] (a) Acquisition of immunomic and transcriptomic data Informed consent was obtained from all participants for all sample collection and processing, and the procedures complied with the relevant ethics committee approval requirements. Samples were collected and processed according to the following procedures: 1. Selection of research subjects The samples were divided into three groups: healthy individuals, patients with stable SLE, and patients with recurrent SLE. At least 30 samples were collected from each group to ensure statistical significance.
[0035] 2. Blood collection procedure Five mL of peripheral venous blood was collected from each subject and stored in EDTA anticoagulant tubes. Samples were immediately placed at 4°C after collection to minimize processing delays (not exceeding 2 hours).
[0036] 3. Extraction of PBMCs by density gradient centrifugation Collected anticoagulated blood samples were mixed with phosphate-buffered saline (PBS) at a 1:1 (V:V) ratio; the mixed blood was slowly added to a pre-prepared Ficoll-Paque density gradient separation solution; centrifuged at 800 × g for 20 minutes at room temperature without braking; mononuclear cells (PBMCs) were extracted from the interface layer formed after centrifugation and placed in a new centrifuge tube. PBMCs were washed twice with PBS (400 × g, 10 min); trypan blue staining was used to count viable cells in PBMCs, ensuring cell viability >90%. Finally, the PBMCs were suspended in a cryoprotectant (such as RPMI-1640 + 10% DMSO), aliquoted, and immediately frozen in liquid nitrogen.
[0037] 3. CD8+ T cell sorting CD8+ cells and CD4+ cells were separated using a magnetic bead sorting method, following the manufacturer's standard procedure. In short, CD8+ T cells were first sorted by labeling them with anti-CD8 magnetic beads (Biolegend MojoSort™ Human CD8+ T Cell Sorting Kit). The sorted CD8+ T cells were immediately lysed to extract mRNA. Simultaneously, the CD8- cell population was also lysed for RNA purification.
[0038] 4. RNA extraction Total RNA was extracted from sorted CD8+ T cells and remaining PBMCs using a commercial RNA extraction kit (TRIzol reagent). The kit instructions were followed to ensure RNA purity and integrity (A260 / A280 ratio between 1.8 and 2.0). Genomic DNA contamination was removed using DNase I; RNA integrity was assessed using agarose gel electrophoresis; and RNA concentration and purity were determined using a NanoDrop or Qubit instrument.
[0039] 5. Immunomics and transcriptome library sequencing The basic steps for immunohistomic amplification and library construction of the extracted samples are: (1) first-strand cDNA synthesis; (2) cDNA amplification; (3) TCR amplification and BCR amplification; (4) immunohistomic library construction.
[0040] Transcriptional library construction was performed using standard experimental methods, with the following basic steps: A sequencing library was constructed from whole transcriptome RNA samples using an RNA-Seq library construction kit; the RNA was further fragmented and reverse transcribed into cDNA; adapter sequences were added and the fragmented library was purified. After library construction, sequencing was performed using either the BGI T7 sequencing platform or the Illumina NovaSeq 6000 sequencing platform; immunohistomic and transcriptome sequencing were performed in 150 bp paired-end (PE150) mode; after sequencing data quality control, the raw sequencing data quality was assessed using FastQC software; low-quality reads and adapter sequences were removed using Trimmomatic software. Qualified sequencing data were stored in FASTQ format for subsequent analysis.
[0041] (II) Analysis of Clonal and Diversity of Immunome Data Immunoclonal diversity of CD4+ T cells, CD8+ T cells, and B cells from peripheral blood of each group was obtained and analyzed. The basic procedure is as follows: 1. Acquisition of raw sequencing data Immunome sequencing data were obtained using the method described in (I). The data included the TCR alpha and beta chains of CD4+ T cells, the TCR alpha and beta chains of CD8+ T cells, and the BCR heavy chain IgH sequence information of B cells. Immunome-specific analysis software (such as MiXCR or IMGT / HighV-QUEST) was used to assemble and quality-filter the sequences: removing low-quality reads (Q value <30); removing short fragments (length <150 bp) and adapter sequences. The assembled sequences were aligned with a reference database (such as the IMGT database), and the V, J, C, and CDR3 regions of the TCR and BCR sequences were annotated.
[0042] 2. Clonality and Diversity Analysis The inverse Simpson index, Shannon's index / normalized Shannon diversity entropy (NSDE), and DE50 value were used to assess the clonal and diversity differences among different groups. Statistical methods were then used to determine the significance of the differences between samples.
[0043] Formula for calculating the Inverse Simpson Diversity Index: ; Where: 1 / D is the pseudo-Simpson index, a larger value indicates higher diversity; N is the total number of different clones in the sample; Pi is the relative abundance of the i-th clone (i.e., the proportion of the count of that clone to the total count). The inverse Simpson diversity index is the reciprocal of the Simpson index; a larger inverse Simpson index indicates higher clonal diversity. It focuses on reflecting the diversity of high-frequency clones.
[0044] The Shannon index is calculated using the following formula: ; Where: H' is the Shannon diversity index; S is the total number of clones; pi is the relative abundance of the i-th clone (pi = ni / N, i.e., ni is the count of the i-th clone, and N is the sum of the counts of all clones).
[0045] Normalized Shannon Diversity Index: The normalized Shannon diversity index is normalized by dividing by the maximum value of the Shannon index, Hmax. Its calculation formula is: ; Where: Hnorm is the normalized Shannon diversity index; Hmax = ln(S), which is the theoretical maximum value of the Shannon index, reaching its maximum when the relative abundance of all species is equal (i.e., pi = 1 / S). The value of Hnorm ranges from [0,1]: close to 0: indicates the presence of a significantly dominant clone in the sample (low diversity). Close to 1: indicates that all clones have equal abundance (high diversity). The advantage of the normalized Shannon index is that it allows for fair comparisons between samples with different numbers of clones because it normalizes the Shannon index to a fixed range. In this study, the normalized Shannon diversity index is labeled as normalized_shannon.
[0046] The formula for calculating DE50 is: ; The formula is calculated as follows: Cloning sequences are sorted from highest to lowest frequency, and the sequences are accumulated starting from the highest frequency. The proportion of clones whose sum of frequencies reaches 50% is the percentage of all clones. DE50, as an indicator of clonalness (clonal degree), measures the uniformity of clonalness. A higher DE50 value indicates a more even distribution of clones, while a lower DE50 value indicates higher clonalness, suggesting that some specific clones have been amplified.
[0047] Furthermore, this study categorized the frequencies of clones in the TCR and BCR libraries, and then compared the differences between different samples after categorization. The frequency categorization of clones in the TCR and BCR libraries was based on their relative abundance ranking in the total clone library for each sample. We divided different clones into the top 100, top 1000, and top 10000 clone groups. In this way, the proportion of the top 100 clone group was defined as the proportion of the 100 most frequent clones in the total clones (denoted as raw_100inall, and so on). Correspondingly, we can calculate diversity parameters such as the quasi-Simpson index for the top 100, top 1000, and top 10000 clone groups.
[0048] This study also used another method to classify clones in the TCR and BCR libraries, mainly based on the cloning frequency of a particular clone in the total clonal population, regardless of clone order. After classification, the differences between different samples were compared. Based on cloning frequency (Fi), the clones were divided as follows: rare clones (Fi < 1e-6), small clones (1e-6 ≤ Fi < 1e-5), medium clones (1e-5 ≤ Fi < 1e-4), large clones (1e-4 ≤ Fi < 1e-3), and hyper-expanded clones (1e-3 ≤ Fi).
[0049] 3. Analysis of the clonality and diversity of CD8+ T cells The clonal and diversity indices of the TCR beta chain (trb) and TCR alpha chain (tra) in CD8+ T cells were calculated using the methods described above.
[0050] The results after analysis show that ( Figure 1 The inverse Simpson diversity, DE50, and normalized Shannon diversity index of the TCR beta chain (trb) were significantly higher in the SLE population than in the non-infected population. A similar trend was observed in the TCR alpha (tra) chain. Therefore, the clonal diversity of the SLE population is significantly higher than that of the non-infected population, indicating that the immune system produces a strong specific response against its own pathogen.
[0051] 4. Analysis of the clonality and diversity of CD4+ T cells: The results after analysis show that ( Figure 2 In the SLE population (A), the inverse Simpson diversity, DE50, and normalized Shannon diversity indices of the TCR beta chain (trb) were significantly higher than those in the non-infected population. A similar trend was observed in the TCR alpha chain (A). Figure 2 (B) Additionally, it was found that the proportion of the top 100 clones and the proportion of superamplified clones in the SLE population were significantly higher than in healthy individuals ( ). Figure 2 (C). Therefore, the CD4+ T cell clonality in the SLE population is significantly higher than that in the healthy population, indicating that the adaptive immune system generates a strong specific immune response against its own pathogen.
[0052] This study extracted the TCR beta chain from CD4+ T cells for diagnostic analysis, exploring its clinical predictive value. The study found significant differences in two parameters of the TCR beta chain between mild-to-moderate SLE patients and severe patients: trb_inv_100 and trb_hyper. This study used ROC curves to evaluate the diagnostic ability of these two parameters: using these two parameters as input variables, a classification model (such as logistic regression) was constructed; the area under the ROC curve (AUC) for both parameters reached above 0.80, indicating that these combinations have good diagnostic performance. Figure 3 Further random clustering analysis was performed using the two parameters mentioned above. The study found that random clustering effectively distinguished between mild to moderate and severe SLE patients, indicating that the combination of the two parameters can effectively differentiate between mild to moderate and severe SLE patients, demonstrating high clinical application value.
[0053] 5. Analysis of B cell clonal diversity: First, the proportions of different BCR IgH subtypes—IgM, IgA, IgD, and IgG—were analyzed. The study found that patients with stable SLE had a significantly higher proportion of IgA compared to healthy individuals and relapsed patients. Conversely, patients with relapsed SLE had a significantly higher proportion of IgG compared to healthy individuals and stable patients. Figure 4 Further, ROC curves were used to evaluate the ability of the two parameters, IgA percentage and IgG percentage, to distinguish between patients in the relapsed and stable phases of SLE. These two parameters were used as input variables to construct a classification model (such as logistic regression). The area under the ROC curve (AUC) for both parameters was above 0.80, indicating that these combinations have good diagnostic performance for SLE relapse.
[0054] Further analysis of the clonal and diversity information of each subtype showed that the inverse Simpson diversity, DE50, and normalized Shannon diversity index of the IgM chain of BCR IgH in the SLE population were significantly higher than those in the non-infected population. Figure 5 In the IgD chain, a similar trend was observed. This study also found that the proportion of the top 100 clones in the IgM chain and the proportion of hyperamplified clones in both the IgD and IgG chains were significantly higher in SLE patients than in healthy individuals. Figure 5 (B). Therefore, the clonal rate in the SLE population is significantly higher than in the healthy population, indicating that the humoral immune system produces a strong specific immune response against the autogenous organism.
[0055] This study also extracted clonal and diversity data of different BCR IgH chain subtypes for diagnostic analysis, exploring their clinical predictive value. The study identified the following diversity and clonal parameters for various BCR IgH subtypes in SLE patients in stable and relapsed phases: iga_counts_ratio, igd_hyper, igg_hyper, igg_small, igg_raw_100inall, and igg_counts_ratio. The differences were significant. ROC curves were used to assess the ability of the two parameters to distinguish between patients in the relapsed and stable phases of SLE. The two parameters were used as input variables to construct a classification model (e.g., logistic regression). The areas under the ROC curves (AUC) for the above parameters were: iga_counts_ratio (AUC = 0.85), igd_hyper (AUC = 0.81), igg_hyper (AUC = 0.86), igg_small (AUC = 0.80), igm_raw_100inall (AUC = 0.79), and igg_counts_ratio (AUC = 0.84), indicating that these combinations have good diagnostic performance for SLE relapse. Figure 6Further random cluster analysis using the above multiple parameters revealed a good distinction between SLE relapse and stable phase patients after random clustering. This indicates that the combination of the above parameters can effectively differentiate between SLE relapse and stable phase patients, demonstrating high clinical application value.
[0056] Further research revealed significant differences in BCR IgH chain parameters between mild-to-moderate SLE patients and severe patients: iga_inv_100, iga_inv_1000, iga_inv_10000, iga_hyper, iga_large, igd_inv_10000, igg_raw_1000inall, igg_inv_100, and igg_medium. The diagnostic ability of the two parameters above was evaluated using ROC curves: These two parameters were used as input variables to construct a classification model (such as logistic regression); the area under the ROC curves (AUC) values for these parameters were: iga_inv_100 (AUC = 0.89), iga_inv_1000 (AUC = 0.88), iga_inv_10000 (AUC = 0.94), iga_hyper (AUC = 0.99), iga_large (AUC = 0.79), igd_inv_10000 (AUC = 0.81), igg_raw_1000inall (AUC = 0.86), igg_inv_100 (AUC = 0.89), and igg_medium (AUC = 0.81), indicating that these combinations have good diagnostic performance. Figure 7 Further random clustering analysis was performed using the above parameters. The study found that random clustering effectively distinguished between mild to moderate and severe SLE patients, indicating that the combination of the above parameters can effectively differentiate between mild to moderate and severe SLE patients, and has high clinical application value.
[0057] (III) Frequency analysis of VJ combination in the immunization group This section aims to identify key differences between healthy individuals and SLE patients by analyzing the frequency of VJ combination usage of TCR and BCR in immunomic sequencing data, providing new molecular markers for SLE diagnosis. This study acquired and analyzed the frequency of VJ combination usage in the immunomic sequence of CD8+ T cells, CD4+ T cells, and B cells.
[0058] 1. Data preprocessing, the basic steps are as follows: (1) Extraction of VJ combination information: Extract the V and J gene sequences of the TCR alpha and TCR beta chains of CD4+ T cells and CD8+ T cells, as well as the various subtypes (IgM, IgD, IgA and IgG) of the BCR IgH chain of B cells from the immunomic sequencing data; use immunomic analysis software (such as MiXCR or IMGT / HighV-QUEST) to compare and annotate the sequences and identify the VJ combination in each sample.
[0059] (2) Data standardization: The VJ combination frequencies of each sample are stored in matrix form: rows represent sample numbers; columns represent specific VJ combinations; cell values are the frequency or absolute reading of that VJ combination. Finally, combinations with sequencing quality lower than Q30 are removed.
[0060] 2. Results Analysis First, the average frequency of VJ combinations in each sample group was calculated, and the frequency data were normalized using a standardization method (such as Z-score). The frequency distribution of each VJ combination in different groups was compared. Statistical methods were used to screen for VJ combinations that showed differential expression between the SLE population and healthy individuals. The criteria for significant difference were: p-value < 0.05 and significant frequency change. A heatmap of VJ combination frequencies was plotted: the horizontal axis represents the sample number, and the vertical axis represents the specific VJ combination.
[0061] (1) Comparative analysis of the frequency of use of VJ combination in CD8+ T cell immunization group Experimental results showed that, compared with healthy individuals, SLE patients had a significant loss of VJ combinations in the TCR beta chain, and the frequency of use of these VJ combinations was extremely low in SLE patients. This study used heatmaps to illustrate the differentially expressed VJ combinations. Figure 8 ).
[0062] (2) Comparative analysis of the frequency of use of VJ combination in CD4+ T cells This study first analyzed the use of variable juxtapositions (VJs) in the TCR alpha chain of CD4+ T cells. The results showed that, compared to healthy individuals, SLE patients lost a significant number of VJs in the TCR alpha chain, and the frequency of use of these VJs was extremely low in SLE patients. This study used heatmaps to illustrate some differentially expressed VJ combinations (…). Figure 9The study found that the following TCR alpha chain VJ combinations were significantly lost in the SLE population: TRAV20_TRAJ6, TRAV4_TRAJ9, TRAV10_TRAJ18, TRAV13−1_TRAJ20, TRAV38−2 / DV8_TRAJ27, TRAV38−2 / DV8_TRAJ39, TRAV9−2_TRAJ44, TRAV38−2 / DV8_TRAJ45, TRAV17_TRAJ47, TRAV12−3_TRAJ48, and TRAV38−2 / DV8_TRAJ49. Further ROC curve analysis was used to assess the diagnostic ability of the lost CD4+ T cell TCR alpha chain VJ combinations: a classification model (such as logistic regression) was constructed using the frequency of CD4+ T cell TCR alpha chain VJ combinations as input variables. Based on the aforementioned differing VJ combinations, the accuracy of predicting SLE relapse was analyzed. The calculated area under the ROC curve (AUC) all reached above 0.85, indicating that these combinations have good diagnostic performance.
[0063] Analysis of VJ combinations in the TCR beta chain of CD4+ T cells revealed that, compared to healthy individuals, SLE patients exhibited a significant loss of VJ combinations in the TCR beta chain, and the frequency of these VJ combinations was extremely low in SLE patients. This study used heatmaps to illustrate some differentially expressed VJ combinations. Figure 10 These VJ combinations have clinical value in diagnosing SLE.
[0064] (3) Comparative analysis of the frequency of use of VJ combination in B cells By identifying and comparing VJ combinations using the BCR IgH chain, the study found that the frequency of some VJ combinations was significantly lower in SLE patients compared to healthy individuals. Furthermore, the study also found that the frequency of certain VJ combinations was high in SLE relapse patients compared to SLE patients in the stable phase. The study used a heatmap to illustrate the differential expression of IgH chain IgM subtypes in VJ combinations among healthy individuals, SLE patients in the stable phase, and SLE relapse patients. Figure 11Box plots were further used to illustrate the frequency distribution of IgH chain IgM subtypes in the three groups. For the VJ combination IGHV1−2_IGHJ1, IGHV3−15_IGHJ1, IGHV3−23_IGHJ1, IGHV3−48_IGHJ2, IGHV4−59_IGHJ2, IGHV3−64_IGHJ3, IGHV3−66_IGHJ3, IGHV3−7_IGHJ3, IGHV1−58_IGHJ4, IGHV4−34_IGHJ4, IGHV1−18_IGHJ5, IGHV2−26_IGHJ5, IGHV1−18_IGHJ6, IGHV1−24_IGHJ6, and IGHV1−3_IGHJ6, the frequency was significantly higher in the SLE relapse group than in the SLE stable group. The diagnostic capability of missing combinations was evaluated using ROC curves: a classification model was constructed using the frequency of missing VJ combinations in the BCR IgM subtype as the input variable. Based on the aforementioned VJ combinations with discrepancies, the accuracy in predicting SLE relapse patients was analyzed. The area under the ROC curve (AUC) was found to be above 0.85, indicating that these combinations have good diagnostic performance.
[0065] (iv) Differential expression analysis of CD8+ T cell transcriptome This section identifies specific genes that are significantly overexpressed in individuals with SLE relapse by analyzing differential gene expression in peripheral blood CD8+ T cell transcriptome sequencing data. These overexpressed genes have the potential to serve as biomarkers for predicting SLE relapse.
[0066] 1. Data Acquisition and Preprocessing Whole transcriptome sequencing data for each sample were obtained using the method described in section (I), in FASTQ file format. The quality of the sequencing data was assessed using FastQC software to remove low-quality reads (Q value < 30) and adapter sequences; data filtering was performed using Trimmomatic or similar tools to ensure the integrity of the sequencing data. The filtered high-quality data was aligned to a reference genome (e.g., the GRCh38 human genome): alignment tool: such as HISAT2 or STAR; output format: gene alignment file (BAM file). Alignment efficiency should reach 95% or higher. Gene expression quantification was performed using FeatureCounts or HTSeq to generate a gene expression matrix (TPM, FPKM, or raw counts format).
[0067] 2. Differential gene expression analysis Gene expression matrices were normalized (e.g., using TPM or DESeq2 normalization methods) to eliminate the effects of sequencing depth and inter-sample variability. Differential expression analysis tools (e.g., DESeq2 or edgeR) were used to compare gene expression levels in healthy individuals, patients with stable SLE, and patients with relapsed SLE. Judgment criteria: Fold Change (FC) > 2 (gene expression level upregulated more than 2-fold); p-value < 0.05 (significance test); FDR < 0.05 (multiple test correction).
[0068] 3. Clinical relevance assessment The study found that compared to healthy individuals and those in the stable phase of SLE, the following genes were significantly overexpressed in individuals experiencing SLE relapse: CD38, S100A11, IFITM3, GPX1, and NCF1. Further analysis of the diagnostic performance of these overexpressed genes for SLE relapse using ROC curves was conducted: gene expression levels were used as input variables to construct a classification model (e.g., logistic regression); the area under the ROC curve (AUC) was calculated for: CD38 (AUC = 0.76), S100A11 (AUC = 0.77), IFITM3 (AUC = 0.77), GPX1 (AUC = 0.72), and NCF1 (AUC = 0.69). Figure 12 ).
[0069] The combined model of the two genes showed an AUC > 0.90, indicating high diagnostic sensitivity and specificity. Based on the high expression characteristics of these genes, a specific molecular diagnostic kit was developed: qRT-PCR primers or probes were designed to detect the expression levels of these two genes in peripheral blood; this could be used for real-time diagnosis of infectious diseases. By combining the expression levels of these genes with immunomic data (such as clonality and diversity parameters), a multi-parameter diagnostic model was constructed; this improved the sensitivity, specificity, and accuracy of SLE relapse diagnosis. The high expression of these genes may play an important role in the pathological process of SLE relapse and could potentially serve as a therapeutic target for drug development in the future.
[0070] (v) Combined multi-parameter analysis of immunomic and transcriptomic data This study combined TCR and BCR specific diversity and clonal parameters and parameter combinations, VJ combination loss, and differentially expressed genes in the transcriptome with a summary analysis to construct a multi-parameter model using machine learning. The model's predictions and ROC curve analysis yielded a higher AUC value (AUC > 0.95), thus significantly improving the sensitivity and specificity of SLE relapse and severity prediction.
[0071] (vi) Summary This invention, through combined immunomic and transcriptomic analysis, innovatively discovers novel features and molecular markers for SLE diagnosis, such as specific diversity and clonal parameters and combinations of TCR and BCR, loss of VJ combinations, and overexpression of genes (CD38 and S100A11, etc.). The multi-parameter model constructed using machine learning significantly improves diagnostic sensitivity and specificity, providing a scientific basis for SLE subtyping and personalized treatment. This technology not only fills a gap in traditional diagnostic methods but also opens up new directions for precision medicine and basic research in SLE, possessing significant clinical application value and potential for technology dissemination.
[0072] Description of the inventive step of this invention: Existing research indicates that systemic lupus erythematosus (SLE) involves T-cell and B-cell abnormalities. Therefore, flow cytometry for cell subtype differentiation or TCR / BCR immunomic sequencing are obvious techniques for identifying SLE biomarkers. However, the adaptive immune system is highly complex. For example, there is significant variation in immune cell subtypes among individuals, and T cells are divided into CD8+ cytotoxic T cells and CD4+ helper T cells with different functions. CD4+ T cells can be further subdivided into even more functional cell subtypes. Therefore, using cell subtype sequencing alone or unclassified TCR sequencing is unlikely to be an accurate diagnostic tool. Furthermore, diagnostic tools for predicting SLE activity are currently lacking.
[0073] Therefore, we need to conduct refined cell subtype analysis and combined immune multi-omics analysis to find a diagnostic tool that can accurately predict the recurrence of systemic lupus erythematosus (SLE). Refined cell subtype analysis and immune multi-omics analysis will generate a large number of clinically relevant parameters, and it remains unknown which of these parameters can be combined to produce an accurate predictive model for SLE.
[0074] In this study, we included patients with stable and active systemic lupus erythematosus (SLE) and, for the first time, analyzed the immunology and clinical relevance of CD4+ T cells (TCR alpha, beta chain), CD8+ cells (TCR alpha, beta chain), and B cells (BCR IgH subtype). Simultaneously, we performed multi-omics immunological analyses: TCR, BCR immunomic, and transcriptomic analyses. Furthermore, we fitted data from TCR sequencing, BCR sequencing, CD8+ T cell transcriptomic sequencing, and flow cytometry to derive a predictive model for disease activity. This enabled the construction and detection of a refined, multi-dimensional relapse prediction model for active SLE.
Claims
1. A composition for detecting systemic lupus erythematosus activity, characterized in that, The systemic lupus erythematosus activity marker composition includes: a combination of the V and J genes of the TCR alpha chain and TCR beta chain of CD4+ T cells and CD8+ T cells, a combination of the V and J genes of the BCR IgH chain of B cells, and the CD38, S100A11, IFITM3, GPX1 and NCF1 genes.
2. The systemic lupus erythematosus activity marker composition according to claim 1, characterized in that, Combinations of the V and J genes in the TCRalpha chain include: TRAV20_TRAJ6, TRAV4_TRAJ9, TRAV10_TRAJ18, TRAV13−1_TRAJ20, TRAV38−2 / DV8_TRAJ27, TRAV38−2 / DV8_TRAJ39, TRAV9−2_TRAJ44, TRAV38−2 / DV8_TRAJ45, TRAV17_TRAJ47, TRAV12−3_TRAJ48, and TRAV38−2 / DV8_TRAJ49.
3. The systemic lupus erythematosus activity marker composition according to claim 1, characterized in that, The combinations of V and J genes in the B cell BCRIGH chain include: IGHV1−2_IGHJ1, IGHV3−15_IGHJ1, IGHV3−23_IGHJ1, IGHV3−48_IGHJ2, IGHV4−59_IGHJ2, IGHV3−64_IGHJ3, IGHV3−66_IGHJ3, IGHV3−7_IGHJ3, IGHV1−58_IGHJ4, IGHV4−34_IGHJ4, IGHV1−18_IGHJ5, IGHV2−26_IGHJ5, IGHV1−18_IGHJ6, IGHV1−24_IGHJ6, and IGHV1−3_IGHJ6.
4. The systemic lupus erythematosus activity marker composition according to claim 1, characterized in that, B cell BCRIgH chains include the following subtypes: IgM, IgD, IgA, and IgG.
5. A systemic lupus erythematosus activity detection model, characterized in that, It comprises the activity detection marker composition as described in any one of claims 1-4.
6. A method for constructing a systemic lupus erythematosus activity detection model, characterized in that, Includes the following steps: (1) Sample collection and sequencing: Peripheral blood samples were collected from healthy individuals, individuals in the stable phase of SLE, and individuals in the active phase. Immunome sequencing was performed on the TCR alpha chain and beta chain of CD4+ T cells and CD8+ T cells, respectively. IgH sequencing of BCR heavy chain of B cells and transcriptome sequencing of CD8+ T cells were also performed. (2) Immunome analysis: By calculating the clonal diversity, VJ combination and inter-sample similarity of BCR heavy chain of CD4+ T cells, CD8+ T cells and B cells respectively, the immunoome characteristics of healthy people, stable SLE and active SLE were revealed; (3) Transcriptome analysis: By analyzing differential gene expression and signaling pathways in CD8+ T cells, we identified genes that were specifically highly expressed in patients with active SLE. (4) Diagnostic model construction: Based on the key parameters of the immunomic and transcriptomic structures of CD4+ T cells, CD8+ T cells, and B cells, a diagnostic model was constructed to distinguish between the stable phase of SLE and the active phase of SLE, as well as the population with severe symptoms among the active SLE population.
7. The method according to claim 6, characterized in that, The diagnostic model was constructed using inverse Simpson diversity of the TCRbeta chain, DE50 index, and normalized Shannon diversity index in the SLE population.
8. The method according to claim 6, characterized in that, The construction of the diagnostic model includes a summary analysis of specific diversity and clonal parameters and parameter combinations of TCR and BCR, loss of VJ combinations, and differentially expressed genes in the transcriptome. A multi-parameter model is constructed by combining machine learning, and the AUC value is obtained by predicting using the model and using ROC curve analysis.
9. The application of the systemic lupus erythematosus activity detection model of claim 5 in the development of reagents for detecting active phase systemic lupus erythematosus.