Methods for treating autoimmune conditions

By isolating and targeting disease-associated TF signatures and pathways, the method addresses the challenge of identifying at-risk individuals and developing effective therapeutics for autoimmune conditions, enabling personalized treatment strategies.

WO2026076231A1PCT designated stage Publication Date: 2026-04-09RGT UNIV OF CALIFORNIA +13
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current methods are inadequate for identifying individuals at risk of developing autoimmune conditions like rheumatoid arthritis and effective therapeutics are lacking due to diverse pathogenic mechanisms and a lack of understanding of disease development pathways.

Method used

The method involves isolating macromolecules from a biological sample, determining disease-associated transcription factor (TF) signatures, and administering therapeutic agents targeting specific cell types or modulating pro-inflammatory downstream genes in pathways such as SUMOylation, RUNX2, YAP1, NOTCH3, and WNT/p-Catenin pathways.

Benefits of technology

This approach allows for personalized treatment strategies by identifying at-risk individuals and targeting specific TF signatures and pathways, potentially preventing or delaying the progression of autoimmune conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025049213_09042026_PF_FP_ABST
    Figure US2025049213_09042026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to methods of treating an autoimmune disease (e.g., RA) in a subject, the methods comprising: determining a disease-associated transcription factor (TF) signature, identifying cell types from a biological sample expressing the TF signature, and / or determining proinflammatory downstream genes that are regulated by the TF signature pathways; and administering to the subject a therapeutic agent targeting the identified cell types and / or the proinflammatory downstream genes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0002] METHODS FOR TREATING AUTOIMMUNE CONDITIONS

[0003] CROSS-REFERENCE TO RELATED APPLICATIONS

[0004] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 702.573, filed October 2. 2024. the contents of which are incorporated herein by reference in their entirety.

[0005] FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0006] This invention was made with Government support under Grant Numbers AR064194 and AR065466 awarded by the National Institutes of Health. The Government has certain rights in the invention.

[0007] TECHNICAL FIELD

[0008] This document relates to methods for treating autoimmune conditions (e.g.. rheumatoid arthritis). For example, this document relates to methods that include administering, to a subject in need thereof, a therapeutic agent targeting disease- associated cell types, transcription factors, or pro-infl ammatory genes.

[0009] BACKGROUND

[0010] Autoimmune conditions are a group of diseases in which the body's immune system mistakenly attacks its own cells, tissues, and organs in any part of the body, w eakening body function and even turning life-threatening. As many as 50 million people in the U.S. have an autoimmune disease, making it the third most prevalent disease category, surpassed only by cancer and heart disease.

[0011] Rheumatoid arthritis (RA) is one of the most prevalent autoimmune diseases, affecting about 1% of the global population. About 18 million people w orldwide w ere living with RA. Untreated, RA can cause severe damage to the joints and their surrounding tissue. It can lead to heart, lung or nervous system problems. Despite this, there remains a dearth of methods of identifying persons at risk of developing RA and methods of identifying helpful therapeutics. Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0012] SUMMARY

[0013] Provided herein are methods for treatment of an autoimmune condition in a subject, the method comprising (a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample; (c) identifying the one or more cell t pes from the biological sample expressing the TF signature; and (d) administering to the subject a therapeutic agent comprising a therapeutic agent targeting the one or more cell types expressing the TF signature.

[0014] Also provided herein are methods for treatment of an autoimmune condition in a subject, the method comprising (a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample; and (c) administering a therapeutic agent to modulate the expression of one or more pro-inflammatory downstream genes regulated by the TF signature, wherein the TF signature comprises elevated expression of one or more TFs of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways.

[0015] In some embodiments, the TF signature comprises elevated expression of one or more TFs in one or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways.

[0016] In some embodiments, the one or more cell types comprise B memory' cells, B intermediate cells, B naive cells, CD14 monocytes, CD16 monocytes, CD4 naive T cells, central memory’ CD4 T cells (CD4 TCM). CD8 naive T cells, effector memory CD8 T cells (CD8 TEM), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright), and / or regulatory' T cells (Treg).

[0017] In some embodiments, the methods for treatment of an autoimmune condition in a subject further comprise determining expression of one or more proinflammatory downstream genes regulated by the TF signature. In some embodiments, the methods further comprise identify ing one or more cell types in the biological sample expressing the disease-associated TF signature, wherein the one or more cell types comprise B memory cells, B intermediate cells, B naive cells, CD14 monocytes, CD 16 monocytes, CD4 naive T cells, central memory CD4 T cells (CD4 TCM), CD8 Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT naive T cells, effector memory CD8 T cells (CD8 TEM), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright), and / or regulatory T cells (Treg).

[0018] In some embodiments, the identified cell type comprises B memory cells and the therapeutic agent targets B memory cells, B intermediate cells, or B naive cells and the therapeutic agent targets B naive cells, or CD14 monocytes and the therapeutic agent targets CD14 monocytes, or CD16 monocytes and the therapeutic agent targets CD 16 monocytes, or CD4 naive T cells and the therapeutic agent targets CD4 naive T cells, or central memory CD4 T cells (CD4 TCM) and the therapeutic agent targets CD4 TCM cells, or CD8 naive T cells and the therapeutic agent targets CD8 naive T cells, or effector memory CD8 T cells (CD8 TEM) and the therapeutic agent targets effector memory CD8 T cells, or natural killer cells (NK) and the therapeutic agent targets NK cells, or regulatory T cells (Treg) and the therapeutic agent targets regulatory Treg cells.

[0019] In some embodiments, the therapeutic agent comprises abatacept, rituximab, ocrelizumab, ofatumumab, epratuzumab. tabalumab. CAR-T cells, CAR- NK cells, tyrosine kinase inhibitors, TLR targeted therapy, memantine, anti-OX40 therapy, chemokine antagonists (e.g., CX3CR1 blockers), anti-NRPl, IDO inhibitors, calcineurin antagonists (e.g., sirolimus), TGF-beta blockade (including biologies and SMAD inhibitors), or PD1 agonists (e.g., peresolimab).

[0020] In some embodiments, the pro-infl ammalory downstream genes comprise MMP23B, XCL2, CCL4, IFNL1, PDGFD, IL12A, ADAMTS10, CCL3, CCL4L2, IL15, MMP24OS, IFNG, TGFB1, CCL5, MMP25-AS1, ADAMTS17, NOTCH1, TNFSF9, MMP25, NOTCH2NL, CCL20, ADAMTSL4, CXCL16, TNFSF8, TGFA, IL18BP, MMP19, TGFB3. XCL1. or ADAMTS1. In some embodiments, the pro- inflammatory downstream genes comprise MMP23B, TGFB1, IFNL1 , PDGFD, or CCL5.

[0021] In some embodiments, administering a therapeutic agent to modulate the expression of the pro-inflammatory downstream gene comprises administering an inhibitor directed to the pro-inflammatory downstream gene, wherein the pro- inflammatory downstream gene comprises one or more cytokines or cytokine signal transduction genes. In some embodiments, the therapeutic agent comprises one or more of a TNF inhibitor, an IL-1 inhibitor, an IL-6 inhibitor, a TGFP inhibitor (e.g., one or more SMAD inhibitors), one or more inhibitors of B cell signaling and Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT activation (e.g., BLyS or BTK), chemokine signal inhibitors (e.g., PI3Ky), rituximab, or a Janus kinase inhibitor.

[0022] In some embodiments, the therapeutic agent comprises vector-based gene therapy, small molecule activators, biologies, CRISPR-based gene editing, epigenetic modulation, TF modulation, or any combinations thereof.

[0023] In some embodiments, the macromolecules comprise RNA, DNA, or proteins. In some embodiments, the biological sample comprises blood, saliva, cerebrospinal fluid, bone marrow, bronchoalveolar lavage, sputum, or biological tissue.

[0024] In some embodiments, the subject is positive for anti-citrullinated protein autoantibodies (ACPA+). In some embodiments, the autoimmune condition comprises rheumatoid arthritis, systemic lupus erythematosus, scleroderma, multiple sclerosis, polymyositis, dermatomyositis, or Sjogren’s syndrome. In some embodiments, the autoimmune condition comprises rheumatoid arthritis.

[0025] In some embodiments, the one or more TFs comprise SP7, SOX9, NKX3-2, DLX6, HAND2, DLX5, HEY1, MSX2, ZNF521, TWIST1. TWIST2, SATB2. GLI3, AR, HEY2, HES 1, NKX2-5, GATA4, TEAD4, TEAD1, TEAD2, TEAD3, PGR, NR5A1, NR5A2, NR1I2, THRB, PPARG, HEYL, HES5, SOX2, SOX6, SOX7, TCF7L1, and / or SOX13.

[0026] In some embodiments, determining the TF signature comprises comparing the expression of the TFs of the biological sample with a reference level. In some embodiments, the reference level comprises the level of the one or more TFs from one or more healthy subjects. In some embodiments, the TF signature comprises the elevated expression of one or more TFs in two or more, three or more, or four or more of the SUMOylation. RUNX2. YAP1. NOTCH3, and / or WNT / [3-Catenin pathways. In some embodiments, the TF signature comprising the elevated expression of one or more TFs in each of the SUMOy lation, RUNX2, YAP1, NOTCH3, and WNT / p- Catenin pathways.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0028] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.

[0029] DESCRIPTION OF DRAWINGS

[0030] FIGs. 1A-1E show the study overview and co-embedding of multi-omics data. FIG. 1A is a schematic illustrating the study workflow. PBMC samples including 35 controls (CON), 26 ACPA positive (At-Risk) and 6 early RA (ERA) were utilized for scRNA-seq and scATAC-seq respectively. For each sample, matched data were coembedded into clusters. Cells in each cluster were aggregated in terms of gene count and open chromatin regions. Then each cluster was used as input of scTaiji to construct a regulatory network and generate the PageRank scores as output. The following unsupervised clustering revealed At-Risk / ERA signatures that were shared across multiple participants and cell types. FIG. IB shows UMAP shaded by major cell types in scRNA-seq cells (left) and scATAC-seq cells (right) respectively for one At-Risk sample. Clusters in both scRNA-seq and scATAC-seq were well separated bycell types. The selected sample represents the typical situation for all the 67 samples. Thirteen cell types include B memory cells, B intermediate cells, B naive cells, CD14 monocytes (CD 14 Mono), CD 16 monocytes (CD 16 Mono), CD4 naive T cells (CD4 T Naive), central memory- CD4 T cells (CD4 TCM). CD8 naive T cells (CD8 T Naive), effector memory CD8 T cells (CD8 TEM), mucosal-associated invariant T cells (MAIT cells), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright) and regulatory T cells (Treg). FIG. 1C shows UMAP shaded bymajor cell types (left) and assays (right) in cells from both scRNA-seq and scATAC- seq for the same sample in FIG. IB. The shading of the left plot is the same as FIG. IB. Clusters in co-embedding space were still separated by cell types while scRNA- seq and scATAC-seq cells were well aligned. FIG. ID shows percent of total cells across cell types. CD4 Naive and CD4 TCM were the most abundant cell type while B memory cells, CD 16 Mono, MAIT, and Treg cells were the relatively rare cell subsets. FIG. IE shows cell type distribution across 3 groups of PBMC samples. The shading is maintained throughout all figures. Centered Log-Ratio (CLR) Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT transformation before Kruskal-Wallis test, *p < 0.1, **p < 0.01. Most cell ty pes showed similar distribution across groups except for B intermediate, B memory’, and NK_CD56bright, which was modestly higher in At-Risk compared to that of two other groups.

[0031] FIGs. 2A-2D are unsupervised clustering showing distinct TF regulatory patterns. FIG. 2A shows PageRank scores heatmap of 5 Kmeans group-specific TFs across 1610 clusters. Top 10 TFs from each Kmeans group are selected as rows and shaded by their group specificity7. Shading for Kmeans groups is RColorBrewer palette Set2. The shading palette is maintained throughout all figures. Shading of the cell indicates the normalized PageRank scores. Each Kmeans group displayed distinct dynamic patterns of TF activity. Side table is the number of the specific TFs for each Kmeans group. G2 has the largest number of specific TFs. FIG. 2B shows cell type distribution across Kmeans groups. The separate top row represents the overall cell ty pe distribution across all the clusters. The bottom five rows are distributions for five Kmeans groups. Shading represents the percentage of clusters of each cell. G2 is a multi-lineage group with distribution similar to the overall distribution. Other 4 groups had predominant cell ty pes. FIG. 2C shows At-Risk / ERA vs CON ratio distribution across Kmeans groups. The first gray bar is the overall ratio adjusted to 1 while the other bars are 5 Kmeans groups. G2 is significantly enriched in At- Risk / ERA while G4 is enriched in CON. Gl, G3, and G5 show no enrichment. FIG. 2D Representative Reactome pathways enriched in each Kmeans group-specific TFs. The horizontal axis represents Kmeans groups and the vertical axis represents pathways. Circle size represents the number of TFs in the pathway and shading represents the adjusted p-values. Bold text represents signature pathways. G2 exhibits unique enrichment of several RA-related pathways e.g. SUMOylation of intracellular receptors (adjusted p-value < 1 e-5), Transcriptional regulation by' RUNX2 (adjusted p-value < le-5), etc.; Chi-squared test, ***p < 0.001, ****p < 0.0001.

[0032] FIGs. 3A-3E show At-Risk / ERA signature that is shared across multiple cell types. FIG. 3A shows heatmap of PageRank scores of all TFs across all clusters with columns ordered by cell types and Kmeans groups. Shading is the same as FIG. 2A. The signature TF group is marked by black box. Each cell type displayed high activity7in signature TFs. FIG. 3B shows G2 clusters per cell type of the total clusters per cell type in CON and At-Risk / ERA respectively. CD4 T Naive, CD4 TCM, CD8 T Naive, Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT and CD8 TEM are mostly enriched in At-Risk / ERA. MAIT cells with the signature TFs were only found in CON clusters. FIG. 3C shows mean PageRank scores of top 50 G2-specific TFs across cell types in G2 and other groups respectively. Rows represent TFs while columns represent cell types in G2 and other groups. Gray represents the average across other 4 groups. FIG. 3D shows representative enriched pathways of G2-specific TFs across cell types. Bold text represents signature pathways. All the cell types were enriched in signature pathways. FIG. 3E shows heatmap of At-Risk / ERA participants in G2 across cell types. The horizontal axis shows the individual participants and the vertical axis shows each cell type. Top bar represents the disease states of participants. Shading represents the number of clusters per cell type for each participant. All the At-Risk and ERA participants had the signature in at least one cell type but the combination and distribution of cell types are highly variable; Chi-squared test, *p < 0.1, **p < 0.05, ***p < 0.01.

[0033] FIGs. 4A-4F show distinct cell-cell communication patterns in At-Risk / ERA. FIG. 4A shows number of cellular interactions within signature clusters in two groups. Edge thickness is proportional to the number of interactions. Thicker edge indicates more interactions. Left and middle circular plots represent networks in At- Risk / ERA and control groups. Shading represents the cell type. Right circular plot represents the differential network between At-Risk / ERA and control. Rightmost panel shows the number of interactions in two groups. At-Risk / ERA group has significantly more interactions than control (p-value=0.06; Wilcoxon rank-sum test). FIG. 4B shows interaction strength within signature clusters in two groups. Edge thickness is proportional to the interaction strength. Thicker edge indicates stronger signals. Left and middle circular plots represent networks in At-Risk / ERA and control groups. Shading represents the cell type. Right circular plot represents the differential network between At-Risk / ERA and control. Rightmost panel shows the interaction strength in tw o groups. At-Risk / ERA group has significantly stronger interactions than control (p-value=0.04; Wilcoxon rank-sum test). FIG. 4C shows representative cellular communication netw orks within signature clusters in control and At-Risk patients. Shading represents the cell type and thickness of edge w eight is proportional to the interaction strength. Thicker edge line indicates stronger signal. Solid and open circles represent source and target respectively. Circle size is proportional to the number of clusters. Both the edge thickness and circle size were normalized and Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT comparable across different networks. At-Risk patient showed much denser and stronger interactions than control across almost all cell types. FIG. 4D shows increased ligand-receptor pairs in At-Risk / ERA group. The rank is based on the difference in total information flow between At-Risk / ERA and control groups. The total information flow7is calculated by summing the probability of all communications between the signature clusters. The left panel showed the relative information flow while the right panel showed the absolute information flow values. FIG. 4E shows representative IL 16 signaling networks within signature clusters in control and At- Risk patients. Each circle represents one Seurat cluster instance with cell ty pe label. Shading represents the cell ty pe and thickness of edge w eight is proportional to the interaction strength. Thicker edge line indicates stronger signal. Solid and open circles represent source and target respectively. Edge thickness was normalized and comparable across different networks. At-Risk patient showed much denser and stronger interactions than controls. FIG. 4F shows outgoing and incoming signaling strength of IL16 pathway across cell ty pes in control and At-Risk / ERA groups. The horizontal axis represents the cell types and vertical axis represents each individual, in which IL 16 signaling pathways is significant. Gradient shading represents the total outgoing signaling strength. Gradient shading represents the total incoming signaling strength. At-Risk / ERA has higher outgoing signals of TGF-P compared to control.

[0034] FIGs. 5A-5F show distinct cell-cell communication mediators in At- Risk / ERA. FIG. 5A shows top 30 predictors of classification model ranked by the average importance across 20 experiments. Exemplary top predictors include MMP23B, TGFB1, IFNL1. IL15, and CCL5. FIG. 5B shows gene expression level of MMP23B in each individual. At-Risk / ERA has significantly higher gene expression level. FIG. 5C shows gene expression level of TGFB1 in each individual. At- Risk / ERA has significantly higher gene expression level. FIG. 5D shows top signature regulators of TGFB1 in At-Risk / ERA signature clusters. Middle node is TGFB1. Other gray nodes are regulators of TGFB1. Gray node size and edge width are proportional to the mean regulatory strength predicted by Taiji. FIG. 5E shows normalized gene expression of top 30 predictors for At-Risk / ERA and control participants in G2 and G4 clusters respectively. For each gene, the maximum gene expression across clusters was taken within each Kmeans group and each individual. Rows represent mediators while columns represent patients. Top 30 predictor Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT cytokines are uniformly more active in At-Risk / ERAs compared to controls. Example genes include MMP23B, CCL4. IL12A, TNFSF14, IL15, NOTCH 1, CC / . and TGFB1. FIG. 5F shows protein expression level of six key mediators in each individual. At-Risk / ERA has significantly higher protein expression level than control. Wilcoxon rank-sum test, **p < 0.05, ***p < 0.01; ****p < 0.001.

[0035] FIGs. 6A-6D shows gene expression level of identified top mediators in established RA synovial tissues. FIG. 6A shows heatmap of normalized gene expression of top 30 mediators across 22 pseudo-bulk clusters. Rows represent genes while columns represent pseudo-bulk clusters. Both rows and columns are hierarchically clustered. Shading represents the average normalized expression across cells in the cluster, scaled for each gene across clusters. Column annotation legend represents cell types. Top mediators displayed gene expression across multiple cell types. FIG. 6B shows heatmap of normalized gene expression of top 30 mediators across synovial tissue samples. Rows represent genes while columns represent samples. Both rows and columns are hierarchically clustered. Shading represents the average normalized expression across cells in the sample, scaled for each gene across samples. Each sample has its own group of highly expressed genes. FIG. 6C show s heatmap of normalized gene expression of cell types across samples. Row s represent cell types while columns represent samples. Both rows and columns are hierarchically clustered. Shading represents the average normalized expression across 30 mediators, scaled for each cell type across samples. Each sample has its own combinations of dominant cell types expressing the top mediators. FIG. 6D show s proposed hypothesis to RA onset. Under the influence of risk factors such as genetics and environmental exposures, epigenetic remodeling took place in multiple cell types involving signature pathways like SUMOylation, RUNX2, YAP1, NOTCH3, and - Catenin Pathways. The signature TFs drive a characteristic set of pro-inflammatory genes in receiver cells that can, in turn, contribute to the onset and perpetuation of RA. Diverse cell types and pathogenic mechanisms can drive a common clinical phenotype known as RA and could explain the wide variation in clinical response to agents that target individual cytokines or cell types.

[0036] FIGs. 7A-7D show At-Risk / ERA signature is shared across multiple cell types. FIG. 7A shows heatmap of all participants in G2 across cell types. The horizontal axis shows the individual participants and the vertical axis shows each cell Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT ty pe. Top bar represents the disease states of participants. Shading represents the percent of clusters per total clusters per cell type for each participant. Almost all the At-Risk and ERA participants had the signature in at least one cell type but not all CON participants had the signature. MAIT and Treg cells showed higher enrichment in CON while other T cells showed higher enrichment in At-Risk / ERA. FIG. 7B shows G2-specific TFs whose regulatees are enriched in At-Risk / ERA signature pathways. The horizontal axis represents TFs. and the vertical axis represents signature pathways. FIG. 7C shows representative Reactome pathways enriched in each Kmeans group-specific regulatees. The horizontal axis represents Kmeans groups and the vertical axis represents pathways. Circle size represents the number of regulatees in the pathway and shading represents the adjusted p-values. FIG. 7D shows intersection of TFs enriched in 5 representative signature pathways. The horizontal axis is the signature pathway and the vertical axis is the set size. The side horizontal bars are the original size of each pathway.

[0037] FIG. 8 shows identified signature pathways with signature TFs and representative downstream genes.

[0038] DETAILED DESCRIPTION

[0039] The present disclosure provided methods for treating autoimmune conditions (e.g., rheumatoid arthritis (RA)) by identifying and targeting disease-associated transcription factor (TF) signatures in one or more cell types, one or more cell types that express the TF signatures, and / or inflammatory downstream genes of the TF signature pathways. Specifically, the present disclosure described methods for determining and modulating a unifying set of TFs, TF downstream pathways that regulate a pro-inflammatory cell communication network, and / or multiple cell types in the network that serve as pathogenic drivers in at-risk individuals or autoimmune conditions (e g., RA).

[0040] These cell-type-specific TF signature pathways explain the personalized pathogenesis of autoimmune conditions (e.g., RA) and contribute to the diversity of clinical responses to targeted therapies. Furthermore, these disclosed methods could provide opportunities for stratifying individuals at-risk for autoimmune conditions (e.g., RA), and selecting therapies tailored for prevention or treatment of autoimmune conditions (e g., RA). Overall, the present disclosure supports a new' paradigm to understand how a common clinical phenotype could arise from diverse pathogenic Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT mechanisms and to generate mechanism-based therapeutic approach targeting case specific TF signatures and cell types.

[0041] Autoimmune conditions

[0042] A person's genes in combination with infections and other environmental exposures likely play a significant role in disease development. In autoimmune disorders, the immune system fails to distinguish between foreign invaders and the body’s own healthy cells, leading to inflammation and damage to various parts of the body. There are more than 80 types of autoimmune disorders, including rheumatoid arthritis, systemic lupus ery thematosus, scleroderma, type 1 diabetes, multiple sclerosis. Hashimoto’s thyroiditis, Graves’ disease, Sjogren’s syndrome, inflammatory bowel disease, spondyloarthritis, celiac disease, myasthenia gravis, polymyositis, dermatomyositis, or Guillain-Barre syndrome. In some embodiments, the autoimmune condition is rheumatoid arthritis.

[0043] Rheumatoid arthritis (RA) is a chronic, systemic immune-mediated disease marked by synovial inflammation, leading to swelling, pain, and joint destruction. RA can happen in most joints, but it’s most common in the small joints of the hands, wrists, and feet. Early signs and symptoms include pain, stiffness, tenderness, swelling or redness in one or more joints, usually in a symmetrical pattern (e.g., both hands and both feet). The symptoms can worsen over time and spread to more joints including the knees, elbows, or shoulders. RA can make it hard to perform daily activities like writing, holding objects with the hands, walking, and climbing stairs. People with RA often feel fatigue and general malaise (e.g., fever, poor sleep quality, loss of appetite) and may experience depressive symptoms.

[0044] The etiology of RA, as well as the timing and anatomic site at which RA- related autoimmunity is initiated, is complex. A complex interplay of inflammatory cells and cytokines contribute to joint inflammation, tissue damage, and the chronic nature of the disease. A model outlining the sequential immune processes in disease progression has emerged, including mucosal initiation, propagation of systemic inflammation and autoimmunity (preclinical RA / at-risk status), and the onset of clinically apparent arthritis (e.g., early clinical RA or clinical RA). Initially, mucosal inflammation and dysbiosis may trigger local autoantibody production, which under normal circumstances would be transient. However, this is followed by a systemic Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT spread of autoimmunity', evidenced by elevated serum autoantibodies, eventually leading to the development of clinically recognizable joint inflammation.

[0045] The majority (-70-80%) of clinical RA is termed ‘seropositive’ because patients exhibit blood elevations of autoantibodies. In some embodiments, the serological tests for autoantibodies include anti-citrullinated protein antibodies (ACPAs), rheumatoid factor (RF), antibodies to modified protein antigens (AMP A), and any other autoantibodies known in art. The elevated levels of ACPAs are strongly associated with the future development of RA in up to 60% of at-risk individuals. The autoantibody elevations can be on average 3-5 years prior to the onset of clinical RA. This prolonged period of autoimmunity and inflammation prior to the onset of arthritis can be designated as preclinical RA in subjects who eventually progress to a clinical diagnosis of RA. Preclinical RA carries an at-risk status of future conversion to clinical RA.

[0046] Systemic inflammation and autoimmunity in RA begin long before the onset of detectable joint inflammation. Interventions prior to the onset of joint inflammation are very likely necessary to modify the course of the disease. Unfortunately, a major limitation to developing effective preventive strategies for RA has been a lack of ability' to detect and classify individuals that are at high risk for RA, as well as limitations in the knowledge of the mechanisms of disease development, including identification of when and at what anatomic site RA begins.

[0047] Currently, no treatments are available to prevent progression to RA in these at- risk individuals. In addition, diverse pathogenic mechanisms underlying a common clinical phenoty pe in RA complicate therapy as no single agent is universally effective. The divergent pathogenic pathways are poorly understood, and there is need to develop reliable tests to predict benefit of targeted therapeutics for individual patients.

[0048] The methods of the disclosure show how to identify and treat a person having an autoimmune condition (e.g., rheumatoid arthritis, systemic lupus erythematosus, scleroderma, type 1 diabetes, multiple sclerosis. Hashimoto’s thyroiditis. Graves’ disease, Sjogren’s syndrome, inflammatory bowel disease, spondyloarthritis, celiac disease, myasthenia gravis, polymyositis, dermatomyositis, or Guillain-Barre syndrome), w herein the etiology of the autoimmune condition is not limited to a single pathogenic process, but rather can involve the dysregulation of one or more Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT transcription factor (TF) pathways in one or more different cell types. It is important to identify and monitor persons at risk of developing autoimmune conditions in order to provide early interventions and delay / shape the course of the disease.

[0049] Method of diagnosis

[0050] Currently, a diagnosis of RA is made when clinically apparent arthritis is present. Typically, the diagnosis is determined by a health-care provider based on the combination of signs and symptoms at typical local joints, serological tests, inflammatory markers (e.g., erythrocyte sedimentation rate (ESR) and C-reactive protein (CRP)), and imaging of joint changes. Furthermore, classification criteria e g., 1987 ACR or 2010 ACR / EULAR classification criteria have been used to classify RA based on points assigned to the number and size of affected joints, serology (RF and ACPA), acute-phase reactants (ESR and CRP), and duration of symptoms. A score of 6 or more out of 10 suggests RA for the latter.

[0051] The present disclosure provides methods for characterizing the immune signatures associated with preclinical RA / at-risk status or clinical RA. Diagnosing preclinical RA / at-risk status or clinical RA can promote the strategies aimed at preventing or delaying the development of RA, thereby improving survival and prognosis outcomes for individuals at high risk of developing RA.

[0052] Provided herein are methods for diagnosis of an autoimmune condition (e.g., RA) in a subject, the method including (a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated TF signature in one or more cell types from the biological sample; and (c) identifying the one or more cell types from the biological sample expressing the TF signature.

[0053] Also provided herein are methods for diagnosis of an autoimmune condition (e.g., RA) in a subject, the method including (a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated TF in one or more cell types from the biological sample; and (c) determining the expression of one or more proinflammatory downstream genes regulated by the TF signature.

[0054] In some embodiments, the subject is positive for anti-citrullinated protein autoantibodies (ACPA+). In some embodiments, the biological sample comprises Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT blood, saliva, cerebrospinal fluid, bone marrow, bronchoalveolar lavage, sputum, or biological tissue. In some embodiments, the macromolecules comprise RNA, DNA, or proteins.

[0055] The analysis of the macromolecules includes any method know n in art that can be used to detect and analyze RNA, DNA, and / or protein. For example, RNA can be analyzed using transcriptomics (e.g.. RNA-seq or single-cell RNA-seq), RT-PCR, northern blotting, in situ hybridization, or microarray. DNA can be analyzed using genomics (e g., DNA sequencing or single-cell ATAC-seq), PCR, gel electrophoresis, southern blotting, or DNA microarray. Protein can be analyzed using proteomics (e.g., mass spectrometry), western blotting, ELISA, protein sequencing, protein crystallography, or flow cytometry. In some embodiments, the analysis of the macromolecules can be any combination of the methods described herein. In some embodiments, the analysis of the macromolecules can be performed utilizing Taiji pipeline.

[0056] As used herein. Taiji pipeline is a software providing an integrative multi - omics data analysis framework invented at UC San Diego. Taiji pipeline can be used as a standalone pipeline to analyze ATAC-seq, RNA-seq, single cell ATAC-seq, or Drop-seq data. The power of Taiji is in its ability to integrate diverse datasets to construct a comprehensive regulatory net ork and identify candidate driver genes. Other data integration techniques known in the art may be used to identify and analyze the macromolecules of a sample.

[0057] Method of treatment

[0058] Provided herein are methods for treatment of an autoimmune condition (e.g., RA) in a subject, the method comprising (a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample; (c) identifying the one or more cell types from the biological sample expressing the disease-associated TF signature; and (d) administering to the subject a therapeutic agent comprising a therapeutic agent targeting the one or more cell types expressing the disease-associated TF signature.

[0059] Also provided herein are methods for treatment of an autoimmune condition (e.g., RA) in a subject, the method comprising (a) isolating macromolecules or having Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT macromolecules isolated from a biological sample obtained from the subject; (b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample; and (c) administering a therapeutic agent to modulate the expression of one or more pro-infl ammatory downstream genes of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways.

[0060] As used in this context, to "‘treat” means to ameliorate at least one symptom or prevent / delay the progression of the disorder associated with an autoimmune condition (e.g., RA). Thus, a treatment can result in preventing or delaying the progression of preclinical autoimmune conditions (e.g., preclinical RA) to clinical autoimmune conditions (e.g., clinical RA). A treatment can also result in preventing or delaying the further progression of early autoimmune symptoms. Generally, the methods include administering a therapeutic agent or any combination of therapeutic agents to a subject who is in need of such treatment.

[0061] A subject can be an individual (e.g., a human) having or suspected of having an autoimmune condition (e.g., RA). In some embodiments, the subject has a preclinical autoimmune condition (e.g., preclinical RA). In some embodiments, the subject has an at-risk status of future conversion to a clinical autoimmune condition (e.g., clinical RA). In some embodiments, the subject has an early stage of an autoimmune condition (e.g., early RA). In some embodiments, the subject is positive for anti-citrullinated protein autoantibodies (ACPA+).

[0062] Transcription factor signature

[0063] Transcription factors (TFs) are proteins that play a critical role in regulating gene expression. By binding to specific DNA sequences near the genes they control, TFs can either promote or inhibit the transcription of the gene. TFs thus play key roles in cell differentiation, cellular response to environmental signals, cell cycle, and cell growth.

[0064] As described herein, disease-associated TF signature for autoimmune conditions (e.g., RA) refers to the elevated expression of one or more TFs in one or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / 0-Catenin pathways. In some embodiments, a disease-associated TF signature comprises elevated expression of TFs in two or more, three or more, four or more, or all five of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways. Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0065] Disease-associated TF signature is the elevated expression of the one or more TFs in one or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways as compared with a reference level. The reference level is the level of the one or more TFs from one or more healthy subjects. In some embodiments, the disease-associated TF signature is a predictor for a preclinical / at risk status (e.g., preclinical RA) or an early autoimmune condition (e.g., early RA), or the progression of the autoimmune disease (e.g.. RA).

[0066] The elevated TFs in the one or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways include, but are not limited to, SP7, SOX9, NKX3-2, DLX6, HAND2, DLX5, HEY1, MSX2, ZNF521, TWIST1, TWIST2. SATB2, GLI3. AR, HEY2, HES 1, NKX2-5, GATA4, TEAD4. TEAD1, TEAD2, TEAD3, PGR, NR5A1, NR5A2, NR1I2, THRB, PPARG, HEYL, HES5, SOX2, SOX6, SOX7, TCF7L1, and SOX13.

[0067] The disease-associated TF signature described herein can occur in multiple cell types. For example, the cell types expressing the disease-associated TF signature can include, but are not limited to, B memory cells, B naive cells, CD 14 monocytes, CD 16 monocytes, CD4 naive T cells, central memory CD4 T cells (CD4 TCM), CD8 naive T cells, effector memory' CD8 T cells (CD8 TEM), natural killer cells (NK), and / or regulatory T cells (Treg). In some embodiments, the cell types expressing the TF signature include CD4 T Naive, CD4 TCM, and / or CD8 T Naive cells.

[0068] There are multiple cell types that can display the TF signature, and the pattern of which cell ty pe or cell t pes with the TF signature is highly variable. Each patient can have their own combination of cell types or TF signature components, thereby contributing to the need for individualized therapies. In some embodiments, different cell types express the disease-associated TF signature in a subject. In some embodiments, different cell types express different TF signatures in a subject. In some embodiments, some cell types express a disease-associated TF signature and other cell types do not express a disease-associated TF signature in a subject. In some embodiments, the disease-associated TF signature of a subject is different than the disease-associated TF signature of a different subject, either because a similar TF pattern is observed in a different cell type between the two subjects, or because each subject has a different TF dysregulation (e.g., different disease-associated TF pattern) in the same cell type. A specific disease-associated TF signature can be unique to a Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT subject either by the particular combination of TFs implicated in the disease- associated TF signature, and / or by the particular cell type harboring the disease- associated TF signature.

[0069] The disease-associated TF signature in turn regulates pro-inflammatory dow nstream genes which contribute to the development or progression of an autoimmune condition (e.g., RA). In some embodiments, the pro-inflammatory downstream genes are a common set of genes shared among subjects having an at-risk status or an early autoimmune condition (e.g., early RA). In some embodiments, the common set of pro-inflammatory dow nstream genes described herein is produced by different cell types among different subjects. As described herein, the pro- inflammatory downstream genes include, but are not limited to. MMP23B, XCL2, CCL4, IFNL1, PDGFD, IL12A, ADAMTS10, CCL3, CCL4L2, IL15, MMP24OS, IFNG, TGFB1, CCL5, MMP25-AS1, ADAMTS17, NOTCH1, TNFSF9, MMP25, NOTCH2NL, CCL20, ADAMTSL4, CXCL16, TNFSF8, TGFA, IL18BP, MMP19, TGFB3, XCL1, and ADAMTS1. In some embodiments, the pro-inflammatory downstream genes include MMP23B, TGFB1, IFNL1, PDGFD, or CCL5.

[0070] As used herein, a “regulatee’’ is a gene that is directly regulated by the TF (e.g., IL-6 is a regulatee of NF-kB). A downstream gene (e.g., CRP, a gene regulated by IL-6) can be any gene that is in the pathway regulated by a regulatee.

[0071] The methods used for determining the disease-associated TF signature described herein can be any method known in art that is utilized to analyze RNA, DNA or protein. For example, the methods include single-cell RNA-seq, single-cell ATAC-seq. Taiji pipeline (a platform to analyze the integration of single-cell RNA- seq and single-cell ATAC-seq), proteomics (e.g., mass spectrometry), microarrays. PCR, RT-PCR, northern blotting, southern blotting, w estern blotting, or any combination thereof.

[0072] Determining the disease-associated TF signature can include comparing the expression of the TFs of the biological sample with a reference level and determining whether one or more TFs are elevated compared to the reference level. The reference level is the level of the one or more TFs from one or more healthy subjects.

[0073] Identifying cell types

[0074] In some embodiments of the diagnosis and / or treatment methods described herein, the subject is positive for one or more cell types expressing a disease- Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT associated TF signature. The cell ty pes are isolated or identified in the subject’s biological sample (e.g., blood, saliva, cerebrospinal fluid, bone marrow, bronchoalveolar lavage, sputum, or biological tissue). The methods used for identifying cell types can be any method known in art, including single-cell RNA-seq, single-cell ATAC-seq, Taiji pipeline (a platform to analyze the integration of singlecell RNA-seq and single-cell ATAC-seq), fluorescence-activated cell sorting, proteomics (e.g.. mass spectrometry), microarrays, or any combination thereof.

[0075] In some embodiments, the one or more cell types expressing a disease- associated TF signature include B memory' cells, B intermediate cells, B naive cells, CD 14 monocytes, CD 16 monocytes, CD4 naive T cells, central memory CD4 T cells (CD4 TCM). CD8 naive T cells, effector memory CD8 T cells (CD8 TEM), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright), and / or regulatory T cells (Treg). In some embodiments, the cell ty pes expressing the TF signature includes CD4 T Naive, CD4 TCM, and / or CD8 T Naive cells.

[0076] Targeting cell types expressing the transcription factor signature

[0077] The disclosed methods for treatment of an autoimmune condition (e.g., RA) can include administering to the subject a therapeutic agent targeting one or more cell types expressing the disease-associated TF signature. For example, the methods may include identifying if the disease-associated TF signature is associated with a particular cell ty pe, and if associated with a particular cell ty pe, administering a therapeutic agent known to target that particular cell type.

[0078] Therefore, in some embodiments of the methods, the identified cell type includes B memory' cells and the therapeutic agent targets B memory cells, or the identified cell type includes B naive cells and the therapeutic agent targets B naive cells, or the identified cell type includes CD 14 monocytes and the therapeutic agent targets CD 14 monocytes, or the identified cell type includes CD 16 monocytes and the therapeutic agent targets CD 16 monocytes, or the identified cell type includes CD4 naive T cells and the therapeutic agent targets CD4 naive T cells, or the identified cell ty pe includes central memory' CD4 T cells (CD4 TCM) and the therapeutic agent targets CD4 TCM cells, or the identified cell type includes CD8 naive T cells and the therapeutic agent targets CD8 naive T cells, or the identified cell type includes effector memory CD8 T cells (CD8 TEM) and the therapeutic agent targets effector memory' CD8 T cells, or the identified cell ty pe includes natural killer cells (NK) and Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT the therapeutic agent targets NK cells, or the identified cell type includes regulatory' T cells (Treg) and the therapeutic agent targets regulatory Treg cells.

[0079] Examples of such therapeutic agents that target a particular cell type include, but are not limited to, abatacept, rituximab, ocrelizumab, ofatumumab, epratuzumab, tabalumab, CAR-T cells, CAR-NK cells, ty rosine kinase inhibitors, TLR targeted therapy, memantine, anti-OX40 therapy, chemokine antagonists (e.g., CX3CR1 blockers), anti-NRPl, IDO inhibitors, calcineurin antagonists (e.g.. sirolimus), TGF- beta blockade (including biologies and SMAD inhibitors), or PD1 agonists (e.g., peresolimab). In some embodiments, the methods for treatment of an autoimmune condition (e.g., RA) include any combination of the therapeutic agents disclosed herein. In some embodiments, the methods for treatment of an autoimmune condition (e.g., RA) include any therapeutic agent known in art that targets the identified cell ty pes.

[0080] As used herein, a therapeutically effective amount of the therapeutic agent is administered to a subject. In some embodiments, the dosage of the therapeutic agent administered to the subject is based on the subject’s weight. In some embodiments, the dosage of the therapeutic agent administered to the subject is independent of the subject’s weight. In some embodiments, the therapeutic agent is administered every' day. In some embodiments, the therapeutic agent is administered less than every day (e.g., every two days, every three days, twice a week, every week, every two weeks, every month, or every year). In some embodiments, the therapeutic agent is administered intravenously, intraarterially, intramuscularly, intradermally, subcutaneously, or intraperitoneally.

[0081] Identifying pro-inflammatory downstream genes

[0082] In some embodiments of the diagnosis and / or treatment methods described herein, the subject is positive for one or more pro-inflammatory downstream genes that are regulated by the disease-associated TF signature pathways disclosed herein. The methods used for determining pro-inflammatory downstream genes can be any method known in art that is utilized to analyze RNA, DNA or protein. For example, the methods include single-cell RNA-seq, single-cell ATAC-seq, Taiji pipeline (a platform to analyze the integration of single-cell RNA-seq and single-cell ATAC- seq), proteomics (e.g., mass spectrometry), microarrays, PCR, RT-PCR, northern blotting, southern blotting, w estern blotting, or any combination thereof. Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0083] In some embodiments, the pro-inflammatory downstream genes include, but are not limited to, MMP23B, XCL2. CCL4, IFNL1. PDGFD, IL12A. ADAMTS10, CCL3, CCL4L2, IL15, MMP24OS, IFNG, TGFB1, CCL5, MMP25-ASL ADAMTS17, NOTCH1, TNFSF9, MMP25, NOTCH2NL, CCL20, ADAMTSL4, CXCL16, TNFSF8, TGFA, IL18BP, MMP19, TGFB3, XCL1, or ADAMTS1. In some embodiments, the pro-inflammatory downstream genes comprise MMP23B, TGFB1, IFNL1, PDGFD, or CCL5. In some embodiments, the pro-inflammatory downstream genes are predictors for progression of a preclinical / at risk status (e.g., preclinical RA) or an early autoimmune condition (e.g., early RA).

[0084] Targeting pro-inflammatory downstream genes

[0085] The disclosed methods for treatment of an autoimmune condition (e.g., RA) include administering to the subject a therapeutic agent targeting one or more pro- inflammatory downstream genes regulated by the disease-associated TF signature pathways. The pro-inflammatory downstream genes and disease-associated TF signature pathways are described herein. In some embodiments, administering a therapeutic agent to modulate the expression of the pro-inflammatory downstream gene includes administering an inhibitor directed to the pro-inflammatory downstream gene, wherein the pro-inflammatory downstream gene includes one or more cytokines or cytokine signal transduction genes.

[0086] As used herein, the therapeutic agent modulating the expression of the pro- inflammatory downstream gene can be a vector-based gene therapy, a small molecule activator, a biologic. CRISPR-based gene editing, epigenetic modulation, transcription factor modulation, or any combination thereof.

[0087] In some embodiments, the therapeutic agent modulating the expression of the pro-inflammatory' downstream gene can be one or more of a TNF inhibitor, an IL-1 inhibitor, an IL-6 inhibitor, a TGFP inhibitor, one or more SMAD inhibitors, one or more inhibitors of B cell signaling and activation (e g., BLyS or BTK), chemokine signal inhibitors (e.g., PI3K), rituximab, or a Janus kinase inhibitor. In some embodiments, the methods for treatment of an autoimmune condition (e.g., RA) include any combinations of the therapeutic agents disclosed herein. In some embodiments, the methods for treatment of an autoimmune condition (e.g., RA) include any therapeutic agent known in art that targets the one or more pro- Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT inflammatory downstream genes regulated by the disease-associated TF signature pathways.

[0088] As used herein, a therapeutically effective amount of the therapeutic agent modulating the expression of the pro-inflammatory downstream gene is administered to a subject. In some embodiments, the dosage of the therapeutic agent administered to the subject is based on the subject’s weight. In some embodiments, the dosage of the therapeutic agent administered to the subject is independent of the subject’s weight. In some embodiments, the therapeutic agent is administered every day. In some embodiments, the therapeutic agent is administered less than every day (e.g., every two days, every' three days, twice a week, every week, every two weeks, every month, or every year). In some embodiments, the therapeutic agent is administered intravenously, intraarterially, intramuscularly, intradermally, subcutaneously, or intraperitoneally.

[0089] Kit for autoimmune condition testing

[0090] Provided herein are kits for use in determining whether a subject is at risk of developing and autoimmune condition (e.g., RA). The kit includes instructions and / or materials for (a) obtaining or having obtained a blood sample from the subject; (b) isolating mononuclear cells from the sample; (c) purifying RNA from the isolated mononuclear cells; (d) determining a level of one or more transcription factors (TFs) or one or more pro-inflammatory' genes, wherein the TFs regulate or interact with SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / 0-Catenin pathways; and (e) comparing the level of the one or more TFs or the one or more pro-inflammatory genes with a reference level, wherein the level of the one or more (TFs) or the one or more pro-inflammatory genes determined in step (d) that is higher than the reference level indicates that the subject is at risk of developing RA.

[0091] Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0092] EXAMPLES

[0093] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

[0094] Materials and Methods

[0095] Clinical cohorts

[0096] Three groups of participants were recruited for this study. The first cohort (At- Risk) included individuals who were at-risk for future clinical RA as indicated by serum ACPA positivity >2x the upper limit of normal2 using the assay anti-cyclic citrullinated peptide-3 anti-CCP3, IgG ELISA (Werfen, San Diego, CA USA). The second cohort (ERA) was comprised of patients who were anti-CCP3 positive and had early RA meeting the 2010 American College of Rheumatology / European Alliance of Associations for Rheumatology (ACR / EULAR) classification criteria for RA and were diagnosed.

[0097] Sample preparation

[0098] Blood was drawn into BD NaHeparin vacutainer tubes (for PBMC; BD #367874) or K2-EDTA vacutainer tubes (for plasma; BD #367863). PBMC isolation and plasma processing were started within 2 hours post draw. For PBMC isolation, the samples in NaEIeparin tubes for each donor were pooled into one common pool and combined with an equivalent volume of room temperature PBS (ThermoFisher #14190235). PBMCs were isolated using Leucosep tubes (Greiner Bio-One #227290) with 15 ml of Ficoll Premium (GE Healthcare #17-5442-03). After centrifugation, the PBMCs were recovered and resuspended with 15 ml cold PBS+0.2% BSA (Sigma #A9576; ‘PBS+BSA”). The cells were pelleted, resuspended in 1 ml cold PBS+BSA per 15 ml whole blood processed and counted with a Cellometer Spectrum (Nexcelom) using Acridine Orange / Propidium Iodide solution. PBMCs were cryopreserved in 90% FBS (ThermoFisher #10438026) / 10% DMSO (Fisher Scientific #D 12345) at a target of 5 x 106cells / ml by slow freezing in a Coolcell LX (VWR #75779-720) overnight in a -80°C freezer followed by transfer to liquid nitrogen. For genomics assays PBMCs were removed from liquid nitrogen storage and immediately thawed in a 37°C water bath. Cells were diluted dropwise into 40 mL AIM V media (Thermo Fisher Scientific #12055091) pre-warmed to 37°C. Cells Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT were pelleted at 400 x g, resuspended in 5 mL cold AIM V media, and recounted using a Cellometer Spectrum. 30 mL cold AIM V media was added to the cells, which were re-pelleted and resuspended to appropriate concentration for the assays. scRNA-seq scRNA-seq was performed on PBMCs as previously described (P. C. Genge, STAR Protoc 2, 100900 (2021)). In brief, scRNA-seq libraries were generated using a modified 10x genomics chromium 3' single cell gene expression assay with Cell Hashing. Sample libraries were constructed across different batches, with the addition of a common control donor leukopak sample in each library as batch control. Libraries were sequenced on the Illumina Nov aseq platform. Hashed 10x Genomics scRNA-seq data processing was earned out using BarWare to generate samplespecific output files. scATAC-seq

[0099] To remove dead cells, debris, and neutrophils prior to scATAC-seq, PBMC samples were sorted by fluorescence-activated cell sorting (FACS) following established protocols. Cells were incubated with Fixable Viability Stain 510 (BD, 564406) for 15 minutes at room temperature and washed with AIM V medium (Gibco. 12055091) before incubating with TruStain FcX (BioLegend, 422302) for 5 minutes on ice, followed by staining with mouse anti-human CD45 FITC (BioLegend, 304038) and mouse anti-human CD15 PE (BD, 562371) antibodies for 20 minutes on ice. After washing, cells were then sorted on a BD FACSAria Fusion with a standard viable CD45+ cell gating scheme. Neutrophils were then excluded in the final sort gate. An aliquot of each post-sort population was used to collect 50,000 events to assess post-sort purity.

[0100] Sample processing

[0101] Permeabilized-cell scATAC-seq was performed as described previously. A 5% w / v digitonin stock was prepared stored at -20°C. To permeabilize, 1 x 106cells were centrifuged and resuspended in cold isotonic Permeabilization Buffer. Then they were diluted with 1 mL of isotonic Wash Buffer and centrifuged, and the supernatant was slowly removed. Cells were resuspended in chilled TD1 buffer (Illumina, Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0102] 15027866) to a target concentration of 2,300- 10,000 cells per pL. Cells were filtered through 35 pm Falcon Cell Strainers (Coming, 352235) before counting on a Cellometer Spectrum Cell Counter (Nexcelom) using ViaStain acridine orange / propidium iodide solution (Nexcelom, C52-0106-5).

[0103] Sequencing library preparation scATAC-seq libraries were prepared following established protocol. In brief. 15,000 cells were combined with TD1 buffer (Illumina, 15027866) and Illumina TDE1 Tn5 transposase (Illumina, 15027916) and incubated at 37°C for 60 minutes. A Chromium NextGEM Chip H (10* Genomics, 2000180) was loaded and a master mix was then added to each sample well. Chromium Single Cell AT AC Gel Beads vl. l (10x Genomics, 2000210) were loaded into the chip, along with Partitioning Oil. The chip was loaded into a Chromium Single Cell Controller instrument (10x Genomics, 120270) for GEM generation. After the run, GEMs were collected and linear amplification was performed on a Cl 000 Touch thermal cycler.

[0104] GEMs were separated into a biphasic mixture with Recovery Agent (10x Genomics, 220016), and the aqueous phase was retained and removed of barcoding reagents using Dynabead MyOne SILANE and SPRIselect reagent bead clean-ups. Sequencing libraries were constructed as described in the 10x scATAC User Guide. Amplification was performed in a C1000 Touch thermal cycler. Final libraries were prepared using a dual-sided SPRIselect size-selection cleanup.

[0105] Quantification and sequencing

[0106] Final libraries were quantified using a Quant-iT PicoGreen dsDNA Assay Kit (Thermo Fisher Scientific, P7589) on a SpectraMax iD3 (Molecular Devices). Library quality and average fragment size were assessed using a Bioanalyzer (Agilent, G2939A) High Sensitivity DNA chip (Agilent, 5067-4626). Libraries were sequenced on the Illumina NovaSeq platform with the following read lengths: 51nt read 1, 8nt i7 index, 16nt i5 index. 51 nt read 2. Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0107] Plasma proteomics

[0108] Plasma samples were run on the Olink Explore 1536 platform. Analytes from the inflammation, oncology, cardiometabolic, and neurology panels were measured. Samples were randomized across plates to achieve a balanced distribution of age and sex. Resulting data were first normalized to an extension control that was included in each sample well. Plates were then standardized by normalizing to inter-plate controls run in triplicate on each plate. Data were then intensity normalized across all samples. Final normalized relative protein quantities were reported as log2 normalized protein expression (NPX) values by Olink. Three protein analytes were repeated across each of the four panels and treated as distinct measurements: TNF, IL-6, and CXCL8. Data, including QC flags, were reviewed for overall quality prior to analysis. Samples were measured across multiple batches.

[0109] To facilitate comparisons between batches, plasma from 12 donors was obtained commercially (BioIVT; Bloodworks Northwest) and randomly interspersed among the above study samples. Samples measured in later batches were bridge normalized to the earliest batch. Bridge offsets were determined for each batch and each analyte separately by taking the median of the persample NPX differences between the later batch result and the earliest (reference) batch result for the 12 commercial samples. Offsets were then subtracted from the analyte measurements of all samples in the later batch to obtain the normalized NPX values.

[0110] Dataset integration

[0111] Paired scRNA-seq and scATAC-seq datasets from each participant were obtained from 26 At-Risk individuals with elevated anti-citrullinated protein antibody (ACPA), 6 seropositive ERA patients and 35 controls (CON).

[0112] 70x scRNA-seq data. scRNA-seq data were aligned using 10x cellranger v3.1.0 and 10x transcriptome vGRCh38-3.0.0. Hashtag Ohgo sequences were processed using CITE- Seq Count vl.4.3, and cells were assigned to sample-linked hashes, split by sample for each well, and merged across wells per sample using an AIFI pipeline. Cells were labeled using Seurat v4 labeling pipeline with default parameters. The reference was Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT customized based on the recently described CITE-seq reference of 162,000 PBMC measured with 228 antibodies.

[0113] 7 Ox scATAC-seq data.

[0114] In the scATAC-seq pipeline, CellRanger alignment was implemented, followed by a rigorous quality control process. Cells with unique fragments between 1000 and 100,000. fragment size between 10 and 2000, >50% of fragments in Altius, >20% of fragments in transcription starting site (TSS), >4 TSS enrichment score were retained. This ensures that cells from the scATAC-seq pipeline are high quality, reduces the number of doublets, and are available in a variety of formats for downstream analysis (.arrow, fragments. tsv.gz, and ,h5- formatted count matrices). scATAC-seq data were aligned using lOx cellranger-atac v 1.1.0, using reference vGRCh38-1.1.0. After alignment, data were processed through a custom QC and counting pipeline to generate a matrix of unique fragment counts in each peak. ArchR vl.0.2 was used to generate Arrow files, doublet filtering (filterRatio=0.5), dimensionality reduction with iterative latent semantic indexing (LSI) (iterations=4), and clustering (resolution=3).

[0115] Integration of scRNA-seq and scATAC-seq data. scATAC-seq data were integrated with the corresponding scRNA-seq using the '‘addGenelntegrationMatrix” function in ArchR with default parameters. After alignment, each cell in the scATAC-seq space was assigned a gene expression signature from the cell in the scRNA-seq that is the most similar. Cells from both scRNA-seq and scATAC-seq were clustered in the same co-embedding space.

[0116] TF regulatory networks construction based on Taiji

[0117] Single cells within the same cluster were treated as one ‘‘pseudo-bulk'’ sample with the annotation as the cell type occurring most frequently in the cluster. The gene counts of scRNAseq were added up and the fragments of scATAC-seq were combined to generate the RNA-seq input and ATAC-seq input for the pseudo-bulk samples respectively. Only pseudo-bulk samples with >2000 open chromatin peaks, >20 scATAC-seq cells and >20 scRNA-seq cells were kept on account of reliability of constructed regulatory networks. Additionally, to link promoters and enhancers, the Attorney Docket No. 15670-0431W01 / UCSD Ref.: SD2025-048-2PCT promoter-enhancer contacts predicted by Epitensor v0.9 was used. Taiji vl.1.0 with default parameters was used for the integrative analysis of RNA-seq and ATAC-seq data. The motif file was downloaded directly from the CIS-BP database containing 1047 human motifs.

[0118] Taiji pipeline overview

[0119] To characterize TF activity in each pseudo-bulk cluster, an integrated multi- omics analysis was performed using the Taiji pipeline. Taiji integrates gene expression and epigenetic modification data to build gene regulatory networks. The algorithm first predicts putative TF binding sites in each open chromatin region that mark active promoters and enhancers using motifs documented in the CIS-BP database. These TFs are then linked to their target genes predicted by EpiTensor. The regulatory interactions are assembled into a genetic network. Finally, the personalized PageRank algorithm is used to assess the global influences of the TFs. In the network, the node weights are determined by the z scores of gene expression levels, allocating higher ranks to the TFs that regulate more differentially expressed genes. Each edge weight is set to be proportional to the TF’s expression level, its binding site’s open chromatin peak intensity7, and the motif binding affinity7, thus representing the regulatory strength. Using this method, Taiji has more power than other methods that identify key regulators in individual transcriptome and chromatin accessibility and has been confirmed using simulated data, literature evidence and experimental validation in numerous studies of various biological problems. For this dataset, the average number of nodes and edges of the networks were 17,046 and 3,002,662, respectively, including 1047 (6.14%) TF nodes. On average, each TF regulates 3417 genes, and each gene is regulated by 184 TFs.

[0120] TF regulatory networks weighting scheme

[0121] As described in the original Taiji paper, a personalized PageRank algorithm was applied to calculate the ranking scores for TFs. The edge weights and node weights in the network were first initialized. The node weight was calculated as ez', where7iis the gene’s relative expression level in cell typeii, which is computed by applying the zz score transformation to its absolute expression levels. The edge weight was determined by ", where p is the peak intensity, Attorney Docket No. 15670-0431W01 / UCSD Ref.: SD2025-048-2PCT calculated asJ , where x is -logio(p), represented by the p-value of the ATAC- seq peak at the predicted TF binding site, rescaled to [0, 1] by a sigmoid function; m is the motif binding affinity, represented by the p-value of the motif binding score, rescaled to [0, 1] by a sigmoid function; g is the TF expression value; n is the number of binding sites linked to gene j. Let s be the vector containing node weights and W be the edge weight matrix. The personalized PageRank score vector v was calculated by solving a system of linear equations v = (1 - d)s + dWv, where d is the damping factor (default to 0.85). The above equation can be solved in an iterative fashion, i.e.. setting vt+i = (1 - d)s + dWvt.

[0122] If the TFs in the same protein family share the same motifs, their PageRank scores are distinguished by their own expression levels because their motifs and the target genes are the same. If a motif is weak, the PageRank score of the TF is decided by whether these motifs occur in the open chromatin regions (measured by the peak intensity of the ATAC-seq data), the TF expression and its target expression levels. The relative difference between the PageRank scores of TFs also helps to uncover important TFs with weak motifs.

[0123] Unsupervised clustering analysis

[0124] To identify the groups of samples showing similar TF activity profile, the samples were clustered based on the normalized PageRank across TFs. First of all, the principal component analysis (PCA) was performed for dimension reduction of the TF score matrix. The first 500 principal components (PCs) were retained for further clustering analysis based on “elbow” method, which explained 85% variance (data not shown). To find the optimal number of groups and similarity metric, the Silhouette analysis was perforated to evaluate the clustering quality using five distance metrics: Euclidean distance, Manhattan distance. Kendall correlation, Pearson correlation, and Spearman correlation (data not shown). Pearson correlation was the most appropriate distance metric since the average Silhouette width was the highest among the five distance metrics. Based on these analyses, 5 Kmeans groups were identified, showing distinct dynamic patterns of TF activity.

[0125] Identification of Kmeans group-specific TFs Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0126] To identify Kmeans group-specific TFs, the clusters were divided into two groups: target group and background group. Target group included the clusters in the Kmeans group of interest and the background group comprised the remaining clusters. Then, the normality test was performed using Shapiro-Wilk’s method to determine whether the two groups were normally distributed and it was found that the PageRank scores of most clusters (95%) didn’t follow normal or log-normal distribution. Thus, Mann-Whitney U Test was used to calculate the P-value. Double cutoffs, i.e. P-value < 0.01 and log2 fold change > 0.5, were used for calling specific TFs (data not shown).

[0127] TF regulatee analysis

[0128] Taiji generated the regulatory network file for each cluster showing the regulatory relationship between TF and regulatees with edge weight, which represents the regulatory' strength. Top 500 regulatees were ranked by mean edge weight across clusters in G2 for each G2-specific TF respectively. To select the representative regulatees. Representative regulatees in FIG. 8 were selected as the top 10 genes regulated by the signature TFs involved in each pathway ranked by the mean edge weight.

[0129] Pathway enrichment analysis

[0130] The enriched functional terms in this study were analyzed by R package clusterProfiler_4.0.5. A cutoff of P-value < 0.05 was used to select the significantly enriched Reactome pathways.

[0131] Cell-cell communication analysis

[0132] The R package CellChat_2.1.2 was used to analyze the intercellular interactions w ithin each individual. First, input scRNA-seq data matrix was normalized by TPM (transcripts per million) method and log-transformed with pseudo count of 1. The assigned cell labels were the cell types identified from co-embedding. Ligand-receptor interaction database was CellChatDB v2 excluding non-protein signaling interactions, which finally includes -2300 validated molecular interactions in the analysis. The default parameters were used following the standard CellChat Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT pipeline. Finally, the intercellular communication networks were obtained for each individual and aggregated together for the downstream visualization.

[0133] Identification of candidate pathogenic genes related to signature group G2

[0134] A customized list of 187 genes was first curated, including all the available cytokines, chemokines, growth factors, NOTCHs, MMPs, and ADMATS with gene expression in this study (data not shown). For each gene, the maximum gene expression across clusters was taken within each Kmeans group and each individual as input. Then, the universal G2-important genes with mean gene expression were identified across all patients ranked as top 50% and Coefficients of variation (CV) less than 2. In total, 63 genes were identified as candidate predictors for the following classification model.

[0135] Classification model construction

[0136] To distinguish the controls from At-Risk / ERA patients, a random forest classification model was developed. The input data was gene expression of identified important genes across patients. For each At-Risk / ERA patient, the maximum gene expression across G2 clusters was taken. For each control, the maximum gene expression across G4 clusters was considered.

[0137] The samples were split into train and test subsets at a 7:3 ratio. The R package Caret_6.0.94 was used for feature importance evaluation based on recursive elimination algorithm implemented in “rfe” function. Only features with positive importance was kept. Random forest model was trained multiple times with an increasing number of predictors, from the most to least important, using 10-fold cross- validation and repeated 5 times. Each trained model was then evaluated on prediction accuracy on the unseen test set. The above process was repeated 20 times with different random seeds from 1 to 20. The mean and standard deviation of the training and testing accuracy was calculated for each number of predictors.

[0138] Comparison with AMP study

[0139] To confirm the expression patterns of newly identified predictors from classification model, we checked the gene expression levels in synovial tissues samples from established RA patients in AMP study. To make it more compatible Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT with cell types in PBMC samples, we only considered 22 clusters defined in original AMP paper that are also present in PBMC populations from 82 synovial tissue samples (Fig. 6A). We collapsed single-cell gene expression profiles into pseudo-bulk count matrices by summing the raw UMI counts for each gene across all cells from the same sample and cluster. For each gene, we normalized counts in each pseudobulk sample into counts per million. We averaged the normalized counts across samples, cell types, and genes and visualized the results as heatmaps in Fig. 6A-C respectively.

[0140] Example 1. Integrative single cell analysis reveals cell types in At-Risk / ERA and CON individuals

[0141] Peripheral blood mononuclear cells (PBMCs) were obtained from 26 ACPA positive (At-Risk) and 6 early RA (ERA) and 35 age and sex-matched 35 controls (CON) and subjected to scATAC-seq and scRNA-seq (FIG. 1A). These data were used to assign each cell to a cell type with Latent Semantic Indexing (LSI) and Principal Component Analysis (PCA) to reduce the dimensionality of the scATAC- seq and scRNA-seq count matrices, respectively. Nearest neighbor graphs in reduced dimensions were built to identify clusters of cells. Uniform Manifold Approximation and Projection (UMAP) was then used to visualize the single cells in reduced dimension space (FIG. IB).

[0142] To integrate scRNA-seq and scATAC-seq for cell type, each cell in the scATAC-seq space was assigned a predicted gene expression profile from the cell in the scRNA-seq that was most similar. Cells from scRNA-seq and scATAC-seq were then clustered in the same co-embedding space for each sample (FIG. 1C). Each coembedded cluster was treated as a pseudo-bulk cluster by summing gene counts from all the scRNA-seq cells and aggregating the raw scATAC-seq peaks. The annotation was defined by the cell fype that occurs most frequently in the cluster. In total, 1610 pseudo-bulk clusters were retained in the final dataset, which included 703.701 scRNA-seq cells and 932,986 scATAC-seq cells, or 1,636,687 cells from 67 samples (median: 25194 cells / sample, 767 cells / cluster; data not shown).

[0143] The cells were assigned to 22 fine-grain transcriptional cell type for each sample. Eleven major cell types, including B memory cells, B intermediate cells, B naive cells. CD 14 monocytes (CD 14 Mono), CD 16 monocytes (CD 16 Mono), CD4 Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT naive T cells (CD4 T Naive), central memory CD4 T cells (CD4 TCM), CD8 naive T cells (CD8 T Naive), effector memory CD8 T cells (CD8 TEM), mucosal-associated invariant T cells (MAIT cells), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright) and regulatory T cells (Treg), accounted for > 99% of total cells and had a sufficient number of cells for subsequent analysis (FIG. ID). Two subtypes of CD4 T cells (CD4 T Naive and CD4 TCM) were the most abundant cell type among all 3 cohorts of PBMC samples with >20% of total cells on average. B intermediate cells, B memory cells, CD16 Mono, NK_CD56bright, and Treg cells were relatively rare cell subsets with each comprising <2% of total cells. The cell ty pes showed similar distribution across At-Risk, ERA and CON groups except for B intermediate. B memory, and NK CD56bright. which were modestly higher in ERA compared to two other groups (Kruskal-Wallis H test, p-value = 0.1, 0.04, and 0.08 respectively) (FIG. IE). The cluster purify was calculated as the percentage of the cells of most abundant cell type for all the 1610 clusters (data not shown), which was 0.72 ± 0.19 across all clusters. The cluster purify showed minor different distributions across cell types. B naive, CD 14 Mono, CD 16 Mono, MAIT, and NK displayed the highest purify scores (mean: 0.86 ± 0.13) while purify scores for T cell subsets were more diverse across clusters and relatively lower (mean: 0.68 ± 0.18). T cell subsets were sometimes included with other T cells. For instance, CD4 TCM cluster showed some other T cells like CD4 T Naive, CD8 T Naive, and CD8 TEM (data not shown).

[0144] Example 2. Taiji analysis reveals distinctive TF patterns

[0145] Single cells within the same cluster are treated as one “pseudo-bulk” sample with the annotation as the cell type occurring most frequently in the cluster. The gene counts of scRNA-seq were added up and the fragments of scATAC-seq were combined to generate the RNA-seq input and ATAC-seq input for the pseudo-bulk samples respectively. Then, the Taiji pipeline was applied to each individual cluster in each patient to evaluate the PageRank scores of TFs, which represents the importance of the TFs. To characterize the global influences of all 1047 TFs across different pseudo-bulk clusters, the clusters were grouped based on the normalized PageRank across TFs. First, PC A was performed for dimension reduction of the TF score matrix with the first 500 principal components (PCs) retained for further analysis based on the “elbow” method, which explained 85% variance. The first several PCs are Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT primarily related to cell type rather than the disease state or the specific cohorts. To determine the optimal number of groups and similarity metrics. Silhouette method was used to evaluate the clustering quality using five distance metrics: Euclidean distance, Manhattan distance, Kendall correlation, Pearson correlation, and Spearman correlation. Pearson correlation was the most appropriate distance metric since the average Silhouette width is the highest among the five distance metrics.

[0146] 5 Kmeans groups were identified by unsupervised clustering, denoted G1 through G5, each of which showed distinct patterns of TF activity (data not shown). The row-wise comparison demonstrates that some TFs have high PageRank scores in one or several Kmeans groups and suggests high TF activity' in specific clusters (FIG. 2A). In total, 640 TFs were identified as Kmeans group-specific TFs by comparing their PageRank scores between a specific group and the background groups (FIG. 2A). These TFs functionally correlated with assigned cell ty pes. For instance, KLF4, which regulates monocyte differentiation, was G1 -specific. G1 was enriched with two subsets of monocytes, including 59.5% CD14 Mono and 31.3% CD16 Mono. T-bet (encoded by 1BX21) and EOMES displayed high activities in G3 where CD8 TEM and NK were the most abundant cell types with 37.9% and 40.3%, respectively. Those two genes are responsible for the cell fates of memory' CD8+T cells and natural killer cells (see FIG. 2A-B; for lineage and group specific TFs that define each Kmeans group). Interestingly, more than half (409 / 640) of the TFs were G2-specific and their z scores were significantly higher in G2 compared to other groups. More than 80% (531 / 640) of the TFs w ere identified as key TFs for only one Kmeans group, suggesting the 5 Kmeans groups had unique active TF patterns.

[0147] Example 3. G2 is a multi-lineage group enriched with At-Risk / ERA and reveals an RA TF signature

[0148] The 5 Kmeans groups generally showed diverse compositions of cell types and disease states (data not shown). As noted above, 4 of the 5 Kmeans groups had their own predominant cell types and accounted for more than 70% of their total clusters. Gl, G3, G4, and G5 were enriched in monocytes; CD8 TEM and NK cells; CD4 T cells; B cells, respectively. However, G2 was unique in that it was mixed and displayed a cell type distribution similar to the overall PBMC distribution and included all 11 major cell types (FIG. 2B). Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0149] G2 was significantly enriched in At-Risk and ERA clusters compared with CON (58% higher in At-Risk and ERA vs. CON, adjusted by the null distribution, p- value < 0.0001; Chi-squared test) and G4 was modestly enriched in CON clusters (24% higher in CON, p-value < 0.0001; Chi-squared test) (FIG. 2C). Many TFs were G2-specific, including zinc finger family members like ZNF304, SP7, GLIS1, ZNF254. For the subsequent analysis, At-Risk and ERA (e.g., At-Risk / ERA) were combined because their TF activity profiles and cell type distributions in G2 were nearly identical (p-value > 0.2; Wilcoxon rank-sum test). Moreover, the identified G2-specific TFs along with the enriched pathways for ERA and At-Risk respectively showed almost complete overlap (p-value < 10'5; data not shown).

[0150] Multiple immunity -related TFs and the downstream genes regulated by those TFs conformed to pathways implicated in the pathogenesis of RA (FIG. 2D). This was particularly true for G2, where 5 relevant and significant pathways were identified, namely SUMOylation of Intracellular Receptors, Transcriptional regulation by RUNX2, YAP1 and WWTRI -stimulated Gene Expression. NOTCH3 Intracellular Domain Regulates Transcription, and Deactivation of the f>-Catenin Transactivating Complex Reactome pathways. The TFs and the representative target genes identified by our analysis are shown in FIG. 8. These TFs and their downstream regulated genes are referred to as the RA TF signature. These TFs were significantly important in the signature pathways and the representative genes were among the top regulated genes by the corresponding TFs predicted by Taiji (Methods).

[0151] Example 4. The G2 RA TF signature is enriched in multiple cell types

[0152] The At-Risk / ERA TFs identified in G2 were present across all 11 major cell types analyzed (FIG. 3A), thereby establishing them as a hallmark ’‘RA TF signature” and their downstream pathways as “signature pathways”. The percentage of G2 clusters per cell type of total global clusters was further calculated for At-Risk / ERA and CON groups (FIG. 3B) Notably, CD4 T Naive, CD4 TCM, and CD8 T Naive showed the greatest enrichment in At-Risk / ERA compared to CON (31% vs 18%, p- value < 0.01; 23% vs 12%, p-value < 0.01; 65% vs 26%, p-value < 0.01, respectively for At-Risk / ERA compared with CON; Chi-squared test). Of interest. MAIT cells with the TF profile were only found in CON clusters (0% vs 43% for At-Risk / ERA Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT and CON, p-value < 0.1; Chi-squared test). Despite the negative correlation between MAIT cell abundance and age, the comparable age of the CON group with At- Risk / ERA suggests that age does not account for these differences and MAIT cells might be protective of conversion / progression of RA. Overall, the top RA signature TFs determined by unsupervised clustering showed significantly higher PageRank scores in G2 compared to other groups across all cell types (FIG. 3C).

[0153] All the major cell types were enriched in this common set of At-Risk / ERA signature pathways while some individual cell types demonstrated specific enriched pathways (FIG. 3D). For example, activation of HOX genes was enriched in B cells, CD4 T cells, CD8 T Naive, and monocytes. RUNX3 regulation is more highly associated with CD8 TEM, NK, CD4 T Naive, and monocytes. Despite individual variations described above, the general pattern of pathways associated with pathogenesis of RA is consistent and extends across the identified cell types.

[0154] Example 5. Patterns of cell types with the G2 RA TF signature are highly variable across individuals

[0155] The cell types that display the TF signature in each member of the At-Risk and ERA cohorts were determined. Multiple combinations of cell ty pes were identified in individual participants (FIG. 3E). Twenty -five out of 26 At-Risk and 6 ERA participants had the signature in at least one cluster and in at least one cell type. However, the distribution of cell types was highly variable among participants. Tn some cases, only one cell type was identified for an individual participant, while in others there were multiple cell types displaying the pattern. For instance, participant 9 had clusters with the signature in all the cell types except NK and Treg. while participant 27 only had CD4 TCM clusters. Some patients displayed more even distribution across multiple cell types like participant 31 while others had predominant signature cell ty pe like participant 3.

[0156] Among all the involved cell types, the signature was most enriched in T cell types including CD4 Naive, CD8 Naive. CD4 TCM, and CD8 TEM (FIG. 3E). Different cell types also displayed diverse distribution patterns across patients. CD4 Naive and CD4 TCM had much wider appearances in many patients while Treg, B cell and monocytes were only found in a few participants. Therefore, the patterns displayed by various individuals were diverse with highly variable cell types. Some Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0157] CONs also displayed these signatures although the number of clusters was significantly less than At-Risk / ERA. particularly for certain T cell subsets (p-value 0.005; Wilcoxon rank-sum test).

[0158] Example 6. Distinct cellular communication networks in At-Risk / ERA and control participants

[0159] After demonstrating individualized patterns of signature cluster cell types in At-Risk / ERA, how the signature cells communicate was investigated to determine how inflammation signals are transmitted. Cell-cell communications (CCC) were analyzed by correlating expression levels of ligands such as cytokines in the source cells with their corresponding receptor expressions in the receiver cells for each individual using CellChat. To compare At-Risk / ERA and CON groups, we first aggregated CCC between the same signature cells across all the ligand-receptor pairs and all the individuals within the group. Distinct CCC patterns were observed: At- Risk / ERA participants displayed significantly more interactions within signature clusters than controls, particularly between T cells and NK cells. Cellular communications with signature monocytes were less common and only observed in the At-Risk / ERA group (FIG. 4A). The difference between the total number of CCC in the two groups was statistically significant (p-value=0.06 using Wilcoxon rank-sum test).

[0160] The cellular communication strength was evaluated. Notably, communication between CD8 T Naive, CD4 TCM, and CD4 T Naive were more pronounced in At- Risk / ERA group, while communications between CD4 T Naive, and CD8 TEM were more intense in controls (FIG. 4B). The total strength in the two groups differed significantly between the two groups (p-value=0.04 using Wilcoxon rank-sum test). As a representative example, participant 53 from control group and participant 9 from At-Risk / ERA group had the most diverse cell type distribution in signature clusters (FIG. 4C), providing an overview of almost all the cell types. It is worth noting that the number and intensity of the total CCC aggregating all the clusters from all the Kmeans groups were comparable between the At- Risk / ERA and CON groups, highlighting the importance of the signature cells differentiating the two groups.

[0161] Similar to the diversity of signature cell types across individuals, the CCC pattern also varied from individual to individual. For example, major senders and Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT receivers were highly variable among individual participants. Some individuals such as participant 5, 26, and 27 used only one cell type as major communicator while others like participant 9, 18, and 23 relied on multiple cell types. Among those with multiple cell types, some displayed more even distributions of signals across cell types like participant 9 and 23 while others exhibited a predominant signature cell type (e.g., CD8 TEM in participant 18).

[0162] Example 7. Identifying key representative mediators regulated by the RA TF signature

[0163] A large number of inflammatory mediators, such as cytokines, chemokines. and growth factors, have been implicated in RA pathogenesis. A gene list of these mediators was evaluated (Methods), which were refer to as “pathogenic genes”. Out of the identified significant ligand-receptor pairs in each participant, twelve ligandreceptor pairs were related to this pathogenic gene set. The important pathways were ranked based on the difference in total information flow within signature clusters when comparing At-Risk / ERA to control samples. The IL 16 - CD4, CD 160 - TNFRSF14, TGF-pi - (TGFBR1+TGFBR2), BTLA - TNFRSF14 were the most prominent ligand-receptor pairs enriched in At-Risk / ERA considering both the difference and absolute information flow values (FIG. 4D).

[0164] The IL 16 - CD4 signaling path ay, which has been implicated in RA, showed significantly stronger signals in At-Risk / ERA group than control group. For instance, participant 31 from At-Risk / ERA group and participant 48 from control group have similar cell type distribution in signature clusters (FIG. 7A) Participant 31 displayed denser and stronger interactions than participant 48 and signature clusters are more likely to act as major senders than receivers (FIG. 4E). We then summarized the outgoing and incoming signals of IL16 - CD4 pair betw een the two groups (FIG. 4F). Multiple cell types send signals of IL 16, including B cells and monocytes that are unique senders in At-Risk / ERA and CD8 TEM and monocytes are unique receivers in At-Risk / ERA group. CD4 T cells are most widely used as communicators across participants. B cells and NK cells only act as senders in IL16 signaling pathway.

[0165] Other interesting and relevant pro-inflammatory pathw ays w ere also enriched in At-Risk / ERAs. For instance, TGF-01, which is an important regulator in RA, showed much denser and stronger intercellular communications in participant 18 from Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT

[0166] At-Risk / ERA than participant 48 from control group. Signature clusters were more likely to act as major senders in TGF-pi signaling pathway, particularly in CD4 T cells and CD8 TEM cells (data not shown). On the other hand, signature clusters mainly acted as major receivers in CD 160 - TNFRSF14 signaling pair. NK cells were the most widely used cell ty pe in CD 160 signal communication.

[0167] We then developed a random forest classification with pathogenic gene expression as features. Sixty-three genes were identified as candidate predictors, which were active across each At-Risk / ERA participant in signature group G2 (Methods). The test accuracy w as monotonically increasing with more predictors, reaching a plateau of 0.93. Top predictors included MMP23B, TGFB1, IFNL1, CCL5, and IL15 (FIG. 4E). MMP23B, which emerged as a top predictor in classification model, showed elevated gene expression level in At-Risk / ERA compared to control (FIG. 5B). MMP23B plays a role in regulating the Kvl.3 potassium channel, which has been implicated in autoimmunity'. TGFB1 also show ed elevated gene expression in At-Risk / ERA (FIG. 5C). Additionally, we explored the relationship between the G2 RA signature TFs identified above and TGFB1 as their target gene. Top 30 signature regulators of TGFB1 included some w ell-known RA-related TFs like RORC. TFAP2A, and KLF1 (FIG. 5D).

[0168] Gene expression was greater for the top 30 predictors in At-Risk / ERA participants compared with controls, including CCL4. IL12A. TNFSF14, II.I5. NOTCH R and CCL5 (FIG. 5E). To validate our predictions, protein expression levels of 6 genes using proteomics w ere assessed, each of which confirmed increased protein expression level in the serum of At-Risk / ERA group compared to controls (CCL3. CCL4, IFN-X1, IL-15, TGF-pi, and TNFSF14) (FIG. 5F).

[0169] Although a common set of pathogenic genes were shared across At-Risk / ERA participants, the cell types that were most likely to produce the specific gene were highly variable. For instance, the top 5 predictors were active in CD8 TEM cells and NK cells in most participants while a few participants expressed the genes through CD4 TCM. and CD8 T Naive cells. NOTCH1 and CXCL16 displayed uniform activity across all cell types with the highest activity in Tregs. Some genes showed exclusively high activity7in specific cell ty pes, such as TNFSF9 in CD4 TCM and ADAMTSL4 in monocytes. Mediator expression patterns were individualized towards specific cell types. For example. TGFB1 was highly expressed in CDS TEM and NK Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT cells in most of the patient while it was more highly expressed in B cells in participant 13 and monocytes in participant 7 (data not shown). These findings suggest that At- Risk / ERA individuals express a common set of pathogenic genes, driven by any cell type possessing RA TF signature.

[0170] We then evaluated the Accelerating Medicine Partnerships (AMP) synovial scRNA-seq data to determine if the RA gene expression signature observed in peripheral blood was confirmed in the inflamed tissue and whether a diversity of cell types was also present. The top mediators were also found in synovial cells and, more strikingly, there was broad diversity of the cell types within the synovial tissues that expressed them (FIG. 6A). The number of genes expressed across all cell types displayed distinct patterns across samples (FIG. 6B). This heterogeneity7suggests that each RA patient, like pre-RA and early RA PBMCs, have this molecular signature. We then determined which cell types display the regular signature for each patient. As with PBMCs, distribution of cell type was highly variable among RA samples (FIG. 6C). In some cases, one dominant cell type was identified for an individual participant, such as monocytes in sample BRI-456, while in others there were multiple cell types, such as instance, BRI-552, displayed expression across all the cell types.

[0171] OTHER EMBODIMENTS

[0172] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCTWHAT IS CLAIMED IS:

1. A method for treatment of an autoimmune condition in a subject, the method comprising:(a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject:(b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample;(c) identifying the one or more cell types from the biological sample expressing the TF signature; and(d) administering to the subject a therapeutic agent comprising a therapeutic agent targeting the one or more cell types expressing the TF signature.

2. A method for treatment of an autoimmune condition in a subject, the method comprising:(a) isolating macromolecules or having macromolecules isolated from a biological sample obtained from the subject;(b) determining a disease-associated transcription factor (TF) signature in one or more cell types from the biological sample; and(c) administering a therapeutic agent to modulate the expression of one or more pro-inflammatory downstream genes regulated by the TF signature, wherein the TF signature comprises elevated expression of one or more TFs of the SUMOylation, RUNX2, YAP1, N0TCH3, and / or WNT / p-Cate n pathways.

3. The method of claim 1, wherein the TF signature comprises elevated expression of one or more TFs in one or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p-Catenin pathways.

4. The method of claims 1 or 3, wherein the one or more cell types comprise B memory cells, B intermediate cells, B naive cells, CD14 monocytes, CD16 monocytes, CD4 naive T cells, central memory CD4 T cells (CD4 TCM), CD8 naive T cells, effector memory CD8 T cells (CD8 TEM), natural killer cells (NK), CD56 bright natural killer cells (NK CD56bright). and / or regulatory T cells (Treg).Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT5. The method of claim 2, further comprising determining expression of one or more proinflammatory downstream genes regulated by the TF signature.

6. The method of claim 2, further comprising identifying one or more cell ty pes in the biological sample expressing the disease-associated TF signature, wherein the one or more cell types comprise B memory’ cells, B intermediate cells, B naive cells. CD 14 monocytes, CD 16 monocytes. CD4 naive T cells, central memory CD4 T cells (CD4 TCM), CD8 naive T cells, effector memory CD8 T cells (CD8 TEM), natural killer cells (NK), CD56 bright natural killer cells (NK_CD56bright), and / or regulatory' T cells (Treg).

7. The method of claim 4, wherein the identified cell type of step (c) comprises:B memory’ cells and the therapeutic agent targets B memory cells, orB naive cells and the therapeutic agent targets B naive cells, orCD 14 monocytes and the therapeutic agent targets CD 14 monocytes, or CD 16 monocytes and the therapeutic agent targets CD 16 monocytes, or CD4 naive T cells and the therapeutic agent targets CD4 naive T cells, or central memory CD4 T cells (CD4 TCM) and the therapeutic agent targets CD4 TCM cells, orCD8 naive T cells and the therapeutic agent targets CD8 naive T cells, or effector memory’ CD8 T cells (CD8 TEM) and the therapeutic agent targets effector memory CD8 T cells, or natural killer cells (NK) and the therapeutic agent targets NK cells, or regulatory T cells (Treg) and the therapeutic agent targets regulatory Treg cells.

8. The method of claim 7, wherein the therapeutic agent comprises abatacept. rituximab, ocrelizumab, ofatumumab. epratuzumab. tabalumab. CAR-T cells, CAR- NK cells, tyrosine kinase inhibitors, TLR targeted therapy, memantine, anti-OX40 therapy, chemokine antagonists (e.g., CX3CR1 blockers), anti-NRPl, IDO inhibitors, calcineurin antagonists (e.g., sirolimus), TGF-beta blockade (including biologies and SMAD inhibitors), or PD1 agonists (e.g., peresolimab).Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT9. The method of claims 2. 5, or 6. wherein the pro-inflammatory downstream genes comprise MMP23B, XCL2, CCL4, IFNL1, PDGFD, IL12A, AD AMTS 10, CCL3, CCL4L2, IL15, MMP24OS, IFNG, TGFB1, CCL5, MMP25-AS1, ADAMTS17, NOTCH1, TNFSF9, MMP25, NOTCH2NL, CCL20, ADAMTSL4, CXCL16, TNFSF8, TGFA, IL18BP, MMP19. TGFB3, XCL1, or ADAMTS1.

10. The method of claim 9, wherein the pro-inflammatory downstream genes comprise MMP23B, TGFB1, IFNL1, PDGFD, or CCL5.

11. The method of claims 2, 5, or 6, wherein administering a therapeutic agent to modulate the expression of the pro-inflammatory downstream gene comprises administering an inhibitor directed to the pro-inflammatory downstream gene, wherein the pro-inflammatory downstream gene comprises one or more cytokines or cytokine signal transduction genes.

12. The method of claim 11, wherein the therapeutic agent comprises one or more of a TNF inhibitor, an IL-1 inhibitor, an IL-6 inhibitor, a TGFP inhibitor (e.g., one or more SMAD inhibitors), one or more inhibitors of B cell signaling and activation (e.g., BLyS or BTK), chemokine signal inhibitors (e.g., PI3Ky), rituximab, or a Janus kinase inhibitor.

13. The method of any of the above claims, wherein the therapeutic agent comprises vector-based gene therapy, small molecule activators, biologies. CRISPR- based gene editing, epigenetic modulation, TF modulation, or any combinations thereof.

14. The method of any of the above claims, wherein the macromolecules comprise RNA, DNA, or proteins.

15. The method of an of the above claims, wherein the biological sample comprises blood, saliva, cerebrospinal fluid, bone marrow, bronchoalveolar lavage, sputum, or biological tissue.Attorney Docket No. 15670-0431 WO 1 / UCSD Ref.: SD2025-048-2PCT16. The method of any one of the above claims, wherein the subject is positive for anti-citrullinated protein autoantibodies (ACPA+).

17. The method of any of the above claims, wherein the autoimmune condition comprises rheumatoid arthritis, systemic lupus erythematosus, scleroderma, multiple sclerosis, polymyositis, dermatomyositis, or Sjogren’s syndrome.

18. The method of any of the above claims, wherein the autoimmune condition comprises rheumatoid arthritis.

19. The method of any of the above claims, wherein the one or more TFs comprise SP7, SOX9, NKX3-2, DLX6, HAND2, DLX5, HEY1, MSX2, ZNF521, TWIST1, TWIST2, SATB2, GLI3, AR, HEY2, HES1, NKX2-5, GATA4, TEAD4, TEAD1, TEAD2, TEAD3. PGR, NR5A1, NR5A2, NR1I2, THRB, PPARG, HEYL, HES5, SOX2, SOX6, SOX7, TCF7L1, and / or SOX13.

20. The method of any of the above claims, wherein determining the TF signature comprises comparing the expression of the TFs of the biological sample with a reference level.

21. The method of claim 20, wherein the reference level comprises the level of the one or more TFs from one or more healthy subjects.

22. The method of any of the above claims, wherein the TF signature comprises the elevated expression of one or more TFs in two or more, three or more, or four or more of the SUMOylation, RUNX2, YAP1, NOTCH3, and / or WNT / p- Catenin pathways.

23. The method of any of the above claims, wherein the TF signature comprising the elevated expression of one or more TFs in each of the SUMOylation, RUNX2, YAP1, NOTCH3, and WNT / -Catenin pathways.