Multiple sclerosis-associated t cell receptor-related methods and compositions
Patent Information
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-01-30
- Publication Date
- 2025-08-07
AI Technical Summary
Current diagnostic methods for multiple sclerosis (MS) lack specific biomarkers, leading to misdiagnosis and difficulty in distinguishing MS from other conditions with overlapping symptoms, due to the heterogeneous clinical and imaging manifestations of the disease.
Assessing T cell receptor beta chain complementary determining region 3 (TCRβ CDR3) sequences using immunosequencing to identify specific TCRβ CDR3 sequences associated with MS, enabling differential diagnosis and predicting MS through the presence of unique TCRβ CDR3 sequences in subjects with non-specific symptoms.
Provides a precise diagnostic tool for MS by identifying MS-associated TCRβ CDR3 sequences, improving diagnostic accuracy and reducing misdiagnosis by distinguishing MS from other demyelinating diseases.
Abstract
Description
[0001] MULTIPLE SCLEROSIS-ASSOCIATED T CELL RECEPTOR-RELATED METHODS AND COMPOSITIONS
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 627,627, filed January 31 , 2024, which application is incorporated herein by reference in its entirety.
[0004] INCORPORATION BY REFERENCE OF SEQUENCE LISTING
[0005] PROVIDED AS A SEQUENCE LISTING XML FILE
[0006] A Sequence Listing is provided herewith as a Sequence Listing XML, “ADBS- 095WO_SEQLIST”, created on January 30, 2025 and having a size of 869,067 bytes. The contents of the Sequence Listing XML are incorporated herein by reference in their entirety.
[0007] INTRODUCTION
[0008] Multiple sclerosis (MS) is an immune-mediated disorder of the central nervous system (CNS). In MS, a subject’s immune system is abnormally activated, traffics to the CNS, and causes injury to myelin sheaths that protect neurons in the brain and / or spinal cord.
[0009] Despite significant refinement in MS diagnosis in recent decades, no specific disease biomarkers or signatures thereof exist. MS has heterogeneous clinical and imaging manifestations, which not only differ between patients, but also vary in individual patients over time. Disease signs and symptoms, presence of oligoclonal bands (OCB) and MRI findings have limited specificity, and misdiagnosis remains a problem with significant clinical and psychosocial implications for both patients as well as health care providers. A wide range of conditions can be mistaken for MS, including: migraine, cerebral small vessel disease, fibromyalgia, functional neurological disorders, and neuromyelitis optica spectrum disorders, along with uncommon inflammatory, infectious and metabolic conditions.
[0010] Antigen-specific cellular immune responses are mediated by a diverse population of T cells and B cells, each bearing immune cell receptors (T cell receptors (TCRs) and B cell receptors (BCRs), respectively) capable of recognizing a specific antigen (in the case of T cells, an antigen peptide bound to a particular major histocompatibility complex (MHO) molecule on the surface of host cells). Encounter with an antigen leads to the clonal expansion, activation, and maturation of T and B cells, resulting in effector populations of cytotoxic (CD8+CTL) and helper (CD4+) T cells, or antibodies and memory B cells, respectively. The presence of antigen-specific effector cells is diagnostic of an immune response specific to that antigen. Activated T cells proliferate by clonal expansion and reside in the memory T cell compartment for many years as a clonal population of cells (clones) with identical-by-descent rearranged TCR genes (Arstila T P, et al. A direct estimate of the human alpha / beta T cell receptor diversity, Science 286: 958-961 ).
[0011] The majority of TCR diversity resides in the beta chain of the TCR a / p heterodimer. Immense diversity is generated by combining noncontiguous TCR|3 variable (V), diversity (D), and joining (J) region gene segments, which collectively encode the CDR3 region, the primary region of the TCR0 locus for determining antigen specificity. Deletion and templateindependent insertion of nucleotides during rearrangement at the Vp-D(3 and D[3-J p junctions further add to the potential diversity of receptors that can be encoded (Cabaniols JP, et al. Most alpha / beta T cell receptor diversity is due to terminal deoxynucleotidyl transferase, J Exp Med 194: 1385-1390, 2001 ). Typically, at a given point in time, an adult with a healthy immune system expresses approximately 10 million unique TCR0 chains on their 1012circulating T cells (Robins HS, et al. (2009) Comprehensive assessment of T-cell receptor beta-chain diversity in alpha / beta T cells, Blood 114: 4099-4107).
[0012] The human T-cell repertoire thus dynamically encodes exposure to disease-related antigens through rearrangements of their receptor-encoding genes and so provides an excellent basis for making diagnostic predictions. It has been demonstrated that TCR0 receptors in peripheral blood samples from human subjects can be employed to predict the status of exposure to a disease; i.e., based on the presence and abundance of such receptors in the training cohort (Emerson et al., Immunosequencing identifies signatures of cytomegalovirus exposure history and human leukocyte antigen (HLA)-mediated effects on the T cell repertoire, Nature Genetics April 2017; doi:10.1038 / ng.3822).
[0013] SUMMARY
[0014] Provided are methods for assessing T cell receptor 0 chain complementary determining region 3 (TCR0 CDR3) sequences. In certain embodiments, prior to the assessing, the subject has been identified as having, or is suspected of having, multiple sclerosis (MS). According to some embodiments, at the time of the assessing, the subject has one or more non-specific symptoms consistent with MS. Also provided are methods comprising administering an MS therapy to a subject identified as comprising T cells that express a T cell receptor 0 chain (TCR0) comprising a TCR0 CDR3 sequence set forth in the present disclosure. Computer readable media and systems for assessing TCR0 CDR3 sequences are also provided. BRIEF DESCRIPTION OF THE FIGURES
[0015] FIG. 1 : Data illustrating the identification of multiple sclerosis-associated TCRp sequences that distinguish multiple sclerosis patients from controls (either healthy or other autoimmune but non-multiple sclerosis-diseased adults) in a holdout set of samples in each genetic background: with HLA DRB1 *15:01 (Panel A) and without HLA allele DRB1 *15:01 (Panel B). Plotted lines show example classification thresholds of 0 (bottom), 2 (middle), and 4 (top) standard deviations above the fitted logistic growth curve.
[0016] FIG. 2: Schematic illustration and data demonstrating examples of specific multiple sclerosis-associated TCRp sequence clusters (Panel A) and show enrichment in multiple sclerosis patients (Panel B).
[0017] DETAILED DESCRIPTION
[0018] Before the methods of the present disclosure are described in greater detail, it is to be understood that the methods are not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the methods will be limited only by the appended claims.
[0019] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the methods. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the methods, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the methods.
[0020] Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0021] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the methods belong. Although any methods similar or equivalent to those described herein can also be used in the practice or testing of the methods, representative illustrative methods are now described.
[0022] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the materials and / or methods in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present methods are not entitled to antedate such publication, as the date of publication provided may be different from the actual publication date which may need to be independently confirmed.
[0023] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0024] It is appreciated that certain features of the methods, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the methods, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed, to the extent that such combinations embrace operable processes and / or compositions. In addition, all sub-combinations listed in the embodiments describing such variables are also specifically embraced by the present methods and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0025] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present methods. Any recited method can be carried out in the order of events recited or in any other order that is logically possible. METHODS FOR ASSESSING TCRp CDR3 SEQUENCES
[0026] The present disclosure provides methods for assessing T cell receptor chain complementary determining region 3 (TCRP CDR3) sequences. In certain embodiments, the methods comprise assessing TCRp sequences determined from a sample obtained from a subject for the presence or absence of one or more TCRp sequences set forth in SEQ ID Nos:1 -997 herein, e.g., SEQ ID Nos:1 -100 herein. The inventors have determined that TCRs comprising such TCRp sequences are associated with multiple sclerosis (MS) by being statistically more prevalent in individuals having MS than those who do not have multiple MS, where the association with multiple sclerosis is observed either among an entire population of individuals or among subsets of individuals inferred as carrying particular human leukocyte antigen (HLA) alleles. In certain embodiments, the methods comprise assessing TCRa sequences determined from a sample obtained from a subject for the presence or absence of one or more TCRa sequences determined to be associated with MS, where the association with MS is observed either among an entire population of individuals or among subsets of individuals inferred as carrying particular human leukocyte antigen (HLA) alleles.
[0027] Accordingly, the methods of the present disclosure find use, for example, in predicting whether a subject has or does not have MS. Prior to the assessing, the subject may have been identified as having, or is suspected of having, demyelinating disease. In certain embodiments, for a subject identified as having demyelinating disease or suspected of having demyelinating disease, the methods find use in diagnosing the subject as having multiple sclerosis. Such a diagnosis may be a differential diagnosis in which a subject exhibiting one or more non-specific symptoms consistent with multiple sclerosis is diagnosed as having multiple sclerosis and not another condition characterized by symptoms which overlap with those of MS, including but not limited to, neuromyelitis optica spectrum disorders, myelin oligodendrocyte glycoprotein antibody-associated disease, Susac syndrome, and / or optic neuritis. Details regarding the methods of the present disclosure will now be described.
[0028] According to some embodiments, the methods of the present disclosure are computer-implemented. By “computer-implemented” is meant at least one step of the method is implemented using one or more processors and one or more non-transitory computer- readable media. For example, in certain embodiments, provided are computer-implemented methods for assessing TCRp CDR3 sequences, the methods being implemented using one or more processors and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to assess TCRp CDR3 sequences determined from a sample obtained from a subject for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 herein. The computer-implemented methods of the present disclosure may further comprise one or more steps that are not computer-implemented, e.g., obtaining a sample (e.g., a blood sample, nervous tissue sample, or the like) from the subject, preparing the sample for immune repertoire nucleic acid sequencing, administering an MS therapy to a subject diagnosed with MS based on the assessment, and / or the like.
[0029] According to some embodiments, the subject has one or more non-specific symptoms consistent with MS at the time of the assessing. Examples of such non-specific symptoms include, but are not limited to, CNS inflammation, demyelination, bladder and / or bowel dysfunction, diplopia, dysphagia, altered sensation or weakness of the face, ataxia, blurry vision with painful eye movements, asymmetric limb weakness, spasticity, impaired mobility, impaired motor dexterity, and cognitive impairment. The methods of the present disclosure find use, e.g., in providing a differential diagnosis based on the assessing in which a subject who has one or more non-specific symptoms consistent with MS is diagnosed as having MS and not another condition characterized by symptoms which overlap with those of MS.
[0030] As summarized above, the methods of the present disclosure comprise assessing the TCRp CDR3 sequences determined from the sample obtained from the subject for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997 set forth herein. As noted above, in certain embodiments, the assessing step may be computer-implemented such that it is performed using one or more processors and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to assess the determined TCRp CDR3 sequences for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997. For example, the instructions may cause the one or more processors to compare each of the determined TCRp CDR3 sequences (e.g., each determined TCRp CDR3 sequence or each unique determined TCRp CDR3 sequence) stored on a computer-readable medium to a database comprising one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 stored on the same or a different computer-readable medium. According to some embodiments, the number of TCRp CDR3 sequences determined from the sample obtained from the subject is from 1 ,000 to 2,000,000. For example, in certain embodiments, the number of determined TCRp CDR3 sequences is 2,000,000 or fewer (e.g., 1 ,500,000 or fewer, 1 ,250,000 or fewer, 1 ,000,000 or fewer, 750,000 or fewer, or 500,000 or fewer), but 1 ,000 or more, 5,000 or more, 10,000 or more, 15,000 or more, 20,000 or more, 25,000 or more, 30,000 or more, 35,000 or more, 40,000 or more, 45,000 or more, 50,000 or more, 55,000 or more, 60,000 or more, 65,000 or more, 70,000 or more, 75,000 or more, 80,000 or more, 85,000 or more, 90,000 or more, 95,000 or more, or 100,000 or more. The number of TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 to which the determined TCRp CDR3 sequences is compared may vary. For example, the determined TCRp CDR3 sequences may be compared to 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 75 or more, 100 or more, 150 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 450 or more, 500 or more, 550 or more, 600 or more, 650 or more, 700 or more, 750 or more, 800 or more, 850 or more, 900 or more, 950 or more, or each of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997. When the determined TCRp CDR3 sequences are compared to fewer than all of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997, the determined TCRp CDR3 sequences may be compared to any desired number (e.g., as set forth above) and any desired combination of TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
[0031] The methods of the present disclosure may include one or more additional steps based on the results of the assessing step. For example, if it is determined from the assessing step that none of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 are present in the TCRp CDR3 sequences determined from the sample obtained from the subject (e.g., a subject exhibiting one or more non-specific symptoms consistent with MS), then the methods may further comprise, e.g., identifying the subject as not having MS. Also by way of example, if it is determined from the assessing step that one or more (e.g., 2 or more, 3 or more, 4 or more, 5 or more, or 10 or more) of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 are present in the TCRp CDR3 sequences determined from the sample obtained from the subject (e.g., a subject exhibiting one or more non-specific symptoms consistent with MS), then the methods may further comprise, e.g., predicting that the subject has MS, diagnosing the subject as having MS, identifying the subject as one who should be administered an MS therapy, and / or administering an MS therapy to the subject, e.g., administering to the subject one or more of any of the MS therapies described elsewhere herein.
[0032] In certain embodiments, the methods further comprise subjecting the results of the assessing step to further analysis, such as subjecting the results of the assessing step to a model. For example, the methods may further comprise subjecting the results of the assessing step to a model in order to classify the subject as having MS or not having MS; and / or to classify the subject as having MS and not having a non-MS demyelinating disease, e.g., neuromyelitis optica spectrum disorder. According to some embodiments, the methods may comprise subjecting the results of the assessing step to multiple models, for instance, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, or 25 or more models. In one non-limiting example, the assessed sequences may be subjected to 2 models - a first model to perform MS classification in samples inferred as carrying the HLA allele DRB1 *15:01 , and a second model to perform MS classification in samples inferred as not carrying the HLA allele DRB1 *15:01. In some instances, the HLA inference of a sample carrying or not carrying the DRB1*15:01 allele may be based on genotyping, while in other embodiments, the HLA inference of a sample carrying or not carrying the DRB1 *15:01 allele may be performed based on applying a classification model to the TCRp repertoire (see, e.g., U.S. Patent Application Publication No. US 2021 / 0381050 A1 , the disclosure of which is incorporated herein by reference in its entirety for all purposes). One of ordinary skill in the art will appreciate that, with the benefit of the TCRp sequences set forth in SEQ ID Nos:1- 997 described herein, a variety of useful models may be applied to the results of the assessment. In one non-limiting example, the methods may further comprise subjecting the results of the assessing step to a two-feature logistic regression with features representing the number of MS-associated TCRp sequences determined from the sample and the total number of unique TCRP sequences determined from the sample.
[0033] In certain embodiments, when the methods further comprise subjecting the results of the assessing step to a model for classification purposes (e.g., as described above), the model may take into account the number of unique MS-associated TCRp sequences that are present in the TCRp sequences determined from the sample, e.g., where the greater the number of unique MS-associated TCRp sequences, the more likely the model is to classify the subject as having MS. According to some embodiments, the number of unique MS- associated TCRp sequences is not a feature utilized by the model to classify the subject.
[0034] In some instances, the model may take into account both the number of MS- associated TCRp sequences determined from the sample and the total number of unique TCRp sequences determined from the sample. In one non-limiting example, a logistic growth curve of best fit and an accompanying error model may be derived based on samples without MS diagnosis, and this fit logistic growth curve and error model may be used to classify samples as having or not having MS. In one non-limiting example of such a logistic growth curve, the logarithm of the total number of unique TCRp sequences may form the domain of the logistic growth function, and the number of unique MS-associated TCRp sequences may form the range of the logistic growth function. In one non-limiting case of an accompanying error model, a binomial error model can be derived to accompany the fit logistic growth curve, such that if a particular point on the curve assumes a fraction F of the upper bound on values taken by the curve, the accompanying standard deviation of this fraction F is equal to a constant times the square root of the product of F and the quantity one minus F, as stated in Greissl et al. (2021 ) medRxiv doi.org / 10.1 101 / 2021 .07.30.21261353. In one non-limiting example of how the logistic growth curve and error model may be used to classify samples as having or not having MS, each sample may be given a score equal to the number of standard deviations greater than the value on the logistic growth curve, with positive scores for points above the curve and negative scores for points below the curve, following which samples exceeding a threshold level of this score can be classified as having MS, while those samples not exceeding this threshold level are classified as not having MS. As demonstrated in the Experimental section below, such a model exhibits high specificity and sensitivity for MS patients inferred as carrying the DRB1 *15:01 allele.
[0035] In certain embodiments, the presence and / or frequency of one or more particular unique MS-associated TCR|3 sequences is a feature(s) used by the model to classify the subject. For example, the presence and / or frequency of one or more particular unique MS- associated TCRp sequences may be given relatively greater weight when classifying the subject as compared to the presence and / or frequency of one or more other unique MS- associated TCRp sequences.
[0036] According to some embodiments, when a classification model weighs particular unique MS-associated TCRp sequences differently than other unique MS-associated TCRp sequences, the model may use convergent recombination to weigh the sequences differently. Different T cells can show convergent recombination where unique DNA sequences were formed in the recombination for a first T cell, a second T cell, a third T cell, etc., but where each leads to the same protein (CDR3 + V-gene + J-gene) which is diagnostic for high likelihood of MS. This convergent recombination may be more likely for certain MS- associated TCRp sequences than others, and the model may take into account these aspects of the signal reflective of the interpretable biology of immune response. Accordingly, in some embodiments, sequences may be given differential weight based on convergent recombination.
[0037] One of ordinary skill in the art will appreciate that, with the benefit of TCRa sequences assessed as being associated with MS, any of the models detailed above can be analogously applied to the results of the assessment, using the assessed TCRa sequences rather than TCRp sequences. In addition, one of ordinary skill in the art will appreciate that, with the benefit of a set of TCRa sequences and a set of TCRp sequences both assessed as being associated with MS, any of the models detailed above can be analogously applied to the results of the assessment, using the combination of assessed TCRa sequences and TCRp sequences rather than a set of TCRp sequences alone. In some instances, by “TCRp sequence” is meant the combination of the TCRp CDR3 sequence, the TCRP V gene (TCR -V), and the TCR J gene (TCRp-J). According to some embodiments, by “TCRa sequence” is meant the combination of the TCRa CDR3 sequence, the TCRa V gene (TCRa-V), and the TCRa J gene (TCRa-J).
[0038] In certain embodiments, prior to the assessing step, the methods may further include one or more steps for determining the TCRp CDR3 sequences from the sample obtained from the subject. For example, the determining may include immunosequencing and evaluation of the T cell repertoire in the biological sample obtained from the subject, e.g., by high-throughput sequencing (HTS) as described elsewhere herein. The determining may be partially implemented using a computer. For example, the analysis of the raw sequencing data may be implemented by a computer. Extraction of DNA or RNA from the biological sample, amplification, and sequencing may be performed manually, using a machine, or a combination thereof. In certain embodiments, the methods may further comprise an initial step of obtaining the biological sample from the subject.
[0039] The biological sample (e.g., peripheral blood, gut tissue, and / or the like) may be obtained from a variety of subjects. Such subjects may be “mammals” or “mammalian,” where these terms are used broadly to describe organisms which are within the class mammalia, including the orders carnivore (e.g., dogs and cats), rodentia (e.g., mice, guinea pigs, and rats), and primates (e.g., humans, non-human primates such as chimpanzees, and monkeys). In some embodiments, the subject is a human subject.
[0040] Biological samples of interest include those that comprise T cells, including but not limited to, whole blood samples, a fraction of whole blood comprising peripheral blood mononuclear cells (e.g., blood plasma), serum, a peripheral blood mononuclear cell (PBMC) sample, a gut tissue sample, urine, buffy coat, synovial fluid, bone marrow, cerebrospinal fluid, saliva, lymph fluid, seminal fluid, vaginal secretions, urethral secretions, exudate, transdermal exudates, pharyngeal exudates, nasal secretions, sputum, sweat, bronchoalveolar lavage, tracheal aspirations, fluid from joints, or vitreous fluid. T cells can also be obtained from biological samples which may be derived from, for example, solid tissue samples. T-cells may be helper T cells (effector T cells or Th cells), cytotoxic T cells (CTLs), memory T cells, and regulatory T cells. In some embodiments, peripheral blood mononuclear cells (PBMC) are isolated by techniques known to those of skill in the art, e.g., by Ficoll-Hypaque® density gradient separation.
[0041] Nucleic acid, such as, genomic DNA or RNA may be extracted from lymphoid cells by methods known to those of skill in the art. Examples include using the QIAamp® DNA blood Mini Kit or a Qiagen DNeasy Blood extraction kit (both commercially available from Qiagen, Gaithersburg, Md., USA) to extract genomic DNA. In some embodiments, 100,000 to 200,000 cells may be used for analysis of diversity, i.e. , about 0.6 to 1 .2 g DNA from diploid T cells. Using PBMCs as a source, the number of T cells can be estimated to be about 30% of total cells. Alternatively, total nucleic acid can be isolated from cells, including both genomic DNA and mRNA. In other embodiments, cDNA is transcribed from mRNA and then used as templates for amplification. The RNA molecules can be transcribed to cDNA using known reverse-transcription kits, such as the SMARTer™ Ultra Low RNA kit for Illumina sequencing (Clontech, Mountain View, Calif.) essentially according to the supplier's instructions.
[0042] Immune Repertoire Sequencing (Multiplex PGR and High Throughput Sequencing)
[0043] According to some embodiments, TCR[3 CDR3 sequences are determined from the sample obtained from the subject by immune cell receptor sequencing, e.g., immune repertoire sequencing.
[0044] By “T cell receptor” or “TCR” is meant a disulfide-linked membrane bound heterodimeric protein normally consisting of the highly variable a and p chains expressed as part of a complex with the invariant CD3 chain molecules. T cells expressing these two chains are referred to as a:(3 (or a[3) T cells, though a minority of T cells express an alternate receptor, formed by variable y and o chains, referred as ya T cells. TCR development occurs through a lymphocyte specific process of gene recombination, which assembles a final sequence from a large number of potential segments. This genetic recombination of TCR gene segments in somatic T cells occurs during the early stages of development in the thymus. The TCRa gene locus contains variable (V) and joining (J) gene segments (Vo and Jo), whereas the TCRp locus contains a D gene segment in addition to Vp and Jp segments. Accordingly, the a chain is generated from VJ recombination and the p chain is involved in VDJ recombination. This is similar for the development of y6 TCRs, in which the TCRy chain is involved in VJ recombination and the TCRS gene is generated from VDJ recombination. The TCR a chain gene locus consists of 46 variable segments, 8 joining segments and the constant region. The TCR p chain gene locus consists of 48 variable segments followed by two diversity segments, 12 joining segments and two constant regions. The D and J segments are located within a relatively short 50 kb region while the variable genes are spread over a large region of 1 .5 mega bases (TCRa) or 0.67 mega bases (TCRp).
[0045] TCRp CDR3 sequence determination may involve quantitative detection of sequences of substantially all possible TCR gene rearrangements that can be present in a sample containing lymphoid cell DNA.
[0046] Amplified nucleic acid molecules comprising rearranged TCR regions obtained from a biological sample are sequenced using high-throughput sequencing. In one embodiment, a multiplex PGR system is used to amplify rearranged TCR loci from genomic DNA as described in U.S. Pub. No. 2010 / 0330571 , filed on Jun. 4, 2010, U.S. Pub. No. 2012 / 0058902, filed on Aug. 24, 2011 , International App. No. PCT / US2013 / 062925, filed on Oct. 1 , 2013, which is each incorporated by reference in its entirety.
[0047] To that end, multiplex PCR is performed using a set of forward primers that specifically hybridize to V segments and a set of reverse primers that specifically hybridize to the J segments of a TCR locus, where a multiplex PCR reaction using the primers allows amplification of all the possible VJ (and VDJ) combinations within a given population of T cells.
[0048] Exemplary V segment primers and J segment primers are described in US2012 / 0058902, US2010 / 033057, W02010 / 151416, WO2011 / 106738, US2015 / 0299785, WO2012 / 027503, US2013 / 0288237, U.S. Pat. No. 9,181 ,590, U.S. Pat. No. 9,181 ,591 , US2013 / 0253842, WC2013 / 188831 , which are each herein incorporated by reference in their entireties.
[0049] A multiplex PCR system can be used to amplify rearranged immune cell receptor loci. In certain embodiments, the CDR3 region is amplified from a TCRB CDR3 region locus. A plurality of V-segment and J-segment primers are used to amplify substantially all (e.g., greater than 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%) rearranged immune cell receptor CDR3-encoding regions to produce a multiplicity of amplified rearranged DNA molecules. In certain embodiments, primers are designed so that each amplified rearranged DNA molecule is less than 600 nucleotides in length, thereby excluding amplification products from non-rearranged immune cell receptor loci.
[0050] In some embodiments, two pools of primers are used in a single, highly multiplexed PCR reaction. The "forward" pool of primers can include a plurality of V segment oligonucleotide primers and the reverse pool can include a plurality of J segment oligonucleotide primers. In some embodiments, there is a primer that is specific to (e.g., having a nucleotide sequence complementary to a unique sequence region of) each V region segment and to each J region segment in the respective TCR or Ig gene locus. In other embodiments, a primer can hybridize to one or more V segments or J segments, thereby reducing the number of primers required in the multiplex PCR. In certain embodiments, the J-segment primers anneal to a conserved sequence in the joining ("J") segment.
[0051] Each primer can be designed such that a respective amplified DNA segment is obtained that includes a sequence portion of sufficient length to identify each J segment unambiguously based on sequence differences amongst known J-region encoding gene segments in the human genome database, and also to include a sequence portion to which a J-segment-specific primer can anneal for resequencing. This design of V- and J-segment- specific primers enables direct observation of a large fraction of the somatic rearrangements present in the immune cell receptor gene repertoire within the subject.
[0052] A multiplex PCR system can use at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, or 25, and in certain embodiments, at least 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, or 39, and in other embodiments at least 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 65, 70, 75, 80, 85, or more forward primers, in which each forward primer specifically hybridizes to (i.e., is complementary to) a sequence corresponding to a V region segment. The multiplex PCR system also uses at least 2, 3, 4, 5, 6, or 7, and in certain embodiments, at least 8, 9, 10, 11 , 12 or 13 reverse primers, or at least 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, or 25 or more reverse primers, in which each reverse primer specifically hybridizes to or is complementary to a sequence corresponding to a J region segment. Various combinations of V and J segment primers can be used to amplify the full diversity of TCR sequences in the immune cell receptor gene repertoire within the subject.
[0053] Further details on multiplex PCR system, including primer oligonucleotide sequences for amplifying TCR sequences are described in Robins et al., 2009 Blood 114, 4099; Robins et al., 2010 Sci. Translat. Med. 2:47ra64; Robins et al., 2011 J. Immunol. Meth. doi:10.1016 / j.jim.2011.09. 001 ; Sherwood et al. 201 1 Sci. Translat. Med. 3:90ra61 ; US2012 / 0058902, US2010 / 033057, WO / 2010 / 151416, WO / 2011 / 106738, US
[0054] 2015 / 0299785, WO2012 / 027503, US2013 / 0288237, U.S. Pat. No. 9,181 ,590, U.S. Pat. No. 9,181 ,591 , US2013 / 0253842, WO2013 / 188831 , which is each incorporated herein by reference in its entirety.
[0055] Oligonucleotides or polynucleotides that are capable of specifically hybridizing or annealing to a target nucleic acid sequence by nucleotide base complementarity can do so under moderate to high stringency conditions. In one embodiment, suitable moderate to high stringency conditions for specific PCR amplification of a target nucleic acid sequence can be between 25 and 80 PCR cycles, with each cycle including a denaturation step (e.g., about 10-30 seconds (s) at greater than about 95°C), an annealing step (e.g., about 10-30s at about 60-68°C), and an extension step (e.g., about 10-60s at about 60-72°C), optionally according to certain embodiments with the annealing and extension steps being combined to provide a two-step PCR. As would be recognized by the skilled person, other PCR reagents can be added or changed in the PCR reaction to increase specificity of primer annealing and amplification, such as altering the magnesium concentration, optionally adding DMSO, and / or the use of blocked primers, modified nucleotides, peptide-nucleic acids, and the like.
[0056] A primer may be a single-stranded DNA. The appropriate length of a primer depends on the intended use of the primer but typically ranges from 6 to 50 nucleotides, or in certain embodiments, from 15-35 nucleotides in length. Short primer molecules generally require cooler temperatures to form sufficiently stable hybrid complexes with the template. A primer need not reflect the exact sequence of the template nucleic acid, but must be sufficiently complementary to hybridize with the template. The design of suitable primers for the amplification of a given target sequence is well known in the art and described in the literature cited herein.
[0057] V- and J-segment primers are used to produce a plurality of amplicons from the multiplex PGR reaction. In certain embodiments, the amplicons range in size from 10, 20, 30, 40, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800 or more nucleotides in length. In certain embodiments, the amplicons have a size between 20-600, 50-600, 20-400, or 50-400 nucleotides in length.
[0058] According to non-limiting theory, these embodiments exploit current understanding in the art (also described above) that once a T lymphocyte has rearranged its TCR-encoding genes, its progeny cells possess the same immune cell receptor-encoding gene rearrangement, thus giving rise to a clonal population (clones) that can be uniquely identified by the presence therein of rearranged (e.g., CDR3-encoding) V- and J-gene segments that can be amplified by a specific pairwise combination of V- and J-specific oligonucleotide primers as herein disclosed.
[0059] The V segment primers and J segment primers will preferably each include a second sequence at the 5'-end of the primer that is not complementary to the target V or J segment. The second sequence can comprise an oligonucleotide having a sequence that is selected from (i) a universal adaptor oligonucleotide sequence, and (ii) a sequencing platform-specific oligonucleotide sequence that is linked to and positioned 5' to a first universal adaptor oligonucleotide sequence. Examples of universal adaptor oligonucleotide sequences can be pGEX forward and pGEX reverse adaptor sequences.
[0060] The resulting amplicons using the V-segment and J-segment primers described above include amplified V and J segments and the universal adaptor oligonucleotide sequences. The universal adaptor sequence can be complementary to an oligonucleotide sequence found in a tailing primer. Tailing primers can be used in a second PGR reaction to generate a second set of amplicons. In some embodiments, tailing primers can have the general formula (I):
[0061] 5'-P-S— B-U-3' (I), where P comprises a sequencing platform-specific oligonucleotide, where S comprises a sequencing platform tag-containing oligonucleotide sequence; where B comprises an oligonucleotide barcode sequence and where the oligonucleotide barcode sequence can be used to identify a sample source, and where U comprises a sequence that is complementary to the universal adaptor oligonucleotide sequence or is the same as the universal adaptor oligonucleotide sequence.
[0062] Additional description about universal adaptor oligonucleotide sequences, barcodes, and tailing primers are found in WO2013 / 188831 , which is incorporated by reference in its entirety.
[0063] Sequencing may be performed using any of a variety of available high throughput single molecule sequencing machines and systems. Illustrative sequence systems include sequence-by-synthesis systems, such as the Illumina Genome Analyzer and associated instruments (Illumina HiSeq) (Illumina, Inc., San Diego, Calif.), Helicos Genetic Analysis System (Helicos BioSciences Corp., Cambridge, Mass.), Pacific Biosciences PacBio RS (Pacific Biosciences, Menlo Park, Calif.), a MinlON™, GridlONx5TM, PromethlON™, or SmidglON™ nanopore-based sequencing system, available from Oxford Nanopore Technologies, or other systems having similar capabilities.
[0064] In certain embodiments, sequencing is achieved using a set of sequencing platformspecific oligonucleotides that hybridize to a defined region within the amplified DNA molecules. The sequencing platform-specific oligonucleotides are designed to sequence amplicons, such that the V- and J-encoding gene segments can be uniquely identified by the sequences that are generated. See, e.g., US2012 / 0058902; US2010 / 033057; WO2011 / 106738; US2015 / 0299785; or WO2012 / 027503, which is each incorporated by reference in its entirety.
[0065] In some embodiments, the raw sequence data is preprocessed to remove errors in the primary sequence of each read and to compress the data. A nearest neighbor algorithm can be used to collapse the data into unique sequences by merging closely related sequences, to remove both PGR and sequencing errors. See, e.g., US2012 / 0058902; US2010 / 033057; WO2011 / 106738; US2015 / 0299785; or WO2012 / 027503, which is each incorporated by reference in its entirety.
[0066] Sequencing the multiplicity of amplified rearranged TCR|3 CDR3-encoding region DNA molecules by high-throughput sequencing (HTS) can be used to produce a TOR clonotype profile comprising at least 10,000 TOR clonotype sequences of 20 to 400 nucleotides in length.
[0067] Amplification Bias Control
[0068] Multiplex PCR assays can result in a bias in the total numbers of amplicons produced from a sample, given that certain primer sets may be more efficient in amplification than others. To overcome the problem of such biased utilization of subpopulations of amplification primers, methods can be used that provide a template composition for standardizing the amplification efficiencies of the members of an oligonucleotide primer set, where the primer set is capable of amplifying rearranged DNA encoding a plurality of TCRs in a biological sample that comprises DNA from lymphoid cells.
[0069] To that end, a template composition is used to standardize the various amplification efficiencies of the primer sets. The template composition can comprise a plurality of diverse template oligonucleotides of general formula (II):
[0070] 5'-U1-B1-V-B2-R-J-B3-U2-3' (II)
[0071] The constituent template oligonucleotides are diverse with respect to the nucleotide sequences of the individual template oligonucleotides. The individual template oligonucleotides can vary in nucleotide sequence considerably from one another as a function of significant sequence variability among the large number of possible TCR variable (V) and joining (J) region polynucleotides. Sequences of individual template oligonucleotide species can also vary from one another as a function of sequence differences in U1 , U2, B (B1 , B2 and B3) and R oligonucleotides that are included in a particular template within the diverse plurality of templates.
[0072] V is a polynucleotide comprising at least 20, 30, 60, 90, 120, 150, 180, or 210, and not more than 1000, 900, 800, 700, 600 or 500 contiguous nucleotides of an adaptive immune receptor variable (V) region encoding gene sequence, or the complement thereof, and in each of the plurality of template oligonucleotide sequences V comprises a unique oligonucleotide sequence.
[0073] J is a polynucleotide comprising at least 15-30, 31 -60, 61-90, 91 -120, or 120-150, and not more than 600, 500, 400, 300 or 200 contiguous nucleotides of an adaptive immune receptor joining (J) region encoding gene sequence, or the complement thereof, and in each of the plurality of template oligonucleotide sequences J comprises a unique oligonucleotide sequence.
[0074] U1 and U2 can be each either nothing or each comprise an oligonucleotide having, independently, a sequence that is selected from (i) a universal adaptor oligonucleotide sequence, and (ii) a sequencing platform-specific oligonucleotide sequence that is linked to and positioned 5' to the universal adaptor oligonucleotide sequence.
[0075] B1 , B2 and B3 can be each either nothing or each comprise an oligonucleotide B that comprises a first and a second oligonucleotide barcode sequence of 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 contiguous nucleotides (including all integer values therebetween), wherein in each of the plurality of template oligonucleotide sequences B comprises a unique oligonucleotide sequence in which (i) the first barcode sequence uniquely identifies the unique V oligonucleotide sequence of the template oligonucleotide and (ii) the second barcode sequence uniquely identifies the unique J oligonucleotide sequence of the template oligonucleotide.
[0076] R can be either nothing or comprises a restriction enzyme recognition site that comprises an oligonucleotide sequence that is absent from V, J, U1 , U2, B1 , B2 and B3.
[0077] Methods are used with the template composition for determining non-uniform nucleic acid amplification potential among members of a set of oligonucleotide amplification primers that are capable of amplifying productively rearranged DNA encoding one or a plurality of TCRs in a biological sample that comprises DNA from lymphoid cells of a subject. The method can include the steps of: (a) amplifying DNA of a template composition for standardizing amplification efficiency of an oligonucleotide primer set in a multiplex polymerase chain reaction (PGR) that comprises: (i) the template composition (II) described above, wherein each template oligonucleotide in the plurality of template oligonucleotides is present in a substantially equimolar amount; (ii) an oligonucleotide amplification primer set that is capable of amplifying productively rearranged DNA encoding one or a plurality of TCRs in a biological sample that comprises DNA from lymphoid cells of a subject.
[0078] The primer set can include: (1) in substantially equimolar amounts, a plurality of V- segment oligonucleotide primers that are each independently capable of specifically hybridizing to at least one polynucleotide encoding a TOR V-region polypeptide or to the complement thereof, wherein each V-segment primer comprises a nucleotide sequence of at least 15 contiguous nucleotides that is complementary to at least one functional TOR V region-encoding gene segment and wherein the plurality of V-segment primers specifically hybridize to substantially all functional TOR V region-encoding gene segments that are present in the template composition, and (2) in substantially equimolar amounts, a plurality of J-segment oligonucleotide primers that are each independently capable of specifically hybridizing to at least one polynucleotide encoding a TOR J-region polypeptide or to the complement thereof, wherein each J-segment primer comprises a nucleotide sequence of at least 15 contiguous nucleotides that is complementary to at least one functional TCR J region-encoding gene segment and wherein the plurality of J-segment primers specifically hybridize to substantially all functional TCR J region-encoding gene segments that are present in the template composition.
[0079] The V-segment and J-segment oligonucleotide primers are capable of promoting amplification in said multiplex polymerase chain reaction (PCR) of substantially all template oligonucleotides in the template composition to produce a multiplicity of amplified template DNA molecules, said multiplicity of amplified template DNA molecules being sufficient to quantify diversity of the template oligonucleotides in the template composition, and wherein each amplified template DNA molecule in the multiplicity of amplified template DNA molecules is less than 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 90, 80 or 70 nucleotides in length.
[0080] Methods for determining non-uniform nucleic acid amplification potential may further include: (b) sequencing all or a sufficient portion of each of said multiplicity of amplified template DNA molecules to determine, for each unique template DNA molecule in said multiplicity of amplified template DNA molecules, (I) a template-specific oligonucleotide DNA sequence and (ii) a relative frequency of occurrence of the template oligonucleotide; and (c) comparing the relative frequency of occurrence for each unique template DNA sequence from said template composition, wherein a non-uniform frequency of occurrence for one or more template DNA sequences indicates non-uniform nucleic acid amplification potential among members of the set of oligonucleotide amplification primers.
[0081] Further details concerning the aforementioned bias control methods are provided in US2013 / 0253842, U.S. Pat. No. 9,150,905, US2015 / 0203897, and WO2013 / 169957, which are incorporated by reference in their entireties.
[0082] PCR Template Abundance Estimation
[0083] To estimate the average read coverage per input template in the multiplex PCR and sequencing approach, a set of synthetic TCR templates (as described above) can be used, comprising each combination of V.beta. and J. beta, gene segments. These synthetic molecules can be those described in general formula (II) above, and in US2013 / 0253842, U.S. Pat. No. 9,150,905, US2015 / 0203897, and WO2013 / 169957, which are incorporated by reference in their entireties.
[0084] These synthetic molecules can be included in each PCR reaction at very low concentration so that only some of the synthetic templates are observed. Using the known concentration of the synthetic template pool, the relationship between the number of observed unique synthetic molecules and the total number of synthetic molecules added to reaction can be simulated (this is very nearly one-to-one at the low concentrations that were used). The synthetic molecules allow calculation for each PCR reaction the mean number of sequencing reads obtained per molecule of PCR template, and an estimation of the number of T cells or B cells in the input material bearing each unique TCR rearrangement or Ig rearrangement, respectively.
[0085] Provided in Table 1 are multiple sclerosis-associated TCRp sequences that distinguish multiple sclerosis patients from controls. Table 1 - MS-Associated TCRB Sequences
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] In Table 1 herein, the amino acid sequence represents the TCR[3 CDR3 segment of the TCR, while V##-## or J##-## refers to a standard two level coding system [family]-[gene] for a particular part of the human genome that can be used as part of a TCR rearrangement formed in response to antigen exposure. The first two digits reflect a member of a family and the second two digits reflect a particular gene from within that family if present. So, by way of example, TCRBV06 would indicate a match of sequence to a specific family of variable (V) chain sequences where TCRBV06-05 indicates a more precise identification to a specific gene from within a family of variable chain sequences. Identities of these V- and J- gene sequences can be found at the international
[0099] ImMunoGeneTics information system (www.imgt.org), including at www.imgt.org / download / V-QUEST / IMGT_V-QUEST_reference_directory / Homo_sapiens / TR / TRBV.fasta. THERAPEUTIC METHODS
[0100] Also provided by the present disclosure are therapeutic methods. According to some embodiments, provided are methods comprising administering a multiple sclerosis (MS) therapy to a subject identified as comprising T cells that express a T cell receptor p chain (TCRP) comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1-997. In certain embodiments, the methods comprise administering an MS therapy to a subject identified as comprising T cells that express two or more (e.g., two or more unique) TCRp comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997.
[0101] In some instances, the methods comprise administering an MS therapy to a subject identified using one or any combination of models / classifiers described elsewhere herein as having MS. In one non-limiting example, the subject is identified using 2 or more models comprising a first model to perform MS classification in samples inferred as carrying the HLA allele DRB1 *15:01 , and a second model to perform MS classification in samples inferred as not carrying the HLA allele DRB1 *15:01. In some instances, the HLA inference of a sample carrying or not carrying the DRB1 *15:01 allele may be based on genotyping, while in other embodiments, the HLA inference of a sample carrying or not carrying the DRB1 *15:01 allele may be performed based on applying a classification model to the TCRp repertoire (see, e.g., U.S. Patent Application Publication No. US 2021 / 0381050 A1 , the disclosure of which is incorporated herein by reference in its entirety for all purposes). One of ordinary skill in the art will appreciate that, with the benefit of the TCRp sequences set forth in SEQ ID Nos:1- 997 described herein, a variety of useful models may be employed to identify the subject as having or not having MS. In one non-limiting example, employed is a model comprising a two-feature logistic regression with features representing the number of MS-associated TCRp sequences determined from the sample and the total number of unique TCRp sequences determined from the sample. In certain embodiments, a model used to identify the subject as having or not having MS may take into account the number of unique MS- associated TCRp sequences that are present in the TCRp sequences determined from the sample, e.g., where the greater the number of unique MS-associated TCRp sequences, the more likely the model is to classify the subject as having MS. According to some embodiments, the number of unique MS-associated TCRp sequences is not a feature utilized by the model to classify the subject. In some instances, a model may take into account both the number of MS-associated TCRp sequences determined from the sample and the total number of unique TCRp sequences determined from the sample. In one non-limiting example, a logistic growth curve of best fit and an accompanying error model may be derived based on samples without MS diagnosis, and this fit logistic growth curve and error model may be used to classify samples as having or not having MS. In one non-limiting example of such a logistic growth curve, the logarithm of the total number of unique TCR|3 sequences may form the domain of the logistic growth function, and the number of unique MS-associated TCRp sequences may form the range of the logistic growth function. In one non-limiting case of an accompanying error model, a binomial error model can be derived to accompany the fit logistic growth curve, such that if a particular point on the curve assumes a fraction F of the upper bound on values taken by the curve, the accompanying standard deviation of this fraction F is equal to a constant times the square root of the product of F and the quantity one minus F, as stated in Greissl et al. (2021 ) medRxiv doi.org / 10.1101 / 2021.07.30.21261353. In one non-limiting example of how the logistic growth curve and error model may be used to classify samples as having or not having MS, each sample may be given a score equal to the number of standard deviations greater than the value on the logistic growth curve, with positive scores for points above the curve and negative scores for points below the curve, following which samples exceeding a threshold level of this score can be classified as having MS, while those samples not exceeding this threshold level are classified as not having MS. As demonstrated in the Experimental section below, such a model exhibits high specificity and sensitivity for MS patients inferred as carrying the DRB1 *15:01 allele. In certain embodiments, the presence and / or frequency of one or more particular unique MS-associated TCR|3 sequences is a feature(s) used by the model to classify the subject. For example, the presence and / or frequency of one or more particular unique MS-associated TCRp sequences may be given relatively greater weight when classifying the subject as compared to the presence and / or frequency of one or more other unique MS-associated TCRp sequences. According to some embodiments, when a classification model weighs particular unique MS-associated TCR|3 sequences differently than other unique MS-associated TCRp sequences, the model may use convergent recombination to weigh the sequences differently. Different T cells can show convergent recombination where unique DNA sequences were formed in the recombination for a first T cell, a second T cell, a third T cell, etc., but where each leads to the same protein (CDR3 + V-gene + J-gene) which is diagnostic for high likelihood of MS. This convergent recombination may be more likely for certain MS-associated TCR|3 sequences than others, and the model may take into account these aspects of the signal reflective of the interpretable biology of immune response. Accordingly, in some embodiments, sequences may be given differential weight based on convergent recombination.
[0102] Any suitable MS therapy may be administered to a subject identified as described above. MS therapies are known and may vary depending upon the age of the patient, stage of the disease, and / or the like. In certain embodiments, the MS therapy comprises administering to the subject a therapeutically effective amount of one or more agents approved by the United States Food and Drug Administration (FDA) and / or the European Medicines Agency (EMA) for treatment of MS at the time of performing the methods of the present disclosure. In some instances, the MS therapy comprises administering to the subject a therapeutically effective amount of natalizumab-sztn (Tyruko), natalizumab (Tysabri), ublituximab-xiiy (Briumvi), ublituximab, ponesimod (Ponvory), ofatumumab (Kesimpta), monomethyl fumarate (Bafiertam), ozanimod (Zeposia), diroximel fumarate (Vumerity), cladribine (Mavenclad), siponimod (Mayzent), ocrelizumab (Ocrevus), daclizumab (Zinbryta), alemtuzumab (Lemtrada), peginterferon beta-1 a (Plegridy), dimethyl fumarate (Tefcidera), diroximel fumarate, teriflunomide (Aubagio), fingolimod (Gilenya), laquinimod, ozanimod, glatiramer acetate, interferon beta 1a, interferon beta 1 b, PIPE-307, or any combination thereof.
[0103] According to some embodiments, the methods are effective in treating the MS of the subject. By “treat” or “treatment” is meant at least an amelioration of the symptoms associated with the MS. Such symptoms may include one or more of CNS inflammation, cortical demyelination, cerebral lesions, cortical lesions, juxtacortical lesions, periventricular lesions, brainstem lesions, cerebellar lesions, spinal cord lesions, optic nerve lesions, bladder and / or bowel dysfunction, diplopia, dysphagia, altered sensation or weakness of the face, ataxia, blurry vision with painful eye movements, asymmetric limb weakness, spasticity, impaired mobility, impaired motor dexterity, and cognitive impairment. Amelioration is used in a broad sense to refer to at least a reduction in the magnitude of a parameter, e.g., symptom, associated with the MS being treated. As such, treatment also includes situations where the MS, or at least symptoms associated therewith, are completely inhibited, e.g., prevented from happening, or stopped, e.g., terminated, such that the individual no longer suffers from the MS, or at least the symptoms that characterize the MS.
[0104] The subject may be identified (e.g., by or subsequent to the methods of the present disclosure) as having Relapsing Remitting MS (RRMS - characterized by flare ups or exacerbations of the neurological symptoms of MS, also known as relapses, followed by periods of recovery or remission), Secondary Progressive MS (SPMS - characterized by a reduction in relapses and a progressive worsening of symptoms (accumulation of disability) over time, with no obvious signs of remission), Primary Progressive MS (PPMS - characterized by a progressive worsening of symptoms and disability from the beginning, without periods of recovery or remission), or clinically isolated syndrome (CIS).
[0105] Dosing may be dependent on severity and responsiveness of the MS to be treated. Optimal dosing schedules can be calculated from measurements of drug accumulation in the body of the individual. The administering physician can determine optimum dosages, dosing methodologies and repetition rates. Optimum dosages may vary depending on the relative potency of individual therapeutic agents, and can generally be estimated based on EC50S found to be effective in in vitro and in vivo animal models, etc. In general, dosage is from about 0.01 pg to about 100 g per kg of body weight, and may be given once or more daily, weekly, monthly or yearly. In some instances, the dosage is from about 1 pg / kg to 100 mg / kg or more, depending on the factors mentioned above. The treating physician can estimate repetition rates for dosing based on measured residence times and concentrations of the therapeutic agent in bodily fluids or tissues. Following successful treatment, it may be desirable to have the subject undergo maintenance therapy to prevent the recurrence of the disease state, where the therapeutic agent is administered in maintenance doses, ranging from about 0.01 pg to about 100 g per kg of body weight, once or more daily, to once every several months, once every six months, once every year, or at any other suitable frequency.
[0106] The therapeutic methods of the present disclosure may include administering a single type of therapeutic agent to the subject, or may include administering two or more types of therapeutic agents to the subject separately or by administration of a cocktail of different therapeutic agents. For example, in certain embodiments, two or more therapeutic agents that find use in treating MS described elsewhere herein may be administered to the subject, e.g., two or more, three or more, four or more, or five or more of such therapeutic agents.
[0107] The one or more therapeutic agents may be administered to the subject using any available method and route suitable for drug delivery, including in vivo and ex vivo methods, as well as systemic and localized routes of administration. Conventional and pharmaceutically acceptable routes of administration include oral and parenteral routes of administration. Parenteral routes of administration of interest include, but are not limited to, injection (e.g., intravenous, intra-arterial, local, subcutaneous, or intramuscular injection), intranasal, intra-tracheal, intradermal, topical application, ocular, nasal, and other parenteral routes of administration. Routes of administration may be combined, if desired, or adjusted depending upon the therapeutic agent and / or the desired effect. The therapeutic agent may be administered in a single dose or in multiple doses. In some embodiments, the therapeutic agent is administered intravenously. In some embodiments, the therapeutic agent is administered by injection, e.g., for systemic delivery (e.g., intravenous infusion) or to a local site.
[0108] A “therapeutically effective amount” or “efficacious amount” refers to the amount of a therapeutic agent that, when administered to a mammal or other subject for treating a disease, is sufficient to effect such treatment for the disease. The “therapeutically effective amount” will vary depending on the therapeutic agent, the disease and its severity and the age, weight, etc., of the subject to be treated. In some embodiments, the MS therapy is an adoptive cell therapy. Non-limiting examples of adoptive cell therapies include those involving administering to the subject an effective amount of recombinant cells (e.g., recombinant immune cells such as T cells) that express a T cell receptor comprising an MS-associated TCRp CDR3 sequence identified as being present in TCRs expressed by T cells in the subject. Similar to CAR therapies, TCR therapies modify the patient’s T lymphocytes ex vivo before being administered back into the patient’s body. The target antigens identified by CAR-T cell therapy are all cell surface proteins, while TCR-T cell therapy can recognize intracellular antigen fragments presented by MHC molecules, so TCR-T cell therapy has a wider range of targets. Approaches for TCR therapy are known and described in, e.g., Zhang et al. (2019) Technol Cancer Res Treat. 18:1533033819831068; Govers et al. (2010) Trends in Molecular Medicine 16(2):77-87; Zhao et al. (2019) Front. Immunol. 10:2250.
[0109] Non-limiting examples of adoptive cell therapies also include those involving administering to the subject an effective amount of recombinant regulatory T cells (Tregs) that express a T cell receptor comprising an MS-associated TCRp CDR3 sequence identified as being present in TCRs expressed by T cells in the subject, e.g., a T cell receptor comprising an MS-associated TCRp CDR3 sequence of one of SEQ ID NOs:1 -997. Accordingly, aspects of the present disclosure also include methods of preventing or inhibiting a multiple sclerosis (MS)-associated adaptive immune response in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the immune regulatory cells of the present disclosure. Such immune regulatory cells include, but are not limited to, those that express a T cell receptor (TCR) comprising a TCRp chain comprising a TCRp CDR3 sequence of one of SEQ ID NOs:1 -997. In some embodiments, the TCRp CDR3 sequence is CAISESWAGGTDTQYF (SEQ ID NO: 458) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO: 998). According to certain embodiments, the TCRp CDR3 sequence is CAISESWTGGSDTQYF (SEQ ID NO. 255) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRPNTGTASKLTF (SEQ ID NO. 999). In some instances, the TCRp CDR3 sequence is CAISESWSGGTDTQYF (SEQ ID NO. 877) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRQNTGTASKLTF (SEQ ID NO. 1000). In some embodiments, the TCRp CDR3 sequence is CAISEGWTGNTDTQYF (SEQ ID NO. 996) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO. 1001). According to certain embodiments, the TCRp CDR3 sequence is CAISESSGSTDTQYF (SEQ ID NO. 997) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRVATGTASKLTF (SEQ ID NO. 1002). In some instances, the TCRp CDR3 sequence is CASSDQGGGYEQYF (SEQ ID NO. 33) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CAASRDNQGGKLIF (SEQ ID NO. 1003). In some embodiments, the immune regulatory cells are regulatory T cells (Tregs). Methods of treating multiple sclerosis in a subject in need thereof are also provided, such methods comprising administering to the subject a therapeutically effective amount of the immune regulatory cells of the present disclosure.
[0110] Nucleic acids that encode a T cell receptor 0 chain comprising a TCR0 CDR3 sequence set forth in SEQ ID Nos:1-997 are also provided. For example, in certain embodiments, provided is an expression vector comprising a nucleic acid sequence that encodes a T cell receptor 0 chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997 operably linked to a nucleic acid expression control sequence. A “vector” is capable of transferring nucleic acid sequences to target cells (e.g., viral vectors, non-viral vectors, particulate carriers, and liposomes). Typically, “vector construct,” “expression vector,” and “gene transfer vector,” mean any nucleic acid construct capable of directing the expression of a nucleic acid of interest and which can transfer nucleic acid sequences to target cells. Thus, the term includes cloning and expression vehicles, as well as viral vectors.
[0111] Because of the knowledge of the codons corresponding to the various amino acids, availability of an amino acid sequence of a polypeptide of interest provides a description of all the polynucleotides capable of encoding the polypeptide of interest. The degeneracy of the genetic code, where the same amino acids are encoded by alternative or synonymous codons allows an extremely large number of nucleic acids to be made, all of which encode the enzymes disclosed herein. Thus, having identified a particular amino acid sequence, those of ordinary skill in the art could make any number of different nucleic acids by simply modifying the sequence of one or more codons in a way which does not change the amino acid sequence of the polypeptide of interest. In this regard, the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based upon the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein, including the amino acid sequences of SEQ ID NOs. 1 -1003.
[0112] The nucleotide sequences of the nucleic acids of the present disclosure may be codon-optimized. “Codon-optimized” refers to changes in the codons of the polynucleotide encoding a polypeptide to those preferentially used in a particular organism such that the encoded protein is efficiently expressed in the organism of interest. Although the genetic code is degenerate in that most amino acids are represented by several codons, called “synonyms” or “synonymous” codons, it is well known that codon usage by particular organisms is nonrandom and biased towards particular codon triplets. This codon usage bias may be higher in reference to a given gene, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and the aggregate protein coding regions of an organism's genome. In some embodiments, a nucleic acid of the present disclosure encoding a polypeptide may be codon-optimized for optimal production from the host organism selected for expression, e.g., human cells, such as human immune cells (e.g., human T cells such a human regulatory T cells (human Tregs)).
[0113] In order to express a desired T cell receptor p chain comprising a TCR CDR3 sequence set forth in SEQ ID Nos:1 -997, a nucleotide sequence encoding the T cell receptor P chain can be inserted into an appropriate vector, e.g., using recombinant DNA techniques known in the art. Exemplary viral vectors include, without limitation, retrovirus (including lentivirus), adenovirus, adeno-associated virus, herpesvirus (e.g., herpes simplex virus), poxvirus, papillomavirus, and papovavirus (e.g., SV40). Illustrative examples of expression vectors include, but are not limited to pCIneo vectors (Promega) for expression in mammalian cells; pLenti4 / V 5-DEST™, pLenti6 / V 5- DEST™, murine stem cell virus (MSCV), MSGV, moloney murine leukemia virus (MMLV), and pLenti6.2 / V5-GW / lacZ (Invitrogen) for lentivirus-mediated gene transfer and expression in mammalian cells. In certain embodiments, a nucleic acid sequence encoding the T cell receptor p chain may be ligated into any such expression vectors for the expression of the T cell receptor p chain in mammalian cells.
[0114] Expression control sequences, control elements, or regulatory sequences present in an expression vector are those non-translated regions of the vector - origin of replication, selection cassettes, promoters, enhancers, translation initiation signals (Shine Dalgarno sequence or Kozak sequence), introns, a polyadenylation sequence, 5' and 3' untranslated regions, and / or the like - which interact with host cellular proteins to carry out transcription and translation. Such elements may vary in their strength and specificity. Depending on the vector system and host utilized, any number of suitable transcription and translation elements, including ubiquitous promoters and inducible promoters may be used.
[0115] Components of the expression vector are operably linked such that they are in a relationship permitting them to function in their intended manner. In some embodiments, the term refers to a functional linkage between a nucleic acid expression control sequence (such as a promoter, and / or enhancer) and a second polynucleotide sequence, e.g., a nucleic acid encoding the T cell receptor p chain, where the expression control sequence directs transcription of the nucleic acid encoding the T cell receptor p chain.
[0116] In some embodiments, the expression vector is an episomal vector or a vector that is maintained extrachromosomally. As used herein, the term “episomal” refers to a vector that is able to replicate without integration into the host cell’s chromosomal DNA and without gradual loss from a dividing host cell also meaning that said vector replicates extrachromosomally or episomally. Such a vector may be engineered to harbor the sequence coding for the origin of DNA replication or “ori” from an alpha, beta, or gamma herpesvirus, an adenovirus, SV40, a bovine papilloma virus, a yeast, or the like. The host cell may include a viral replication transactivator protein that activates the replication. Alpha herpes viruses have a relatively short reproductive cycle, variable host range, efficiently destroy infected cells and establish latent infections primarily in sensory ganglia. Illustrative examples of alpha herpes viruses include HSV 1 , HSV 2, and VZV. Beta herpesviruses have long reproductive cycles and a restricted host range. Infected cells often enlarge. Non-limiting examples of beta herpes viruses include CMV, HHV-6 and HHV-7. Gamma-herpesviruses are specific for either T or B lymphocytes, and latency is often demonstrated in lymphoid tissue. Illustrative examples of gamma herpes viruses include EBV and HHV-8.
[0117] Also provided are recombinant cells that comprise any of the expression vectors of the present disclosure comprising a nucleic acid that encodes a T cell receptor p chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997. In certain aspects, provided are cells that express a TOR comprising a T cell receptor p chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997 on the surface of the cell.
[0118] In some embodiments, the cells of the present disclosure are eukaryotic cells. Eukaryotic cells of interest include, but are not limited to, yeast cells, insect cells, mammalian cells, and the like. Mammalian cells of interest include, e.g., murine cells, non-human primate cells, human cells, and the like.
[0119] “Recombinant host cells,” “host cells,” “cells,” “cell lines,” “cell cultures,” and other such terms denoting microorganisms or higher eukaryotic cell lines, refer to cells which can be, or have been, used as recipients for a recombinant vector or other transferred DNA, and include the progeny of the cell which has been transfected. Host cells may be cultured as unicellular or multicellular entities (e.g., tissue, organs, or organoids) including an expression vector of the present disclosure.
[0120] In some embodiments, the cells provided herein are immune cells. Non-limiting examples of recombinant immune cells which may include any of the expression vectors of the present disclosure include T cells, B cells, natural killer (NK) cells, macrophages, monocytes, neutrophils, dendritic cells, mast cells, basophils, and eosinophils. In some embodiments, the immune cell is a T cell. Examples of T cells include naive T cells (TN), cytotoxic T cells (TCTL), memory T cells (TMEM), T memory stem cells (TSCM), central memory T cells (TCM), effector memory T cells (TEM), tissue resident memory T cells (TRM), effector T cells (TEFF), regulatory T cells (TREGS), helper T cells (TH, TH 1 , TH2, TH17) CD4+ T cells, CD8+ T cells, virus-specific T cells, alpha beta T cells (Tap), and gamma delta T cells (TYB). In another aspect, the cells provided herein comprise stem cells, e.g., an embryonic stem cell or an adult stem cell.
[0121] Also provided are methods of making the cells of the present disclosure. In some embodiments, such methods include transfecting or transducing cells with a nucleic acid or expression vector of the present disclosure, e.g., an expression vector comprising a nucleic acid that encodes a T cell receptor p chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos: 1-997. The term “transfection” or “transduction” is used to refer to the introduction of foreign DNA into a cell. A cell has been “transfected” when exogenous DNA has been introduced inside the cell membrane. A number of transfection techniques are generally known in the art. See, e.g., Sambrook et al. (2001 ) Molecular Cloning, a laboratory manual, 3rdedition, Cold Spring Harbor Laboratories, New York, Davis et al. (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw- Hill, and Chu et al. (1981 ) Gene 13:197. Such techniques can be used to introduce one or more exogenous DNA moieties into suitable host cells. The term refers to both stable and transient uptake of the genetic material.
[0122] In some embodiments, a cell of the present disclosure is produced by transfecting the cell with a viral vector encoding the T cell receptor p chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997. In some embodiments, such methods include activating a population of T cells (e.g., T cells obtained from an individual to whom a TCR T cell therapy will be administered), stimulating the population of T cells to proliferate, and transducing the T cell with a viral vector encoding the T cell receptor chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997. In some embodiments, the T cells are transduced with a retroviral vector, e.g., a gamma retroviral vector or a lentiviral vector, encoding the T cell receptor p chain comprising a TCR CDR3 sequence set forth in SEQ ID Nos:1 -997. In some embodiments, the T cells are transduced with a lentiviral vector encoding the T cell receptor p chain comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1- 997.
[0123] Cells of the present disclosure may be autologous / autogeneic (“self”) or non- autologous (“non-self,” e.g., allogeneic, syngeneic or xenogeneic). “Autologous” as used herein, refers to cells from the same individual. “Allogeneic” as used herein refers to cells of the same species that differ genetically from the cell in comparison. “Syngeneic,” as used herein, refers to cells of a different individual that are genetically identical to the cell in comparison. In some embodiments, the cells are T cells obtained from a mammal. In some embodiments, the mammal is a primate. In some embodiments, the primate is a human.
[0124] T cells may be obtained from a number of sources including, but not limited to, peripheral blood, peripheral blood mononuclear cells, bone marrow, lymph node tissue, cord blood, thymus tissue, tissue from a site of infection, ascites, pleural effusion, spleen tissue, and tumors. In certain embodiments, T cells can be obtained from a unit of blood collected from an individual using any number of known techniques such as sedimentation, e.g., FICOLL™ separation.
[0125] In some embodiments, an isolated or purified population of T cells is used. In some embodiments, TCTL and TH lymphocytes are purified from PBMCs. In some embodiments, the TCTL and TH lymphocytes are sorted into naive (TN), memory (TMEM), and effector (TEFF) T cell subpopulations either before or after activation, expansion, and / or genetic modification. Suitable approaches for such sorting are known and include, e.g., magnetic-activated cell sorting (MACS), where TN are CD45RA+ CD62L+ CD95-; TSCM are CD45RA+ CD62L+ CD95+; TCM are CD45RO+ CD62L+ CD95+; and TEM are CD45RO+ CD62L- CD95+. An example approach for such sorting is described in Wang et al. (2016) Blood 127(24):2980- 90.
[0126] A specific subpopulation of T cells expressing one or more of the following markers: CD3, CD4, CD8, CD28, CD45RA, CD45RO, CD62, CD127, and HLA-DR can be further isolated by positive or negative selection techniques. In some embodiments, a specific subpopulation of T cells, expressing one or more of the markers selected from the group consisting of CD62L, CCR7, CD28, CD27, CD122, CD127, CD197; or CD38 or CD62L, CD127, CD197, and CD38, is further isolated by positive or negative selection techniques. In some embodiments, the manufactured T cell compositions do not express one or more of the following markers: CD57, CD244, CD 160, PD-1 , CTLA4, TIM3, and LAG3. In some embodiments, the manufactured T cell compositions do not substantially express one or more of the following markers: CD57, CD244, CD 160, PD-1 , CTLA4, TIM3, and LAG3.
[0127] In order to achieve therapeutically effective doses of T cell compositions, the T cells may be subjected to one or more rounds of stimulation, activation and / or expansion. T cells can be activated and expanded generally using methods as described, for example, in U.S. Patents 6,352,694; 6,534,055; 6,905,680; 6,692,964; 5,858,358; 6,887,466; 6,905,681 ; 7,144,575; 7,067,318; 7,172,869; 7,232,566; 7,175,843; 5,883,223; 6,905,874; 6,797,514; and 6,867,041 , each of which is incorporated herein by reference in its entirety for all purposes. In some embodiments, T cells are activated and expanded for about 1 to 21 days, e.g., about 5 to 21 days. In some embodiments, T cells are activated and expanded for about 1 day to about 4 days, about 1 day to about 3 days, about 1 day to about 2 days, about 2 days to about 3 days, about 2 days to about 4 days, about 3 days to about 4 days, or about 1 day, about 2 days, about 3 days, or about 4 days prior to introduction of a nucleic acid (e.g., expression vector) encoding the polypeptide into the T cells.
[0128] In some embodiments, T cells are activated and expanded for about 6 hours, about 12 hours, about 18 hours or about 24 hours prior to introduction of a nucleic acid (e.g., expression vector) encoding the T cell receptor p chain comprising a TCR CDR3 sequence set forth in SEQ ID Nos:1 -997 into the T cells. In some embodiments, T cells are activated at the same time that a nucleic acid (e.g., an expression vector) encoding the T cell receptor P chain is introduced into the T cells.
[0129] In some embodiments, conditions appropriate for T cell culture include an appropriate media (e.g., Minimal Essential Media or RPMI Media 1640 or, X-vivo 15, (Lonza)) and one or more factors necessary for proliferation and viability including, but not limited to serum (e.g., fetal bovine or human serum), interleukin-2 (IL-2), insulin, IFN-y, IL-4, IL-7, IL-21 , GM- CSF, IL-10, IL-12, IL-15, TGFp, and TNF-a or any other additives suitable for the growth of cells known to the skilled artisan. Further illustrative examples of cell culture media include, but are not limited to RPMI 1640, Clicks, AEVI-V, DMEM, MEM, a-MEM, F-12, X-Vivo 15, and X-Vivo 20, Optimizer, with added amino acids, sodium pyruvate, and vitamins, either serum-free or supplemented with an appropriate amount of serum (or plasma) or a defined set of hormones, and / or an amount of cytokine(s) sufficient for the growth and expansion of T cells.
[0130] In some embodiments, the nucleic acid (e.g., an expression vector) encoding the T cell receptor p chain is introduced into the cell (e.g., a T cell) by microinjection, transfection, lipofection, heat-shock, electroporation, transduction, gene gun, microinjection, DEAE- dextran-mediated transfer, and the like. In some embodiments, the nucleic acid (e.g., expression vector) encoding the T cell receptor p chain is introduced into the cell (e.g., a T cell) by AAV transduction. The AAV vector may comprise ITRs from AAV2, and a serotype from any one of AAV1 , AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAV 10. In some embodiments, the AAV vector comprises ITRs from AAV2 and a serotype from AAV6. In some embodiments, the nucleic acid (e.g., expression vector) encoding the T cell receptor p chain is introduced into the cell (e.g., a T cell) by lentiviral transduction. The lentiviral vector backbone may be derived from HIV-1 , HIV-2, visna-maedi virus (VMV) virus, caprine arthritis-encephalitis virus (CAEV), equine infectious anemia virus (EIAV), feline immunodeficiency virus (FIV), bovine immune deficiency virus (BIV), or simian immunodeficiency virus (SIV). The lentiviral vector may be integration competent or an integrase deficient lentiviral vector (TDLV). In one embodiment, IDLV vectors including an HIV-based vector backbone (i.e., HIV cis-acting sequence elements) are employed.
[0131] COMPUTER-READABLE MEDIA AND SYSTEMS
[0132] Also provided by the present disclosure are computer-readable media and systems.
[0133] In certain aspects, provided are one or more computer-readable media having stored thereon one or more TCR CDR3 sequences set forth in SEQ ID Nos:1 -997. The number of TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 stored on the one or more computer- readable media may vary. For example, the one or more computer-readable media may have stored thereon 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 75 or more, 100 or more, 150 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 450 or more, 500 or more, 550 or more, 600 or more, 650 or more, 700 or more, 750 or more, 800 or more, 850 or more, 900 or more, 950 or more, or each of the TCRp CDR3 sequences set forth in SEQ ID Nos:1- 997. When fewer than all of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 are stored on the one or more computer-readable media, the one or more computer-readable media may have stored thereon any desired number (e.g., as set forth above) and combination of TCRp CDR3 sequences set forth in SEQ ID Nos:1-997. In some embodiments, the one or more computer-readable media may have stored thereon 997 or fewer, 950 or fewer, 900 or fewer, 850 or fewer, 800 or fewer, 750 or fewer, 700 or fewer,
[0134] 650 or fewer, 600 or fewer, 550 or fewer, 500 or fewer, 450 or fewer, 400 or fewer, 350 or fewer, 300 or fewer, 250 or fewer, 200 or fewer, 190 or fewer, 180 or fewer, 170 or fewer,
[0135] 160 or fewer, 150 or fewer, 140 or fewer, 130 or fewer, 120 or fewer, 110 or fewer, 100 or fewer, 90 or fewer, 80 or fewer, 70 or fewer, 60 or fewer, 50 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, or 10 or fewer of the TCRp CDR3 sequences set forth in SEQ ID Nos:1- 997, in any desired combination.
[0136] Also provided are systems for assessing TCRp CDR3 sequences. According to some embodiments, provided are systems for assessing TCRp CDR3 sequences, such systems comprising one or more processors and one or more computer-readable media. The one or more computer-readable media comprise instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to assess TCRp CDR3 sequences determined from a sample obtained from a subject (e.g., a subject identified as having MS or suspected of having MS, including but not limited to a subject exhibiting one or more non-specific symptoms consistent with MS) for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997. According to some embodiments, the number of TCRp CDR3 sequences determined from the sample obtained from the subject is from 1 ,000 to 2,000,000. For example, in certain embodiments, the number of determined TCRp CDR3 sequences is 2,000,000 or fewer (e.g., 1 ,500,000 or fewer, 1 ,250,000 or fewer, 1 ,000,000 or fewer, 750,000 or fewer, or 500,000 or fewer), but 1 ,000 or more, 5,000 or more, 10,000 or more, 15,000 or more, 20,000 or more, 25,000 or more, 30,000 or more, 35,000 or more, 40,000 or more, 45,000 or more, 50,000 or more, 55,000 or more, 60,000 or more, 65,000 or more, 70,000 or more, 75,000 or more, 80,000 or more, 85,000 or more, 90,000 or more, 95,000 or more, or 100,000 or more. The number of TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 to which the determined TCRp CDR3 sequences is compared may vary. For example, the determined TCRp CDR3 sequences may be compared to 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, 50 or more, 75 or more, 100 or more, 150 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 450 or more, 500 or more, 550 or more, 600 or more, 650 or more, 700 or more, 750 or more, 800 or more, 850 or more, 900 or more, 950 or more, or each of the TCRp CDR3 sequences set forth in SEQ ID Nos:1- 997. When the determined TCRp CDR3 sequences are compared to fewer than all of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997, the determined TCRp CDR3 sequences may be compared to any desired number (e.g., as set forth above) and combination of TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
[0137] The one or more computer-readable media may further comprise instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to perform one or more additional steps based on the results of the assessing step. For example, if it is determined from the assessing step that none of the TCRp CDR3 sequences set forth in SEQ ID Nos:1-997 are present in the TCRp CDR3 sequences determined from the sample obtained from the subject, then the instructions may further cause the one or more processors to, e.g., identify the subject as not having MS, identify the subject as one who should not be administered an MS therapy, and / or the like. Also, by way of example, if it is determined from the assessing step that one or more (e.g., 2 or more, 3 or more, 4 or more, 5 or more, or 10 or more) of the TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997 are present in the TCRp CDR3 sequences determined from the sample obtained from the subject (e.g., a subject suspected of having MS, including but not limited to a subject exhibiting one or more non-specific symptoms consistent with MS), then the instructions may further cause the one or more processors to, e.g., predict that the subject has MS, diagnose the subject as having MS, identify the subject as one who should be administered an MS therapy, and / or the like.
[0138] In certain embodiments, the one or more computer-readable media may further comprise instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to subject the results of the assessing step to further analysis, such as subjecting the results of the assessing step to a model. One or more of any of the models described elsewhere herein may be employed. For example, the model may take into account the number of unique MS-associated TCRp sequences that are present in the TCRp sequences determined from the sample, e.g., where the greater the number of unique MS-associated TCRp sequences, the more likely the model is to classify the subject as having MS. According to some embodiments, the number of unique MS-associated TCR|3 sequences is not a feature utilized by the model to classify the subject. In some instances, the model may take into account both the number of MS-associated TCRp sequences determined from the sample and the total number of unique TCRp sequences determined from the sample. In one non-limiting example, a logistic growth curve of best fit and an accompanying error model may be derived based on samples without MS diagnosis, and this fit logistic growth curve and error model may be used to classify samples as having or not having MS. In one non-limiting example of such a logistic growth curve, the logarithm of the total number of unique TCRp sequences may form the domain of the logistic growth function, and the number of unique MS-associated TCRp sequences may form the range of the logistic growth function. In one non-limiting case of an accompanying error model, a binomial error model can be derived to accompany the fit logistic growth curve, such that if a particular point on the curve assumes a fraction F of the upper bound on values taken by the curve, the accompanying standard deviation of this fraction F is equal to a constant times the square root of the product of F and the quantity one minus F, as stated in Greissl et al. (2021) medRxiv doi.org / 10.1101 / 2021 .07.30.21261353. In one non-limiting example of how the logistic growth curve and error model may be used to classify samples as having or not having MS, each sample may be given a score equal to the number of standard deviations greater than the value on the logistic growth curve, with positive scores for points above the curve and negative scores for points below the curve, following which samples exceeding a threshold level of this score can be classified as having MS, while those samples not exceeding this threshold level are classified as not having MS. As demonstrated in the Experimental section below, such a model exhibits high specificity and sensitivity for MS patients inferred as carrying the DRB1 *15:01 allele.
[0139] A variety of processor-based systems may be employed to implement the embodiments of the present disclosure. Such systems may include system architecture wherein the components of the system are in electrical communication with each other using a bus. System architecture can include a processing unit (CPU or processor), as well as a cache, that are variously coupled to the system bus. The bus couples various system components including system memory, (e.g., read only memory (ROM) and random access memory (RAM), to the processor.
[0140] System architecture can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of the processor. System architecture can copy data from the memory and / or the storage device to the cache for quick access by the processor. In this way, the cache can provide a performance boost that avoids processor delays while waiting for data. These and other modules can control or be configured to control the processor to perform various actions. Other system memory may be available for use as well. Memory can include multiple different types of memory with different performance characteristics. Processor can include any general purpose processor and a hardware module or software module, such as first, second and third modules stored in the storage device, configured to control the processor as well as a special-purpose processor where software instructions are incorporated into the actual processor design. The processor may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0141] To enable user interaction with the computing system architecture, an input device can represent any number of input mechanisms, such as a microphone for speech, a touch- sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. An output device can also be one or more of a number of output mechanisms. In some instances, multimodal systems can enable a user to provide multiple types of input to communicate with the computing system architecture. A communications interface can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0142] The storage device is typically a non-volatile memory and can be a hard disk or other types of computer-readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs), read only memory (ROM), and hybrids thereof.
[0143] The storage device can include software modules for controlling the processor. Other hardware or software modules are contemplated. The storage device can be connected to the system bus. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as the processor, bus, output device, and so forth, to carry out various functions of the disclosed technology.
[0144] Embodiments within the scope of the present disclosure may also include tangible and / or non-transitory computer-readable storage media or devices for carrying or having computer-executable instructions or data structures stored thereon. Such tangible computer- readable storage devices can be any available device that can be accessed by a general purpose or special purpose computer, including the functional design of any special purpose processor as described above. By way of example, and not limitation, such tangible computer-readable devices can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other device which can be used to carry or store desired program code in the form of computer-executable instructions, data structures, or processor chip design. When information or instructions are provided via a network or another communications connection (either hardwired, wireless, or combination thereof) to a computer, the computer properly views the connection as a computer-readable medium. Thus, any such connection is properly termed a computer- readable medium. Combinations of the above should also be included within the scope of the computer-readable storage devices.
[0145] Computer-executable instructions include, for example, instructions and data which cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Computer-executable instructions also include program modules that are executed by computers in stand-alone or network environments. Generally, program modules include routines, programs, components, data structures, objects, and the functions inherent in the design of specialpurpose processors, etc. that perform tasks or implement abstract data types. Computerexecutable instructions, associated data structures, and program modules represent examples of the program code means for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps.
[0146] Other embodiments of the disclosure may be practiced in network computing environments with many types of computer system configurations, including personal computers, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. Embodiments may also be practiced in distributed computing environments where tasks are performed by local and remote processing devices that are linked (either by hardwired links, wireless links, or by a combination thereof) through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices. For purposes of completeness, non-limiting aspects and embodiments of the present disclosure are further disclosed in the following numbered clauses.
[0147] 1. A computer-implemented method for assessing T cell receptor chain complementary determining region 3 (TCRp CDR3) sequences, the method comprising: assessing TCRp CDR3 sequences determined from a sample obtained from a subject for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID NOs:1 -997.
[0148] 2. The computer-implemented method of clause 1 , wherein the subject has one or more non-specific symptoms consistent with multiple sclerosis (MS) at the time of the assessing.
[0149] 3. The computer-implemented method of clause 2, wherein the one or more nonspecific symptoms are selected from the group consisting of: CNS inflammation, demyelination, bladder and / or bowel dysfunction, diplopia, dysphagia, altered sensation or weakness of the face, ataxia, blurry vision with painful eye movements, asymmetric limb weakness, spasticity, impaired mobility, impaired motor dexterity, cognitive impairment, and any combination thereof.
[0150] 4. The computer-implemented method of any one of clauses 1 to 3, wherein the TCRp CDR3 sequences determined from the sample obtained from the subject comprise 10,000 or more TCRp CDR3 sequences.
[0151] 5. The computer-implemented method of any one of clauses 1 to 4, wherein the TCRp CDR3 sequences determined from the sample obtained from the subject were determined by performing amplification and high throughput sequencing of genomic DNA present in the sample obtained from the subject.
[0152] 6. The computer-implemented method of any one of clauses 1 to 5, wherein the sample obtained from the subject is a peripheral blood sample.
[0153] 7. The computer-implemented method of any one of clauses 1 to 5, wherein the sample obtained from the subject is a neural tissue sample.
[0154] 8. The computer-implemented method of any one of clauses 1 to 7, wherein the assessing comprises subjecting the TCRp CDR3 sequences determined from the sample to a model.
[0155] 9. A method comprising administering a multiple sclerosis (MS) therapy to a subject identified as comprising T cells that express a T cell receptor p chain (TCRP) comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997. 10. The method of clause 9, wherein the method comprises administering a multiple sclerosis therapy to a subject predicted to have MS using a model.
[0156] 11 . The method of clause 9 or 10, wherein the MS therapy comprises administering to the subject a therapeutically effective amount of natalizumab-sztn (Tyruko), natalizumab (Tysabri), ublituximab-xiiy (Briumvi), ublituximab, ponesimod (Ponvory), ofatumumab (Kesimpta), monomethyl fumarate (Bafiertam), ozanimod (Zeposia), diroximel fumarate (Vumerity), cladribine (Mavenclad), siponimod (Mayzent), ocrelizumab (Ocrevus), daclizumab (Zinbryta), alemtuzumab (Lemtrada), peginterferon beta-1 a (Plegridy), dimethyl fumarate (Tefcidera), diroximel fumarate, teriflunomide (Aubagio), fingolimod (Gilenya), laquinimod, ozanimod, glatiramer acetate, interferon beta 1a, interferon beta 1b, PIPE-307, or any combination thereof.
[0157] 12. A non-transitory computer readable medium having stored thereon one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
[0158] 13. The non-transitory computer readable medium of clause 12 having stored thereon 5 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
[0159] 14. The non-transitory computer readable medium of clause 12 having stored thereon 10 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
[0160] 15. The non-transitory computer readable medium of clause 12 having stored thereon 100 or more TCRP CDR3 sequences set forth in SEQ ID Nos:1 -997.
[0161] 16. A system for assessing TCRp CDR3 sequences, comprising: one or more processors; and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to assess TCRp CDR3 sequences determined from a sample obtained from a subject for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
[0162] 17. The system of clause 16, wherein the TCRp CDR3 sequences determined from the sample obtained from the subject comprise 10,000 or more TCRp CDR3 sequences.
[0163] 18. The system of clause 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 5 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
[0164] 19. The system of clause 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 10 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
[0165] 20. The system of clause 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 100 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
[0166] 21 . The system of any one of clauses 16-20, wherein the assessing comprises subjecting the TCRp CDR3 sequences determined from the sample to a model.
[0167] 22. A genetically modified cell that expresses a T cell receptor (TCR) comprising a TCRp chain comprising a TCRp CDR3 sequence of one of SEQ ID NOs:1 -997.
[0168] 23. The cell of clause 22, wherein the TCRp CDR3 sequence is CAISESWAGGTDTQYF (SEQ ID NO: 458) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO: 998).
[0169] 24. The cell of clause 22, wherein the TCRp CDR3 sequence is CAISESWTGGSDTQYF (SEQ ID NO. 255) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRPNTGTASKLTF (SEQ ID NO. 999).
[0170] 25. The cell of clause 22, wherein the TCRp CDR3 sequence is CAISESWSGGTDTQYF (SEQ ID NO. 877) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRQNTGTASKLTF (SEQ ID NO. 1000).
[0171] 26. The cell of clause 22, wherein the TCRp CDR3 sequence is CAISEGWTGNTDTQYF (SEQ ID NO. 996) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO. 1001 ).
[0172] 27. The cell of clause 22, wherein the TCRp CDR3 sequence is CAISESSGSTDTQYF (SEQ ID NO. 997) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRVATGTASKLTF (SEQ ID NO. 1002).
[0173] 28. The cell of clause 22, wherein the TCRp CDR3 sequence is CASSDQGGGYEQYF (SEQ ID NO. 33) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CAASRDNQGGKLIF (SEQ ID NO. 1003).
[0174] 29. The cell of any one of clauses 22-28, wherein the cell is an immune regulatory cell.
[0175] 30. The cell of clause 29, wherein the immune regulatory cell is a regulatory T cell (Treg).
[0176] 31 . A method of preventing or inhibiting a multiple sclerosis (MS)-associated adaptive immune response in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of immune regulatory cells as defined in clause 29 or 30. 32. A method of treating multiple sclerosis in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of immune regulatory cells as defined in clause 29 or 30.
[0177] 33. The method of any one of clauses 9-11 , 31 or 32, wherein the MS is relapsing-remitting MS (RRMS), secondary-progressive MS (SPMS), primary-progressive MS (PPMS), progression-relapsing MS (PRMS), Marburg variant multiple sclerosis, tumefactive multiple sclerosis, neuromyelitis optica (Devic's disease), or Balo's concentric sclerosis.
[0178] 34. The method of any one of clauses 9-11 or 31 -33, wherein prior to the administering, the subject has been identified as HLA-DRB1*15 positive.
[0179] The following examples are offered by way of illustration and not by way of limitation.
[0180] EXPERIMENTAL
[0181] Example 1 - Prediction of Multiple Sclerosis from TCR Repertoire Data
[0182] Described in this example is the identification of multiple sclerosis-associated TCRp sequences, sometimes referred to herein as “enhanced sequences for multiple sclerosis”, “enhanced sequences”, or the like.
[0183] T-cell receptor repertoires were generated by immunosequencing of whole blood samples, peripheral blood mononuclear cells (PBMC), or buffy coat preparations of blood. Briefly, genomic DNA was extracted from the blood or cell samples using standard extraction kits. As much as 18 pg of genomic DNA was then input into a multiplex PGR reaction to amplify the CDR3 regions of TCRfJ chains followed by high-throughput sequencing (the immunoSEQ Assay as described above).
[0184] An initial diagnostic model to predict multiple sclerosis from TCR repertoire data was run on peripheral samples from multiple sclerosis patients and healthy controls. The cases in the training set included 1 ,979 multiple sclerosis samples. The controls in the training set included 4,236 healthy blood donors collected via contract research organization and other studies. The model first subsets the samples by the status of Class II HLA allele DRB1 *15:01 using HLA inference from each individual’s TCR repertoire data (described in U.S. Patent Application Publication No. US 2021 / 0381050 A1) to account for enrichment of this HLA allele in multiple sclerosis patients. Following this stratification, the training samples with HLA allele DRB1 *15:01 included 883 multiple sclerosis patients and 702 healthy controls and the training samples without HLA allele DRB1 *15:01 included 1 ,096 multiple sclerosis patients and 3,534 healthy controls. The model then uses one-tailed Fisher’s exact tests within each genetic background (with and without HLA allele DRB1 *15:01) to identify unique TCR sequences that are elevated in the multiple sclerosis case samples versus the controls, which are referred to herein as “enhanced sequences”. Unique sequences are identified by their V gene, J gene, and TCRp CDR3 amino acid sequences. For each genetic background, a logistic growth curve was fit to training samples with that genetic background based on variables N and E, where N is the total number of unique TCRp DNA sequences in a subject and E is the number of unique TCRp DNA sequences that encode an enhanced sequence identified in that genetic background. For each genetic background, a binomial error model was also derived based on the logistic growth curve (following the methodology detailed in Griessl et al. medRxiv 2021 - doi.org / 10.1101 / 2021 .07.30.21261353). These logistic growth curves and binomial error models were used to predict multiple sclerosis disease status, where each sample obtains a score equal to the number of standard deviations above the logistic growth curve, and a threshold value of this score is used for classification. As illustrated in FIG. 1 , such a model was used to predict multiple sclerosis disease status in Holdout Data sets not used in training. In samples with DRB1 *15:01 , the model’s AUC (area under the receiver operating characteristic curve) was 0.809 and sensitivity was 62.5% at 85% specificity; in samples without DRB1 *15:01 , the model’s AUC was 0.607 and sensitivity was 27.8% at 85% specificity.
[0185] Other models were trained in multiple sclerosis, including different selected cohorts and statistical models, that identified other multiple sclerosis-associated sequences. Together, 997 unique TCRp sequences were identified to be significantly associated with multiple sclerosis (Table 1 herein). This list showed a high degree of overlap of sequences identified in various approaches.
[0186] It was determined that the multiple sclerosis disease-associated TCRp sequences exhibit clusters associated with HLA alleles, and high prevalence in multiple sclerosis patients but not controls, as shown in FIG. 2. The analysis shown in FIG. 2B. was performed with the most public members of Cluster ID NO 1 from FIG. 2A. as summarized in the inlaid motif plot.
[0187] High-throughput pairing of T cell receptor alpha and beta sequences (pairSEQ for short, see Howie et al. Science Translational Medicine 2015 and PCT / US2015 / 058035) was performed in 47 samples of patients with multiple sclerosis. The resulting dataset of pairs were searched for TCRp sequences that were highly similar to our set of enhanced sequences for multiple sclerosis, where similarity was defined using a threshold on Levenshtein distance (minimum number of insertions, deletions, or substitutions to change one string into another). Of the matched pairs, one cluster of 5 pairs was identified where both the TCRp and TCRa sequences were all highly similar, and where the enhanced sequence was associated to the DR15 HLA haplotype. One pair with the TCRp was identified matching an enhanced sequence associated to the DQ2.5 HLA haplotype from Cluster ID NO 2 in FIG. 2A (see Table 2). Table 2 provides paired TCRp / TCRa sequences found in multiple sclerosis samples and of high sequence similarity to the multiple sclerosis disease- associated sequences. Additionally listed are V and J genes separately for a and p as well as the Levenshtein distance between the p in the pair and the closest multiple sclerosis disease-associated sequences.
[0188] Table 2
[0189] Accordingly, the preceding merely illustrates the principles of the present disclosure. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein.
Claims
WHAT is CLAIMED is:
1. A computer-implemented method for assessing T cell receptor (3 chain complementary determining region 3 (TCR[3 CDR3) sequences, the method comprising: assessing TCR[3 CDR3 sequences determined from a sample obtained from a subject for the presence or absence of one or more TCR£ CDR3 sequences set forth in SEQ ID NOs:1 -997.
2. The computer-implemented method of claim 1 , wherein the subject has one or more non-specific symptoms consistent with multiple sclerosis (MS) at the time of the assessing.
3. The computer-implemented method of claim 2, wherein the one or more nonspecific symptoms are selected from the group consisting of: CNS inflammation, demyelination, bladder and / or bowel dysfunction, diplopia, dysphagia, altered sensation or weakness of the face, ataxia, blurry vision with painful eye movements, asymmetric limb weakness, spasticity, impaired mobility, impaired motor dexterity, cognitive impairment, and any combination thereof.
4. The computer-implemented method of any one of claims 1 to 3, wherein the TCR0 CDR3 sequences determined from the sample obtained from the subject comprise 10,000 or more TCR[3 CDR3 sequences.
5. The computer-implemented method of any one of claims 1 to 4, wherein the TCR[3 CDR3 sequences determined from the sample obtained from the subject were determined by performing amplification and high throughput sequencing of genomic DNA present in the sample obtained from the subject.
6. The computer-implemented method of any one of claims 1 to 5, wherein the sample obtained from the subject is a peripheral blood sample.
7. The computer-implemented method of any one of claims 1 to 5, wherein the sample obtained from the subject is a neural tissue sample.
8. The computer-implemented method of any one of claims 1 to 7, wherein the assessing comprises subjecting the TCR|3 CDR3 sequences determined from the sample to a model.
9. A method comprising administering a multiple sclerosis (MS) therapy to a subject identified as comprising T cells that express a T cell receptor p chain (TCRP) comprising a TCRp CDR3 sequence set forth in SEQ ID Nos:1 -997.
10. The method of claim 9, wherein the method comprises administering a multiple sclerosis therapy to a subject predicted to have MS using a model.11 . The method of claim 9 or 10, wherein the MS therapy comprises administering to the subject a therapeutically effective amount of natalizumab-sztn (Tyruko), natalizumab (Tysabri), ublituximab-xiiy (Briumvi), ublituximab, ponesimod (Ponvory), ofatumumab (Kesimpta), monomethyl fumarate (Bafiertam), ozanimod (Zeposia), diroximel fumarate (Vumerity), cladribine (Mavenclad), siponimod (Mayzent), ocrelizumab (Ocrevus), daclizumab (Zinbryta), alemtuzumab (Lemtrada), peginterferon beta-1 a (Plegridy), dimethyl fumarate (Tefcidera), diroximel fumarate, teriflunomide (Aubagio), fingolimod (Gilenya), laquinimod, ozanimod, glatiramer acetate, interferon beta 1a, interferon beta 1b, PIPE-307, or any combination thereof.
12. A non-transitory computer readable medium having stored thereon one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
13. The non-transitory computer readable medium of claim 12 having stored thereon 5 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
14. The non-transitory computer readable medium of claim 12 having stored thereon 10 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
15. The non-transitory computer readable medium of claim 12 having stored thereon 100 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1 -997.
16. A system for assessing TCRp CDR3 sequences, comprising: one or more processors; and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to assess TCRp CDR3 sequences determined from a sample obtained from asubject for the presence or absence of one or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
17. The system of claim 16, wherein the TCRp CDR3 sequences determined from the sample obtained from the subject comprise 10,000 or more TCRp CDR3 sequences.
18. The system of claim 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 5 or more TCR CDR3 sequences set forth in SEQ ID Nos:1-997.
19. The system of claim 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 10 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.
20. The system of claim 16 or 17, wherein the instructions, when executed by the one or more processors, cause the one or more processors to assess the determined TCRp sequences for the presence or absence of 100 or more TCRp CDR3 sequences set forth in SEQ ID Nos:1-997.21 . The system of any one of claims 16-20, wherein the assessing comprises subjecting the TCRp CDR3 sequences determined from the sample to a model.
22. A genetically modified cell that expresses a T cell receptor (TCR) comprising a TCRp chain comprising a TCRp CDR3 sequence of one of SEQ ID NOs:1 -997.
23. The cell of claim 22, wherein the TCRp CDR3 sequence is CAISESWAGGTDTQYF (SEQ ID NO: 458) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO: 998).
24. The cell of claim 22, wherein the TCRp CDR3 sequence is CAISESWTGGSDTQYF (SEQ ID NO. 255) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRPNTGTASKLTF (SEQ ID NO. 999).
25. The cell of claim 22, wherein the TCR[3 CDR3 sequence is CAISESWSGGTDTQYF (SEQ ID NO. 877) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRQNTGTASKLTF (SEQ ID NO. 1000).
26. The cell of claim 22, wherein the TCRp CDR3 sequence is CAISEGWTGNTDTQYF (SEQ ID NO. 996) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRGNTGTASKLTF (SEQ ID NO. 1001 ).
27. The cell of claim 22, wherein the TCR[3 CDR3 sequence is CAISESSGSTDTQYF (SEQ ID NO. 997) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CIVRVATGTASKLTF (SEQ ID NO. 1002).
28. The cell of claim 22, wherein the TCR CDR3 sequence is CASSDQGGGYEQYF (SEQ ID NO. 33) and the cell further comprises a TCRa comprising a TCRa CDR3 sequence comprising CAASRDNQGGKLIF (SEQ ID NO. 1003).
29. The cell of any one of claims 22-28, wherein the cell is an immune regulatory cell.
30. The cell of claim 29, wherein the immune regulatory cell is a regulatory T cell (Treg).31 . A method of preventing or inhibiting a multiple sclerosis (MS)-associated adaptive immune response in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of immune regulatory cells as defined in claim 29 or 30.
32. A method of treating multiple sclerosis in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of immune regulatory cells as defined in claim 29 or 30.
33. The method of any one of claims 9-11 , 31 or 32, wherein the MS is relapsing-remitting MS (RRMS), secondary-progressive MS (SPMS), primary-progressive MS (PPMS), progression-relapsing MS (PRMS), Marburg variant multiple sclerosis, tumefactive multiple sclerosis, neuromyelitis optica (Devic's disease), or Balo's concentric sclerosis.
34. The method of any one of claims 9-11 or 31 -33, wherein prior to the administering, the subject has been identified as HLA-DRB1*15 positive.