IMMUNE REPERTOIRE MONITORING
Patent Information
- Application Number
- DE602019069854
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-06-07
- Filing Date
- 2019-03-22
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2039-03-22
AI Technical Summary
Current methods for analyzing the immune repertoire at high resolution are limited by low throughput and the introduction of sequence errors, making it difficult to accurately reflect the true immune receptor repertoire.
The method involves performing next-generation sequencing of target immune receptor nucleic acid template molecules derived from a biological sample, determining the immune receptor haplotype, and predicting a subject's potential or predisposition to adverse events following immunotherapy by comparing the haplotype to a reference set.
This approach allows for the accurate prediction of a subject's response to immunotherapy, enabling personalized treatment decisions and minimizing adverse events.
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 681,734 filed June 7, 2018 and U.S. Provisional Application No. 62 / 647,566 filed March 23, 2018.BACKGROUND
[0002] Adaptive immune response comprises selective response of B and T cells recognizing antigens. The immunoglobulin genes encoding antibody (Ab, in B cell) and T-cell receptor (TCR, in T cell) antigen receptors comprise complex loci wherein extensive diversity of receptors is produced as a result of recombination of the respective variable (V), diversity (D), and joining (J) gene segments, as well as subsequent somatic hypermutation events during early lymphoid differentiation. The recombination process occurs separately for both subunit chains of each receptor and subsequent heterodimeric pairing creates still greater combinatorial diversity. Calculations of the potential combinatorial and junctional possibilities that contribute to the human immune receptor repertoire have estimated that the number of possibilities greatly exceeds the total number of peripheral B or T cells in an individual. See, for example, Davis and Bjorkman (1988) Nature 334:395-402; Arstila et al. (1999) Science 286:958-961; van Dongen et al., In: Leukemia, Henderson et al. (eds) Philadelphia: WB Saunders Company, 2002, pp 85-129.
[0003] Extensive efforts have been made over years to improve analysis of the immune repertoire at high resolution. WO2015 / 058159 and WO2014 / 055561 provide methods of sequencing an immune receptor repertoire in the context of predicting the response to immunotherapy. LOONEY et al. "Long-amplicon TCR beta repertoire sequencing to reveal human T-cell receptor variable gene polymorphism: Implications for the prediction and interpretation of immunotherapy outcome. "(JOURNAL OF CLINICAL ONCOLOGY,vol. 36, no. 4, 2018, pages 129-129) discloses methods of sequencing an immune receptor repertoire in the context of immunotherapy, namely T-cell receptor alleles are profiled. WO2016 / 197131 discloses the sequencing of immunoglobulins and scoring for non-synonymous-to-synonymous mutations. SACHA GNJATIC ET AL: "Identifying baseline immune-related biomarkers to predict clinical outcome of immunotherapy", JOURNAL FOR IMMUNOTHERAPY OF CANCER, BIOMED CENTRAL LTD, LONDON, UK, vol. 5, no. 1, 16 May 2017 (2017-05-16), pages 1-18, reviews methods to predict the response to immunotherapy. Disclosed are i.a. T-cell diversity in the context of toxicity and haplotypes (SNPs) that may influence the clinical response. However, there is no mention of immune receptors having CDR elements. Means for specific detection and monitoring of expanded clones of lymphocytes would provide significant opportunities for characterization and analysis of normal and pathogenic immune reactions and responses. Despite efforts, effective high resolution analysis has provided challenges. Low throughput techniques such as Sanger sequencing may provide resolution, but are limited to provide efficient means to broadly capture the entire immune repertoire. Advances in next generation sequencing (NGS) have provided access to capturing the repertoire, however, due to the nature of the numerous related sequences and introduction of sequence errors as a result of the technology, efficient and effective reflection of the true repertoire has proven difficult. Thus, improved sequencing methodologies and workflows capable of resolving complex populations of highly variable immune cell receptor sequences are being developed. There remains a need for new methods for effective profiling of vast repertoires of immune cell receptors to better understand immune cell response, enhance diagnostic and treatment capabilities, and devise new therapeutics.SUMMARY OF THE INVENTION
[0004] The invention provides methods or predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy.. In some embodiments, provided method comprise performing sequencing of target immune receptor nucleic acid template molecules derived from a biological sample from a subject candidate for an immunotherapy, wherein the target immune receptor nucleic acid template molecules comprise FR1, CDR1, FR2, CDR2, FR3, and CDR3 coding regions of the target immune receptor and the sequencing is by next generation sequencing; determining the sequence of the target immune receptor repertoire of the sample based on the sequencing; identifying the immune receptor haplotype of the subject from the determined sequences; and predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy by comparing the identified immune receptor haplotype of the subject to a reference set of immune receptor haplotypes of individuals with annotated adverse events following therapy treatments. In some embodiments, the method, prior to the sequencing, further comprises performing a multiplex amplification reaction to amplify target immune receptor nucleic acid template molecules, wherein the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of V genes of the target immune receptor. In some embodiments, the subject has cancer, an autoimmune disease or a chronic viral infection.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] FIG. 1 is a diagram of an exemplary workflow for removal of PCR or sequencing-derived errors using stepwise clustering of similar CDR3 nucleotides sequences with steps: (A) very fast heuristic clustering into groups based on similarity (cd-hit-est); (B) cluster representative chosen as most common sequence, randomly picked for ties; (C) merge reads into representatives; (D) compare representatives and if within allotted hamming distance, merge clusters. FIGS. 2A-2B depict correlation plots comparing TCR V gene usage characterization from the same peripheral blood mononuclear cell RNA sample prepared for sequencing using three different methodologies: single primer 5'-RACE, the presently provided primers and workflows, and the BIOMED-2 primer set. FIG. 2A depicts correlation plots of TCR V gene usage comparing 5'-RACE and BIOMED-2 primer set for sample preparation. FIG. 2B depicts correlation plots of TCR V gene usage comparing 5'-RACE and the presently provided primers and workflows for sample preparation. FIG. 3 is a graph comparing the TCR clone frequency in peripheral blood (x-axis) and the TCR clone frequency in tumor (y-axis) from an individual with squamous cell carcinoma of the lung. Sequencing of the TCR beta immune repertoire identified 219 TCR clones shared between peripheral blood and tumor and 370 TCR clones unique to the tumor. FIGS. 4A-4B depict convergent TCR beta frequencies in peripheral blood lymphocyte samples from subjects prior to immunotherapy. FIGS. 5A-5C depict convergent TCR beta frequencies in peripheral blood lymphocyte samples from subjects with melanoma prior to immunotherapy. FIG. 6 depicts convergent TCR frequencies in samples from healthy subjects and subjects with cancer. FIGS. 7A-7B depict TCRV haplotype group analysis for subjects who experienced adverse events following immunotherapy. FIG. 8 depicts within cluster sum of squares follow subdivision of data into up to 15 clusters via k-means analysis for 81 samples (cohort 1 and 2). The curver of sum of squares values is used to estimate the optimal number of clusters via the "elbow" method. FIG. 9 depicts correlation between the number of uncommon alleles and the mean frequency of adverse events per haplotype group (Spearman cor = .83). The number inside chart indicates haplotype group. FIGS. 10A-10B are graphs depicting the receiver operator curves for K-Nearest Neighbor classifier (FIG. 10A) and logistic regression classifier (FIG. 10B). Sensitivity and specificity for performance of classifiers trained on cohort 1 then tested on cohort 2 samples. Area under the curve is displayed. DESCRIPTION OF THE INVENTION
[0006] We have developed methods of predicting a subject's response based on characterizing the immune repertoire of the subject before receiving treatment.
[0007] In one aspect, the present invention provides methods for predicting a subjects potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy.
[0008] Previous groups have identified 15% or more of the total sequences in the peripheral blood as appearing to derive from convergent TCR groups (Ruggerio et al. (2015) Nat. Commun. 6:8081 and Emerson et al. (2017) Nature Genetics 49(5):659). By contrast, sequencing of TCR repertoires of healthy peripheral blood samples using the high accuracy amplification and sequencing assays and sequence data analysis provided herein indicates that convergent TCRs represent only about 0.2% - 2% of the total sequences in the healthy peripheral blood TCR repertoire. One possibility for this substantial difference may involve amplification and sequencing errors that may have existed in the data of the previous studies. For example, base substitution sequencing errors and PCR errors can create artifacts that resemble TCR convergence, thus inflating convergence estimates. An inflated estimate in TCR convergence frequency could potentially mask subtle true differences in convergence frequencies between two or more samples, such as samples from immunotherapy responders and non-responders.
[0009] Methods, compositions and analysis provided herein are for use in predicting clinical responsiveness to a therapy comprises identifying convergent immune receptor groups in a pre-treatment sample from a subject using methodology for high accuracy amplification and sequencing of immune receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) in the subject's sample. The immune receptor sequencing data is used to identify immune receptor clones and the frequency of all clones having a convergent immune receptor in the sample is a predictor of the subject's clinical response to a therapy. The subject may be treated with a therapy in a manner dependent on the frequency of the convergent immune receptor clones. For example, a subject having a convergent immune receptor clone frequency greater than a convergent frequency cutoff indicates that the subject is candidate for the therapy whereas a subject having a convergent immune receptor clone frequency less than a convergent frequency cutoff indicates that the subject is not candidate for the therapy. In some embodiments, provided methods comprise identifying convergent immune receptor clones from the immune receptor clones present in the sample at a frequency of greater than 1 in 50,000. In some embodiments, the convergent frequency cutoff is a frequency of greater than 0.01. The subject may have cancer and is a candidate for an immunotherapy. The subject may be a candidate for a vaccination against an infectious agent or disease or a candidate for autoimmune suppressant treatment.
[0010] IProvided methods comprise identifying convergent immune receptor clones using V gene identity and sequences comprising CDR3 amino acid sequences. Provided methods comprise identifying convergent immune receptor clone using sequences that comprise CDR3 sequences, CDR1 and CDR3 sequences, or CDR2 and CDR3 sequences.
[0011] Provided methods comprise identifying convergent TCR clones as those comprising TCR variable and CDR3 rearrangements that are similar or identical in amino acid sequence but different in nucleotide sequence. For example, a significant fraction of the TCRs that differ from one another by one amino acid residue may nonetheless have similar or identical specificity for an antigen and so such TCRs may be considered convergent.
[0012] In some embodiments, methods are provided for predicting clinical response of a subject to immunotherapy by identifying the frequency of convergent TCR groups in a biological sample from a subject prior to receiving immunotherapy. In some embodiments, such methods comprise a) performing sequencing of a TCR repertoire from a subject's peripheral blood sample at a pre-treatment time point and identifying TCR clones based on the sequence; b) identifying convergent TCR clones as those comprising TCR V gene and CDR3 sequences that are similar or identical in amino acid sequence but different in nucleotide sequence; and c) calculating the sum of the frequency of all the convergent TCR clones in the subject's TCR repertoire.
[0013] The following embodiments relating to methods for treating a subject are disclosed but not claimed. In some embodiments, methods are provided for treating a subject based on characterizing the immune repertoire of the subject before receiving the treatment. In some embodiments, methods are provided for treating a subject based on characterizing the frequency of convergent immune receptor clones in a sample from the subject before receiving the treatment. In embodiments, provided methods comprise performing a multiplex amplification reaction to amplify target immune receptor nucleic acid template molecules derived from a biological sample from a subject where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of V genes of the target immune receptor, thereby generating target immune receptor amplicon molecules comprising the target immune receptor repertoire. In embodiments, such provided methods further comprise performing sequencing of the target immune receptor repertoire amplicons; identifying immune receptor clones from the sequencing and identifying convergent immune receptor clones among the immune receptor clones, wherein the convergent immune receptor clones have a similar or identical amino acid sequence and a different nucleotide sequence; determining the frequency of convergent immune receptor clones in the sample; and treating the subject with a therapy in a manner dependent on the frequency of the convergent immune receptor clones. In some embodiments, provided methods comprise treating the subject with a particular therapy when the frequency of convergent immune receptor clones in the sample is greater than a convergent frequency cutoff. In some embodiments, provided methods comprise treating the subject with an alternative therapy when the frequency of convergent immune receptor clones in the sample is less than a convergent frequency cutoff.
[0014] The following embodiments relating to methods for treating a subject are disclosed but not claimed. In some embodiments, methods are provided for treating a subject with cancer comprising performing a multiplex amplification reaction to amplify target TCR nucleic acid template molecules derived from a biological sample from the subject where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of TCR V genes, thereby generating target TCR amplicon molecules comprising the target TCR repertoire of the subject. Such methods further comprise performing sequencing of the target TCR repertoire amplicons; identifying TCR clones from the sequencing and identifying convergent TCR clones among the immune receptor clones, wherein the convergent TCR clones have a similar or identical amino acid sequence and a different nucleotide sequence; determining the frequency of convergent TCR clones in the sample; and treating the subject in a manner dependent on the frequency of the convergent TCR clones. In some embodiments, provided methods comprise treating the subject with an immunotherapy when the frequency of convergent TCR clones in the sample is greater than a convergent frequency cutoff. In some embodiments, provided methods comprise treating the subject with an alternate immunotherapy or with a non-immunotherapy treatment when the frequency of convergent immune receptor clones in the sample is less than a convergent frequency cutoff. In some embodiments, the target cancer types for these methods include any cancer that may harbor immunogenic antigens that may be targeted by therapies that enhance T cell ability to destroy damaged cells. Such therapies include for example immune checkpoint blockade agents, but also T cell agonists and agents that indirectly modulate the activity of T cells, for example dendritic cell vaccination. In some embodiments, the target cancer types include without limitation melanoma, adenocarcinoma, non-small cell lung cancer, prostate cancer and others.
[0015] Provided methods may comprise an initial step in calculating the sum of the frequency of convergent TCR clones is the elimination of clones of a frequency below a given threshold. In some embodiments, TCR clones of a frequency of >1 in 50,000 may be used to calculate the convergent TCR frequency. In other embodiments, TCR clones of a frequency of >1 in 10,000, > 1 in 25,000, > 1 in 75,000, or >1 in 100,000 may be used to calculate the convergent TCR frequency. Such elimination of low frequency clones prior to calculating the frequency of convergent TCR clones may help make the calculation more robust to amplification and sequencing error.
[0016] The set of frequencies of identified convergent TCR clones may be transformed by a mathematical function prior to calculating the sum of the frequencies of the convergent TCR clones in the sample. For example, this transformation involves raising each convergent TCR clone frequency to an exponential power greater than 1 (for example, squaring each value), thereby increasing the relative contribution of higher frequency clones to the aggregate convergent TCR clone frequency. In another example, this transformation involves raising each convergent TCR clone frequency to an exponential less than 1 (for example, taking the square root of each frequency), thereby increasing the relative contribution of lower frequency clones to the aggregate convergent TCR clone frequency. The choice of transformation depends on the relative contribution of the T cell subsets to potential immune response, such as for example, potential anti-tumor or anti-chronic viral infection responses. For example, the mathematical transformation may depend on the immunotherapy treatment (for example without limitation checkpoint-blockade therapy) and the extent to which the treatment is able to reactivate dysfunctional T cell subsets. Convergent TCR clones are identified in the subject's repertoire which are known or suspected to be irrelevant to the disease or condition of the subject. Accordingly, irrelevant convergent TCR clones are excluded from the convergent TCR clone frequency calculation for the sample.
[0017] A convergent frequency cutoff for separating therapy responders from non-responders may depend on the disease or disorder for treatment of prevention, the intended therapy, and / or a desire to minimize the false positive and false negative rate for prediction of response. An immune receptor convergent frequency cutoff may be identified by measuring the convergent immune receptor frequency in pre-treatment (baseline) samples from subjects having received a therapy and comparing the measured convergent immune receptor frequencies with the clinical response of the subjects to the therapy. A TCR convergent frequency cutoff for separating responders from non-responders may depend on the cancer type of the subject, the intended immunotherapy, and / or a desire to minimize the false positive and false negative rate for prediction of response.
[0018] The convergent immune receptor clone frequency cutoff may be a frequency of about 0.01. In some examples, the convergent immune receptor clone frequency cutoff ranges from about 0.001 to about 0.03or the convergent TCR clone frequency cutoff is within the range of about 0.001 to about 0.03. The convergent TCR clone frequency cutoff may be a frequency of about 0.002, 0.004, 0.006, 0.008, 0.011, 0.012, 0.013, 0.014, 0.016, 0.018, 0.020, 0.022, 0.025, or 0.03. In some embodiments, detection of a convergent immune receptor clone frequency greater than a particular frequency cutoff predicts objective clinical response of the subject to a therapy. Detection of a convergent immune receptor clone frequency less than a particular frequency cutoff may predict no objective clinical response of the subject to a therapy. A convergent TCR frequency in a pre-immunotherapy sample of > 0.01 may be predictive of the subject having an objective clinical response following immunotherapy. A convergent TCR frequency in a pre-immunotherapy sample of < 0.01 may be predictive of the subject having no objective clinical response following immunotherapy. As shown herein, individuals having a TCR beta convergent clone frequency >.01 are more likely to be responders to immunotherapy, such as an immune checkpoint blockade agent and / or anti-cancer vaccine, while those having a frequency of <.01 are more likely to be non-responders.
[0019] A change in convergent TCR clone frequency over the course of a therapy treatment may be used as a predictor of response to the therapy. In a manner dependent on disease type and treatment, responders may be distinguished from non-responders by an increase in the frequency of convergent TCR clones over the course of a therapy. For example, in cancers (or chronic viral infections) in which convergent TCR clones of the T cell population primarily consist of effector T cells of a progenitor exhausted T cell phenotype, a terminally exhausted phenotype or an effector phenotype among other T cell phenotypes, an increase in the frequency of convergent TCR clones over the course of a treatment may be indicative of an increase in the activity of anti-cancer (or anti-viral) T cells. In other cancers, convergent TCR clones may primarily be of T regulatory phenotype and an increase in the frequency of convergent TCR clones over the course of a therapy may indicate a poor prognosis.
[0020] Measurement or determination of the frequency of convergent TCR clones may be combined with other T cell repertoire features, such as for example, measurements of T cell clonal expansion, to improve the prediction of clinical responsiveness. Measurement or determination of the frequency of convergent TCR clones may be combined with B cell repertoire features, such as for example, measurements of B cell clonal expansion, to improve the prediction of clinical responsiveness. Measurement or determination of the frequency of convergent TCR clones may be combined with measurement or detection of expression of one or more genes relevant to immune response to improve the prediction of clinical responsiveness. Such immune response relevant genes include without limitation PD-1 and / or PD-L1 genes, interferon gamma pathway genes, and myeloid derived suppressor cell related genes. Procedures and reagents for detecting or measuring such gene expression are known in the art and include without limitation quantitative or semi-quantitative PCR analysis, comparative hybridization methods, or sequencing procedures and reagents and kits for use in same including without limitation TaqMan ™< assays and the Oncomine ™< Immune Response Research Assay (Thermo Fisher Scientific).
[0021] As used herein, subjects with "objective clinical response" or "responders" are individuals who had stable disease (SD), partial response (PR), or complete response (CR) following immunotherapy. As used herein, subjects with "no objective clinical response" or "no objective response" or "non-responders" are individuals who have progressive disease (PD) following immunotherapy. Use of these terms is in keeping with the RECIST grading guideline (Eisenhauer et al. (2009) European Journal of Cancer 45:228-247.)
[0022] As used herein, a "convergent TCR group" is a set of T cell receptors (TCRs) that are similar in amino acid sequence and functionally equivalent, or are identical or assumed to be identical in amino acid sequence. It is generally assumed, owing to the amino acid similarity, that a convergent TCR group recognizes the same antigen. In some embodiments, convergent TCR group members are identical or assumed to be identical in the variable gene and CDR3 amino acid sequence despite having a different nucleotide sequence. Convergent TCR group members may result from differences in non-templated nucleotide bases at the VDJ junction that arise during the generation of a productive TCR gene rearrangement.
[0023] Sequences identifying convergent immune receptor clones may comprise CDR3 sequences, CDR1 and CDR3 sequences, or CDR2 and CDR3 sequences. The convergent immune receptor clones are identified using V gene identity and sequences comprising CDR3 amino acid sequences. Convergent immune receptor clones may have identical or similar CDR3 amino acid sequences.
[0024] The frequency of convergent TCRs may have utility as an indicator of T cell responses to tumor antigen, auto-antigen associated with chronic autoimmune disease (including without limitation type I diabetes and rheumatoid arthritis) or an antigen associated with chronic viral infection. Accordingly, determining the frequency of convergent TCR clones in a subject may aid in predicting the emergence of any such chronic diseases or disorders.
[0025] The following embodiments relating to methods for treating a subject are disclosed but not claimed. In some embodiments, methods are provided for treating a subject with an autoimmune disease or disorder comprising performing a multiplex amplification reaction to amplify target TCR nucleic acid template molecules derived from a biological sample from the subject where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of TCR V genes, thereby generating target TCR amplicon molecules comprising the target TCR repertoire of the subject. Such methods further comprise performing sequencing of the target TCR repertoire amplicons; identifying TCR clones from the sequencing and identifying convergent TCR clones among the immune receptor clones, wherein the convergent TCR clones have a similar or identical amino acid sequence and a different nucleotide sequence; determining the frequency of convergent TCR clones in the sample; and treating the subject in a manner dependent on the frequency of the convergent TCR clones. In some embodiments, provided methods comprise treating the subject with an immunosuppressant therapy (for example, without limitation, methotrexate, rituximab or a biologic based therapy e.g., regulatory T cell therapy (Bluestone et al. (2018) Science 362:154-155)) when the frequency of convergent TCR clones in the sample is greater than a convergent frequency cutoff. In some embodiments, provided methods comprise treating the subject with an alternate immunotherapy or with a non-immunotherapy treatment when the frequency of convergent TCR clones in the sample is less than a convergent frequency cutoff. In some embodiments, detecting a frequency of convergent TCR clones less than a convergent frequency cutoff is an indication to administer a T regulatory cell based therapy when the identified convergent TCR clones are known or expected to be of T regulatory type and protective.
[0026] Methods are provided for predicting efficacy of vaccination against an infectious disease for a subject by identifying the frequency of convergent TCR groups in a biological sample from a subject prior to receiving the vaccination. Such methods comprise a) performing sequencing of a TCR repertoire from a subject's peripheral blood sample at a pre-treatment time point and identifying TCR clones based on the sequence; b) identifying convergent TCR clones as those having TCR variable and CDR3 rearrangements that are similar or identical in amino acid sequence but different in nucleotide sequence; and c) calculating the sum of the frequency of all the convergent TCR clones. In some embodiments, TCR convergence is measured in a peripheral blood sample from the subject taken days to weeks after the vaccination, for example without limitation about 7-14 days after vaccination, and compared to the pre-vaccination TCR convergence frequency levels. Such a comparison may be used to show efficacy or lack of efficacy of the vaccination.
[0027] Methods are provided for vaccinating a subject against an infectious disease comprising performing a multiplex amplification reaction to amplify target TCR nucleic acid template molecules derived from a pre-vaccination biological sample from the subject where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of TCR V genes, thereby generating target TCR amplicon molecules comprising the target TCR repertoire of the subject. The method further comprises performing sequencing of the target TCR repertoire amplicons; identifying TCR clones from the sequencing and identifying convergent TCR clones among the immune receptor clones, wherein the convergent TCR clones have a similar or identical amino acid sequence and a different nucleotide sequence; determining the frequency of convergent TCR clones in the sample; and vaccinating the subject with a vaccine against an infectious disease when the frequency of convergent TCR clones in the sample is greater than a convergent frequency cutoff.
[0028] Methods are provided for detecting and / or identifying TCR clones directed to antigens associated with a chronic disease or condition such as, for example, tumor antigen, auto-antigen associated with chronic autoimmune disease and antigen associated with chronic viral infection. Methods for detecting and / or identifying TCR clones directed to chronic antigen(s) comprise performing a multiplex amplification reaction to amplify target TCR nucleic acid template molecules derived from a biological sample from a subject having a chronic disease or condition where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of TCR V genes, thereby generating target TCR amplicon molecules comprising the target TCR repertoire of the subject. The method further comprises performing sequencing of the target TCR repertoire amplicons; identifying TCR clones from the sequencing and identifying convergent TCR clones among the immune receptor clones, wherein the convergent TCR clones have a similar or identical amino acid sequence and a different nucleotide sequence; and determining the frequency convergent TCR clones in the sample, wherein convergent TCR clones are responsive to chronic antigens.
[0029] Methods provided may be of use for improved production of antigen specific TCRs and engineering antigen reactive T cell populations, for example for therapeutic applications. Nonlimiting examples of antigen specific TCR beta (TCRB) chains and engineered T cells for therapeutic applications include 1) TCRBs and T cells that target cancer and virus-associated antigens for use in treating cancer and virus-associated conditions and 2) TCRBs and regulatory T cells that target autoantigens for use in treating for example severe autoimmune disease. Convergent TCRB chains are hypothesized to be beta chain dominant meaning that they can be paired with many different TCR alpha rearrangements without affecting the antigen specificity of the receptor. This property may be used to circumvent many of the laborious steps currently undertaken to create antigen specific TCRs for therapeutic uses. Accordingly, methods are provided for engineering antigen reactive T cells from convergent TCR beta clones, such methods comprise performing a multiplex amplification reaction to amplify target TCR beta nucleic acid template molecules derived from a biological sample from the subject where the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of TCR beta V genes, thereby generating target TCR beta amplicon molecules comprising the target TCR beta repertoire of the subject; performing sequencing of the target TCR beta repertoire amplicons; identifying TCR beta clones from the sequencing and identifying convergent TCR beta clones among the TCR beta clones, wherein the convergent TCR beta clones have a similar or identical amino acid sequence and a different nucleotide sequence. Such methods further comprise cloning sequences of the convergent TCRB chain(s) into an expression vector that enables expression of a convergent TCRB polypeptide in T cells. In some embodiments, such methods further include transducing cloned TCRB chain(s) into T cells isolated from a donor (with or without native TCRB chain removed or inactivated) and screening the engineered T cell population to identity cells in the population that are antigen reactive. In other embodiments, the methods include transducing cloned TCRB chain(s) into T cells with inactivated endogenous TCRB and transducing the same cells with a separate expression vector encoding a TCR alpha chain polypeptide and screening the engineered T cell population to identify cells in the population that are antigen reactive. Procedures and reagents for screening for antigen reactive T cells are known in the art, and include without limitation tests for T cell IFN-gamma or granzyme B secretion (e.g. ELISPOT), flow cytometry using labeled tetramers or dextramers, and in vitro tests of cytotoxicity using T cells co-cultured with cancer cell lines, among other methods. In some embodiments, methods for engineering antigen reactive T cells exclude cloning convergent TCRB sequences that are known to target irrelevant or off-target antigens.
[0030] In another aspect, provided methods are for predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following a therapy by identifying the haplotype of the subject's immune repertoire prior to receiving the therapy. For example, as described herein, identifying TRBV alleles or haplotype group of a subject provides a biomarker predictive of therapy-associated adverse event(s) or autoimmune reactivity.
[0031] Knowing the likelihood that a recipient of an immunotherapy, such as, for example, a checkpoint blockade agent, will suffer an adverse event following the immunotherapy may allow a healthcare or drug provider to optimize the therapeutic dose and subject monitoring to improve efficacy and safety of the therapy. In some embodiments, provided methods are for predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy by identifying the TCR haplotype of the subject's immune repertoire prior to receiving the immunotherapy. In some embodiments, methods are provided for predicting a subject's potential or predisposition to be protected from or vulnerable to a therapy-associated adverse event by identifying the haplotype associated with or causative of risk-associative TRBV alleles, or risk-associated TRB locus haplotypes using TRB repertoire sequencing of a sample from the subject. In some embodiments, the sample is obtained from the subject prior to, during or after administration of the therapy to the subject.
[0032] In some embodiments, provided methods, and analyses are for predicting a subject's potential or predisposition to one or more adverse events following immunotherapy comprising identifying the TCR V gene haplotype group in a sample from a subject using methodology for high accuracy amplification and sequencing of TCR sequences. In embodiments, TCR sequencing data is used to identify the TCR V haplotype of the subject and TCR V haplotype group to which the subject belongs. In embodiments, a TCR V haplotype group is a predictor of the subject's likelihood of being vulnerable to an adverse event following immunotherapy such as for example checkpoint blockade immunotherapy. In embodiments, a TCR V haplotype group is also a predictor of the subject's likelihood of being protected from adverse events following immunotherapy such as for example checkpoint blockade immunotherapy.
[0033] The following embodiments relating to methods for treating a subject are disclosed but not claimed. In some embodiments, methods are provided for treating a subject based on characterizing the immune repertoire haplotype of the subject before receiving the treatment. In some embodiments, provided methods comprise performing sequencing of target immune receptor nucleic acid template molecules derived from a biological sample from a subject with cancer, wherein the target immune receptor nucleic acid template molecules comprise FR1, CDR1, FR2, CDR2, FR3, and CDR3 coding regions of the target immune receptor; determining the sequence of the target immune receptor repertoire of the sample; identifying the immune receptor haplotype of the subject from the determined sequences; and treating the subject with an immunotherapy associated with no or low grade adverse events in individuals having the immune receptor haplotype of the subject. In some embodiments, provided methods comprise de-selecting a subject as a candidate for an immunotherapy associated with moderate or severe grade adverse events in individuals having the immune receptor haplotype of the subject.
[0034] As used herein, "haplotype" refers to a set of variable alleles that tend to be inherited together owing to genetic linkage and population structure.
[0035] As used herein, the terms "haplotype group" and "haplogroup" refer to a set of haplotypes that are similar or identical to one another in terms of co-inherited variable gene alleles. The haplotype grouping is robust to minor differences arising from recombination (which can break up a haplotype and blend it with a different haplotype), noise in a sequencing assay (which has the potential lead to a failure to detect an allele in a sample), or random mutation or genetic events that lead to the emergence of a novel allele in an individual, without affecting the other alleles in the haplotype.
[0036] In addition to sequencing cDNA or mRNA of expressed TRB mRNA or gDNA of rearranged TRB genes, TRBV gene alleles and haplotype groups may be determined using other techniques including, but not limited to, real-time PCR analysis, whole genome sequencing, restriction fragment length analysis, and / or application of such methods to identify polymorphic sites that are genetically linked to TRBV alleles but are not within TRBV genes. TRBV gene alleles and haplotype groups may also be determined by combining TRB sequence information with allele characterization derived from a combination of such techniques.
[0037] In some embodiments, methods are provided for predicting the likelihood of an immune system-mediated adverse event of a subject to immunotherapy by identifying the TRBV haplotype group of the subject. In some embodiments, such methods comprise performing sequencing of a TRB repertoire from a sample of a subject, identifying the set of TRBV gene alleles in the sequence data and detecting the TRBV haplotype for the subject from the identified TRBV gene alleles. In some embodiments, methods are provided for identifying TRBV haplotype groups that have a protective effect against immune system-mediated adverse events following immunotherapy.
[0038] For determining a TRBV haplotype of a sample, a set of sequencing reads representing TCR beta chains of T cells derived from sequencing of expressed TCR beta mRNA or rearranged TCR beta gDNA, where the set of sequencing reads includes coverage of the FR1, CDR1, FR2, CDR2, FR3, and CDR3 domains of the rearranged TCR beta chain. In some embodiments, provided compositions and methods for amplification and sequencing the TCR beta repertoire in a sample from the FR1 through the CDR3 domains are of use in determining a TRBV haplotype of a sample. In other embodiments, amplification and sequencing techniques which provide coverage for FR1 through CDR3 domains, including for example 5' RACE methods, are of use in determining a TRBV haplotype of a sample.
[0039] In some embodiments, identifying TRBV haplotypes involves generating a clone summary file containing the sequence and features of all the TRB clonotypes detected in the sample. The first step in the procedure uses this file as input to identify the set of TRBV gene alleles present in the sample. It includes of the following operations: 1. Count the number of clones possessing each unique V gene sequence in the clone summary file. Each unique V gene sequence potentially represents a different V gene allele, subject to further qualification. The V gene sequence is defined as the portion of the reported TRB sequence 5' of the CDR3 region encompassing the FR1, CDR1, FR2, CDR2, and FR3 regions of the TRB V gene. (In Ion Reporter, the V gene sequence is provided in the "sequence" column of the clone summary file and the CDR3 region is defined in the "CDR3 NT" column.) 2. Aggregate the unique, counted V gene sequences from 1) into groups based on their annotated V gene identity. Ion Reporter informatics pipeline annotates the V gene identity via BLAST-based alignment to V gene sequences in the IMGT database, but other equivalent methods could be used. 3. For each V gene sequence group, perform the following steps: a. Identify the top two most frequent V gene sequences, using the clone counting results from 1). Use these two most frequent sequences as input to step 3b. If there is only one unique sequence detected then use that single sequence as input to step 3b. b. Filter the sequences from 3a) based on the level of support for that sequence in the data. This includes the total number of clones having that sequence as well as the fraction of clones having the annotated variable gene that also possess that variable gene sequence. In one embodiment, a qualified V gene sequence must be supported by a minimum number of 5 clones to be found at a minimum frequency of .01 within sequences having the same annotated V gene identity. For example, if there are 1000 clones having a V gene annotated as "TRBV5", then for a V gene sequence to be qualified it must be present in at least 10 of the 1000 clones (10 / 1000 =.01 frequency). 4. The set of sequences retained after step 3b represent the set of TRBV alleles detected in a sample, also known as the TRBV haplotype. This haplotype will be compared to a reference set of haplotypes produced for example using the procedure below. 5. Generation of reference TRBV haplotype set. Prepare and sequence TRB chains from a set of samples representing individuals of known adverse event status to obtain the sequence of at least 1000 clones, per the output of the Ion Reporter workflow (other appropriate values for this minimum number of clones could be, without limitation, 500, 2000 or 5000 clones). Any sample containing a plurality of T cells is appropriate for library generation and sequencing, though a peripheral blood lymphocyte (PBL) sample is a particularly suitable input type. Using steps 1-4 above, determine the TRBV allele haplotype of each sample. Write the TRBV allele haplotypes in a table format such the each row represents a different sample and each column indicates a unique V gene sequence (allele). If a given allele was detected in a sample, indicate via "1" in the table; else indicate with 0. a. Note: We produced a reference TRBV haplotype set by sequencing of PBL from 54 individuals with annotated adverse events following checkpoint blockade immunotherapy. 6. Perform principal component analysis using the table produced in 5) and extract the top two components. 7. Using the top two component values from 6), perform k-means clustering to identify the number of haplotype groups in the data. In one embodiment, the number of groups used for k-means clustering was 4. In other embodiments, this number may differ depending on the nature of the sample set. 8. For each haplotype group identified in 7), determine the frequency and severity of adverse events for samples within that group based on the prior annotations. 9. Assign the query TRBV haplotype from 4) to the most similar haplotype group from 8) using k-nearest neighbors classification, or other suitable machine learning approach. 10. The estimated likelihood of adverse events for the query sample is indicated by the frequency of adverse events within the assigned haplotype group. 11. In other embodiments, the accuracy of the estimate from step 10 may be further improved by incorporation of HLA typing data for the samples, such as for example is produced by the One Lambda HLA typing assay using the S5 530 chip.
[0040] As indicated above in (7), the number of haplotype groups identified using k-means clustering may differ depending on the nature of the sample set. In some embodiments, 4 haplotype groups (or clusters) are identified in a sample set. In other embodiments, 5 haplotype groups, 6 haplotype groups, 7 haplotype groups, 8 haplotype groups, 9 haplotype groups or 10 haplotype groups are identified. In some embodiments, 10-15 haplotype groups, 12-18 haplotype groups, 5-10 haplotype groups or 10-20 haplotype groups are identified.
[0041] As shown in Example 11, TRBV allele typing indicated the presence of four main haplotype groups in a set of 55 Caucasian individuals who experienced adverse events following cancer immunotherapy with checkpoint blockade agents. All samples were graded for adverse events by using standard criteria as defined for example in Common Terminology Criteria for Adverse Events, version 3.0, from the Cancer Therapy Evaluation Program (ctep.cancer.gov). Haplotype Group 2, accounting for 37% of the cohort, appeared to be protected from severe adverse events (grade 3 or 4) following the immunotherapy. Stratifying the results by the checkpoint blockade agent further supports Haplotype Group 2 as protective against adverse events following treatment with Ipilimumab and Nivolumab.
[0042] As shown in Example 12, TRBV allele typing indicated the presence of six main haplotype groups in a set of 81 Caucasian individuals who experienced adverse events following cancer immunotherapy with checkpoint blockade agents. This sample set combines samples analyzed in Example 11 (cohort 1) with an additional 27 samples (cohort 2). All samples were graded for adverse events by using standard criteria as defined for example in Common Terminology Criteria for Adverse Events, version 3.0, from the Cancer Therapy Evaluation Program. From this analysis, haplotype group 2, accounting for 33% of the cohort, appeared to be protected from severe adverse events (grade 3 or 4) following the immunotherapy. Using two model approaches, principal component analysis and k-means clustering with cohort 1 samples was able to predict adverse events in cohort 2 as demonstrated by analysis of receiver-operator characteristic curves. Haplotype group 2 members have fewer unique alleles and fewer uncommon alleles (present in <50% of the population) than members of other haplotype groups. There was a significant positive correlation between the number of uncommon alleles and the frequency of severe immune-related adverse events.
[0043] As described herein, method and compositions provided are used to identify and characterize novel or non-canonical TCR alleles, such as TRBV alleles, of a subject's immune repertoire. Novel or non-canonical TRBV alleles can help define differences in haplotype groups. In some embodiments, novel or non-canonical TRBV alleles are identified and / or characterized prior to performing haplotype analysis.
[0044] If an individual possesses a putatively novel or non-canonical variable allele, clones utilizing the allele will present as having a systematic mismatch to the IMGT database. Given that each clone is readily distinguishable from one another in sequence space owing to the diversity of the CDR3 region, the number of clones having a particular systematic mismatch is indicative of the minimum number of unique template molecules supporting a putative non-IMGT allele. Bone fide novel alleles will be found on a plurality of clones, each possessing a distinct CDR3 nucleotide sequence, while mismatches owing to random PCR error or sequencing error will not be found on multiple clones within a repertoire. In some embodiments, to report an allele for downstream haplotype analysis, either a putative novel allele or canonical IMGT allele, the allele should be present on a minimum of 5 clones (clone support) and make up at least 5% of the sequences obtained for that variable gene (frequency support). Up to two alleles of a particular variable gene may be detected in a single sample. If more than two potential alleles are detected for a particular variable gene, only the two alleles having the greatest clone support are reported for the sample.
[0045] The following embodiments relating to methods for treating a subject are disclosed but not claimed. In some embodiments, methods are provided for treating a subject based on characterizing the immune repertoire haplotype of the subject before receiving the treatment. In some embodiments, provided methods comprise performing sequencing of target immune receptor nucleic acid template molecules derived from a biological sample from a subject with cancer, wherein the target immune receptor nucleic acid template molecules comprise FR1, CDR1, FR2, CDR2, FR3, and CDR3 coding regions of the target immune receptor and the sequencing is by next generation sequencing; determining the sequence of the target immune receptor repertoire of the sample based on the sequencing; identifying the immune receptor haplotype of the subject from the determined sequences; and treating the subject with an immunotherapy associated with no or low grade adverse events in individuals having the immune receptor haplotype of the subject. In some embodiments, provided methods further comprise comparing the identified immune receptor haplotype of the subject to a reference set of immune receptor haplotypes of individuals with annotated adverse events following immunotherapy treatments. In some embodiments, provided methods, prior to the sequencing, further comprise performing a multiplex amplification reaction to amplify target immune receptor nucleic acid template molecules, wherein the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of V gene primers directed to a majority of V genes of the target immune receptor.
[0046] In some embodiments, provided methods comprise performing sequencing of target immune receptor nucleic acid template molecules derived from a biological sample from a subject with cancer, wherein the target immune receptor nucleic acid template molecules comprise FR1, CDR1, FR2, CDR2, FR3, and CDR3 coding regions of the target immune receptor and the sequencing is by next generation sequencing; determining the sequence of the target immune receptor repertoire of the sample based on the sequencing; identifying the immune receptor haplotype of the subject from the determined sequences; comparing the identified immune receptor haplotype of the subject to a reference set of immune receptor haplotypes of individuals with annotated adverse events following immunotherapy treatments; and predicting the susceptibility of the subject to experiencing severe immune-related adverse events following immunotherapy, such as for example an immunotherapy comprising one or more checkpoint blockade agents. In some embodiments, provided methods comprise predicting or assessing the likelihood of severe adverse events following immunotherapy by determining the number of uncommon alleles (eg., <50% frequency in the population) in a subject's TRB V repertoire. In some embodiments, a TRBV haplotype with fewer unique alleles and / or fewer uncommon alleles in a pre-immunotherapy sample is predictive of the subject avoiding severe adverse events following immunotherapy. As shown herein, at least one haplotype group is associated with fewer unique alleles and fewer uncommon alleles than members of other haplotype groups and subjects of this haplotype group appeared to be protected from severe adverse events following immunotherapy with checkpoint blockade agents.
[0047] In other embodiments, methods are provided for treating a subject based on detecting in a sample from a subject at least one allele or gene identified as a member of an immune repertoire haplotype predictive of severe adverse event susceptibility. In some embodiments, following characterization of an immune repertoire haplotype predictive of severe adverse event susceptibility as described herein, predicting or assessing the likelihood severe adverse events associated with immunotherapy may be determined by detecting one or more alleles or genes of the haplotype in the subject. Detecting one or more alleles or genes of the haplotype can be through use of methods provided herein or by canonical methods for detecting alleles or genes, including without limitation real-time quantitative or semi-quantitative PCR analysis, comparative hybridization, Sanger sequencing, RFLP analysis, de novo whole or local genome assembly using next generation sequencing.
[0048] As used herein, "immunotherapy" refers to a type of therapeutic treatment that uses agents which directly or indirectly stimulate or suppress an immune response to treat or prevent a disease or disorder, such as without limitation cancer, infection, autoimmune disease. Examples of immunotherapeutic agents include cytokines and other nonspecific immune stimulators (e.g., interleukins, interferons, BCG), vaccines (e.g., dendritic cell vaccines, tumor cell vaccines, antigen vaccines), monoclonal antibody-based agents, immune checkpoint inhibitors or blockade agents, CAR (chimeric antigen receptor)-T cells, adoptive T-cells, and other cell-based immunotherapeutics. Immune checkpoint inhibitors or blockade agents target molecules on certain immune cells that need to be activated or inactivated to start an immune response. Immune checkpoint proteins include PD-1, PD-L1, and CTLA-4. Monoclonal antibody-based immune checkpoint blockade agents include, without limitation, PD-1 inhibitors pembrolizumab, nivolumab, cemiplimab; PD-L1 inhibitors atezolizumab, avelumab, durvalumab; and CTLA-4 inhibitor ipilimumab. Agents that block the activity of these checkpoint proteins can lead to side effects.
[0049] In some embodiments of provided methods, the target immune receptor is a T cell receptor (TCR) selected from TCR alpha, TCR beta, TCR gamma, and TCR delta. In some embodiments, target immune receptor nucleic acid template molecules are derived from RNA from the subject sample. In other embodiments, target immune receptor nucleic acid template molecules comprise genomic DNA from the subject sample having rearranged VDJ or VJ gene segments. In some embodiments, the biological sample is a peripheral blood sample. In some embodiments, the immunotherapy comprises a checkpoint blockade agent. In some embodiments, the immunotherapy comprises a dendritic cell vaccine or a tumor cell vaccine.
[0050] In some embodiments of provided methods, determining the target sequence includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads. In some embodiments, the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads. In other embodiments, the combination of productive reads and rescued productive reads is at least 60% of the sequencing reads.
[0051] In certain embodiments of provided methods, plurality of amplification primer pairs are used and the plurality of primer pairs includes one or more primers that anneal to at least a portion of the C gene portion of the target immune receptor nucleic acid template molecules. In other embodiments, the plurality of amplification primer pairs includes at least 10 primers that anneal to at least a portion of the J gene portion of the target immune receptor nucleic acid template molecules. In some embodiments, the plurality of amplification primers includes a plurality of V gene primers that anneal to at least a portion of the FR1 regions of the target immune receptor nucleic acid template molecules. In some embodiments, the plurality of amplification primers includes a plurality of V gene primers that anneal to at least a portion of the FR3 regions of the target immune receptor nucleic acid template molecules.
[0052] In some embodiments, a multiplex next generation sequencing workflow is used for effective detection and analysis of the immune repertoire in a subject's sample. Provided methods, compositions, systems, and kits are for use in high accuracy amplification and sequencing of immune cell receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) in monitoring and resolving complex immune cell repertoire(s) in a subject. The target immune cell receptor genes have undergone rearrangement (or recombination) of the VDJ or VJ gene segments, the gene segments depending on the particular receptor gene (e.g., TCR beta or TCR alpha). In certain embodiments, the present disclosure provides methods, compositions, and systems that use nucleic acid amplification, such as polymerase chain reaction (PCR), to enrich expressed variable regions of immune receptor target nucleic acid for subsequent sequencing. In certain embodiments, the present disclosure provides methods, compositions, and systems that use nucleic acid amplification, such as PCR, to enrich rearranged target immune cell receptor gene sequences from gDNA for subsequent sequencing. In certain embodiments, the present disclosure also provides methods and systems for effective identification and removal of amplification or sequencing-derived error(s) to improve read assignment accuracy and lower the false positive rate. In particular, provided methods described herein may improve accuracy and performance in sequencing applications with nucleotide sequences associated with genomic recombination and high variability. In some embodiments, methods, compositions, systems, and kits provided herein are for use in amplification and sequencing of the complementarity determining regions (CDRs) of an expressed immune receptor in a sample. In some embodiments, methods, compositions, systems, and kits provided herein are for use in amplification and sequencing of the CDRs of rearranged immune cell receptor gDNA in a sample. Thus, provided herein are multiplex immune cell receptor expression compositions and immune cell receptor gene-directed compositions for multiplex library preparation, use in conjunction with next generation sequencing technologies and workflow solutions (e.g., manual or automated), for effective detection and characterization of the immune repertoire in a sample.
[0053] The CDRs of a TCR or BCR results from genomic DNA undergoing recombination of the V(D)J gene segments as well as addition and / or deletion of nucleotides at the gene segment junctions. Recombination of the V(D)J gene segments and subsequent hypermutation events leads to extensive diversity of the expressed immune cell receptors. With the stochastic nature of V(D)J recombination, it is often the case that rearrangement of the T or B cell receptor genomic DNA will fail to produce a functional receptor, instead producing what is termed an "unproductive" rearrangement. Typically, unproductive rearrangements have out-of-frame Variable and Joining coding segments, and lead to the presence of premature stop codons and synthesis of irrelevant peptides. Unproductive TCR or BCR gene rearrangements are generally rare in cDNA-based repertoire sequencing for a number of biological or physiological reasons such as: 1) nonsense-mediated decay, which destroys mRNA containing premature stop codons, 2) B and T cell selection, where only B and T cells with a functional receptor survive, and 3) allelic exclusion, where only a single rearranged receptor allele is expressed in any given B or T cell.
[0054] Accordingly, in some embodiments, methods and compositions provided herein are used for amplifying the recombined, expressed variable regions of immune cell receptor mRNA, eg TCR and BCR mRNA. In some embodiments, RNA extracted from biological samples is converted to cDNA. Multiplex amplification is used to enrich for a portion of TCR or BCR cDNA which includes at least a portion of the variable region of the receptor. In some embodiments, the amplified cDNA includes one or more complementarity determining regions CDR1, CDR2, and / or CDR3 for the target receptor. In some embodiments, the amplified cDNA includes one or more complementarity determining regions CDR1, CDR2, and / or CDR3 for TCR beta.
[0055] TCR and BCR sequences can also appear as unproductive rearrangements from errors introduced during amplification reactions or during sequencing processes. For example, an insertion or deletion (indel) error during a target amplification or sequencing reaction can cause a frameshift in the reading frame of the resulting coding sequence. Such a change may result in a target sequence read of a productive rearrangement being interpreted as an unproductive rearrangement and discarded from the group of identified clonotypes. Accordingly, in some embodiments, methods and systems provided herein include processes for identification and / or removing PCR or sequencing-derived error from the determined immune receptor sequence.
[0056] In some embodiments, methods and compositions provided are used for amplifying the rearranged variable regions of immune cell receptor gDNA, e.g., rearranged TCR and BCR gene DNA. Multiplex amplification is used to enrich for a portion of rearranged TCR or BCR gDNA which includes at least a portion of the variable region of the receptor. In some embodiments, the amplified gDNA includes one or more complementarity determining regions CDR1, CDR2, and / or CDR3 for the target receptor. In some embodiments, the amplified gDNA includes one or more complementarity determining regions CDR1, CDR2, and / or CDR3 for TCR beta. In some embodiments, the amplified gDNA includes primarily CDR3 for the target receptor, e.g., CDR3 for TCR beta.
[0057] As used herein, "immune cell receptor" and "immune receptor" are used interchangeably.
[0058] As used herein, the terms "complementarity determining region" and "CDR" refer to regions of a T cell receptor or an antibody where the molecule complements an antigen's conformation, thereby determining the molecule's specificity and contact with a specific antigen. In the variable regions of T cell receptors and antibodies, the CDRs are interspersed with regions that are more conserved, termed framework regions (FR). Each variable region of a T cell receptor and an antibody contains 3 CDRs, designated CDR1, CDR2 and CDR3, and also contains 4 framework sub-regions, designated FR1, FR2, FR3 and FR4.
[0059] As used herein, the term "framework" or "framework region" or "FR" refers to the residues of the variable region other than the CDR residues as defined herein. There are four separate framework sub-regions that make up the framework: FR1, FR2, FR3, and FR4.
[0060] The particular designation in the art for the exact location of the CDRs and FRs within the receptor molecule (TCR or immunoglobulin) varies depending on what definition is employed. Unless specifically stated otherwise, the IMGT designations are used herein in describing the CDR and FR regions (see Brochet et al. (2008) Nucleic Acids Res. 36:W503-508). As one example of CDR / FR amino acid designations, the residues that make up the FRs and CDRs of T cell receptor beta have been characterized by IMGT as follows: residues 1-26 (FR1), 27-38 (CDR1), 39-55 (FR2), 56-65 (CDR2), 66-104 (FR3), 105-117 (CDR3), and 118-128 (FR4).
[0061] Other well-known standard designations for describing the regions include those found in Kabat et al., (1991) Sequences of Proteins of Immunological Interest, 5th Ed. Public Health Service, National Institutes of Health, Bethesda, Md., and in Chothia and Lesk (1987) J. Mol. Biol. 196:901-917. As one example of CDR designations, the residues that make up the six immunoglobulin CDRs have been characterized by Kabat as follows: residues 24-34 (CDRL1), 50-56 (CDRL2) and 89-97 (CDRL3) in the light chain variable region and 31-35 (CDRH1), 50-65 (CDRH2) and 95-102 (CDRH3) in the heavy chain variable region; and by Chothia as follows: residues 26-32 (CDRL1), 50-52 (CDRL2) and 91-96 (CDRL3) in the light chain variable region and 26-32 (CDRH1), 53-55 (CDRH2) and 96-101 (CDRH3) in the heavy chain variable region.
[0062] The term "T cell receptor" or "T cell antigen receptor" or "TCR," as used herein, refers to the antigen / MHC binding heterodimeric protein product of a vertebrate, e.g. mammalian, TCR gene complex, including the human TCR alpha, beta, gamma and delta chains. For example, the complete sequence of the human TCR beta locus has been sequenced, see, for example, Rowen et al. (1996) Science 272:1755-1762; the human TCR alpha locus has been sequenced and resequenced, see, for example, Mackelprang et al. (2006) Hum Genet. 119:255-266; and see, for example, Arden (1995) Immunogenetics 42:455-500 for a general analysis of the T-cell receptor V gene segment families.
[0063] The term "antibody" or immunoglobulin" or "B cell receptor" or "BCR," as used herein, is intended to refer to immunoglobulin molecules comprised of four polypeptide chains, two heavy (H) chains and two light (L) chains (lambda or kappa) inter-connected by disulfide bonds. An antibody has a known specific antigen with which it binds. Each heavy chain of an antibody is comprised of a heavy chain variable region (abbreviated herein as HCVR, HV or VH) and a heavy chain constant region. The heavy chain constant region is comprised of three domains, CH1, CH2 and CH3. Each light chain is comprised of a light chain variable region (abbreviated herein as LCVR or VL or KV or LV to designate kappa or lambda light chains) and a light chain constant region. The light chain constant region is comprised of one domain, CL.
[0064] As noted, the diversity of the TCR and BCR chain CDRs is created by recombination of germline variable (V), diversity (D), and joining (J) gene segments, as well as by independent addition and deletion of nucleotides at each of the gene segment junctions during the process of TCR and BCR gene rearrangement. In the rearranged nucleic acid encoding a TCR beta and a TCR delta, for example, CDR1 and CDR2 are found in the V gene segments and CDR3 includes some of the V gene segment, and the D and J gene segments. In the rearranged nucleic acid encoding a TCR alpha and a TCR gamma, CDR1 and CDR2 are found in the V gene segments and CDR3 includes some of the V gene segment and the J gene segment. In the rearranged nucleic acid encoding a BCR heavy chain, CDR1 and CDR2 are found in the V gene segment and CDR3 includes some of the V gene segment and the D and J gene segments. In the rearranged nucleic acid encoding a BCR light chain, CDR1 and CDR2 are found in the V gene segment and CDR3 includes some of the V gene segment and the J gene segment.
[0065] In some embodiments, a multiplex amplification reaction is used to amplify cDNA derived from mRNA expressed from rearranged TCR or BCR genomic DNA. In some embodiments, a multiplex amplification reaction is used to amplify at least a portion of a TCR or BCR CDR from cDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify at least two CDRs of a TCR or BCR from cDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify at least three CDRs of a TCR or BCR from cDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of the TCR or BCR CDRs expressed in the sample. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least 3 CDRs is used to identify and characterize novel TCR or BCR alleles. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least 3 CDRs is used to identify and characterize novel TCR or BCR alleles.
[0066] In some embodiments, a multiplex amplification reaction is used to amplify TCR or BCR genomic DNA having undergone V(D)J rearrangement. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecule(s) comprising at least a portion of a TCR or BCR CDR from gDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecule(s) comprising at least two CDRs of a TCR or BCR from gDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecules comprising at least three CDRs of a TCR or BCR from gDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of the rearranged TCR or BCR CDRs in the sample. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least CDR3 is used to identify and characterize novel TCR or BCR alleles. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least 3 CDRs is used to identify and characterize novel TCR or BCR alleles.
[0067] In the multiplex amplification reactions, each primer set used target a same TCR or BCR region however the different primers in the set permit targeting the gene's different V(D)J gene rearrangements. For example, the primer set for amplification of the expressed TCR beta or the rearranged TCR beta gDNA are all designed to target the same region(s) from TCR beta mRNA or TCR beta gDNA, respectively, but the individual primers in the set lead to amplification of the various TCR beta VDJ gene combinations. In some embodiments, at least one primer or primer set is directed to a relatively conserved region (eg, a portion of the C gene) of an immune receptor gene and the other primer set includes a variety of primers directed to a more variable region of the same gene (eg, a portion of the V gene). In other embodiments, at least one primer set includes a variety of primers directed to at least a portion of J gene segments of an immune receptor gene and the other primer set includes a variety of primers directed to at least a portion of V gene segments of the same gene.
[0068] In some embodiments, a multiplex amplification reaction is used to amplify cDNA derived from mRNA expressed from rearranged TCR genomic DNA, including rearranged TCR beta, TCR alpha, TCR gamma, and TCR delta genomic DNA. In some embodiments, at least a portion of a TCR CDR, for example CDR3, is amplified from cDNA in a multiplex amplification reaction. In some embodiments, at least two CDR portions of TCR are amplified from cDNA in a multiplex amplification reaction. In certain embodiments, a multiplex amplification reaction is used to amplify at least the CDR1, CDR2, and CDR3 regions of a TCR cDNA. In some embodiments, the resulting amplicons are used to determine the expressed TCR CDR nucleotide sequence.
[0069] In some embodiments, a multiplex amplification reaction is used to amplify rearranged TCR genomic DNA, including rearranged TCR beta, TCR alpha, TCR gamma, and TCR delta genomic DNA. In some embodiments, at least a portion of a TCR CDR, for example CDR3, is amplified from gDNA in a multiplex amplification reaction. In some embodiments, at least two CDR portions of TCR are amplified from gDNA in a multiplex amplification reaction. In certain embodiments, a multiplex amplification reaction is used to amplify at least the CDR1, CDR2, and CDR3 regions of a rearranged TCR gDNA. In some embodiments, the resulting amplicons are used to determine the expressed TCR CDR nucleotide sequence.
[0070] In some embodiments, multiplex amplification reactions are performed with primer sets designed to generate amplicons which include the expressed CDR1, CDR2, and / or CDR3 regions of the target immune receptor mRNA. In some embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR1 of a V gene and (ii) at least one primer directed to a portion of the C gene of the target immune receptor. In other embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR2 of a V gene and (ii) at least one primer directed to a portion of the C gene of the target immune receptor. In other embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR3 of a V gene and (ii) at least one primer directed to a portion of the C gene of the target immune receptor. In some embodiments, the C gene-directed primer is directed C gene coding sequences within about 200 nucleotides of the 5' end of the C gene. In some embodiments, the C gene-directed primer is directed C gene coding sequences within about 150 nucleotides of the 5' end of the C gene. In some embodiments, the C gene-directed primer is directed C gene coding sequences within about 100 nucleotides of the 5' end of the C gene. In some embodiments, the C gene-directed primer is directed C gene coding sequences within about 50 nucleotides, within about 50 to about 150, within about 75 to about 175, or within about 100 to about 200 nucleotides of the 5' end of the C gene.
[0071] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR1 region and (ii) at least one primer which anneals to a portion of the constant (C) gene to amplify TCR cDNA such that the resultant amplicons include the CDR1, CDR2, and CDR3 coding portions of the TCR mRNA. In certain embodiments, an FR1-directed primer set is combined with a set of at least two C gene-directed primers to generate amplicons which include at least the CDR1, CDR2, and CDR 3 coding portions of a TCR mRNA. For example, exemplary primers specific for TCR beta (TRB) V gene FR1 regions are shown in Table 2 and exemplary primers specific for TRB C genes are shown in Table 4.
[0072] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR2 region and (ii) at least one primer which anneals to a portion of the C gene to amplify TCR cDNA such that the resultant amplicons include the CDR2 and CDR3 coding portions of the TCR mRNA. In certain embodiments, such a FR2-directed primer set is combined with at least two C gene-directed primers to generate amplicons which include the CDR2 and CDR3 coding portions of a TCR mRNA. Exemplary FR2-directed primers include the BIOMED-2 primers developed and standardized by a consortium of European academic laboratories and research hospitals (van Dongen et al. (2003) Leukemia 17:2257-2327) and shown in Table 6. Exemplary primers specific for TRB C genes are shown in Table 4.
[0073] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR3 region and (ii) at least one primer which anneals to a portion of the C gene to amplify TCR cDNA such that the resultant amplicons include primarily the CDR3 coding portion of the TCR mRNA. In certain embodiments, such a FR3-directed primer set is combined with at least two C gene-directed primers to generate amplicons with the CDR 3 coding portion of a TCR mRNA. For example, exemplary primers specific for TCR beta (TRB) V gene FR3 regions are shown in Table 3 and exemplary primers specific for TRB C genes are shown in Table 4.
[0074] In some embodiments, multiplex amplification reactions are performed with primer sets designed to generate amplicons which include the CDR1, CDR2, and / or CDR3 regions of the target immune receptor mRNA or rearranged gDNA. In some embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR1 of a V gene and (ii) one set of primers in which each primer is directed to at least a portion of the J gene of the target immune receptor. In other embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR2 of a V gene and (ii) one set of primers in which each primer is directed to at least a portion of the J gene of the target immune receptor. In other embodiments, multiplex amplification reactions are performed using (i) one set of primers in which each primer is directed to at least a portion of the framework region FR3 of a V gene and (ii) one set of primers in which each primer is directed to at least a portion of the J gene of the target immune receptor.
[0075] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR1 region and (ii) a set of primers which anneal to a portion of the J gene to amplify TCR nucleic acid such that the resultant amplicons include the CDR1, CDR2, and CDR3 coding portions of the TCR mRNA or rearranged gDNA. For example, exemplary primers specific for TCR beta (TRB) V gene FR1 regions are shown in Table 2 and exemplary primers specific for TRB J genes are shown in Table 5.
[0076] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR2 region and (ii) a set of primers which anneal to a portion of the J gene to amplify TCR nucleic acid such that the resultant amplicons include the CDR2 and CDR3 coding portions of the TCR mRNA or rearranged gDNA. For example, exemplary primers specific for TRB V gene FR2 regions are shown in Table 6 and exemplary primers specific for TRB J genes are shown in Table 5.
[0077] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the V gene FR3 region and (ii) a set of primers which anneal to a portion of the J gene to amplify TCR nucleic acid such that the resultant amplicons include primarily the CDR3 coding portion of the TCR mRNA or rearranged gDNA. For example, exemplary primers specific for the TRB V gene FR3 regions are shown in Table 3 and exemplary primers specific for TRB J genes are shown in Table 5.
[0078] In some embodiments, provided are compositions for multiplex amplification of at least a portion of an expressed TCR or BCR variable region. In some embodiments, the composition comprises a plurality of sets of primer pair reagents directed to a portion of a V gene framework region and a portion of a constant (C) gene of rearranged target immune receptor genes selected from the group consisting of TCR beta, TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, and immunoglobulin light chain kappa. In some embodiments, the composition comprises a plurality of sets of primer pair reagents directed to a portion of a V gene framework region and a portion of a J gene of rearranged target immune receptor genes selected from the group consisting of TCR beta, TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, and immunoglobulin light chain kappa.
[0079] Amplification by PCR is performed with at least two primers. For the methods provided herein, a set of primers is used that is sufficient to amplify all or a defined portion of the variable sequences at the locus of interest, which locus may include any or all of the aforementioned TCR and Immunoglobulin loci. In some embodiments, various parameters or criteria outlined herein may be used to select the set of target-specific primers for the multiplex amplification.
[0080] In some embodiments, primer sets used in the multiplex reactions are designed to amplify at least 50% of the known expressed or gDNA rearrangements at the locus of interest. In certain embodiments, primer sets used in the multiplex reactions are designed to amplify at least 75%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or more of the known expressed or gDNA rearrangements at the locus of interest. For example, use of at least 49 forward primers of Table 2, each directed to a portion of the FR1 region from different TCR beta V genes, in combination with at least one of the reverse primers of Table 4 directed to a portion of the TCR beta C gene will amplify at least 50% of the known expressed TCR beta rearrangements. For another example, use of 64 forward primers of Table 2, each directed to a portion of the FR1 region from different TCR beta V genes, in combination with two reverse primers of Table 4, each directed to a portion of the TCR beta C genes, will amplify all of the currently known expressed TCR beta rearrangements. For another example, use of 59 forward primers of Table 3, each directed to a portion of the FR3 region from different TCR beta V genes, in combination with two reverse primers of Table 4, each directed to a portion of the TCR beta C genes, will amplify all of the currently known expressed TCR beta rearrangements. For another example, use of 59 forward primers of Table 3, each directed to a portion of the FR3 region from different TCR beta V genes, in combination with 16 reverse primers of Table 5, each directed to a portion of different TCR beta J genes, will amplify all of the currently known expressed or gDNA TCR beta rearrangements. In some embodiments, use of 59 forward primers of Table 3, each directed to a portion of the FR3 region from different TCR beta V genes, in combination with 14 reverse primers of Table 5, each directed to a portion of different TCR beta J genes, will amplify all of the currently known expressed or gDNA TCR beta rearrangements For another example, use of 64 forward primers of Table 2, each directed to a portion of the FR1 region from different TCR beta V genes, in combination with 16 reverse primers of Table 5, each directed to a portion of different TCR beta J genes, will amplify all of the currently known expressed or gDNA TCR beta rearrangements. In other embodiments, use of 64 forward primers of Table 2, each directed to a portion of the FR1 region from different TCR beta V genes, in combination with 14 reverse primers of Table 5, each directed to a portion of different TCR beta J genes, will amplify all of the currently known expressed or gDNA TCR beta rearrangements.
[0081] For example, such a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 49, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR1 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR1 regions is combined with at least 1 forward primer directed to a sequence corresponding to at least a portion of the constant gene of the same TCR gene. In some embodiments, the plurality of reverse primers directed to the TCR V gene FR1 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 forward primers each directed to a sequence corresponding to at least a portion to the constant gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR1 directed primers may be the forward primers and the TCR C gene-directed primer(s) may be the reverse primer(s). Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 49, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR1 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR1 regions is combined with at least 1 reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same TCR gene. In some embodiments, the plurality of forward primers directed to the TCR V gene FR1 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 reverse primers each directed to a sequence corresponding to at least a portion to the C gene of the same TCR gene. In some embodiments, such FR1 and C gene amplification primer sets may be directed to TCR beta gene sequences. In some preferred embodiments, about 60 to about 70 forward primers directed to different TRB V gene FR1 regions are combined with 2 reverse primers directed to a portion of the TRB C gene. In some preferred embodiments, the forward primers directed to TRB V gene FR1 regions are selected from those listed in Table 2 and the reverse primers directed to the TRB C gene are selected from those listed in Table 4. In other embodiments, the FR1 and C gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, or immunoglobulin light chain kappa gene sequences.
[0082] In some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR2 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR2 regions is combined with at least 1 forward primer directed to a sequence corresponding to at least a portion of the C gene of the same TCR gene. In some embodiments, the plurality of reverse primers directed to the TCR V gene FR2 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 forward primers each directed to a sequence corresponding to at least a portion to the C gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR2 directed primers may be the forward primers and the TCR C gene-directed primer(s) may be the reverse primer(s). Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR2 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR2 regions is combined with at least 1 reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same TCR gene. In some embodiments, the plurality of forward primers directed to the TCR V gene FR2 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 reverse primers each directed to a sequence corresponding to at least a portion to the C gene of the same TCR gene. In some embodiments, such FR2 and C gene amplification primer sets may be directed to TCR beta gene sequences. In some embodiments, about 20 to about 30 forward primers directed to different TRB V gene FR2 regions are combined with 2 reverse primers directed to a portion of the TRB C gene. In some preferred embodiments, the forward primers directed to TRB V gene FR2 regions are selected from those listed in Table 6 and the reverse primers directed to the TRB C gene are selected from those listed in Table 4. In other embodiments, the FR2 and C gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, or immunoglobulin light chain kappa gene sequences.
[0083] In some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR3 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR3 regions is combined with at least 1 forward primer directed to a sequence corresponding to at least a portion of the C gene of the same TCR gene. In some embodiments, the plurality of reverse primers directed to the TCR V gene FR3 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 forward primers each directed to a sequence corresponding to at least a portion to the C gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR3 directed primers may be the forward primers and the TCR C gene -directed primer(s) may be the reverse primer(s). Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR3 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR3 regions is combined with at least 1 reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same TCR gene. In some embodiments, the plurality of forward primers directed to the TCR V gene FR3 regions is combined with at least 2, at least 3, at least 4, at least 5, or about 2 to about 6 reverse primers each directed to a sequence corresponding to at least a portion to the C gene of the same TCR gene. In some embodiments, such FR3 and C gene amplification primer sets may be directed to TCR beta gene sequences. In some preferred embodiments, about 55 to about 65 forward primers directed to different TRB V gene FR3 regions are combined with 2 reverse primers directed to a portion of the TRB C gene. In some preferred embodiments, the forward primers directed to TRB V gene FR3 regions are selected from those listed in Table 3 and the reverse primers directed to the TRB C gene are selected from those listed in Table 4. In other embodiments, the FR3 and C gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, and immunoglobulin light chain kappa gene sequences.
[0084] In some embodiments, such a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 49, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR1 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR1 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 forward primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR1-directed primers may be the forward primers and the TCR J gene-directed primers may be the reverse primers. Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 49, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR1 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR1 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 reverse primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments, such FR1 and J gene amplification primer sets may be directed to TCR beta gene sequences. In some preferred embodiments, about 60 to about 70 forward primers directed to different TRB V gene FR1 regions are combined with about 15 to about 20 reverse primers directed to different TRB J genes. In some preferred embodiments, about 60 to about 70 forward primers directed to different TRB V gene FR1 regions are combined with about 12 to about 18 reverse primers directed to different TRB J genes. In some preferred embodiments, the forward primers directed to TRB V gene FR1 regions are selected from those listed in Table 2 and the reverse primers directed to the TRB J gene are selected from those listed in Table 5. In other embodiments, the FR1 and J gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, or immunoglobulin light chain kappa gene sequences.
[0085] In some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR2 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR2 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 forward primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR2-directed primers may be the forward primers and the TCR J gene-directed primers may be the reverse primers. Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR2 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR2 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 reverse primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments, such FR2 and J gene amplification primer sets may be directed to TCR beta gene sequences. In some preferred embodiments, about 20 to about 30 forward primers directed to different TRB V gene FR2 regions are combined with about 15 to about 20 reverse primers directed to different TRB J genes. In some preferred embodiments, about 20 to about 30 forward primers directed to different TRB V gene FR2 regions are combined with about 12 to about 18 reverse primers directed to different TRB J genes. In some preferred embodiments, the forward primers directed to TRB V gene FR2 regions are selected from those listed in Table 6 and the reverse primers directed to the TRB J gene are selected from those listed in Table 5. In other embodiments, the FR2 and J gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, or immunoglobulin light chain kappa gene sequences.
[0086] In some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers in which each reverse primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR3 regions. In such embodiments, the plurality of reverse primers directed to the TCR V gene FR3 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 forward primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments of the multiplex amplification reactions, the TCR V gene FR3-directed primers may be the forward primers and the TCR J gene-directed primers may be the reverse primers. Accordingly, in some embodiments, a multiplex amplification reaction includes at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers in which each forward primer is directed to a sequence corresponding to at least a portion of one or more TCR V gene FR3 regions. In such embodiments, the plurality of forward primers directed to the TCR V gene FR3 regions is combined with at least 10, 12, 14, 16, 18, 20, or about 15 to about 20 reverse primers directed to a sequence corresponding to at least a portion of a J gene of the same TCR gene. In some embodiments, such FR3 and J gene amplification primer sets may be directed to TCR beta gene sequences. In some preferred embodiments, about 55 to about 65 forward primers directed to different TRB V gene FR3 regions are combined with about 15 to about 20 reverse primers directed to different TRB J genes. In some preferred embodiments, about 55 to about 65 forward primers directed to different TRB V gene FR3 regions are combined with about 12 to about 18 reverse primers directed to different TRB J genes. In some preferred embodiments, the forward primers directed to TRB V gene FR3 regions are selected from those listed in Table 3 and the reverse primers directed to the TRB J gene are selected from those listed in Table 5. In other embodiments, the FR3 and J gene amplification primer sets may be directed to TCR alpha, TCR gamma, TCR delta, immunoglobulin heavy chain, immunoglobulin light chain lambda, and immunoglobulin light chain kappa gene sequences.
[0087] In some embodiments, the concentration of the forward primer is about equal to that of the reverse primer in a multiplex amplification reaction. In other embodiments, the concentration of the forward primer is about twice that of the reverse primer in a multiplex amplification reaction. In other embodiments, the concentration of the forward primer is about half that of the reverse primer in a multiplex amplification reaction. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 5 nM to about 2000 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 50 nM to about 800 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 5 nM, about 10 nM, about 50 nM, about 100 nM, about 150 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each of the primers targeting the V gene FR region is about 50 nM to about 800 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 5 nM to about 2000 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 5 nM, about 10 nM, about 50 nM, about 100 nM, about 150 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each of the primers targeting the J gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 50 nM, about 100 nM, about 200 nM, or about 400 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 5 nM to about 2000 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 50 nM to about 800 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 600 nM, about 800 nM, about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each forward and reverse primer in a multiplex reaction is about 5 nM, about 10 nM, about 150 nM or 50 nM to about 800 nM.
[0088] In some embodiments, the V gene FR and C gene target-directed primers combine as amplification primer pairs to amplify target immune receptor cDNA sequences and generate target amplicons. Generally, the length of a target amplicon will depend upon which V gene primer set (eg, FR1, FR2, or FR3 directed primers) is paired with the C gene primer(s). Accordingly, in some embodiments, target amplicons can range from about 100 nucleotides (or bases or base pairs) in length to about 600 nucleotides (or bases or base pairs) in length. In some embodiments, target amplicons can range from about 80 nucleotides to about 600 nucleotides in length. In some embodiments, target amplicons are from about 200 to about 600 or about 300 to about 600 nucleotides in length. In some embodiments, target amplicons are about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides in length. In some embodiments, target amplicons are about 250 to about 275, about 250 to about 350, about 300 to about 350, about 310 to about 330, about 325 to about 375, about 300 to about 400, about 350 to about 400, about 350 to about 425, about 350 to about 450, about 380 to about 410, about 375 to about 425, about 400 to about 500, about 425 to about 500, about 450 to about 550, about 500 to about 600, about 400 to about 500, or about 400 to about 600 nucleotides in length. In some embodiments, target amplicons are about 80, about 100, about 120, about 140, about 200, about 250, about 275, about 300, about 320, about 350, about 375, about 400, about 425, about 450, about 500, about 550, or about 600 nucleotides in length. In some embodiments, TCR beta amplicons are about 100, about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides in length. In some embodiments, TCR beta amplicons are about 320, about 300 to about 350 or about 310 to about 330 nucleotides in length. In some embodiments, TCR beta amplicons are about 400, about 375 to about 425 or about 390 to about 410 nucleotides in length.
[0089] In some embodiments, the V gene FR and J gene target-directed primers combine as amplification primer pairs to amplify target immune receptor cDNA or rearranged gDNA sequences and generate target amplicons. Generally, the length of a target amplicon will depend upon which V gene primer set (eg, FR1, FR2, or FR3 directed primers) is paired with the J gene primers. Accordingly, in some embodiments, target amplicons can range from about 50 nucleotides to about 350 nucleotides in length. In some embodiments, target amplicons are about 50 to about 200, about 70 to about 170, about 200 to about 350, about 250 to about 320, about 270 to about 300, about 225 to about 300, about 250 to about 275, about 200 to about 235, about 200 to about 250, or about 175 to about 275 nucleotides in length. In some embodiments, TCR beta amplicons are about 80, about 60 to about 100, or about 70 to about 90 nucleotides in length. In some embodiments, TCR beta amplicons, such as those generated using V gene FR3- and J gene-directed primer pairs, are about 50 to about 200 nucleotides in length, preferably about 60 to about 160, about 65 to about 120, about 70 to about 90 nucleotides, or about 80 nucleotides in length. In some embodiments, generating amplicons of such short lengths allows the provided methods and compositions to effectively detect and analyze the immune repertoire from highly degraded gDNA template material, such as that derived from an FFPE sample.
[0090] In some embodiments, amplification primers may include a barcode sequence, for example to distinguish or separate a plurality of amplified target sequences in a sample. In some embodiments, amplification primers may include two or more barcode sequences, for example to distinguish or separate a plurality of amplified target sequences in a sample. In some embodiments, amplification primers may include a tagging sequence that can assist in subsequent cataloguing, identification or sequencing of the generated amplicon. In some embodiments, the barcode sequence(s) or the tagging sequence(s) is incorporated into the amplified nucleotide sequence through inclusion in the amplification primer or by ligation of an adapter. Primers may further comprise nucleotides useful in subsequent sequencing, e.g. pyrosequencing. Such sequences are readily designed by commercially available software programs or companies.
[0091] In some embodiments, multiplex amplification is performed with target-directed amplification primers which do not include a tagging sequence. In other embodiments, multiplex amplification is performed with amplification primers each of which include a target-directed sequence and a tagging sequence such as, for example, the forward primer or primer set includes tagging sequence 1 and the reverse primer or primer set includes tagging sequence 2. In still other embodiments, multiplex amplification is performed with amplification primers where one primer or primer set includes target directed sequence and a tagging sequence and the other primer or primer set includes a target-directed sequence but does not include a tagging sequence, such as, for example, the forward primer or primer set includes a tagging sequence and the reverse primer or primer set does not include a tagging sequence.
[0092] Accordingly, in some embodiments, a plurality of target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture with TCR or BCR directed amplification primers in which the forward and / or reverse primers include a tagging sequence and the resultant amplicons include the target TCR or BCR sequence and a tagging sequence on one or both ends. In some embodiments, the forward and / or reverse amplification primer or primer sets may also include a barcode and the one or more barcode is then included in the resultant amplicon.
[0093] In some embodiments, a plurality of target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture with TCR or BCR directed amplification primers and the resultant amplicons contain only TCR or BCR sequences. In some embodiments, a tagging sequence is added to the ends of such amplicons through, for example, adapter ligation. In some embodiments, a barcode sequence is added to one or both ends of such amplicons through, for example, adapter ligation.
[0094] Nucleotide sequences suitable for use as barcodes and for barcoding libraries are known in the art. Adapters and amplification primers and primer sets including a barcode sequence are commercially available. Oligonucleotide adapters containing a barcode sequence are also commercially available including, for example, IonXpress ™< , IonCode ™< , Ion Torrent ™< Dual Barcode, Ion AmpliSeq ™< HD Dual Barcode, and Ion Select barcode adapters (Thermo Fisher Scientific). Similarly, additional and other universal adapter / primer sequences described and known in the art (e.g., Illumina universal adapter / primer sequences, PacBio universal adapter / primer sequences, etc.) can be used in conjunction with the methods and compositions provided herein and the resultant amplicons sequenced using the associated analysis platform.
[0095] In some embodiments, two or more barcodes are added to amplicons when sequencing multiplexed samples. In some embodiments, at least two barcodes are added to amplicons prior to sequencing multiplexed samples to reduce the frequency of artefactual results (e.g., immune receptor gene rearrangements or clone identification) derived from barcode cross-contamination or barcode bleed-through between samples. In some embodiments, at least two bar codes are used to label samples when tracking low frequency clones of the immune repertoire. In some embodiments, at least two barcodes are added to amplicons when the assay is used to detect clones of frequency less than 1: 1,000. In some embodiments, at least two barcodes are added to amplicons when the assay is used to detect clones of frequency less than 1: 10,000. In other embodiments, at least two barcodes are added to amplicons when the assay is used to detect clones of frequency less than 1:20,000, less than 1:40,000, less than 1:100,000, less than 1:200,000, less than 1:400,000, less than 1:500,00, or less than 1: 1,000,000. Methods for characterizing the immune repertoire which benefit from a high sequencing depth per clone and / or detection of clones at such low frequencies include, but are not limited to, monitoring a patient with a hyperproliferative disease undergoing treatment and testing for minimal residual disease following treatment.
[0096] In some embodiments, target-specific primers (e.g., the V gene FR1-, FR2- and FR3-directed primers, the J gene directed primers, and the C gene directed primers) used in the methods of the invention are selected or designed to satisfy any one or more of the following criteria: (1) includes two or more modified nucleotides within the primer sequence, at least one of which is included near or at the termini of the primer and at least one of which is included at, or about the center nucleotide position of the primer sequence; (2) length of about 15 to about 40 bases in length; (3) Tm of from above 60°C to about 70°C; (4) has low cross-reactivity with non-target sequences present in the sample of interest; (5) at least the first four nucleotides (going from 3' to 5' direction) are non-complementary to any sequence within any other primer present in the same reaction; and (6) non-complementarity to any consecutive stretch of at least 5 nucleotides within any other produced target amplicon. In some embodiments, the target-specific primers used in the methods provided are selected or designed to satisfy any 2, 3, 4, 5, or 6 of the above criteria.
[0097] In some embodiments, the target-specific primers used in the methods of the invention include one or more modified nucleotides having a cleavable group. In some embodiments, the target-specific primers used in the methods of the invention include two or more modified nucleotides having cleavable groups. In some embodiments, the target-specific primers comprise at least one modified nucleotide having a cleavable group selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine-dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine or 5-methylcytidine.
[0098] In some embodiments, target amplicons using the amplification methods (and associated compositions, systems, and kits) disclosed herein, are used in the preparation of an immune receptor repertoire library. In some embodiments, the immune receptor repertoire library includes introducing adapter sequences to the termini of the target amplicon sequences. In certain embodiments, a method for preparing an immune receptor repertoire library includes generating target immune receptor amplicon molecules according to any of the multiplex amplification methods described herein, treating the amplicon molecule by digesting a modified nucleotide within the amplicon molecules' primer sequences, and ligating at least one adapter to at least one of the treated amplicon molecules, thereby producing a library of adapter-ligated target immune receptor amplicon molecules comprising the target immune receptor repertoire. In some embodiments, the steps of preparing the library are carried out in a single reaction vessel involving only addition steps. In certain embodiments, the method further includes clonally amplifying a portion of the at least one adapter-ligated target amplicon molecule.
[0099] In some embodiments, target amplicons using the methods (and associated compositions, systems, and kits) disclosed herein, are coupled to a downstream process, such as but not limited to, library preparation and nucleic acid sequencing. For example, target amplicons can be amplified using bridge amplification, emulsion PCR or isothermal amplification to generate a plurality of clonal templates suitable for nucleic acid sequencing. In some embodiments, the amplicon library is sequenced using any suitable DNA sequencing platform such as any next generation sequencing platform, including semi-conductor sequencing technology such as the Ion Torrent sequencing platform. In some embodiments, an amplicon library is sequenced using an Ion Torrent S5 520 ™< System or an Ion Torrent S5 530 ™< System or an Ion Torrent PGM 318 ™< System. In some embodiments, an amplicon library is sequenced using an Ion Torrent S5 540 ™< System or an Ion Torrent S5 550 ™< System.
[0100] In some embodiments, sequencing of immune receptor amplicons generated using the methods (and associated compositions and kits) disclosed herein, produces contiguous sequence reads from about 200 to about 600 nucleotides in length. In some embodiments, contiguous read lengths are from about 300 to about 400 nucleotides. In some embodiments, contiguous read lengths are from about 350 to about 450 nucleotides. In some embodiments, read lengths average about 300 nucleotides, about 350 nucleotides, or about 400 nucleotides. In some embodiments, contiguous read lengths are from about 250 to about 350 nucleotides, about 275 to about 340, or about 295 to about 325 nucleotides in length. In some embodiments, read lengths average about 270, about 280, about 290, about 300, or about 325 nucleotides in length. In other embodiments, contiguous read lengths are from about 180 to about 300 nucleotides, about 200 to about 290 nucleotides, about 225 to about 280 nucleotides, or about 230 to about 250 nucleotides in length. In some embodiments, read lengths average about 200, about 220, about 230, about 240, or about 250 nucleotides in length. In other embodiments, contiguous read lengths are from about 70 to about 200 nucleotides, about 80 to about 150 nucleotides, about 90 to about 140 nucleotides, or about 100 to about 120 nucleotides in length. In some embodiments, contiguous read lengths are from about 50 to about 170 nucleotides, about 60 to about 160 nucleotides, about 60 to about 120 nucleotides, about 70 to about 100 nucleotides, about 70 to about 90 nucleotides, or about 80 nucleotides in length. In some embodiments, read lengths average about 70, about 80, about 90, about 100, about 110, or about 120 nucleotides. In some embodiments, the sequence read length include the amplicon sequence and a barcode sequence. In some embodiments, the sequence read length does not include a barcode sequence.
[0101] In some embodiments, the amplification primers and primer pairs are target-specific sequences that can amplify specific regions of a nucleic acid molecule. In some embodiments, the target-specific primers can amplify expressed RNA or cDNA. In some embodiments, the target-specific primers can amplify mammalian RNA, such as human RNA or cDNA prepared therefrom, or murine RNA or cDNA prepared therefrom. In some embodiments, the target-specific primers can amplify DNA, such as gDNA. In some embodiments, the target-specific primers can amplify mammalian DNA, such as human DNA or murine DNA.
[0102] In methods and compositions provided herein, for example those for determining, characterizing, and / or tracking the immune repertoire in a biological sample, the amount of input RNA or gDNA required for amplification of target sequences will depend in part on the fraction of immune receptor bearing cells (e.g., T cells or B cells) in the sample. For example, a higher fraction of T cells in the sample, such as samples enriched for T cells, permits use of a lower amount of input RNA or gDNA for amplification. In some embodiments, the amount of input RNA for amplification of one or more target sequences can be about 0.05 ng to about 10 micrograms. In some embodiments, the amount of input RNA used for multiplex amplification of one or more target sequences can be from about 5 ng to about 2 micrograms. In some embodiments, the amount of RNA used for multiplex amplification of one or more target sequences can be from about 5 ng to about 1 microgram or about 10 ng to about 1 microgram. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 1.5 micrograms, about 2 micrograms, about 2.5 micrograms, about 3 micrograms, about 3.5 micrograms, about 4.0 micrograms, about 5 micrograms, about 6 micrograms, about 7 micrograms, or about 10 micrograms. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is from about 25 ng to about 500 ng RNA or from about 50 ng to about 200 ng RNA. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is from about 0.05 ng to about 10 ng RNA, from about 0.1 ng to about 5 ng RNA, from about 0.2 ng to about 2 ng RNA, or from about 0.5 ng to about 1 ng RNA. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0103] As described herein, RNA from a biological sample is converted to cDNA, typically using reverse transcriptase in a reverse transcription reaction, prior to the multiplex amplification. In some embodiments, a reverse transcription reaction is performed with the input RNA and a portion of the cDNA from the reverse transcription reaction is used in the multiplex amplification reaction. In some embodiments, substantially all of the cDNA prepared from the input RNA is added to the multiplex amplification reaction. In other embodiments, a portion, such as about 80%, about 75%, about 66%, about 50%, about 33%, or about 25% of the cDNA prepared from the input RNA is added to the multiplex amplification reaction. In other embodiments, about 15%, about 10%, about 8%, about 6%, or about 5% of the cDNA prepared from the input RNA is added to the multiplex amplification reaction.
[0104] In some embodiments, the amount of cDNA from a sample added to the multiplex amplification reaction can be about 0.001 ng to about 5 micrograms. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences can be from about 0.01 ng to about 2 micrograms. In some embodiments, the amount of cDNA used for multiplex amplification of one or more target sequences can be from about 0.1 ng to about 1 microgram or about 1 ng to about 0.5 microgram. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is from about 0.01 ng to about 10 ng cDNA, from about 0.05 ng to about 5 ng cDNA, from about 0.1 ng to about 2 ng cDNA, or from about 0.01 ng to about 1 ng cDNA. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.005 ng, about 0.01 ng, about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0105] In some embodiments, mRNA is obtained from a biological sample and converted to cDNA for amplification purposes using conventional methods. Methods and reagents for extracting or isolating nucleic acid from biological samples are well known and commercially available. In some embodiments, RNA extraction from biological samples is performed by any method described herein or otherwise known to those of skill in the art, e.g., methods involving proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with DNAse to digest contaminating DNA, and RNA purification using silica-gel-membrane technology, or any combination thereof. Exemplary methods for RNA extraction from biological samples using commercially available kits including RecoverAll ™< Multi-Sample RNA / DNA Workflow (Invitrogen), RecoverAll ™< Total Nucleic Acid Isolation Kit (Invitrogen), NucleoSpin ®< RNA blood (Macherey-Nagel), PAXgene ®< Blood RNA system, TRI Reagent ™< (Invitrogen), PureLink ™< RNA Micro Scale kit (Invitrogen), MagMAX ™< FFPE DNA / RNA Ultra Kit (Applied Biosystems) ZR RNA MicroPrep ™< kit (Zymo Research), RNeasy Micro kit (Qiagen), and ReliaPrep ™< RNA Tissue miniPrep system (Promega).
[0106] In some embodiments, the amount of input gDNA for amplification of one or more target sequences can be about 0.1 ng to about 10 micrograms. In some embodiments, the amount of gDNA required for amplification of one or more target sequences can be from about 0.5 ng to about 5 micrograms. In some embodiments, the amount of gDNA required for amplification of one or more target sequences can be from about 1 ng to about 1 microgram or about 10 ng to about 1 microgram. In some embodiments, the amount of gDNA required for amplification of one or more immune repertoire target sequences is from about 10 ng to about 500 ng, about 25 ng to about 400 ng, or from about 50 ng to about 200 ng. In some embodiments, the amount of gDNA required for amplification of one or more target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 20 ng, about 50 ng, about 100 ng, or about 200 ng. In some embodiments, the amount of gDNA required for amplification of one or more immune repertoire target sequences is about 1 microgram, about 2 micrograms, about 3 micrograms, about 4.0 micrograms, or about 5 micrograms.
[0107] In some embodiments, gDNA is obtained from a biological sample using conventional methods. Methods and reagents for extracting or isolating nucleic acid from biological samples are well known and commercially available. In some embodiments, DNA extraction from biological samples is performed by any method described herein or otherwise known to those of skill in the art, e.g., methods involving proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with RNAse to digest contaminating RNA, and DNA purification using silica-gel-membrane technology, or any combination thereof. Exemplary methods for DNA extraction from biological samples using commercially available kits including Ion AmpliSeq ™< Direct FFPE DNA Kit, MagMAX ™< FFPE DNA / RNA Ultra Kit, TRI Reagent ™< (Invitrogen), PureLink ™< Genomic DNA Mini kit (Invitrogen), RecoverAll ™< Total Nucleic Acid Isolation Kit (Invitrogen), MagMAX ™< DNA Multi-Sample Kit (Invitrogen) and DNA extraction kits from BioChain Institute Inc. (e.g., FFPE Tissue DNA Extraction Kit, Genomic DNA Extraction Kit, Blood and Serum DNA Isolation Kit).
[0108] A sample or biological sample, as used herein, refers to a composition from an individual that contains or may contain cells related to the immune system. Exemplary biological samples, include without limitation, tissue (for example, lymph node, organ tissue, bone marrow), whole blood, synovial fluid, cerebral spinal fluid, tumor biopsy, and other clinical specimens containing cells. The sample may include normal and / or diseased cells and be a fine needle aspirate, fine needle biopsy, core sample, or other sample. In some embodiments, the biological sample may comprise hematopoietic cells, peripheral blood mononuclear cells (PBMCs), T cells, B cells, tumor infiltrating lymphocytes ("TILs") or other lymphocytes. In some embodiments, the sample may be fresh (e.g., not preserved), frozen, or formalin-fixed paraffin-embedded tissue (FFPE). Some samples comprise cancer cells, such as carcinomas, melanomas, sarcomas, lymphomas, myelomas, leukemias, and the like, and the cancer cells may be circulating tumor cells.
[0109] The biological sample can be a mix of tissue or cell types, a preparation of cells enriched for at least one particular category or type of cell, or an isolated population of cells of a particular type or phenotype. Samples can be separated by centrifugation, elutriation, density gradient separation, apheresis, affinity selection, panning, FACS, centrifugation with Hypaque, etc. prior to analysis. Methods for sorting, enriching for, and isolating particular cell types are well-known and can be readily carried out by one of ordinary skill. In some embodiments, the sample may a preparation enriched for T cells, for example CD3+ T cells.
[0110] In some embodiments, the provided methods and systems include processes for analysis of immune repertoire receptor cDNA or gDNA sequence data and for identification and / or removing PCR or sequencing-derived error(s) from the determined immune receptor sequence.
[0111] In some embodiments, the error correction strategy includes the following steps: 1) Align the sequenced rearrangement to a reference database of variable, diversity and joining / constant genes to produce a query sequence / reference sequence pair. Many alignment procedures may be used for this purpose including, for example, IgBLAST, a freely-available tool from the NCBI, and custom computer scripts. 2) Realign the reference and query sequences to each other, taking into account the flow order used for sequencing. The flow order provides information that allows one to identify and correct some types of erroneous alignments. 3) Identify the borders of the CDR3 region by their characteristic sequence motifs. 4) Over the aligned portion of the rearrangement corresponding to the variable gene and joining / constant genes, excluding the CDR3 region, identify indels in the query with respect to the reference and alter the mismatching query base position so that it is consistent with the reference. 5) For the CDR3 region, if the CDR3 length is not a multiple of three (indicative of an indel error): (a) Search the CDR3 for the homopolymer stretch having the highest probability of containing a sequence error, based on PHRED score (denoted e). (b) Obtain the probability of error over the entire CDR3 region based on PHRED score (denoted t) (c) If e / t is greater than a defined threshold, edit the homopolymer by either increasing or decreasing the length of the homopolymer by one base such that the CDR3 nucleotide length is a multiple of three. (d)As an alternative to steps a-c, search the CDR3 for the longest homopolymer, and if the length of the homopolymer is above a defined threshold, edit the homopolymer by either increasing or decreasing the length of the homopolymer by one base such that the CDR3 nucleotide length is a multiple of three.
[0112] In some embodiments, methods are provided to identify T cell or B cell clones in repertoire data that are robust to PCR and sequencing error. Accordingly, the following describes steps that may be employed in such methods to identify T cell or B cell clones in a manner that is robust to PCR and sequencing error. Table 1 a diagram of an exemplary workflow for use in identifying and removing PCR or sequencing-derived errors from immune receptor sequencing data. Exemplary portions and embodiments of this workflow are also represented in FIG. 1.
[0113] For a set of TCR or BCR sequences derived from mRNA, where 1) each sequence has been annotated as a productive rearrangement, either natively or after error correction, such as previously described, and 2) each sequence has an identified V gene and CDR3 nucleotide region, in some embodiments, methods include the following: 1) Identify and exclude chimeric sequences. For each unique CDR3 nucleotide sequence present in the dataset, tally the number of reads having that CDR3 nucleotide sequence and any of the possible V genes. Any V gene-CDR3 combination making up less than 10% of total reads for that CDR3 nucleotide sequence is flagged as a chimera and eliminated from downstream analyses. As an example, for the sequences below having the same CDR3 nucleotide sequence, e.g., the sequences having TRBV3 and TRBV6 paired with CDR3nt sequence AATTGGT will be flagged as chimeric. V geneCDR3ntRead countsTRBV2AATTGGT1000TRBV3AATTGGT10TRBV6AATTGGT3 2) Identify and exclude sequences containing simple indel errors. For each read in the dataset, obtain the homopolymer-collapsed representation of the CDR3 sequence of that read. For each set of reads having the same V gene and collapsed-CDR3 combination, tally the number of occurrences of each non-collapsed CDR3 nucleotide sequence. Any non-collapsed CDR3 sequence making up < 1 0% of total reads for that read set is flagged as having a simple homopolymer error. As an example, three different V gene-CDR3 nucleotide sequences are presented that are identical after homopolymer collapsing of the CDR3 nucleotide sequence. The two less frequent V gene-CDR3 combinations make up < 10% of total reads for the read set and will be flagged as containing a simple indel error. For example: V geneCDR3ntHomopolymer collapsed CDR3ntRead countsTRBV2AATTGGTATGT1000TRBV2AAATGGTATGT10TRBV2AAAATTTGGT (SEQ ID NO: 521)ATGT3 3) Identify and exclude singleton reads. For each read in the dataset, tally the number of times that the exact read sequence is found in the dataset. Reads that appear only once in the dataset will be flagged as singleton reads. 4) Identify and exclude truncated reads. For each read in the dataset, determine whether the read possesses an annotated V gene FR1, CDR1, FR2, CDR2, and FR3 region, as indicated by the IgBLAST alignment of the read to the IgBLAST reference V gene set. Reads that do not possess the above regions are flagged as truncated if the region(s) is expected based on the particular V gene primer used for amplification. 5) Identify and exclude rearrangements lacking bidirectional support. For each read in the dataset, obtain the V gene and CDR3 sequence of the read as well as the strand orientation of the read (plus or minus strand). For each V gene-CDR3 combination in the dataset, tally the number of plus and minus strand reads having that V gene-CDR3nt combination. V gene-CDR3nt combinations that are only present in reads of one orientation will be deemed to be a spurious. All reads having a spurious V gene-CDR3nt combination will be flagged as lacking bidirectional support. 6) For genes that have not been flagged, perform stepwise clustering based on CDR3 nucleotide similarity. Separate the sequences into groups based on the V gene identity of the read, excluding allele information (v-gene groups). For each group: a. Arrange reads in each group into clusters using cd-hit-est and the following parameters: cd-hit-est -i vgene_groups.fa -o clustered _vgene_groups.cdhit -T 24 -d 0 -M 100000 -B 0 -r 0 -g 1 -S 0 -U 2 -uL .05 -n 10 -1 7 Where vgene_groups.fa is a fasta format file of the CDR3 nucleotide regions of sequences having the same V gene and clustered vgene_groups.cdhit is the output, containing the subdivided sequences. b. Assign each sequence in a cluster the same clone ID, used to denote that members of the subgroup are believed to represent the same T cell clone or B cell clone. c. Chose a representative sequence for each cluster, such that the representative sequence is the sequence that appears the greatest number of times, or, in cases of a tie, is randomly chosen. d. Merge all other reads in the cluster into the representative sequence such that the number of reads for the representative sequence is increased according to the number of reads for the merged sequences. e. Compare the representative sequences within a v-gene group to each other on the basis of hamming distance. If a representative sequence is within a hamming distance of 1 to a representative sequence that is >50 times more abundant, merge that sequence into the more common representative sequence. If a representative sequence is within a hamming distance of 2 to a representative sequence that is >10000 times more abundant, merge that sequence into the more common representative sequence. f. Identify complex sequence errors. Homopolymer-collapse the representative sequences within each V gene group, then compare to each other using Levenshtein distances. If a representative sequence is within a Levenshtein distance of 1 to a representative sequence that is >50 times more abundant, merge that sequence into the more common representative sequence. g. Identify CDR3 misannotation errors. Homopolymer-collapse the representative sequences within each V gene group, then perform a pairwise comparison of each homopolymer-collapsed sequence. For each pair of sequences, determine whether one sequence is a subset of the other sequence. If so, merge the less abundant sequence into the more abundant sequence if the more abundance sequence is >500 fold more abundant. 7) Report cluster representatives to user.
[0114] In some embodiments, the provided workflow is not limited to the frequency ratios listed in the various steps, and other frequency ratios may be substituted for the representative ratios included above. For example, in some embodiments, comparing the representative sequences within a v-gene group to each other on the basis of hamming distance may use a frequency ratio other than those listed in step (e) above. For example and without limitation, frequency ratios of 1000, 5000, 20,000, etc may be used if a representative sequence is within a hamming distance of 2 to a representative sequence. For example and without limitation, frequency ratios of 20, 100, 200, etc may be used if a representative sequence is within a hamming distance of 1 to a representative sequence. The frequency ratios provided are representative of the general process of labeling the more abundant sequence of a similar pair as a correct sequence.
[0115] Similarly, when comparing the frequencies of two sequences at other steps in the workflow, eg, step (1), step (2), step (6f) and step (6g), frequency ratios other than those listed in the step above may be used.
[0116] As used herein, the term "homopolymer-collapsed sequence" is intended to represent a sequence where repeated bases are collapsed to a single base representative. As an example, for the non-collapsed sequence AAAATTTTTATCCCCCCCCGGG (SEQ ID NO: 522), the homopolymer-collapsed sequence is ATATCG.
[0117] As used herein, the terms "clone," "clonotype," "lineage," or "rearrangement" are intended to describe a unique V gene nucleotide combination for an immune receptor, such as a TCR or BCR. For example, a unique V gene-CDR3 nucleotide combination.
[0118] As used herein, the term "productive reads" refers to a TCR or BCR sequence reads that have no stop codon and have in-frame variable gene and joining gene segments. Productive reads are biologically plausible in coding for a polypeptide.
[0119] As used herein, "chimeras" or chimeric sequences" refer to artefactual sequences that arise from template switching during target amplification, such as PCR. Chimeras typically present as a CDR3 sequence grafted onto an unrelated V gene, resulting in a CDR3 sequence that is associated with multiple V genes within a dataset. The chimeric sequence is usually far less abundant than the true sequence in the dataset.
[0120] As used herein, the term "indel" refers to an insertion and / or deletion of one or more nucleotide bases in a nucleic acid sequence. In coding regions of a nucleic acid sequence, unless the length of an indel is a multiple of 3, it will produce a frameshift when the sequence is translated. As used herein, "simple indel errors" are errors that do not alter the homopolymer-collapsed representation of the sequence. As used herein, "complex indel errors" are indel sequencing errors that alter the homopolymer-collapsed representation of the sequence and include, without limitation, errors that eliminate a homopolymer, insert a homopolymer into the sequence, or create a dyslexic-type error.
[0121] As used herein, "singleton reads" refer to sequence reads whose indel-corrected sequence appears only once in a dataset. Typically, singleton reads are enriched for reads containing a PCR or sequencing error.
[0122] As used herein, "truncated reads" refer to immune receptor sequence reads that are missing annotated V gene regions. For example, truncated reads include, without limitation, sequence reads that are missing annotated TCR or BCR V gene FR1, CDR1, FR2, CDR2, or FR3 regions. Such reads typically are missing a portion of the V gene sequence due to quality trimming. Truncated reads can give rise to artifacts if the truncation leads one to misidentify the V gene.
[0123] In the context of identified V gene-CDR3 sequences (clonotypes), "bidirectional support" indicates that a particular V gene-CDR3 sequence is found in at least one read that maps to the plus strand (proceeding from the V gene to constant gene) and at least one reads that maps to the minus strand (proceeding form the constant gene to the V gene). Systematic sequencing errors often lead to identification of V gene-CDR3 sequences having unidirectional support.
[0124] For a set of sequences that have been grouped according to a predetermined sequence similarity threshold to account for variation due to PCR or sequencing error, the "cluster representative" is the sequence that is chosen as most likely to be error free. This is typically the most abundant sequence.
[0125] As used herein, "IgBLAST annotation error" refers to rare events where the border of the CDR3 is identified to be in an incorrect adjacent position. These events typically add three bases to the 5' or 3' end of a CDR3 nucleotide sequence.
[0126] For two sequences of equal length, the "Hamming distance" is the number of positions at which the corresponding bases are different. For any two sequences, the "Levenshtein distance" or the "edit distance" is the number of single base edits required to make one sequence into another sequence.
[0127] In some embodiments in which J gene-directed primers are used in amplification of the immune receptor sequences, for example multiplex amplification with primers directed to V gene FR3 regions and primers directed to J genes, raw sequence reads derived from the assay undergo a J gene sequence inference process before any downstream analysis. In this process, the beginning and end of raw read sequences are interrogated for the presence of characteristic sequences of 10-30 nucleotides corresponding to the portion of the J gene sequences expected to exist after amplification with the J primer and any subsequent manipulation or processing (for example, digestion) of the amplicon termini prior to sequencing. The characteristic nucleotide sequences permit one to infer the sequence of the J primer, as well as the remaining portion of the J gene that was targeted since the sequence of each J gene is known. To complete the J gene sequence inference process, the inferred J gene sequence is added to the raw read to create an extended read that then spans the entire J gene. The extended read then contains the entire J gene sequence, the entire sequence of the CDR3 region, and at least a portion of the V gene sequence, which will be reported after downstream analysis. The portion of V gene sequence in the extended read will depend on the V gene-directed primers used for the multiplex amplification, for example, FR3-, FR2-, or FR1-directed primers.
[0128] Use of V gene FR3 and J gene primers to amplify expressed immune receptor sequences or rearranged immune receptor gDNA sequences yields a minimum length amplicon (for example, about 60-100 or about 80 nucleotides in length) while still producing data that allows for reporting of the entire CDR3 region. With the expectation of short amplicon length, reads of amplicons <100 nucleotides in length are not eliminated as low-quality and / or off target products during the sequence analysis workflow. However, the explicit search for the expected J gene sequences in the raw reads allows one to eliminate amplicons deriving from off-target amplifications by the J gene primers. In addition, this short amplicon length improves the performance of the assay on highly degraded template material, such as that derived from an FFPE sample.
[0129] In some embodiments, provided methods comprise sequencing an immune receptor library and subjecting the obtained sequence data to error identification and correction processes to generate rescued productive reads, and identifying productive and rescued productive sequence reads. In some embodiments, provided methods comprise sequencing an immune receptor library and subjecting the obtained sequence dataset to error identification and correction processes, identifying productive and rescued productive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes in the library.
[0130] In some embodiments, provided methods comprise sequencing a rearranged immune receptor DNA library and subjecting the obtained sequence data to error identification and correction processes for the V gene portions to generate rescued productive reads, and identifying productive, rescued productive, and unproductive sequence reads. In some embodiments, provided methods comprise sequencing a rearranged immune receptor DNA library and subjecting the obtained sequence dataset to error identification and correction processes for the V gene portions, identifying productive, rescued productive, and unproductive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes in the library. In some embodiments, both productive and unproductive sequence reads of rearranged immune receptor DNA are separately reported.
[0131] In some embodiments, the provided error identification and correction workflow is used for identifying and resolving PCR or sequencing-derived errors that lead to a sequence read being identified as from an unproductive rearrangement. In some embodiments, the provided error identification and correction workflow is applied to immune receptor sequence data generated from a sequencing platform in which indel or other frameshift-causing errors occur while generating the sequence data.
[0132] In some embodiments, the provided error identification and correction workflow is applied to sequence data generated by an Ion Torrent sequencing platform. In some embodiments, the provided error identification and correction workflow is applied to sequence data generated by Roche 454 Life Sciences sequencing platforms, PacBio sequencing platforms, and Oxford Nanopore sequencing platforms.
[0133] In some embodiments, provided methods comprise preparation and formation of a plurality of immune receptor-specific amplicons. In some embodiments, the method comprises hybridizing a plurality of V gene-specific primers and at least one C gene-specific primer to a cDNA molecule, extending a first primer (e.g., a V gene-specific primer) of the primer pair, denaturing the extended first primer from the cDNA molecule, hybridizing to the extended first primer product, a second primer (e.g., a C gene-specific primer) of the primer pair and extending the second primer, digesting the target-specific primer pairs to generate a plurality of target amplicons. In other embodiments, the method comprises hybridizing a plurality of V gene gene-specific primers and a plurality of J gene-specific primers to a cDNA molecule, extending a first primer (e.g., a V gene-specific primer) of the primer pair, denaturing the extended first primer from the cDNA molecule, hybridizing to the extended first primer product, a second primer (e.g., a J gene-specific primer) of the primer pair and extending the second primer, digesting the target-specific primer pairs to generate a plurality of target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons prior to performing a nick translation reaction to generate a plurality of target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters includes at least one barcode sequence. In some embodiments, each adapter ligated to the ends of the target amplicons includes a barcode sequence. In some embodiments, the one or more target amplicons can be amplified using bridge amplification, emulsion PCR or isothermal amplification to generate a plurality of clonal templates suitable for nucleic acid sequencing.
[0134] In some embodiments, provided methods comprise preparation and formation of a plurality of immune receptor-specific amplicons. In some embodiments, the method comprises hybridizing a plurality of V gene gene-specific primers and a plurality of J gene-specific primers to a gDNA molecule, extending a first primer (eg, a V gene-specific primer) of the primer pair, denaturing the extended first primer from the gDNA molecule, hybridizing to the extended first primer product, a second primer (e.g., a J gene-specific primer) of the primer pair and extending the second primer, digesting the target-specific primer pairs to generate a plurality of target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons prior to performing a nick translation reaction to generate a plurality of target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters includes at least one barcode sequence. In some embodiments, each adapter ligated to the ends of the target amplicons includes a barcode sequence. In some embodiments, the one or more target amplicons can be amplified using bridge amplification or emulsion PCR to generate a plurality of clonal templates suitable for nucleic acid sequencing.
[0135] In some embodiments, the disclosure provides methods for sequencing target amplicons and processing the sequence data to identify productive immune receptor rearrangements expressed in the biological sample from which the cDNA was derived. In other embodiments, the disclosure provides methods for sequencing target amplicons and processing the sequence data to identify productive immune receptor gene rearrangements gDNA from a biological sample. In embodiments in which J gene-directed primers are used to amplify the expressed immune receptor sequences or rearranged immune receptor gDNA sequences, processing the sequence data includes inferring the nucleotide sequence of the J gene primer used for amplification as well as the remaining portion of the J gene that was targeted, as described herein. In some embodiments, processing the sequence data includes performing provided error identification and correction steps to generate rescued productive sequences. In some embodiments, use of the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads being at least 50% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, use of the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads being at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, use of the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads being about 50-60%, about 60-70%, about 70-80%, about 80-90%, about 50-80%, or about 60-90% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, use of the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads averaging about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% of the sequencing reads for an immune receptor cDNA or gDNA sample.
[0136] With particular samples, the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads being less than 50% of the sequencing reads for an immune receptor cDNA or gDNA sample when particular samples are used. Such samples include, for example, those in which the RNA or gDNA is highly degraded such as FFPE samples, and those in which the number of target immune cells is very low such as, for example, samples with very low T cell count or samples from subjects experiencing severe leukopenia. Accordingly, in some embodiments, use of the provided error identification and correction workflow can result in a combination of productive reads and rescued productive reads being about 30-50%, about 40-50%, about 30-40%, about 40-60%, at least 30%, or at least 40% of the sequencing reads for an immune receptor cDNA or gDNA sample.
[0137] In certain embodiments, methods of the invention comprise the use of target immune receptor primer sets wherein the primers are directed to sequences of the same target immune receptor gene. Immune receptors are selected from T cell receptors and antibody receptors. In some embodiments a T cell receptor is a T cell receptor selected from the group consisting of TCR alpha, TCR beta, TCR gamma, and TCR delta. In some embodiments the immune receptor is an antibody receptor selected from the group consisting of heavy chain alpha, heavy chain delta, heavy chain epsilon, heavy chain gamma, heavy chain mu, light chain kappa, and light chain lambda.
[0138] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of an immune receptor coding sequence comprising at least a portion of a framework region within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target constant gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region.
[0139] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the framework region 1 of the template molecules. In certain embodiments the one or more C gene primers of ii) comprises at least two primers that anneal to at least a portion of the C gene portion of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining regions CDR1, CDR2, and CDR3 of an immune receptor expression sequence. In some embodiments the amplicons are about 300 to about 600 nucleotides in length or at least about 350 to about 500 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0140] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying a productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the average sequence read length is between 300 and 600 nucleotides, or is between 350 and 550 nucleotides, or is between 330 and 425 nucleotides, or is about 350 to about 425 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining regions CDR1, CDR2, and CDR3 of an immune receptor expression sequence.
[0141] In some embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR1 region about 70 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR1 region about 50 nucleotides in length. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 90 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 50 to about 80 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 55 to about 75 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 60 to about 70 different FR1-directed primers. In some embodiments the target immune receptor primer set comprises one or more C gene primers. In particular embodiments a target immune receptor primer set comprises at least two C gene primers wherein each is directed to at least a portion of the same 50 nucleotide region within the target C gene.
[0142] In particular embodiments, methods of the invention comprise use of at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Tables 2 and 4, respectively. In other certain embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-89 and 181-184 or selected from SEQ ID NOs: 90-180 and 181-184. In some embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-155 and 181-182 or selected from SEQ ID NOs: 90-155 and 183-184. In some embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-89 and 181-182 or selected from SEQ ID NOs: 1-89 and 183-184. In some embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-180 and 181-182 or selected from SEQ ID NOs: 90-180 and 183-184. In other certain embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 183-184. In other certain embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 181-182. In still other certain embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-153 and 181-182. In certain embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-155, and 181-182 or at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-155, and 183-184. In still other certain embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-153 and 183-184. In still other certain embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180, and 181-182. In still other certain embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180, and 183-184.
[0143] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least one primer selected from SEQ ID NOs: 183-184. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least one primer selected from SEQ ID NOs: 183-184.
[0144] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least one primer selected from SEQ ID NOs: 183-184. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least one primer selected from SEQ ID NOs: 183-184.
[0145] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 40 to about a 60 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 3 region of the template molecules. In certain embodiments the one or more C gene primers of ii) comprises at least two primers that anneal to at least a portion of the C gene of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining region CDR3 of an immune receptor expression sequence. In some embodiments the amplicons are about 80 to about 200 nucleotides in length, about 80 to about 140 nucleotides in length, about 90 to about 130 nucleotides in length or at least about 100 to about 120 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0146] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying a productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the average sequence read length is between 80 and 185 nucleotides, is between 115 and 200 nucleotides, is between 90 and 130 nucleotides, or is between about 100 and about 120 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining region CDR3 of an immune receptor expression sequence.
[0147] In certain embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 70 nucleotides in length. In particular embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 50 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 40 to about 60 nucleotides in length. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 80 different FR3-directed primers. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 50 to about 70 different FR3-directed primers. In some embodiments, a target immune receptor primer set comprises V gene primers comprising about 55 to about 65 different FR3-directed primers. In some embodiments, a target immune receptor primer set comprises V gene primers comprising about 58, 59, 60, 61, or 62 different FR3-directed primers. In some embodiments the target immune receptor primer set comprises one or more C gene primers. In particular embodiments a target immune receptor primer set comprises at least two C gene primers wherein each is directed to at least a portion of the same 50 nucleotide region within the target C gene.
[0148] In particular embodiments, methods of the invention comprise the use of at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Tables 3 and 4, respectively. In other certain embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-248 and 181-184 or selected from SEQ ID NOs: 249-312 and 181-184. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-248 and 183-184 or selected from SEQ ID NOs: 185-248 and 181-182. In other certain embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-243 and 183-184. In other certain embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-243 and 181-182. In other certain embodiments methods of the invention comprise the use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 181-182 or selected from SEQ ID NOs: 249-312 and 183-184. In still other certain embodiments methods of the invention comprise the use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 249-307 and 181-182. In still other certain embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 249-307 and 183-184.
[0149] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least one primer selected from SEQ ID NOs: 183-184. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least one primer selected from SEQ ID NOs: 183-184.
[0150] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene, and ii) one or more C gene primers directed to at least a portion of the C gene of the respective immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the FR2 region of the template molecules. In certain embodiments the one or more C gene primers of ii) comprises at least two primers that anneal to at least a portion of the constant portion C gene of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining regions CDR2 and CDR3 of an immune receptor expression sequence. In some embodiments the amplicons are about 180 to about 375 nucleotides in length, about 200 to about 350 nucleotides, about 225 to about 325 nucleotides, or about 250 to about 300 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0151] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 40%, at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the average sequence read length is between about 200 and about 375 nucleotides, between about 250 and about 350 nucleotides, or between about 275 and about 350 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining regions CDR2 and CDR3 of an immune receptor expression sequence.
[0152] In particular embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR2 region about 70 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR2 region about 50 nucleotides in length. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 90 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 30 to about 60 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 20 to about 50 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 60 to about 70 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 20 to about 30 different FR2-directed primers. In some embodiments the target immune receptor primer set comprises one or more C gene primers. In particular embodiments a target immune receptor primer set comprises at least two C gene primers wherein each is directed to at least a portion of the same 50 nucleotide region within the target C gene.
[0153] In particular embodiments, methods of the invention comprise use of at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Tables 6 and 4, respectively. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 483-505 and 181-182. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 483-505 and 183-184.
[0154] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least one primer selected from SEQ ID NOs: 181-182. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least one primer selected from SEQ ID NOs: 183-184.
[0155] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of an immune receptor coding sequence comprising at least a portion of a framework region within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In particular embodiments the one or more plurality of J gene primers of ii) are directed to sequences over about a 50 nucleotide portion of the J gene. In more particular embodiments the one or more plurality of J gene primers of ii) are directed to sequences over about a 30 nucleotide portion of the J gene. In certain embodiments, the one or more plurality of J gene primers of ii) are directed to sequences completely within the J gene.
[0156] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 40 to about a 60 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 3 region of the template molecules. In certain embodiments the plurality of J gene primers of ii) comprises at least ten primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 14 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 16 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 10 to about 20 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 12 to about 18 primers that anneal to at least a portion of the J gene portion of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining region CDR3 of an immune receptor expression sequence. In some embodiments the amplicons are about 60 to about 160 nucleotides in length, about 70 to about 100 nucleotides in length, at least about 70 to about 90 nucleotides in length, about 80 to about 90 nucleotides in length, or about 80 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0157] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In some embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence, identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the sequence read lengths are about 60 to about 185 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In some embodiments the average sequence read length is between 70 and 90 nucleotides, or is between about 75 and about 85 nucleotides, or is about 80 nucleotides. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining region CDR3 of an immune receptor expression sequence.
[0158] In particular embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 50 nucleotides in length. In other embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 70 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR3 region about 40 to about 60 nucleotides in length. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 80 different FR3-directed primers. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 50 to about 70 different FR3-directed primers. In some embodiments, a target immune receptor primer set comprises V gene primers comprising about 55 to about 65 different FR3-directed primers. In some embodiments, a target immune receptor primer set comprises V gene primers comprising about 58, 59, 60, 61, or 62 different FR3-directed primers. In some embodiments the target immune receptor primer set comprises a plurality of J gene primers. In some embodiments a target immune receptor primer set comprises at least ten J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises at least 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 10 to about 20 different J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 12, 13, 14, 15, 16, 17 or 18 different J gene primers. In particular embodiments a target immune receptor primer set comprises about 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In particular embodiments a target immune receptor primer set comprises about 14 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides.
[0159] In particular embodiments, methods of the invention comprise the use of at least one set of primers comprising V gene primers i) and J gene primers ii) selected from Tables 3 and 5, respectively. In certain other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-248 and 313-397 or selected from SEQ ID NOs: 185-248 and 398-482. In certain other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-248 and 313-329 or selected from SEQ ID NOs: 185-248 and 329-342. In still other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-248 and 398-414 or selected from SEQ ID NOs: 185-248 and 414-427. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-243 and 313-328. In still other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 185-243 and 398-413. In certain other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 313-397 or selected from SEQ ID NOs: 249-312 and 398-482. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 313-329 or selected from SEQ ID NOs: 249-312 and 329-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 398-414 or selected from SEQ ID NOs: 249-312 and 414-427. In certain other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 398-413. In still other embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 249-312 and 313-328.
[0160] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482.
[0161] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 249-312 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 185-248 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427.
[0162] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 1 region of the template molecules. In certain embodiments the plurality of J gene primers of ii) comprise at least ten primers that anneal to at least a portion of the J gene of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 14 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 16 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 10 to about 20 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 12 to about 18 primers that anneal to at least a portion of the J gene portion of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining regions CDR1, CDR2, and CDR3 of an immune receptor expression sequence. In some embodiments the amplicons are about 220 to about 350 nucleotides in length, about 225 to about 300 nucleotides, about 250 to about 325 nucleotides, about 250 to about 275 nucleotides, or about 270 to about 300 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0163] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In some embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence, identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the average sequence read length is between 200 and 350 nucleotides, between 225 and 325 nucleotides, between 250 and 300 nucleotides, between 270 and 300 nucleotides, or is between 295 and 325 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining regions CDR1, CDR2, and CDR3 of an immune receptor expression sequence.
[0164] In particular embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR1 region about 70 nucleotides in length. In other certain embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR1 region about 80 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR1 region about 50 nucleotides in length. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 90 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 50 to about 80 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 55 to about 75 different FR1-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 60 to about 70 different FR1-directed primers. In some embodiments the target immune receptor primer set comprises a plurality of J gene primers. In some embodiments a target immune receptor primer set comprises at least ten J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In particular embodiments a target immune receptor primer set comprises at least 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 10 to about 20 different J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 12, 13, 14, 15, 16, 17 or 18 different J gene primers. In particular embodiments a target immune receptor primer set comprises about 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In particular embodiments a target immune receptor primer set comprises about 14 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides.
[0165] In particular embodiments, methods of the invention comprise use of at least one set of primers comprising V gene primers i) and J gene primers ii) selected from Tables 2 and 5, respectively. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-89 and 313-397 or selected from SEQ ID NOs: 90-180 and 398-482. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-89 and 398-482 or selected from SEQ ID NOs: 90-180 and 313-397. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 313-397 or selected from SEQ ID NOs: 1-64 and 398-482. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 313-329 or selected from SEQ ID NOs: 1-64 and 329-342. In still other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 398-414 or selected from SEQ ID NOs: 1-64 and 414-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 313-328. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 1-64 and 398-413. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-180 and 313-342 or selected from SEQ ID NOs: 90-180 and 398-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-155 and 313-342 or selected from SEQ ID NOs: 90-155 and 398-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-155 and 398-414 or selected from SEQ ID NOs: 90-155 and 414-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-155 and 313-329 or selected from SEQ ID NOs: 90-155 and 329-342. In still other embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-153 and 398-414. In still other embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-153 and 313-328. In still other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180 and 329-342 or selected from SEQ ID NOs: 90-92, 95-180 and 313-329. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180 and 398-414 or selected from SEQ ID NOs: 90-92, 95-180 and 414-427. In certain other embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180 and 398-413. In still other embodiments methods of the invention comprise use of at least one set of primers of i) and ii) comprising primers selected from SEQ ID NOs: 90-92, 95-180, and 313-328.
[0166] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482.
[0167] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 1-89 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427.
[0168] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482.
[0169] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 50 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 60 primers selected from SEQ ID NOs: 90-180 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427.
[0170] In certain embodiments, provided is a method for amplification of expression nucleic acid sequences of an immune receptor repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor and wherein performing amplification using each set results in amplicons representing the entire repertoire of the respective immune receptor in the sample; thereby generating immune receptor amplicons comprising the repertoire of the immune receptor. In particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about an 80 nucleotide portion of the framework region. In more particular embodiments the one or more plurality of V gene primers of i) are directed to sequences over about a 50 nucleotide portion of the framework region. In some embodiments the one or more plurality of V gene primers of i) anneal to at least a portion of the FR2 region of the template molecules. In certain embodiments the plurality of J gene primers of ii) comprise at least ten primers that anneal to at least a portion of the J gene of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 14 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 16 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 10 to about 20 primers that anneal to at least a portion of the J gene portion of the template molecules. In some embodiments the plurality of J gene primers of ii) comprises about 12 to about 18 primers that anneal to at least a portion of the J gene portion of the template molecules. In particular embodiments at least one set of the generated amplicons includes complementarity determining regions CDR2 and CDR3 of an immune receptor gene sequence. In some embodiments the amplicons are about 160 to about 270 nucleotides in length, about 180 to about 250 nucleotides, or about 195 to about 225 nucleotides in length. In some embodiments the nucleic acid template used in methods is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0171] In certain embodiments, methods are provided for providing sequence of the immune repertoire in a sample, comprising performing a multiplex amplification reaction to amplify immune receptor nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. Sequencing of resulting immune receptor amplicon molecules is then performed and the sequences of the immune receptor amplicon molecules determined thereby provides sequence of the immune repertoire in the sample. In some embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence, identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments, determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In particular embodiments the combination of productive reads and rescued productive reads is at least 40%, at least 50%, at least 60% at least 70% or at least 75% of the sequencing reads for the immune receptors. In additional embodiments the method further comprises sequence read clustering and immune receptor clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments the average sequence read length is between 160 and 300 nucleotides, between 180 and 280 nucleotides, between 200 and 260 nucleotides, or between 225 and 270 nucleotides, depending in part on inclusion of any barcode sequence in the read length. In certain embodiments at least one set of the sequenced amplicons includes complementarity determining regions CDR2 and CDR3 of an immune receptor expression sequence.
[0172] In particular embodiments, methods provided utilize target immune receptor primer sets comprising V gene primers wherein the one or more of a plurality of V gene primers are directed to sequences over an FR2 region about 70 nucleotides in length. In other particular embodiments the one or more of a plurality of V gene primers are directed to sequences over an FR2 region about 50 nucleotides in length. In certain embodiments a target immune receptor primer set comprises V gene primers comprising about 45 to about 90 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 30 to about 60 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 20 to about 50 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 60 to about 70 different FR2-directed primers. In some embodiments a target immune receptor primer set comprises V gene primers comprising about 20 to about 30 different FR2-directed primers. In some embodiments the target immune receptor primer set comprises a plurality of J gene primers. In some embodiments a target immune receptor primer set comprises at least ten J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In particular embodiments a target immune receptor primer set comprises at least 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 10 to about 20 different J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In some embodiments a target immune receptor primer set comprises about 12, 13, 14, 15, 16, 17 or 18 different J gene primers. In particular embodiments a target immune receptor primer set comprises about 16 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides. In particular embodiments a target immune receptor primer set comprises about 14 J gene primers wherein each is directed to at least a portion of a J gene within target polynucleotides.
[0173] In particular embodiments, methods of the invention comprise use of at least one set of primers comprising V gene primers i) and J gene primers ii) selected from Tables 6 and 5, respectively. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primer selected from SEQ ID NOs: 483-505 and 313-397 or selected from SEQ ID NOs: 483-505 and 398-482. In some embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primer selected from SEQ ID NOs: 483-505 and 313-342 or selected from SEQ ID NOs: 483-505 and 398-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primer selected from SEQ ID NOs: 483-505 and 313-329 or selected from SEQ ID NOs: 483-505 and 329-342. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primer selected from SEQ ID NOs: 483-505 and 398-414 or selected from SEQ ID NOs: 483-505 and 414-427. In other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 483-505 and 313-328. In certain other embodiments methods of the invention comprise use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 483-505 and 398-413.
[0174] In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-397. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-482. In some embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 313-342. In other embodiments methods of the invention comprise the use of at least one set of primers i) and ii) comprising at least 20 primers selected from SEQ ID NOs: 483-505 and at least 10 primers, at least 12 primers, at least 14 primers, at least 16 primers, at least 18 primers, or at least 20 primers selected from SEQ ID NOs: 398-427.
[0175] In certain embodiments, methods of the invention comprise use of a biological sample selected from the group consisting of hematopoietic cells, lymphocytes, and tumor cells. In some embodiments the biological sample is selected from the group consisting of peripheral blood mononuclear cells (PBMCs), T cells, B cells, circulating tumor cells, and tumor infiltrating lymphocytes (herein "TILs" or "TIL"). In some embodiments, the biological sample comprises T cells undergoing ex vivo activation and / or expansion.
[0176] In some embodiments, methods, compositions, and systems are provided for determining the immune repertoire of a biological sample by assessing both expressed immune receptor RNA and rearranged immune receptor genomic DNA (gDNA) from a biological sample. Expression nucleic acid sequences of a sample may be assessed using the methods, compositions, and systems provided herein. The sample gDNA may be assessed for rearranged immune receptor gene sequences using the methods, composition, and systems described in the co-owned U.S. Provisional Application No. 62 / 553,736, filed September 1, 2017, entitled "Compositions and Methods for Immune Repertoire Sequencing". In some embodiments, the sample RNA and gDNA may be assessed concurrently and following reverse transcription of the RNA to form cDNA, the cDNA and gDNA may be amplified in the same multiplex amplification reaction. In some embodiments, cDNA from the sample RNA and the sample gDNA may undergo multiplex amplification in separate reactions. In some embodiments, cDNA from the sample RNA and sample gDNA may under multiplex amplification with parallel primer pools. In some embodiments, the same immune receptor-directed primer pools are used to assess the immune repertoire of gDNA and RNA from the sample. In some embodiments, the different immune receptor-directed primer pools are used to assess the immune repertoire of gDNA and RNA from the sample. In some embodiments, multiplex amplification reactions are performed separately with cDNA from the sample RNA and with sample gDNA to amplify target immune receptor molecules from the sample and the resulting immune receptor ampicons are sequenced, thereby providing sequence of the expressed immune receptor RNA and rearranged immune receptor gDNA of a biological sample.
[0177] In some embodiments, the methods and compositions provided are used to identify and / or characterize an immune repertoire of a subject. In some embodiments, methods and compositions provided are used to identify and characterize novel or non-canonical TCR or BCR alleles of a subject's immune repertoire. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database and the sequence of at least one allelic variant absent from that IMGT database is identified. In some embodiments, identified allelic variants absent from the IMGT database are subjected to evidence-based filtering using, for example, criteria such as clone number support, sequence read support and / or number of individuals having the allelic variant. Allelic variants identified and reported as absent from IMGT may be compared to other databases containing immune repertoire sequence information, such as NCBI NR database and Lym1K database, to cross-validate the reported novel or non-canonical TCR or BCR alleles. Characterizing the existence of undocumented or non-canonical TRB polymorphism, for example, may help with understanding factors that influence autoimmune disease and response to immunotherapy. Thus, in some embodiments, methods and compositions are provided to identify novel or non-canonical TRBV gene allele polymorphisms and allelic variants that may predict or detect autoimmune disease or immune-mediated adverse events. In other embodiments, provided are methods for making recombinant nucleic acids encoding identified novel TRBV allelic variants. In some embodiments, provided are methods for making recombinant TRBV allelic variant molecules and for making recombinant cells which express the same.
[0178] In some embodiments, methods and compositions provided are used to identify and characterize novel or non-canonical TCR or BCR alleles of a subject's immune repertoire. In some embodiments, a patient's immune repertoire may be identified or characterized before and / or after a therapeutic treatment, for example treatment for a cancer or immune disorder. In some embodiments, identification or characterization of an immune repertoire may be used to assess the effect or efficacy of a treatment, to modify therapeutic regimens, and to optimize the selection of therapeutic agents. In some embodiments, identification or characterization of the immune repertoire may be used to assess a patient's response to an immunotherapy, e.g., CAR (chimeric antigen receptor)-T cell therapy, a cancer vaccine and / or other immune-based treatment or combination(s) thereof. In some embodiments, identification or characterization of the immune repertoire may indicate a patient's likelihood to respond to a therapeutic agent or may indicate a patient's likelihood to not be responsive to a therapeutic agent.
[0179] In some embodiments, a patient's immune repertoire may be identified or characterized to monitor progression and / or treatment of hyperproliferative diseases, including detection of residual disease following patient treatment, monitor progression and / or treatment of autoimmune disease, transplantation monitoring, and to monitor conditions of antigenic stimulation, including following vaccination, exposure to bacterial, fungal, parasitic, or viral antigens, or infection by bacteria, fungi, parasites or virus. In some embodiments, identification or characterization of the immune repertoire may be used to assess a patient's response to an anti-infective or anti-inflammatory therapy.
[0180] In certain embodiments, the methods and compositions provided are used to monitor changes in immune repertoire clonal populations, for example changes in clonal expansion, changes in clonal contraction, and changes in relative ratios of clones or clonal populations. In some embodiments, the provided methods and compositions are used to monitor changes in immune repertoire clonal populations (e.g., clonal expansion, clonal contraction, changes in relative ratios) in response to tumor growth. In some embodiments, the provided methods and compositions are used to monitor changes in immune repertoire clonal populations (e.g., clonal expansion, clonal contraction, changes in relative ratios) in response to tumor treatment. In some embodiments, the provided methods and compositions provided are used to monitor changes in immune repertoire clonal populations (e.g., clonal expansion, clonal contraction, changes in relative ratios) during a remission period. For many lymphoid malignancies, a clonal B cell receptor or T cell receptor sequence can be used a biomarker for the malignant cells of the particular cancer (e.g., leukemia) and to monitor residual disease, tumor expansion, contraction, and / or treatment response. In certain embodiments a clonal B cell receptor or T cell receptor may be identified and further characterized to confirm a new utility in therapeutic, biomarker and / or diagnostic use.
[0181] In some embodiments, methods and compositions are provided for identifying and / or characterizing immune repertoire clonal populations in a sample from a subject, comprising performing one or more multiplex amplification reactions with the sample or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying one or more immune repertoire clonal populations for the target immune receptor from the sample. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the one or more multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the one or more multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0182] In some embodiments, methods and compositions are provided for identifying and / or characterizing immune repertoire clonal populations in a sample from a subject, comprising performing one or more multiplex amplification reactions with the sample or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying one or more immune repertoire clonal populations for the target immune receptor from the sample. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0183] In some embodiments, methods and compositions are provided for monitoring changes in immune repertoire clonal populations in a subject, comprising performing one or more multiplex amplification reaction with a subject's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire clonal populations for the target immune receptor from the sample, and comparing the identified immune repertoire clonal populations to those identified in samples obtained from the subject at a different time. In some embodiments, methods and compositions are provided for monitoring changes in immune repertoire clonal populations in a subject, comprising performing one or more multiplex amplification reaction with a subject's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire clonal populations for the target immune receptor from the sample, and comparing the identified immune repertoire clonal populations to those identified in samples obtained from the subject at a different time. In various embodiments, the one or more multiplex amplification reactions performed in such methods may be a single multiplex amplification reaction or may be two or more multiplex amplification reactions performed in parallel, for example parallel, highly multiplexed amplification reactions performed with different primer pools. Samples for use in monitoring changes in immune repertoire clonal populations include, without limitation, samples obtained prior to a diagnosis, samples obtained at any stage of diagnosis, samples obtained during a remission, samples obtained at any time prior to a treatment (pre-treatment sample), samples obtained at any time following completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0184] In certain embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a patient to monitor progression and / or treatment of the patient's hyperproliferative disease. In some embodiments, the methods and compositions provided are used for minimal residual disease (MRD) monitoring for a patient following treatment. In some embodiments, the methods and compositions are used to identify and / or track B cell lineage malignancies or T cell lineage malignancies. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients diagnosed with leukemia or lymphoma, including without limitation, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, cutaneous T cell lymphoma, B cell lymphoma, mantle cell lymphoma, and multiple myeloma. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients diagnosed with solid tumors, including without limitation, breast cancer, lung cancer, colorectal, and neuroblastoma. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients following cancer treatment including without limitation bone marrow transplant, lymphocyte infusion, adoptive T-cell therapy, other cell-based immunotherapy, and antibody-based immunotherapy.
[0185] In some embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a patient to monitor progression and / or treatment of the patient's hyperproliferative disease, comprising performing one or more multiplex amplification reactions with a sample from the patient or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0186] In some embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a patient to monitor progression and / or treatment of the patient's hyperproliferative disease, comprising performing one or more multiplex amplification reaction with a sample from the patient or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0187] In some embodiments, methods and compositions are provided for MRD monitoring for a patient having a hyperproliferative disease, comprising performing one or more multiplex amplification reaction with a patient's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and detecting the presence or absence of immune receptor sequence(s) in the sample associated with the hyperproliferative disease. In some embodiments, methods and compositions are provided for MRD monitoring for a patient having a hyperproliferative disease, comprising performing one or more multiplex amplification reaction with a patient's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and detecting the presence or absence of immune receptor sequence(s) in the sample associated with the hyperproliferative disease. In various embodiments, the one or more multiplex amplification reactions performed in such methods may be a single multiplex amplification reaction or may be two or more multiplex amplification reactions performed in parallel, for example parallel, highly multiplexed amplification reactions performed with different primer pools. Samples for use in MRD monitoring include, without limitation, samples obtained during a remission, samples obtained at any time following completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0188] In certain embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a subject in response to a treatment. In some embodiments, the methods and compositions are used to characterize and / or monitor populations or clones of tumor infiltrating lymphocytes (TILs) before, during, and / or following tumor treatment. In some embodiments, profiling immune receptor repertoires of TILs provides characterization and / or assessment of the tumor microenvironment and T cell expansion permissiveness within the tumor microenvironment. For example, a dearth of highly expanded TIL clones within the tumor, for example as indicated by higher evenness of T cell clone sizes through characterization of the TCR repertoire, may indicate a repressive tumor microenvironment. On the other hand, identification of multiple highly expanded T cell clones and less evenness of T cell clone sizes may indicate a tumor microenvironment permissive for T cell expansion. In some embodiments, the methods and compositions for determining immune repertoire are used to identify and / or track therapeutic T cell population(s) and B cell population(s). In some embodiments, the methods and compositions provided are used to identify and / or monitor the persistence of cell-based therapies following patient treatment, including but not limited to, presence (e.g., persistent presence) of engineered T cell populations including without limitation CAR-T cell populations, TCR engineered T cell populations, persistent CAR-T expression, presence (e.g., persistent presence) of administered TIL populations, TIL expression (e.g., persistent expression) following adoptive T-cell therapy, and / or immune reconstitution after allogeneic hematopoietic cell transplantation.
[0189] In some embodiments, the methods and compositions provided are used to characterize and / or monitor T cell clones or populations present in patient sample following administration of cell-based therapies to the patient, including but not limited to, e.g., cancer vaccine cells, CAR-T, TIL, and / or other engineered T cell-based therapy. In some embodiments, the provided methods and compositions are used to characterize and / or monitor immune repertoire in a patient sample following cell-based therapies in order to assess and / or monitor the patient's response to the administered cell-based therapy. Samples for use in such characterizing and / or monitoring following cell-based therapy include, without limitation, circulating blood cells, circulating tumor cells, TILs, tissue, and tumor sample(s) from a patient.
[0190] In some embodiments, methods and compositions are provided for monitoring T cell-based therapy for a patient receiving such therapy, comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and detecting the presence or absence of immune receptor sequence(s) in the sample associated with the T cell-based therapy. In some embodiments, methods and compositions are provided for monitoring T cell-based therapy for a patient receiving such therapy, comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and detecting the presence or absence of immune receptor sequence(s) in the sample associated with the T cell-based therapy.
[0191] In some embodiments, methods and compositions are provided for monitoring a patient's response following administration of a T cell-based therapy, comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and comparing the identified immune repertoire to the immune receptor sequence(s) identified in samples obtained from the patient at a different time. In some embodiments, methods and compositions are provided for monitoring a patient's response following administration of a T cell-based therapy, comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor, and comparing the identified immune repertoire to the immune receptor sequence(s) identified in samples obtained from the patient at a different time. T cell-based therapies suitable for such monitoring include, without limitation, CAR-T cells, TCR engineered T cells, TILs, and other enriched autologous T cells. In various embodiments, the one or more multiplex amplification reactions performed in such methods may be a single multiplex amplification reaction or may be two or more multiplex amplification reactions performed in parallel, for example parallel, highly multiplexed amplification reactions performed with different primer pools. Samples for use in such monitoring include, without limitation, samples obtained prior to a diagnosis, samples obtained at any stage of diagnosis, samples obtained during a remission, samples obtained at any time prior to a treatment (pre-treatment sample), samples obtained at any time following completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0192] In some embodiments, the methods and compositions for determining T cell and / or B cell receptor repertoires are used to measure and / or assess immunocompetence before, during, and / or following a treatment, including without limitation, solid organ transplant or bone marrow transplant. For example, the diversity of the T cell receptor beta repertoire can be used to measure immunocompetence and immune cell reconstitution following a hematopoietic stem cell transplant treatment. Also, the rate of change in diversity of the TRB repertoire between time points following a transplant can be used to modify patient treatment.
[0193] In some embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a subject in response to a treatment, comprising obtaining a sample from the subject following initiation of a treatment, performing one or more multiplex amplification reactions with the sample or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V gene of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In some embodiments, the method further comprises comparing the identified immune repertoire from the sample obtained following treatment initiation to the immune repertoire from a sample of the patient obtained prior to treatment. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0194] In some embodiments, methods and compositions are provided for identifying and / or characterizing the immune repertoire of a subject in response to a treatment, comprising obtaining a sample from the subject following initiation of a treatment, performing one or more multiplex amplification reactions with the sample or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In some embodiments, the method further comprises comparing the identified immune repertoire from the sample obtained following treatment initiation to the immune repertoire from a sample of the patient obtained prior to treatment. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0195] In some embodiments, methods and compositions are provided for monitoring changes in the immune repertoire of a subject in response to a treatment, comprising performing one or more multiplex amplification reactions with a subject's or patient's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor from the sample, and comparing the identified immune repertoire to those identified in samples obtained from the subject at a different time. In some embodiments, methods and compositions are provided for monitoring changes in the immune repertoire of a subject in response to a treatment, comprising performing one or more multiplex amplification reactions with a subject's or patient's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor from the sample, and comparing the identified immune repertoire to those identified in samples obtained from the subject at a different time. In various embodiments, the one or more multiplex amplification reactions performed in such methods may be a single multiplex amplification reaction or may be two or more multiplex amplification reactions performed in parallel, for example parallel, highly multiplexed amplification reactions performed with different primer pools. Samples for use in monitoring changes in immune repertoire include, without limitation, samples obtained prior to a diagnosis, samples obtained at any stage of diagnosis, samples obtained during a remission, samples obtained at any time prior to a treatment (pre-treatment sample), samples obtained at any time following completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0196] In certain embodiments, the methods and compositions provided are used to characterize and / or monitor immune repertoires associated with immune system-mediated adverse event(s), including without limitation, those associated with inflammatory conditions, autoimmune reactions, and / or autoimmune diseases or disorders. In some embodiments, the methods and compositions provided are used to identify and / or monitor T cell and / or B cell immune repertoires associated with chronic autoimmune diseases or disorders including, without limitation, multiple sclerosis, Type I diabetes, narcolepsy, rheumatoid arthritis, ankylosing spondylitis, asthma, and SLE. In some embodiments, a systemic sample, such as a blood sample, is used to determine the immune repertoire(s) of an individual with an autoimmune condition. In some embodiments, a localized sample, such as a fluid sample from an affected joint or region of swelling, is used to determine the immune repertoire(s) of an individual with an autoimmune condition. In some embodiments, comparison of the immune repertoire found in a localized or affected area sample to the immune repertoire found in the systemic sample can identify clonal T or B cell populations to be targeted for removal.
[0197] In some embodiments, methods and compositions are provided for identifying and / or monitoring an immune repertoire associated with a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a sample from the patient or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In some embodiments, the method further comprises comparing the identified immune repertoire from the sample to an identified immune repertoire from a sample from the patient obtained at a different time. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence and identifying a productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0198] In some embodiments, methods and compositions are provided for identifying and / or monitoring an immune repertoire associated with a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a sample from the patient or with cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor thereby generating immune receptor amplicon molecules. The method further comprises sequencing the resulting immune receptor amplicon molecules, determining the sequences of the immune receptor amplicon molecules, and identifying immune repertoire for the target immune receptor from the sample. In some embodiments, the method further comprises comparing the identified immune repertoire from the sample to an identified immune repertoire from a sample from the patient obtained at a different time. In particular, embodiments determining the sequence of the immune receptor amplicon molecules includes obtaining initial sequence reads, adding the inferred J gene sequence to the sequence read to create an extended sequence read, aligning the extended sequence read to a reference sequence and identifying productive reads, correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising i) a plurality of V gene primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective immune receptor coding sequence, wherein each set of i) and ii) primers directed to the same target immune receptor sequences is selected from the group consisting of a T cell receptor and an antibody receptor.
[0199] In some embodiments, methods and compositions are provided for identifying and / or monitoring an immune repertoire associated with progression and / or treatment of a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1, FR2 or FR3 within the V gene, and ii) one or more C gene primers directed to at least a portion of the respective target C gene of the immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor from the sample, and comparing the identified immune repertoire to the immune repertoire(s) identified in samples obtained from the patient at a different time. In some embodiments, methods and compositions are provided for identifying and / or monitoring an immune repertoire associated with progression and / or treatment of a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a patient's sample to amplify immune repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers directed to a majority of different V genes of at least one immune receptor coding sequence comprising at least a portion of FR1 or FR3 within the V gene, and ii) a plurality of J gene primers directed to a majority of different J genes of the respective target immune receptor coding sequence, sequencing the resultant immune receptor amplicons, identifying immune repertoire sequences for the target immune receptor from the sample, and comparing the identified immune repertoire to the immune repertoire(s) identified in samples obtained from the patient at a different time. In various embodiments, the one or more multiplex amplification reactions performed in such methods may be a single multiplex amplification reaction or may be two or more multiplex amplification reactions performed in parallel, for example parallel, highly multiplexed amplification reactions performed with different primer pools. Samples for use in monitoring changes in immune repertoire associated with immune system-mediated adverse event(s) include, without limitation, samples obtained prior to a diagnosis, samples obtained at any stage of diagnosis, samples obtained during a remission, samples obtained at any time prior to a treatment (pre-treatment sample), samples obtained at any time following completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0200] In some embodiments, the methods and compositions provided are used to characterize and / or monitor immune repertoires associated with passive immunity, including naturally acquired passive immunity and artificially acquired passive immunity therapies. For example, the methods and compositions provided may be used to identify and / or monitor protective antibodies that provide passive immunity to the recipient following transfer of antibody-mediated immunity to the recipient, including without limitation, antibody-mediated immunity conveyed from a mother to a fetus during pregnancy or to an infant through breast-feeding, or conveyed via administration of antibodies to a recipient. In another example, the methods and compositions provided may be used to identify and / or monitor B cell and / or T cell immune repertoires associated with passive transfer of cell-mediated immunity to a recipient, such as the administration of mature circulating lymphocytes to a recipient histocompatible with the donor. In some embodiments, the methods and compositions provided are used to monitor the duration of passive immunity in a recipient.
[0201] In some embodiments, the methods and compositions provided are used to characterize and / or monitor immune repertoires associated with active immunity or vaccination therapies. For example, following exposure to a vaccine or infectious agent, the methods and compositions provided may be used to identify and / or monitor protective antibodies or protective clonal B cell or T cell populations that may provide active immunity to the exposed individual. In some embodiments, the methods and compositions provided are used to monitor the duration of B or T cell clones which contribute to immunity in an exposed individual. In some embodiments, the methods and compositions provided are used to identify and / or monitor B cell and / or T cell immune repertoires associated with exposure to bacterial, fungal, parasitic, or viral antigens. In some embodiments, the methods and compositions provided are used to identify and / or monitor B cell and / or T cell immune repertoires associated with bacterial, fungal, parasitic, or viral infection.
[0202] In some embodiments, the methods and compositions provided are used to screen or characterize lymphocyte populations which are grown and / or activated in vitro for use as immunotherapeutic agents or in immunotherapeutic-based regimens. In some embodiments, the methods and compositions provided are used to screen or characterize TIL populations or other harvested T cell populations which are grown and / or activated in vitro, for example, TILs or other harvested T cells grown and / or activated for use in adoptive immunotherapy. In some embodiments, the methods and compositions provided are used to screen or characterize CAR-T populations or other engineered T cell populations which are grown and / or activated in vitro, for use, for example, in immunotherapy.
[0203] In some embodiments, the methods and compositions provided are used to assess cell populations by monitoring immune repertoires during ex vivo workflows for manufacturing engineered T cell preparations, for example, for quality control or regulatory testing purposes.
[0204] In some embodiments, the sequences of novel or non-canonical TCR or BCR alleles identified as described herein may be used to generate recombinant TCR or BCR nucleic acids or molecules. For example, as described herein, used of the provided methods and compositions led to the identification of fifteen TRB allelic variants not found in the IMGT database. Such novel or non-canonical allele sequence information and amplicons can be used to generate new recombinant TRB allelic variants and / or nucleic acids encoding the same.
[0205] In some embodiments, the methods and compositions provided are used in the screening and / or production of recombinant antibody libraries. Compositions provided which directed to identifying BCRs can be used to rapidly evaluate recombinant antibody library size and composition to identify antibodies of interest.
[0206] In some embodiments, profiling immune receptor repertoires as provided herein may be combined with profiling immune response gene expression to provide characterization of the tumor microenvironment. In some embodiments, combining or correlating a tumor sample's immune receptor repertoire profile with a targeted immune response gene expression profile provides a more thorough analysis of the tumor microenvironment and may suggest or provide guidance for immunotherapy treatments.
[0207] Suitable cells for analysis include, without limitation, various hematopoietic cells, lymphocytes, and tumor cells, such as peripheral blood mononuclear cells (PBMCs), T cells, B cells, circulating tumor cells, and tumor infiltrating lymphocytes (TILs). Lymphocytes expressing immunoglobulin include pre-B cells, B-cells, e.g. memory B cells, and plasma cells. Lymphocytes expressing T cell receptors include thymocytes, NK cells, pre-T cells and T cells, where many subsets of T cells are known in the art, e.g. Th1, Th2, Th17, CTL, T reg, etc. For example, in some embodiments, a sample comprising PBMCs may be used as a source for TCR and / or antibody immune repertoire analysis. The sample may contain, for example, lymphocytes, monocytes, and macrophages as well as antibodies and other biological constituents.
[0208] Analysis of the immune repertoire is of interest for conditions involving cellular proliferation and antigenic exposure, including without limitation, the presence of cancer, exposure to cancer antigens, exposure to antigens from an infectious agent, exposure to vaccines, exposure to allergens, exposure to food stuffs, presence of a graft or transplant, and the presence of autoimmune activity or disease. Conditions associated with immunodeficiency are also of interest for analysis, including congenital and acquired immunodeficiency syndromes.
[0209] B cell lineage malignancies of interest include, without limitation, multiple myeloma; acute lymphocytic leukemia (ALL); relapsed / refractory B cell ALL, chronic lymphocytic leukemia (CLL); diffuse large B cell lymphoma; mucosa-associated lymphatic tissue lymphoma (MALT); small cell lymphocytic lymphoma; mantle cell lymphoma (MCL); Burkitt lymphoma; mediastinal large B cell lymphoma; Waldenström macroglobulinemia; nodal marginal zone B cell lymphoma (NMZL); splenic marginal zone lymphoma (SMZL); intravascular large B-cell lymphoma; primary effusion lymphoma; lymphomatoid granulomatosis, etc. Non-malignant B cell hyperproliferative conditions include monoclonal B cell lymphocytosis (MBL).
[0210] T cell lineage malignancies of interest include, without limitation, precursor T-cell lymphoblastic lymphoma; T-cell prolymphocytic leukemia; T-cell granular lymphocytic leukemia; aggressive NK cell leukemia; adult T-cell lymphoma / leukemia (HTLV 1-positive); extranodal NK / T-cell lymphoma; enteropathy-type T-cell lymphoma; hepatosplenic γδ T-cell lymphoma; subcutaneous panniculitis-like T-cell lymphoma; mycosis fungoides / Sezary syndrome; anaplastic large cell lymphoma, T / null cell; peripheral T-cell lymphoma; angioimmunoblastic T-cell lymphoma; chronic lymphocytic leukemia (CLL); acute lymphocytic leukemia (ALL); prolymphocytic leukemia; and hairy cell leukemia.
[0211] Other malignancies of interest include, without limitation, acute myeloid leukemia, head and neck cancers, brain cancer, breast cancer, ovarian cancer, cervical cancer, colorectal cancer, endometrial cancer, gallbladder cancer, gastric cancer, bladder cancer, prostate cancer, testicular cancer, liver cancer, lung cancer, kidney (renal cell) cancer, esophageal cancer, pancreatic cancer, thyroid cancer, bile duct cancer, pituitary tumor, wilms tumor, kaposi sarcoma, osteosarcoma, thymus cancer, skin cancer, heart cancer, oral and larynx cancer, neuroblastoma and non-hodgkin lymphoma.
[0212] Neurological inflammatory conditions are of interest, e.g. Alzheimer's Disease, Parkinson's Disease, Lou Gehrig's Disease, etc. and demyelinating diseases, such as multiple sclerosis, chronic inflammatory demyelinating polyneuropathy, etc. as well as inflammatory conditions such as rheumatoid arthritis. Systemic lupus erythematosus (SLE) is an autoimmune disease characterized by polyclonal B cell activation, which results in a variety of anti-protein and non-protein autoantibodies (see Kotzin et al. (1996) Cell 85:303-306). These autoantibodies form immune complexes that deposit in multiple organ systems, causing tissue damage. An autoimmune component may be ascribed to atherosclerosis, where candidate autoantigens include Hsp60, oxidized LDL, and 2-Glycoprotein I (2GPI).
[0213] A sample for use in the methods described herein may be one that is collected from a subject with a malignancy or hyperproliferative condition, including lymphomas, leukemias, and plasmacytomas. A lymphoma is a solid neoplasm of lymphocyte origin, and is most often found in the lymphoid tissue. Thus, for example, a biopsy from a lymph node, e.g. a tonsil, containing such a lymphoma would constitute a suitable biopsy. Samples may be obtained from a subject or patient at one or a plurality of time points in the progression of disease and / or treatment of the disease.
[0214] In some embodiments, the disclosure provides methods for performing target-specific multiplex PCR on a cDNA sample having a plurality of expressed immune receptor target sequences using primers having a cleavable group.
[0215] In certain embodiments, library and / or template preparation to be sequenced are prepared automatically from a population of nucleic acid samples using the compositions provided herein using an automated systems, e.g., the Ion Chef ™< system.
[0216] As used herein, the term "subject" includes a person, a patient, an individual, someone being evaluated, etc.
[0217] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of features is not necessarily limited only to those features but may include other features not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive-or and not to an exclusive-or.
[0218] As used herein, "antigen" refers to any substance that, when introduced into a body, e.g., of a subject, can stimulate an immune response, such as the production of an antibody or T cell receptor that recognizes the antigen. Antigens include molecules such as nucleic acids, lipids, ribonucleoprotein complexes, protein complexes, proteins, polypeptides, peptides and naturally occurring or synthetic modifications of such molecules against which an immune response involving T and / or B lymphocytes can be generated. With regard to autoimmune disease, the antigens herein are often referred to as autoantigens. With regard to allergic disease the antigens herein are often referred to as allergens. Autoantigens are any molecule produced by the organism that can be the target of an immunologic response, including peptides, polypeptides, and proteins encoded within the genome of the organism and post-translationally-generated modifications of these peptides, polypeptides, and proteins. Such molecules also include carbohydrates, lipids and other molecules produced by the organism. Antigens also include vaccine antigens, which include, without limitation, pathogen antigens, cancer associated antigens, allergens, and the like.
[0219] As used herein, "amplify", "amplifying" or "amplification reaction" and their derivatives, refer to any action or process whereby at least a portion of a nucleic acid molecule (referred to as a template nucleic acid molecule) is replicated or copied into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally includes sequence that is substantially identical or substantially complementary to at least some portion of the template nucleic acid molecule. The template nucleic acid molecule can be single-stranded or double-stranded and the additional nucleic acid molecule can independently be single-stranded or double-stranded. In some embodiments, amplification includes a template-dependent in vitro enzyme-catalyzed reaction for the production of at least one copy of at least some portion of the nucleic acid molecule or the production of at least one copy of a nucleic acid sequence that is complementary to at least some portion of the nucleic acid molecule. Amplification optionally includes linear or exponential replication of a nucleic acid molecule. In some embodiments, such amplification is performed using isothermal conditions; in other embodiments, such amplification can include thermocycling. In some embodiments, the amplification is a multiplex amplification that includes the simultaneous amplification of a plurality of target sequences in a single amplification reaction. At least some of the target sequences can be situated on the same nucleic acid molecule or on different target nucleic acid molecules included in the single amplification reaction. In some embodiments, "amplification" includes amplification of at least some portion of DNA- and RNA-based nucleic acids alone, or in combination. The amplification reaction can include single or double-stranded nucleic acid substrates and can further including any of the amplification processes known to one of ordinary skill in the art. In some embodiments, the amplification reaction includes polymerase chain reaction (PCR).
[0220] As used herein, "amplification conditions" and its derivatives, refers to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, the amplification conditions can include isothermal conditions or alternatively can include thermocycling conditions, or a combination of isothermal and thermocycling conditions. In some embodiments, the conditions suitable for amplifying one or more nucleic acid sequences includes polymerase chain reaction (PCR) conditions. Typically, the amplification conditions refer to a reaction mixture that is sufficient to amplify nucleic acids such as one or more target sequences, or to amplify an amplified target sequence ligated to one or more adapters, e.g., an adapter-ligated amplified target sequence. Amplification conditions include a catalyst for amplification or for nucleic acid synthesis, for example a polymerase; a primer that possesses some degree of complementarity to the nucleic acid to be amplified; and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs) to promote extension of the primer once hybridized to the nucleic acid. The amplification conditions can require hybridization or annealing of a primer to a nucleic acid, extension of the primer and a denaturing step in which the extended primer is separated from the nucleic acid sequence undergoing amplification. Typically, but not necessarily, amplification conditions can include thermocycling; in some embodiments, amplification conditions include a plurality of cycles where the steps of annealing, extending and separating are repeated. Typically, the amplification conditions include cations such as Mg 2+< or Mn 2+< (e.g., MgCl 2 , etc) and can also include various modifiers of ionic strength.
[0221] As used herein, "target sequence" or "target sequence of interest" and its derivatives, refers to any single or double-stranded nucleic acid sequence that can be amplified or synthesized according to the disclosure, including any nucleic acid sequence suspected or expected to be present in a sample. In some embodiments, the target sequence is present in double-stranded form and includes at least a portion of the particular nucleotide sequence to be amplified or synthesized, or its complement, prior to the addition of target-specific primers or appended adapters. Target sequences can include the nucleic acids to which primers useful in the amplification or synthesis reaction can hybridize prior to extension by a polymerase. In some embodiments, the term refers to a nucleic acid sequence whose sequence identity, ordering or location of nucleotides is determined by one or more of the methods of the disclosure.
[0222] As defined herein, "sample" and its derivatives, is used in its broadest sense and includes any specimen, culture and the like that is suspected of including a target. In some embodiments, the sample comprises cDNA, RNA, PNA, LNA, chimeric, hybrid, or multiplex-forms of nucleic acids. The sample can include any biological, clinical, surgical, agricultural, atmospheric or aquatic-based specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample such as expressed RNA, fresh-frozen or formalin-fixed paraffin-embedded nucleic acid specimen.
[0223] As used herein, "contacting" and its derivatives, when used in reference to two or more components, refers to any process whereby the approach, proximity, mixture or commingling of the referenced components is promoted or achieved without necessarily requiring physical contact of such components, and includes mixing of solutions containing any one or more of the referenced components with each other. The referenced components may be contacted in any particular order or combination and the particular order of recitation of components is not limiting. For example, "contacting A with B and C" encompasses embodiments where A is first contacted with B then C, as well as embodiments where C is contacted with A then B, as well as embodiments where a mixture of A and C is contacted with B, and the like. Furthermore, such contacting does not necessarily require that the end result of the contacting process be a mixture including all of the referenced components, as long as at some point during the contacting process all of the referenced components are simultaneously present or simultaneously included in the same mixture or solution. Where one or more of the referenced components to be contacted includes a plurality (e.g., "contacting a target sequence with a plurality of target-specific primers and a polymerase"), then each member of the plurality can be viewed as an individual component of the contacting process, such that the contacting can include contacting of any one or more members of the plurality with any other member of the plurality and / or with any other referenced component (e.g., some but not all of the plurality of target specific primers can be contacted with a target sequence, then a polymerase, and then with other members of the plurality of target-specific primers) in any order or combination.
[0224] As used herein, the term "primer" and its derivatives refer to any polynucleotide that can hybridize to a target sequence of interest. In some embodiments, the primer can also serve to prime nucleic acid synthesis. Typically, the primer functions as a substrate onto which nucleotides can be polymerized by a polymerase; in some embodiments, however, the primer can become incorporated into the synthesized nucleic acid strand and provide a site to which another primer can hybridize to prime synthesis of a new strand that is complementary to the synthesized nucleic acid molecule. The primer may be comprised of any combination of nucleotides or analogs thereof, which may be optionally linked to form a linear polymer of any suitable length. In some embodiments, the primer is a single-stranded oligonucleotide or polynucleotide. (For purposes of this disclosure, the terms "polynucleotide" and "oligonucleotide" are used interchangeably herein and do not necessarily indicate any difference in length between the two). In some embodiments, the primer is single-stranded but it can also be double-stranded. The primer optionally occurs naturally, as in a purified restriction digest, or can be produced synthetically. In some embodiments, the primer acts as a point of initiation for amplification or synthesis when exposed to amplification or synthesis conditions; such amplification or synthesis can occur in a template-dependent fashion and optionally results in formation of a primer extension product that is complementary to at least a portion of the target sequence. Exemplary amplification or synthesis conditions can include contacting the primer with a polynucleotide template (e.g., a template including a target sequence), nucleotides and an inducing agent such as a polymerase at a suitable temperature and pH to induce polymerization of nucleotides onto an end of the target-specific primer. If double-stranded, the primer can optionally be treated to separate its strands before being used to prepare primer extension products. In some embodiments, the primer is an oligodeoxyribonucleotide or an oligoribonucleotide. In some embodiments, the primer can include one or more nucleotide analogs. The exact length and / or composition, including sequence, of the target-specific primer can influence many properties, including melting temperature (T m ), GC content, formation of secondary structures, repeat nucleotide motifs, length of predicted primer extension products, extent of coverage across a nucleic acid molecule of interest, number of primers present in a single amplification or synthesis reaction, presence of nucleotide analogs or modified nucleotides within the primers, and the like. In some embodiments, a primer can be paired with a compatible primer within an amplification or synthesis reaction to form a primer pair consisting or a forward primer and a reverse primer. In some embodiments, the forward primer of the primer pair includes a sequence that is substantially complementary to at least a portion of a strand of a nucleic acid molecule, and the reverse primer of the primer of the primer pair includes a sequence that is substantially identical to at least of portion of the strand. In some embodiments, the forward primer and the reverse primer are capable of hybridizing to opposite strands of a nucleic acid duplex. Optionally, the forward primer primes synthesis of a first nucleic acid strand, and the reverse primer primes synthesis of a second nucleic acid strand, wherein the first and second strands are substantially complementary to each other, or can hybridize to form a double-stranded nucleic acid molecule. In some embodiments, one end of an amplification or synthesis product is defined by the forward primer and the other end of the amplification or synthesis product is defined by the reverse primer. In some embodiments, where the amplification or synthesis of lengthy primer extension products is required, such as amplifying an exon, coding region, or gene, several primer pai...
Claims
1. A method for predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy comprising: a) performing sequencing of target immune receptor nucleic acid template molecules derived from a biological sample from a subject candidate for an immunotherapy, wherein the target immune receptor nucleic acid template molecules comprise FR1, CDR1, FR2, CDR2, FR3, and CDR3 coding regions of the target immune receptor and wherein the sequencing is by next generation sequencing; b) determining the sequence of the target immune receptor repertoire of the sample based on the sequencing; c) identifying the immune receptor haplotype of the subject from the determined sequences; and d) predicting a subject's potential or predisposition to be protected from or vulnerable to an adverse event following an immunotherapy by comparing the identified immune receptor haplotype of the subject to a reference set of immune receptor haplotypes of individuals with annotated adverse events following immunotherapy treatments.
2. The method of claim 1, further comprising performing a multiplex amplification reaction to amplify target immune receptor nucleic acid template molecules before step a), wherein the multiplex amplification reaction comprises a plurality of amplification primer pairs including a plurality of variable (V) gene primers directed to a majority of V genes of the target immune receptor.
3. The method of claim 1, wherein step b) includes obtaining initial sequence reads, aligning the initial sequence read to a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads.
4. The method of claim 3, wherein the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads or at least 60% of the sequencing reads.
5. The method of claim 2, wherein the plurality of V gene primers anneal to at least a portion of the FR1 regions of the target immune receptor nucleic acid template molecules.
6. The method of claim 5, wherein the plurality of V gene primers are selected from the primers of Table 2.
7. The method of claim 2, wherein the plurality of amplification primer pairs includes one or more primers that anneal to at least a portion of the C gene portion of the target immune receptor nucleic acid template molecules.
8. The method of claim 7, wherein the one or more primers is selected from the primers of Table 4.
9. The method of claim 2, wherein the plurality of amplification primer pairs includes at least 10 primers that anneal to at least a portion of the J gene portion of the target immune receptor nucleic acid template molecules.
10. The method of claim 9, wherein the at least 10 primers are selected from the primers of Table 5.
11. The method of any preceding claim, wherein the subject has a cancer, autoimmune disease or chronic viral infection.
12. The method of any preceding claim, wherein target immune receptor is a T cell receptor selected from the group consisting of TCR alpha, TCR beta, TCR gamma and TCR delta, preferably wherein the target immune receptor is TCR beta.
13. The method of any preceding claim, wherein the immunotherapy comprises a checkpoint blockade agent, a tumor cell vaccine or a dendritic cell vaccine.
14. The method of any preceding claim, wherein the target immune receptor nucleic acid template molecules are derived from RNA from the sample or comprise genomic DNA having rearranged VDJ or VJ gene segments.
15. The method of any preceding claim, wherein the biological sample is selected from the group consisting of peripheral blood mononuclear cells (PBMCs), T cells, tumor infiltrating lymphocytes (TILs), and peripheral blood.