Immune repertoire biomarkers in autoimmune and immunodeficiency diseases
The method of multiplex amplification and sequencing of B-cell immune repertoire nucleic acid sequences addresses the challenge of high-resolution immune repertoire analysis, allowing for precise prediction of treatment response and diagnosis of immune disorders, enhancing therapeutic efficacy.
Patent Information
- Application Number
- JP2022544702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-22
- Filing Date
- 2021-01-22
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-01-22
AI Technical Summary
Current methods for high-resolution analysis of the immune repertoire are limited in capturing the entire repertoire efficiently and accurately, hindering diagnostic and therapeutic capabilities for autoimmune diseases and immunodeficiencies.
A method involving multiplex amplification and sequencing of B-cell immune repertoire nucleic acid sequences, including the use of V gene and C gene primers to generate amplicon molecules, followed by sequencing and correction of indels, to determine somatic hypermutation and class switch recombination frequencies, enabling prediction of treatment response and diagnosis of immune disorders.
Enables accurate prediction of treatment response and prognosis for autoimmune diseases and immunodeficiencies, reducing the time to identify effective therapies and improving patient outcomes.
Smart Images

Figure 0007787078000034 
Figure 0007787078000035 
Figure 0007787078000036
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority under and the benefit under 35 U.S.C. §119(e) of each of U.S. Provisional Application No. 62 / 964,524, filed January 22, 2020, and U.S. Provisional Application No. 62 / 964,550, filed January 22, 2020. The entire contents of each of the foregoing applications are incorporated herein by reference in their entirety.
[0002] Sequence Listing This application incorporates by reference the material in the concurrently filed Electronic Sequence Listing, which was submitted as a text (.txt) file entitled "LT01523_ST25.txt," created on January 21, 2021, has a file size of 185 KB, and is incorporated herein by reference in its entirety.
[0003] The present invention relates to methods for preparing libraries of target immune repertoire nucleic acid sequences, as well as compositions and uses therefor. [Background technology]
[0004] The adaptive immune response involves the selective response of B cells and T cells that recognize antigens. The immunoglobulin genes encoding the antigen receptors for antibodies (Ab, in B cells) and T cell receptors (TCR, in T cells) comprise complex genetic loci, and a wide variety of receptors is generated as a result of recombination of the respective variable (V), diversity (D), and joining (J) gene segments, followed by somatic hypermutation events during early lymphocyte differentiation. The recombination process occurs separately for both subunit chains of each receptor, and subsequent heterodimer pairing generates even greater recombination diversity. Calculations of the potential combinations and junction possibilities contributing to the human immune receptor repertoire have estimated that the number of possibilities greatly exceeds the total number of peripheral B or T cells in an individual. See, for example, Davis and Bjorkman (1988) Nature 334:395-402, Arstila et al. (1999) Science 286:958-961, van Dongen et al., In: Leukemia, Henderson et al. (eds) Philadelphia: WB Saunders Company, 2002, pp 85-129.
[0005] Significant efforts have been made over the years to improve high-resolution analysis of the immune repertoire. A means for the specific detection and monitoring of expanded lymphocyte clones would provide important opportunities for the characterization and analysis of normal and pathogenic immune responses. Despite these efforts, effective high-resolution analysis remains a challenge. Low-throughput techniques such as Sanger sequencing can provide resolution but are limited in providing an efficient means for broadly capturing the entire immune repertoire. Advances in next-generation sequencing (NGS) have provided access to capturing the repertoire, but due to the nature of the large number of sequences involved and the resulting introduction of sequencing errors, efficient and effective reflection of the true repertoire has proven difficult. Improved sequencing methodologies and workflows capable of analyzing complex populations of highly variable immune cell receptor sequences have been developed. Novel methods for effectively profiling the vast repertoire of immune cell receptors continue to be needed to better understand immune cell responses, enhance diagnostic and therapeutic capabilities, and devise novel therapies. Summary of the Invention
[0006] Provided herein are methods for predicting the clinical response of a subject with an autoimmune disease or disorder to a therapy based on characterizing the subject's B-cell immune repertoire. Further provided herein are methods for predicting the prognosis of a subject with leukemia based on characterizing the subject's B-cell immune repertoire. Such methods include using a single multiplex amplification reaction to amplify target BCR nucleic acid template molecules obtained from a subject's sample: i) (a) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR1 within the V gene; (b) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR2 within the V gene; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence that includes at least a portion of an FR3 within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of at least one BCR coding sequence; Each set of primers i) and ii) is directed to the coding sequence of the same target IgH BCR gene, and amplification using at least one set of primers i) and ii) results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target BCR amplicon molecules containing the target BCR repertoire. The method further includes sequencing the target BCR amplicon molecules to determine the sequence of the molecules, where sequencing includes obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, and correcting one or more indels to create rescued productive sequence reads; identifying a BCR repertoire clonal population from the determined target BCR sequences; and determining the frequency of somatic hypermutation (SHM) within the variable gene portions between the clones. The method further includes determining the frequency of ongoing SHM, in which the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and / or determining the frequency of ongoing class switch recombination (CSR). In some embodiments, at least one and / or combination of switched isotypes are identified with the same variable lineage of B cell immune receptor clones in the sample. In some embodiments, a method is provided that then includes identifying the subject having an autoimmune disease or disorder as a likely responder to chemotherapy if the SHM frequency is below a frequency threshold in non-switched IgM / IgD-expressing B cells in the sample and the immune repertoire is dominated by a high frequency of switched isotype IgG, IgA, or IgE-expressing B cells in the sample, a likely non-responder to chemotherapy if the SHM frequency is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells in the sample, and / or a likely responder to immunotherapy if the SHM frequency is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells in the sample.In some embodiments, methods are provided that include then classifying the repertoire clones according to the following subclasses: Class I: no ongoing CSR or SHM, no V-gene SHM; Class II: no ongoing CSR or SHM, V-gene SHM greater than 0 and less than 6%; Class III: no ongoing CSR or SHM, V-gene SHM greater than about 6%; Class IV: ongoing CSR and / or SHM; and assigning a prognosis for the subject with immunodeficiency based on the subclassification of the B-cell immune repertoire as follows: Class I: worst prognosis; Class II: poor prognosis; Class III: good prognosis; Class IV: best prognosis.
[0007] Additionally provided are methods for diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immune deficiency disorder based on characterizing the subject's B cell immune repertoire. In some embodiments, such methods comprise: a single multiplex amplification reaction to amplify target BCR nucleic acid template molecules obtained from a sample of the subject; i) (a) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR1 within the V gene; (b) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR2 within the V gene; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence that includes at least a portion of an FR3 within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of at least one BCR coding sequence; Each set of primers i) and ii) is directed to the coding sequence of the same target IgH BCR gene, and amplification using at least one set of primers i) and ii) results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target BCR amplicon molecules comprising the target BCR repertoire. The method further includes sequencing the target BCR amplicon molecules to determine the sequence of the molecules, where sequencing comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, and correcting one or more indel errors to create rescued productive sequence reads; identifying a BCR repertoire clonal population from the sequencing to determine a level of somatic hypermutation (SHM) within the variable gene portions among the immune receptor clones, where the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences; and determining the class switch recombination frequency of B cell immune receptor clones in the sample. The method further includes identifying the subject as having a primary immunodeficiency disorder if the SHM frequency is below a frequency threshold in the switched isotype, the frequency of class switch recombination (CSR) in the sample is below a frequency threshold, and the immune repertoire is dominated by non-switched IgM / IgD-expressing B cells in the sample, thereby diagnosing the subject as having a chronic variable immunodeficiency disorder.
[0008] Further provided herein are methods for treating a subject with an autoimmune disease or disorder or an immunodeficiency based on the subject's B cell immune repertoire. In some embodiments, such methods comprise: a single multiplex amplification reaction to amplify target BCR nucleic acid template molecules obtained from a sample of the subject; i) (a) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR1 within the V gene; (b) a plurality of V gene primers for a majority of V genes that differ in at least one BCR coding sequence that includes at least a portion of FR2 within the V gene; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence that includes at least a portion of an FR3 within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of at least one BCR coding sequence; Each set of primers i) and ii) is directed to the coding sequence of the same target IgH BCR gene, and performing amplification using at least one set of primers i) and ii) results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target BCR amplicon molecules comprising the target BCR repertoire. The method further includes sequencing the target BCR amplicon molecules to determine the sequences of the molecules, where sequencing includes obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, and correcting one or more indel errors to create rescued productive sequence reads; identifying a BCR repertoire clonal population from the sequencing; identifying immune receptor clones from the sequencing; and identifying a level of somatic hypermutation (SHM) within the variable gene portion between the immune receptors. The method further comprises determining an ongoing SHM frequency in which the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and / or determining an ongoing class switch recombination (CSR) frequency in which at least one or a combination of switched isotypes are identified in the same variable lineage of B cell immune receptor clones in the sample. In some embodiments, methods are provided that further comprise treating a subject with an autoimmune disease or disorder (i) with chemotherapy if the SHM frequency in the sample is below a frequency threshold in non-switched IgM / IgD-expressing B cells and the immune repertoire in the sample is dominated by a high frequency of switched isotype IgG, IgA, or IgE-expressing B cells, and / or (ii) with immunotherapy if the SHM frequency in the sample is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells.In some embodiments, methods are provided that further include classifying the repertoire clones according to the following subclasses: Class I: no ongoing CSR or SHM, no V-gene SHM; Class II: no ongoing CSR or SHM, V-gene SHM greater than 0 and less than 6%; Class III: no ongoing CSR or SHM, V-gene SHM greater than about 6%; Class IV: ongoing CSR and / or SHM; and treating the immunocompromised subject based on the subclassification of the B-cell immune repertoire with Class I: stem cell therapy with the optional addition of any of chemotherapy, radiation, or DNA repair inducers; Class II: chemotherapy, radiation, or immunotherapy with the optional addition of DNA repair inducers; Class III: chemotherapy or immunotherapy; and Class IV: standard chemotherapy or immunotherapy. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram of an exemplary workflow for removing errors from PCR or sequencing using stepwise clustering of similar CDR3 nucleotide sequences, including: (A) very fast heuristic clustering (cd-hit-est) into groups based on similarity; (B) a representative cluster is selected as the most common sequences picked randomly for connection; (C) merging reads into the representatives; and (D) comparing the representatives and merging clusters if they are within an assigned Hamming distance. [Figure 2] Diagram of an exemplary workflow for removal of residual insertion / deletion (indel) errors by comparing homopolymer-collapsed CDR3 sequences using Levenshtein distance, with steps of (A) collapsing homopolymers and calculating Levenshtein distances between cluster representatives, (B) merging reads that now cluster together and represent complex indel errors, and (C) reporting the lineage to the user. [Figure 3]Graphs showing the results of characterization of the number of reads and the quality of the reads (productive vs. off-target or non-productive) in BCR assays only, BCR and TCR assays amplified in one pool, and BCR and TCR assays amplified in two separate pools. [Figure 4A-4B] (FIG. 4A) The total number of clones detected and (FIG. 4B) the populations of BCR and TCR clones (by percentage of the total) in the combined assay are shown. [Figures 5A-5G] Histograms showing sequence read lengths of the IgH repertoire from RNA of various cell or tissue samples: (Figure 5A) PBL, (Figure 5B) CD19+ cells, (Figure 5C) tonsil FFPE, (Figure 5D) lung tumor FFPE, (Figure 5E) bone marrow, (Figure 5F) normal spleen, and (Figure 5G) normal brain. [Figure 6] Sequence read lengths obtained after multiplex amplification of PBL cDNA using exemplary IgH V gene FR1-C gene primer sets 1 to 7 are shown. [Figures 7A-7B] Bar graphs show total isotype representation within PBL samples as sequence reads per isotype (Figure 7A) and clones detected per isotype (Figure 7B) obtained from assays using an exemplary IgH V gene FR1-C gene multiplex amplification reaction. [Figure 8A-8B] Histograms of IgH V gene mutation rates in PBL samples are shown for (FIG. 8A) all IgH isotypes and (FIG. 8B) IgD only. [Figure 9] Shown are the total productive reads from the sample (rightmost point of each graph) and IgH clonal analysis results in eight low-resolution processed datasets derived from the total productive reads. [Figure 10] 1 shows a graph depicting the linearity of plasmid detection in an IgH library generated from a pool of 20 control plasmids at equimolar concentrations mixed with leukocyte cDNA. Plasmids associated with the plasmid ID numbers are shown in Table 15. [Figures 11A-11B]Figure 11 shows the IgH clonal analysis results of a retrospective analysis of samples from rheumatoid arthritis patients treated with methotrexate. Figure 11A shows the frequency of IGG-, IGA-, and IGE-expressing B cells over the course of treatment, and Figure 11B shows the frequency of somatic hypermutation of IGM- and IGD-expressing B cells in responders and non-responders. [Figure 12] Figure 1 shows IgH clonal analysis results from a retrospective analysis of somatic hyperrecombination (SHM) and class switch recombination (CSR) in samples from lymphoma patients. Conventional classification by percentage of V gene mutations at 2 or 3 percent is also plotted. DETAILED DESCRIPTION OF THE INVENTION
[0010] Methotrexate is commonly used as a first-line treatment for rheumatoid arthritis, yet only a subset of recipients experience durable remission of disease symptoms. Those who do not respond favorably have the option of undergoing alternative therapies (e.g., TNF antagonists or B-cell depleting agents such as rituximab; see the expanded list on the website below). Due to the slow kinetics of methotrexate activity, the first signs of disease relief are typically observed 4–6 weeks after initiation of therapy, and assessing response requires long-term monitoring, during which time individuals may continue to suffer from disease symptoms. Furthermore, a subset of individuals may eventually revert to their pretreatment disease state after experiencing temporary remission of symptoms for several months. Therefore, there is a critical need for biomarkers that predict response to treatment (e.g., methotrexate). Such biomarkers would reduce the time required to identify effective therapies, thereby improving patient outcomes and lowering treatment costs.
[0011] B cell somatic hypermutation (SHM) and class switch recombination (CSR) are mechanistically related but distinct processes that require precisely targeted generation and repair of single- and double-stranded DNA breaks. Without being bound by theory, we hypothesize that in the context of leukemia (such as CLL), the presence of ongoing SHM or CSR may reveal the functionality of DNA damage repair pathways, potentially linked to therapeutic strategies involving DNA break generation (e.g., chemoimmunotherapy) or inhibition of DNA repair mechanisms (e.g., PARP inhibition). To date, prognosis of leukemia (e.g., CLL) has relied solely on quantification of SHM within the IGHV region. There remains an unmet need for biomarkers that predict prognosis and / or response to treatment in immunodeficiency disorders. Such biomarkers would reduce the time required to identify effective treatments, thereby improving patient outcomes and lowering treatment costs.
[0012] The present inventors have developed methods for predicting a subject's prognosis and / or response to treatment based on characterizing the subject's B cell immune repertoire before receiving the treatment. In one embodiment, the present invention provides a method for predicting the clinical response of a subject with an autoimmune disease to treatment by determining the somatic hypermutation and class switch recombination frequencies of the subject's B cell immune repertoire before receiving the treatment. In one embodiment, the present invention provides a method for predicting the clinical response of a subject with an immune deficiency (e.g., leukemia) to treatment by determining the somatic hypermutation and class switch recombination frequencies of the subject's B cell immune repertoire before receiving the treatment.
[0013] In some embodiments, the methods, compositions, and assays provided herein for use in methods for predicting and / or treating a subject with an autoimmune disease's clinical responsiveness to a therapy include using methodologies for high-precision amplification and sequencing of immune receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) in a subject's sample to identify the levels and frequencies of somatic hyperrecombination (SHM) immune receptor subclass groups in a pre-treatment sample from the subject. The immune receptor sequencing data is used to identify immune receptor clones in the sample and the frequencies of all clones with somatic hyperrecombination mutations, along with the frequencies of switched and non-switched subclass types in the sample, as predictors of the subject's clinical response to the therapy. In some embodiments, the subject is treated with the therapy in a manner that depends on the SHM frequencies of immune receptor clones in the sample and the levels of switched and non-switched isotypes. For example, in some embodiments, a subject having an SHM frequency of IgM- and / or IgD-expressing B cells below a specified threshold (e.g., 8%) and an immunoreceptor clonal frequency dominated by a high frequency of switched isotype IgG-, IgA-, or IgE-expressing B cells indicates that the subject is likely to be responsive to chemotherapy (e.g., methotrexate) and is a candidate for chemotherapy (e.g., methotrexate), whereas a subject having an SHM frequency of IgM- and / or IgD-expressing B cells above a specified threshold (e.g., 8%) and an immunoreceptor clonal frequency dominated by a high frequency of non-switched IgM / IgD-expressing B cells indicates that the subject is unlikely to be responsive to chemotherapy (e.g., methotrexate) and is not a candidate for chemotherapy (e.g., methotrexate); rather, such subjects are likely to be more responsive to immunotherapy and / or B cell-depleting therapy, and such subjects may preferably be candidates for immunotherapy and / or B cell-depleting therapy.
[0014] In some embodiments, methods, compositions, and assays provided herein for use in predicting prognosis and / or clinical response to therapy in subjects with leukemia disease include using methodologies for high-precision amplification and sequencing of immune receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) to identify the levels and frequency of somatic hyperrecombination (SHM) and class-subtype recombination (CSR) in a sample from the subject. The immune receptor sequencing data is used to identify immune receptor clones in the immune receptors in the sample, as well as the frequency of all clones with ongoing somatic hyperrecombination mutations (SHM) and / or class-subtype recombination (CSR), as well as the frequency of SHM of clones in the sample as a predictor of the subject's prognosis and / or clinical response to therapy. In some embodiments, the subject is predicted to have a poor or good prognosis in a manner dependent on the frequency of ongoing SHM / CSR and SHM in the immune receptor clones in the sample. In some embodiments, subjects are classified according to the resulting SHM / CRS profile, e.g., Class I: no ongoing CSR or SHM, no V-gene SHM; Class II: no ongoing CSR or SHM, V-gene SHM greater than 0 and less than 6%; Class III: no ongoing CSR or SHM, V-gene SHM greater than about 6%; Class IV: ongoing CSR and / or SHM. In certain embodiments, subjects are predicted to have a particular prognosis based on such classification, e.g., Class I: worst prognosis; Class II: poor prognosis; Class III: good prognosis; Class IV: best prognosis.In some embodiments, subjects are alternatively and / or additionally treated with therapy in a manner dependent on the ongoing SHM / CRS and SHM frequency in the immune receptor clones in the sample, optionally following reclassification, e.g., in some embodiments, they may be candidates for stem cell therapy with the optional addition of any of chemotherapy, radiation, or DNA repair inducers; Class II: they may be candidates for chemotherapy, radiation, or immunotherapy with the optional addition of DNA repair inducers; Class III: they may be candidates for chemotherapy or immunotherapy; and Class IV: they may be candidates for standard chemotherapy or immunotherapy.
[0015] In additional embodiments of the present invention, the methods, compositions, and assays provided herein for use in diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immune deficiency disorder include using a methodology for high-precision amplification and sequencing of immune receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) in a subject's sample to identify the levels and frequencies of somatic hyperrecombination (SHM) immune receptor subclass groups in a pre-treatment sample from the subject. Using the immune receptor sequencing data, the frequencies of immune receptor clones and all clones with somatic hyperrecombination mutations in the immune receptors in the sample, along with the frequencies of switched and non-switched subclass types in the sample, are identified as diagnostic for a primary immune deficiency disorder. In some embodiments, the subject is diagnosed based on the SHM frequencies of immune receptor clones in the sample and the levels of switched and non-switched isotypes. For example, in some embodiments, subjects with very low and / or almost negligible levels of switched isotype IgG, IgA, or IgE-expressing B cells and an SHM frequency of IgM- and / or IgD-expressing B cells below a specified threshold (e.g., 8%) are diagnosed as having a primary immunodeficiency disorder, e.g., a chronic variable immunodeficiency disorder.
[0016] In some embodiments, provided methods involve identifying SHM and Ig isotype and / or CSR of immune receptor clones using V gene identity and sequence and C gene identity sequence. In some embodiments, provided methods involve analysis of immune receptor clones using sequences including CDR3 sequences, CDR1 and CDR3 sequences, or CDR2 and CDR3 sequences or CDR, CDR2, and CDR3 sequences, and C gene sequences. In some embodiments, provided methods involve identifying BCR clones as containing BCR variable and C gene rearrangements that are similar or identical in nucleotide sequence. For example, a significant portion of BCRs that differ from each other by one or a few residues may nevertheless have similar or identical specificity for an antigen, and therefore such BCRs may be considered related.
[0017] In one aspect of the present invention, methods are provided for predicting the clinical response to treatment of a subject with an autoimmune disorder and / or the prognosis of a subject with an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire. The provided methods include performing a multiplex amplification reaction to amplify target B cell immunoreceptor nucleic acid template molecules derived from a biological sample from the subject. In some embodiments, the provided methods include at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene; or (c) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of a C gene of at least one BCR coding sequence. Each set of primers i) and ii) is directed to a coding sequence of a target IgH BCR gene, and amplification results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target B cell immunoreceptor amplicon molecules comprising the target immune receptor repertoire. In some embodiments, the provided method further includes sequencing the target immune receptor repertoire amplicons, identifying immune receptor clones from the sequencing and determining a level of somatic hypermutation (SHM) within variable gene portions among the immune receptor clones, wherein the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and determining the subclass of the B cell immunoreceptor clones in the sample, as well as the frequency of non-switched IgM and / or IgD and switched IgG, IgA, and / or IgE-expressing cells in the sample.The method provides for identifying a subject as a likely responder to chemotherapy if the SHM frequency in the sample is below a frequency threshold in non-switched IgM / IgD-expressing B cells and the immune repertoire is dominated by a high frequency of switched isotype IgG, IgA, or IgE-expressing B cells; (ii) a non-responder to chemotherapy if the SHM frequency in the sample is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells; and / or (iii) a responder to immunotherapy if the SHM frequency in the sample is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells. In some embodiments, the provided methods further comprise sequencing the target immune receptor repertoire amplicon, identifying immune receptor clones from the sequencing and determining the level of somatic hypermutation (SHM) within variable gene portions among the immune receptor clones, wherein the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences; and further determining an ongoing SHM frequency where the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and / or an ongoing class switch recombination (CSR) frequency where at least one and / or combination of switched isotypes are identified in similar variable lineages of B cell immune receptor clones in the sample. The method provides for classifying the repertoire clones according to the following subclasses: Class I: no ongoing CSR or SHM, no V-gene SHM; Class II: no ongoing CSR or SHM, V-gene SHM greater than 0 and less than 6%; Class III: no ongoing CSR or SHM, V-gene SHM greater than about 6%; Class IV: ongoing CSR and / or SHM; and identifying the subject's immunodeficiency prognosis based on the subclassification of the B-cell immune repertoire as Class I: worst prognosis, Class II: poor prognosis, Class III: good prognosis, Class IV: best prognosis.
[0018] In some embodiments, methods are provided for predicting the clinical response of a subject having an autoimmune disease or disorder to therapy and / or predicting the prognosis of a subject having an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire, wherein each of a plurality of V gene primers and / or one or more C gene primers meets the following criteria: (1) contain two or more modified nucleotides within the primer, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) are about 15 to about 40 bases in length; (3) have a T of greater than 60°C and up to about 70°C; m (4) have low cross-reactivity to non-target sequences present in the sample; (5) have at least the first four nucleotides (in the 3' to 5' direction) that are non-complementary to any sequence in any other primer present in the same reaction; and (6) are non-complementary to any contiguous stretch of at least five nucleotides in any other generated target amplicon.
[0019] In some embodiments, methods are provided for predicting the clinical response to therapy of a subject having an autoimmune disease or disorder and / or predicting the prognosis of a subject having an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire, wherein each of the plurality of V gene primers and / or one or more C gene primers preferably comprises one or more cleavable groups located (i) near or at the terminus of the primer, or (ii) near or near the central nucleotide of the primer. In some embodiments of the provided methods, each of the plurality of V gene primers and / or one or more C gene primers comprises two or more modified nucleotides having cleavable groups selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine. In some embodiments of the provided methods, the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule, and the one or more C gene primers include at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
[0020] In some embodiments, methods are provided for predicting the clinical response to therapy of a subject with an autoimmune disease or disorder and / or predicting the prognosis of a subject with an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire, wherein the generated target BCR amplicon molecule includes complementarity determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence. In certain embodiments, at least one set of i) and ii) is selected from the primers in Table 3 and Tables 6-10, respectively. In other specific embodiments, at least one set of i) and ii) is selected from the primer sets in Table 11.
[0021] In some embodiments, a method is provided for predicting the clinical response to therapy of a subject with an autoimmune disease or disorder and / or predicting the prognosis of a subject with an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire, wherein at least one set of i) and ii) is i)(c) and ii)(a), wherein the plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule, and the one or more C gene primers include at least five primers annealing to at least a portion of the C gene portion of the template molecule. In certain embodiments, the generated target BCR amplicon molecule comprises a complementarity-determining region CDR3 of the target BCR gene sequence. In other specific embodiments, at least one set of i) and ii) is selected from the primers in Table 2 and Tables 6-10, respectively.
[0022] In certain embodiments of the methods provided herein for predicting the clinical response to therapy of a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, the preferred SHM frequency cutoff is 8%.
[0023] In certain embodiments of the methods provided herein for predicting the prognosis of a subject with leukemia based on characterizing the subject's B cell immune repertoire, the average class III SHM frequency is about 2%.
[0024] In some embodiments, provided methods for predicting the clinical response to a therapy of a subject having an autoimmune disease or disorder and / or predicting the prognosis of a subject having an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire further comprise, prior to step b), adding at least one adapter to at least one of the target immune receptor amplicon molecules, thereby generating a library of adapter-modified target immune receptor amplicon molecules. In certain embodiments, the at least one adapter is added by ligation.
[0025] In some embodiments, provided methods for predicting the clinical response to therapy of a subject with an autoimmune disease or disorder and / or predicting the prognosis of a subject with an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire include: obtaining initial sequence reads; aligning the initial sequence reads with reference reads; identifying productive reads; and correcting one or more indel errors to generate rescued productive sequence reads. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads.
[0026] In another aspect of the present invention, methods are provided for diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immunodeficiency disorder based on characterizing the subject's B cell immune repertoire. The provided methods include performing a multiplex amplification reaction to amplify target B cell immunoreceptor nucleic acid template molecules from a biological sample from the subject. In some embodiments, the provided methods include at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene; or (c) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of a C gene of at least one BCR coding sequence. Each set of primers i) and ii) is directed to a coding sequence of a target IgH BCR gene, and amplification results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target B cell immunoreceptor amplicon molecules comprising the target immune receptor repertoire. The provided method further includes sequencing the target immune receptor repertoire amplicons, identifying immune receptor clones from the sequencing, determining a level of somatic hypermutation (SHM) within variable gene portions among the immune receptor clones, wherein the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and determining the subclass of the B cell immunoreceptor clones in the sample, as well as the frequency of class switch recombination in the sample.The method includes identifying the subject as having a primary immunodeficiency disorder if the SHM frequency is below a frequency threshold in the switched isotype, the frequency of class switch recombination (CSR) in the sample is below a frequency threshold, and the immune repertoire is dominated by non-switched IgM / IgD-expressing B cells in the sample, thereby diagnosing the subject as having a chronic variable immunodeficiency disorder.
[0027] In some embodiments, methods are provided for diagnosing a subject having symptoms of an autoimmune disease or disorder as having a chronic variable immunodeficiency disorder, wherein each of the plurality of V gene primers and / or one or more C gene primers meets the following criteria: (1) contain two or more modified nucleotides within the primer, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) are about 15 to about 40 bases in length; (3) have a T of greater than 60°C and up to about 70°C; m (4) have low cross-reactivity to non-target sequences present in the sample; (5) have at least the first four nucleotides (in the 3' to 5' direction) that are non-complementary to any sequence in any other primer present in the same reaction; and (6) are non-complementary to any contiguous stretch of at least five nucleotides in any other generated target amplicon.
[0028] In some embodiments of the provided methods of diagnosing a subject with an autoimmune disease or symptoms of an autoimmune disease as having a chronic variable immune deficiency disorder, each of the plurality of V gene primers and / or one or more C gene primers preferably comprises one or more cleavable groups located (i) near or at the terminus of the primer, or (ii) near or near the central nucleotide of the primer. In some embodiments of the provided methods, each of the plurality of V gene primers and / or one or more C gene primers comprises two or more modified nucleotides having cleavable groups selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine. In some embodiments of the provided methods, the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule, and the one or more C gene primers include at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
[0029] In some embodiments of the provided methods of diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immunodeficiency disorder, the generated target BCR amplicon molecule comprises complementarity-determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence. In certain embodiments, at least one set of i) and ii) is selected from the primers in Table 3 and Tables 6-10, respectively. In other specific embodiments, at least one set of i) and ii) is selected from the primer sets in Table 11.
[0030] In some embodiments of the provided methods, at least one set of i) and ii) is i)(c) and ii)(a), wherein the plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule, and the one or more C gene primers include at least five primers that anneal to at least a portion of the C gene portion of the template molecule. In certain embodiments, the generated target BCR amplicon molecule comprises a complementarity-determining region CDR3 of the target BCR gene sequence. In other specific embodiments, at least one set of i) and ii) is selected from the primers in Table 2 and Tables 6-10, respectively.
[0031] In some embodiments, the provided methods of diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immune deficiency disorder further comprise, prior to step b), adding at least one adapter to at least one of the target immune receptor amplicon molecules, thereby generating a library of adapter-modified target immune receptor amplicon molecules. In certain embodiments, the at least one adapter is added by ligation.
[0032] In some embodiments, a method is provided for diagnosing a subject with symptoms of an autoimmune disease or disorder as having a chronic variable immune deficiency disorder, wherein sequencing comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads.
[0033] In another aspect of the present invention, methods are provided for treating a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire. The provided methods include performing a multiplex amplification reaction to amplify target B cell immunoreceptor nucleic acid template molecules from a biological sample from the subject. In some embodiments, the provided methods include at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene; or (c) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of a C gene of at least one BCR coding sequence. Each set of primers i) and ii) is directed to a coding sequence of a target IgH BCR gene, and amplification results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target B cell immunoreceptor amplicon molecules comprising the target immune receptor repertoire. The provided method further includes sequencing the target immune receptor repertoire amplicons, identifying immune receptor clones from the sequencing and determining a level of somatic hypermutation (SHM) within variable gene portions among the immune receptor clones, wherein the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and determining the subclass of the B cell immunoreceptor clones in the sample, as well as the frequency of non-switched IgM and / or IgD and switched IgG, IgA, and / or IgE-expressing cells in the sample.The method provides for treating the subject (i) with chemotherapy if the SHM frequency in the sample is below a frequency threshold in non-switched IgM / IgD-expressing B cells and the immune repertoire is dominated by a high frequency of switched isotype IgG, IgA, or IgE-expressing B cells in the sample, or (ii) with immunotherapy if the SHM frequency is above a frequency threshold and the immune repertoire is dominated by a high frequency of non-switched IgM / IgD-expressing B cells.
[0034] In some embodiments, a method is provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, wherein each of a plurality of V gene primers and / or one or more C gene primers meets the following criteria: (1) contain two or more modified nucleotides within the primer, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) are about 15 to about 40 bases in length; (3) have a T of greater than 60°C and up to about 70°C; m (4) have low cross-reactivity to non-target sequences present in the sample; (5) have at least the first four nucleotides (in the 3' to 5' direction) that are non-complementary to any sequence in any other primer present in the same reaction; and (6) are non-complementary to any contiguous stretch of at least five nucleotides in any other generated target amplicon.
[0035] In some embodiments of the provided methods for treating a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, each of the plurality of V gene primers and / or one or more C gene primers preferably comprises one or more cleavable groups located (i) near or at the terminus of the primer, or (ii) near or near the central nucleotide of the primer. In some embodiments of the provided methods, each of the plurality of V gene primers and / or one or more C gene primers comprises two or more modified nucleotides having cleavable groups selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine. In some embodiments of the provided methods, the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule, and the one or more C gene primers include at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
[0036] In some embodiments of the methods provided for treating a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, the generated target BCR amplicon molecule comprises complementarity determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence. In certain embodiments, at least one set of i) and ii) is selected from the primers in Table 3 and Tables 6-10, respectively. In other specific embodiments, at least one set of i) and ii) is selected from the primer sets in Table 11.
[0037] In some embodiments of the provided methods for treating a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, at least one set of i) and ii) is i)(c) and ii)(a), and includes at least five primers, where a plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule and one or more C gene primers anneal to at least a portion of the C gene portion of the template molecule. In certain embodiments, the generated target BCR amplicon molecule comprises a complementarity-determining region CDR3 of the target BCR gene sequence. In other specific embodiments, at least one set of i) and ii) is selected from the primers in Table 2 and Tables 6-10, respectively.
[0038] In certain embodiments of the methods provided for treating a subject with an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, the preferred SHM frequency cutoff is 8%.
[0039] In some embodiments, a method is provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, further comprising, prior to step b), adding at least one adapter to at least one of the target immune receptor amplicon molecules, thereby generating a library of adapter-modified target immune receptor amplicon molecules. In certain embodiments, the at least one adapter is added by ligation.
[0040] In some embodiments, a method is provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, wherein sequencing comprises obtaining initial sequence reads, aligning the initial sequence reads with reference reads, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads.
[0041] In some embodiments, methods are provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, wherein the subject has rheumatoid arthritis.
[0042] In some embodiments, methods are provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, wherein the immunotherapy comprises a checkpoint blockade or a B cell depleting agent.
[0043] In some embodiments, methods are provided for treating a subject having an autoimmune disease or disorder based on characterizing the subject's B cell immune repertoire, wherein the chemotherapy comprises methotrexate.
[0044] In another aspect of the present invention, methods are provided for treating a subject with an immune deficiency (e.g., leukemia) based on characterizing the subject's B cell immune repertoire. The provided methods include performing a multiplex amplification reaction to amplify target B cell immunoreceptor nucleic acid template molecules derived from a biological sample from the subject. In some embodiments, the provided methods include a multiplex amplification reaction including at least one set of: (a) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene; or (c) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of a C gene of at least one BCR coding sequence. Each set of primers i) and ii) is directed to a coding sequence of a target IgH BCR gene, and amplification results in amplicon molecules representing the target BCR repertoire in the sample, thereby generating target B cell immunoreceptor amplicon molecules comprising the target immune receptor repertoire. In some embodiments, the provided method further comprises sequencing the target immune receptor repertoire amplicon, identifying immune receptor clones from the sequencing and determining a level of somatic hypermutation (SHM) within variable gene portions among the immune receptor clones, wherein the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and further determining an ongoing SHM frequency in which the VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, and / or an ongoing class switch recombination (CSR) frequency in which at least one and / or combination of switched isotypes is identified in similar variable lineage B cell immunoreceptor clones in the sample.Methods are provided for classifying repertoire clones according to the following subclasses: Class I: no ongoing CSR or SHM, no V-gene SHM; Class II: no ongoing CSR or SHM, V-gene SHM greater than 0 and less than 6%; Class III: no ongoing CSR or SHM, V-gene SHM greater than about 6%; Class IV: ongoing CSR and / or SHM; and treating subjects based on the subclassification of the B-cell immune repertoire with Class I: stem cell therapy with the optional addition of any of chemotherapy, radiation, or DNA repair inducers; Class II: chemotherapy, radiation, or immunotherapy with the optional addition of DNA repair inducers; Class III: chemotherapy or immunotherapy; and Class IV: standard chemotherapy or immunotherapy.
[0045] In some embodiments, a method is provided for treating a subject with leukemia based on the subject's B-cell immune repertoire, wherein each of a plurality of V gene primers and / or one or more C gene primers meets the following criteria: (1) comprises two or more modified nucleotides within the primer, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) is about 15 to about 40 bases in length; (3) has a T of greater than 60°C and up to about 70°C; m (4) have low cross-reactivity to non-target sequences present in the sample; (5) have at least the first four nucleotides (in the 3' to 5' direction) that are non-complementary to any sequence in any other primer present in the same reaction; and (6) are non-complementary to any contiguous stretch of at least five nucleotides in any other generated target amplicon.
[0046] In some embodiments of the provided methods for treating a subject with leukemia based on the subject's B cell immune repertoire, each of the plurality of V gene primers and / or one or more C gene primers preferably comprises one or more cleavable groups located (i) near or at the terminus of the primer, or (ii) near or near the central nucleotide of the primer. In some embodiments of the provided methods, each of the plurality of V gene primers and / or one or more C gene primers comprises two or more modified nucleotides having cleavable groups selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine. In some embodiments of the provided methods, the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule, and the one or more C gene primers include at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
[0047] In some embodiments of the provided methods for treating a subject with leukemia based on the subject's B cell immune repertoire, the generated target BCR amplicon molecule comprises complementarity determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence. In certain embodiments, at least one set of i) and ii) is selected from the primers in Table 3 and Tables 6-10, respectively. In other specific embodiments, at least one set of i) and ii) is selected from the primer sets in Table 11.
[0048] In some embodiments of the provided methods for treating a subject with leukemia based on the subject's B cell immune repertoire, at least one set of i) and ii) is i)(c) and ii)(a), and includes at least five primers, where a plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule and one or more C gene primers anneal to at least a portion of the C gene portion of the template molecule. In certain embodiments, the generated target BCR amplicon molecule includes a complementarity-determining region CDR3 of the target BCR gene sequence. In other specific embodiments, at least one set of i) and ii) is selected from the primers in Table 2 and Tables 6-10, respectively.
[0049] In certain embodiments of the methods provided herein for treating a subject with leukemia based on the subject's B cell immune repertoire, the average class III SHM frequency is about 2%.
[0050] In some embodiments, a method for treating a subject with leukemia based on the subject's B-cell immune repertoire is provided, further comprising, prior to step b), adding at least one adapter to at least one of the target immune receptor amplicon molecules, thereby generating a library of adapter-modified target immune receptor amplicon molecules. In certain embodiments, the at least one adapter is added by ligation.
[0051] In some embodiments, a method is provided for treating a subject with leukemia based on the subject's B cell immune repertoire, wherein sequencing comprises obtaining initial sequence reads, aligning the initial sequence reads with reference reads, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50% of the sequencing reads.
[0052] In some embodiments of the methods provided for treating a subject with leukemia based on the subject's B cell immune repertoire, the subject has chronic lymphocytic leukemia.
[0053] In some embodiments of the methods provided for treating a subject with leukemia based on the subject's B cell immune repertoire, the immunotherapy comprises a targeted biologic agent or a B cell depleting agent.
[0054] In some embodiments, multiplexed next-generation sequencing workflows are used in combination with the provided methods for effective detection and analysis of immune repertoires in samples. The provided methods utilize workflows, compositions, systems, and kits for use in high-precision amplification and sequencing of immune cell receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) in monitoring and analyzing complex immune cell repertoire(s) in a subject. The target immune cell receptor genes have undergone rearrangement (or recombination) of VDJ or VJ gene segments, and the gene segments depend on the specific receptor gene (e.g., IgH, IgK, TCR beta, or TCR alpha). In certain embodiments, the present disclosure provides methods using workflows, compositions, and systems that employ nucleic acid amplification, such as polymerase chain reaction (PCR), to enrich expressed variable regions of immune receptor target nucleic acids for subsequent sequencing. In certain embodiments, the present disclosure provides methods using workflows, compositions, and systems that employ nucleic acid amplification, such as PCR, to enrich rearranged target immune cell receptor gene sequences from gDNA for subsequent sequencing. In certain embodiments, the present disclosure also provides methods for improving read assignment accuracy and reducing false positive rates using workflows and systems for effectively identifying and removing error(s) from amplification or sequencing. In particular, the provided methods described herein can improve accuracy and performance in sequencing applications using nucleotide sequences associated with genomic recombination and high variability. In some embodiments, the methods, compositions, systems, and kits provided herein are for use in amplifying and sequencing the complementarity-determining regions (CDRs) of expressed immune receptors in a sample. In some embodiments, the methods, compositions, systems, and kits provided herein are for use in amplifying and sequencing the CDRs of rearranged immune cell receptor gDNA in a sample.Thus, multiplexed immune cell receptor expression compositions and immune cell receptor gene-directed compositions for multiplexed library preparation, used in conjunction with next-generation sequencing technologies and workflow solutions (e.g., manual or automated), can be used in conjunction with the methods provided herein for effective detection and characterization of the immune repertoire in a sample.
[0055] The CDRs of TCRs or BCRs arise from genomic DNA that undergoes recombination of V(D)J gene segments and the addition and / or deletion of nucleotides at gene segment junctions. V(D)J gene segment recombination and subsequent hypermutation events lead to a wide diversity of expressed immune cell receptors. Due to the stochastic nature of V(D)J recombination, rearrangements of T cell receptor or B cell receptor genomic DNA frequently fail to produce functional receptors and instead generate what are called "non-productive" rearrangements. Non-productive rearrangements typically have out-of-frame variable and joining coding segments, leading to the presence of premature stop codons and the synthesis of unrelated peptides. Non-productive TCR or BCR gene rearrangements are generally rare in cDNA-based repertoire sequencing due to several biological or physiological reasons, such as: 1) nonsense-mediated decay, which destroys mRNAs containing premature stop codons; 2) B cell and T cell selection, which means that only B and T cells with functional receptors survive; and 3) allelic exclusion, which means that only a single rearranged receptor allele is expressed in any given B or T cell.
[0056] Thus, in some embodiments, the methods and compositions provided herein are used to amplify recombinantly expressed variable regions of immune cell receptor mRNA, such as BCR and / or TCR mRNA. In some embodiments, RNA extracted from a biological sample is converted to cDNA. Multiplex amplification is used to enrich for portions of BCR or TCR cDNA that contain at least a portion of the receptor's variable region. In some embodiments, the amplified cDNA comprises one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of the target receptor. In some embodiments, the amplified cDNA comprises one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of an immunoglobulin heavy chain (IgH).
[0057] BCR and TCR sequences may also appear as non-productive rearrangements due to errors introduced during the amplification reaction or sequencing process. For example, insertion or deletion (indel) errors during target amplification or sequencing reactions can cause frameshifts in the reading frame of the resulting coding sequence. Such changes can result in target sequence reads of productive rearrangements being interpreted as non-productive rearrangements and discarded from the identified clonotype group. Thus, in some embodiments, the methods and systems provided herein include processes for identifying and / or removing errors from the determined immune receptor sequences derived from PCR or sequencing.
[0058] In some embodiments, the provided methods and compositions are used to amplify rearranged variable regions of immune cell receptor gDNA, such as rearranged BCR and / or TCR gene DNA. Multiplex amplification is used to enrich for portions of rearranged BCR or TCR gDNA that contain at least a portion of the receptor's variable region. In some embodiments, the amplified gDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of the target receptor. In some embodiments, the amplified gDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of IgH. In some embodiments, the amplified gDNA primarily contains CDR3 of the target receptor, for example, CDR3 of IgH.
[0059] As used herein, "immune cell receptor" and "immunoreceptor" are used interchangeably.
[0060] As used herein, the terms "complementarity-determining region" and "CDR" refer to the region of a T cell receptor or antibody (immunoglobulin) in which the molecule complements the three-dimensional structure of an antigen, thereby determining the specificity of the molecule and its contact with a particular antigen. In the variable regions of T cell receptors and antibodies, CDRs are interspersed with more conserved regions called framework regions (FRs). Each variable region of a T cell receptor and antibody contains three CDRs called CDR1, CDR2, and CDR3, and four framework subregions called FR1, FR2, FR3, and FR4.
[0061] As used herein, the term "framework" or "framework region" or "FR" refers to the residues of the variable region other than the CDR residues as defined herein. There are four distinct framework subregions that make up the framework: FR1, FR2, FR3, and FR4.
[0062] Specific designations in the art regarding the exact locations of CDRs and FRs within receptor molecules (TCRs or immunoglobulins) vary depending on the definitions used. Unless otherwise specified, the IMGT designations are used herein to describe CDR and FR regions (see Brochet et al. (2008) Nucleic Acids Res. 36:W503-508, specifically incorporated herein by reference). As an example of a CDR / FR amino acid designation, the residues constituting the FRs and CDRs of T-cell receptor beta are characterized by IMGT as residues 1-26 (FR1), 27-38 (CDR1), 39-55 (FR2), 56-65 (CDR2), 66-104 (FR3), 105-117 (CDR3), and 118-128 (FR4).
[0063] Other well-known standard designations for describing regions include those found in Kabat et al., (1991) Sequences of Proteins of Immunological Interest, 5th Ed. Public Health Service, National Institutes of Health, Bethesda, Md., and Chothia and Lesk (1987) J. Mol. Biol. 196:901-917, which are specifically incorporated herein by reference. As an example of CDR designation, the residues that make up the six immunoglobulin CDRs are characterized by Kabat as residues 24-34 (CDRL1), 50-56 (CDRL2), and 89-97 (CDRL3) in the light chain variable region and 31-35 (CDRH1), 50-65 (CDRH2), and 95-102 (CDRH3) in the heavy chain variable region, and by Chothia as residues 26-32 (CDRL1), 50-52 (CDRL2), and 91-96 (CDRL3) in the light chain variable region and 26-32 (CDRH1), 53-55 (CDRH2), and 96-101 (CDRH3) in the heavy chain variable region.
[0064] As used herein, the term "T cell receptor" or "T cell antigen receptor" or "TCR" refers to the antigen / MHC-binding heterodimeric protein product of the TCR gene complex in a vertebrate, e.g., a mammal, including the human TCR alpha, beta, gamma, and delta chains. For example, the complete sequence of the human TCR beta locus has been sequenced, see e.g., Rowen et al. (1996) Science 272:1755-1762; the human TCR alpha locus has been sequenced and resequenced, see e.g., Mackelprang et al. (2006) Hum Genet. 119:255-266; and for a general analysis of the T cell receptor V gene segment family, see e.g., Arden (1995) Immunogenetics 42:455-500, each of which is specifically incorporated herein by reference for the sequence information provided and referenced in the publication.
[0065] As used herein, the term "antibody" or "immunoglobulin" or "B cell receptor" or "BCR" refers to four polypeptide chains, two heavy (H) chains and two light (L) chains (lambda or kappa), interconnected by disulfide bonds. An antibody has a known specific antigen to which it binds. Each heavy chain of an antibody consists of a heavy chain variable region (abbreviated herein as HCVR, HV, or VH) and a heavy chain constant region. The heavy chain constant region consists of three domains: CH1, CH2, and CH3. Each light chain consists of a light chain variable region (abbreviated herein as LCVR or VL or KV or LV to indicate kappa or lambda light chains) and a light chain constant region. The light chain constant region consists of one domain, CL. The heavy chain determines the class or isotype to which the immunoglobulin belongs. For example, in mammals, the five major immunoglobulin isotypes are IgA, IgD, IgG, IgE, and IgM, which are classified according to the alpha, delta, epsilon, gamma, or mu heavy chains they contain, respectively.
[0066] As described, diversity in TCR and BCR chain CDRs is created by recombination of germline variable (V), diversity (D), and joining (J) gene segments, as well as by independent additions and deletions at each gene segment junction during each TCR and BCR gene rearrangement. In rearranged nucleic acids encoding BCR heavy chains, CDR1 and CDR2 are found in the V gene segment, and CDR3 comprises the V gene segment and some of the D and J gene segments. In rearranged nucleic acids encoding BCR light chains, CDR1 and CDR2 are found in the V gene segment, and CDR3 comprises the V gene segment and some of the J gene segments. In rearranged nucleic acids encoding TCR beta and TCR delta, for example, CDR1 and CDR2 are found in the V gene segment, and CDR3 comprises some of the V gene segment and the D and J gene segments. In the rearranged nucleic acids encoding TCR alpha and TCR gamma, CDR1 and CDR2 are found in the V gene segments, and CDR3 comprises the V gene segment and some of the J gene segments.
[0067] In some embodiments, a multiplex amplification reaction is used to amplify cDNA derived from mRNA expressed from rearranged BCR and / or TCR genomic DNA. In some embodiments, a multiplex amplification reaction is used to amplify at least a portion of BCR and / or TCR CDRs from cDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify at least two CDRs of BCR and / or TCR from cDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify at least three CDRs of BCR and / or TCR from cDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of BCR and / or TCR CDRs expressed in the sample. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least three CDRs is used to identify and characterize novel BCR and / or TCR alleles.
[0068] In some embodiments, a multiplex amplification reaction is used to amplify BCR and / or TCR genomic DNA that has undergone V(D)J rearrangement. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecule(s) comprising at least a portion of a BCR and / or TCR CDR from gDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecule(s) comprising at least two CDRs of a BCR and / or TCR from gDNA derived from a biological sample. In some embodiments, a multiplex amplification reaction is used to amplify nucleic acid molecules comprising at least three CDRs of a BCR and / or TCR from gDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of the rearranged BCR and / or TCR CDRs in the sample. In some embodiments, determining the nucleotide sequences of such amplicons comprising at least the CDRs is used to identify and characterize novel BCR and / or TCR alleles.
[0069] In some embodiments of multiplex amplification reactions, each primer set used targets the same BCR or TCR region, but different primers within the set target different V(D)J gene rearrangements of the gene. For example, primer sets for amplification of expressed IgH or rearranged IgH gDNA are all designed to target the same region(s) of IgH mRNA or IgH gDNA, respectively, but individual primers within the set result in amplification of different IgH VDJ gene combinations. In some embodiments, at least one primer or primer set is directed to a relatively conserved region of an immune receptor gene (e.g., a portion of a C gene), while other primer sets contain various primers directed to more variable regions of the same gene (e.g., a portion of a V gene). In other embodiments, at least one primer set contains various primers directed to at least a portion of a J gene segment of an immune receptor gene, and other primer sets contain various primers directed to at least a portion of a V gene segment of the same gene.
[0070] In some embodiments, a multiplex amplification reaction is used to amplify cDNA derived from mRNA expressed from rearranged BCR genomic DNA, including rearranged IgH, IgK, and IgL genomic DNA. In some embodiments, at least a portion of a BCR CDR, e.g., CDR3, is amplified from cDNA in a multiplex amplification reaction. In some embodiments, at least two CDR portions of a BCR are amplified from cDNA in a multiplex amplification reaction. In certain embodiments, a multiplex amplification reaction is used to amplify at least the CDR1, CDR2, and CDR3 regions of a BCR cDNA. In some embodiments, the resulting amplicons are used to determine the expressed BCR CDR nucleotide sequences. In some embodiments, the resulting amplicons are used to determine the expressed BCR CDR nucleotide sequences and the Ig isotype of the sequences. In some embodiments, the resulting amplicons are used to determine the expressed IgH CDR nucleotide sequences and the Ig isotype and Ig subisotype.
[0071] In some embodiments, a multiplex amplification reaction is used to amplify rearranged BCR genomic DNA, including rearranged IgH, IgK, and IgL genomic DNA. In some embodiments, at least a portion of the BCR CDR, for example, CDR3, is amplified from gDNA in a multiplex amplification reaction. In some embodiments, at least two CDR portions of the BCR are amplified from gDNA in a multiplex amplification reaction. In certain embodiments, a multiplex amplification reaction is used to amplify at least the CDR1, CDR2, and CDR3 regions of the BCR gDNA. In some embodiments, the resulting amplicons are used to determine the rearranged BCR CDR nucleotide sequence. In some embodiments, the resulting amplicons are used to determine the rearranged BCR CDR nucleotide sequence and the Ig isotype of the sequence.
[0072] In some embodiments, a multiplex amplification reaction is performed using primer sets designed to generate amplicons comprising the expressed CDR1, CDR2, and / or CDR3 regions of the target immune receptor mRNA. In some embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each directed to at least a portion of the framework region FR1 of a V gene, and (ii) at least one primer directed to a portion of at least one C gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each directed to at least a portion of the framework region FR2 of a V gene, and (ii) at least one primer directed to a portion of at least one C gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each directed to at least a portion of the framework region FR3 of a V gene, and (ii) at least one primer directed to a portion of at least one C gene of the target immune receptor. In some embodiments, a multiplex amplification reaction is performed using a primer set designed to generate an amplicon containing one or more expressed IgH isotypes of target mRNA, and such reaction is performed using (i) one of the FR1, FR2, or FR3 primer sets described above, and (ii) a set of primers, each primer directed to a portion of at least one C gene of IgA, IgD, IgE, IgG, and / or IgM. In some embodiments, the C gene-directed primer(s) are directed to a C gene coding sequence within about 200 nucleotides of the 5' end of the C gene(s). In some embodiments, the C gene-directed primer(s) are directed to a C gene coding sequence within about 150 nucleotides of the 5' end of the C gene(s). In some embodiments, the C gene-directed primer(s) are directed to a C gene coding sequence within about 100 nucleotides of the 5' end of the C gene(s).In some embodiments, the C gene-directed primer(s) are directed to the C gene coding sequence within about 50 nucleotides, about 50 to about 150 nucleotides, about 75 to about 175 nucleotides, or about 100 to about 200 nucleotides of the 5' end of the C gene(s). In some embodiments, the C gene-directed primer(s) are directed to C gene coding sequences that allow not only isotype differentiation but also subisotype determination. For example, in some embodiments, the C gene-directed primer(s) generate a sufficient portion of the constant region within the amplicon so that subisotypes can be determined based on the determined sequence data. In some embodiments, the C gene-directed primer(s) include primers to IgG and / or IgA C gene coding sequences that allow identification of IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2 subisotypes.
[0073] In some embodiments, the multiplex amplification reaction amplifies BCR cDNA using (i) a set of primers that each anneal to at least a portion of the V gene FR1 region and (ii) at least one primer that anneals to a portion of the constant (C) gene, whereby the resulting amplicon contains the CDR1, CDR2, and CDR3-encoding portions of BCR mRNA. In certain embodiments, the FR1-directed primer set is combined with at least two C gene-directed primer sets to generate an amplicon containing at least the CDR1, CDR2, and CDR3-encoding portions of BCR mRNA. In some embodiments, the IgH FR1-directed primer set is combined with at least two C gene primer sets directed to the coding portions of two different IgH isotypes to generate an amplicon containing at least the CDR1, CDR2, and CDR3-encoding portions of IgH mRNA. In some embodiments, the IgH FR1-directed primer set is combined with at least three, at least four, or at least five primers directed to the coding portions of different IgH isotypes. For example, exemplary primers specific for the IgH V gene FR1 region are shown in Table 3, and exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0074] In some embodiments, the multiplex amplification reaction amplifies BCR cDNA using (i) a set of primers that each anneal to at least a portion of the V gene FR2 region and (ii) at least one primer that anneals to a portion of the C gene, thereby resulting in an amplicon that contains the CDR2- and CDR3-encoding portions of BCR mRNA. In certain embodiments, such a FR2-directed primer set is combined with at least two C gene-directed primers to generate an amplicon containing the CDR2- and CDR3-encoding portions of BCR mRNA. In some embodiments, an IgH FR2-directed primer set is combined with at least two C gene primer sets directed to the coding portions of two different IgH isotypes to generate an amplicon containing the CDR2- and CDR3-encoding portions of IgH mRNA. In some embodiments, an IgH FR2-directed primer set is combined with at least three, at least four, or at least five C gene primers directed to the coding portions of different IgH isotypes. Exemplary FR2-directed primers include the BIOMED-2 primers developed and standardized by a consortium of European academic laboratories and research hospitals (van Dongen et al. (2003) Leukemia 17:2257-2327) and are shown in Table 4. Exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0075] In some embodiments, the multiplex amplification reaction amplifies BCR cDNA using (i) a set of primers that each anneal to at least a portion of the V gene FR3 region and (ii) at least one primer that anneals to a portion of the C gene, such that the resulting amplicon primarily comprises the CDR3-encoding portion of BCR mRNA. In certain embodiments, such a FR3-directed primer set is combined with at least two C gene-directed primers to generate an amplicon having the CDR3-encoding portion of BCR mRNA. In some embodiments, an IgH FR3-directed primer set is combined with at least two C gene primer sets directed to the coding portions of two different IgH isotypes to generate an amplicon having the CDR3-encoding portion of IgH mRNA. In some embodiments, an IgH FR3-directed primer set is combined with at least three, at least four, or at least five C gene primers directed to the coding portions of different IgH isotypes. For example, exemplary primers specific for the IgH V gene FR3 region are shown in Table 2, and exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0076] In some embodiments, a multiplex amplification reaction is performed using primer sets designed to generate amplicons comprising the CDR1, CDR2, and / or CDR3 regions of a target immune receptor mRNA or rearranged gDNA. In some embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each of which is directed to at least a portion of the framework region FR1 of a V gene, and (ii) a set of primers, each of which is directed to at least a portion of the J gene of a target immune receptor. In other embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each of which is directed to at least a portion of the framework region FR2 of a V gene, and (ii) a set of primers, each of which is directed to at least a portion of the J gene of a target immune receptor. In other embodiments, the multiplex amplification reaction is performed using (i) a set of primers, each of which is directed to at least a portion of the framework region FR3 of a V gene, and (ii) a set of primers, each of which is directed to at least a portion of the J gene of a target immune receptor.
[0077] In some embodiments, the multiplex amplification reaction amplifies BCR nucleic acids using (i) a set of primers that each anneal to at least a portion of the V gene FR1 region and (ii) a set of primers that anneal to a portion of the J gene, such that the resulting amplicon contains the CDR1, CDR2, and CDR3-encoding portions of the BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the IgH V gene FR1 region are shown in Table 3, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0078] In some embodiments, the multiplex amplification reaction amplifies BCR cDNA using (i) a set of primers that each anneal to at least a portion of the V gene FR2 region and (ii) a set of primers that anneal to a portion of the J gene, such that the resulting amplicon contains the CDR2- and CDR3-encoding portions of BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the IgH V gene FR2 region are shown in Table 4, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0079] In some embodiments, the multiplex amplification reaction amplifies BCR nucleic acids using (i) a set of primers that each anneal to at least a portion of the V gene FR3 region and (ii) a set of primers that anneal to a portion of the J gene, such that the resulting amplicon primarily comprises the CDR3-encoding portion of the BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the IgH V gene FR3 region are shown in Table 2, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0080] In some embodiments, compositions for multiplex amplification of at least a portion of expressed BCR variable regions are provided. In some embodiments, the compositions comprise a plurality of primer pair reagent sets directed to a portion of a V gene framework region and a portion of a constant (C) gene of a rearranged target immune receptor gene selected from the group consisting of immunoglobulin heavy chain (IgH), immunoglobulin light chain lambda (IgL), and immunoglobulin light chain kappa (IgK). In some embodiments, the compositions comprise a plurality of primer pair reagent sets directed to a portion of a V gene framework region and a portion of a J gene of a rearranged target immune receptor gene selected from the group consisting of IgH, IgL, and IgK.
[0081] In some embodiments, the composition comprises (i) a plurality of primer pair reagent sets for a portion of the IgH V gene framework regions and a portion of the IgH C gene of a rearranged IgH gene, and (ii) a plurality of primer pair reagent sets for a portion of the TCR beta V gene framework regions and a portion of the TCR beta C gene of a rearranged TCR beta gene. In some embodiments, the composition comprises (i) a plurality of primer pair reagent sets for a portion of the IgH V gene framework regions and a portion of the IgH J gene of a rearranged IgH gene, and (ii) a plurality of primer pair reagent sets for a portion of the TCR beta V gene framework regions and a portion of the TCR beta J gene of a rearranged TCR beta gene.
[0082] PCR amplification is carried out using at least two primers.In the method provided herein, a set of primers sufficient to amplify all or a specified portion of the variable sequence at the locus of interest is used, and the locus can include any or all of the TCR and immunoglobulin loci described above.In some embodiments, various parameters or criteria as outlined herein can be used to select a set of target-specific primers for multiplex amplification.
[0083] In some embodiments, the primer set used in the multiplex reaction is designed to amplify at least 50% of the known expressed gDNAs or gDNA rearrangements at the locus of interest. In certain embodiments, the primer set used in the multiplex reaction is designed to amplify at least 75%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or more of the known expressed gDNAs or gDNA rearrangements at the locus of interest. In another example, the use of 27 forward primers from Table 3, each directed to a portion of the FR1 region from a different IgH V gene, in combination with at least one reverse primer from Tables 6-10, each directed to a portion of a different IgH C gene, amplifies all of the currently known expressed IgH rearrangements for a given isotype. In another example, the use of 68 forward primers from Table 2, each directed to a portion of the FR3 region from a different IgH V gene, in combination with at least one reverse primer from Tables 6-10, each directed to a portion of a different IgH C gene, amplifies all of the currently known expressed IgH rearrangements for a given isotype. As another example, the 68 forward primers in Table 2, each directed to a portion of the FR3 region from a different IgH V gene, combined with the four reverse primers in Table 5, each directed to a portion of a different IgH J gene, will amplify all of the currently known expressed gDNA IgHs or gDNA IgH rearrangements. As another example, the 27 forward primers in Table 3, each directed to a portion of the FR1 region from a different IgH V gene, combined with the four reverse primers in Table 5, each directed to a portion of a different IgH J gene, will amplify all of the currently known expressed gDNA IgHs or gDNA IgH rearrangements.
[0084] For example, such multiplex amplification reactions include at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR1 regions. In such embodiments, the multiple reverse primers directed to the BCR V gene FR1 regions are combined with at least one forward primer directed to a sequence corresponding to at least a portion of a constant gene of the same BCR gene. In some embodiments, multiple reverse primers directed to the BCR V gene FR1 region are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 forward primers, each directed to a sequence corresponding to at least a portion of at least one of the constant genes of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR1-directed primer can be the forward primer, and the BCR C gene-directed primer(s) can be the reverse primer(s). Thus, in some embodiments, the multiplex amplification reaction comprises at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 forward primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR1 regions. In such embodiments, the multiple forward primers directed to the BCR V gene FR1 regions are combined with at least one reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene.In some embodiments, multiple forward primers directed to the FR1 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12, reverse primers each directed to a sequence corresponding to at least a portion of at least one C gene of the same BCR gene. In some embodiments, such FR1 and C gene amplification primer sets can be directed to IgH gene sequences. In some preferred embodiments, about 22 to about 35 reverse primers directed to the FR1 regions of different IgH V genes are combined with about 2 to about 8 forward primers directed to portions of the IgH C gene. In another preferred embodiment, about 22 to about 35 reverse primers directed to the FR1 regions of different IgH V genes are combined with about 5 to about 15 forward primers directed to portions of the IgH C gene. In another preferred embodiment, about 48 to about 60 reverse primers for the FR1 regions of different IgH V genes are combined with about 5 to about 15 forward primers for a portion of the IgH C gene. In some preferred embodiments, about 22 to about 35 forward primers for the FR1 regions of different IgH V genes are combined with about 2 to about 8 reverse primers for a portion of the IgH C gene. In another preferred embodiment, about 22 to about 35 forward primers for the FR1 regions of different IgH V genes are combined with about 5 to about 15 reverse primers for a portion of the IgH C gene. In yet another preferred embodiment, about 48 to about 60 forward primers for the FR1 regions of different IgH V genes are combined with about 5 to about 15 reverse primers for a portion of the IgH C gene.In some preferred embodiments, the forward primer for the IgH V gene FR1 region is selected from those listed in Table 3, and the reverse primer for the IgH C gene is selected from those listed in Tables 6-10. In other embodiments, the FR1 and C gene amplification primer set can be for Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, or TCR beta gene sequences.
[0085] In some embodiments, the multiplex amplification reaction includes at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR2 regions. In such embodiments, the multiple reverse primers directed to the BCR V gene FR2 regions are combined with at least one forward primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the multiple reverse primers directed to the BCR V gene FR2 regions are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 forward primers, each directed to a sequence corresponding to at least a portion of at least one of the C genes of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR2-directed primer can be a forward primer, and the BCR C gene-directed primer(s) can be a reverse primer(s). Thus, in some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 different forward primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR2 regions. In such embodiments, multiple forward primers directed to the BCR V gene FR2 regions are combined with at least one reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene.In some embodiments, multiple forward primers directed to the FR2 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12, reverse primers each directed to a sequence corresponding to at least a portion of at least one C gene of the same BCR gene. In some embodiments, such FR2 and C gene amplification primer sets can be directed to IgH gene sequences. In some embodiments, about 5 to about 15 reverse primers directed to the FR2 regions of different IgH V genes are combined with about 2 to about 8 forward primers directed to portions of the IgH C gene. In some embodiments, about 5 to about 15 reverse primers directed to the FR2 regions of different IgH V genes are combined with about 5 to about 15 forward primers directed to portions of the IgH C gene. In some embodiments, about 5 to about 15 forward primers for different IgH V gene FR2 regions are combined with about 2 to about 8 reverse primers for portions of the IgH C gene. In some embodiments, about 5 to about 15 forward primers for different IgH V gene FR2 regions are combined with about 5 to about 15 reverse primers for portions of the IgH C gene. In some preferred embodiments, the forward primers for the IgH V gene FR2 regions are selected from those listed in Table 4, and the reverse primers for the IgH C gene are selected from those listed in Tables 6 to 10. In other embodiments, the FR2 and C gene amplification primer sets can be directed to Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, or TCR beta gene sequences.
[0086] In some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR3 regions. In such embodiments, the multiple reverse primers directed to the BCR V gene FR3 regions are combined with at least one forward primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the multiple reverse primers directed to the BCR V gene FR3 regions are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 forward primers, each directed to a sequence corresponding to at least a portion of at least one of the C genes of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR3-directed primer can be a forward primer, and the BCR C gene-directed primer(s) can be a reverse primer(s). Thus, in some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers, each forward primer directed to a sequence corresponding to at least a portion of one or more BCR V gene FR3 regions. In such embodiments, multiple forward primers directed to the BCR V gene FR3 regions are combined with at least one reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene.In some embodiments, multiple forward primers directed to the FR3 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12, reverse primers each directed to a sequence corresponding to at least a portion of at least one C gene of the same BCR gene. In some embodiments, such FR3 and C gene amplification primer sets can be directed to IgH gene sequences. In some preferred embodiments, about 62 to about 75 reverse primers directed to the FR3 regions of different IgH V genes are combined with about 2 to about 8 forward primers directed to portions of the IgH C gene. In another preferred embodiment, about 62 to about 75 reverse primers directed to the FR3 regions of different IgH V genes are combined with about 5 to about 15 forward primers directed to portions of the IgH C gene. In some preferred embodiments, about 62 to about 75 forward primers for different IgH V gene FR3 regions are combined with about 2 to about 8 reverse primers for portions of the IgH C gene. In other preferred embodiments, about 62 to about 75 forward primers for different IgH V gene FR3 regions are combined with about 5 to about 15 reverse primers for portions of the IgH C gene. In some preferred embodiments, the forward primers for the IgH V gene FR3 regions are selected from those listed in Table 2, and the reverse primers for the IgH C gene are selected from those listed in Tables 6 to 10. In other embodiments, the FR3 and C gene amplification primer sets can be directed to Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, and TCR beta gene sequences.
[0087] In some embodiments, such multiplex amplification reactions include at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR1 regions. In such embodiments, multiple reverse primers directed to BCR V gene FR1 regions are combined with at least 2, 3, 4, 5, 6, 8, or about 3-6 forward primers directed to a sequence corresponding to at least a portion of a J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR1-directed primers can be forward primers, and the BCR J gene-directed primers can be reverse primers. Thus, in some embodiments, the multiplex amplification reaction includes at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 forward primers, each directed to a sequence corresponding to at least a portion of the FR1 region of one or more BCR V genes. In such embodiments, the multiple forward primers directed to the FR1 region of the BCR V genes are combined with at least 2, 3, 4, 5, 6, 8, or about 3-6 reverse primers directed to a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments, such FR1 and J gene amplification primer sets may be directed to IgH gene sequences. In some preferred embodiments, about 22 to about 35 reverse primers directed to different IgH V gene FR1 regions are combined with about 3 to about 6 forward primers directed to different IgH J genes.In some preferred embodiments, about 22 to about 35 forward primers for different IgH V gene FR1 regions are combined with about 3 to about 6 reverse primers for different IgH J genes. In some preferred embodiments, the forward primers for the IgH V gene FR1 regions are selected from those listed in Table 3, and the reverse primers for the IgH J genes are selected from those listed in Table 5. In other embodiments, the FR1 and J gene amplification primer sets can be for Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, or TCR beta gene sequences.
[0088] In some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR2 regions. In such embodiments, the reverse primers directed to the BCR V gene FR2 regions are combined with at least 2, 3, 4, 5, 6, 8, or about 3-6 forward primers directed to a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR2-directed primer can be a forward primer, and the BCR J gene-directed primer can be a reverse primer. Thus, in some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 forward primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR2 regions. In such embodiments, multiple forward primers directed to the BCR V gene FR2 region are combined with at least 2, 3, 4, 5, 6, 8, or about 3 to 6 reverse primers directed to sequences corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments, such FR2 and J gene amplification primer sets can be directed to IgH gene sequences. In some preferred embodiments, about 5 to about 15 reverse primers directed to different IgH V gene FR2 regions are combined with about 3 to about 6 forward primers directed to different IgH J genes. In some preferred embodiments, about 5 to about 15 forward primers directed to different IgH V gene FR2 regions are combined with about 3 to about 6 reverse primers directed to different IgH J genes. In some preferred embodiments, the forward primers directed to the IgH V gene FR2 region are selected from those listed in Table 4, and the reverse primers directed to the IgH J gene are selected from those listed in Table 5.In other embodiments, the FR2 and J gene amplification primer set can be directed to an Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, or TCR beta gene sequence.
[0089] In some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 reverse primers, each directed to a sequence corresponding to at least a portion of one or more BCR V gene FR3 regions. In such embodiments, the multiple reverse primers directed to the BCR V gene FR3 regions are combined with at least 2, 3, 4, 5, 6, 8, or about 3-6 forward primers directed to a sequence corresponding to at least a portion of a J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the BCR V gene FR3-directed primers can be forward primers, and the BCR J gene-directed primers can be reverse primers. Thus, in some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85, or 90 forward primers, each directed to a sequence corresponding to at least a portion of the FR3 region of one or more BCR V genes. In such embodiments, multiple forward primers directed to the FR3 region of the BCR V genes are combined with at least 2, 3, 4, 5, 6, 8, or about 3 to 6 reverse primers directed to sequences corresponding to at least a portion of the J genes of the same BCR gene. In some embodiments, such FR3 and J gene amplification primer sets can be directed to IgH gene sequences. In some preferred embodiments, about 62 to about 75 reverse primers directed to different IgH V gene FR3 regions are combined with about 3 to about 6 forward primers directed to different IgH J genes. In some preferred embodiments, about 62 to about 75 forward primers for different IgH V gene FR3 regions are combined with about 3 to about 6 reverse primers for different IgH J genes. In some preferred embodiments, the forward primers for the IgH V gene FR3 regions are selected from those listed in Table 2, and the reverse primers for the IgH J genes are selected from those listed in Table 5.In other embodiments, the FR3 and J gene amplification primer sets can be directed to Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, and TCR beta gene sequences.
[0090] In some embodiments, the concentration of the forward primer is approximately equal to the concentration of the reverse primer in the multiplex amplification reaction. In other embodiments, the concentration of the forward primer is about twice the concentration of the reverse primer in the multiplex amplification reaction. In other embodiments, the concentration of the forward primer is about half the concentration of the reverse primer in the multiplex amplification reaction. In some embodiments, the concentration of each primer targeting a V gene FR region is about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting a V gene FR region is about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting a J gene is about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting a J gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting a J gene is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each of the primers targeting the J genes is about 200 nM, about 400 nM, about 600 nM, or about 800 nM, hi some embodiments, the concentration of each of the primers targeting the J genes is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM.In some embodiments, the concentration of each primer targeting the J gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting the J gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting the C gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is about 50 nM to about 400 nM or about 100 nM to about 500 nM. In some embodiments, the concentration of each primer targeting the C gene is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM. In some embodiments, the concentration of each primer targeting the C gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting the C gene is about 50 nM to about 800 nM. In some embodiments, the concentration of each forward and reverse primer in the multiplex reaction is about 50 nM, about 100 nM, about 200 nM, or about 400 nM. In some embodiments, the concentration of each forward and reverse primer in the multiplex reaction is about 5 nM to about 2000 nM. In some embodiments, the concentration of each forward and reverse primer in the multiplex reaction is about 50 nM to about 800 nM. In some embodiments, the concentration of each forward and reverse primer in the multiplex reaction is about 50 nM to about 400 nM or about 100 nM to about 500 nM, or about 600 nM, 800 nM, 1000 nM, 1250 nM, 1500 nM, 1750 nM, or 2000 nM.In some embodiments, the concentration of each forward and reverse primer in the multiplex reaction is about 5 nM, about 10 nM, about 150 nM, or between 50 nM and about 800 nM.
[0091] In some embodiments, V gene FR and C gene targeting primers are combined as an amplification primer pair to amplify the target immune receptor cDNA sequence and generate a target amplicon. Generally, the length of the target amplicon depends on the V gene primer set (e.g., FR1, FR2, or FR3-directed primer) paired with the C gene primer(s). Thus, in some embodiments, the target amplicon can range from about 100 nucleotides (or bases or base pairs) to about 600 nucleotides (or bases or base pairs) in length. In some embodiments, the target amplicon can range from about 80 nucleotides to about 600 nucleotides in length. In some embodiments, the target amplicon is about 200 to about 600 or about 300 to about 600 nucleotides in length. In some embodiments, the target amplicon is about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides in length. In some embodiments, the target amplicon is about 250 to about 275, about 250 to about 350, about 300 to about 350, about 310 to about 330, about 325 to about 375, about 300 to about 400, about 350 to about 400, about 350 to about 425, about 350 to about 450, about 380 to about 410, about 375 to about 425, about 400 to about 500, about 425 to about 500, about 450 to about 550, about 500 to about 600, about 400 to about 500, or about 400 to about 600 nucleotides in length. In some embodiments, the target amplicon is about 80, about 100, about 120, about 140, about 200, about 250, about 275, about 300, about 320, about 350, about 375, about 400, about 425, about 450, about 500, about 550, or about 600 nucleotides in length. In some embodiments, the IgH amplicon is about 100, about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides in length. In some embodiments, the IgH amplicon is about 320, about 300 to about 350, or about 310 to about 330 nucleotides in length.In some embodiments, the IgH amplicon is about 400, about 375 to about 425, or about 390 to about 410 nucleotides in length.
[0092] In some embodiments, V gene FR and J gene targeting primers are combined as an amplification primer pair to amplify a target immune receptor cDNA sequence or rearranged gDNA sequence to generate a target amplicon. Generally, the length of the target amplicon depends on the V gene primer set (e.g., FR1, FR2, or FR3-directed primer) paired with the J gene primer. Thus, in some embodiments, the target amplicon can range from about 50 nucleotides to about 350 nucleotides in length. In some embodiments, the target amplicon is about 50 to about 200, about 70 to about 170, about 200 to about 350, about 250 to about 320, about 270 to about 300, about 225 to about 300, about 250 to about 275, about 200 to about 235, about 200 to about 250, or about 175 to about 275 nucleotides in length. In some embodiments, IgH amplicons are about 80, about 60 to about 100, or about 70 to about 90 nucleotides in length. In some embodiments, IgH amplicons, such as those generated using V gene FR3-directed and J gene-directed primer pairs, are about 50 to about 200 nucleotides in length, preferably about 60 to about 160, about 65 to about 120, about 90 to about 90, about 70 to about 90, or about 80 nucleotides in length. In some embodiments, generating amplicons of such short length enables the provided methods and compositions to effectively detect and analyze immune repertoires from highly degraded gDNA template material, such as that derived from FFPE samples or cell-free DNA (cfDNA).
[0093] In some embodiments, amplification primers may contain barcode sequences, for example, to distinguish or separate multiple amplified target sequences in a sample. In some embodiments, amplification primers may contain two or more barcode sequences, for example, to distinguish or separate multiple amplified target sequences in a sample. In some embodiments, amplification primers may contain tagging sequences that can assist in subsequent cataloging, identification, or sequencing of the generated amplicons. In some embodiments, barcode sequence(s) or tagging sequence(s) are incorporated into the amplified nucleotide sequence by inclusion in the amplification primer or by adapter ligation. Primers may further contain nucleotides useful for subsequent sequencing, for example, pyrosequencing. Such sequences are easily designed by commercially available software programs or companies.
[0094] In some embodiments, multiplex amplification is performed using target amplification primers that do not contain tagging sequences. In other embodiments, multiplex amplification is performed using amplification primers that each contain a targeting sequence and a tagging sequence, e.g., the forward primer or primer set contains tagging sequence 1 and the reverse primer or primer set contains tagging sequence 2. In yet other embodiments, multiplex amplification is performed using amplification primers where one primer or primer set contains a targeting sequence and a tagging sequence, and the other primer or primer set contains a targeting sequence but does not contain a tagging sequence, e.g., the forward primer or primer set contains a tagging sequence and the reverse primer or primer set does not contain a tagging sequence.
[0095] Thus, in some embodiments, multiple target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture using BCR- and / or TCR-directed amplification primers, where the forward and / or reverse primers contain tagging sequences, and the resulting amplicons contain the target BCR and / or TCR sequences and tagging sequences at one or both ends. In some embodiments, the forward and / or reverse amplification primer or primer sets may also contain barcodes, such that one or more barcodes are included in the resulting amplicons.
[0096] In some embodiments, multiple target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture using BCR and / or TCR-directed amplification primers, and the resulting amplicons contain only BCR and / or TCR sequences. In some embodiments, tagging sequences are added to the ends of such amplicons, e.g., via adapter ligation. In some embodiments, barcode sequences are added to one or both ends of such amplicons, e.g., via adapter ligation.
[0097] Nucleotide sequences suitable for use as barcodes and for barcode libraries are known in the art. Adapters and amplification primers and primer sets containing barcode sequences are commercially available. Oligonucleotide adapters containing barcode sequences are also commercially available, including, for example, IonXpress™, IonCode™, and Ion Select barcode adapters (Thermo Fisher Scientific). Similarly, additional and other universal adapter / primer sequences reported and known in the art (e.g., Illumina universal adapter / primer sequences, PacBio universal adapter / primer sequences, etc.) can be used in conjunction with the resulting amplicons sequenced using the methods and compositions and supporting analysis platforms provided herein.
[0098] In some embodiments, two or more barcodes are added to the amplicon when sequencing multiplexed samples. In some embodiments, at least two barcodes are added to the amplicon before sequencing the multiplexed samples to reduce the frequency of artifacts (e.g., immune receptor gene rearrangements or clonal identification) resulting from barcode cross-contamination or barcode bleed-through between samples. In some embodiments, at least two barcodes are used to label samples when tracking low-frequency clones in the immune repertoire. In some embodiments, at least two barcodes are added to the amplicon when the assay is used to detect clones at a frequency of less than 1:1,000. In some embodiments, at least two barcodes are added to the amplicon when the assay is used to detect clones at a frequency of less than 1:10,000. In other embodiments, at least two barcodes are added to the amplicon when the assay is used to detect clones at frequencies less than 1:20,000, less than 1:40,000, less than 1:100,000, less than 1:200,000, less than 1:400,000, less than 1:500,000, or less than 1:1,000,000. Methods for characterizing the immune repertoire that benefit from high sequencing depth per clone and / or detection of clones at such low frequencies include, but are not limited to, monitoring patients with hyperproliferative diseases undergoing treatment and testing for minimal residual disease after treatment.
[0099] In some embodiments, the target-specific primers used in the methods of the present invention (e.g., V gene FR1, FR2, and FR3-directed primers, J gene-directed primers, and C gene-directed primers) meet the following criteria: (1) contain two or more modified nucleotides within the primer sequence, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) are about 15 to about 40 bases in length; and (3) have a T of greater than 60°C and up to about 70°C.m (4) have low cross-reactivity to non-target sequences present in the sample of interest, (5) have at least the first four nucleotides (in the 3' to 5' direction) that are non-complementary to any sequence in any other primer present in the same reaction, and (6) are non-complementary to any contiguous stretch of at least five nucleotides in any other generated target amplicon. In some embodiments, the target-specific primers used in the provided methods are selected or designed to meet any two, three, four, five, or six of the above criteria.
[0100] In some embodiments, the target-specific primers used in the methods of the present invention comprise one or more modified nucleotides having a cleavable group. In some embodiments, the target-specific primers used in the methods of the present invention comprise two or more modified nucleotides having a cleavable group. In some embodiments, the target-specific primers comprise at least one modified nucleotide having a cleavable group selected from methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine.
[0101] In some embodiments, target amplicons generated using the amplification methods (and related compositions, systems, and kits) disclosed herein are used to prepare an immune receptor repertoire library. In some embodiments, the immune receptor repertoire library comprises introducing adapter sequences onto the ends of the target amplicon sequences. In certain embodiments, a method for preparing an immune receptor repertoire library comprises generating target immune receptor amplicon molecules according to any of the multiplex amplification methods described herein, processing the amplicon molecules by digesting modified nucleotides within the primer sequences of the amplicon molecules, and ligating at least one adapter to at least one of the processed amplicon molecules, thereby generating a library of adapter-linked target immune receptor amplicon molecules comprising the target immune receptor repertoire. In some embodiments, the library preparation step is performed in a single reaction vessel that includes only the addition step. In certain embodiments, the method further comprises clonally amplifying a portion of at least one adapter-linked target amplicon molecule.
[0102] In some embodiments, target amplicons generated using the methods disclosed herein (and related compositions, systems, and kits) are coupled to downstream processes, including, but not limited to, library preparation and nucleic acid sequencing. For example, target amplicons can be amplified using bridge amplification, emulsion PCR, or isothermal amplification to generate multiple clonal templates suitable for nucleic acid sequencing. In some embodiments, the amplicon library is sequenced using any suitable DNA sequencing platform, such as any next-generation sequencing platform, including semiconductor sequencing technologies such as Ion Torrent sequencing platforms. In some embodiments, the amplicon library is sequenced using the Ion GeneStudio S5 540™ system, the Ion GeneStudio S5 520™ system, the Ion GeneStudio S5 530™ system, or the Ion PGM 318™ system.
[0103] In some embodiments, sequencing of immune receptor amplicons generated using the methods disclosed herein (and related compositions and kits) generates contiguous sequence reads about 200 to about 600 nucleotides in length. In some embodiments, the contiguous read length is about 300 to about 400 nucleotides. In some embodiments, the contiguous read length is about 350 to about 450 nucleotides. In some embodiments, the read length averages about 300 nucleotides, about 350 nucleotides, or about 400 nucleotides. In some embodiments, the contiguous read length is about 250 to about 350 nucleotides, about 275 to about 340 nucleotides, or about 295 to about 325 nucleotides in length. In some embodiments, the read length averages about 270, about 280, about 290, about 300, or about 325 nucleotides in length. In other embodiments, the read length is about 180 to about 300 nucleotides, about 200 to about 290 nucleotides, about 225 to about 280 nucleotides, or about 230 to about 250 nucleotides. In some embodiments, the read length is, on average, about 200, about 220, about 230, about 240, or about 250 nucleotides. In other embodiments, the continuous read length is about 70 to about 200 nucleotides, about 80 to about 150 nucleotides, about 90 to about 140 nucleotides, or about 100 to about 120 nucleotides. In some embodiments, the continuous read length is about 50 to about 170 nucleotides, about 60 to about 160 nucleotides, about 60 to about 120 nucleotides, about 70 to about 100 nucleotides, about 70 to about 90 nucleotides, or about 80 nucleotides. In some embodiments, the read length averages about 70, about 80, about 90, about 100, about 110, or about 120 nucleotides. In some embodiments, the sequence read length includes the amplicon sequence and the barcode sequence. In some embodiments, the sequence read length does not include the barcode sequence.
[0104] In some embodiments, primers and primer pairs are target-specific nucleic acid molecules that can amplify specific regions of nucleic acid molecules. In some embodiments, the target-specific primers can amplify expressed RNA or cDNA. In some embodiments, the target-specific primers can amplify mammalian RNA, such as human RNA or cDNA prepared therefrom, or mouse RNA or cDNA prepared therefrom. In some embodiments, the target-specific primers can amplify DNA, such as gDNA. In some embodiments, the target-specific primers can amplify mammalian DNA, such as human DNA or mouse DNA.
[0105] In the methods and compositions provided herein, e.g., for determining, characterizing, and / or tracking immune repertoires in biological samples, the amount of input RNA or gDNA required for amplification of target sequences depends in part on the proportion of immune receptor-bearing cells (e.g., T cells or B cells) in the sample. For example, a higher proportion of B cells in a sample, such as a B cell-enriched sample, allows for the use of lower amounts of input RNA or gDNA for amplification. In some embodiments, the amount of input RNA for amplification of one or more target sequences can be from about 0.05 ng to about 10 micrograms. In some embodiments, the amount of input RNA used for multiplex amplification of one or more target sequences can be from about 5 ng to about 2 micrograms. In some embodiments, the amount of RNA used for multiplex amplification of one or more target sequences can be from about 5 ng to about 1 microgram or from about 10 ng to about 1 microgram. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 1.5 micrograms, about 2 micrograms, about 2.5 micrograms, about 3 micrograms, about 3.5 micrograms, about 4.0 micrograms, about 5 micrograms, about 6 micrograms, about 7 micrograms, or about 10 micrograms. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 25 ng to about 500 ng of RNA or about 50 ng to about 200 ng of RNA. In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is from about 0.05 ng to about 10 ng of RNA, from about 0.1 ng to about 5 ng of RNA, from about 0.2 ng to about 2 ng of RNA, or from about 0.5 ng to about 1 ng of RNA.In some embodiments, the amount of RNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0106] As described herein, RNA from a biological sample is converted to cDNA prior to multiplex amplification, typically using a reverse transcriptase in a reverse transcription reaction. In some embodiments, a reverse transcription reaction is performed using input RNA, and a portion of the cDNA from the reverse transcription reaction is used in the multiplex amplification reaction. In some embodiments, substantially all of the cDNA prepared from the input RNA is added to the multiplex amplification reaction. In other embodiments, only a portion of the cDNA prepared from the input RNA, such as about 80%, about 75%, about 66%, about 50%, about 33%, or about 25%, is added to the multiplex amplification reaction. In other embodiments, about 15%, about 10%, about 8%, about 6%, or about 5% of the cDNA prepared from the input RNA is added to the multiplex amplification reaction.
[0107] In some embodiments, the amount of cDNA from a sample added to a multiplex amplification reaction can be about 0.001 ng to about 5 micrograms. In some embodiments, the amount of cDNA used for the multiplex amplification of one or more immune repertoire target sequences can be about 0.01 ng to about 2 micrograms. In some embodiments, the amount of cDNA used for the multiplex amplification of one or more target sequences can be about 0.1 ng to about 1 microgram or about 1 ng to about 0.5 micrograms. In some embodiments, the amount of cDNA used for the multiplex amplification of one or more immune repertoire target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.01 ng to about 10 ng of cDNA, about 0.05 ng to about 5 ng of cDNA, about 0.1 ng to about 2 ng of cDNA, or about 0.01 ng to about 1 ng of cDNA. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.005 ng, about 0.01 ng, about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0108] In some embodiments, mRNA is obtained from biological samples and converted into cDNA using conventional methods for amplification purposes.Methods and reagents for extracting or isolating nucleic acid from biological samples are well known and commercially available.In some embodiments, RNA extraction from biological samples is carried out by any method described herein or otherwise known to those skilled in the art, such as proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with DNase to digest contaminating DNA, and RNA purification using silica-gel-membrane technology, or any combination thereof. Exemplary methods for RNA extraction from biological samples using commercially available kits include the RecoverAll™ Multi-Sample RNA / DNA Workflow (Invitrogen), RecoverAll™ Total Nucleic Acid Isolation Kit (Invitrogen), NucleoSpin® RNA Blood (Macherey-Nagel), PAXgene® Blood RNA System, TRI Reagent™ (Invitrogen), PureLink™ RNA Microscale Kit (Invitrogen), MagMAX™ FFPE DNA / RNA Ultra Kit (Applied Biosystems), ZR RNA MicroPrep™ Kit (Zymo Research), RNeasy Micro Kit (Qiagen), and ReliaPrep™ RNA Tissue miniPrep System (Promega).
[0109] In some embodiments, the amount of input gDNA for amplification of one or more target sequences can be about 0.1 ng to about 10 micrograms. In some embodiments, the amount of gDNA required for amplification of one or more target sequences can be about 0.5 ng to about 5 micrograms. In some embodiments, the amount of gDNA required for amplification of one or more target sequences can be about 1 ng to about 1 microgram or about 10 ng to about 1 microgram. In some embodiments, the amount of gDNA required for amplification of one or more immune repertoire target sequences is about 10 ng to about 500 ng, about 25 ng to about 400 ng, or about 50 ng to about 200 ng. In some embodiments, the amount of gDNA required for amplification of one or more target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 20 ng, about 50 ng, about 100 ng, or about 200 ng. In some embodiments, the amount of gDNA required for amplification of one or more immune repertoire target sequences is about 1 microgram, about 2 micrograms, about 3 micrograms, about 4.0 micrograms, or about 5 micrograms.
[0110] In some embodiments, gDNA is obtained from biological samples using conventional methods.Methods and reagents for extracting or isolating nucleic acid from biological samples are well known and commercially available.In some embodiments, DNA extraction from biological samples is carried out by any method described herein or otherwise known to those skilled in the art, such as proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with RNase to digest contaminating RNA, and DNA purification using silica-gel-membrane technology, or any combination thereof. Exemplary methods of DNA extraction from biological samples using commercially available kits include the Ion AmpliSeq™ Direct FFPE DNA Kit, MagMAX™ FFPE DNA / RNA Ultra Kit, TRI Reagent™ (Invitrogen), PureLink™ Genomic DNA Mini Kit (Invitrogen), RecoverAll™ Total Nucleic Acid Isolation Kit (Invitrogen), MagMAX™ DNA Multi-Sample Kit (Invitrogen), and DNA extraction kits from BioChain Institute Inc. (e.g., FFPE Tissue DNA Extraction Kit, Genomic DNA Extraction Kit, Blood and Serum DNA Isolation Kit).
[0111] As used herein, a sample or biological sample refers to a composition from an individual that contains or may contain cells related to the immune system. Exemplary biological samples include, but are not limited to, tissues (e.g., lymph nodes, organ tissues, bone marrow), whole blood, synovial fluid, cerebrospinal fluid, tumor biopsies, and other clinical specimens containing cells. The sample may contain normal and / or diseased cells and may be a fine-needle aspirate, fine-needle biopsy, core sample, or other sample. In some embodiments, the biological sample may contain hematopoietic cells, peripheral blood mononuclear cells (PBMCs), T cells, B cells, tumor-infiltrating lymphocytes ("TILs"), or other lymphocytes. In some embodiments, the sample may be fresh (e.g., not preserved), frozen, or formalin-fixed, paraffin-embedded tissue (FFPE). Some samples contain cancer cells, such as carcinoma, melanoma, sarcoma, lymphoma, myeloma, leukemia, etc., and the cancer cells may be circulating tumor cells. In some embodiments, the biological sample contains cfDNA, for example, as found in blood or plasma.
[0112] The biological sample can be a mixture of tissues or cell types, a cell preparation enriched for at least one specific category or type of cells, or an isolated population of cells of a specific type or phenotype. The sample can be separated before analysis by centrifugation, elutriation, density gradient separation, apheresis, affinity selection, panning, FACS, centrifugation with Hypaque, etc. Methods for sorting, enriching, and isolating specific cell types are well known and can be easily performed by those skilled in the art. In some embodiments, the sample can be a preparation enriched for B cells.
[0113] In some embodiments, the methods and systems provided include analysis of immune repertoire receptor cDNA or gDNA sequence data and processes for identifying and / or removing PCR- or sequencing-derived error(s) from the determined immune receptor sequences.
[0114] In some embodiments, the error correction strategy comprises: 1) aligning the sequenced rearrangements to a reference database of variable, diverse, and linkage / constant genes to generate query / reference sequence pairs; many alignment procedures can be used for this purpose, including, for example, IgBLAST, a free publicly available tool from NCBI, and custom computer scripts; 2) realigning the reference and query sequences to each other, taking into account the flow order used for sequencing, which flow order allows for identifying and correcting some types of misalignments; 3) identifying the boundaries of the CDR3 region by their characteristic sequence motifs; 4) identifying indels in the query relative to the reference across the aligned portions of the rearrangement corresponding to the variable and binding / constant genes, excluding the CDR3 region, and changing the mismatched query base positions to match the reference; 5) For the CDR3 region, if the CDR3 length is not a multiple of 3 (indicating an indel error): (a) searching the CDR3 for homopolymer stretches that have the highest probability of containing a sequence error based on the PHRED score (denoted by e); (b) obtaining the probability of error across the CDR3 region based on the PHRED score (denoted t); (c) if e / t is higher than a defined threshold, editing the homopolymer by either increasing or decreasing the length of the homopolymer by one base so that the CDR3 nucleotide length is a multiple of three; (d) as an alternative to steps a-c, searching for the longest homopolymer in the CDR3, and if the length of the homopolymer exceeds a defined threshold, editing the homopolymer by either increasing or decreasing the length of the homopolymer by one base so that the CDR3 nucleotide length is a multiple of three.
[0115] In some embodiments, methods are provided for identifying B cell and / or T cell clones in repertoire data that are robust to PCR and sequencing errors. Accordingly, steps that may be used in such methods to identify B cell and / or T cell clones in a manner that is robust to PCR and sequencing errors are described below. Table 1 is a diagram of an exemplary workflow for use in identifying and removing PCR- or sequencing-derived errors from immune receptor sequencing data. Exemplary portions and embodiments of this workflow are also shown in Figures 1-2. [Table 1]
[0116] In a set of mRNA or gDNA-derived TCR or BCR sequences, 1) each sequence is annotated as either native or as a productive rearrangement, either after error correction as described above, and 2) each sequence has identified V gene and CDR3 nucleotide regions, in some embodiments, the method comprises: 1) Identify and remove chimeric sequences. For each unique CDR3 nucleotide sequence present in the dataset, the number of reads of that CDR3 nucleotide sequence and potential V genes is tallied. V gene-CDR3 combinations that account for less than 10% of the total reads in that CDR3 nucleotide sequence are flagged as chimeric and excluded from downstream analysis. For example, for the following sequences with the same CDR3 nucleotide sequence, for example, sequences with TRBV3 and TRBV6 paired with the CDR3nt sequence AATTGGT are flagged as chimeric. [Table 2]
[0117] 2) Identify and exclude sequences containing simple indel errors. For each read in the dataset, obtain a homopolymer-collapsed representation of the CDR3 sequence of that read. For each set of reads with the same V gene and collapsed CDR3 combination, tally the occurrence of each uncollapsed CDR3 nucleotide sequence. Uncollapsed CDR3 sequences that make up <10% of the total reads in that read set are flagged as having a simple homopolymer error. As an example, three different V gene-CDR3 nucleotide sequences are shown as being identical after homopolymer collapse of the CDR3 nucleotide sequence. Two less frequent V gene-CDR3 combinations make up less than <10% of the total reads in the read set and are flagged as containing a simple indel error. For example, [Table 3]
[0118] 3) Identify and filter out singleton reads. For each read in the dataset, count the number of times the exact read sequence is found in the dataset. Reads that appear only once in the dataset are flagged as singleton reads. 4) Identify and filter out truncated reads. For each read in the dataset, determine whether the read has the annotated V gene FR1, CDR1, FR2, CDR2, and FR3 regions as indicated by an IgBLAST alignment of the read against the IgBLAST reference V gene set. Reads that do not have the above regions are flagged as truncated if the region(s) are expected based on the specific V gene primers used for amplification. 5) Identify and filter out rearrangements that lack bidirectional support. For each read in the dataset, obtain the V gene and CDR3 sequences of the read, as well as the strand orientation (plus or minus strand) of the read. For each V gene-CDR3 combination in the dataset, tally the number of reads on the plus and minus strands that have that V gene-CDR3nt combination. V gene-CDR3nt combinations that are present in reads in only one direction are considered spurious. All reads with spurious V gene-CDR3nt combinations are flagged as lacking bidirectional support. 6) For unflagged genes, perform stepwise clustering based on CDR3 nucleotide similarity. Separate sequences into groups based on the V gene identity of the reads, excluding allele information (v gene groups). For each group, a. The readouts in each group were analyzed using cd-hit-est and the following parameters: Cluster them using cd-hit-est-i vgene_groups.fa-o clustered_vgene_groups.cdhit-T24-d0-M100000-B0-r0-g1-S0-U2-uL.05-n10-l7. (The freely available software program cd-hit-est clusters nucleotide datasets into clusters that meet a user-defined similarity threshold. (See https: / / github.com / weizhongli / cdhit / wiki / 3.-User%27s-Guide#CDHITEST for code and instructions on cd-hit-est.) Here, vgene_groups.fa is a fasta format file of the CDR3 nucleotide regions of sequences with the same V gene, and clustered_vgene_groups.cdhit is the output containing the refined sequences. b. The same clone ID is assigned to each sequence within a cluster and is used to indicate that members of a subgroup are likely to represent the same T-cell or B-cell clone. c. For each cluster, a representative sequence is selected such that it is the sequence that appears the most times, or in the case of ties, is selected randomly. d. Merge all other reads in the cluster into the representative sequence, thereby increasing the number of reads in the representative sequence according to the number of reads in the merged sequence. Representative sequences within the ev gene cluster are compared to each other based on Hamming distance. If a representative sequence is within a Hamming distance of 1 to a representative sequence that is >50-fold more abundant, the sequence is merged with the more common representative sequence. If a representative sequence is within a Hamming distance of 2 to a representative sequence that is >10,000-fold more abundant, the sequence is merged with the more common representative sequence. f. Identify complex sequence errors. Representative sequences within each V gene group are homopolymer-collapsed and then compared to each other using the Levenshtein distance. If a representative sequence is within a Levenshtein distance of 1 to a representative sequence that is >50-fold more abundant, that sequence is merged with the more common representative sequence. g. Identify CDR3 misannotation errors. Representative sequences within each V gene cluster are homopolymer-disrupted, followed by stepwise comparison of each homopolymer-disrupted sequence. For each pair of sequences, determine whether one sequence is a subset of the other. If so, merge the less abundant sequence with the more abundant sequence if the more abundant sequence is >500-fold more abundant. 7) Report the cluster representative to the user.
[0119] In some embodiments, step 6 of the above workflow groups the rearranged sequences based on V gene identity (excluding allele information) and CDR3 nucleotide length. In other embodiments, J gene identity and / or isotype identity are also used as part of the grouping criteria. Thus, in some embodiments, step 6 of the above workflow comprises the following steps: a. The readouts in each group were analyzed using cd-hit-est and the following parameters: Place in a cluster using cd-hit-est-i vgene_groups.fa-o clustered_vgene_groups.cdhit-T24-l9-d0-M100000-B0-r0-g1-S15-U2-uL.05-n9. where vgene_groups.fa is a fasta format file of the sequenced portion of the VDJ rearrangement. In some embodiments, the entire sequence of the VDJ is considered for clustering, as somatic hypermutation can occur throughout the VDJ region. b. The same clone ID is assigned to each sequence within a cluster and is used to indicate that members of a subgroup are likely to represent the same T-cell or B-cell clone. c. For each cluster, a representative sequence is selected such that it is the sequence that appears the most times, or in the case of ties, is selected randomly. d. Merge all other reads in the cluster into the representative sequence, thereby increasing the number of reads in the representative sequence according to the number of reads in the merged sequence. Representative sequences within the ev gene cluster are compared to each other based on Hamming distance. If a representative sequence is within a Hamming distance of 1 to a representative sequence that is >50 times more abundant, the sequence is merged into the more common representative sequence. If a representative sequence is within a Hamming distance of 2 to a representative sequence that is >10,000 times more abundant, the sequence is merged into the more common representative sequence. In some embodiments, a fold threshold of >50 / 3 and >1,000 / 3 is used to merge sequences with a Hamming distance of 1 or 2, respectively. Since the longer the sequence, the higher the chance of accumulating amplification and / or sequence errors, lowering the fold threshold can be useful when comparing sequences of the entire VDJ region, rather than sequences of only the CDR3 region. f. Identify complex sequence errors. Representative sequences within each V gene group are homopolymer-collapsed and then compared to each other using the Levenshtein distance. If a representative sequence is within a Levenshtein distance of 1 to a representative sequence that is >50-fold more abundant, that sequence is merged with the more common representative sequence. g. Identify CDR3 misannotation errors. Representative sequences within each V gene cluster are homopolymer-disrupted, followed by stepwise comparison of each homopolymer-disrupted sequence. For each pair of sequences, determine whether one sequence is a subset of the other. If so, merge the less abundant sequence with the more abundant sequence if the more abundant sequence is >500-fold more abundant.
[0120] In some embodiments, the provided workflow is not limited to the frequency ratio thresholds listed in various steps, and other frequency ratio thresholds can be substituted for the representative frequency ratio thresholds included above. The frequency ratio refers to the ratio of the abundance value of the more common representative sequence to the abundance value of the less common representative sequence. The frequency ratio threshold provides a threshold at which the less common representative sequence is merged into the more common representative sequence. For example, in some embodiments, comparing representative sequences within a v gene group to each other based on Hamming distance can use frequency ratio thresholds other than those listed in step (e) above. For example, but not limited to, if a representative sequence is within a Hamming distance of 2 to the representative sequence, a frequency ratio threshold of 1000, 5000, 20,000, etc. can be used. For example, but not limited to, if a representative sequence is within a Hamming distance of 1 to the representative sequence, a frequency ratio threshold of 20, 100, 200, etc. can be used. The frequency ratio thresholds provided represent a general process of denoting the more abundant sequence of a similar pair as the correct sequence.
[0121] Similarly, when comparing the frequencies of two sequences in other steps in the workflow, such as step (1), step (2), step (6f), and step (6g), frequency ratio thresholds other than those listed in the steps above may be used.
[0122] As used herein, the term "homopolymer disruption sequence" is intended to refer to a sequence in which repeated bases are disrupted to single base representatives. For example, in the undisrupted sequence AAAATTTTTATCCCCCCCCGGG (SEQ ID NO: 603), the homopolymer disruption sequence is ATATCG.
[0123] As used herein, the terms "clone," "clonotype," "strain," or "rearrangement" are intended to describe a unique V gene nucleotide combination in an immune receptor, such as a TCR or BCR, e.g., a unique V gene-CDR3 nucleotide combination.
[0124] As used herein, the term "productive read" refers to a TCR or BCR sequence read that does not have a stop codon and has in-frame variable gene and joining gene segments. A productive read is biologically reliable in encoding a polypeptide.
[0125] As used herein, "chimera" or "chimeric sequence" refers to an artificial sequence resulting from template switching during target amplification, such as PCR. Chimeras typically exist as CDR3 sequences grafted onto unrelated V genes, resulting in a CDR3 sequence associated with multiple V genes in a dataset. Chimeric sequences are usually much less abundant than true sequences in a dataset.
[0126] As used herein, the term "indel" refers to the insertion and / or deletion of one or more nucleotide bases in a nucleic acid sequence. In the coding region of a nucleic acid sequence, if the length of the indel is not a multiple of three, a frameshift occurs when the sequence is translated. As used herein, a "simple indel error" is an error that does not change the homopolymer fragment representation of the sequence. As used herein, a "complex indel error" is an indel sequencing error that changes the homopolymer fragment representation of the sequence, including, but not limited to, errors that eliminate homopolymers, insert homopolymers into the sequence, or create unreadable errors.
[0127] As used herein, a "singleton read" refers to a sequence read in which the indel-corrected sequence appears only once in a dataset. Typically, singleton reads are enriched with reads containing PCR or sequencing errors.
[0128] As used herein, " truncated reads " refers to immune receptor sequence reads that lack annotated V gene regions. For example, truncated reads include, but are not limited to, sequence reads that lack annotated TCR or BCRV gene FR1, CDR1, FR2, CDR2, or FR3 regions. Such reads generally lack part of the V gene sequence due to quality trimming. Truncated reads can produce artifacts if the truncation causes misidentification of V genes.
[0129] In the context of identified V gene-CDR3 sequences (clonotypes), "bidirectional support" indicates that a particular V gene-CDR3 sequence is found in at least one read mapping to the plus strand (going from the V gene toward the constant gene) and in at least one read mapping to the minus strand (going from the constant gene toward the V gene). Systematic sequencing errors often result in the identification of V gene-CDR3 sequences with unidirectional support.
[0130] In a set of sequences that have been grouped according to a predetermined sequence similarity threshold to account for variation due to PCR or sequencing errors, the "cluster representative" is the sequence selected as most likely to be error-free. This is typically the most abundant sequence.
[0131] As used herein, "IgBLAST annotation errors" refers to rare events in which the CDR3 boundary is identified as being in the wrong adjacent position. These events typically add three bases to the 5' or 3' end of the CDR3 nucleotide sequence.
[0132] In two sequences of equal length, the "Hamming distance" is the number of positions where corresponding bases or amino acids differ. In any two sequences, the "Levenshtein distance" or "edit distance" is the number of single base or amino acid edits required to make one nucleotide or amino acid sequence into another.
[0133] In some embodiments, in which J gene-directed primers are used for amplifying immune receptor sequences, e.g., multiplex amplification with primers for the V gene FR3 region and primers for the J gene, the raw sequence reads derived from the assay undergo J gene sequence inference before any downstream analysis. In this process, the beginning and end of the raw read sequence are examined for the presence of a 10-30 nucleotide signature corresponding to a portion of the J gene sequence predicted to be present after amplification with the J primer and subsequent manipulation or treatment (e.g., digestion) of the amplicon ends before sequencing. The signature nucleotide sequence allows the sequence of the J primer and the remainder of the targeted J gene to be inferred because the sequence of each J gene is known. To complete the J gene sequence inference process, the inferred J gene sequence is added to the raw read, and an extended read spanning the entire J gene is then created. The extended read, which includes the entire J gene sequence, the entire CDR3 region, and at least a portion of the V gene sequence, is then reported after downstream analysis. The portion of the V gene sequence in the extended read depends on the V gene-directed primer, e.g., FR3, FR2, or FR1-directed primer, used in the multiplex amplification.
[0134] The use of V gene FR3 primers and J gene primers to amplify expressed or rearranged immune receptor gDNA sequences generates minimally long amplicons (e.g., approximately 60–100 or approximately 80 nucleotides in length) while still generating data that allow reporting of the entire CDR3 region. Due to the expected short amplicon length, amplicon reads <100 nucleotides in length are not eliminated as low-quality and / or off-target products during the sequence analysis workflow. However, explicit searching for expected J gene sequences in raw reads allows for the elimination of amplicons derived from off-target amplification by J gene primers. Additionally, this short amplicon length improves assay performance on highly degraded template material, such as those derived from FFPE or cfDNA samples.
[0135] In some embodiments, provided methods include sequencing an immune receptor library, subjecting the resulting sequence data to an error identification and correction process to generate rescued productive reads, and identifying the productive and rescued productive sequence reads. In some embodiments, provided methods include sequencing an immune receptor library, subjecting the resulting sequence data set to an error identification and correction process, identifying the productive and rescued productive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes in the library.
[0136] In some embodiments, provided methods include sequencing a rearranged immune receptor DNA library, subjecting the resulting sequence data to an error identification and correction process for the V gene portions to generate rescued productive reads, and identifying productive, rescued productive, and non-productive sequence reads. In some embodiments, provided methods include sequencing a rearranged immune receptor DNA library, subjecting the resulting sequence data set to an error identification and correction process for the V gene portions, identifying productive, rescued productive, and non-productive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes within the library. In some embodiments, both productive and non-productive sequence reads of the rearranged immune receptor DNA are reported separately.
[0137] In some embodiments, the provided error identification and correction workflows are used to identify and resolve errors from PCR or sequencing that lead to sequence reads identified as resulting from non-productive rearrangements. In some embodiments, the provided error identification and correction workflows are applied to immune receptor sequence data generated from sequencing platforms that introduce errors that cause indels or other frameshifts when generating the sequence data.
[0138] In some embodiments, the provided error identification and correction workflows are applied to sequencing data generated by an Ion Torrent sequencing platform. In some embodiments, the provided error identification and correction workflows are applied to sequence data generated by a Roche 454 Life Sciences sequencing platform, a PacBio sequencing platform, and an Oxford Nanopore sequencing platform.
[0139] In some embodiments, the BCR repertoire analysis workflow includes an additional final step of identifying clonal lineages in the sample. Clonal lineages represent sets of B cell clones (e.g., identified as having unique VDJ sequences) that originate from a common VDJ rearrangement but differ in somatic hypermutation and / or class switch recombination. It is generally expected that members of a clonal lineage may be more likely to target the same antigen than members of different clonal lineages.
[0140] In some embodiments, the process of clonal lineage identification involves using a set of identified BCR clones (e.g., IgH clones) (e.g., as described herein) to: 1. Divide the clone sequences into groups whose members share the same variable genes (excluding allelic information), the same CDR3 nucleotide length, and the same joining genes (excluding allelic information). In some embodiments, the J gene criteria above can be omitted. 2. The clone sequences in each group are arranged into clusters based on the CDR3 nucleotide similarity of the clone sequences. The threshold for CDR3 nucleotide similarity is about 0.70 to about 0.99. In some embodiments, the threshold for CDR3 nucleotide similarity is about 0.80 to about 0.99. In some embodiments, the threshold for CDR3 nucleotide similarity is about 0.80 to about 0.90. In certain embodiments, the threshold for CDR3 nucleotide similarity is about 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99. a. In some embodiments, the clustering comprises: This was performed using cd-hit-est as described in cd-hit-est-i vgene_groups.fa-o clustered_vgene_groups.cdhit-T24-l9-d0-M100000-B0-r0-g1-S0-c.85-n5, where vgene_groups.fa consists of the set of CDR3 nucleotide sequences for each clone in the group. Clones in the same cluster are considered members of the same clonal lineage. b. In some cases, somatic hypermutation may be so extensive that the described clustering criteria cannot separate all clonal lineage members. In such cases, some embodiments perform an additional step to merge the clusters identified in (a). The additional step consists of searching for instances of mutations resulting from shared somatic hypermutation in variable genes between clonal lineages and then merging clonal lineages if the percentage and / or number of shared mutations exceeds a certain threshold. Variable gene mutations are identified by comparing variable gene sequences to the best matching variable gene sequences in the IMGT database, as described. In some embodiments, the threshold for the number of shared mutations is 2 or greater. In some embodiments, the threshold for the number of shared mutations is 3 or greater. In other embodiments, the threshold for the number of shared mutations is 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the percentage of shared mutations is about 0.15 to about 0.95. In some embodiments, the percentage of shared mutations is about 0.75 to about 0.85. In other embodiments, the proportion of shared mutations is about 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 0.95.
[0141] In some cases, some alleles of variable genes not represented in the IMGT database may be identified. In such cases, alignment with the IMGT database shows discrepancies that are not due to somatic hypermutation. To avoid noise caused by such unannotated genetic variations, in some embodiments, a first step is performed before (b), which identifies all putative novel variable gene alleles in the sample and notes each position that differs from the reference. In some embodiments, such positions are excluded from consideration in the analysis described in (b). Methods for identifying novel alleles from immune repertoire sequencing data are described, for example, in Gadala-Maria et al. (2015) Proc. Natl. Acad. Sci. USA 112:E862-E870 and PCT Publication No. 2018 / 136562.
[0142] At the end of this clonal lineage identification process, each clone is assigned to a clonal lineage. BCR repertoire characteristics such as diversity, uniformity, and convergence can be calculated using the clonal lineage as the unit of analysis. In some embodiments, characteristics of clonal lineages such as the number of clones belonging to a lineage, the isotypes of those clones, the maximum and minimum frequency of clones in the lineage, the maximum and minimum somatic hypermutation of variable genes in the lineage, and others are calculated and reported to the user.
[0143] In the absence of somatic hypermutation, BCR convergence can be calculated as the frequency of clones that are identical or functionally identical in amino acid sequence but differ in nucleotide sequence. These clones independently undergo VDJ recombination and generally represent clones that are assumed to have proliferated in response to a common antigen. However, somatic hypermutation can create distinct VDJ sequences that do not represent B cells independently undergoing VDJ recombination. To account for this, a definition of convergence that takes into account the identification of clonal lineages is used. For this purpose, "BCR convergence" is defined as the frequency of B cell clones that are members of different clonal lineages but have similar or identical amino acid sequences, as determined above. In some embodiments, two IGH rearrangements are considered convergent if they are assigned to separate clonal lineages but have the same variable genes (excluding allelic information) and the same or similar CDR3 amino acid sequences. In other embodiments, where sequencing covers all three CDR domains of the IGH chain, two IGH rearrangements can be considered convergent if they are assigned to separate clonal lineages but have the same variable genes (excluding allelic information) and the same or similar CDR1, 2, and 3 amino acid sequences. In some embodiments, similar CDR amino acid sequences are within a Hamming or Levenshtein edit distance of 1. In other embodiments, similar CDR amino acid sequences are within a Hamming or Levenshtein edit distance of 2.
[0144] Thus, in some embodiments, functionally equivalent B cells are identified by searching for BCR clones with the same variable gene and CDR amino acid sequences within a Hamming or Levenshtein edit distance of 1 or 2. In some embodiments, the program cd-hit is used to identify clones with similar but functionally equivalent amino acid sequences. (See https: / / github.com / weizhongli / cdhit / wiki / 3.-User%27s-Guide for code and information about the program cd-hit.) In some embodiments, cd-hit uses the following command: This was performed using cd-hit-i vgene_groups.fa-o clustered_vgene_groups.cdhit-T24-l5-d0-M100000-B0-g1-S1-U1-n5, where vgene_groups.fa consists of a set of CDR3 amino acid sequences of clones with the same variable gene. Clones within the same cluster are considered functionally equivalent. In some embodiments, the value of the parameter -S may be 0, 1, 2, or 3. In some embodiments, the value of the parameter -U may be 0, 1, 2, or 3. In some embodiments, vgene_groups.fa consists of a set of CDR1, 2, and 3 amino acid sequences of clones that have the same variable genes. In some embodiments, vgene_groups.fa consists of a set of clones that have both the same variable genes and the same CDR3 length.
[0145] In some embodiments, the provided sequence analysis workflow includes downsampling analysis. In immune repertoire sequencing and subsequent analysis, the use of downsampling analysis can help eliminate variations due to, for example, differences in sequencing depth across assays. For example, an exemplary downsampling analysis for use in RNA or cDNA sequencing and analysis workflows applies the following procedure to the data: a) start with the entire set of productive and rescued productive reads, and randomly reduce and remove sequence reads to one of several fixed read intensities; b) use this subset of reads to perform downstream calculations (e.g., clonotyping and calculation of secondary repertoire characteristics, including but not limited to, uniformity, convergence, diversity, the number and identity of detected clones, and clonal lineage).
[0146] In some embodiments, downsampling analysis identifies the point at which a particular sample is sequenced to saturation, e.g., the point at which additional reads do not identify additional clones or lineages or add additional diversity to the detected repertoire. In some embodiments, downsampling allows for refinement of sequencing intensity or multiplexing between assays (three or more sample types) or between assays (two sample types) using similar sample types.
[0147] In some embodiments, the set of variable gene alleles detected by the provided assay methods and compositions can be used for the de novo identification of haplotype groups within a human population. In certain embodiments, the provided assay methods and compositions, including the use of multiple V gene-specific primers and at least one C gene-specific primer to amplify IgH CDR1, 2, and 3 nucleotide sequences, can be used to identify the IgH haplotypes of a subject's BCR repertoire. For example, in some embodiments, the provided methods and compositions can be used to identify the IgH haplotypes of a subject's BCR repertoire, using at least one set of primers including multiple V gene FR1 primers selected from Table 3 and at least one C gene primer selected from Tables 6-10. Methods for identifying TCR haplotype groups are described in PCT Application No. PCT / US2019 / 023731, filed March 22, 2019, which is incorporated herein by reference in its entirety, and can similarly be used in conjunction with the methods and compositions provided herein to identify IgH haplotype groups. In some embodiments, a set of variable gene alleles detected by amplifying and sequencing IgH CDR1, 2, and 3 nucleotide sequences can be used to assign a sample to one of several existing haplotype groups as part of a number of procedures for predicting the risk of autoimmune disease or adverse events following immunotherapy. Methods for assigning samples to haplotype groups in procedures for predicting the risk of autoimmune disease or adverse events following immunotherapy are also described in PCT Application No. PCT / US2019 / 023731, filed March 22, 2019, and incorporated herein by reference, and can be used in combination with the methods and compositions provided herein to assign samples to IgH haplotype groups, for example, to predict such risk. In some embodiments, IgH CDR1, 2, and 3 sequence data obtained using the provided assay methods and compositions can be used to infer a staged IgH locus haplotype (e.g., Kidd et al. (2012) J. Immunol. 188(3):1333-1340).
[0148] In some embodiments, provided methods include preparing and forming a plurality of immune receptor-specific amplicons. In some embodiments, the methods include hybridizing a plurality of V gene-specific primers and at least one C gene-specific primer to a cDNA molecule, extending a first primer of the primer pair (e.g., the V gene-specific primer), denaturing the extended first primer from the cDNA molecule, hybridizing a second primer of the primer pair (e.g., the C gene-specific primer) to the extended first primer product and extending the second primer, and digesting the target-specific primer pair to generate a plurality of target amplicons. In other embodiments, the method includes hybridizing a plurality of V gene-specific primers and a plurality of J gene-specific primers to a cDNA molecule, extending a first primer of the primer pair (e.g., the V gene-specific primer), denaturing the extended first primer from the cDNA molecule, hybridizing a second primer of the primer pair (e.g., the J gene-specific primer) to the extended first primer product to extend the second primer, and digesting the target-specific primer pair to generate a plurality of target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons before performing a nick translation reaction to generate a plurality of target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters comprises at least one barcode sequence. In some embodiments, each adapter ligated to the end of a target amplicon comprises a barcode sequence. In some embodiments, one or more target amplicons can be amplified using bridge amplification, emulsion PCR, or isothermal amplification to generate a plurality of clonal templates suitable for nucleic acid sequencing.
[0149] In some embodiments, provided methods include preparing and forming multiple immune receptor-specific amplicons. In some embodiments, the methods include hybridizing multiple V gene-specific primers and multiple J gene-specific primers to a gDNA molecule, extending a first primer (e.g., the V gene-specific primer) of the primer pair, denaturing the extended first primer from the gDNA molecule, hybridizing a second primer (e.g., the J gene-specific primer) of the primer pair to the extended first primer product and extending the second primer, and digesting the target-specific primer pair to create multiple target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons before performing a nick translation reaction to generate multiple target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters comprises at least one barcode sequence. In some embodiments, each adapter ligated to the end of a target amplicon comprises a barcode sequence. In some embodiments, one or more target amplicons can be amplified using bridge amplification or emulsion PCR to generate multiple clonal templates suitable for nucleic acid sequencing.
[0150] In some embodiments, the present disclosure provides methods for sequencing target amplicons and processing the sequence data to identify productive immune receptor rearrangements expressed in a biological sample from which the cDNA was derived. In other embodiments, the present disclosure provides methods for sequencing target amplicons and processing the sequence data to identify productive immune receptor gene-rearranged gDNA from a biological sample. In embodiments in which J gene-directed primers are used to amplify expressed immune receptor sequences or rearranged immune receptor gDNA sequences, processing the sequence data includes inferring the nucleotide sequence of the J gene primer used for amplification as well as the remainder of the targeted J gene, as described herein. In some embodiments, processing the sequence data includes performing provided error identification and correction steps to generate rescued productive sequences. In some embodiments, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that are at least 50% of the sequencing reads in an immune receptor cDNA or gDNA sample. In some embodiments, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that are at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the sequencing reads in an immune receptor cDNA or gDNA sample. In some embodiments, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that are about 50-60%, about 60-70%, about 70-80%, about 80-90%, about 50-80%, or about 60-90% of the sequencing reads in an immune receptor cDNA or gDNA sample.In some embodiments, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that average about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90% of the sequencing reads in an immune receptor cDNA or gDNA sample.
[0151] For certain samples, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that are less than 50% of the sequencing reads in an immune receptor cDNA or gDNA sample when the particular sample is used. Such samples include, for example, FFPE samples and cfDNA samples in which the RNA or gDNA is highly degraded, as well as samples in which the number of target immune cells is very low, such as samples with very low B cell counts or samples from subjects experiencing severe leukopenia. Thus, in some embodiments, use of the provided error identification and correction workflows can result in a combination of productive reads and rescued productive reads that are about 30-50%, about 40-50%, about 30-40%, about 40-60%, at least 30%, or at least 40% of the sequencing reads in an immune receptor cDNA or gDNA sample.
[0152] In certain embodiments, the methods of the present invention involve the use of a target immune receptor primer set, in which the primers are directed to the same target immune receptor gene sequence, e.g., BCR (immunoglobulin) and TCR gene sequence. In some embodiments, the immune receptor is an antibody receptor selected from the group consisting of heavy chain alpha, heavy chain delta, heavy chain epsilon, heavy chain gamma, heavy chain mu, light chain kappa, and light chain lambda. In some embodiments, the T cell receptor is a T cell receptor selected from the group consisting of TCR alpha, TCR beta, TCR gamma, and TCR delta. In some embodiments, the methods of the present invention involve the use of a target immune receptor primer set, in which at least one of the primer sets is directed to a BCR sequence and another primer set is directed to a TCR sequence, and both BCR and TCR target nucleic acids from a sample are amplified in a single multiplex amplification reaction.
[0153] In certain embodiments, a method for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of a framework region within the V gene; and ii) one or more C gene primers for at least a portion of each target constant gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating immune receptor amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region.
[0154] In certain embodiments, a method for amplifying expressed nucleic acid sequences of an immune receptor repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 1 (FR1) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating immune receptor amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In some embodiments, the one or more V gene primers of i) anneal to at least a portion of framework region 1 of the template molecule. In specific embodiments, the one or more C gene primers of ii) comprise at least two primers that anneal to at least a portion of a C gene portion of the template molecule. In some embodiments, the one or more C gene primers of ii) comprise at least two primers, each annealing to at least a portion of a C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers of ii) separately comprise at least one primer for a portion of each of the C genes of an IgA, IgD, IgG, IgM, and IgE template molecule. In specific embodiments, at least one set of generated amplicons comprises complementarity determining regions CDR1, CDR2, and CDR3 of a BCR expression sequence.In some embodiments, the amplicon is about 300 to about 600 nucleotides in length, or at least about 350 to about 500 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0155] In certain embodiments, a method for determining the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers, the set including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 1 (FR1) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene(s) of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting BCR amplicon molecules are then sequenced, and the sequences of the BCR amplicon molecules thus determined provide the sequence of the BCR repertoire in the sample. In certain embodiments, sequencing the BCR amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads represents at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the BCR. In additional embodiments, the method further includes sequence read clustering and BCR clonotype reporting. In some embodiments, the identified immune repertoire sequences are compared with a contemporaneous or most recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the average sequence read length is 300 to 600 nucleotides, or 350 to 550 nucleotides, or 330 to 425 nucleotides, or about 350 to about 425 nucleotides, depending in part on the inclusion of any barcode sequences in the read length.In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining regions CDR1, CDR2, and CDR3 of a BCR expressed sequence.
[0156] In some embodiments, the provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR1 region about 70 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR1 region about 50 nucleotides in length. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 18 to about 45 different FR1-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 22 to about 35 different FR1-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 25 to about 35 different FR1-directed primers. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 40 to about 65 different FR1-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 48 to about 60 different FR1-directed primers. In some embodiments, the target BCR primer set comprises one or more C gene primers. In certain embodiments, the target immune receptor primer set includes at least 5 to about 15 C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In certain embodiments, the target immune receptor primer set includes at least 2 to about 8 C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In some embodiments, the target BCR primer set includes two or more C gene primers directed to different Ig isotype molecules, e.g., IgA, IgD, IgG, IgM, and IgE. In some embodiments, the target BCR primer set includes at least five C gene primers, each directed to a C gene of a different Ig isotype molecule.
[0157] In certain embodiments, the methods of the present invention involve the use of at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Table 3 and Tables 6-10, respectively. In certain embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), comprising about 15 to about 35 primers selected from Table 3 and about 5 to about 20 primers selected from Tables 6-10, respectively. In some embodiments, the methods provided involve the use of at least one set of primers comprising i) about 22 to about 35 primers selected from Table 3 and ii) one or more primers selected from each of Tables 6-10. In certain embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), comprising about 40 to about 65 primers selected from Table 3 and about 5 to about 20 primers selected from Tables 6-10, respectively. In some embodiments, the methods provided involve the use of at least one set of primers comprising i) about 48 to about 60 primers selected from Table 3 and ii) one or more primers selected from each of Tables 6-10. In other specific embodiments, the methods of the invention comprise the use of at least one set of primers comprising i) a primer selected from SEQ ID NOs: 137-283, and ii) a primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In other embodiments, the methods provided comprise the use of at least one set of primers comprising i) a primer selected from SEQ ID NOs: 284-430, and ii) a primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.In some embodiments, the methods of the invention comprise the use of at least one set of primers comprising i) a primer selected from SEQ ID NOs: 137-283, and ii) a primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601, or comprising i) a primer selected from SEQ ID NOs: 284-430 and ii) a primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582.
[0158] In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 20 or at least 25 primers selected from SEQ ID NOs: 137-283, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include about 15 to about 35 primers selected from SEQ ID NOs: 137-283, and about 5 to about 15 primers selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers comprising i) about 22 to about 35 primers selected from SEQ ID NOs: 137-283, and ii) at least one primer selected from SEQ ID NOs: 448-459, at least one primer selected from SEQ ID NOs: 472-479, at least one primer selected from SEQ ID NOs: 488-513, at least one primer selected from SEQ ID NOs: 540-551, and at least one primer selected from SEQ ID NOs: 564-582. In other embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 284-430, and at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601. In some embodiments, the provided methods include the use of at least one set of primers i) and ii), including about 15 to about 35 primers selected from SEQ ID NOs: 284-430 and about 5 to about 15 primers selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.In some embodiments, provided methods involve the use of at least one set of primers comprising i) about 22 to about 35 primers selected from SEQ ID NOs: 284-430, and ii) at least one primer selected from SEQ ID NOs: 460-471, at least one primer selected from SEQ ID NOs: 480-487, at least one primer selected from SEQ ID NOs: 514-539, at least one primer selected from SEQ ID NOs: 552-563, and at least one primer selected from SEQ ID NOs: 583-601. In some embodiments, methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 284-430, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In other embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii), comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 137-283, and at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.
[0159] In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 40 or at least 50 primers selected from SEQ ID NOs: 137-283, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include about 40 to about 65 primers selected from SEQ ID NOs: 137-283, and about 5 to about 15 primers selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers including i) about 48 to about 60 primers selected from SEQ ID NOs: 137-283, and ii) at least one primer selected from SEQ ID NOs: 448-459, at least one primer selected from SEQ ID NOs: 472-479, at least one primer selected from SEQ ID NOs: 488-513, at least one primer selected from SEQ ID NOs: 540-551, and at least one primer selected from SEQ ID NOs: 564-582. In other embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), including at least 40 or at least 50 primers selected from SEQ ID NOs: 284-430, and at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601. In some embodiments, the provided methods include the use of at least one set of primers i) and ii), including about 40 to about 65 primers selected from SEQ ID NOs: 284-430 and about 5 to about 15 primers selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.In some embodiments, provided methods involve the use of at least one set of primers comprising i) about 48 to about 60 primers selected from SEQ ID NOs: 284-430, and ii) at least one primer selected from SEQ ID NOs: 460-471, at least one primer selected from SEQ ID NOs: 480-487, at least one primer selected from SEQ ID NOs: 514-539, at least one primer selected from SEQ ID NOs: 552-563, and at least one primer selected from SEQ ID NOs: 583-601. In some embodiments, methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 40 or at least 50 primers selected from SEQ ID NOs: 284-430, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In other embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii), comprising at least 40 or at least 50 primers selected from SEQ ID NOs: 137-283, and at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.
[0160] In certain embodiments, a method for amplifying expressed nucleic acid sequences of BCRs in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating immune receptor amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning about a 40 to about 60 nucleotide portion of the framework region. In some embodiments, the one or more V gene primers in i) anneal to at least a portion of the framework 3 region of the template molecule. In specific embodiments, the one or more C gene primers in ii) comprise at least two primers that anneal to at least a portion of the C gene portion of the BCR template molecule. In some embodiments, the one or more C gene primers in ii) comprise at least two primers that each anneal to at least a portion of a C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers in ii) comprise at least one primer, separately directed to a portion of each of the C genes of an IgA, IgD, IgG, IgM, and IgE template molecule.In certain embodiments, at least one set of generated amplicons comprises the complementarity determining region CDR3 of the BCR expressed sequence. In some embodiments, the amplicons are about 80 to about 200 nucleotides in length, about 80 to about 140 nucleotides in length, about 90 to about 130 nucleotides in length, or at least about 100 to about 120 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0161] In certain embodiments, a method for determining the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers, the set including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene(s) of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting BCR amplicon molecules are then sequenced, and the determined sequences of the BCR amplicon molecules provide the sequence of the BCR in the sample. In certain embodiments, determining the sequence of the BCR amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequence of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads represents at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the BCR. In additional embodiments, the method further includes sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified BCR repertoire are compared with a contemporaneous or most recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the average sequence read length is 80 to 185 nucleotides, or 115 to 200 nucleotides, or 90 to 130 nucleotides, or about 100 to about 120 nucleotides, depending in part on the inclusion of any barcode sequences in the read length.In certain embodiments, at least one set of sequenced amplicons comprises the complementarity determining region CDR3 of a BCR expressed sequence.
[0162] In certain embodiments, provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 70 nucleotides in length. In certain embodiments, provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 50 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 40 to about 60 nucleotides in length. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 50 to about 85 different FR3-directed primers. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 55 to about 80 different FR3-directed primers. In some embodiments, the target immune receptor primer set comprises V gene primers comprising about 62 to about 75 different FR3-directed primers. In some embodiments, the target BCR primer set comprises V gene primers, including about 65, 66, 67, 68, 69, or about 70 different FR3-directed primers. In some embodiments, the target BCR primer set comprises one or more C gene primers. In specific embodiments, the target immune receptor primer set comprises at least 5 to about 15 C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In specific embodiments, the target BCR primer set comprises at least 2 to about 8 C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In some embodiments, the one or more C gene primers in ii) comprise at least two primers, each annealing to at least a portion of a C gene of an IgA, IgD, IgG, IgM, or IgE template molecule.In some embodiments, the one or more C gene primers of ii) comprise at least one primer for a portion of a C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules separately.
[0163] In certain embodiments, the methods of the invention comprise the use of at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Table 2 and Tables 6-10, respectively. In certain embodiments, the methods of the invention comprise the use of at least one set of primers i) and ii), comprising about 55 to about 80 primers selected from Table 2 and about 5 to about 20 primers selected from Tables 6-10, respectively. In some embodiments, the methods provided comprise the use of at least one set of primers comprising i) about 62 to about 75 primers selected from Table 2, and ii) one or more primers selected from each of Tables 6-10. In other specific embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii) selected from SEQ ID NOs: 1-68, and 448-459, 472-479, 488-513, 540-551, and 564-582, or comprising primers selected from SEQ ID NOs: 69-136, and 460-471, 480-487, 514-539, 552-563, and 583-601. In some embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii) selected from SEQ ID NOs: 1-68, and 460-471, 480-487, 514-539, 552-563, and 583-601, or comprising primers selected from SEQ ID NOs: 69-136, and 448-459, 472-479, 488-513, 540-551, and 564-582.
[0164] In some embodiments, the methods of the invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-68, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-68, and about 5 to about 15 primers selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the provided methods involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-68, as well as at least one primer selected from SEQ ID NOs: 448-459, at least one primer selected from SEQ ID NOs: 472-479, at least one primer selected from SEQ ID NOs: 488-513, at least one primer selected from SEQ ID NOs: 540-551, and at least one primer selected from SEQ ID NOs: 564-582. In other embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-136, as well as at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601. In some embodiments, the provided methods include the use of at least one set of primers i) and ii), including at least 60 primers selected from SEQ ID NOs: 69-136 and about 5 to about 15 primers selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.In some embodiments, the provided methods involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 69-136, as well as at least one primer selected from SEQ ID NOs: 460-471, at least one primer selected from SEQ ID NOs: 480-487, at least one primer selected from SEQ ID NOs: 514-539, at least one primer selected from SEQ ID NOs: 552-563, and at least one primer selected from SEQ ID NOs: 583-601. In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-68, as well as at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601. In another embodiment, the method of the present invention comprises the use of at least one set of primers i) and ii), comprising at least 60 primers selected from SEQ ID NOs: 69-136, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582.
[0165] In certain embodiments, a method for amplifying expressed nucleic acid sequences of BCRs in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a V gene portion using at least one set of: i) a plurality of V gene primers directed to a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 2 (FR2) within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of each BCR coding sequence, wherein each set of primers in i) and ii) directed to the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In some embodiments, the one or more V gene primers in i) anneal to at least a portion of the FR2 region of the BCR template molecule. In specific embodiments, the one or more C gene primers in ii) comprise at least two primers that anneal to at least a portion of the constant portion C gene portion of the BCR template molecule. In some embodiments, the one or more C gene primers in ii) comprise at least two primers, each annealing to at least a portion of a C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers in ii) separately comprise at least one primer for a portion of each C gene of an IgA, IgD, IgG, IgM, and IgE template molecule. In specific embodiments, at least one set of generated amplicons comprises the complementarity-determining regions CDR2 and CDR3 of a BCR expression sequence.In some embodiments, the amplicon is about 180 to about 375 nucleotides in length, about 200 to about 350 nucleotides, about 225 to about 325 nucleotides, or about 250 to about 300 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0166] In certain embodiments, a method for determining the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers, including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of FR2 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting BCR amplicon molecules are then sequenced, and the sequences of the BCR amplicon molecules thus determined provide the sequence of the BCR repertoire in the sample. In certain embodiments, determining the sequence of the BCR amplicon molecules includes obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequence of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads represents at least 40%, at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the BCR. In additional embodiments, the method further includes sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared with a contemporaneous or most recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the average sequence read length is about 200 to about 375 nucleotides, about 250 to about 350 nucleotides, or about 275 to about 350 nucleotides, depending in part on the inclusion of any barcode sequences in the read length.In certain embodiments, at least one set of sequenced amplicons comprises the complementarity determining regions CDR2 and CDR3 of the BCR expressed sequences.
[0167] In certain embodiments, the provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR2 region about 70 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR2 region about 50 nucleotides in length. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 4 to about 20 different FR2-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 5 to about 15 different FR2-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 5, 6, 7, 8, 9, 10, 11, or 12 different FR2-directed primers. In some embodiments, the target BCR primer set comprises one or more C gene primers. In certain embodiments, the target immune receptor primer set comprises at least 5 to about 15 C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In certain embodiments, the target BCR primer set includes at least two to about eight C gene primers, each directed to at least a portion of the same 50-nucleotide region within each of the target C genes. In some embodiments, the one or more C gene primers in ii) include at least two primers that each anneal to at least a portion of a C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers in ii) include at least one primer that is separately directed to a portion of a C gene of each of an IgA, IgD, IgG, IgM, and IgE template molecule.
[0168] In certain embodiments, the methods of the invention comprise the use of at least one set of primers comprising V gene primer i) and C gene primer ii) selected from Tables 4 and 6-10, respectively. In certain other embodiments, the methods of the invention comprise the use of at least one set of primers i) and ii), comprising primers selected from SEQ ID NOs: 431-437, and 448-459, 472-479, 488-513, 540-551, and 564-582. In other embodiments, the methods of the invention comprise primers selected from SEQ ID NOs: 431-437, and 460-471, 480-487, 514-539, 552-563, and 583-601. In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least five primers selected from SEQ ID NOs: 431-437, and at least one primer selected from SEQ ID NOs: 448-459, 472-479, 488-513, 540-551, and 564-582. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include at least five primers selected from SEQ ID NOs: 431-437, and at least one primer selected from SEQ ID NOs: 448-459, at least one primer selected from SEQ ID NOs: 472-479, at least one primer selected from SEQ ID NOs: 488-513, at least one primer selected from SEQ ID NOs: 540-551, and at least one primer selected from SEQ ID NOs: 564-582. In other embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii), comprising at least five primers selected from SEQ ID NOs: 431-437, and at least one primer selected from SEQ ID NOs: 460-471, 480-487, 514-539, 552-563, and 583-601.In some embodiments, the provided methods involve the use of at least one set of primers i) and ii), including at least five primers selected from SEQ ID NOs: 431-437, as well as at least one primer selected from SEQ ID NOs: 460-471, at least one primer selected from SEQ ID NOs: 480-487, at least one primer selected from SEQ ID NOs: 514-539, at least one primer selected from SEQ ID NOs: 552-563, and at least one primer selected from SEQ ID NOs: 583-601.
[0169] In certain embodiments, a method for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of a BCR coding sequence, including at least a portion of a framework region within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In certain embodiments, the one or more J gene primers in ii) are directed to a sequence spanning about a 50 nucleotide portion of the J gene. In more particular embodiments, the one or more J gene primers in ii) are directed to a sequence spanning about a 30 nucleotide portion of the J gene. In certain embodiments, the one or more J gene primers in ii) are directed to a sequence entirely within the J gene.
[0170] In certain embodiments, a method for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating BCR amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning about a 40 to about 60 nucleotide portion of the framework region. In some embodiments, the one or more V gene primers in i) anneal to at least a portion of the framework 3 region of the template molecule. In specific embodiments, the J gene primers in ii) include at least two primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) include at least two to about eight primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) include about four primers that anneal to at least a portion of the J gene portion of the template molecule.In some embodiments, the plurality of J gene primers in ii) comprises about three to about six primers that anneal to at least a portion of the J gene portion of the template molecule. In certain embodiments, at least one set of generated amplicons comprises the complementarity-determining region CDR3 of a BCR expressed sequence. In some embodiments, the amplicons are about 60 to about 160 nucleotides in length, about 70 to about 100 nucleotides in length, about 100 to about 120 nucleotides in length, at least about 70 to about 90 nucleotides in length, about 80 to about 90 nucleotides in length, or about 80 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0171] In certain embodiments, a method for determining the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers, the set including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 3 (FR3) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting BCR amplicon molecules are then sequenced, and the sequences of the determined immune receptor amplicon molecules provide the sequence of the BCR repertoire in the sample. In some embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting immune receptor molecules. In certain embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, adding predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the BCR. In additional embodiments, the method further comprises sequence read clustering and BCR clonotype reporting.In some embodiments, the sequences of the identified BCR repertoire are compared to a contemporaneous or more recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the sequence read length is about 60 to about 185 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In some embodiments, the average sequence read length is 90 to 120 nucleotides, 70 to 90 nucleotides, about 75 to about 85 nucleotides, or about 80 nucleotides. In certain embodiments, at least one set of sequenced amplicons comprises the complementarity determining region (CDR3) of a BCR-expressed sequence.
[0172] In certain embodiments, the provided methods utilize a targeted BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 50 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 70 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR3 region about 40 to about 60 nucleotides in length. In certain embodiments, the targeted BCR primer set comprises V gene primers comprising about 50 to about 85 different FR3-directed primers. In certain embodiments, the targeted BCR primer set comprises V gene primers comprising about 55 to about 80 different FR3-directed primers. In some embodiments, the targeted immune receptor primer set comprises V gene primers comprising about 62 to about 75 different FR3-directed primers. In some embodiments, the targeted BCR primer set comprises V gene primers comprising about 65, 66, 67, 68, 69, or about 70 different FR3-directed primers. In some embodiments, the targeted BCR primer set comprises multiple J gene primers. In some embodiments, the targeted BCR primer set comprises at least two J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises two to about eight J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about three to about six different J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about two, three, four, five, six, seven, or eight different J gene primers. In certain embodiments, the targeted immune receptor primer set comprises about four J gene primers, each directed to at least a portion of a J gene in the target polynucleotide.
[0173] In certain embodiments, the methods of the invention comprise the use of at least one set of primers comprising V gene primer i) and J gene primer ii) selected from Tables 2 and 5, respectively. In certain other embodiments, the methods of the invention comprise the use of at least one set of primers i) and ii), comprising primers selected from SEQ ID NOs: 1-68 and 438-442, or selected from SEQ ID NOs: 69-136 and 443-447. In certain other embodiments, the methods of the invention comprise the use of at least one set of primers i) and ii), comprising primers selected from SEQ ID NOs: 1-68 and 443-447, or selected from SEQ ID NOs: 69-136 and 438-442.
[0174] In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 1-68, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 69-136, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 60 primers selected from SEQ ID NOs: 69-136, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, the methods of the invention involve the use of at least one set of primers i) and ii), including at least 60 primers selected from SEQ ID NOs: 1-68, and at least two primers, at least three primers, or at least four primers selected from SEQ ID NOs: 443-447.
[0175] In certain embodiments, a method for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating BCR amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In some embodiments, the one or more V gene primers in i) anneal to at least a portion of the framework 1 region of the template molecule. In certain embodiments, the J gene primers in ii) include at least two primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) include at least two to about eight primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) include about four primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) include about three to about six primers that anneal to at least a portion of the J gene portion of the template molecule.In certain embodiments, at least one set of generated amplicons comprises complementarity determining regions CDR1, CDR2, and CDR3 of the BCR expressed sequence. In some embodiments, the amplicons are about 220 to about 350 nucleotides in length, about 225 to about 300 nucleotides in length, about 250 to about 325 nucleotides in length, about 250 to about 270 nucleotides in length, or about 270 to about 300 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0176] In certain embodiments, a method for determining the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 1 (FR1) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting immune receptor amplicon molecules are then sequenced, and the sequences of the BCR amplicon molecules determined thereby provide the sequence of the BCR repertoire in the sample. In some embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting immune receptor molecules. In certain embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, adding predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the immune receptor. In additional embodiments, the method further comprises sequence read clustering and BCR clonotype reporting.In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or more recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the average sequence read length is 200-350 nucleotides, 225-325 nucleotides, 250-300 nucleotides, 270-300 nucleotides, or about 295 to about 325 nucleotides, depending in part on the inclusion of any barcode sequences. In certain embodiments, at least one set of sequenced amplicons comprises the complementarity determining regions CDR1, CDR2, and CDR3 of the BCR-expressing sequences.
[0177] In certain embodiments, the provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR1 region about 70 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR1 region about 80 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR1 region about 50 nucleotides in length. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 18 to about 45 different FR1-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 22 to about 35 different FR1-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 25 to about 35 different FR1-directed primers. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 40 to about 65 different FR1-directed primers. In some embodiments, the targeted BCR primer set comprises V gene primers, including about 48 to about 60 different FR1-directed primers. In some embodiments, the targeted BCR primer set comprises a plurality of J gene primers. In some embodiments, the targeted BCR primer set comprises at least two J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises two to about eight J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about three to about six different J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about 2, 3, 4, 5, 6, 7, or 8 different J gene primers.In certain embodiments, the target immune receptor primer set comprises about four J gene primers, each directed to at least a portion of a J gene within the target polynucleotide.
[0178] In certain embodiments, the methods of the invention involve the use of at least one set of primers comprising V gene primer i) and J gene primer ii) selected from Tables 3 and 5, respectively. In certain other embodiments, the methods of the invention involve the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 137-283 and 438-442, or selected from SEQ ID NOs: 284-430 and 443-447. In other embodiments, the methods of the invention involve the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 137-283 and 443-447, or selected from SEQ ID NOs: 284-430 and 438-442. In some embodiments, the methods of the invention involve the use of at least one set of primers i) and ii) comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 137-283 and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, provided methods involve the use of at least one set of primers i) and ii), comprising about 15-35 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, provided methods involve the use of at least one set of primers i) and ii), comprising about 22-35 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In other embodiments, methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447.In some embodiments, provided methods involve the use of at least one set of primers i) and ii), comprising about 15-35 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In some embodiments, provided methods involve the use of at least one set of primers i) and ii), comprising about 22-35 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In some embodiments, methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In other embodiments, the methods of the present invention comprise the use of at least one set of primers i) and ii), comprising at least 20, or at least 25, primers selected from SEQ ID NOs: 284-430, and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 438-442.
[0179] In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which include at least 40 or at least 50 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include approximately 40-65 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), which include approximately 48-60 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In other embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), comprising at least 40 or at least 50 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), comprising approximately 40-65 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In some embodiments, the methods provided involve the use of at least one set of primers i) and ii), comprising approximately 48-60 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447.In some embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which includes at least 40 or at least 50 primers selected from SEQ ID NOs: 137-283, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447. In other embodiments, the methods of the present invention involve the use of at least one set of primers i) and ii), which includes at least 40 or at least 50 primers selected from SEQ ID NOs: 284-430, and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442.
[0180] In certain embodiments, a method for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample is provided, comprising performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework 2 (FR2) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and performing amplification using each set results in amplicons representing the entire repertoire of each immune receptor in the sample, thereby generating amplicons comprising the BCR repertoire. In certain embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 80-nucleotide portion of the framework region. In more specific embodiments, the one or more V gene primers in i) are directed to sequences spanning an approximately 50-nucleotide portion of the framework region. In some embodiments, the one or more V gene primers in i) anneal to at least a portion of the FR2 region of the template molecule. In certain embodiments, the J gene primers in ii) comprise at least 10 primers that anneal to at least a portion of the J gene of the template molecule. In some embodiments, the J gene primers in ii) comprise about 14 primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) comprise at least two primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the J gene primers in ii) comprise at least two to about eight primers that anneal to at least a portion of the J gene portion of the template molecule.In some embodiments, the plurality of J gene primers in ii) comprises about four primers that anneal to at least a portion of the J gene portion of the template molecule. In some embodiments, the plurality of J gene primers in ii) comprises about three to about six primers that anneal to at least a portion of the J gene portion of the template molecule. In certain embodiments, at least one set of generated amplicons comprises the complementarity-determining regions CDR2 and CDR3 of the BCR gene sequence. In some embodiments, the amplicons are about 160 to about 270 nucleotides, about 180 to about 250 nucleotides, or about 195 to about 225 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0181] In certain embodiments, a method for providing the sequence of a BCR repertoire in a sample is provided, comprising: performing a multiplex amplification reaction to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence including at least a portion of FR2 within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target immune receptor coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The resulting immune receptor amplicon molecules are then sequenced, and the sequences of the BCR amplicon molecules determined thereby provide the sequence of the BCR repertoire in the sample. In some embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting immune receptor molecules. In certain embodiments, sequencing the BCR amplicon molecules comprises obtaining initial sequence reads, adding predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 40%, at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads in the BCR. In additional embodiments, the method further comprises sequence read clustering and BCR clonotype reporting.In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or more recent version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, the average sequence read length is 160-300 nucleotides, 180-280 nucleotides, 200-260 nucleotides, or about 225 to about 270 nucleotides, depending in part on the inclusion of any barcode sequences. In certain embodiments, at least one set of sequenced amplicons comprises the complementarity determining regions CDR2 and CDR3 of the BCR-expressing sequences.
[0182] In certain embodiments, the provided methods utilize a target BCR primer set comprising V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence spanning an FR2 region about 70 nucleotides in length. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence spanning an FR2 region about 50 nucleotides in length. In certain embodiments, the target BCR primer set comprises V gene primers comprising about 4 to about 20 different FR2-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 5 to about 15 different FR2-directed primers. In some embodiments, the target BCR primer set comprises V gene primers comprising about 5, 6, 7, 8, 9, 10, 11, or 12 different FR2-directed primers. In some embodiments, the target BCR primer set comprises a plurality of J gene primers. In some embodiments, the target BCR primer set comprises at least two J gene primers, each directed to at least a portion of a J gene within the target polynucleotide. In some embodiments, the targeted BCR primer set comprises two to about eight J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about three to about six different J gene primers, each directed to at least a portion of a J gene in the target polynucleotide. In some embodiments, the targeted BCR primer set comprises about two, three, four, five, six, seven, or eight different J gene primers. In certain embodiments, the targeted immune receptor primer set comprises about four J gene primers, each directed to at least a portion of a J gene in the target polynucleotide.
[0183] In certain embodiments, the methods of the invention involve the use of at least one set of primers comprising V gene primer i) and J gene primer ii) selected from Tables 4 and 5, respectively. In certain other embodiments, the methods of the invention involve the use of at least one set of primers i) and ii) comprising primers selected from SEQ ID NOs: 431-437 and 438-442, or selected from SEQ ID NOs: 431-437 and 443-447. In some embodiments, the methods of the invention involve the use of at least one set of primers i) and ii), comprising at least five primers selected from SEQ ID NOs: 431-437 and at least two, at least three, or at least four primers selected from SEQ ID NOs: 438-442. In other embodiments, the methods of the invention involve the use of at least one set of primers i) and ii), comprising at least five primers selected from SEQ ID NOs: 431-437 and at least two, at least three, or at least four primers selected from SEQ ID NOs: 443-447.
[0184] In certain embodiments, the methods of the present invention involve the use of a biological sample selected from the group consisting of hematopoietic cells, lymphocytes, and tumor cells. In some embodiments, the biological sample is selected from the group consisting of peripheral blood mononuclear cells (PBMCs), T cells, B cells, circulating tumor cells, and tumor-infiltrating lymphocytes (herein "TILs" or "TILs"). In some embodiments, the biological sample contains B cells that have undergone ex vivo activation and / or proliferation. In some embodiments, the biological sample contains cfDNA, for example, as found in blood or plasma. In some embodiments, the biological sample is selected from the group consisting of tissue (e.g., lymph nodes, organ tissue, bone marrow), whole blood, synovial fluid, cerebrospinal fluid, tumor biopsy, and other clinical specimens containing cells.
[0185] In some embodiments, methods, compositions, and systems are provided for determining the immune repertoire of a biological sample by assessing both expressed immune receptor RNA and rearranged immune receptor genomic DNA (gDNA) from the biological sample. In some embodiments, the sample RNA and gDNA can be assessed simultaneously, forming cDNA by reverse transcription of the RNA, and the cDNA and gDNA can be amplified in the same multiplex amplification reaction. In some embodiments, the cDNA from the sample RNA and the sample gDNA can undergo multiplex amplification in separate reactions. In some embodiments, the cDNA from the sample RNA and the sample gDNA can undergo multiplex amplification using parallel primer pools. In some embodiments, the same BCR-directed primer pool is used to assess the BCR repertoire of the gDNA and RNA from the sample. In some embodiments, different immune receptor-directed primer pools are used to assess the immune repertoire of the gDNA and RNA from the sample. In some embodiments, multiplex amplification reactions are performed separately using cDNA from the sample RNA and sample gDNA to amplify the same or different target immune receptor molecules from the sample, and the resulting immune receptor amplicons are sequenced, thereby providing the sequences of the expressed immune receptor RNA and rearranged immune receptor gDNA of the biological sample.
[0186] In some embodiments, different immune receptor-directed primer pools are used to assess the immune repertoire of gDNA and / or RNA from the sample. In some embodiments, a multiplex amplification reaction is performed using the IgH primer set provided herein and a TCR beta-directed primer set as described, for example, in PCT Application Nos. PCT / US2018 / 014111, filed January 17, 2018, and PCT Application No. PCT / US2018 / 049259, filed August 31, 2018, each of which is incorporated herein by reference in its entirety, or a TCR beta-directed primer set commercially available as Oncomine™ TCR beta-SR Assay DNA, Oncomine™ TCR beta-SR Assay RNA, and Oncomine™ TCR beta-LR Assay (Thermo Fisher Scientific). The ability to assess both BCR (e.g., IgH) and TCR (e.g., TCR beta) repertoires from a sample using a single multiplex amplification reaction is useful for saving time and limited biological samples and is applicable to many of the methods described herein, including those related to allergy and autoimmunity, vaccine development and use, and immuno-oncology. For example, B cell repertoire analysis may be used in combination with T cell repertoire analysis to improve detection of changes in the immune repertoire following administration of immunotherapy, such as checkpoint blockade or checkpoint inhibitor immunotherapy, potentially indicating a response to immunotherapy. B cell repertoire analysis may also be used in combination with T cell repertoire analysis to improve assessment of vaccine efficacy. Exemplary immune repertoire changes in response to immunotherapy or vaccine administration include, but are not limited to, a decrease in T cell and B cell homogeneity after treatment (e.g., but not limited to, 7-14 days after treatment) compared to pre-treatment homogeneity values, and an increase in the presence of IgG1-expressing B cells after treatment(s) compared to pre-treatment values.
[0187] In some embodiments, the provided methods and compositions are used to identify and / or characterize a subject's immune repertoire. In some embodiments, the provided methods and compositions are used to identify and characterize novel or non-canonical BCR alleles in a subject's immune repertoire. In some embodiments, the sequences of the identified immune repertoire are compared with a contemporaneous or latest version of the IMGT database to identify the sequence of at least one allelic variant not present in the IMGT database. In some embodiments, identified allelic variants not present in the IMGT database are subjected to evidence-based screening using criteria such as clone number support, sequence read support, and / or the number of individuals carrying the allelic variant. Allelic variants identified and reported as not present in IMGT can be compared with other databases containing immune repertoire sequence information, such as the NCBI NR database and the Lym1K database, to cross-validate the reported novel or non-canonical BCR allele. Characterizing the presence of unreported or non-canonical IgH polymorphisms can be useful, for example, in understanding factors affecting autoimmune diseases, infectious diseases, and response to immunotherapy. In some embodiments, the sequences of the novel or non-standard BCR alleles identified as described herein may be used to generate recombinant BCR nucleic acids or molecules. Thus, in other embodiments, methods are provided for producing recombinant nucleic acids encoding the identified novel IgH allelic variants. In some embodiments, methods are provided for producing recombinant IgH allelic variant molecules and for producing recombinant cells that express them.
[0188] In some embodiments, the provided methods and compositions are used to identify and characterize novel or non-canonical BCR alleles in a subject's immune repertoire. In some embodiments, a patient's immune repertoire can be identified or characterized before and / or after a therapeutic treatment, e.g., treatment for cancer or an immune disorder. In some embodiments, the identification or characterization of the immune repertoire can be used to evaluate the effectiveness or efficacy of a treatment to modify a treatment regimen and / or optimize the selection of a therapeutic agent. In some embodiments, the identification or characterization of the immune repertoire can be used to evaluate a patient's response to immunotherapy, cancer vaccines, and / or other immune-based therapies or combination(s) thereof. In some embodiments, the identification or characterization of the immune repertoire can indicate the patient's likelihood of responding to a therapeutic agent or the patient's likelihood of not responding to a therapeutic agent.
[0189] In some embodiments, a patient's BCR repertoire may be identified or characterized to monitor the progression and / or treatment of a hyperproliferative disease, including detecting residual disease after treatment of the patient, to monitor the progression and / or treatment of an autoimmune disease, to monitor transplantation, and to monitor antigenic stimulation conditions, including post-vaccination, exposure to bacterial, fungal, parasitic, or viral antigens, or infection by bacteria, fungi, parasites, or viruses. In some embodiments, identification or characterization of the BCR repertoire may be used to assess a patient's response to anti-infective or anti-inflammatory therapy.
[0190] In some embodiments, methods and compositions are provided for identifying and / or characterizing immune repertoire clonal populations in a sample from a subject, comprising performing one or more multiplex amplification reactions using the sample or cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having constant and variable portions using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the immune repertoire coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying one or more immune repertoire clonal populations for the target BCRs from the sample. In particular, embodiments of sequencing immune repertoire amplicon molecules include obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting immune receptor molecules. In other embodiments of such methods and compositions, one or more multiplex amplification reactions are performed using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.In other embodiments of such methods and compositions, the one or more multiplex amplification reactions are performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene(s) of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0191] In some embodiments, methods and compositions are provided for identifying and / or characterizing immune repertoire clonal populations in a sample from a subject, comprising performing one or more multiplex amplification reactions using the sample or cDNA prepared from the sample to amplify immune repertoire nucleic acid template molecules having J and V gene portions using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers for a majority of different J genes of each BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence are selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying one or more immune repertoire clonal populations for the target BCRs from the sample. In certain embodiments, determining the sequence of the immune receptor amplicon molecule comprises obtaining initial sequence reads, appending predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequence of the resulting immune receptor molecule. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence that comprises at least a portion of framework region 2 (FR2) within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0192] Thus, in some embodiments, the provided methods, compositions, and workflows are for use in assessing the clonality, diversity, and abundance of B cell populations, including, but not limited to, those described herein. For example, clonal expansion can identify B cells responding to antigen challenge, and longitudinal analysis can be used to evaluate the effectiveness of vaccination. In some embodiments, the provided methods, compositions, and workflows are for use in identifying clonal lineages with many members. For example, clonal lineages with many members can represent B cells responding to long-term antigen stimulation. In some embodiments, the provided methods, compositions, and workflows are for use in identifying antigen-specific B cells. For example, comparing IgH repertoires across groups of individuals exposed to the same antigen can reveal shared IgH amino acid motifs indicative of antigen-specific IgH chains. In some embodiments, the provided methods, compositions, and workflows are for use in assessing clonal overlap. For example, clonal overlap analysis can reveal B cell trafficking and developmental relationships between B cell populations. In some embodiments, the provided methods, compositions, and workflows are for use in determining the VDJ sequences of dominant clones included in longitudinal analysis. In some embodiments, the methods, compositions, and workflows provided are for use in identifying malignant subclones via clonal lineage analysis. For example, in some B-cell malignancies (e.g., follicular lymphoma), somatic hypermutation occurs, resulting in the existence of malignant subclones with distinct but related IgH sequences that can be tracked with the methods, compositions, and workflows provided.
[0193] In some embodiments, provided methods, compositions, and workflows are for use in assessing clonal evolution. For example, analysis of clonal lineages can reveal isotype switches and IgH residues important for antigen binding. In some embodiments, provided methods, compositions, and workflows are for use in assessing isotype abundance. For example, over- or under-representation of a particular isotype can indicate a disease or immune deficiency, such as, but not limited to, elevated IgG1 in response to a viral infection or elevated IgE in allergies, while missing or under-represented isotypes can indicate a primary immune deficiency. In some embodiments, provided methods, compositions, and workflows are for use in quantifying somatic hypermutation. For example, the frequency of somatic hypermutation provides insight into the stage of B-cell development at which malignant transformation occurs.
[0194] In some embodiments, the provided methods and compositions are used to identify and / or characterize somatic hypermutation (SHM) within a BCR repertoire or clonal population. In some embodiments, the provided methods and compositions are used to identify and / or screen for rare BCR clones or subclones, for example, those with somatically hypermutated VDJ rearrangements. In some embodiments, the identification, quantification, and / or characterization of rare BCR clones can provide biomarkers for a given condition or therapeutic response. Thus, in some embodiments, the methods and compositions provided herein are used to identify, screen, and / or characterize BCR clones as biomarkers, for example, using samples obtained from retrospective or longitudinal subject studies.
[0195] In some embodiments, a method for identifying and / or characterizing BCR clonal lineage and SHM includes performing one or more multiplex amplification reactions with a subject's sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; and performing VDJ sequence analysis as provided herein to identify and / or quantify SMH and clonal lineage in target BCRs from the sample. In other embodiments, a method for identifying and / or characterizing BCR clonal lineage and SHM includes performing one or more multiplex amplification reactions with a subject's sample to amplify a BCR nucleic acid template molecule having a J gene portion and a variable portion using primers for a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a set of multiple J gene primers for a majority of the different J genes of each target BCR coding sequence; sequencing the resulting BCR amplicons; and performing VDJ sequence analysis as provided herein to identify SHM and clonal lineage in target BCRs from the sample.
[0196] In some embodiments, provided methods and compositions are used to identify, quantify, characterize, and / or monitor isotype (or subisotype) class or isotype class switching within a BCR repertoire or B cell clonal lineage. In some embodiments, such methods include performing one or more multiplex amplification reactions with a subject's sample to amplify an IgH nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers for a majority of different IgH V genes, including at least a portion of FR1, FR2, or FR3 within the V gene, and one or more C gene primers for at least a portion of a C gene of the IgH coding sequence; sequencing the resulting amplicons; and performing a sequence analysis as provided herein to identify the IgH isotype class(es) of the BCR repertoire or clonal lineage of the sample. In some embodiments, the primer set includes one or more primers for at least a portion of a C gene of a single isotype, e.g., IgE. In other embodiments, the primer set includes at least two primers, each for at least a portion of a C gene of two different isotypes. In other embodiments, the primer set comprises at least one primer separately for at least a portion of the C gene of the IgA, IgD, IgG, IgM, and IgE isotype classes.
[0197] In certain embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations and clonal lineages, such as changes in clonal expansion, changes in clonal contraction, changes in the relative proportions of clones or clonal populations within the BCR repertoire, changes in clonal lineage expansion or contraction, changes in somatic hypermutation, and / or isotype class switching within the repertoire. In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations or clonal lineages in response to tumor growth (e.g., clonal population or lineage expansion, clonal population or lineage contraction, changes in clonal populations or lineages in relative proportions, changes in somatic hypermutation and / or class switching). In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations or clonal lineages in response to tumor therapy (e.g., clonal population or lineage expansion, clonal population or lineage contraction, changes in clonal populations or lineages in relative proportions, changes in somatic hypermutation and / or class switching). In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations or clonal lineages during remission (e.g., clonal population or lineage expansion, clonal population or lineage contraction, clonal population or lineage shifts in relative proportions, changes in somatic hypermutation and / or class switching). In many lymphoid malignancies, clonal B cell receptor sequences can be used as biomarkers in malignant cells of certain cancers (e.g., leukemia) to monitor residual disease, tumor expansion, contraction, and / or treatment response. In certain embodiments, clonal B cell receptors can be identified and further characterized to identify novel utilities for therapeutic, biomarker, and / or diagnostic uses.
[0198] In some embodiments, methods and compositions are provided for monitoring changes in BCR clonal populations in a subject, comprising performing one or more multiplex amplification reactions with a sample from the subject to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying immune repertoire clonal populations in the target BCRs from the sample; and comparing the identified BCR repertoire clonal populations to those identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for monitoring changes in BCR clonal populations in a subject, comprising performing one or more multiplex amplification reactions with a sample from the subject to amplify an immune repertoire nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of the different J genes of each target BCR coding sequence, sequencing the resulting BCR amplicons, identifying immune repertoire clonal populations in the target BCR from the sample, and comparing the identified BCR repertoire clonal populations with those identified in samples obtained from the subject at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction, or can be two or more multiplex amplification reactions performed in parallel, e.g., parallel highly multiplexed amplification reactions performed using different primer pools.Samples for use in monitoring changes in BCR repertoire clonal populations include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained any time before treatment (pre-treatment sample), samples obtained any time after completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0199] In certain embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the progression and / or treatment of the patient's hyperproliferative disease. In some embodiments, the provided methods and compositions are used for minimal residual disease (MRD) monitoring of patients after treatment. In some embodiments, the provided methods and compositions enable deep sequencing of a patient's BCR repertoire, useful for MRD measurement and identifying rare BCR clones. In some embodiments, monitoring MRD includes assessing somatic hypermutation of the BCR repertoire. In some embodiments, the methods and compositions are used to identify and / or track B-cell lineage malignancies or T-cell lineage malignancies. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients diagnosed with leukemia or lymphoma, including, but not limited to, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, chronic myelogenous leukemia, cutaneous T-cell lymphoma, B-cell lymphoma, mantle cell lymphoma, and multiple myeloma. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients diagnosed with solid tumors, including, but not limited to, breast cancer, lung cancer, colorectal cancer, and neuroblastoma. In some embodiments, the methods and compositions are used to detect and / or monitor MRD in patients after cancer treatment, including, but not limited to, bone marrow transplant, lymphocyte infusion, adoptive T cell therapy, other cell-based immunotherapies, and antibody-based immunotherapies.
[0200] In some embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the patient's prognosis and / or treatment of a hyperproliferative disease, comprising performing one or more multiplex amplification reactions using a sample from the patient or cDNA prepared from the sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 1 (FR1) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying the immune repertoire for the target BCR from the sample. In particular, embodiments of determining the sequence of immune repertoire amplicon molecules include obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequence of the resulting immune receptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of an FR3 within the V gene, and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of FR2 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0201] In some embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the patient's prognosis and / or treatment of a hyperproliferative disease, comprising performing one or more multiplex amplification reactions using a sample from the patient or cDNA prepared from the sample to amplify an immune repertoire nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers including: i) a plurality of V gene primers for a majority of the distinct V genes of at least one BCR coding sequence that includes at least a portion of framework region 3 (FR3) within the V gene; and ii) a plurality of J gene primers for a majority of the distinct J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence are selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying the immune repertoire for the target BCR from the sample. In certain embodiments, determining the sequence of the immune receptor amplicon molecule comprises obtaining initial sequence reads, appending predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and determining the sequence of the resulting immune receptor molecule. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of FR2 within the V gene; and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0202] In some embodiments, methods and compositions are provided for MRD monitoring in patients with hyperproliferative diseases, comprising performing one or more multiplex amplification reactions with a patient sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying immune repertoire sequences in the target BCRs; and detecting the presence or absence of the BCR sequence(s) in the sample associated with the hyperproliferative disease. In some embodiments, methods and compositions are provided for MRD monitoring in patients with hyperproliferative diseases, comprising performing one or more multiplex amplification reactions with a patient sample to amplify an immune repertoire nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of the different J genes of each target BCR coding sequence, sequencing the resulting BCR amplicons, identifying the immune repertoire sequence in the target BCR, and detecting the presence or absence of immune receptor sequence(s) in the sample associated with the hyperproliferative disease. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, e.g., parallel highly multiplexed amplification reactions performed using different primer pools. Samples for use in MRD monitoring include, but are not limited to, samples obtained during remission, samples obtained any time after completion of treatment (post-treatment samples), and samples obtained during the course of treatment.
[0203] In certain embodiments, methods and compositions are provided for identifying and / or characterizing the BCR repertoire of a subject that responds to therapy. In some embodiments, the methods and compositions are used to characterize and / or monitor tumor-infiltrating lymphocyte (TIL) populations or clones before, during, and / or after tumor therapy. In some embodiments, characterization and / or evaluation of the tumor microenvironment is provided by profiling the immune receptor repertoire of TILs. In some embodiments, the methods and compositions for determining immune repertoires are used to identify and / or track therapeutic T cell population(s) and B cell population(s). In some embodiments, the provided methods and compositions are used to identify and / or monitor the persistence of a cell-based therapy after patient treatment, including, but not limited to, the presence (e.g., persistent presence) of a CAR-T cell population, a TCR-engineered T cell population, persistent CAR-T expression, the presence (e.g., persistent presence) of an administered TIL population, TIL expression (e.g., persistent expression) after adoptive T cell therapy, and / or the presence (e.g., persistent presence) of an engineered T cell population, including immune reconstitution after allogeneic hematopoietic cell transplantation.
[0204] In some embodiments, the provided methods and compositions are used to characterize and / or monitor B cell clones or populations present in a patient sample after administration of a cell-based therapy to a patient, for example, but not limited to, cancer vaccine cells, CAR-T, TIL, and / or other engineered cell-based therapy. In some embodiments, the provided methods and compositions are used to characterize and / or monitor the BCR repertoire in a patient sample after cell-based therapy, and to evaluate and / or monitor the patient's response to the administered cell-based therapy. Samples for use in such characterization and / or monitoring after cell-based therapy include, but are not limited to, circulating blood cells, circulating tumor cells, TIL, tissue, cfDNA, and tumor sample(s) from the patient.
[0205] In some embodiments, methods and compositions are provided for monitoring cell-based therapies for patients receiving such therapies, comprising performing one or more multiplex amplification reactions with a patient sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying immune repertoire sequences in the target BCRs; and detecting the presence or absence of the BCR sequence(s) in the sample associated with the cell-based therapy. In some embodiments, methods and compositions are provided for monitoring cell-based therapies for patients receiving such therapies, comprising performing one or more multiplex amplification reactions with a patient sample to amplify an immune repertoire nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of the different J genes of each target BCR coding sequence; sequencing the resulting BCR amplicons; identifying the immune repertoire sequence in the target BCR; and detecting the presence or absence of the BCR sequence(s) in the sample associated with the cell-based therapy.
[0206] In some embodiments, methods and compositions are provided for monitoring a patient's response following administration of a cell-based therapy, comprising performing one or more multiplex amplification reactions with a patient's sample to amplify a BCR repertoire nucleic acid template molecule having a constant portion and a variable portion using at least one set of: i) primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying an immune repertoire sequence for the target BCR; and comparing the identified BCR repertoire to immune repertoire sequence(s) identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for monitoring a patient's response after administration of a cell-based therapy, comprising performing one or more multiplex amplification reactions with a patient sample to amplify a BCR repertoire nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for a majority of distinct V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of distinct J genes of each BCR coding sequence, sequencing the resulting BCR amplicons, identifying an immune repertoire sequence in the target BCR, and comparing the identified BCR repertoire with immune repertoire sequence(s) identified in samples obtained from the patient at different times. Cell-based therapies suitable for such monitoring include, but are not limited to, CAR-T cells, TCR-engineered T cells, TILs, and other enriched autologous cells. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction, or can be two or more multiplex amplification reactions performed in parallel, e.g., parallel highly multiplexed amplification reactions performed using different primer pools.Samples for use in such monitoring include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained any time before treatment (pre-treatment sample), samples obtained any time after completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0207] In some embodiments, the methods and compositions for determining B cell receptor repertoire, or B cell and T cell receptor repertoire, are used to measure and / or assess immune competence before, during, or after treatment, including but not limited to, organ transplant or bone marrow transplant.
[0208] In certain embodiments, the provided methods and compositions are used to identify and / or characterize the BCR repertoire of a subject that responds to a therapeutic treatment, including, but not limited to, immunotherapy, anti-allergy therapy, and anti-infective therapy. Thus, in some embodiments, the provided methods and compositions are used to identify BCR repertoires or clonal lineage biomarkers or signatures of a therapeutic response, such as a favorable response (e.g., successful vaccination) or an adverse response (e.g., immune system-mediated adverse event) to a therapeutic treatment. In some embodiments, methods and compositions are provided for identifying and / or characterizing a BCR repertoire in a subject who responds to treatment, comprising obtaining a sample from the subject after the start of treatment, and performing one or more multiplex amplification reactions using the sample or cDNA prepared from the sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers including: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence, the V gene primers including at least a portion of framework region 1 (FR1) within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method also includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying the immune repertoire for the target BCR from the sample. In some embodiments, the method further comprises comparing the BCR repertoire identified from a sample obtained after treatment initiation with the BCR repertoire from a patient sample obtained before treatment. In particular, embodiments for sequencing BCR amplicon molecules comprise obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting immune receptor molecules.In other embodiments of such methods and compositions, the multiplex amplification reaction is carried out using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR3 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene(s) of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, the multiplex amplification reaction is carried out using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR2 within the V gene; and ii) one or more C gene primers for at least a portion of each target C gene(s) of the BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0209] In some embodiments, methods and compositions are provided for identifying and / or characterizing a subject's BCR repertoire that responds to treatment, comprising obtaining a sample from the subject after the start of treatment, and performing one or more multiplex amplification reactions using the sample or cDNA prepared from the sample to amplify BCR nucleic acid template molecules having a J gene portion and a V gene portion using at least one set comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers for a pairwise majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method also includes sequencing the resulting BCR amplicon molecules, determining the sequence of the BCR amplicon molecules, and identifying the immune repertoire for the target BCR from the sample. In some embodiments, the method further comprises comparing the BCR repertoire identified from a sample obtained after treatment initiation with the BCR repertoire from a patient sample obtained before treatment. In particular, embodiments for sequencing BCR amplicon molecules include obtaining initial sequence reads, appending predicted J gene sequences to the sequence reads to create extended sequence reads, aligning the extended sequence reads to a reference sequence to identify productive reads, correcting one or more indel errors to generate rescued productive sequence reads, and sequencing the resulting BCR molecules.In other embodiments of such methods and compositions, the multiplex amplification reaction is carried out using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, the multiplex amplification reaction is carried out using at least one set of primers comprising: i) a plurality of V gene primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers for a majority of different J genes of each target BCR coding sequence, wherein each set of primers in i) and ii) for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0210] In some embodiments, methods and compositions are provided for monitoring changes in a subject's BCR repertoire in response to treatment, comprising performing one or more multiplex amplification reactions with a subject or patient sample to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers for most of the different V genes of at least one BCR coding sequence, including at least a portion of FR1, FR2, or FR3 within the V gene, and one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying immune repertoire sequences in the target BCRs from the sample; and comparing the identified BCR repertoire with that identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for monitoring changes in a subject's BCR repertoire in response to treatment, comprising performing one or more multiplex amplification reactions with a subject or patient sample to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for most of the different V genes of at least one BCR coding sequence, including at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for most of the different J genes of each target BCR coding sequence, sequencing the resulting BCR amplicons, identifying immune repertoire sequences in the target BCR from the sample, and comparing the identified BCR repertoire with that identified in samples obtained from the subject at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, e.g., parallel highly multiplexed amplification reactions performed using different primer pools.Samples for use in monitoring changes in the BCR repertoire include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained any time before treatment (pre-treatment sample), samples obtained any time after completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0211] In certain embodiments, the provided methods and compositions are used to characterize and / or monitor BCR repertoires associated with immune system-mediated adverse events, including, but not limited to, those associated with inflammatory conditions, autoimmune responses, and / or autoimmune diseases or disorders. In some embodiments, the provided methods and compositions are used to identify and / or monitor B cell, or B cell and T cell, immune repertoires associated with chronic autoimmune diseases or disorders, including, but not limited to, multiple sclerosis, type 1 diabetes, paralysis, rheumatoid arthritis, ankylosing spondylitis, asthma, and SLE. In some embodiments, a systemic sample, such as a blood sample, is used to determine the immune repertoire(s) of an individual with an autoimmune condition. In some embodiments, a local sample, such as a fluid sample from an affected joint or area of swelling, is used to determine the immune repertoire(s) of an individual with an autoimmune condition. In some embodiments, a comparison of the immune repertoire found in a localized or affected area sample with the immune repertoire found in a systemic sample can identify clonal T cell or B cell populations to target for elimination.
[0212] In some embodiments, methods and compositions are provided for identifying and / or monitoring a BCR repertoire associated with the progression and / or treatment of a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a patient sample to amplify a BCR repertoire nucleic acid template molecule having a constant portion and a variable portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying an immune repertoire sequence for the target BCR from the sample; and comparing the identified BCR repertoire to BCR repertoire sequence(s) identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for identifying and / or monitoring a BCR repertoire associated with the progression and / or treatment of a patient's immune system-mediated adverse event(s), comprising performing one or more multiplex amplification reactions with a patient sample to amplify a BCR nucleic acid template molecule having a J gene portion and a V gene portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of the different J genes of each target BCR coding sequence, sequencing the resulting BCR amplicons, identifying BCR sequences in the target immune repertoire from the sample, and comparing the identified BCR repertoire with BCR repertoire(s) identified in samples obtained from the patient at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, e.g., parallel highly multiplexed amplification reactions performed using different primer pools.Samples for use in monitoring changes in the immune repertoire associated with immune system-mediated adverse events include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained any time before treatment (pre-treatment sample), samples obtained any time after completion of treatment (post-treatment sample), and samples obtained during the course of treatment.
[0213] In some embodiments, the provided methods and compositions are used to characterize and / or monitor immune repertoires associated with passive immunity, including naturally acquired and artificially acquired passive immunotherapy. For example, the provided methods and compositions may be used to identify and / or monitor protective antibodies that provide passive immunity to a recipient after, but not limited to, the transfer of antibody-mediated immunity from the mother to the fetus during pregnancy or to the infant via breastfeeding, or via the administration of antibodies to the recipient. In another example, the provided methods and compositions may be used to identify and / or monitor B cell and / or T cell immune repertoires associated with the passive transfer of cell-mediated immunity to a recipient, such as the administration of mature circulating lymphocytes to a recipient histocompatible with the donor. In some embodiments, the provided methods and compositions are used to monitor the duration of passive immunity in a recipient.
[0214] In some embodiments, the provided methods and compositions are used to characterize and / or monitor immune repertoires associated with active immunization or vaccination therapy. For example, after exposure to a vaccine or infectious pathogen, the provided methods and compositions can be used to identify and / or monitor protective antibodies or clonal B cell populations, or clonal B cell and T cell populations, that can provide active immunity to the exposed individual. In some embodiments, the provided methods and compositions are used to monitor the duration of B cell clones, or B cell and T cell clones, that contribute to immunity in the exposed individual. In some embodiments, the provided methods and compositions are used to identify and / or monitor B cell and / or T cell immune repertoires associated with exposure to bacterial, fungal, parasitic, or viral antigens. In some embodiments, the provided methods and compositions are used to identify and / or monitor B cell and / or T cell immune repertoires associated with bacterial, fungal, parasitic, or viral infection. Thus, in some embodiments, the methods and compositions provided are for use in vaccine development, including but not limited to, identifying and / or characterizing one or more responses to a vaccine candidate, as well as evaluating one or more responses to a vaccine for quality or regulatory purposes.
[0215] In some embodiments, methods and compositions are provided for monitoring changes in BCR repertoires after exposure to a vaccine or infectious pathogen, comprising performing one or more multiplex amplification reactions with a sample from the exposed subject to amplify BCR repertoire nucleic acid template molecules having constant and variable portions using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and one or more C gene primers for at least a portion of each target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying immune repertoire sequences in the target BCRs from the sample; and comparing the identified BCR repertoire with BCR repertoire(s) identified in a sample obtained from the subject at a different time (e.g., before exposure or after the sample to be tested is obtained). In some embodiments, methods and compositions are provided for monitoring changes in BCR repertoires after vaccination or exposure to an infectious pathogen, comprising performing one or more multiplex amplification reactions with a sample from an exposed subject to amplify BCR repertoire nucleic acid template molecules having a J gene portion and a V gene portion using at least one set of primers for a majority of different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and a plurality of J gene primers for a majority of the different J genes of each target BCR coding sequence; sequencing the resulting BCR amplicons; identifying BCR sequences in the target immune repertoire from the sample; and comparing the identified BCR repertoire with BCR repertoire(s) identified in samples obtained from the patient at different times.In certain embodiments, methods and compositions are provided for monitoring changes in BCR repertoires after exposure to a vaccine or infectious pathogen, comprising: performing one or more multiplex amplification reactions using cDNA prepared from a sample from an exposed subject to amplify IgH nucleic acid template molecules having constant and variable portions using at least one set of primers, the set including: a plurality of V gene primers for a majority of different IgH V genes, including at least a portion of FR1, FR2, or FR3 within the V genes, and one or more C gene primers for at least a portion of an IgH C gene; sequencing the resulting BCR amplicons; identifying expressed IgH repertoire sequences from the sample containing repertoire isotype information; and comparing the identified IgH repertoire with IgH repertoire(s) identified in samples obtained from the subject at different times (e.g., before exposure or after the sample to be tested was obtained). In some embodiments, the primer set includes one or more primers for at leas...
Claims
1. A method for assessing the clinical response of a subject having a primary immunodeficiency disorder or an autoimmune disease to a therapy based on characterizing the subject's B-cell immune repertoire, comprising: a) performing a multiplex amplification reaction to amplify target B-cell immunoreceptor nucleic acid template molecules derived from a biological sample from a subject, wherein: The multiplex amplification reaction i) (a) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of said V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within said V genes; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of said at least one BCR coding sequence; each set of primers i) and ii) directed against a coding sequence of a target IgH BCR gene; wherein said amplifying step results in amplicon molecules representative of said target BCR repertoire in said sample, thereby producing target B cell immunoreceptor amplicon molecules comprising said target immune receptor repertoire. b) sequencing said target immune receptor repertoire amplicon; c) identifying immune receptor clones from said sequencing and identifying the level of somatic hypermutation (SHM) within variable gene segments among said immune receptor clones, wherein: The VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, determining the level of somatic hypermutation (SHM) within the variable gene portion of each clone by comparing the sequence of said variable genes to the sequence of the closest matching germline variable gene in commonly used public reference databases; and d) determining the subclasses of the B cell immunoreceptor clones in the sample and the clonal frequencies of non-switched IgM-expressing cells, non-switched IgD-expressing cells, switched IgG-expressing cells, switched IgA-expressing cells, and switched IgE-expressing cells in the sample; (i) a result in which the SHM frequency is below a frequency threshold for B cells expressing non-switched IgM or non-switched IgD in the sample, and the immune repertoire is dominated by B cells expressing switched isotypes IgG, IgA, or IgE at a high frequency in the sample, indicates a likely responder to chemotherapy; (ii) the SHM frequency is greater than a frequency threshold, and the immune repertoire is dominated by B cells expressing a high frequency of unswitched IgM or unswitched IgD in the sample, indicating a high likelihood of a non-responder to chemotherapy; and / or (iii) the SHM frequency is greater than a frequency threshold, a result in which the immune repertoire is dominated by B cells expressing a high frequency of non-switched IgM or non-switched IgD in the sample indicates a high likelihood of responding to immunotherapy.
2. Each of the plurality of V gene primers and / or the one or more C gene primers meets the following criteria: (1) comprising two or more modified nucleotides within the primer, at least one of which is near or at the end of the primer, and at least one of which is at or near the central nucleotide position of the primer; (2) the length is 15 to 40 bases in length; (3) T exceeding 60°C and up to 70°C m , (4) having low cross-reactivity to non-target sequences present in the sample; (5) at least the first four nucleotides (in the 3' to 5' direction) are non-complementary to any sequence in any other primer present in the same reaction; and (6) non-complementary to any contiguous stretch of at least five nucleotides within any other generated target amplicon.
3. 3. The method of claim 1 or 2, wherein each of the plurality of V gene primers and / or the one or more C gene primers comprises two or more modified nucleotides having a cleavable group selected from methylguanine, 8-oxo-guanosine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine.
4. the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule; 3. The method of claim 1, wherein the one or more C gene primers comprise at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
5. 5. The method of claim 4, wherein the generated target BCR amplicon molecule comprises complementarity determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence.
6. The method of claim 4, wherein at least one of the sets of i) and ii) is selected from the primers in Table 3 and Tables 6 to 10, respectively.
7. At least one set of i) and ii) is i)(c) and ii)(a); the plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule; 3. The method of claim 1, wherein the one or more C gene primers comprise at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
8. the generated target BCR amplicon molecule comprises a complementarity determining region CDR3 of the target BCR gene sequence; or The method of claim 7, wherein at least one of the sets of i) and ii) is selected from the primers in Table 2 and Tables 6 to 10, respectively.
9. 3. The method of claim 1 or 2, wherein the SHM frequency cutoff is 8%.
10. 10. The method of claim 1, further comprising, prior to step b), adding at least one adapter to at least one of said target immune receptor amplicon molecules, thereby generating a library of adapter-modified target immune receptor amplicon molecules.
11. 2. The method of claim 1, wherein the sequencing comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads.
12. 1. A method of assessing a subject having symptoms of an autoimmune disease or disorder as having a chronic variable immune deficiency disorder based on characterizing the subject's B-cell immune repertoire, comprising: a) performing a multiplex amplification reaction to amplify target B-cell immunoreceptor nucleic acid template molecules derived from a biological sample from a subject, wherein: The multiplex amplification reaction i) (a) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of said V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within said V genes; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of said at least one BCR coding sequence; each set of primers i) and ii) directed against a coding sequence of a target IgH BCR gene; wherein said amplifying step results in amplicon molecules representative of said target BCR repertoire in said sample, thereby producing target B cell immunoreceptor amplicon molecules comprising said target immune receptor repertoire. b) sequencing said target immune receptor repertoire amplicon; c) identifying immune receptor clones from said sequencing and identifying the level of somatic hypermutation (SHM) within variable gene segments among said immune receptor clones, wherein: The VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, determining the level of somatic hypermutation (SHM) within the variable gene portion of each clone by comparing the sequence of said variable genes to the sequence of the closest matching germline variable gene in commonly used public reference databases; and d) determining the class switch recombination frequency of said B cell immunoreceptor clone in said sample; wherein the SHM frequency is below a frequency threshold for switched isotypes, the frequency of class switch recombination (CSR) in the sample is below a frequency threshold, and the immune repertoire is dominated by B cells expressing non-switched IgM or non-switched IgD in the sample, a result indicating the subject has a chronic variable immune deficiency disorder.
13. 13. The method of claim 12, wherein each of the plurality of V gene primers and / or the one or more C gene primers has any one or more of the criteria defined in claims 2 to 11.
14. 14. The method of any one of claims 1, 2, 12, or 13, wherein the subject has rheumatoid arthritis.
15. the immunotherapy comprises a checkpoint blockade or a B cell depleting agent, and / or 3. The method of claim 1 or 2, wherein the chemotherapy comprises methotrexate.
16. 1. A method of assessing a subject having leukemia based on characterizing the subject's B-cell immune repertoire, comprising: a) performing a multiplex amplification reaction to amplify target B-cell immunoreceptor nucleic acid template molecules derived from a biological sample from a subject; The multiplex amplification reaction i) (a) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene; (b) a plurality of V gene primers for a majority of said V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 2 (FR2) within said V genes; or (c) a plurality of V gene primers for a majority of the V genes that differ in at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) one or more C gene primers directed to at least a portion of a C gene of said at least one BCR coding sequence; each set of primers i) and ii) directed against a coding sequence of a target IgH BCR gene; wherein said amplifying step results in amplicon molecules representative of said target BCR repertoire in said sample, thereby producing target B cell immunoreceptor amplicon molecules comprising said target immune receptor repertoire. b) sequencing said target immune receptor repertoire amplicon; c) identifying immune receptor clones from said sequencing and identifying the level of somatic hypermutation (SHM) within variable gene segments among said immune receptor clones, wherein: determining the level of somatic hypermutation (SHM) within the variable gene portion of each clone by comparing the sequence of said variable genes to the sequence of the closest matching germline variable gene in commonly used public reference databases; d) determining the ongoing SHM frequency and / or the ongoing class switch recombination (CSR) frequency, wherein: the ongoing SHM frequency is determined by detecting the presence of distinct subclones within VDJ region sequences compared to other clonal lineage members; The VDJ regions of clonal lineage immune receptor clones exhibiting SHM have similar nucleotide sequences, The ongoing CSR frequency is determined by detecting the presence of IgM or IgD, and at least one switched isotype (IgG, IgA, or IgE) or a combination of switched isotypes within the same lineage of the B cell immunoreceptor clone in the sample. e) classifying said repertoire clones according to the following subclasses: Class I: No ongoing CSR or SHM, no V gene SHM; Class II: No ongoing CSR or SHM, V gene SHM greater than 0 and below a mutation frequency threshold of 6%; Class III: no ongoing CSR or SHM, V gene SHM above a mutation frequency threshold of 6%; Class IV: including ongoing CSR and / or SHM; a result of the B cell immune repertoire subclass classification being Class I indicates that the subject has the worst-classified leukemia; a result that the B cell immune repertoire subclassification is Class II indicates that the subject has a poorly classified leukemia; a result that the B cell immune repertoire subclassification is Class III indicates that the subject has a well-classified leukemia; A result of the B cell immune repertoire subclass classification being Class IV indicates that the subject has a best-classified leukemia.
17. 17. The method of claim 16, wherein each of the plurality of V gene primers and / or the one or more C gene primers has any one or more of the criteria defined in claims 2 to 11.
18. 18. The method of claim 16 or 17, wherein the average Class III SHM frequency is 2%.
Citation Information
Patent Citations
Multiwindow control system
JP1988083789A
Sequence analysis of complex amplicons
JP2013524849A
Detection of isotype profiles as disease signatures
JP2014527831A
Monitoring of transformation from follicular lymphoma to diffuse large B-cell lymphoma using immunorepertory analysis.
JP2015504661A
Methods and compositions for detecting recombination and rearrangement events
JP2020507326A