Compositions and Methods for Immune Repertoire Sequencing
By using multiple primers to different regions of the immune receptor, the problem of difficult to efficiently analyze the immune cell receptor library in the prior art is solved, and high-resolution immune cell library analysis is achieved, which enhances the understanding and diagnostic ability of immune cell responses.
Patent Information
- Application Number
- CN201980058009.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-26
- Filing Date
- 2019-07-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-07-18
AI Technical Summary
The prior art is difficult to achieve high-resolution analysis of immune cell receptor libraries, especially when dealing with highly variable immune cell receptor sequences, there are problems of sequence errors and inefficient reflections.
A composition is provided, including a plurality of primers, different variable regions and target constant regions of the immune receptor coding sequence for amplifying and sequencing of the immunospoke library in a single-stream assay sample. The composition includes primers for the V gene, J gene and C gene, and amplifies the target immune receptor sequence through a multiple amplification reaction to generate amplicons representing the immune receptor pool in the sample.
This method can effectively analyze and analyze immune cell receptor sequences, improve the high-resolution analysis ability of the immune group library, reduce sequence errors, enhance the understanding of immune cell responses, and improve the ability to diagnose and treat design.
Smart Images

Figure CN112654720B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 62 / 839,505, filed Apr. 26, 2019, and U.S. Provisional Application No. 62 / 700,168, filed Jul. 18, 2018. The entire contents of each of the foregoing applications are incorporated herein by reference. BACKGROUND OF THE INVENTION
[0003] The adaptive immune response involves the selective response of B and T cells that recognize antigens. Immunoglobulin genes encoding the antigen receptors of antibodies (Abs, in B cells) and T cell receptors (TCRs, in T cells) include complex loci, where extensive receptor diversity is generated due to the recombination of the corresponding variable (V), diversity (D), and joining (J) gene segments, and subsequent somatic hypermutation events during early lymphoid differentiation. The recombination process occurs independently for each of the two subunit chains of each receptor, and subsequently, heterodimer pairing generates still greater combinatorial diversity. Calculations of the potential combinatorial and joining possibilities contributing to the human immune repertoire have estimated that the number of possibilities greatly exceeds the total number of peripheral B or T cells in an individual. See, e.g., Davis and Bjorkman (1988) Nature 334:395-402; Arstila et al. (1999) Science 286:958-961; van Dongen et al., Leukemia, Henderson et al. (Eds.) Philadelphia: W.B. Saunders Company, 2002, pp. 85-129.
[0004] For many years, extensive efforts have been made to improve the analysis of the immune repertoire at high resolution. Means for specifically detecting and monitoring the expanded clones of lymphocytes will provide important opportunities for characterizing and analyzing normal and pathogenic immune responses. Despite the efforts, effective high-resolution analysis presents challenges. Low-throughput techniques such as Sanger sequencing can provide resolution but are limited in providing an effective means for comprehensively capturing the entire immune repertoire. Advances in next-generation sequencing (NGS) have provided a way to capture the repertoire; however, due to the nature of the large number of related sequences and the introduction of sequence errors caused by the technology, it has proven difficult to efficiently and effectively reflect the true repertoire. Therefore, improved sequencing methods and workflows capable of resolving complex populations of highly variable immune cell receptor sequences are being developed. There is still a need for new methods for effectively analyzing large immune cell receptor repertoires to better understand immune cell responses, enhance diagnostic and treatment capabilities, and design new therapies. Summary of the Invention
[0005] In one aspect of the present invention, there is provided a composition for determining an immune repertoire in a single-stream assay sample. In some embodiments, the composition comprises at least one set of primers i) and ii), wherein i) consists of a plurality of variable (V) gene primers directed against most of the different variable regions of an immune receptor coding sequence; and ii) consists of one or more constant (C) gene primers directed against at least a portion of the corresponding target constant region of the corresponding immune receptor coding sequence. In some embodiments, the composition comprises at least one set of primers i) and ii), wherein i) consists of a plurality of V gene primers directed against most of the different V genes of an immune receptor coding sequence; and ii) consists of a plurality of J gene primers directed against most of the different joining (J) genes of the corresponding target immune receptor coding sequence. In some embodiments, the composition for analyzing the B cell receptor (BCR) repertoire in a sample comprises at least one set of primers i) and ii), wherein i) consists of (a) a plurality of V gene primers directed against most of the different V genes of at least one BCR coding sequence comprising at least a portion of framework region 1 (FR1) within the V gene or (b) a plurality of V gene primers directed against most of the different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene; and ii) consists of (a) one or more C gene primers or (b) a plurality of J gene primers, the one or more C gene primers being directed against at least a portion of the C gene of the at least one BCR coding sequence, the plurality of J gene primers being directed against at least a portion of most of the different J genes of the at least one BCR coding sequence; wherein each set of i) primers and ii) primers is directed against the coding sequence of the same target BCR gene selected from IgH, IgL, and IgK; and wherein each set of i) primers and ii) primers directed against the same target BCR is configured to amplify the target BCR repertoire.
[0006] In some embodiments, the V gene primers recognize at least a portion of FR1 within the V gene. In some embodiments, the V gene primers recognize at least a portion of FR2 within the V gene. In some embodiments, the V gene primers recognize at least a portion of FR3 within the V gene. Each set of i) primers and ii) primers targets the same target BCR sequence selected from the group consisting of IgH, IgL, and IgK and is configured such that the resulting amplicons generated using such a composition represent a repertoire of sequences of the corresponding receptors in the sample. In certain embodiments, the provided composition comprises a plurality of primer pair reagents selected from Tables 2 and 5. In certain embodiments, the provided composition comprises a plurality of primer pair reagents selected from Tables 3 and 5. In certain embodiments, the provided composition comprises a plurality of primer pair reagents selected from Tables 3 and 6 - 10. In certain embodiments, the provided composition comprises a plurality of primer pair reagents selected from Tables 2 and 6 - 10. In some embodiments, the provided composition comprises a plurality of primer pair reagents selected from Table 4 and Table 5 or selected from Table 4 and Tables 6 - 10. In some embodiments, a multiplex assay comprising the composition of the present invention is provided. In some embodiments, a test kit comprising the composition of the present invention is provided.
[0007] In other aspects of the present invention, methods for assaying immune repertoire activity in a biological sample are provided. Such methods include performing multiplex amplification with a primer set that targets two different types of immune receptors in a single reaction, such as multiplex amplification of TCR targets and BCR targets.
[0008] In some embodiments, a method for amplifying the expressed nucleic acid sequences of an immune receptor repertoire in a sample comprises performing a single multiplex amplification reaction using at least one set of the following to amplify the expressed target immune receptor nucleic acid template molecules:
[0009] i) (a) a plurality of V gene primers for most different V genes targeting at least one immune receptor coding sequence comprising at least a portion of FR1 within the V gene,
[0010] (b) a plurality of V gene primers for most different V genes targeting at least one immune receptor coding sequence comprising at least a portion of FR2 within the V gene, or
[0011] (c) a plurality of V gene primers for most different V genes targeting at least one immune receptor coding sequence comprising at least a portion of FR3 within the V gene; and
[0012] ii) (a) one or more C gene primers that target at least a portion of the C gene of the at least one immune receptor coding sequence, or
[0013] (b) A plurality of J gene primers, the plurality of J gene primers being directed to at least a portion of most of the different J genes of the at least one immune receptor coding sequence;
[0014] Wherein each set of i) primers and ii) primers is directed to the coding sequence of the same target immune receptor gene selected from a T cell receptor gene or an antibody receptor gene, and wherein performing the amplification using the at least one set of i) and ii) primers produces amplicon molecules representative of the target immune receptor repertoire in the sample; thereby producing immune receptor amplicon molecules comprising the target immune receptor repertoire.
[0015] In certain embodiments, at least a portion of the first framework region (FR1) of the V gene of the immune receptor sequence to at least a portion of the C gene is encompassed within the amplified target immune receptor sequence. In certain embodiments, at least a portion of the second framework region (FR2) of the V gene of the immune receptor sequence to at least a portion of the C gene is encompassed within the amplified target immune receptor sequence. In certain embodiments, at least a portion of the third framework region (FR3) of the V gene of the immune receptor sequence to at least a portion of the C gene is encompassed within the amplified target immune receptor sequence. In other embodiments, the method comprises amplifying the expressed nucleic acid sequences of the immune receptor repertoire in a sample, which comprises performing a multiplex amplification reaction in the presence of a polymerase under amplification conditions to produce a plurality of amplified target expressed sequences, the plurality of amplified target expressed sequences comprising one or more immune receptors of interest having variable, diversity, and joining (VDJ) gene segments or one or more immune receptors of interest having variable and joining (VJ) gene segments. In certain embodiments, at least a portion of the first framework region (FR1) of the V gene of the immune receptor sequence to at least a portion of the joining (J) gene is encompassed within the amplified target immune receptor sequence. In certain embodiments, at least a portion of the second framework region (FR2) of the V gene of the immune receptor sequence to at least a portion of the joining (J) gene is encompassed within the amplified target immune receptor sequence. In certain embodiments, at least a portion of the third framework region (FR3) of the V gene of the immune receptor sequence to at least a portion of the joining (J) gene is encompassed within the amplified target immune receptor sequence.
[0016] The method of the present invention further comprises preparing a BCR repertoire library using the amplified target immune receptor sequences by introducing adapter sequences to the ends of the amplified target sequences. In some embodiments, the adapter-modified immune receptor repertoire library is clonally amplified.
[0017] The method further includes detecting the sequence of the immunoglobulin repertoire of each immunoglobulin receptor in the sample and / or the expression of each target immunoglobulin receptor sequence among the plurality of target immunoglobulin receptor sequences, wherein a change in the level of the repertoire sequence and / or the expression of one or more target immunoglobulin receptor markers, compared to a second sample or a control sample, determines a change in the immunoglobulin repertoire activity in the sample. In certain embodiments, sequencing of immunoglobulin receptor amplicon molecules is performed using next-generation sequence analysis to determine the sequence of the immunoglobulin receptor amplicon. In a particular embodiment, determining the sequence of an immunoglobulin receptor amplicon molecule includes obtaining initial sequence reads, aligning and identifying productive reads, correcting errors to generate rescued productive reads, and determining the sequence of the resulting total productive reads, thereby providing the sequence of the immunoglobulin repertoire in the sample. The provided methods described herein utilize the compositions of the invention provided herein. In still other aspects of the invention, specific analytical methods are provided for error correction in order to generate comprehensive and valid sequence information from the methods provided herein.
[0018] In another aspect, a method is provided for identifying or screening for biomarkers of a disease or condition in a subject. In some embodiments, such methods include performing a single multiplex amplification reaction using at least one set of the following to amplify target BCR nucleic acid template molecules obtained from a subject sample:
[0019] i) (a) a plurality of V gene primers for a majority of different V genes for at least one BCR coding sequence that includes at least a portion of FR1 within the V gene,
[0020] (b) a plurality of V gene primers for a majority of different V genes for at least one BCR coding sequence that includes at least a portion of FR2 within the V gene, or
[0021] (c) a plurality of V gene primers for a majority of different V genes for at least one BCR coding sequence that includes at least a portion of FR3 within the V gene; and
[0022] ii) (a) one or more C gene primers that target at least a portion of the C gene of the at least one BCR coding sequence, or
[0023] (b) a plurality of J gene primers that target at least a portion of a majority of different J genes of the at least one BCR coding sequence;
[0024] Each set of i) primers and ii) primers targets the coding sequence of the same target BCR gene selected from the IgH, IgL, and IgK genes, and wherein performing the amplification using said at least one set of i) primers and ii) primers generates amplicon molecules representative of the target BCR repertoire in the sample; thereby generating target BCR amplicon molecules comprising the target BCR repertoire. The method further comprises: performing sequencing on the target BCR amplicon molecules and determining the sequences of said molecules, wherein determining the sequences comprises obtaining initial sequence reads, aligning the initial sequence reads to a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads; identifying a BCR repertoire clone population from the determined target BCR sequences; and identifying the sequence of at least one BCR clone to be used as a biomarker for the disease or condition. In some embodiments, the disease or condition biomarker is identified or screened from cancer, autoimmune diseases, infectious diseases, allergies, response to vaccination, and response to immunotherapy treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is an exemplary workflow for removing PCR- or sequencing-derived errors using progressive clustering of similar CDR3 nucleotide sequences, where the steps are: (A) very fast heuristic clustering into groups based on similarity (cd-hit-est); (B) selecting the most common sequence as the representative of the cluster, randomly picking to break ties; (C) merging reads into the representative; (D) comparing representatives and merging clusters if within the assigned Hamming distance.
[0026] Figure 2 is an exemplary workflow for removing residual insertion / deletion (indel) errors by comparing homopolymer-collapsed CDR3 sequences using the Levenshtein distance, where the steps are: (A) collapsing homopolymers and calculating the Levenshtein distance between cluster representatives; (B) merging reads that are now clustered together, which represent compound indel errors; (C) reporting the lineage to the user.
[0027] Figure 3 is a graph depicting the results of the number of reads and the characterization of read quality (productive vs. off-target or non-productive) for individual BCR assays, BCR and TCR assays amplified in one pool, and BCR and TCR assays amplified in two separate pools.
[0028] Figures 4A - 4B depicts the total number of clones detected in the combined assay ( Figure 4A) and populations of BCR clones and TCR clones (as a percentage of the total) ( Figure 4B ).
[0029] Figures 5A - 5F is a histogram depicting the sequence read lengths of the IgH repertoire of RNA from various cell or tissue samples including: ( Figure 5A ) PBL, ( Figure 5B ) CD19+ cells, ( Figure 5C ) tonsil FFPE, ( Figure 5D ) lung tumor FFPE, ( Figure 5E ) bone marrow, ( Figure 5F ) normal spleen, and ( Figure 5G ) normal brain.
[0030] Figure 6 Depicts sequence read lengths obtained after multiplex amplification of PBL cDNA using exemplary IgH V gene FR1 - C gene primer sets 1 - 7.
[0031] Figures 7A - 7B The bar graph depicts the total isotype representation as sequence reads ( Figure 7A ) for each isotype and clones detected ( Figure 7B ) for each isotype determined from an exemplary IgH V gene FR1 - C gene multiplex amplification reaction within the PBL sample.
[0032] Figures 8A - 8B Depicts histograms of the IgH V gene mutation rates for all IgH isotypes ( Figure 8A ) and for the PBL sample of only IgD ( Figure 8B ).
[0033] Figure 9 Depicts the IgH clone analysis results for the total productive reads from the sample (the right - most point on each graph) and 8 down - sampled data sets derived from the total productive reads.
[0034] Figure 10 Depicts a graph showing the linearity of plasmid detection in an IgH library generated from a pool of 20 control plasmids mixed with leukocyte cDNA at equimolar concentrations. The plasmids associated with plasmid ID numbers are shown in Table 15. Detailed Description
[0035] Multiplexed next-generation sequencing workflows have been developed for the efficient detection and analysis of immune repertoires in samples. The provided methods, compositions, systems, and kits are used for highly accurate amplification and sequencing of immune cell receptor sequences (e.g., T cell receptor (TCR), B cell receptor (BCR or Ab) targets) when monitoring and resolving one or more complex immune cell repertoires of a subject. The target immune cell receptor genes have undergone rearrangement (or recombination) of VDJ or VJ gene segments, which depend on the specific receptor gene (e.g., IgH, IgK, TCRβ, or TCRα). In certain embodiments, the present disclosure provides methods, compositions, and systems that use nucleic acid amplification, such as polymerase chain reaction (PCR), to enrich the expressed variable regions of immune receptor target nucleic acids for subsequent sequencing. In certain embodiments, the present disclosure provides methods, compositions, and systems that use nucleic acid amplification, such as PCR, to enrich the rearranged target immune cell receptor gene sequences from gDNA for subsequent sequencing. In certain embodiments, the present disclosure also provides methods and systems for effectively identifying and removing one or more errors derived from amplification or sequencing to improve read assignment accuracy and reduce false positive rates. Specifically, the provided methods described herein can improve accuracy and performance in sequencing applications where nucleotide sequences are related to genomic recombination and high variability. In some embodiments, the methods, compositions, systems, and kits provided herein are used for amplifying and sequencing the complementarity-determining regions (CDRs) of expressed immune receptors in a sample. In some embodiments, the methods, compositions, systems, and kits provided herein are used for amplifying and sequencing the CDRs of rearranged immune cell receptor gDNA in a sample. Accordingly, the present disclosure provides multiplexed immune cell receptor expression compositions and compositions directed to immune cell receptor genes for multiplexed library preparation for use in combination with next-generation sequencing technologies and workflow solutions (e.g., manual or automated) to effectively detect and characterize immune repertoires in samples.
[0036] The CDRs of TCRs or BCRs are generated from genomic DNA that undergoes recombination of V(D)J gene segments and addition and / or deletion of nucleotides at the junctions of gene segments. Recombination of V(D)J gene segments and subsequent hypermutation events result in extensive diversity of expressed immune cell receptors. Due to the random nature of V(D)J recombination, it is common for rearrangements of T or B cell receptor genomic DNA to fail to produce a functional receptor and instead produce so-called "non-productive" rearrangements. Typically, non-productive rearrangements have out-of-frame variable and joining coding segments and result in the presence of premature stop codons and synthesis of unrelated peptides. However, for many biological or physiological reasons such as the following, non-productive TCR or BCR gene rearrangements are generally rare in cDNA-based repertoire sequencing: 1) nonsense-mediated decay, which destroys mRNAs containing premature stop codons, 2) B and T cell selection, in which only B cells and T cells with functional receptors survive, and 3) allelic exclusion, in which only a single rearranged receptor allele is expressed in any given B or T cell.
[0037] Thus, in some embodiments, the methods and compositions provided herein are used to amplify the rearranged variable regions of immune cell receptor mRNAs (e.g., BCR and / or TCR mRNAs). In some embodiments, RNA extracted from a biological sample is converted to cDNA. Multiplex amplification is used to enrich a portion of the BCR or TCR cDNA that contains at least a portion of the variable region of the receptor. In some embodiments, the amplified cDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of the target receptor. In some embodiments, the amplified cDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of the immunoglobulin heavy chain (IgH).
[0038] TCR sequences and BCR sequences can also appear as non-productive rearrangements due to errors introduced during the amplification reaction or during the sequencing process. For example, insertion or deletion (indel) errors during the target amplification or sequencing reaction can result in a frameshift in the reading frame of the resulting coding sequence. Such alterations can cause productive rearrangement target sequence reads to be interpreted as non-productive rearrangements and discarded from the set of identified clonotypes. Thus, in some embodiments, the methods and systems provided herein include methods for identifying and / or removing PCR- or sequencing-derived errors from determined immune receptor sequences.
[0039] In some embodiments, the provided methods and compositions are used to amplify rearranged variable regions of immune cell receptor gDNA (e.g., rearranged BCR and / or TCR gene DNA). Multiplex amplification is used to enrich a portion of rearranged BCR or TCR gDNA that contains at least a portion of the variable region of the receptor. In some embodiments, the amplified gDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of the target receptor. In some embodiments, the amplified gDNA contains one or more complementarity-determining regions CDR1, CDR2, and / or CDR3 of IgH. In some embodiments, the amplified gDNA predominantly contains the CDR3 of the target receptor, e.g., the CDR3 of IgH.
[0040] As used herein, the terms "immune cell receptor" and "immune receptor" are used interchangeably.
[0041] As used herein, the terms "complementarity-determining region" and "CDR" refer to regions of a T cell receptor or antibody (immunoglobulin) where the molecule complements the conformation of an antigen, thereby determining the molecule's specificity and contact with a particular antigen. In the variable regions of T cell receptors and antibodies, CDRs are interspersed with more conserved regions called framework regions (FRs). Each variable region of a T cell receptor and antibody contains 3 CDRs, named CDR1, CDR2, and CDR3, and also contains 4 framework subregions, named FR1, FR2, FR3, and FR4.
[0042] As used herein, the term "framework" or "framework region" or "FR" refers to the residues of the variable region other than the CDR residues defined herein. There are four separate framework subregions that make up the framework: FR1, FR2, FR3, and FR4.
[0043] In the art, the specific names for the exact positions of CDRs and FRs within a receptor molecule (TCR or immunoglobulin) vary according to the definitions employed. Unless otherwise specifically stated, IMGT nomenclature is used herein to describe CDR and FR regions (see Brochet et al. (2008) Nucleic Acids Res. 36:W503-508, which is incorporated herein by reference in its entirety). As an example of CDR / FR amino acid names, the residues that make up the FR and CDRs of T cell receptor β have been characterized by IMGT as follows: residues 1-26 (FR1), 27-38 (CDR1), 39-55 (FR2), 56-65 (CDR2), 66-104 (FR3), 105-117 (CDR3), and 118-128 (FR4).
[0044] Other well-known standard names used to describe the regions include those that are present in Kabat et al., (1991) Sequences of Proteins of Immunological Interest, 5th ed. Public Health Service, National Institutes of Health, Bethesda, Md. and Chothia and Lesk 1987, J. Mol. Biol. 196:901-917), which are specifically incorporated herein by reference. As an example of CDR names, the residues that make up the six immunoglobulin CDRs have been characterized by Kabat as follows: residues 24-34 (CDRL1), 50-56 (CDRL2), and 89-97 (CDRL3) in the light chain variable region and 31-35 (CDRH1), 50-65 (CDRH2), and 95-102 (CDRH3) in the heavy chain variable region; and by Chothia as follows: residues 26-32 (CDRL1), 50-52 (CDRL2), and 91-96 (CDRL3) in the light chain variable region and 26-32 (CDRH1), 53-55 (CDRH2), and 96-101 (CDRH3) in the heavy chain variable region.
[0045] As used herein, the term "T cell receptor" or "T cell antigen receptor" or "TCR" refers to the antigen / MHC-binding heterodimeric protein product of a vertebrate, such as a mammal, and the TCR gene complex, including the human TCR alpha, beta, gamma, and delta chains. For example, the complete sequence of the human TCR beta locus has been sequenced, see, e.g., Rowen et al. (1996) Science 272:1755-1762; the human TCR alpha locus has been sequenced and re-sequenced, see, e.g., Mackelprang et al. (2006) Hum Genet. 119:255-266; and for a general analysis of the T cell receptor V gene segment families, see, e.g., Arden (1995) Immunogenetics 42:455-500; each of which is specifically incorporated herein by reference for the sequence information provided and cited in the publications.
[0046] As used herein, the term "antibody" or "immunoglobulin" or "B cell receptor" or "BCR" refers to an immunoglobulin molecule composed of four polypeptide chains (two heavy (H) chains and two light (L) chains) (λ or κ) that are interconnected by disulfide bonds. An antibody has a known specific antigen to which it binds. Each heavy chain of an antibody contains a heavy chain variable region (abbreviated herein as HCVR, HV, or VH) and a heavy chain constant region. The heavy chain constant region includes three domains, CH1, CH2, and CH3. Each light chain contains a light chain variable region (abbreviated herein as LCVR or VL or KV or LV to designate a κ or λ light chain) and a light chain constant region. The light chain constant region contains one domain, CL. The heavy chain determines the class or isotype to which the immunoglobulin belongs. For example, in mammals, the five major immunoglobulin isotypes are IgA, IgD, IgG, IgE, and IgM, and the immunoglobulin isotypes are classified according to the α, δ, ε, γ, or μ heavy chains they contain, respectively.
[0047] As mentioned, the diversity of TCR and BCR chain CDRs is generated by the recombination of germline variable (V), diversity (D), and joining (J) gene segments and by the independent addition and deletion of nucleotides at each gene segment junction during gene segment ligation in the process of TCR and BCR gene rearrangement. In a rearranged nucleic acid encoding a BCR heavy chain, CDR1 and CDR2 are found in the V gene segment, and CDR3 contains some gene segments from the V gene segment as well as D and J gene segments. In a rearranged nucleic acid encoding a BCR light chain, CDR1 and CDR2 are found in the V gene segment, and CDR3 contains some gene segments from the V gene segment and J gene segment. For example, in rearranged nucleic acids encoding TCRβ and TCRδ, CDR1 and CDR2 are found in the V gene segment, and CDR3 contains some gene segments from the V gene segment as well as D and J gene segments. In rearranged nucleic acids encoding TCRα and TCRγ, CDR1 and CDR2 are found in the V gene segment region, and CDR3 contains some gene segments from the V gene segment and J gene segment.
[0048] In some embodiments, multiplex amplification reactions are used to amplify cDNA that is derived from mRNA expressed from rearranged BCR and / or TCR genomic DNA. In some embodiments, multiplex amplification reactions are used to amplify at least a portion of BCR and / or TCR CDRs from cDNA derived from a biological sample. In some embodiments, multiplex amplification reactions are used to amplify at least two CDRs of BCR and / or TCR from cDNA derived from a biological sample. In some embodiments, multiplex amplification reactions are used to amplify at least three CDRs of BCR and / or TCR from cDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of the BCR and / or TCR CDRs expressed in the sample. In some embodiments, determining the nucleotide sequences of such amplicons that include at least 3 CDRs is used to identify and characterize novel BCR and / or TCR alleles.
[0049] In some embodiments, multiplex amplification reactions are used to amplify BCR and / or TCR genomic DNA that has undergone V(D)J rearrangement. In some embodiments, multiplex amplification reactions are used to amplify one or more nucleic acid molecules that include at least a portion of BCR and / or TCR CDRs from gDNA derived from a biological sample. In some embodiments, multiplex amplification reactions are used to amplify one or more nucleic acid molecules that include at least two CDRs of BCR and / or TCR from gDNA derived from a biological sample. In some embodiments, multiplex amplification reactions are used to amplify nucleic acid molecules that include at least three CDRs of BCR and / or TCR from gDNA derived from a biological sample. In some embodiments, the resulting amplicons are used to determine the nucleotide sequences of the rearranged BCR and / or TCR CDRs in the sample. In some embodiments, determining the nucleotide sequences of such amplicons that include at least CDR3 is used to identify and characterize novel BCR and / or TCR alleles.
[0050] In some embodiments of the multiplex amplification reactions, each primer set used targets the same BCR or TCR region, however different primers within the set allow targeting of different V(D)J gene rearrangements of the gene. For example, primer sets for amplifying expressed IgH or rearranged IgH gDNA are all designed to target one or more identical regions from IgH mRNA or IgH gDNA respectively, but individual primers within the set result in amplification of various IgH VDJ gene combinations. In some embodiments, at least one primer or primer set is directed to a relatively conserved region of the immunoreceptor gene (e.g., a portion of the C gene), and another primer set includes various primers directed to a more variable region of the same gene (e.g., a portion of the V gene). In other embodiments, at least one primer set includes various primers directed to at least a portion of the J gene segment of the immunoreceptor gene, and another primer set includes various primers directed to at least a portion of the V gene segment of the same gene.
[0051] In some embodiments, multiplex amplification reactions are used to amplify cDNA of mRNA expressed from rearranged BCR genomic DNA that is derived from genomic DNA comprising rearranged IgH, IgK, and IgL. In some embodiments, at least a portion of the BCR CDR, such as CDR3, is amplified from the cDNA in the multiplex amplification reaction. In some embodiments, at least two CDR portions of the BCR are amplified from the cDNA in the multiplex amplification reaction. In certain embodiments, the multiplex amplification reaction is used to amplify at least the CDR1 region, CDR2 region, and CDR3 region of the BCR cDNA. In some embodiments, the resulting amplicons are used to determine the nucleotide sequence of the expressed BCR CDR. In some embodiments, the resulting amplicons are used to determine the nucleotide sequence of the expressed BCR CDR and the Ig isotype of the sequence. In some embodiments, the resulting amplicons are used to determine the nucleotide sequence of the expressed IgH CDR and Ig isotype and Ig sub-isotype.
[0052] In some embodiments, multiplex amplification reactions are used to amplify rearranged BCR genomic DNA, which comprises rearranged IgH, IgK, and IgL genomic DNA. In some embodiments, at least a portion of the BCR CDR, such as CDR3, is amplified from the gDNA in the multiplex amplification reaction. In some embodiments, at least two CDR portions of the BCR are amplified from the gDNA in the multiplex amplification reaction. In certain embodiments, the multiplex amplification reaction is used to amplify at least the CDR1 region, CDR2 region, and CDR3 region of the rearranged BCR gDNA. In some embodiments, the resulting amplicons are used to determine the nucleotide sequence of the rearranged BCR CDR. In some embodiments, the resulting amplicons are used to determine the nucleotide sequence of the rearranged BCR CDR and the Ig isotype of the sequence.
[0053] In some embodiments, a multiplex amplification reaction is performed with a primer set designed to generate amplicons that include the CDR1 region, CDR2 region, and / or CDR3 region of the expression of a target immune receptor mRNA. In some embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each targeting at least a portion of the framework region FR1 of a V gene and (ii) at least one primer targeting a portion of at least one C gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each targeting at least a portion of the framework region FR2 of a V gene and (ii) at least one primer targeting a portion of at least one C gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each targeting at least a portion of the framework region FR3 of a V gene and (ii) at least one primer targeting a portion of at least one C gene of the target immune receptor. In some embodiments, a multiplex amplification reaction is performed with a primer set designed to generate amplicons that include one or more expressed IgH isotypes of the target mRNA, and such reaction is performed using: (i) one of the above-described FR1 primer set, FR2 primer set, or FR3 primer set and (ii) a set of primers each targeting a portion of at least one C gene of IgA, IgD, IgE, IgG, and / or IgM. In some embodiments, one or more primers targeting a C gene are coding sequences of the C gene within about 200 nucleotides of the 5' end of the one or more C genes. In some embodiments, one or more primers targeting a C gene are coding sequences of the C gene within about 150 nucleotides of the 5' end of the one or more C genes. In some embodiments, one or more primers targeting a C gene are coding sequences of the C gene within about 100 nucleotides of the 5' end of the one or more C genes. In some embodiments, one or more primers targeting a C gene are coding sequences of the C gene within about 50 nucleotides, about 50 to about 150 nucleotides, about 75 to about 175 nucleotides, or about 100 to about 200 nucleotides of the 5' end of the one or more C genes. In some embodiments, one or more primers targeting a C gene target gene-coding sequencing of the C gene that not only differentiates isotypes but also allows determination of sub-isotypes. For example, in some embodiments, one or more primers targeting a C gene generate a sufficient portion of the constant region in the amplicon such that the sub-isotype can be determined based on the sequenced data. In some embodiments, one or more primers targeting a C gene include primers targeting IgG and / or IgA C gene coding sequences that allow discrimination of IgG1 sub-isotype, IgG2 sub-isotype, IgG3 sub-isotype, IgG4 sub-isotype, IgA1 sub-isotype, and IgA2 sub-isotype.
[0054] In some embodiments, a multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR1 region of the V gene and (ii) at least one primer that anneals to a portion of the constant (C) gene to amplify BCR cDNA such that the resulting amplicons contain the CDR1-encoding portion, the CDR2-encoding portion, and the CDR3-encoding portion of the BCR mRNA. In certain embodiments, the set of primers for FR1 is combined with a set of at least two primers for the C gene to produce amplicons that contain at least the CDR1-encoding portion, the CDR2-encoding portion, and the CDR3-encoding portion of the BCR mRNA. In some embodiments, the set of primers for IgH FR1 is combined with a set of at least two C gene primers that target the coding portions of two different IgH isotypes to produce amplicons that contain at least the CDR1-encoding portion, the CDR2-encoding portion, and the CDR3-encoding portion of the IgH mRNA. In some embodiments, the set of primers for IgH FR1 is combined with at least three, at least four, or at least five primers that target the coding portions of different IgH isotypes. For example, exemplary primers specific for the FR1 region of the IgH V gene are shown in Table 3, and exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0055] In some embodiments, a multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR2 region of the V gene and (ii) at least one primer that anneals to a portion of the C gene to amplify BCR cDNA such that the resulting amplicons contain the CDR2-encoding portion and the CDR3-encoding portion of the BCR mRNA. In certain embodiments, such a set of primers for FR2 is combined with at least two primers for the C gene to produce amplicons that contain the CDR2-encoding portion and the CDR3-encoding portion of the BCR mRNA. In some embodiments, the set of primers for IgH FR2 is combined with a set of at least two C gene primers that target the coding portions of two different IgH isotypes to produce amplicons that have the CDR2 and CDR3-encoding portions of the IgH mRNA. In some embodiments, the set of primers for IgH FR2 is combined with at least three, at least four, or at least five C gene primers that target the coding portions of different IgH isotypes. Exemplary primers for FR2 include the BIOMED-2 primers (van Dongen et al. (2003) Leukemia 17:2257-2327) developed and standardized by a consortium of European academic laboratories and research hospitals and shown in Table 4. Exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0056] In some embodiments, a multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR3 region of a V gene and (ii) at least one primer that anneals to a portion of a C gene to amplify BCR cDNA such that the resulting amplicons predominantly contain the CDR3-encoding portion of BCR mRNA. In certain embodiments, such a set of primers for FR3 is combined with at least two primers for the C gene to produce amplicons having the CDR3-encoding portion of BCR mRNA. In some embodiments, a set of primers for IgH FR3 is combined with a set of at least two C gene primers that target the coding portions of two different IgH isotypes to produce amplicons having the CDR3-encoding portion of IgH mRNA. In some embodiments, a set of primers for IgH FR3 is combined with at least three, at least four, or at least five C gene primers that target the coding portions of different IgH isotypes. For example, exemplary primers specific for the IgH V gene FR3 region are shown in Table 2, and exemplary primers specific for the IgH C gene are shown in Tables 6-10.
[0057] In some embodiments, a multiplex amplification reaction is performed with a set of primers designed to produce amplicons that contain the CDR1 region, CDR2 region, and / or CDR3 region of a target immune receptor mRNA or rearranged gDNA. In some embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each of which targets at least a portion of the framework region FR1 of a V gene and (ii) a set of primers each of which targets at least a portion of the J gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each of which targets at least a portion of the framework region FR2 of a V gene and (ii) a set of primers each of which targets at least a portion of the J gene of the target immune receptor. In other embodiments, the multiplex amplification reaction is performed using: (i) a set of primers each of which targets at least a portion of the framework region FR3 of a V gene and (ii) a set of primers each of which targets at least a portion of the J gene of the target immune receptor.
[0058] In some embodiments, a multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR1 region of a V gene and (ii) a set of primers that anneals to a portion of a J gene to amplify BCR nucleic acid such that the resulting amplicons contain the CDR1-encoding portion, CDR2-encoding portion, and CDR3-encoding portion of BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the IgH V gene FR1 region are shown in Table 3, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0059] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR2 region of the V gene and (ii) a set of primers that anneals to a portion of the J gene to amplify the BCR nucleic acid such that the resulting amplicons contain the CDR2-encoding portion and the CDR3-encoding portion of the BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the FR2 region of the IgH V gene are shown in Table 4, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0060] In some embodiments, the multiplex amplification reaction uses (i) a set of primers each of which anneals to at least a portion of the FR3 region of the V gene and (ii) a set of primers that anneals to a portion of the J gene to amplify the BCR nucleic acid such that the resulting amplicons predominantly contain the CDR3-encoding portion of the BCR mRNA or rearranged gDNA. For example, exemplary primers specific for the FR3 region of the IgH V gene are shown in Table 2, and exemplary primers specific for the IgH J gene are shown in Table 5.
[0061] In some embodiments, there are provided compositions for multiplex amplification of at least a portion of the expressed BCR variable region. In some embodiments, the compositions include multiple primer pair reagents directed to a portion of the V gene framework region and a portion of the constant (C) gene of a rearranged target immune receptor gene selected from the group consisting of: immunoglobulin heavy chain (IgH), immunoglobulin light chain lambda (IgL), and immunoglobulin light chain kappa (IgK). In some embodiments, the compositions include multiple primer pair reagents directed to a portion of the V gene framework region and a portion of the J gene of a rearranged target immune receptor gene selected from the group consisting of IgH, IgL, and IgK.
[0062] In some embodiments, the compositions include (i) multiple primer pair reagents directed to a portion of the IgH V gene framework region and a portion of the IgH C gene of the rearranged IgH gene and (ii) multiple primer pair reagents directed to a portion of the TCRβ V gene framework region and a portion of the TCRβ C gene of the rearranged TCRβ gene. In some embodiments, the compositions include (i) multiple primer pair reagents directed to a portion of the IgH V gene framework region and a portion of the IgH J gene of the rearranged IgH gene and (ii) multiple primer pair reagents directed to a portion of the TCRβ J gene framework region and a portion of the TCR V gene of the rearranged TCRβ gene.
[0063] Amplification by PCR is performed with at least two primers. For the methods provided herein, a set of primers is used that is sufficient to amplify all or a defined portion of the variable sequences at the locus of interest, which may include any or all of the TCR and immunoglobulin loci described above. In some embodiments, various parameters or criteria outlined herein can be used to select a target-specific primer set for multiplex amplification.
[0064] In some embodiments, the primer sets used in the multiplex reactions are designed to amplify at least 50% of the known expressed rearrangements or gDNA rearrangements at the locus of interest. In certain embodiments, the primer sets used in the multiplex reactions are designed to amplify at least 75%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or more of the known expressed rearrangements or gDNA rearrangements at the locus of interest. As another example, combining 27 forward primers from Table 3 (each primer targeting a portion of the FR1 region from a different IgH V gene) with at least one reverse primer from Tables 6 - 10 (each primer targeting a portion of a different IgH C gene) will amplify all of the currently known expressed IgH rearrangements for a given isotype. As another example, combining 68 forward primers from Table 2 (each primer targeting a portion of the FR3 region from a different IgH V gene) with at least one reverse primer from Tables 6 - 10 (each primer targeting a portion of a different IgH C gene) will amplify all of the currently known expressed IgH rearrangements for a given isotype. As another example, combining 68 forward primers from Table 2 (each primer targeting a portion of the FR3 region from a different IgH V gene) with 4 reverse primers from Table 5 (each primer targeting a portion of a different IgH J gene) will amplify all of the currently known expressed IgH rearrangements or gDNA IgH rearrangements. As another example, combining 27 forward primers from Table 3 (each primer targeting a portion of the FR1 region from a different IgH V gene) with 4 reverse primers from Table 5 (each primer targeting a portion of a different IgH J gene) will amplify all of the currently known expressed IgH rearrangements or gDNA IgH rearrangements.
[0065] For example, such multiplex amplification reactions comprise at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 reverse primers, wherein each reverse primer is directed to a sequence corresponding to at least a portion of the FR1 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers directed to the FR1 region of the BCR V gene are combined with at least 1 forward primer directed to a sequence corresponding to at least a portion of the constant gene of the same BCR gene. In some embodiments, the plurality of reverse primers directed to the FR1 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or from about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 forward primers, each primer being directed to a sequence corresponding to at least a portion of at least one gene of the constant gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer directed to the FR1 of the BCR V gene can be a forward primer, and one or more primers directed to the BCR C gene can be one or more reverse primers. Thus, in some embodiments, the multiplex amplification reaction comprises at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 forward primers, wherein each forward primer is directed to a sequence corresponding to at least a portion of the FR1 region of one or more BCR V genes. In such embodiments, the plurality of forward primers directed to the FR1 region of the BCR V gene are combined with at least 1 reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the plurality of forward primers directed to the FR1 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or from about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 reverse primers, each primer being directed to a sequence corresponding to at least a portion of at least one gene of the C gene of the same BCR gene. In some embodiments, such FR1 and C gene amplification primer sets can be directed to IgH gene sequences.In some preferred embodiments, about 22 to about 35 reverse primers directed to FR1 regions of different IgH V genes are combined with about 2 to about 8 forward primers directed to a portion of the IgH C gene. In other preferred embodiments, about 22 to about 35 reverse primers directed to FR1 regions of different IgH V genes are combined with about 5 to about 15 forward primers directed to a portion of the IgH C gene. In other preferred embodiments, about 48 to about 60 reverse primers directed to FR1 regions of different IgH V genes are combined with about 5 to about 15 forward primers directed to a portion of the IgH C gene. In some preferred embodiments, about 22 to about 35 forward primers directed to FR1 regions of different IgH V genes are combined with about 2 to about 8 reverse primers directed to a portion of the IgH C gene. In other preferred embodiments, about 22 to about 35 forward primers directed to FR1 regions of different IgH V genes are combined with about 5 to about 15 reverse primers directed to a portion of the IgH C gene. In still other preferred embodiments, about 48 to about 60 forward primers directed to FR1 regions of different IgH V genes are combined with about 5 to about 15 reverse primers directed to a portion of the IgH C gene. In some preferred embodiments, the forward primers directed to the FR1 region of the IgH V gene are selected from those listed in Table 3, and the reverse primers directed to the IgH C gene are selected from those listed in Tables 6-10. In other embodiments, the FR1 and C gene amplification primer sets can be directed to Ig light chain λ, Ig light chain κ, TCRα, TCRγ, TCRδ or TCRβ gene sequences.
[0066] In some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 reverse primers, wherein each reverse primer is directed to a sequence corresponding to at least a portion of the FR2 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers directed to the FR2 region of the BCR V gene are combined with at least 1 forward primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the plurality of reverse primers directed to the FR2 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 forward primers, each primer being directed to a sequence corresponding to at least a portion of at least one gene of the C gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer directed to the FR2 of the BCR V gene can be a forward primer, and one or more primers directed to the BCR C gene can be one or more reverse primers. Thus, in some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 forward primers, wherein each forward primer is directed to a sequence corresponding to at least a portion of the FR2 region of one or more BCR V genes. In such embodiments, the plurality of forward primers directed to the FR2 region of the BCR V gene are combined with at least 1 reverse primer directed to a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the plurality of forward primers directed to the FR2 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15, or about 7 to about 12 reverse primers, each primer being directed to a sequence corresponding to at least a portion of at least one gene of the C gene of the same BCR gene. In some embodiments, such FR2 and C gene amplification primer sets can be directed to IgH gene sequences. In some embodiments, about 5 to about 15 reverse primers directed to different IgH V gene FR2 regions are combined with about 2 to about 8 forward primers directed to a portion of the IgH C gene. In some embodiments, about 5 to about 15 reverse primers directed to different IgH V gene FR2 regions are combined with about 5 to about 15 forward primers directed to a portion of the IgH C gene. In some embodiments, about 5 to about 15 forward primers directed to different IgH V gene FR2 regions are combined with about 2 to about 8 reverse primers directed to a portion of the IgH C gene.In some embodiments, from about 5 to about 15 forward primers directed to FR2 regions of different IgHV genes are combined with from about 5 to about 15 reverse primers directed to a portion of the IgH C gene. In some preferred embodiments, the forward primers directed to the FR2 region of the IgH V gene are selected from those forward primers listed in Table 4, and the reverse primers directed to the IgH C gene are selected from those reverse primers listed in Tables 6 - 10. In other embodiments, the FR2 and C gene amplification primer sets can be directed to Ig light chain lambda, Ig light chain kappa, TCR alpha, TCR gamma, TCR delta, or TCR beta gene sequences.
[0067] In some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85 or 90 reverse primers, wherein each reverse primer targets a sequence corresponding to at least a portion of the FR3 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers targeting the FR3 region of the BCR V gene are combined with at least 1 forward primer targeting a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the plurality of reverse primers targeting the FR3 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15 or about 7 to about 12 forward primers, each primer targeting a sequence corresponding to at least a portion of at least one gene of the C gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer targeting the FR3 of the BCR V gene can be a forward primer, and one or more primers targeting the BCR C gene can be one or more reverse primers. Thus, in some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85 or 90 forward primers, wherein each forward primer targets a sequence corresponding to at least a portion of the FR3 region of one or more BCR V genes. In such embodiments, the plurality of forward primers targeting the FR3 region of the BCR V gene are combined with at least 1 reverse primer targeting a sequence corresponding to at least a portion of the C gene of the same BCR gene. In some embodiments, the plurality of forward primers targeting the FR3 region of the BCR V gene are combined with at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15, or about 2 to about 7, about 5 to about 20, about 5 to about 15 or about 7 to about 12 reverse primers, each primer targeting a sequence corresponding to at least a portion of at least one gene of the C gene of the same BCR gene. In some embodiments, such FR3 and C gene amplification primer sets can target IgH gene sequences. In some preferred embodiments, about 62 to about 75 reverse primers targeting different IgH V gene FR3 regions are combined with about 2 to about 8 forward primers targeting a portion of the IgH C gene. In other preferred embodiments, about 62 to about 75 reverse primers targeting different IgH V gene FR3 regions are combined with about 5 to about 15 forward primers targeting a portion of the IgH C gene.In some preferred embodiments, from about 62 to about 75 forward primers directed to the FR3 regions of different IgH V genes are combined with from about 2 to about 8 reverse primers directed to a portion of the IgH C gene. In other preferred embodiments, from about 62 to about 75 forward primers directed to the FR3 regions of different IgH V genes are combined with from about 5 to about 15 reverse primers directed to a portion of the IgH C gene. In some preferred embodiments, the forward primers directed to the FR3 regions of the IgH V genes are selected from those listed in Table 2, and the reverse primers directed to the IgH C gene are selected from those listed in Tables 6-10. In other embodiments, the FR3 and C gene amplification primer sets can be directed to Ig light chain λ, Ig light chain κ, TCRα, TCRγ, TCRδ, and TCRβ gene sequences.
[0068] In some embodiments, such multiplex amplification reactions comprise at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 reverse primers, wherein each reverse primer is directed to a sequence corresponding to at least a portion of the FR1 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers directed to the FR1 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8, or about 3 - 6 forward primers directed to a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer directed to the FR1 of the BCR V gene can be a forward primer, and the primer directed to the BCR J gene can be a reverse primer. Thus, in some embodiments, the multiplex amplification reaction comprises at least 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90, preferably 22, 23, 24, 25, 26, 27, 28, 29, 30, 34, 38, 42, 46, 50, 54, 58, or 62 forward primers, wherein each forward primer is directed to a sequence corresponding to at least a portion of the FR1 region of one or more BCR V genes. In such embodiments, the plurality of forward primers directed to the FR1 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8, or about 3 - 6 reverse primers directed to a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments, such FR1 and J gene amplification primer sets can be directed to IgH gene sequences. In some preferred embodiments, about 22 to about 35 reverse primers directed to different IgH V gene FR1 regions are combined with about 3 to about 6 forward primers directed to different IgH J genes. In some preferred embodiments, about 22 to about 35 forward primers directed to different IgH V gene FR1 regions are combined with about 3 to about 6 reverse primers directed to different IgH J genes. In some preferred embodiments, the forward primers directed to the FR1 region of the IgH V gene are selected from those listed in Table 3, and the reverse primers directed to the IgH J gene are selected from those listed in Table 5. In other embodiments, the FR1 and J gene amplification primer sets can be directed to Ig light chain λ, Ig light chain κ, TCRα, TCRγ, TCRδ, or TCRβ gene sequences.
[0069] In some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 reverse primers, wherein each reverse primer targets a sequence corresponding to at least a portion of the FR2 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers targeting the FR2 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8, or about 3 - 6 forward primers targeting a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer targeting the FR2 region of the BCR V gene can be a forward primer, and the primer targeting the BCR J gene can be a reverse primer. Thus, in some embodiments, the multiplex amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 forward primers, wherein each forward primer targets a sequence corresponding to at least a portion of the FR2 region of one or more BCR V genes. In such embodiments, the plurality of forward primers targeting the FR2 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8, or about 3 - 6 reverse primers targeting a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments, such FR2 and J gene amplification primer sets can target IgH gene sequences. In some preferred embodiments, about 5 to about 15 reverse primers targeting different IgH V gene FR2 regions are combined with about 3 to about 6 forward primers targeting different IgH J genes. In some preferred embodiments, about 5 to about 15 forward primers targeting different IgH V gene FR2 regions are combined with about 3 to about 6 reverse primers targeting different IgH J genes. In some preferred embodiments, the forward primers targeting the FR2 region of the IgH V gene are selected from those listed in Table 4, and the reverse primers targeting the IgH J gene are selected from those listed in Table 5. In other embodiments, the FR2 and J gene amplification primer sets can target Ig light chain λ, Ig light chain κ, TCRα, TCRγ, TCRδ, or TCRβ gene sequences.
[0070] In some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85 or 90 reverse primers, wherein each reverse primer targets a sequence corresponding to at least a portion of the FR3 region of one or more BCR V genes. In such embodiments, the plurality of reverse primers targeting the FR3 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8 or about 3 - 6 forward primers targeting a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments of the multiplex amplification reaction, the primer targeting the FR3 region of the BCR V gene can be a forward primer, and the primer targeting the BCR J gene can be a reverse primer. Thus, in some embodiments, the multiplex amplification reaction comprises at least 20, 25, 30, 40, 45, preferably 50, 55, 60, 65, 70, 75, 80, 85 or 90 forward primers, wherein each forward primer targets a sequence corresponding to at least a portion of the FR3 region of one or more BCR V genes. In such embodiments, the plurality of forward primers targeting the FR3 region of the BCR V gene are combined with at least 2, 3, 4, 5, 6, 8 or about 3 - 6 reverse primers targeting a sequence corresponding to at least a portion of the J gene of the same BCR gene. In some embodiments, such FR3 and J gene amplification primer sets can target IgH gene sequences. In some preferred embodiments, about 62 to about 75 reverse primers targeting different IgH V gene FR3 regions are combined with about 3 to about 6 forward primers targeting different IgH J genes. In some preferred embodiments, about 62 to about 75 forward primers targeting different IgH V gene FR3 regions are combined with about 3 to about 6 reverse primers targeting different IgH J genes. In some preferred embodiments, the forward primers targeting the FR3 region of the IgH V gene are selected from those listed in Table 2, and the reverse primers targeting the IgH J gene are selected from those listed in Table 5. In other embodiments, the FR3 and J gene amplification primer sets can target Ig light chain λ, Ig light chain κ, TCRα, TCRγ, TCRδ and TCRβ gene sequences.
[0071] In some embodiments, the concentration of the forward primer in the multiplex amplification reaction is approximately equal to the concentration of the reverse primer. In other embodiments, the concentration of the forward primer in the multiplex amplification reaction is approximately twice the concentration of the reverse primer. In other embodiments, the concentration of the forward primer in the multiplex amplification reaction is approximately half the concentration of the reverse primer. In some embodiments, the concentration of each primer targeting the FR region of the V gene is from about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is from about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is from about 50 nM to about 400 nM or from about 100 nM to about 500 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting the FR region of the V gene is from about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the J gene is from about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting the J gene is from about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the J gene is from about 50 nM to about 400 nM or from about 100 nM to about 500 nM. In some embodiments, the concentration of each primer targeting the J gene is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each primer targeting the J gene is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM. In some embodiments, the concentration of each primer targeting the J gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting the J gene is from about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is from about 5 nM to about 2000 nM. In some embodiments, the concentration of each primer targeting the C gene is from about 50 nM to about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is from about 50 nM to about 400 nM or from about 100 nM to about 500 nM.In some embodiments, the concentration of each primer targeting the C gene is about 200 nM, about 400 nM, about 600 nM, or about 800 nM. In some embodiments, the concentration of each primer targeting the C gene is about 5 nM, about 10 nM, about 50 nM, about 100 nM, or about 150 nM. In some embodiments, the concentration of each primer targeting the C gene is about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, the concentration of each primer targeting the C gene is from about 50 nM to about 800 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is about 50 nM, about 100 nM, about 200 nM, or about 400 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is from about 5 nM to about 2000 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is from about 50 nM to about 800 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is from about 50 nM to about 400 nM or from about 100 nM to about 500 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is about 600 nM, about 800 nM, about 1000 nM, about 1250 nM, about 1500 nM, about 1750 nM, or about 2000 nM. In some embodiments, in a multiplex reaction, the concentration of each forward and reverse primer is about 5 nM, about 10 nM, about 150 nM, or from about 50 nM to about 800 nM.
[0072] In some embodiments, the V gene FR and primers for the C gene target are combined in the form of an amplification primer pair to amplify the target immune receptor cDNA sequence and generate a target amplicon. Generally, the length of the target amplicon will depend on which V gene primer set (e.g., FR1, FR2, or primers for FR3) is paired with one or more C gene primers. Thus, in some embodiments, the length of the target amplicon can range from about 100 nucleotides (or bases or base pairs) to about 600 nucleotides (or bases or base pairs). In some embodiments, the length of the target amplicon can range from about 80 nucleotides to about 600 nucleotides. In some embodiments, the length of the target amplicon is about 200 to about 600 or about 300 to about 600 nucleotides. In some embodiments, the length of the target amplicon is about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides. In some embodiments, the length of the target amplicon is about 250 to about 275, about 250 to about 350, about 300 to about 350, about 310 to about 330, about 325 to about 375, about 300 to about 400, about 350 to about 400, about 350 to about 425, about 350 to about 450, about 380 to about 410, about 375 to about 425, about 400 to about 500, about 425 to about 500, about 450 to about 550, about 500 to about 600, about 400 to about 500, or about 400 to about 600 nucleotides. In some embodiments, the length of the target amplicon is about 80, about 100, about 120, about 140, about 200, about 250, about 275, about 300, about 320, about 350, about 375, about 400, about 425, about 450, about 500, about 550, or about 600 nucleotides. In some embodiments, the length of the IgH amplicon is about 100, about 80 to about 140, about 90 to about 130, or about 100 to about 120 nucleotides. In some embodiments, the length of the IgH amplicon is about 320, about 300 to about 350, or about 310 to about 330 nucleotides. In some embodiments, the length of the IgH amplicon is about 400, about 375 to about 425, or about 390 to about 410 nucleotides.
[0073] In some embodiments, the V gene FR and primers targeting J gene targets are combined in the form of an amplification primer pair to amplify the target immune receptor cDNA or rearranged gDNA sequence and generate a target amplicon. Generally, the length of the target amplicon will depend on which V gene primer set (e.g., FR1, FR2, or primers targeting FR3) is paired with the J gene primer. Thus, in some embodiments, the length of the target amplicon can range from about 50 nucleotides to about 350 nucleotides. In some embodiments, the length of the target amplicon is about 50 to about 200, about 70 to about 170, about 200 to about 350, about 250 to about 320, about 270 to about 300, about 225 to about 300, about 250 to about 275, about 200 to about 235, about 200 to about 250, or about 175 to about 275 nucleotides. In some embodiments, the length of the IgH amplicon is about 80, about 60 to about 100, or about 70 to about 90 nucleotides. In some embodiments, the length of the IgH amplicon, such as those generated using V gene FR3 and primers targeting the J gene, is about 50 to about 200 nucleotides, preferably about 60 to about 160, about 65 to about 120, about 90 to about 120, about 70 to about 90 nucleotides, or about 80 nucleotides. In some embodiments, generating such short-length amplicons enables the provided methods and compositions to effectively detect and analyze the immune repertoire from highly degraded gDNA template materials, such as template materials derived from FFPE samples or cell-free DNA (cfDNA).
[0074] In some embodiments, the amplification primer can comprise a barcode sequence, for example to distinguish or separate multiple amplified target sequences in a sample. In some embodiments, the amplification primer can comprise two or more barcode sequences, for example to distinguish or separate multiple amplified target sequences in a sample. In some embodiments, the amplification primer can comprise a tag sequence, which can assist in subsequent cataloging, identification, or sequencing of the generated amplicons. In some embodiments, the barcode sequence or tag sequence is incorporated into the amplified nucleotide sequence by inclusion in the amplification primer or by ligation of an adaptor. The primer can further include nucleotides that can be used for subsequent sequencing, such as pyrosequencing. Such sequences are readily designed by commercially available software programs or companies.
[0075] In some embodiments, multiplex amplification is performed using targeted amplification primers that do not contain tag sequences. In other embodiments, multiplex amplification is performed with amplification primers, each of which contains a targeting sequence and a tag sequence. For example, the forward primer or primer set includes tag sequence 1, and the reverse primer or primer set includes tag sequence 2. In still other embodiments, multiplex amplification is performed with amplification primers, where one primer or primer set contains a sequence for the target and a tag sequence, and the other primer or primer set contains a sequence for the target but does not contain a tag sequence. For example, the forward primer or primer set contains a tag sequence, and the reverse primer or primer set does not contain a tag sequence.
[0076] Thus, in some embodiments, multiple target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture with amplification primers for BCR and / or TCR, where the forward and / or reverse primers contain tag sequences, and the resulting amplicons contain target BCR and / or TCR sequences and tag sequences at one or both ends. In some embodiments, the forward and / or reverse amplification primers or primer sets may also contain barcodes, and then one or more barcodes are included in the resulting amplicons.
[0077] In some embodiments, multiple target cDNA or gDNA template molecules are amplified in a single multiplex amplification reaction mixture with amplification primers for BCR and / or TCR, and the resulting amplicons contain only BCR and / or TCR sequences. In some embodiments, tag sequences are added to the ends of such amplicons by, for example, adapter ligation. In some embodiments, barcode sequences are added to one or both ends of such amplicons by, for example, adapter ligation.
[0078] Nucleotide sequences suitable for use as barcodes and barcode libraries are known in the art. Adapters and amplification primers and primer sets containing barcode sequences are commercially available. Oligonucleotide adapters containing barcode sequences are also commercially available, including, for example, IonXpress TM 、IonCode TM and Ion Select barcode adapters (Thermo Fisher Scientific). Similarly, additional and other common adapter / primer sequences described and known in the art (e.g., Illumina common adapter / primer sequences, PacBio common adapter / primer sequences, etc.) can be used in combination with the methods and compositions provided herein, and the resulting amplicons are sequenced using the relevant analysis platforms.
[0079] In some embodiments, when sequencing multiplexed samples, two or more barcodes are added to the amplicons. In some embodiments, at least two barcodes are added to the amplicons prior to sequencing multiplexed samples to reduce the frequency of artifacts (e.g., immune receptor gene rearrangements or clonal discrimination) resulting from barcode cross-contamination or barcode leakage between samples. In some embodiments, when tracking low-frequency clones of the immune repertoire, at least two barcodes are used to label the samples. In some embodiments, when using an assay to detect clones with a frequency less than 1:1,000, at least two barcodes are added to the amplicons. In some embodiments, when using an assay to detect clones with a frequency less than 1:10,000, at least two barcodes are added to the amplicons. In other embodiments, when using an assay to detect clones with a frequency less than 1:20,000, less than 1:40,000, less than 1:100,000, less than 1:200,000, less than 1:400,000, less than 1:500,00, or less than 1:1,000,000, at least two barcodes are added to the amplicons. Methods for characterizing the immune repertoire that benefit from high sequencing depth per clone and / or clonal detection at such low frequencies include, but are not limited to, monitoring patients with hyperproliferative diseases during treatment and testing for minimal residual disease after treatment.
[0080] In some embodiments, target-specific primers (e.g., V gene FR1, FR2, and primers for FR3, primers for J gene, and primers for C gene) used in the methods of the invention are selected or designed to meet any one or more of the following criteria: (1) contain two or more modified nucleotides within the primer sequence, at least one of the nucleotides being near or at the end of the primer and at least one of the nucleotides being at or around the central nucleotide position of the primer sequence; (2) be about 15 to about 40 bases in length; (3) have a Tm of above 60°C to about 70°C; (4) have low cross-reactivity with non-target sequences present in the sample of interest; (5) at least the first four nucleotides (in the 3' to 5' direction) not be complementary to any sequence within any other primer present in the same reaction; and (6) not be complementary to any continuous segment of at least 5 nucleotides within any other resulting target amplicon. In some embodiments, target-specific primers used in the provided methods are selected or designed to meet any 2, 3, 4, 5, or 6 of the above criteria.
[0081] In some embodiments, the target-specific primers used in the methods of the present invention comprise one or more modified nucleotides having a cleavable group. In some embodiments, the target-specific primers used in the methods of the present invention comprise two or more modified nucleotides having a cleavable group. In some embodiments, the target-specific primer comprises at least one modified nucleotide having a cleavable group selected from the following: methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine or 5-methylcytidine.
[0082] In some embodiments, target amplicons using the amplification methods (and related compositions, systems and kits) disclosed herein are used to prepare immune receptor repertoire libraries. In some embodiments, the immune receptor repertoire library comprises introducing an adapter sequence to the ends of the target amplicon sequences. In certain embodiments, the method for preparing an immune receptor repertoire library comprises generating target immune receptor amplicon molecules according to any of the multiplex amplification methods described herein, processing the amplicon molecules by digesting modified nucleotides within the primer sequences of the amplicon molecules, and ligating at least one adapter to at least one processed amplicon molecule, thereby generating a library of adapter-ligated target immune receptor amplicon molecules comprising the target immune receptor repertoire. In some embodiments, the steps of preparing the library are performed in a single reaction vessel involving only addition steps. In certain embodiments, the method further comprises clonal amplification of a portion of at least one adapter-ligated target amplicon molecule.
[0083] In some embodiments, target amplicons using the methods (and related compositions, systems and kits) disclosed herein are coupled to downstream processes, such as but not limited to library preparation and nucleic acid sequencing. For example, bridge amplification, emulsion PCR or isothermal amplification can be used to amplify the target amplicons to generate multiple clonal templates suitable for nucleic acid sequencing. In some embodiments, the amplicon library is sequenced using any suitable DNA sequencing platform, such as any next-generation sequencing platform, including semiconductor sequencing technologies, such as the Ion Torrent sequencing platform. In some embodiments, the IonGeneStudio S5 540 TM system or Ion GeneStudio S5 520 TM system or Ion GeneStudio S5 530 TM system or Ion PGM 318 TM system is used to sequence the amplicon library.
[0084] In some embodiments, sequencing of the immune receptor amplicons generated using the methods (and related compositions and kits) disclosed herein yields contiguous sequence reads that are about 200 to about 600 nucleotides in length. In some embodiments, the contiguous read length is about 300 to about 400 nucleotides. In some embodiments, the contiguous read length is about 350 to about 450 nucleotides. In some embodiments, the average read length is about 300 nucleotides, about 350 nucleotides, or about 400 nucleotides. In some embodiments, the contiguous read length is about 250 to about 350 nucleotides, about 275 to about 340, or about 295 to about 325 nucleotides. In some embodiments, the average read length is about 270, about 280, about 290, about 300, or about 325 nucleotides. In other embodiments, the contiguous read length is about 180 to about 300 nucleotides, about 200 to about 290 nucleotides, about 225 to about 280 nucleotides, or about 230 to about 250 nucleotides. In some embodiments, the average read length is about 200, about 220, about 230, about 240, or about 250 nucleotides. In other embodiments, the contiguous read length is about 70 to about 200 nucleotides, about 80 to about 150 nucleotides, about 90 to about 140 nucleotides, or about 100 to about 120 nucleotides in length. In other embodiments, the contiguous read length is about 50 to about 170 nucleotides, about 60 to about 160 nucleotides, about 60 to about 120 nucleotides, about 70 to about 100 nucleotides, about 70 to about 90 nucleotides, or about 80 nucleotides in length. In some embodiments, the average read length is about 70, about 80, about 90, about 100, about 110, or about 120 nucleotides. In some embodiments, the sequence read length includes the amplicon sequence and the barcode sequence. In some embodiments, the sequence read length does not include the barcode sequence.
[0085] In some embodiments, the amplicon primers and primer pairs are target-specific sequences that can amplify a specific region of a nucleic acid molecule. In some embodiments, the target-specific primers can amplify expressed RNA or cDNA. In some embodiments, the target-specific primers can amplify mammalian RNA, such as human RNA or cDNA prepared therefrom, or murine RNA or cDNA prepared therefrom. In some embodiments, the target-specific primers can amplify DNA, such as gDNA. In some embodiments, the target-specific primers can amplify mammalian DNA, such as human DNA or murine DNA.
[0086] In the methods and compositions provided herein, such as those for assaying, characterizing, and / or tracking the immune repertoire in a biological sample, the amount of input RNA or gDNA required to amplify a target sequence will depend in part on the fraction of cells bearing immune receptors (e.g., T cells or B cells) in the sample. For example, a higher fraction of B cells in a sample such as a B cell-enriched sample allows for the use of lower amounts of input RNA or gDNA for amplification. In some embodiments, the amount of input RNA for amplifying one or more target sequences can be from about 0.05 ng to about 10 micrograms. In some embodiments, the amount of input RNA for multiplex amplification of one or more target sequences can be from about 5 ng to about 2 micrograms. In some embodiments, the amount of RNA for multiplex amplification of one or more target sequences can be from about 5 ng to about 1 microgram or from about 10 ng to about 1 microgram. In some embodiments, the amount of RNA for multiplex amplification of one or more immune repertoire target sequences is about 1.5 micrograms, about 2 micrograms, about 2.5 micrograms, about 3 micrograms, about 3.5 micrograms, about 4.0 micrograms, about 5 micrograms, about 6 micrograms, about 7 micrograms, or about 10 micrograms. In some embodiments, the amount of RNA for multiplex amplification of one or more immune repertoire target sequences is about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of RNA for multiplex amplification of one or more immune repertoire target sequences is from about 25 ng to about 500 ng RNA or from about 50 ng to about 200 ng RNA. In some embodiments, the amount of RNA for multiplex amplification of one or more immune repertoire target sequences is from about 0.05 ng to about 10 ng RNA, from about 0.1 ng to about 5 ng RNA, from about 0.2 ng to about 2 ng RNA, or from about 0.5 ng to about 1 ng RNA. In some embodiments, the amount of RNA for multiplex amplification of one or more immune repertoire target sequences is about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0087] As described herein, prior to multiplex amplification, reverse transcriptase is typically used in a reverse transcription reaction to convert RNA from a biological sample into cDNA. In some embodiments, the reverse transcription reaction is performed with the input RNA, and a portion of the cDNA from the reverse transcription reaction is used in the multiplex amplification reaction. In some embodiments, substantially all of the cDNA prepared from the input RNA is added to the multiplex amplification reaction. In other embodiments, a portion of the cDNA prepared from the input RNA, such as about 80%, about 75%, about 66%, about 50%, about 33%, or about 25% is added to the multiplex amplification reaction. In other embodiments, about 15%, about 10%, about 8%, about 6%, or about 5% of the cDNA prepared from the input RNA is added to the multiplex amplification reaction.
[0088] In some embodiments, the amount of cDNA from a sample added to the multiplex amplification reaction can be from about 0.001 ng to about 5 micrograms. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences can be from about 0.01 ng to about 2 micrograms. In some embodiments, the amount of cDNA used for multiplex amplification of one or more target sequences can be about 0.1 ng to about 1 microgram or about 1 ng to about 0.5 microgram. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 25 ng, about 50 ng, about 100 ng, about 200 ng, about 250 ng, about 500 ng, about 750 ng, or about 1000 ng. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is from about 0.01 ng to about 10 ng cDNA, from about 0.05 ng to about 5 ng cDNA, from about 0.1 ng to about 2 ng cDNA, or from about 0.01 ng to about 1 ng cDNA. In some embodiments, the amount of cDNA used for multiplex amplification of one or more immune repertoire target sequences is about 0.005 ng, about 0.01 ng, about 0.05 ng, about 0.1 ng, about 0.2 ng, about 0.5 ng, about 1.0 ng, about 2.0 ng, or about 5.0 ng.
[0089] In some embodiments, mRNA is obtained from a biological sample using conventional methods and converted to cDNA for amplification purposes. Methods and reagents for extracting or isolating nucleic acids from biological samples are well known and commercially available. In some embodiments, RNA is extracted from a biological sample by any method described herein or otherwise known to those of skill in the art, such as methods involving proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with DNAse to digest contaminating DNA, and RNA purification using silica membrane technology or any combination thereof. Exemplary methods for extracting RNA from biological samples use commercially available kits, including RecoverAll TM Multi-Sample RNA / DNA Workflow (Invitrogen), RecoverAll TM Total Nucleic Acid Isolation Kit (Invitrogen), Blood (Macherey-Nagel), Blood RNA System, TRI Reagent TM (Invitrogen), PureLink TM RNA Microscale Kit (Invitrogen), MagMAX TM FFPE DNA / RNA Ultra Kit (Applied Biosystems), ZR RNA MicroPrep TM Kit (ZymoResearch), RNeasy Micro Kit (Qiagen), and ReliaPrep TM RNA Tissue Miniprep System (Promega).
[0090] In some embodiments, the amount of input gDNA for amplifying one or more target sequences can be from about 0.1 ng to about 10 micrograms. In some embodiments, the amount of gDNA required for amplifying one or more target sequences can be from about 0.5 ng to about 5 micrograms. In some embodiments, the amount of gDNA required for amplifying one or more target sequences can be from about 1 ng to about 1 microgram or from about 10 ng to about 1 microgram. In some embodiments, the amount of gDNA required for amplifying one or more immune repertoire target sequences is from about 10 ng to about 500 ng, from about 25 ng to about 400 ng, or from about 50 ng to about 200 ng. In some embodiments, the amount of gDNA required for amplifying one or more target sequences is about 0.5 ng, about 1 ng, about 5 ng, about 10 ng, about 20 ng, about 50 ng, about 100 ng, or about 200 ng. In some embodiments, the amount of gDNA required for amplifying one or more immune repertoire target sequences is about 1 microgram, about 2 micrograms, about 3 micrograms, about 4.0 micrograms, or about 5 micrograms.
[0091] In some embodiments, gDNA is obtained from a biological sample using conventional methods. Methods and reagents for extracting or isolating nucleic acids from biological samples are well known and commercially available. In some embodiments, DNA is extracted from a biological sample by any method described herein or otherwise known to those skilled in the art, such as methods involving proteinase K tissue digestion and alcohol-based nucleic acid precipitation, treatment with RNAse to digest contaminating RNA, and DNA purification using silica-membrane technology, or any combination thereof. Exemplary methods for extracting DNA from biological samples use commercially available kits, including the FFPE DNA kit for Ion AmpliSeq TM , the MagMAX TM FFPE DNA / RNA Ultra kit, TRIReagent TM (Invitrogen), the PureLink TM Genomic DNA Mini kit (Invitrogen), the RecoverAll TM Total Nucleic Acid Isolation kit (Invitrogen), the MagMAX TM DNA Multi-Sample kit (Invitrogen), and DNA extraction kits from BioChain Institute Inc. (e.g., FFPE Tissue DNA Extraction kit, Genomic DNA Extraction kit, Blood and Serum DNA Isolation kit).
[0092] A sample or biological sample as used herein refers to a composition from an individual containing or potentially containing cells related to the immune system. Exemplary biological samples include, but are not limited to, tissues (e.g., lymph nodes, organ tissue, bone marrow), whole blood, synovial fluid, cerebrospinal fluid, tumor biopsies, and other cellular clinical specimens. The sample can contain normal cells and / or diseased cells, and can be a fine needle aspirate, fine needle biopsy, core sample, or other sample. In some embodiments, the biological sample can include hematopoietic cells, peripheral blood mononuclear cells (PBMCs), T cells, B cells, tumor infiltrating lymphocytes (“TILs”), or other lymphocytes. In some embodiments, the sample can be fresh (e.g., not preserved), frozen, or formalin-fixed paraffin-embedded tissue (FFPE). Some samples include cancer cells, such as carcinoma, melanoma, sarcoma, lymphoma, myeloma, leukemia, etc., and the cancer cells can be circulating tumor cells. In some embodiments, the biological sample includes cfDNA as found, for example, in blood or plasma.
[0093] The biological sample can be a mixture of tissue or cell types, a cell preparation enriched for at least one specific category or cell type, or a separated cell population of a specific type or phenotype. Prior to analysis, the sample can be separated by centrifugation, elutriation, density gradient separation, apheresis, affinity selection, panning, FACS, centrifugation with Hypaque, etc. Methods for sorting, enriching, and separating specific cell types are well known and can be readily performed by one of ordinary skill in the art. In some embodiments, the sample can be a preparation enriched for B cells.
[0094] In some embodiments, the provided methods and systems include methods for analyzing immune repertoire receptor cDNA or gDNA sequence data and for identifying and / or removing one or more PCR- or sequencing-derived errors from the determined immune receptor sequences.
[0095] In some embodiments, the error correction strategy includes the steps of:
[0096] 1) Align the sequenced rearrangements to a reference database of variable, diversity, and joining / constant genes to generate query sequence / reference sequence pairs. Many alignment programs can be used for this purpose, including, for example, IgBLAST, a freely available tool from NCBI, and custom computer scripts.
[0097] 2) Re-align the reference and query sequences to each other, taking into account the flow order used for sequencing. The flow order provides information that allows one to identify and correct certain types of misalignments.
[0098] 3) Identify the boundaries of the CDR3 region by its characteristic sequence motifs.
[0099] 4) In the alignment portion corresponding to the rearrangement of the variable gene and the joining / constant gene (excluding the CDR3 region), insertions and deletions are made in the query relative to the reference identity and the positions of mismatched query bases are changed to be consistent with the reference.
[0100] 5) For the CDR3 region, if the CDR3 length is not a multiple of three (indicating an indel error):
[0101] (a) Based on the PHRED score (denoted as e), search for the homopolymer extension with the highest probability of containing a sequence error in the CDR3.
[0102] (b) Obtain the error probability for the entire CDR3 region based on the PHRED score (denoted as t)
[0103] (c) If e / t is greater than a determined threshold, edit the homopolymer by increasing or decreasing the length of the homopolymer by one base so that the CDR3 nucleotide length is a multiple of three.
[0104] (d) As an alternative to steps a - c, search for the longest homopolymer in the CDR3 and, if the length of the homopolymer is above a defined threshold, edit the homopolymer by increasing or decreasing the length of the homopolymer by one base so that the CDR3 nucleotide length is a multiple of three.
[0105] In some embodiments, methods are provided for identifying B - cell and / or T - cell clones in repertoire data that is robust to PCR and sequencing errors. Accordingly, steps that can be used in such methods to identify B - cell and / or T - cell clones in a manner that is robust to PCR and sequencing errors are described below. Table 1 is an exemplary workflow for identifying and removing PCR - or sequencing - derived errors from immune receptor sequencing data. Exemplary portions and embodiments of this workflow are also presented in Figures 1 - 2 are presented.
[0106] Table 1: Sequence correction workflow
[0107]
[0108] For a set of TCR or BCR sequences derived from mRNA or gDNA, where 1) each sequence has been annotated as a productive rearrangement, either natively or after error correction, as previously described, and 2) each sequence has an identified V gene and CDR3 nucleotide region, in some embodiments, the method comprises the following:
[0109] 1) Identify and exclude chimeric sequences. For each unique CDR3 nucleotide sequence present in the dataset, calculate the number of reads having that CDR3 nucleotide sequence and any possible V gene. Any V gene-CDR3 combination that constitutes less than 10% of the total reads for that CDR3 nucleotide sequence is labeled as chimeric and eliminated from downstream analysis. For example, for the following sequences having the same CDR3 nucleotide sequence, sequences having TRBV3 and TRBV6 paired with the CDR3nt sequence AATTGGT, for instance, will be labeled as chimeric.
[0110] V gene CDR3nt Read count TRBV2 AATTGGT 1000 TRBV3 AATTGGT 10 TRBV6 AATTGGT 3
[0111] 2) Identify and exclude sequences containing simple indel errors. For each read in the dataset, obtain the homopolymer-collapsed representation of the CDR3 sequence of that read. For each group of reads having the same V gene and collapsed-CDR3 combination, calculate the number of occurrences of each uncollapsed CDR3 nucleotide sequence. Any uncollapsed CDR3 sequence that constitutes <10% of the total reads for that set of reads is labeled as having a simple homopolymer error. As an example, three different V gene-CDR3 nucleotide sequences are presented that are identical after homopolymer collapse of the CDR3 nucleotide sequence. The two less frequent V gene-CDR3 combinations constitute <10% of the total reads for the set of reads and will be labeled as containing simple indel errors. For example:
[0112]
[0113]
[0114] 3) Identify and exclude singleton reads. For each read in the dataset, record the number of times the exact read sequence is found in the dataset. Reads that occur only once in the dataset will be labeled as singleton reads.
[0115] 4) Identify and exclude truncated reads. For each read in the dataset, determine whether the read has the annotated V gene FR1, CDR1, FR2, CDR2, and FR3 regions as indicated by IgBLAST alignment of the read against the IgBLAST reference V gene set. Reads that do not have the above regions will be labeled as truncated if the regions are expected based on the specific V gene primers used for amplification.
[0116] 5) Identify and exclude rearrangements lacking bidirectional support. For each read in the dataset, obtain the V gene and CDR3 sequence of the read and the strand orientation (positive or negative strand) of the read. For each V gene-CDR3 combination in the dataset, calculate the number of positive and negative strand reads with that V gene-CDR3 nt combination. V gene-CDR3 nt combinations present only in reads from one orientation will be considered false. All reads with false V gene-CDR3 nt combinations will be labeled as lacking bidirectional support.
[0117] 6) For unlabeled genes, perform stepwise clustering based on CDR3 nucleotide similarity. Divide the sequences into groups based on the V gene identity of the reads, excluding allelic information (v gene set). For each group:
[0118] a. Use cd-hit-est and the following parameters to cluster the reads in each group:
[0119] cd-hit-est -i vgene_groups.fa -o clustered_vgene_groups.cdhit -T 24 -d 0 -M 100000 -B 0 -r 0 -g 1 -S 0 -U 2 -uL.05 -n 10 –l 7. (The freely available software program cd-hit-est clusters nucleotide datasets into clusters that meet a user-defined similarity threshold). (For the code and instructions for cd-hit-est, see https: / / github.com / weizhongli / cdhit / wiki / 3.-User%27s-Guide#CDHITEST).
[0120] where vgene_groups.fa is a fasta format file of the CDR3 nucleotide regions of sequences with the same V gene, and clustered_vgene_groups.cdhit is the output containing the subdivided sequences.
[0121] b. Assign the same clone ID to each sequence in the cluster, indicating that members of the subgroup are considered to represent the same T cell clone or B cell clone.
[0122] c. Select the representative sequence for each cluster such that the representative sequence is the sequence that appears the most number of times, or in case of a tie, is randomly selected.
[0123] d. Merge all other reads in the cluster into the representative sequence such that the number of reads of the representative sequence increases according to the number of reads of the merged sequences.
[0124] e. Compare the representative sequences within the v genome to each other based on the Hamming distance. If a representative sequence is within Hamming distance 1 of a representative sequence that is >50-fold more abundant, merge the sequence into the more common representative sequence. If a representative sequence is within Hamming distance 2 of a representative sequence that is >10,000-fold more abundant, merge the sequence into the more common representative sequence.
[0125] f. Identify compound sequence errors. Perform homopolymer-collapse on the representative sequences within each V genome, then compare them to each other using the Levenshtein distance. If a representative sequence is within 1 Levenshtein distance of a representative sequence that is >50-fold more abundant, merge the sequence into the more common representative sequence.
[0126] g. Identify CDR3 misannotation errors. Perform homopolymer-collapse on the representative sequences within each V genome, then perform pairwise comparisons on each homopolymer-collapsed sequence. For each pair of sequences, determine whether one sequence is a subset of the other. If so, if the more abundant sequence is >500-fold more abundant, merge the less abundant sequence into the more abundant sequence.
[0127] 7) Report the clustering representatives to the user.
[0128] In some embodiments, step 6 of the above workflow divides the rearranged sequences into groups based on V gene identity (excluding allele information) and CDR3 nucleotide length. In other embodiments, J gene identity and / or isotype identity are also used as part of the grouping criteria. Thus, in some embodiments, step 6 of the above workflow includes the following steps:
[0129] a. Use cd-hit-est and the following parameters to cluster the reads within each group:
[0130] cd-hit-est -i vgene_groups.fa -o clustered_vgene_groups.cdhit -T 24 -l 9 -d 0 -M 100000 -B 0 -r 0 -g 1 -S 15 -U 2 -uL.05 -n 9.
[0131] Where vgene_groups.fa is a fasta format file of the sequenced portion of the VDJ rearrangement.
[0132] In some embodiments, the full sequence of the VDJ is considered for clustering because somatic hypermutation can occur throughout the VDJ region.
[0133] b. Assign the same clone ID to each sequence in a cluster, where the clone ID is used to indicate that members of a subgroup are considered to represent the same T cell clone or B cell clone.
[0134] c. Select a representative sequence for each cluster such that the representative sequence is the sequence that occurs the most number of times, or in case of a tie, is randomly selected.
[0135] d. Merge all other reads in the cluster into the representative sequence such that the read count of the representative sequence is increased according to the read count of the merged sequences.
[0136] e. Compare the representative sequences within a v genome based on Hamming distance. If a representative sequence is within Hamming distance 1 of a representative sequence that is >50-fold more abundant, then merge the sequence into the more common representative sequence. If a representative sequence is within Hamming distance 2 of a representative sequence that is >10,000-fold more abundant, then merge the sequence into the more common representative sequence. In some embodiments, multiple thresholds of >50 / 3 and >10,000 / 3 are used to merge sequences with Hamming distances of 1 or 2, respectively. Reducing the multiple threshold may be useful when comparing sequences of the entire VDJ region rather than just the CDR3 region, as longer sequences are more likely to accumulate amplification and / or sequencing errors.
[0137] f. Identify compound sequence errors. Perform homopolymer-collapse on the representative sequences within each v genome and then compare them to each other using the Levenshtein distance. If a representative sequence is within 1 Levenshtein distance of a representative sequence that is >50-fold more abundant, then merge the sequence into the more common representative sequence.
[0138] g. Identify CDR3 misannotation errors. Perform homopolymer-collapse on the representative sequences within each v genome and then perform pairwise comparisons on each homopolymer-collapsed sequence. For each pair of sequences, determine whether one sequence is a subset of the other. If so, if the more abundant sequence is >500-fold more abundant, then merge the less abundant sequence into the more abundant sequence.
[0139] In some embodiments, the provided workflows are not limited to the frequency ratio thresholds listed in the respective steps, and other frequency ratio thresholds may be used in place of the representative frequency ratio thresholds included above. The frequency ratio refers to the ratio of the abundance value of the more common representative sequence to the abundance value of the less common representative sequence. The frequency ratio threshold gives a threshold for merging the less common representative sequence into the more common representative sequence. For example, in some embodiments, comparing representative sequences within the v genome based on Hamming distance may use frequency ratios other than those listed in step (e) above. For example but not limited to, if one representative sequence is within Hamming distance 2 of another representative sequence, frequency ratio thresholds such as 1000, 5000, 20,000, etc. may be used. For example but not limited to, if one representative sequence is within Hamming distance 1 of another representative sequence, frequency ratio thresholds such as 20, 100, 200, etc. may be used. The provided frequency ratio thresholds represent the general process of labeling the more abundant sequences of a similar pair as the correct sequence.
[0140] Similarly, when comparing the frequencies of two sequences in other steps of the workflow (such as step (1), step (2), step (6f), and step (6g)), frequency ratios other than those listed in the above steps may be used.
[0141] The term "homopolymer-collapsed sequence" as used herein is intended to represent a sequence in which repeated bases are collapsed into a single base representation. As an example, for the non-collapsed sequence AAAATTTTTATCCCCCCCCGGG (SEQ ID NO:603), the homopolymer-collapsed sequence is ATATCG.
[0142] The terms "clone", "clonotype", "lineage", or "rearrangement" as used herein are intended to describe the unique V gene nucleotide combinations for immune receptors such as TCR or BCR. For example, unique V gene - CDR3 nucleotide combinations.
[0143] The term "productive read" as used herein refers to a TCR or BCR sequence read that does not have a stop codon and has an in-frame variable gene segment and joining gene segment. When encoding a polypeptide, a productive read is biologically reasonable.
[0144] As used herein, "chimera" or "chimeric sequence" refers to an artificial sequence generated by template switching during target amplification such as PCR. Chimeras typically exist as CDR3 sequences transplanted onto unrelated V genes, resulting in CDR3 sequences associated with multiple V genes within the dataset. Chimeric sequences are typically far less abundant than the true sequences in the dataset.
[0145] As used herein, the term "indel" refers to the insertion and / or deletion of one or more nucleotide bases in a nucleic acid sequence. In the coding region of a nucleic acid sequence, unless the length of the indel is a multiple of 3, it will result in a frameshift when the sequence is translated. As used herein, a "simple indel error" is an error that does not change the homopolymer-collapsed representation of the sequence. As used herein, a "complex indel error" is an indel sequencing error that changes the homopolymer-collapsed representation of the sequence and includes, but is not limited to, errors that eliminate homopolymers, insert homopolymers into the sequence, or create reading-frame-disrupting errors.
[0146] As used herein, a "singleton read" refers to a sequence read whose corrected sequence appears only once in the dataset. Typically, singleton reads are enriched for reads containing PCR or sequencing errors.
[0147] As used herein, a "truncated read" refers to an immunoreceptor sequence read that lacks the annotated V gene region. For example, truncated reads include, but are not limited to, sequence reads that lack the annotated TCR or BCR V gene FR1, CDR1, FR2, CDR2, or FR3 regions. Due to quality trimming, such reads typically lack a portion of the V gene sequence. If the truncation results in misidentification of the V gene, truncated reads may produce artifacts.
[0148] In the context of an identified V gene-CDR3 sequence (clonotype), "bidirectional support" means that a particular V gene-CDR3 sequence is found in at least one read mapped to the positive strand (going from the V gene to the constant gene) and at least one read mapped to the negative strand (going from the constant gene to the V gene). Systematic sequencing errors typically result in the identification of V gene-CDR3 sequences with unidirectional support.
[0149] For a set of sequences that have been grouped based on a predefined sequence similarity threshold to account for variations due to PCR or sequencing errors, a "cluster representative" is the sequence that is selected as being most likely error-free. This is typically the most abundant sequence.
[0150] As used herein, an "IgBLAST annotation error" refers to a rare event in which the boundaries of the CDR3 are identified as being in incorrect adjacent positions. These events typically add three bases to the 5' or 3' end of the CDR3 nucleotide sequence.
[0151] For two sequences of equal length, the "Hamming distance" is the number of positions at which the corresponding bases or amino acids are different. For any two sequences, the "Levenshtein distance" or "edit distance" is the number of single-base or amino acid edits required to make one nucleotide or amino acid sequence into the other nucleotide or amino acid sequence.
[0152] In some embodiments where primers directed to the J gene are used to amplify immune receptor sequences (e.g., multiplex amplification with primers directed to the FR3 region of the V gene and primers directed to the J gene), a J gene sequence inference process is performed on the raw sequence reads derived from the assay prior to any downstream analysis. In this process, it is interrogated whether a characteristic sequence of 10 - 30 nucleotides exists at the start and end of the raw read sequence, which corresponds to the portion of the J gene sequence expected to be present after amplification with the J primer and any subsequent manipulation or processing (e.g., digestion) at the ends of the amplicons prior to sequencing. The characteristic nucleotide sequence allows for the inference of the J primer sequence and the remaining portion of the J gene that is targeted as the sequence of each J gene is known. To complete the J gene sequence inference process, the inferred J gene sequence is added to the raw read to produce an extended read that then spans the entire J gene. The extended read then contains the entire J gene sequence, the complete sequence of the CDR3 region, and at least a portion of the V gene sequence, which will be reported after downstream analysis. The portion of the V gene sequence in the extended read will depend on the primers directed to the V gene used for multiplex amplification, such as primers for FR3, FR2, or FR1.
[0153] Amplification of expressed immune receptor sequences or rearranged immune receptor gDNA sequences using V gene FR3 and J gene primers produces amplicons of a minimum length (e.g., length of about 60 - 100 or about 80 nucleotides), while still producing data that allows for reporting of the entire CDR3 region. Due to the expected short amplicon length, reads of amplicons with a length of <100 nucleotides are not eliminated as low-quality and / or off-target products during the sequence analysis workflow. However, explicit search for the expected J gene sequence in the raw reads allows for the elimination of amplicons derived from off-target amplification by the J gene primers. Additionally, this short amplicon length improves the performance of assays on highly degraded template materials (e.g., template materials derived from FFPE or cfDNA samples).
[0154] In some embodiments, the provided methods include sequencing an immune receptor library and subjecting the obtained sequence data to an error discrimination and correction process to produce rescued productive reads, and discriminating productive sequence reads and rescued productive sequence reads. In some embodiments, the provided methods include sequencing an immune receptor library and subjecting the obtained sequence data set to an error discrimination and correction process, discriminating productive sequence reads and rescued productive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes in the library.
[0155] In some embodiments, the provided method includes sequencing a rearranged immune receptor DNA library and, for the V gene segment, subjecting the obtained sequence data to an error discrimination and correction process to generate rescued productive reads, and discriminating productive sequence reads, rescued productive sequence reads, and unproductive sequence reads. In some embodiments, the provided method includes sequencing a rearranged immune receptor DNA library and performing an error discrimination and correction process on the obtained sequence data set for the V gene segment, discriminating productive, rescued productive, and unproductive sequence reads, and grouping the sequence reads by clonotype to identify immune receptor clonotypes in the library. In some embodiments, both productive and unproductive sequence reads of the rearranged immune receptor DNA are reported separately.
[0156] In some embodiments, the provided error discrimination and correction workflow is used to identify and resolve PCR- or sequencing-derived errors that cause sequence reads to be identified as from unproductive rearrangements. In some embodiments, the provided error discrimination and correction workflow is applied to immune receptor sequence data generated from a sequencing platform, where indels or other frameshift-causing errors occur during the generation of the sequence data.
[0157] In some embodiments, the provided error discrimination and correction workflow is applied to sequence data generated by the Ion Torrent sequencing platform. In some embodiments, the provided error discrimination and correction workflow is applied to sequence data generated by the Roche 454 Life Sciences sequencing platform, the Pacific Biosciences sequencing platform, and the Oxford Nanopore sequencing platform.
[0158] In some embodiments, the BCR repertoire analysis workflow includes an additional final step of identifying the clonal lineages in the sample. A clonal lineage represents a set of B cell clones (e.g., identified as having unique VDJ sequences) that are derived from a common VDJ rearrangement but differ due to somatic hypermutation and / or class switch recombination. It is generally assumed that members of a clonal lineage are more likely to target the same antigen than members of different clonal lineages.
[0159] In some embodiments, the process of clonal lineage identification includes using a set of BCR clones (e.g., IgH clones) identified (e.g., as described herein) to perform the following:
[0160] 1. Divide the clonal sequences into groups, where group members share the same variable gene (excluding allele information), the same CDR3 nucleotide length, and the same joining gene (excluding allele information). In some embodiments, the above J gene criterion may be omitted.
[0161] 2. Based on the CDR3 nucleotide similarity of the cloned sequences, the cloned sequences in each group are arranged into clusters. The threshold of the CDR3 nucleotide similarity is from about 0.70 to about 0.99. In some embodiments, the threshold of the CDR3 nucleotide similarity is between about 0.80 and about 0.99. In some embodiments, the threshold of the CDR3 nucleotide similarity is between about 0.80 and about 0.90. In certain embodiments, the threshold of the CDR3 nucleotide similarity is about 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98 or 0.99.
[0162] a. In some embodiments, clustering is performed using cd-hit-est as described below:
[0163] cd-hit-est -i vgene_groups.fa -o clustered_vgene_groups.cdhit -T 24 -l 9 -d 0 -M 100000 -B 0 -r 0 -g 1 -S 0 -c.85 -n 5, where vgene_groups.fa consists of a set of CDR3 nucleotide sequences of each clone within a group. Clones within the same cluster are considered members of the same clone lineage.
[0164] b. In some cases, somatic hypermutation may be extensive enough that the described clustering criteria may not group all members of a clone lineage. For such cases, in some embodiments, additional steps are performed to merge the clusters identified in (a). The additional steps consist of: searching for instances of somatic hypermutation-derived mutations shared in the variable genes between clone lineages, and then merging the clone lineages if the fraction and / or number of shared mutations is above a certain threshold. As described above, variable gene mutations are identified by comparing the variable gene sequences with the most matching variable gene sequences in the IMGT database. In some embodiments, the threshold for the number of shared mutations is 2 or greater. In some embodiments, the threshold for the number of shared mutations is 3 or greater. In other embodiments, the threshold for the number of shared mutations is 4, 5, 6, 7, 8, 9, 10 or greater. In some embodiments, the fraction of shared mutations is from about 0.15 to about 0.95. In some embodiments, the fraction of shared mutations is about 0.75 or about 0.85. In other embodiments, the fraction of shared mutations is about 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or 0.95.
[0165] In some cases, variable gene alleles that are not expressed in the IMGT database can be identified. In such cases, the alignment with the IMGT database will indicate mismatches that are not due to somatic hypermutation. To avoid noise caused by such unannotated gene variants, in some embodiments, an initial step is performed before (b), where all putative novel variable gene alleles in the sample are identified, noting each position that differs from the reference. In some embodiments, such positions are then not considered in the analysis described in (b). For example, Gadala-Maria et al. (2015) Proceedings of the National Academy of Sciences of the United States of America 112: E862-E870 and PCT application publication no. WO2018 / 136562 have described methods for identifying novel alleles from immunoglobulin repertoire sequencing data.
[0166] At the end of this clonal lineage identification process, each clone has been assigned to a clonal lineage. The clonal lineages can be used as the unit of analysis to calculate characteristics of the BCR repertoire, such as diversity, evenness, and convergence. In some embodiments, clonal lineage characteristics, such as the number of clones belonging to the lineage, the isotypes of those clones, the maximum and minimum frequencies of clones in the lineage, the maximum and minimum variable gene somatic hypermutations in the lineage, etc., are calculated and reported to the user.
[0167] In the absence of somatic hypermutation, BCR convergence can be calculated as the frequency of clones that are amino acid sequence identical or functionally identical but have different nucleotide sequences. These represent clones that have independently undergone VDJ recombination and are generally assumed to have proliferated in response to a common antigen. However, somatic hypermutation can generate different VDJ sequences that do not represent B cells that have independently undergone VDJ recombination. To address this situation, a definition of convergence is used that takes into account clonal lineage identification. For this purpose, "BCR convergence" is defined as the frequency of B cell clones that are members of different clonal lineages as described above but have similar or identical amino acid sequences. In some embodiments, if two IGH rearrangements are assigned to separate clonal lineages but have the same variable gene (without allele information) and the same or similar CDR3 amino acid sequence, the rearrangements are considered to be convergent. In other embodiments where the sequencing covers all three CDR domains of the IGH chain, if two IGH rearrangements are assigned to separate clonal lineages but have the same variable gene (without allele information) and the same or similar CDR1, 2, and 3 amino acid sequences, the rearrangements can be considered to be convergent. In some embodiments, similar CDR amino acid sequences are within a Hamming or Levenshtein edit distance of 1. In other embodiments, similar CDR amino acid sequences are within a Hamming or Levenshtein edit distance of 2.
[0168] Thus, in some embodiments, functionally equivalent B cells are identified by searching for BCR clones having the same variable gene and CDR amino acid sequences within a Hamming or Levenshtein edit distance of 1 or 2. In some embodiments, the program cd-hit can be used to identify clones having similar but functionally equivalent amino acid sequences. (For the code and information about the program cd-hit, see https: / / github.com / weizhongli / cdhit / wiki / 3.-User%27s-Guide) In some embodiments, cd-hit is run using the following command:
[0169] cd-hit -i vgene_groups.fa -o clustered_vgene_groups.cdhit -T 24 -l 5 -d 0 -M 100000 -B 0 -g 1 -S 1 -U 1 -n 5, where vgene_groups.fa consists of a set of CDR3 amino acid sequences of clones having the same variable gene. Clones within the same cluster are considered to be functionally equivalent.
[0170] In some embodiments, the value of the parameter -S can be 0, 1, 2, or 3. In some embodiments, the value of the parameter -U can be 0, 1, 2, or 3.
[0171] In some embodiments, vgene_groups.fa consists of a set of CDR1, 2, and 3 amino acid sequences of clones having the same variable gene. In some embodiments, vgene_groups.fa consists of a set of clones having both the same variable gene and the same CDR3 length.
[0172] In some embodiments, the provided sequence analysis workflow includes subsampling analysis. For immunosequencing and subsequent analysis, using subsampling analysis can help, for example, to eliminate variability due to differences in sequencing depth across assays. For example, an exemplary subsampling analysis for use with RNA or cDNA sequencing and analysis workflows applies the following procedures to the data: a) starting from the total set of productive + rescued productive reads, removing sequence reads randomly until one of several read depths; and b) performing all downstream calculations (e.g., clonotype and calculation of secondary repertoire characteristics, the secondary repertoire characteristics including but not limited to evenness, convergence, diversity, number of detected clones and identities, and clone lineages) using this subset of reads.
[0173] In some embodiments, downsampling analysis identifies the point at which a particular sample has been sequenced to saturation, e.g., the point at which additional reads do not identify additional clones or lineages or add additional diversity to the detected repertoire. In some embodiments, downsampling allows for refinement or multiplexing of sequencing depth between or among assays performed using similar sample types.
[0174] In some embodiments, a set of variable gene alleles detected by the provided assay methods and compositions can be used to re-identify haplotype groups within a human population. In certain embodiments, the provided assay methods and compositions that include using multiple V gene-specific primers and at least one C gene-specific primer to amplify IgH CDR1, 2, and 3 nucleotide sequences can be used to identify the IgH haplotype of a subject's BCR repertoire. For example, in some embodiments, the provided methods and compositions can be used to identify the IgH haplotype of a subject's BCR repertoire, the methods and compositions using at least one set of primers that includes multiple V gene FR1 primers selected from Table 3 and at least one C gene primer selected from Tables 6-10. Methods for identifying TCR haplotype groups are described in PCT Application No. PCT / US2019 / 023731, filed Mar. 22, 2019, the entire contents of which are incorporated herein by reference and can be similarly used in combination with the methods and compositions provided herein to identify IgH haplotype groups. In some embodiments, a set of variable gene alleles detected by amplifying and sequencing IgH CDR1, 2, and 3 nucleotide sequences can be used to assign a sample to one of a number of pre-existing haplotype groups as part of a larger procedure for predicting the risk of autoimmune disease or adverse events following immunotherapy. In the procedure for predicting the risk of autoimmune disease or adverse events following immunotherapy, the method for assigning a sample to a haplotype group is also described in PCT Application No. PCT / US2019 / 023731, filed Mar. 22, 2019 and incorporated herein by reference, and can be similarly used in combination with the methods and compositions provided herein to assign a sample to an IgH haplotype group, e.g., to predict such risk. In some embodiments, the IgH CDR1, 2, 3 sequence data obtained using the provided assay methods and compositions can be used to infer phased IgH locus haplotypes (e.g., Kidd et al. (2012) J. Immunol. 188(3):1333-1340).
[0175] In some embodiments, the provided method includes preparing and forming a plurality of immune receptor-specific amplicons. In some embodiments, the method includes: hybridizing a plurality of V gene-specific primers and at least one C gene-specific primer to cDNA molecules; extending a first primer of the primer pair (e.g., a V gene-specific primer); denaturing the extended first primer from the cDNA molecule; hybridizing a second primer of the primer pair (e.g., a C gene-specific primer) to the extended first primer product and extending the second primer; digesting the target-specific primer pair to produce a plurality of target amplicons. In other embodiments, the method includes: hybridizing a plurality of V gene-specific primers and a plurality of J gene-specific primers to cDNA molecules; extending a first primer of the primer pair (e.g., a V gene-specific primer); denaturing the extended first primer from the cDNA molecule; hybridizing a second primer of the primer pair (e.g., a J gene-specific primer) to the extended first primer product and extending the second primer; digesting the target-specific primer pair to produce a plurality of target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons prior to performing a nick translation reaction to generate a plurality of target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters includes at least one barcode sequence. In some embodiments, each adapter ligated to the ends of the target amplicons includes a barcode sequence. In some embodiments, bridge amplification, emulsion PCR, or isothermal amplification can be used to amplify the one or more target amplicons to produce a plurality of clone templates suitable for nucleic acid sequencing.
[0176] In some embodiments, the provided method includes preparing and forming a plurality of immune receptor-specific amplicons. In some embodiments, the method includes hybridizing a plurality of V gene-specific primers and a plurality of J gene-specific primers to gDNA molecules, extending a first primer of the primer pair (e.g., a V gene-specific primer), denaturing the extended first primer from the gDNA molecule, hybridizing a second primer of the primer pair (e.g., a J gene-specific primer) to the extended first primer product, and extending the second primer, digesting the target-specific primer pair to generate a plurality of target amplicons. In some embodiments, adapters are ligated to the ends of the target amplicons prior to performing a nick translation reaction to generate a plurality of target amplicons suitable for nucleic acid sequencing. In some embodiments, at least one of the ligated adapters includes at least one barcode sequence. In some embodiments, each adapter ligated to the ends of the target amplicons includes a barcode sequence. In some embodiments, bridge amplification or emulsion PCR can be used to amplify the one or more target amplicons to produce a plurality of clone templates suitable for nucleic acid sequencing.
[0177] In some embodiments, the present disclosure provides methods for sequencing target amplicons and processing sequence data to identify productive immune receptor rearrangements expressed in a biological sample from which the cDNA is derived. In some embodiments, the present disclosure provides methods for sequencing target amplicons and processing sequence data to identify productive immune receptor gene rearrangements gDNA from a biological sample. In embodiments where primers targeting J genes are used to amplify expressed immune receptor sequences or rearranged immune receptor gDNA sequences, processing the sequence data includes the nucleotide sequence of the J gene primers used for amplification and the remainder of the targeted J gene, as described herein. In some embodiments, processing the sequence data includes performing the provided error identification and correction steps to generate rescued productive sequences. In some embodiments, using the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being at least 50% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, using the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, using the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being about 50 - 60%, about 60 - 70%, about 70 - 80%, about 80 - 90%, about 50 - 80%, or about 60 - 90% of the sequencing reads for an immune receptor cDNA or gDNA sample. In some embodiments, using the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% of the sequencing reads for an immune receptor cDNA or gDNA sample.
[0178] In the case of certain samples, the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being less than 50% of the sequencing reads for an immune receptor cDNA or gDNA sample when using the certain sample. Such samples include, for example, those with highly degraded RNA or gDNA (such as FFPE samples and cfDNA samples), and those with a very low number of target immune cells, such as samples with a very low B cell count or samples from subjects experiencing severe leukopenia. Thus, in some embodiments, using the provided error identification and correction workflow can result in the combination of productive reads and rescued productive reads being about 30 - 50%, about 40 - 50%, about 30 - 40%, about 40 - 60%, at least 30%, or at least 40% of the sequencing reads for an immune receptor cDNA or gDNA sample.
[0179] In certain embodiments, the methods of the invention include using a set of target immunoreceptor primers, where the primers target sequences of the same target immunoreceptor genes, such as BCR (immunoglobulin) genes and TCR genes. In some embodiments, the immunoreceptor is an antibody receptor selected from the group consisting of: heavy chain alpha, heavy chain delta, heavy chain epsilon, heavy chain gamma, heavy chain mu, light chain kappa, and light chain lambda. In some embodiments, the T cell receptor is a T cell receptor selected from the group consisting of: TCR alpha, TCR beta, TCR gamma, and TCR delta. In some embodiments, the methods of the invention include using a set of target immunoreceptor primers, where at least one set of primers in the set of primers targets the sequence of the BCR and another set of primers targets the sequence of the TCR, and both the BCR target nucleic acid and the TCR target nucleic acid from the sample are amplified in a single multiplex amplification reaction.
[0180] In certain embodiments, a method for amplifying an expressed nucleic acid sequence of a BCR repertoire in a sample is provided, the method comprising performing a multiplex amplification reaction using at least one of the following sets to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion: i) a plurality of V gene primers that target most different V genes of at least one BCR coding sequence that includes at least a portion of the framework region within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target constant gene of the BCR coding sequence, where each set of i) primers and ii) primers targeting the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, and where performing the amplification using each set produces amplicons representative of the entire repertoire of the corresponding immunoreceptor in the sample; thereby producing immunoreceptor amplicons that include the repertoire of the BCR. In a particular embodiment, the one or more plurality of V gene primers of i) target a sequence that is about 80 nucleotide portions above the framework region. In a more particular embodiment, the one or more plurality of V gene primers of i) target a sequence that is about 50 nucleotide portions above the framework region.
[0181] In certain embodiments, methods are provided for amplifying expressed nucleic acid sequences of an immune receptor repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of the following to amplify a BCR nucleic acid template molecule having a constant portion and a variable portion: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing the amplification using each set produces amplicons representative of the entire repertoire of the corresponding immune receptor in the sample; thereby generating immune receptor amplicons that include a repertoire of BCRs. In a particular embodiment, the one or more plurality of V gene primers of i) target a sequence of about 80 nucleotide portions above the framework region. In a more particular embodiment, the one or more plurality of V gene primers of i) target a sequence of about 50 nucleotide portions above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of framework region 1 of the template molecule. In certain embodiments, one or more C gene primers of ii) include at least two primers that anneal to at least a portion of the C gene portion of the template molecule. In some embodiments, one or more C gene primers of ii) include at least two primers, each of the at least two primers annealing to at least a portion of the C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, one or more C gene primers of ii) include at least one primer that targets a portion of the C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules, respectively. In a particular embodiment, the amplicons produced by at least one set contain complementarity determining regions CDR1, CDR2, and CDR3 of the BCR expression sequence. In some embodiments, the amplicon length is from about 300 to about 600 nucleotides or is at least about 350 to about 500 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA produced by reverse transcribing a nucleic acid molecule extracted from a biological sample.
[0182] In certain embodiments, methods are provided for providing sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence comprising at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers that target at least a portion of one or more corresponding target C genes of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting BCR amplicon molecules is then performed, and the sequences of the BCR amplicon molecules so determined provide the sequences of the BCR repertoire in the sample. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the BCR. In further embodiments, the method further comprises sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database, and the sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the average sequence read length is between 300 and 600 nucleotides, or between 350 and 550 nucleotides, or between 330 and 425 nucleotides, or between about 350 and about 425 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining regions CDR1, CDR2, and CDR3 of the BCR expression sequence.
[0183] In some embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the plurality of V gene primers are directed to a sequence that is about 70 nucleotides longer than the FR1 region. In other specific embodiments, one or more of the plurality of V gene primers are directed to a sequence that is about 50 nucleotides longer than the FR1 region. In certain embodiments, the target BCR primer set includes V gene primers that include from about 18 to about 45 different primers directed to FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 22 to about 35 different primers directed to FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 25 to about 35 different primers directed to FR1. In certain embodiments, the target BCR primer set includes V gene primers that include from about 40 to about 65 different primers directed to FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 48 to about 60 different primers directed to FR1. In some embodiments, the target BCR primer set includes one or more C gene primers. In specific embodiments, the target immune receptor primer set includes at least 5 to about 15 C gene primers, wherein each gene primer is directed to at least a portion of a 50 nucleotide identical region within each target C gene of the target C genes. In specific embodiments, the target immune receptor primer set includes at least 2 to about 8 C gene primers, wherein each gene primer is directed to at least a portion of a 50 nucleotide identical region within each target C gene of the target C genes. In some embodiments, the target BCR primer set includes two or more C gene primers that are directed to different Ig isotype molecules, such as IgA, IgD, IgG, IgM, and IgE. In some embodiments, the target BCR primer set includes at least five C gene primers, each primer being directed to the C gene of a different Ig isotype molecule.
[0184] In certain embodiments, the method of the present invention comprises using at least one set of primers, said at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Table 3 and Tables 6 - 10, respectively. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 15 to about 35 primers selected from Table 3 and from about 5 to about 20 primers selected from Tables 6 - 10, respectively. In some embodiments, the provided method comprises using at least one set of primers, said at least one set of primers comprising i) from about 22 to about 35 primers selected from Table 3, and ii) one or more primers selected from each of Tables 6 - 10. In certain embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 40 to about 65 primers selected from Table 3 and from about 5 to about 20 primers selected from Tables 6 - 10, respectively. In some embodiments, the provided method comprises using at least one set of primers, said at least one set of primers comprising i) from about 48 to about 60 primers selected from Table 3, and ii) one or more primers selected from each of Tables 6 - 10. In certain other embodiments, the method of the present invention comprises using at least one set of primers, said at least one set of primers comprising i) primers selected from SEQ ID NO: 137 - 283 and ii) primers selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551 and 564 - 582. In other embodiments, the provided method comprises using at least one set of primers, said at least one set of primers comprising i) primers selected from SEQ ID NO: 284 - 430 and ii) primers selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563 and 583 - 601. In some embodiments, the method of the present invention comprises using at least one set of primers, said at least one set of primers comprising i) primers selected from SEQ ID NO: 137 - 283 and ii) primers selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563 and 583 - 601 or comprising i) primers selected from SEQ ID NO: 284 - 430 and ii) primers selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551 and 564 - 582.
[0185] In some embodiments, the method of the present invention includes using at least one set of primers i) and ii), the at least one set of primers including at least 20 or at least 25 primers selected from SEQ ID NOs: 137 - 283 and at least one primer selected from SEQ ID NOs: 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In some embodiments, the provided method includes using at least one set of primers i) and ii), the at least one set of primers including about 15 to about 35 primers selected from SEQ ID NOs: 137 - 283 and about 5 to about 15 primers selected from SEQ ID NOs: 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In some embodiments, the provided method includes using at least one set of primers, the at least one set of primers including i) about 22 to about 35 primers selected from SEQ ID NOs: 137 - 283 and ii) at least one primer selected from SEQ ID NOs: 448 - 459, at least one primer selected from SEQ ID NOs: 472 - 479, at least one primer selected from SEQ ID NOs: 488 - 513, at least one primer selected from SEQ ID NOs: 540 - 551, and at least one primer selected from SEQ ID NOs: 564 - 582. In other embodiments, the method of the present invention includes using at least one set of primers i) and ii), the at least one set of primers including at least 20 or at least 25 primers selected from SEQ ID NOs: 284 - 430 and at least one primer selected from SEQ ID NOs: 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the provided method includes using at least one set of primers i) and ii), the at least one set of primers including about 15 to about 35 primers selected from SEQ ID NOs: 284 - 430 and about 5 to about 15 primers selected from SEQ ID NOs: 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the provided method includes using at least one set of primers, the at least one set of primers including i) about 22 to about 35 primers selected from SEQ ID NOs: 284 - 430 and ii) at least one primer selected from SEQ ID NOs: 460 - 471, at least one primer selected from 480 - 487, at least one primer selected from 514 - 539, at least one primer selected from 552 - 563, and at least one primer selected from 583 - 601.In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 20 or at least 25 primers selected from SEQ ID NO: 284 - 430 and at least one primer selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551 and 564 - 582. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 20 or at least 25 primers selected from SEQ ID NO: 137 - 283 and at least one primer selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563 and 583 - 601.
[0186] In some embodiments, the method of the present invention includes using at least one set of primers i) and ii), the at least one set of primers including at least 40 or at least 50 primers selected from SEQ ID NO: 137 - 283 and at least one primer selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In some embodiments, the provided method includes using at least one set of primers i) and ii), the at least one set of primers including about 40 to about 65 primers selected from SEQ ID NO: 137 - 283 and about 5 to about 15 primers selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In some embodiments, the provided method includes using at least one set of primers, the at least one set of primers including i) about 48 to about 60 primers selected from SEQ ID NO: 137 - 283 and ii) at least one primer selected from SEQ ID NO: 448 - 459, at least one primer selected from SEQ ID NO: 472 - 479, at least one primer selected from SEQ ID NO: 488 - 513, at least one primer selected from SEQ ID NO: 540 - 551, and at least one primer selected from SEQ ID NO: 564 - 582. In other embodiments, the method of the present invention includes using at least one set of primers i) and ii), the at least one set of primers including at least 40 or at least 50 primers selected from SEQ ID NO: 284 - 430 and at least one primer selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the provided method includes using at least one set of primers i) and ii), the at least one set of primers including about 40 to about 65 primers selected from SEQ ID NO: 284 - 430 and about 5 to about 15 primers selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the provided method includes using at least one set of primers, the at least one set of primers including i) about 48 to about 60 primers selected from SEQ ID NO: 284 - 430 and ii) at least one primer selected from SEQ ID NO: 460 - 471, at least one primer selected from 480 - 487, at least one primer selected from 514 - 539, at least one primer selected from 552 - 563, and at least one primer selected from 583 - 601.In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 284-430 and at least one primer selected from SEQ ID NO: 448-459, 472-479, 488-513, 540-551 and 564-582. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 137-283 and at least one primer selected from SEQ ID NO: 460-471, 480-487, 514-539, 552-563 and 583-601.
[0187] In certain embodiments, provided are methods for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of the following to amplify BCR nucleic acid template molecules having a constant portion and a variable portion: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing amplification using each set produces amplicons representative of the entire repertoire of the corresponding immune receptor in the sample; thereby generating immune receptor amplicons comprising the BCR repertoire. In certain embodiments, the one or more plurality of V gene primers of i) target a sequence that is about 80 nucleotides above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence that is about 50 nucleotides above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence that is about 40 to about 60 nucleotides above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 3 region of the template molecule. In certain embodiments, the one or more C gene primers of ii) include at least two primers that anneal to at least a portion of the C gene of the BCR template molecule. In some embodiments, the one or more C gene primers of ii) include at least two primers, each of the at least two primers annealing to at least a portion of the C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers of ii) include at least one primer that targets a portion of the C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules, respectively. In certain embodiments, the amplicons produced by at least one set contain the complementarity-determining region CDR3 of the BCR expression sequence. In some embodiments, the length of the amplicons is about 80 to about 200 nucleotides, about 80 to about 140 nucleotides, about 90 to about 130 nucleotides, or at least about 100 to about 120 nucleotides. In some embodiments, the nucleic acid template used in the method is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0188] In certain embodiments, methods are provided for providing sequences of BCR repertoires in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers that target at least a portion of one or more corresponding target C genes of the BCR coding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting BCR amplicon molecules is then performed, and the sequences of the BCR amplicon molecules so determined provide the sequences of the BCRs in the sample. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the BCR. In further embodiments, the methods further comprise sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified BCR repertoires are compared to a contemporaneous or current version of the IMGT database, and sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the average sequence read length is between 80 and 185 nucleotides, between 115 and 200 nucleotides, between 90 and 130 nucleotides, or between about 100 and about 120 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining region CDR3 of the BCR expressed sequence.
[0189] In certain embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the plurality of V gene primers target a sequence that is approximately 70 nucleotides longer than the FR3 region. In particular embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the plurality of V gene primers target a sequence that is approximately 50 nucleotides longer than the FR3 region. In other particular embodiments, one or more of the plurality of V gene primers target a sequence that is approximately 40 to approximately 60 nucleotides longer than the FR3 region. In certain embodiments, the target BCR primer set includes V gene primers that include from about 50 to about 85 different primers that target FR3. In certain embodiments, the target BCR primer set includes V gene primers that include from about 55 to about 80 different primers that target FR3. In some embodiments, the target immune receptor primer set includes V gene primers that include from about 62 to about 75 different primers that target FR3. In some embodiments, the target BCR primer set includes V gene primers that include about 65, 66, 67, 68, 69, or 70 different primers that target FR3. In some embodiments, the target BCR primer set includes one or more C gene primers. In particular embodiments, the target immune receptor primer set includes at least 5 to about 15 C gene primers, wherein each gene primer targets at least a portion of a 50 nucleotide identical region within each of the target C genes in the target C gene. In particular embodiments, the target BCR primer set includes at least 2 to about 8 C gene primers, wherein each gene primer targets at least a portion of a 50 nucleotide identical region within each of the target C genes in the target C gene. In some embodiments, one or more of the C gene primers of ii) include at least two primers, each of the at least two primers annealing to at least a portion of the C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, one or more of the C gene primers of ii) include at least one primer that targets a portion of the C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules, respectively.
[0190] In certain embodiments, the method of the invention comprises using at least one set of primers, said at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Table 2 and Tables 6-10, respectively. In some embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 55 to about 80 primers selected from Table 2 and from about 5 to about 20 primers selected from Tables 6-10. In some embodiments, the provided method comprises using at least one set of primers, said at least one set of primers comprising i) from about 62 to about 75 primers selected from Table 2, and ii) one or more primers selected from each of Tables 6-10. In certain other embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NO: 1-68 and 448-459, 472-479, 488-513, 540-551 and 564-582 or selected from SEQ ID NO: 69-136 and 460-471, 480-487, 514-539, 552-563 and 583-601. In some embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NO: 1-68 and 460-471, 480-487, 514-539, 552-563 and 583-601 or selected from SEQ ID NO: 69-136 and 448-459, 472-479, 488-513, 540-551 and 564-582.
[0191] In some embodiments, the method of the present invention includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 1-68 and at least one primer selected from SEQ ID NO: 448-459, 472-479, 488-513, 540-551 and 564-582. In some embodiments, the provided method includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 1-68 and about 5 to about 15 primers selected from SEQ ID NO: 448-459, 472-479, 488-513, 540-551 and 564-582. In some embodiments, the provided method includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 1-68 and at least one primer selected from SEQ ID NO: 448-459, at least one primer selected from SEQ ID NO: 472-479, at least one primer selected from SEQ ID NO: 488-513, at least one primer selected from SEQ ID NO: 540-551 and at least one primer selected from SEQ ID NO: 564-582. In other embodiments, the method of the present invention includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 69-136 and at least one primer selected from SEQ ID NO: 460-471, 480-487, 514-539, 552-563 and 583-601. In some embodiments, the provided method includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 69-136 and about 5 to about 15 primers selected from SEQ ID NO: 460-471, 480-487, 514-539, 552-563 and 583-601. In some embodiments, the provided method includes using at least one set of primers i) and ii), said at least one set of primers including at least 60 primers selected from SEQ ID NO: 69-136 and at least one primer selected from SEQ ID NO: 460-471, at least one primer selected from SEQ ID NO: 480-487, at least one primer selected from SEQ ID NO: 514-539, at least one primer selected from SEQ ID NO: 552-563 and at least one primer selected from SEQ ID NO: 583-601.In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 1-68 and at least one primer selected from SEQ ID NO: 460-471, 480-487, 514-539, 552-563 and 583-601. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 69-136 and at least one primer selected from SEQ ID NO: 448-459, 472-479, 488-513, 540-551 and 564-582.
[0192] In certain embodiments, provided are methods for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of the following to amplify BCR nucleic acid template molecules having a constant portion and a V gene portion: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 2 (FR2) within the V gene, and ii) one or more C gene primers that target at least a portion of the C gene of the corresponding BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing the amplification using each set produces amplicons representative of the entire repertoire of the corresponding immune receptor in the sample; thereby producing amplicons comprising the BCR repertoire. In certain embodiments, the one or more plurality of V gene primers of i) target a sequence that is about 80 nucleotide portions above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence that is about 50 nucleotide portions above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of the FR2 region of the BCR template molecule. In certain embodiments, the one or more C gene primers of ii) comprise at least two primers that anneal to at least a portion of the constant portion C gene of the BCR template molecule. In some embodiments, the one or more C gene primers of ii) comprise at least two primers, each of the at least two primers annealing to at least a portion of the C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, the one or more C gene primers of ii) comprise at least one primer that targets a portion of the C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules, respectively. In certain embodiments, the amplicons produced by at least one set comprise complementarity determining regions CDR2 and CDR3 of the BCR expression sequence. In some embodiments, the length of the amplicons is from about 180 to about 375 nucleotides, from about 200 to about 350 nucleotides, from about 225 to about 325 nucleotides, or from about 250 to about 300 nucleotides in length. In some embodiments, the nucleic acid template used in the method is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0193] In certain embodiments, methods are provided for providing sequences of BCR repertoires in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of FR2 within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting BCR amplicon molecules is then performed, and the sequences of the BCR amplicon molecules thus determined provide the sequences of the BCR repertoire in the sample. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 40%, at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the BCR. In further embodiments, the methods further comprise sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified immune repertoires are compared to a contemporaneous or current version of the IMGT database, and sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the average sequence read length is between about 200 and about 375 nucleotides, between about 250 and about 350 nucleotides, or between about 275 and about 350 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining regions CDR2 and CDR3 of the BCR expressed sequence.
[0194] In certain embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the plurality of V gene primers target a sequence that is approximately 70 nucleotides longer than the FR2 region. In other specific embodiments, one or more of the plurality of V gene primers target a sequence that is approximately 50 nucleotides longer than the FR2 region. In some embodiments, the target BCR primer set includes V gene primers that include from about 4 to about 20 different primers targeting FR2. In some embodiments, the target BCR primer set includes V gene primers that include from about 5 to about 15 different primers targeting FR2. In some embodiments, the target BCR primer set includes V gene primers that include about 5, 6, 7, 8, 9, 10, 11, or 12 different primers targeting FR2. In some embodiments, the target BCR primer set includes one or more C gene primers. In certain embodiments, the target immune receptor primer set includes at least 5 to about 15 C gene primers, wherein each gene primer targets at least a portion of a 50 nucleotide identical region within each target C gene of the target C genes. In certain embodiments, the target BCR primer set includes at least 2 to about 8 C gene primers, wherein each gene primer targets at least a portion of a 50 nucleotide identical region within each target C gene of the target C genes. In some embodiments, one or more of the C gene primers of ii) include at least two primers, each of the at least two primers annealing to at least a portion of the C gene of an IgA, IgD, IgG, IgM, or IgE template molecule. In some embodiments, one or more of the C gene primers of ii) include at least one primer that targets a portion of the C gene of each of the IgA, IgD, IgG, IgM, and IgE template molecules, respectively.
[0195] In certain embodiments, the method of the invention comprises using at least one set of primers, said at least one set of primers comprising V gene primers i) and C gene primers ii) selected from Tables 4 and 6 - 10, respectively. In certain other embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NO: 431 - 437 and 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In other embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NO: 431 - 437 and 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 1 primer selected from SEQ ID NO: 448 - 459, 472 - 479, 488 - 513, 540 - 551, and 564 - 582. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 1 primer selected from SEQ ID NO: 448 - 459, at least 1 primer selected from SEQ ID NO: 472 - 479, at least 1 primer selected from SEQ ID NO: 488 - 513, at least 1 primer selected from SEQ ID NO: 540 - 551, and at least 1 primer selected from SEQ ID NO: 564 - 582. In other embodiments, the method of the invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 1 primer selected from SEQ ID NO: 460 - 471, 480 - 487, 514 - 539, 552 - 563, and 583 - 601. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 1 primer selected from SEQ ID NO: 460 - 471, at least 1 primer selected from SEQ ID NO: 480 - 487, at least 1 primer selected from SEQ ID NO: 514 - 539, at least 1 primer selected from SEQ ID NO: 552 - 563, and at least 1 primer selected from SEQ ID NO: 583 - 601.
[0196] In certain embodiments, methods are provided for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one of the following sets to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment: i) a plurality of V gene primers that target a majority of the different V genes of the BCR coding sequences that include at least a portion of the framework region within the V gene, and ii) a plurality of J gene primers that target a majority of the different J genes of the respective target immunoreceptor coding sequences, wherein each set of primers of i) and ii) that targets the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing the amplification using each set produces amplicons that represent the entire repertoire of the respective immunoreceptors in the sample; thereby producing amplicons that include a repertoire of BCRs. In certain embodiments, the one or more of the plurality of V gene primers of i) target a sequence that is about 80 nucleotides above the framework region. In more certain embodiments, the one or more of the plurality of V gene primers of i) target a sequence that is about 50 nucleotides above the framework region. In certain embodiments, the one or more of the plurality of J gene primers of ii) target a sequence that is about 50 nucleotides above the J gene. In more certain embodiments, the one or more of the plurality of J gene primers of ii) target a sequence that is about 30 nucleotides above the J gene. In certain embodiments, the one or more of the plurality of J gene primers of ii) target a sequence that is entirely within the J gene.
[0197] In certain embodiments, provided are methods for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of the following to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers that target a majority of different J genes of the corresponding target BCR-encoding sequence, wherein each set of primers i) and ii) targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing amplification using each set produces amplicons representative of the entire repertoire of the corresponding immune receptor in the sample; thereby producing BCR amplicons comprising a repertoire of BCRs. In certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 80 nucleotide portions above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 50 nucleotide portions above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 40 to about 60 nucleotide portions above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 3 region of the template molecule. In certain embodiments, the plurality of J gene primers of ii) comprise at least two primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise at least 2 to about 8 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 4 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 3 to about 6 primers that anneal to at least a portion of the J gene segment of the template molecule. In certain embodiments, the amplicons produced by at least one set contain the complementarity determining region CDR3 of the BCR expression sequence. In some embodiments, the length of the amplicon is about 60 to about 160 nucleotides, about 70 to about 100 nucleotides, about 100 to about 120 nucleotides, at least about 70 to about 90 nucleotides, about 80 to about 90 nucleotides, or about 80 nucleotides. In some embodiments, the nucleic acid template used in the method is cDNA produced by reverse transcription from nucleic acid molecules extracted from a biological sample.
[0198] In certain embodiments, methods are provided for providing sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers that target a majority of the different J genes of the respective target immunoreceptor-encoding sequences, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting BCR amplicon molecules is then performed, and the sequences of the immunoreceptor amplicon molecules thus determined provide the sequences of the BCR repertoire in the sample. In some embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence; identifying productive reads; correcting one or more indel errors to produce rescued productive sequence reads; and determining the sequences of the resulting immunoreceptor molecules. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; adding an inferred J gene sequence to the sequence reads to produce extended sequence reads; aligning the extended sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to produce rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive sequence reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the BCR. In further embodiments, the method further comprises sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified BCR repertoire are compared to a contemporaneous or current version of the IMGT database, and sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the sequence read lengths are from about 60 to about 185 nucleotides, depending in part on including any barcode sequences in the read lengths. In some embodiments, the average sequence read length is between 90 and 120 nucleotides, between 70 and 90 nucleotides, or between about 75 and about 85 nucleotides or about 80 nucleotides. In certain embodiments, at least one set of sequenced amplicons includes complementarity determining region CDR3 of the BCR expressed sequence.
[0199] In certain embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the V gene primers among the plurality of V gene primers target a sequence that is approximately 50 nucleotides longer than the FR3 region. In other embodiments, one or more of the V gene primers among the plurality of V gene primers target a sequence that is approximately 70 nucleotides longer than the FR3 region. In other specific embodiments, one or more of the V gene primers among the plurality of V gene primers target a sequence that is approximately 40 to approximately 60 nucleotides longer than the FR3 region. In certain embodiments, the target BCR primer set includes V gene primers that include from about 50 to about 85 different primers targeting FR3. In certain embodiments, the target BCR primer set includes V gene primers that include from about 55 to about 80 different primers targeting FR3. In some embodiments, the target immune receptor primer set includes V gene primers that include from about 62 to about 75 different primers targeting FR3. In some embodiments, the target BCR primer set includes V gene primers that include about 65, 66, 67, 68, 69, or 70 different primers targeting FR3. In some embodiments, the target BCR primer set includes a plurality of J gene primers. In certain embodiments, the target BCR primer set includes at least two J gene primers, wherein each gene primer targets at least a portion of the J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from 2 to about 8 J gene primers, wherein each gene primer targets at least a portion of the J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from about 3 to about 6 different J gene primers, wherein each gene primer targets at least a portion of the J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes about 2, 3, 4, 5, 6, 7, or 8 different J gene primers. In some embodiments, the target immune receptor primer set includes about 4 J gene primers, wherein each gene primer targets at least a portion of the J gene within the target polynucleotide.
[0200] In certain embodiments, the method of the invention includes using at least one set of primers, the at least one set of primers including V gene primers (i) and J gene primers (ii) selected from Tables 2 and 5, respectively. In certain other embodiments, the method of the invention includes using at least one set of primers (i) and (ii), the at least one set of primers including primers selected from SEQ ID NO: 1-68 and 438-442 or selected from SEQ ID NO: 69-136 and 443-447. In certain other embodiments, the method of the invention includes using at least one set of primers (i) and (ii), the at least one set of primers including primers selected from SEQ ID NO: 1-68 and 443-447 or selected from SEQ ID NO: 69-136 and 438-442.
[0201] In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 1-68 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 438-442. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 69-136 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 443-447. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 69-136 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 438-442. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 60 primers selected from SEQ ID NO: 1-68 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 443-447.
[0202] In certain embodiments, provided are methods for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of the following to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers that target a majority of different J genes of the corresponding target immunoreceptor-encoding sequences, wherein each set of primers i) and ii) for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing amplification using each set produces amplicons representative of the entire repertoire of the corresponding immunoreceptor in the sample; thereby producing BCR amplicons that comprise a repertoire of BCRs. In certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 80 nucleotide portions above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 50 nucleotide portions above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of the framework 1 region of the template molecule. In certain embodiments, the plurality of J gene primers of ii) comprise at least two primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise at least 2 to about 8 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 4 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 3 to about 6 primers that anneal to at least a portion of the J gene segment of the template molecule. In certain embodiments, the amplicons produced by at least one set comprise complementarity-determining regions CDR1, CDR2, and CDR3 of the BCR expression sequence. In some embodiments, the length of the amplicon is about 220 to about 350 nucleotides, about 225 to about 300 nucleotides, about 250 to about 325 nucleotides, about 250 to about 275 nucleotides, or a length of about 270 to about 300 nucleotides. In some embodiments, the nucleic acid template used in the method is cDNA produced by reverse transcribing nucleic acid molecules extracted from a biological sample.
[0203] In certain embodiments, methods are provided for providing sequences of BCR repertoires in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having J gene segments and V gene segments, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers that target a majority of different J genes of a corresponding target immune receptor-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting immune receptor amplicon molecules is then performed, and the sequences of the BCR amplicon molecules thus determined provide the sequences of the BCR repertoire in the sample. In some embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence; identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; adding an inferred J gene sequence to the sequence reads to generate extended sequence reads; aligning the extended sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive sequence reads is at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the immune receptor. In additional embodiments, the method further comprises sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified immune repertoires are compared to a contemporaneous or current version of the IMGT database, and the sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the average sequence read length is between 200 and 350 nucleotides, between 225 and 325 nucleotides, between 250 and 300 nucleotides, between 270 and 300 nucleotides, or between 295 and 325 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining regions CDR1, CDR2, and CDR3 of the BCR expressed sequence.
[0204] In certain embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the plurality of V gene primers target a sequence that is about 70 nucleotides longer than the FR1 region. In certain other embodiments, one or more of the plurality of V gene primers target a sequence that is about 80 nucleotides longer than the FR1 region. In certain other embodiments, one or more of the plurality of V gene primers target a sequence that is about 50 nucleotides longer than the FR1 region. In certain embodiments, the target BCR primer set includes V gene primers that include from about 18 to about 45 different primers targeting FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 22 to about 35 different primers targeting FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 25 to about 35 different primers targeting FR1. In certain embodiments, the target BCR primer set includes V gene primers that include from about 40 to about 65 different primers targeting FR1. In some embodiments, the target BCR primer set includes V gene primers that include from about 48 to about 60 different primers targeting FR1. In some embodiments, the target BCR primer set includes a plurality of J gene primers. In certain embodiments, the target BCR primer set includes at least two J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from 2 to about 8 J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from about 3 to about 6 different J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes about 2, 3, 4, 5, 6, 7, or 8 different J gene primers. In some embodiments, the target immune receptor primer set includes about 4 J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide.
[0205] In certain embodiments, the method of the present invention comprises using at least one set of primers, said at least one set of primers comprising V gene primers i) and J gene primers ii) selected from Tables 3 and 5, respectively. In certain other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NOs: 137 - 283 and 438 - 442 or selected from SEQ ID NOs: 284 - 430 and 443 - 447. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising primers selected from SEQ ID NOs: 137 - 283 and 443 - 447 or selected from SEQ ID NOs: 284 - 430 and 438 - 442. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 438 - 442. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 15 to about 35 primers selected from SEQ ID NOs: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 438 - 442. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 22 to about 35 primers selected from SEQ ID NOs: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 438 - 442. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 20 or at least 25 primers selected from SEQ ID NOs: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 443 - 447. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 15 to about 35 primers selected from SEQ ID NOs: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 443 - 447. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 22 to about 35 primers selected from SEQ ID NOs: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NOs: 443 - 447.In some embodiments, the method of the present invention includes using at least one set of primers i) and ii), said at least one set of primers including at least 20 or at least 25 primers selected from SEQ ID NO: 137 - 283 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 443 - 447. In other embodiments, the method of the present invention includes using at least one set of primers i) and ii), said at least one set of primers including at least 20 or at least 25 primers selected from SEQ ID NO: 284 - 430 and at least 2, at least 3 or at least 4 primers selected from SEQ ID NO: 438 - 442.
[0206] In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 438 - 442. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 40 to about 65 primers selected from SEQ ID NO: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 438 - 442. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 48 to about 60 primers selected from SEQ ID NO: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 438 - 442. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 443 - 447. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 40 to about 65 primers selected from SEQ ID NO: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 443 - 447. In some embodiments, the provided method comprises using at least one set of primers i) and ii), said at least one set of primers comprising from about 48 to about 60 primers selected from SEQ ID NO: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 443 - 447. In some embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 137 - 283 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 443 - 447. In other embodiments, the method of the present invention comprises using at least one set of primers i) and ii), said at least one set of primers comprising at least 40 or at least 50 primers selected from SEQ ID NO: 284 - 430 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 438 - 442.
[0207] In certain embodiments, provided are methods for amplifying expressed nucleic acid sequences of a BCR repertoire in a sample, the methods comprising performing a multiplex amplification reaction using at least one of the following sets to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment: i) a plurality of V gene primers that target most different V genes of a BCR-encoding sequence that includes at least a portion of framework region 2 (FR2) within the V gene, and ii) a plurality of J gene primers that target most different J genes of a corresponding target immune receptor-encoding sequence, wherein each set of primers i) and ii) targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, and wherein performing amplification using each set produces amplicons representative of the entire repertoire of the corresponding immune receptor in the sample; thereby generating immune receptor amplicons comprising a repertoire of BCRs. In certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 80 nucleotide portions above the framework region. In more certain embodiments, the one or more plurality of V gene primers of i) target a sequence of about 50 nucleotide portions above the framework region. In some embodiments, the one or more plurality of V gene primers of i) anneal to at least a portion of the FR2 region of the template molecule. In certain embodiments, the plurality of J gene primers of ii) comprise at least ten primers that anneal to at least a portion of the J gene of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 14 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise at least two primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise at least 2 to about 8 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 4 primers that anneal to at least a portion of the J gene segment of the template molecule. In some embodiments, the plurality of J gene primers of ii) comprise about 3 to about 6 primers that anneal to at least a portion of the J gene segment of the template molecule. In certain embodiments, the amplicons produced by at least one set contain complementarity-determining regions CDR2 and CDR3 of the BCR gene sequence. In some embodiments, the length of the amplicons is about 160 to about 270 nucleotides, about 180 to about 250 nucleotides, or a length of about 195 to about 225 nucleotides. In some embodiments, the nucleic acid template used in the method is cDNA generated by reverse transcription from nucleic acid molecules extracted from a biological sample.
[0208] In certain embodiments, methods are provided for providing sequences of BCR repertoires in a sample, the methods comprising performing a multiplex amplification reaction using at least one set of primers to amplify BCR nucleic acid template molecules having a J gene segment and a V gene segment, the at least one set of primers comprising: i) a plurality of V gene primers that target most different V genes of at least one BCR-encoding sequence that includes at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers that target most different J genes of corresponding target immunoreceptor-encoding sequences, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. Sequencing of the resulting immunoreceptor amplicon molecules is then performed, and the sequences of the BCR amplicon molecules thus determined provide the sequences of the BCR repertoire in the sample. In some embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence; identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immunoreceptor molecules. In certain embodiments, determining the sequences of the BCR amplicon molecules comprises: obtaining initial sequence reads; adding an inferred J gene sequence to the sequence reads to generate extended sequence reads; aligning the extended sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In certain embodiments, the combination of productive reads and rescued productive reads is at least 40%, at least 50%, at least 60%, at least 70%, or at least 75% of the sequencing reads of the BCR. In further embodiments, the method further comprises sequence read clustering and BCR clonotype reporting. In some embodiments, the sequences of the identified immunological repertoires are compared to a contemporaneous or current version of the IMGT database, and sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the average sequence read length is between 160 and 300 nucleotides, between 180 and 280 nucleotides, between 200 and 260 nucleotides, or between 225 and 270 nucleotides, depending in part on the inclusion of any barcode sequences in the read length. In certain embodiments, at least one set of sequenced amplicons comprises complementarity determining regions CDR2 and CDR3 of the BCR expression sequence.
[0209] In certain embodiments, the provided method utilizes a target BCR primer set that includes V gene primers, wherein one or more of the V gene primers among the plurality of V gene primers target a sequence that is approximately 70 nucleotides longer than the FR2 region. In other certain embodiments, one or more of the V gene primers among the plurality of V gene primers target a sequence that is approximately 50 nucleotides longer than the FR2 region. In some embodiments, the target BCR primer set includes V gene primers that include from about 4 to about 20 different primers targeting FR2. In some embodiments, the target BCR primer set includes V gene primers that include from about 5 to about 15 different primers targeting FR2. In some embodiments, the target BCR primer set includes V gene primers that include about 5, 6, 7, 8, 9, 10, 11, or 12 different primers targeting FR2. In some embodiments, the target BCR primer set includes a plurality of J gene primers. In certain embodiments, the target BCR primer set includes at least two J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from 2 to about 8 J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes from about 3 to about 6 different J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide. In some embodiments, the target BCR primer set includes about 2, 3, 4, 5, 6, 7, or 8 different J gene primers. In some embodiments, the target immune receptor primer set includes about 4 J gene primers, wherein each gene primer targets at least a portion of a J gene within the target polynucleotide.
[0210] In certain embodiments, the method of the invention includes using at least one set of primers, the at least one set of primers including V gene primers (i) and J gene primers (ii) selected from Tables 4 and 5, respectively. In certain other embodiments, the method of the invention includes using at least one set of primers (i) and (ii), the at least one set of primers including primers selected from SEQ ID NO: 431 - 437 and 438 - 442 or selected from SEQ ID NO: 431 - 437 and 443 - 447. In some embodiments, the method of the invention includes using at least one set of primers (i) and (ii), the at least one set of primers including at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 438 - 442. In other embodiments, the method of the invention includes using at least one set of primers (i) and (ii), the at least one set of primers including at least 5 primers selected from SEQ ID NO: 431 - 437 and at least 2, at least 3, or at least 4 primers selected from SEQ ID NO: 443 - 447.
[0211] In certain embodiments, the methods of the present invention include using a biological sample selected from the group consisting of hematopoietic cells, lymphocytes, and tumor cells. In some embodiments, the biological sample is selected from the group consisting of: peripheral blood mononuclear cells (PBMCs), T cells, B cells, circulating tumor cells, and tumor infiltrating lymphocytes (referred to herein as "TILs" or "TIL"). In some embodiments, the biological sample includes B cells that are activated and / or expanded ex vivo. In some embodiments, the biological sample includes cfDNA as found, for example, in blood or plasma. In some embodiments, the biological sample is selected from the group consisting of: tissue (e.g., lymph node, organ tissue, bone marrow), whole blood, synovial fluid, cerebrospinal fluid, tumor biopsy, and other cellular-containing clinical specimens.
[0212] In some embodiments, methods, compositions, and systems are provided for determining the immune repertoire of a biological sample by evaluating both expressed immune receptor RNA and rearranged immune receptor genomic DNA (gDNA) from the biological sample. In some embodiments, sample RNA and gDNA can be evaluated simultaneously, and after reverse transcription of the RNA to form cDNA, the cDNA and gDNA can be amplified in the same multiplex amplification reaction. In some embodiments, cDNA from sample RNA and sample gDNA can be multiplex amplified in separate reactions. In some embodiments, cDNA from sample RNA and sample gDNA can be multiplex amplified with a parallel primer pool. In some embodiments, the BCR repertoire of gDNA and RNA from a sample is evaluated using a primer pool directed to the same BCR. In some embodiments, the immune repertoire of gDNA and RNA from a sample is evaluated using primer pools directed to different immune receptors. In some embodiments, multiplex amplification reactions are performed separately with cDNA from sample RNA and sample gDNA to amplify the same or different target immune receptor molecules from the sample, and the resulting immune receptor amplicons are sequenced, thereby providing the sequences of the expressed immune receptor RNA and rearranged immune receptor gDNA of the biological sample.
[0213] In some embodiments, the immune repertoire of gDNA and / or RNA from a sample is evaluated using primer pools directed to different immune receptors. In some embodiments, multiplex amplification reactions are performed with a set of IgH primers provided herein and, for example, a set of primers directed to TCRβ as described in, for example, PCT Application No. PCT / US2018 / 014111, filed January 17, 2018, and PCT Application No. PCT / US2018 / 049259, filed August 31, 2018, the entire contents of each of which are incorporated herein by reference or as Oncomine TM TCRβ-SR Assay DNA, OncomineTM TCRβ-SR Assay for RNA and Oncomine TM The TCRβ-LR Assay (Thermo Fisher Scientific) is commercially available. The ability to assess both the BCR (e.g., IgH) repertoire and the TCR (e.g., TCRβ) repertoire from a sample using a single multiplex amplification reaction can be used to save time and limited biological samples and is applicable to many of the methods described herein, including methods related to allergy and autoimmunity, vaccine development and use, and immuno-oncology. For example, combining B cell repertoire analysis with T cell repertoire analysis can be used to improve the detection of changes in the immune repertoire after administration of immunotherapy (such as checkpoint blockade or checkpoint inhibitor immunotherapy), which potentially indicates a response to immunotherapy. Moreover, combining B cell repertoire analysis with T cell repertoire analysis can be used to improve the assessment of vaccine efficacy. Exemplary immune repertoire changes in response to immunotherapy or in response to vaccine administration include, but are not limited to, a decrease in T cell and B cell evenness after treatment (e.g., but not limited to at days 7-14 after treatment) compared to pre-treatment evenness values and an increase in the representation of IgG1-expressing B cells after one or more treatments compared to pre-treatment values.
[0214] In some embodiments, the provided methods and compositions are used to identify and / or characterize the immune repertoire of a subject. In some embodiments, the provided methods and compositions are used to identify and characterize novel or atypical BCR alleles of the immune repertoire of a subject. In some embodiments, the sequences of the identified immune repertoire are compared to a contemporaneous or current version of the IMGT database, and the sequences of at least one allelic variant not present in the IMGT database are identified. In some embodiments, the identified allelic variants not present in the IMGT database are subjected to evidence-based filtering using criteria such as clone number support, sequence read support, and / or number of individuals with allelic variants. The identified and reported allelic variants not present in the IMGT can be compared to other databases containing immune repertoire sequence information such as the NCBI NR database and the Lym1K database to cross-validate the reported novel or non-classical BCR alleles. For example, characterizing the presence of unrecorded or non-classical IgH polymorphisms may help to understand factors affecting autoimmune diseases, infectious diseases, and responses to immunotherapy. In some embodiments, the sequences of the identified novel or non-classical BCR alleles as described herein can be used to generate recombinant BCR nucleic acids or molecules. Accordingly, in other embodiments, methods for preparing recombinant nucleic acids encoding the identified novel IgH allelic variants are provided. In some embodiments, methods for preparing recombinant IgH allelic variant molecules and for preparing recombinant cells expressing the molecules are provided.
[0215] In some embodiments, the provided methods and compositions are used to identify and characterize novel or atypical BCR alleles of a subject's immune repertoire. In some embodiments, a patient's immune repertoire can be identified or characterized before and / or after a therapeutic treatment, such as a treatment for cancer or an immune disorder. In some embodiments, identifying or characterizing the immune repertoire can be used to evaluate the effectiveness or efficacy of a treatment to modify the treatment regimen and / or optimize the selection of therapeutic agents. In some embodiments, identifying or characterizing the immune repertoire can be used to evaluate a patient's response to immunotherapy, cancer vaccines, and / or other immune-based treatments or a combination of one or more thereof. In some embodiments, identifying or characterizing the immune repertoire can indicate the likelihood that a patient will respond to a therapeutic agent or can indicate the likelihood that a patient will not respond to a therapeutic agent.
[0216] In some embodiments, a patient's BCR repertoire can be identified or characterized to monitor the progression and / or treatment of a hyperproliferative disease (including detecting residual disease after treatment of the patient), monitor the progression and / or treatment of an autoimmune disease, transplant monitoring, and monitor antigen-stimulated conditions, including after vaccination, exposure to bacterial, fungal, parasitic, or viral antigens, or infection with bacteria, fungi, parasites, or viruses. In some embodiments, identifying or characterizing the BCR repertoire can be used to evaluate a patient's response to anti-infection or anti-inflammatory therapies.
[0217] In some embodiments, methods and compositions are provided for identifying and / or characterizing immunoglobulin repertoire clonal populations in a sample from a subject, the methods and compositions comprising performing one or more multiplex amplification reactions with the sample or cDNA prepared from the sample to amplify immunoglobulin repertoire nucleic acid template molecules having a constant portion and a variable portion using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the immunoglobulin receptor-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immunoglobulin receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further comprises: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying one or more immunoglobulin repertoire clonal populations of the target BCR from the sample. In certain embodiments, determining the sequence of the immunoglobulin receptor amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequence of the resulting immunoglobulin receptor molecules. In other embodiments of such methods and compositions, the one or more multiplex amplification reactions are performed using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immunoglobulin receptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, the one or more multiplex amplification reactions are performed using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 2 (FR2) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immunoglobulin receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0218] In some embodiments, methods and compositions are provided for identifying and / or characterizing immunoglobulin repertoire clonal populations in a sample from a subject, the methods and compositions including performing one or more multiplex amplification reactions on the sample or cDNA prepared from the sample using at least one set of primers to amplify immunoglobulin repertoire nucleic acid template molecules having a J gene segment and a V gene segment, the at least one set of primers including: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers that target a majority of different J genes of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying one or more immunoglobulin repertoire clonal populations of the target BCR from the sample. In certain embodiments, determining the sequence of the immunoreceptor amplicon molecules comprises: obtaining initial sequence reads; adding an inferred J gene sequence to the sequence reads to generate extended sequence reads; aligning the extended sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequence of the resulting immunoreceptor molecules. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers that includes: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) a plurality of J gene primers that target a majority of different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers that includes: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 2 (FR2) within the V gene, and ii) a plurality of J gene primers that target a majority of different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0219] Thus, in some embodiments, the provided methods, compositions, and workflows are used for, but not limited to, assessing the clonality, diversity, and abundance of B cell populations. For example, clonal expansion can identify B cells responding to antigen challenge, and longitudinal analysis can be used to evaluate the efficacy of vaccination. In some embodiments, the provided methods, compositions, and workflows are used to identify clonal lineages with many members. For example, clonal lineages with many members can represent B cells responding to chronic antigen stimulation. In some embodiments, the provided methods, compositions, and workflows are used to identify antigen-specific B cells. For example, comparing IgH repertoires across groups of individuals who have been exposed to the same antigen can reveal shared IgH amino acid motifs indicative of antigen-specific IgH chains. In some embodiments, the provided methods, compositions, and workflows are used to assess clonal overlap. For example, clonal overlap analysis can reveal developmental relationships between B cell trafficking and B cell populations. In some embodiments, the provided methods, compositions, and workflows are used to determine dominant clonal VDJ sequences included in longitudinal analysis. In some embodiments, the provided methods, compositions, and workflows are used to identify malignant subclones by clonal lineage analysis. For example, for some B cell malignancies (e.g., follicular lymphoma), somatic hypermutation is ongoing, resulting in the presence of malignant subclones with different but related IgH sequences, which can be tracked using the provided methods, compositions, and workflows.
[0220] In some embodiments, the provided methods, compositions, and workflows are used to assess clonal evolution. For example, analysis of clonal lineages can reveal isotype switching and IgH residues important for antigen binding. In some embodiments, the provided methods, compositions, and workflows are used to assess isotype abundance. For example, overexpression or underexpression of certain isotypes may indicate disease or immunodeficiency, such as but not limited to elevated IgG1 in response to viral infection, elevated IgE in allergy, and isotype deletions or underexpression that may indicate primary immunodeficiency. In some embodiments, the provided methods, compositions, and workflows are used to quantify somatic hypermutation. For example, the frequency of somatic hypermutation provides insight into the developmental stage of B cells undergoing malignant transformation.
[0221] In some embodiments, the provided methods and compositions are used to identify and / or characterize somatic hypermutation (SHM) within a BCR repertoire or clonal population. In some embodiments, the provided methods and compositions are used to identify and / or screen for rare BCR clones or subclones, such as those BCR clones or subclones having VDJ rearrangements with somatic hypermutation. In some embodiments, identifying, quantifying, and / or characterizing rare BCR clones can provide biomarkers for a given medical condition or treatment response. Thus, in some embodiments, the methods and compositions provided herein are used to identify, screen for, and / or characterize BCR clones as biomarkers using, for example, samples obtained from retrospective or longitudinal subject studies.
[0222] In some embodiments, methods for identifying and / or characterizing BCR clone lineages and SHM include: performing one or more multiplex amplification reactions on a sample from a subject using at least one set of primers and one or more C gene primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting most different V genes of at least one BCR coding sequence including at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; and performing the VDJ sequence analysis provided herein to identify and / or quantify the SMH and clone lineage of the target BCRs from the sample. In other embodiments, methods for identifying and / or characterizing BCR clone lineages and SHM include: performing one or more multiplex amplification reactions on a sample from a subject using at least one set of primers and multiple J gene primers to amplify BCR nucleic acid template molecules having a J gene portion and a variable portion, the at least one set of primers targeting most different V genes of at least one BCR coding sequence including at least a portion of FR1, FR2, or FR3 within the V gene, the multiple J gene primers targeting most different J genes of the corresponding target BCR coding sequence; sequencing the resulting BCR amplicons; and performing the VDJ sequence analysis provided herein to identify the SHM and clone lineage of the target BCRs from the sample.
[0223] In some embodiments, the provided methods and compositions are used to identify, quantify, characterize, and / or monitor isotype (or sub-isotype) classes or isotype class switching within a BCR repertoire or B cell clonal lineage. In some embodiments, such methods include: performing one or more multiplex amplification reactions on a sample from a subject using at least one set of primers and one or more C gene primers to amplify IgH nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of different IgH V gene coding sequences including at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the C gene of the IgH coding sequence; sequencing the resulting amplicons; performing sequence analysis provided herein to identify one or more IgH isotype classes of the BCR repertoire or clonal lineage of the sample. In some embodiments, the primer set includes one or more primers targeting at least a portion of the C gene for a single isotype, such as IgE. In other embodiments, the primer set includes at least two primers, each targeting at least a portion of the C gene for two different isotypes. In other embodiments, the primer set includes at least one primer separately targeting at least a portion of the C genes for the IgA, IgD, IgG, IgM, and IgE isotype types.
[0224] In certain embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations and clonal lineages, such as changes in clonal expansion, changes in clonal contraction, changes in the relative proportions of clones or clonal populations within the BCR repertoire, changes in the expansion or contraction of clonal lineages, and changes in somatic hypermutation and / or isotype class switching within the repertoire. In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations or clonal lineages in response to tumor growth (e.g., clonal population or lineage expansion, clonal population or lineage contraction, changes in clonal populations or lineages at relative ratios, somatic hypermutation and / or class switching). In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations in response to tumor treatment (e.g., clonal population or lineage expansion, clonal population or lineage contraction, changes in clonal populations or lineages at relative ratios, somatic hypermutation and / or class switching). In some embodiments, the provided methods and compositions are used to monitor changes in BCR repertoire clonal populations or clonal lineages during remission (e.g., clonal population or lineage expansion, clonal population or lineage contraction, changes in clonal populations or lineages at relative ratios, somatic hypermutation and / or class switching). For many lymphoid malignancies, clonal B cell receptor sequences can be used as biomarkers for malignant cells of a particular cancer (e.g., leukemia) and for monitoring residual disease, tumor expansion, contraction, and / or treatment response. In certain embodiments, clonal B cell receptors can be identified and further characterized to confirm new utilities in therapeutic, biomarker, and / or diagnostic applications.
[0225] In some embodiments, methods and compositions are provided for monitoring changes in BCR clonal populations in a subject, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the subject using at least one set of primers and ii) one or more C gene primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting most different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire clonal populations of the target BCRs from the sample; and comparing the identified BCR repertoire clonal populations with those identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for monitoring changes in BCR clonal populations in a subject, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the subject using at least one set of primers and a plurality of J gene primers to amplify immunoglobulin repertoire nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting most different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the plurality of J gene primers targeting most different J genes of the corresponding target BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire clonal populations of the target BCRs from the sample; and comparing the identified immunoglobulin repertoire clonal populations with those identified in samples obtained from the subject at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, such as parallel highly multiplexed amplification reactions performed with different primer pools. Samples for monitoring changes in BCR repertoire clonal populations include but are not limited to samples obtained prior to diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained at any time prior to treatment (pre-treatment samples), samples obtained at any time after treatment completion (post-treatment samples), and samples obtained during the course of treatment.
[0226] In certain embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the progression and / or treatment of a hyperproliferative disorder in the patient. In some embodiments, the provided methods and compositions are for minimal residual disease (MRD) monitoring following treatment of a patient. In some embodiments, the provided methods and compositions allow for deep sequencing of a patient's BCR repertoire that can be used for MRD measurement and for identifying rare BCR clones. In some embodiments, monitoring MRD includes assessing somatic hypermutation of the BCR repertoire. In some embodiments, the methods and compositions are for identifying and / or tracking B cell lineage malignancies or T cell lineage malignancies. In some embodiments, the methods and compositions are for detecting and / or monitoring MRD in patients diagnosed with leukemia or lymphoma, including but not limited to acute lymphoblastic leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, cutaneous T cell lymphoma, B cell lymphoma, mantle cell lymphoma, and multiple myeloma. In some embodiments, the methods and compositions are for detecting and / or monitoring MRD in patients diagnosed with solid tumors, including but not limited to breast cancer, lung cancer, colorectal cancer, and neuroblastoma. In some embodiments, the methods and compositions are for detecting and / or monitoring MRD in patients following cancer treatment, including but not limited to bone marrow transplantation, lymphocyte infusion, adoptive T cell therapy, other cell-based immunotherapies, and antibody-based immunotherapies.
[0227] In some embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the progression and / or treatment of a hyperproliferative disorder in the patient, the methods and compositions comprising: performing one or more multiplex amplification reactions using at least one set of primers on a sample from the patient or cDNA prepared from the sample to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR coding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR coding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further comprises: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying the immune repertoire of the target BCR from the sample. In certain embodiments, determining the sequence of the immune receptor amplicon molecules comprises: obtaining initial sequence reads; aligning the initial sequence reads to a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequence of the resulting immune receptor molecules. In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR coding sequence that includes at least a portion of FR3 within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR coding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR coding sequence that includes at least a portion of FR2 within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR coding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0228] In some embodiments, methods and compositions are provided for identifying and / or characterizing a patient's BCR repertoire to monitor the progression and / or treatment of a hyperproliferative disorder in the patient, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the patient or cDNA prepared from the sample using at least one set of primers to amplify immunoglobulin repertoire nucleic acid template molecules having a J gene segment and a V gene segment, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers that target a majority of the different J genes of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further comprises: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying the immunoglobulin repertoire of the target BCR from the sample. Specifically, embodiments of determining the sequence of an immunoreceptor amplicon molecule comprise: obtaining initial sequence reads; adding an inferred J gene sequence to the sequence reads to generate extended sequence reads; aligning the extended sequence reads with a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequence of the resulting immunoreceptor molecule. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers that target a majority of the different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, the multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers that target a majority of the different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immunoreceptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0229] In some embodiments, methods and compositions are provided for MRD monitoring of patients with hyperproliferative diseases, the methods and compositions comprising: performing one or more multiplex amplification reactions on a patient sample using at least one set of primers and ii) one or more C gene primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire sequences of the target BCR; and detecting the presence or absence of one or more BCR sequences in a sample associated with a hyperproliferative disease. In some embodiments, methods and compositions are provided for MRD monitoring of patients with hyperproliferative diseases, the methods and compositions comprising: performing one or more multiplex amplification reactions on a patient sample using at least one set of primers and multiple J gene primers to amplify immunoglobulin repertoire nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the multiple J gene primers targeting a majority of the different J genes of the corresponding target BCR coding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire sequences of the target BCR; and detecting the presence or absence of one or more immunoglobulin receptor sequences in a sample associated with a hyperproliferative disease. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, such as parallel highly multiplexed amplification reactions performed with different primer pools. Samples for MRD monitoring include, but are not limited to, samples obtained during remission, samples obtained at any time after completion of treatment (post-treatment samples), and samples obtained during the course of treatment.
[0230] In certain embodiments, methods and compositions are provided for identifying and / or characterizing the BCR repertoire of a subject in response to treatment. In some embodiments, the methods and compositions are used to characterize and / or monitor populations or clones of tumor infiltrating lymphocytes (TILs) before, during, and / or after tumor treatment. In some embodiments, analyzing the immune receptor repertoire of TILs provides a characterization and / or assessment of the tumor microenvironment. In some embodiments, the methods and compositions for determining the immune repertoire are used to identify and / or track one or more populations of therapeutic T cells and B cells. In some embodiments, the provided methods and compositions are used to identify and / or monitor the persistence of cell-based therapies after patient treatment, including but not limited to the presence (e.g., persistence) of engineered T cell populations, such as CAR-T cell populations, TCR-engineered T cell populations, persistent CAR-T expression, the presence (e.g., persistence) of administered TIL populations, TIL expression (e.g., persistence) after adoptive T cell therapy, and / or immune reconstitution after allogeneic hematopoietic cell transplantation.
[0231] In some embodiments, the provided methods and compositions are used to characterize and / or monitor B cell clones or populations present in a patient sample after administration of a cell-based therapy to the patient, the cell-based therapy including but not limited to, for example, cancer vaccine cells, CAR-T, TIL, and / or other engineered cell-based therapies. In some embodiments, the provided methods and compositions are used to characterize and / or monitor the BCR repertoire in a patient sample after cell-based therapy in order to evaluate and / or monitor the patient's response to the administered cell-based therapy. Samples for such characterization and / or monitoring after cell-based therapy include but are not limited to circulating blood cells, circulating tumor cells, TILs, tissue, cfDNA, and one or more tumor samples from the patient.
[0232] In some embodiments, methods and compositions are provided for monitoring cell-based therapies in patients receiving such therapies, the methods and compositions comprising: performing one or more multiplex amplification reactions on a patient sample using at least one set of primers and ii) one or more C gene primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR coding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire sequences of the target BCRs; and detecting the presence or absence of one or more BCR sequences in a sample associated with the cell-based therapy. In some embodiments, methods and compositions are provided for monitoring cell-based therapies in patients receiving such therapies, the methods and compositions comprising: performing one or more multiplex amplification reactions on a patient sample using at least one set of primers and a plurality of J gene primers to amplify BCR repertoire nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR coding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the plurality of J gene primers targeting a majority of the different J genes of the corresponding target BCR coding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire sequences of the target BCRs; and detecting the presence or absence of one or more BCR sequences in a sample associated with the cell-based therapy.
[0233] In some embodiments, methods and compositions are provided for monitoring a patient's response after administration of a cell-based therapy, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from a subject using at least one set of primers and ii) one or more C gene primers to amplify BCR repertoire nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunological repertoire sequences of the target BCRs; and comparing the identified BCR repertoire with one or more immunoreceptor sequences identified in samples obtained from the patient at different times. In some embodiments, methods and compositions are provided for monitoring a patient's response after administration of a cell-based therapy, the methods and compositions comprising: performing one or more multiplex amplification reactions on a patient's sample using at least one set of primers and multiple J gene primers to amplify BCR repertoire nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the multiple J gene primers targeting a majority of the different J genes of the corresponding target BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunological repertoire sequences of the target BCRs; and comparing the identified BCR repertoire with one or more immunoreceptor sequences identified in samples obtained from the patient at different times. Cell-based therapies suitable for such monitoring include, but are not limited to, CAR-T cells, TCR-engineered T cells, TILs, and other enriched autologous cells. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, such as parallel highly multiplexed amplification reactions performed with different primer pools. Samples for such monitoring include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained at any time before treatment (pre-treatment samples), samples obtained at any time after treatment completion (post-treatment samples), and samples obtained during the course of treatment.
[0234] In some embodiments, methods and compositions for determining B cell receptor repertoires or B cell and T cell receptor repertoires are used to measure and / or evaluate immunocompetence before, during, and / or after a treatment, the treatment including, but not limited to, solid organ transplantation or bone marrow transplantation.
[0235] In certain embodiments, the provided methods and compositions are used to identify and / or characterize a subject's BCR repertoire in response to a therapeutic treatment, the treatment including but not limited to immunotherapy, anti-allergy treatment, and anti-infective agent treatment. Thus, in some embodiments, the provided methods and compositions are used to identify BCR repertoire or clonal lineage biomarkers or signatures of a treatment response, such as a favorable response to a therapeutic treatment (e.g., successful vaccination) or an adverse response (e.g., an immune system-mediated adverse event). In some embodiments, methods and compositions are provided for identifying and / or characterizing a subject's BCR repertoire in response to a treatment, the methods and compositions including: obtaining a sample from the subject after initiation of the treatment; performing one or more multiplex amplification reactions on the sample or cDNA prepared from the sample using at least one set of primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers including: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further includes: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying the immune repertoire of the target BCRs from the sample. In some embodiments, the method further includes comparing the identified BCR repertoire from a sample obtained after initiation of the treatment with the BCR repertoire from a patient sample obtained prior to the treatment. Specifically, embodiments of determining the sequences of the BCR amplicon molecules include: obtaining initial sequence reads; aligning the initial sequence reads with a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting immune receptor molecules. In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers, the at least one set of primers including: i) a plurality of V gene primers that target a majority of different V genes of at least one BCR-encoding sequence that includes at least a portion of FR3 within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence that includes at least a portion of FR2 within the V gene, and ii) one or more C gene primers that target at least a portion of the corresponding target C gene of the BCR-encoding sequence, wherein each set of i) primers and ii) primers targeting the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0236] In some embodiments, methods and compositions are provided for identifying and / or characterizing the BCR repertoire of a subject in response to treatment, the methods and compositions comprising: obtaining a sample from a subject after initiation of treatment; performing one or more multiplex amplification reactions on the sample or cDNA prepared from the sample using at least one set of primers to amplify BCR nucleic acid template molecules having J gene segments and V gene segments, the at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of framework region 3 (FR3) within the V gene, and ii) a plurality of J gene primers that target a majority of the different J genes of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK, thereby generating BCR amplicon molecules. The method further comprises: sequencing the resulting BCR amplicon molecules; determining the sequences of the BCR amplicon molecules; and identifying the immune repertoire of the target BCR from the sample. In some embodiments, the method further comprises comparing the identified BCR repertoire from a sample obtained after initiation of treatment with the BCR repertoire from a patient sample obtained prior to treatment. Specifically, embodiments of determining the sequences of the BCR amplicon molecules comprise: obtaining initial sequence reads; adding inferred J gene sequences to the sequence reads to generate extended sequence reads; aligning the extended sequence reads with a reference sequence and identifying productive reads; correcting one or more indel errors to generate rescued productive sequence reads; and determining the sequences of the resulting BCR molecules. In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1 within the V gene, and ii) a plurality of J gene primers that target a majority of the different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK. In other embodiments of such methods and compositions, a multiplex amplification reaction is performed using at least one set of primers comprising: i) a plurality of V gene primers that target a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR2 within the V gene, and ii) a plurality of J gene primers that target a majority of the different J gene primers of the corresponding target BCR-encoding sequence, wherein each set of i) primers and ii) primers for the same target immune receptor sequence is selected from the group consisting of IgH, IgL, and IgK.
[0237] In some embodiments, methods and compositions are provided for monitoring changes in the BCR repertoire of a subject in response to treatment, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the subject or patient using at least one set of primers and one or more C gene primers to amplify BCR nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunological repertoire sequences of the target BCRs from the sample; and comparing the identified BCR repertoire with those identified in samples obtained from the subject at different times. In some embodiments, methods and compositions are provided for monitoring changes in the BCR repertoire of a subject in response to treatment, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the subject or patient using at least one set of primers and multiple J gene primers to amplify BCR nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the multiple J gene primers targeting a majority of the different J genes of the corresponding target BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunological repertoire sequences of the target BCRs from the sample; and comparing the identified BCR repertoire with the repertoires identified in samples obtained from the subject at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, such as parallel highly multiplexed amplification reactions performed with different primer pools. Samples for monitoring BCR repertoire changes include, but are not limited to, samples obtained prior to diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained at any time prior to treatment (pre-treatment samples), samples obtained at any time after treatment completion (post-treatment samples), and samples obtained during the course of treatment.
[0238] In certain embodiments, the provided methods and compositions are used to characterize and / or monitor BCR repertoires associated with one or more immune system-mediated adverse events, including but not limited to those associated with inflammatory conditions, autoimmune responses, and / or autoimmune diseases or disorders. In some embodiments, the provided methods and compositions are used to identify and / or monitor B cell immunological repertoires or B cell and T cell immunological repertoires associated with chronic autoimmune diseases or disorders, including but not limited to multiple sclerosis, type I diabetes, narcolepsy, rheumatoid arthritis, ankylosing spondylitis, asthma, and SLE. In some embodiments, one or more immunological repertoires of an individual having an autoimmune condition are determined using a systemic sample such as a blood sample. In some embodiments, one or more immunological repertoires of an individual having an autoimmune condition are determined using a local sample such as a fluid sample from an affected joint or swollen area. In some embodiments, comparing the immunological repertoire found in a local or affected area sample with the immunological repertoire found in a systemic sample can identify targeted clonal T or B cell populations for depletion.
[0239] In some embodiments, methods and compositions are provided for identifying and / or monitoring a BCR repertoire associated with the progression and / or treatment of one or more immune system-mediated adverse events in a patient. The methods and compositions include: performing one or more multiplex amplification reactions on a sample from a subject using at least one set of primers and one or more C gene primers to amplify BCR repertoire nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immune repertoire sequences of the target BCRs from the sample; and comparing the identified BCR repertoire to one or more BCR repertoires identified in samples obtained from the patient at different times. In some embodiments, methods and compositions are provided for identifying and / or monitoring a BCR repertoire associated with the progression and / or treatment of one or more immune system-mediated adverse events in a patient. The methods and compositions include: performing one or more multiplex amplification reactions on a sample from the patient using at least one set of primers and multiple J gene primers to amplify BCR nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, and the multiple J gene primers targeting a majority of the different J genes of the corresponding target BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the BCR sequences of the target immune receptors from the sample; and comparing the identified BCR repertoire to one or more BCR repertoires identified in samples obtained from the patient at different times. In various embodiments, the one or more multiplex amplification reactions performed in such methods can be a single multiplex amplification reaction or can be two or more multiplex amplification reactions performed in parallel, such as parallel highly multiplexed amplification reactions performed with different primer pools. Samples for monitoring changes in the immune repertoire associated with immune system-mediated (one or more) adverse events include, but are not limited to, samples obtained before diagnosis, samples obtained at any stage of diagnosis, samples obtained during remission, samples obtained at any time before treatment (pre-treatment samples), samples obtained at any time after treatment completion (post-treatment samples), and samples obtained during treatment.
[0240] In some embodiments, the provided methods and compositions are used to characterize and / or monitor the immunoglobulin repertoire associated with passive immunity, which includes naturally acquired passive immunity and artificially acquired passive immunotherapy. For example, the provided methods and compositions can be used to identify and / or monitor protective antibodies that provide passive immunity to a recipient after transfer of antibody-mediated immunity to the recipient, including but not limited to antibody-mediated immunity transferred from a mother to a fetus during pregnancy or to an infant through breastfeeding or via administration of antibodies to a recipient. In another example, the provided methods and compositions can be used to identify and / or monitor the B-cell and / or T-cell immunoglobulin repertoire associated with passive transfer of cell-mediated immunity to a recipient, such as administration of mature circulating lymphocytes to a recipient compatible with donor tissue. In some embodiments, the provided methods and compositions are used to monitor the duration of passive immunity in a recipient.
[0241] In some embodiments, the provided methods and compositions are used to characterize and / or monitor the immunoglobulin repertoire associated with active immunity or vaccination therapy. For example, after exposure to a vaccine or an infectious agent, the provided methods and compositions can be used to identify and / or monitor protective antibodies or a population of clonal B cells or clonal B and T cells that can provide active immunity to the exposed individual. In some embodiments, the provided methods and compositions are used to monitor the duration of B-cell clones or B- and T-cell clones that contribute to the immunity of the exposed individual. In some embodiments, the provided methods and compositions are used to identify and / or monitor the B-cell and / or T-cell immunoglobulin repertoire associated with exposure to bacterial, fungal, parasitic, or viral antigens. In some embodiments, the provided methods and compositions are used to identify and / or monitor the B-cell and / or T-cell immunoglobulin repertoire associated with bacterial, fungal, parasitic, or viral infections. Thus, in some embodiments, the provided methods and compositions are used in vaccine development, including but not limited to identifying and / or characterizing one or more responses to a vaccine candidate for quality or regulatory purposes and assessing one or more responses to a vaccine.
[0242] In some embodiments, methods and compositions are provided for monitoring changes in the BCR repertoire following exposure to a vaccine or infectious agent, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from an exposed subject using at least one set of primers and one or more C gene primers to amplify BCR repertoire nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the corresponding target C gene of the BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the immunoglobulin repertoire sequences of the target BCRs from the sample; and comparing the identified BCR repertoire to one or more BCR repertoires identified in samples obtained from the subject at different times (e.g., prior to exposure or after obtaining the tested sample). In some embodiments, methods and compositions are provided for monitoring changes in the BCR repertoire following exposure to a vaccine or infectious agent, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from an exposed subject using at least one set of primers and multiple J gene primers to amplify BCR repertoire nucleic acid template molecules having a J gene portion and a V gene portion, the at least one set of primers targeting a majority of the different V genes of at least one BCR-encoding sequence comprising at least a portion of FR1, FR2, or FR3 within the V gene, the multiple J gene primers targeting a majority of the different J genes of the corresponding target BCR-encoding sequence; sequencing the resulting BCR amplicons; identifying the BCR sequences of the target immunoglobulin receptors from the sample; and comparing the identified BCR repertoire to one or more BCR repertoires identified in samples obtained from the patient at different times. In certain embodiments, methods and compositions are provided for monitoring changes in the BCR repertoire following exposure to a vaccine or infectious agent, the methods and compositions comprising: performing one or more multiplex amplification reactions on cDNA prepared from a sample from an exposed subject using at least one set of primers to amplify IgH nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers comprising: i) multiple V gene primers targeting a majority of the different IgH V genes comprising at least a portion of FR1, FR2, or FR3 within the V gene and one or more C gene primers targeting at least a portion of the IgH C gene; sequencing the resulting BCR amplicons; identifying the expressed IgH repertoire sequences from the sample that contain repertoire isotype information; and comparing the identified IgH repertoire to one or more IgH repertoires identified in samples obtained from the subject at different times (e.g., prior to exposure or after obtaining the tested sample). In some embodiments, the primer set comprises one or more primers targeting at least a portion of the C gene for a single isotype such as IgG. In other embodiments, the primer set comprises at least two primers, each targeting at least a portion of the C gene for two different isotypes.In other embodiments, the primer set includes at least one primer that specifically targets at least a portion of the C gene for each of the IgA, IgD, IgG, IgM, and IgE isotype types. Accordingly, the methods and compositions can be used to monitor changes (including isotype class switching) in the B cell repertoire and to evaluate a subject's response to vaccine exposure.
[0243] In certain embodiments, methods and compositions are provided for identifying and / or characterizing an IgE repertoire of a subject following exposure to an allergen or agent that induces an allergic reaction or response, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the exposed subject using at least one set of primers and one or more C gene primers to amplify IgH repertoire nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting most of the different IgH V genes including at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the IgE gene coding sequence; sequencing the resulting IgH amplicons; and identifying expressed IgE immunoglobulin repertoire sequences from the sample. In some embodiments, methods and compositions are provided for monitoring changes in an IgE repertoire of a subject following exposure to an allergen or agent that induces an allergic reaction or response, the methods and compositions comprising: performing one or more multiplex amplification reactions on a sample from the exposed subject using at least one set of primers and one or more C gene primers to amplify IgH repertoire nucleic acid template molecules having a constant portion and a variable portion, the at least one set of primers targeting most of the different IgH V genes including at least a portion of FR1, FR2, or FR3 within the V gene, the one or more C gene primers targeting at least a portion of the IgE gene coding sequence; sequencing the resulting IgH amplicons; identifying expressed IgE immunoglobulin repertoire sequences from the sample; and comparing the identified IgE repertoire to one or more IgE repertoires identified in samples obtained from the subject at different times (e.g., before exposure or after obtaining the sample being tested). In some embodiments, at least one primer set of such methods and compositions includes additional C gene primers that target at least a portion of other IgH isotypes such as primers for IgG, IgM, IgA, and / or IgD. In other embodiments, the primer set includes at least one primer that specifically targets at least a portion of the C gene for each of the IgA, IgD, IgG, IgM, and IgE isotype types. Accordingly, the methods and compositions can be used to monitor changes (including isotype class switching) in the IgE repertoire within the total BCR repertoire and to evaluate an allergic reaction or response of a subject to allergen exposure. In some embodiments, such methods and compositions are used to determine and / or monitor the isotype switching origin of IgE-expressing B cells within the repertoire.
[0244] In some embodiments, the provided methods and compositions are used to screen or characterize lymphocyte populations that are grown and / or activated in vitro and used as immunotherapeutic agents or in immunotherapy-based regimens. In some embodiments, the provided methods and compositions are used to screen or characterize tumor-infiltrating lymphocyte (TIL) populations or other harvested B cell populations that are grown and / or activated in vitro. In some embodiments, determining the IgH sequence of the BCR facilitates the identification and generation of antigen-specific B cells. In some embodiments, the provided methods and compositions are used to screen or characterize engineered B cell populations that are grown and / or activated in vitro, such as for immunotherapy or antibody production. In some embodiments, the provided methods and compositions are used to evaluate a cell population by monitoring the BCR repertoire during an ex vivo workflow for manufacturing engineered cell preparations, for example, for quality control or regulatory testing purposes.
[0245] In some embodiments, the sequences of the novel or non-classical BCR alleles identified as described herein can be used to generate recombinant BCR nucleic acids or molecules. In some embodiments, the provided methods and compositions are used to screen and / or generate recombinant antibody libraries. The provided compositions related to the identification of BCRs can be used to rapidly evaluate recombinant antibody library size and composition to identify antibodies of interest.
[0246] In some embodiments, analyzing the immune receptor repertoire as provided herein can be combined with analyzing immune response gene expression to provide a characterization of the tumor microenvironment. In some embodiments, combining or correlating the BCR repertoire profile of a tumor sample with a targeted immune response gene expression profile provides a more thorough analysis of the tumor microenvironment and can suggest or provide guidance for immunotherapy treatment.
[0247] Suitable cells for analysis include, but are not limited to, various hematopoietic cells, lymphocytes, and tumor cells, such as peripheral blood mononuclear cells (PBMCs), T cells, B cells, circulating tumor cells, and tumor-infiltrating lymphocytes (TILs). Lymphocytes expressing immunoglobulins include pre-B cells, B cells, such as memory B cells and plasma cells. Lymphocytes expressing T cell receptors include thymocytes, NK cells, pre-T cells, and T cells, and many subsets of T cells are known in the art, such as Th1, Th2, Th17, CTL, T reg, etc. For example, in some embodiments, a sample comprising PBMCs can be used as a source for antibody immunoglobulin repertoire analysis. The sample can contain, for example, lymphocytes, monocytes, and macrophages, as well as antibodies and other biological components.
[0248] Analysis of BCR repertoire is relevant for conditions involving cell proliferation and antigen exposure, including but not limited to the presence of cancer, exposure to cancer antigens, exposure to antigens from infectious agents, exposure to vaccines, exposure to allergens, exposure to food, the presence of grafts or implants, and the presence of autoimmune activity or disease. Conditions associated with immunodeficiency are also relevant for analysis, including congenital and acquired immunodeficiency syndromes.
[0249] B-cell lineage malignancies of interest include but are not limited to multiple myeloma; acute lymphoblastic leukemia (ALL); relapsed / refractory B-cell ALL, chronic lymphocytic leukemia (CLL); diffuse large B-cell lymphoma; mucosa-associated lymphoid tissue lymphoma (MALT); small lymphocytic lymphoma; mantle cell lymphoma (MCL); Burkitt lymphoma; mediastinal large B-cell lymphoma; Waldenström macroglobulinemia ( macroglobulinemia); nodal marginal zone B-cell lymphoma (NMZL); splenic marginal zone lymphoma (SMZL); intravascular large B-cell lymphoma; primary effusion lymphoma; lymphomatoid granulomatosis, etc. Non-malignant B-cell hyperproliferative conditions include monoclonal B-cell lymphocytosis (MBL).
[0250] T-cell lineage malignancies of interest include but are not limited to precursor T-cell lymphoblastic lymphoma; T-cell prolymphocytic leukemia; T-cell granular lymphocytic leukemia; aggressive NK-cell leukemia; adult T-cell lymphoma / leukemia (HTLV 1-positive); extranodal NK / T-cell lymphoma; enteropathy-type T-cell lymphoma; hepatosplenic γδ T-cell lymphoma; subcutaneous panniculitis-like T-cell lymphoma; mycosis fungoides / Sezary syndrome; anaplastic large cell lymphoma, T / null cell; peripheral T-cell lymphoma; angioimmunoblastic T-cell lymphoma; chronic lymphocytic leukemia (CLL); acute lymphoblastic leukemia (ALL); prolymphocytic leukemia; and hairy cell leukemia.
[0251] Other malignancies of interest include but are not limited to acute myeloid leukemia, head and neck cancer, brain cancer, breast cancer, ovarian cancer, cervical cancer, colorectal cancer, endometrial cancer, gallbladder cancer, gastric cancer, bladder cancer, prostate cancer, testicular cancer, liver cancer, lung cancer, kidney (renal cell) cancer, esophageal cancer, pancreatic cancer, thyroid cancer, cholangiocarcinoma, pituitary tumors, Wilms tumor, Kaposi sarcoma, osteosarcoma, thymic cancer, skin cancer, heart cancer, oral and pharyngeal cancers, neuroblastoma, and non-Hodgkin lymphoma.
[0252] Attention is paid to neuropathic conditions such as Alzheimer's Disease, Parkinson's Disease, Lou Gehrig's Disease, etc., and demyelinating diseases such as multiple sclerosis, chronic inflammatory demyelinating polyneuropathy, etc., and inflammatory conditions such as rheumatoid arthritis. Systemic lupus erythematosus (SLE) is an autoimmune disease characterized by polyclonal B cell activation, which produces a variety of anti-protein and non-protein autoantibodies (see Kotzin et al. (1996) Cell 85:303-306). These autoantibodies form immune complexes that deposit in multiple organ systems, leading to tissue damage. The autoimmune components can be attributed to atherosclerosis, where candidate autoantigens include Hsp60, oxidized LDL, and 2-glycoprotein I (2GPI).
[0253] Samples for use in the methods described herein can be samples collected from a subject having a malignancy or a hyperproliferative condition, including lymphoma, leukemia, and plasmacytoma. Lymphoma is a solid tumor of lymphocytic origin and is most commonly found in lymphoid tissue. Thus, for example, a biopsy from a lymph node (such as the tonsil) containing such lymphoma would constitute a suitable biopsy. Samples can be obtained from a subject or patient at one or more time points during disease progression and / or disease treatment.
[0254] In some embodiments, the present disclosure provides a method for performing target-specific multiplex PCR on a cDNA sample having multiple expressed immune receptor target sequences using primers having a cleavable group.
[0255] In certain embodiments, the library to be sequenced and / or the template preparation uses an automated system, such as an Ion Chef TM system, and is automatically prepared from a nucleic acid sample population using the compositions provided herein.
[0256] As used herein, the term "subject" includes humans, patients, individuals, persons being evaluated, etc.
[0257] As used herein, the terms "comprises / comprising", "includes / including", "has / having" or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of features is not necessarily limited to those features, but may include other features not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" means inclusive or rather than exclusive or.
[0258] As used herein, "antigen" refers to any substance that can stimulate an immune response when introduced, for example, into a subject, such as the production of antibodies or T cell receptors that recognize the antigen. Antigens include molecules such as nucleic acids, lipids, ribonucleoprotein complexes, protein complexes, proteins, polypeptides, peptides, and naturally occurring or synthetic modifications of such molecules, to which an immune response involving T and / or B lymphocytes can be generated. With respect to autoimmune diseases, the antigens herein are typically referred to as autoantigens. With respect to allergic diseases, the antigens herein are typically referred to as allergens. An autoantigen is any molecule produced by an organism that can be the target of an immune response, including peptides, polypeptides, and proteins encoded within the genome of the organism, as well as post-translationally generated modifications of these peptides, polypeptides, and proteins. Such molecules also include carbohydrates, lipids, and other molecules produced by the organism. Antigens also include vaccine antigens, which include, but are not limited to, pathogen antigens, cancer-associated antigens, allergens, and the like.
[0259] As used herein, "amplify / amplifying" or "amplification reaction" and their derivatives refer to any action or process of replicating or copying at least a portion of a nucleic acid molecule (referred to as a template nucleic acid molecule) into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally contains a sequence that is substantially the same as or substantially complementary to at least some portions of the template nucleic acid molecule. The template nucleic acid molecule can be single-stranded or double-stranded and the other nucleic acid molecule can independently be single-stranded or double-stranded. In some embodiments, amplification comprises a template-dependent in vitro enzymatic reaction that is used to prepare at least one copy of at least a portion of a nucleic acid molecule or to prepare at least one copy of a nucleic acid sequence that is complementary to at least a portion of a nucleic acid molecule. Amplification optionally comprises linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification is carried out using isothermal conditions; in other embodiments, such amplification can comprise thermal cycling. In some embodiments, the amplification is multiplex amplification, which comprises simultaneously amplifying multiple target sequences in a single amplification reaction. At least some of the target sequences can be on the same nucleic acid molecule or on different target nucleic acid molecules included in a single amplification reaction. In some embodiments, "amplification" comprises amplification of at least a portion of DNA- and RNA-based nucleic acids, either alone or in combination. The amplification reaction can comprise single-stranded or double-stranded nucleic acid substrates and can further comprise any of the amplification methods known to those of ordinary skill in the art. In some embodiments, the amplification reaction comprises PCR.
[0260] As used herein, "amplification conditions" and derivatives thereof refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, the amplification conditions can include isothermal conditions or alternatively can include thermal cycling conditions or a combination of isothermal and thermal cycling conditions. In some embodiments, the conditions suitable for amplifying one or more nucleic acid sequences include PCR conditions. Generally, amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid such as one or more target sequences or an amplified target sequence (e.g., an adaptor-ligated amplified target sequence) ligated to one or more adaptors. Amplification conditions include a catalyst for amplification or for nucleic acid synthesis, such as a polymerase; primers having a degree of complementarity to the nucleic acid to be amplified; and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs), to facilitate primer extension after hybridization to the nucleic acid. Amplification conditions require hybridization or annealing of the primers to the nucleic acid, primer extension, and denaturation steps, where the extended primers are separated from the nucleic acid sequence undergoing amplification. Generally but not necessarily, amplification conditions can include thermal cycling, and in some embodiments, the amplification conditions include multiple cycles where the annealing, extension, and separation steps are repeated. Generally, amplification conditions include cations, such as Mg 2+ or Mn 2+ (e.g., MgCl 2 etc.), and can also include various ionic strength modifiers.
[0261] As used herein, "target sequence" or "target target sequence" and derivatives thereof refer to any single-stranded or double-stranded nucleic acid sequence that can be amplified or synthesized according to the present disclosure, including any nucleic acid sequence suspected or expected to be present in a sample. In some embodiments, prior to the addition of target-specific primers or attachment of adaptors, the target sequence exists in double-stranded form and includes at least a portion of a specific nucleotide sequence to be amplified or synthesized or its complement. The target sequence can include a nucleic acid that can hybridize to a primer suitable for an amplification or synthesis reaction prior to polymerase extension. In some embodiments, the term refers to a nucleic acid sequence whose sequence identity, nucleotide order, or position is determined by one or more of the methods of the present disclosure.
[0262] As defined herein, "sample" and derivatives thereof are used in their broadest sense and include any specimen, culture, etc. suspected of containing a target. In some embodiments, the sample includes nucleic acids in cDNA, RNA, PNA, LNA, chimeric, hybrid, or multiplex forms. The sample can include any biological-, clinical-, surgical-, agricultural-, atmospheric-, or aquatic-based specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as an expressed RNA, a freshly frozen or formalin-fixed paraffin-embedded nucleic acid specimen.
[0263] As used herein, when used in reference to two or more components, "contact" and its derivatives refer to any process for facilitating or effecting the proximity, closeness, admixture, or blending of the reference components without necessarily requiring physical contact of such components, and includes the mixing of solutions containing any one or more of the reference components with each other. The components so referred to may be contacted in any particular order or combination, and the particular order of component recitation is not limiting. For example, "contacting A with B and C" encompasses embodiments in which A is first contacted with B and then with C, and embodiments in which C is contacted with A and then with B, and embodiments in which a mixture of A and C is contacted with B, etc. Further, such contact does not necessarily require that the end result of the contacting process be a mixture containing all of the recited components, so long as, at some point in the contacting process, all of the recited components are present simultaneously or are contained in the same mixture or solution simultaneously. When one or more of the recited components to be contacted comprises a plurality (e.g., "contacting a target sequence with a plurality of target-specific primers and a polymerase"), each member of the plurality may be considered a single component of the contacting process, such that the contact may include any one or more members of the plurality contacting any other member of the plurality and / or any other recited component in any order or combination (e.g., some but not all of the plurality of target-specific primers may contact the target sequence, then the polymerase, and then contact the other members of the plurality of target-specific primers).
[0264] As used herein, the term "primer" and its derivatives refer to any polynucleotide that can hybridize to a relevant target sequence. In some embodiments, a primer can also be used to initiate nucleic acid synthesis. Generally, a primer serves as a substrate upon which nucleotides can be polymerized by a polymerase; however, in some embodiments, a primer can become incorporated into the synthesized nucleic acid strand and provide a site to which another primer can hybridize to initiate synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can comprise any combination of nucleotides or nucleotide analogs, which can optionally be linked to form a linear polymer of any suitable length. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide. (In the present disclosure, the terms "polynucleotide" and "oligonucleotide" can be used interchangeably herein and do not necessarily indicate any difference in length between two nucleotides). In some embodiments, a primer is single-stranded, but it can also be double-stranded. A primer is optionally naturally occurring, such as in a purified restriction digest, or can be produced synthetically. In some embodiments, when exposed to amplification or synthesis conditions, a primer serves as a starting point for amplification or synthesis; such amplification or synthesis can be carried out in a template-dependent manner and optionally forms a primer extension product complementary to at least a portion of the target sequence. Exemplary amplification or synthesis conditions can comprise contacting the primer with a polynucleotide template (e.g., a template comprising the target sequence), nucleotides, and an inducer (such as a polymerase) at a suitable temperature and pH to induce polymerization of nucleotides at the terminus of the target-specific primer. If double-stranded, the primer can optionally be treated to separate its strands prior to use in preparing the primer extension product. In some embodiments, a primer is an oligodeoxynucleotide or oligoribonucleotide. In some embodiments, a primer can comprise one or more nucleotide analogs. The exact length and / or composition (including sequence) of a target-specific primer can affect a number of properties, including melting temperature (T m) GC content, formation of secondary structures, repetitive nucleotide backbone, length of the predicted primer extension product, degree of coverage across the relevant nucleic acid molecule, number of primers in a single amplification or synthesis reaction, presence of nucleotide analogs or modified nucleotides within the primer, etc. In some embodiments, a primer can pair with a compatible primer within an amplification or synthesis reaction to form a primer pair consisting of a forward primer and a reverse primer. In some embodiments, the forward primer of the primer pair comprises a sequence that is substantially complementary to at least a portion of one strand of the nucleic acid molecule, and the reverse primer of the primer pair comprises a sequence that is substantially identical to at least a portion of said strand. In some embodiments, the forward primer and the reverse primer are capable of hybridizing to opposite strands of a nucleic acid duplex. Optionally, the forward primer primes the synthesis of a first nucleic acid strand, and the reverse primer primes the synthesis of a second nucleic acid strand, wherein the first and second strands are substantially complementary to each other or can hybridize to form a double-stranded nucleic acid molecule. In some embodiments, one end of the amplification or synthesis product is defined by the forward primer and the other end of the amplification or synthesis product is defined by the reverse primer. In some embodiments, when it is desired to amplify or synthesize a long primer extension product (such as amplifying an exon, coding region, or gene), several primer pairs can be generated rather than spanning the desired length to achieve sufficient amplification of the region. In some embodiments, a primer can comprise one or more cleavable groups. In some embodiments, the primer length ranges from about 10 to about 60 nucleotides, from about 12 to about 50 nucleotides, and from about 15 to about 40 nucleotides in length. Generally, when exposed to amplification conditions in the presence of dNTPs and polymerase, the primer is capable of hybridizing to the corresponding target sequence and undergoing primer extension. In some embodiments, the primer contains one or more cleavable groups at one or more positions within the primer.
[0265] As used herein, "target-specific primer" and derivatives thereof refer to single-stranded or double-stranded polynucleotides, typically oligonucleotides, that comprise at least one sequence that is at least 50% complementary, typically at least 75% complementary or at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% or at least 99% complementary or identical to at least a portion of a nucleic acid molecule comprising a target sequence. In such cases, the target-specific primer and the target sequence are described as "corresponding" to each other. In some embodiments, the target-specific primer is capable of hybridizing to at least a portion of its corresponding target sequence (or the complement of the target sequence); such hybridization can optionally be carried out under standard hybridization conditions or under stringent hybridization conditions. In some embodiments, the target-specific primer is not capable of hybridizing to the target sequence or its complement, but is capable of hybridizing to a portion of the nucleic acid strand comprising the target sequence or its complement. In some embodiments, the target-specific primer comprises at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary or more typically at least 99% complementary to at least a portion of the target sequence itself; in other embodiments, the target-specific primer comprises at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary or more typically at least 99% complementary to at least a portion of a nucleic acid molecule other than the target sequence. In some embodiments, the target-specific primer is substantially non-complementary to other target sequences present in a sample; optionally, the target-specific primer is substantially non-complementary to other nucleic acid molecules present in a sample. In some embodiments, nucleic acid molecules present in a sample that do not contain or correspond to the target sequence (or the complement of the target sequence) are referred to as "non-specific" sequences or "non-specific nucleic acids". In some embodiments, the target-specific primer is designed to comprise a nucleotide sequence that is substantially complementary to at least a portion of its corresponding target sequence. In some embodiments, the target-specific primer is at least 95% complementary or at least 99% complementary or identical (over its entire length) to at least a portion of the nucleic acid molecule comprising its corresponding target sequence. In some embodiments, the target-specific primer is at least 90%, at least 95%, at least 98% or at least 99% complementary or identical (over its entire length) to at least a portion of its corresponding target sequence. In some embodiments, a forward target-specific primer and a reverse target-specific primer define a target-specific primer pair that is used to amplify a target sequence via template-dependent primer extension. Typically, each primer of the target-specific primer pair comprises at least one sequence that is substantially complementary to at least a portion of a nucleic acid molecule comprising the corresponding target sequence, but less than 50% complementary to at least one other target sequence in the sample.In some embodiments, the amplification is performed using multiple target-specific primer pairs in a single amplification reaction, where each primer pair comprises a forward target-specific primer and a reverse target-specific primer, each comprising at least one sequence that is substantially complementary or substantially identical to a corresponding target sequence in the sample, and each primer pair has a different corresponding target sequence. In some embodiments, the target-specific primers are substantially non-complementary to any other target-specific primer in the amplification reaction at their 3'-end or their 5'-end. In some embodiments, the target-specific primers may comprise minimal cross-hybridization with other target-specific primers in the amplification reaction. In some embodiments, the target-specific primers comprise minimal cross-hybridization with non-specific sequences in the amplification reaction mixture. In some embodiments, the target-specific primers comprise minimal self-complementarity. In some embodiments, the target-specific primers may comprise one or more cleavable groups located at the 3'-end. In some embodiments, the target-specific primers may comprise one or more cleavable groups located near or around the central nucleotide of the target-specific primer. In some embodiments, one of the more target-specific primers comprises only non-cleavable nucleotides at the 5'-end of the target-specific primer. In some embodiments, optionally in the same amplification reaction, the target-specific primers comprise minimal nucleotide sequence overlap at the 3'-end or 5'-end of the primer compared to one or more different target-specific primers. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more target-specific primers in a single reaction mixture comprise one or more of the above embodiments. In some embodiments, substantially all of the multiple target-specific primers in a single rea...
Claims
1. A method for amplifying expressed nucleic acid sequences of a B cell receptor (BCR) repertoire in a sample, the method comprising: performing a single multiplex amplification reaction using at least one set of the following to amplify expressed target BCR nucleic acid template molecules: i) (a) Multiple V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of framework region 1 (FR1) within the V gene, (b) Multiple V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of framework region 2 (FR2) within the V gene, or (c) Multiple V gene primers for a majority of different V genes of at least one BCR coding sequence that includes at least a portion of framework region 3 (FR3) within the V gene; and ii) (a) One or more C gene primers that target at least a portion of the C gene of the at least one BCR coding sequence, or (b) Multiple J gene primers that target at least a portion of a majority of different J genes of the at least one BCR coding sequence; wherein each set of i) primers and ii) primers targets the coding sequence of the same target BCR gene selected from IgH, IgL, and IgK genes, wherein the multiple V gene primers, the multiple J gene primers, and / or the one or more C gene primers include one or more cleavable groups that are located (i) near or at the end of the primer, or (ii) near or around the central nucleotide of the primer, and wherein performing the amplification using the at least one set of i) primers and ii) primers generates amplicon molecules representative of the target BCR repertoire in the sample; wherein each set of i) primers and ii) primers is: A) SEQ ID NOs: 69 - 136, and 443 - 446; B) SEQ ID NOs: 311 - 336, and 461, 462, 481, 515 - 519, 554, 585, 586; C) SEQ ID NOs: 321 - 328, 337 - 354, and 463, 464, 482, 520 - 524, 555, 587, 588; D) SEQ ID NOs: 321 - 328, 335, 336, 347 - 352, 355 - 364, and 465, 466, 483, 525 - 529, 556, 589, 590; E) SEQ ID NOs: 321 - 328, 347 - 354, 365 - 374, and 467, 468, 484, 530 - 534, 557, 591, 592; F) SEQ ID NOs: 321 - 328, 375 - 393, and 469, 485, 535, 558, 559, 593 - 595; G) SEQ ID NOs: 386 - 391, 394 - 410, and 470, 486, 536, 537, 560, 561, 596 - 598; H) SEQ ID NOs: 295 - 301, 411 - 430, and 471, 487, 538, 539, 562, 563, 599 - 601; I) SEQ ID NOs: 321 - 328, 375 - 393, and 469, 485, 535, 558, 559, 593, 595; or J) SEQ ID NOs: 295 - 301, 321 - 328, 375 - 393, 411 - 430, and 469, 485, 535, 558, 559, 593, 595; and treating the amplification reaction product with any suitable reagent that will selectively cleave or otherwise selectively disrupt the nucleotide bonds of the one or more cleavable groups, thereby producing target BCR amplicon molecules comprising the expressed target BCR repertoire.
2. The method according to claim 1, wherein at least one of the plurality of V gene primers, the plurality of J gene primers, and / or the one or more C gene primers has any one or more of the following criteria: (1) comprising two or more modified nucleotides within the primer, at least one of the nucleotides being near or at the end of the primer and at least one of the nucleotides being at or around the central nucleotide position of the primer; (2) having a length of 15 to 40 bases; (3)T m is: higher than 60°C, up to 70°C; (4) having low cross - reactivity with non - target sequences present in the sample; (5) at least the first four nucleotides in the 3' to 5' direction not being complementary to any sequence within any other primer present in the same reaction; and (6) not being complementary to any continuous segment of at least 5 nucleotides within any other resulting target amplicon.
3. The method according to claim 1 or claim 2, wherein at least one of the plurality of V gene primers, the plurality of J gene primers, and / or the one or more C gene primers comprises one or more cleavable groups located at (i) near or at the end of the primer or (ii) near or around the central nucleotide of the primer.
4. The method according to claim 1 or claim 2, wherein at least one of the plurality of V gene primers, the plurality of J gene primers, and / or the one or more C gene primers comprises two or more modified nucleotides having a cleavable group selected from: methylguanine, 8 - oxo - guanine, xanthine, hypoxanthine, 5,6 - dihydro - uracil, uracil, 5 - methylcytosine, thymine dimer, 7 - methylguanosine, 8 - oxo - deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5 - methylcytidine.
5. The method according to claim 1 or claim 2, wherein the at least one set of i) and ii) is i)(a) and ii)(a), wherein the plurality of V gene primers anneal to at least a portion of the FR1 portion of the template molecule, and wherein the one or more C gene primers comprise at least five primers that anneal to at least a portion of the C gene portion of the template molecule.
6. The method according to claim 5, wherein the resulting target BCR amplicon molecules comprise complementarity determining regions CDR1, CDR2, and CDR3 of the target BCR gene sequence.
7. The method according to claim 5, wherein the at least one set of i) and ii) are primers selected from Tables 3 and 6 - 10, respectively.
8. The method according to claim 5, wherein the at least one set of i) and ii) is selected from the primer sets of Table 13.
9. The method according to claim 1 or claim 2, wherein the at least one set of i) and ii) is i)(c) and ii)(b), wherein the plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule, and wherein the plurality of J gene primers comprise at least two primers that anneal to at least a portion of the J gene portion of the template molecule.
10. The method according to claim 9, wherein the resulting target BCR amplicon molecules comprise the complementarity determining region CDR3 of the target BCR gene sequence.
11. The method according to claim 9, wherein the at least one set of i) and ii) is selected from the primers of Tables 2 and 5.
12. A method for preparing an expressed BCR repertoire library, the method comprising: i) treating the target BCR amplicon molecules according to any one of claims 1 to 11 to form blunt - ended amplicon molecules; and ii) ligating at least one adaptor to at least one of the target BCR amplicon molecules in the treated target BCR amplicon molecules, thereby producing a library of adaptor - ligated target BCR amplicon molecules comprising a target BCR repertoire.
13. The method according to claim 12, wherein the step of preparing the library is performed in a single reaction vessel with only addition steps.
14. The method according to claim 12 or 13, wherein the adaptor is a single - stranded or double - stranded adaptor.
15. The method according to claim 12 or claim 13, wherein the adaptor comprises a barcode, a tag, or a universal primer sequence.
16. The method according to claim 12 or claim 13, wherein the ligation comprises ligating different adaptors to each end of at least one of the treated amplicon molecules.
17. The method according to claim 16, wherein each of the two different adaptors comprises a different barcode sequence.
18. The method according to claim 12 or claim 13, wherein the ligation is performed by blunt - end ligation.
19. The method according to claim 12 or claim 13, wherein the method further comprises clonal amplification of at least a portion of an adaptor-ligated target immunoreceptor amplicon molecule.
20. A method for providing sequences of an expressed BCR repertoire in a sample, the method comprising: i) performing sequencing on a target BCR library according to any one of claims 12 to 19; ii) determining the sequences of library molecules, wherein determining the sequences comprises obtaining initial sequence reads, aligning the initial sequence reads with a reference sequence, identifying productive reads, and correcting one or more indel errors to produce rescued productive sequence reads; and iii) reporting the sequences determined for the library molecules, thereby providing the sequences of the expressed BCR repertoire in the sample.
21. The method according to claim 20, wherein when the primer set in the amplification reaction comprises the plurality of J gene primers of ii)(b), determining the sequences of ii) further comprises inferring the sequences of the J gene primers and the target J gene, and adding the inferred J gene sequences to the initial sequence reads prior to the alignment.
22. The method according to claim 20 or claim 21, further comprising sequence read clustering and BCR clonotype reporting.
23. The method according to claim 20 or claim 21, wherein the combination of productive reads and rescued productive reads is at least 50% of the reported sequencing reads.
24. The method according to claim 20 or claim 21, wherein the nucleic acid is cDNA produced from RNA molecules extracted from a biological sample by reverse transcription.
25. A method for amplifying rearranged genomic DNA (gDNA) sequences of a B cell receptor (BCR) repertoire in a sample, the method comprising: performing a single multiplex amplification reaction using at least one of the following sets to amplify a target BCR gDNA template molecule having a J gene portion and a V gene portion, the target BCR gDNA having a rearranged VDJ or VJ gene segment: i) (a) a plurality of V gene primers for most different V genes of at least one BCR coding sequence comprising at least a portion of FR1 within the V gene, (b) a plurality of V gene primers for most different V genes of at least one BCR coding sequence comprising at least a portion of FR2 within the V gene, or (c) a plurality of V gene primers for most different V genes of at least one BCR coding sequence comprising at least a portion of FR3 within the V gene; and ii) a plurality of J gene primers, the plurality of J gene primers targeting at least a portion of most different J genes of the at least one BCR coding sequence. Each set of (i) primers and (ii) primers is directed against the coding sequence of the same target BCR gene selected from IgH, IgL, and IgK genes, wherein the plurality of V gene primers and / or the plurality of J gene primers comprise one or more cleavable groups positioned near or at the end of (i) the primer, or near or around the central nucleotide of (ii) the primer, and wherein performing the amplification using said at least one set of (i) primers and (ii) primers generates amplicon molecules representative of the target BCR repertoire in the sample. Wherein the (i) primers and (ii) primers are respectively the following primer sets: SEQ ID NOs: 69 - 136, and 443 - 446; and Treat the amplification reaction product with any suitable reagent that will selectively cleave or otherwise selectively disrupt the nucleotide bonds of the one or more cleavable groups. Thereby generating target BCR amplicon molecules comprising the target BCR repertoire.
26. The method according to claim 25, wherein at least one of the plurality of V gene primers and the plurality of J gene primers has any one or more of the following criteria: (1) Contains two or more modified nucleotides within the primer, at least one of said nucleotides being near or at the end of the primer, and at least one of said nucleotides being at or around the central nucleotide position of the primer; (2) Is 15 to 40 bases in length; (3)T m is: higher than 60 °C, up to 70 °C; (4) Has low cross-reactivity with non-target sequences present in the sample; (5) At least the first four nucleotides from the 3' to 5' direction are not complementary to any sequence within any other primer present in the same reaction; and (6) Is not complementary to any continuous segment of at least 5 nucleotides within any other generated target amplicon.
27. The method according to claim 25 or claim 26, wherein at least one of the plurality of V gene primers and / or the plurality of J gene primers comprises one or more cleavable groups positioned (i) near or at the end of the primer or (ii) near or around the central nucleotide of the primer.
28. The method according to claim 25 or claim 26, wherein at least one of the plurality of V gene primers and / or the plurality of J gene primers comprises two or more modified nucleotides with cleavable groups selected from: methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine.
29. The method according to claim 25 or claim 26, wherein the at least one set of i) and ii) is i)(c) and ii), wherein the plurality of V gene primers anneal to at least a portion of the FR3 portion of the template molecule, and wherein the plurality of J gene primers comprise at least two primers that anneal to at least a portion of the J gene portion of the template molecule.
30. The method according to claim 29, wherein the resulting target BCR amplicon molecules comprise the complementarity determining region CDR3 of the target BCR gene sequence.
31. The method according to claim 29, wherein the at least one set of i) and ii) are primers selected from Table 2 and Table 5, respectively.
32. A method for preparing a library of rearranged gDNA BCR repertoires, the method comprising: i) treating the target BCR amplicon molecules according to any one of claims 25 to 31 to form blunt-ended amplicon molecules; and ii) ligating at least one adaptor to at least one of the target BCR amplicon molecules in the treated target BCR amplicon molecules, thereby generating a library of adaptor-ligated target BCR amplicon molecules comprising a target BCR repertoire.
33. The method according to claim 32, wherein the step of preparing the library is performed in a single reaction vessel with only addition steps.
34. The method according to claim 32 or claim 33, wherein the adaptor is a single-stranded or double-stranded adaptor.
35. The method according to claim 32 or claim 33, wherein the adaptor comprises a barcode, a tag, or a universal primer sequence.
36. The method according to claim 32 or claim 33, wherein the ligation comprises ligating different adaptors to each end of at least one of the amplified molecules in the treated amplified molecules.
37. The method according to claim 36, wherein each of the two different adaptors comprises a different barcode sequence.
38. The method according to claim 32 or claim 33, wherein the ligation is performed by blunt-end ligation.
39. The method according to claim 32 or claim 33, wherein the method further comprises clonal amplification of a portion of at least one adaptor-ligated target immune receptor amplicon molecule.
40. A method for providing sequences of rearranged gDNA BCR repertoires in a sample, the method comprising: i) sequencing the target BCR library according to any one of claims 32 to 39; ii) determining the sequences of the library molecules, wherein determining the sequences comprises obtaining initial sequence reads, aligning the initial sequence reads with a reference sequence, identifying productive reads, and correcting one or more indel errors to generate rescued productive sequence reads; and iii) reporting the sequences determined for the library molecules, thereby providing the sequences of the rearranged gDNA BCR repertoires in the sample.
41. The method according to claim 40, wherein when the primer set in the amplification reaction comprises the plurality of J gene primers, determining the sequence of ii) further comprises inferring the sequences of the J gene primers and the target J gene, and adding the inferred J gene sequence to the initial sequence reads before the alignment.
42. The method according to claim 40 or claim 41, further comprising sequence read clustering and BCR clonotype reporting.
43. The method according to claim 40 or claim 41, wherein the combination of productive reads and rescued productive reads is at least 50% of the reported sequencing reads.
44. A method for screening biomarkers for a disease or condition in a subject, the method comprising: performing a single multiplex amplification reaction according to claim 1 or claim 25 to amplify target BCR nucleic acid template molecules from a sample of the subject; performing sequencing on the target BCR amplicons molecules and determining the sequences of the molecules, wherein determining the sequences comprises obtaining initial sequence reads, aligning the initial sequence reads with a reference sequence, identifying productive reads, and correcting one or more indel errors to produce rescued productive sequence reads; identifying BCR repertoire clone populations from the determined target BCR sequences; and identifying the sequences of at least one BCR clone to serve as a biomarker for the disease or condition of the subject.
45. The method according to claim 44, wherein the disease or condition is selected from cancer, autoimmune diseases, infectious diseases, allergies, response to vaccination, and response to immunotherapy treatment.
46. The method according to claim 1 or 25, wherein the target BCR gene is IgH.
47. The method according to claim 1 or 25, wherein the sample comprises hematopoietic cells, lymphocytes, tumor cells, or cell-free DNA (cfDNA).
48. The method according to claim 1 or 25, wherein the sample is selected from the group consisting of peripheral blood mononuclear cells (PBMCs), B cells, circulating tumor cells, and tumor infiltrating lymphocytes.
49. The method according to claim 1 or 25, wherein the sample is formalin-fixed paraffin-embedded (FFPE) tissue, fresh tissue, frozen tissue, a blood sample, or a plasma sample.
50. A composition for analyzing the B cell receptor (BCR) repertoire in a sample, the composition comprising at least one set of: i) (a) a plurality of V gene primers for a majority of different V genes for at least one BCR-encoding sequence targeting at least a portion of framework region 1 (FR1) within the V gene, or (b) a plurality of V gene primers for a majority of different V genes for at least one BCR-encoding sequence targeting at least a portion of framework region 3 (FR3) within the V gene ; and ii) (a) one or more C gene primers, the one or more C gene primers targeting at least a portion of the C gene of the at least one BCR-encoding sequence, or (b) A plurality of J gene primers, said plurality of J gene primers being directed to at least a portion of most of the different J genes of said at least one BCR coding sequence; wherein each set of i) primers and ii) primers is directed to the coding sequence of the same target BCR gene selected from IgH, IgL, and IgK; wherein the plurality of V gene primers, the plurality of J gene primers, and / or one or more C gene primers comprise one or more cleavable groups located (i) near or at the end of the primer, or (ii) near or around the central nucleotide of the primer, wherein each set of i) primers and ii) primers directed to the same target BCR gene is configured to amplify the target BCR repertoire in a single multiplex amplification reaction, and wherein the amplification reaction product is treated with any suitable reagent that will selectively cleave or otherwise selectively disrupt the nucleotide bond of said one or more cleavable groups; and wherein the sets of i) primers and ii) primers are respectively: A) SEQ ID NOs: 69 - 136, and 443 - 446; B) SEQ ID NOs: 311 - 336, and 461, 462, 481, 515 - 519, 554, 585, 586; C) SEQ ID NOs: 321 - 328, 337 - 354, and 463, 464, 482, 520 - 524, 555, 587, 588; D) SEQ ID NOs: 321 - 328, 335, 336, 347 - 352, 355 - 364, and 465, 466, 483, 525 - 529, 556, 589, 590; E) SEQ ID NOs: 321 - 328, 347 - 354, 365 - 374, and 467, 468, 484, 530 - 534, 557, 591, 592; F) SEQ ID NOs: 321 - 328, 375 - 393, and 469, 485, 535, 558, 559, 593 - 595; G) SEQ ID NOs: 386 - 391, 394 - 410, and 470, 486, 536, 537, 560, 561, 596 - 598; H) SEQ ID NOs: 295 - 301, 411 - 430, and 471, 487, 538, 539, 562, 563, 599 - 601; I) SEQ ID NOs: 321 - 328, 375 - 393, and 469, 485, 535, 558, 559, 593, 595 or J) SEQ ID NOs: 295 - 301, 321 - 328, 375 - 393, 411 - 430, and 469, 485, 535, 558, 559, 593, 595.
51. The composition according to claim 50, wherein at least one of the plurality of V gene primers, the plurality of J gene primers, and / or the one or more C gene primers comprises one or more cleavable groups, and the one or more cleavable groups are located at (i) near or at the end of the primer or (ii) near or around the central nucleotide of the primer.
52. The composition according to claim 50 or claim 51, wherein at least one of the plurality of V gene primers, the one or more C gene primers, and / or the plurality of J gene primers comprises two or more modified nucleotides having a cleavable group selected from the following: methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil, uracil, 5-methylcytosine, thymine dimer, 7-methylguanosine, 8-oxo-deoxyguanosine, xanthosine, inosine, dihydrouridine, bromodeoxyuridine, uridine, or 5-methylcytidine.
53. The composition according to claim 50 or claim 51, wherein the primers of (i) and (ii) are configured to amplify the IgH repertoire.
54. The composition according to claim 50 or claim 51, wherein the at least one set of (i) and (ii) are respectively selected from the primers in Table 2 and Table 5.
55. The composition according to claim 50 or claim 51, wherein the at least one set of (i) and (ii) are respectively selected from the primers in Table 3 and Tables 6-10.
Citation Information
Patent Citations
Method for multiplexed nucleic acid patch polymerase chain reaction
US20100129874A1
Scaffolded nucleic acid polymer particles and methods of making and using
US20100304982A1
Process for amplifying, detecting, and / or-cloning nucleic acid sequences
US4683195A
Process for amplifying nucleic acid sequences
US4683202A
Fluorescent nucleotide analogs and uses therefor
US7405281B2