Methods, kits and compositions for multiplex preparation of nucleic acids for next generation sequencing

WO2026169400A1PCT designated stage Publication Date: 2026-08-13CELLECTA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-08-13

Smart Images

  • Figure US2026011133_13082026_PF_FP_ABST
    Figure US2026011133_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are methods of producing amplified next generation sequencing (NGS) nucleic acid libraries. Aspects of the methods include contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, where the multiplex collection of nucleic acid primers includes at least one replicate set of primers, where the replicate set of primers includes primers having a common domain and a different replicate barcode (RBC) domain. Aspects of the methods further include amplifying the template extension product composition to produce the amplified NGS nucleic acid library and sequencing the NGS library. Also provided are kits and compositions, e.g., for use in performing embodiments of the methods as described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Atty. Docket No.: CLCT-011 WO

[0002] METHODS, KITS AND COMPOSITIONS FOR MULTIPLEX PREPARATION OF NUCLEIC ACIDS FOR NEXT GENERATION SEQUENCING

[0003] CROSS-REFERENCE TO RELATED APPLICATION

[0004] Pursuant to 35 U.S.C. § 119 (e), this application claims priority to the filing date of United States Provisional Patent Application Serial No. 63 / 756,706 filed February 10, 2025, the disclosure of which application is incorporated herein by reference in its entirety.

[0005] INCORPORATION BY REFERENCE OF SEQUENCE LISTING XML FILE A Sequence Listing is provided herewith as a Sequence Listing XML, “CLCT-011 WO SEQ LISTING”, created on January 13, 2026, and having a size of 16,784 bytes. The contents of the Sequence Listing XML are incorporated herein by reference in their entirety.

[0006] INTRODUCTION

[0007] Next generation sequencing (NGS) technology provides high-throughput sequencing that enables the rapid and cost-effective analysis of genetic material. Unlike traditional sequencing methods, NGS can sequence millions of DNA or RNA fragments simultaneously, providing comprehensive insights into the genome and transcriptome. NGS allows for profiling of gene expression and identification of mutations and genetic variations and with high resolution and scale.

[0008] To carry out NGS, an amplified library suitable for NGS must first be prepared from a nucleic acid sample (e.g., a DNA sample or an RNA sample). Multiplex polymerase chain reaction (multiplex PCR) provides a high-throughput method for amplifying targets of interest. Multiplex PCR involves simultaneous amplification of multiple target sequences in a single PCR reaction using a variety of primers, making it ideal for detecting a range of genetic variants, pathogens, or biomarkers in a single sample. By amplifying several regions of interest at once, multiplex PCR reduces time, cost, and sample requirements compared to traditional PCR approaches.

[0009] A widely used method for mitigating errors that occur during PCR (e.g., multiplex PCR) and NGS library preparation involves amplification that uses primers or adapters that incorporate Unique Molecular Identifiers (UMIs). The UMI approach involves incorporatingAtty. Docket No.: CLCT-011 WO

[0010] specific uniquely tagged sequences into each nucleic acid molecule using primer extension or adapter ligation prior to the amplification step. In order to achieve specific labeling of each template molecule, the number of tags incorporated in primers or adapters must significantly exceed (e.g., at least 5-fold) the number of sample template molecules. In order to achieve this high complexity, the UMIs are designed as random sequences of all four A, G, C and T nucleotides that are typically at least 10+ nucleotides in length. Following amplification, the number of reads for each unique UMI is analyzed and those with low reads per UMI are typically considered artifacts of PGR and are excluded from further analysis.

[0011] SUMMARY

[0012] The inventors have realized that UMI-based strategies have notable disadvantages. UMI-based strategies employ non-specific sequences, and thus require complex bioinformatic analyses to prevent misclassifying mutated UMIs as genuine variations. Additionally, as UMIs are generally random sequences, UMIs may interact with other nucleic acids in undesirable ways (e.g. with other primers). For example, receptor-specific primers containing UMIs often exhibit reduced amplification efficiency and generate higher levels of background noise, including primer-dimer products, which can compromise assay performance.

[0013] Accordingly, methods are needed that allow for validation of NGS sequencing assays without the use of unique molecular identifiers. There is need for high-throughput sequencing analysis that does not involve complex bioinformatic analysis. Further, methods are needed to distinguish true biological variants from mutations introduced during the PGR and NGS processes in order to improve reproducibility, accuracy and efficiency of NGS sequencing assays.

[0014] Embodiments of the present invention satisfy the above, and other, needs in the art by providing an approach to distinguish errors from true sequencing reads in NGS assays by employing replicate barcode (RBC) domains. RBC domains are specific sequences present in replicate sets of primers, and thus are not each a random sequence. RBCs allow for a distinct number of replicate libraries to be produced in a multiplex reaction assay, where each library is labeled with a particular RBC. As such, this design allows for straightforward normalization and assessment of the reproducibility of each target sequence while facilitating the identification and exclusion of mutated sequence variants that may arise during PGR and NGS. Furthermore, employing RBCs in accordance with embodiments of the invention allows for rare sequenceAtty. Docket No.: CLCT-011 WO

[0015] variants to be validated as true reads, and differentiated from artifacts of the PCR or NGS preparation processes.

[0016] Provided are methods of producing amplified next generation sequencing (NGS) nucleic acid libraries. Aspects of the methods include contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, where the multiplex collection of nucleic acid primers includes at least one replicate set of primers, where the replicate set of primers includes primers having a common domain and a different replicate barcode (RBC) domain. Aspects of the methods further include amplifying the template extension product composition to produce the amplified NGS nucleic acid library and sequencing the NGS library. Also provided are kits and compositions, e.g., for use in performing embodiments of the methods as described herein.

[0017] BRIEF DESCRIPTION OF THE FIGURES FIG. 1 provides a schematic illustrating position and design of primers used to amplify a CDR3 region and / or a CDR1-CDR2-CDR3 region from RNA in accordance with an embodiment of the invention.

[0018] FIG. 2 provides a schematic illustrating position and design of primers used to amplify a CDR3 region from DNA in accordance with an embodiment of the invention.

[0019] FIG. 3 provides a schematic of a method of preparing amplified indexed nucleic acid products from RNA in accordance with an embodiment of the invention.

[0020] FIG. 4 provides a schematic of a method of preparing amplified indexed nucleic acid products from DNA in accordance with an embodiment of the invention.

[0021] FIG. 5A provides a schematic of a workflow used to generate and analyze an NGS library in accordance with an embodiment of the invention.

[0022] FIG. 5B provides a schematic showing a replicate barcode (RBC) primer strategy according to an embodiment of the invention. A set of eight reverse gene-specific primers (Rev-GSPs) are each tagged with a unique six-nucleotide barcode to amplify each individual clonotype of TCR or BCR gene.

[0023] FIG. 5C shows an example of a final indexed NGS amplicon structure with RBCs.

[0024] FIG. 6 depicts a graph showing a read alignment of RBC-based assay data.

[0025] FIG. 7 shows plots of the read distribution of 8 RBCs across different sample types. FIG. 8A - FIG. 8B show repertoire profiling of a human peripheral blood mononuclear cell (PBMC) sample using an RBC or UMI strategy for adaptive immune receptor (AIR) DNAAtty. Docket No.: CLCT-011 WO

[0026] (FIG. 8A) and AIR RNA (FIG. 8B). The reproducibility of the NGS read count number among 8 RBCs is shown. The read counts (column normalized z-scores) of the most abundant T cell TRB receptor clonotypes are shown for the human CDR3 amplification. BC1 - BC8 are the 8 different barcodes, RBC column is the total of the reads for all 8 RBC barcodes and UMI is the number of NGS read counts for a UMI-based assay. The coefficient of variation (CoV) of the abundance measurement of each clonotype is shown in the last column.

[0027] FIG. 9 depicts a graph showing the dispersion of read counts among the internal replicates. The coefficient of variation (CoV) is used to measure the dispersion among the internal replicates. An exponential decay function is used to model the dispersion in the data. FIG. 9 corresponds to the samples shown in FIGS. 8A and 8B.

[0028] FIG. 10 depicts the sensitivity and specificity of AIR assay with RevGSP-RBC and RevGSP-UMI primers. Gel electrophoresis comparative analysis of CDR3 PCR products (after second PCR with index primers) amplified from 50 ng PBMC RNA (1 ,2,5,6) or 2 ug PBMC DNA (3, 4, 7, 8) or negative controls (water, C-, 9,10,11 ,12) using AIR-TCR or AIR-BCR assays comprising Rev GSP with RBC (1 ,3, 5, 7, 9,11) or Rev GSP with UMI (2,4,6,8,10,12) is shown. In order to detect primer-dimers all negative control samples were amplified for an additional 5 cycles (in a second PCR).

[0029] FIG. 11 depicts a graph showing the increased sensitivity of the AIR RBC assay in detecting clonotypes compared to the AIR UMI assay.

[0030] FIG. 12A - FIG. 12B depicts graphs showing the analysis of NGS read numbers for clones that occur in 1 to 8 RBC barcodes. This analysis shows the reproducibility in clonotype detection and identifies the number of reads for clonotypes amplified from single template molecules. Analysis of the distribution of NGS reads for clonotypes of different abundances is used to generate a cut-off line (depicted as a dashed line) for excluding clonotypes associated with mutations generated in the amplification and NGS steps.

[0031] FIG. 13 provides a schematic showing the strategies for TCR / BCR CDR3 region repertoire analysis.

[0032] FIG. 14 provides a schematic of a method of preparing amplified indexed nucleic acid products from single cells in accordance with an embodiment of the invention.

[0033] FIG. 15 shows a gel electrophoresis of amplicons after an index primer PCR reaction. Lanes 1-3: Libraries prepared from total whole blood RNA of different samples; Lane 4: Positive control (C+) PBMC RNA; Lane 5: Negative Control (C-) water.Atty. Docket No.: CLCT-011 WO

[0034] FIG. 16 depicts a graph of the size distribution of the amplified indexed library after primer removal step analyzed by fragment analyzer.

[0035] FIG. 17 shows a gel electrophoresis of amplicons after an index primer PCR reaction for scAIR-TCR- Mark37 (“T”) and scAIR-BCR-Mark30 (“B”) libraries generated from positive control RNA.

[0036] FIG. 18 depicts a graph of the size distribution the amplified indexed libraries of scAIR-TCR-Mark37 (top) and scAIR-BCR-Mark30 (bottom) from positive control RNAs.

[0037] FIG. 19 shows a schematic of a reagent cartridge for sequencing pooled amplified indexed libraries on a NextSeq 1000 / 2000.

[0038] FIG. 20 shows a gel electrophoresis of amplicons after a second primer index PCR reaction. TCR / BCR: Lanes 1 (1.5 ug), Lane 2 (3 ug), Lane 3 (6 ug) of gDNA. Lane 4: positive control DNA from AIR DNA kit (5 ug).

[0039] FIG. 21 depicts a graph of the size distribution of the amplified indexed library after primer removal step analyzed by fragment analyzer for TCR (left) and BCR (right).

[0040] FIG. 22 shows a gel electrophoresis of amplicons after index primer PCR reaction. The full-length receptor region is ~600bp. Lanes 1-4: Libraries prepared from 100 ng of total RNA.

[0041] FIG. 23 depicts a graph of the size distribution of the amplified indexed library after primer removal step.

[0042] FIG. 24A - FIG. 24C: Immunoprofiling with replicate barcodes (RBCs). FIG. 24A:

[0043] Molecular workflow starts with the DNA and RNA templates. The RBCs are introduced either during gene-specific primer barcode extension in the DNA templates or during cDNA synthesis with gene-specific primers in the RNA templates. Both processes are followed by gene-specific forward primer extension and two PCRs. The resulting product for NGS contains the Illumina sequencing adapters (P5 and P7), Illumina dual index primers (UDP1 and UDP2), universal anchor primers (A1 and A2), the replicate barcode (RBC), and the TCR or BCR amplicon. FIG.

[0044] 24B: Bioinformatics workflow starts with the initial read alignment, followed by clonotype assembly using tools such as MiXCR. Clonotypes are filtered and passing clonotypes read counts are converted to estimated template count. FIG. 24C: The distribution of reads per clonotype shows a bimodal distribution where artifactual and real clonotypes are distinguished. Rare clonotypes that appear with only one replicate barcode can be used to identify a normalization value. This value is then used as the factor to convert reads to estimated template counts.Atty. Docket No.: CLCT-011 WO

[0045] FIG. 25A - FIG. 25D: Repertoire from RBC and UMI based assays. FIG. 25A: PCR amplification of RBC- and UMI-based samples. FIG. 25B: Number of clonotypes identified. FIG.

[0046] 25C: Reproducibility between RBC and UMI assays in IGH and TRB chains. RBC is on the x-axis and UMI is on the y-axis. FIG. 25D: Heatmap of read counts across RBCs from RNA templates. Left panel shows reads from B-cells and right panel shows reads from T-cells.

[0047] FIG. 26A - FIG. 26D: Repertoire metrics. FIG. 26A: Z-score normalized read counts of V genes in the repertoire for TRB (left panel) and IGH (right panel). FIG. 26B: Clonotype occupancy of the repertoire divided by ranked read counts (e.g. top 1 - 10, top 11 - 100). TRB is shown in the top panel and IGH is shown in the bottom panel. FIG. 26C: Diversity metric measured with the D50 index. IGH is shown in the left panel and TRB is shown in the right panel. FIG. 26D: Diversity metric measured with the True diversity index. TRB is shown in the left panel and IGH is shown in the right panel.

[0048] FIG. 27A - FIG. 27B: Template estimation using RBCs. FIG. 27A: Histograms of the number of read counts for each clonotype only found in one RBC. The threshold value identifies the cutoff to distinguish real vs artifactual clonotypes. The average value of the left peak is used as the normalization value to convert read counts to UMIs. FIG. 27B: Scatterplot of the UMI count (in the UMI experiments) vs the estimated number of templates (in the RBC experiments).

[0049] FIG. 28 shows plots of the read distribution of 8 RBCs. Each plot shows a different sample type. Barcodes 1-8 are plotted in order from L to R for each plot.

[0050] FIG. 29 depicts the reproducibility between RBC and UMI assays in IGK, IGL and TRA chains. RBC is shown on the x-axis and UMI is shown on the y-axis.

[0051] FIG. 30A - FIG. 30D: Template reproducibility between RBC barcodes in IGH and TRB chains. FIGS. 30A-30B use DNA templates. FIGS. 30C-30D use RNA templates. FIGS. 30A, 30C are T-cell samples. FIGS. 30B, 30D are B-cell samples.

[0052] DEFINITIONS

[0053] As used herein, the term “hybridization conditions” means conditions in which a primer, or other polynucleotide, specifically hybridizes to a region of a target nucleic acid with which the primer or other polynucleotide shares some complementarity. Whether a primer specifically hybridizes to a target nucleic acid is determined by such factors as the degree of complementarity between the primer and a region or domain of the target nucleic acid and the temperature at which the hybridization occurs, which may be informed by the melting temperature (TM) of the primer. The melting temperature refers to the temperature at which halfAtty. Docket No.: CLCT-011 WO

[0054] of the primer-target nucleic acid duplexes remain hybridized and half of the duplexes dissociate into single strands. The Tm of a duplex may be experimentally determined or predicted using the following formula Tm = 81.5 + 16.6(log10[Na+]) + 0.41 (fraction G+C) - (60 / N), where N is the chain length and [Na+] is less than 1 M. See Sambrook and Russell (2001 ; Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor N.Y., Ch.

[0055] 10). Other more advanced models that depend on various parameters may also be used to predict Tm of primer / target duplexes depending on various hybridization conditions. Approaches for achieving specific nucleic acid hybridization may be found in, e.g., Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology-Hybridization with Nucleic Acid Probes, part I, chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assays,” Elsevier (1993).

[0056] The terms “complementary” and “complementarity” as used herein refer to a nucleotide sequence that base-pairs by non-covalent bonds to all or a region of a target nucleic acid (e.g., a region of the nucleic acid product). In the canonical Watson-Crick base pairing, adenine (A) forms a base pair with thymine (T), as does guanine (G) with cytosine (C) in DNA. In RNA, thymine is replaced by uracil (U). As such, A is complementary to T and G is complementary to C. In RNA, A is complementary to U and vice versa. Typically, “complementary” refers to a nucleotide sequence that is at least partially complementary. The term “complementary” may also encompass duplexes that are fully complementary such that every nucleotide in one strand is complementary to every nucleotide in the other strand in corresponding positions. In certain cases, a nucleotide sequence may be partially complementary to a target, in which not all nucleotides are complementary to every nucleotide in the target nucleic acid in all the corresponding positions. For example, a primer may be perfectly (i.e., 100%) complementary to a region or domain of the target nucleic acid, or the primer and the target nucleic acid may share some degree of complementarity which is less than perfect (e.g., 70%, 75%, 85%, 90%, 95%, 99%).

[0057] The percentage identity of two nucleotide sequences can be determined by aligning the sequences for optimal comparison purposes (e.g., gaps can be introduced in the sequence of a first sequence for optimal alignment). The nucleotides at corresponding positions are then compared, and the percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity= # of identical positions / total # of positionsx100). When a position in one sequence is occupied by the same nucleotide as the corresponding position in the other sequence, then the molecules are identical at that position.Atty. Docket No.: CLCT-011 WO

[0058] A non-limiting example of such a mathematical algorithm is described in Karlin et al., Proc. Natl. Acad. Sci. USA 90:5873-5877 (1993). Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) as described in Altschul et al., Nucleic Acids Res. 25:389-3402 (1997). When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., NBLAST) can be used. In one aspect, parameters for sequence comparison can be set at score=100, wordlength=12, or can be varied (e.g., wordlength=5 or wordlength=20).

[0059] As used herein, an “oligonucleotide” is a single-stranded multimer of nucleotides from 2 to 500 nucleotides, e.g., 2 to 200 nucleotides. Oligonucleotides may be synthetic or may be made enzymatically, and, in some embodiments, are 25 to 70 nucleotides in length.

[0060] Oligonucleotides may contain ribonucleotide monomers (i.e., may be oligoribonucleotides or “RNA oligonucleotides”) or deoxyribonucleotide monomers (i.e., may be oligodeoxyribonucleotides or “DNA oligonucleotides”). In some cases, oligonucleotides may contain a mixture of ribonucleotides and deoxyribonucleotides. In some cases, Oligonucleotides may contain modified, i.e., non-natural nucleotides or modifications, including for example, LNA, FANA, 2’-O-Me RNA, 2’-fluoro RNA, or the like, linkage modifications (e.g., phosphorothioates, 3’-3’ and 5’-5’ reversed linkages), 5’ and / or 3’ end modifications (e.g., 5’ and / or 3’ amino, biotin, DIG, phosphate, thiol, dyes, quenchers, etc.), one or more fluorescently labeled nucleotides, or any other feature that provides a desired functionality to the oligonucleotides. Oligonucleotides may be 6 to10, 10 to 20, 21 to 30, 31 to 40, 41 to 50, 51 to 60, 61 to 70, 71 to 80, 80 to 100, 100 to 150 or 150 to 200, up to 500 or more nucleotides in length, for example.

[0061] A “domain” when used in reference to nucleic acids refers to a stretch or length of a nucleic acid made up of a plurality of nucleotides, where the stretch or length provides a defined function to the nucleic acid. Examples of domains include capture domains, barcode domains, primer binding domains, hybridization domains, replicate barcode (RBC) domains, unique molecular identifier (UMI) domains, adapter domains (e.g., Next Generation Sequencing (NGS) adapter domains), template switching domains, NGS indexing domains, etc. In some instances, the terms “domain” and “region” may be used interchangeably. While the length of a given domain may vary, in some instances the length ranges from 2 to 100 nt, such as 5 to 50 nt, e.g., 5 to 30 nt. Amplification primer binding domains (e.g., PGR binding domains) are domains that are configured to bind via hybridization to an amplification primer (e.g., PGR primer). Adapter domains may be used as an amplification primer binding domain.Atty. Docket No.: CLCT-011 WO

[0062] As used herein, the expression “derived from” describes a composition that results from a process whereby a first component (e.g., a first nucleic acid molecule), or information from that first component, is used to isolate, derive or construct a different second component (e.g., a second nucleic acid molecule that is different in structure, sequence or character from the first nucleic acid molecule from which it was derived). For example, a cDNA molecule is derived from a corresponding DNA template (e.g., genomic DNA) or RNA template (e.g., mRNA).

[0063] Similarly, a cDNA library is derived from RNA or DNA that is collected from a cell or population of cells. Also for example, a cDNA library can be derived from mRNA that is collected from a cell or population of cells.

[0064] By "template extension product composition" is meant a nucleic acid composition that includes nucleic acids that are template extension products. Template extension products are deoxyribonucleic acids that include a primer domain at the 5' end covalently bonded to a synthesized domain at the 3' end, which synthesized domain is a domain of base residues added by a polymerase mediated reaction to the 3' end of the primer domain in a sequence that is dictated by a template nucleic acid to which the primer domain is hybridized during production of the template extension product. Template extension product compositions may include double stranded nucleic acids that include a template nucleic acid strand hybridized to a template extension product strand, e.g., as described above. Template extension product compositions may further include a template switch oligonucleotide (TSO) hybridized to the template extension product strand. The length of the template extension products and / or double stranded nucleic acids that incorporate the same in the template extension product compositions may vary, wherein in some instances the nucleic acids have a length ranging from 50 to 1000 nt, such as 60 to 800 nt and including 70 to 700 nt. The number of distinct nucleic acids that differ from each other by sequence in the template extension product compositions produced via methods of the invention may also vary, ranging in some instances from 10 to 100,000,000, such as 100 to 1 ,000,000 and including 1 ,000 to 500,000, and 5,000 to 500,000.

[0065] DETAILED DESCRIPTION

[0066] Provided are methods of producing amplified next generation sequencing (NGS) nucleic acid libraries. Aspects of the methods include contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, where the multiplex collection of nucleic acid primers includes at least one replicate set of primers, where the replicate set of primers includes primers having a commonAtty. Docket No.: CLCT-011 WO

[0067] domain and a different replicate barcode (RBC) domain. Aspects of the methods further include amplifying the template extension product composition to produce the amplified NGS nucleic acid library and sequencing the NGS library. Also provided are kits and compositions, e.g., for use in performing embodiments of the methods as described herein.

[0068] Before the present invention is described in greater detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0069] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0070] Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.

[0071] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.

[0072] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. TheAtty. Docket No.: CLCT-011 WO

[0073] citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.

[0074] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements or use of a “negative” limitation.

[0075] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.

[0076] While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112.

[0077] METHODS

[0078] As summarized above, provided are methods of producing amplified next generation sequencing (NGS) nucleic acid libraries. Aspects of the methods include contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, where the multiplex collection of nucleic acid primers includes at least one replicate set of primers, where the replicate set of primers includes primers having a common domain and a different replicate barcode (RBC) domain. Aspects of the methods further include amplifying the template extension product composition to produce the amplified NGS nucleic acid library and sequencing the NGS library. For ease of description, embodiments of primers as further described below include forward primers, reverse primers and templateAtty. Docket No.: CLCT-011 WO

[0079] switch oligonucleotides (TSOs). Such primers, which may also be referred to as “primer extension nucleic acids” are nucleic acids that are used in primer extension reactions. In some embodiments, the primer extension nucleic acids themselves prime the primer extension reactions. In some embodiments, the primer extension nucleic acids do not prime a primer extension reaction, but nonetheless participate such in a reaction such that they are incorporated into a primer extension product. For example, in some instances, the primer extension nucleic acids are in the form of template switch oligonucleotides (TSOs) (e.g., TSOs that include an RBC domain). In other words, the primer extension nucleic acids are used as TSOs in a primer extension reaction (e.g., a template-switching primer extension reaction). In other instances, primers with RBC domains may be incorporated in an adapter which may then be ligated to single-stranded or double-stranded extension products or fragments of extension products by a wide range of conventional protocols using ligases or transposases. In another embodiment, adapters comprising primers with RBC domains are ligated directly to DNA or cDNA or fragments of nucleic acid template samples without prior extension step.

[0080] Multiplex Collection of Nucleic Acid Primers

[0081] As summarized above, aspects of the methods of the disclosure include contacting a nucleic acid sample with a multiplex collection of nucleic acid primers. By multiplex collection of nucleic acid primers, it is meant a collection of nucleic acid primers that are contacted with the nucleic acid sample simultaneously.

[0082] The nucleic acid primers may include DNA, RNA or a combination thereof. In some embodiments, the nucleic acid primers are DNA primers. In some embodiments, the nucleic acid primers are RNA primers. The nucleic acid primers may include one or more nucleotides (or analogs thereof) that are modified or otherwise non-naturally occurring. For example, the nucleic acid primer may include one or more nucleotide analogs (e.g., LNA, FANA, 2’-O-Me RNA, 2’-fluoro RNA, or the like), linkage modifications (e.g., phosphorothioates, 3’-3’ and 5’-5’ reversed linkages), 5’ and / or 3’ end modifications (e.g., 5’ and / or 3’ amino, biotin, DIG, phosphate, thiol, dyes, quenchers, etc.), one or more fluorescently labeled nucleotides, or any other feature that provides a desired functionality to the nucleic acid primer.

[0083] Nucleic acid primers may be of any convenient length. In some instances, the length ranges the length of the nucleic acid primers ranges from 10 to 80 nt, such as 15 to 70 nt, e.g., 20 to 60 nt, including 25 to 50 nt.Atty. Docket No.: CLCT-011 WO

[0084] In some embodiments, the number of primers in each replicate set of primers is 100-fold or less than the number of template molecules in the nucleic acid template sample. In some embodiments, the number of primers in each replicate set of primers is 1000-fold or less than the number of template molecules in the nucleic acid template sample.

[0085] In some embodiments, the primers of the multiplex collection of nucleic acid primers do not include unique molecular identifier (UMI) domains. As reviewed above, UMI domains are stretches of random or semi-random nucleotides of varying length, e.g., ranging in length in some instances from 6 to 16 nucleotides. In other words, each UMI has a unique sequence. UMIs are described at, e.g., Smith et al., Genome Res. 2017. 27:491-499; Andersson, et al., Molecular Aspects of Medicine. 2024. 96:101253; Single Cell Sequencing and Systems Immunology. (2015). Netherlands: Springer Netherlands.

[0086] In some embodiments, the multiplex collection of primers (e.g., gene specific primers) includes forward primers, reverse primers or both. In some embodiments, multiplex collection of primers (e.g., gene specific primers) includes a forward primer set (i.e., a set of forward primers), a reverse primer set (i.e., a set of reverse primers) or both. By reverse primer, it is meant a primer that binds to the sense strand of a nucleic acid molecule (e.g., a nucleic acid template). In some embodiments, the reverse primer binds to RNA (e.g., mRNA, such as protein-coding sequences of mRNA). In some embodiments, the reverse primer binds to DNA (e.g., genomic DNA). By forward primer, it is meant a primer that binds to the antisense strand (e.g., the nucleic acid strand which is complementary to the sense strand). In some instances, the antisense strand is an extension product generated by extension of a reverse primer where the sense strand (e.g., mRNA) is a template. In some instances, the antisense strand is DNA (e.g., gDNA). In some embodiments, the forward primer binds to DNA, including, but not limited to cDNA (e.g., an extension product) and genomic DNA.

[0087] Any convenient number of forward primers and reverse primers may be used. In some instances, the number of forward primers is 1 or more, including, e.g., 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 250 or more, 500 or more and 1000 or more. In some cases, the forward primer may be a single universal oligonucleotide. In some instances, the number of reverse primers is 1 or more, including, e.g., 2 or more, 5 or more, 10 or more, 50 or more, 100 or more, and 1000 or more. In some instances, the reverse primer may be a single universal oligo dT primer. In some embodiments, the multiplex collection of primers (e.g., gene specific primers) includes a different number of forward and reverse primers. In some cases, there are aAtty. Docket No.: CLCT-011 WO

[0088] greater number of forward primers than reverse primers. In some cases, there are a greater number of reverse primers than forward primers. In some cases, there are an equal number of forward primers and reverse primers. In some cases (e.g., in cases of immune receptor repertoire profiling), the number of forward primers (e.g., primers which are complementary to highly variable V regions) ranges from 100 to 400 primers and the number of reverse primers (e.g., primers which are complementary to conservative C or J regions) ranges from 1 to 24 primers. In some cases (e.g., in cases of expression profiling of human genes), the number of forward and reverse primers are the same. For example, in some instances, the number of forward and reverse primers is equal to the number of target genes.

[0089] In some cases, there are one or more forward primers than reverse primers (e.g., 1 more forward primer, 5 more forward primers, 10 more forward primers, 50 more forward primers, 100 more forward primers, or 1000 more forward primers than reverse primers) including, e.g., 5 or more forward primers, 10 or more forward primers, 50 or more forward primers, 100 or more forward primers, and 1000 or more forward primers than reverse primers.

[0090] In some cases there are 1 .1x or more forward primers than reverse primers (e.g., 1.5x more forward primers, 2x more forward primers, 5x more forward primers, 10x more forward primers, or 100x more forward primers than reverse primers), including, e.g., 1 ,5x or more forward primers, 2x or more forward primers, 5x or more forward primers, 10x or more forward primers, and 100x or more forward primers than reverse primers.

[0091] In some cases, there are one or more reverse primers than forward primers (e.g., 1 more reverse primer, 5 more reverse primers, 10 more reverse primers, 50 more reverse primers, 100 more reverse primers, or 1000 more reverse primers than forward primers) including, e.g., 5 or more reverse primers, 10 or more reverse primers, 50 or more reverse primers, 100 or more reverse primers, and 1000 or more reverse primers than forward primers.

[0092] In some cases there are 1 .1x or more reverse primers than forward primers (e.g., 1.5x more reverse primers, 2x more reverse primers, 5x more reverse primers, 10x more reverse primers, or 100x more reverse primers than forward primers), including, e.g., 1.5x or more reverse primers, 2x or more reverse primers, 5x or more reverse primers, 10x or more reverse primers, and 100x or more reverse primers than forward primers.

[0093] The primers of the multiplex collection of primers including gene-specific primers may be designed using any convenient algorithm and / or software tool, e.g., such as the Primer3 algorithm, Primer Design Tool from NCI, etc. The melting temperature between the selected gene-specific primers with gene-specific binding domain and target templates may vary, rangingAtty. Docket No.: CLCT-011 WO

[0094] in some instances from 60QC to 859C, such as 659C to 809C. Furthermore, the primers may be selected that lack significant secondary structures, or self-complementarity (e.g., primers may be selected with less than 4-bp complementary regions) and cross-complementarity to each other of less than 10 nt complementarity region. For example, primers may be selected such that they do not form primer-dimers. In order to avoid primer-dimer formation in a multiplex RT-PCR assay, the selected primers in some embodiments are designed with the nucleotide A at the 3’-end and biased GCA-rich composition with reduced percentage of T nucleotides, where in some instances the percentage of T is 20% or less, such as 15% or less, including 10% or less, down to 0%. In some instances, the primers have a GC-content of between 45% to 85%, such as 50% to 75%.

[0095] In some embodiments, the multiplex collection of nucleic acid primers includes a multiplex collection of gene specific primers. Gene specific primers include gene specific domains. In other words, a gene specific primer includes a domain used to target a specific gene. Gene specific primers may be used to amplify one or more specific genes. In some instances, gene specific primers are used to amplify one or more genes (e.g., 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 genes) including, e.g., 2 or more genes, 3 or more genes, 4 or more genes, 5 or more genes, 6 or more genes, 7 or more genes, 8 or more genes, 9 or more genes, 10 or more genes, 20 or more genes, 30 or more genes, 40 or more genes, 50 or more genes, 60 or more genes, 70 or more genes, 80 or more genes, 90 or more genes, 100 or more genes, 200 or more genes, 500 or more genes or 100 or more genes. Gene specific primers are known in the art and can be found at, e.g., U.S. Patent No. 10,975,440, U.S. Patent No. 11 ,655,510 and U.S. Patent No. 11 ,274,334, the disclosures of which are herein incorporated by reference.

[0096] The multiplex collection of gene specific primers may be configured to generate template extension products from any convenient target nucleic acid. The nucleic acid primers may bind to (e.g., hybridize with) nucleic acid targets that includes DNA, RNA or a combination thereof. In some embodiments, binding is defined by complementarity between nucleic acid primers and nucleic acid targets. In some embodiments, the nucleic acid primers bind to a DNA target. In some embodiments, the nucleic acid primers bind to an RNA target.

[0097] Target nucleic acids include those that correspond to a wide range of mammalian genes, and pathogenic genes from a wide range of pathogenic organisms, such as viruses, bacteria, fungi, etc. which could be present in the human or mammalian bodies. Of interest in certain applications are human, mammalian species commonly used as a model organisms to studyAtty. Docket No.: CLCT-011 WO

[0098] human diseases, such as mouse, rat, or monkey, and pathogenic organisms associated with human diseases. In some embodiments, the targeted genes may be protein coding, or may express non-coding RNAs, micro RNAs, mitochondrial RNAs, regulatory RNAs, etc. In some embodiments, the genes are selected from the genes that could be transcribed or expressed in the organism and present in the biological samples in the form of RNA.

[0099] In some embodiments, the multiplex collection of gene specific primers may be configured to generate template extension products of a subset of cell-specific, tissue-specific or state-specific genes. These genes encode marker products (e.g. proteins, peptides or RNAs) that are specifically expressed in different cell types (e.g. markers for T, B, NK, stromal, cancer, epithelial, neuronal, etc. cells), different tissues or different cell states, e.g. marker products induced by treatment (e.g. drugs) or changes in conditions (e.g. heat shock), disease states (e.g. cancer, infection), or natural biological processes (e.g. differentiation, apoptosis, aging, etc.).

[0100] In some embodiments, the multiplex collection is configured to generate template extension products from adaptive immune receptor (AIR) nucleic acids. An AIR nucleic acid is a nucleic acid (e.g., gene) that encodes a protein on an immune cell that recognizes antigens. In some embodiments, the adaptive immune receptor nucleic acids include RNA and / or DNA of T cell receptor (TCR) or B cell receptor (BCR) genes. In some embodiments, the adaptive immune receptor nucleic acids include RNA and / or DNA of T cell receptor (TCR) and B cell receptor (BCR) genes. TCR genes include, e.g., TRA, TRB, TRD and TRG gene families. BCR genes include, e.g., IGH, IGK and IGL gene families. Each TCR or BCR gene family includes individual genes having the structure V-(D)-J-C, where each V region, J region, D region, and C region may differ. In some instances, the gene has the structure V-J-C. In some instances, the gene has the structure V-D-J-C. C regions of immune receptor genes have the most conservative sequences specific to each type of gene and may be used for gene-specific primer design. Each V region, D region, and J region may include variable subregions and conservative subregions, where the conservative subregions may be used for primer design. V regions have the structure UTR-FR1-CDR1-FR2-CDR2-FR3-CDR3, where UTR, FR1, FR2 and FR3 are conservative regions which may be used for primer design. CDR1 , CDR2 and CDR3 are variable subregions involved in the recognition of specific antigens by immune receptors and antibodies. The CDR3 region is the most variable subregion. The multiplex collection of primers may be configured to generate template extension products from the CDR1 , GDR2, CDR3, CDR1-CDR2, CDR2-CDR3 or CDR1 -CDR2-CDR3 region of the AIR nucleic acids, orAtty. Docket No.: CLCT-011 WO

[0101] any combination thereof. In some embodiments, the multiplex collection of primers is configured to generate template extension products from CDR3 or CDR1-CDR2-CDR3 portion of TCR genes TRA, TRB, TRD, TRG and / or BCR genes IGH, IGK, and IGL. In some embodiments, the multiplex collection of primers is configured to generate template extension products from CDR3 and CDR1 -CDR2-CDR3 portion of TCR genes TRA, TRB, TRD, TRG and / or BCR genes IGH, IGK, and IGL.

[0102] Forward and reverse primers (e.g., forward and reverse gene-specific primers) may be designed to bind to regions flanking the CDR3 and / or CDR1-CDR2-CDR3 regions, such that the CDR3 and / or CDR1-CDR2-CDR3 regions may be amplified. See, e.g., FIG. 13. Reverse primers (e.g., reverse primers for immune repertoire analysis) may be designed to bind to different regions depending on if the template nucleic acid is a template DNA or a template RNA. In cases where primers bind to an RNA template, the primers (e.g., reverse gene-specific primers) may be designed to bind to the conservative C region. In cases where primers bind to a DNA template, the primers (e.g., reverse gene-specific primers) may be designed to bind to the J region. In DNA templates, the C region is separated from V-D-J region by an intron. In some cases (e.g., in cases where the primer binds to an RNA template or a DNA template), primers (e.g., forward gene-specific primers) may be designed to bind to the FR3 region or the FR1 region. By binding to the FR3 region, the primers may be used to amplify the CDR3 region. By binding to the UTR and / or FR1 region, the primers may be used to amplify the CDR1-CDR2-CDR3 regions.

[0103] In some embodiments, the multiplex collection of nucleic acid primers includes template switch oligonucleotides (e.g., TSOs) or replicate set of TSO with RBC domain. In other words, the primers may be template switch oligonucleotides (TSOs). In some embodiments, the multiplex collection of nucleic acid primers includes TSOs and does not include forward primers. In some embodiments, the multiplex collection of nucleic acid primers includes TSOs and reverse primers. In some cases, TSOs may be used to amplify the UTR-ODR1-CDR2-CDR3 regions in a single multiplex reaction in combination with a reverse primer (e.g., a C-region specific reverse gene specific primer). In some embodiments, the multiplex collection of nucleic acid primers includes replicate set of TSO with RBC domain and universal oligo dT reverse primers.

[0104] Primers present in the multiplex collection of nucleic acid primers (e.g., gene specific primers) of the invention may be experimentally validated using any convenient protocol. In some instances, the experimentally validated gene specific primers are validated in a multiplexAtty. Docket No.: CLCT-011 WO

[0105] amplification assay with a synthetic control template mix which mimics the natural target template sequences and includes binding sites for the whole set of gene-specific primer pairs and / or a universal natural template mix derived from multiple different mammalian tissues, blood or cell types. Specifically, as a template for multiplex RT-PCR or PCR assay, a set (usually between 3 to 6) of natural total universal RNAs or DNAs, e.g., including a mix of several RNAs or DNAs isolated from human or mouse cell lines or tissue samples (e.g., available from Takara-Clontech, Agilent, Qiagen, Origene, etc.) may be employed as a natural nucleic acid control. In addition (or alternatively) to the set of the natural control template nucleic acids, a mix of the synthetic control template nucleic acids, e.g., one that has been synthesized on the surface of custom microarrays (e.g., Custom Array, TWIST or Agilent), conventional oligonucleotide synthesis or assembled by ligation of oligonucleotides (gene fragments) and designed for each target amplicon, may be employed. Synthetic control templates could be ssDNA, dsDNA or RNA, wherein RNA is usually synthesized from dsRNA using RNA polymerase. In such synthetic control templates, the templates include the sequence of the both PCR primer-binding site domains and the full-length or truncated in the middle cDNA region between PCR primers that corresponds to the primer extension domain. In some functional validation assays, two synthetic template concentrations (e.g., 10-fold difference) may be employed to measure abundance level (number of specific reads) in a manner that is not dependent on the amount of starting universal RNA or DNA template. The length of synthetic control templates may vary, ranging in some instances from 100 to 1 ,000, such as 110 to 800, including 120 to 700 nt. The amplification products generated in the multiplex PCR assays may be quantitatively analyzed by sequence analysis using conventional NGS instruments (e.g., available from Illumina, ThermoFisher, Nanopore and other commercial vendors). The NGS data generated for different templates and experimental conditions may be scaled to the same number of total reads (usually total 10,000,000 reads), aligned with the sequences of PCR primer domain and downstream extended domain sequences for each target amplicon. The number of specific reads corresponding to each target amplicon may be measured as the number of correctly aligned sequences for each primer and downstream extended domain sequences.

[0106] Replicate Sets of Primers

[0107] As reviewed above, the multiplex collection of nucleic acid primers includes at least one replicate set of primers. By "primers" it is meant nucleic acids that are used in primer extension reactions. In some embodiments, the primers themselves prime the primer extension reactions.Atty. Docket No.: CLCT-011 WO

[0108] In some embodiments, the primers do not prime a primer extension reaction, but nonetheless participate such in a reaction such that they are incorporated into a primer extension product. For example, in some instances, the primers are in the form of template switch oligonucleotides (TSOs). In other words, the primers are used as TSOs in a primer extension reaction (e.g., a template-switching primer extension reaction). A replicate set of primers includes primers (e.g., at least 3 primers) having a common domain and a different replicate barcode (RBC) domain. There may be multiple copies or sets of the replicate of primers in the multiplex collection of nucleic acid primers. The common domain is a domain (e.g., a target-binding domain such as a gene-specific binding domain) that is present in each primer in the replicate set of primers. In some cases, all primers in a replicate set of primers are designed with gene-specific binding domain complementary to the specific target gene sequence. However, all members of the replicate set of primers do not include the same RBC domain. In other words, each primer in the redundant subset of primers has a different RBC domain. The number of primers in a replicate set of primers may be any integer over 1. However, the number of primers in a replicate set of primers is less than the number of primers needed to label each individual nucleic acid molecule in the nucleic acid template sample with a unique barcode. In some embodiments, the number of primers in each replicate set of primers is 2 or more primers (e.g., 2 primers, 3 primers, 4 primers, 5 primers, 6 primers, 7 primers, 8 primers, 9 primers, 10 primers, 11 primers, 12 primers, 13 primers, 14 primers, 15 primers, 16 primers, 17 primers, 18 primers, 19 primers, 20 primers, 21 primers, 22 primers, 23 primers, 24 primers, 25 primers, 30 primers, 35 primers, 40 primers, 45 primers, 50 primers, 55 primers, 60 primers, 65 primers, 70 primers, 75 primers, 80 primers, 85 primers, 90 primers, 95 primers, or 100 primers) including, e.g., 3 or more primers, 4 or more primers, 5 or more primers, 6 or more primers, 7 or more primers, 8 or more primers, 9 or more primers, 10 or more primers, 11 or more primers, 12 or more primers, 13 or more primers, 14 or more primers, 15 or more primers, 16 or more primers, 17 or more primers, 18 or more primers, 19 or more primers, 20 or more primers, 21 or more primers, 22 or more primers, 23 or more primers, 24 or more primers, 25 or more primers, 30 or more primers, 35 or more primers, 40 or more primers, 45 or more primers, 50 or more primers, 55 or more primers, 60 or more primers, 65 or more primers, 70 or more primers, 75 or more primers, 80 or more primers, 85 or more primers, 90 or more primers, 95 or more primers, 100 or more primers, 200 or more primers, 300 or more primers, 400 or more primers, and 500 or more primers. In some embodiments, the number of primers in a given redundant subset of primers ranges from 2 to 50 including, e.g., 3 to 50, 3 to 40, 3 to 30, 3 to 25, 3 to 24, 3 to 22, 3 to 20, 3 to 18, 3 to 16, 3 toAtty. Docket No.: CLCT-011 WO

[0109] 14, 3 to 12, 3 to 10, 3 to 8, 3 to 6, 4 to 50, 4 to 40, 4 to 30, 4 to 25, 4 to 24, 4 to 22, 4 to 20, 4 to 18, 4 to 16, 4 to 14, 4 to 12, 4 to 10 and 4 to 8. In some embodiments, the number of primers in each replicate set is at least 3 primers. In some embodiments, the number of primers in each replicate set of primers ranges from 3 to 24. In some embodiments, the number of primers in each replicate set of primers ranges from 4 to 8. In some embodiments, the number of primers in each replicate set of primers is 8.

[0110] The multiplex collection of nucleic acid primers includes at least one replicate set of primers. In some cases, the multiplex collection of nucleic acid primers includes two or more replicate sets of primers (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 75, 100, 150, 200, 250, 500, or 1000 replicate sets of primers) including, e.g., 10 or more replicate sets of primers, 50 or more replicate sets of primers, 100 or more replicate sets of primers, and 1000 or more replicate sets of primers.

[0111] In some embodiments, the same collection of RBCs is present in each replicate set of primers. In other words, each replicate set of primers includes the same RBCs. For example, in one instance, a multiplex collection of nucleic acid primers includes 2 replicate sets of primers, set A and set B. Set A consists of 8 primers, where each primer has a common domain that binds to target A. Set A consists of a first primer that includes RBC1 , a second primer that includes RBC2, a third primer that includes RBC3, a fourth primer that includes RBC4, a fifth primer that includes RBC5, a sixth primer that includes RBC6, a seventh primer that includes RBC7 and an eighth primer that includes RBC8. Set B consists of 8 primers, where each primer has a common domain that binds to target B. Set B consists of a first primer that includes RBC1 , a second primer that includes RBC2, a third primer that includes RBC3, a fourth primer that includes RBC4, a fifth primer that includes RBC5, a sixth primer that includes RBC6, a seventh primer that includes RBC7 and an eighth primer that includes RBC8. In this example, the same collection of RBCs (i.e., RBC1, RBC2, RBC3, RBC4, RBC5, RBC6, RBC7 and RBC8) is present in each replicate set of primers.

[0112] In some embodiments, different collections of RBCs are present in each replicate set of primers. In other words, replicate sets of primers do not include the same RBCs. For example, in one instance, a multiplex collection of nucleic acid primers includes 2 replicate sets of primers, set C and set D. Set C consists of 8 primers, where each primer has a common domain that binds to target C. Set C consists of a first primer that includes RBC1 , a second primer that includes RBC2, a third primer that includes RBC3, a fourth primer that includes RBC4, a fifth primer that includes RBC5, a sixth primer that includes RBC6, a seventh primer that includesAtty. Docket No.: CLCT-011 WO

[0113] RBC7 and an eighth primer that includes RBC8. Set D consists of 8 primers, where each primer has a common domain that binds to target D. Set D consists of a first primer that includes RBC9, a second primer that includes RBC10, a third primer that includes RBC11 , a fourth primer that includes RBC12, a fifth primer that includes RBC13, a sixth primer that includes RBC14, a seventh primer that includes RBC15 and an eighth primer that includes RBC16. In this example, different sets of RBCs are present in each replicate set of primers (e.g., replicate set C consists of RBC1, RBC2, RBC3, RBC4, RBC5, RBC6, RBC7 and RBC8, while replicate set D consists of RBC9, RBC10, RBG11 , RBC12, RBC13, RBC14, RBG15 and RBC16).

[0114] In some cases, RBCs in different replicate sets may overlap. For example, in one instance, a multiplex collection of primers includes 2 replicate sets of primers, set E and set F. Set E consists of 3 primers, where each primer has a common domain that binds to target E. Set E consists of a first primer that includes RBC1 , a second primer that includes RBC2 and a third primer that includes RBC10. Set F consists of 3 primers, where each primer has a common domain that binds to target F. Set F consists of a first primer that includes RBC1 , a second primer that includes RBC2 and a third primer that includes RBC20. In this example, set E includes RBC1, RBC2 and RBC10, while set F includes RBC1 , RBC2 and RBC20, such that the RBC sets for targets E and F are partially overlapping (i.e., they share one or more, but not all RBCs).

[0115] In some cases, there are a different number of RBCs in each replicate set. For example, a multiplex collection of primers includes 2 replicate sets of primers, set G and set H. Set G consists of 3 primers, where each primer has a common domain that binds to target G. Set G consists of a first primer that includes RBC1 , a second primer that includes RBC2 and a third primer that includes RBC3. Set H consists of 5 primers, where each primer has a common domain that binds to target H. Set H consists of a first primer that includes RBC10, a second primer that includes RBC20, a third primer that includes RBC30, a fourth primer that includes RBC40 and a fifth primer that includes RBC50. In this example, set G includes three RBCs (e.g., RBC1 , RBC2 and RBC3), which is a different number of RBCs than set H, where set H includes five RBCs (e.g., RBC10, RBC20, RBC30, RBC40 and RBC50).

[0116] In embodiments where the multiplex collection of gene specific primers includes a forward primer set and a reverse primer set, one or both of the forward and reverse primer sets may include at least one replicate set of primers, where the replicate set of primers includes primers having a common gene specific primer (GSP) domain and a different RBC domain. In some embodiments, the forward primer set includes at least one replicate set of primers. InAtty. Docket No.: CLCT-011 WO

[0117] some embodiments, the reverse primer set includes at least one replicate set of primers. In some embodiments, both the forward primer set and the reverse primer set include at least one replicate set of primers. As such, in these embodiments, both the forward and the reverse primers are present as replicate sets of primers, such that the multiplex collection of gene specific primers includes at least one replicate set of forward primers and at least one replicate set of reverse primers.

[0118] In embodiments where the multiplex collection of gene specific primers includes a forward and a reverse primer set, the forward and / or reverse primer set may include primers that are not part of the at least one replicate set of primers. For example, in the case of multiplex amplification of different gene families, one gene family may be amplified using a replicate set of primers, while other gene families are amplified using primers that are not part of a replicate set of primers. As another example, in cases where the multiplex amplification is designed for analysis of immune receptor genes, T cell marker genes and B cell marker genes, only primers designed for the immune receptor genes may be amplified using a replicate set of primers, while the T cell marker genes and B cell marker genes are amplified using primers that are not part of a replicate set of primers.

[0119] In embodiments where the multiplex collection of gene specific primers includes a forward primer set and a reverse primer set, contacting the nucleic acid sample with the multiplex collection of nucleic acid primers may include first contacting the nucleic acid sample with the reverse primer set to produce a reverse primer template extension product composition, and then contacting the reverse primer template extension product composition with the forward primer set to produce a forward primer extension product composition. In other words, reverse primers of the reverse primer set hybridize to nucleic acid templates (e.g., RNA templates or DNA templates) of the nucleic acid template sample and are extended (e.g., by a polymerase) along the nucleic acid templates to form reverse primer extension products. Next, forward primers of the forward primer set hybridize to the reverse primer extension products (e.g., cDNAs) and are extended (e.g., by a polymerase) along the reverse primer extension products to form forward primer extension products.

[0120] In embodiments where the multiplex collection of gene specific primers includes a forward primer set and a reverse primer set, contacting the nucleic acid sample with the multiplex collection of nucleic acid primers may include first contacting the nucleic acid sample with the forward primer set to produce a forward primer template extension product composition, and then contacting the forward primer template extension product composition with the reverseAtty. Docket No.: CLCT-011 WO

[0121] primer set to produce a reverse primer extension product composition. In other words, forward primers of the forward primer set hybridize to nucleic acid templates (e.g., DNA templates) of the nucleic acid template sample and are extended (e.g., by a polymerase) along the nucleic acid templates to form forward primer extension products. Next, reverse primers of the reverse primer set hybridize to the forward primer extension products (e.g., cDNAs) and are extended (e.g., by a polymerase) along the forward primer extension products to form reverse primer extension products.

[0122] In some embodiments, each replicate set of at least two replicate sets binds to a different target. For instance, each replicate set of primers is specific for a particular gene. For example, a multiplex collection of nucleic acid primers may include four replicate sets of primers consisting of set A, set B, set C and set D. In this case, the primers of set A are specific for gene A, the primers of set B are specific for gene B, the primers of set C are specific for gene C and the primers of set D are specific for gene D. In other words, each replicate set of primers targets a particular gene.

[0123] In some embodiments, the primers may be in the form of TSOs. In such embodiments, the replicate set of primers includes TSOs, where the TSOs have a common domain (e.g., a common TSO domain responsible for template switching reaction) and a different RBC domains. In some instances, the TSOs may further include a capture domain, an anchor domain or any other desired domain, e.g., a domain employed in a next generation sequencing protocol. In some embodiments, the common domain is a capture domain. In other words, TSOs of a replicate set have a common capture domain and a different RBC domain. The capture domain binds to a non-templated stretch of nucleotides added by a polymerase (e.g., a reverse transcriptase) at the 3’ end of the cDNA strand generated in a primer extension reaction. The capture domain of the TSO may be referred to as a 3’ capture domain, where the 3' capture domain may vary in length, and in some instances ranges from 2 to 10 nts in length, such as 3 to 7 nts in length. The sequence of the 3' capture domain may be any convenient sequence, e.g., an arbitrary sequence, a heteropolymeric sequence (e.g., a hetero-trinucleotide) or homopolymeric sequence (e.g., a homo-trinucleotide, such as rG-rG-rG), or the like. Examples of 3' capture domains and template switch oligonucleotides are further described in U.S. Patent No. 5,962,272, the disclosure of which is herein incorporated by reference.

[0124] In embodiments where the at least one replicate set of primers includes TSOs, the TSOs may be employed with one or more primers (e.g., reverse primers) to produce a template extension product composition. In other words, a primer (e.g., a reverse primer) may prime aAtty. Docket No.: CLCT-011 WO

[0125] primer extension reaction and be extended by a polymerase along the nucleic acid template. In some cases, reverse primers could be gene-specific primers and in other cases are universal primers, e.g., an oligo dT primer. Once the polymerase reaches the end of the nucleic acid template, the polymerase may add a stretch of non-templated nucleic acids to the 3’ end of the cDNA strand. The TSO (e.g., the TSO that includes the RBC), may then hybridize with the non-templated nucleic acids at the 3’ end of the cDNA strand and the polymerase may template switch such that the TSO is the template. Reverse primers used in combination with TSOs include, but are not limited to, gene-specific reverse primers, reverse primers that include an oligodT region, or reverse primers that include a domain that binds to a universal sequence.

[0126] Anchor domains

[0127] In some embodiments, the primers employed in embodiments of the invention further include anchor domains. In some embodiments, the primers of the forward primer set include anchor domains. In some embodiments, the primers of the reverse primer set include anchor domains. In some embodiments, the primers of the forward and the reverse primer sets include anchor domains. In embodiments where the primers are in the form of template switch oligonucleotides (TSOs), the TSOs may include anchor domains.

[0128] The anchor domains may be common anchor domains. By common anchor domains, it is meant an anchor domain that is common between primers. In other words, all forward primers and / or reverse primers may include the same common anchor domain.

[0129] Anchor domains are domains that are employed in nucleic acid amplification, such as polymerase chain reaction (PCR) (e.g., multiplex PCR), steps of the methods, where they serve as primer binding sites for the primers employed in such amplification steps. Where the amplification employed is PCR, the anchor domains may also be referred to as PCR primer binding domains. The length of the anchor domains may vary, as desired. In some instances, the anchor domains of each primer range in length from 10 to 50 nucleotides, such as 15 to 45 nucleotides, e.g., 18 to 40 nucleotides, including 18 to 30 nucleotides. Where desired, the anchor domains may include PCR suppression sequences. PCR suppression sequences are sequences configured to suppress the formation of non-target DNA during PCR amplification reactions, e.g., via the production of pan-like structures. Such sequences, when present, may vary in length, ranging in some instances from 5 to 25 nucleotides, such as 7 to 21 nucleotides, including 7 to 20 nucleotides. PCR suppression sequences of interest include, but are not limited to, those sequences described in U.S. Patent No. 5,565,340; the disclosure of which isAtty. Docket No.: CLCT-011 WO

[0130] herein incorporated by references. An example of forward and reverse anchor domains that include PGR suppression sequences are: AGCACCGACCAGCAGACA (SEQ ID NO:01).

[0131]

[0132] As reviewed above, methods of disclosure include contacting a nucleic acid sample with a multiplex collection of nucleic acid primers to produce a template extension product composition. As defined above, a template extension product composition is a nucleic acid composition that includes nucleic acids that are template extension products (i.e., nucleic acids that are produced by template extension reactions, e.g., reactions where primers are annealed to a template and then extended).

[0133] Nucleic acid samples suitable for the methods of the disclosure include any convenient nucleic acid sample. By nucleic acid sample, it is meant a sample comprising nucleic acids. In some embodiments, the nucleic acid sample is a ribonucleic acid (RNA) sample. In some embodiments, the nucleic acid sample includes RNA (e.g., solely RNA). In some embodiments, the nucleic acid sample is a deoxyribonucleic acid (DNA) sample. In some embodiments, the nucleic acid sample includes DNA (e.g., solely DNA). In some embodiments, the nucleic acid sample includes both RNA and DNA.

[0134] The RNA may be any type of RNA (or sub-type thereof) including, but not limited to, a messenger RNA (mRNA), a microRNA (miRNA), a small interfering RNA (siRNA), a transacting small interfering RNA (ta-siRNA), a natural small interfering RNA (nat-siRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), a small nucleolar RNA (snoRNA), a small nuclear RNA (snRNA), a long non-coding RNA (IncRNA), a non-coding RNA (ncRNA), a transfer-messenger RNA (tmRNA), a precursor messenger RNA (pre-mRNA), a small Cajal body-specific RNA (scaRNA), a piwi-interacting RNA (piRNA), an endoribonuclease-prepared siRNA (esiRNA), a small temporal RNA (stRNA), a signal recognition RNA, a telomere RNA, a ribozyme, or any combination of RNA types thereof or subtypes thereof. In some embodiments, the RNA is a messenger RNA (mRNA).

[0135] The DNA may be any type of DNA (or sub-type thereof) including, but not limited to, cDNA, genomic DNA (e.g., prokaryotic genomic DNA (e.g., bacterial genomic DNA, archaea genomic DNA, etc.), eukaryotic genomic DNA (e.g., plant genomic DNA, fungi genomic DNA, animal genomic DNA (e.g., mammalian genomic DNA (e.g., human genomic DNA, rodent genomic DNA (e.g., mouse, rat, etc.), etc.), insect genomic DNA (e.g., drosophila), amphibian genomic DNA (e.g., Xenopus), etc.)), viral genomic DNA, mitochondrial DNA, or anyAtty. Docket No.: CLCT-011 WO

[0136] combination of DNA types thereof or subtypes thereof. In some embodiments, the DNA is mammalian genomic DNA.

[0137] Nucleic acid samples may be obtained from any convenient sample of cells. Samples of cells may include any type of cell including, but not limited to, immune system cells, neuronal cells, cardiac cells, endothelial cells, fibroblasts, liver cells, tumor cells, or combinations thereof. Immune cells include, but are not limited to, lymphocytes, e.g., a T cell (e.g., a cytotoxic T cell (e.g., a CD8+ T cell), a helper T cells (e.g., a GD4+ T cell), a regulatory T cells (“Treg”), etc.) a natural killer (NK) cells, a B cells, and the like. In some cases, samples may include fractions of specific immune cells (e.g., fractions of naive, memory, or cytotoxic immune cells, etc.). Subject immune cells may also include e.g., whole blood, peripheral blood mononuclear cells (PBMC), macrophages, dendritic cells, monocytes, etc. Samples of cells may include stem cells, immortalized cells or primary cells. In some instances, the sample of cells utilized in the subject methods may be a sample of mammalian cells, such as rodent (e.g., mouse or rat) cells, a nonhuman primate cells, human cells, or the like. In some instances, the sample of mammalian cells may be mammalian blood cells, including but not limited to e.g., rodent (e.g., mouse or rat) blood cells, non-human primate cells, human blood cells, or the like. In some embodiments, the sample of cells are human cells. In some embodiments, the sample of blood cells is human blood cells.

[0138] In some embodiments, the nucleic acid samples are nucleic acid samples (e.g., RNA and / or DNA samples) obtained from B cells. In some embodiments, the nucleic acid samples are nucleic acid samples (e.g., RNA and / or DNA samples) obtained from T cells. In some embodiments, the nucleic acid samples are nucleic acid samples (e.g., RNA and / or DNA samples) obtained from B cells and T cells. In some embodiments, nucleic acid samples are nucleic acid samples obtained from low lymphocyte samples (i.e. , samples comprising a low number of lymphocytes). A low lymphocyte sample includes, but is not limited to, a cancer biopsy sample, a non-lymphoid tissue sample, a biological fluid (e.g., saliva, urine, semen, vaginal fluids, mucus, sweat, menstrual fluid, etc.) sample, a formalin-fixed paraffin-embedded (FPPE) sample, a tissue microsample, and a blood microsamples. By microsample, it is generally meant a sample of less than 50 pL (e.g., 25 pL, 30 pL, 35 pL, 40 pL, 45 pL, etc.). In some embodiments, nucleic acid samples are nucleic acid samples obtained from high lymphocyte samples (i.e., samples comprising a high number of lymphocytes). A high lymphocyte sample includes, but is not limited to, a whole blood sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a lymphoid tissue sample, an immune cellAtty. Docket No.: CLCT-011 WO

[0139] fraction sample and a bone marrow aspirate sample. Immune cell fractions include fractions of particular immune cells including, but not limited to, naive T and / or B cells, memory T and / or B cells, effector T and / or B cells, and exhausted T and / or B cells.

[0140] In some cases, nucleic acid samples may be obtained from a pool of cells. In some embodiments, the pool of cells are directly obtained from samples of cells (e.g., whole blood, PBMC cells). In some embodiments, where the cells originate from a sample where the cells are not initially isolated, e.g., when the cell is part of a tissue, the cells may be obtained from an initial cell sample, e.g., using any convenient cell isolation protocol. In some instances, cells may be obtained using standard methods known in the art including, for example, enzymatically using trypsin or papain to digest proteins connecting cells in tissue samples or releasing adherent cells in culture, or mechanically separating cells in a sample. In some instances, cells (e.g., fractions of cells, such as immune cells) may be obtained by sorting a cellular sample using a cell sorter instrument. By “cell sorter” as used herein is meant any instrument that allows for the sorting of individual cells into an appropriate vessel for downstream processes, such as those processes of library preparation as described herein. Useful cell sorters include flow cytometers, such as those instruments utilized in fluorescence activated cell sorting (FACS).

[0141] In some cases, nucleic acid samples may be obtained from individual cells (i.e. , single cells). Single cells, for use in the herein described methods relating thereto, may be obtained by any convenient method. For example, in some instances, single cells may be obtained through limiting dilution of cellular sample. In some instances, the present methods may include a step of obtaining single cells. A single cell suspension can be obtained using standard methods known in the art including, for example, enzymatically using trypsin or papain to digest proteins connecting cells in tissue samples or releasing adherent cells in culture, or mechanically separating cells in a sample. Single cells can be placed in any suitable reaction vessel in which single cells can be treated individually. For example, a 96-well plate, 384 well plate, or a plate with any number of wells such as 2000, 4000, 6000, or 10000 or more. The multi-well plate can be part of a chip and / or device. The present disclosure is not limited by the number of wells in the multi-well plate. In various embodiments, the total number of wells on the plate is from 100 to 200,000, or from 5000 to 10,000. In other embodiments the plate comprises smaller chips, each of which includes 5,000 to 20,000 wells. For example, a square chip may include 125 by 125 nanowells, with a diameter of 0.1 mm. Such methods are further described in greater detail below.Atty. Docket No.: CLCT-011 WO

[0142] In some instances, single cells may be obtained by sorting a cellular sample using a cell sorter instrument. By “cell sorter” as used herein is meant any instrument that allows for the sorting of individual cells into an appropriate vessel for downstream processes, such as those processes of library preparation as described herein. Useful cell sorters include flow cytometers, such as those instruments utilized in fluorescence activated cell sorting (FACS). Flow cytometry is a well-known methodology using multi-parameter data for identifying and distinguishing between different particle (e.g., cell) types i.e., particles that vary from one another terms of label (wavelength, intensity), size, etc., in a fluid medium. In flow cytometrically analyzing a sample, an aliquot of the sample is first introduced into the flow path of the flow cytometer. When in the flow path, the cells in the sample are passed substantially one at a time through one or more sensing regions, where each of the cells is exposed separately individually to a source of light at a single wavelength (or in some instances two or more distinct sources of light) and measurements of scatter and / or fluorescent parameters, as desired, are separately recorded for each cell. The data recorded for each cell is analyzed in real time or stored in a data storage and analysis means, such as a computer, for later analysis, as desired.

[0143] Cells sorted using a flow cytometer may be sorted into a common vessel (i.e., a single tube), or may be separately sorted into individual vessels. For example, in some instances, cells may be sorted into individual wells of a multi-well plate, as described below.

[0144] Useful cell sorters also include multi-well-based systems that do not employ flow cytometry. Such multi-well-based systems include essentially any system where cells may be deposited into individual wells of a multi-well container by any convenient means, including e.g., through the use of Poisson distribution (i.e., limiting dilution) statistics, individual placement of cells (e.g., through manual cell picking or dispensing using a robotic arm or pipettor). In some instances, useful multi-well systems include a multi-well wafer or chip, where cells are deposited into the wells or the wafer / chip and individually identified by a microscopic analysis system. In some instances, an automated microscopic analysis system may be employed in conjunction with a multi-well wafer / chip to automatically identify individual cells to be subjected to downstream analyses, including library preparation, as described herein.

[0145] In some instances, one or more cells may be sorted into or otherwise transferred to an appropriate reaction vessel. Reaction components may be added to reaction vessels, including e.g., components for preparing a template nucleic acid component, components for generating a product double stranded cDNA, components for one or more library preparation reactions, etc.Atty. Docket No.: CLCT-011 WO

[0146] The wells of a multi-well device can be designed such that a single well includes a single cell or a single droplet. An individual cell or droplet may also be isolated in any other suitable container, e.g., microfluidic chamber, droplet, nanowell, tube, etc. Any convenient method for manipulating single cells or droplets may be employed, where such methods include fluorescence activated cell sorting (FACS), robotic device injection, gravity flow, or micromanipulation and the use of semi-automated cell pickers (e.g. the Quixell™ cell transfer system from Stoelting Co.), etc. In some instances, single cells or droplets can be deposited in wells of a plate according to Poisson statistics (e.g., such that approximately 10%, 20%, 30% or 40% or more of the wells contain a single cell or droplet - which number can be defined by adjusting the number of cells or droplets in a given unit volume of fluid that is to be dispensed into the containers). In some instances, a suitable reaction vessel comprises a droplet (e.g., a microdroplet). Individual cells or droplets can, for example, be individually selected based on features detectable by microscopic observation, such as location, morphology, the presence of a reporter gene (e.g., expression), the presence of a bound antibody (e.g., antibody labelling), FISH, the presence of an RNA (e.g., intracellular RNA labelling), or qPCR.

[0147] Following obtainment of a desired cell population or single cells, e.g., as described above, nucleic acids can be released from the cells by lysing the cells. Lysis can be achieved by, for example, heating or freeze-thaw of the cells, or by the use of detergents or other chemical methods, or by a combination of these. However, any suitable lysis method can be used. In some instances, a mild lysis procedure can advantageously be used to prevent the release of nuclear chromatin, thereby avoiding genomic contamination of a cDNA library, and to minimize degradation of mRNA. For example, heating the cells at 72QC for 2 minutes in the presence of Tween-20 is sufficient to lyse the cells while resulting in no detectable genomic contamination from nuclear chromatin. Alternatively, cells can be heated to 65eC for 10 minutes in water (Esumi et al., Neurosci Res 60(4):439-51 (2008)); or 70QC for 90 seconds in PCR buffer II (Applied Biosystems) supplemented with 0.5% NP-40 (Kurimoto et al., Nucleic Acids Res 34(5) :e42 (2006)); or lysis can be achieved with a protease such as Proteinase K or by the use of chaotropic salts such as guanidine isothiocyanate (U.S. Patent No. 8,802,367).

[0148] FIG. 1 depicts an embodiment where a nucleic acid (e.g., RNA) is contacted by one or more forward primers and a reverse primer. As shown in FIG. 1 , the nucleic acid that is contacted by the primers is an mRNA that corresponds to a TOR or a BCR (e.g., a TCR mRNA or a BCR mRNA). The reverse primer (e.g., a reverse primer of a replicate set of reverse primers) includes a Rev domain, an RBC domain and an Anchor 2 domain. The Rev domain isAtty. Docket No.: CLCT-011 WO

[0149] designed to hybridize to the conserved C region of different TCR or BCR mRNA isoforms. In other words, the C region remains consistent for different TCR or BCR mRNAs, while the V, D and J regions may vary between different TCR or BCR mRNAs. A first forward primer (e.g., a forward primer of a first replicate set of forward primers) includes a gene specific Fwd domain and an Anchor 1 domain. The first forward primer hybridizes to the FR3 variable region via the Fwd domain of the template extension product generated from Rev C primers, which allows (after extension using the forward primer) amplification of the CDR3 region of the TCR or BCR nucleic acid. The resultant template extension product composition includes the CDR3 region flanked by two anchor domains (e.g., Anchor 1 and Anchor 2) and an RBC domain. The mRNA (e.g., TCR mRNA or BCR mRNA) may be contacted with a second forward primer (e.g., a forward primer of a second replicate set of forward primers) in addition to or instead of the first forward primer. The second forward primer includes a Fwd domain and an Anchor 1 domain. The second forward primer hybridizes to the FR1 variable region via the Fwd domain of the template extension product generated from Rev C primers, which allows (after extension using the forward primer) amplification of the full length CDR1-CDR2-CDR3 variable region of the TCR or BCR nucleic acid. The resultant template extension product composition includes the CDR1-CDR2-CDR3 region flanked by two anchor domains (e.g., Anchor 1 and Anchor 2) and an RBC domain.

[0150] FIG. 2 depicts an embodiment where a nucleic acid (e.g., DNA) is contacted by a forward primer and a reverse primer. As shown in FIG. 2, the nucleic acid that is contacted by the primers is a DNA that corresponds to a TCR or a BCR gene (e.g., TCR DNA or BCR DNA). The reverse primer (e.g., a reverse primer of a replicate set of reverse primers) includes a Rev domain, an RBC domain and an Anchor 2 domain. The Rev domain is designed to hybridize to the J region of the sense strand of the TCR or BCR DNA. The forward primer (e.g., a forward primer of a replicate set of forward primers) includes a Fwd domain and an Anchor 1 domain. The forward primer hybridizes to the FR3 variable region of the template extension product generated from the reverse primer via the Fwd domain, which allows (after extension using the forward primer) amplification of the CDR3 region of the TCR or BCR nucleic acid. The resultant template extension product composition includes the CDR3 region flanked by two anchor domains (e.g., Anchor 1 and Anchor 2) and an RBC domain.

[0151] Methods of producing a template extension product composition may be carried out under primer extension reaction conditions. By “primer extension reaction conditions” it is meant reaction conditions that permit polymerase-mediated extension of a 3’ end of a nucleic acidAtty. Docket No.: CLCT-011 WO

[0152] strand, i.e., primer, hybridized to a template nucleic acid (e.g., a template DNA or a template RNA). Achieving suitable reaction conditions may include selecting reaction mixture components, concentrations thereof, and a reaction temperature to create an environment in which the polymerase is active and the relevant nucleic acids in the reaction interact (e.g., hybridize) with one another in the desired manner.

[0153] In some cases, the primer extension reaction and the hybridization reaction may be carried out independently of each other. For example, in some cases, a set of reverse gene specific primers may hybridize with a template (e.g., a mRNA template) under hybridization conditions, and then the reverse primer-mRNA hybrids may be purified by separating the reverse primer-mRNA hybrids from primers that are not hybridized. After purification, the reverse primer-mRNA hybrids may then be incubated under primer extension reaction conditions to generate template extension product compositions.

[0154] In cases where hybridization reactions and primer extension reactions are carried out as separate steps, the reverse primers and nucleic acid template sample (e.g., RNA template sample) may be mixed and incubated together under hybridization conditions that are optimized to achieve a high yield of reverse primer-mRNA hybrids. In such instances, the amount of the template nucleic acid (e.g., target nucleic acid) may vary depending on the source of the available biological sample (e.g., single cell or tissue sample).

[0155] The concentration of primers in the primer extension reaction mixture produced upon combination of the template nucleic acid and primers may vary, as desired. The amount of target template nucleic acid that is combined with the primers and other reagents, e.g., as described below, to produce a primer extension reaction mixture may vary. The amount of template nucleic acid may range from 1 pg to 5,000 ng, such as 10 pg to 2,000 ng, including 1 ng to 1 ,000 ng. The reverse primer concentration may range from 1 nM to 500 nM, such as 5 nM to 200 nM, including 10 nM to 100 nM. The hybridization reaction may be performed in a buffer with neutral pH in the range of pH 6 to pH 8, a high salt concentration (e.g. NaCI) in the range 0.2 M to 1 M, a hybridization temperature defined by the melting temperature of the reverse primers as described above, and an incubation time in the range between 10 min to 2 hours. In some instances, the target nucleic acid template composition is combined into the reaction mixture such that the final concentration of nucleic acid in the reaction mixture ranges from 1 fg / pL to 10 pg / pL, such as from 1 pg / pL to 5 pg / pL, such as from 0.1 ng / pL to 50 ng / pL, such as from 0.5 ng / pL to 20 ng / pL, including from 1 ng / pL to 10 ng / pL.Atty. Docket No.: CLCT-011 WO

[0156] In producing the primer extension reaction mixture, the primers and target template nucleic acid composition are combined with a number of additional reagents (e.g., to increase specificity, uniformity, yield, etc. of extension products), which may vary as desired. A variety of polymerases may be employed when practicing the subject methods. Reference to a particular polymerase, such as those exemplified below, will be understood to include functional variants thereof unless indicated otherwise. Examples of useful polymerases include DNA polymerases, e.g., where the template nucleic acid is DNA. In some instances, DNA polymerases of interest include, but are not limited to: thermostable DNA polymerases, such as may be obtained from a variety of bacterial species, including Thermus aquaticus (Taq), Thermus thermophilus (Tth), Thermus filiformis, Thermus flavus, Thermococcus literalis, and Pyrococcus furiosus (Pfu) or modified and mutated versions of these DNA polymerases (e.g. Phusion DNA polymerase, Q5 DNA polymerase, etc.). Alternatively, where the target template nucleic acid composition is made up of RNA, the polymerase may be a reverse transcriptase (RT), where examples of reverse transcriptases include Moloney Murine Leukemia Virus reverse transcriptase (MMLV RT), e.g., SuperScript II, SuperScript III, MaxiScript reverse transcriptase (Thermo-Fisher), SMARTScribe™ reverse transcriptase (Takara), AMV reverse transcriptase, Bombyx mori reverse transcriptase (e.g., Bombyx mori R2 non-LTR element reverse transcriptase), etc. In one embodiment, the enzymes with DNA polymerase activity are designed for hot-start primer extension reaction, e.g., used as a complex with specific antibody or chemical compound which blocks enzymatic activity at low temperature but fully releases the activity at reaction conditions. For example, in some instances a hot-start reverse transcriptase composition, e.g. complex between MMLV RT and Therma-Stop RT reagent (Thermagenix) is employed.

[0157] Primer extension reaction mixtures also include dNTPs. In certain aspects, each of the four naturally-occurring dNTPs (dATP, dGTP, dCTP and dTTP) are added to the reaction mixture. For example, dATP, dGTP, dCTP and dTTP may be added to the reaction mixture such that the final concentration of each dNTP is from 0.05 to 10 mM, such as from 0.1 to 2 mM, including 0.2 to 1 mM. According to one embodiment, at least one type of nucleotide added to the reaction mixture is a non-naturally occurring nucleotide, e.g., a modified nucleotide having a binding or other moiety (e.g., a fluorescent moiety) attached thereto, a nucleotide analog, or any other type of non-naturally occurring nucleotide that finds use in the subject methods or a downstream application of interest.

[0158] In addition to the template nucleic acid, primers, the polymerase, and dNTPs, the reaction mixture may include buffer components that establish an appropriate pH, saltAtty. Docket No.: CLCT-011 WO

[0159] concentration (e.g., KCI concentration), metal cofactor concentration (e.g., Mg2+or Mn2+concentration), and the like, for the extension reaction and template switching to occur. Other components may be included, such as one or more nuclease inhibitors (e.g., an RNase inhibitor and / or a DNase inhibitor), one or more additives for facilitating amplification / replication of GC rich sequences (e.g., GC-Melt™ reagent (Clontech Laboratories, Inc. (Mountain View, GA)), betaine, single-stranded binding proteins (e.g., T4 Gene 32, cold shock protein A (CspA), recA protein, and / or the like) DMSO, ethylene glycol, 1,2-propanediol, or combinations thereof), one or more molecular crowding agents (e.g., polyethylene glycol, or the like), one or more enzymestabilizing components (e.g., DTT present at a final concentration ranging from 1 to 10 mM (e.g., 5 mM)), and / or any other reaction mixture components useful for facilitating polymerase-mediated extension reactions.

[0160] The primer extension reaction mixture can have a pH suitable for the primer extension reaction. In certain embodiments, the pH of the reaction mixture ranges from 5 to 9, such as from 7 to 9, including from 8 to 9, e.g., 8 to 8.5. In some instances, the reaction mixture includes a pH adjusting agent. pH adjusting agents of interest include, but are not limited to, sodium hydroxide, hydrochloric acid, phosphoric acid buffer solution, citric acid buffer solution, and the like. For example, the pH of the reaction mixture can be adjusted to the desired range by adding an appropriate amount of the pH adjusting agent.

[0161] The temperature range suitable for production of the product nucleic acid may vary according to factors such as the particular polymerase employed, the melting temperatures of any optional primers employed, etc. According to one embodiment, the primer extension reaction conditions include bringing the reaction mixture to a temperature ranging from 4 to 72 °C, such as from 16 to 70QC, e.g., 37 to 65QC, such as 60 -C to 65QC. The temperature of the reaction mixture may be maintained for a sufficient period of time for polymerase mediated, template directed primer extension to occur. While the period of time may vary, in some instances the period of time ranges from 5 to 60 minutes, such as 15 to 45 minutes, e.g., 30 minutes.

[0162] In a given primer extension reaction condition, where desired, hybridization complexes of template and primer may be purified, e.g., via separation from excess of non-bound primers, e.g., by nuclease treatment or binding to solid support, e.g., such as beads, e.g., as described above. In this way, excess of primers, such as oligo dT primers and / or gene-specific primers, may be removed in order to achieve a high specificity of primer extension reaction from the target template sequences.Atty. Docket No.: CLCT-011 WO

[0163] Where desired, the primer extension reaction conditions may include one or more temperature cycling steps. For example, in some instances, the primer extension product composition is produced by a method that includes first contacting the target nucleic acid template composition with a first primer subset that includes for example the forward primers of the set of primer pairs under primer extension reaction conditions to produce a forward primer extension product composition; increasing the temperature to denature the resultant product and template strands and inactivate any additional enzymatic activity (e.g., exonuclease I activity added after extension step to degrade PCR primers) present in the forward primer extension product composition (where the elevated temperature may vary, ranging in some instances from 90 to 100 °C, such as 95 °C) and then contacting the resultant denatured forward primer extension product composition with a second primer subset that includes the reverse primers of the set of primer pairs under primer extension reaction conditions to produce the desired primer extension product composition. Where desired, the primer extension products and template nucleic acids may be separated from any free forward primers prior to contact with the set of reverse primers. The extended DNA products after the first and second extension steps may be purified from the excess of the primers using any convenient protocol, including primer digestion with exonucleases (exonuclease I) or purification, such as Magnetic beads or spin columns, etc.

[0164] Amplifying the Template Extension Product Composition

[0165] As summarized above, methods of producing an amplified next generation sequencing (NGS) nucleic acid library include amplifying the template extension product composition to produce the amplified NGS nucleic acid library. Amplification is distinct from a template extension reaction in that amplification refers to the process of creating multiple copies of a specific nucleic acid sequence. Meanwhile, template extension reactions involve extending a primer along a nucleic acid template (e.g., to generate a cDNA). For instance, template extension reactions may first be carried out to produce template extension products from the nucleic acid template, and then the template extension products may be amplified.

[0166] In some embodiments, the amplifying may include one or more rounds of amplification, e.g., first and second rounds of amplification, etc. In some embodiments, the amplifying includes a first round of amplification with anchor primers and a second round of amplification with index primers to generate the amplified NGS nucleic acid library.Atty. Docket No.: CLCT-011 WO

[0167] An embodiment of the methods disclosed herein is depicted in FIG. 3. As shown in step 300 of FIG. 3, an mRNA is contacted with a reverse gene specific primer (e.g., a reverse gene specific primer of a replicate set of reverse gene specific primers). The reverse gene specific primer binds to the C region of TCR mRNA and BCR mRNA. The reverse gene specific primer includes a RevGSP domain, an RBC domain and an Anchor 2 domain. The RevGSP domain hybridizes to the C region present in the TCR or BCR mRNA. Next, at step 302, the reverse primer is extended by a polymerase (e.g., reverse transcriptase), to create a first strand antisense cDNA that includes the RBC domain and the Anchor 2 domain. Next, at step 304, a forward gene specific primer (e.g., a forward gene specific primer of a replicate set of forward gene specific primers) is annealed to the first strand cDNA. The forward gene specific primer includes a FwdGSP domain that hybridizes to the V region and an Anchor 1 domain. The FwdGSP domain may be complementary to the FR3 V gene region to amplify the CDR3 region or complementary to the UTR-FR1 region to amplify the CDR1-CDR2-CDR3 region of the TCR or BCR genes. The forward gene specific primer is then extended by a polymerase (e.g., a DNA polymerase) to create a DNA strand that is flanked by the Anchor 1 and Anchor 2 domains and includes the RBC domain. Step 306 shows the first PCR step (i.e. , the first round of amplification), where the template extension product composition is amplified using an Anchor 1 primer (i.e., a primer that hybridizes to the Anchor 1 domain) and an Anchor 2 primer (i.e., a primer that hybridizes to the Anchor 2 domain). Next, at step 308, a second PCR step (i.e., second round of amplification) is carried out using index primers. The forward index primer includes a domain that hybridizes to the Anchor 1 domain, an index domain and P5 domain. The second index primer includes a domain that hybridizes to the Anchor 2 domain, an index domain and P7 domain. Step 310 shows an amplified nucleic acid that includes P5 and P7 domains, index domains (e.g., UDP1 and LIDP2) and an RBC domain. The amplified nucleic acids can then be sequenced using NGS at step 312.

[0168] Another embodiment of the method is depicted in FIG. 4. As shown in step 400 of FIG.

[0169] 4, a DNA molecule is contacted with a reverse gene specific primer (e.g., a reverse gene specific primer of a replicate set of reverse gene specific primers). The reverse gene specific primer binds to the J region of TCR DNA and BCR DNA). The reverse gene specific primer includes a RevGSP domain, an RBC domain and an Anchor 2 domain. The RevGSP domain hybridizes to the J region present in the TCR or BCR DNA. The reverse primer is extended by a polymerase (e.g., DNA polymerase), to create an antisense DNA strand that includes the RBC domain and the Anchor 2 domain. Next, at step 402, a forward gene specific primer (e.g., aAtty. Docket No.: CLCT-011 WO

[0170] forward gene specific primer of a replicate set of forward gene specific primers) are annealed to the antisense DNA. The forward gene specific primer includes a FwdGSP domain that hybridizes to the FR3 region and an Anchor 1 domain. The forward gene specific primer is then extended by a polymerase (e.g., a DNA polymerase) to create a sense DNA strand that is flanked by the Anchor 1 and Anchor 2 domains and includes the RBC domain. Step 404 shows the first PCR step (i.e., the first round of amplification), where the template extension product composition is amplified using an Anchor 1 primer (i.e., a primer that hybridizes to the Anchor 1 domain) and an Anchor 2 primer (i.e., a primer that hybridizes to the Anchor 2 domain). Next, at step 406, a second PCR step (i.e., second round of amplification) is carried out using index primers. The forward index primer includes a domain that hybridizes to the Anchor 1 domain, an index domain (e.g., UDP1) and P5 domain. The second index primer includes a domain that hybridizes to the Anchor 2 domain, an index domain (e.g., UDP2) and P7 domain. Step 408 shows an amplified nucleic acid that includes P5 and P7 domains, index domains (e.g., UDP1 and UDP2) and an RBC domain. The amplified nucleic acids can then be sequenced using NGS at step 410.

[0171] Another embodiment of the methods disclosed herein using single cells is depicted in FIG. 14. As shown in step 1400 of FIG. 14, single cells (e.g., T cells and B cells) are sorted into a 96-well plate with reverse gene specific primers. The cells in the 96-well plate are lysed. Next, as shown in step 1402, reverse gene specific primers (e.g., reverse gene specific primers of a replicate sets of reverse gene specific primers) hybridize to the mRNA in the wells of the 96-well plate. The reverse gene specific primers bind to the C region of TCR mRNA and BCR mRNA. The reverse gene specific primer includes a RevGSP domain, an RBC domain, a cell BC (i.e., cell barcode) domain and an Anchor 2 domain. The RevGSP domain hybridizes to the C region present in the TCR or BCR mRNA. Next, at step 1404, the reverse primer-mRNA hybrids are pooled, collected in a single tube and purified to remove unhybridized primers. Next, at step 1406, the reverse primer is extended by a polymerase (e.g., reverse transcriptase), to create a first strand antisense cDNA that includes the RBC domain and the Anchor 2 domain. Next, at step 1408, a forward gene specific primer (e.g., a forward gene specific primer of a replicate set of forward gene specific primers) is annealed to the first strand cDNA. The forward gene specific primer includes a FwdGSP domain that hybridizes to the V region and an Anchor 1 domain. The FwdGSP domain may be complementary to the FR3 V gene region to amplify the CDR3 region or complementary to the UTR-FR1 region to amplify the CDR1-CDR2-CDR3 region of the TCR or BCR genes. The forward gene specific primer is then extended by a polymerase (e.g., a DNAAtty. Docket No.: CLCT-011 WO

[0172] polymerase) to create a DNA strand that is flanked by the Anchor 1 and Anchor 2 domains and includes the RBC domain. Step 1410 shows the first PGR step (i.e., the first round of amplification), where the template extension product composition is amplified using an Anchor 1 primer (i.e., a primer that hybridizes to the Anchor 1 domain) and an Anchor 2 primer (i.e., a primer that hybridizes to the Anchor 2 domain). Next, at step 1412, a second PGR step (i.e., second round of amplification) is carried out using index primers. The forward index primer includes a domain that hybridizes to the Anchor 1 domain, an index domain and P5 domain. The second index primer includes a domain that hybridizes to the Anchor 2 domain, an index domain and P7 domain. Step 1414 shows an amplified nucleic acid that includes P5 and P7 domains, index domains (e.g., UDP1 and LIDP2) and an RBC domain. The amplified nucleic acids can then be sequenced using NGS.

[0173] As reviewed above, in some instances, template extension product compositions are amplified, where amplicons are produced from the template extension products. The term "amplicon" is employed in its conventional sense to refer to a piece of DNA that is the product of artificial amplification or replication events, e.g., as produced using various methods including polymerase chain reactions (PGR), ligase chain reactions (LCR), etc. Where template extension product compositions are amplified, the template extension product compositions, e.g., as described above, may include additional domains that are employed in subsequent amplification steps to produce a desired amplicon composition. For example, as illustrated in FIGS. 3 and 4 above, flanking anchor domains are provided in the primer extension products, where the flanking anchor domains include universal priming sites which may be employed in PGR amplification.

[0174] As such, embodiments of the methods may include combining a template extension product composition with universal forward and reverse primers under amplification conditions sufficient to produce a desired product amplicon composition. The forward and reverse universal primers may be configured to bind to the common forward and reverse anchor domains and thereby nucleic acids present in the template extension product compositions. The universal forward and reverse primers may vary in length, ranging in some instances from 10 to 75 nt, such as 20 to 60 nt.

[0175] In some instances, the universal forward and reverse primers include one or more additional domains, such as but not limited to: an indexing domain, a clustering domain, a Next Generation Sequencing (NGS) adapter domain (i.e., high-throughput sequencing (HTS) adapter domain), etc. Alternatively, these domains may be introduced during one or more subsequentAtty. Docket No.: CLCT-011 WO

[0176] steps, such as one or more subsequent amplification reactions, e.g., as described in greater detail below. The amplification reaction mixture will include, in addition to the primer extension product composition and universal forward and reverse primers, other reagents, as desired, such polymerase, dNTPs, buffering agents, etc., e.g., as described above.

[0177] Amplification conditions may vary. In some instances, the reaction mixture is subjected to polymerase chain reaction (PCR) conditions. PCR conditions include a plurality of reaction cycles, where each reaction cycle includes: (1) a denaturation step, (2) an annealing step, and (3) a polymerization step. The number of reaction cycles will vary depending on the application being performed, and may be 1 or more, including 2 or more, 3 or more, four or more, and in some instances may be 15 or more, such as 20 or more and including 30 or more, where the number of different cycles will typically range from about 12 to 24. The denaturation step includes heating the reaction mixture to an elevated temperature and maintaining the mixture at the elevated temperature for a period of time sufficient for any double stranded or hybridized nucleic acid present in the reaction mixture to dissociate. For denaturation, the temperature of the reaction mixture may be raised to, and maintained at, a temperature ranging from 85 to 100 -C, such as from 90 to 98 -C and including 94 to 98 -C for a period of time ranging from 3 to 120 sec, such as 5 to 30 sec. Following denaturation, the reaction mixture will be subjected to conditions sufficient for primer annealing to template DNA present in the mixture. The temperature to which the reaction mixture is lowered to achieve these conditions may be chosen to provide optimal efficiency and specificity, and in some instances ranges from about 50 to 75QC, such as 60 to 74QC and including 68 to 72QC. Annealing conditions may be maintained for a sufficient period of time, e.g., ranging from 10 sec to 30 min, such as from 10 sec to 5 min. Following annealing of the primer to a template DNA or during annealing of the primer to a template DNA, the reaction mixture may be subjected to conditions sufficient to provide for polymerization of nucleotides to the primer ends in manner such that the primer is extended in a 5' to 3' direction using the DNA to which it is hybridized as a template, i.e. conditions sufficient for enzymatic production of primer extension product. To achieve polymerization conditions, the temperature of the reaction mixture may be raised to or maintained at a temperature ranging from 65 to 75, such as from about 68 to 72QC and maintained for a period of time ranging from 15 sec to 20 min, such as from 20 sec to 5 min. In some embodiments, the annealing stage could be avoided, and protocol could include only denaturation and polymerization steps as described above. The above cycles of denaturation, annealing and polymerization may be performed using an automated device, typically known asAtty. Docket No.: CLCT-011 WO

[0178] a thermal cycler. Thermal cyclers that may be employed are described in U.S. Pat. Nos.

[0179] 5,612,473; 5,602,756; 5,538,871 ; and 5,475,610, the disclosures of which are herein incorporated by reference.

[0180] The product amplicon composition of this first amplification reaction will include amplicons corresponding to the gene specific domains that are present in the initial nucleic acid sample and are bounded by primer pairs present in the employed set of gene specific primers and RBG from one side of the amplicon. In some instances, the number of distinct amplicons of differing sequence in this initial amplicon composition ranges from 10 to 10,000,000, 20 to 1,000,000, 50 to 500,000, including 100 to 400,00 and 100 to 300,000, where in some instances the number of distinct amplicons present in this initial amplicon composition is 25 or more, 500 or more, 5,000 or more, 50,000 or more, 500,000 or more, 1 ,000,000 or more. A subject amplicon composition may include or exclude multiple different product amplicons corresponding to same gene as amplified by two or more different primer pairs directed to the gene. The multiple product amplicons making up the amplicon composition may vary in length, ranging in length in some instances from 50 to 1000, such as 60 to 800, including 80 to 700 nt.

[0181] The amplicon composition may be employed in a variety of different applications, including evaluation of the expression profile of the sample from which the RNA template target nucleic acid was obtained. In such instances, the expression profile may be obtained from the amplicon composition using any convenient protocol, such as but not limited to differential gene expression analysis, array-based gene expression analysis, NGS sequencing, etc.

[0182] In some cases, the amplicon composition may be employed to measure copy number of cells based on representation of different amplicons and copy number of DNA molecules used in multiplex PCR assay.

[0183] In some embodiments, the method further includes sequencing the multiple barcoded product amplicons, e.g., by using a Next Generation Sequencing (NGS) protocol. In such instances, if not already present, the methods may include modifying the initial amplicon composition to include one or more components employed in a given NGS protocol, e.g., sequencing platform adapter constructs, indexing domains, clustering domains, etc.

[0184] By “sequencing platform adapter construct” is meant a nucleic acid construct that includes at least a portion of a nucleic acid domain (e.g., a sequencing platform adapter nucleic acid sequence) or complement thereof utilized by a sequencing platform of interest, such as a sequencing platform provided by Illumina® (e.g., the NovaSeq™, NextSeq™, HiSeq™, MiSeq™ and / or Genome Analyzer™ sequencing systems); Thermo Fisher (e.g., Ion Torrent™ (such asAtty. Docket No.: CLCT-011 WO

[0185] the Ion PGM™ and / or Ion Proton™ sequencing systems) and Life Technologies™ ( such as a SOLiD sequencing system)); Pacific Biosciences (e.g., the PACBIO RS II sequencing system); Roche (e.g., the 454 GS FLX+ and / or GS Junior sequencing systems); Oxford Nanopore technologies (e.g., MinlON™, GridlON™, PrometlON™ sequencing systems) or any other sequencing platform of interest.

[0186] In certain aspects, the sequencing platform adapter construct includes a nucleic acid domain selected from: a domain (e.g., a “capture site” or “capture sequence”) that specifically binds to a surface-attached sequencing platform oligonucleotide (e.g., the P5 / i5 or P7 / i7 oligonucleotides attached to the surface of a flow cell in an Illumina® sequencing system); where the construct may include one or more additional domains, such as but not limited to: a sequencing primer binding domain or clustering domain (e.g., a domain to which the Read 1 or Read 2 primers of the Illumina® platform may bind); a indexing domain (e.g., a domain that uniquely identifies the sample source of the nucleic acid being sequenced to enable sample multiplexing by marking every molecule from a given sample with a specific index or “tag”), a replicate barcode (RBG) domain , and a barcode sequencing primer binding domain (a domain to which a primer used for sequencing a barcode binds).

[0187] The sequencing platform adapter constructs may include nucleic acid domains (e.g., “sequencing adapters”) of any length and sequence suitable for the sequencing platform of interest. In certain aspects, the nucleic acid domains are from 4 to 200 nucleotides in length. For example, the nucleic acid domains may be from 4 to 100 nucleotides in length, such as from 6 to 75, from 8 to 50, or from 10 to 40 nucleotides in length. According to certain embodiments, the sequencing platform adapter construct includes a nucleic acid domain that is from 2 to 8 nucleotides in length, such as from 9 to 15, from 16-22, from 23-29, or from 30-36 nucleotides in length.

[0188] The nucleic acid domains may have a length and sequence that enables a polynucleotide (e.g., an oligonucleotide) employed by the sequencing platform of interest to specifically bind to the nucleic acid domain, e.g., for solid phase amplification and / or sequencing by synthesis of the cDNA insert flanked by the nucleic acid domains. Example nucleic acid domains include the P5 (5’-AATGATACGGCGACCACCGA-3’) (SEQ ID NO:02), P7 (5’-CAAGCAGAAGACGGCATACGAGAT-3’)(SEQ ID NO:03), Read 1 primer (5’-ACACTCTTTCCCTACACGACGGTCTTCGGATCT-3’) (SEQ ID NO:04) and Read 2 primer (5’-GTGACTGGAGTTCAGAGGTGTGCTGTTCCGATCT-3’) (SEQ ID NQ:05) domains employed on the lllumina®-based sequencing platforms. Other example nucleic acid domains include theAtty. Docket No.: CLCT-011 WO

[0189] A adapter (5’-CCATCTCATCGCTGCGTGTGTCCGACTCAG-3’)(SEQ ID NO:06) and P1 adapter (5’-CCTCTCTATGGGCAGTCGGTGAT-3’)(SEQ ID NO:07) domains employed on the Ion TorrentTM-based sequencing platforms.

[0190] The nucleotide sequences of nucleic acid domains useful for sequencing on a sequencing platform of interest may vary and / or change over time. Adapter sequences are typically provided by the manufacturer of the sequencing platform (e.g., in technical documents provided with the sequencing system and / or available on the manufacturer’s website). Based on such information, the sequence of the sequencing platform adapter construct of the template switch oligonucleotide (and optionally, a first strand synthesis primer, amplification primers, and / or the like) may be designed to include all or a portion of one or more nucleic acid domains in a configuration that enables sequencing the nucleic acid insert (corresponding to the template nucleic acid) on the platform of interest.

[0191] The sequencing adapters may be added to the amplicons of the initial amplicon composition using, e.g., an amplification protocol. In such instances, the initial amplicon composition may be combined with forward and reverse sequencing adapter primers that include one or more sequencing adapter domains, e.g., as described above, as well as domains that bind to universal primer sites found in all of the amplicons in the composition, e.g., the forward and reverse anchor domains, such as described above. As reviewed above, amplification conditions may include the addition of forward and reverse sequencing adapter primers configured to bind to the common forward and reverse anchor domains and thereby amplify all or a desired portion of the product nucleic acid, dNTPs, and a polymerase suitable for effecting the amplification (e.g., a thermostable polymerase for polymerase chain reaction), where examples of such conditions are further described above. The forward and reverse sequencing adapter primers employed in these embodiments may vary in length, ranging in length in some instances from 20 to 60 nt, such as 25 to 50 nt. Addition of NGS sequencing adapters results in the production of a composition which is configured for sequencing by an NGS sequencing protocol, i.e., an NGS library.

[0192]

[0193] Following prescribed library (e.g., NGS library) preparation and / or amplification steps, e.g., as described above, prepared libraries may be considered ready for sequencing. In some embodiments, the method further includes sequencing the library (e.g., NGS library). In certain embodiments, the methods provided may further include subjecting an NGS library to an NGSAtty. Docket No.: CLCT-011 WO

[0194] protocol. The protocol may be carried out on any suitable NGS sequencing platform. NGS sequencing platforms of interest include, but are not limited to, a sequencing platform provided by Illumina® (e.g., the iSeq™, MiniSeq™, HiSeq™, MiSeq™, NextSeq™, NovaSeqTMand / or Genome Analyzer™ sequencing systems); Ion Torrent™ (e.g., the Ion PGM™ and / or Ion Proton™ sequencing systems); Pacific Biosciences (e.g., the PACBIO RS II Sequel sequencing system); Oxford Nanopore Technologies (ONT); Life Technologies™ (e.g., a SOLiD sequencing system); Roche (e.g., the 454 GS FLX+ and / or GS Junior sequencing systems); Element AVITI systems; Complete Genomics DNBSEQ Platforms (e.g., E25, G99, G400 and T7); or any other sequencing platform of interest. The NGS protocol will vary depending on the particular NGS sequencing system employed. Detailed protocols for sequencing an NGS library, e.g., which may include further amplification (e.g., solid-phase amplification), sequencing the amplicons, and analyzing the sequencing data are available from the manufacturer of the NGS sequencing system employed.

[0195] Other variations include, e.g., replacing lllumina®-specific sequencing domains in the various primers / oligonucleotides with sequencing domains required by sequencing systems from, e.g., Ion Torrent™ (e.g., the Ion PGM™ and Ion Proton™ sequencing systems); Pacific Biosciences (e.g., the PACBIO RS II sequencing system); Life Technologies™ (e.g., a SOLiD sequencing system); Roche (e.g., the 454 GS FLX+ and GS Junior sequencing systems); or any other sequencing platform of interest.

[0196] In some embodiments, the methods described herein are used in combination with other kits (e.g., the 10x Genomics Chromium V(D)J assay, the Cellecta DriverMap™ AIR profiling assay, and the Cellecta DriverMap™ EXP profiling Assay.

[0197] UTILITY

[0198] Methods of producing amplified next generation sequencing (NGS) nucleic acid libraries in accordance with the invention, e.g., as described above, may be employed to produce NGS libraries for a variety of different purposes. The subject methods include at least one replicate set of primers that include a replicate barcode (RBC) domain. As such, the RBC domain is not a unique molecular identifier (UMI).

[0199] The methods of the present application find use in producing accurate, high-throughput NGS sequencing. RBC domains allow for production of internal replicates during amplification, allowing for the application of conventional statistical tools to calculate accuracy in measuring of each specific target nucleic sequences. Therefore, RBC technology provides a new innovativeAtty. Docket No.: CLCT-011 WO

[0200] strategy for quantitative analysis of abundance level of different clonotypes, gene expression level, and copy number of specific genes or cells. Unlike methods that use UMIs, RBC technology, which employs a limited set of specific barcode sequences, does not require complex barcode error correction data analysis methods, and thus allows for simpler and more accurate sequencing analysis. In addition, using gene specific primers with an RBC instead of a UMI allows for a significantly improved yield of amplified target nucleic acid products (at least 2-4 fold) and reduces the background of non-specific products (e.g., primer dimers). As a result, RBCs improve sensitivity and specificity of multiplex PCR assays. Further, RBC technology allows for the correction of mutations arising from PCR to be differentiated from true sequencing reads of correct, non-mutated amplified nucleic acid sequences (e.g., genes or clonotypes). Analysis of reads number for different classes of highly abundant, medium abundant and low abundant clonotypes identified by RBC technology allows for effective normalization of sequencing to the number of template molecules used in multiplex PCR assay.

[0201] The methods of the present application are compatible with a wide variety of sample types. As such, the amplified NGS library may be utilized in various downstream analyses and, in some instances, the preparation of the library may be specifically reconfigured for a desired type of downstream analysis.

[0202] In some embodiments, the methods of the present application find use in adaptive immune receptor (AIR) repertoire profiling (e.g., quantitative analysis of a set of immune receptor clonotypes in a sample). The methods of the present application find use in both AIR RNA repertoire profiling and AIR DNA repertoire profiling. AIR profiling is a powerful tool for characterizing the adaptive immune system’s responses to cancer, autoimmune and infectious diseases, allergies, vaccinations, and therapeutic treatments. The unique sequences of the T-cell and B-cell receptors (TCRs and BCRs) and antibody variable regions (e.g., CDR3) that recognize foreign antigens define the individual differences in adaptive immune responses. Profiling the TCR and BCR variable regions using RT-PCR and NGS provides critical data for the discovery of novel, disease-associated immunity biomarkers and development of novel disease-specific diagnostic assays.

[0203] KITS, COMPOSITIONS AND DEVICES

[0204] Aspects of the present disclosure also include compositions and kits as well as devices for use therewith or therein.Atty. Docket No.: CLCT-011 WO

[0205] Most generally, the term "kit" is used to describe any assemblage of articles that facilitate the execution of a process, method, assay, analysis, manipulation of a sample, or the like. Kits can contain written instructions describing how to use the kit (e.g., instructions describing the methods of the present invention), chemical reagents or enzymes required for the method, primers, probes, buffer solutions, any type of containers (for example, containers for sample collection or sample manipulation) or reaction vessels, or any other components. A kit need not contain every component necessary to execute a method of the invention. The compositions and kits of the invention may include, e.g., one or more of any of the reaction components described above with respect to the subject methods.

[0206] In some embodiments, kits of the invention include (a) a multiplex collection of nucleic acid primers that includes at least one replicate set of primers, where the replicate set of primers includes primers having a common domain and a different replicate barcode (RBC) domain, and (b) a DNA polymerase. DNA polymerases include those recited above. In some cases, the DNA polymerase may be an RNA-dependent DNA polymerase or a DNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a polymerase which has templateswitching properties. In some embodiments, the number of primers in each replicate set of primers ranges from 3 to 24. In some embodiments, the number of primers in each replicate set of primers ranges from 4 to 8. In some embodiments, the number of primers in each replicate set of primers is 8. In some embodiments, the primers of the multiplex collection of nucleic acid primers do not include unique molecular identifier (UMI) domains. In some embodiments, the multiplex collection of nucleic acid primers includes a multiplex collection of gene specific primers. In some embodiments, the multiplex collection of nucleic acid primers includes at least one replicate set of primers in the form of template switch oligonucleotides (TSOs), where the TSOs in each replicate set have a common TSO domain and a different RBC domain.

[0207] In some embodiments, the multiplex collection of gene specific primers includes a forward primer set and a reverse primer set, where one or both forward and reverse primer sets includes at least one replicate set of primers, where the replicate set of primers includes primers having a common gene specific primer (GSP) domain and a different RBC domain. In some embodiments, the same collection of RBCs is present in each replicate set of primers. In some embodiments, different collections of RBCs are present in each replicate set of primers. In some embodiments, the reverse primer set includes at least one replicate set of primers. In some embodiments, the forward primer set includes at least one replicate set of primers. In some embodiments, the primers of the forward and reverse primer subsets further include anchorAtty. Docket No.: CLCT-011 WO

[0208] domains. In some embodiments, the multiplex collection includes a different number of forward and reverse primers. In some embodiments, the multiplex collection includes a greater number of forward primers than reverse primers.

[0209] In some embodiments, the multiplex collection is configured to generate template extension products from adaptive immune receptor nucleic acids. In some embodiments, the adaptive immune receptor nucleic acids include RNA or DNA of T cell receptor (TCR) or B cell receptor (BCR) genes. In some embodiments, the multiplex collection is configured to generate template extension products from CDR3 or GDR1-CDR2-CDR3 portion of TCR genes TRA, TRB, TRD, TRG and / or BCR genes IGH, IGK, and IGL. In some embodiments, full-length UTR-CDR1-CDR2-CDR3 template extension products are generated for both TCR and BCR genes using TSOs and RNA as a template.

[0210] The kits may further include one or more additional reagents employed in embodiments of the invention, e.g., as described above, where such reagents may include, but are not limited to: ligases, transposases, nucleases, polymerases, primers, buffers (e.g., hybridization buffers or buffers necessary for enzymatic reactions), dNTPs (including e.g., dATP, dCTP, dGTP, dTTP, dUTP, etc. or any one or any combination thereof), and the like. The subject kits may include, or the compositions and devices may be provided with, one or more test reagents, including e.g., control nucleic acids (e.g., control nucleic acid templates), and the like. In some instances, the reagents may be provided in lyophilized form, such as lyophilized enzymes, e.g., lyophilized reverse transcriptase, lyophilized DNA polymerase, etc.

[0211] Compositions of the present invention include compositions including primers that share a common domain and differ from each other by a replicate barcode (RBC) domain. Also provided are NGS libraries, where the NGS library includes amplified nucleic acids that include a domain derived from redundant subsets of primers which share a common domain and differ from each other by a replicate barcode (RBC) domain.

[0212] In some instances, components of the subject kits and / or compositions may be presented as a “cocktail” where, as used herein, a cocktail refers to a collection or combination of two or more different but similar components in a single vessel. Components of the kits may be present in separate containers, or multiple components may be present in a single container, as desired. The subject compositions may be present in any suitable environment. According to one embodiment, the composition is present in a reaction tube (e.g., a 0.2 mL tube, a 0.5 mL tube, a 1.5 mL tube, a 2 mL tube or the like), multiwell plate or a well or microfluidic chamber or droplet or other suitable container. In certain aspects, the composition is present in two or moreAtty. Docket No.: CLCT-011 WO

[0213] (e.g., a plurality of) reaction tubes or wells (e.g., a plate, such as a 96-well plate, a multi-well plate, e.g., containing about 1000, 5000, or 10,000 or more wells). The tubes and / or plates may be made of any suitable material, e.g., polypropylene, or the like, PDMS, or aluminum. The containers may also be treated to reduce adsorption of nucleic acids to the walls of the container. In certain aspects, the tubes and / or plates in which the composition is present provide for efficient heat transfer to the composition (e.g., when placed in a heat block, water bath, thermocycler, and / or the like), so that the temperature of the composition may be altered within a short period of time, e.g., as necessary for a particular enzymatic reaction to occur. According to certain embodiments, the composition is present in a thin-walled polypropylene tube, or a plate having thin-walled polypropylene wells or materials such as aluminum having high heat conductance.

[0214] In some instances, the components or reagents of the kit may be provided in a collection of individual vessels (e.g., separate tubes) or multiple vessels, e.g., a multi-well device, in which the components or reagents may be provided in liquid or dried form.

[0215] Any suitable reaction vessel(s) may be employed in the subject kits or devices and / or to contain a subject composition. Useful reaction vessels include but are not limited to e.g., tubes (e.g., single tubes, multi-tube strips, etc.), wells (e.g., of a multi-well plate (e.g., a 96-well plate, 384 well plate, or a plate with any number of wells such as 2000, 4000, 6000, or 10000 or more). Multi-well plates may be independent or may be part of a chip and / or device, e.g., as described in greater detail below. As such, in certain embodiments, the reaction vessel employed is a well or wells of a multi-well device. The present disclosure is not limited by the type of multi-well devices (e.g., plates or chips) employed. In general, such devices have a plurality of wells that contain, or are dimensioned to contain, liquid (e.g., liquid that is trapped in the wells such that gravity alone cannot make the liquid flow out of the wells). One exemplary chip is the 5184-well SMARTGHIP™ (Takara Bio USA, San Jose CA). Other exemplary chips are provided in U.S. Patents 8,252,581 ; 7,833,709; and 7,547,556, all of which are herein incorporated by reference in their entireties including, for example, for the teaching of chips, wells, thermocycling conditions, and associated reagents used therein). Other exemplary chips include the OPENARRAY™ plates used in the QUANTSTUDIO™ real-time PCR system (sold by Applied Biosystems). Another exemplary multi-well device is a 96-well or 384-well plate.

[0216] In addition to the above-mentioned components, a subject kit may further include instructions for using the components of the kit, e.g., to practice the subject methods as described above. The instructions are generally recorded on a suitable recording medium. TheAtty. Docket No.: CLCT-011 WO

[0217] instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or sub-packaging) etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g., portable flash drive, CD-ROM, diskette, Hard Disk Drive (HDD) etc. In yet other embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g., via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate.

[0218] The following examples are offered by way of illustration and not by way of limitation.

[0219] EXPERIMENTAL

[0220] Example 1: Generation of NGS AIR RBC and AIR UMl Amplified Libraries

[0221] To enhance the quantification, sensitivity, specificity and accuracy of adaptive immune receptor (AIR) profiling, an approach employing replicate barcodes (RBCs) has been developed. These consist of a replicate set of eight TCR / BCR gene-specific primers (GSPs), where such sets are designed against all functional, known TCR / BCR gene isoforms (clonotypes) and tagged with a unique six-nucleotide barcode for both AIR-RNA and AIR-DNA assays (FIG. 5B). The RBCs are incorporated into cDNA / DNA during primer extension to create internal replicates, which are then amplified during PCR. The subsequent amplification steps yield eight replicate NGS libraries, each labeled with a distinct RBC. This design allows quantitative analysis of each receptor clonotype sequence while facilitating the identification and exclusion of mutated sequence variants that may arise during PCR and NGS (error correction). Furthermore, AIR RBC technology is designed for straightforward normalization of NGS reads to a number of template immune receptor molecules used in assay. A schematic of the workflow is shown in FIG. 5A.

[0222] An example of a final indexed NGS amplicon structure for the AIR assay with RBCs is shown in FIG. 5C. The sequences shown in FIG. 5C are listed in Table 1 A below.

[0223]

[0224] Atty. Docket No.: CLCT-011 WO

[0225]

[0226] Table 1A. Sequences shown in FIG. 5C.

[0227] adapter

[0228] “NNNNNN” shown in Table 1 A can be eight different 6-nucleotide RBC barcodes with the sequences shown in Table 1B. As shown in FIG. 5C, SeqDNA-Fwd and SeqDNA-Rev are NGS primers for sequencing DNA inserts. SeqIND-Fwd and SeqIND-Rev are NGS primers for sequencing UDP 10-n indexes. FP5 and RP7 are Illumina adapters necessary for cluster generation in Illumina’s NGS flow cells, and GSP is the position of Forward (Fwd) and Reverse (Rev) GSPs.

[0229]

[0230] Table 1B. Sequences of RBCs.

[0231] For validation of AIR-RBC technology, both AIR-TOR and AIR-BCR for amplification of CDR3 or CDR1 -CDR2-CDR3 (full-length) were tested using different reverse (Rev) GSP primerAtty. Docket No.: CLCT-011 WO

[0232] sets designed with RBC or UMI tags. Table 1C below shows various types of nucleic acids that were tested.

[0233]

[0234] Table 1C. Nucleic acid sample types.

[0235] The bioinformatic analyses for RBC-based AIR assays may be carried out using a MiXGR pipeline. The first step involves read alignment to a reference gene database (e.g. the IMGT database). In addition, during this step, the pre-determined RBC sequence barcodes are detected. Each read is pre-allocated to a specific RBC prior to downstream analysis. The reads then undergo an error correction step. During this step, potential PCR or sequencing errors are detected via either of two approaches. The first approach is through accounting for the Phred quality score of reads whereby advanced correction techniques are used to correct for the sequences in low quality reads. The second approach is through the use of multi-layer clustering to distinguish between real mutations in the sample vs PCR or sequencing generated mutations.

[0236] After error correction, clonotypes are assembled for each sample for a specified region of the TCR or BCR (e.g. CDR3-only or FR1-C region). Reads with genes matching the specified region are identified and compiled into a core set of clonotypes. The output of this analysis is a tabular set of clonotypes with a description of the number of reads aligning to this clonotype, the V / D / J / C genes used, V / D / J / C gene alignment statistics, the nucleotide / amino acid sequence, etc.Atty. Docket No.: CLCT-011 WO

[0237] Note that for UMI-based assays, the UMIs themselves undergo error correction. This is because UMIs have great diversity and could possibly have a Hamming distance of a few nucleotides between unique UMI sequences. PCR or sequencing errors could easily result in incorrect UMI assignment. In contrast, the limited number of RBCs with a greater Hamming distance between the barcode sequences can easily be resolved. In UMI-based assays, the reads are flattened based on their UMIs whereby all reads with the same clonotype sequences and UMI sequence are counted as 1. In this case, UMI-based assays count the total number of uniquely identified molecules.

[0238] Comparative immune receptor repertoire profiling of RNA or DNA isolated from normal PBMC samples was carried out using a set of primers designed with eight 6-nucleotide RBCs or N14 UMI barcodes. In other words, each primer included either an RBC sequence or a UMI sequence. AIR-RBC technology demonstrates significantly improved yield (2-4-fold) of both CDR3 and CDR123 products for both AIR-RNA and AIR-DNA assays against AIR-UMI technology as illustrated in FIG. 10. These results clearly demonstrate benefits of using Rev GSPs with specific rather than random barcode sequences. Primers with UMI random tag sequences are less efficient in primer extension and follow-up amplification reaction due to nonspecific extension of different primers with random UMI sequences, problems with UMI-induced secondary primer structure and non-specific annealing of UMI-based primers to other template molecules. Moreover, the background level resulting from primer-dimers is significantly higher for AIR assay employed primers with the UMI design. Primer-dimers are mainly formed due to non-specific extension of forward GSPs annealed to random UMI sequences present in Rev GSP.

[0239] AIR RBC-based NGS data was processed with the MiXCR software package. The RBC-based NGS data in FIG. 6 shows successful alignment to reference immune receptor genes with MiXCR. FIG. 7 shows the performance of the different barcodes across different sample types. The number of reads associated with the 8 different barcodes for a particular NGS library is relatively similar across the different samples and specific AIR assay. This is consistent across different gene-specific primer sets. Sample names are denoted as: R_ = RNA, D_ = DNA, T- = T Cell, B- = B Cell (Human), mT- = T Cell, mB- = B Cell (Mouse), CDR3 = CDR3 amplicon, CDR123 = full length CDR1 -CDR2-CDR3 amplicon, and ##g = amount of starting material in either pg or ng.

[0240] Using eight internal replicates in the AIR assay allows the calculation of the reproducibility of quantitation of NGS read numbers for each clonotype. In FIGS. 8A and 8B,Atty. Docket No.: CLCT-011 WO

[0241] both the DNA and RNA based assays which use RBCs show lower coefficient of variation (CoV) for highly abundant clonotypes. As expected, the CoV increases at less abundant clonotypes with a smaller number of starting template molecules. Importantly, immune repertoire profiles for both RBC-based and UMI-based assays are correlated with each other for the medium-high abundant clonotypes. However, the RBC-based assay provides additional important information for data reproducibility, allowing for measuring the accuracy in quantifying different clonotypes.

[0242] There is inherent noise in generating NGS read counts coming from the profiling of low and medium abundant clonotypes. The dispersion in the data can be modeled relative to the number of read counts (FIG. 9).

[0243] Having eight RBCs allows for measuring the error rate and identifying mutated clonotypes characterized by low read number in comparison with correct sequence clonotypes. Furthermore, NGS analysis of the eight replicates using MiXCR software allows for normalization of the read numbers to the number of template molecules employed in the assay for each clonotype.

[0244] In summary, the RBC technology shows:

[0245] 1. Accurate clonotype quantitation-. Unlike UMI’s that use random nucleotide sequences, RBC’s use a limited number of unique, specific barcode sequences designed for each target TCR and BCR isoform that allows an accurate measure of the abundance level of each individual clonotype. These barcodes act as internal replicates in each AIR reaction and allow for application of conventional statistical tools to calculate accuracy in measuring of each specific clonotype abundance level. The barcodes further allow for normalization of the NGS read number to the number of template molecules employed in the assay, allowing for identification of rare, mutated clonotypes which need to be excluded from quantitative AIR analysis.

[0246] 2. Simplified Data Analysis'. Using a specific, unique RBCs allows simple deconvolution and binning of NGS reads for eight specific subsets associated with specific RBC. GSPs with UM I sequences do not have specific tag sequence and thus require a special, complex error-correction program for connecting NGS reads with specific template molecules due to deletions / mutations which are present in original UMI sequence or generated in amplification / NGS steps. Until now, no program exists which could effectively correct deletions in UMI sequences, and as a result UMI-based assay could artificially increase the number of detected clonotypes.Atty. Docket No.: CLCT-011 WO

[0247] 3. Error correction-. Straightforward data analysis allows a statistical measure to quantify the number of clonotypes based on the number of molecules. Statistical analysis of RBC-labeled replicates allows for normalization and quantitative assessment of clonotypes. Internal RBC replicate libraries facilitate accurate error correction.

[0248] 4. Increased Sensitivity due to Improved Design of Primers-. RevGSP with RBC are more efficient primers than primers with UMIs due to primer dimers, secondary structures etc. AIR-RBC technology increases sensitivity of amplification of clonotypes. There is a 1.5- 2X improvement in efficiency for AIR RNA assay and 2-4X in AIR-DNA assay compared to their UMI counterparts (FIG. 10).

[0249] 5. Increased clonotype detection-. Increasing in sensitivity of AIR-RBC assay allows an increased number of detected clonotypes in comparison with AIR-UMI assay (FIG. 11).

[0250] 6. Normalization and error correction-. As shown in FIGS. 12A and 12B, analysis of NGS read numbers for different high, medium and low abundant clonotypes allows for identifying the number of reads for clonotypes amplified from single template molecules. As a result, normalizing the NGS read number to single template molecule read number allows one to normalize / convert the NGS read number to the number of template molecules for all clonotypes identified in the AIR assay. Analysis of the distribution of NGS reads for different abundance clonotypes allows one to set up a cut-off (e.g., dashed lines in FIGS. 12A and 12B) for excluding clonotypes associated with mutations generated in the amplification and NGS steps. This error correction step allows for the generation of high quality AIR repertoire profiling data.

[0251]

[0252] The protocol below describes how to prepare NGS samples using AIR-RBC technology starting from total RNA from blood, tissue, cells or other biological samples, including Positive Control RNA.

[0253] A. List of reagents is shown in Table 2.

[0254] <

[0255]

[0256] Atty. Docket No.: CLCT-011 WO

[0257]

[0258] Table 2. List of reagents.Atty. Docket No.: CLCT-011 WO

[0259] B. Hybridization of mRNA with Reverse C-region Gene-Specific Primers

[0260] In this step, mRNA-Rev C GSP hybrids are generated and then purified from non-hybridized primers by nuclease treatment and AM Pure magnetic beads. The protocol is written assuming the reactions will be set off in a 96-well plate.

[0261] 1. Prepare a Hybridization Master Mix as described in Table 3 for all samples and controls (make 5% extra to account for pipetting error) and aliquot 7 pl in the wells of a 96-well plate:

[0262]

[0263] Table 3. Hybridization Master Mix

[0264] If running only TOR or BCR repertoire analysis, use only TCR-C or BCR-C primer mix and adjust the total volume to 7 pl by water. If running biological or amplification triplicates, increase the number of tubes, volume, and amount of RNA allocated for each sample respectively. For the DirectCell protocol, add 1 pl of 2% N-Lauroyl Sarcosine (final concentration in hybridization mix will be 0.1%). It is recommended to run positive control RNA (50 ng), and negative control (water) for troubleshooting and comparing batch effects in different experiments.

[0265] 2. Adjust the volume of each RNA Sample (e.g., 100 ng of whole blood or 50 ng of PBMC RNA) to 14 pl with water as shown in the Table 4.

[0266]

[0267] Table 4. Components of RNA sample.

[0268] Note: For tissue samples with a low content of immune cells, the recommended amount is 200-1,000 ng of total RNA). For AIR profiling directly in immune cell fractions, use 5,000-50,000 cells in 14 pl of 1xPBS buffer.Atty. Docket No.: CLCT-011 WO

[0269] 3. Add 14 pl of RNA Sample to each well with pre-aliquoted 7 pl of Hybridization Master Mix and mix contents by pipetting 3 times. Seal the plate with adhesive film, and spin down to collect droplets. Load the plate in the thermal cycler, and run the program in Table 5 to hybridize mRNAs with Reverse TCR / BCR-C primers:

[0270]

[0271] Table 5. Program.

[0272] 4. Add 24 pl (1 ,2x volume) of Agencourt AMPure® XP Reagent (adjusted to room temperature) to each reaction well, and pipet up and down 5 times to thoroughly mix the bead suspension with the hybridization reaction mix. Check that the whole volume in each well has a uniform brown color.

[0273] 5. Incubate the mixture for 5 minutes at room temperature. While waiting, prepare the Reverse Transcriptase Buffer Master Mix based on protocol below. Store on ice before use.

[0274] 6. Place the plate in the Magnetic Stand for 96-well plates for 1 -2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0275] 7. Add 200 pl of freshly prepared 80% ethanol in each reaction well without removal of plate from Magnetic Stand, wait for 2 minutes, carefully remove, and discard the supernatant without disturbing the bead pellets.

[0276] 8. Repeat 80% ethanol washing in step 9.

[0277] 9. Briefly centrifuge the plate at low speed and place the plate in the Magnetic Stand. Use a 20-pl pipette to remove the residual ethanol droplets from reaction wells and air-dry the beads at room temperature for approximately 5 minutes.

[0278] C. cDNA Synthesis

[0279] In this step, the purified mRNA-Reverse GS primer hybrids eluted from AMPure beads are extended by Reverse Transcriptase to generate cDNA (antisense strand of mRNA).Atty. Docket No.: CLCT-011 WO

[0280] 1. Prepare the Reverse Transcriptase Buffer Master Mix as shown in Table 6 for each sample plus 5% extra volume of all components:

[0281]

[0282] Table 6. RT Buffer Master Mix.

[0283] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Add 22 pl of the RT Buffer Master Mix in each reaction well and resuspend AMPure beads attached to the well surface by pipetting or in the plate using an Eppendorf shaker. Briefly centrifuge the plate at low speed and place the plate in Magnetic Stand for 1 minute. Transfer 20 pl of clear supernatant without beads from each reaction well to the new plate and seal the plate.

[0284] 3. Load the plate in the thermal cycler and start running the program in Table 7:

[0285]

[0286] Table 7. Program.

[0287] D. Forward Gene-Specific Primer Extension

[0288] In this step, the pool of Forward Gene-Specific Primers with adjoining Anchor 1 sequences generates sense strands of the target amplicons flanked from both sides by Anchor 1 and Anchor 2 sequences using the cDNA generated in the previous step as a template and purified from non-extended primers by nuclease treatment.

[0289] 1. Prepare the Forward GS Primer Extension Master Mix as shown in Table 8 for all samples plus 5% extra volume of all components:

[0290]

[0291] Atty. Docket No.: CLCT-011 WO

[0292]

[0293] Table 8. Forward GS Primer Extension Master Mix.

[0294] Note: If running only TCR or BCR repertoire profiling, use only a single corresponding forward TCR or BCR FR3 / FR1 primer mix and adjust by water total volume to 20 pL.

[0295] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Spin down the cDNA plate, remove the seal, then add 20 pl of the Forward GS Primer Extension Master Mix to each reaction well of the plate as shown in Table 9.

[0296]

[0297] Table 9. Components.

[0298] 3. Mix contents by pipetting three times, seal the plate with a new adhesive film, and spin down to collect droplets.

[0299] 4. Load the plate in the thermal cycler, and run the program in Table 10:

[0300]

[0301] Table 10. Program.

[0302] 5. Spin down the plate, remove the seal from the plate, then add 2 pl of the Primer Removal Enzyme to each reaction well of the plate. Mix contents by pipetting 3 times, seal the plate, and spin down to collect droplets.

[0303] 6. Load the plate in the thermal cycler, and run the program in Table 11 :Atty. Docket No.: CLCT-011 WO

[0304]

[0305] Table 11. Program.

[0306] A. First PCR with Anchor Primers

[0307] This step utilizes universal Anchor PCR primers to amplify the target cDNA fragments flanked with the Anchor 1 and Anchor 2 sequences generated during the previous Forward Gene-Specific Primer Extension step.

[0308] 1. Prepare the Anchor PCR Master Mix as shown in Table 12 for all samples and controls plus 5% extra volume of all components:

[0309]

[0310] Table 12. Anchor PCR Master Mix.

[0311] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Spin down the Forward GS Primer Extension plate, remove the seal, then add 60 pl of Anchor PCR Master Mix to each reaction well as shown in Table 13.

[0312] Component Volume per sample, l

[0313]

[0314] Forward GS Primer Extension DNA (after primer removal step) 42 (Anchor PCR Master Mix (prepared above) 60 (Total

[0315]

[0316] 102

[0317]

[0318] Table 13. Components.

[0319] 3. Mix content by pipetting 3 times. Seal the plate with new adhesive film and spin down to collect droplets.Atty. Docket No.: CLCT-011 WO

[0320] 4. Load the plate in the thermal cycler in a location dedicated to PGR work. Run the program shown in Table 14 using the recommended number of PGR cycles:

[0321]

[0322] Table 14. Program.

[0323] Note: To avoid bias in gene expression levels by over-cycling samples, it is recommended to start with 18 PCR cycles for samples that contain 25 ng or more RNA (e.g., 50 ng of PBMC or 100 ng of whole blood or lymphocyte-rich samples) or at least 25,000 immune cells. For samples with less RNA or cells (or low lymphocyte RNA content), extra cycles may be added, but in general, it is not recommended to exceed 20 cycles. For very small RNA samples (1-5 ng) or experiments with less than -1000 immune cells, 22-26 cycles may be required. The recommended number of cycles needs to be optimized and adjusted based on specific cell types, sample types, RNA quality, and the level of TCR / BCR mRNAs, etc. In order to optimize cycle number, it is recommended to analyze 5 ul of amplified products in a 3% agarose gel or fragment analyzer. For optimal cycle number, weak cDNA amplification products in the range of 300-500 bp will be detected. If no products are seen, add 3 more cycles and analyze the yield of PCR products again.

[0324] B. Second PCR with Indexed Primers

[0325] This step adds a dual unique DNA / RNA UDP index combination to each Anchored PCR Product generated in the previous PCR with Anchor Primers step as well as universal flanking P5 and P7 sequences needed for cluster formation on the Illumina NGS flow cell. The Index PCR Plate in the kit contains a unique combination of Forward and Reverse DNA / RNA UDP index primers in each well. The primers have been dried onto the bottom of each well and will be dissolved when the PCR reaction mix with the sample is added. One well should be used for each sample (triplicate samples are different samples) being sequenced.Atty. Docket No.: CLCT-011 WO

[0326] 1. Prepare enough of the Index PCR Master Mix, following the formulation in Table 15 for all samples and controls plus 5% extra volume of all components:

[0327]

[0328] Table 15. Index PCR Master Mix.

[0329] 2. Gently vortex the Index PCR Master Mix, and spin down briefly to collect droplets. Remove the plate seal from the Index PCR Plate. Set up the Index Primer PCR Reactions as follows.

[0330] 3. Aliquot 50 pl of the Index PCR Master Mix into appropriate wells of the 96-well Index PCR Plate (or cut-off portion of the plate) provided in the kit. To avoid index- to-index contamination, add Index PCR Master Mix using a new tip for each well.

[0331] 4. Spin down, then remove the seal from the Anchor Primer PCR plate (plate from PCR with Anchor Primers step). Transfer 2 pl of Anchored PCR products to each of the Index Primer PCR reactions on the Index PCR Plate (see Table 16). To avoid mistakes, ensure that samples in the Anchor Primer PCR plate are arranged in the same format as the Index PCR Primer pair mixes in the Index PCR Plate (e.g. Sample 1 A is aliquoted to well 1A). It is recommended to record the sample name and well number (e.g., Sample 1 in well 1A) for all the samples, including the positive control. This will help minimize mistakes in the NGS deconvolution step.

[0332]

[0333] Table 16. Components.

[0334] 5. Seal the plate with new adhesive film and spin down to collect droplets.

[0335] 6. Load the plate in the thermal cycler, and run the program in Table 17.Atty. Docket No.: CLCT-011 WO

[0336]

[0337] Table 17. Program.

[0338] C. NGS Prep and Sequencing

[0339] The product from the PGR with Indexed Primers step contains both P5 and P7 sequences and dual DNA / RNA UDP Indexes for Next-Gen Sequencing (NGS) on Illumina instruments. The protocol in this section provides the instructions to normalize the amount of each Amplified Indexed Library to obtain similar reads for each sample, to combine and clean up samples before loading onto the instrument, and guidelines for sequencing and data analysis.

[0340] QC, Quantify and Combine Samples for NGS

[0341] In this step, the yield of products from the Index Primer PGR Reaction — the Amplified Indexed Libraries of transcripts from each sample — are analyzed (adjusted if necessary), measured, and then pooled in equimolar amounts for sequencing.

[0342] 1. Analyze the Amplified Indexed Libraries using one of the following methods:

[0343] • Standard Method: Separate 5 pl of Amplified Indexed Libraries on a 3% agarose-TAE gel and analyze the size distribution of NGS probes by UV transilluminator. To minimize the sample number, one sample from each triplicate set may be run. See FIG. 15 for the expected results of amplified libraries generated from good-quality whole blood RNA samples. For AIR TCR-BCR assay, the smear with several bright bands should be in the 220-420 bp range.

[0344] • Alternative Method: Analyze 1 pl of each of the Amplified Indexed Libraries on either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol.Atty. Docket No.: CLCT-011 WO

[0345] 2. Analyze yields of the Amplified Indexed Libraries. The yield should be roughly the same for all experimental samples within + / - 2-3-fold levels and similar to the Positive Control RNA sample. Negative control sample should not generate any significant yield of amplified products. If some samples show a significantly lower yield of amplification products, it could indicate differences in the amount, quality of RNA, or content of TCR / BCR mRNAs used in AIR assay. For the experimental RNA samples with a significantly lower yield of PCR product (e.g., >5-10-fold) than other samples or positive control RNA, the lower yield samples may be re-run in the thermal cycler for 2-5 additional cycles.

[0346] Remove Excess PCR Primers and Combine Samples

[0347] 1. Remove excess primers from the completed PCR reactions by adding 1 pl of Primer Removal Enzyme to each of the Amplified Indexed Libraries and the Negative Control sample, then incubate at 37°C for 30 minutes.

[0348] 2. To ensure accurate quantification for sequencing, the quantification procedure of the Amplified Indexed Libraries and the Positive Control RNA should be repeated. The preferred method is to analyze 2 pl of each of the Amplified Indexed Libraries using either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol. Quantifying the PCR products after removing PCR primers is more accurate than quantifying before the primer removal clean-up step. (See, e.g., FIG. 16).

[0349] 3. After primer removal and quantification, use the yield assessment of the Amplified Indexed Libraries as a basis to combine equimolar amounts of each of the Amplified Index Libraries into a single pool for NGS. For example, if the yield of Library 1 is twice that of Library 2, then mix 5 pl of Library 1 with 10 pl of Library 2. To minimize sample-to- sample sequencing variations, combine and load all experimental samples onto one flow cell.

[0350] For the AIR-RBC Assay, which generates a complex pool of amplified TCR / BCR CDR regions (100K-200K from 50 ng of total whole blood, PBMC RNA), refer to Table 18 for guidelines on how many samples may be combined for different instruments and read depths. Generally, aim for 5-10 million reads per AIR sample (starting from 50 ng of PBMC RNA) or at least 20 readsAtty. Docket No.: CLCT-011 WO

[0351] per template molecule. For Illumina instruments, AIR profiling assays require 300-n paired-end reagent kits for CDR3 and 600-n paired-end reagent kit for CDR1-CDR2-CDR3.

[0352]

[0353] Table 18. Number of samples for instruments.

[0354] Purification & Quantification of Amplified Indexed Library

[0355] The purpose of this step is to remove any residual primers and reagents from the pooled Amplified Indexed Libraries so that the preparations are ready for NGS.

[0356] 1. Add 1.5x volume of Agencourt® AMPure® XP Reagent (at room temperature) to pooled Amplified Indexed Libraries, mix in an Eppendorf tube, and pipet up and down 5 times to thoroughly mix the bead suspension with the pooled Amplified Indexed Libraries.

[0357] 2. Incubate the mixture for 5 minutes at room temperature.

[0358] 3. Place the tube in the Magnetic Stand for 1 minute or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0359] 4. Add 500 pl of freshly prepared 80% ethanol to the tube.

[0360] 5. Place the tube in the Magnetic Stand for 2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0361] 6. Repeat steps 4 and 5 for a second wash.Atty. Docket No.: CLCT-011 WO

[0362] 7. Briefly centrifuge tubes at low speed and place the tubes in the Magnetic Stand. Use a 200-pl pipette to remove the residual ethanol droplets from the tube and air-dry the beads at room temperature for 2 minutes.

[0363] 8. Add 45 pl of TE buffer to the pellet to disperse the beads and let stand for 1 minute.

[0364] 9. Place the tube on the Magnetic Stand for 1 minute. Transfer 40 pl of the supernatant to a new Eppendorf tube.

[0365] 10. Measure the concentration of pooled Amplified Indexed Library sample using the Qubit dsDNA High Sensitivity Assay.

[0366] 11. Dilute the pooled Amplified Indexed Library sample to 1.7 ng / pl, which corresponds to a concentration of 10nM for the following NGS step.

[0367] D. Next Generation Sequencing

[0368] The Amplified AIR NGS Index Libraries made from each total RNA sample should be run on an Illumina sequencer following the manufacturer’s instructions. Generally, it is recommended to set up the sequencing reactions to generate 5-10 million reads per sample (read depth per sample). This depth works out to produce 20-50 reads, on average, per template molecule.

[0369] Note: The procedure below is based on the NextSeq 500 / 550 using the 300-cycle NextSeq 500 / 550 High Output Kit v2.5 (Illumina Cat.# 20024908). For other Illumina instruments, please contact Illumina.

[0370] Tech Support Team for optimal NGS instructions.

[0371] The multiplexing level may be modified to meet experimental needs. For example, multiplex fewer samples together to generate more sequencing reads per sample (i.e., more depth) for increased and more sensitive detection of clonotypes present across a broader dynamic range of abundance / expression levels, or, if mostly interested only in highly abundant / expressed clonotypes, more samples may be sequenced together which will be less expensive.Atty. Docket No.: CLCT-011 WO

[0372] Follow the standard Illumina procedures for Cluster Generation, starting with 10 nM of the purified PGR sample. It is recommended to add 15% of PhiX to the library.

[0373] 1. Add 6 pl of each of the custom sequencing primers into the appropriate wells of the Illumina reagent cartridge, as follows:

[0374] • Forward SeqDNA NGS Primer (Read 1 Sequencing Primer) into well #20

[0375] • Reverse SeqDNA NGS Primer (Read 2 Sequencing Primer) into well #21

[0376] • Forward SeqIND NGS Primer (Index 1 Sequencing Primer) into well #22

[0377] • Reverse SeqIND NGS Primer (Index 2 Sequencing Primer) into well #22

[0378] 2. Perform the NGS run using 300-nt paired-end reads or 600-nt paired-end reads on the NextSeq. Proceed to cluster amplification using the appropriate Illumina Paired-End Cluster Generation Kit; refer to the manufacturer’s instructions for this step. The optimal seeding concentration for cluster amplification of indexed libraries is approximately 1.8 pM for NextSeq500 / 550. Use the program in Table 19 for the sequence run (for NGS analysis of CDR3 / CDR1-CDR2-CDR3 NGS libraries):

[0379]

[0380] Table 19. Program.

[0381] Example 3: Generation of NGS TCR and BCR CDR1 -CDR2-CDR3 libraries from single-cells using single-cell AIR-RBC technology.

[0382] The protocol below describes how to prepare NGS samples using single-cell AIR-RBC technology starting from single cells sorted in 96 well plates.

[0383] A. List of Reagents is shown in Table 20.

[0384]

[0385] Atty. Docket No.: CLCT-011 WO

[0386]

[0387] Table 20. List of reagents.Atty. Docket No.: CLCT-011 WO

[0388] B. Protocol for Sorting T and B Cells

[0389] The scAIR Master Plates (scAIR TCR-Mark37 and scAIR BCR-Mark30) contain 50 pl per well of pre- aliquoted Hybridization buffer containing a mix of barcoded Reverse TCR (designed for alpha, beta, gamma, and delta) or BCR (designed or H, K and L chains) C-region primers and barcoded Reverse GSPs designed for T cell or B cell subtyping / phenotyping marker mRNAs. The scAIR Master Plates contain enough reagents for the preparation of up to eight 96-well sorting PCR plates.

[0390] i. Before use, thaw the Master Plates with Hybridization Master Mix, shake the plate in a plate shaker, and centrifuge the plate with a benchtop mini plate-spinning centrifuge or any bucket centrifuge (e.g., 2,000 rpm, 1 min) to collect droplets to the bottom of the wells.

[0391] ii. Remove the seal from a plate with the Hybridization Master Mix and transfer 5 pl of Hybridization Master Mix from each well of the Master plate into 1 to 8 new 96-well PCR plate(s) compatible with cell sorter. Importantly, use a new tip for each well of the Master plate to avoid cross-contamination of barcoded reverse GSPs between different wells. A Master plate with wells of unused Hybridization Master Mix can be stored at +4°C for up to 3 months.

[0392] ill. For sorting on the FACS sorter instrument: load the 96-well sorting PCR plate with 5 pl of pre-aliquoted Hybridization Master Mix into the motorized stage holder. Follow the single-cell sorting protocol as per the specific instrument to sort single cells for experimental samples. Depending on the experimental design, more than one cell may be sorted into some wells. For example, up to 500 bulk control cells may be sorted in one well as a strategy to identify the most abundant clonotypes and / or to use as a positive control to troubleshoot the protocol.

[0393] iv. Immediately after sorting, remove the plate, seal the plate, and use a benchtop mini plate-spinning centrifuge (or any other centrifuge) to spin the plate for 10-15 secs. This will propel any drops with cells that are on the side walls of the wells down into the buffer. Freeze and store the plate(s) at -80°C prior to use for the scAIR profiling experiment. To reduce sequencing cost, up to 8 plates (e.g., from different experiments) could be run together for follow-up scAIR-NGS protocol.

[0394] C. Hybridization of mRNA from Sorted Cells with Reverse Gene-Specific Primers

[0395] In this step, T / B cells sorted in the 96-well PCR plate(s) are incubated at high temperatures in a hybridization buffer containing Rev GSPs with well barcodes and 4x RBC. In this hybridization step, mRNA-Rev GSP hybrids are generated, and purified from non- hybridized primers by nuclease treatment, 96 barcoded mRNA-RevGSP hybrids are combined and purified withAtty. Docket No.: CLCT-011 WO

[0396] SPRIselect magnetic beads. If running more than one 96-well plate with sorted single T / B cells, each plate gives one pooled reaction tube, i.e. , don’t combine pooled samples from different plates. The protocol below specifies the volume of reagents for processing one 96-well plate with sorted single T / B cells. The same protocol could be used for processing the plate with spiked (in each well) scAIR TCR / BCR Positive Control RNA sample. In current and all follow-up steps proportionally increase the volume of reagents to assemble Master mixes if using more than one plate.

[0397] Please, note that the proposed volumes of Master Mixes for all steps already include 10-20% excess of reagents, e.g., it is not necessary to increase the volume of reagents to prepare Master Mixes.

[0398] i. Thaw the plate with sorted cells (stored at -80°C, Section 5.1) and briefly centrifuge the plate to collect droplets. Load the plate in the thermocycler, and run the program in Table 21 to hybridize mRNAs with Reverse GSPs:

[0399]

[0400] Table 21. Program.

[0401] ii. During the hybridization step, prepare the Primer Removal Master Mix as shown in Table 22.

[0402]

[0403] Table 22. Primer Removal Master Mix.

[0404] ill. Aliquot 66 pl of prepared Master Mix in each tube of the 8-tube strip, and store the 8-tube strip at +4°C:Atty. Docket No.: CLCT-011 WO

[0405] iv. After completion of the Hybridization step, using an 8-channel pipet add 5 pl of Primer Removal Master Mix into each plate well. Use new tips for each well to avoid contamination between samples. Seal the plate, mix the reagents in the plate using an Eppendorf plate shaker, and briefly spin the plate to collect the contents at the bottom of the wells. Incubate plate in a thermocycler at 37°C for 20 min.

[0406] v. During the Primer Removal step (Step 3), prepare Diluted SPRIselect beads (see Table 23) in the reservoir designed for an 8-channel pipet, mix well viscous beads and dilution buffer, and keep it at room temperature.

[0407]

[0408] Table 23. Amounts of beads.

[0409] vi. Add 10 pl (1x volume) of Diluted SPRIselect Beads into each plate well using an 8-channel pipet, seal the plate, and mix SPRIselect beads and solution in wells using Eppendorf plate shaker. Importantly, the speed of the plate shaker will need to be optimized to completely mix viscous SPRIselect beads with a solution in each well. Check visually that each well contains a homogeneous brown color solution in the whole volume.

[0410] vii. Shake the plate in the Eppendorf plate shaker at 1400 rpm for 10 minutes at room temperature.

[0411] viii. Using an 8-channel pipet combine the content of all wells of the 96-well plate together in a reservoir designed for an 8-channel pipet. Transfer the whole volume of the pooled sample into a 2 ml test tube. Pipet slowly and try to collect all viscous solutions from wells and reservoirs. Place the 2 ml test tube with the pooled sample in a magnetic separation device designed for 1.5-2 ml test tubes for approximately 5 min until the liquid appears completely clear. Carefully remove and discard the clear supernatant from the test tube without disturbing the magnetic bead pellet.

[0412] ix. Add 2 ml of freshly prepared 80% ethanol in the pooled sample test tube without removing the test tube from the Magnetic Stand, wait for 2 minutes, carefully remove, and discard the supernatant without disturbing the bead pellet.Atty. Docket No.: CLCT-011 WO

[0413] x. Repeat 80% ethanol washing in Step 8.

[0414] xi. Briefly centrifuge the pooled sample test tube at low speed and place the test tube in the Magnetic Stand. Use a 200-pl pipette to remove the residual ethanol droplets from the bottom of the test tube and air-dry the beads at room temperature for approximately 5 minutes until the disappearance of ethanol droplets on the test tube surface and the magnetic bead pellet appears dry and no longer shines.

[0415] D. cDNA Synthesis

[0416] In this step, the purified mRNA-Reverse GSP hybrids eluted from SPRIselect beads are extended by Reverse Transcriptase to generate cDNA (antisense strand of mRNA). The protocol below specifies volumes of reagents and Master Mixes for one pooled sample purified from one 96-well plate.

[0417] i. Prepare the Reverse Transcriptase Buffer Master Mix as shown in Table 24.

[0418]

[0419] Table 24. RT buffer master mix.

[0420] ii. Gently vortex Master Mix and spin down briefly to collect droplets. Add 27 pl of the RT Master Mix into each pooled sample 2 ml test tube with magnetic beads and gently resuspend SPRIselect beads attached to the tube surface by pipetting. Avoid the formation of bubbles. Briefly centrifuge the tube at low speed and place the tube in a Magnetic Stand for 1-2 minutes until the supernatant is clear. Transfer 25 pl of clear supernatant without beads from the tube to the new 0.2-0.5 ml test tube compatible with the used thermocycler.

[0421] iii. Load the test tube in the thermal cycler and start running the program in Table 25 (use standard 105°C heating lid temperature):Atty. Docket No.: CLCT-011 WO

[0422]

[0423] Table 25. Program

[0424] E. Forward Gene-Specific Primer Extension

[0425] In this step, the pool of Forward Gene-Specific Primers with adjoining Anchor 1 sequences is annealed to cDNA template generated in previous step, extended by DNA polymerase and generates sense strands of the target amplicons flanked from both sides by Anchor 1 and Anchor 2 sequences. The protocol below specifies volumes of reagents and Master Mixes for one pooled cDNA sample.

[0426] i. Prepare the Forward GS Primer Extension Master Mix as shown in Table 26.

[0427]

[0428] Table 26. Forward GS primer extension master mix.

[0429] ii. Gently vortex Master Mix and spin down briefly to collect droplets. Spin down the sample test tube and add 25 pl of the Forward GSP Extension Master Mix (see Table 27).

[0430]

[0431] Table 27. Components.

[0432] ill. Mix contents by pipetting three times, briefly spin down the tube, load the tube in the thermal cycler, and run the program in Table 28.

[0433]

[0434] Atty. Docket No.: CLCT-011 WO

[0435]

[0436] Table 28. Program.

[0437] iv. Add 2 pl of the Primer Removal Enzyme to the sample, mix contents by pipetting 3 times, spin down to collect droplets, load the sample test tube in the thermal cycler, and run the program in Table 29.

[0438]

[0439] Table 29. Program.

[0440] F. First PGR with Anchor Primers

[0441] This step utilizes universal Anchor PGR primers to amplify the target DNA fragments flanked with the Anchor 1 and Anchor 2 sequences generated during the previous Forward Gene-Specific Primer Extension step. The protocol below specified volumes of reagents and Master Mixes for one pooled DNA sample.

[0442] i. Prepare the Anchor PGR Master Mix as shown in Table 30.

[0443]

[0444] Table 30. Anchor PGR master mix.

[0445] ii. Gently vortex Master Mix and spin down briefly to collect droplets. Add 50 pl of Anchor PGR Master Mix to each sample test tube (see Table 31).

[0446] Volume per sample, pl

[0447]

[0448] Atty. Docket No.: CLCT-011 WO

[0449]

[0450] Table 31. Components.

[0451] iii. Mix content by pipetting 3 times and spin down to collect droplets.

[0452] iv. Load the plate in the thermal cycler in a location dedicated to PCR work. Run the program in Table 32 using the recommended number of PCR cycles:

[0453]

[0454] Table 32. Program

[0455] Note: For Pooled sorted single-cell sample run 26 cycles. For scAIR Positive Control RNA sample run 22 cycles. The number of cycles in 1st PCR in many cases needs to be optimized due to differences in transcriptional activity of sorted cells. E.g., activated cells are more transcriptionally active than naTve / dysfunctional cells. For plasma B cells, it is recommended to use 20 cycles.

[0456] G. Second PCR with Indexed Primers

[0457] This step adds a unique UDP dual index combination to each Anchored PCR Product generated in the previous PCR with Anchor Primers step as well as universal flanking P5 and P7 sequences needed for cluster formation on the Illumina NGS flow cell. The Index PCR Plate contains a complete set of 96 unique combinations of Forward and Reverse DNA / RNA UDP index primers in each well. The protocol below specifies the volumes of reagents and Master Mix for each sample which can be 1 to 8 Pooled samples and positive control sample. In order to avoid low complexity reads in index NGS reads, it is recommended to amplify all samples in triplicates using 3 different indexes.Atty. Docket No.: CLCT-011 WO

[0458] i. Prepare the PCR Master Mix, following the formulation in Table 33 for one sample, and increase proportionally if running more samples:

[0459]

[0460] Table 33. PCR master mix.

[0461] ii. Gently vortex the PCR Master Mix, and spin down briefly to collect droplets. Transfer 5 pl of Anchored PCR products after the first PCR to test tube with a PCR Master Mix (see Table 34).

[0462]

[0463] Table 34. Components.

[0464] ill. Mix the test tube contents by pipetting, and spin down to collect droplets. Remove the plate seal from the Index PCR Plate (provided in the kit), and aliquot 50 pl of the Anchored PCR product in PCR Master Mix into three different wells of the 96-well Index PCR Plate (or cut-off portion of the plate). It is recommended to record the sample / replicate name and well number (e.g., Sample 1 in well A1 , A2 and A3) for all the samples. This will help minimize mistakes in the NGS deconvolution step.

[0465] iv. Seal the plate (or portion of plate) with new adhesive film, and spin down to collect droplets. Load the plate in the thermal cycler, and run the program in Table 35.

[0466]

[0467] Atty. Docket No.: CLCT-011 WO

[0468] Table 35. Program.

[0469] H. QC, Quantify and Combine Samples for NGS

[0470] The product from the Second PCR with Indexed Primers step contains both P5 and P7 sequences and dual UDP Indexes designed for Next-Gen Sequencing (NGS) on Illumina instruments. The protocol in this section provides the instructions to analyze and adjust the amount of each Amplified Indexed Library to obtain similar reads for each sample, to combine and clean up samples before loading onto the instrument, and guidelines for sequencing and data analysis. In this step, the yield of products from the Index Primer PCR Reaction — the Amplified Indexed Libraries of transcripts from each sorted plate sample — is analyzed, measured, adjusted if necessary, and then pooled in equimolar amounts for sequencing.

[0471] Perform QC and Quantify Amplified Indexed Libraries

[0472] i. Analyze the Amplified Indexed Libraries using one of the following methods:

[0473] I . Standard Method: Separate 5 pl of Amplified Indexed Libraries on a 3% agarose-TAE gel and analyze the size distribution of NGS probes by UV transilluminator. See FIG. 17 for the expected results of the amplified library generated from scAIR positive control RNA (1 plate) after 2nd PCR step with indexed primers.

[0474] 2. Alternative Method: Analyze 1 pl of each of the Amplified Indexed Libraries on either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol. FIG. 18 shows the results of the analysis of amplified 2nd PCR products using Fragment Analyzer. For the scAIR TCR-Mark37 assay, T cell marker amplicons generate the smear with several bright bands in the 220-420 bp range and CDR1-2-3 TCR amplicons have a size of approximately 550 bp. For the scAIR BCR-Mark30 assay, B cell marker amplicons generate the smear with several bright bands in the 250-420 bp range and CDR1-CDR2-CDR3 BCR amplicons have a size of approximately 600 bp.

[0475] ii. Adjust Second PCR cycle number (optional step)

[0476] The yield of amplified products after the second amplification step significantly depends on the transcriptional activity of sorted cells. For example, activated T or B cells are moreAtty. Docket No.: CLCT-011 WO

[0477] transcriptionally active than naive cells. Amplified products generated from Positive Control RNA (included in the kit) could be a useful reference for analysis of PCR products and adjustment / troubleshooting of the scAIR protocol. If the analysis of amplified products didn’t reveal any visible yield of expected PCR products, run the sample with the amplified indexed library for an additional 3-5 PCR cycles and repeat the analysis. Optimize the number of cycles in 2nd PCR step to get visible PCR products in all experimental samples (generated from different plates). It is not necessary to achieve the same yield of PCR products for different amplified indexed libraries, differences in the yield between samples will be adjusted when the samples are combined in the next step.

[0478] 1. Quantify amplified indexed libraries using Fragment Analyzer / Bioanalyzer: Analyze yields of the Amplified Indexed Libraries using Fragment Analyzer / Bioanalyzer software and combine different Amplified Indexed Library products (from all replicate samples) together in equal amounts using quantitation gates set-up between 200-600bp (for scAIR-TCR-Mark37) or 200-700bp (for scAIR-BCR-Mark30). Please refer to Section 8 for recommendations on how many samples (replicates) can be combined together for different NGS flow cells. If PCR primers complicate data analysis, it is recommended to remove PCR primers by incubation of amplified products for each sample with Primer Removal Enzyme (see step below) and repeat analysis / quantification of Amplified Indexed Libraries in Bioanalyzer or Fragment Analyzer.

[0479] ill. Remove excess primers from the pooled Amplified Indexed Library by adding 2 pl of Primer Removal Enzyme to the combined sample (e.g., 100 pl), then incubate at 37°C for 30 minutes.

[0480] Purification & Quantification of Amplified Indexed Library

[0481] The purpose of this step is to remove any residual primers and reagents from the pooled Amplified Indexed Library so that the preparations are ready for NGS.

[0482] iv. Add 1.5x volume of SPRIselect beads (adjusted to room temperature) to pooled Amplified Indexed Library, mix in an Eppendorf tube, and pipet up and down 5 times to thoroughly mix the bead suspension with the pooled Amplified Indexed Library.

[0483] v. Incubate the mixture for 5 minutes at room temperature.Atty. Docket No.: CLCT-011 WO

[0484] vi. Place the tube in the Magnetic Stand for 1 minute or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0485] vii. Add 500 pl of freshly prepared 80% ethanol to the tube.

[0486] viii. Place the tube in the Magnetic Stand for 2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0487] ix. Repeat steps 4 and 5 for a second wash.

[0488] x. Briefly centrifuge tubes at low speed and place the tubes in the Magnetic Stand. Use a 200-pl pipette to remove the residual ethanol droplets from the tube and air-dry the beads at room temperature for 2 minutes.

[0489] xi. Add 45 pl of TE buffer to the pellet to disperse the beads and let stand for 1 minute.

[0490] xii. Place the tube on the Magnetic Stand for 1 minute. Transfer 40 pl of the supernatant to a new Eppendorf tube.

[0491] xiii. Measure the concentration of pooled Amplified Indexed Library sample using the Qubit® dsDNA High Sensitivity Assay or equivalent.

[0492] xiv. Dilute the pooled Amplified Indexed Library sample to 2.1 ng / pl, which corresponds to a concentration of 10nM for the following NGS step.

[0493] I. Next Generation Sequencing of Pooled Amplified Indexed Libraries

[0494] Pooling of Amplified indexed libraries

[0495] For the DriverMap scAIR TCR-Mark37 or scAIR BCR-Mark30 Assays, which generates a low complexity pool of amplified products for TCR / BCR and 30-37 T / B cell marker genes for 96 single cells (per plate) sequencing depth is not a critical issue for most Illumina’s NGS kits. Generally, it is recommended to use low-output 600-cycle Paired-End NGS reagent kits and aim for 5-10 million reads per one sample or 100 million reads for eight pooled Amplified IndexedAtty. Docket No.: CLCT-011 WO

[0496] Library samples. Under these NGS sequencing depth conditions, at least 20 reads per template molecule (for TOR or BCR chains) may be generated which is an optimal condition for downstream analysis of NGS data. Please, refer to Table 36 for guidelines on how many samples may be combined for different instruments.

[0497]

[0498] Table 36. Guidelines for combining samples for different instruments.

[0499] Sequencing on NextSeq 1000 / 2000. The instructions below are provided for sequencing of pooled Amplified Indexed Libraries using the NextSeq 1000 / 2000 XLEAP-SBS P1 reagent kit (600 cycles) (Cat.# 20100981) and NextSeq 1000 / 2000 XLEAP-SBS Read Primer Kit (Cat.# 20112859).

[0500] Sequencing of Pooled Amplified Indexed Libraries on NextSeq 1000 / 2000

[0501] The instructions below are provided for sequencing of pooled Amplified Indexed Libraries using the NextSeq 1000 / 2000 XLEAP-SBS P1 reagent kit (600 cycles) Cat.# 20100981 and NextSeq 1000 / 2000 XLEAP-SBS Read Primer Kit (Cat.# 20112859).

[0502] Follow the standard Illumina protocol for the XLEAP-SBS kit on how to prepare XLEAP-SBS consumables, dilute libraries, and set up a sequencing run. Modifications of the standard Illumina protocol, specific for the use of AIR assay custom sequencing primers are the following:

[0503] 1. Preparation of diluted pooled Amplified Indexed Library. Dilute purified pooled Amplified Indexed library (from Section 7.2) to 1 ,000 pM loading concentration by mixing 21.6 pl of RSB buffer with 2.4 pl of purified pooled Indexed Library (10 nM). Add 2 pl of PhiX Control (1 nM concentration) to the diluted indexed library.

[0504] 2. Preparation of diluted Custom AIR Sequencing Primers. Custom AIR sequencing primers provided in the kit need to be diluted using HT1 buffer (from NextSeq 1000 / 2000 XLEAP-SBS Read Primer Kit)Atty. Docket No.: CLCT-011 WO

[0505] • Prepare Custom DNA Read primer mix: combine 1.8 pl of Forward SeqDNA primer (100 pM) and 1.8 pl of Reverse SeqDNA primer (100 pM) with 600 pl of HT1 buffer.

[0506] • Prepare Custom Index primer mix: Combine 3.6 pl of Forward SeqIND primer (100 pM) and 3.6 pl of Reverse SeqIND primer (100 pM) with 600 pl of HT1 buffer.

[0507] 3. Before loading the diluted Amplified Indexed Library and Custom Read and Custom Index primer mixes, make sure that the reagent cartridge is thawed and inspected before proceeding. Invert the cartridge several times to ensure complete reagent mixing. This cannot be done once the indexed library and primers are loaded.

[0508] 4. T ransfer 600 pl of the custom read primer into the “Custom 1 ” well and 600 pl of the custom index primer into the “Custom 2” well. (See, e.g., FIG. 19).

[0509] 5. Load 26 pl of diluted Amplified Indexed Library in Library port well. (See, e.g., FIG. 19).

[0510] 6. Log in to BaseSpace for run monitoring and select manual setup mode. When in manual setup mode, enter the read lengths, then specify the custom primer wells from the drop-down under each read to the correct well. See Table 37.

[0511] Program Cycles

[0512]

[0513] Custom 1: 308

[0514] Custom 2: 10

[0515] Custom 2: 10

[0516]

[0517] Custom 1 :

[0518]

[0519] 308

[0520] Table 37. Custom programs.

[0521] This protocol will direct the instrument to perform NGS using the custom primers, not the default stock material inside the cartridge. Once that is complete, follow the standard protocol to start the run.Atty. Docket No.: CLCT-011 WO

[0522] Example 4: Generation of NGS TCR and BCR CDR3 libraries from DNA samples using AIR-RBC technology.

[0523] The protocol below describes how to prepare NGS samples using AIR-RBC technology starting from genomic DNA from blood, tissue, cells, or other biological samples, including Positive Control PBMC DNA.

[0524] A. List of Reagents is shown in Table 28.

[0525]

[0526] Atty. Docket No.: CLCT-011 WO

[0527]

[0528] Table 38. List of reagents.

[0529] B. Reverse J Gene-Specific (GS) Primers Extension

[0530] In this step, the pool of Reverse J GS primers with adjoining RBCs and AP2 sequences is annealed to J regions (top, sense strand) of TOR or BCR genes, extended by DNA polymerase to generate Reverse GSP extended cDNA and purified from non-extended primers by nuclease treatment. The protocol assumes the reactions will be set off in a 96-well plate. For small numbers of samples, reactions can be done in tubes or strips also.

[0531] 1. Prepare the Reverse J GS Primer Extension Master Mix as described in Table 39 for all samples and controls (make 5% extra to account for pipetting error) and aliquot 20 pl in the wells of a 96-well plate:

[0532]

[0533] Table 39. Reverse J GS primer extension master mix.

[0534] 2. Adjust the volume of each DNA Sample to 30 pl with water, as shown in Table 40.

[0535]

[0536] Table 40. Volumes of DNA sample with water.Atty. Docket No.: CLCT-011 WO

[0537] 3. Add 30 pl of DNA Sample to each well with pre-aliquoted 20 pl of Reverse J GS Primer Extension Master Mix. See Table 41.

[0538]

[0539] Table 41. Components.

[0540] 4. Mix contents by pipetting 3 times. Seal the plate with adhesive film, and spin down to collect droplets. Load the plate in the thermal cycler and run the program in Table 42.

[0541]

[0542] Table 42. Program.

[0543] 5. Spin down the plate, remove the seal from the plate, then add 2 pl of the Primer Removal Enzyme to each reaction well of the plate. Mix contents by pipetting three times, seal the plate and spin down to collect droplets.

[0544] 6. Load the plate in the thermal cycler, and run the program in Table 43.

[0545]

[0546] Table 43. RT

[0547] C. Forward FR3 Gene-Specific (GS) Primer Extension

[0548] In this step, the pool of Forward FR3 GS Primers with adjoining Anchor 1 sequences is annealed to Reverse GSP extended cDNA template generated in the previous step, extended by DNA polymerase, and generates the target CDR3 DNA region (top, sense strand) flanked from both sides by Anchor 1 and Anchor 2 sequences.Atty. Docket No.: CLCT-011 WO

[0549] 1. Spin down the plate from the previous step, remove the seal from the plate, then add 2.5 pl of the Forward FR3 GS Primer Mix to each reaction well of the plate. See Table 44.

[0550]

[0551] Table 44. Components.

[0552] 2. Mix contents by pipetting 3 times, seal the plate with a new adhesive film, and spin down to collect droplets.

[0553] 3. Load the plate in the thermal cycler, and run the program in Table 45.

[0554]

[0555] Table 45. Program.

[0556] 4. Spin down the plate, remove the seal from the plate, then add 2 pl of the Primer Removal Enzyme to each reaction well of the plate. Mix contents by pipetting 3 times, seal the plate, and spin down to collect droplets.

[0557] 5. Load the plate in the thermal cycler, and run the program in Table 46.

[0558]

[0559] Table 46. Program.

[0560] D. First PGR with Anchor Primers

[0561] This step utilizes universal Anchor PCR primers to amplify the target CDR3 DNA fragments flanked with the universal Anchor 1 and Anchor 2 sequences generated during the previous Forward FR3 Gene-Specific Primer Extension step.Atty. Docket No.: CLCT-011 WO

[0562] 1. Prepare the Anchor PCR Master Mix as shown in Table 47 for all samples and controls plus 5% extra volume of all components.

[0563]

[0564] Table 47. Anchor PCR master mix.

[0565] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Spin down the Forward GS Primer Extension plate, remove the seal, then add 50 pl of Anchor PCR Master Mix to each reaction well. See Table 48.

[0566]

[0567] Table 48. Components.

[0568] 3. Mix content by pipetting three times. Seal the plate with new adhesive film and spin it to collect droplets.

[0569] 4. Load the plate in the thermal cycler in a location dedicated to PCR work. Run the program in Table 49 using the recommended number of PCR cycles:

[0570]

[0571] Table 49. Program.Atty. Docket No.: CLCT-011 WO

[0572] Note: To avoid over-cycling biases, it is recommended to start with 22 PCR cycles for TCR or 25 PCR cycles for BCR assay for samples that contain 5 pg of DNA isolated from lymphocyterich samples (e.g., whole blood, PBMC, lymphoid tissues, sorted cells, and positive control PBMC DNA). For similar samples with less DNA amount, add extra cycles (e.g., if using 0.5 ug DNA, do three extra cycles). For small DNA samples isolated from sorted cells (or in DirectCell protocol), it is recommended to start from 25 cycles for 250 ng of DNA (or 50,000 sorted cells) and use proportionally more cycles for using less starting DNA. For most samples, it is not recommended to exceed 28 cycles. For DNA samples with low content of immune cells, use 5-10 ug DNA and 28 PCR cycles. The recommended number of cycles, in many cases, needs to be optimized and adjusted based on specific cell types, quality of DNA, type of tissue sample, the content of T / B cells, etc.

[0573] E. Second PCR with Indexed Primers

[0574] This step adds a dual unique DNA / RNA UD index combination to each Anchored PCR Product generated in the previous PCR with Anchor Primers step as well as universal flanking P5 and P7 sequences needed for cluster formation on the Illumina NGS flow cell. The Index PCR Plate in the kit contains a unique combination of Forward and Reverse DNA / RNA UD index primers in each well. The primers have been dried onto the bottom of each well and will be dissolved when the Index PCR Master Mix is added.

[0575] 1. Prepare enough of the Index PCR Master Mix, following the formulation in Table 50 for all samples and controls plus 5% extra volume of all components.

[0576]

[0577] Table 50. Index PCR master mix.

[0578] 2. Gently vortex the Index PCR Master Mix, and spin down briefly to collect droplets.

[0579] Remove the plate seal from the Index PCR Plate. Set up the Index Primer PCR Reactions as follows:Atty. Docket No.: CLCT-011 WO

[0580] 3. Aliquot 50 pl of the Index PCR Master Mix into appropriate wells of the 96-well Index PCR Plate (or cut-off portion of the plate) provided in the kit. To avoid index- to-index contamination, add Index PCR Master Mix using a new tip for each well.

[0581] 4. Spin down, then remove the seal from the Anchor Primer PCR plate (plate from PCR with Anchor Primers step). Transfer 2 pl of Anchored PCR products to each of the Index Primer PCR reactions on the Index PCR Plate. See Table 51. To avoid mistakes, ensure that samples in the Anchor Primer PCR plate are arranged in the same format as the Index PCR Primer pair mixes in the Index PCR Plate (e.g. Sample 1 A is aliquoted to well 1A). It is recommended to record the sample name and well number (e.g., Sample 1 in well 1 A) for all the samples, including the positive control. This will help minimize mistakes in the NGS deconvolution step.

[0582]

[0583] Table 51. Components.

[0584] 5. Seal the plate with new adhesive film and spin down to collect droplets.

[0585] 6. Load the plate in the thermal cycler, and run the program in Table 52.

[0586] <

[0587]

[0588] Table 52. Program.

[0589] F. Quantify and Combine Samples for NGS

[0590] The product from the PCR with Indexed Primers step contains both P5 and P7 sequences and dual DNA / RNA UD Indexes for Next-Gen Sequencing (NGS) on Illumina instruments. The protocol in this section provides the instructions to normalize the amount of each AmplifiedAtty. Docket No.: CLCT-011 WO

[0591] Indexed Library to obtain similar reads for each sample, to combine and clean up samples before loading onto the instrument, and guidelines for sequencing and data analysis.

[0592] In this step, the yield of products from the Index Primer PCR Reaction — the Amplified Indexed Libraries of transcripts from each sample — are measured and then pooled in equimolar amounts for sequencing.

[0593] Quantify the Amplified Indexed Libraries

[0594] 1. Analyze the Amplified Indexed Libraries using one of the following methods:

[0595] • Standard Method: Separate 5 pl of Amplified Indexed Libraries on a 3% agarose-TAE gel and analyze the size distribution of NGS probes by UV transilluminator. See FIG. 20 for the expected results of amplified libraries generated from good-quality PBMC DNA samples. For both AIR TCR and AIR BCR assay the smear with several bright bands should be in the 200-300 bp range.

[0596] • Alternative Method: Analyze 1 pl of each of the Amplified Indexed Libraries on either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Gat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol.

[0597] 2. Analyze yields of the Amplified Indexed Libraries. The yield should all be roughly the same for all experimental samples within + / - 2-3-fold levels and should be similar to Positive Control PBMC DNA sample. A negative control sample should not generate any significant yield of amplified products. If some samples show significantly lower yield (e.g., >5-10-fold) of amplification products, it could be an indication of differences in the amount, quality of DNA or content of lymphocyte DNA used in AIR assay. For experimental DNA samples with significantly less yield of PCR products than other samples or the Positive Control DNA sample (i.e., >5-10-fold difference), rerun the Amplified Indexed Libraries in the thermal cycler (see Note below) for 2-5 additional cycles. After cycling, quantify the products again relative to the Positive Control DNA. Note: Do not include the Positive Control DNA sample in additional cycles. Remove the positive control sample from the plate and keep it as a reference to assess the quantity of the PCR samples.

[0598] Remove Excess PCR Primers and Combine SamplesAtty. Docket No.: CLCT-011 WO

[0599] 1. Remove excess primers from the completed PCR reactions by adding 1 pl of Primer Removal Enzyme to each of the Amplified Indexed Libraries and the Negative Control sample, then incubate at 37°C for 30 minutes.

[0600] 2. To ensure accurate quantification for sequencing, quantitate the yield of the Amplified Indexed Libraries. The preferred method is to analyze 2 pl of each of the Amplified Indexed Libraries using either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol. Quantifying the PCR products after the removal of PCR primers is more accurate than quantifying before the primer removal clean up step. (See, e.g., FIG. 21).

[0601] 3. After primer removal and quantification, use the yield assessment of the Amplified Indexed Libraries as a basis to combine equimolar amounts of each of the Amplified Index Libraries into a single pool for NGS. For example, if the yield of Library 1 is twice that of Library 2, then mix 5 pl of Library 1 with 10 pl of Library 2. To minimize sample-to- sample sequencing variations, combine and load all experimental samples onto one NGS flow cell.

[0602] For the AIR-RBC Assay, which generates a complex pool of amplified TCR or BCR CDR3 regions (40K-100K from 5 pg of PBMC DNA), refer to Table 53 for guidelines on how many samples may be combined for different instruments and read depths. Generally, aim for 5-10 million reads per AIR sample (starting from 5 pg of PBMC DNA) or achieve at least an average of 20 reads per DNA molecule. For Illumina instruments, AIR profiling assays require the use of 300-n paired-end reagent kits.

[0603]

[0604] Table 53. Guidelines for how many samples can be combined for different instruments.Atty. Docket No.: CLCT-011 WO

[0605] E. Purification & Quantification of Amplified Indexed Library

[0606] The purpose of this step is to remove any residual primers and reagents from the pooled Amplified Indexed Libraries so that the preparations are ready for NGS. To remove residual amount of genomic DNA, it is recommended to use double (0.6V > 1 ,5V) AMPure XP purification strategy.

[0607] 1. Add 0.6x volume of Agencourt® AMPure® XP Reagent (at room temperature) to pooled Amplified Indexed Libraries mix in an Eppendorf tube, and pipet up and down 5 times to thoroughly mix the bead suspension with the pooled Amplified Indexed Libraries.

[0608] 2. Incubate the mixture for 5 minutes at room temperature.

[0609] 3. Place the tube in the Magnetic Stand for 2 minutes or until the solution is clear. Carefully collect the supernatant without disturbing the bead pellet.

[0610] 4. Add 0.9x volume of Agencourt® AMPure® XP Reagent (at room temperature) to pooled Amplified Indexed Libraries mix in an Eppendorf tube, and pipet up and down 5 times to thoroughly mix the bead suspension with the pooled Amplified Indexed Libraries.

[0611] 5. Place the tube in the Magnetic Stand for 2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the bead pellet.

[0612] 6. Add 500 pl of freshly prepared 80% ethanol to the tube.

[0613] 7. Place the tube in the Magnetic Stand for 1 minute or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0614] 8. Repeat steps 6 and 7 for a second wash.

[0615] 9. Briefly centrifuge tubes at low speed and place the tube in the Magnetic Stand. Use a 200-pl pipette to remove the residual ethanol droplets from the tube and air-dry the beads at room temperature for 2 minutes.Atty. Docket No.: CLCT-011 WO

[0616] 10. Add 45 pl of TE buffer to the pellet to disperse the beads and let stand for 1 minute.

[0617] 11. Place the tube on the Magnetic Stand for 1 minute. Transfer 40 pl of the supernatant to a new Eppendorf tube.

[0618] 12. Measure the concentration of both pooled Amplified Indexed Libraries sample and Negative Control sample using the Qubit® dsDNA High Sensitivity Assay.

[0619] 13. Dilute the pooled Amplified Indexed Libraries probe sample to 1.6 ng / pl, which corresponds to a concentration of 10 nM for the following NGS step (for NextSeq).

[0620] F. Next Generation Sequencing.

[0621] The Amplified Index Libraries made from each DNA sample can be run on an Illumina NextSeq550 sequencer following the manufacturer’s instructions. Generally, it is recommended to set up the sequencing reactions to generate 5-10 million reads per sample (read depth per sample). This depth produces 20-50 reads, on average, per DNA molecule.

[0622] The multiplexing level may be modified to meet experimental needs. For example, multiplex fewer samples together to generate more sequencing reads per sample (i.e., more depth) for increased and more sensitive detection of clonotypes present across a broader dynamic range of abundance levels or, if mostly interested only in highly abundant clonotypes, more samples may be sequenced together which will be more cost-effective.

[0623] Follow the standard Illumina procedures for Cluster Generation, starting with 10 nM of the purified PCR sample.

[0624] 1. Add 6 pl of each of the custom sequencing primers into the appropriate wells of the Illumina reagent cartridge, as follows:

[0625] • Forward SeqDNA NGS Primer (Read 1 Sequencing Primer) into well #20

[0626] • Reverse SeqDNA NGS Primer (Read 2 Sequencing Primer) into well #21

[0627] • Forward SeqIND NGS Primer (Index 1 Sequencing Primer) into well #22

[0628] • Reverse SeqIND NGS Primer (Index 2 Sequencing Primer) into well #22Atty. Docket No.: CLCT-011 WO

[0629] 2. Perform the NGS run using 300-nt paired-end reads on the NextSeq. Proceed to cluster amplification using the appropriate Illumina Paired-End Cluster Generation Kit; refer to the manufacturer’s instructions for this step. The optimal seeding concentration for cluster amplification of indexed libraries is approximately 1.8 pM for NextSeq500 /

[0630] 550. Use the program in Table 54 for the sequence run.

[0631]

[0632] Table 54. Program.

[0633] Example 5: Generation of NGS TOR and BCR full-length CDR1-CDR2-CDR3 libraries from RNA samples using AIR technology with TSO-RBC.

[0634] The protocol below describes in detail how to prepare NGS samples using RBC-barcoded template switching oligonucleotides and RBC-barcoded Reverse gene-specific primers (GSP) starting from total RNA from blood, tissue, cells, and other biological samples, including Positive Control RNA.

[0635] A. Reagents are listed in Table 55.

[0636]

[0637] Atty. Docket No.: CLCT-011 WO

[0638]

[0639] Table 55. List of reagents.

[0640] B. Hybridization of mRNA with Reverse C-region Gene- Specific Primers

[0641] In this step, mRNA-Rev TCR / BCR-C GSP hybrids are generated and then purified from non-hybridized primers by nuclease treatment and AMPure magnetic beads. The protocol is written assuming the reactions will be set off in a 96-well plate.

[0642] 1. Prepare a Hybridization Master Mix as described in Table 56 for all samples and controls (make 5% extra to account for pipetting error) and aliquot 7 pl in the wells of a 96-well plate:

[0643]

[0644] Atty. Docket No.: CLCT-011 WO

[0645]

[0646] Table 56. Hybridization master mix.

[0647] Note: If running only TOR or BCR repertoire analysis, use only TCR-C or BCR-C primer mix and adjust the total volume to 7 pl by water. If running biological or amplification triplicates, increase the number of tubes, volume, and amount of RNA allocated for each sample respectively. For the DirectCell protocol, add 1 pl of 2% N-Lauroyl Sarcosine (final concentration in hybridization mix will be 0.1%). It is recommended to run a positive control RNA (50 ng), and negative control (water) for troubleshooting and comparing batch effects in different experiments.

[0648] 2. Adjust the volume of each RNA Sample (e.g., 100 ng of whole blood or 50 ng of PBMC RNA) to 14 pl with water as shown in Table 57.

[0649]

[0650] Table 57. Components.

[0651] Note: For tissue samples with a low content of immune cells, the recommended amount is 200-1 ,000 ng of total RNA). For AIR profiling directly in immune cell fractions, use 5,000-50,000 cells in 14 pl of 1xPBS buffer.

[0652] 3. Add 14 pl of RNA Sample to each well with pre-aliquoted 7 pl of Hybridization Master Mix and mix contents by pipetting 3 times. Seal the plate with adhesive film, and spin down to collect droplets. Load the plate in the thermal cycler, and run the program in Table 58 to hybridize mRNAs with Reverse TCR / BCR-C primers:

[0653]

[0654] Table 58. Program.

[0655] 4. Add 24 pl (1 ,2x volume) of Agencourt AMPure® XP Reagent (adjusted to room temperature) to each reaction well, and pipet up and down 5 times to thoroughly mix the beadAtty. Docket No.: CLCT-011 WO

[0656] suspension with the hybridization reaction mix. Check that the whole volume in each well has a uniform brown color.

[0657] 5. Incubate the mixture for 5 minutes at room temperature. While waiting, prepare the Reverse Transcriptase Buffer Master Mix based on protocol below. Store on ice before use.

[0658] 6. Place the plate in the Magnetic Stand for 96-well plates for 1 -2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0659] 7. Add 200 pl of freshly prepared 80% ethanol in each reaction well without removal of plate from Magnetic Stand, wait for 2 minutes, carefully remove, and discard the supernatant without disturbing the bead pellets.

[0660] 8. Repeat 80% ethanol washing in step 9.

[0661] 9. Briefly centrifuge the plate at low speed and place the plate in the Magnetic Stand. Use a 20-pl pipette to remove the residual ethanol droplets from reaction wells and airdry the beads at room temperature for approximately 5 minutes.

[0662] C. cDNA Synthesis

[0663] In this step, the purified mRNA-Reverse GS primer hybrids eluted from AMPure beads are extended by Reverse Transcriptase and incorporate at 3’-end sequences complementary to mix of RBC-barcoded TSOs present in reaction to generate RBC-barcoded from both 5’ and 3’ end cDNA (antisense strand of mRNA).

[0664] 1. Prepare the Reverse Transcriptase Buffer Master Mix as shown in Table 59 for each sample plus 5% extra volume of all components:

[0665]

[0666] Atty. Docket No.: CLCT-011 WO

[0667] Table 59. Reverse transcriptase buffer master mix.

[0668] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Add 22 pl of the RT Buffer Master Mix in each reaction well and resuspend AMPure beads attached to the well surface by pipetting or in the plate using an Eppendorf shaker. Briefly centrifuge the plate at low speed and place the plate in Magnetic Stand for 1 minute. Transfer 20 pl of clear supernatant without beads from each reaction well to the new plate and seal the plate.

[0669] 3. Load the plate in the thermal cycler and start running the program in Table 60.

[0670]

[0671] Table 60. Program.

[0672] D. First PCR with Anchor Primers

[0673] This step utilizes universal Anchor PCR primers to amplify the target RBC-barcoded cDNA fragments flanked with the Anchor 1 and Anchor 2 sequences generated during the previous cDNA synthesis step.

[0674] 1. Prepare the Anchor PCR Master Mix as shown in Table 61 for all samples and controls plus 5% extra volume of all components:

[0675]

[0676] Table 61. Anchor PCR master mix.

[0677] 2. Gently vortex Master Mix and spin down briefly to collect droplets. Spin down the Forward GS Primer Extension plate, remove the seal, then add 60 pl of Anchor PCR Master Mix to each reaction well as shown in Table 62.Atty. Docket No.: CLCT-011 WO

[0678]

[0679] Table 62. Components.

[0680] 3. Mix content by pipetting 3 times. Seal the plate with new adhesive film and spin down to collect droplets.

[0681] 4. Load the plate in the thermal cycler in a location dedicated to PCR work. Run the program in Table 63 using the recommended number of PCR cycles.

[0682] <

[0683]

[0684] Table 63. Program.

[0685] Note: To avoid bias in gene expression levels by over-cycling samples, it is recommended to start with 18 PCR cycles for samples that contained 25 ng or more RNA (e.g., 50 ng of PBMC or 100 ng of whole blood or lymphocyte-rich samples) or at least 25,000 immune cells. For samples with less RNA or cells (or low lymphocyte RNA content), extra cycles may have to be added, but in general, it is recommended not to exceed 20 cycles. For very small RNA samples (1 -5 ng) or experiments with less than -1000 immune cells, 22-26 cycles may be required. The recommended number of cycles needs to be optimized and adjusted based on specific cell types, sample types, RNA quality, and the level of TCR / BCR mRNAs, etc. In order to optimize cycle number it is recommended to analyze 5 ul of amplified products in a 3% agarose gel or fragment analyzer. For optimal cycle number, weak cDNA amplification products in the range of 300-500 bp will be detected. If no products are seen, add 3 more cycles and analyze the yield of PCR products again.

[0686] E. Second PCR with Indexed PrimersAtty. Docket No.: CLCT-011 WO

[0687] This step adds a dual unique DNA / RNA UDP index combination to each Anchored PGR Product generated in the previous PGR with Anchor Primers step as well as universal flanking P5 and P7 sequences needed for cluster formation on the Illumina NGS flow cell.

[0688] The Index PGR Plate in the kit contains a unique combination of Forward and Reverse DNA / RNA UDP index primers in each well. The primers have been dried onto the bottom of each well and will be dissolved when the PGR reaction mix with the sample is added. One well should be used for each sample (triplicate samples are different samples) being sequenced.

[0689] 1. Prepare enough of the Index PGR Master Mix, following the formulation below for all samples and controls plus 5% extra volume of all components shown in Table 64.

[0690]

[0691] Table 64. Index PGR master mix.

[0692] 2. Gently vortex the Index PGR Master Mix, and spin down briefly to collect droplets. Remove the plate seal from the Index PGR Plate. Set up the Index Primer PGR Reactions as follows:

[0693] 3. Aliquot 50 pl of the Index PGR Master Mix into appropriate wells of the 96-well Index PGR Plate (or cut-off portion of the plate) provided in the kit. To avoid index- to-index contamination, add Index PGR Master Mix using a new tip for each well.

[0694] 4. Spin down, then remove the seal from the Anchor Primer PGR plate (plate from PGR with Anchor Primers step). Transfer 2 pl of Anchored PGR products to each of the Index Primer PGR reactions on the Index PGR Plate as shown in Table 65. To avoid mistakes, ensure that samples in the Anchor Primer PGR plate are arranged in the same format as the Index PGR Primer pair mixes in the Index PGR Plate (e.g. Sample 1 A is aliquoted to well 1 A). It is recommended to record the sample name and well number (e.g., Sample 1 in well 1 A) for all the samples, including the positive control. This will help minimize mistakes in the NGS deconvolution step.Atty. Docket No.: CLCT-011 WO

[0695]

[0696] Table 65. Components.

[0697] 5. Seal the plate with new adhesive film and spin down to collect droplets.

[0698] 6. Load the plate in the thermal cycler, and run the program in Table 66.

[0699] <

[0700]

[0701] Table 66. Program.

[0702] F. NGS Prep and Sequencing

[0703] The product from the PCR with Indexed Primers step contains both P5 and P7 sequences and dual DNA / RNA UDP Indexes for Next-Gen Sequencing (NGS) on Illumina instruments. The protocol in this section provides the instructions to normalize the amount of each Amplified Indexed Library to obtain similar reads for each sample, to combine and clean up samples before loading onto the instrument, and guidelines for sequencing and data analysis.

[0704] QC, Quantify and Combine Samples for NGS

[0705] In this step, the yield of products from the Index Primer PCR Reaction — the Amplified Indexed Libraries of transcripts from each sample — are analyzed (adjusted if necessary), measured, and then pooled in equimolar amounts for sequencing.

[0706] 1. Analyze the Amplified Indexed Libraries using one of the following methods: • Standard Method: Separate 5 pl of Amplified Indexed Libraries on a 3% agarose-TAE gel and analyze the size distribution of NGS probes by UV transilluminator. To minimizeAtty. Docket No.: CLCT-011 WO

[0707] the sample number, only one sample from each triplicate set may be run. See FIG. 22 for the expected results of amplified libraries generated from good-quality whole blood RNA samples. For AIR TCR-BCR assay, the smear with several bright bands should be in the 220-420 bp range.

[0708] • Alternative Method: Analyze 1 pl of each of the Amplified Indexed Libraries on either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol.

[0709] 2. Analyze yields of the Amplified Indexed Libraries. The yield should be roughly the same for all experimental samples within + / - 2-3-fold levels and similar to the Positive Control RNA sample. Negative control sample should not generate any significant yield of amplified products. If some samples show a significantly lower yield of amplification products, it could indicate differences in the amount, quality of RNA, or content of TCR / BCR mRNAs used in AIR assay. For the experimental RNA samples with a significantly lower yield of PCR product (e.g., >5-10-fold) than other samples or positive control RNA, re-run the lower yield samples in the thermal cycler for 2-5 additional cycles.

[0710] Remove Excess PCR Primers and Combine Samples

[0711] 1. Remove excess primers from the completed PCR reactions by adding 1 pl of Primer Removal Enzyme to each of the Amplified Indexed Libraries and the Negative Control sample, then incubate at 37°C for 30 minutes.

[0712] 2. To ensure accurate quantification for sequencing, repeat the quantification procedure of the Amplified Indexed Libraries and the Positive Control RNA. The preferred method is to analyze 2 pl of each of the Amplified Indexed Libraries using either an Agilent Bioanalyzer with the Agilent High Sensitivity DNA Kit (Cat.# 5067-4626) or Fragment Analyzer using the High Sensitivity NGS Analysis Kit (Cat.# DNF-473-1000) using the manufacturer’s protocol. Quantifying the PCR products after removing PCR primers is more accurate than quantifying before the primer removal clean-up step. (See, e.g., FIG. 23).

[0713] 3. After primer removal and quantification, use the yield assessment of the Amplified Indexed Libraries as a basis to combine equimolar amounts of each of the AmplifiedAtty. Docket No.: CLCT-011 WO

[0714] Index Libraries into a single pool for NGS. For example, if the yield of Library 1 is twice that of Library 2, then mix 5 pl of Library 1 with 10 pl of Library 2. To minimize sample-to- sample sequencing variations, combine and load all experimental samples onto one flow cell.

[0715] For the AIR-RBC Assay, which generates a complex pool of amplified TCR / BCR CDR regions (100K-200K from 50 ng of total whole blood, PBMC RNA), refer to Table 67 for guidelines on how many samples may be combined for different instruments and read depths. Generally, aim for 5-10 million reads per AIR sample (starting from 50 ng of PBMC RNA) or at least 20 reads per template molecule. For Illumina instruments, AIR profiling assays require 300-n paired-end reagent kits for CDR3 and 600-n paired-end reagent kit for CDR1-CDR2-CDR3.

[0716]

[0717] Table 67. Guidelines for how many samples can be combined for different instruments.

[0718] Purification & Quantification of Amplified Indexed Library

[0719] The purpose of this step is to remove any residual primers and reagents from the pooled Amplified Indexed Libraries so that the preparations are ready for NGS.

[0720] 1. Add 1 ,5x volume of Agencourt® AMPure® XP Reagent (at room temperature) to pooled Amplified Indexed Libraries, mix in an Eppendorf tube, and pipet up and down 5 times to thoroughly mix the bead suspension with the pooled Amplified Indexed Libraries.

[0721] 2. Incubate the mixture for 5 minutes at room temperature.

[0722] 3. Place the tube in the Magnetic Stand for 1 minute or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.Atty. Docket No.: CLCT-011 WO

[0723] 4. Add 500 pl of freshly prepared 80% ethanol to the tube.

[0724] 5. Place the tube in the Magnetic Stand for 2 minutes or until the solution is clear. Carefully remove and discard the supernatant without disturbing the pellet.

[0725] 6. Repeat steps 4 and 5 for a second wash.

[0726] 7. Briefly centrifuge tubes at low speed and place the tubes in the Magnetic Stand. Use a 200-pl pipette to remove the residual ethanol droplets from the tube and air-dry the beads at room temperature for 2 minutes.

[0727] 8. Add 45 pl of TE buffer to the pellet to disperse the beads and let stand for 1 minute.

[0728] 9. Place the tube on the Magnetic Stand for 1 minute. Transfer 40 pl of the supernatant to a new Eppendorf tube.

[0729] 10. Measure the concentration of pooled Amplified Indexed Library sample using the Qubit dsDNA High Sensitivity Assay.

[0730] 11. Dilute the pooled Amplified Indexed Library sample to 1.7 ng / pl, which corresponds to a concentration of 10nM for the following NGS step.

[0731] G. Next Generation Sequencing

[0732] The Amplified AIR NGS Index Libraries made from each total RNA sample should be run on an Illumina sequencer following the manufacturer’s instructions. Generally, it is recommended to set up the sequencing reactions to generate 5-10 million reads per sample (read depth per sample). This depth works out to produce 20-50 reads, on average, per template molecule. The procedure below is based on the NextSeq 500 / 550 using the 600-cycle NextSeq 500 / 550 High Output Kit v2.5 (Illumina Cat.# 20024908).Atty. Docket No.: CLCT-011 WO

[0733] The multiplexing level may be modified to meet experimental needs. For example, multiplex fewer samples together to generate more sequencing reads per sample (i.e. , more depth) for increased and more sensitive detection of clonotypes present across a broader dynamic range of abundance / expression levels, or, if mostly interested only in highly abundant / expressed clonotypes, more samples together may be sequenced together which will be less expensive.

[0734] Follow the standard Illumina procedures for Cluster Generation, starting with 10 nM of the purified PCR sample. It is recommended to add 15% of PhiX to the library.

[0735] 1. Add 6 pl of each of the custom sequencing primers into the appropriate wells of the Illumina reagent cartridge, as follows:

[0736] • Forward SeqDNA NGS Primer (Read 1 Sequencing Primer) into well #20

[0737] • Reverse SeqDNA NGS Primer (Read 2 Sequencing Primer) into well #21

[0738] • Forward SeqIND NGS Primer (Index 1 Sequencing Primer) into well #22

[0739] • Reverse SeqIND NGS Primer (Index 2 Sequencing Primer) into well #22

[0740] 2. Perform the NGS run using 600-nt paired-end reads on the NextSeq. Proceed to cluster amplification using the appropriate Illumina Paired-End Cluster Generation Kit; refer to the manufacturer’s instructions for this step. The optimal seeding concentration for cluster amplification of indexed libraries is approximately 1.8 pM for NextSeq500 / 550. Use the program in Table 68 for the sequence run (for NGS analysis of CDR3 / CDR1-CDR2-CDR3 NGS libraries):

[0741]

[0742] Table 68. Program.

[0743] Example 6: Molecule template estimation using RBCs in multiplex PCR for AIR profiling

[0744] T-cell receptor (TCR) and B-cell receptor (BCR) sequencing (TCR-seq and BCR-seq) are two important technologies in studying the immune repertoire of samples such as peripheral blood mononuclear cells (PBMCs) or tumors. In their most common form, BCR-seq and TCR-seq combine multiplex polymerase chain reaction (PCR) of the repertoire using primersAtty. Docket No.: CLCT-011 WO

[0745] targeting regions of the V(D)J and the constant region with next-generation sequencing (NGS). Data produced by these methods provides information regarding immune repertoire(s) including the presence critical clonotypes, repertoire diversity, variable (V) gene usage, analysis of public clonotypes, etc. One issue that can arise during generation of the TCR / BCR-seq data is sequence bias during the PCR or NGS steps. To combat this, unique molecular identifiers (UMIs) have been used to identify and eliminate sequence bias. However, UMI fragments can be long and very diverse, resulting in the UMI sequences interfering with any of the multitude of primers during multiplex PCR. Here, replicate barcodes (RBCs), a set of eight short barcodes (e.g., 6-9 nucleotides in length), are introduced. This compact set of barcodes improves PCR efficiency and facilitates PCR primer designs. Also, like UMIs, the RBCs may be used to estimate the number of template molecules (RNA or DNA). Using RBC-labeled primers for TCR and BCR repertoire profiling from PBMCs produces highly comparable results and similarly template values to those obtained through UMI-based assay counts. Overall, RBCs are a useful and simpler alternative to UMIs in assaying TCR and BCR repertoires.

[0746] Described herein is the use of replicate barcodes (RBCs) as an alternative to UMIs, along with a normalization strategy to estimate the number of molecular templates. RBCs are incorporated either during cDNA synthesis for RNA templates or barcode extension for DNA templates, both prior to PCR amplification. It was found that RBC-based immunoprofiling yields improved PCR amplification compared to UMI-based immunoprofiling using the same set of gene-specific primers. RBC-based and UMI-based immunoprofiling generates similar repertoires from the same biological samples. It was found that RBC-based assays yield similar quantification of molecular templates as UMI-based assays. These findings suggest that RBCs can act as an alternative to UMIs in immunoprofiling.

[0747] Replicate barcodes for adaptive immunoprofilinci

[0748] While UMIs have been used to improve accuracy in the quantification of clonotypes, there are multiple issues with their usage including the length and randomness of UMI sequences. RBCs aim to address the later issues. Similar to UMIs, RBC sequences are incorporated during the initial steps of the molecular biology workflow (FIG. 24A). For DNA templates, the RBCs are incorporated in the 3’ end of the amplicons (adjacent to the J region of the clonotypes) during a primer extension step. For RNA templates, the RBCs are incorporated in the 3’ end of the amplicons (adjacent to the C region of the clonotypes) during reverse transcription. In both cases, the result of the molecular workflow and next generation sequencing results in a paired end sequencing with read 2 containing RBCs.Atty. Docket No.: CLCT-011 WO

[0749] The reads are then aligned to reference immune receptor genes and clonotypes are assembled (FIG. 24B) with software such as MiXCR. Post-alignment, RBCs are used in two ways. First, the PCR and NGS steps can generate errors in the final sequences. RBCs can be used to identify sequences that are real vs artifactual. In this process, the assumption is that artifactual sequences have much lower abundance than real sequences. Second, RBCs can be used to identify a normalization factor to convert read counts to the number of DNA or RNA molecule templates in the sample. During this process, it is critical to identify clonotypes which have only one template (i.e. rare clonotypes). The distribution of read counts of this set of clonotypes can be plotted (FIG. 24C) to determine the normalization factor. From there, the normalization is used as a denominator to convert the reads to the estimated number of templates.

[0750] With the reads aligned and clonotypes quantified, downstream analysis procedures for immunoprofiling can be performed on the data. Repertoire statistics such as diversity metrics, gene usage, and clonotype overlaps can be measured. Clonotypes of interests such as peptide-activated clonotypes, vaccine-associated clonotypes, sequences of tumor infiltrating lymphocytes, etc. can be identified.

[0751] Clonotypes in RBC and UM I assays

[0752] To test the performance of RBCs, immunoprofiling for (A) TCR and BCR, (B) RNA and DNA template based assays, (C) CDR3 amplification and full length amplification and (D) human and mouse PBMCs was performed (see Table 69 for full list of samples). In addition, four samples were generated using UMIs for comparison. These four UMI samples matches one of the RBC samples in exact conditions including biological sample used, amount of nucleic acid material, and molecular workflow. During the PCR amplification steps, it was found that the RBC showed improved efficiency (FIG. 25A) where stronger bands are seen, in particular for the DNA assay. RBCs have shorter sequences compared UMIs and could, in part, contribute to this difference.

[0753]

[0754] Atty. Docket No.: CLCT-011 WO

[0755]

[0756] Table 69. List of samples in the experiment. Sample names are formatted as $template $cellType-$amplicon-$version-$templateAmount. Template is either R for RNA or D for DNA. Mouse cell types (mT and mB) are distinguished from human cell types (T and B) with an ’m’ before the cell type. Amplicon is either CDR3 for CDR3 or CDR123 for full length. Version is either RBC or UMI.

[0757] These samples were analyzed using standard MiXCR presets (see Methods below). Both RBC and UMI samples yielded alignment rates at 90% or more for samples with CDR3 amplicons. Unsurprisingly, the full-length amplicon samples yielded lower alignment rates since they required longer sequencing (600 NT PE vs 300 NT PE). Since the barcode sequences present in the RBC experiments were pre-selected, the performance of each barcode during amplification was investigated. However, there was no evidence that any of the selected barcodes showed any bias as the reads in the samples showed similar distribution across eight barcodes (FIG. 28). The eight barcodes have sequences shown in Table 70.

[0758]

[0759] Atty. Docket No.: CLCT-011 WO

[0760]

[0761] Table 70. Barcode sequences.

[0762] The repertoire in the UMI and the RBC assays were generally similar. The number of clonotypes in corresponding RBC and UMI samples were similar in number (FIG. 25B). The read counts were similar between the UMI and RBC samples as well across multiple receptor chains IGH and TRB (FIG. 25C); and IGK, IGL and TRA (FIG. 29). Read counts were well correlated between the repertoires of corresponding samples. This is especially evident for the clonotypes of high abundance.

[0763] Since the RBCs categorize reads into one of the unique sequences, the setup effectively simulates 8 technical replicates during the molecular workflow. The octuplicates can be used to assess variation in clonotype quantification. At least for the top clonotypes, the read quantitation is reproducible between the technical replicates (FIGS. 25D and 30A-30D). This pattern is more apparent in the RNA samples where there are more clonotypes observed.

[0764] Repertoire metrics in RBC and UMI assays

[0765] Analyzing both TCR and BCR repertoire metrics is fundamental for understanding immune system dynamics and responses. Changes in the gene usage (i.e. , the frequency of V, D, and J gene segments) could be indicative of pathogenic diseases, autoimmune diseases or cancer. For example, particular gene segments have been associated with systemic lupus erythematosus, COVID-19, chronic hepatitis B infection, lymphocytic choriomeningitis virus infection, murine cytomegalovirus infection and rheumatoid arthritis. Similarly, the diversity of TCR and BCR repertoires provide insights into how the immune system adapts to these challenges such as that seen in malaria infection, acute respiratory symptoms in infants, aging and the immune system’s ability to fight CMV infection.

[0766] In these experiments, the V gene segment usage is similar between the UMI and RBC assays of the same biological sample (FIG. 26A). A handful of V genes are robustly present in each sample, while other genes are less so. Clonotype occupancy is similar between the UMI and RBC samples, as well (FIG. 26B). The DNA samples tend to have the top clonotypesAtty. Docket No.: CLCT-011 WO

[0767] occupy more of the repertoire than the RNA samples. This is consistent with the D50 diversity metric, which calculates the number of clonotypes that occupy the top 50% of the repertoire space (FIG. 26C). The true diversity metric calculates diversity by counting the number of clonotypes along with their relative abundances to account for both the richness and evenness of clonotypes in the repertoire. This metric also shows consistency between the UMI and RBC samples (FIG. 26D). Overall, the UMI and RBC assays recapitulate each other’s results at the clonotype-level and at the repertoire-level.

[0768] Template estimation using RBCs

[0769] The number of templates can be determined from the UMI-based assay through counting of unique UMI sequences. For the RBC-based assay, the number of templates is estimated. For this procedure, clonotypes that are only found in a single RBC bin are first identified. Statistically, these clonotypes are rare and, thus, the reads from these clonotypes likely originated from a single template. After filtering out clonotype sequences that are artifactual, the average of the distribution of the reads originating from a single template were identified (FIG. 27A). This value is the number of reads per one template. This is then used as the denominator to convert read counts to estimated number of DNA or RNA templates.

[0770] The values estimated from the RBC-based assay were compared to the values determined from the UMI-based assay for the cases where a corresponding sample exists (FIG.

[0771] 27B). For the four cases that were compared, it was found that the number of UMI and the estimated number of templates correspond extremely well, particularly in the 100+ counts range. Good correlation is still seen in the 10 - 100 counts range, while at the < 10 counts range, more variation between the two assays is observed. This variation at lower abundant clonotypes is unsurprising as it can be attributed to less robust sampling of template material when preparing the sample for the pre-NGS molecular workflow. Additionally, it could be due to dropout events during template capture in the cDNA synthesis in the RNA-based assay or the primer elongation step in the DNA-based assay. This data suggests that for practical purposes, the RBCs are sufficient to determine the number of RNA and DNA templates during immunoprofiling.

[0772] Discussion

[0773] Immune repertoire profiling has become indispensable in biomedical research, providing insights into adaptive immune responses across diverse conditions. By analyzing TCR and BCR repertoires, clonal patterns in cancer, autoimmune disorders, and infectious diseases have been uncovered. This in turn enables the development of targeted diagnostics and therapies. Here,Atty. Docket No.: CLCT-011 WO

[0774] replicate barcodes (RBCs) have been developed as an alternative to UMIs. RBCs are low diversity, short nucleotide fragments that can be used to estimate the number of molecules in the data. It was found that the use of RBCs generate similar clonotypes, repertoire metrics, and molecular template quantification as UMIs with the same set of TCR and BCR gene segment primers. In addition, it was found that the shorter RBC fragments allow for better PCR amplification of clonotypes.

[0775] Methods

[0776] Data Availability: The datasets generated for this study are available under accession number GSE297422.

[0777] Nucleotide extraction: QIAGEN RNAeasy Micro / Mini Plus Kit was used for total RNA isolation. DNeasy Blood & Tissue kit was used for DNA extraction.

[0778] DNA- and RNA-based immunoprofilinq: Cellecta DriverMap Adaptive Immune Receptor (AIR) was used to perform TCR and BCR repertoire profiling.

[0779] Sample Demultiplexing: bcl2fastq (Illumina, Inc. bcl2fastq Conversion Software v2.20. Illumina, Inc., San Diego, CA, USA (2019)) was used to demultiplex and convert the sequencing results to individual sample fastqs.

[0780] QC Analysis: fastqc (Andrews S. FastQC: a quality control tool for high throughput sequence data. (2010)) was used to assess the quality of the sequencing reads prior to downstream analysis.

[0781] Read Alignment: mixer (Bolotin, D., Poslavsky, S., Mitrophanov, I. et al. MiXCR: software for comprehensive adaptive immunity profiling. Nat Methods 12, 380-381 (2015)) was used to align the reads to the reference immune repertoire genes. In addition, this command also generates the clonotype tables along with relevant metadata information.

[0782] The command used for the RBC-based assays is:

[0783] mixer analyze local:cellecta-new-bulk-kit \ -species $SPECIES \ — assemble-clonotypes-by $ASSEMBLY TYPE \ $SPLIT CLONES CONFIG \ -$MOLECULE \ -tag-pattern $TAG_PATTERN \ -rigid-left-alignment-boundary \ -floating-right-alignment-boundary $BOUNDARY \ -export-productive-clones-only \ $READ_1 $READ_2 $OUTPUT_NAME

[0784] For UMI-based assays, the command used is:

[0785] mixer analyze $PRESET \ -species $SPECIES \ -export-productive-clones-only \ $READ 1 $READ 2 $OUTPUT_NAME

[0786] whereAtty. Docket No.: CLCT-011 WO

[0787] $SPECIES = hsa (human) or mmu (mouse)

[0788] $ASSEMBLY_TYPE = CDR3 (CDR3 amplicon) or {CDR1 Begin:FR4End} (Full length amplicon)

[0789] $SPLIT_CLONES_CONFIG = --split-clones-by C (B cells) or leave blank (T cells) $MOLECULE = rna (for RNA) or dna (for DNA)

[0790] $BOUNDARY = J (CDR3 amplicon) or C (Full length amplicon)

[0791] $READ_1 , $READ_2 = paired end read file names

[0792] $OUTPUT_NAME = name of output files

[0793] $PRESET = cellecta-human-rna-xcr-umi-drivermap-air (RNA-based) or cellecta-human-dna-xcr-umi-drivermap-air (DNA-based)

[0794] The $TAG_PATTERN is set as:

[0795] ‘A(R1 :*)\A(MIRBC:GCATCAN)(R2:*)|A(MIRBC:AGTCGTN)(R2:*)|A(MIRBC:TCGCATCN)( R2:*)|A(MIRBC:AGCGTAGN)(R2:*)|A(MIRBC:TACGACTN)(R2:*)|A(MIRBC:CTGATGAN)(R2:*)|A(M IRBC:GATAGCATN)(R2:*)|A(MIRBC:GTAGGCTAN)(R2:*)’

[0796] The preset for the RBC-based assay is specified by a cellecta-new-bulk-kit.yaml file. Template Estimation: The number of molecule templates is estimated using a normalization factor. Briefly, rare clonotypes or clonotypes originating from a single template are identified from each dataset. First, the clonotypes are binned by as to how many RBCs they are identified with (from 1 -8 RBCs). The clonotypes that were only identified in a single RBC (nRBC = 1) are further processed. The read counts from this group are plotted whereby two distributions are identified. The first distribution is an exponential decay originating from zero, and the second distribution is a normal distribution at some positive value. To distinguish the real clonotypes (second distribution) and the artifactual clonotypes (first distribution), kernel density estimation is used to smoothen the distribution. Then the two maximas are identified corresponding to the two distributions. Right in between the maximas is the lowest point, which is used as a filtering cutoff to eliminate the artifactual clonotypes. The filtered set of nRBC = 1 are then determined as the set of rare clonotypes. The second peak, which corresponds to the average of the filtered nRBC = 1 , is used as the normalization factor to determine the estimated number of templates. This procedure is performed all types of RBC-based immunoprofiling assays in this study.Atty. Docket No.: CLCT-011 WO

[0797] Downstream Analysis: The immunarch (ImmunoMind Team, immunarch: An R Package for Painless Bioinformatics Analysis of T-Cell and B-Cell Immune Repertoires. Zenodo (2019)) package is used for downstream analysis to study gene usage, repertoire diversity and clonotype occupancy. R (R Core Team, 2023) is used for various statistical analyses and data visualization.

[0798] Notwithstanding the appended claims, the disclosure is also defined by the following clauses:

[0799] 1. A method of producing an amplified next generation sequencing (NGS) nucleic acid library, the method comprising:

[0800] contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, wherein the multiplex collection of nucleic acid primers comprises at least one replicate set of primers, wherein the replicate set of primers comprises at least 3 primers having a common domain and a different replicate barcode (RBC) domain; and

[0801] amplifying the template extension product composition to produce the amplified NGS nucleic acid library.

[0802] 2. The method according to Clause 1 , wherein the number of primers in each replicate set of primers ranges from 3 to 24.

[0803] 3. The method according to Clause 2, wherein the number of primers in each replicate set of primers ranges from 4 to 8.

[0804] 4. The method according to Clause 3, wherein the number of primers in each replicate set of primers is 8.

[0805] 5. The method according to any of the preceding clauses, wherein the number of primers in the multiplex collection of nucleic acid primers is 100 fold or less than the number of template molecules in the nucleic acid template sample.

[0806] 6. The method according to any of the preceding clauses, wherein the nucleic acid template sample comprises a ribonucleic acid (RNA) sample.

[0807] 7. The method according to any of Clauses 1 to 5, wherein the nucleic acid template sample comprises a deoxyribonucleic acid (DNA) sample.

[0808] 8. The method according to any of the preceding clauses, wherein the multiplex collection of nucleic acid primers comprises a multiplex collection of gene specific primers.Atty. Docket No.: CLCT-011 WO

[0809] 9. The method according to Clause 8, wherein the multiplex collection of gene specific primers comprises:

[0810] a forward primer set; and

[0811] a reverse primer set;

[0812] wherein one or both of the forward and reverse primer sets comprises the at least one replicate set of primers, wherein the replicate set of primers comprises primers having a common gene specific primer (GSP) domain and a different RBC domain.

[0813] 10. The method according to any of Glauses 8 to 9, wherein the multiplex collection of gene specific primers is configured to amplify 100 or more different genes.

[0814] 11. The method according to any of Clauses 9 to 10, wherein the contacting comprises first contacting the nucleic acid sample with the reverse primer set to produce a reverse primer template extension product composition, and then contacting the reverse primer template extension product composition with the forward primer set to produce a forward primer extension product composition.

[0815] 12. The method according to any of Clauses 9 to 11 , wherein the same collection of RBCs is present in each replicate set of primers.

[0816] 13. The method according to any of Clauses 9 to 11 , wherein different collections of RBCs are present in each replicate set of primers.

[0817] 14. The method according to any of Clauses 9 to 13, wherein the reverse primer set comprises the at least one replicate set of primers.

[0818] 15. The method according to any of Clauses 9 to 14, wherein the forward primer set comprises the at least one replicate set of primers.

[0819] 16. The method according to any of Clauses 9 to 15, wherein primers of the forward and reverse primer sets further comprise anchor domains.

[0820] 17. The method according to any of Clauses 9 to 16, wherein the multiplex collection comprises a different number of forward and reverse primers.

[0821] 18. The method according to Clause 17, wherein the multiplex collection comprises a greater number of forward than reverse primers.

[0822] 19. The method according to any of the preceding clauses, wherein the multiplex collection is configured to generate template extension products from adaptive immune receptor nucleic acids.

[0823] 20. The method according to Clause 19, wherein the adaptive immune receptor nucleic acids comprise RNA or DNA of T cell receptor (TCR) or B cell receptor (BCR) genes.Atty. Docket No.: CLCT-011 WO

[0824] 21. The method according to Clause 20, wherein the multiplex collection is configured to generate template extension products from CDR3 or CDR1 -CDR2-GDR3 portion of TOR genes TRA, TRB, TRD, TRG and / or BCR genes IGH, IGK, and IGL.

[0825] 22. The method according to any of Clauses 1 to 6, wherein the multiplex collection of nucleic acid primers comprises at least one replicate set of primers in the form of template switch oligonucleotides (TSOs), wherein the TSOs in the replicate set of primers have a common TSO domain and a different RBC domain.

[0826] 23. The method according to any of the preceding clauses, wherein the amplifying comprises a first round of amplification with anchor primers and a second round of amplification with index primers to generate the amplified NGS nucleic acid library.

[0827] 24. The method according to Clause 23, wherein the method further comprises sequencing the NGS library.

[0828] 25. A kit comprising:

[0829] (a) a multiplex collection of nucleic acid primers comprising at least one replicate set of primers, wherein the replicate set of primers comprises at least 3 primers having a common domain and a different replicate barcode (RBC) domain; and

[0830] (b) a DNA polymerase.

[0831] 26. The kit according to Clause 25, wherein the number of primers in each replicate set of primers ranges from 3 to 24.

[0832] 27. The kit according to Clause 26, wherein the number of primers in each replicate set of rimers ranges from 4 to 8.

[0833] 28. The kit according to Clause 27, wherein the number of primers in each replicate set of primers is 8.

[0834] 29. The kit according to any of Clauses 25 to 28, wherein primers of the multiplex collection of nucleic acid primers do not include unique molecular identifier (UMI) domains.

[0835] 30. The kit according to any of Clauses 25 to 29, wherein the multiplex collection of nucleic acid primers comprises a multiplex collection of gene specific primers.

[0836] 31. The kit according to Clause 30, wherein the multiplex collection of gene specific primers comprises:

[0837] a forward primer set; and

[0838] a reverse primer set;Atty. Docket No.: CLCT-011 WO

[0839] wherein one or both of the forward and reverse primer sets comprises at least one replicate set of primers, wherein the replicate set of primers comprises primers having a common gene specific primer (GSP) domain and a different RBG domain.

[0840] 32. The kit according to Clause 31 , wherein the same collection of RBCs is present in each replicate set of primers.

[0841] 33. The kit according to Clause 31 , wherein different collections of RBCs are present in each replicate set of primers.

[0842] 34. The kit according to any of Clauses 31 to 33, wherein the reverse primer set comprises the at least one replicate set of primers.

[0843] 35. The kit according to any of Clauses 31 to 34, wherein the forward primer set comprises the at least one replicate set of primers.

[0844] 36. The kit according to Clauses 31 to 35, wherein primers of the forward and reverse primer sets further comprise anchor domains.

[0845] 37. The kit according to any of Clauses 31 to 36, wherein the multiplex collection comprises a different number of forward and reverse primers.

[0846] 38. The kit according to Clause 37, wherein the multiplex collection comprises a greater number of forward than reverse primers.

[0847] 39. The kit according to any of Clauses 25 to 38, wherein the multiplex collection is configured to generate template extension products from adaptive immune receptor nucleic acids.

[0848] 40. The kit according to Clause 39, wherein the adaptive immune receptor nucleic acids comprise RNA or DNA of T cell receptor (TCR) or B cell receptor (BCR) genes.

[0849] 41. The kit according to Clause 40, wherein the multiplex collection is configured to generate template extension products from CDR3 or CDR1-CDR2-CDR3 portion of TCR genes TRA, TRB, TRD, TRG and / or BCR genes IGH, IGK, and IGL.

[0850] 42. The kit according to any of Clauses 25 to 41 , wherein the polymerase comprises an RNA-dependent DNA polymerase.

[0851] 43. The kit according to any of Clauses 25 to 42, wherein the polymerase comprises a DNA-dependent DNA polymerase.

[0852] 44. The kit according to any of Clauses 25 to 29, wherein the multiplex collection of nucleic acid primers comprises at least one replicate set of primers in the form of template switch oligonucleotides (TSOs), wherein the TSOs in the replicate set of primers have a common TSO domain and a different RBC domain.Atty. Docket No.: CLCT-011 WO

[0853] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.

[0854] Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention is embodied by the appended claims.

Claims

Atty. Docket No.: CLCT-011WOWHAT IS CLAIMED IS:

1. A method of producing an amplified next generation sequencing (NGS) nucleic acid library, the method comprising:contacting a nucleic acid template sample with a multiplex collection of nucleic acid primers to produce a template extension product composition, wherein the multiplex collection of nucleic acid primers comprises at least one replicate set of primers, wherein the replicate set of primers comprises at least 3 primers having a common domain and a different replicate barcode (RBC) domain; andamplifying the template extension product composition to produce the amplified NGS nucleic acid library.

2. The method according to Claim 1 , wherein the number of primers in each replicate set of primers ranges from 4 to 8.

3. The method according to Claim 2, wherein the number of primers in each replicate set of primers is 8.

4. The method according to any of the preceding claims, wherein the number of primers in the multiplex collection of nucleic acid primers is 100 fold or less than the number of template molecules in the nucleic acid template sample.

5. The method according to any of the preceding claims, wherein the nucleic acid template sample comprises a ribonucleic acid (RNA) sample.

6. The method according to any of Claims 1 to 4, wherein the nucleic acid template sample comprises a deoxyribonucleic acid (DNA) sample.

7. The method according to any of the preceding claims, wherein the multiplex collection of nucleic acid primers comprises a multiplex collection of gene specific primers.

8. The method according to Claim 7, wherein the multiplex collection of gene specific primers comprises:a forward primer set; andAtty. Docket No.: CLCT-011WOa reverse primer set;wherein one or both of the forward and reverse primer sets comprises the at least one replicate set of primers, wherein the replicate set of primers comprises primers having a common gene specific primer (GSP) domain and a different RBC domain.

9. The method according to Claim 8, wherein the same collection of RBCs is present in each replicate set of primers.

10. The method according to Claims 8 or 9, wherein the reverse primer set comprises the at least one replicate set of primers.

11. The method according to any of Claims 8 to 10, wherein the forward primer set comprises the at least one replicate set of primers.

12. The method according to any of the preceding claims, wherein the multiplex collection of nucleic acid primers is configured to generate template extension products from adaptive immune receptor nucleic acids.

13. The method according to any of Claims 1 to 5, wherein the multiplex collection of nucleic acid primers comprises at least one replicate set of primers in the form of template switch oligonucleotides (TSOs), wherein the TSOs in the replicate set of primers have a common TSO domain and a different RBC domain.

14. A kit comprising:(a) a multiplex collection of nucleic acid primers comprising at least one replicate set of primers, wherein the replicate set of primers comprises at least 3 primers having a common domain and a different replicate barcode (RBC) domain; and(b) a DNA polymerase.

15. The kit according to Claim 14, wherein primers of the multiplex collection of nucleic acid primers do not include unique molecular identifier (UMI) domains.