Methods for sequencing immune cell receptors

JP2025504146A5Pending Publication Date: 2026-02-12JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024546236
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-03
Filing Date
2023-02-03
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing B-cell (BCR) and T-cell (TCR) receptor sequence sequencing methods have high error rates, system bias, low sensitivity and complex library construction steps, making it difficult to accurately detect and quantify low-frequency mutations and expressions.

Method used

The double-stranded sequence sequencing method is used to reduce the error rate by adding the same exogenous barcodes to both ends of the double-stranded DNA molecule, and ensure accurate quantities through PCR amplification and clustering steps, including generating adapted double-stranded DNA molecules, amplifying and clustering analysis of double-stranded sequence reads.

Benefits of technology

It significantly reduces sequencing error rates, improves sequencing accuracy and sensitivity, and enables reliable detection and quantification of TCR and BCR sequences, especially in low-frequency mutations and low DNA volume samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are methods for determining the sequence of double-stranded DNA molecules of immune cell receptors (e.g., T cell receptors, B cell receptors).
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 306,439, filed February 3, 2022, which is incorporated herein by reference in its entirety.

[0002] Federally Sponsored Research or Development This invention was made with Government support under Grants CA006973, GM008752, and GM136577 awarded by the National Institutes of Health. The Government has certain rights in this invention.

[0003] Sequence Listing This application contains a sequence listing that has been submitted electronically as an XML file titled "44807_0406WO1_ST.26.XML". The XML file was created on February 3, 2023 and is 59,594 bytes in size. The material in the XML file is hereby incorporated by reference in its entirety. [Technical field]

[0004] The present disclosure relates to the field of nucleic acid analysis, in particular to nucleic acid sequence analysis that can determine the sequence of immune cell receptors (e.g., B cell receptors, T cell receptors) and detect mutations in nucleic acid sequences. [Background technology]

[0005] B cell (BCR) and T cell (TCR) receptors underlie the function of the adaptive immune system. A large and diverse repertoire of BCR and TCR receptors is generated through somatic recombination by imprecise joining of variable (V), diversity (D) and joining (J) genes. Comprehensive characterization of the BCR and TCR repertoire is important for applications including understanding immune responses to pathogens, malignancies, and self-antigens. Tracking specific BCR and TCR sequences is also important for understanding clonal cell dynamics and responses in health and disease. Because individual clones can be rare, methods that allow precise determination of sequences and accurate quantification of sequence abundance are essential.

[0006] High-throughput sequencing can be used for characterization of TCR and BCR repertoires. Existing methods for library preparation starting with RNA as template generally use adapter ligation or 5'RACE strategies. These methods can incorporate unique identifiers (UIDs) to increase accuracy. However, cells can contain diverse BCR or TCR transcripts, confounding quantification of clonal abundance. In addition, RNA templates may not be obtained from samples with reduced nucleotide quality, including fixed specimens. Methods starting with DNA as template for library preparation use multiplex PCR or gene capture schemes. These methods are subject to biases from sources including primer competition and differences in amplification efficiency. Complex methods are required to account for biases, such as computational corrections, use of spike-in standards, and primer balancing. Thus, existing methods for BCR and TCR sequencing are expensive, complex, require sophisticated or elaborate library preparation methods, or exhibit all of these elements of limitations. Moreover, even more advanced methods still exhibit systematic biases along with limitations in sensitivity, reproducibility, and quantitative precision.

[0007] Next-generation sequencing is in principle well suited for the confirmation and quantification of TCR and BCR sequences, but in practice the error rate of sequencing itself is too high to confidently detect TCR or BCR sequences present at low frequencies in the original sample. One type of strategy to overcome this obstacle involves bioinformatics analysis to calculate the probability that the observed sequence is not a technical artifact and is more likely to be present in the original sample. However, this strategy alone is often insufficient to detect rare sequences with optimal reliability for clinical use, leading to the use of molecular barcodes to tag all the original template molecules. Molecular barcodes provide redundant sequencing of the progeny of each tagged molecule generated by PCR, and sequencing errors are easily recognized.

[0008] Two types of molecular barcodes have been described: exogenous and endogenous. Exogenous barcodes consist of pre-specified or random nucleotides and are added during library preparation or PCR. Endogenous barcodes are formed by sequences at the 5' and 3' ends of the template fragment. Endogenous barcodes allow for "duplex sequencing," where each of the two strands (Watson and Crick) of the original DNA duplex can be distinguished by their 5' to 3' orientation that is revealed upon sequencing. Double-stranded sequencing can reduce sequencing errors because if a mutation is accidentally introduced during library preparation or sequencing, it is highly unlikely that both strands of DNA contain that same mutation. A variety of molecular barcode approaches based on either endogenous or exogenous barcodes, or a combination thereof, have been developed and applied to a wide range of clinical applications.

[0009] The barcoding strategy, which adds identical exogenous barcodes to the Watson and Crick strands of the template molecule, allows the identity of the two strands of the template to be unambiguously determined without reference to the endogenous sequence ends. Also, because the method involves double-stranded sequencing, its error rate is minimal. Although this method has the lowest error rate of any sequencing technique described to date, two problems limit its clinical application. First, it is difficult to convert most of the initial template molecules into adapter-ligated fragments with the same barcodes on each strand. This problem is particularly problematic when the amount of initial DNA is limited, such as found in cell-free plasma DNA used for liquid biopsy. Second, hybridization-based capture is used to enrich for desired regions of the genome. Although effective for enriching large regions of interest, hybridization capture is not suitable for TCR or BCR applications, does not scale well to small target regions, and exhibits low double-stranded recovery. Although successive rounds of capture can partially overcome these limitations, existing hybridization capture-based methods typically recover only a small number of input molecules with sequence information from both strands. When the region to be targeted is very small (e.g., one or a few positions in a genome of particular interest) or when the amount of DNA available is limited (e.g., <33 ng, as is often the case in plasma), capture-based approaches are not optimal. Thus, there is a need for methods that can reliably confirm and quantitate TCR and BCR sequences. Summary of the Invention

[0010] Provided herein is a method for determining the sequence of an immune cell receptor double-stranded DNA molecule, the method comprising: (a) attaching a 3' adaptor fragment to each 3' end of the double-stranded DNA molecule and a 5' adaptor fragment to each 5' end of the double-stranded DNA molecule to generate adapted double-stranded DNA molecules, where the adapted double-stranded DNA molecules comprise an adapted Watson strand and an adapted Crick strand, where the 3' adaptor fragment comprises a molecular barcode, a primer sequence, and an adapter sequence, and where the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double-stranded DNA molecule, where copying comprises performing a round of linear extension of the adapted double-stranded DNA molecule, where an adapted double-stranded Watson template and an adapted double-stranded Crick template are generated; (c) generating a first population of analyte DNA fragments from the adapted double-stranded Watson template, and sequencing the analyte DNA fragments. (d) generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (e) grouping the first sequencing reads according to the molecular barcodes present on at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (f) grouping the second sequencing reads according to the molecular barcodes present on at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family; (g) analyzing the first sequencing reads of the first analyte DNA family; and (h) analyzing the second sequencing reads of the second analyte DNA family, thus determining a sequence of the double-stranded DNA molecule.

[0011] In some embodiments, the 3' adapter fragment comprises a partial double-stranded molecular barcode. In some embodiments, the partial double-stranded molecular barcode comprises an endogenous barcode, an exogenous barcode, or both.

[0012] In some embodiments, the copying step (b) further comprises performing a round of linear extension of the adapted double-stranded DNA molecule using (i) a first primer that is complementary to the 3' adapter sequence, and (ii) a second primer that is complementary to the complementary strand of the 5' adapter sequence.

[0013] In some embodiments, generating steps (c) and (d) are performed under PCR conditions. In some embodiments, generating step (c) further comprises amplifying the adapted double-stranded Watson template with a first set of Watson target selective primer pairs, where the first set of Watson target selective primer pairs includes (i) a first Watson target selective primer that includes a sequence complementary to the 3' adapter sequence, and (ii) a second Watson target selective primer that includes a target selective sequence.

[0014] In some embodiments, the second Watson target selective primer comprises a sequence selected from the group consisting of: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO: 30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, or SEQ ID NO:65.

[0015] In some embodiments, the second Watson target selective primer comprises a sequence selected from the group consisting of: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0016] In some embodiments, the generating step (d) further comprises amplifying the adapted double-stranded click template with a first set of click target selective primer pairs, where the first set of click target selective primer pairs comprises (i) a first click target selective primer that comprises a sequence complementary to the 3' adapter sequence, and (ii) a second click target selective primer that comprises a target selective sequence.

[0017] In some embodiments, the second click target selective primer comprises a sequence selected from the group consisting of: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO: 30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, or SEQ ID NO:65.

[0018] In some embodiments, the second click target selective primer comprises a sequence selected from the group consisting of: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:26.

[0019] In some embodiments, the double-stranded DNA molecule comprises a V(D)J sequence of an immune cell receptor. In some embodiments, the target selective sequence comprises a sequence complementary to the V(D)J sequence of the immune cell receptor. In some embodiments, the immune cell receptor comprises a B cell receptor. In some embodiments, the immune cell receptor comprises a T cell receptor.

[0020] In some embodiments, the method further comprises identifying (i) a mutation in the adapted double-stranded Watson template of the first analyte DNA family, (ii) a mutation in the adapted double-stranded Crick template of the second analyte DNA family, or (iii) a mutation in both the adapted double-stranded Watson template and the adapted double-stranded Crick template. In some embodiments, the mutation is selected from the group consisting of an insertion, deletion, substitution, deletion-insertion, duplication, inversion, frameshift, repeat expansion, translocation, and combinations thereof.

[0021] In some embodiments, the method determines the sequence of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying both strands of the double-stranded DNA molecule. In some embodiments, mutations in both the adapted double-stranded Watson template and the adapted double-stranded Crick template are identified.

[0022] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. Although similar or equivalent methods and materials to those described herein can be used to practice the present invention, the preferred methods and materials are described below. All publications, patent applications, patents, and other documents mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0023] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will become apparent from the description and drawings, and from the claims. [Brief description of the drawings]

[0024] [Figure 1]An exemplary workflow for the identification and analysis of double-stranded DNA molecules in immune cells is presented. [Diagram 2] 1 is a bar graph showing the median fraction of on-target reads (i.e., reads that consist of the intended amplicon) across 13 J-segment targets derived from Watson and Crick strands. [Diagram 3] A bar graph showing targets exhibiting relatively uniform amplification, with coefficients of variation of 29% and 24% for Watson- and Crick-derived reads. [Figure 4] 1 is a bar graph showing that the number of double-stranded UID families (i.e., each UID family represents an original molecule present in the DNA sample) is exceptionally uniform across each of the 13 targets. [Diagram 5] Each TRBJ primer set was shown to recover nearly equal numbers of corresponding synthetic construct molecules (median: 833.5, range: 587-1783 for the average of the Watson and Crick strands), with minimal cross-reactivity of non-corresponding synthetic constructs. [Figure 6] FIG. 13 is a graph showing that the number of synthetic construct molecules identified, as measured by the number of molecules identified using a CMV-specific primer set, was highly correlated with orthogonal determinations of synthetic control construct concentrations. [Figure 7] FIG. 13 is a graph showing that the number of synthetic construct molecules identified was highly correlated with orthogonal determinations of synthetic control construct concentrations, as measured by the number of molecules identified by concentration in a ThermoFisher Qubit dsDNA HS assay. [Figure 8] 1 is a bar graph showing that the fraction of correct clonotypes identified by each primer set was high (median: 0.999; range 0.998-1.000). [Figure 9] 1 is a bar graph showing that the percentage of sequencing reads assignable to TCR clonotypes for all TRBJ segments was low (range 0-0.06%). [Figure 10] 1 is a bar graph showing that the number of identified clonotypes was similarly low for all TRBJ segments (range 0-6). [Figure 11] 1 is a bar graph showing that the percentage of sequencing reads assignable to TCR clonotypes was again low for all TRBJ segments (range 0-3.2%). [Figure 12] Bar graph showing that the number of identified clonotypes was also low in all TRBJ segments (range 0-82). [Figure 13] 1 is a bar graph showing the performance of primer sets Set 1 and Set 2 evaluated on DNA from T cells from normal healthy donors, where Set 1 primers yielded a higher percentage of sequencing reads that could be assigned to TCR clonotypes than Set 2 primers. [Figure 14] 1 is a bar graph showing the performance of primer sets Set 1 and Set 2 evaluated on DNA from T cells from normal healthy donors, where the Set 1 primers identified a greater number of clonotypes than the Set 2 primers. [Figure 15] 1 is a bar graph showing the performance of primer sets Set 1 and Set 3 evaluated on DNA from T cells from normal healthy donors, where primers in Set 1 yielded a greater percentage of sequencing reads that could be assigned to TCR clonotypes than primers in Set 3. [Figure 16] 1 is a bar graph showing the performance of primer sets Set 1 and Set 3 evaluated on DNA from T cells from normal healthy donors, where the Set 1 primers identified a greater number of clonotypes than the Set 3 primers. [Figure 17] 1 is a bar graph showing that Set 1 primers had a greater percentage of sequencing reads that were assignable to TCR clonotypes than Set 3 primers. [Figure 18]13 is a bar graph showing that the coefficient of variation for the number of on-target reads to each TRBJ segment for multiplex pool 1 was 103.5%. [Figure 19] 13 is a bar graph showing that multiplex pool 2 showed a more balanced recovery of each TRBJ segment, with a coefficient of variation for the number of on-target reads to the TRBJ segments of 17.5%. [Figure 20] 13 is a bar graph showing that the coefficient of variation for the number of on-target reads for each TRBJ segment for multiplex pool 3 was 19.4% for Watson and 21.4% for Crick, whereas multiplex pool 4 showed even more balanced recovery of each TRBJ segment with a coefficient of variation for the number of on-target reads for each TRBJ segment of 13.2% for Watson and 18.1% for Crick. [Figure 21] FIG. 13 shows the results of determining the yield from varying amounts of input DNA from healthy donor T cells, where the number of TCRs recovered, averaged across donors and replicates, was linear over input amounts from 25 ng to 400 ng. [Figure 22] Bar graph showing estimated yields averaged over donors and replicates and were consistent across input amounts from 25 ng to 400 ng. [Diagram 23] 1 is a graph showing the results of assessing the TCR repertoire of cells, where analysis demonstrated a decrease in clonal diversity over the course of their expansion. [Figure 24] We present the results of assessing the TCR repertoire of cells, and the results identified the expansion of specific clones over the course of expansion. [Diagram 25] We present the results of assessing the TCR repertoire of cells, and the results identified the expansion of specific clones over the course of expansion. [Figure 26] We present the results of assessing the TCR repertoire of cells, and the results identified the expansion of specific clones over the course of expansion. [Figure 27] 1 is a graph showing the results of extracting DNA and analyzing the TCR repertoire from T cells from two healthy donors, designated "AB02" and "AB04." The results show that the number of TCRs recovered and the diversity of the recovered TCRs were high for both donors across all replicates. [Figure 28] FIG. 13 is a graph showing that pairwise distance correlation between replicates from each donor was consistently high. [Figure 29] FIG. 13 is a graph showing that for representative donor AB02, clonotype frequencies in samples from DNA input amounts of 400 ng, 100 ng, and 25 ng were well correlated. [Diagram 30] FIG. 13 is a graph showing that for representative donor AB02, clonotype frequencies in samples from DNA input amounts of 400 ng, 100 ng, and 25 ng were well correlated. [Diagram 31] Figure 1 shows the analysis of TCR V segment gene usage in T cell populations by flow cytometry using the Beckman Coulter IOTest Beta Mark TCR VB Repertoire Kit. The ratio of V gene segment usage correlated well with the ratio of V gene segment usage measured by flow cytometry. [Diagram 32] FIG. 13 is a graph showing that the proportion of TCR reads corresponding to Jurkat clone TCR correlated well with the proportion of input DNA. [Diagram 33] 1 is a graph showing the results of analyzing DNA from Jurkat clonal T cell lines, in which one TCR clone was identified in four replicates and two clones were called in two replicates in the Jurkat samples. [Diagram 34] 1 is a bar graph showing the results of DNA isolation from plasma samples from patients with colorectal cancer, where analysis reveals the TCR repertoire in the plasma. [Diagram 35] 1 is a bar graph showing the results of DNA isolation from white blood cell samples from patients with colorectal cancer, where the analysis shows the TCR repertoire in the white blood cell samples. [Diagram 36] 1 is a bar graph showing the results of DNA isolation from tumor samples from patients with colorectal cancer, where analysis reveals the TCR repertoire in the tumor samples. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Detailed Description

[0026] B cell (BCR) and T cell (TCR) receptors underlie the function of the adaptive immune system. A large and diverse repertoire of BCR and TCR receptors is generated through somatic recombination by imprecise joining of variable (V), diversity (D) and joining (J) genes. Comprehensive characterization of the BCR and TCR repertoire is important for applications including understanding immune responses to pathogens, malignancies, and self-antigens. Tracking specific BCR and TCR sequences is also important for understanding clonal cell dynamics and responses in health and disease. Because individual clones can be rare, methods that allow precise determination of sequences and accurate quantification of sequence abundance are essential.

[0027] High-throughput sequencing can be used for characterization of TCR and BCR repertoires. Existing methods for library preparation starting with RNA as template generally use adapter ligation or 5'RACE strategies. These methods can incorporate unique identifiers (UIDs) to increase accuracy. However, cells can contain diverse BCR or TCR transcripts, confounding quantification of clonal abundance. In addition, RNA templates may not be obtained from samples with reduced nucleotide quality, including fixed specimens. Methods starting with DNA as template for library preparation use multiplex PCR or gene capture schemes. These methods are subject to biases from sources including primer competition and differences in amplification efficiency. Complex methods are required to account for biases, such as computational corrections, use of spike-in standards, and primer balancing. Thus, existing methods for BCR and TCR sequencing are expensive, complex, require sophisticated or elaborate library preparation methods, or exhibit all of these elements of limitations. Moreover, even more advanced methods still exhibit systematic biases along with limitations in sensitivity, reproducibility, and quantitative precision.

[0028] Thus, there is a need for improvements to sequencing library preparation and workflows to enable accurate identification of rare mutations, as well as epigenetic changes, from the same aliquot of DNA purified from clinically relevant samples.

[0029] Provided herein is a method for sequencing an immune cell receptor double stranded DNA molecule, the method comprising: (a) attaching a 3' adapter fragment to each 3' end of the double stranded DNA molecule and a 5' adapter fragment to each 5' end of the double stranded DNA molecule to generate an adapted double-stranded DNA molecule, where the adapted double stranded DNA molecule comprises an adapted Watson strand and an adapted Crick strand, where the 3' adapter fragment comprises a molecular barcode, a primer sequence, and an adapter sequence, and where the molecular barcode of the adapted Watson strand is a reverse complement of the molecular barcode of the adapted Crick strand; (b) copying both strands of the adapted double stranded DNA molecule, where copying comprises a linear extension of the adapted double stranded DNA molecule. (c) generating a first population of analyte DNA fragments from the adapted double-stranded Watson template and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (d) generating a second population of analyte DNA fragments from the adapted double-stranded Click template and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (e) generating a first analyte DNA fragment from the adapted double-stranded Click template and generating a second sequencing read for at least one member of the second population of analyte DNA fragments. (f) grouping the first sequencing reads according to the molecular barcodes present on at least one member of the first population of analyte DNA fragments to generate a second analyte DNA family; (g) analyzing the first sequencing reads of the first analyte DNA family; and (h) analyzing the second sequencing reads of the second analyte DNA family, thus determining the sequence of the double-stranded DNA molecule.

[0030] Various non-limiting embodiments of these methods are described herein and may be used in any combination without limitation. Additional embodiments of the various components of the methods for identifying the presence or absence of mutations and methylation are known in the art.

[0031] It must be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0032] As used herein, "adaptor," "adapter," and "tag" are terms used interchangeably and refer to a species that can be joined to a polynucleotide sequence (e.g., in a process called "tagging") using any one of a number of different techniques, including, but not limited to, ligation, hybridization, and tagmentation. In some embodiments, an adapter can also be a nucleic acid sequence that adds a function, such as a spacer sequence, a primer sequence / site, a barcode sequence, or a unique molecular identifier sequence.

[0033] As used herein, the term "barcode" refers to a label, or identifier, that conveys or is capable of conveying information (e.g., information about an analyte in a sample). A barcode can be part of an analyte and can be independent of the analyte. In some embodiments, a barcode can be attached to an analyte. In some embodiments, a particular barcode can be unique with respect to other barcodes. In some embodiments, a barcode can have a variety of different formats. For example, a barcode can include non-random, semi-random, and / or random nucleic acid and / or amino acid sequences, as well as synthetic nucleic acid and / or amino acid sequences. In some embodiments, a barcode can be attached to an analyte or to another moiety or structure in a reversible or irreversible manner. In some embodiments, a barcode can be added, for example, to fragments of a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample prior to or during sequencing of the sample. In some embodiments, a barcode can allow for identification and / or quantification of individual sequencing-reads. In some embodiments, a barcode may refer to a unique identifier (UID), and the terms "barcode" and "UID" may be used interchangeably.

[0034] As used herein, the terms "nucleotide" and "nt" are used interchangeably herein to generally refer to biological molecules that are composed of nucleic acids. Nucleotides can have moieties that include known purine and pyrimidine bases. Nucleotides can also have other heterocyclic bases that are modified. Such modifications include, for example, methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses, or other heterocycles. The terms "polynucleotide", "nucleic acid", and "oligonucleotide" can be used interchangeably and refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise a sequence that is not naturally occurring. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides or nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.

[0035] As used herein, "primer" generally refers to a polynucleotide molecule comprising a nucleotide sequence (e.g., an oligonucleotide) having a free 3'-OH group, which is capable of hybridizing to a template sequence (such as a target polynucleotide, or a primer extension product, etc.) and promoting polymerization of a polynucleotide complementary to the template. In some embodiments, the primer is a biotinylated primer.

[0036] Summary

[0037] This document relates to methods and materials useful for accurately identifying TCR / BCR receptor sequences present in a nucleic acid sample. In some embodiments, the method includes identifying TCR / BCR receptor sequences by using both Watson and Crick strands of a double-stranded nucleic acid template. Such methods are particularly useful for characterizing and quantifying TCR / BCR receptor sequences and allow identification of TCR and BCR repertoires with high confidence.

[0038] In some cases, the methods and materials described herein can determine TCR / BCR receptor sequences with a low error rate. For example, the methods and materials described herein can be used to determine TCR / BCR receptor sequences in a nucleic acid template with an error rate of less than about 1% (e.g., less than about 0.1%, less than about 0.05%, or less than about 0.01%). In some cases, the methods and materials described herein can be used to determine TCR / BCR receptor sequences in a nucleic acid template with an error rate of about 0.001% to about 0.01%. In some cases, the error rate associated with identifying TCR / BCR receptor sequences in analyte DNA fragments by the methods described herein is less than 1x10 -2 Not exceeding 1x10 -3 Not exceeding 1x10 -4 Not exceeding 1x10 -5 Not exceeding 1x10 -6 Not exceeding 5x10 -6 Not to exceed 1x10 -7In some cases, the error rate associated with identifying TCR / BCR receptor sequences in analyte DNA fragments by the methods described herein is reduced by at least 2-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, or 100-fold as compared to alternative methods of identifying TCR / BCR receptor sequences that do not require the use of both the Watson and Crick strands of the analyte DNA fragment.

[0039] In some embodiments, the alternative method includes standard molecular barcoding or standard PCR-based molecular barcoding followed by sequencing. In certain embodiments, the alternative method includes: (a) attaching adapters to a population of double-stranded DNA fragments in an analyte DNA sample, where the adapters include unique exogenous UIDs; (b) performing an initial amplification to amplify the adapter-ligated double-stranded DNA fragments to produce amplicons; (c) determining sequence reads of one or more amplicons of the adapter-ligated double-stranded DNA fragments; (d) assigning the sequence reads to a UID family, where each member of the UID family includes an identical exogenous UID sequence; (e) a nucleotide sequence is identified as accurately representing the analyte DNA fragment when a threshold percentage of members of the UID family include the sequence; and (f) identifying a TCR / BCR receptor sequence in the analyte DNA fragment.

[0040] In some cases, the methods and materials described herein can be used to achieve efficient duplex recovery. For example, the methods described herein can be used to recover PCR amplification products derived from both the Watson and Crick strands of a double-stranded nucleic acid template. In some cases, the methods described herein can be used to achieve a duplex recovery rate of at least 50% (e.g., about 50%, about 60%, about 70%, about 75%, about 80%, about 82%, about 85%, about 88%, about 90%, about 93%, about 95%, about 97%, about 99%, or 100%).

[0041] In some cases, the methods and materials described herein can be used to determine TCR / BCR receptor sequences with low allele frequencies. For example, the methods described herein can be used to determine TCR / BCR receptor sequences with low allele frequencies of less than about 1% (e.g., less than about 0.1%, less than about 0.05%, or less than about 0.01%). In some cases, the methods described herein can be used to determine TCR / BCR receptor sequences with low allele frequencies of less than about 0.001%.

[0042] In some cases, the methods described herein can be used to determine TCR / BCR receptor sequences present in an analyte nucleic acid sample at a frequency of 0.1% or less. In some embodiments, the methods described herein can be used to determine TCR / BCR receptor sequences present in an analyte nucleic acid sample at a frequency of 0.1% to 0.00001%. In some embodiments, the methods described herein can be used to determine TCR / BCR receptor sequences present in an analyte nucleic acid sample at a frequency of 0.1% to 0.01%.

[0043] In some cases, a method for determining the TCR / BCR receptor sequence of a double-stranded nucleic acid can include generating a double-stranded sequencing library having a double-stranded molecular barcode on each end (e.g., 5' and 3' ends) of each nucleic acid in the library, generating a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences from the double-stranded sequencing library, and determining the TCR / BCR receptor sequence of the double-stranded nucleic acid in each single-stranded library. The presence of a first molecular barcode in the 3' duplex adaptor and a second molecular barcode present in the 5' adaptor can be used to distinguish amplification products derived from the Watson strand from amplification products derived from the Crick strand.

[0044] In some cases, the method for identifying a TCR / BCR receptor sequence includes: (a) partially attaching a double-stranded 3' adapter to the 3' ends of both the Watson and Crick strands of a population of double-stranded DNA fragments in an analyte DNA sample, where a first strand of the partially double-stranded 3' adapter includes in a 5'-3' direction the following: (i) a first segment, (ii) an exogenous UID sequence, (iii) a 5' adapter annealing site, and (iv) an R2 sequencing primer site; and and wherein a second strand of the partially double-stranded 3' adapter comprises, in a 5'-3' direction, (i) a segment complementary to the first segment, and (ii) a 3' blocking group, optionally wherein the second strand is degradable; (b) annealing the 5' adapter to the 3' adapter via an annealing site, wherein the 5' adapter comprises, in a 5' to 3' direction, (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and that includes an R1 sequencing primer site, and (ii) a 3' blocking group for the 5' adapter. (c) performing a nick translation reaction to extend a 5' adapter across the exogenous UID sequence of the 3' adapter and covalently attach the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment; (d) performing an initial amplification to amplify the adapter-ligated double-stranded DNA fragment to produce an amplicon; (e) determining sequence reads of one or more amplicons of the one or more adapter-ligated double-stranded DNA fragments; (f) assigning the sequence reads to UID families, where each member of the UID family contains the same exogenous UID sequence; (g) assigning the sequence reads of each UID family to Watson and Crick subfamilies based on the spatial relationship of the exogenous UID sequence to the R1 and R2 read sequences; (h) identifying a nucleotide sequence as accurately representing the Watson strand of the analyte DNA fragment when a threshold percentage of members of the Watson subfamily contain that sequence;(i) a nucleotide sequence is identified as exactly representing the Crick strand of an analyte DNA fragment when a threshold percentage of members of the Crick subfamily contain that sequence; (j) identifying a TCR / BCR receptor sequence in a nucleotide sequence that exactly represents the Watson strand; (k) identifying a TCR / BCR receptor sequence in a nucleotide sequence that exactly represents the Crick strand; and (l) identifying a TCR / BCR receptor sequence in an analyte DNA fragment when the TCR / BCR receptor sequence in the nucleotide sequence that exactly represents the Watson strand and the TCR / BCR receptor sequence in the nucleotide sequence that exactly represents the Crick strand are the same TCR / BCR receptor sequence.

[0045] In some cases, a method for identifying a TCR / BCR receptor sequence includes: (a) attaching an adapter to a population of double-stranded DNA fragments, where the adapter includes a double-stranded portion that includes an exogenous UID and a forked portion that includes: (i) a single-stranded 3' adapter sequence that includes an R2 sequencing primer site, and (ii) a single-stranded 5' adapter sequence that includes an R1 sequencing primer site; (b) performing an initial amplification to amplify the adapter-ligated double-stranded DNA fragments to produce amplicons; (c) selectively amplifying the Watson strand amplicons that include a target polynucleotide sequence with a first set of Watson target selective primer pairs, the first set of Watson target selective primer pairs including: (i) a first Watson target selective primer that includes a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a second Watson target selective primer that includes a target selective sequence, thereby generating a target Watson amplification product; (d) a Crick strand that includes the same target polynucleotide sequence. (i) selectively amplifying the amplicons of (a) and (b) with a first set of click target selective primer pairs, the first set of click target selective primer pairs comprising: (i) a first click target selective primer comprising a sequence complementary to the R1 sequencing primer portion of the universal 5' adapter sequence, and (ii) a second click target selective primer comprising the same target selective sequence as the second click target selective primer sequence, thereby producing target click amplification products; (e) determining sequence reads of the target Watson amplification products and the target click amplification products; (f) assigning the sequence reads to UID families, where each member of the UID family comprises the same exogenous UID sequence; (g) assigning the sequence reads of each UID family to Watson subfamilies and Crick subfamilies based on the spatial relationship of the exogenous UID sequence to the R1 and R2 read sequences; (h) identifying a nucleotide sequence as accurately representing a Watson strand of an analyte DNA fragment when a threshold percentage of members of the Watson family comprise that sequence;(i) a nucleotide sequence is identified as exactly representing the Crick strand of an analyte DNA fragment when a threshold percentage of members of the Crick family contain that sequence; and (j) a TCR / BCR receptor sequence in an analyte DNA fragment is identified when both a nucleotide sequence exactly representing the Watson strand and a nucleotide sequence exactly representing the Crick strand contain the same TCR / BCR receptor sequence;

[0046] In some cases, the methods and materials described herein can be used to independently evaluate each strand of a double-stranded nucleic acid. For example, when a nucleic acid mutation is identified in an independently evaluated strand of a double-stranded nucleic acid as described herein, the materials and methods described herein can be used to determine which strand of the double-stranded nucleic acid the nucleic acid mutation originated from.

[0047] Any suitable method can be used to generate a double-stranded sequencing library. As used herein, a double-stranded sequencing library is a plurality of nucleic acid fragments that includes a double-stranded molecular barcode at one end (e.g., 5'-end and / or 3'-end) of each nucleic acid fragment in the library, and both strands of the double-stranded nucleic acid can be sequenced. In some cases, a nucleic acid sample can be fragmented to generate nucleic acid fragments, and the generated nucleic acid fragments can be used to generate a double-stranded sequencing library. The nucleic acid fragments used to generate a double-stranded sequencing library can also be referred to herein as input nucleic acids. For example, when the nucleic acid fragments used to generate a double-stranded sequencing library are DNA fragments, the DNA fragments can also be referred to herein as input DNA. A double-stranded sequencing library can include any suitable number of nucleic acid fragments. In some cases, generating a double-stranded sequencing library can include fragmenting a nucleic acid template and ligating an adaptor to each end of each nucleic acid fragment in the library.

[0048] analyte nucleic acid

[0049] The nucleic acid template in the analyte nucleic acid sample can include any type of nucleic acid (e.g., DNA, RNA, and DNA / RNA hybrid). In some cases, the nucleic acid template can be a double-stranded DNA template. Examples of nucleic acids that can be used as template for the methods described herein include, but are not limited to, genomic DNA, circulating free DNA (cfDNA; e.g., circulating tumor DNA (ctDNA), and cell-free fetal DNA (cffDNA)).

[0050] In some embodiments, the nucleic acid template in the nucleic acid sample is a nucleic acid fragment, e.g., a DNA fragment. In some embodiments, the ends of the DNA fragments represent unique sequences that can be used as endogenous unique identifiers of the fragments. In some embodiments, the fragments are produced manually. In some embodiments, the fragments are produced by shearing, such as enzymatic shearing, shearing by chemical means, acoustic shearing, nebulization, centrifugal shearing, point-sink shearing, needle shearing, sonication, restriction endonucleases, non-specific nucleases (e.g., DNase I), and the like. In some embodiments, the fragments are not produced manually. In some embodiments, the fragments are from a cfDNA sample.

[0051] In some embodiments, the nucleic acid fragments being analyzed have a length of from about 4 to about 1000 nucleotides (e.g., from about 10 to about 1000, from about 20 to about 1000, from about 30 to about 1000, from about 40 to about 1000, from about 50 to about 1000, from about 60 to about 1000, from about 70 to about 1000, from about 80 to about 1000, from about 90 to about 1000, from about 100 to about 1000, from about 250 to about 1000, from about 500 to about 1000, from about 750 to about 1000, from about 4 to about 750, from about 10 to about 10 up to about 750, from about 20 to about 750, from about 30 to about 750, from about 40 to about 750, from about 50 to about 750, from about 60 to about 750, from about 70 to about 750, from about 80 to about 750, from about 90 to about 750, from about 100 to about 750, from about 250 to about 750, from about 500 to about 750, from about 4 to about 500, from about 10 to about 500, from about 20 to about 500, from about 30 to about 500, from about 40 to about 500, from about 50 to about 500, from about 60 to about 500, from about 70 to about 500, from about 80 to about 500 , from about 90 to about 500, from about 100 to about 500, from about 250 to about 500, from about 4 to about 250, from about 10 to about 250, from about 20 to about 250, from about 30 to about 250, from about 40 to about 250, from about 50 to about 250, from about 60 to about 250, from about 70 to about 250, from about 80 to about 250, from about 90 to about 250, from about 100 to about 250, from about 4 to about 100, from about 10 to about 100, from about 20 to about 100, from about 30 to about 100, from about 40 to about 100, from about 50 to about 100, from about 60 to about 100, from about 70 to about 100, from about 80 to about 100, from about 90 to about 100, from about 4 to about 90, from about 10 to about 90, from about 20 to about 90, from about 30 to about 90, from about 40 to about 90, from about 50 to about 90, from about 60 to about 90, from about 70 to about 90, from about 80 to about 90, from about 4 to about 80, from about 10 to about 80, from about 20 to about 80, from about 30 to about 80, from about 40 to about 80, from about 50 to about 80, from about 60 to about 80, from about 70 to about 80, from about 4 to about 70, from about 10 to about 70,from about 20 to about 70, from about 30 to about 70, from about 40 to about 70, from about 50 to about 70, from about 60 to about 70, from about 4 to about 60, from about 10 to about 60, from about 20 to about 60, from about 30 to about 60, from about 40 to about 60, from about 50 to about 60, from about 4 to about 50, from about 10 to about 50, from about 20 to about 50, from about 30 to about 50, from about 40 to about 50, from about 4 to about 40, from about 10 to about 40, from about 20 to about 40, from about 30 to about 40, from about 4 to about 30, from about 10 to about 30, from about 20 to about 30, from about 4 to about 20, from about 10 to about 20, or from about 4 to about 10). In some embodiments, the length of the nucleic acid fragments analyzed can be less than 1000 (e.g., less than 750, less than 500, less than 250, less than 100, less than 50, or less than 20) nucleotides.

[0052] In some embodiments, the ends of the nucleic acid template are used as endogenous UIDs. A skilled artisan may determine the length of the endogenous UID required to uniquely identify a nucleic acid template using factors such as, for example, the length of the entire template, the complexity of the nucleic acid template in the partition or starting nucleic acid sample, etc. In some embodiments, 10-500 nucleotides of the end of the nucleic acid template are used as endogenous UIDs. In some embodiments, 15-100 nucleotides of the end of the nucleic acid template are used as endogenous UIDs. In some embodiments, 15-40 nucleotides of the end of the nucleic acid template are used as endogenous UIDs. In some embodiments, at least 10 nucleotides of the end of the nucleic acid template are used as endogenous UIDs. In some embodiments, at least 15 nucleotides of the end of the nucleic acid template are used as endogenous UIDs. In some embodiments, only one end of the nucleic acid template is used as an endogenous UID.

[0053] In some embodiments, the nucleic acid template comprises one or more target polynucleotides. The terms "target polynucleotide", "target region", "nucleic acid template of interest", "desired locus", "desired template", or "target" are used interchangeably herein to refer to a polynucleotide of interest under study. In certain embodiments, the target polynucleotide comprises one or more sequences of interest and under study. The target polynucleotide can include, for example, a genomic sequence. The target polynucleotide can include a target sequence for which it is desired to determine the presence, amount, and / or nucleotide sequence, or changes therein.

[0054] The target polynucleotide can be a region of a gene associated with a disease. In some embodiments, the gene is a druggable target. As used herein, the term "druggable target" generally refers to a gene or cellular pathway that is regulated by a disease treatment. The disease can be cancer. Thus, the gene can be a known cancer-associated gene.

[0055] In some embodiments, the input nucleic acid, also referred to herein as a nucleic acid sample, is obtained from a biological sample. The biological sample may be obtained from a subject. In some embodiments, the subject is a mammal. Examples of mammals from which nucleic acid can be obtained and used as a nucleic acid template in the methods described herein include, but are not limited to, humans, non-human primates (e.g., monkeys), dogs, cats, sheep, rabbits, mice, hamsters, and rats. In some embodiments, the subject is a human subject. In some embodiments, the subject is a plant.

[0056] Biological samples include, but are not limited to, plasma, serum, blood, tissue, tumor samples, stool, sputum, saliva, urine, sweat, tears, ascites, bronchoaveolar lavage, semen, archaeological specimens, and forensic samples. In certain embodiments, the biological sample is a solid biological sample, e.g., a tumor sample. In some embodiments, the solid biological sample is processed. The solid biological sample may be processed by fixing in a formalin solution and then embedding in paraffin (e.g., an FFPE sample). Alternatively, the processing may include freezing the sample before performing the probe-based assay. In some embodiments, the sample is neither fixed nor frozen. The unfixed, unfrozen sample may simply be stored in a storage liquid configured for the preservation of nucleic acids, for example.

[0057] In some embodiments, the biological sample is a liquid biological sample. Liquid biological samples include, but are not limited to, plasma, serum, blood, sputum, saliva, urine, sweat, tears, peritoneal fluid, bronchoalveolar lavage, and semen. In some embodiments, the liquid biological sample is acellular or substantially acellular. In certain embodiments, the biological sample is a plasma or serum sample. In some embodiments, the liquid biological sample is a whole blood sample. In some embodiments, the liquid biological sample comprises peripheral mononuclear blood cells.

[0058] In some embodiments, the nucleic acid sample is separated and purified from the biological sample. Nucleic acid can be separated and purified from the biological sample using any means known in the art. For example, the biological sample can be treated to release nucleic acid from cells or to separate nucleic acid from unwanted components of the biological sample (e.g., proteins, cell walls, other contaminants). For example, nucleic acid can be extracted from the biological sample using liquid extraction (e.g., Trizol, DNAzol) techniques. Nucleic acid can also be extracted using commercial kits (e.g., Qiagen DNeasy kit, QIAamp kit, Qiagen Midi kit, QIAprep spin kit).

[0059] In some embodiments, the biological sample contains low amounts of nucleic acid. In some embodiments, the biological sample contains less than about 500 nanograms (ng) of nucleic acid. For example, the biological sample contains about 30 ng to about 40 ng of nucleic acid.

[0060] Nucleic acids can be concentrated by known methods, including, by way of example only, centrifugation. Nucleic acids can be bound to selective membranes (e.g., silica) for purposes of purification. Nucleic acids can also be enriched for fragments of a desired length, e.g., fragments less than 1000, 500, 400, 300, 200, or 100 base pairs in length. Such enrichment based on size can be performed, for example, using PEG-induced precipitation, electrophoretic gels or chromatographic materials (Huber et al. (1993) Nucleic Acids Res. 21:1061-6), gel filtration chromatography, TSK gels (Kato et al. (1984) J. Biochem, 95:83-86), the publications of which are incorporated herein by reference.

[0061] Polynucleotides extracted from a biological sample can be selectively precipitated or concentrated using any method known in the art.

[0062] In some embodiments, the nucleic acid sample comprises less than about 35 ng of nucleic acid. For example, the nucleic acid sample can comprise from about 1 ng to about 35 ng of nucleic acid (e.g., from about 1 ng to about 30 ng, from about 1 ng to about 25 ng, from about 1 ng to about 20 ng, from about 1 ng to about 15 ng, from about 1 ng to about 10 ng, from about 1 ng to about 5 ng, from about 5 ng to about 35 ng, from about 10 ng to about 35 ng, from about 15 ng to about 35 ng, from about 20 ng to about 35 ng, from about 25 ng to about 35 ng, from about 30 ng to about 35 ng, from about 5 ng to about 30 ng, from about 10 ng to about 25 ng, from about 15 ng to about 20 ng, from about 5 ng to about 10 ng, from about 10 ng to about 15 ng, from about 15 ng to about 20 ng, from about 20 ng to about 25 ng, or from about 25 ng to about 30 ng of nucleic acid). In some cases, the nucleic acid sample may include nucleic acid from a genome that contains nucleic acid of more than about several hundred nucleotides.

[0063] In some cases, the nucleic acid sample can be essentially free of contamination. For example, when the nucleic acid sample is a cfDNA template, the cfDNA can be essentially free of genomic DNA contamination. In some cases, a cfDNA sample that is essentially free of genomic DNA contamination can contain minimal (or no) high molecular weight (e.g., >1000bp) DNA. In some cases, the methods described herein can include determining whether the nucleic acid sample is essentially free of contamination. Any suitable method can be used to determine whether the nucleic acid sample is essentially free of contamination. Examples of methods that can be used to determine whether the nucleic acid sample is essentially free of contamination include, for example, a TapeStation system, and a Bioanalyzer. For example, when using a TapeStation system and / or a Bioanalyzer to determine whether a cfDNA sample is essentially free of genomic DNA contamination, a prominent peak at ~180bp (e.g., corresponding to mononucleosomal DNA) can be used to indicate that the nucleic acid sample is essentially free of genomic DNA contamination.

[0064] In some cases, the nucleic acid fragments that can be used to generate a double-stranded sequencing library (e.g., before attaching a 3' double-stranded adaptor to the 3' end of the nucleic acid fragment) can be end-repaired. Any suitable method can be used to end-repair the nucleic acid template. For example, a blunt-end reaction (e.g., blunt-end ligation) and / or a dephosphorylation reaction can be used to end-repair the nucleic acid template. In some cases, the blunt-end reaction can include filling in the single-stranded region. In some cases, the blunt-end reaction can include degrading the single-stranded region. In some cases, the blunt-end reaction and the dephosphorylation reaction can be used to end-repair the nucleic acid template.

[0065] adapter

[0066] As used herein, "adapter" and "adapter fragment" can refer to a species that can be coupled to a polynucleotide sequence using any one of a number of different techniques, including, but not limited to, ligation, hybridization, and tagmentation. In some embodiments, an adapter fragment can also be a nucleic acid sequence that adds a function, such as a spacer sequence, a primer sequence / site, or a barcode sequence (e.g., a UID sequence).

[0067] In some embodiments, the method includes attaching an adaptor to a population of double-stranded DNA molecules to produce a population of adaptor-attached double-stranded DNA molecules, where the adapted double-stranded DNA molecules include an adapted Watson strand and an adapted Crick strand, where the adaptor fragment includes a molecular barcode, a primer sequence, and an adaptor sequence, and where the molecular barcode of the adapted Watson strand is the reverse complement of the molecular barcode of the adapted Crick strand. In some embodiments, the primer sequence can be the reverse complement of the adaptor sequence. In some embodiments, the adaptor sequence can include a specific sequence to enable sequencing when generating a sequence library. In some embodiments, the adaptor sequence includes a sequencing primer sequence (e.g., R1, R2).

[0068] In some embodiments, the adapter comprises a double-stranded portion that includes an exogenous UID and a forked portion that includes: (i) a single-stranded 3' adapter sequence and (ii) a single-stranded 5' adapter sequence. In some embodiments, the single-stranded 3' adapter sequence is not complementary to the single-stranded 5' adapter sequence. In some embodiments, the 3' adapter sequence comprises a second (e.g., R2) sequencing primer site, and the 5' adapter sequence comprises a first (e.g., R1) sequencing primer site. It is understood that the "R1" and "R2" sequencing primer sites are used by a sequencing system to produce paired-end reads, e.g., reads from opposite ends of a DNA fragment to be sequenced. In some embodiments, the R1 sequencing primer is used to produce a first population of reads from a first end of a DNA fragment, and the R2 sequencing primer is used to produce a second population of reads from the opposite end of the DNA fragment. The first population is referred to herein as "R1" or "read 1" reads. The second population is referred to herein as "R2" or "read 2" reads. The R1 and R2 reads can be aligned as a "read pair" or "mate pair" that corresponds to each strand of a double-stranded analyte DNA fragment.

[0069] Certain sequencing systems, such as Illumina, utilize what are referred to as "R1" and "R2" primers, and "R1" and "R2" reads. For purposes of this application, it should be noted that the terms "R1" and "R2", and "read 1" and "read 2" are not limited to the manner in which they are referenced in relation to a particular sequencing platform. For example, if an Illumina sequencer is used, the "R2" primer and corresponding R2 read disclosed herein may refer to the Illumina "R2" primer and read, or may refer to the Illumina "R1" primer and read insofar as the "R1" primer and corresponding R1 read disclosed herein refer to other Illumina primers and reads. For clarity, in some embodiments in which the "R2" primer provided herein is an Illumina "R1" primer that produces an "R1" read, the corresponding "R1" primer provided herein is an Illumina "R2" primer that produces an "R2" read. For clarity, in some embodiments, the "R2" primer provided herein is an Illumina "R2" primer that provides an "R2" read, and the "R1" primer provided herein is an Illumina "R1" primer that provides an R1 read.

[0070] In some embodiments, the exogenous UID is unique to each double-stranded DNA fragment in a nucleic acid sample. In some embodiments, the exogenous UID is not unique to each double-stranded DNA fragment.

[0071] In some embodiments, the exogenous UID has a length. The length can be about 2-4000 nt. The length can be about 6-100 nt. The length can be about 8-50 nt. The length can be about 10-20 nt. The length can be about 12-14 nt. In some embodiments, the length of the exogenous UID is sufficient to uniquely barcode the molecule, and the length / sequence of the exogenous UID does not interfere with downstream amplification steps.

[0072] In some embodiments, the exogenous UID sequence is not present in the nucleic acid template. In some embodiments, the exogenous UID sequence is not present in the desired template carrying the desired locus. Such unique sequences can be randomly generated, for example, by a computer-readable medium, and selected, for example, by BLAST against known nucleotide databases, such as EMBL (European Molecular Biology Laboratory), GenBank, or DDBJ (DNA Data Bank of Japan). In some embodiments, the exogenous UID sequence is present in the nucleic acid template. In such cases, the position of the exogenous UID sequence in the sequence reads is used to distinguish the exogenous UID sequence from sequences within the nucleic acid template.

[0073] In some embodiments, the exogenous UID sequence is random. In some embodiments, the exogenous UID sequence is a random N-mer. For example, if the exogenous UID sequence has a length of 6 nt, then it can be a random 6-mer. If the exogenous UID sequence has a length of 12 nt, then it can be a random 12-mer.

[0074] Exogenous UIDs may be created using random addition of nucleotides to form a sequence having a length that is used as an identifier. At each addition position, a selection of one of four deoxyribonucleotides may be used. Alternatively, a selection of one of three, two, or one deoxyribonucleotides may be used. Thus, UIDs may be fully random, somewhat random, or non-random at certain positions.

[0075] In some embodiments, the exogenous UIDs are not random N-mers, but are selected from a predetermined set of exogenous UID sequences.

[0076] Exemplary exogenous UIDs suitable for use in the methods disclosed herein are described in PCT / US2012 / 033207, which is incorporated by reference in its entirety.

[0077] The forked adaptors described herein can be attached to double-stranded DNA fragments by any means known in the art.

[0078] In some embodiments, the forked adaptors are attached to the double-stranded DNA fragments by: (a) attaching a partially double-stranded 3' adaptor to the 3' ends of both the Watson and Crick strands of the population of double-stranded DNA fragments, where a first strand of the partially double-stranded 3' adaptor comprises, in a 5'-3' direction, a universal 3' adaptor sequence that comprises the following: (i) a first segment, (ii) an exogenous UID sequence, (iii) an annealing site for the 5' adaptor, and (iv) an R2 sequencing primer site, and where a second strand of the partially double-stranded 3' adaptor comprises, in a 5'-3' direction, a segment complementary to the first segment, and (ii) an exogenous UID sequence. ) a 3' blocking group, optionally where the second strand is degradable; (b) annealing a 5' adapter to the 3' adapter via an annealing site, where the 5' adapter comprises, in a 5' to 3' direction, the following: (i) a universal 5' adapter sequence that is not complementary to the universal 3' adapter sequence and includes an R1 sequencing primer site, and (ii) a sequence that is complementary to the annealing site of the 5' adapter; and (c) performing a nick translation reaction to extend the 5' adapter across the exogenous UID sequence of the 3' adapter and covalently link the extended 5' adapter to the 5' ends of the Watson and Crick strands of the double-stranded DNA fragment.

[0079] In some embodiments, the forked adaptors are attached to the double-stranded DNA fragments by: (a) attaching a 3' duplex adaptor to the 3' end of both the Watson strand and the Crick strand of the population of double-stranded DNA fragments. As described herein, a 3' duplex adaptor, also referred to herein as a partially double-stranded 3' adaptor, is an oligonucleotide complex that includes a molecular barcode that allows a first oligonucleotide (also referred to herein as the "first strand") to anneal (hybridize) to a second oligonucleotide (also referred to herein as the "second strand") such that a portion (e.g., a first portion) of the 3' duplex DNA fragment is double-stranded and a portion (e.g., a second portion) of the 3' duplex adaptor is single-stranded. In some cases, the first oligonucleotide of the 3' duplex adaptor described herein includes a first segment that includes nucleotides that are complementary to nucleotides present in the second oligonucleotide of the 3' duplex adaptor (e.g., the first oligonucleotide of the 3' duplex adaptor and the second oligonucleotide of the 3' duplex adaptor are capable of annealing at complementary regions).

[0080] The first oligonucleotide of the 3' duplex adapter described herein can be an oligonucleotide comprising a 5' phosphate and a molecular barcode. The first oligonucleotide of the 3' duplex adapter described herein can comprise any suitable number of nucleotides. Any suitable molecular barcode can be comprised in the first oligonucleotide of the 3' duplex adapter described herein. In some cases, the molecular barcode can comprise a random sequence. In some cases, the molecular barcode can comprise a fixed sequence. Examples of molecular barcodes that can be comprised in the first oligonucleotide of the 3' duplex adapter described herein include, without limitation, IDT 8, IDT 10, ILMN 8, ILMN 10 available from Integrated DNA technologies. Any suitable type of molecular barcode can be used. In some cases, the molecular barcode comprises an exogenous UID sequence. Exogenous UIDs are described herein. An example of an oligonucleotide that includes a 5' phosphate and a molecular barcode and can be included in the first oligonucleotide of a 3' duplex adapter described herein includes, but is not limited to, ATAAAACGACGGCNNNNNNNNNNNNNNAGATCGGAAGAGCACACGTCTGAACTCCAG*T*C (the asterisk represents a phosphorothioate linkage; SEQ ID NO: 1), where NNNNNNNNNNNNNN (SEQ ID NO: 2) is the molecular barcode, and where the number of nucleotides in the molecular barcode can be from 0 to about 25.

[0081] In some embodiments, the first oligonucleotide of the 3' duplex adapter contains an annealing site for the 5' adapter.

[0082] In some embodiments, the first oligonucleotide of the 3' duplex adapter comprises a universal 3' adapter sequence. In some embodiments, the universal 3' adapter sequence comprises an R2 sequencing primer site.

[0083] In some cases, the first oligonucleotide of the 3' duplex adapter described herein can also include one or more features to prevent or reduce extension during PCR. The feature that can prevent or reduce extension during PCR can be any type of feature (e.g., chemical modification). Examples of features that can prevent or reduce extension during PCR and can be included in the first oligonucleotide of the 3' duplex adapter described herein include, but are not limited to, 3SpC3 and 3Phos. The feature that can prevent or reduce extension during PCR can be incorporated into the first oligonucleotide of the 3' duplex adapter described herein at any suitable position within the oligonucleotide. In some cases, the molecule that prevents or reduces extension during PCR can be incorporated internally within the oligonucleotide. In some cases, the molecule that prevents or reduces extension during PCR can be incorporated into the oligonucleotide and at its end (e.g., 5' end).

[0084] In certain embodiments, the first oligonucleotide of the 3' duplex adapter comprises a 5' phosphate, a first segment comprising a nucleotide complementary to a nucleotide present in the second oligonucleotide of the 3' duplex adapter, an exogenous UID sequence, an annealing site for the 5' adapter, and a universal 3' adapter sequence.

[0085] The second oligonucleotide of the 3' duplex adapter described herein can be an oligonucleotide that includes a blocked 3' group (e.g., to reduce or eliminate dimerization of the two adapters). The second oligonucleotide of the 3' duplex adapter described herein can include any suitable number of nucleotides. In some embodiments, the second oligonucleotide of the 3' duplex adapter is complementary to the first segment of the first oligonucleotide of the 3' duplex adapter. Exemplary oligonucleotides that include a blocked 3' group and can be included in the second oligonucleotide of the 3' duplex adapter described herein include, but are not limited to, GCCGUCGUUUUAdT (SEQ ID NO: 3).

[0086] The second oligonucleotide of the 3' duplex adapter described herein can be degradable. Any suitable method can be used to degrade the second oligonucleotide of the 3' duplex adapter described herein. For example, UDG can be used to degrade the second oligonucleotide of the 3' duplex adapter described herein.

[0087] In some embodiments, the 3' duplex adapter described herein can comprise a first oligonucleotide comprising the sequence ATAAAACGACGGCNNNNNNNNNNNNNNAGATCGGAAGAGCACACGTCTGAACTCCAG*T*C / 3SpC3 (SEQ ID NO:1) annealed to a second oligonucleotide comprising the sequence GCCGUCGUUUUAdT (SEQ ID NO:3).

[0088] In some cases, the 3' duplex adapters described herein can include commercially available adapters. Exemplary commercially available adapters that can be used (or used to generate) as the 3' duplex adapters described herein include, but are not limited to, the adapters in the Accel-NGS 2S DNA Library Kit (Swift Biosciences, cat. #21024).

[0089] The 3' adaptors can be attached (e.g., covalently attached) to the 3' ends of the double-stranded DNA fragments using any suitable method. In some embodiments, the 3' adaptors are attached by ligation. In some embodiments, the ligation comprises the use of a ligase. Examples of ligases that can be used to attach the 3' adaptors to the 3' ends of each nucleic acid fragment include, but are not limited to, T4 DNA ligase, E. coli ligase (e.g., enzyme Y3), CircLigase I, CircLigase II, Taq-ligase, T3 ligase, T7 ligase, and 9N ligase.

[0090] Once the 3' duplex adapter is attached (e.g., covalently attached) to the 3' end of each nucleic acid fragment, the second oligonucleotide of the 3' duplex adapter described herein can be degraded, and a 5' adapter can be attached (e.g., covalently attached) to the 5' end of each nucleic acid fragment. In some embodiments, the 5' adapter sequence is not complementary to the first oligonucleotide of the 3' adapter. In some embodiments, the 5' adapter sequence comprises, in the 5' to 3' direction, an R1 sequencing primer site and a sequence complementary to the annealing site of the 3' adapter.

[0091] In some embodiments, attaching the 5' adaptor comprises annealing the 5' adaptor to the 3' adaptor via an annealing site.

[0092] The 5' adaptor can be annealed to the nucleic acid fragment upstream of the molecular barcode on the 3' duplex adaptor such that a gap (e.g., a single-stranded nucleic acid fragment) is present on the nucleic acid fragment that includes a portion of the 3' duplex adaptor (e.g., a molecular barcode). The gap including a portion of the 3' duplex adaptor can be filled (e.g., to generate a double-stranded nucleic acid fragment). Any suitable method can be used to fill the single-stranded gap. Examples of methods that can be used to fill the single-stranded gap on the nucleic acid fragment include, but are not limited to, a polymerase such as a DNA polymerase (e.g., Taq polymerase such as Taq-B polymerase) and a nick translation reaction (including both a ligase such as E. coli ligase and a polymerase such as a DNA polymerase). When filling the single-stranded gap on the nucleic acid fragment includes providing a polymerase, the method can also include providing deoxyribonucleotide triphosphates (dNTPs; e.g., dATP, dGTP, dCTP, and dTTP). In some cases, the attachment of 5' adaptors to the 5' end of each nucleic acid fragment and the filling of single-stranded gaps can be performed simultaneously (eg, in a single reaction tube).

[0093] In some cases, alternative methods can be used to attach adapters to templates. For example, nucleic acid fragments can be treated with single-stranded nucleases (e.g., to digest overhangs), and then ligation can be used to prepare a double-stranded sequencing library. For example, a single nucleotide can be added to the 3' end of each nucleic acid fragment, and an adapter (e.g., containing a molecular barcode) containing a complementary base at the 5' end can be ligated to each nucleic acid fragment to prepare a double-stranded sequencing library of adapter-attached templates.

[0094] Molecular barcoding

[0095] As used herein, "molecular barcode" refers to a barcode that serves to identify individual nucleic acid fragments in an original sample prior to barcoding and amplification. In some embodiments, each individual nucleic acid fragment has a unique molecular barcode. In some embodiments, the barcode may be a randomly generated nucleotide sequence or a purposefully chosen nucleotide run. In particular, when attaching molecular barcodes, the number of individual molecular barcodes in the reaction mixture exceeds the number of nucleic acid fragments.

[0096] In some embodiments, the molecular barcode is unique to each double-stranded DNA fragment in the nucleic acid sample. In some embodiments, the molecular barcode comprises an intrinsic barcode, an exogenous barcode, or both.

[0097] In some embodiments, the molecular barcodes are from about 2 to about 4000 in length (e.g., from about 2 to about 3500, from about 2 to about 3000, from about 2 to about 2500, from about 2 to about 2000, from about 2 to about 1500, from about 2 to about 1000, from about 2 to about 500, from about 2 to about 100, from about 2 to about 50, from about 2 to about 20, from about 2 to about 10, from about 10 to about 4000, from about 10 to about 3500, from about 10 to about 3000, from about 10 to about 2500, from about 10 to about 2000, from about 10 to about 1500, from about 10 to about 2000, from about 1000, from about 10 to about 500, from about 10 to about 100, from about 10 to about 50, from about 10 to about 20, from about 20 to about 4000, from about 20 to about 3500, from about 20 to about 3000, from about 20 to about 2500, from about 20 to about 2000, from about 20 to about 1500, from about 20 to about 1000, from about 20 to about 500, from about 20 to about 100, from about 20 to about 50, from about 50 to about 4000, from about 50 to about 3500, from about 50 to about 3000, from about 50 to about 2500, from about 50 to about 2000, 0 to about 1500, about 50 to about 1000, about 50 to about 500, about 50 to about 100, about 100 to about 4000, about 100 to about 3500, about 100 to about 3000, about 100 to about 2500, about 100 to about 2000, about 100 to about 1500, about 100 to about 1000, about 100 to about 500, about 500 to about 4000, about 500 to about 3500, about 500 to about 3000, about 500 to about 2500, about 500 to about 2000, about 500 to about 1500, about 500 from about 1000 to about 4000, from about 1000 to about 3500, from about 1000 to about 3000, from about 1000 to about 2500, from about 1000 to about 2000, from about 1000 to about 1500, from about 1500 to about 4000, from about 1500 to about 3500, from about 1500 to about 3000, from about 1500 to about 2500, from about 1500 to about 2000, from about 2000 to about 4000, from about 2000 to about 3500, from about 2000 to about 3000, from about 2000 to about 2500, from about 2500 to about 4000,From about 2500 to about 3500, from about 2500 to about 3000, from about 3000 to about 4000, from about 3000 to about 3500, or from about 3500 to about 4000 nucleotides. In some embodiments, the length of the molecular barcode is sufficient to uniquely barcode a molecule, and the length / sequence of the molecular barcode does not interfere with downstream amplification steps.

[0098] In some embodiments, the molecular barcode sequence can be random. In some embodiments, the molecular barcode sequence can be a random N-mer. For example, if the molecular barcode sequence has a length of 6 nt, then it can be a random 6-mer. If the molecular barcode sequence has a length of 12 nt, then it can be a random 12-mer.

[0099] In some embodiments, molecular barcodes can be created using random addition of nucleotides to form a sequence having a length that is used as an identifier. At each addition position, a selection from one of four deoxyribonucleotides can be used. Alternatively, a selection from one of three, two, or one deoxyribonucleotides can be used. Thus, molecular barcodes can be fully random, somewhat random, or non-random at certain positions. In some embodiments, the molecular barcodes are not random N-mers, but are selected from a predefined set of molecular barcode sequences. Exemplary molecular barcodes suitable for use in the methods disclosed herein are described in PCT / US2012 / 033207, which is incorporated herein by reference in its entirety.

[0100] Attachment of molecular barcodes to nucleic acid fragments can be performed by any means known in the art, including enzymatic, chemical, or biological. In some embodiments, one means employs polymerase chain reaction. In some embodiments, another means uses a ligase enzyme. For example, the ligase enzyme may be mammalian or bacterial. Other enzymes that may be used for attachment are other polymerase enzymes. Molecular barcodes may be added to one or both ends of the fragment, preferably both ends. In some embodiments, molecular barcodes may be included within nucleic acid molecules that contain other regions for other intended functionality. For example, universal priming sites may be added to allow for later amplification. In some embodiments, another additional site may be a region of complementarity to a specific region or gene in the nucleic acid fragment.

[0101] Initial amplification of adaptor-attached templates

[0102] After adapter attachment, the adapter-attached template can be amplified in an initial amplification reaction (e.g., PCR amplification). Any suitable method can be used to amplify the adapter-attached template. Exemplary methods that can be used to amplify the adapter-attached template include, without limitation, whole genome PCR.

[0103] Any suitable primer pair can be used for amplifying the adaptor-attached template. In some cases, a universal primer pair can be used. The primers can include, without limitation, from about 12 nucleotides to about 30 nucleotides. Examples of primer pairs that can be used to amplify the adaptor-attached templates described herein include, without limitation, those described in Example 4.

[0104] Any suitable PCR conditions can be used for the initial amplification. PCR amplification can include a denaturation step, an annealing step, and an extension step. Each step of the amplification cycle can include any suitable conditions. In some cases, the denaturation step can include a temperature of about 90°C to about 105°C (e.g., about 94°C to about 98°C) and a time of about 1 second to about 5 minutes (e.g., about 10 seconds to about 1 minute). For example, the denaturation step can include a temperature of about 98°C for about 10 seconds. In some cases, the annealing step can include a temperature of about 50°C to about 72°C and a time of about 30 seconds to about 90 seconds. In some cases, the extension step can include a temperature of about 55°C to about 80°C and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the annealing and extension steps can be performed in a single cycle. For example, the annealing and phase extension phase can include a temperature of about 65° C. for about 75 seconds.

[0105] The PCR conditions used in the initial amplification can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can include from about 1 to about 50 cycles. In some embodiments, the PCR amplification includes no more than 11 cycles. In some embodiments, the PCR amplification includes no more than 7 cycles. In some embodiments, the PCR amplification includes no more than 5 cycles.

[0106] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0107] In some cases, the PCR amplification can also include a hold step. For example, the PCR amplification can include a hold step after performing PCR amplification cycles and, optionally, after performing any final extension step. In some cases, the hold step can include a temperature of from about 4° C. to about 15° C. for an indefinite period of time.

[0108] In some cases, the double-stranded sequencing library (e.g., the amplified double-stranded sequencing library) generated as described herein can be purified. Any suitable method can be used to purify the double-stranded sequencing library. Exemplary methods that can be used to purify the double-stranded sequencing library include, without limitation, magnetic beads (e.g., solid-phase reversible immobilization (SPRI) magnetic beads).

[0109] Optional ssDNA library preparation

[0110] In some cases, the double-stranded sequencing library can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands. By generating a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands, non-specific amplification (e.g., from a primer complementary to a linking sequence such as a 3' double-stranded adaptor or a 5' adaptor) can be minimized. Any suitable method can be used to generate a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands (e.g., from a double-stranded sequencing library generated as described herein). In some cases, a library of sequences derived from single-stranded Watson strands and a library of sequences derived from single-stranded Crick strands can be generated from an amplified double-stranded sequencing library by dividing the amplification product into at least two aliquots and subjecting each aliquot to PCR amplification, in which the Watson strand is amplified from the first aliquot and the Crick strand is amplified from the second aliquot. For example, a first aliquot of the amplified products from the amplified double-stranded sequencing library can be subjected to PCR amplification with a primer pair, where the first primer is biotinylated and the second primer is non-biotinylated to generate a single-stranded library of Watson strands, and a second aliquot of the amplified products from the amplified double-stranded sequencing library can be subjected to PCR amplification with a primer pair, where the first primer is non-biotinylated and the second primer is biotinylated to generate a single-stranded library of Crick strands. In some cases, a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences can be generated.

[0111] Any suitable method can be used to generate a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand from the amplified double-stranded sequencing library. For example, the amplification products from the amplified double-stranded sequencing library can be divided into a first PCR amplification and a second PCR amplification in which only one of the two primers of the PCR primer pair is tagged. For example, the first PCR amplification can use a primer pair including a tagged primer (e.g., the first primer) and an untagged primer (e.g., the second primer), and the second PCR amplification can use a primer pair including an untagged primer (e.g., the first primer) and a tagged primer (e.g., the second primer). The primer tag can be any tag that allows the PCR amplification product generated from the tagged primer to be recovered. In some cases, the tagged primer can be a biotinylated primer, and the PCR amplification product generated from the biotinylated primer can be recovered using streptavidin. For example, in PCR amplification using a primer pair that comprises biotinylated primer and non-biotinylated primer, a library of sequences from single-stranded Watson strand and a library of sequences from single-stranded Crick strand can be generated.In some cases, tagged primer can be phosphorylated primer, and the PCR amplification product generated from phosphorylated primer can be recovered using lambda nuclease.For example, in PCR amplification using a primer pair that comprises phosphorylated primer and non-phosphorylated primer, a library of sequences from single-stranded Watson strand and a library of sequences from single-stranded Crick strand can be generated.

[0112] Any suitable primer pair can be used to generate a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences (e.g., from a double-stranded sequencing library generated as described herein). The primers can include, without limitation, from about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair can include at least one primer that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in the amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in a double-stranded sequencing library prior to amplification). Examples of primer pairs that can be used to generate a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences described herein include, without limitation, P5 primer and P7 primer.

[0113] Any suitable PCR conditions can be used to generate a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand (e.g., from a double-stranded sequencing library generated as described herein). PCR amplification can include a denaturation step, an annealing step, and an extension step. Each step of an amplification cycle can include any suitable conditions. In some cases, the denaturation step can include a temperature of about 90°C to about 105°C and a time of about 1 second to about 5 minutes. For example, the denaturation step can include a temperature of about 98°C for 10 seconds. In some cases, the annealing step can include a temperature of about 50°C to about 72°C and a time of about 30 seconds to about 90 seconds. In some cases, the extension step can include a temperature of about 55°C to about 80°C and a time of about 15 seconds per kb of amplicon generated to about 30 seconds per kb of amplicon generated. In some cases, the extension step reflects the processivity of the polymerase used. In some cases, the annealing and extension steps can be performed in a single cycle. For example, the annealing and phase extension steps can include a temperature of about 65° C. for about 75 seconds.

[0114] The PCR conditions used to generate a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences (e.g., from a double-stranded sequencing library generated as described herein) can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can include, without limitation, from about 1 to about 50 cycles. For example, the PCR amplification can include about 4 amplification cycles.

[0115] In some cases, when the PCR conditions include a heat-activated polymerase, the PCR amplification can also include an initialization step. For example, the PCR amplification can include an initialization step before performing a PCR amplification cycle. In some cases, the initialization step can include a temperature of about 94°C to about 98°C and a time period of about 15 seconds to about 1 minute. For example, the initialization step can include a temperature of about 98°C for about 30 seconds.

[0116] In some cases, the PCR amplification may also include a hold step. For example, the PCR amplification may include a hold step after performing PCR amplification cycles, and optionally after any final extension step. In some cases, the hold step may include a temperature of from about 4° C. to about 15° C. for an indefinite period of time.

[0117] Any suitable method can be used to separate double-stranded amplification products into single-stranded amplification products.In some cases, double-stranded amplification products can be denatured to separate double-stranded amplification products into two single-stranded amplification products.Examples of the method that can be used to separate double-stranded amplification products into single-stranded amplification products include, without limitation, heat denaturation, chemical (e.g., NaOH) denaturation, and salt denaturation.

[0118] After PCR amplification, the tagged PCR amplification product can be collected. Any suitable method can be used to collect the tagged PCR amplification product generated using the tagged primer. When the tagged primer is a biotinylated primer, the biotinylated amplification product (e.g., generated from the biotinylated primer) can be collected using streptavidin (e.g., streptavidin-functionalized beads). For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, the biotinylated amplification products generated from the first PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the biotinylated amplification products generated from the second PCR amplification can be bound to streptavidin-functionalized beads (e.g., a first set of streptavidin-functionalized beads), and the double-stranded amplification products can be separated (e.g., denatured) into single strands of amplification products.In some cases, recovering biotinylated PCR amplification products can also include releasing biotinylated PCR amplification products from streptavidin (e.g., streptavidin-functionalized beads). Separating the double-stranded amplification products generated by a first PCR amplification using a primer pair comprising a first biotinylated primer and a second non-biotinylated primer, and a second PCR amplification using a primer pair comprising a first non-biotinylated primer and a second biotinylated primer, can allow the single-stranded amplification products generated from the biotinylated primer to remain bound to the streptavidin-functionalized beads, while the single-stranded amplification products generated from the non-biotinylated primer can be denatured (e.g., denatured and degraded), thereby generating a library of single-stranded Watson strand-derived sequences and a library of single-stranded Crick strand-derived sequences of the double-stranded sequencing library.

[0119] When the tagged primer is a phosphorylated primer, the phosphorylated amplification product (e.g., generated from the phosphorylated primer) can be recovered using an exonuclease (e.g., lambda exonuclease).For example, when the amplified double-stranded sequencing library is further amplified in a first PCR amplification using a primer pair comprising a first phosphorylated primer and a second non-phosphorylated primer, and in a second PCR amplification using a primer pair comprising a first non-phosphorylated primer and a second phosphorylated primer, the double-stranded amplification product can be separated into single-stranded amplification products. Separating the double-stranded amplification products generated by the first PCR amplification using a primer pair comprising a first phosphorylated primer and a second unphosphorylated primer, and the second PCR amplification using a primer pair comprising a first unphosphorylated primer and a second phosphorylated primer, allows the single-stranded amplification products generated from the unphosphorylated primer to be recovered, while the single-stranded amplification products generated from the phosphorylated primer can be degraded by lambda exonuclease, thereby generating a library of sequences derived from the single-stranded Watson strand and a library of sequences derived from the single-stranded Crick strand of the double-stranded sequencing library.

[0120] Target Enrichment

[0121] In some embodiments of any one of the methods herein, the amplicons produced by the initial amplification are enriched for one or more target polynucleotides. In some embodiments, prior to target enrichment, a single-stranded DNA library is prepared from the amplicons produced by the initial amplification. Exemplary methods for producing a single-stranded DNA library are described herein.

[0122] Any suitable method can be used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand generated as described herein). In some cases, a target region can be amplified from a library of amplification products by subjecting the library of amplification products to PCR amplification with a primer pair, where the primer (e.g., a first primer) can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in the double-stranded sequencing library prior to amplification), and the primer (e.g., a second primer) can target (e.g., target and bind to) a target region (e.g., a region of interest).

[0123] In some cases, the target region can be amplified from a library of amplification products (e.g., a duplex sequencing library, a library of sequences derived from a single stranded Watson strand, or a library of sequences derived from a single stranded Crick strand generated as described herein) in a single PCR amplification. For example, the target region can be amplified from a library of amplification products in a single PCR amplification using a primer pair that includes a first primer that can target an adapter sequence (e.g., an adapter sequence that includes a molecular barcode) present in an amplification product generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification) and a second primer that can target the target region.

[0124] In some cases, the target region can be amplified from a library of amplification products (e.g., a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand generated as described herein) in a multiplexed PCR amplification. Multiplexed PCR amplification (e.g., a first PCR amplification followed by a nested PCR amplification) can be used to increase the specificity of amplifying the target region. For example, a target region can be amplified from a library of amplification products in a series of PCR amplifications, where a first PCR amplification uses a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting the target region, and the amplification products generated in the first PCR amplification are subjected to a subsequent nested PCR amplification using a primer pair comprising a first primer capable of targeting an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in the amplification products generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a second primer capable of targeting a nucleic acid sequence from the target region present in the amplification products generated in the first PCR amplification.

[0125] Any suitable primer pair can be used to amplify a target region from a library of amplification products (e.g., a duplex sequencing library, a library of sequences derived from a single stranded Watson strand, or a library of sequences derived from a single stranded Crick strand generated as described herein). The primers can comprise, without limitation, from about 12 nucleotides to about 30 nucleotides. In some cases, the primer pair comprises a primer (e.g., a first primer) that can target (e.g., target and bind to) an adapter sequence (e.g., an adapter sequence comprising a molecular barcode) present in an amplification product generated as described herein (e.g., by ligating a 3' duplex adapter comprising a first molecular barcode and a 5' adapter comprising a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification), and a primer (e.g., a second primer) that can target (e.g., target and bind to) a target region (e.g., a region of interest). Examples of primers that can target adapter sequences that include molecular barcodes present in amplification products generated as described herein (e.g., by ligating a 3' duplex adapter that includes a first molecular barcode and a 5' adapter that includes a second molecular barcode to a nucleic acid fragment in a duplex sequencing library prior to amplification) include, without limitation, i5 index primers and i7 index primers. A primer that can target a target region can include a sequence complementary to the target region.

[0126] In some embodiments, when the target region is a nucleic acid encoding an immune cell receptor, the primer capable of targeting the target region comprises a sequence complementary to the sequence of the immune cell receptor. In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region comprises a sequence complementary to the sequence of the T cell receptor. In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 4 (e.g., a nucleic acid sequence comprising SEQ ID NO: 4). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 5 (e.g., a nucleic acid sequence comprising SEQ ID NO: 5). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 6 (e.g., a nucleic acid sequence comprising SEQ ID NO: 6). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 7 (e.g., a nucleic acid sequence comprising SEQ ID NO: 7). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO:8 (e.g., a nucleic acid sequence comprising SEQ ID NO:8). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO:9 (e.g., a nucleic acid sequence comprising SEQ ID NO:9). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO:10 (e.g., a nucleic acid sequence comprising SEQ ID NO:10). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO:11 (e.g., a nucleic acid sequence comprising SEQ ID NO:11).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 12 (e.g., a nucleic acid sequence comprising SEQ ID NO: 12). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 13 (e.g., a nucleic acid sequence comprising SEQ ID NO: 13). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 14 (e.g., a nucleic acid sequence comprising SEQ ID NO: 14). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 15 (e.g., a nucleic acid sequence comprising SEQ ID NO: 15). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 16 (e.g., a nucleic acid sequence comprising SEQ ID NO: 16). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 17 (e.g., a nucleic acid sequence comprising SEQ ID NO: 17). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 18 (e.g., a nucleic acid sequence that includes SEQ ID NO: 18). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 19 (e.g., a nucleic acid sequence that includes SEQ ID NO: 19). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 20 (e.g., a nucleic acid sequence that includes SEQ ID NO: 20).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 21 (e.g., a nucleic acid sequence comprising SEQ ID NO: 21). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 22 (e.g., a nucleic acid sequence comprising SEQ ID NO: 22). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 23 (e.g., a nucleic acid sequence comprising SEQ ID NO: 23). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 24 (e.g., a nucleic acid sequence comprising SEQ ID NO: 24). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 25 (e.g., a nucleic acid sequence comprising SEQ ID NO: 25). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 26 (e.g., a nucleic acid sequence comprising SEQ ID NO: 26). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 27 (e.g., a nucleic acid sequence that includes SEQ ID NO: 27). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 28 (e.g., a nucleic acid sequence that includes SEQ ID NO: 28). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 29 (e.g., a nucleic acid sequence that includes SEQ ID NO: 29).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 30 (e.g., a nucleic acid sequence comprising SEQ ID NO: 30). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 31 (e.g., a nucleic acid sequence comprising SEQ ID NO: 31). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 32 (e.g., a nucleic acid sequence comprising SEQ ID NO: 32). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 33 (e.g., a nucleic acid sequence comprising SEQ ID NO: 33). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 34 (e.g., a nucleic acid sequence comprising SEQ ID NO: 34). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 35 (e.g., a nucleic acid sequence comprising SEQ ID NO: 35). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 36 (e.g., a nucleic acid sequence that includes SEQ ID NO: 36). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 37 (e.g., a nucleic acid sequence that includes SEQ ID NO: 37). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 38 (e.g., a nucleic acid sequence that includes SEQ ID NO: 38).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 39 (e.g., a nucleic acid sequence comprising SEQ ID NO: 39). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 40 (e.g., a nucleic acid sequence comprising SEQ ID NO: 40). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 41 (e.g., a nucleic acid sequence comprising SEQ ID NO: 41). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 42 (e.g., a nucleic acid sequence comprising SEQ ID NO: 42). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 43 (e.g., a nucleic acid sequence comprising SEQ ID NO: 43). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 44 (e.g., a nucleic acid sequence comprising SEQ ID NO: 44). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 45 (e.g., a nucleic acid sequence that includes SEQ ID NO: 45). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 46 (e.g., a nucleic acid sequence that includes SEQ ID NO: 46). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 47 (e.g., a nucleic acid sequence that includes SEQ ID NO: 47).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 48 (e.g., a nucleic acid sequence comprising SEQ ID NO: 48). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 49 (e.g., a nucleic acid sequence comprising SEQ ID NO: 49). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 50 (e.g., a nucleic acid sequence comprising SEQ ID NO: 50). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 51 (e.g., a nucleic acid sequence comprising SEQ ID NO: 51). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 52 (e.g., a nucleic acid sequence comprising SEQ ID NO: 52). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 53 (e.g., a nucleic acid sequence comprising SEQ ID NO: 53). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 54 (e.g., a nucleic acid sequence comprising SEQ ID NO: 54). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 55 (e.g., a nucleic acid sequence comprising SEQ ID NO: 55). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 56 (e.g., a nucleic acid sequence comprising SEQ ID NO: 56). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 57 (e.g., a nucleic acid sequence comprising SEQ ID NO: 57). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 58 (e.g., a nucleic acid sequence comprising SEQ ID NO: 58). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 59 (e.g., a nucleic acid sequence comprising SEQ ID NO: 59). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 60 (e.g., a nucleic acid sequence that includes SEQ ID NO: 60). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 61 (e.g., a nucleic acid sequence that includes SEQ ID NO: 61). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, a primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 62 (e.g., a nucleic acid sequence that includes SEQ ID NO: 62).In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 63 (e.g., a nucleic acid sequence comprising SEQ ID NO: 63). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 64 (e.g., a nucleic acid sequence comprising SEQ ID NO: 64). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 65 (e.g., a nucleic acid sequence comprising SEQ ID NO: 65). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 66 (e.g., a nucleic acid sequence comprising SEQ ID NO: 66). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 67 (e.g., a nucleic acid sequence comprising SEQ ID NO: 67). In some embodiments, when the target region is a nucleic acid encoding a T cell receptor, the primer capable of targeting the target region can comprise the nucleic acid sequence of SEQ ID NO: 68 (e.g., a nucleic acid sequence comprising SEQ ID NO: 68).

[0127] [Table 1-1] [Table 1-2]

[0128] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand generated as described herein) can include one or more molecular barcodes.

[0129] In some cases, one or both primers of a primer pair used to amplify a target region from a library of amplification products (e.g., a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand generated as described herein) can include one or more graft sequences (e.g., graft sequences for next-generation sequencing).

[0130] In one embodiment, target enrichment comprises (a) selectively amplifying a Watson strand amplicon comprising a target polynucleotide sequence with a first set of Watson target selective primer pairs, the first set of Watson target selective primer pairs comprising: (i) a first Watson target selective primer comprising a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a second Watson target selective primer comprising a target selective sequence, thereby generating a target Watson amplification product; and (b) selectively amplifying a Click strand amplicon comprising the same target polynucleotide sequence with a first set of click target selective primer pairs, the first set of click target selective primer pairs comprising: (i) a first click target selective primer comprising a sequence complementary to the R1 sequencing primer site of the universal 5' adapter sequence, and (ii) a second click target selective primer comprising the same target selective sequence as the second Watson target selective primer sequence, thereby generating a target click amplification product.

[0131] In some embodiments, the method further comprises purifying the target Watson amplification products and the target click amplification products from non-target polynucleotides. In some embodiments, the purification comprises attaching the target Watson amplification products and the target click amplification products to a solid support. In some embodiments, the first Watson target selective primer and the first click target selective primer comprise a first member of an affinity binding pair, and where the solid support comprises a second member of the affinity binding pair. In some embodiments, the first member is biotin, and the second member is streptavidin. In some embodiments, the solid support comprises a bead, a well, a membrane, a tube, a column, a plate, sepharose, a magnetic bead, or a chip. In some embodiments, the method comprises removing polynucleotides that are not attached to the solid support.

[0132] In some embodiments, the method includes (a) further amplifying the target Watson amplification product with a second Watson target selective primer set, the second Watson target selective primer set including: (i) a third Watson target selective primer that includes a sequence complementary to the R2 sequencing primer site of the universal 3' adapter sequence, and (ii) a fourth Watson target selective primer that includes, in the 5' to 3' direction, an R1 sequencing primer site and a target selective sequence selective for the same target polynucleotide, thereby creating a targeted Watson library member; (b) ) further amplifying the target click amplification products with a second click target selective primer set, the second click target selective primer set being: (i) a third click target selective primer that comprises a sequence complementary to the R1 sequencing primer site of the universal 3' adapter sequence, and (ii) a fourth click target selective primer that comprises, in the 5' to 3' direction, the R2 sequencing primer site and a target selective sequence selective for the same target polynucleotide of the fourth Watson target selective primer, thereby creating a targeted click library member.

[0133] In some embodiments, the third Watson and Click target selective primer further comprises a sample barcode sequence. In some embodiments, the third Watson target selective primer further comprises a first grafting sequence that allows hybridization to the first grafting primer on the sequencer, and where the third Click target selective primer further comprises a second grafting sequence that allows hybridization to the second grafting primer on the sequencer. In some embodiments, the fourth Watson target selective primer further comprises a second grafting sequence, and where the fourth Click target selective primer further comprises a first grafting sequence. In some embodiments, the first grafting sequence is a P7 sequence, and where the second grafting sequence is a P5 sequence.

[0134] Any suitable PCR conditions can be used to generate an amplified target region as described herein (e.g., from a library of amplification products such as a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand, etc.). Exemplary PCR conditions are described herein. PCR conditions used to generate an amplified target region as described herein (e.g., from a library of amplification products such as a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand, etc.) can include any suitable number of PCR amplification cycles. In some cases, the PCR amplification can include, without limitation, from about 1 to about 50 cycles. For example, when the PCR amplification of the amplified target region includes a single PCR amplification, the PCR amplification can include about 18 amplification cycles. For example, when the PCR amplification of the amplified target region includes a first PCR amplification and a subsequent nested PCR amplification, the first PCR amplification can include about 18 amplification cycles, and the subsequent nested PCR amplification can include about 10 amplification cycles.

[0135] exemplary target

[0136] Any suitable target region (e.g., a region of interest) can be amplified from a library of amplification products (e.g., a double-stranded sequencing library, a library of sequences derived from a single-stranded Watson strand, or a library of sequences derived from a single-stranded Crick strand generated as described herein) and evaluated for TCR / BCR receptor sequences. In some cases, the target region can be a region of a nucleic acid encoding an immune cell receptor. Examples of target regions that can be amplified and evaluated to determine a sequence include, but are not limited to, nucleic acids encoding pattern recognition receptors (PRRs), toll-like receptors (TLRs), c-type lectin receptors (CLRs), NOD-like receptors (NLRs), RIG-I-like receptors, killer activating receptors (KARs), killer inhibitor receptors (KIRs), complement receptors, Fc receptors, B cell receptors, T cell receptors, and cytokine receptors. In some embodiments, the target region that can be amplified and evaluated can include nucleic acids encoding T cell receptors. In some embodiments, the target region that can be amplified and evaluated can include nucleic acid encoding a B cell receptor.

[0137] Any suitable method can be used to evaluate the target region (e.g., an amplified target region) for determining the TCR / BCR receptor sequence. In some cases, one or more sequencing methods can be used to evaluate the amplified target region for determining the TCR / BCR receptor sequence.

[0138] Sequencing

[0139] In some cases, one or more sequencing methods can be used to evaluate the amplified target regions and determine the TCR / BCR receptor sequence. In some cases, sequencing reads can be used to evaluate the amplified target regions for the TCR / BCR receptor sequence, and both Watson and Crick strands can be used to determine the TCR / BCR receptor sequence. Examples of sequencing methods that can be used to evaluate the amplified target regions of the TCR / BCR receptor sequence as described herein include, without limitation, single-read sequencing, paired-end sequencing, NGS, and deep sequencing. In some embodiments, single-read sequencing includes sequencing over the entire length of the template to generate sequence reads. In some embodiments, the sequencing includes paired-end sequencing. In some embodiments, the sequencing is performed using a massively parallel sequencer. In some embodiments, the massively parallel sequencer is configured to determine sequence reads from both ends of the template polynucleotide.

[0140] Analysis of sequence reads

[0141] In some embodiments, the sequence reads are mapped to a reference genome.

[0142] In some embodiments, sequence reads are assigned to a UID family. A UID family can include sequence reads from amplicons derived from an original template, e.g., an original double-stranded DNA fragment from a nucleic acid sample.

[0143] In some embodiments, each member of a UID family comprises the same exogenous UID sequence. In some embodiments, each member of a UID family further comprises the same endogenous UID sequence. Endogenous UIDs are described herein.

[0144] In some embodiments, each member of a UID family further comprises the same exogenous UID sequence and the same endogenous UID sequence. In some embodiments, the combination of the exogenous UID sequence and the endogenous UID sequence is unique to the UID family. In some embodiments, the combination of the exogenous UID sequence and the endogenous UID sequence is not present in another UID family represented in the nucleic acid sample.

[0145] The number of members of a UID family can depend on the depth of sequencing. In some embodiments, a UID family comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860 In some embodiments, a UID family includes about 2-1000 members, about 2-500 members, about 2-100 members, about 2-50 members, or about 2-20 members.

[0146] In some embodiments, sequence reads for each UID family are assigned to a Watson subfamily and a Crick subfamily. In some embodiments, sequence reads for each UID family are assigned to a Watson and Crick subfamily based on the orientation of the insert with respect to the adapter sequence. In some embodiments, the orientation of the insert with respect to the adapter sequence is determined by how the sequence reads are aligned as a "read pair" or a "mate pair."

[0147] In some embodiments, the assignment of sequence reads to Watson and Crick subfamilies is based on the spatial relationship of the exogenous UID sequence to the R1 and R2 read sequences. In some embodiments, members of the Watson subfamily are characterized by the exogenous UID sequence being downstream of the R2 sequence and upstream of the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous UID sequence being downstream of the R1 sequence and upstream of the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by the exogenous UID sequence being more closely adjacent to the R2 sequence and less closely adjacent to the R1 sequence. In some embodiments, members of the Crick subfamily are characterized by the exogenous UID sequence being more closely adjacent to the R1 sequence and less closely adjacent to the R2 sequence. In some embodiments, members of the Watson subfamily are characterized by an exogenous UID sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of the R2 sequence. In some embodiments, members of the Crick subfamily are characterized by an exogenous UID sequence immediately downstream or within 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, or 1-5 nucleotides of the R1 sequence.

[0148] In some embodiments, a UID subfamily (e.g., a Watson subfamily and / or a Crick subfamily) is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 310, 320, 330, 340, 350, 36 , 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 members. In some embodiments, the UID subfamily (e.g., the Watson subfamily and / or the Crick subfamily) includes about 2-500 members, about 2-100 members, about 2-50 members, about 2-20 members, or about 2-10 members.

[0149] In some embodiments, a nucleotide sequence is defined as accurately representing the Watson strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, when a threshold percentage (or a percentage above a threshold) of members of the Watson subfamily contain the sequence. In some embodiments, a nucleotide sequence is defined as accurately representing the Crick strand of an analyte DNA fragment, e.g., a double-stranded DNA fragment from a nucleic acid sample, when a threshold percentage (or a percentage above a threshold) of members of the Crick subfamily contain the sequence.

[0150] The threshold value can be determined by a skilled practitioner based on, for example, the number of members of the subfamily, the specific purpose of the sequencing experiment, and the specific parameters of the sequencing experiment. In some embodiments, the threshold value is set at 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In certain embodiments, the threshold value is set at 50%. By way of example only, in an embodiment where the threshold value is set at 50%, a nucleotide sequence is determined to accurately represent an analyte DNA fragment, e.g., a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, when at least 50% of the subfamily members contain the sequence. Simply as another example, in an embodiment where the threshold is set at 50%, a nucleotide sequence is defined as accurately representing an analyte DNA fragment, e.g., a Watson strand or a Crick strand of a double-stranded DNA fragment from a nucleic acid sample, when more than 50% of the subfamily members contain the sequence.

[0151] In some embodiments, a sequence that accurately represents the Watson strand of the analyte DNA fragment is defined as comprising a TCR / BCR receptor sequence.

[0152] In some embodiments, a sequence that accurately represents the Crick strand of the analyte DNA fragment is determined to include a TCR / BCR receptor sequence.

[0153] In some embodiments, the analyte DNA fragment is used to determine the TCR / BCR receptor sequence when the sequence that exactly represents the Watson strand and the sequence that exactly represents the Crick strand contain the same sequence.

[0154] In some cases, the location of the molecular barcode in the paired-end sequencing read of the amplified target region can be used to distinguish which strand of the double-stranded nucleic acid template the amplified target region is from. For example, when the first paired-end sequencing read of the amplified target region indicates that the molecular barcode is read last, the amplified target region can be identified as being from the sense strand of the nucleic acid template, and when the first paired-end sequencing read of the amplified target region indicates that the molecular barcode is read first, the amplified target region can be identified as being from the antisense strand of the nucleic acid template. For example, when the second paired-end sequencing read of the amplified target region indicates that the molecular barcode is read first, the amplified target region can be identified as being from the antisense strand of the nucleic acid template, and when the second paired-end sequencing read of the amplified target region indicates that the molecular barcode is read last, the amplified target region can be identified as being from the sense strand of the nucleic acid template. In some cases, paired-end sequencing can be used to distinguish the amplification products from the Watson strand from the amplification products from the Crick strand.

[0155] After sequencing of the target region (e.g., the target region amplified as described herein), the sequencing reads can be aligned to a reference genome and grouped by the molecular barcode present in each sequencing read.In some cases, the sequencing reads that contain the same molecular barcode and map to both the Watson strand and the Crick strand of the double-stranded nucleic acid template (e.g., both the Watson strand and the Crick strand of the target region) can be identified as having double-stranded support.For example, when the sequencing reads that indicate the presence of one or more mutations in the target region contain the same molecular barcode and map to both the Watson strand and the Crick strand of the target region, the mutation(s) can be identified as having double-stranded support.

[0156] Immune Cell Receptors

[0157] As used herein, "immune cell receptor" refers to a receptor, usually on a cell membrane, that binds to a substance (e.g., a cytokine) and triggers a response in the immune system. For example, immune cell receptors in the immune system include, but are not limited to, pattern recognition receptors (PRRs), toll-like receptors (TLRs), killer activating and inhibitor receptors (KARs and KIRs), complement receptors, Fc receptors, B cell receptors, and T cell receptors. In some embodiments, the immune cell receptor is a T cell receptor. In some embodiments, the immune cell receptor is a B cell receptor.

[0158] T cell receptors (TCRs) are protein complexes found on the surface of T cells, or T lymphocytes, that are responsible for recognizing fragments of antigens as peptides bound to major histocompatibility complex (MHC) molecules. The generation of diversity in TCRs is similar to that for antibodies and B cell antigen receptors. In some embodiments, it results primarily from genetic recombination of DNA-encoding segments in individual somatic T cells by somatic V(D)J recombination using RAG1 and RAG2 recombinases. However, unlike immunoglobulins, TCR genes do not undergo somatic hypermutation, and T cells do not express activation-induced cytidine deaminase (AID). The recombination process that creates diversity in BCRs (e.g., antibodies) and TCRs is unique to lymphocytes (T cells and B cells) at early stages of their development in primary lymphoid organs (thymus for T cells, bone marrow for B cells).

[0159] The B cell receptor (BCR) is a transmembrane protein on the surface of B cells. It is composed of a membrane-bound immunoglobulin molecule and a signaling portion. The former forms a type 1 transmembrane receptor protein and is typically located on the outer surface of these lymphocyte cells. The BCR controls the activation of B cells through biochemical signaling and by physically acquiring antigens from the immune synapse. The portion of the BCR that recognizes antigens is composed of three different genetic regions called V, D, and J. All of these regions are recombined and spliced ​​at the genetic level in a combinatorial process that is exceptional to the immune system. There are many genes that code for each of these regions in the genome, and they can be combined in various ways to generate a wide variety of receptor molecules. This generation of diversity is crucial, since the body may encounter more antigens than there are genes available. Through such a process, the body finds ways to produce receptor molecules that recognize antigens in a wide variety of different combinations. Heavy chain rearrangement of the BCR entails an early stage of B cell development. EXAMPLES

[0160] The present disclosure is further described in the following examples, which do not limit the scope of the disclosure as claimed.

[0161] Example 1 - PCR-based enrichment of BCR and TCR sequences

[0162] The method described here involves three key steps: i) library construction by in situ generation of double-stranded molecular barcodes (Fig. 1a), ii) target enrichment by anchored PCR (Fig. 1b), and iii) in silico reconstruction of the template molecule (Fig. 1c). Bona fide sequences present in the original starting template are identified by requiring the same sequence to be found on both strands of the same initial DNA molecule. This strategy minimizes DNA damage, PCR, and sequencing artifacts, and allows rare sequences to be identified with high confidence.

[0163] To address the inefficiencies and introduced errors typically associated with library construction, we designed a strategy that relies on sequential ligation of adapter sequences to the 3' and 5' DNA fragment ends and the in situ generation of double-stranded molecular barcodes (Fig. 1a). The in situ generation of molecular barcodes is a key innovation in library preparation methods. The enzymes used for the in situ generation of double-stranded molecular barcodes uniquely barcode each DNA fragment and circumvented the need to enzymatically prepare double-stranded adapters, which have been noted to have a detrimental effect on the recovery of input DNA (Fig. 1a, steps 2 and 3). After adapter ligation, the fragments are subjected to a limited number of PCR cycles, which creates redundant copies of the two original DNA strands (Fig. 1a, step 4).

[0164] Another innovation of the protocol disclosed here is the use of a hemi-nested PCR-based approach for enrichment. Because the combinatorial diversity introduced by V(D)J recombination is immense, hemi-nested PCR is theoretically suitable for enriching TCR or BCR sequences. Only a restricted set of primers targeting a limited number of J or V segments needs to be designed (as opposed to traditional PCR-based methods that use primers targeting all paired VJ segment combinations). Hemi-nested PCR has been used for target enrichment before, but significant changes were required to apply it with high efficiency to duplex sequencing. Previous descriptions of hemi-nested PCR have either not retained the strand information necessary to reconstruct the original duplex molecule or have not recovered a high enough fraction of template molecules to detect variants present at frequencies less than 0.1% in a limited amount of DNA. The hemi-nested approach described here employs two separate PCRs - one for the Watson strand and one for the Crick strand (Figure 1b). After sequencing, the reads corresponding to each strand of the original DNA duplex are grouped into Watson and Crick families. Each family member has an identical intrinsic barcode representing the sequence at one end of the initial template fragment and an identical exogenous barcode introduced in situ during library construction. Mutations present in the Watson strand family are called "Watson supermutants." Mutations present in the Crick strand family are called "Crick supermutants." Those present in both the Watson strand family and the Crick strand family with the same molecular barcode ("duplex family") are called "supercalifragilisticexpialidocious mutants," hereafter referred to as "supercalimutants" (Figure 1c).

[0165] The TCR and BCR sequences can be analyzed using custom or publicly available software packages. In the example demonstrated here, the sequences are grouped and aligned for UID error correction using MIGEC, and grouped for clonotypes using MiXCR, and further analyzed using VDJtools.

[0166] Example 2 - Hemi-nested primers targeting TCRs

[0167] The performance of the method described here was evaluated on DNA samples derived from human fibroblasts. For this purpose, hemi-nested primers targeting 13 TCR J segments were designed. Since fibroblasts do not undergo V(D)J recombination, this substrate could be used as a suitable template to evaluate the performance of primers designed to enrich for different TCRs.

[0168] Across the 13 J-segment targets, the median fraction of on-target reads derived from the Watson strand (i.e., reads that constitute the intended amplicon) was 94% (range: 66-96%) (Figure 2). Similarly, the median fraction of on-target reads derived from the Crick strand was 94% (range: 67-85%) (Figure 2). Each target also showed relatively uniform amplification, with coefficients of variation for Watson- and Crick-derived reads of 29% and 24%, respectively (Figure 3). Finally, the number of double-stranded UID families (i.e., each UID family represents an original molecule present in the DNA sample) was exceptionally uniform across each of the 13 targets (median: 5,681, range: 5,317-5,835). The coefficient of variation was 2.8% (Figure 4). Such uniformity is important for accurate quantification of TCR and BCR sequences.

[0169] Example 3 - Synthetic constructs in human fibroblast DNA

[0170] Synthetic constructs were designed to consist of 58bp of the TRBV2 gene, the CDR3 sequence of Jurkat clone E6-1, a barcode specific to each TRBJ, 50bp of one of the TRBJ genes (TRBJ1-1 through TRBJ2-7), and 50bp of CMV promoter sequence, as listed from 5' to 3'. These constructs were spiked into normal human fibroblast DNA. Libraries were then prepared for sequencing using each TRBJ primer set in addition to a primer set specific to the CMV sequence. Sequence reads were grouped by UIDs, sequences were aligned, and the number of barcode molecules identified for each corresponding TRBJ gene synthetic construct was counted.

[0171] Each TRBJ primer set recovered nearly equal numbers of corresponding synthetic construct molecules (median: 833.5, range: 587-1783 for the average of Watson and Crick strands) (Figure 5). Identification of cross-reactivity of non-corresponding synthetic constructs was minimal (Figure 5). The number of synthetic construct molecules identified correlated highly with orthogonal determinations of synthetic control construct concentration as measured by the number of molecules identified using the CMV-specific primer set (Figure 6) and by concentration in a ThermoFisher Qubit dsDNA HS assay (Figure 7). The percentage of correct clonotypes identified by each primer set was high (median: 0.999; range 0.998-1.000) (Figure 8).

[0172] Example 4 - Primers for sequencing TCR receptors

[0173] Primer sets (Table 2) were designed to sequence TCR receptors using any one of the methods described herein. Primers were designed to require only one round of PCR target enrichment using gene-specific primers. Additionally, primers were designed and incorporated into the protocol to include RNase H cleavage by the IDT rhAmpSeq system. The performance of these primers on DNA was evaluated from pooled plasma of normal healthy donors. For all TRBJ segments, the percentage of sequencing reads assignable to TCR clonotypes was low (range 0-0.06%) (Figure 9). The number of clonotypes identified was similarly low for all TRBJ segments (range 0-6) (Figure 10). The performance of the primers and protocol was also evaluated using DNA derived from T cells from normal healthy donors. The percentage of sequencing reads assignable to TCR clonotypes was again low for all TRBJ segments (range 0-3.2%) (Figure 11). The number of clonotypes identified was also low for all TRBJ segments (range 0-82) (Figure 12).

[0174] [Table 2]

[0175] Two different primer sets were designed for sequencing TCR receptors using any one of the methods disclosed herein, designated herein as "Set 1" and "Set 2" (Table 3). The performance of these primer sets on DNA from T cells of normal healthy donors was evaluated. Primers in Set 1 yielded a higher percentage of sequencing reads that were assignable to TCR clonotypes than primers in Set 2 (Figure 13), and discriminated many more clonotypes (Figure 14).

[0176] [Table 3-1] [Table 3-2] [Table 3-3]

[0177] Two different primer sets were also designed for sequencing TCR receptors using any one of the methods disclosed herein, designated herein as "Set 1" and "Set 3" (Table 4). The performance of these primer sets was evaluated on DNA from T cells from normal healthy donors. Set 1 primers had a higher percentage of sequencing reads assignable to TCR clonotypes than Set 3 primers (Figure 15), and discriminated even more clonotypes (Figure 16). The performance of these primer sets was also evaluated using synthetic control constructs that modeled the TCR repertoire. Set 1 primers had a greater percentage of sequencing reads assignable to TCR clonotypes than Set 3 primers (Figure 17).

[0178] [Table 4-1] [Table 4-2] [Table 4-3]

[0179] Multiplex primer sets were also created in which the TRBJ primers were pooled in equimolar ratios or in ratios that yielded balanced reads per TRBJ segment, and were named "multiplex pool 1" and "multiplex pool 2," respectively (Table 5). The performance of these primer sets was evaluated on DNA from fibroblasts of normal healthy donors. The coefficient of variation for the number of on-target reads for each TRBJ segment for multiplex pool 1 was 103.5% (Figure 18). Multiplex pool 2 showed a much more balanced recovery of each TRBJ segment, with a coefficient of variation for the number of on-target reads for the TRBJ segments of 17.5% (Figure 19).

[0180] [Table 5]

[0181] Example 5 - Multiplex primer set using TRBJ primers

[0182] The TRBJ primers were pooled in ratios to obtain balanced reads per TRBJ segment to create a multiplex primer set. The ratio of primers in each mix was adjusted based on the ratio of reads from the above mixes. The ratio of primers was adjusted separately for Watson and Crick GSP reactions. The performance of these primer sets was evaluated with DNA from fibroblasts of normal healthy donors. The penultimate set of primer pools is referred to as "multiplex pool 3" and the last set of primer pools is referred to as "multiplex pool 4" (Table 6). The coefficient of variation for the number of on-target reads for each TRBJ segment for multiplex pool 3 was 19.4% for Watson and 21.4% for Crick (Figure 20). Multiplex pool 4 showed a much more balanced recovery of each TRBJ segment, with the coefficient of variation for the number of on-target reads for each TRBJ segment being 13.2% for Watson and 18.1% for Crick (Figure 20).

[0183] [Table 6]

[0184] Example 6 - Performance evaluation by yield determination

[0185] The performance of the method described here was evaluated by varying the amount of input DNA from healthy donor T cells and determining the yield. The number of recovered TCRs was linear across input amounts from 25 ng to 400 ng, averaged across donors and replicates (Figure 21). Yields were also consistent across input amounts from 25 ng to 400 ng, averaged across donors and replicates (Figure 22).

[0186] Example 7 - Epstein-Barr Virus (EBV)-specific T cells

[0187] Epstein-Barr Virus (EBV)-specific T cells were expanded with EBV peptides. T cells were harvested on days 0, 9, 16, and 27 of expansion. The TCR repertoire of those cells was assessed using the methods described here. The analysis accurately demonstrated a decrease in clonal diversity over the course of expansion (Figure 23). The method also identified the outgrowth of specific clones over the course of expansion (Figure 24, Figure 25, Figure 26).

[0188] Example 8 - Identification of TCR sequences from extracted DNA

[0189] DNA was extracted from T cells from two healthy donors, designated "AB02" and "AB04". TCR repertoires were analyzed using the methods described here from various replicates and DNA input amounts from these samples. The number and diversity of recovered TCRs was high for both donors across all replicates (Figure 27). Pairwise distance correlations between replicates for each donor were consistently high (Figure 28). For a representative donor, AB02, clonotype frequencies of samples from DNA input amounts of 400ng, 100ng, and 25ng correlated well (Figure 29, Figure 30). In addition, TCR V segment gene usage was analyzed in the T cell populations by flow cytometry using the Beckman Coulter IOTest Beta Mark TCR VB Repertoire Kit. The proportion of V gene segment usage correlated well with the proportion of V gene segment usage measured by flow cytometry (Figure 31).

[0190] DNA was isolated from Jurkat clonal T cell lines. This DNA was spiked in various amounts into DNA from healthy donor T cells. TCR repertoires were analyzed using methods described here. The proportion of TCR reads corresponding to Jurkat clonal TCRs correlated well with the proportion of input DNA (Figure 32).

[0191] The methods described here were used to analyze DNA from Jurkat clonal T cell lines, and the method correctly identified exactly one TCR clone in four replicates in the Jurkat sample, and two clones in two replicates were called (Figure 33).

[0192] DNA was isolated from plasma, leukocyte, and tumor samples of patients with colorectal cancer. The methods described here were used to analyze the TCR repertoire in each compartment. The results show the diversity of TCRs in plasma (Figure 34), leukocyte (Figure 35), and tumor (Figure 36) samples.

Claims

1. 1. A method for determining the sequence of an immune cell receptor double-stranded DNA molecule, comprising: (a) attaching a 3' adapter fragment to each 3' end of the double-stranded DNA molecule and a 5' adapter fragment to each 5' end of the double-stranded DNA molecule to generate adapted double-stranded DNA molecules; The adapted double-stranded DNA molecule comprises an adapted Watson strand and an adapted Crick strand; the 3' adapter fragment comprises a molecular barcode, a primer sequence, and an adapter sequence; and The adapted Watson strand molecular barcode is the reverse complement of the adapted Crick strand molecular barcode, To attach and (b) copying both strands of the adapted double-stranded DNA molecule, where the copying comprises performing a round of linear extension of the adapted double-stranded DNA molecule to generate an adapted double-stranded Watson template and an adapted double-stranded Crick template; (c) generating a first population of analyte DNA fragments from the adapted double-stranded Watson template and generating a first sequencing read for at least one member of the first population of analyte DNA fragments; (d) generating a second population of analyte DNA fragments from the adapted double-stranded Click template and generating a second sequencing read for at least one member of the second population of analyte DNA fragments; (e) grouping the first sequencing reads according to molecular barcodes present on at least one member of the first population of analyte DNA fragments to generate a first analyte DNA family; (f) grouping the second sequencing reads according to the molecular barcodes present on at least one member of the second population of analyte DNA fragments to generate a second analyte DNA family; (g) analyzing the first sequencing reads of the first analyte DNA family; (h) analyzing the second sequencing reads of the second analyte DNA family, thereby determining the sequences of the double-stranded DNA molecules; A method comprising:

2. 2. The method of claim 1, wherein the 3' adapter fragment comprises a partial double-stranded molecular barcode.

3. 3. The method of claim 2, wherein the partially double-stranded molecular barcode comprises an endogenous barcode, an exogenous barcode, or both.

4. 2. The method of claim 1, wherein the copying step (b) further comprises performing a round of linear extension of the adapted double-stranded DNA molecule using (i) a first primer complementary to the 3' adapter sequence and (ii) a second primer complementary to the complementary strand of the 5' adapter sequence.

5. 2. The method of claim 1, wherein generating steps (c) and (d) are performed under PCR conditions.

6. the generating step (c) further comprises amplifying the adapted double-stranded Watson template with a first set of Watson target-selective primer pairs; 6. The method of claim 5, wherein the first set of Watson target selective primer pairs includes (i) a first Watson target selective primer comprising a sequence complementary to the 3' adapter sequence, and (ii) a second Watson target selective primer comprising a target selective sequence.

7. The second Watson target selective primers are SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:6 4, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, or SEQ ID NO:

65.

8. 7. The method of claim 6, wherein the second Watson target-selective primer comprises a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:

26.

9. the generating step (d) further comprises amplifying the adapted double-stranded Click template with a first set of Click target-selective primer pairs; 6. The method of claim 5, wherein the first set of click target selective primer pairs includes (i) a first click target selective primer comprising a sequence complementary to the 3' adapter sequence, and (ii) a second click target selective primer comprising a target selective sequence.

10. The second click target selective primers are SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:

9. The method of claim 8, wherein the sequence comprises a sequence selected from the group consisting of SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, or SEQ ID NO:

65.

11. 9. The method of claim 8, wherein the second click target selective primer comprises a sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, or SEQ ID NO:

26.

12. The method of claim 1, wherein the double-stranded DNA molecule comprises a V(D)J sequence of an immune cell receptor.

13. 13. The method of claim 12, wherein the target-selective sequence comprises a sequence complementary to a V(D)J sequence of an immune cell receptor.

14. The method of claim 1 , wherein the immune cell receptor comprises a B cell receptor.

15. The method of claim 1 , wherein the immune cell receptor comprises a T cell receptor.

16. 10. The method of claim 1, further comprising identifying (i) mutations in the adapted double-stranded Watson template of the first analyte DNA family, (ii) mutations in the adapted double-stranded Crick template of the second analyte DNA family, or (iii) mutations in both the adapted double-stranded Watson template and the adapted double-stranded Crick template.

17. 17. The method of claim 16, wherein the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof.

18. 18. The method of any one of claims 1-17, wherein the method determines the sequence of a double-stranded DNA molecule in a population of double-stranded DNA molecules by assaying both strands of the double-stranded DNA molecule.

19. 20. The method of claim 18, wherein mutations are identified in both the adapted double-stranded Watson template and the adapted double-stranded Crick template.