Methods and systems for detecting genetic variants
By differentially tagging both halves of double-stranded DNA molecules and analyzing paired and unpaired sequences, the method addresses the challenge of estimating unseen molecules, enhancing the accuracy and sensitivity of genetic variant detection in cell-free DNA.
Patent Information
- Application Number
- JP2025172644
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2014-03-05
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-21
AI Technical Summary
Existing methods for detecting and quantifying genetic variants in cell-free DNA are limited by the inability to accurately estimate the count of molecules that have been converted but not sequenced, leading to variable sensitivity and inaccuracies in disease detection.
A method involving differential tagging of both halves of double-stranded DNA molecules using hairpin, bubble, or fork-shaped adapters, followed by sequencing and bioinformatics to record and estimate the number of paired and unpaired molecules, allowing for accurate quantification of rare DNA in a heterogeneous population.
Enables the detection and quantification of rare DNA with high specificity and sensitivity, reducing sequencing noise and improving the accuracy of genetic disease monitoring and characterization.
Smart Images

Figure 2026010082000007 
Figure 2026010082000008 
Figure 2026010082000009
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit under U.S.C. §119(e) of U.S. Provisional Application No. 61 / 921,456, filed December 28, 2013, and U.S. Provisional Application No. 61 / 948,509, filed March 5, 2014, each of which is incorporated herein by reference in its entirety. [Background technology]
[0002] The detection and quantification of polynucleotides is important for molecular biology and medical applications, such as diagnostics. Genetic testing is particularly useful for many diagnostic methods. For example, disorders caused by rare genetic alterations (e.g., sequence variants) or alterations in epigenetic markers, such as cancer and partial or complete aneuploidy, can be detected or more accurately characterized with DNA sequence information.
[0003] Early detection and monitoring of genetic diseases, such as cancer, is often useful and necessary for successful treatment or management of the disease. One approach is to use cell-free nucleic acid sequencing, a collection of polynucleotides that can be found in different types of body fluids. Cell-free DNA (cfDNA) can contain genetic abnormalities associated with specific diseases.With the improvement in sequencing and nucleic acid manipulation techniques, there is a need in the art for improved methods and systems for detecting and monitoring diseases using cell-free DNA.In some cases, disease can be characterized or detected based on the detection of genetic abnormalities, such as copy number variation and / or sequence variation of one or more nucleic acid sequences, or the occurrence of other specific rare genetic alterations.Cell-free DNA (cfDNA) can contain genetic abnormalities associated with specific diseases.With the improvement in sequencing and nucleic acid manipulation techniques, there is a need in the art for improved methods and systems for detecting and monitoring diseases using cell-free DNA.
[0004] Specifically, many methods have been developed for accurate copy number variation estimation, especially for heterogeneous genomic samples such as tumor-derived gDNA or cfDNA for many applications (e.g., prenatal, transplantation, immunology, metagenomics, or cancer diagnostics). Most of these methods involve sample preparation to convert the original nucleic acid into a sequenceable library, followed by massively parallel sequencing and finally bioinformatics to estimate copy number variation at one or more loci. Summary of the Invention [Means for solving the problem]
[0005] While many of these methods can reduce or combat errors introduced by the sample preparation and sequencing process for every molecule converted and sequenced, they cannot estimate the count of molecules that have been converted but not sequenced, which can dramatically and detrimentally affect the sensitivity that can be achieved because such counts of molecules that have been converted but not sequenced can be highly variable between genomic regions.
[0006] To address this problem, input double-stranded deoxyribonucleic acid (DNA) can be converted by a process that differentially tags both halves of each double-stranded molecule, in some cases. This can be done by using hairpin, bubble, or fork-shaped adapters or by combining double-stranded and single-stranded segments (bubble, fork, or hairpin adapters). The unhybridized portion of the adapter can be ligated using a variety of techniques, including ligation of another adapter with a single strand (considered single-stranded herein). Once correctly tagged, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule is differentially tagged and can be identified by sequencer and subsequent bioinformatics. For every molecule in a particular region, the count of molecules in which both Watson and Crick sides are recovered ("Pair") versus molecules in which only one half is recovered ("Singlet") is recorded. The number of unseen molecules can be estimated based on the number of pairs and singlets detected.
[0007] An embodiment of the present disclosure provides a method for detecting and / or quantifying rare deoxyribonucleic acid (DNA) in a heterogeneous population of original DNA fragments, comprising tagging original DNA fragments in a single reaction using a library of multiple different tags such that more than 30% of the fragments are tagged at both ends, each tag comprising a molecular barcode. The single reaction can be carried out in a single reaction vessel. More than 50% of the fragments can be tagged at both ends. The multiple different tags can be any of 100, 500, 1000, 10,000, or 100,000 different tags or less.
[0008] Another aspect provides a set of library adaptors that can be used to tag molecules of interest (e.g., by ligation, hybridization, etc.). The set of library adaptors can include a plurality of polynucleotide molecules having molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, the molecular barcodes are at least 4 nucleotide bases in length, and (a) the molecular barcodes are different from one another and have an edit distance of at least 1 between one another, (b) the molecular barcodes are located at least one nucleotide base away from a terminal end of their respective polynucleotide molecules, (c) optionally, at least one terminal base is identical in all of the polynucleotide molecules, and (d) none of the polynucleotide molecules contains a complete sequencer motif.
[0009] In some embodiments, the library adaptors (or adaptors) are identical to one another except for the molecular barcodes. In some embodiments, each of the plurality of library adaptors comprises at least one double-stranded portion and at least one single-stranded portion (e.g., a non-complementary portion or an overhang). In some embodiments, the double-stranded portion has a molecular barcode selected from a collection of different molecular barcodes. In some embodiments, the given molecular barcode is a randomer. In some embodiments, each of the library adaptors further comprises a strand-identification barcode in at least one single-stranded portion. In some embodiments, the strand-identification barcode comprises at least 4 nucleotide bases. In some embodiments, the single-stranded portion has a partial sequencer motif. In some embodiments, the library adaptor does not comprise a complete sequencer motif.
[0010] In some embodiments, none of the library adaptors contain sequences for hybridizing to a flow cell or for forming a hairpin for sequencing.
[0011] In some embodiments, the library adaptors all have ends with the same nucleotide(s). In some embodiments, the identical terminal nucleotide(s) spans a length of 2 nucleotide bases or more.
[0012] In some embodiments, each of the library adapters is Y-shaped, bubble-shaped, or hairpin-shaped. In some embodiments, none of the library adapters contain a sample identification motif. In some embodiments, each of the library adapters comprises a sequence that can selectively hybridize to a universal primer. In some embodiments, each of the library adapters comprises a molecular barcode that is at least 5, 6, 7, 8, 9, and 10 nucleotide bases in length. In some embodiments, each of the library adapters is 10 to 80 nucleotide bases in length, or 30 to 70 nucleotide bases in length, or 40 to 60 nucleotide bases in length. In some embodiments, at least 1, 2, 3, or 4 terminal bases are identical in all of the library adapters. In some embodiments, at least 4 terminal bases are identical in all of the library adapters.
[0013] In some embodiments, the edit distance of the molecular barcodes of the library adapters is the Hamming distance. In some embodiments, the edit distance is at least 1, 2, 3, 4, or 5. In some embodiments, the edit distance relates to individual bases of the plurality of polynucleotide molecules. In some embodiments, the molecular barcodes are located at least 10 nucleotide bases away from the end of the adapter. In some embodiments, the plurality of library adapters comprises at least 2, 4, 6, 8, 10, 20, 30, 40, or 50 different molecular barcodes, or 2 to 100, 4 to 80, 6 to 60, or 8 to 40 different molecular barcodes. In any of the embodiments herein, there are more polynucleotides (e.g., cfDNA fragments) to be tagged than there are different molecular barcodes, such that tagging is not unique.
[0014] In some embodiments, the ends of the adaptors are configured for ligation (e.g., to a target nucleic acid molecule). In some embodiments, the ends of the adaptors are blunt ends.
[0015] In some embodiments, the adaptors are purified and isolated. In some embodiments, the library comprises one or more non-naturally occurring bases.
[0016] In some embodiments, the polynucleotide molecule comprises a primer sequence positioned 5' with respect to the molecular barcode.
[0017] In some embodiments, the set of library adaptors consists essentially of a plurality of polynucleotide molecules.
[0018] In another aspect, the method includes (a) tagging a collection of polynucleotides with a plurality of polynucleotide molecules from a library of adaptors to create a collection of tagged polynucleotides; and (b) amplifying the collection of tagged polynucleotides in the presence of sequencing adaptors, wherein the sequencing adaptors have primers having nucleotide sequences that are selectively hybridizable to complementary sequences in the plurality of polynucleotide molecules. The library of adaptors can be as described above or elsewhere herein. In some embodiments, each of the sequencer adaptors further comprises an index tag, which can be a sample identification motif.
[0019] Another aspect provides a method for detecting and / or quantifying rare DNA in a heterogeneous population of original DNA fragments, wherein the rare DNA has a concentration that is less than 1%, the method comprising: (a) tagging the original DNA fragments in a single reaction such that greater than 30% of the original DNA fragments are tagged at both ends with library adaptors that comprise molecular barcodes, thereby providing tagged DNA fragments; (b) performing high-fidelity amplification on the tagged DNA fragments; (c) optionally, selectively enriching a subset of the tagged DNA fragments; (d) sequencing one or both strands of the tagged, amplified, and optionally selectively enriched DNA fragments to obtain sequence reads that comprise the nucleotide sequences of the molecular barcodes and at least a portion of the original DNA fragments; (e) determining consensus reads from the sequence reads that are representative of a single strand of the original DNA fragments; and (f) quantifying the consensus reads to detect and / or quantify the rare DNA with greater than 99.9% specificity.
[0020] In some embodiments, (e) comprises comparing sequence reads having the same or similar molecular barcodes and the same or similar end of fragment sequences. In some embodiments, the comparing step further comprises performing a phylogentic analysis on sequence reads having the same or similar molecular barcodes. In some embodiments, the molecular barcodes comprise barcodes with an edit distance of up to 3. In some embodiments, the ends of the fragment sequences comprise fragment sequences with an edit distance of up to 3.
[0021] In some embodiments, the method further comprises sorting the sequence reads into paired reads and unpaired reads, and quantifying the number of paired reads and unpaired reads that map to each of the one or more genetic loci.
[0022] In some embodiments, tagging occurs by having an excess of library adaptors compared to the original DNA fragments. In some embodiments, the excess is at least a 5-fold excess. In some embodiments, tagging comprises the use of a ligase. In some embodiments, tagging comprises attachment to blunt ends.
[0023] In some embodiments, the method further comprises binning the sequence reads according to the molecular barcodes and sequence information from at least one end of each of the original DNA fragments to create bins of single-stranded reads. In some embodiments, the method further comprises determining the sequence of a given original DNA fragment among the original DNA fragments by analyzing the sequence reads in each bin. In some embodiments, the method further comprises detecting and / or quantifying rare DNA by comparing the number of times each base occurs at each position of the genome represented by the tagged, amplified, and optionally enriched DNA fragments.
[0024] In some embodiments, the library adaptors do not contain a complete sequencer motif. In some embodiments, the method further comprises selectively enriching a subset of the tagged DNA fragments. In some embodiments, the method further comprises, after enrichment, amplifying the enriched tagged DNA fragments in the presence of sequencing adaptors comprising primers. In some embodiments, (a) results in tagged DNA fragments having 2 to 1000 different combinations of molecular barcodes.
[0025] In some embodiments, the DNA fragments are tagged with polynucleotide molecules derived from a library of adaptors described above or elsewhere herein.
[0026] In another embodiment, a method for processing and / or analyzing a nucleic acid sample of a subject includes (a) exposing polynucleotide fragments from the nucleic acid sample to a set of library adaptors to generate tagged polynucleotide fragments, and (b) subjecting the tagged polynucleotide fragments to a nucleic acid amplification reaction under conditions that yield amplified polynucleotide fragments as amplification products of the tagged polynucleotide fragments. The set of library adaptors includes a plurality of polynucleotide molecules having molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, the molecular barcodes are at least 4 nucleotide bases in length, (1) the molecular barcodes are different from one another and have an edit distance of at least 1 between one another, (2) the molecular barcodes are located at least one nucleotide base away from a terminal end of each polynucleotide molecule, (3) optionally, at least one terminal base is identical in all of the polynucleotide molecules, and (4) none of the polynucleotide molecules contains a complete sequencer motif.
[0027] In some embodiments, the method further comprises determining the nucleotide sequence of the amplified tagged polynucleotide fragments. In some embodiments, the nucleotide sequence of the amplified tagged polynucleotide fragments is determined without polymerase chain reaction (PCR). In some embodiments, the method further comprises analyzing the nucleotide sequence with a programmed computer processor to identify one or more genetic variants in the subject's nucleotide sample. In some embodiments, the one or more genetic variants are selected from the group consisting of base change(s), insertion(s), repeat(s), deletion(s), copy number variation(s), and transversion(s). In some embodiments, the one or more genetic variants comprise one or more tumor-associated genetic alterations.
[0028] In some embodiments, the subject is suffering from or suspected of suffering from a disease. In some embodiments, the disease is cancer. In some embodiments, the method further comprises collecting a nucleic acid sample from the subject. In some embodiments, the nucleic acid sample is collected from a location selected from the group consisting of blood, plasma, serum, urine, saliva, mucosal excretion, sputum, feces, cerebrospinal fluid, and tears of the subject. In some embodiments, the nucleic acid sample is a cell-free nucleic acid sample. In some embodiments, the nucleic acid sample is collected from 100 nanograms (ng) or less of double-stranded polynucleotide molecules of the subject.
[0029] In some embodiments, the polynucleotide fragments comprise double-stranded polynucleotide molecules. In some embodiments, in (a), the plurality of polynucleotide molecules are coupled to the polynucleotide fragments by blunt-end ligation, sticky-end ligation, molecular inversion probes, PCR, ligation-based PCR, multiplex PCR, single-stranded ligation, and single-stranded circularization. In some embodiments, exposing the polynucleotide fragments of the nucleic acid sample to the plurality of polynucleotide molecules produces tagged polynucleotide fragments with a conversion efficiency of at least 10%. In some embodiments, at least 5%, 6%, 7%, 8%, 9%, 10%, 20%, or 25% of the tagged polynucleotide fragments share a common polynucleotide molecule or sequence. In some embodiments, the method further comprises generating polynucleotide fragments from the nucleic acid sample.
[0030] In some embodiments, the subjecting step comprises the step of subjecting to one or more of the following: ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PD and amplifying tagged polynucleotide fragments from sequences corresponding to genes selected from the group consisting of GFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
[0031] In another aspect, a method includes (a) generating a plurality of sequence reads from a plurality of polynucleotide molecules, wherein the plurality of polynucleotide molecules covers genomic loci of a target genome, wherein the genomic loci are selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF 1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH 1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CC (b) grouping, by a computer processor, the plurality of sequence reads into families, each family including sequence reads derived from one of the template polynucleotides; (c) for each of the families, merging the sequence reads to generate a consensus sequence; (d) calling the consensus sequence at a given genomic locus among the genomic loci; and (e) detecting, at the given genomic locus, any of genetic variants among the calls, frequency of genetic alterations among the calls, total number of calls, and total number of alterations among the calls.
[0032] In some embodiments, each family comprises sequence reads derived from only one of the template polynucleotides. In some embodiments, a given genomic locus comprises at least one nucleic acid base. In some embodiments, a given genomic locus comprises multiple nucleic acid bases. In some embodiments, the calling step comprises calling at least one nucleic acid base at a given genomic locus. In some embodiments, the calling step comprises calling multiple nucleic acid bases at a given genomic locus. In some embodiments, the calling step comprises any one of phylogenetic analysis, voting, weighing, assigning a probability to each read at a locus in a family, and calling the base with the highest probability.
[0033] In some embodiments, the method further comprises performing (d)-(e) at an additional genomic locus among the genomic loci. In some embodiments, the method further comprises determining copy number variation at one of the given genomic locus and the additional genomic locus based on the counts at the given genomic locus and the additional genomic locus.
[0034] In some embodiments, the grouping step comprises classifying the plurality of sequence reads into families by identifying similarities between (i) different molecular barcodes coupled to the plurality of polynucleotide molecules and (ii) the plurality of sequence reads, wherein each family comprises a plurality of nucleic acid sequences associated with different combinations of molecular barcodes and similar or identical sequence reads. Different molecular barcodes have different sequences.
[0035] In some embodiments, the consensus sequence is generated by evaluating the quantitative measure or statistical significance level of each sequence read. In some embodiments, the quantitative measure comprises using a binomial distribution, an exponential distribution, a beta distribution, or an empirical distribution. In some embodiments, the method further comprises mapping the consensus sequence to a target genome. In some embodiments, the plurality of genes comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or all of the plurality of genes selected from the group.
[0036] Another aspect of the present disclosure provides a method comprising: (a) providing in a single reaction vessel template polynucleotide molecules and a set of library adaptors, wherein the library adaptors are polynucleotide molecules having different molecular barcodes (e.g., between 2 and 1,000 different molecular barcodes), and none of the library adaptors contain a complete sequencer motif; (b) coupling the library adaptors to the template polynucleotide molecules in the single reaction vessel at an efficiency of at least 10%, thereby tagging each template polynucleotide with a tagging combination among a plurality of different tagging combinations (e.g., between 4 and 1,000,000 different tagging combinations) to produce tagged polynucleotide molecules; (c) subjecting the tagged polynucleotide molecules to an amplification reaction under conditions that yield amplified polynucleotide molecules as amplification products of the tagged polynucleotide molecules; and (d) sequencing the amplified polynucleotide molecules.
[0037] In some embodiments, the template polynucleotide molecules are blunt-ended or sticky-ended. In some embodiments, the library adaptors are identical except for the molecular barcodes. In some embodiments, each of the library adaptors has a double-stranded portion and at least one single-stranded portion. In some embodiments, the double-stranded portion has a molecular barcode from among the plurality of molecular barcodes. In some embodiments, each of the library adaptors further comprises a strand-identification barcode on at least one single-stranded portion. In some embodiments, the single-stranded portion has a partial sequencer motif. In some embodiments, the library adaptors have the same sequence of terminal nucleotides. In some embodiments, the template polynucleotide molecules are double-stranded. In some embodiments, the library adaptors couple to both ends of the template polynucleotide molecule.
[0038] In some embodiments, subjecting the tagged polynucleotide molecules to an amplification reaction comprises non-specifically amplifying the tagged polynucleotide molecules.
[0039] In some embodiments, the amplification reaction includes using a priming site to amplify each of the tagged polynucleotide molecules. In some embodiments, the priming site is a primer. In some embodiments, the primer is a universal primer. In some embodiments, the priming site is a nick.
[0040] In some embodiments, prior to (e), the method further comprises the steps of: (i) separating polynucleotide molecules comprising one or more given sequences from the amplified polynucleotide molecules to produce enriched polynucleotide molecules; and (ii) amplifying the enriched polynucleotide molecules with sequencing adaptors.
[0041] In some embodiments, the efficiency is at least 30%, 40%, or 50%. In some embodiments, the method further comprises identifying genetic variants upon sequencing the amplified polynucleotide molecules. In some embodiments, the sequencing step comprises (i) subjecting the amplified polynucleotide molecules to an additional amplification reaction under conditions that generate additional amplified polynucleotide molecules as amplification products of the amplified polynucleotide molecules, and (ii) sequencing the additional amplified polynucleotide molecules. In some embodiments, the additional amplification is performed in the presence of sequencing adaptors.
[0042] In some embodiments, (b) and (c) are performed without aliquoting the tagged polynucleotide molecules. In some embodiments, the tagging is non-unique tagging.
[0043] Another embodiment is a system for analyzing target nucleic acid molecules of a subject, the system comprising: a communication interface that receives nucleic acid sequence reads for a plurality of polynucleotide molecules covering genomic loci of a target genome; a computer memory that stores the nucleic acid sequence reads for the plurality of polynucleotide molecules received by the communication interface; and a computer processor operably coupled to the communication interface and the memory and programmed to: (i) group the plurality of sequence reads into families, each family including sequence reads from one of the template polynucleotides; (ii) for each of the families, merge the sequence reads to generate a consensus sequence; (iii) call the consensus sequence at a given genomic locus among the genomic loci; and (iv) detect any of genetic variants among the calls, frequency of genetic alterations among the calls, total number of calls, and total number of alterations among the calls at the given genomic locus, wherein the genomic loci are selected from the group consisting of ALK, A, and B. PC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNN B1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFR A system is provided that corresponds to a plurality of genes selected from the group consisting of A, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
[0044] In another embodiment, ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC , PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1.
[0045] In some embodiments, the oligonucleotide molecule is 10 to 200 bases in length. In some embodiments, the oligonucleotide molecule selectively hybridizes to exon regions of at least five genes. In some embodiments, the oligonucleotide molecule selectively hybridizes to at least 30 exons in at least five genes. In some embodiments, multiple oligonucleotide molecules selectively hybridize to each of the at least 30 exons. In some embodiments, the oligonucleotide molecule hybridizing to each exon has a sequence that overlaps with at least one other oligonucleotide molecule.
[0046] In another embodiment, the kit includes a first container containing a plurality of library adaptors, each having a different molecular barcode, and a second container containing a plurality of sequencing adaptors, each sequencing adaptor comprising at least a portion of a sequencer motif and optionally a sample barcode. The library adaptors can be those described above or elsewhere herein.
[0047] In some embodiments, the sequencing adaptor comprises a sample barcode. In some embodiments, the library adaptor is blunt-ended and Y-shaped, and is less than or equal to 80 nucleic acid bases in length. In some embodiments, the sequencing adaptor is up to 70 bases from end to end.
[0048] In another aspect, a method is provided for detecting sequence variants in a cell-free DNA sample, the method comprising detecting rare DNA at a concentration of less than 1% with greater than 99.9% specificity.
[0049] In another aspect, the method comprises detecting genetic variants in a sample comprising DNA with a detection limit of at least 1% and a specificity of greater than 99.9%. In some embodiments, the method further comprises converting cDNA (e.g., cfDNA) into adaptor-tagged DNA with a conversion efficiency of at least 30%, 40%, or 50%, thereby reducing sequencing noise (or distortion) by eliminating false-positive sequence reads.
[0050] Another aspect includes a method for generating a sequence read from a sample, the method comprising: (a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule comprising a first and a second complementary strand; (b) tagging the double-stranded polynucleotide molecules with a set of duplex tags, each duplex tag differently tags a first and a second complementary strand of a double-stranded polynucleotide molecule in the set; (c) sequencing at least a portion of the tagged strands to produce a set of sequence reads; (d) reducing and / or tracking redundancy in the set of sequence reads; and (e) sorting the sequence reads into paired reads and unpaired reads, wherein (i) each paired read is derived from a first tagged tag derived from a double-stranded polynucleotide molecule in the set. (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads; (f) determining a quantitative measure of (i) paired reads and (ii) unpaired reads that map to each of the one or more genetic loci; and (g) estimating with a programmed computer processor a quantitative measure of total double-stranded polynucleotide molecules in the set that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each locus.
[0051] In some embodiments, the method further comprises (h) detecting copy number variation in the sample by determining the normalized total quantitative measure determined in step (g) at each of one or more loci and determining copy number variation based on the normalized measure. In some embodiments, the sample comprises double-stranded polynucleotide molecules substantially sourced from cell-free nucleic acids. In some embodiments, the duplex tags are not sequencing adaptors.
[0052] In some embodiments, reducing redundancy in the set of sequence reads comprises collapsing sequence reads generated from amplified products of original polynucleotide molecules in the sample back to the original polynucleotide molecules. In some embodiments, the method further comprises determining a consensus sequence of the original polynucleotide molecules. In some embodiments, the method further comprises identifying polynucleotide molecules at one or more loci that contain sequence variants. In some embodiments, the method further comprises determining a quantitative measure of paired reads mapping to the loci, wherein both strands of the pair contain the sequence variant. In some embodiments, the method further comprises determining a quantitative measure of paired molecules, wherein only one member of the pair has the sequence variant, and / or determining a quantitative measure of unpaired molecules that have the sequence variant. In some embodiments, the sequence variant is selected from the group consisting of single nucleotide variants, indels, transversions, translocations, inversions, deletions, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, and chromosomal lesions.
[0053] Another aspect includes, after execution by a computer processor, steps of: (a) receiving into memory a set of sequence reads of polynucleotides tagged with duplex tags; (b) reducing and / or tracking redundancy in the set of sequence reads; and (c) sorting the sequence reads into paired reads and unpaired reads, wherein: (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second, differently tagged, complementary strand derived from a double-stranded polynucleotide molecule in the set; and (ii) each unpaired read is represented among the sequence reads in the set of sequence reads. A system is provided that includes a computer-readable medium that includes machine-executable code for implementing a method that includes representing first tagged strands that have no second, differently tagged, complementary strands derived from double-stranded polynucleotide molecules; (d) determining a quantitative measure of (i) paired reads and (ii) unpaired reads that map to each of one or more genetic loci; and (e) estimating a quantitative measure of total double-stranded polynucleotide molecules in the set that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each locus.
[0054] Another aspect includes the steps of: (a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule comprising first and second complementary strands; (b) tagging the double-stranded polynucleotide molecules with a set of duplex tags, each duplex tag differently tags a first and second complementary strand of a double-stranded polynucleotide molecule in the set; (c) sequencing at least a portion of the tagged strands to produce a set of sequence reads; (d) reducing and / or tracking redundancy in the set of sequence reads; and (e) sorting the sequence reads into paired and unpaired reads, wherein: (i) each paired (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand that is derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads; and (f) determining quantitative measures of at least two of (i) the paired reads, (ii) the unpaired reads that map to each of one or more genetic loci, (iii) the read depth of the paired reads, and (iv) the read depth of the unpaired reads.
[0055] In some embodiments, (f) comprises determining at least three of the quantitative measures (i)-(iv). In some embodiments, (f) comprises determining all of the quantitative measures (i)-(iv). In some embodiments, the method further comprises (g) estimating, with the programmed computer processor, a quantitative measure of the total double-stranded polynucleotide molecules in the set mapping to each of the one or more genetic loci based on the paired and unpaired reads that map to each locus and the quantitative measure of their read depth.
[0056] In another aspect, the method includes the steps of: (a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides, wherein the first tag set comprises a plurality of tags, each tag in the first tag set comprises the same control tag and an identifying tag, and wherein the tag set comprises a plurality of different identifying tags; (b) tagging test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides, wherein the second tag set comprises a plurality of tags, each tag in the second tag set comprises the same test tag that is distinguishable from the control tag and the identifying tag, and wherein the second tag set comprises a plurality of different identifying tags; (c) mixing the tagged control parent polynucleotides with the tagged test parent polynucleotides to form pools; (d) amplifying the tagged parent polynucleotides in the pools to form pools of amplified, tagged polynucleotides; and (e) amplifying the amplified, tagged polynucleotides in the amplified pools. (f) grouping the sequence reads into families, each family containing sequence reads generated from the same parent polynucleotide, the grouping optionally being based on information from an identifying tag and the start / end sequences of the parent polynucleotides; and, optionally, determining a consensus sequence for each of the parent polynucleotides from the plurality of sequence reads in the group; (g) classifying each family or consensus sequence as a control parent polynucleotide or a test parent polynucleotide based on having a test tag or a control tag; (h) determining quantitative measures of the control parent polynucleotides and the control test polynucleotides that map to each of at least two genetic loci; and (i) determining copy number variation in the test parent polynucleotide at at least one genetic locus based on the relative amounts of the test parent polynucleotides and the control parent polynucleotides that map to the at least one genetic locus.
[0057] In another aspect, a method includes: (a) generating a plurality of sequence reads from a plurality of template polynucleotides, wherein each polynucleotide is mapped to a genomic locus; (b) grouping the sequence reads into families, wherein each family comprises sequence reads generated from one of the template polynucleotides; (c) calling bases (or sequences) at the genomic locus for each of the families; and (d) detecting, at the genomic locus, any of: genomic alterations among the calls; frequency of genetic alterations among the calls; total number of calls; and total number of alterations among the calls.
[0058] In some embodiments, the calling comprises any of phylogenetic analysis, voting, weighing, assigning a probability to each read at a locus in a family, and calling the base with the highest probability. In some embodiments, the method is performed at two loci and comprises determining a CNV at one of the loci based on counts at each of the loci.
[0059] Another aspect provides a method for determining a quantitative measure indicative of the number of double-stranded DNA fragments in a sample, the method comprising: (a) determining a quantitative measure of individual DNA molecules for which both strands are detected; (b) determining a quantitative measure of individual DNA molecules for which only one of the DNA strands is detected; (c) inferring from (a) and (b) above a quantitative measure of individual DNA molecules for which neither strand is detected; and (d) using (a)-(c) to determine a quantitative measure indicative of the number of individual double-stranded DNA fragments in the sample.
[0060] In some embodiments, the method further comprises detecting copy number variation in the sample by determining a normalized quantitative measure determined in step (d) at each of the one or more loci and determining copy number variation based on the normalized measure. In some embodiments, the sample comprises double-stranded polynucleotide molecules sourced substantially from cell-free nucleic acid.
[0061] In some embodiments, determining quantitative measures of individual DNA molecules comprises tagging the DNA molecules with a set of duplex tags, each duplex tag differently tagging a complementary strand of a double-stranded DNA molecule in the sample to provide tagged strands. In some embodiments, the method further comprises sequencing at least some of the tagged strands to generate a set of sequence reads. In some embodiments, the method comprises sorting the sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads. In some embodiments, the method further comprises determining a quantitative measure of (i) paired reads and (ii) unpaired reads that map to each of the one or more genetic loci, and determining a quantitative measure of total double-stranded DNA molecules in the sample that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each locus.
[0062] In another aspect, a method for reducing distortion in a sequencing assay comprises: (a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides; (b) tagging test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides; (c) mixing the tagged control parent polynucleotides with the tagged test parent polynucleotides to form a pool; (d) determining the quantities of the tagged control parent polynucleotides and the tagged test parent polynucleotides; and (e) using the quantities of the tagged control parent polynucleotides to reduce distortion in the quantities of the tagged test parent polynucleotides.
[0063] In some embodiments, the first tag set comprises a plurality of tags, each tag in the first tag set comprises the same control tag and identifying tag, and the first tag set comprises a plurality of different identifying tags. In some embodiments, the second tag set comprises a plurality of tags, each tag in the second tag set comprises the same test tag and identifying tag, the test tag is distinguishable from the control tag, and the second tag set comprises a plurality of different identifying tags. In some embodiments, (d) comprises amplifying the tagged parent polynucleotides in the pool to form a pool of amplified tagged polynucleotides, and sequencing the amplified tagged polynucleotides in the amplified pool to produce a plurality of sequence reads. In some embodiments, the method further comprises grouping the sequence reads into families, each family comprising sequence reads generated from the same parent polynucleotide, and this grouping optionally comprises a step based on information derived from the identifying tag and the start / end sequence of the parent polynucleotide, and optionally a step of determining a consensus sequence for each of the plurality of parent polynucleotides from the plurality of sequence reads in the group.
[0064] In some embodiments, (d) comprises determining copy number variation in the test parent polynucleotides at more than or equal to one locus based on the relative amounts of the test parent polynucleotides and control parent polynucleotides mapping to the locus.
[0065] Another aspect provides a method comprising: (a) ligating adaptors to double-stranded DNA polynucleotides to produce a tagged library containing inserts from the double-stranded DNA polynucleotides and having between 4 and 1 million different tags, wherein the ligation is performed in a single reaction vessel and the adaptors contain molecular barcodes; (b) generating multiple sequence reads for each of the double-stranded DNA polynucleotides in the tagged library; (c) grouping the sequence reads into families based on information in the tags and information at the ends of the inserts, each family containing sequence reads generated from a single DNA polynucleotide among the double-stranded DNA polynucleotides; and (d) calling a base at each position in the double-stranded DNA molecules based on the base at that position in members of the family. In some embodiments, (b) comprises amplifying each of the double-stranded DNA polynucleotide molecules in the tagged library to generate an amplification product and sequencing the amplification product. In some embodiments, the method further comprises sequencing the double-stranded DNA polynucleotide molecules multiple times. In some embodiments, (b) comprises sequencing the entire insert. In some embodiments, (c) further comprises collapsing the sequence reads in each family to generate a consensus sequence. In some embodiments, (d) comprises calling multiple consecutive bases from at least a subset of the sequence reads to identify single nucleotide variations (SNVs) in the double-stranded DNA molecule.
[0066] Another aspect provides a method for detecting disease cell heterogeneity from a sample containing polynucleotides derived from somatic cells and disease cells, comprising quantifying polynucleotides in the sample having a nucleotide sequence variant at each of a plurality of genetic loci, determining copy number variation (CNV) at each of the plurality of genetic loci, where the CNV indicates a genetic dosage of the locus in the disease cell polynucleotides, and determining, with a programmed computer processor, a relative measure of the quantity of polynucleotides having the sequence variant at the locus per genetic dosage at the locus for each of the plurality of genetic loci, and comparing the relative measures at each of the plurality of genetic loci, where different relative measures indicate tumor heterogeneity.
[0067] In another aspect, a method includes subjecting a subject to one or more pulsed therapy cycles, each pulsed therapy cycle including (a) a first period during which a drug is administered at a first amount and (b) a second period during which the drug is administered at a second, reduced amount, wherein (i) the first period is characterized by a tumor burden detected above a first clinical level, and (ii) the second period is characterized by a tumor burden detected below a second clinical level. The present invention provides, for example, the following items. (Item 1) 1. A method for determining a quantitative measure indicative of the number of individual double-stranded deoxyribonucleic acid (DNA) molecules in a sample, comprising: (a) determining a quantitative measure of individual DNA molecules for which both strands are detected; (b) determining a quantitative measure of individual DNA molecules in which only one of the DNA strands is detected; (c) inferring from (a) and (b) above a quantitative measure of individual DNA molecules for which neither strand was detected; (d) using (a)-(c) to determine the quantitative measure indicative of the number of individual double-stranded DNA molecules in the sample; A method comprising: (Item 2) 2. The method of claim 1, further comprising detecting copy number variation in the sample by determining a normalized quantitative measure determined in step (d) at each of one or more genetic loci and determining copy number variation based on the normalized measure. (Item 3) 2. The method of claim 1, wherein the sample comprises double-stranded polynucleotide molecules substantially sourced from cell-free nucleic acid. (Item 4) 2. The method of claim 1, wherein determining the quantitative measure of individual DNA molecules comprises tagging the DNA molecules with a set of duplex tags, each duplex tag differently tagging a complementary strand of a double-stranded DNA molecule in the sample to provide tagged strands. (Item 5) 5. The method of claim 4, further comprising sequencing at least a portion of the tagged strands to produce a set of sequence reads. (Item 6) 6. The method of claim 5, further comprising sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads. (Item 7) 7. The method of claim 6, further comprising determining a quantitative measure of (i) the paired reads and (ii) the unpaired reads that map to each of one or more genetic loci, and determining a quantitative measure of total double-stranded DNA molecules in the sample that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each locus. (Item 8) 1. A method for reducing distortion in a sequencing assay, comprising: (a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides; (b) tagging the test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides; (c) mixing tagged control parent polynucleotides with tagged test parent polynucleotides to form a pool; (d) determining the quantities of tagged control parent polynucleotides and tagged test parent polynucleotides; (e) using said quantity of tagged control parent polynucleotides to reduce distortion in said quantity of tagged test parent polynucleotides; A method comprising: (Item 9) 9. The method of claim 8, wherein the first tag set comprises a plurality of tags, each tag in the first tag set comprises the same control tag and an identifying tag, and the first tag set comprises a plurality of different identifying tags. (Item 10) 10. The method of claim 9, wherein the second tag set comprises a plurality of tags, each tag in the second tag set comprises the same test tag and an identifying tag, the test tag is distinguishable from the control tag, and the second tag set comprises a plurality of different identifying tags. (Item 11) 10. The method of claim 9, wherein (d) comprises amplifying tagged parent polynucleotides in the pool to form a pool of amplified, tagged polynucleotides; and sequencing the amplified, tagged polynucleotides in the amplified pool to produce a plurality of sequence reads. (Item 12) 12. The method of claim 11, further comprising: grouping sequence reads into families, each family containing sequence reads generated from the same parent polynucleotide, the grouping optionally being based on information derived from an identifying tag and start / end sequences of the parent polynucleotides; and optionally determining a consensus sequence for each of a plurality of parent polynucleotides derived from the plurality of sequence reads in a group. (Item 13) 9. The method of claim 8, wherein (d) comprises determining copy number variation in the test parent polynucleotides at more than one locus based on the relative amounts of test parent polynucleotides and control parent polynucleotides mapping to the locus. (Item 14) a set of library adaptors comprising a plurality of polynucleotide molecules having molecular barcodes, wherein said plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, and said molecular barcodes are at least 4 nucleotide bases in length; (a) the molecular barcodes are different from one another and have an edit distance of at least 1 between one another; (b) the molecular barcodes are located at least one nucleotide base away from a terminus of each polynucleotide molecule; (c) optionally, at least one terminal base is identical in all of said polynucleotide molecules; (d) none of said polynucleotide molecules contains a complete sequencer motif A set of library adapters. (Item 15) 15. The set of library adaptors of item 14, wherein the polynucleotide molecules are identical except for the molecular barcodes. (Item 16) 15. The set of library adaptors of item 14, wherein each of the plurality of polynucleotide molecules has a double-stranded portion and at least one single-stranded portion. (Item 17) 17. The set of library adaptors of Item 16, wherein the double-stranded portion has a molecular barcode among the plurality of molecular barcodes. (Item 18) 18. The set of library adaptors of item 17, wherein the given molecular barcode is a randomer. (Item 19) 17. The set of library adaptors of item 16, wherein each of the plurality of polynucleotide molecules further comprises a strand-identification barcode on the at least one single-stranded portion. (Item 20) 20. The set of library adaptors of item 19, wherein said strand-identification barcode comprises at least 4 nucleotide bases. (Item 21) 17. The set of library adaptors of item 16, wherein the single-stranded portion has a partial sequencer motif. (Item 22) 15. The set of library adaptors of item 14, wherein the polynucleotide molecules have a sequence of terminal nucleotides that are the same. (Item 23) 15. The set of library adaptors of item 14, wherein each of the plurality of polynucleotide molecules is Y-shaped, bubble-shaped, or hairpin-shaped. (Item 24) 15. The set of library adaptors of item 14, wherein none of the polynucleotide molecules contains a sample identification motif. (Item 25) 15. The set of library adaptors of item 14, wherein said molecular barcodes are at least 10 nucleotide bases in length. (Item 26) 15. The set of library adaptors of Item 14, wherein each of the plurality of polynucleotide molecules is between 10 nucleotide bases and 60 nucleotide bases in length. (Item 27) 15. The set of library adaptors of item 14, wherein the at least one terminal base is identical in all of the polynucleotide molecules. (Item 28) 15. The set of library adaptors of item 14, wherein the molecular barcodes are located at least 10 nucleotide bases away from a terminus of each polynucleotide molecule. (Item 29) 15. The set of library adaptors of item 14, consisting essentially of said plurality of polynucleotide molecules. (Item 30) (a) tagging a collection of polynucleotides with a plurality of polynucleotide molecules from a library of adaptors according to item 14 to create a collection of tagged polynucleotides; (b) amplifying the collection of tagged polynucleotides in the presence of sequencing adaptors, wherein the sequencing adaptors have primers having nucleotide sequences that are selectively hybridizable to complementary sequences in the plurality of polynucleotide molecules; A method comprising: (Item 31) 1. A method for detecting or quantifying rare deoxyribonucleic acid (DNA) in a heterogeneous population of original DNA fragments, wherein the rare DNA has a concentration that is less than 1%, the method comprising: (a) tagging the original DNA fragments in a single reaction such that greater than 30% of the original DNA fragments are tagged at both ends with library adaptors that comprise molecular barcodes, thereby providing tagged DNA fragments; (b) performing high-fidelity amplification on the tagged DNA fragments; (c) optionally, selectively enriching a subset of said tagged DNA fragments; (d) sequencing one or both strands of the tagged, amplified, and optionally selectively enriched DNA fragments to obtain sequence reads comprising the nucleotide sequences of the molecular barcodes and at least a portion of the original DNA fragments; (e) determining a consensus read from the sequence reads that is representative of a single strand of the original DNA fragment; (f) quantifying the consensus reads to detect or quantify the rare DNA with greater than 99.9% specificity; A method comprising: (Item 32) 32. The method of claim 31, wherein step (e) comprises comparing sequence reads having the same or similar molecular barcodes and the same or similar ends of the fragment sequences. (Item 33) 33. The method of claim 32, wherein the comparing step further comprises performing a phylogenetic analysis on the sequence reads that have the same or similar molecular barcodes. (Item 34) 33. The method of claim 32, wherein the molecular barcodes comprise barcodes having an edit distance of up to 3. (Item 35) 32. The method of item 31, wherein the ends of the fragment sequences comprise fragment sequences having an edit distance of up to 3. (Item 36) 32. The method of claim 31, further comprising sorting sequence reads into paired reads and unpaired reads, and quantifying the number of paired reads and unpaired reads that map to each of the one or more genetic loci. (Item 37) 32. The method of item 31, wherein the tagging occurs by having an excess amount of library adaptors compared to the original DNA fragments. (Item 38) 32. The method of claim 31, further comprising binning the sequence reads according to the molecular barcodes and sequence information from at least one end of each of the original DNA fragments to create bins of single-stranded reads. (Item 39) 39. The method of claim 38, further comprising determining, in each bin, the sequence of a given original DNA fragment among the original DNA fragments by analyzing sequence reads. (Item 40) 40. The method of claim 39, further comprising detecting or quantifying the rare DNA by comparing the number of times each base occurs at each position in the genome represented by the tagged, amplified, and optionally enriched DNA fragments. (Item 41) 32. The method of claim 31, further comprising selectively enriching a subset of the tagged DNA fragments. (Item 42) 42. The method of claim 41, further comprising, after enrichment, amplifying the enriched tagged DNA fragments in the presence of sequencing adaptors comprising primers. (Item 43) 32. The method of claim 31, wherein the DNA fragments are tagged with polynucleotide molecules derived from the library of adaptors of claim 1. (Item 44) 1. A method for processing and / or analyzing a nucleic acid sample of a subject, comprising: (a) exposing polynucleotide fragments from the nucleic acid sample to a set of library adaptors to generate tagged polynucleotide fragments; (b) subjecting the tagged polynucleotide fragments to a nucleic acid amplification reaction under conditions that yield amplified polynucleotide fragments as amplification products of the tagged polynucleotide fragments; the set of library adaptors comprises a plurality of polynucleotide molecules having molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, and the molecular barcodes are at least 4 nucleotide bases in length; (1) the molecular barcodes are different from one another and have an edit distance of at least 1 between one another; (2) the molecular barcode is located at least one nucleotide base away from the end of each polynucleotide molecule; (3) optionally, at least one terminal base is identical in all of said polynucleotide molecules; (4) A method wherein none of the polynucleotide molecules contains a complete sequencer motif. (Item 45) 45. The method of claim 44, further comprising determining the nucleotide sequence of the amplified tagged polynucleotide fragments. (Item 46) 46. The method of claim 45, wherein the nucleotide sequences of the amplified tagged polynucleotide fragments are determined without polymerase chain reaction (PCR). (Item 47) 46. The method of claim 45, further comprising analyzing the nucleotide sequence with a programmed computer processor to identify one or more genetic variants in the nucleotide sample of the subject. (Item 48) 45. The method of claim 44, wherein the nucleic acid sample is a cell-free nucleic acid sample. (Item 49) 45. The method of claim 44, wherein exposing the polynucleotide fragments of the nucleic acid sample to the plurality of polynucleotide molecules produces the tagged polynucleotide fragments with a conversion efficiency of at least 10%. (Item 50) The subjecting step may comprise administering to a subject a gene encoding one of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PT 45. The method of claim 44, comprising amplifying the tagged polynucleotide fragment from a sequence corresponding to a gene selected from the group consisting of PN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. (Item 51) (a) generating a plurality of sequence reads from a plurality of polynucleotide molecules, wherein the plurality of polynucleotide molecules spans genomic loci of a target genome, the genomic loci being selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, and HRAS; , IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1; (b) grouping, with a computer processor, the plurality of sequence reads into families, each family comprising sequence reads derived from one of the template polynucleotides; (c) for each of the families, merging sequence reads to generate a consensus sequence; (d) calling the consensus sequence at a given genomic locus among the genomic loci; (e) at said given genomic locus: i. genetic variants within the call; ii. the frequency of the genetic alteration among said calls; iii. Total number of calls; and iv. The total number of changes in the call detecting one of A method comprising: (Item 52) 52. The method of claim 51, wherein each family comprises sequence reads derived from only one of the template polynucleotides. (Item 53) 52. The method of Item 51, further comprising performing (d) to (e) at additional genomic loci among the genomic loci. (Item 54) 54. The method of claim 53, further comprising determining copy number variation at one of the given genomic locus and the additional genomic locus based on the counts at the given genomic locus and the additional genomic locus. (Item 55) 52. The method of claim 51, wherein the grouping step comprises classifying the plurality of sequence reads into families by identifying (i) distinct molecular barcodes coupled to the plurality of polynucleotide molecules and (ii) similarities between the plurality of sequence reads, wherein each family comprises a plurality of nucleic acid sequences associated with a distinct combination of molecular barcodes and similar or identical sequence reads. (Item 56) 52. The method of claim 51, wherein the consensus sequence is generated by assessing a quantitative measure or statistical significance level of each of the sequence reads. (Item 57) 52. The system of item 51, wherein the plurality of genes comprises at least 10 of the plurality of genes selected from the group. (Item 58) (a) providing in a single reaction vessel template polynucleotide molecules and a set of library adaptors, wherein said library adaptors are polynucleotide molecules having different molecular barcodes, and none of said library adaptors contains a complete sequencer motif; (b) coupling the library adaptors to the template polynucleotide molecules in the single reaction vessel at an efficiency of at least 10%, thereby tagging each template polynucleotide with a tagging combination from among a plurality of different tagging combinations to produce tagged polynucleotide molecules; (c) subjecting the tagged polynucleotide molecules to an amplification reaction under conditions that yield amplified polynucleotide molecules as amplification products of the tagged polynucleotide molecules; (d) sequencing the amplified polynucleotide molecules; A method comprising: (Item 59) 59. The method of claim 58, wherein the library adaptors are identical except for the molecular barcodes. (Item 60) 59. The method of claim 58, wherein each of the library adaptors has a double-stranded portion and at least one single-stranded portion, wherein the single-stranded portion has a partial sequencer motif. (Item 61) 59. The method of claim 58, wherein the library adaptors couple to both ends of the template polynucleotide molecule. (Item 62) Item 59. The method of item 58, wherein the efficiency is at least 30%. (Item 63) 60. The method of claim 58, further comprising identifying genetic variants upon sequencing the amplified polynucleotide molecules. (Item 64) 59. The method of claim 58, wherein the sequencing step comprises (i) subjecting the amplified polynucleotide molecules to an additional amplification reaction under conditions that produce additional amplified polynucleotide molecules as amplification products of the amplified polynucleotide molecules; and (ii) sequencing the additional amplified polynucleotide molecules. (Item 65) Item 66. The method of Item 64, wherein the additional amplification is performed in the presence of sequencing adaptors. 59. The method of claim 58, wherein (b) and (c) are performed without aliquoting the tagged polynucleotide molecules. (Item 67) 1. A system for analyzing a target nucleic acid molecule of a subject, comprising: a communication interface for receiving nucleic acid sequence reads for a plurality of polynucleotide molecules covering genomic loci of a target genome; a computer memory that stores the nucleic acid sequence reads for the plurality of polynucleotide molecules received by the communication interface; and a computer processor operably coupled to the communication interface and the memory, and programmed to: (i) group the plurality of sequence reads into families, each family including sequence reads derived from one of the template polynucleotides; (ii) for each of the families, merge sequence reads to generate a consensus sequence; (iii) call the consensus sequence at a given genomic locus among the genomic loci; and (iv) detect, at the given genomic locus, any of genetic variants among the calls, frequencies of genetic alterations among the calls, total number of calls, and total number of alterations among the calls. and wherein the genomic locus is selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, and MLH1. , MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. system. (Item 68) ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1 , ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, P A set of oligonucleotide molecules that selectively hybridize to at least five genes selected from the group consisting of TPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. (Item 69) 69. The set according to Item 68, wherein the oligonucleotide molecules are 10 to 200 bases in length. (Item 70) 69. The kit of item 68, wherein the oligonucleotide molecules selectively hybridize to exon regions of the at least five genes. (Item 71) 71. The kit of item 70, wherein the oligonucleotide molecules selectively hybridize to at least 30 exons in the at least 5 genes. (Item 72) 72. The kit of claim 71, wherein a plurality of oligonucleotide molecules selectively hybridize to each of the at least 30 exons. (Item 73) 73. The kit of item 72, wherein the oligonucleotide molecules that hybridize to each exon have sequences that overlap with at least one other oligonucleotide molecule. (Item 74) a first container containing a plurality of library adaptors, each having a different molecular barcode; a second container containing a plurality of sequencing adaptors, each sequencing adaptor comprising at least a portion of a sequencer motif and optionally a sample barcode; Kit including: (Item 75) 75. The kit of item 74, wherein the sequencing adaptor comprises the sample barcode. (Item 76) 1. A method for detecting sequence variants in a cell-free DNA sample, comprising detecting rare DNA at concentrations less than 1% with greater than 99.9% specificity. (Item 77) (a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule comprising a first and a second complementary strand; (b) tagging the double-stranded polynucleotide molecules with a set of duplex tags, each duplex tag differently tags the first and second complementary strands of a double-stranded polynucleotide molecule in the set; (c) sequencing at least a portion of the tagged strands to produce a set of sequence reads; (d) reducing and / or tracking redundancy in said set of sequence reads; (e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads; (f) determining a quantitative measure of (i) the paired reads and (ii) the unpaired reads that map to each of one or more genetic loci; and (g) estimating, with a programmed computer processor, a quantitative measure of total double-stranded polynucleotide molecules in the set that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each genetic locus. A method comprising: (Item 78) (h) detecting copy number variation in the sample by determining a normalized total quantitative measure determined in step (g) at each of the one or more loci and determining copy number variation based on the normalized measure. (Item 79) 78. The method of claim 77, wherein the sample comprises double-stranded polynucleotide molecules sourced substantially from cell-free nucleic acid. (Item 80) 78. The method of claim 77, wherein the duplex tag is not a sequencing adaptor. (Item 81) 78. The method of claim 77, wherein reducing redundancy in the set of sequence reads comprises collapsing sequence reads produced from amplified products of original polynucleotide molecules in the sample back to the original polynucleotide molecules. (Item 82) 82. The method of claim 81, further comprising determining a consensus sequence of the original polynucleotide molecule. (Item 83) 83. The method of claim 82, further comprising identifying polynucleotide molecules at one or more loci that contain sequence variants. (Item 84) 83. The method of claim 82, further comprising determining a quantitative measure of paired reads mapping to a locus, wherein both strands of the pair comprise a sequence variant. (Item 85) 85. The method of claim 84, further comprising determining a quantitative measure of paired molecules, where only one member of the pair has a sequence variant, and / or determining a quantitative measure of unpaired molecules that have a sequence variant. (Item 86) (a) receiving from a sequencer into a memory a set of sequence reads of polynucleotides tagged with duplex tags; (b) reducing and / or tracking redundancy in said set of sequence reads; (c) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads; (d) determining a quantitative measure of (i) the paired reads and (ii) the unpaired reads that map to each of one or more genetic loci; and (e) estimating a quantitative measure of total double-stranded polynucleotide molecules in the set that map to each of the one or more genetic loci based on the quantitative measure of paired reads and unpaired reads that map to each genetic locus. A method comprising: (Item 87) (a) providing a sample comprising a set of double-stranded polynucleotide molecules, each double-stranded polynucleotide molecule comprising a first and a second complementary strand; (b) tagging the double-stranded polynucleotide molecules with a set of duplex tags, each duplex tag differently tags the first and second complementary strands of a double-stranded polynucleotide molecule in the set; (c) sequencing at least a portion of the tagged strands to produce a set of sequence reads; (d) reducing and / or tracking redundancy in said set of sequence reads; (e) sorting sequence reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads; (f) determining quantitative measures of at least two of: (i) the paired reads, (ii) the unpaired reads mapping to each of one or more genetic loci, (iii) the read depth of the paired reads, and (iv) the read depth of the unpaired reads; A method comprising: (Item 88) (a) tagging control parent polynucleotides with a first tag set to produce tagged control parent polynucleotides, wherein the first tag set comprises a plurality of tags, each tag in the first tag set comprising the same control tag and an identifying tag, and wherein the tag set comprises a plurality of different identifying tags; (b) tagging the test parent polynucleotides with a second tag set to produce tagged test parent polynucleotides, wherein the second tag set comprises a plurality of tags, each tag in the second tag set comprising the same test tag distinguishable from the control tag and the identifying tag, and wherein the second tag set comprises a plurality of different identifying tags; (c) mixing tagged control parent polynucleotides with tagged test parent polynucleotides to form a pool; (d) amplifying tagged parent polynucleotides in the pool to form a pool of amplified tagged polynucleotides; (e) sequencing the amplified, tagged polynucleotides in the amplified pool to produce a plurality of sequence reads; (f) grouping sequence reads into families, each family including sequence reads generated from the same parent polynucleotide, the grouping optionally being based on information from identifying tags and start / end sequences of said parent polynucleotides, and optionally determining a consensus sequence for each of a plurality of parent polynucleotides from said plurality of sequence reads in a group; (g) classifying each family or consensus sequence as a control parent polynucleotide or a test parent polynucleotide based on having a test tag or a control tag; (h) determining quantitative measures of control parent polynucleotides and control test polynucleotides that map to each of the at least two genetic loci; (i) determining copy number variation in the test parent polynucleotides at at least one genetic locus based on the relative abundances of test parent polynucleotides and control parent polynucleotides mapping to the at least one genetic locus; A method comprising: (Item 89) (a) generating a plurality of sequence reads from a plurality of template polynucleotides, each polynucleotide being mapped to a genomic locus; (b) grouping the sequence reads into families, each family comprising sequence reads generated from one of the template polynucleotides; (c) calling nucleotide bases or sequences at said genomic loci for each of said families; (d) at said genomic locus: i. genomic alterations among the calls; ii. the frequency of the genetic alteration among said calls; iii. total number of calls; iv. The total number of changes in the call detecting one of A method comprising: (Item 90) 90. The method of claim 89, wherein calling comprises any of phylogenetic analysis, voting, weighing, assigning a probability to each read at the locus in a family and calling the nucleotide base with the highest probability. (Item 91) 90. The method of claim 89, performed at two loci, comprising determining the CNV at one of the loci based on counts at each of the loci. (Item 92) (a) ligating adaptors to double-stranded deoxyribonucleic acid (DNA) polynucleotides to produce a tagged library comprising inserts from the double-stranded DNA polynucleotides and having between 4 and 1 million different tags, wherein the ligation is performed in a single reaction vessel and the adaptors comprise molecular barcodes; (b) generating a plurality of sequence reads for each of the double-stranded DNA polynucleotides in the tagged library; (c) grouping sequence reads into families based on information in the tags and information at the ends of the inserts, each family including sequence reads generated from a single DNA polynucleotide among the double-stranded DNA polynucleotides; (d) calling the nucleotide base at each position in the double-stranded DNA molecule based on the nucleotide base at that position in members of the family; A method comprising: (Item 93) 94. The method of claim 93, wherein (d) comprises calling multiple consecutive bases from at least a subset of the sequence reads to identify single nucleotide variations (SNVs) in the double-stranded DNA molecule.
[0068] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. Incorporation by Reference
[0069] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0070] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "figure" and "FIG.") [Brief explanation of the drawings]
[0071] [Figure 1] FIG. 1 is a flowchart representation of the disclosed method for determining copy number variations (CNVs).
[0072] [Figure 2] FIG. 2 depicts the mapping of pairs and singlets to locus A and locus B in the genome.
[0073] [Figure 3] FIG. 3 shows the reference sequence encoding locus A.
[0074] [Figure 4A] 4A-4C show amplification, sequencing, redundancy reduction and pairing of complementary molecules. [Figure 4B] 4A-4C show amplification, sequencing, redundancy reduction and pairing of complementary molecules. [Figure 4C] 4A-4C show amplification, sequencing, redundancy reduction and pairing of complementary molecules.
[0075] [Figure 5] FIG. 5 shows the increased confidence in detecting sequence variants by pairing reads from Watson and Crick strands.
[0076] [Figure 6] FIG. 6 illustrates a computer system that is programmed or otherwise configured to implement various methods of the present disclosure.
[0077] [Figure 7] Figure 7 is a schematic representation of a system for analyzing samples containing nucleic acids from a user, including a sequencer; bioinformatics software for reporting analysis by, for example, a handheld device or desktop computer; and an internet connection.
[0078] [Figure 8] FIG. 8 is a flow chart representation of a method of the present invention for determining CNV using pooled test and control pools.
[0079] [Figure 9] 9A-9C schematically illustrate a method for tagging polynucleotide molecules with library adaptors and subsequently sequencing adaptors. DETAILED DESCRIPTION OF THE INVENTION
[0080] While various embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed.
[0081] The term "genetic variant" as used herein generally refers to an alteration, variant, or polymorphism in a nucleic acid sample or genome of a subject. Such alteration, variant, or polymorphism can be relative to a reference genome, which can be the reference genome of the subject or other individual. A single nucleotide polymorphism (SNP) is one form of polymorphism. In some examples, one or more polymorphisms include one or more single nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable length tandem repeats, and / or flanking sequences. Copy number variants (CNVs), transversions, and other rearrangements are also forms of genetic variation. A genomic alteration can be a base change, an insertion, a deletion, a repeat, a copy number variation, or a sequence fragment. It can be a translation or a transversion.
[0082] The term "polynucleotide" as used herein generally refers to a molecule comprising one or more nucleic acid subunits. A polynucleotide can comprise one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. A nucleotide can comprise A, C, G, T, or U, or variants thereof. A nucleotide can comprise any subunit that can be incorporated into a growing nucleic acid chain. Such subunits can be A, C, G, T, or U, or any other subunit specific to one or more complementary A, C, G, T, or U, or complementary to a purine (i.e., A or G, or variants thereof) or pyrimidine (i.e., C, T, or U, or variants thereof). A subunit can resolve individual nucleic acid bases or groups of bases (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or their uracil counterparts). In some instances, the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), or a derivative thereof. The polynucleotide can be single-stranded or double-stranded.
[0083] The term "subject," as used herein, generally refers to animals, such as mammalian species (e.g., humans) or avian (e.g., avian) species, or other organisms, such as plants. More specifically, a subject can be a vertebrate, mammal, mouse, primate, monkey, or human. Animals include, but are not limited to, farm animals, sport animals, and pets. A subject can be a healthy individual, an individual who is diseased or suspected of having a disease or who is predisposed to having a disease, or an individual who is in need of therapy or suspected of being in need of therapy. A subject can be a patient.
[0084] The term "genome" generally refers to the entirety of an organism's genetic information. A genome can be coded in either DNA or RNA. A genome can include coding regions that code for proteins as well as non-coding regions. A genome can include all chromosomal sequences in an organism together. For example, the human genome has a total of 46 chromosomes. All of these sequences together make up the human genome.
[0085] The terms "adapter(s)," "adapter(s)," and "tag(s)" are used interchangeably throughout this specification. An adapter or tag can be coupled to a polynucleotide sequence to "tag" it by any approach, including ligation, hybridization, or other approaches.
[0086] The term "library adaptor" or "library adaptor" as used herein generally refers to a molecule (e.g., polynucleotide) whose identity (e.g., sequence) can be used to distinguish polynucleotides in a biological sample (also referred to herein as a "sample").
[0087] The term "sequencing adaptor," as used herein, generally refers to a molecule (e.g., a polynucleotide) adapted to allow a sequencing instrument to sequence a target polynucleotide, such as by interacting with the target polynucleotide to enable sequencing. The sequencing adaptor enables sequencing of the target polynucleotide by the sequencing instrument. In one example, the sequencing adaptor comprises a nucleotide sequence that hybridizes or binds to a capture polynucleotide attached to a solid support of a sequencing system, such as a flow cell. In another example, the sequencing adaptor comprises a nucleotide sequence that hybridizes or binds to a polynucleotide to generate a hairpin loop that enables sequencing of the target polynucleotide by the sequencing system. The sequencing adaptor can comprise a sequencer motif that is complementary to the flow cell sequence of another molecule (e.g., a polynucleotide) and can be a nucleotide sequence that can be used by the sequencing system to sequence the target polynucleotide. The sequencer motif can also comprise a primer sequence for use in sequencing, such as sequencing by synthesis. The sequencer motif can comprise a sequence(s) required for coupling the library adaptor to a sequencing system and sequencing the target polynucleotide.
[0088] As used herein, the terms "at least," "at most," or "about," when preceding a numerical series, refer to every member of the series, unless otherwise specified.
[0089] The term "about" and its grammatical equivalents in reference to a reference numerical value can include a range of values from that value up to plus or minus 10%. For example, the amount "about 10" can include amounts from 9 to 11. In other embodiments, the term "about" in reference to a reference numerical value can include a range of values from that value plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.
[0090] The term "at least" and its grammatical equivalents in reference to a reference numerical value can include the reference numerical value and more than that value. For example, the amount "at least 10" can include the value 10, as well as any number greater than 10, such as 11, 100, and 1,000.
[0091] The term "at most" and its grammatical equivalents in reference to a reference numerical value can include the reference numerical value and less than that value. For example, the amount "at most 10" can include the value 10 and any number less than 10, such as 9, 8, 5, 1, 0.5, and 0.1.
[0092] 1. Methods for Processing and / or Analyzing Nucleic Acid Samples
[0093] An embodiment of the present disclosure provides a method for determining genomic alterations in a nucleic acid sample of a subject. Figure 1 shows a method for determining copy number variations (CNVs). This method can be performed to determine other genomic alterations, such as SNVs.
[0094] A. Polynucleotide Isolation
[0095] The methods disclosed herein can include isolating one or more polynucleotides. A polynucleotide can include any type of nucleic acid, such as a genomic nucleic acid sequence or an artificial sequence (e.g., a sequence not present in genomic nucleic acid). For example, an artificial sequence can contain non-naturally occurring nucleotides. A polynucleotide can also include both genomic nucleic acid and an artificial sequence in any portion. For example, a polynucleotide can include 1-99% genomic nucleic acid and 99%-1% artificial sequence, totaling up to 100%. Thus, fractional percentages are also contemplated. For example, a ratio of 99.1% to 0.9% is contemplated.
[0096] A polynucleotide can comprise any type of nucleic acid, such as DNA and / or RNA. For example, if the polynucleotide is DNA, it can be genomic DNA, complementary DNA (cDNA), or any other deoxyribonucleic acid. A polynucleotide can also be cell-free DNA (cfDNA). For example, a polynucleotide can be circulating DNA. Circulating DNA can include circulating tumor DNA (ctDNA). A polynucleotide can be double-stranded or single-stranded. Alternatively, a polynucleotide can comprise a combination of double-stranded and single-stranded portions.
[0097] Polynucleotides do not have to be cell-free. In some cases, polynucleotides can be isolated from a sample. For example, in step (102) (FIG. 1), double-stranded polynucleotides are isolated from the sample. The sample can be any biological sample isolated from a subject. For example, the sample can include, without limitation, bodily fluids, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsies, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid, fluid in the intercellular spaces including gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine, or any other bodily fluid. Bodily fluids can include saliva, blood, or serum. For example, polynucleotides can be cell-free DNA isolated from bodily fluids, such as blood or serum. The sample may also be a tumor sample that can be obtained from a subject by a variety of approaches, including but not limited to, venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage, scraping, surgical incision or intervention, or other approaches.
[0098] A sample can contain various amounts of nucleic acid containing a genome equivalent. For example, a sample of about 30 ng DNA can contain about 10,000 (10 4 ) haploid human genome equivalent, while cfDNA can contain approximately 200 billion (2 × 10 11) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.
[0099] The sample can contain nucleic acids from different sources. For example, the sample can contain germline DNA or somatic DNA. The sample can contain nucleic acids carrying mutations. For example, the sample can contain DNA carrying germline mutations and / or somatic mutations. The sample can also contain DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations).
[0100] B. Tagging
[0101] The polynucleotides disclosed herein can be tagged. For example, in step (104) (FIG. 1), a double-stranded polynucleotide is tagged with a duplex tag, a tag that differentially labels the complementary strands (i.e., the "Watson" and "Crick" strands) of the double-stranded molecule. In one embodiment, the duplex tag is a polynucleotide having complementary and non-complementary portions.
[0102] A tag can be any type of molecule attached to a polynucleotide, including, but not limited to, a nucleic acid, a chemical compound, a fluorescent probe, or a radioactive probe. A tag can be an oligonucleotide (e.g., DNA or RNA). A tag can include a known sequence, an unknown sequence, or both. A tag can include a random sequence, a predetermined sequence, or both. A tag can be double-stranded or single-stranded. A double-stranded tag can be a duplex tag. A double-stranded tag can include two complementary strands. Alternatively, a double-stranded tag can include a hybridized portion and an unhybridized portion. A double-stranded tag can be Y-shaped, e.g., a hybridized portion is present at one end of the tag and an unhybridized portion is present at the opposite end of the tag. One such example is the "Y adapter" used in Illumina sequencing. Other examples include hairpin adapters or bubble adapters. Bubble adapters have a non-complementary sequence flanked on both sides by complementary sequences.
[0103] The tagging disclosed herein can be performed using any method. A polynucleotide can be tagged with an adaptor by hybridization. For example, the adaptor can have a nucleotide sequence complementary to at least a portion of the sequence of the polynucleotide. Alternatively, a polynucleotide can be tagged with an adaptor by ligation.
[0104] For example, tagging can include the use of one or more enzymes. The enzyme can be a ligase. The ligase can be a DNA ligase. For example, the DNA ligase can be T4 DNA ligase, E. coli DNA ligase, and / or mammalian ligase. The mammalian ligase can be DNA ligase I, DNA ligase III, or DNA ligase IV. The ligase can be a thermostable ligase. The tag can be ligated to a blunt end of a polynucleotide (blunt-end ligation). Alternatively, the tag can be ligated to a sticky end of a polynucleotide (sticky-end ligation). The efficiency of ligation can be increased by optimizing various conditions. The efficiency of ligation can be increased by optimizing the reaction time of ligation. For example, the ligation reaction time can be less than 12 hours, e.g., less than 1 hour, less than 2 hours, less than 3 hours, less than 4 hours, less than 5 hours, less than 6 hours, less than 7 hours, less than 8 hours, less than 9 hours, less than 10 hours, less than 11 hours, less than 12 hours, less than 13 hours, less than 14 hours, less than 15 hours, less than 16 hours, less than 17 hours, less than 18 hours, less than 19 hours, or less than 20 hours. In a specific example, the ligation reaction time is less than 20 hours. The efficiency of ligation can be increased by optimizing the ligase concentration in the reaction. For example, the ligase concentration can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, or at least 600 units per microliter. Efficiency can also be optimized by adding or varying the concentration of enzymes, enzyme cofactors, or other additives suitable for ligation, and / or optimizing the temperature of the solution containing the enzymes. Efficiency can also be optimized by varying the order of addition of the various components of the reaction. The ends of the tag sequences can contain dinucleotides to increase ligation efficiency.When a tag includes a non-complementary portion (e.g., a Y-shaped adaptor), the sequence of the complementary portion of the tag adaptor can include one or more selected sequences that promote ligation efficiency. Preferably, such sequences are located at the end of the tag. Such sequences can include 1, 2, 3, 4, 5, or 6 terminal bases. Ligation efficiency can also be increased by using a reaction solution with high viscosity (e.g., a low Reynolds number). For example, the solution can have a Reynolds number of less than 3000, less than 2000, less than 1000, less than 900, less than 800, less than 700, less than 600, less than 500, less than 400, less than 300, less than 200, less than 100, less than 50, less than 25, or less than 10. It is also contemplated that a roughly uniform distribution of fragments (e.g., a tight standard deviation) can be used to increase ligation efficiency. For example, the variation in fragment size can vary by less than 20%, less than 15%, less than 10%, less than 5%, or less than 1%.Tagging can also include, for example, primer extension by polymerase chain reaction (PCR).Tagging can also include ligation-based PCR, multiplex PCR, single-strand ligation, or single-strand circularization.
[0105] In some cases, the tags herein include molecular barcodes. Such molecular barcodes can be used to distinguish polynucleotides in a sample. Preferably, the molecular barcodes are different from each other. For example, the molecular barcodes can have a difference between them that can be characterized by a predetermined edit distance or Hamming distance. In some cases, the molecular barcodes herein have a minimum edit distance of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. To further improve the efficiency of converting untagged molecules to tagged molecules (e.g., tagging), short tags are preferably utilized. For example, in some embodiments, the library adapter tag can be up to 65, 60, 55, 50, 45, 40, or 35 nucleotide bases in length. Such a collection of short library barcodes preferably includes a large number of different molecular barcodes, for example, at least 2, 4, 6, 8, 10, 12, 14, 16, 18, or 20 different barcodes, with a minimum edit distance of 1, 2, 3, or more.
[0106] Thus, a collection of molecules can include one or more tags. In some cases, some molecules in the collection can include an identifying tag ("identifier"), such as a molecular barcode, that is not shared by any other molecules in the collection. For example, in some cases of a collection of molecules, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least At least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% can comprise an identifier or molecular barcode that is not shared by any other molecule in the collection. As used herein, a collection of molecules is considered "uniquely tagged" if at least 95% of the molecules in the collection each have an identifier (a "unique tag" or "unique identifier") that is not shared by any other molecule in the collection.A collection of molecules is considered to be "non-uniquely tagged" if at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% or about 50% of the molecules in the collection each have an identification tag or molecular barcode (" non-unique tag " or " non-unique identifier ") shared by at least one other molecule in the collection.Therefore, in a non-uniquely tagged group, 1% or less of the molecules are uniquely tagged.For example, in a non-uniquely tagged group, 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45% or 50% or less of the molecules can be uniquely tagged.
[0107] Based on the estimated number of molecules in sample, can use a large number of different tags.In some tagging methods, the number of different tags can be at least the same as the estimated number of molecules in sample.In other tagging methods, the number of different tags can be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 100 or 1000 times more than the estimated number of molecules in sample.In specific tagging, can use at least 2 times (or more than) the number of different tags than the estimated number of molecules in sample.
[0108] The molecules in the sample can be tagged non-uniquely.In this case, the number of tags or molecular barcodes used is less than the number of molecules to be tagged in the sample.For example, 100, 50, 40, 30, 20 or 10 or less unique tags or molecular barcodes are used to tag complex samples, such as cell-free DNA samples that have many different fragments.
[0109] Polynucleotides to be tagged can be fragmented naturally or using other approaches, such as shearing. Polynucleotides can be fragmented by certain methods, including, but not limited to, mechanical shearing, passing the sample through a syringe, sonication, heat treatment (e.g., 90°C for 30 minutes), and / or nuclease treatment (e.g., using DNase, RNase, endonuclease, exonuclease, and / or restriction enzymes).
[0110] The polynucleotide fragments (prior to tagging) can comprise sequences of any length. For example, the polynucleotide fragments (prior to tagging) can be at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, The polynucleotide fragments can comprise lengths of 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more nucleotides. The polynucleotide fragments are preferably about the average length of cell-free DNA. For example, the polynucleotide fragments can comprise a length of about 160 bases. The polynucleotide fragments can also be fragmented from larger fragments into smaller fragments, down to a length of about 160 bases.
[0111] The tagged polynucleotide can include a sequence associated with cancer. The cancer-associated sequence can include a single nucleotide variation (SNV), a copy number variation (CNV), an insertion, a deletion, and / or a rearrangement.
[0112] The polynucleotides are used to identify acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brainstem glioma, brain tumor, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumor, breast cancer, bronchial tumor, Burkitt lymphoma, non-Hodgkin's lymphoma, carcinoid tumor, cervical cancer, chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing's sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrous histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin's lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell lung cancer, small cell lung cancer, melanoma, oral cancer, myelodysplastic syndrome, multiple myeloma The sequences may include sequences associated with cancer, such as medulloblastoma, nasal cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papilloma, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumor, plasma cell neoplasm, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome, skin cancer, non-melanoma, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's hypergammaglobulinemia, and / or Wilms' tumor.
[0113] A haploid human genome equivalent contains approximately 3 picograms of DNA. A sample of approximately 1 microgram of DNA contains approximately 300,000 haploid human genome equivalents. Improvements in sequencing can be achieved as long as at least some of the overlapping or cognate polynucleotides have unique identifiers relative to each other, i.e., have different tags. However, in certain embodiments, the number of tags used is selected so that there is at least a 95% probability that all overlapping molecules starting at any one position will have a unique identifier. For example, in a sample containing approximately 10,000 haploid human genome equivalents of fragmented genomic DNA, e.g., cfDNA, z is expected to be between 2 and 8. Such a population can be tagged with between about 10 and 100 different identifiers, e.g., about 2, 4, 9, 16, 25, 36, 49, 64, 81, or 100 different identifiers.
[0114] Nucleic acid barcodes with identifiable sequences, including molecular barcodes, can be used for tagging. For example, multiple DNA barcodes can contain various numbers of nucleotide sequences. Multiple DNA barcodes with 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more identifiable nucleotide sequences can be used. When attached to only one end of a polynucleotide, multiple DNA barcodes can produce 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more different identifiers. Alternatively, when attached to both ends of a polynucleotide, multiple DNA barcodes can generate 4, 9, 16, 25, 36, 49, 64, 81, 100, 121, 144, 169, 196, 225, 256, 289, 324, 361, 400 or more different identifiers (which is 2 times the number when DNA barcodes are attached to only one end of a polynucleotide). In one example, multiple DNA barcodes having 6, 7, 8, 9, or 10 identifiable nucleotide sequences can be used. When attached to both ends of a polynucleotide, these generate 36, 49, 64, 81, or 100 possible different identifiers, respectively. In a particular example, multiple DNA barcodes can include 8 identifiable nucleotide sequences. When attached to only one end of a polynucleotide, multiple DNA barcodes can generate 8 different identifiers. Alternatively, when attached to both ends of a polynucleotide, multiple DNA barcodes can generate 64 different identifiers. A sample tagged in this manner can be a sample having a range of about 10 ng to about 100 ng, about 1 μg, or about 10 μg of fragmented polynucleotides, e.g., genomic DNA, e.g., cfDNA.
[0115] Polynucleotides can be uniquely identified in various ways. Polynucleotides can be uniquely identified by unique DNA barcodes. For example, any two polynucleotides in a sample are attached to two different DNA barcodes. Alternatively, polynucleotides can be uniquely identified by a combination of DNA barcodes and one or more endogenous sequences of the polynucleotide. For example, any two polynucleotides in a sample can be attached to the same DNA barcode, but the two polynucleotides can still be identified by different endogenous sequences. The endogenous sequence can be located at the end of the polynucleotide. For example, the endogenous sequence can be adjacent to (e.g., between) the attached DNA barcode. In some cases, the endogenous sequence can be at least 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases in length. Preferably, the endogenous sequence is the end sequence of the fragment / polynucleotide to be analyzed. The endogenous sequence can be the length of the sequence. For example, multiple DNA barcodes, including 8 different DNA barcodes, can be attached to both ends of each polynucleotide in sample.Each polynucleotide in sample can be identified by the combination of DNA barcode and the endogenous sequence of about 10 base pairs at the end of polynucleotide.Without being limited by theory, the endogenous sequence of polynucleotide can also be the entire polynucleotide sequence.
[0116] Also disclosed herein are compositions of tagged polynucleotides. The tagged polynucleotides can be single-stranded. Alternatively, the tagged polynucleotides can be double-stranded (e.g., double-stranded tagged polynucleotides). Thus, the present invention also provides compositions of double-stranded tagged polynucleotides. The polynucleotides can comprise any type of nucleic acid (DNA and / or RNA). The polynucleotides include any type of DNA disclosed herein. For example, the polynucleotides can comprise DNA, such as fragmented DNA or cfDNA. The set of polynucleotides in the composition that are mapped to mappable base positions in a genome can be non-uniquely tagged, i.e., the number of different identifiers can be at least two and less than the number of polynucleotides that are mapped to mappable base positions. The number of different identifiers can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 and less than the number of polynucleotides mapped to mappable base positions.
[0117] In some cases, as compositions increase in volume from about 1 ng to about 10 μg or more, a larger set of different molecular barcodes can be used, for example, between 5 and 100 different library adaptors can be used to tag polynucleotides in a cfDNA sample.
[0118] The systems and methods disclosed herein can be used in applications involving the assignment of molecular barcodes. Molecular barcodes can be assigned to any type of polynucleotide disclosed herein. For example, molecular barcodes can be assigned to cell-free polynucleotides (e.g., cfDNA). In many cases, the identifiers disclosed herein can be barcode oligonucleotides used to tag polynucleotides. Barcode identifiers can be nucleic acid oligonucleotides (e.g., DNA oligonucleotides). Barcode identifiers can be single-stranded. Alternatively, barcode identifiers can be double-stranded. Barcode identifiers can be attached to polynucleotides using any of the methods disclosed herein. For example, barcode identifiers can be attached to polynucleotides by enzymatic ligation. Barcode identifiers can also be incorporated into polynucleotides by PCR. In other cases, the reaction can include the addition of a metal isotope to the analyte directly or via an isotope-labeled probe. In general, the assignment of unique or non-unique identifiers or molecular barcodes in reactions of the present disclosure can follow, for example, the methods and systems described in U.S. Patent Application Publication Nos. 2001 / 0053519, 2003 / 0152490, 2011 / 0160078, and U.S. Patent No. 6,582,908, each of which is incorporated by reference in its entirety.
[0119] The identifiers or molecular barcodes used herein can be entirely endogenous, allowing for circular ligation of individual fragments followed by random shearing or targeted amplification, in which case the combination of the new molecular start and stop points and the original intramolecular ligation points can form a specific identifier.
[0120] As used herein, an identifier or molecular barcode can include any type of oligonucleotide. In some cases, the identifier can be a predetermined, random, or semi-random sequence oligonucleotide. The identifier can be a barcode. For example, multiple barcodes can be used such that the barcodes are not necessarily unique to each other within the plurality. Alternatively, multiple barcodes can be used such that each barcode is unique to any other barcode within the plurality. A barcode can include a specific sequence (e.g., a predetermined sequence) that can be tracked individually. Furthermore, a barcode can be attached to an individual molecule (e.g., by ligation) such that the combination of the barcode and the sequence to which it can be ligated creates a specific sequence that can be tracked individually. As described herein, detection of a barcode in combination with sequence data at the beginning (start) and / or end (stop) of a sequence read can enable the assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read can also be used to assign a unique identity to such a molecule. As described herein, fragments derived from a single strand of nucleic acid can be assigned unique identities, thereby allowing subsequent identification of fragments derived from the parent strand.In this way, polynucleotides in a sample can be uniquely or substantially uniquely tagged.Double-stranded tags can include degenerate or semi-degenerate nucleotide sequences, for example, random degenerate sequences.Nucleotide sequences can include any number of nucleotides.For example, nucleotide sequences can include 1 (when using non-natural nucleotides), 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In a particular example, the sequence can include 7 nucleotides, hi another example, the sequence can include 8 nucleotides.The sequence can also comprise 9 nucleotides. The sequence can comprise 10 nucleotides.
[0121] Barcode can comprise adjacent or non-adjacent sequences.When 4 nucleotides are not interrupted by any other nucleotide, barcodes that comprise at least 1, 2, 3, 4, 5 or more nucleotides are adjacent or non-adjacent sequences.For example, when barcode comprises sequence TTGC, if barcode is TTGC, barcode is adjacent.On the other hand, if barcode is TTXGC (wherein X is nucleobase), barcode is non-adjacent.
[0122] Identifiers or molecular barcodes can have n-mer sequences that can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides in length. Tags herein can include any range of nucleotide lengths. For example, sequences can be between 2-100, 10-90, 20-80, 30-70, 40-60, or about 50 nucleotides in length.
[0123] The tag can include a double-stranded fixed reference sequence downstream of the identifier or molecular barcode. Alternatively, the tag can include a double-stranded fixed reference sequence upstream or downstream of the identifier or molecular barcode. Each strand of the double-stranded fixed reference sequence can be, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length.
[0124] C. Adapter
[0125] A library of polynucleotide molecules can be synthesized for use in sequencing.For example, a library of polynucleotides can be created that includes a plurality of polynucleotide molecules that are each less than or equal to 100, 90, 80, 70, 60, 50, 45, 40 or 35 nucleic acid (or nucleotide) bases in length.Each of the plurality of polynucleotide molecules can be less than or equal to 35 nucleic acid bases in length.Each of the plurality of polynucleotide molecules can be less than or equal to 30 nucleic acid bases in length.The plurality of polynucleotide molecules can also be less than or equal to 250, 200, 150, 100 or 50 nucleic acid bases in length. Additionally, the plurality of polynucleotide molecules may include 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 56, 57, 58, 59 ... It can also be less than or equal to 4, 53, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11 or 10 nucleobases.
[0126] A polynucleotide library comprising a plurality of polynucleotide molecules can also have molecular barcode sequences (or molecular barcodes) that are distinct (with respect to each other) for at least four nucleic acid bases. A molecular barcode (also referred to herein as "barcode" or "identifier") sequence is a nucleotide sequence that distinguishes one polynucleotide from another. In other embodiments, polynucleotide molecules can have different barcode sequences for 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid bases.
[0127] A polynucleotide library containing a plurality of polynucleotide molecules can also have a plurality of different barcode sequences. For example, the plurality of polynucleotide molecules can have at least four different molecular barcode sequences. In some cases, the plurality of polynucleotide molecules has 2 to 100, 4 to 50, 4 to 30, 4 to 20, or 4 to 10 different molecular barcode sequences. The plurality of polynucleotide molecules can also have other ranges of different barcode sequences, such as 1 to 4, 2 to 5, 3 to 6, 4 to 7, 5 to 8, 6 to 9, 7 to 10, 8 to 11, 9 to 12, 10 to 13, 11 to 14, 12 to 15, 13 to 16, 14 to 17, 15 to 18, 16 to 19, 17 to 20, 18 to 21, 19 to 22, 20 to 23, 21 to 24, or 22 to 25 different barcode sequences. In other cases, the plurality of polynucleotide molecules comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21, The library adapters can have 4, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more different barcode sequences. In a particular example, the plurality of library adapters comprises at least 8 different sequences.
[0128] The position of different barcode sequences can vary within a plurality of polynucleotides.For example, different barcode sequences can be within 20, 15, 10, 9, 8, 7, 6, 5, 4, 3 or 2 nucleic acid bases from the end of each of a plurality of polynucleotide molecules.In one example, a plurality of polynucleotide molecules have different barcode sequences within 10 nucleic acid bases from the end.In another example, a plurality of polynucleotide molecules have different barcode sequences within 5 or 1 nucleic acid bases from the end.In other cases, different barcode sequences can be present at the end of each of a plurality of polynucleotide molecules.In other variations, distinct molecular barcode sequences are inserted at 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 from the end of each one of the plurality of polynucleotide molecules. , 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 1 11, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 1 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200 or more nucleobases.
[0129] The ends of the polynucleotide molecules can be adapted for ligation to a target nucleic acid molecule. For example, the ends can be blunt-ended. In other cases, the ends are adapted for hybridization to complementary sequences of the target nucleic acid molecule.
[0130] A polynucleotide library comprising a plurality of polynucleotide molecules can also have an edit distance of at least 1. In some cases, the edit distance relates to the individual bases of the plurality of polynucleotide molecules. In other cases, the plurality of polynucleotide molecules can have an edit distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more. The edit distance can be a Hamming distance.
[0131] In some cases, the plurality of polynucleotides does not contain a sequencing adaptor. The sequence adaptor can be a polynucleotide comprising a sequence that hybridizes to one or more sequencing adaptors or primers. The sequencing adaptor can further comprise a sequence that hybridizes to a solid support, for example, a flow cell sequence. The term "flow cell sequence" and its grammatical equivalents herein refer to a sequence that allows hybridization to a substrate, for example, by a primer attached to the substrate. The substrate can be a bead or a planar surface. In some embodiments, the flow cell sequence can allow attachment of a polynucleotide to a flow cell or surface (e.g., the surface of a bead, e.g., an Illumina flow cell).
[0132] If the plurality of polynucleotide molecules does not contain a sequencing adaptor or primer, each polynucleotide molecule in the plurality does not contain a nucleic acid sequence or other moiety adapted to allow sequencing of the target nucleic acid molecule by a given sequencing approach, such as Illumina, SOLiD, Pacific Biosciences, GeneReader, Oxford Nanopore, Complete Genomics, Gnu-Bio, Ion Torrent, Oxford Nanopore, or Genia. In some examples, if the plurality of polynucleotide molecules does not contain a sequencing adaptor or primer, the plurality of polynucleotide molecules does not contain a flow cell sequence. For example, the plurality of polynucleotide molecules cannot be attached to a flow cell, such as that used in an Illumina flow cell sequencer. However, these flow cell sequences can be added to the plurality of polynucleotide molecules, if necessary, by methods such as PCR amplification or ligation. Currently, an Illumina flow cell sequencer can be used. Alternatively, if the plurality of polynucleotide molecules does not contain sequencing adaptors or primers, the plurality of polynucleotide molecules does not contain hairpin adaptors or adaptors for generating hairpin loops in the target nucleic acid molecule, such as Pacific Bioscience SMRTbell™ adaptors. However, such hairpin adaptors can be added to the plurality of polynucleotide molecules, if necessary, by methods such as PCR amplification or ligation. The plurality of polynucleotide molecules can be circular or linear.
[0133] The plurality of polynucleotide molecules can be double-stranded. In some cases, the plurality of polynucleotide molecules can be single-stranded or can include hybridized and unhybridized regions. The plurality of polynucleotide molecules can be non-naturally occurring polynucleotide molecules.
[0134] The adaptor can be a polynucleotide molecule. The polynucleotide molecule can be Y-shaped, bubble-shaped, or hairpin-shaped. The hairpin adaptor can contain a restriction site(s) or a uracil-containing base. The adaptor can include a complementary portion and a non-complementary portion. The non-complementary portion can have an edit distance (e.g., Hamming distance). For example, the edit distance can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30. The complementary portion of the adaptor can comprise a sequence selected to allow and / or promote ligation to a polynucleotide, e.g., a sequence that allows and / or promotes ligation to a polynucleotide in high yield.
[0135] The polynucleotide molecules disclosed herein can be purified. In some cases, the polynucleotide molecules disclosed herein can be isolated polynucleotide molecules. In other cases, the polynucleotide molecules disclosed herein can be purified and isolated polynucleotide molecules.
[0136] In certain embodiments, each of the plurality of polynucleotide molecules is Y-shaped or hairpin-shaped. Each of the plurality of polynucleotide molecules can comprise a different barcode. The different barcode can be a randomer in the complementary portion (e.g., double-stranded portion) of the Y-shaped or hairpin-shaped adaptor. Alternatively, the different barcode can be present on one strand of the non-complementary portion (e.g., one of the Y-shaped arms). As noted above, the different barcodes can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more nucleobases (or any length described throughout this application), e.g., 7 bases. The barcodes can be contiguous or non-contiguous sequences, as described above. The plurality of polynucleotide molecules is 10 nucleobases to 35 nucleobases in length (or any length described above). Additionally, the plurality of polynucleotide molecules can comprise an edit distance (described above) that is a Hamming distance. The plurality of polynucleotide molecules can have distinct barcode sequences within 10 nucleobases of their termini.
[0137] In another embodiment, the plurality of polynucleotide molecules can be sequencing adaptors. The sequencing adaptors can include a sequence that hybridizes to one or more sequencing primers. The sequencing adaptors can further include a sequence that hybridizes to a solid support, e.g., a flow cell sequence. For example, the sequencing adaptors can be flow cell adaptors. The sequencing adaptors can be attached to one or both ends of a polynucleotide fragment. In another example, the sequencing adaptors can be hairpin-shaped. For example, the hairpin-shaped adaptors can include a complementary double-stranded portion and a loop portion, and the double-stranded portion can be attached (e.g., ligated) to a double-stranded polynucleotide. The hairpin-shaped sequencing adaptors can be attached to both ends of a polynucleotide fragment to generate a circular molecule that can be sequenced multiple times. Sequencing adapters are arranged end-to-end, with up to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54 , 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more bases. For example, the sequencing adaptor can be up to 70 bases end-to-end. The sequencing adaptor can include 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, or 50-70 bases end-to-end. In a particular example, the sequencing adaptor can comprise 20-30 bases from end to end. In another example, the sequencing adaptor can comprise 50-60 bases from end to end. The sequencing adaptor can comprise one or more barcodes. For example, the sequencing adaptor can comprise a sample barcode. The sample barcode can comprise a predetermined sequence.Sample barcode can be used to identify the source of polynucleotide.Sample barcode can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more (or any length described throughout this application) nucleic acid base, for example, at least 8 bases.As mentioned above, barcode can be contiguous or non-contiguous sequence.
[0138] The polynucleotide molecules described herein can be used as adaptors. An adaptor can include one or more identifiers. An adaptor can include an identifier with a random sequence. Alternatively, an adaptor can include an identifier with a predetermined sequence. Some adaptors can include an identifier with a random sequence and another identifier with a predetermined sequence. An adaptor including an identifier can be a double-stranded or single-stranded adaptor. An adaptor including an identifier can be a Y-shaped adaptor. A Y-shaped adaptor can include one or more identifiers with a random sequence. The one or more identifiers can be present in the hybridized and / or non-hybridized portions of the Y-shaped adaptor. A Y-shaped adaptor can include one or more identifiers with a predetermined sequence. The one or more identifiers with a predetermined sequence can be present in the hybridized and / or non-hybridized portions of the Y-shaped adaptor. A Y-shaped adaptor can include one or more identifiers with a random sequence and one or more identifiers with a predetermined sequence. For example, one or more identifiers having a random sequence can be present in the hybridized portion of the Y-shaped adaptor and / or the unhybridized portion of the Y-shaped adaptor. One or more identifiers having a predetermined sequence can be present in the hybridized portion of the Y-shaped adaptor and / or the unhybridized portion of the Y-shaped adaptor. In a particular example, a Y-shaped adaptor can include an identifier having a random sequence in its hybridized portion and an identifier having a predetermined sequence in its unhybridized portion. The identifiers can be any length disclosed herein. For example, a Y-shaped adaptor can include an identifier having a 7-nucleotide random sequence in its hybridized portion and an identifier having an 8-nucleotide predetermined sequence in its unhybridized portion.
[0139] The adaptor can comprise a double-stranded portion having a molecular barcode and at least one or two single-stranded portions. For example, the adaptor can be Y-shaped and comprise a double-stranded portion and two single-stranded portions. The single-stranded portions can comprise sequences that are not complementary to each other.
[0140] The adaptor can include an end having a sequence selected to allow the adaptor to be efficiently ligated or otherwise coupled to the polynucleotide (e.g., with at least about 20%, 30%, 40%, 50% efficiency). In some cases, the terminal nucleotides in the double-stranded portion of the adaptor are selected from a combination of purines and pyrimidines to provide efficient ligation.
[0141] In some examples, the set of library adaptors includes a plurality of polynucleotide molecules (library adaptors) having molecular barcodes. The library adaptors are less than or equal to 80, 70, 60, 50, 45, or 40 nucleotide bases in length. The molecular barcodes can be at least 4 nucleotide bases in length, but can be 4 to 20 nucleotide bases in length. The molecular barcodes can be different from each other and have an edit distance of at least 1, 2, 3, 4, or 5 between each other. The molecular barcodes are located at least 1, 2, 3, 4, 5, 10, or 20 nucleotide bases away from the end of each library adaptor. In some cases, at least one terminal base is identical in all library adaptors.
[0142] The library adaptors can be identical except for the molecular barcodes, for example, the library adaptors can have identical sequences but differ only with respect to the nucleotide sequence of the molecular barcodes.
[0143] Each of the library adaptors can have a double-stranded portion and at least one single-stranded portion. "Single-stranded portion" refers to an area of non-complementarity or overhang. In some cases, each of the library adaptors has a double-stranded portion and two single-stranded portions. The double-stranded portion can have a molecular barcode. In some cases, the molecular barcode is a randomer. Each of the library adaptors can further include a strand-identification barcode in the single-stranded portion. The strand-identification barcode can include at least 4 nucleotide bases, and in some cases, 4 to 20 nucleotide bases.
[0144] In some examples, each of the library adaptors has a double-stranded portion having a molecular barcode and two single-stranded portions. The single-stranded portions may not hybridize with each other. The single-stranded portions may not be fully complementary to each other.
[0145] The library adaptors can have the same sequence of terminal nucleotides in the double-stranded portion. The sequence of terminal nucleotides can be at least 2, 3, 4, 5, or 6 nucleotide bases in length. For example, one strand of the double-stranded portion of the library adaptor can have the sequence ACTT, TCGC, or TACC at its end, while the other strand can have a complementary sequence. In some cases, such a sequence is selected to optimize the efficiency of ligation of the library adaptor to the target polynucleotide. Such a sequence can be selected to optimize the binding interaction between the end of the library adaptor and the target polynucleotide.
[0146] In some cases, none of the library adaptors contain a sample identification motif (or sample molecular barcode). Such a sample identification motif can be provided by a sequencing adaptor. A sample identification motif can comprise at least 4, 5, 6, 7, 8, 9, 10, 20, 30 or 40 nucleotide bases of sequencer, which allows the polynucleotide molecules from a given sample to be identified from the polynucleotide molecules from other samples. For example, this can allow the polynucleotide molecules from two subjects to be sequenced in the same pool, and the sequence reads of the subjects can then be identified.
[0147] The sequencer motif comprises a nucleotide sequence(s) required for coupling the library adaptor to a sequencing system and sequencing the target polynucleotide coupled to the library adaptor. The sequencer motif can comprise a sequence complementary to a flow cell sequence and a sequence (sequencing initiation sequence) that can selectively hybridize to a primer (or priming sequence) for use in sequencing. For example, such a sequencing initiation sequence can be complementary to a primer used in sequencing by synthesis (e.g., Illumina). Such a primer can be included in the sequencing adaptor. The sequencing initiation sequence can be a primer hybridization site.
[0148] In some cases, none of the library adaptors contain a complete sequencer motif. The library adaptors can contain a partial sequencer motif or no sequencer motif. In some cases, the library adaptors include a sequencing initiation sequence. The library adaptors can include a sequencing initiation sequence but do not include a flow cell sequence. The sequencing initiation sequence can be complementary to a primer for sequencing. The primer can be a sequence-specific primer or a universal primer. Such a sequencing initiation sequence can be located in a single-stranded portion of the library adaptor. Alternatively, such a sequencing initiation sequence can be a priming site (e.g., a kink or nick) to allow a polymerase to couple to the library adaptor during sequencing.
[0149] In some cases, the partial or complete sequencer motif is provided by a sequencing adaptor. The sequencing adaptor can include a sample molecular barcode and a sequencer motif. The sequencing adaptors can be provided in a set separate from the library adaptors. The sequencing adaptors in a given set can be identical—i.e., contain the same sample barcode and sequencer motif.
[0150] The sequencing adaptor can include a sample identification motif and a sequencer motif. The sequencer motif can include a primer complementary to a sequencing initiation sequence. In some cases, the sequencer motif also includes a flow cell sequence or other sequence that allows the polynucleotide to be configured or arranged in a manner that allows the polynucleotide to be sequenced by a sequencer.
[0151]
[00130] The library adaptors and sequencing adaptors can each be partial adaptors, i.e., contain some, but not all, of the sequence necessary to enable sequencing by a sequencing platform. Together, they provide a complete adaptor. For example, the library adaptor can include a partial sequencer motif or no sequencer motif, but such sequencer motif is provided by the sequencing adaptor.
[0152] 9A-9C schematically illustrate a method for tagging a target polynucleotide molecule with a library adaptor. FIG. 9A shows the library adaptor as a partial adaptor containing a primer hybridization site on one strand and a molecular barcode toward the other end. The primer hybridization site can be a sequencing initiation sequence for subsequent sequencing. The library adaptor is less than or equal to 80 nucleotide bases in length. In FIG. 9B, the library adaptor is ligated at both ends of the target polynucleotide molecule to result in a tagged target polynucleotide molecule. The tagged target polynucleotide molecule can be subjected to nucleic acid amplification to generate copies of the target. Next, in FIG. 9C, a sequencing adaptor containing a sequencer motif is provided and hybridized to the tagged target polynucleotide molecule. The sequencing adaptor contains a sample identification motif. The sequencing adaptor can contain a sequence to enable sequencing of the tagged target by a given sequencer.
[0153] D. Sequencing
[0154] The tagged polynucleotides can be sequenced to generate sequence reads (e.g., as shown in step (106) in FIG. 1). For example, tagged double-stranded polynucleotides can be sequenced. Sequence reads can be generated from only one strand of the tagged double-stranded polynucleotide. Alternatively, both strands of the tagged double-stranded polynucleotide can generate sequence reads. The two strands of the tagged double-stranded polynucleotide can include the same tag. Alternatively, the two strands of the tagged double-stranded polynucleotide can include different tags. When the two strands of the tagged double-stranded polynucleotide are differently tagged, the sequence read generated from one strand (e.g., the Watson strand) can be distinguished from the sequence read generated from the other strand (e.g., the Crick strand). Sequencing can involve generating multiple sequence reads per molecule. This can occur, for example, as a result of amplification of individual polynucleotide strands, e.g., by PCR, during the sequencing process.
[0155] The methods disclosed herein can include polynucleotide amplification. Polynucleotide amplification results in the incorporation of nucleotides into a nucleic acid molecule or primer, thereby forming a new nucleic acid molecule complementary to a template nucleic acid. The newly formed polynucleotide molecule and its template can be used as a template for synthesizing additional polynucleotides. The amplified polynucleotide can be any nucleic acid, such as deoxyribonucleic acid, including genomic DNA, cDNA (complementary DNA), cfDNA, and circulating tumor DNA (ctDNA). The amplified polynucleotide can also be RNA. As used herein, a single amplification reaction can include many rounds of DNA replication. A DNA amplification reaction can include, for example, polymerase chain reaction (PCR). A single PCR reaction can include 2 to 100 "cycles" of denaturation, annealing, and synthesis of DNA molecules. For example, the amplification step can be performed for 2 to 7, 5 to 10, 6 to 11, 7 to 12, 8 to 13, 9 to 14, 10 to 15, 11 to 16, 12 to 17, 13 to 18, 14 to 19, or 15 to 20 cycles. PCR conditions can be optimized based on the GC content of the sequences, including the primers.
[0156] Nucleic acid amplification techniques can be used in conjunction with the assays described herein. Some amplification techniques are PCR methodologies, examples of which include solution PCR and in situ PCR. Examples of amplification methods include, but are not limited to, PCR. For example, amplification can include PCR-based amplification. Alternatively, amplification can include non-PCR-based amplification. Amplification of the template nucleic acid can include the use of one or more polymerases. For example, the polymerase can be a DNA polymerase or an RNA polymerase. In some cases, high-fidelity amplification is performed, such as by using a high-fidelity polymerase (e.g., Phusion® High-Fidelity DNA Polymerase) or a PCR protocol. In some cases, the polymerase can be a high-fidelity polymerase. For example, the polymerase can be KAPA HiFi DNA polymerase. The polymerase can also be Phusion DNA polymerase. The polymerase can be used under reaction conditions that reduce or minimize amplification bias due to, for example, fragment length, GC content, etc.
[0157] PCR amplification of a single strand of a polynucleotide will generate copies of both this strand and its complement.During sequencing, both strands and their complements will generate sequence reads.However, for example, the sequence read generated from the complement of a Watson strand can be identified as such because it has the complement of the double-stranded tag portion tagged on the original Watson strand.In contrast, the sequence read generated from a Crick strand or its amplification product will have the double-stranded tag portion tagged on the original Crick strand.In this way, the sequence read generated from the amplified product of the complement of a Watson strand can be distinguished from the complementary sequence read generated from the amplification product of the Crick strand of the original molecule.
[0158] All amplified polynucleotides can be submitted to a sequencing device for sequencing. Alternatively, a sampling or subset of all amplified polynucleotides can be submitted to a sequencing device for sequencing. For any original double-stranded polynucleotide, there are three possible sequencing outcomes. First, sequence reads can be generated from both complementary strands of the original molecule (i.e., from both the Watson strand and the Crick strand). Second, sequence reads can be generated from only one of the two complementary strands (i.e., from either the Watson strand or the Crick strand, but not both). Third, sequence reads cannot be generated from either of the two complementary strands. Consequently, counting unique sequence reads that map to a locus will underestimate the number of double-stranded polynucleotides in the original sample that map to this locus. Methods for estimating unseen and uncounted polynucleotides are described herein.
[0159] The sequencing method can be massively parallel sequencing, i.e., simultaneously (or in rapid succession) sequencing at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules. Sequencing methods can include, but are not limited to, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, single-molecule sequencing-by-synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single-molecule arrays (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms, and any other sequencing method known in the art.
[0160] For example, double-stranded tagged polynucleotides can be amplified, for example, by PCR (see, e.g., Figure 4A; the double-stranded tagged polynucleotides are referred to as mm' and nn'). In Figure 4A, the strand of a double-stranded polynucleotide containing sequence m has sequence tags w and y, while the strand of a double-stranded polynucleotide containing sequence m' has sequence tags x and z. Similarly, the strand of a double-stranded polynucleotide containing sequence n has sequence tags a and c, while the strand of a double-stranded polynucleotide containing sequence n' has sequence tags b and d. During amplification, each strand produces its own and its complementary sequence. However, the amplified progeny of original strand m containing complementary sequence m' is distinguishable from the amplified progeny of original strand m', for example, because the progeny of original strand m has the sequence 5'-y'm'w'-3', and the progeny of the original single-stranded m' strand has the sequence 5'-zm'x-3'. Figure 4B shows amplification in more detail. During amplification, errors represented by dots may be introduced into the amplification progeny. The applied progeny are sampled for sequencing, resulting in the sequence reads shown, so that not every strand produces a sequence read. Because a sequence read can originate from either a strand or its complement, both the sequence and the complement sequence will be included in the set of sequence reads. Note that a polynucleotide can have the same tag at each end. Thus, for a tag "a" and a polynucleotide "m," the first strand can be tagged as am-a', and the complement can be tagged as a-m'-a.
[0161] E. Determining Consensus Sequence Reads
[0162] The methods disclosed herein may include a step of determining a consensus sequence read among the sequence reads, such as by reducing or tracking redundancy (e.g., step (108), as shown in FIG. 1). Sequencing amplified polynucleotides can produce reads of several amplification products derived from the same original polynucleotide, referred to as "redundant reads." Identifying redundant reads allows for the determination of unique molecules in the original sample. If molecules in a sample are uniquely tagged, reads generated from the amplification of a single unique original molecule can be identified based on their distinct barcodes. Ignoring barcodes, reads derived from unique original molecules can be determined based on the sequences at the beginning and end of the read, optionally in combination with the length of the read. However, in certain cases, a sample may be expected to have multiple original molecules with the same start-stop sequence and the same length. Without barcoding, these molecules are difficult to distinguish from one another. However, when a collection of polynucleotides is non-uniquely tagged (i.e., when an original molecule shares the same identifier with at least one other original molecule), the combination of start / stop sequence and / or polynucleotide length with information from the barcode significantly increases the probability that any sequence read can be traced to an original polynucleotide, in part because, even in the absence of unique tagging, any two original polynucleotides with the same start / stop sequence and length are unlikely to be similarly tagged with the same identifier.
[0163] F. Collapse
[0164] Collapsing allows for a reduction in noise (i.e., background) generated at each step of the process. The methods disclosed herein can include collapsing a consensus sequence, e.g., generating it by comparing multiple sequence reads. For example, sequence reads generated from a single original polynucleotide can be used to generate a consensus sequence for such original polynucleotide. Repeated rounds of amplification can introduce errors into progeny polynucleotides. Also, because sequencing typically does not require perfect fidelity, sequencing errors are also introduced at this stage. However, comparison of sequence reads of molecules derived from a single original molecule, including molecules with sequence variants, can be analyzed to determine the original or "consensus" sequence. This can be done phylogenetically. A consensus sequence can be generated from a family of sequence reads by any of a variety of methods. Such methods include, for example, linear or nonlinear methods of consensus sequence construction derived from digital communication theory, information theory, or bioinformatics (such as voting (e.g., biased voting), averaging, statistical, maximum a posteriori or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods). For example, if all or most of the sequence reads tracing back to the original molecule have the same sequence variant, this variant was likely present in the original molecule. On the other hand, if a sequence variant is present in a subset of redundant sequence reads, this variant may have been introduced during amplification / sequencing and represents an artifact not present in the original. Furthermore, if only sequence reads derived from the Watson or Crick strand of the original polynucleotide contain the variant, the variant may have been introduced by single-sided DNA damage, a first-cycle PCR error, or contamination with polynucleotides amplified from different samples.
[0165] After the fragments are amplified and the sequences of the amplified fragments are read and aligned, the fragments are subjected to base calling, e.g., to determine the most likely nucleotide for each locus. However, variations in the number of amplified fragments and unseen amplified fragments (e.g., fragments whose sequences have not been read; there can be numerous reasons for this, such as amplification errors, sequencing read errors, being too long, too short, being truncated, etc.) can introduce errors in base calling. If there are too many unseen amplified fragments relative to the observed amplified fragments (the amplified fragments that are actually read), the reliability of base calling can be reduced.
[0166] Therefore, a method for correcting the number of unseen fragments in base calling is disclosed herein. For example, in the case of base calling of locus A (any locus), it is first assumed that there are N amplified fragments. The sequence readout can be derived from two types of fragments: double-stranded fragments and single-stranded fragments. Therefore, N1, N2, and N3 are assigned as the numbers of double-stranded, single-stranded, and unseen fragments, respectively. Thus, N = N1 + N2 + N3 (N1 and N2 are known from the sequence readout, and N and N3 are unknown). When the equation is solved for N (or N3), N3 (or N) is estimated.
[0167] Probability is used to estimate N, for example, assigning "p" to be the probability of detecting (or reading) a nucleotide at locus A in a single-stranded sequence readout.
[0168] For sequence readouts derived from duplexes, a nucleotide call from a double-stranded amplified fragment has a probability p*p=p^2, and the observation of all N1 duplexes has the following equation: N1=N*(p^2).
[0169] Regarding sequence readouts derived from single strands. Assuming one of the two strands is observed and the other is unseen, the probability of observing one strand is "p", but the probability of missing the other strand is (1-p). Furthermore, there is a factor of 2 due to not distinguishing between single strands of 5-primer origin and 3-primer origin. Therefore, a nucleotide call derived from a single-stranded amplified fragment has probability 2 x p x (1-p). Thus, the observation of all N2 single strands has the following equation: N2 = N x 2 x p x (1-p).
[0170] 'p' is also unknown. To solve for p, use the ratio of N1 to N2 and solve for 'p':
number
[0171] In addition to the ratio of paired to unpaired strands (a post-collapse measure), there is useful information in the pre-collapse read depth at each locus, which can be used to further improve the total molecule count call and / or increase the confidence in variant calls.
[0172] For example, Figure 4C demonstrates sequence reads that have been corrected for complementary sequences. The sequences generated from the original Watson strand or the original Crick strand can be distinguished based on their duplex tags. The sequences generated from the same original strand can be grouped. Examination of the sequences can allow the sequence of the original strand ("consensus sequence") to be inferred. In this case, for example, the sequence variant in the nn' molecule is included in the consensus sequence because it is included in all sequence reads, while other variants are observed as stray errors. After sequence collapse, the original polynucleotide pair can be identified based on their complementary sequences and duplex tags.
[0173] Figure 5 demonstrates the increased confidence in detecting sequence variants by pairing reads from Watson and Crick strands. The sequence nn' can contain sequence variants indicated by dots. In some cases, the sequence pp' does not contain sequence variants. Amplification, sequencing, redundancy reduction, and pairing can result in both Watson and Crick strands of the same original molecule containing sequence variants. In contrast, as a result of errors introduced during sampling in amplification and sequencing, the consensus sequence of Watson strand p can contain sequence variants, while the consensus sequence of Crick strand p' does not. Amplification and sequencing are less likely to introduce the same variant into both strands of a duplex (nn' sequence) than into one strand (pp' sequence). Therefore, variants in the pp' sequence are more likely to be artifacts, and variants in the nn' sequence are more likely to be present in the original molecule.
[0174] The methods disclosed herein can be used to correct errors resulting from experiments, such as PCR, amplification, and / or sequencing. For example, such methods can include the steps of attaching one or more double-stranded adapters to both ends of a double-stranded polynucleotide, thereby preparing a tagged double-stranded polynucleotide; amplifying the double-stranded tagged polynucleotide; sequencing both strands of the tagged polynucleotide; comparing the sequence of one strand with its complement to determine any errors introduced during sequencing; and correcting the sequence errors based on (d). The adapters used in this method can be any adapters disclosed herein, such as Y-shaped adapters. The adapters can include any barcodes (e.g., separate barcodes) disclosed herein.
[0175] G. Mapping
[0176] The sequence reads or consensus sequences can be mapped to one or more selected loci (e.g., as shown in step (110) in FIG. 1). A locus can be, for example, a specific nucleotide position within a genome, a sequence of nucleotides (e.g., an open reading frame), a fragment of a chromosome, an entire chromosome, or an entire genome. A locus can be a polymorphic locus. A polymorphic locus can be a locus where sequence variation exists in a population and / or in a subject and / or sample. A polymorphic locus can be generated by two or more distinct sequences coexisting at the same location in the genome. Distinct sequences can differ from each other by one or more nucleotide substitutions, deletions / insertions and / or duplications of any number of nucleotides, generally a relatively small number of nucleotides, such as less than 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide(s), among others. Polymorphic loci can be created by a single nucleotide position that varies within a population, for example, a single nucleotide variation (SNV) or single nucleotide polymorphism (SNP).
[0177] Reference genomes for mapping can include the genome of any species of interest. Human genome sequences useful as references can include the hg19 assembly or any previous or available hg assembly. Such sequences can be collated using the genome browser available at genome.ucsc.edu / index.html. Genomes of other species include, for example, PanTro2 (chimpanzee) and mm9 (mouse).
[0178] In the methods disclosed herein, collapsing can be performed before or after mapping. In some embodiments, collapsing can be performed before mapping. For example, sequence reads can be grouped into families based on their tags and one or more endogenous sequences, without considering the location where the reads are mapped in the genome. Then, family members can be collapsed into a consensus sequence. The consensus sequence can be generated using any of the collapsing methods disclosed herein. The consensus sequence can then be mapped to a location in the genome. The reads mapped to the locus can be quantified (e.g., counted). The percentage of reads carrying mutations at the locus can also be determined. Alternatively, collapsing can be performed after mapping. For example, all reads can first be mapped to the genome. Then, the reads can be grouped into families based on their tags and one or more endogenous sequences. Once the reads are mapped to the genome, a consensus base can be determined for each family at each locus. In other embodiments, a consensus sequence can be generated for one strand (e.g., Watson strand or Crick strand) of a DNA molecule. Mapping can be performed before or after the consensus sequence of one strand of the DNA molecule is determined. The number of doublets and singlets can be determined. These numbers can be used to calculate the number of unseen molecules. For example, the number of unseen molecules can be calculated using the following equation: N = D + S + U; D = Np(2), S = N2pq (where p = 1 - q, p is the probability of observation; q is the probability of missing a strand).
[0179] H. Grouping
[0180] The methods disclosed herein can also include a step of grouping sequence reads. Sequence reads can be grouped based on various types of sequences, such as the sequences of oligonucleotide tags (e.g., barcodes), the sequences of polynucleotide fragments, or any combination. For example, as shown in step (112) (FIG. 1), sequence reads can be grouped as follows: sequence reads generated from the "Watson" strand and the "Crick" strand of a double-stranded polynucleotide in a sample can be identified based on their duplex tags. In this manner, a sequence read or consensus sequence derived from the Watson strand of a double-stranded polynucleotide can be paired with a sequence read or consensus sequence derived from its complementary Crick strand. Paired sequence reads are referred to as "pairs."
[0181] A sequence read in which no sequence read corresponding to the complementary strand is found is termed a "singlet."
[0182] A double-stranded polynucleotide for which no sequence reads have been generated for either of the two complementary strands is referred to as an "unseen" molecule.
[0183] I. Quantification
[0184] The methods disclosed herein also include a step of quantifying sequence reads. For example, as shown in step (114) (FIG. 1), pairs and singlets mapping to a selected locus or each of a plurality of selected loci are quantified, e.g., counted.
[0185] Quantification can include estimating the number of polynucleotides (e.g., paired polynucleotides, singlet polynucleotides, or unseen polynucleotides) in a sample. For example, as shown in step (116) (FIG. 1), the number of double-stranded polynucleotides in a sample for which no sequence reads were generated ("unseen" polynucleotides) is estimated. The probability that a double-stranded polynucleotide will not generate a sequence read can be determined based on the relative number of pairs and singlets at any locus. This probability can be used to estimate the number of unseen polynucleotides.
[0186] In step (118), the estimate of the total number of double-stranded polynucleotides in the sample that map to the selected locus is the sum of the number of pairs, the number of singlets, and the number of unseen molecules that map to the locus.
[0187] The number of unseen original molecules in a sample can be estimated based on the relative number of pairs and singlets (Figure 2). Referring to Figure 2, as an example, counts for a particular genomic locus, locus A, are recorded, which shows that 1000 molecules are paired and 1000 molecules are unpaired. Assuming a uniform probability, p, for each Watson or Crick strand to undergo the transformation, the proportion of molecules that fail to undergo the transformation (unseen) can be calculated as follows: R = ratio of paired to unpaired molecules = 1, where R = 1 = p 2 / (2p(1-p)). This is because p=2 / 3 and the amount of molecules lost is (1-p) 2 = 1 / 9. So in this example, approximately 11% of the converted molecules are lost and not detected. Consider another genomic locus in the same sample, Locus B, which shows that 1440 molecules are paired and 720 are unpaired. Using the same method, we can infer that the number of missing molecules is only 4%. Comparing the two regions, we can assume that Locus A had 2000 unique molecules compared to 2160 molecules at Locus B - A difference of approximately 8%. However, by accurately adding up the missing molecules in each region, we estimate that there are 2000 / (8 / 9) = 2250 molecules in locus A and 2160 / 0.96 = 2250 molecules in locus B. Therefore, the counts in both regions are actually equal. This correction, and therefore even greater sensitivity, can be achieved by converting the original double-stranded nucleic acid molecules and keeping bioinformatically track of all paired and unpaired molecules at the end of the process. Similarly, the same procedure can be used to infer true copy number variation in regions that are likely to have similar counts of observed unique molecules. By taking into account the number of unseen molecules in two or more regions, copy number variation becomes apparent.
[0188] In addition to using the binomial distribution, other methods for estimating the number of unseen molecules include exponential, beta, gamma, or empirical distributions based on the redundancy of observed sequence reads. In the latter case, the distribution of read counts of paired and unpaired molecules can be derived from such redundancy to infer the underlying distribution of the original polynucleotide molecules at a particular locus. This can often result in a better estimate of the number of unseen molecules.
[0189] J. CNV detection
[0190] The methods disclosed herein also include a step of detecting CNVs. For example, once the total number of polynucleotides mapping to a locus is determined, as shown in step (120) (FIG. 1), this number can be used in a standard method for determining CNVs at that locus. The quantitative measure can be normalized to a standard. The standard can be the amount of any polynucleotide. In one method, the quantitative measure at a test locus can be normalized to a quantitative measure of polynucleotides mapping to a control locus in the genome, such as a gene of known copy number. The quantitative measure can be compared to the amount of nucleic acid in any sample disclosed herein. For example, in another method, the quantitative measure can be compared to the amount of nucleic acid in the original sample. For example, if the original sample contained 10,000 haploid gene equivalents, the quantitative measure can be compared to the measure expected for diploidy. In another method, the quantitative measure can be normalized to the measure from a control sample, and the normalized measures at different loci can be compared.
[0191] In some cases where copy number variation analysis is desired, sequence data can be: 1) aligned to a reference genome; 2) filtered and mapped; 3) partitioned into sequence windows or bins; 4) coverage reads counted per window; 5) coverage reads can then be normalized using a probabilistic or statistical modeling algorithm; 6) output files can be generated that reflect the distinct copy number states at various locations within the genome. In other cases where rare mutation analysis is desired, sequence data can be: 1) aligned to a reference genome; 2) filtered and mapped; 3) variant base frequencies can be calculated based on coverage reads for this specific base; 4) variant base frequencies can be normalized using a probabilistic, statistical, or probabilistic modeling algorithm; 5) output files can be generated that reflect the mutation states at various locations within the genome.
[0192] Once the sequence read coverage ratios have been determined, a probabilistic modeling algorithm can optionally be applied to convert the normalized ratios for each window region into distinct copy number states. In some cases, the algorithm can include a hidden Markov model. In other cases, the probabilistic model can include dynamic programming, support vector machines, Bayesian modeling, probabilistic modeling, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodology, or neural networks.
[0193] The methods disclosed herein can include detecting SNVs, CNVs, insertions, deletions, and / or rearrangements in specific regions of the genome, including ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, It may include sequences in genes such as MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, or NTRK1.
[0194] In some cases, the method uses a panel that includes the exons of one or more genes. The panel can also include the introns of one or more genes. The panel can also include the exons and introns of one or more genes. The one or more genes can be the genes disclosed above. The panel can include about 80,000 bases that cover the panel of genes. The panel can comprise about 1000, 2000, 3000, 4000, 5000, 10000, 15000, 20000, 25000, 30000, 35000, 40000, 45000, 50000, 55000, 60000, 65000, 70000, 75000, 80000, 85000, 90000, 95000, 100000, 105000, 110000, 115000, 120000, 125000 or more bases.
[0195] In some embodiments, the copy number of a gene can be reflected in the frequency of the genetic form of the gene in the sample.For example, in healthy individuals, copy number variation is not reflected in the variant in the gene in one chromosome, which is detected in about 50% of the molecules detected in the sample (for example, heterozygosity).Also, in healthy individuals, the duplication of a gene with a variant can be reflected in the variant detected in about 66% of the molecules detected in the sample.Therefore, if the tumor burden in DNA sample is 10%, without CNV, the frequency of somatic mutation in the gene in one chromosome of cancer cells can be about 5%.In the case of aneuploidy, the reverse can also be true.
[0196] The method disclosed herein can be used to determine whether sequence variants are present at germline level or are more likely to be caused by somatic mutations in, for example, cancer cells.For example, the sequence variants in genes that are detected at a level that is almost consistent with heterozygosity in germline are more likely to be the product of somatic mutations when CNVs are also detected in the gene.In some cases, to the extent that gene duplication in germline is expected to have variants consistent with gene dosage (for example, 66% for trisomy in locus), gene amplification detection with sequence variant dosage that is significantly different from this expected dosage indicates that CNVs are more likely to exist as a result of somatic mutations.
[0197] The method disclosed herein can also be used to infer tumor heterogeneity in the situation where sequence variants in two genes are detected at different frequencies.For example, when two genes are detected at different frequencies but their copy numbers are relatively equal, tumor heterogeneity can be inferred.Alternatively, when the difference in frequency between two sequence variants is consistent with the difference in copy number of the two genes, tumor homogeneity can be inferred.Therefore, for example, when EGFR variants are detected at 11% and KRAS variants are detected at 5%, and no CNV is detected in these genes, the difference in frequency may reflect tumor heterogeneity (for example, all tumor cells carry EGFR mutations, and half of tumor cells also carry KRAS mutations).Alternatively, when the EGFR gene carrying mutations is detected at twice the normal copy number, one interpretation is that it is a homogeneous population of tumor cells, and each cell carries mutations in EGFR and KRAS genes, but this KRAS gene is overlapping.
[0198] In response to chemotherapy, dominant tumor types may eventually be replaced by cancer cells carrying mutations that render the cancer unresponsive to the treatment regimen through Darwinian selection. The emergence of these resistant mutants can be delayed by the methods of the present invention. In one embodiment of this method, a subject is subjected to one or more pulsed therapy cycles, each pulsed therapy cycle comprising a first period in which a drug is administered at a first amount and a second cycle in which the drug is administered at a second, reduced amount. The first period can be characterized by a tumor burden detected above a first clinical level. The second period can be characterized by a tumor burden detected below a second clinical level. The first and second clinical levels can be different in different pulsed therapy cycles. For example, the first clinical level can be lower in subsequent cycles. The multiple cycles can include at least 2, 3, 4, 5, 6, 7, 8, or more cycles. For example, BRAF mutation V600E can be detected in the polynucleotides of diseased cells at an amount indicating a tumor burden of 5% in cfDNA. Chemotherapy can be initiated with dabrafenib. Subsequent testing can show that the amount of BRAF mutation in cfDNA falls below 0.5% or becomes undetectable. At this point, dabrafenib therapy can be stopped or significantly shortened. Further, subsequent testing can find that DNA with BRAF mutation has risen to 2.5% of polynucleotides in cfDNA. At this point, dabrafenib therapy can be resumed, for example, at the same level as the initial treatment. Subsequent testing can find that DNA with BRAF mutation has decreased to 0.5% of polynucleotides in cfDNA. Dabrafenib therapy can be stopped or reduced again. The cycle can be repeated multiple times.
[0199] Treatment intervention can also be changed by detecting the rise of mutations that are resistant to the original drug.For example, cancer with EGFR mutation L858R responds to treatment with erlotinib.However, cancer with EGFR mutation T790M is resistant to erlotinib.However, it is responsive to ruxolitinib.The method of the present invention involves monitoring changes in tumor profile, and changing treatment intervention when genetic variants associated with drug resistance rise to a predetermined clinical level.
[0200] The present invention discloses a method for detecting disease cell heterogeneity in a sample containing polynucleotides derived from somatic cells and disease cells, comprising the steps of: (a) quantifying polynucleotides in the sample that have sequence variants at each of a plurality of loci; (b) determining CNVs at each of the plurality of loci, the different relative amounts of disease molecules at the loci, where the CNVs indicate the gene dosage of the loci in the disease cell polynucleotides; (c) determining a relative measure of the amount of polynucleotides that have sequence variants at the loci per gene dosage at each of the plurality of loci; and (d) comparing the relative measures at each of the plurality of loci, where different relative measures indicate tumor heterogeneity. In the methods disclosed herein, gene dosage can be determined on a total molecule basis. For example, if there are 1x total molecules at a first locus and 1.2x molecules mapped to a second locus, the gene dosage is 1.2. The variants at this locus can be divided by 1.2. In some embodiments, the method disclosed herein can be used to detect any disease cell heterogeneity, for example, tumor cell heterogeneity.This method can be used to detect disease cell heterogeneity from the sample that contains any kind of polynucleotide, for example, cfDNA, genomic DNA, cDNA or ctDNA.In this method, quantification can include, for example, determining the number or relative amount of polynucleotide.Determining CNV can include mapping and normalizing the total molecules that have different relative amounts to gene locus.
[0201] In another embodiment, in response to chemotherapy, dominant tumor types may eventually be replaced by cancer cells carrying mutations that render the cancer unresponsive to the treatment regimen through Darwinian selection. The emergence of these resistant mutants can be delayed by the methods disclosed throughout the present specification. The methods disclosed herein can include the following steps: a) subjecting a subject to one or more pulsed treatment cycles, each pulsed treatment cycle comprising (i) a first period during which a drug is administered at a first amount and (ii) a second period during which the drug is administered at a second, reduced amount; (A) the first period is characterized by a tumor burden detected above a first clinical level, and (B) the second period is characterized by a tumor burden detected below a second clinical level.
[0202] K. Sequence Variant Detection
[0203] The systems and methods disclosed herein can be used to detect sequence variants, eg, SNVs. For example, sequence variants may be identified from multiple sequence reads, e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000,A consensus sequence derived from at least 7,000, at least 8,000, at least 9,000, or at least 10,000 or more sequence reads can be detected. The consensus sequence can be derived from sequence reads of a single-stranded polynucleotide. The consensus sequence can also be derived from sequence reads of one strand of a double-stranded polynucleotide (e.g., paired reads). In an exemplary method, paired reads allow for increased confidence in identifying the presence of a sequence variant in a molecule. For example, if both strands of a pair contain the same variant, it can be reasonably certain that the variant was present in the original molecule, since the probability that the same variant will be introduced into both strands during amplification / sequencing is rare. In contrast, if only one strand of a pair contains a sequence variant, this is more likely to be an artifact. Similarly, because the probability that a variant will be introduced once during amplification / sequencing is higher than twice, the confidence that a singlet with a sequence variant was present in the original molecule is less than the confidence that the variant is present in a double strand.
[0204] Other methods of copy number variation detection and sequence variant detection are described in PCT / US2013 / 058061, which is incorporated by reference herein in its entirety.
[0205] Sequence readings can be collapsed to generate consensus sequences, which can be mapped to reference sequences to identify genetic variants such as CNV or SNV.Alternatively, sequence readings can be mapped in advance or without mapping.In this case, sequence readings can be individually mapped to reference sequences to identify CNV or SNV.
[0206] Figure 3 shows a reference sequence encoding locus A. The polynucleotide in Figure 3 can be Y-shaped or have other shapes, such as a hairpin.
[0207] In some cases, SNV or multi-nucleotide variant (MNV) can be determined across multiple sequence reads at a given locus (for example, nucleotide base) by aligning the sequence reads corresponding to the locus.Then, multiple consecutive nucleotide bases from at least a subset of sequence reads are mapped to reference to determine the SNV or MNV in the polynucleotide molecule or its part corresponding to the read.The multiple consecutive nucleotide bases can span the actual, predicted or suspected position of SNV or MNV.The multiple consecutive nucleotide bases can span at least 3, 4, 5, 6, 7, 8, 9 or 10 nucleotide bases.
[0208] L. Nucleic Acid Detection / Quantification
[0209] The methods described herein can be used to tag nucleic acid fragments, such as deoxyribonucleic acid (DNA), with extremely high efficiency. This efficient tagging allows for efficient and accurate detection of rare DNA fragments in heterogeneous populations of original DNA fragments (such as cfDNA). A rare polynucleotide (e.g., rare DNA) can be a polynucleotide containing a genetic variant that occurs in a population of polynucleotides at a frequency of less than 10%, 5%, 4%, 3%, 2%, 1%, or 0.1%. A rare DNA can be a polynucleotide with a detectable characteristic at a concentration of less than 50%, 25%, 10%, 5%, 1%, or 0.1%.
[0210] Tagging can occur in a single reaction. In some cases, two or more reactions can be carried out and pooled together. Tagging each original DNA fragment in a single reaction results in more than 50% (for example, 60%, 70%, 80%, 90%, 95% or 99%) of the original DNA fragments being tagged at both ends with tags that comprise molecular barcodes, thereby providing tagged DNA fragments. Tagging can also result in more than 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the original DNA fragments being tagged on both ends with tags comprising molecular barcodes. Tagging can result in 100% of the original DNA fragments being tagged at both ends with tags that contain molecular barcodes. Tagging can also result in single-end tagging.
[0211] Tagging can also occur by using an excess amount of tag compared to the original DNA fragment.For example, the excess can be at least 5 times excess.In other cases, the excess can be at least 1.25, 1.5, 1.75, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 times or more excess.Tagging can include attachment to blunt ends or sticky ends.Tagging can also be performed by hybridization PCR. Tagging is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 pico- and / or microliter reaction volumes can also be performed.
[0212] The method can also include performing high-fidelity amplification on the tagged DNA fragments.Any high-fidelity DNA polymerase can be used.For example, the polymerase can be KAPA HiFi DNA polymerase or Phusion DNA polymerase.
[0213] Further, the method can include a step of selectively enriching a subset of tagged DNA fragments. For example, selective enrichment can be achieved by hybridization or amplification techniques. Selective enrichment can be achieved using a solid support (e.g., beads). The solid support (e.g., beads) can include a probe (e.g., an oligonucleotide that specifically hybridizes to a specific sequence). For example, the probe can hybridize to a specific genomic region, e.g., a gene. In some cases, the genomic region, e.g., a gene, can be a region associated with a disease, e.g., cancer. After enrichment, the selected fragments can be attached with any of the sequencing adaptors disclosed in the present invention. For example, the sequence adaptor can include a flow cell sequence, a sample barcode, or both. In another example, the sequence adaptor can be a hairpin-shaped adaptor and / or include a sample barcode. The resulting fragments can then be amplified and sequenced. In some cases, the adaptor does not include a sequencing primer region.
[0214] The method can include sequencing one or both strands of the DNA fragments. In one example, both strands of the DNA fragments are independently sequenced. The tagged, amplified, and / or selectively enriched DNA fragments are sequenced to obtain sequence reads that include molecular barcodes and sequence information of at least a portion of the original DNA fragments.
[0215] The method can include reducing or tracking redundancy in sequence reads (as described above) to determine a consensus read that is representative of a single strand of the original DNA fragment. For example, to reduce or track redundancy, the method can include comparing sequence reads with the same or similar molecular barcodes and the same or similar ends of fragment sequences. The method can include performing a phylogenetic analysis on sequence reads with the same or similar molecular barcodes. The molecular barcodes can have barcodes with varying edit distances (including any edit distances described throughout this application), for example, an edit distance of up to 3. The ends of the fragment sequences can include fragment sequences with varying edit distances (including any edit distances described throughout this application), for example, an edit distance of up to 3.
[0216] The method can include binning sequence reads according to molecular barcodes and sequence information.For example, the binning of sequence reads according to molecular barcodes and sequence information can be performed from at least one end of each original DNA fragment to generate a bin of single-stranded reads.The method can further include analyzing sequence reads in each bin to determine the sequence of a given original DNA fragment among the original DNA fragments.
[0217] In some cases, the sequence reads in each bin can be collapsed into a consensus sequence and then mapped to the genome. Alternatively, the sequence reads can be mapped to the genome prior to binning and then collapsed into a consensus sequence.
[0218] The method can also include sorting the sequence reads into paired reads and unpaired reads. After sorting, the number of paired reads and unpaired reads that map to each of one or more genetic loci can be quantified.
[0219] The method can include quantifying the consensus reads to detect and / or quantify rare DNA as described throughout the present application. The method can include detecting and / or quantifying rare DNA by comparing the number of times each base occurs at each position in the genome represented by the tagged, amplified, and / or enriched DNA fragments.
[0220] The method can include tagging original DNA fragments in a single reaction using a library of tags. The library can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10000, or any number of tags disclosed throughout this application. For example, the tag library can include at least 8 tags. The tag library can include 8 tags (which can generate 64 different possible combinations). The method can be performed so that a high percentage of the fragments, for example, greater than 50% (or any percentage described throughout this application), are tagged on both ends, and each tag includes a molecular barcode.
[0221] M. Nucleic Acid Processing and / or Analysis
[0222] The methods described throughout this application can be used to process and / or analyze a nucleic acid sample of a subject. The methods can include exposing polynucleotide fragments of the nucleic acid sample to a plurality of polynucleotide molecules to obtain tagged polynucleotide fragments. The plurality of polynucleotide molecules that can be used are described throughout this application.
[0223] For example, each of the plurality of polynucleotide molecules can be less than or equal to 40 nucleic acid bases in length, have distinct barcode sequences and an edit distance of at least 1 for at least 4 nucleic acid bases, each distinct barcode sequence is within 20 nucleic acid bases of an end of a respective one of the plurality of polynucleotide molecules, and the plurality of polynucleotide molecules are not sequencing adaptors.
[0224] The tagged polynucleotide fragments can be subjected to a nucleic acid amplification reaction under conditions that produce amplified polynucleotide fragments as amplification products of the tagged polynucleotide fragments. After amplification, the nucleotide sequence of the amplified tagged polynucleotide fragments is determined. In some cases, the nucleotide sequence of the amplified tagged polynucleotide fragments is determined without using polymerase chain reaction (PCR).
[0225] The method can include analyzing nucleotide sequences using a programmed computer processor to identify one or more genetic variants in the subject's nucleotide sample. Any genetic alteration can be identified, including but not limited to, base change(s), insertion(s), repeat(s), deletion(s), copy number variation(s), epigenetic modification(s), nucleosome binding site(s), copy number change(s) due to replication origin(s), and transversion(s). Other genetic alterations can include, but are not limited to, one or more tumor-associated genetic alterations.
[0226] The subject of the method may be suspected of having a disease. For example, the subject may be suspected of having cancer. The method may include collecting a nucleic acid sample from the subject. The nucleic acid sample may be collected from blood, plasma, serum, urine, saliva, mucosal excretion, sputum, feces, cerebrospinal fluid, skin, hair, sweat, and / or tears. The nucleic acid sample may be a cell-free nucleic acid sample. In some cases, the nucleic acid sample is collected from 100 nanograms (ng) or less of double-stranded polynucleotide molecules from the subject.
[0227] The polynucleotide fragment can comprise a double-stranded polynucleotide molecule. In some cases, multiple polynucleotide molecules are coupled to the polynucleotide fragment by blunt-end ligation, sticky-end ligation, molecular inversion probing, polymerase chain reaction (PCR), ligation-based PCR, multiplex PCR, single-stranded ligation, or single-stranded circularization.
[0228] The methods described herein provide for highly efficient tagging of nucleic acids. For example, exposing polynucleotide fragments of a nucleic acid sample to a plurality of polynucleotide molecules produces tagged polynucleotide fragments with a conversion efficiency of at least 30%, for example, at least 50% (e.g., 60%, 70%, 80%, 90%, 95% or 99%). Conversion efficiencies of at least 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% can be achieved.
[0229] The method can result in tagged polynucleotide fragments that share a common polynucleotide molecule. For example, at least 5%, 6%, 7%, 8%, 9%, 10%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the tagged polynucleotide fragments share a common polynucleotide molecule. The method can include generating polynucleotide fragments from a nucleic acid sample.
[0230] In some cases, the subjecting step of the method comprises administering to a patient a gene encoding a gene encoding a marker selected from the group consisting of ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA , PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. Additionally, any combination of these genes can be amplified. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or all 54 of these genes can be amplified.
[0231] The methods described herein can include generating a plurality of sequence reads from a plurality of polynucleotide molecules. The plurality of polynucleotide molecules can cover genomic loci of a target genome. For example, the genomic loci can correspond to a plurality of genes listed above. Furthermore, the genomic loci can be any combination of these genes. Any given genomic locus can include at least two nucleic acid bases. Any given genomic locus can also include multiple nucleic acid bases, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid bases.
[0232] The method can include grouping a plurality of sequence reads into families using a computer processor. Each family can include sequence reads derived from one of the template polynucleotides. Each family can include sequence reads derived from only one of the template polynucleotides. For each family, the sequence reads can be merged to generate a consensus sequence. The grouping can include classifying the plurality of sequence reads into families by (i) identifying distinct molecular barcodes coupled to a plurality of polynucleotide molecules and (ii) identifying similarities between the plurality of sequence reads, where each family includes a plurality of nucleic acid sequences associated with a distinct combination of molecular barcodes and similar or identical sequence reads.
[0233] Once integrated, a consensus sequence can be called at a given genomic locus among the genomic loci. For any given genomic locus, any of the following can be determined: i) the genetic variants among the calls; ii) the frequency of genetic alterations among the calls; iii) the total number of calls; and iv) the total number of alterations among the calls. The calls can include calling at least one nucleic acid base at a given genomic locus. The calls can include calling multiple nucleic acid bases at a given genomic locus. In some cases, the calls can include phylogenetic analysis, voting (e.g., biased voting), weighing, assigning a probability to each read at a locus in a family, or calling the base with the highest probability. The consensus sequence can be generated by evaluating a quantitative measure or statistical significance level for each of the sequence reads. When a quantitative measure is performed, the method can include using a binomial distribution, an exponential distribution, a beta distribution, or an empirical distribution. However, the frequency of the base at a particular position can also be used to make the call, for example, if 51% or more of the reads are "A" at this position, the base can be called "A" at that particular position. The method can further include mapping the consensus sequence to the target genome.
[0234] The method can further include making a consensus call at an additional genomic locus among the genomic loci. The method can include determining copy number variation at one of the given genomic locus and the additional genomic locus based on the counts at the given genomic locus and the additional genomic locus.
[0235] The methods described herein can include providing a library of template polynucleotide molecules and adaptor polynucleotide molecules in a reaction vessel. The adaptor polynucleotide molecules can have between 2 and 1,000 different barcode sequences and, in some cases, are not sequencing adaptors. Other variations of adaptor polynucleotide molecules are described throughout this application and can also be used in the methods.
[0236] The adaptor polynucleotide molecules can have the same sample tag. The adaptor polynucleotide molecules can be coupled to both ends of the template polynucleotide molecule. The method can include coupling the adaptor polynucleotide molecules to the template polynucleotide molecules with an efficiency of at least 30%, e.g., at least 50% (e.g., 60%, 70%, 80%, 90%, 95%, or 99%), thereby tagging each template polynucleotide with a tagging combination from 4 to 1,000,000 different tagging combinations to produce tagged polynucleotide molecules. In some cases, the reaction can occur in a single reaction vessel. The coupling efficiency can also be at least 30%, 35%, 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%. The tagging can be non-specific tagging.
[0237] The tagged polynucleotide molecules can then be subjected to an amplification reaction under conditions that produce amplified polynucleotide molecules as amplification products of the tagged polynucleotide molecules. The template polynucleotide molecule can be double-stranded. Furthermore, the template polynucleotide molecule can be blunt-ended. In some cases, the amplification reaction includes non-specifically amplifying the tagged polynucleotide molecules. The amplification reaction can also include using a priming site to amplify each of the tagged polynucleotide molecules. The priming site can be a primer, for example, a universal primer. The priming site can also be a nick.
[0238] The method can also include sequencing the amplified polynucleotide molecules. The sequencing step can include (i) subjecting the amplified polynucleotide molecules to an additional amplification reaction under conditions that generate additional amplified polynucleotide molecules as amplification products of the amplified polynucleotide molecules, and / or (ii) sequencing the additional amplified polynucleotide molecules. The additional amplification can be performed in the presence of primers containing flow cell sequences that produce polynucleotide molecules that can bind to a flow cell. The additional amplification can also be performed in the presence of primers containing sequences for hairpin-shaped adapters. Hairpin-shaped adapters can be attached to both ends of the polynucleotide fragments to generate circular molecules that can be sequenced multiple times. The method can further include identifying genetic variants during sequencing of the amplified polynucleotide molecules.
[0239] The method can further include separating polynucleotide molecules containing one or more given sequences from the amplified polynucleotide molecules to produce enriched polynucleotide molecules. The method can also include amplifying the enriched polynucleotide molecules with primers containing flow cell sequences. This amplification with primers containing flow cell sequences will produce polynucleotide molecules that can bind to a flow cell. Amplification can also be performed in the presence of primers containing sequences for hairpin adapters. Hairpin adapters can be attached to both ends of polynucleotide fragments to generate circular molecules that can be sequenced multiple times.
[0240] Flow cell sequences or hairpin adaptors can be added by non-amplification methods, such as ligation of such sequences. Other techniques, such as hybridization methods, for example, nucleotide overhangs, can be used.
[0241] The method can be performed without aliquoting the tagged polynucleotide molecules, for example, once the tagged polynucleotide molecules are generated, amplification and sequencing can occur in the same tube without further preparation.
[0242] The methods described herein can be useful in detecting single nucleotide variations (SNVs), copy number variations (CNVs), insertions, deletions and / or rearrangements. In some cases, SNVs, CNVs, insertions, deletions and / or rearrangements can be associated with diseases, such as cancer.
[0243] N. Monitoring the patient's condition
[0244] The method disclosed herein can also be used to monitor the disease state of patients.The disease of the subject can be monitored over time to determine the progress (for example, regression) of the disease.The marker that indicates the disease can be monitored in the biological sample of the subject, such as cell-free DNA sample.
[0245] For example, monitoring a subject's cancer status can include (a) determining the amount of one or more SNVs or copy number of multiple genes (e.g., in exons), (b) repeating such determination at different time points, and (c) determining whether there is a difference in the number of SNVs, level of SNVs, number or level of genomic rearrangements, or copy number between (a) and (b). Genes include ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, and the like. , MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. The gene can be selected from any 5, 10, 15, 20, 30, 40, 50, or all of the genes in this group.
[0246] O. Sensitivity and Specificity
[0247] The methods disclosed herein can be used to detect cancer polynucleotides in a sample and cancer in a subject with a high degree of concordance, for example, high sensitivity and / or specificity. For example, such methods can detect cancer polynucleotides (e.g., rare DNA) in a sample at a concentration of less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% with a specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. Such polynucleotides can indicate cancer or other diseases. Furthermore, such methods can detect cancer polynucleotides in a sample with a positive predictive value of at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, 99.99%, 99.999%, or 99.9999%.
[0248] Subjects who are actually positive and identified as positive in the test are called true positives (TP). Subjects who are actually negative and identified as positive in the test are called false positives (FP). Subjects who are actually negative and identified as negative in the test are called true negatives (TN). Subjects who are actually positive and identified as negative in the test are called false negatives (FN). Sensitivity is the percentage of actual positives that are identified as positive in the test. This includes, for example, cases where a cancer genetic variant should be found and was found (sensitivity = TP / (TP + FN)). Specificity is the percentage of actual negatives that are identified as negative in the test. This includes, for example, cases where a cancer genetic variant should not be found and was not found. Specificity can be calculated using the following equation: specificity = TN / (TN + FP). Positive predictive value (PPV) can be measured by the percentage of subjects who test positive and are true positives. PPV can be calculated using the following equation: PPV = TP / (TP + FP). Positive predictive value can be increased by increasing sensitivity (e.g., the probability of detecting an actual positive) and / or specificity (e.g., the probability of not mistaking an actual negative for a positive).
[0249] A low conversion rate from polynucleotides to adaptor-tagged polynucleotides can reduce the probability of converting and thus detecting rare polynucleotide targets, thereby reducing sensitivity. Noise in the test can reduce specificity by increasing the number of false positives detected in the test. Both low conversion rates and noise reduce the percentage of true positives and increase the percentage of false positives, thereby reducing positive predictive value.
[0250] The methods disclosed herein can achieve high levels of concordance, e.g., sensitivity and specificity, resulting in high positive predictive value. Methods for increasing sensitivity include highly efficient conversion of polynucleotides in a sample to adaptor-tagged polynucleotides. Methods for increasing specificity include, for example, reducing sequencing errors by molecular tracking.
[0251] The method of the present disclosure can be used to detect genetic variations (e.g., rare DNA) in non-uniquely tagged initial starting genetic material at a concentration of less than 5%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% with a specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. In some embodiments, the method can further include converting polynucleotides in the initial starting material with an efficiency of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. The sequence reads of the tagged polynucleotides can then be tracked to generate a consensus sequence of the polynucleotide with an error rate of 2%, 1%, 0.1%, or 0.01% or less.
[0252] 2. Pooling Method
[0253] Disclosed herein is a method for detecting copy number variations and / or sequence variants at one or more loci in a test sample. One embodiment is shown in Figure 8. Typically, detecting copy number variations involves determining a quantitative measure (e.g., absolute or relative number) of polynucleotides that map to a locus of interest in the genome of the test sample, and comparing this number to a quantitative measure of polynucleotides that map to the locus in a control sample. In certain methods, the quantitative measure is determined by comparing the number of molecules in the test sample that map to the locus of interest with the number of molecules in the test sample that map to a reference sequence, for example, a sequence that is expected to exist at wild-type ploidy. In some examples, the reference sequence is HG19, build 37, or build 38. The comparison can involve, for example, determining a ratio. This measure is then compared to a similar measure determined in a control sample. Thus, for example, if a test sample has a ratio of 1.5:1 for a locus of interest to a reference locus and a control sample has a ratio of 1:1 for the same locus, it can be concluded that the test sample exhibits polyploidy at the locus of interest.
[0254] If the test and control samples are analyzed separately, the workflow may introduce distortions between the final counts in the control and test samples.
[0255] In one method disclosed herein (e.g., flowchart 800), polynucleotides are provided from test and control samples (802). The polynucleotides in the test sample and the polynucleotides in the control sample are tagged with a tag (source tag) that identifies the polynucleotide as originating from the test or control sample (804). The tag can be, for example, a polynucleotide sequence or a barcode that unambiguously identifies the source.
[0256] The polynucleotides in each of the control and test samples can also be tagged with an identifier tag that is carried by every amplification progeny of the polynucleotide.The information from the start and end sequences of the polynucleotide and the identifier tag can identify the sequence read from the polynucleotide amplified from the original parent molecule.Each molecule can be tagged uniquely compared to other molecules in the sample.Alternatively, each molecule does not need to be tagged uniquely compared to other molecules in the sample.In other words, the number of different identifier sequences can be less than the number of molecules in the sample.By combining identifier information with start / stop sequence information, the probability of confusing two molecules with the same start / stop sequence is significantly reduced.
[0257] The number of different identifiers used for tagging nucleic acid (for example, cfDNA) can depend on the number of different haploid genome equivalents.Different identifiers can be used to tag at least 2, at least 10, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, at least 10,000 or more different haploid genome equivalents. Thus, the number of different identifiers used to tag 500-10,000 different haploid genome equivalent nucleic acid samples, e.g., cell-free DNA, can be between 1, 2, 3, 4, and 5, and any of the following: 100, 90, 80, 70, 60, 50, 40, or 30. For example, the number of different identifiers used to tag 500-10,000 different haploid genome equivalent nucleic acid samples can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 6, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or less.
[0258] Polynucleotides can be tagged by ligating an adaptor containing a tag or identifier before amplification. Ligation can be performed using an enzyme, such as a ligase. For example, tagging can be performed using a DNA ligase. The DNA ligase can be T4 DNA ligase, E. coli DNA ligase, and / or mammalian ligase. The mammalian ligase can be DNA ligase I, DNA ligase III, or DNA ligase IV. The ligase can also be a thermostable ligase. The tag can be ligated to a blunt end of the polynucleotide (blunt-end ligation). Alternatively, the tag can be ligated to a sticky end of the polynucleotide (sticky-end ligation). Polynucleotides can be tagged by blunt-end ligation using an adaptor (e.g., an adaptor with a forked end). Highly efficient ligation can be achieved using a large excess of adapters (e.g., greater than 1.5x, greater than 2x, greater than 3x, greater than 4x, greater than 5x, greater than 6x, greater than 7x, greater than 8x, greater than 9x, greater than 10x, greater than 11x, greater than 12x, greater than 13x, greater than 14x, greater than 15x, greater than 20x, greater than 25x, greater than 30x, greater than 35x, greater than 40x, greater than 45x, greater than 50x, greater than 55x, greater than 60x, greater than 65x, greater than 70x, greater than 75x, greater than 80x, greater than 85x, greater than 90x, greater than 95x, or greater than 100).
[0259] Once tagged with a tag that identifies the source of the polynucleotide, polynucleotides from different sources (e.g., different samples) can be pooled. After pooling, polynucleotides from different sources (e.g., different samples) can be distinguished by any measurement using the tag, including any process of quantitative measurement. For example, as shown in (806) (FIG. 8), polynucleotides from a control sample and a test sample can be pooled. The pooled molecules can be subjected to sequencing (808) and bioinformatics workflow. Both are subjected to the same variations in the process, thus reducing any differential bias. Because molecules originating from the control and test samples are differentially tagged, they can be distinguished in any process of quantitative measurement.
[0260] The relative amounts of the pooled control and test samples can vary. The amount of the control sample can be the same as the amount of the test sample. The amount of the control sample can also be greater than the amount of the test sample. Alternatively, the amount of the control sample can be less than the amount of the test sample. The smaller the relative amount of one sample relative to the total, the fewer identifying tags are required in the actual tagging process. The value can be selected to reduce to an acceptable level the probability that two parent molecules with the same start / end sequence will have the same identifying tag. This probability can be less than 10%, less than 1%, less than 0.1%, or less than 0.01%. The probability can be less than 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.
[0261] The methods disclosed herein can also include a step of grouping sequence reads. For example, a bioinformatics workflow can include grouping sequence reads produced from the progeny of a single parent molecule, as shown in (810) (FIG. 8). This can involve any of the redundancy reduction methods described herein. Molecules sourced from test and control samples can be distinguished based on the source tags they carry (812). Molecules mapping to target loci are quantified for both the test and control source molecules (812). This can include, for example, the normalization methods described herein, in which counts at target loci are normalized to counts at reference loci.
[0262] The normalized (or raw) abundances at the target locus from the test and control samples are compared to determine the presence of copy number variations (814).
[0263] 3. Computer Control System
[0264] The present disclosure provides a computer control system programmed to implement the methods of the present disclosure. Figure 6 shows a computer system 1501 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 1501 can regulate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 1501 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing. The computer system 1501 can be a user's electronic device or a computer system located remotely from the electronic device. The electronic device can be a mobile electronic device.
[0265] The computer system 1501 includes a central processing unit (CPU, collectively referred to herein as "processor" and "computer processor") 1505, which can be a single-core or multi-core processor or multiple processors for parallel processing. The computer system 1501 also includes memory or memory locations 1510 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 1515 (e.g., a hard disk), a communication interface 1520 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 1525, such as cache, other memory, data storage, and / or electronic display adapters. The memory 1510, storage unit 1515, interface 1520, and peripheral devices 1525 communicate with the CPU 1505 via a communication bus (solid lines), such as a motherboard. The storage unit 1515 can be a data storage unit (or data repository) for storing data. Computer system 1501 can be operatively coupled to a computer network ("network") 1530 with the aid of communication interface 1520. Network 1530 can be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. Network 1530, in some cases, is a telecommunications and / or data network. Network 1530 can include one or more computer servers, which can enable distributed computing, such as cloud computing. Network 1530 can, in some cases, implement a peer-to-peer network, which can enable devices coupled to computer system 1501 to act as clients or servers, with the aid of computer system 1501.
[0266] The CPU 1505 can execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 1510. The instructions may be directed to the CPU 1505, which can then program or otherwise configure the CPU 1505 to perform the methods of the present disclosure. Examples of operations performed by the CPU 1505 may include fetch, decode, execute, and writeback.
[0267] The CPU 1505 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1501 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0268] The storage unit 1515 can store files such as drivers, libraries, and saved programs. The storage unit 1515 can store user data, such as user preferences and user programs. The computer system 1501 can, in some cases, include one or more additional data storage units that are external to the computer system 1501, such as located on a remote server in communication with the computer system 1501 via an intranet or the Internet.
[0269] Computer system 1501 can communicate with one or more remote computer systems via network 1530. For example, computer system 1501 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate, or a tablet PC (e.g., an Apple (registered trademark)). The computer system 1501 may be a personal digital assistant (e.g., a Samsung® Galaxy Tab, a Samsung® iPad®, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user may access the computer system 1501 via a network 1530.
[0270] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored in electronic storage locations of computer system 1501, such as memory 1510 or electronic storage unit 1515. The machine-executable or machine-readable code may be provided in the form of software. In use, the code may be executed by processor 1505. In some cases, the code may be retrieved from storage unit 1515 and stored in memory 1510 for immediate access by processor 1505. In some circumstances, electronic storage unit 1515 may be precluded, and machine-executable instructions may be stored in memory 1510.
[0271] The code may be pre-compiled and configured for use by a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be executed in a pre-compiled or as-compiled manner. It may be supplied in a programming language that may be selected to allow the execution of the code.
[0272] Aspects of the systems and methods provided herein, such as computer system 1501, can be embodied in programming. Various aspects of the technology can be considered to be a “product” or “article of manufacture,” typically in the form of machine (or processor) executable code and / or associated data carried or embodied in some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of a computer, processor, or other, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage of software programming at any time. The software, in whole or in part, can sometimes be communicated via the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include optical, electrical, and electromagnetic waves used through physical interfaces between local devices, such as through wired and optical landline networks, and over various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, and the like, may also be considered media bearing software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0273] Thus, machine-readable media, such as computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or otherwise, such as may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated in radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example: a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punched card paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0274] The computer system 1501 can include or be in communication with an electronic display 1535 that includes a user interface (UI) 1540. The UI allows a user to set various conditions, e.g., PCR or sequencing conditions, for the methods described herein. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0275] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software through execution by the central processing unit 1505. The algorithms can, for example, process reads to generate a resulting sequence.
[0276] 7 illustrates a schematic of another system for analyzing a sample containing nucleic acids from a subject. The system includes a sequencer, bioinformatics software, and an internet connection for reporting and analysis by, for example, a handheld device or a desktop computer.
[0277] 1. A system for analyzing a target nucleic acid molecule of a subject, the system comprising: a communication interface that receives nucleic acid sequence reads for a plurality of polynucleotide molecules covering genomic loci of a target genome; a computer memory that stores the nucleic acid sequence reads for the plurality of polynucleotide molecules received by the communication interface; and a computer processor operatively coupled to the communication interface and the memory and programmed to: (i) group the plurality of sequence reads into families, each family including sequence reads derived from one of the template polynucleotides; (ii) for each of the families, merge the sequence reads to generate a consensus sequence; (iii) call the consensus sequence at a given genomic locus among the genomic loci; and (iv) detect, at the given genomic locus, any of genetic variants among the calls, frequency of genetic alterations among the calls, total number of calls, and total number of alterations among the calls, wherein the genomic loci are selected from the group consisting of ALK, APC, BRAF, and BRAF. , CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB 4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, Disclosed herein is a system corresponding to a plurality of genes selected from the group consisting of PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. Different variations of each component of this system are described throughout the disclosure of methods and compositions. These individual components and variations thereof can also be applied in this system.
[0278] 4. Kit
[0279] Kits containing the compositions described herein may be useful in practicing the methods described herein.
[0014] ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC , PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, and NTRK1. Disclosed herein is a kit comprising a plurality of oligonucleotide probes that selectively hybridize to genes. The number of genes to which the oligonucleotide probes can selectively hybridize can vary. For example, the number of genes can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, or 54. The kit can include a container containing a plurality of oligonucleotide probes and instructions for carrying out any of the methods described herein.
[0280] The oligonucleotide probe can selectively hybridize to the exon region of a gene, for example, at least five genes. In some cases, the oligonucleotide probe can selectively hybridize to at least 30 exons of a gene, for example, at least five genes. In some cases, multiple probes can selectively hybridize to each of the at least 30 exons. The probe hybridizing to each exon can have a sequence that overlaps with at least one other probe. In some embodiments, the oligoprobe can selectively hybridize to the non-coding region of a gene disclosed herein, for example, the intron region of a gene. The oligoprobe can also selectively hybridize to the region of a gene, including both the exon and intron regions of the gene disclosed herein.
[0281] Any number of exons can be targeted by the oligonucleotide probes, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160 , 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1,000 or more exons can be targeted.
[0282] The kit can include at least 4, 5, 6, 7, or 8 different library adaptors having distinct molecular barcodes and identical sample barcodes. The library adaptors do not have to be sequencing adaptors. For example, the library adaptors do not include flow cell sequences or sequences that enable the formation of hairpin loops for sequencing. Different variations and combinations of molecular barcodes and sample barcodes are described throughout this application and are applicable to the kit. Furthermore, in some cases, the adaptors are not sequencing adaptors. Moreover, the adaptors provided by the kit can also include sequencing adaptors. The sequencing adaptors can include sequences that hybridize to one or more sequencing primers. The sequencing adaptors can further include sequences that hybridize to a solid support, e.g., flow cell sequences. For example, the sequencing adaptors can be flow cell adaptors. The sequencing adaptors can be attached to one or both ends of polynucleotide fragments. In some cases, the kit can include at least 8 different library adaptors having distinct molecular barcodes and identical sample barcodes. The library adaptors do not have to be sequencing adaptors. The kit can further include a sequencing adaptor having a first sequence that selectively hybridizes to the library adaptor and a second sequence that selectively hybridizes to the flow cell sequence. In another example, the sequencing adaptor can be hairpin-shaped. For example, the hairpin-shaped adaptor can include a complementary double-stranded portion and a loop portion, and the double-stranded portion can be attached (e.g., ligated) to a double-stranded polynucleotide. The hairpin-shaped sequencing adaptor can be attached to both ends of a polynucleotide fragment to generate a circular molecule that can be sequenced multiple times.Sequencing adapters are arranged end-to-end up to 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, The sequencing adaptor can be 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more bases. The sequencing adaptor can comprise 20-30, 20-40, 30-50, 30-60, 40-60, 40-70, 50-60, 50-70 bases from end to end. In particular examples, the sequencing adaptor can comprise 20-30 bases from end to end. In another example, the sequencing adaptor can comprise 50-60 bases from end to end. The sequencing adaptor can comprise one or more barcodes. For example, the sequencing adaptor can comprise a sample barcode. The sample barcode can comprise a predetermined sequence. The sample barcode can be used to identify the source of the polynucleotide. The sample barcode can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more nucleic acid bases (or any length described throughout this application), e.g., at least 8 bases. The barcode can be a contiguous or non-contiguous sequence, as described above.
[0283] The library adaptors can be blunt-ended and Y-shaped and can be less than or equal to 40 nucleobases in length. Other variations can be found throughout this application and are applicable to the kits. [Example]
[0284] Example 1 Methods for copy number variation detection
[0285] Blood collection
[0286] Collect 10-30 mL blood samples at room temperature. Centrifuge the samples to remove cells. Collect plasma after centrifugation.
[0287] cfDNA extraction
[0288] The sample is subjected to proteinase K digestion. The DNA is precipitated with isopropanol. The DNA is captured on a DNA purification column (e.g., QIAamp DNA Blood Mini Kit) and eluted in 100 μl of solution. DNA smaller than 500 bp is selected by Ampure SPRI magnetic bead capture (PEG / salt). The resulting product is suspended in 30 μl of H2O. The size distribution is checked (major peak = 166 nucleotides; minor peak = 330 nucleotides) and quantified. 5 ng of extracted DNA contains approximately 1700 haploid genome equivalents ("HGE"). The general correlation between the amount of DNA and HGE is as follows: 3 pg DNA = 1 HGE; 3 ng DNA = 1K HGE; 3 μg DNA = 1M HGE; 10 pg DNA = 3 HE; 10 ng DNA = 3K HGE; 10 μg DNA = 3M HGE.
[0289] "Single Molecule" Library Prep
[0290] Highly efficient DNA synthesis was achieved by blunt-end repair and ligation with eight different octomers (i.e., 64 combinations) with overloaded hairpin adapters. A tagging (>80%) is performed. 2.5 ng DNA (i.e., approximately 800 HGE) is used as starting material. Each hairpin adapter contains a random sequence in its non-complementary portion. Hairpin adapters are attached to both ends of each DNA fragment. Each tagged fragment can be identified by the random sequence in the hairpin adapter and the 10p endogenous sequence in the fragment.
[0291] Amplified the tagged DNA by 10 cycles of PCR to produce approximately 1-7 µg DNA containing roughly 500 copies of each of the 800 HGEs in the starting material.
[0292] PCR reaction can be optimized by buffer optimization, polymerase optimization and cycle reduction. Amplification bias, such as non-specific bias, GC bias and / or size bias, can also be reduced by optimization. Noise(s) (e.g., polymerase-introduced error) can be reduced by using high-fidelity polymerase.
[0293] Libraries can be prepared using Verniata or Sequenom methods.
[0294] Sequences can be enriched as follows: DNA containing a region of interest (ROI) is captured using biotin-labeled beads with a probe for the ROI. The ROI is amplified by 12 cycles of PCR to generate a 2000-fold amplification. The resulting DNA is then denatured, diluted to 8 pM, and loaded onto an Illumina sequencer.
[0295] Massively parallel sequencing
[0296] Use 0.1-1% of the sample (approximately 100 pg) for sequencing.
[0297] Digital Bioinformatics
[0298] Sequence reads are grouped into families, with each family having approximately 10 sequence reads. Families are collapsed into consensus sequences by voting (e.g., biased voting) at each position in the family. If 8 or 9 members match, the base is called against the consensus sequence. If 60% or less of the members match, the base is not called against the consensus sequence.
[0299] The resulting consensus sequence is mapped to the reference genome. Each base in the consensus sequence is covered by approximately 3,000 different families. A quality score is calculated for each sequence, and sequences are filtered based on their quality scores.
[0300] Sequence variations are detected by counting the distribution of bases at each locus: if 98% of the reads have the same base (homozygous) and 2% have a different base, the locus likely harbors a sequence variant derived from cancer DNA.
[0301] CNVs are detected by counting the total number of sequences (bases) mapping to the locus and comparing them with the reference locus. To increase CNV detection, the following genes are used: ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL Perform CNV analysis in specific regions, including regions in the , NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA, or NTRK1 genes.
[0302] Example 2 Method for correcting base calling by determining the total number of unseen molecules in a sample
[0303] After amplifying the fragments and reading and aligning the sequences of the amplified fragments, the fragments are subjected to base calling. Variations in the number of amplified fragments and unobserved amplified fragments may introduce errors into base calling. Such variations are corrected by calculating the number of unobserved amplified fragments.
[0304] For base calling of locus A (any locus), it is first assumed that there are N amplified fragments. The sequence readout can come from two types of fragments: double-stranded fragments and single-stranded fragments. Below is a theoretical example of calculating the total number of unseen molecules in a sample.
[0305] N is the total number of molecules in the sample. Assume 1000 is the number of duplexes detected. Assume 500 is the number of single-stranded molecules detected. P is the probability of observing a chain. Q is the probability of not detecting a strand.
[0306] Since Q=1-P, 1000=NP(2) 500=N2PQ 1000 / P(2)=N 500÷2PQ=N 1000 / P(2)=500÷2PQ 1000*2PQ=500P(2) 2000PQ=500P(2) 2000Q=500P 2000(1-P)=500P 2000-2000P=500P 2000=500P+2000P 2000=2500P 2000÷2500=P 0.8=P 1000 / P(2)=N 1000÷0.64=N 1562=N Number of unseen fragments = 62.
[0307] Example 3 Identification of genetic variants in cancer-associated somatic variants in patients
[0308] The assay is used to analyze a panel of genes to identify genetic variants in cancer-associated somatic variants with high sensitivity.
[0309] Cell-free DNA was extracted from patient plasma and amplified by PCR. Genetic variants were analyzed by massively parallel sequencing of the amplified target genes. For one set of genes, all exons were sequenced because such sequencing coverage has been shown to have clinical utility (Table 1). For another set of genes, sequencing coverage included exons with previously reported somatic mutations (Table 2). The minimum detectable mutant allele (detection limit) depended on the cell-free DNA concentration of the patient sample, which varied from less than 10 to more than 1,000 genome equivalents per mL of peripheral blood. With smaller amounts of cell-free DNA and / or low-level gene copy amplification, amplification may not be detected in the sample. Certain sample or variant characteristics, such as poor sample quality or inadequate collection, resulted in reduced analytical sensitivity.
[0310] The percentage of genetic variants found in circulating cell-free DNA in the blood is related to the specific tumor biology of this patient. Factors that influenced the amount / percentage of genetic variants detected in circulating cell-free DNA in the blood include tumor growth, turnover, size, heterogeneity, angiogenesis, disease progression, or treatment. Table 3 annotates the percentage or allele frequency (%cfDNA) of altered circulating cell-free DNA detected in this patient. Some of the detected genetic variants are listed in descending order by %cfDNA.
[0311] Genetic variants are detected in circulating cell-free DNA isolated from this patient's blood sample. These genetic variants are cancer-associated somatic variants, some of which have been associated with either increased or decreased clinical response to specific treatments. "Minor alterations" are defined as alterations detected at less than 10% of the allele frequency of "major alterations." The detected allele frequencies of these alterations (Table 3) and the associated treatment for this patient are annotated.
[0312] All genes listed in Tables 1 and 2 are analyzed as part of the Guardant360™ test. No amplification of ERBB2, EGFR, or MET is detected in circulating cell-free DNA isolated from this patient's blood specimen.
[0313] Patient test results, including genetic variants, are listed in Table 4. [Table 1] [Table 2] [Table 3] [Table 4]
[0314] Example 4 Determination of patient-specific detection limits for genes analyzed by the Guardant360™ assay
[0315] The method of Example 3 is used to detect genetic alterations in the cell-free DNA of a patient. The sequence reads of these genes include exon and / or intron sequences.
[0316] The detection limits of the test are shown in Table 5. The detection limit value depends on the cell-free DNA concentration and the sequencing coverage per gene. [Table 5]
[0317] Example 5 Correction of sequence errors by comparing Watson and Crick sequences
[0318] Double-stranded cell-free DNA is isolated from patient plasma.Cell-free DNA fragments are tagged with 16 different bubble-containing adaptors, each of which contains a unique barcode.By ligation, bubble-containing adaptors are attached to both ends of each cell-free DNA fragment.After ligation, each cell-free DNA fragment can be identified separately by the sequence of the distinct barcode and the two 20bp endogenous sequences at each end of the cell-free DNA fragment.
[0319] The tagged cell-free DNA fragments are amplified by PCR. The amplified fragments are enriched using beads containing oligonucleotide probes that specifically bind to a group of cancer-related genes. Therefore, the cell-free DNA fragments derived from a group of cancer-related genes are selectively enriched.
[0320] Sequencing primer binding sites, sample barcodes, and cell-flow sequences Each containing a sequencing adaptor is attached to the enriched DNA molecules, and the resulting molecules are amplified by PCR.
[0321] Both strands of the amplified fragments are sequenced. Because each bubble-containing adaptor contains a non-complementary portion (e.g., a bubble), the sequence of one strand of the bubble-containing adaptor is different from the sequence of the other strand (complement). Therefore, the sequence reads of amplicons derived from the Watson strand of the original cfDNA can be distinguished from amplicons derived from the Crick strand of the original cfDNA by the attached bubble-containing adaptor sequence.
[0322] The sequence reads from the strands of the original cell-free DNA fragments are compared with the sequence reads from the other strands of the original cell-free DNA fragments. If a variant occurs only in the sequence reads from one strand of the original cell-free DNA fragments but not in the other strand, this variant will be identified as an error (e.g., due to PCR and / or amplification) rather than a true genetic variant.
[0323] Group sequence reads into families. Correct errors in sequence reads. Generate a consensus sequence for each family by collapsing.
[0324] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited to the specific examples set forth herein. While the present invention has been described with reference to the above specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Accordingly, numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It will be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the invention. It is therefore contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of the claims and their equivalents be covered thereby.
Claims
[Claim 1] A composition as described in the specification.