Compositions and methods for identifying nucleic acid molecules
Patent Information
- Application Number
- JP2025114064
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-12-07
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-21
AI Technical Summary
Existing methods for next-generation sequencing face challenges in distinguishing errors introduced during sample preparation and base calling from actual SNPs or mutations, particularly in complex samples like mammalian cDNA or circulating DNA, due to the need for lengthy unique identifiers that reduce read length and increase costs.
The use of molecular beacon tags (MITs) to tag nucleic acid molecules, with specific ratios and diversities, allows for effective identification of amplification and base-calling errors by sequencing tagged nucleic acid molecules, and determining copy numbers of chromosomes or chromosomal segments.
This approach enables accurate differentiation of errors from actual sample differences, particularly in high-throughput sequencing, improving error identification and copy number determination in complex samples with reduced costs and increased read length.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application No. 15 / 372,279, filed December 7, 2016, which is incorporated herein by reference in its entirety.
[0002] Array List This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on November 14, 2017, is named N_018_WO_01_SL.txt and is 5,069 bytes in size.
[0003] FIELD OF THE INVENTION TECHNICAL FIELD The present disclosure relates generally to methods for analyzing nucleic acids. [Background technology]
[0004] Next-generation sequencing has significantly increased sequencing throughput and has led to new applications of sequencing with important practical implications, such as improved cancer diagnosis and non-invasive prenatal testing for disorders such as Down's syndrome. There are various techniques for performing next-generation sequencing, each associated with specific types of errors. Furthermore, these methods share common sources of error, such as errors that occur during sample preparation.
[0005] Sample preparation for next-generation sequencing typically involves multiple amplification steps, each of which introduces errors. Amplification reactions, such as PCR, used in sample preparation for high-throughput sequencing can include initial amplification of nucleic acids in the sample to generate the library to be sequenced, clonal amplification of the library, typically onto a solid support, and additional amplification reactions to add additional information or functionality, such as sample identification barcodes. Errors can be introduced during any of the amplification reactions, for example, by base misincorporation by the polymerase used for amplification. It can be difficult to distinguish these errors introduced during sample preparation and errors generated during the sequencing reaction from actual and informative SNPs or mutations present in the initial sample, especially when SNPs or mutations are present at low frequencies. Furthermore, base calling at each nucleotide can also introduce errors, usually caused by low signal intensity and / or the surrounding nucleic acid sequence.
[0006] There are several known methods for identifying errors caused by sample preparation. One method is to obtain a greater sequencing depth so that a sample nucleic acid segment is read multiple times from the same molecule or from different copies of the same nucleic acid molecule. These multiple reads can be aligned to generate a consensus sequence. However, low-frequency SNPs or mutations in a population of nucleic acid molecules will appear similar to errors introduced during amplification or base calling. Another method for identifying these errors involves tagging nucleic acid molecules so that each nucleic acid molecule incorporates a unique identifier before being sequenced. Sequencing results from identically tagged nucleic acid molecules are pooled, and the consensus sequence from these pooled results is likely to be the true sequence of the nucleic acid from the sample. If some of the identically tagged nucleic acid molecules have different sequences, amplification errors can be identified.
[0007] Despite these conventional methods, there remains a need to discover advantageous combinations of parameters for highly effective and easily manufacturable methods of tagging nucleic acid molecules, particularly for analyzing complex samples, including genomic samples such as mammalian cDNA or circulating DNA samples. Many conventional methods require the generation of a large number of unique identifiers, which can also lead to the need for longer unique identifiers. The reaction mixtures in such methods are designed so that there is a large excess of unique identifiers relative to the sample nucleic acid molecules. In addition to the high cost of creating such libraries of unique identifiers, increasing the length of the unique identifiers reduces the amount of sample nucleic acid sequence that can be read with the already limited read length of most next-generation sequencers. Other conventional disclosures, sometimes merely prophetic, do not provide detailed combinations of parameters for combinations such as identifier diversity relative to the copy number of the region of interest or the diversity of any two identifiers, identifier diversity relative to the total number of sample nucleic acid molecules, and total number of identifiers relative to the total number of sample nucleic acid molecules. This is particularly true for complex, naturally isolated samples, such as cDNA or genomic samples, including fragmented genomic samples, such as circulating free DNA in mammalian blood. Summary of the Invention [Problem to be solved by the invention]
[0008] There remains a need for low-cost tagging methods and the identification of key parameter combinations for tagging complex samples isolated from nature. Such methods would be beneficial, for example, for detecting amplification and base-calling errors when used in high-throughput sequencing workflows, particularly in the analysis of complex and clinically important samples.
[0009] (Summary of the Invention) The present disclosure provides improved methods and compositions for tagging nucleic acid molecules using molecular beacon tags ("MITs") to identify amplification products resulting from individual sample nucleic acids after amplification of a population of sample nucleic acid molecules. Additionally, methods are provided herein for using MITs to sequence sample nucleic acid molecules, identify errors made during sample preparation or base calling, and determine the copy number of chromosomes or chromosomal segments. Also provided herein are compositions comprising a reaction mixture of sample nucleic acid molecules and MITs, populations of tagged nucleic acid molecules, libraries of MITs, and kits for generating tagged nucleic acid molecules using MITs. Thus, the present disclosure provides methods and compositions for distinguishing errors introduced during sample preparation and base calling, particularly during high-throughput sequencing workflows, from actual differences present in nucleic acid molecules in the starting sample.
[0010] Thus, in one embodiment, provided herein is a method for sequencing a population of sample nucleic acid molecules, comprising the steps of: forming a reaction mixture containing a population of sample nucleic acid molecules and a set of molecular beacon tags (MITs), wherein the MITs are nucleic acid molecules, the number of different MITs in the set of MITs is 10 to 1,000, and the ratio of the total number of sample nucleic acid molecules in the population of sample nucleic acid molecules to the diversity of MITs in the set of MITs or the ratio of the diversity of any two types of MITs in the set of MITs is at least 500:1, 1,000:1, 10,000:1, or 100,000:1; attaching at least one MIT from the set of MITs to at least 50% of the sample nucleic acid segments of the sample nucleic acid molecules to generate a population of tagged nucleic acid molecules, wherein the at least one MIT is located 5' and / or 3' to the sample nucleic acid segment on each tagged nucleic acid molecule, and the population of tagged nucleic acid molecules comprises at least one copy of each MIT from the set of MITs; amplifying the population of tagged nucleic acid molecules to create a library of tagged nucleic acid molecules; and Determining the sequences of the bound MITs of the tagged nucleic acid molecules in the library of tagged nucleic acid molecules and at least a portion of the sample nucleic acid segments, thereby sequencing the population of sample nucleic acid molecules. The total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules in the reaction mixture.
[0011] In some embodiments, the method can include using the sequence of at least one MIT on each tagged nucleic acid molecule to identify the individual sample nucleic acid molecule that generated the tagged nucleic acid molecule.In some embodiments, the method can further include, before identifying the individual sample nucleic acid molecule, mapping the determined sequence of at least one sample nucleic acid segment to a location in the genome of the source from which the sample is derived, and using the mapped genome location together with the sequence of at least one MIT to identify the individual sample nucleic acid molecule that generated the tagged nucleic acid molecule.Furthermore, in such an embodiment, the mutation in the nucleic acid segment or in the allele of the nucleic acid segment can be identified.
[0012] In some embodiments, the sample may be a mammalian sample, such as a human sample, and the sample may be, for example, a blood sample. The diversity of any two MIT combinations in the set of MITs can exceed the total number of sample nucleic acid molecules spanning each target locus of the multiple target loci in the genome of the mammal from which the mammalian sample is sourced.
[0013] In some embodiments, the MIT can be bound during ligation. In some embodiments, tagged nucleic acid molecules can be enriched using hybrid capture. In some embodiments, enriched tagged nucleic acid molecules can be clonally amplified on a solid support or multiple solid supports before their sequences are determined using high-throughput sequencing.
[0014] In some embodiments, the method can include using a sample, at least some of the sample nucleic acids comprising at least one target locus from a plurality of target loci from a chromosome or chromosome segment of interest. In some embodiments, the method can further include using the identified sample nucleic acid molecules to measure the amount of DNA for each target locus by counting the number of sample nucleic acid molecules comprising each target locus, and determining, on a computer, the copy number of one or more chromosomes or chromosome segments of interest using the amount of DNA at each target locus in the sample nucleic acid molecules.
[0015] In some embodiments, the sample can contain circulating cell-free human DNA, including circulating tumor DNA, and the diversity of the combination of any two MITs in the set of MITs exceeds the total number of circulating cell-free DNA fragments or sample nucleic acid molecules spanning the target locus in the human genome.
[0016] In another aspect, a method is provided for identifying amplification errors from sample preparation for high throughput sequencing or for identifying base calling errors in a high throughput sequencing reaction of a population of tagged nucleic acid molecules derived from a sample, the method comprising the steps of: forming a reaction mixture comprising a population of sample nucleic acid molecules and a set of molecular beacon tags (MITs), wherein the MITs are double-stranded nucleic acid molecules, the number of different MITs in the set of MITs is 10 to 100, 250, 500, 1,000, 2,000, 2,500, or 5,000, and the ratio of the total number of sample nucleic acid molecules in the population of sample nucleic acid molecules to the diversity of MITs in the set of MITs is greater than 500:1, 1,000:1, 10,000:1, or 100,000:1; attaching at least one MIT from the set of MITs to a sample nucleic acid segment of at least one sample nucleic acid molecule of the population of sample nucleic acid molecules to generate a population of tagged nucleic acid molecules, wherein the at least one MIT is located 5' and / or 3' to the sample nucleic acid segment on each tagged nucleic acid molecule, and the population of tagged nucleic acid molecules comprises at least one copy of each MIT in the set of MITs; amplifying the population of tagged nucleic acid molecules to create a library of tagged nucleic acid molecules; determining the sequences of the linked MITs and at least a portion of the sample nucleic acid segments of tagged nucleic acid molecules in the library of tagged nucleic acid molecules using high-throughput sequencing, wherein the sequence of at least one MIT on each tagged nucleic acid molecule identifies the individual sample nucleic acid molecule that gave rise to the tagged nucleic acid molecule; and Identifying tagged nucleic acid molecules with amplification errors by identifying nucleic acid segments having a nucleotide sequence that is found in less than 25% of tagged nucleic acid molecules derived from the same initial sample nucleic acid molecule, wherein the total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules in the reaction mixture.
[0017] In some embodiments, the method can further comprise a sample having genomic DNA fragments that are more than 20 nucleotides and less than 1,000 nucleotides in length, or more than 50 nucleotides and less than 500 nucleotides in length, wherein the diversity of any two MIT combinations in the set of MITs exceeds the total number of DNA fragments or sample nucleic acid molecules that span the target locus in genome.In some embodiments, the method can be used for example with maternal blood samples, and the copy number determination is for non-invasive prenatal testing.In some embodiments, the method can be used with blood samples from individuals who are suffering from cancer or suspected of suffering from cancer.
[0018] In another aspect, provided herein is a method for determining the copy number of one or more chromosomes or chromosome segments of interest from a target individual in a sample of blood or a fraction thereof from the target individual or from the target individual's mother, the method comprising the steps of: generating a population of tagged nucleic acid molecules by reacting a population of sample nucleic acid molecules with a set of nucleic acid molecule indicator tags (MITs), wherein the number of different MITs in the set of MITs is between 10 and 10,000, or between 10 and 1,000, the ratio of the total number of sample nucleic acid molecules in the population of sample nucleic acid molecules to the diversity of MITs in the set of MITs is greater than 500:1, 1,000:1, 10,000:1, or 100,000:1, at least some of the sample nucleic acid molecules comprise one or more target loci among a plurality of target loci on a chromosome or chromosome segment of interest, and the sample is 1.0 ml or less of blood or a fraction of blood derived from 1.0 ml or less of blood; amplifying the enriched population of tagged nucleic acid molecules to generate a library of tagged nucleic acid molecules; determining the sequence of the bound MIT of a tagged nucleic acid molecule in the library of tagged nucleic acid molecules and the sequence of at least a portion of the sample nucleic acid segment to determine the identity of the sample nucleic acid molecule that gave rise to the tagged nucleic acid; Using the determined identities, measuring the amount of DNA for each target locus by counting the number of sample nucleic acid molecules containing each target locus; and and determining, on a computer, the copy number of one or more chromosomes or chromosome segments of interest using the amount of DNA at each target locus in the sample nucleic acid molecule. The total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules in the reaction mixture.
[0019] In some embodiments, the number of target loci and the volume of the sample provide an effective amount of all target loci to achieve the desired sensitivity and specificity for copy number determination. In some embodiments, the method may further include using the number of target loci and the total number of sample nucleic acid molecules spanning the target loci to provide an effective amount of total sequencing reads to achieve the desired sensitivity and specificity for copy number determination. In some embodiments, this may be at least 10, 25, 50, 100, 250, 500, 1,000, 1,500, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 1,0000, 15,000, 2,0000, 25,000, 30,000, 40,000, or 50,000 target loci. In some embodiments, the method can include at least 10,000, 100,000, 500,000, or 1,000,000 total target loci in the sample, wherein the set of MITs includes at least 25, 30, 32, 50, 64, 100, 200, 250, 500, or 1,000 MITs, wherein the sample is from a mother and contains at least 1%, 2%, 3%, 4%, or 5% fetal nucleic acid compared to maternal nucleic acid, the desired specificity is 95%, 96%, 97%, 98%, or 99%, and the desired sensitivity is 95%, 96%, 97%, 98%, or 99%.
[0020] In some embodiments, the methods can include a ligation reaction to generate a population of tagged nucleic acid molecules, where the population of tagged nucleic acid molecules is enriched using hybrid capture prior to amplification, and the number of total target loci in the sample is at least 4, 5, 6, 7, 8, 9, 10, 15, or 20 times greater than the number of total target loci necessary to meet the desired specificity and the desired sensitivity.
[0021] In some embodiments, the method can further comprise using the amount of DNA at each target locus to determine the probability of each copy number hypothesis from a set of copy number hypotheses for one or more chromosomes or chromosome segments of interest, and selecting the copy number hypothesis with the highest probability.
[0022] In some embodiments, the method may include determining the probability of each copy number hypothesis by comparing the amount of DNA at a plurality of target loci with the amount of DNA at the disomic loci using a plurality of disomic loci from one or more chromosomes or chromosomal segments predicted to be disomic on the sample nucleic acid molecule.
[0023] In some embodiments, the method can be used on a maternal blood sample where copy number determination is for non-invasive prenatal testing. In some embodiments, the method can be used on a blood sample from an individual suffering from or suspected of suffering from cancer.
[0024] Another embodiment provided herein is a reaction mixture comprising: a population of at least 100,000, 200,000, 250,000, 500,000, or 1,000,000 sample nucleic acid molecules, between 10, 20, 25, 50, or 100 and 200, 250, 500, 1,000, 2,000, or 2,500 nucleotides in length; between 10 and 100, 200, 250, 500, 1,000, or 10,000, between 3, 4, 5, 6, or 7 nucleotides in length at the lower end of the range and 8, 9, 10, 11, 12, 15, or 20 nucleotides in length at the upper end of the range. a set of molecular indicator tags (MITs); and a ligase, wherein the MITs are nucleic acid molecules separated from the sample nucleic acid molecules, the total number of MIT molecules in the reaction mixture is greater than the total number of sample nucleic acid molecules in the reaction mixture, the ratio of the total number of sample nucleic acid molecules in the reaction mixture to the diversity of MITs in the set of MITs in the reaction mixture is at least 1,000:1, 10,000:1, or 100,000:1, the sequence of each MIT in the set of MITs differs by at least two nucleotides from all other MIT sequences in the set, and the reaction mixture contains at least two copies of each MIT.
[0025] In another aspect, the present disclosure provides a method for determining the copy number of one or more chromosomes or chromosome segments of interest in a sample of blood or a fraction thereof from a target individual, the method comprising the steps of: generating a reaction mixture comprising a population of sample nucleic acid molecules derived from a sample and a set of at least 32 molecular beacon tags (MITs), wherein each MIT in the set of MITs is a double-stranded nucleic acid molecule comprising a different nucleic acid sequence, the sample being derived from 1.0 ml or less of blood, the ratio of the total number of sample nucleic acid molecules in the population of sample nucleic acid molecules to the diversity of MITs in the set of MITs being greater than 1,000:1, and at least some of the sample nucleic acid molecules comprise one or more target loci of at least 1,000 target loci on a chromosome or chromosome segment of interest; attaching at least two MITs from the set of MITs to a sample nucleic acid segment of each sample nucleic acid molecule of the population of sample nucleic acid molecules to generate a population of tagged nucleic acid molecules, wherein each of the at least two MITs is located 5' and / or 3' to the sample nucleic acid segment on each tagged nucleic acid molecule, and the population of tagged nucleic acid molecules comprises at least one copy of each MIT of the set of MITs; amplifying the population of tagged nucleic acid molecules to create a library of tagged nucleic acid molecules; determining the sequence of the bound MIT of tagged nucleic acid molecules in the library of tagged nucleic acid molecules and the sequence of at least a portion of the sample nucleic acid segment, wherein the sequence of the bound MIT on each tagged nucleic acid molecule and the sequence of at least a portion of the nucleic acid segment are used to identify tagged nucleic acid molecules belonging to the same paired MIT nucleic acid segment family, wherein at least two MITs on each member of the paired MIT nucleic acid segment family are identical or complementary, the nucleic acid molecule segment of each member of the MIT nucleic acid segment family is mapped to the same coordinates on the genome of the source of the population of sample nucleic acid molecules, and at least 25% of the sample nucleic acid molecules are represented in the library of tagged nucleic acid molecules whose sequences are determined; Determining the DNA content of each target locus for the sample nucleic acid molecule by counting the number of MIT nucleic acid segment families that span each target locus; and The process of determining the copy number of one or more chromosomes or chromosome segments of interest on a computer using the amount of DNA at each target locus in the sample nucleic acid molecule. The total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules in the reaction mixture. MIT nucleic acid segment families share identical MITs at the same relative position relative to the nucleic acid segment, the same fragment end position, and the same sequence orientation (positive or negative relative to the human genome). Each sample nucleic acid molecule entering the MIT library preparation process can generate two families (one can be mapped to each of the positive and negative orientations). When an MIT nucleic acid segment family contains complementary MITs at the same relative position relative to the same nucleic acid segment and complementary fragment end position, the two MIT nucleic acid segment families can be paired, one with a positive orientation and the other with a negative orientation. In some embodiments, paired MIT nucleic acid segment families can be used to confirm the presence of sequence differences in the sample nucleic acid molecules.
[0026] In some embodiments, the method can further comprise analyzing single nucleotide polymorphism loci for one or more target loci on one or more chromosomes or chromosome segments. In a further embodiment, before determining the copy number of one or more chromosomes or chromosome segments of interest, the proportion of sample nucleic acid molecules containing different alleles at each locus can be estimated by counting the number of MIT nucleic acid segment families containing each allele at each locus, and the estimated proportion of sample nucleic acid molecules containing different alleles at each locus can be used to determine the copy number of one or more chromosomes or chromosome segments of interest.
[0027] In some embodiments, the method can include a sample of circulating cell-free human DNA, wherein the diversity of possible combinations of any two MITs in the set of MITs exceeds the number of circulating cell-free DNA fragments or sample nucleic acid molecules in the reaction mixture that span one or more target loci in the human genome.
[0028] In some embodiments, the method can further include analyzing multiple disomic loci on a chromosome or chromosome segment predicted to be disomic, wherein the method further includes determining the DNA dosage for each disomic locus for the sample nucleic acid molecule by counting the number of MIT nucleic acid segment families spanning each disomic locus, and determining the copy number of one or more chromosomes or chromosome segments of interest uses the DNA dosage for each target locus and the DNA dosage for each disomic locus.
[0029] In some embodiments, the method may further include generating, on a computer, a plurality of ploidy hypotheses each associated with a different possible ploidy state of the chromosome or chromosome segment of interest, and identifying the copy number for the individual by determining, on a computer, the relative probability of each ploidy hypothesis using the DNA amount for each target locus and selecting the ploidy state corresponding to the hypothesis with the greatest probability.
[0030] In some embodiments, the methods can be used on maternal samples where copy number determination is for non-invasive prenatal testing. In some embodiments, the methods can be used on samples from individuals with or suspected of having cancer.
[0031] In another aspect, provided herein is a method for determining the copy number of one or more chromosomes or chromosome segments in a sample of blood or a fraction thereof from a target individual, the method comprising the steps of: generating a population of tagged nucleic acid molecules by reacting a population of sample nucleic acid molecules with a set of molecular beacon tags (MITs), wherein the sample is 2.5, 2.0, 1.0, or 0.5 ml or less, and the number of different MITs in the set of MITs is between 10 and 100, 200, 250, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, and the total number of sample nucleic acid molecules in the population of sample nucleic acid molecules and the number of different MITs in the set of MITs are is at least 100:1, 500:1, 1,000:1, 10,000:1, or 100,000:1, wherein each tagging nucleic acid molecule comprises one or two MITs located 5' and 3' relative to a nucleic acid segment from the population of nucleic acid molecules, e.g., two MITs located 5' and 3', respectively, and wherein a portion of the sample nucleic acid molecules comprise one or more target loci among a plurality of loci on a chromosome or chromosome segment of interest; amplifying the population of tagged nucleic acid molecules to create a library of tagged nucleic acid molecules; determining the sequences of the bound MITs of the tagged nucleic acid molecules and at least a portion of the sample nucleic acid segments in the library of tagged nucleic acid molecules, for example, determining the sequences of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 95%, or 100%, wherein the sequences of the bound MITs and at least a portion of the nucleic acid segments on each tagged nucleic acid molecule are used to identify tagged nucleic acid molecules belonging to the same paired MIT nucleic acid segment family, wherein at least two MITs on each member of the paired MIT nucleic acid segment family are identical or complementary, and the nucleic acid molecule segments of each member of the MIT nucleic acid segment family are mapped to the same coordinates on the genome of the source of the population of sample nucleic acid molecules; determining the DNA content for each target locus by counting the number of MIT nucleic acid segment families spanning each target locus for the sample nucleic acid molecules; and In silico, determining the copy number of one or more chromosomes or chromosome segments of interest using the amount of DNA at each target locus in the sample nucleic acid molecule. The total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules in the reaction mixture.
[0032] In some embodiments, the method may further include generating, on a computer, a plurality of ploidy hypotheses each associated with a different possible ploidy state of the chromosome or chromosome segment of interest, and identifying the copy number for the individual by determining, on a computer, the relative probability of each ploidy hypothesis using the DNA amount for each target locus and selecting the ploidy state corresponding to the hypothesis with the greatest probability.
[0033] In some embodiments, the methods can be used on maternal samples where copy number determination is for non-invasive prenatal testing. In some embodiments, the methods can be used on samples from individuals with or suspected of having cancer.
[0034] In another aspect, provided herein is a reaction mixture comprising a population of between 500,000,000 and 1,000,000,000,000 sample nucleic acid molecules each having a length of 10 to 1,000 nucleotides, a set of 10 to 1,000 molecular indicator tags (MITs) each having a length of 4 to 8 nucleotides, and a ligase, wherein the MITs are nucleic acid molecules, the ratio of the total number of sample nucleic acid molecules in the reaction mixture to the diversity of the MITs in the set of MITs is 1,000:1 to 1,000,000:1, the sequence of each MIT in the set of MITs differs by at least two nucleotides from all other MIT sequences in the set, and the set includes at least two copies of each MIT.
[0035] In some embodiments, the method may further comprise using a sample nucleic acid molecule that has not been amplified in vitro. In some embodiments, the method may be used for maternal samples, where copy number determination is for non-invasive prenatal testing. In some embodiments, the method may be used for samples from individuals with or suspected of having cancer.
[0036] In another aspect, provided herein is a reaction mixture comprising: a population of 500,000,000 to 5,000,000,000,000 sample nucleic acid molecules; and a set of primers having sequences designed to bind to internal sequences of the sample nucleic acid molecules, wherein the primers further comprise molecular indicator tags (MITs) from a set of 10 to 500 MITs, where the MITs are nucleic acid molecules having a length of 4 to 8 nucleotides, the ratio of the diversity of the sample nucleic acid molecules in the reaction mixture to the diversity of the MITs in the set of MITs in the reaction mixture is 10,000:1 to 1,000,000:1, and the sequence of each MIT in the set of MITs differs from all other MIT sequences in the set by at least 2 nucleotides.
[0037] In some embodiments, the method may further comprise having more primers in the reaction mixture than the total number of sample nucleic acid molecules.
[0038] In another aspect, provided herein is a population of tagged nucleic acid molecules comprising 500,000,000 to 5,000,000,000,000 different tagged nucleic acid molecules 10 to 1,000 nucleotides in length, wherein each of the tagged nucleic acid molecules comprises at least one molecular indicator tag (MIT) located 5' and / or 3' to a sample nucleic acid segment, and wherein the at least one MIT is a member of a set of 10 to 500 different MITs, each 4 to 20 nucleotides in length, and the population of tagged nucleic acid molecules comprises each member of the set of MITs, and at least two tagged nucleic acid molecules of the population comprise at least one identical MIT and a sample nucleic acid segment that differs by 50% or more, and the ratio of the number of sample nucleic acid segments to the number of MITs in the population is 1,000:1 to 1,000,000,000:1.
[0039] In some embodiments, the population of tagged nucleic acid molecules can be part of a reaction mixture that further comprises a polymerase or ligase. In various embodiments, the population of nucleic acid molecules can be used to create a library, wherein the library comprises from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, 250, 500, and 1,000 copies of some or all of the population of nucleic acid molecules at the lower end of the range, to 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, 250, 500, 1,000, 2,500, 5,000, and 10,000 copies of some or all of the population of nucleic acid molecules at the upper end of the range. In some embodiments, the library can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1,000 tagged nucleic acid molecules having an identical sequence to a MIT and a sample nucleic acid segment that is identical from 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, and 99.9% identity at the low end of the range to 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, and 100% identity at the high end of the range. In various embodiments, the library can comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, or 1,000 tagged nucleic acid molecules having an identical sequence between the MIT and the sample nucleic acid segment with a difference of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, or 25 nucleotides. In some embodiments, the library of nucleic acid molecules can be clonally amplified onto a solid support or multiple solid supports.
[0040] In another aspect, provided herein is a population of tagged nucleic acid molecules, the population formed by a method comprising attaching at least one molecular beacon tag (MIT) to a population of 500,000,000 to 5,000,000,000,000 sample nucleic acid molecules comprising sample nucleic acid segments of 50 to 500 nucleotides in length to form tagged nucleic acid molecules comprising at least one MIT located 5' and / or 3' to the sample nucleic acid segment, wherein the MIT is a nucleic acid molecule and the MIT is a member of a set of 10 to 500 different MITs, each MIT being 4 to 20 nucleotides in length, the population of tagged nucleic acid molecules comprising each member of the set of MITs, at least two tagged nucleic acid molecules of the population comprising at least one identical MIT and a sample nucleic acid segment that differs by more than 50%, and the ratio of the diversity of the sample nucleic acid molecule segments in the population to the diversity of the MITs in the set of MITs is 1,000:1 to 1,000,000,000:1.
[0041] In another aspect, provided herein is a first container containing a ligase and a second container containing a set of molecular beacon tags (MITs), wherein each MIT in the set of MITs comprises a portion of a Y adaptor nucleic acid molecule of a set of Y adaptor nucleic acid molecules, each Y adaptor of the set comprises a base-paired double-stranded polynucleotide segment and at least one unbase-paired single-stranded polynucleotide segment, the sequences of each of the Y adaptor nucleic acid molecules in the set other than the MIT sequence are identical, the MIT is a double-stranded sequence that is a portion of the base-paired double-stranded polynucleotide segment, the set of MITs comprises 10 to 500 MITs, the MITs are 4 to 8 nucleotides in length, and the sequence of each MIT in the set of MITs differs by at least 2 nucleotides from every other MIT sequence in the set. The kit can further include a polymerase.
[0042] In some embodiments disclosed herein, the present disclosure provides a reaction mixture in which a population of sample nucleic acid molecules is combined with a set of MITs under appropriate conditions to bind the MITs to nucleic acid molecules or nucleic acid segments of nucleic acid molecules, thereby generating a population of tagged nucleic acid molecules. In some embodiments disclosed herein, the population of tagged nucleic acid molecules can be processed by amplification, which can be part of a high-throughput sequencing sample preparation workflow, for example, and used for downstream analysis, such as high-throughput sequencing. The MIT can be attached via direct ligation or as part of an amplification primer, such as a PCR primer. Typically, the MIT is 5' to the sequence-specific binding region of the primer, but the primer can be designed to be between the universal binding region and the sequence-specific binding region, or the MIT is within the sequence-specific binding region and forms a loop upon hybridization with the sample nucleic acid molecule. In some embodiments, the MIT can be present on the forward primer, such that amplification using the primer generates tagged nucleic acid molecules with the MIT 5' to the target locus. In some embodiments, the MIT can be present on the reverse primer, such that amplification using the primer generates tagged nucleic acid molecules with the MIT 3' to the target locus. In some embodiments, MIT can be present on both the forward and reverse primers, such that amplification with the primers generates a tagged nucleic acid molecule with MIT both 5' and 3' to the target locus.
[0043] In some embodiments disclosed herein, MIT can be a single-stranded or double-stranded nucleic acid molecule. In some embodiments, the sequence of MIT can be different from the sequence of all other MITs in the set of MIT by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments, the MITs in the set of MIT are typically the same length. In other embodiments, the MITs in the set of MIT are different lengths. In any of the embodiments disclosed herein, the length of MIT is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.
[0044] In some embodiments, MIT can be a Y adaptor, or a single-stranded oligonucleotide, or at least a portion of a double-stranded nucleic acid, e.g., a double-stranded adaptor. In some embodiments, MIT can be part of a Y adaptor nucleic acid molecule of a set of Y adaptor nucleic acid molecules, each Y adaptor of the set comprising a base-paired double-stranded polynucleotide segment and at least one unbase-paired single-stranded polynucleotide segment, the sequence of each Y adaptor nucleic acid molecule in the set other than the MIT sequence is identical, and MIT is a double-stranded sequence that is part of the base-paired double-stranded polynucleotide segment. In some embodiments, the double-stranded polynucleotide segment is between 5, 10, 15, and 20 nucleotides in length at the lower end of the range and 10, 15, 20, 25, 30, 35, 40, 45, and 50 nucleotides in length at the upper end of the range, and does not include the MIT, and the single-stranded polynucleotide segment can be between 5, 10, 15, and 20 nucleotides in length at the lower end of the range and 10, 15, 20, 25, 30, 35, 40, 45, and 50 nucleotides in length at the upper end of the range. In some embodiments, the MIT can be between 3, 4, 5, 6, 7, 8, 9, 10, or 15 nucleotides in length at the lower end of the range and 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 nucleotides in length at the upper end of the range. In some embodiments disclosed herein, MIT can be part of an oligonucleotide that further comprises a sequence designed to bind to a sample nucleic acid molecule, a universal primer binding sequence, and / or an adaptor sequence, particularly an adaptor sequence useful for high-throughput sequencing.In some embodiments, the total length of the oligonucleotide can be between 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides at the lower end of the range and 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides at the upper end of the range.In some embodiments, one or more MITs can bind to a sample nucleic acid molecule.For example, in some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 MITs can bind to a sample nucleic acid molecule.In some embodiments disclosed herein, MIT can be attached to the 5' and / or 3' of the sample nucleic acid segment, which can be part or all of the sample nucleic acid molecule. In some embodiments, two MITs can be attached to each sample nucleic acid molecule, for example, each sample nucleic acid molecule, and each tagged nucleic acid molecule comprises two MITs located 5' and 3', respectively, relative to the nucleic acid segment from a population of nucleic acid molecules.
[0045] In some embodiments disclosed herein, the sample nucleic acid molecules can be used in a reaction mixture before any other in vitro amplification is performed. In some embodiments, the total number of sample nucleic acid molecules in the population of nucleic acid molecules is greater than or equal to the lower end of the range: 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1 x 10 6 , 2.5×10 6 , 5×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , and 1 × 10 10 of sample nucleic acid molecules and the upper end of the range: 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1 x 10 6 , 2.5×10 6 , 5×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , 1×10 10 , 1×10 11 , and 1 × 10 12In some embodiments disclosed herein, the total number of sample nucleic acid molecules in the reaction mixture may be greater than the diversity of MITs in the set of MITs. For example, the ratio of the total number of sample nucleic acid molecules to the diversity of MITs in the set of MITs is at least 2:1, 10:1, 100:1, 1,000:1, 5,000:1, 10,000:1, 25,000:1, 50,000:1, 100,000:1, 250,000:1, 500,000:1, 1,000,000:1, 5,000,000:1, 10,000,000:1, 1×10 8 :1, 1×10 9 :1, 1×10 10 :1 or more. In some embodiments, the diversity of the possible combinations of bound MITs may be greater than the total number of sample nucleic acid molecules in the reaction mixture that span the target loci. For example, the ratio of the diversity of the possible combinations of bound MITs (e.g., any combination of 2, 3, 4, 5, etc., depending on the number of MITs bound to the sample nucleic acid molecules) to the total number of sample nucleic acid molecules that span the target loci may be at least 1.0:1, 1.1:1, 1.5:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 25:1, 50:1, 100:1, 500:1, or 1,000:1. In some embodiments, the MITs in the set of MITs are at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1 x 10 6 , 2.5×10 6 , 5×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , 1×10 10 , 1×10 11 , or 1×10 12 different sample nucleic acid molecules to generate a population of tagged nucleic acid molecules.
[0046] In some embodiments disclosed herein, the antibody concentration is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1 x 10 6 , 2.5×10 6 , 5×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , 1×10 10 , 1×10 11 , and 1 × 10 12 of the sample nucleic acid molecules in the reaction mixture can have bound MIT. In some embodiments, at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the sample nucleic acid molecules in the reaction mixture can have bound MIT.
[0047] In some embodiments disclosed herein, the reaction mixture may contain more MIT molecules than the sample nucleic acid molecules. For example, in some embodiments, the total number of MIT molecules in the reaction mixture may be at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 times the total number of sample nucleic acid molecules in the reaction mixture. In some respects, this fold difference depends on the number of MITs added. For example, when two MITs are combined, the total number of MIT molecules in the reaction mixture may be at least two times greater than the total number of sample nucleic acid molecules in the reaction mixture. When three MITs are combined, the total number of MIT molecules in the reaction mixture may be at least three times greater than the total number of sample nucleic acid molecules in the reaction mixture, and so on. In some embodiments, the ratio of the total number of MITs having identical sequences in the reaction mixture to the total number of nucleic acid molecules in the reaction mixture can be between 0.1:1, 0.2:1, 0.3:1, 0.4:1, 0.5:1, 1:1, 1.5:1, 2:1 at the lower end of the range and 0.3:1, 0.4:1, 0.5:1, 1:1, 1.5:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, and 10:1 at the upper end of the range.
[0048] In some embodiments, the sequence of the combined MIT and nucleic acid segment in a population of tagged nucleic acid molecules can be determined by sequencing, particularly high-throughput sequencing. In some embodiments, the tagged nucleic acid molecules can be clonally amplified, particularly on a solid support or multiple solid supports, for sequencing. In some embodiments, the determined sequence of the MIT on the tagged nucleic acid molecule can be used to identify the sample nucleic acid molecule from which the tagged nucleic acid molecule originates, particularly using the sequence of the end of the nucleic acid segment or the end of a fragment-specific insert disclosed herein. In some embodiments, the determined sequence of the nucleic acid segment on the tagged nucleic acid molecule can be used to help identify the sample nucleic acid molecule from which the tagged nucleic acid molecule originates. In some embodiments, the determined sequence of the nucleic acid segment can be mapped to the location in the genome of the source of the sample nucleic acid molecule, and this information can be used to help identify.
[0049] In some embodiments, the lower end of the range is 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1×10 6 , 2.5×10 6 , 5×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , and 1 × 10 10 tagged nucleic acids and at the upper end of the range: 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1x10 6 , 2.5x10 6 , 5x10 6 , 1x10 7 , 1x10 8 , 1x10 9 , 1x10 10 , 1x10 11 , 1x10 12 In some embodiments, tagged nucleic acid molecules from the two strands of a single sample nucleic acid molecule can be identified and used to generate paired MIT families. Typically, in downstream sequencing reactions where single-stranded nucleic acid molecules are sequenced, MIT families can be identified by identifying tagged nucleic acid molecules with identical or complementary MIT sequences. In these embodiments, paired MIT families can be used to confirm the presence of sequence differences in sample nucleic acid molecules. In some further embodiments, the determined sequences of nucleic acid segments can be used to generate paired MIT nucleic acid segment families with complementary or identical MIT and nucleic acid segment sequences. In these embodiments, paired MIT nucleic acid segment families can be used to confirm the presence of sequence differences in sample nucleic acid molecules.
[0050] In some embodiments, tagged nucleic acid molecules with specific target loci can be enriched. In some embodiments, one-sided or two-sided PCR can be used to enrich these target loci on one or more chromosomes. In some embodiments, hybrid capture can be used. In some embodiments, enrichment can be targeted between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 15,000, or 20,000 target loci at the lower end of the range and 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 250, 500, 1,000, 2,500, 5,000, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, and 250,000 target loci at the upper end of the range. In some embodiments, the target locus can be between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, and 100 nucleotides in length at the lower end of the range, and 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, and 1,000 nucleotides in length at the upper end of the range. In some embodiments, target loci on different sample nucleic acid molecules can be at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% identical or share at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% sequence identity.
[0051] In some embodiments disclosed herein, the sample may be derived from a mammal. In some embodiments, the sample may be derived from a human, particularly a sample of human blood or a fraction thereof. In any of the disclosed embodiments, the sample may be less than 0.1, 0.2, 0.25, 0.5, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 3.5, 4, 4.5, or 5 ml of blood or plasma. In some embodiments disclosed herein, the sample may contain circulating cell-free human DNA. In some embodiments, the sample containing circulating cell-free human DNA may be derived from a mother and may contain maternal and fetal DNA. In some embodiments, the sample containing circulating cell-free human DNA may be a blood sample from a person with or suspected of having cancer and may contain normal and tumor DNA.
[0052] Other features and advantages of the present disclosure will become apparent from the following detailed description and claims. [Brief explanation of the drawings]
[0053] [Figure 1] 1 is a schematic diagram showing the attachment of two MITs to a nucleic acid molecule or nucleic acid segment using ligation. In order of appearance, SEQ ID NOs: 1-2, 2, 2, 1, 3-4, 4, and 3 are disclosed, respectively. [Figure 2] 1 is a schematic diagram showing the incorporation of two MITs into a nucleic acid molecule or nucleic acid segment using PCR with primers containing the MIT sequences. Disclosed in order of appearance are SEQ ID NOS: 5-6, 6, 5, 7-8, 8, 7, and 9-14, respectively. [Figure 3]The structures of amplicons generated by different exemplary methods provided herein are shown. The amplicon generated after single-sided STAR (FIG. 3A) has an MIT on one side, and the first base of MIT is the first base of Read 1 or Read 2, depending on how the single-sided STAR is performed. In FIG. 3A, the first base of MIT will be the first base of Read 1. The amplicon generated after hybrid capture (FIG. 3B) has MIT on both sides of the amplicon, with the first base of Read 1 being the first base of MIT1 and the first base of Read 2 being the first base of MIT. [Figure 4] 1 is a table showing the results of a sequencing experiment using MIT. [Figure 5] 5 is a bar graph showing the average error rate and the average error rate of paired MIT nucleic acid segment families for two samples in three different experiments (data from FIG. 4).
[0054] The above-identified figures are provided by way of illustration and not by way of limitation.
[0055] (Detailed Description of the Invention) The present disclosure relates to methods and compositions comprising oligonucleotide tags, referred to herein as molecular indicator tags (MITs), which are attached to a population of nucleic acid molecules derived from a sample to identify individual sample nucleic acid molecules (i.e., members of the population) from the population of nucleic acid molecules after sample processing for a sequencing reaction. In some embodiments, the sequencing reaction is a high-throughput sequencing reaction performed on tagged nucleic acid molecules derived from the sample nucleic acid molecules. Unlike prior art methods that teach tagging each sample nucleic acid molecule with a unique identifier, with a diversity of unique identifiers greater than the number of sample nucleic acid molecules in the sample, the present disclosure typically includes more sample nucleic acid molecules than the diversity of MITs in the set of MITs. In fact, the methods and compositions herein provide greater than 1,000 unique identifiers for each different MIT in the set of MITs, up to 1 x 10 6 Super, 1×10 9The method can include more than one or even more starting molecules, and still identify individual sample nucleic acid molecules that give rise to tagged nucleic acid molecules after amplification.
[0056] In the methods and compositions herein, the diversity of the set of MITs is advantageously smaller than the total number of sample nucleic acid molecules spanning the target locus, while the diversity of the possible combinations of bound MITs using the set of MITs is greater than the total number of sample nucleic acid molecules spanning the target locus. Typically, to improve the identification ability of the set of MITs, at least two MITs are bound to the sample nucleic acid molecule to form a tagged nucleic acid molecule. The sequence of the bound MITs determined from the sequencing read can be used to identify clonally amplified identical copies of the same sample nucleic acid molecule bound to different solid supports or different regions of a solid support during sample preparation for the sequencing reaction. The sequences of the tagged nucleic acid molecules can be compiled, compared, and used to distinguish nucleotide mutations that occurred during amplification from nucleotide differences that were present in the original sample nucleic acid molecule.
[0057] The set of MITs in the present disclosure typically has a diversity less than the total number of sample nucleic acid molecules, whereas many conventional methods utilize a set of "unique identifiers" in which the diversity of unique identifiers is greater than the total number of sample nucleic acid molecules. However, the MITs of the present disclosure maintain sufficient tracking power by using a set of MITs greater than the total number of sample nucleic acid molecules spanning the target locus, thereby encompassing the diversity of possible combinations of combined MITs. This smaller diversity in the set of MITs of the present disclosure significantly reduces the cost and manufacturing complexity associated with generating and / or obtaining a set of tracking tags. While the total number of MIT molecules in the reaction mixture is typically greater than the total number of sample nucleic acid molecules, the diversity of the set of MITs is much less than the total number of sample nucleic acid molecules, which substantially reduces costs and simplifies manufacturing compared to prior art methods. Thus, the set of MITs can encompass a diversity between the lower end of the range (3, 4, 5, 10, 25, 50, or 100 distinct MITs) and the upper end of the range (10, 25, 50, 100, 200, 250, and 1,000 distinct MITs). Therefore, in the present disclosure, this relatively low diversity of MITs results in a diversity of MITs that is much smaller than the total number of sample nucleic acid molecules, which, when combined with the total number of MITs in the reaction mixture being greater than the total number of sample nucleic acid molecules, and the greater diversity in the possible combinations of any two MITs in the set of MITs than the number of sample nucleic acid molecules spanning the target locus, and the greater diversity in the possible combinations of any two MITs in the set of MITs than the number of sample nucleic acid molecules spanning the target locus, provides a particularly advantageous embodiment that is cost-effective and highly effective with complex samples isolated from nature. Furthermore, mapping sequenced nucleic acid molecules to a genome provides additional advantages, such as simpler analysis and identification information regarding the sequence of sample nucleic acid molecules compared to a reference genome.
[0058] Brief Description of Exemplary Methods Thus, in one embodiment, provided herein is a method for sequencing a population of sample nucleic acid molecules, which may optionally further comprise using sequencing to identify individual sample nucleic acid molecules from the population of sample nucleic acid molecules. In some embodiments, the population of nucleic acid molecules has not been in vitro amplified before binding MIT, and is 1 x 10 8 ~1×10 13 , or in some embodiments, 1 x 10 9 ~1×10 12 , or 1×10 10 ~1×10 12 The sample nucleic acid molecules may comprise 10 to 500 MITs having different sequences. In certain methods and compositions herein, the ratio of the total number of nucleic acid molecules in a population of nucleic acid molecules to the diversity of MITs in the set of MITs may be 1,000:1 to 1,000,000,000:1. The ratio of the diversity of the possible combinations of MITs in a set of MITs to the total number of sample nucleic acid molecules spanning the target locus may be 1.01:1 to 10:1. As discussed in further detail herein, MITs are typically composed, at least in part, of oligonucleotides 4 to 20 nucleotides in length. The set of MITs can be designed such that the sequences of all MITs in the set differ from each other by at least 2, 3, 4, or 5 nucleotides.
[0059] In some embodiments provided herein, at least one (e.g., two) MITs from a set of MITs are attached to each nucleic acid molecule or segment of each nucleic acid molecule in a population of nucleic acid molecules to generate a population of tagged nucleic acid molecules. As further discussed herein, MITs can be attached to sample nucleic acid molecules in various configurations. For example, after attachment, one MIT can be located at the 5' end of the tagged nucleic acid molecule, or 5' to some, most, or typically each sample nucleic acid segment of the tagged nucleic acid molecule, and / or another MIT can be located 3' to some, most, or typically each sample nucleic acid segment of the tagged nucleic acid molecule. In other embodiments, at least two MITs are located 5' and / or 3' to the sample nucleic acid segment of the tagged nucleic acid molecule, or 5' and / or 3' to some, most, or typically each sample nucleic acid segment of each tagged nucleic acid molecule. The two MITs can be added at either 5' or 3' by including them on the same polynucleotide segment before attachment or by performing separate reactions. For example, PCR can be performed using primers that bind to specific sequences within a sample nucleic acid molecule and include a region 5' to a sequence-specific region encoding two MITs. In some embodiments, at least one copy of each MIT of a set of MITs is bound to a sample nucleic acid molecule, and each of two copies of at least one MIT is bound to a different sample nucleic acid molecule, and / or at least two nucleic acid molecules having the same or substantially the same sequence have at least one different bound MIT. Those skilled in the art will identify methods for binding MITs to nucleic acid molecules of a population of nucleic acid molecules. For example, MITs can be bound via linkage or added 5' to an internal sequence binding site of a PCR primer and bound during a PCR reaction, as discussed in more detail herein.
[0060] After or while MIT binds to the sample nucleic acids to form tagged nucleic acid molecules, the population of tagged nucleic acid molecules is typically amplified to create a library of tagged nucleic acid molecules. Amplification methods for creating libraries, including those particularly relevant to high-throughput sequencing workflows, are known in the art. For example, such amplification can be PCR-based library preparation. These methods can further include clonally amplifying the library of tagged nucleic acid molecules onto one or more solid supports using PCR or another amplification method (such as an isothermal method). Methods for creating clonally amplified libraries on solid supports in high-throughput sequencing sample preparation workflows are known in the art. Additional amplification steps, such as multiplex amplification reactions in which a subset of the population of sample nucleic acid molecules is amplified, can also be included in the methods for identifying sample nucleic acids provided herein.
[0061] In some embodiments of the methods provided herein, the nucleotide sequence of MIT and some, most, or all (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 25, 50, 75, 100, 150, 200, 250, 500, 1,000, 2,500, 5,000, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, 5,000,000, 10,000,000, 25,000,000, 50,000,000, 100,000,000, 250,000,000, 500,000,000, 1 x 10 9 , 1×10 10 , 1×10 11 , 1×10 12 , or 1×10 13The nucleotide sequence of at least a portion of the sample nucleic acid molecule segments (between 10, 20, 25, 30, 40, 50, 60, 70, 80, or 90% of the tagged nucleic acid molecules at the low end of the range and 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, 96, 97, 98, 99, and 100% of the tagged nucleic acid molecules at the high end of the range) is determined. The sequences of the first MIT and optionally the second MIT or more MITs on the clonally amplified copies of the tagged nucleic acid molecules can be used to identify the individual sample nucleic acid molecules that gave rise to the clonally amplified tagged nucleic acid molecules in the library.
[0062] In some embodiments, sequences determined from tagged nucleic acid molecules sharing the same first MIT and, optionally, the same second MIT can be used to distinguish amplification errors from true sequence differences at target loci in sample nucleic acid molecules, thereby identifying amplification errors. For example, in some embodiments, the set of MITs is a double-stranded MIT, which can be part of a partially or completely double-stranded adapter, such as a Y adapter. In these embodiments, for every starting molecule, the Y adapter preparation generates two daughter molecule types (one in the + direction and one in the - direction). A true mutation in the sample molecule will have both daughter molecules paired with the same two MITs in these embodiments where the MIT is a double-stranded adapter or part thereof. Furthermore, when tagged nucleic acid molecules are sequenced and grouped into MIT nucleic acid segment families by the MIT in the sequence, considering the MIT sequence and optionally its complement to the double-stranded MIT, and optionally considering at least a portion of the nucleic acid segment, if the starting molecule from which the tagged nucleic acid molecule is generated contains a mutation, most, and typically at least 75% of the nucleic acid segments in the MIT nucleic acid segment family in double-stranded MIT embodiments will contain the mutation. In the case of an amplification (e.g., PCR) error, the worst-case scenario is that the error occurs in cycle 1 of the first PCR. In these embodiments, the amplification error causes 25% of the final product to contain errors (plus any additional cumulative errors, which should be <1%). Thus, in some embodiments, for example, if an MIT nucleic acid segment family contains at least 75% of reads for a particular mutation or polymorphic allele, it can be concluded that the mutation or polymorphic allele is truly present in the sample nucleic acid molecule that gave rise to the tagged nucleic acid molecule. The later the error occurs in the sample preparation process, the lower the proportion of sequence reads containing errors in the set of sequencing reads grouped (i.e., bucketed) into MIT nucleic acid segment families paired by MIT.For example, an error in library preparation amplification will result in a higher percentage of sequences with errors in the paired MIT nucleic acid segment family than an error in a subsequent amplification step in a workflow such as targeted multiplex amplification. An error in the final clonal amplification in a sequencing workflow will produce the lowest percentage of nucleic acid molecules in the paired MIT nucleic acid segment family that contain the error.
[0063] Any sequencing method can be used to perform the methods provided herein, particularly methods for determining the sequence of a sample nucleic acid molecule, or particularly a plurality of sample nucleic acid molecules, using multiple amplified copies of the sample nucleic acid molecule. Furthermore, sample nucleic acid segments and tagged nucleic acid molecules that yield substantially identical (e.g., at least 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96, 97, 98, or 99% identical) sequences for different MIT tags can be compared to determine sequence diversity in a population of sample nucleic acid molecules and distinguish true variants or mutations from errors occurring during sample preparation, even at low allele frequencies. Embodiments of the disclosed methods include methods for sequencing a population of sample nucleic acid molecules. Such methods are particularly useful for high-throughput sequencing methods. Such methods are discussed in more detail herein.
[0064] The methods described above and disclosed herein can be used for many purposes, as will be recognized by those skilled in the art in light of this disclosure. For example, the methods can be used to determine the nucleic acid sequence of a population of nucleic acid molecules in a sample, to identify the sample nucleic acid molecule that gave rise to a tagged nucleic acid molecule, to identify a sample nucleic acid molecule from a population of sample nucleic acid molecules, to identify amplification errors, to measure amplification bias, and to characterize the mutation rate of a polymerase. Additional uses will be apparent to those skilled in the art. In these methods, after determining the sequence of the tagged nucleic acid segments, nucleic acid segments having substantially the same nucleic acid segment sequence and the same two MIT tags, or nucleic acid segments having substantially the same or the same nucleic acid segment sequence and at least one different MIT tag, can be used for comparison and further analysis.
[0065] Sample and library preparation In various embodiments provided herein, a sample can be derived from natural or non-natural sources. In some embodiments, the nucleic acid molecules in the sample can be derived from an organism or cell. Any nucleic acid molecule can be used; for example, the sample can include genomic DNA, mRNA, or miRNA covering a portion of the entire genome from the organism or cell. In some respects, the total length of the entire genome or DNA sequence in the sample divided by the average size of the nucleic acid molecules can be used to determine the number of nucleic acid molecules in the sample, representing the entire genome or DNA sequence. In further respects, this number can be used to determine the number of nucleic acid molecules spanning a target locus in the sample. A locus can comprise a single nucleotide or a segment of 1 to 1,000, 10,000, 100,000, 1 million, or more nucleotides. As a non-limiting example, a locus can be a single nucleotide polymorphism, an intron, or an exon. In some embodiments, a locus can comprise an insertion, deletion, or rearrangement. In some embodiments, the sample can include a blood, serum, or plasma sample. In some embodiments, the sample may contain free-floating DNA (e.g., circulating cell-free tumor DNA or circulating cell-free fetal DNA) in blood, serum, or plasma. In these embodiments, the sample is typically from an animal, such as a mammal or human, and is typically present in fragments of about 160 nucleotides in length. In some embodiments, free-floating DNA is isolated from blood using EDTA-2Na tubes after removal of cellular debris and platelets by centrifugation. Plasma samples can be stored at -80°C until DNA is extracted, for example, using a QIAamp DNA Mini Kit (Qiagen, Hilden, Germany) (e.g., Hamakawa et al., Br J Cancer. 2015; 112:352-356). However, samples may be derived from other sources, and nucleic acid molecules from any organism can be used in this method. In some embodiments, DNA from bacteria and / or viruses can be used to analyze true sequence variants within mixed populations, particularly in environmental and biodiversity sampling.
[0066] Some embodiments disclosed herein are typically carried out using sample nucleic acid molecules produced in and by living cells. Such nucleic acid molecules are typically isolated directly from natural sources, such as cells or body fluids, without any in vitro amplification before MIT binding. Thus, the sample nucleic acid molecules are directly used in the reaction mixture for MIT binding. This avoids the potential introduction of amplification errors before the sample nucleic acid molecules are tagged. This, in turn, improves the ability to distinguish actual sequence variants from amplification errors. However, in some embodiments, the sample nucleic acid molecules can be amplified before MIT binding. Those skilled in the art will understand the best method to use when amplification is necessary before MIT binding. For example, a high-fidelity polymerase with proofreading capabilities can be used for amplification to help reduce the number of amplification errors that may occur before the nucleic acid molecules bind MIT. Furthermore, a small number of amplification cycles (e.g., between 2, 3, 4, and 5 cycles at the lower end of the range and 3, 4, 5, 6, 7, 8, 9, or 10 cycles at the upper end of the range) can be used.
[0067] In some embodiments, nucleic acid molecules in a sample can be fragmented to generate nucleic acid molecules of any selected length before being tagged with MIT.Those skilled in the art will recognize the method for performing such fragmentation and the length to be selected, as discussed in more detail herein.For example, nucleic acid fragmentation can be performed using physical methods such as sonication, enzymatic methods such as digestion with DNase I or restriction endonucleases, or chemical methods such as applying heat in the presence of divalent metal cations.As discussed in more detail herein, fragmentation can be performed to leave nucleic acid molecules of a selected size range.In other embodiments, nucleic acid molecules can be selected to have a specific size range using methods known in the art.
[0068] After fragmentation, sample nucleic acid molecules may have 5' and / or 3' overhangs that need to be repaired prior to further library preparation. In some embodiments, prior to the attachment of MIT or other tags, sample nucleic acid molecules with 5' and 3' overhangs can be repaired to generate blunt-ended sample nucleic acid molecules using methods known in the art. For example, the polymerase activity and exonuclease activity of Klenow large fragment polymerase can be used in an appropriate buffer to fill in 5' overhangs and remove 3' overhangs on the nucleic acid molecules. In some embodiments, a phosphate can be added to the 5' end of the repaired nucleic acid molecule using polynucleotide kinase (PNK) and reaction conditions understood by those skilled in the art. In further embodiments, a single nucleotide or multiple nucleotides can be added to one strand of a double-stranded molecule to generate a "sticky end." For example, adenosine (A) can be added to the 3' end of the nucleic acid molecule (A-tailing). In some embodiments, sticky ends other than A-overhangs can be used. In some embodiments, other adapters, such as looped ligation adapters, can be added. None of these modifications, or all or any combination thereof, may be performed in any of the embodiments disclosed herein.
[0069] Many kits and methods are known in the art for generating libraries of nucleic acid molecules for subsequent sequencing. Kits specifically modified for preparing libraries from small nucleic acid fragments, particularly circulating cell-free DNA, can be useful in performing the methods provided herein. For example, the NEXTflex Cell Free Kit (Bioo Scientific, Austin, TX) or the Natera Library Prep Kit (Natera, San Carlos, CA). Such kits will typically be modified to include adapters customized for the amplification and sequencing steps of the methods provided herein. Adapter ligation can also be performed using commercially available kits, such as the ligation kit found in the Agilent SureSelect Kit (Agilent, Santa Clara, CA).
[0070] The sample nucleic acid molecule is composed of natural or unnatural ribonucleotides or deoxyribonucleotides linked via phosphodiester bonds. Furthermore, the sample nucleic acid molecule comprises a nucleic acid segment that is the target for sequencing. The sample nucleic acid molecule can be or comprise a nucleic acid segment that is at least 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotides in length. In any of the embodiments disclosed herein, the sample nucleic acid molecule or nucleic acid segment may be at the lower end of the range of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, and 500 nucleotides in length; At the upper end of the range, the length may be between 10, 11, 12, 13, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, and 10,000 nucleotides. In some embodiments, the nucleic acid molecules can be fragments of genomic DNA and can range in length from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, and 500 nucleotides at the lower end of the range and from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, and 500 nucleotides at the upper end of the range. The length of the nucleic acid molecule may be between 0, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, and 10,000 nucleotides. For clarity, nucleic acids initially isolated from biological tissues, fluids, or cultured cells may be much longer than the sample nucleic acid molecules processed using the methods herein. As discussed herein, for example, such initially isolated nucleic acid molecules can be fragmented to generate nucleic acid segments before use in the methods herein.In some embodiments, the nucleic acid molecule and the nucleic acid segment can be identical. The sample nucleic acid molecule or sample nucleic acid segment can comprise a target locus that contains one or more nucleotides being queried, particularly a single nucleotide polymorphism or single nucleotide variant. In any of the disclosed embodiments, the target locus can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotides in length and can comprise a portion or the entire sample nucleic acid molecule and / or sample nucleic acid segment. In other embodiments, the target locus is between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, and 500 nucleotides in length at the lower end of the range and 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, and 500 nucleotides in length at the upper end of the range. In some embodiments, the target loci on different sample nucleic acid molecules can be at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% identical. In some embodiments, target loci on different sample nucleic acid molecules can share at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% sequence identity.
[0071] In some embodiments, the entire sample nucleic acid molecule is a sample nucleic acid segment.For example, in certain embodiments where MIT is directly linked to the end of the sample nucleic acid molecule, or linked to a nucleic acid linked to the end of the sample nucleic acid molecule, or linked as part of a primer that binds to the sequence at the end of the sample nucleic acid segment, or an adapter such as a universal adapter that is added thereto, as further discussed herein, the entire nucleic acid molecule can be a sample nucleic acid segment.In other embodiments, such as certain embodiments where MIT is linked to the sample nucleic acid molecule as part of a primer that targets an internal binding site at the end of the sample nucleic acid molecule, a portion of the sample nucleic acid molecule can be a sample nucleic acid segment that is targeted for downstream sequencing.For example, at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the sample nucleic acid molecule can be a nucleic acid segment.
[0072] In some embodiments, the sample nucleic acid molecules are a mixture of nucleic acids isolated from natural sources, and some sample nucleic acid molecules have the same sequence, with some sample nucleic acid molecules having at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% sequence identity, and some having less than 50%, 40%, 30%, 20%, 10%, or 5% sequence identity, ranging from 20, 25, 50, 75, 100, 125, 150, 200, 250, 300, 400, or 500 nucleotides at the lower end of the range. Such sample nucleic acid molecules can be nucleic acid samples isolated from mammalian tissues or body fluids, such as humans, without enriching certain sequences over other sequences. In other embodiments, target sequences, such as those derived from genes of interest, can be enriched before performing the methods provided herein.
[0073] In certain embodiments, some or all of the sample nucleic acid molecules in a population of nucleic acid molecules can have identical or substantially identical nucleic acid segments.If the sequences of the nucleic acid segments share at least 90% sequence identity, the nucleic acid molecules can be said to be substantially identical.In some exemplary embodiments, the sample nucleic acid molecules can share nucleic acid segments with 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9% sequence identity, ranging from 20, 25, 50, 75, 100, 125, 150, 200, or 250 nucleotides at the lower end of the range to 50, 75, 100, 125, 150, 200, 250, 300, 400, or 500 nucleotides at the upper end of the range. The methods provided herein are effective in distinguishing between sample nucleic acid molecules that share at least 90%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity within a sample.
[0074] In some embodiments, the 5' and 3' ends of the nucleic acid segment adjacent to the linked MIT can be used to help identify and distinguish sample nucleic acid molecules. These sequences are referred to herein as fragment-specific insert ends. After MIT binding as discussed elsewhere herein, the combination of the MIT and the fragment-specific insert end can uniquely identify sample nucleic acid molecules. This is because a sufficiently high ratio of MIT to sample nucleic acid molecules can be selected so that the probability that two different sample nucleic acid molecules have the same fragment-specific insert end in the same orientation and the same linked MIT is extremely low. For example, the probability is 1, 0.5, 0.1, 0.05, 0.01, 0.005, 0.001, or less. For example, identifying each sample nucleic acid molecule from a set of 200 MITs using only MIT provides 40,000 (200 x 200) possible combinations of identifiers. Using the additional information provided by the fragment-specific insert end, the number of possible combinations can be rapidly increased. For example, including two nucleotides from the 5' and 3' fragment-specific insert ends in identifying a nucleic acid molecule increases the 40,000 possible combinations to 10,240,000 possible combinations, where each nucleotide is equally likely to occur in a dinucleotide sequence. The length of the fragment-specific insert ends, when used in the methods provided herein, can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 nucleotides at the lower end of the range, and 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 nucleotides at the upper end of the range. , 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50 nucleotides in length. In some embodiments, the fragment-specific ends used in combination with MIT to identify sample nucleic acid molecules are 1, 2, 3, or 4 nucleotides in length.
[0075] In further embodiments, the determined sequences of the fragment-specific insert ends can be used to map each end of the nucleic acid molecule to a specific location (i.e., genomic coordinate) within the genome of the organism from which the sample was isolated. The mapped locations provide a separate identifier for each tagged nucleic acid molecule. Mapping each end significantly increases the number of identifiers available for each tagged nucleic acid molecule. In these embodiments, the mapped locations of each end of the nucleic acid molecule can be used in combination with MIT to identify the individual sample nucleic acid molecule that gave rise to the tagged nucleic acid molecule. For example, for a given target base in mononucleosomal circular cell-free DNA (cfDNA), the 5' fragment end can be anywhere from about 0 to 199 bases upstream. Similarly, the 3' fragment end can be 0 to 199 bases downstream. Theoretically, this would yield 40,000 possible final combinations. In practice, most molecules are 100 to 200 bases in length, so the total number of possible combinations is approximately 15,000 (a maximum, although not all combinations occur with equal probability). This means 40,000 MIT combinations x 15,000 possible fragment ends = 600,000,000 possible end combinations. Furthermore, when a nucleic acid segment is mapped to a genome, mutations in that segment or alleles of that segment can be identified.
[0076] The total number of sample nucleic acid molecules can vary widely depending on the sample source and preparation and the needs of the method. For example, the total sample nucleic acid molecules can be greater than 1 x 10 at the low end of the range. 10 , 2 × 10 10 , 2.5×10 10 , 5×10 10 , and 1 × 10 11 and 5×10 at the top end of the range 10 , 1×10 11 , 2 × 10 11 , 2.5×10 11 , 5×10 11 , 1×10 12 , 2 × 10 12 , 2.5×10 12 , 5×10 12 , and 1 × 10 13For example, because mononucleosomal cfDNA is a set of approximately 100-200 bp nucleic acid fragments with highly variable fragmentation patterns, 10,000 copies of the genome from human circulating cell-free DNA can be as high as 2 x 10 11 total sample nucleic acid molecules (3,000,000,000 bp / genome copy × 10,000 genome copies / 150 bp / sample nucleic acid molecule = 2 × 10 11 (sample nucleic acid molecules).
[0077] In some embodiments provided herein, the total number of sample nucleic acid molecules can range from 50, 100, 200, 250, 500, 750, 1,000, 2,000, 2,500, 5,000, and 10,000 copies of the human genome at the lower end of the range to 1,000, 2,000, 2,500, 5,000, 10,000, 20,000, 25,000, 50,000, and 100,000 copies of the human genome at the upper end of the range. In other embodiments, the total number of sample nucleic acid molecules is the number of nucleic acid molecules of 100 to 500 nucleotides in length, e.g., 200 nucleotides, in cfDNA at the lower end of the range to 2.5, 3, 4, 5, 10, 20, or 25 nM at the upper end of the range.
[0078] The diversity of a set or population of nucleic acid molecules is the number of unique sequences among the nucleic acid molecules in that set or population. The diversity of a sample nucleic acid molecule is the number of unique sequences among the sample nucleic acid molecules. Even if the nucleic acid molecules in a sample have not been subjected to amplification, it is common to have two or more copies of identical or nearly identical nucleic acid sequences in the sample. Current nucleic acid sample preparation and DNA isolation procedures typically result in multiple copies of every nucleic acid molecule in a sample.
[0079] In any of the embodiments disclosed herein, the nucleotide sequence diversity of the sample nucleic acid molecules in the population can be at the lower end of the range of 100, 1,000, 10,000, 1 x 10 5 , 1×10 6 , and 1 × 10 7 different nucleic acid sequences and 1 x 10 at the upper end of the range5 , 1×10 6 , 1×10 7 , 1×10 8 , 1×10 9 , and 1 × 10 10 In some embodiments, the diversity of nucleotide sequences in the population of sample nucleic acid molecules can be between 1 x 10 at the low end of the range and 1 x 10 at the low end of the range. 6 , 5×10 6 , and 1 × 10 7 different nucleic acid sequences and 1 x 10 at the upper end of the range 7 , 1×10 8 , 1×10 9 , and 1 × 10 10 between different nucleotide sequences.
[0080] For human cfDNA sample, there are about 3 billion nucleotides in human genome, nucleic acid fragment size is about 150 nucleotides, and fragmentation pattern is not random but not fixed, so there are about 20 million (3 billion / 150) to about 3 billion different nucleic acid fragments in human cfDNA sample.Therefore, in some embodiments, sample is human cfDNA sample, such as purified sample, or serum or plasma sample, and the diversity of sample is 20 million to 3 billion.
[0081] In certain embodiments of the present disclosure, sample nucleic acid molecules can be approximately the same length.For example, sample nucleic acid molecules can be about 200 nucleotides for circulating cell-free DNA samples, or for certain samples such as blood, serum, plasma samples that contain circulating cell-free DNA, the length of sample nucleic acid molecules can be between the lower end of 50, 75, 100, 125 or 150 nucleotides and the upper end of 150, 200, 250 or 300 nucleotides.
[0082] In other embodiments, the sample nucleic acid molecules can have different starting lengths. The length of the sample nucleic acid molecules, with or without fragmentation, can be any size appropriate for subsequent method steps. For example, the sample nucleic acid molecules can have a starting length of at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 8 ... 900, 1,000, 1,250, 1,500, 1,750, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000 nucleotides and the top 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000 nucleotides. 0, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,250, 1,500, 1750, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000 , 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, and 100,000 nucleotides.
[0083] In some respects, the size range selected for the initial length of the sample nucleic acid segment molecule depends on the binding method.When PCR is used, a longer range of nucleic acid molecule length is selected, because it increases the possibility that two primers will bind to the same nucleic acid molecule.When ligation is used, a shorter range of nucleic acid molecule length is selected, because a shorter range of nucleic acid molecule length will shorten the length of the amplicon produced by PCR in the later steps of the method, especially when PCR is performed using a universal primer that binds to the outside of the nucleic acid segment.Therefore, when ligation is used to bind MIT, the sample nucleic acid molecule will generally be shorter than when PCR is used to bind MIT. For example, in some embodiments, the sample nucleic acid molecule may have a lower end of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, and 1,000 nucleotides and an upper end of 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, and 1,000 nucleotides. 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 9,000, and 10,000 nucleotides, and the MITs are joined by ligation.In certain embodiments, the sample nucleic acid molecule is selected from the lower end of 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2250, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 90 ... 00, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000 nucleotides and the top 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,500, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 11,000 The fragments are between 00, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, and 100,000 nucleotides, and the MITs are joined by PCR.
[0084] In some embodiments, the nucleic acid molecules in the sample can be synthesized using a machine. In some embodiments, the nucleic acid molecules are produced by living cells. In some embodiments, nucleic acid molecules produced by living cells and nucleic acid molecules synthesized using a machine can be combined and used as the sample nucleic acid molecules. This combination can be useful for quantification purposes. In some embodiments, the sample nucleic acid molecules are not amplified in vitro.
[0085] MIT and MIT reaction mixture In the methods provided herein, the step of binding MIT to a sample nucleic acid molecule or nucleic acid segment typically includes forming a reaction mixture. The reaction mixture formed during such a method can itself be a specific aspect of the present disclosure. The reaction mixture provided herein can include a sample nucleic acid molecule as specifically disclosed herein, and can include a set of MITs as specifically disclosed herein, wherein the total number of nucleic acid molecules in the sample is greater than the diversity of MITs in the set of MITs. In some embodiments, the total number of nucleic acid molecules in the sample is also greater than the diversity of possible combinations of bound MITs.
[0086] In some embodiments disclosed herein, the ratio of the total number of sample nucleic acid molecules to the diversity of MITs in the set of MITs, or the diversity of possible combinations of combined MITs using the set of MITs, is at the lower end of the range of 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 1200:1, 1300:1, 1400:1, 1500:1, 1600:1, 1700:1, 1800:1, 1900:1, 2000:1, 2100:1, 2200:1, 2300:1, 2400:1, 2500:1, 2600:1, 2700:1, 2800:1, 2900:1, 3000:1, 3100:1, 3200:1, 3300:1, 3400:1, 3500:1, 3600:1, 3700:1, 3800:1, 3900:1, 4000:1, 4100:1, 4200:1, 4300:1, 4400:1, 4500:1, 4600:1, 4700:1, 4800:1, 4900:1, 5000:1, 5100:1, 5200 500:1, 600:1, 700:1, 800:1, 900:1, 1,000:1, 2,000:1, 3,000:1, 4,000:1, 5,000:1, 6,000:1, 7,000:1, 8,000:1, 9,000:1, 10,000:1, 15,000:1, 20,000:1, 25,000:1, 30,000:1, 50,000:1, 60,000:1, 70,000:1, 80,000:1, 90,000:1 100,000:1, 200,000:1, 300,000:1, 500,000:1, 600,000:1, 700,000:1, 800,000:1, 900,000:1, and 1,000,000:1 at the upper end of the range, and 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800: 1, 900:1, 1,000:1, 2,000:1, 3,000:1, 4,000:1, 5,000:1, 6,000:1, 7,000:1, 8,000:1, 9,000:1, 10,000:1, 15,000:1, 20,000:1, 25,000:1, 30,000:1, 40,000:1, 50,000:1, 60,000:1 0:1, 70,000:1, 80,000:1, 90,000:1, 100,000:1, 200,000:1, 300,000:1, 400,000:1, 500,000:1, 600,000:1, 700,000:1, 800,000:1, 900,000:1, 1,000,000:1, 2,000,000:1, 3, 000,000:1, 4,000,000:1, 5,000,000:1, 6,000,000:1, 7,000,000:1, 8,000,000:1, 9,000,000:1, 10,000,000:1, 50,000,000:1, 100,000,000:1, and 1,000,000,000:1.
[0087] In some embodiments, sample is human cfDNA sample.In such method, as disclosed herein, diversity is about 20 million to about 3 billion.In these embodiments, the ratio of the total number of sample nucleic acid molecules and the diversity of MIT set is 100,000:1, 1x10 at the lower end of the range. 6 :1, 1×10 7 :1, 2 x 10 7 :1 and 2.5 x 10 7 :1 and 2x10 at the top end of the range 7 :1, 2.5x10 7 :1, 5x10 7 :1, 1x10 8 :1, 2.5x10 8 :1, 5x10 8 :1 and 1x10 9 :1 and can be between.
[0088] In some embodiments, the diversity of possible combinations of MIT using MIT set is preferably greater than the total number of sample nucleic acid molecules that span target loci.For example, if there are 100 copies of the human genome, all of which are fragmented into 200bp fragments, so that there are about 15,000,000 fragments for each genome, the diversity of possible combinations of MIT is preferably greater than 100 (the number of copies of each target locus), but less than 1,500,000,000 (the total number of nucleic acid molecules).For example, the diversity of possible combinations of MIT is preferably greater than 100, but much less than 1,500,000,000, such as 200, 300, 400, 500, 600, 700, 800, 900 or 1,000 possible combinations of MIT.The diversity of MIT in MIT set is less than the total number of nucleic acid molecules, but the total number of MIT in reaction mixture exceeds the total number of nucleic acid molecules or nucleic acid molecule segments in reaction mixture. For example, if there are 1,500,000,000 total nucleic acid molecules or nucleic acid molecule segments, there will be more than 1,500,000,000 total MIT molecules in the reaction mixture. In some embodiments, the ratio of MIT diversity in the set of MITs may be lower than the number of nucleic acid molecules in the sample that span the target loci, while the diversity of possible combinations of combined MITs using the set of MITs may be greater than the number of nucleic acid molecules in the sample that span the target loci. For example, the ratio of the number of nucleic acid molecules in a sample spanning a target locus to the diversity of MITs in the set of MITs may be at least 10:1, 25:1, 50:1, 100:1, 125:1, 150:1, or 200:1, and the ratio of the diversity of possible combinations of combining MITs using the set of MITs to the number of nucleic acid molecules in a sample spanning a target locus may be at least 1.01:1, 1.1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 25:1, 50:1, 100:1, 250:1, 500:1, or 1,000:1.
[0089] Typically, the diversity of MITs in a set of MITs is less than the total number of sample nucleic acid molecules that span the target loci, but the diversity of the possible combinations of bound MITs is greater than the total number of sample nucleic acid molecules that span the target loci.In an embodiment in which two MITs are bound to sample nucleic acid molecules, the diversity of MITs in a set of MITs is less than the total number of sample nucleic acid molecules that span the target loci, but greater than the square root of the total number of sample nucleic acid molecules that span the target loci.In some embodiments, the diversity of MITs is less than the total number of sample nucleic acid molecules that span the target loci, but greater than the square root of the total number of sample nucleic acid molecules that span the target loci by 1, 2, 3, 4, or 5.Therefore, the diversity of MITs is less than the total number of sample nucleic acid molecules that span the target loci, but the total number of combinations of any two MITs is greater than the total number of sample nucleic acid molecules that span the target loci.The diversity of MITs in a set is typically less than half the number of sample nucleic acid molecules that span the target loci in a sample that has at least 100 copies of each target locus. In some embodiments, the diversity of MITs in a set can be at least 1, 2, 3, 4, or 5 or more times greater than the square root of the total number of sample nucleic acid molecules spanning the target loci, but can be less than 1 / 5, 1 / 10, 1 / 20, 1 / 50, or 1 / 100 of the total number of sample nucleic acid molecules spanning the target loci. For a sample having 2,000 to 1,000,000 sample nucleic acid molecules spanning the target loci, the number of MITs in the set does not exceed 1,000. For example, for a sample having 10,000 copies of the genome in a genomic DNA sample, such as a circulating cell-free DNA sample, such that the sample has 10,000 sample nucleic acid molecules spanning the target loci, the diversity of MITs can be 10 to 1,000, 10 to 500, or 10 to 250. In some embodiments, the diversity of MITs in the set of MITs is between the square root of the total number of sample nucleic acid molecules spanning the target loci and 1, 10, 25, 50, 100, 125, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, or 1,000 less than the total number of sample nucleic acid molecules spanning the target loci.In some embodiments, the diversity of MITs in a set of MITs may be between 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, and 80% of the number of sample nucleic acid molecules spanning target loci at the lower end of the range and 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, and 99% of the number of sample nucleic acid molecules spanning target loci at the upper end of the range.
[0090] In some embodiments, the ratio of the total number of MITs in the reaction mixture to the total number of sample nucleic acid molecules in the reaction mixture is between 1.01, 1.1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 25:1 50:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1,000:1, 2,000:1, 3,000:1, 4,000:1, 5,000:1, 6,000:1, 7,000:1, 8,000:1, 9,000:1, 10,000:1 at the lower end of the range and 25:1 at the upper end of the range. It can be between 50:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1,000:1, 2,000:1, 3,000:1, 4,000:1, 5,000:1, 6,000:1, 7,000:1, 8,000:1, 9,000:1, 10,000:1, 15,000:1, 20,000:1, 25,000:1, 30,000:1, 40,000:1, and 50,000:1. In some embodiments, the total number of MITs in the reaction mixture is at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% of the total number of sample nucleic acid molecules in the reaction mixture. In other embodiments, the ratio of the total number of MITs in the reaction mixture to the total number of sample nucleic acid molecules in the reaction mixture can be at least sufficient for each sample nucleic acid molecule to have an appropriate number of bound MITs, i.e., 2:1 when 2 types of MITs are bound, 3:1 when 3 types of MITs, 4:1 when 4 types of MITs, 5:1 when 5 types of MITs, 6:1 when 6 types of MITs, 7:1 when 7 types of MITs, 8:1 when 8 types of MITs, 9:1 when 9 types of MITs, and 10:1 when 10 types of MITs.
[0091] In some embodiments, the ratio of the total number of MITs having identical sequences in the reaction mixture to the total number of nucleic acid segments in the reaction mixture is at the lower end of the range, 0.1:1, 0.2:1, 0.3:1, 0.4:1, 0.5:1, 0.6:1, 0.7:1, 0.8:1, 0.9:1, 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.25:1, 2.5:1, 2.75:1, 3:1, 3.5:1, 4:1, 4.5:1, and 5:1. 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.25:1, 2.5:1, 2.75:1, 3:1, 3.5:1, 4:1, 4.5:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, and 100:1 at the upper end of the range.
[0092] The set of MITs can include, for example, at least three MITs or 10 to 500 MITs. In some embodiments, as discussed herein, nucleic acid molecules derived from a sample are added directly to the binding reaction mixture without amplification. These sample nucleic acid molecules can be purified from a source, such as a living cell or organism, as disclosed herein, and then MITs can be bound without amplifying the nucleic acid molecules. In some embodiments, the sample nucleic acid molecules or nucleic acid segments can be amplified before MITs are bound. As discussed herein, in some embodiments, nucleic acid molecules derived from a sample can be fragmented to generate sample nucleic acid segments. In some embodiments, other oligonucleotide sequences can be bound (e.g., ligated) to the ends of the sample nucleic acid molecules before MITs are bound.
[0093] In some embodiments disclosed herein, the ratio of sample nucleic acid molecules, nucleic acid segments, or fragments comprising the target locus to MIT in the reaction mixture is at the lower end of the range, 1.01:1, 1.05, 1.1:1, 1.2:1, or 1.3:1. 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 25:1, 30:1, 35:1, 40:1, 45:1, and 50:1 at the upper end of the range. The ratio can be between 60:1, 70:1, 80:1, 90:1, 100:1, 125:1, 150:1, 175:1, 200:1, 300:1, 400:1, and 500:1. For example, in some embodiments, the ratio of sample nucleic acid molecules, nucleic acid segments, or fragments having a particular target locus to MIT in the reaction mixture is between the lower end of 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 25:1, 30:1, 35:1, 40:1, 45:1, and 50:1 and the upper end of 20:1, 25:1, 30:1, 35:1, 40:1, 45:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, and 200:1. In some embodiments, the ratio of sample nucleic acid molecules or nucleic acid segments to MIT in the reaction mixture can be between a lower limit of 25:1, 30:1, 35:1, 40:1, 45:1, or 50:1 and an upper limit of 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. In some embodiments, the diversity of possible combinations of bound MIT can be greater than the number of sample nucleic acid molecules, nucleic acid segments, or fragments spanning the target loci. For example, in some embodiments, the ratio of the diversity of possible combinations of bound MIT to the number of sample nucleic acid molecules, nucleic acid segments, or fragments spanning the target loci can be at least 1.01, 1.1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 25:1, 50:1, 100:1, 250:1, 500:1, or 1,000:1.
[0094] As provided herein, a reaction mixture for tagging nucleic acid molecules with MITs (i.e., binding nucleic acid molecules to MITs) can contain additional reagents in addition to a population of sample nucleic acid molecules and a set of MITs. For example, a reaction mixture for tagging can include a ligase or polymerase containing an appropriate buffer at an appropriate pH, adenosine triphosphate (ATP) for ATP-dependent ligases, nicotinamide adenine dinucleotide for NAD-dependent ligases, deoxynucleoside triphosphates (dNTPs) for polymerases, and optionally a molecular crowding agent such as polyethylene glycol. In certain embodiments, the reaction mixture can include a population of sample nucleic acid molecules, a set of MITs, and a polymerase or ligase, wherein the ratio of the number of sample nucleic acid molecules, nucleic acid segments, or fragments having a particular target locus to the number of MITs in the reaction mixture can be any of the ratios disclosed herein, for example, 2:1 to 100:1, or 10:1 to 100:1, or 25:1 to 75:1, or 40:1 to 60:1, or 45:1 to 55:1, or 49:1 to 51:1.
[0095] In some embodiments disclosed herein, the number of different MITs in the set of MITs (i.e., diversity) can be, at the low end, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,0 ...10,000, 10,000, 10,000, 10,000, 10,000, 10,000, 10,000, 10,000, 10,000 The range can be between 1,000, 2,500, and 3,000 MITs and, at the upper end, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, and 5,000 MITs with different sequences. For example, the diversity of different MITs in the set of MITs can be between 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, and 100 different MIT sequences on the low end and 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, and 300 different MIT sequences on the high end. In some embodiments, the diversity of different MITs in the set of MITs can be between 50, 60, 70, 80, 90, 100, 125, and 150 different MIT sequences on the low end and 100, 125, 150, 175, 200, and 250 different MIT sequences on the high end. In some embodiments, the diversity of different MITs in the set of MITs can be between 3 and 1,000, or between 10 and 500, or between 50 and 250 different MIT sequences.In some embodiments, the diversity of possible combinations of binding MITs using a set of MITs is 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 150, 200, 250, 300, 400, 500, and Possible combinations are 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 250,000, 500,000, and 1,000,000, with combined MIT at the upper end of the range: 10, 15, 20, 25, 30, 40, 50, 75, 100, 150, 200, 250, 300, 400, 500, 1,000, 2,000, 3,000, The number of possible combinations may be between 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 250,000, 500,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, and 10,000,000 possible combinations.
[0096] The MITs in a set of MITs are typically all the same length.For example, in some embodiments, MITs can be between 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 and 20 nucleotides at the bottom and 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 and 30 nucleotides at the top.In certain embodiments, MITs can be any length from 3, 4, 5, 6, 7 or 8 nucleotides at the bottom to 5, 6, 7, 8, 9, 10 or 11 nucleotides at the top.In some embodiments, the length of MITs can be any length from 4, 5 or 6 nucleotides at the bottom to 5, 6 or 7 nucleotides at the top. In some embodiments, the length of the MIT is 5, 6, or 7 nucleotides.
[0097] As will be understood, a set of MITs typically includes many identical copies of each MIT member of the set. In some embodiments, the set of MITs includes from 10, 20, 25, 30, 40, 50, 100, 500, 1,000, 10,000, 50,000, and 100,000 times more copies than the total number of sample nucleic acid molecules spanning the target locus, at the lower end of the range, to 100, 500, 1,000, 10,000, 50,000, 100,000, 250,000, 500,000, and 1,000,000 times more copies than the total number of sample nucleic acid molecules spanning the target locus, at the upper end of the range. For example, in a human circulating cell-free DNA sample isolated from plasma, there may be an amount of DNA fragments, including, for example, 1,000 to 100,000 circulating fragments spanning any target locus in the genome. In certain embodiments, there are no more than 1 / 10, 1 / 4, 1 / 2, or 3 / 4 copies of a given MIT among all unique MITs in the set. Between members of the set, there may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 differences between any sequence and the remaining sequences. In some embodiments, the sequence of each MIT in the set differs from all other MITs by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. To reduce the possibility of misidentifying MITs, the set of MITs can be designed using methods that will be recognized by those skilled in the art, such as taking into account the Hamming distance between all MITs in the set. The Hamming distance measures the minimum number of substitutions required to change one string or nucleotide sequence into another. Here, the Hamming distance measures the minimum number of amplification errors required to convert one MIT sequence in a set into another MIT sequence from the same set. In certain embodiments, different MITs in the set of MITs have a Hamming distance of less than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 between each other.
[0098] In certain embodiments, a set of isolated MITs as provided herein is one embodiment of the present disclosure. The set of isolated MITs may be a set of single-stranded, or partially or completely double-stranded, nucleic acid molecules, with each MIT being part or all of a nucleic acid molecule in the set. In a particular example, a set of Y adaptor (i.e., partially double-stranded) nucleic acids, each containing a different MIT, is provided herein. The set of Y adaptor nucleic acids can be identical to each other except for the MIT portion. Multiple copies of the same Y adaptor MIT can be included in the set. The set can have the number and diversity of nucleic acid molecules as disclosed herein for the set of MITs. As a non-limiting example, the set can include 2, 5, 10, or 100 copies of 50 to 500 MIT-containing Y adaptors, each MIT segment being 4 to 8 nucleic acids in length, and each MIT segment differing from other MIT segments by at least two nucleotides but containing identical sequences other than the MIT sequence. Further details regarding the Y adaptor portion of the set of Y adaptors are provided herein.
[0099] In other embodiments, a reaction mixture comprising a set of MITs and a population of sample nucleic acid molecules is an embodiment of the present disclosure. Furthermore, such compositions can be part of many of the methods and other compositions provided herein. For example, in further embodiments, the reaction mixture can comprise a polymerase or ligase, an appropriate buffer, and auxiliary components as discussed in more detail herein. For any of these embodiments, the set of MITs can range from 25, 50, 100, 200, 250, 300, 400, 500, or 1,000 MITs at the lower end of the range to 100, 200, 250, 300, 400, 500, 1,000, 1,500, 2,000, 2,500, 5,000, 10,000, or 25,000 MITs at the upper end of the range. For example, in some embodiments, the reaction mixture comprises a set of 10 to 500 MITs.
[0100] MIT Combine The molecular beacon tags (MITs) discussed in more detail herein can be attached to sample nucleic acid molecules in a reaction mixture using methods recognized by those skilled in the art. In some embodiments, the MITs can be attached alone, i.e., without additional oligonucleotide sequences. In some embodiments, the MITs can be part of a larger oligonucleotide that can further include other nucleotide sequences, as discussed in more detail herein. For example, the oligonucleotides can also include primers specific to the nucleic acid segment or universal primer binding sites, sequencing adapters such as Y adapters, library tags, adapters such as ligated adapter tags, and combinations thereof. Those skilled in the art will recognize how to incorporate various tags into oligonucleotides to generate tagged nucleic acid molecules useful for sequencing, particularly high-throughput sequencing. The MITs of the present disclosure are advantageous in that, due to the small diversity of nucleic acid molecules, they can be more easily used with additional sequences, such as Y adapters and / or universal sequences, and therefore can be more easily combined with additional sequences on the adapters to create smaller, and therefore more cost-effective, sets of MIT-containing adapters.
[0101] In some embodiments, MITs are linked in tagged nucleic acid molecules such that one MIT is 5' to the sample nucleic acid segment and one MIT is 3' to the sample nucleic acid segment. For example, in some embodiments, MITs can be directly linked to the 5' and 3' ends of the sample nucleic acid molecule using ligation. In some embodiments disclosed herein, ligation typically involves forming a reaction mixture with an appropriate buffer, ions, and appropriate pH, in which a population of sample nucleic acid molecules, a set of MITs, adenosine triphosphate, and a ligase are combined. Those skilled in the art will understand how to form a reaction mixture and the various ligases available for use. In some embodiments, the nucleic acid molecule can have a 3' adenosine overhang, and the MIT can be located on a double-stranded oligonucleotide with a 5' thymidine overhang, for example, directly adjacent to the 5' thymidine.
[0102] In further embodiments, the MITs provided herein can be included as part of a Y adaptor before they are ligated to a sample nucleic acid molecule. Y adaptors are known in the art and are used, for example, to provide more efficient primer binding sequences at the two ends of a nucleic acid molecule prior to high-throughput sequencing. A Y adaptor is formed by annealing a first oligonucleotide and a second oligonucleotide, wherein the 5' segment of the first oligonucleotide and the 3' segment of the second oligonucleotide are complementary, and the 3' segment of the first oligonucleotide and the 5' segment of the second oligonucleotide are not complementary. In some embodiments, the Y adaptor comprises a base-paired double-stranded polynucleotide segment and a non-base-paired single-stranded polynucleotide segment distal to the ligation site. The double-stranded polynucleotide segment can be between 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length at the lower end of the range, and 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 nucleotides in length at the upper end of the range. The single-stranded polynucleotide segments on the first and second oligonucleotides can be between 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length at the lower end of the range and 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, and 30 nucleotides in length at the upper end of the range. In these embodiments, the MIT is typically a double-stranded sequence added to the end of a Y adaptor, which is ligated to the sample nucleic acid segment to be sequenced. An exemplary Y adaptor is shown in Figure 1. In some embodiments, the non-complementary segments of the first and second oligonucleotides can be of different lengths.
[0103] In some embodiments, the double-stranded MIT linked by ligation will have the same MIT in both strands of the sample nucleic acid molecule.At some point, the tagged nucleic acid molecules obtained from these two strands will be identified and used to generate paired MIT families.In the downstream sequencing reaction, in which single-stranded nucleic acids are typically sequenced, MIT families can be identified by identifying tagged nucleic acid molecules with identical or complementary MIT sequences.In these embodiments, paired MIT families can be used to confirm the existence of sequence differences in the initial sample nucleic acid molecules as discussed herein.
[0104] As shown in Figure 2, in some embodiments, MIT can be bound to a sample nucleic acid segment by being incorporated into the 5' of a forward and / or reverse PCR primer that binds to a sequence in the sample nucleic acid segment. In some embodiments, MIT can be incorporated into a universal forward and / or reverse PCR primer that binds to a universal primer binding sequence previously bound to the sample nucleic acid molecule. In some embodiments, MIT can be bound using a combination of a universal forward or reverse primer having a 5' MIT sequence and a forward or reverse PCR primer that binds to an internal binding sequence in the sample nucleic acid segment that has a 5' MIT sequence. After two cycles of PCR, the sample nucleic acid molecule amplified using both a forward primer and a reverse primer with an incorporated MIT sequence has MIT bound to the 5' of the sample nucleic acid segment and to the 3' of the sample nucleic acid segment in each tagged nucleic acid molecule. In some embodiments, PCR is performed for 2, 3, 4, 5, 6, 7, 8, 9, or 10 cycles in the binding step.
[0105] In some embodiments disclosed herein, two MITs on each tagging nucleic acid molecule can be attached using similar techniques, so that both MITs are 5' to the sample nucleic acid segment, or so that both MITs are 3' to the sample nucleic acid segment. For example, two MITs can be incorporated into the same oligonucleotide and linked to one end of the sample nucleic acid molecule, or two MITs can be present on the forward or reverse primer, and the paired reverse or forward primer can have zero MIT. In other embodiments, any combination of MITs attached to the 5' and / or 3' positions relative to the nucleic acid segment can be attached, and three or more MITs can be attached.
[0106] As discussed herein, other sequences can be attached to the sample nucleic acid molecule before, after, during, or together with the MIT. For example, a ligated adapter, often referred to as a library tag or ligated adapter tag (LT), can be added with or without a universal primer binding sequence for use in the subsequent universal amplification step. In some embodiments, the length of the oligonucleotide comprising the MIT and other sequences is at the lower end of the range of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 29, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, and 100 nucleotides. and 10, 11, 12, 13, 14, 15, 16, 17, 18, 29, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, and 200 nucleotides at the upper end of the range. In some respects, the number of nucleotides in the MIT sequence may be a percentage of the number of nucleotides in the entire sequence of the oligonucleotide that includes the MIT. For example, in some embodiments, the MIT can be at most 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the total nucleotides of the oligonucleotide linked to the sample nucleic acid molecule.
[0107] After MIT is attached to the sample nucleic acid molecule by ligation or PCR reaction, it may be necessary to clean up the reaction mixture to remove unwanted components that may affect subsequent method steps. In some embodiments, the sample nucleic acid molecule can be purified from primers or ligase. In other embodiments, proteins and primers can be digested with proteases and exonucleases using methods known in the art.
[0108] After MIT is bound to the sample nucleic acid molecule, a population of tagged nucleic acid molecules is generated, which itself forms an embodiment of the present disclosure.In some embodiments, the size range of tagged nucleic acid molecules can be between 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400 and 500 nucleotides at the lower end of the range and 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 2,000, 3,000, 4,000 and 5,000 nucleotides at the upper end of the range.
[0109] Such populations of tagged nucleic acid molecules may be at the lower end of the range: 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9, 000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000 00, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 10,000,000, 20,000,000, 30,000,000, 40,00,000, 50,000,00 The ranges are from 0, 50,000,000, 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, and 1,000,000,000 tagged nucleic acid molecules at the upper end of the range to 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4 ,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, 10,000,000, 20,000,000, 30,000,000, 40,00,000, 50,000,000, 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900 ,000,000, 1,000,000,000, 2,000,000,000, 3,000,000,000, 4,000,000,000, 5,000,000,000, 6,000,000,000, 7,000,000,000, 8,000,000,000, 9,000,000,000, and up to 10,000,000,000 tagged nucleic acid molecules. In some embodiments, the population of tagged nucleic acid molecules comprises at the lower end of the range 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, and 1,000,000,000 tagged nucleic acids. The number of tagged nucleic acid molecules can range from 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, 1,000,000,000, 2,000,000,000, 3,000,000,000, 4,000,000,000, and 5,000,000,000 tagged nucleic acid molecules.
[0110] In some respects, the goal can be to have a certain percentage of all sample nucleic acid molecules in a population of sample nucleic acid molecules bind MIT. In some embodiments, the goal can be to have at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% of the sample nucleic acid molecules bind MIT. In other respects, the goal can be to have a certain percentage of the sample nucleic acid molecules in the population successfully bind MIT. In any of the embodiments disclosed herein, at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% of the sample nucleic acid molecules can have successfully bound MITs to generate a population of tagged nucleic acid molecules. In any of the embodiments disclosed herein, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 200, 300, 500 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 sample nucleic acid molecules can successfully bind MIT to generate a population of tagged nucleic acid molecules.
[0111] In some embodiments disclosed herein, MIT can be an oligonucleotide sequence of ribonucleotides or deoxyribonucleotides linked via phosphodiester bonds. The nucleotides disclosed herein can refer to both ribonucleotides and deoxyribonucleotides, and those skilled in the art will recognize which form is relevant for a particular application. In certain embodiments, the nucleotide can be selected from the group of natural nucleotides consisting of adenosine, cytidine, guanosine, uridine, 5-methyluridine, deoxyadenosine, deoxycytidine, deoxyguanosine, deoxythymidine, and deoxyuridine. In some embodiments, MIT can be a non-natural nucleotide. Non-natural nucleotides can include: sets of nucleotides that bind to each other, such as d5SICS and dNaM; metal-coordinating bases, such as 2,6-bis(ethylthiomethyl)pyridine (SPy) with a silver ion and monodentate pyridine (Py) with a copper ion; universal bases that can pair with two or more or any other base, such as 2'-deoxyinosine derivatives, nitroazole analogs, and hydrophobic aromatic non-hydrogen-bonding bases; and xDNA nucleobases with extended bases. In certain embodiments, the oligonucleotide sequence can be predetermined, while in other embodiments, the oligonucleotide sequence can be degenerate.
[0112] In some embodiments, MITs contain phosphodiester bonds between the natural sugars ribose and / or deoxyribose attached to the nucleobases. In some embodiments, non-natural linkages can be used. These linkages include, for example, phosphorothioate, boranophosphate, phosphonate, and triazole linkages. In some embodiments, combinations of non-natural linkages and / or phosphodiester linkages can be used. In some embodiments, peptide nucleic acids can be used, in which the sugar backbone is instead made from repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In any of the embodiments disclosed herein, non-natural sugars can be used in place of ribose or deoxyribose sugars. For example, threose can be used to generate α-(L)-threofuranosyl-(3'-2') nucleic acid (TNA). Other linkage types and sugars will be apparent to those skilled in the art and can be used in any of the embodiments disclosed herein.
[0113] In some embodiments, nucleotides with extra bonds between sugar atoms can be used. For example, bridged or locked nucleic acids can be used in MIT. These nucleic acids contain a bond between the 2' and 4' positions of the ribose sugar.
[0114] In certain embodiments, reactive linker can be added to the nucleotide incorporated in the sequence of MIT.Afterwards, reactive linker can be mixed with the molecule that is appropriately tagged under the appropriate conditions for reaction to occur.For example, aminoallyl nucleotide can be added, which can react with the molecule that is linked to reactive leaving group such as succinimidyl ester, and thiol-containing nucleotide can be added, which can react with the molecule that is linked to reactive leaving group such as maleimide.In other embodiments, biotin-linked nucleotide can be used in the sequence of MIT, which can be linked to streptavidin-tagged molecule.
[0115] Various combinations of natural nucleotides, non-natural nucleotides, phosphodiester linkages, non-natural linkages, natural sugars, non-natural sugars, peptide nucleic acids, bridged nucleic acids, locked nucleic acids, and nucleotides with reactive linkers will be recognized by those of skill in the art and can be used to form MITs in any of the embodiments disclosed herein.
[0116] Amplification of tagged nucleic acid molecules In some embodiments, the methods of the present disclosure include amplifying the tagged nucleic acid molecules before determining their sequences. Typically, multiple rounds of amplification are performed during sample preparation for high-throughput sequencing, as is known in the art. These amplification steps generally occur after MIT binds to the nucleic acid molecule, although amplification of the sample nucleic acid molecule may occur before MIT binding in some embodiments. In certain embodiments, at least one, two, three, four, five, or six amplification reactions are performed after MIT binds to the sample nucleic acid segment of the sample nucleic acid molecule. In high-throughput sequencing methods, for example, amplification reactions may include amplifying the initial nucleic acid in the sample to generate a library to be sequenced, typically clonal amplification of the library on a solid support, and adding additional information or functionality, such as a sample identification barcode, through additional amplification reactions. As described below, barcodes can be added at any time during the amplification process and before and / or after target enrichment. Tagged sample nucleic acid molecules can have one or more barcodes at one or both ends. Each amplification reaction typically involves multiple cycles of amplification (e.g., from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles on the lower end of the range to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 75, or 100 cycles on the higher end of the range), either by thermal cycling or by natural biochemical reaction cycling as occurs during isothermal amplification. In some examples, any of the methods of embodiments provided herein can include an amplification step in which at least 10, 15, 20, 25, or 30 cycles of amplification (e.g., thermal cycling in PCR amplification) are performed.
[0117] In some embodiments, after binding the MIT, the tagged nucleic acid molecules can be amplified using a universal primer that binds to the pre-bound universal amplification primer binding sequence to generate a library of sample nucleic acid molecules. Specific target nucleic acids in the library of nucleic acid molecules can be enriched, for example, through multiplex PCR, particularly one-sided PCR, or through hybrid capture. The enrichment step can be followed by another universal amplification reaction. Any barcode amplification reaction, regardless of whether or not there is a targeted amplification step, can be used to barcode tagged nucleic acid molecules generated from sample nucleic acid molecules from separate samples or subpools, and then pool the products from multiple reaction mixtures or subpools. As is known, such barcodes allow the identification of the sample from which the tagged nucleic acid molecules were generated. This can be used to identify multiple starting samples, and can be useful when dividing the sample nucleic acid molecules after labeling to increase the total number of tag combinations. Such barcodes differ from the MIT of the present disclosure because they do not identify individual sample nucleic acid molecules, but rather identify the sample from the nucleic acid molecules generated in the sample mixture. The tagged nucleic acid molecules or amplified tagged nucleic acid molecules are typically templated on one or more solid supports and may be clonal amplified or clonal amplified during a template amplification reaction. It is important to note that amplification errors may be introduced at any amplification step during the process. Using the methods disclosed herein, it is possible to determine which amplification step an error occurs in, or whether an error occurs during a subsequent sequencing reaction. For example, if a sample is split into multiple PCRs, each adding a new, different MIT, it is possible to determine whether an error occurred during a particular PCR step.
[0118] In some embodiments, the sample nucleic acid molecule is unchanged before MIT is bound; after MIT is bound, the tagged nucleic acid molecule is amplified using universal primers to create a library or population of tagged nucleic acid molecules; the amplified library of tagged nucleic acid molecules undergoes target enrichment via multiplex PCR (e.g., one-sided multiplex PCR); the enriched tagged nucleic acid molecules undergo an optional barcode amplification step; clonal amplification onto one or more solid supports is performed; the sequence of the tagged nucleic acid molecule is determined; and the sample nucleic acid molecule is identified using the determined sequence of the bound MIT.
[0119] In any of the embodiments disclosed herein, these amplification steps can be performed using methods well known in the art, such as isothermal amplification, such as PCR amplification using thermal cycling or recombinase polymerase amplification. For any of the amplification steps disclosed herein, one of skill in the art will understand how to adapt the method for isothermal amplification.
[0120] In some embodiments, the tagged nucleic acid molecules can be used to generate libraries for sequencing, particularly high-throughput sequencing. Typically, the tagged nucleic acid molecules are amplified using a universal primer that binds to a universal primer binding sequence incorporated into the tagged nucleic acid molecule, as discussed elsewhere herein. In some embodiments, the universal amplification is performed for multiple cycles, for example, between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 cycles at the lower end of the range and 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50 cycles at the upper end of the range. In some embodiments, the amplification is performed such that each of the tagged nucleic acid molecules is copied to a number of copies at the lower end of the range, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000 , 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 10,000,000, 20,000,000, 30,000,000, 40,000,000, and 50 ,000,000 copies at the upper end of the range, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7, ,000,000, 8,000,000, 9,000,000, 10,000,000, 20,000,000, 30,000,000, 40,00,000, 50,000,000, 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, and up to 1,000,000,000 copies.
[0121] Target enrichment In certain embodiments, the method of the present disclosure can include a target enrichment step before determining the sequence of the sample nucleic acid molecule. In some embodiments, target enrichment is performed using a multiplex PCR reaction, particularly a one-sided PCR reaction. In these embodiments, a universal primer and multiple target-specific primers that bind to internal sequences of the target sample nucleic acid segment are used to generate amplicons from tagged nucleic acid molecules using both the universal primer binding sequence and the target-specific sequence, but no amplicons are generated from tagged nucleic acid molecules that lack either or both of these sequences. In some embodiments, the universal primer can bind to the 5' universal primer binding site on one strand of DNA, and the target-specific primer can bind to the complement of the DNA strand within the nucleic acid segment 3' to the universal primer binding site on the other strand of complementary DNA. The binding direction can be reversed, with the universal primer binding to the 3' universal primer binding site on one strand and the target-specific primer binding to the complement of the DNA strand within the nucleic acid segment 5' to the universal primer binding site on the other strand of complementary DNA.
[0122] In some embodiments of the present disclosure, preferentially enriching DNA comprises obtaining a plurality of hybrid capture probes targeting a desired sequence, hybridizing the hybrid capture probes to DNA in a sample, and physically removing some or all of the unhybridized DNA from the DNA sample. Thus, a sequence complementary to the targeting tagged nucleic acid molecule is attached to a solid support, and tagged nucleic acid molecules are added under conditions such that the targeting tagged nucleic acid molecule anneals to the complementary sequence and the non-targeting tagged nucleic acid molecule does not anneal. After removing the non-targeting tagged nucleic acid molecule, reaction conditions can be adjusted to allow the targeting tagged nucleic acid molecule to be dissociated and isolated from the solid support. In some embodiments, an amplification step can be performed after hybrid capture using universal amplification primers.
[0123] A hybrid capture probe refers to any nucleic acid sequence, possibly modified, generated by various methods, such as PCR or direct synthesis, intended to be complementary to one strand of a specific target DNA sequence in a sample. An exogenous hybrid capture probe can be added to a prepared sample and hybridized through a denaturation-reannealing process to form an exogenous-endogenous fragment duplex. These duplexes can then be physically separated from the sample by various means. Hybrid capture probes were originally developed to target and enrich large portions of the genome using relative uniformity between targets. In this application, it was important that all targets were amplified with sufficient uniformity so that all target loci could be detected by sequencing. However, no consideration was given to maintaining the allele ratios in the original sample. After capture, the alleles present in the sample can be determined by direct sequencing of the captured molecules. These sequencing reads can be analyzed and counted according to allele type.
[0124] As discussed herein, in some embodiments, the methods of the present disclosure include single-sided multiplex PCR. Such methods can use tagged nucleic acid molecules with one or more adapters at one or more ends. Single-sided PCR can be performed in two stages. For example, a first single-sided PCR can be performed on the targeted tagged nucleic acid molecules using multiple forward primers specific to each targeted tagged nucleic acid molecule and a reverse primer that binds to a universal primer binding site present on the ligated adapters on all tagged nucleic acid molecules. A second single-sided PCR can then be performed on the products of the first single-sided PCR using multiple forward primers specific to each targeted tagged nucleic acid molecule and a reverse primer that binds to the same or a different universal primer binding site from the universal primer binding site used in the initial single-sided PCR reaction.
[0125] In some embodiments, tagged nucleic acid molecules undergo templated on one or more solid supports through clonal amplification in one or two reactions.Methods for performing templated and / or clonal amplification are known in the art and depend on the sequencing method used for analysis.Those skilled in the art will recognize the method used to perform clonal amplification.
[0126] Amplification reaction mixture In some embodiments, amplifying nucleic acid molecules can include forming an amplification reaction mixture. Amplification reaction mixtures useful in the present disclosure can include components known in the art, particularly for PCR amplification. For example, the reaction mixture typically includes a source of nucleotides, such as nucleotide triphosphates, a polymerase, magnesium, and primers, and optionally one or more tagged nucleic acid molecules. In certain embodiments, the reaction mixture is formed by combining a polymerase, nucleotide triphosphates, tagged nucleic acid molecules, and a set of forward and / or reverse primers. Thus, in certain embodiments, a reaction mixture is provided herein that includes a population of tagged nucleic acid molecules and a pool of primers, at least some of which bind to tagged nucleic acid molecules within the population of tagged nucleic acid molecules. In addition to the MIT sequence, the tagged nucleic acid molecules can include adapter sequences for binding primers, for example, for sequencing reactions and / or universal amplification reactions. In some embodiments, the forward and reverse primers for amplifying tagged nucleic acid sequences can be designed to bind to universal primer binding sequences attached to the tagged nucleic acid molecules so that all tagged nucleic acid sequences are amplified. In some embodiments, the forward and reverse primers can be designed so that one binds to a universal primer binding sequence and the other binds to a target-specific sequence within the sample nucleic acid segment, e.g., in one-sided PCR. In other embodiments, the forward and reverse primers can both be designed to bind to a target-specific sequence within the sequence of the sample nucleic acid segment, e.g., in two-sided PCR.
[0127] In any of the embodiments disclosed herein, the reaction mixture may be at or near the lower end of the range: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 410, 420, 430, 440, 450, 460, 470, 480, 490, 510, 520, 530, 540, 550, 560, 570, 580, 590, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 810, 820, 830, 840, 8 60, 170, 180, 190, 200, 225, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 10,000 ,000, 20,000,000, 30,000,000, 40,00,000, 50,000,000, 100,000,000, 200,000,000, 300,000 ,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 800,000,000, 900,000,000, and From 1,000,000,000 tagged nucleic acid molecules at the upper end of the range, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 2,000,000, 2,500,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, 10,000,000, 20,000,000, 30,000,000, 40,00,000, 50,000,000, 100,000,000, 200,000,000, 300,000,000, 400,000,000, 500,000,000, 600,000,000, 700,000,000, 80 The number of tagged nucleic acid molecules can include up to 0,000,000, 900,000,000, 1,000,000,000, 2,000,000,000, 3,000,000,000, 4,000,000,000, 5,000,000,000, 6,000,000,000, 7,000,000,000, 8,000,000,000, 9,000,000,000, and 10,000,000,000 tagged nucleic acid molecules. In some embodiments, the reaction mixture may be at the lower end of the range of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, and It can include 10,000 copies of each tagged nucleic acid molecule up to the upper end of the range of 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, and 100,000 copies of each tagged nucleic acid molecule.
[0128] In any of the embodiments disclosed herein, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% of the tagged nucleic acid molecules are successfully amplified, where successful amplification is defined as PCR having an efficiency of at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100%.
[0129] In further embodiments, the reaction mixture can include a population of 100 to 1,000,000 tagged nucleic acid molecules, each 50 to 500 nucleotides in length, having 10 to 100,000 different sample nucleic acid segments, and a set of 10 to 500 MITs, each 4 to 20 nucleotides in length, wherein the ratio of the number of sample nucleic acid segments to the number of MITs in the population is 2:1 to 100:1. In certain embodiments, each member of the set of MITs is bound to at least one tagged nucleic acid molecule of the population. In certain embodiments, at least two tagged nucleic acid molecules of the population contain at least one identical MIT and a sample nucleic acid segment that differs by more than 50%. In some embodiments, the reaction mixture can include a polymerase or ligase.
[0130] In some embodiments, the reaction mixture contains 25, 50, 100, 200, 250, 300, 400, 500, 1,000, 2,500, 5,000, 10,000, 20,000, 25,000, or 50,000 primers or primer pairs at the lower end of the range to 200, 250, 300, 400, 500, 1,000, 2,500, 5,000, 10,000, 20,000, 25,000, 50,000, 6 The set, library, or pool of primers can include up to 0,000, 70,000, 80,000, 90,000, 100,000, 125,000, 150,000, 200,000, 250,000, 300,000, 4,000, or 500,000 primers or primer pairs, each binding to a primer binding sequence located in one or more of the plurality of tagged nucleic acid molecules.
[0131] In some embodiments, a library of nucleic acid molecules useful for sequencing is formed. In some embodiments, the library comprises from 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, and 1,000 copies of each tagged nucleic acid molecule at the lower end of the range, to 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, and 1,000 copies of each tagged nucleic acid molecule at the upper end of the range. It can include up to 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, and 10,000 copies.
[0132] In some embodiments, the library of nucleic acid molecules can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, and 1,000 tagged nucleic acid molecules, including sample nucleic acid segments having an identical first MIT attached to the 5' end of the nucleic acid segment and an identical second MIT attached to the 3' end of the nucleic acid segment, with at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotide difference.
[0133] In some embodiments, a library of nucleic acid molecules can comprise multiple clonal populations of each tagged nucleic acid molecule on one solid support or multiple solid supports.
[0134] In some embodiments, the amplification reaction mixture herein contains a polymerase with proofreading activity, a polymerase without (or negligible) proofreading activity, or a mixture of a polymerase with proofreading activity and a polymerase without (or negligible) proofreading activity. In some embodiments, a hot-start polymerase, a non-hot-start polymerase, or a mixture of a hot-start polymerase and a non-hot-start polymerase is used. In some embodiments, HotStar Taq DNA polymerase is used (see, e.g., Qiagen, Hilden, Germany). In some embodiments, AmpliTaq Gold® DNA polymerase is used (Thermo Fisher, Carlsbad, CA). In some embodiments, PrimeSTAR GXL DNA polymerase, a high-fidelity polymerase that provides efficient PCR amplification when there is an excess of template in the reaction mixture and when amplifying long products, is used (Takara Clontech, Mountain View, CA). In some embodiments, KAPA Taq DNA polymerase or KAPA Taq HotStart DNA polymerase is used; these are based on the single-subunit wild-type Taq DNA polymerase from the thermophilic bacterium Thermus aquaticus and have 5'-3' polymerase and 5'-3' exonuclease activity, but lack 3'-5' exonuclease (proofreading) activity (Kapa Biosystems, Wilmington, MA). In some embodiments, Pfu DNA polymerase is used; this is a highly thermostable DNA polymerase from the hyperthermophilic archaeon Pyrococcus furiosus. Pfu catalyzes the template-dependent polymerization of nucleotides into double-stranded DNA in the 5'->3' direction and also exhibits 3'->5' exonuclease (proofreading) activity, which allows the polymerase to correct nucleotide incorporation errors.It has no 5' to 3' exonuclease activity (Thermo Fisher Scientific, Waltham, MA). In some embodiments, Klentaq 1 is used. It is a Klenow fragment analog of Taq DNA polymerase that has no exonuclease or endonuclease activity (DNA Polymerase Technology, St. Louis, MO). In some embodiments, the polymerase is a Phusion DNA polymerase, such as Phusion High-Fidelity DNA polymerase or Phusion Hot Start Flex DNA polymerase (New England BioLabs, Ipswich, MA). In some embodiments, the polymerase is a Q5® DNA polymerase, such as Q5® High-Fidelity DNA polymerase or Q5® Hot Start High-Fidelity DNA polymerase (New England BioLabs). In some embodiments, the polymerase is T4 DNA polymerase (New England BioLabs).
[0135] In some embodiments, 5 to 600 units / mL (units per mL reaction volume), such as 5 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, or 500 to 600 units / mL (inclusive), of polymerase is used.
[0136] PCR method In some embodiments, hot-start PCR is used to reduce or prevent polymerization before PCR thermal cycling. An exemplary hot-start PCR method involves initial inhibition of DNA polymerase or physical separation of reaction components until the reaction mixture reaches a higher temperature. In some embodiments, gradual release of magnesium is used. DNA polymerase requires magnesium ions for activity, so magnesium is chemically separated from the reaction by binding to a compound and released into solution only at elevated temperatures. In some embodiments, non-covalent binding of an inhibitor is used. In this method, a peptide, antibody, or aptamer can be non-covalently attached to the enzyme at low temperatures to inhibit its activity. After incubation at elevated temperatures, the inhibitor is released and the reaction begins. In some embodiments, a cold-sensitive Taq polymerase, such as a modified DNA polymerase that shows little activity at low temperatures, is used. In some embodiments, chemical modification is used. In this method, a molecule is covalently attached to the side chain of an amino acid in the active site of the DNA polymerase. The molecule is released from the enzyme by incubating the reaction mixture at elevated temperatures. Once the molecule is released, the enzyme is activated.
[0137] In some embodiments, the amount of template nucleic acid (such as an RNA or DNA sample) is 20 to 5,000 ng, for example, 20 to 200 ng, 200 to 400, 400 to 600, 600 to 1,000, 1,000 to 1,500, or 2,000 to 3,000 ng (inclusive).
[0138] Methods for performing PCR are known in the art and typically include cycles of denaturing, annealing, and extension steps (which may be the same as or different from the annealing step).
[0139] An exemplary set of conditions involves a semi-nested PCR approach. The first PCR reaction uses a 20 μl reaction volume containing a 2× Qiagen MM final concentration, 1.875 nM of each primer (outer forward and reverse primers) in the library, and the DNA template. The thermal cycling parameters include 95°C for 10 minutes; 25 cycles of 96°C for 30 seconds, 65°C for 1 minute, 58°C for 6 minutes, 60°C for 8 minutes, 65°C for 4 minutes, and 72°C for 30 seconds; followed by a 2-minute hold at 72°C and then a 4°C hold. Next, 2 μl of the resulting product, diluted 1:200, is used as input in the second PCR reaction. This reaction uses a 10 μl reaction volume with a 1× Qiagen MM final concentration, 20 nM of each inner forward primer, and 1 μM reverse primer tag. Thermal cycling parameters included 95°C for 10 minutes; 15 cycles of 95°C for 30 seconds, 65°C for 1 minute, 60°C for 5 minutes, 65°C for 5 minutes, and 72°C for 30 seconds; followed by 72°C for 2 minutes, followed by a hold at 4°C. As discussed herein, the annealing temperature may optionally be higher than the melting temperature of some or all of the primers, as discussed herein (see U.S. Patent Application No. 14 / 918,544, filed October 20, 2015, which is incorporated herein by reference in its entirety).
[0140] Melting temperature (T m The annealing temperature (T) is the temperature at which half (50%) of a DNA duplex between an oligonucleotide (such as a primer) and its perfect complement dissociates into single-stranded DNA. A ) is the temperature at which the PCR protocol is run. In traditional methods, this is usually the lowest T of the primers used. m Because the T is 5°C lower than the T, nearly all possible duplexes are formed (essentially all primer molecules bind to the template nucleic acid). This is very efficient, but low temperatures can lead to nonspecific reactions. AOne consequence of having T is that a primer may anneal to a sequence other than the true target, since an internal single base mismatch or partial annealing may be tolerated. A (T m ), and at a given moment, only a small fraction of the targets (e.g., about 1-5%) have primers annealed to them. As they are extended, they are removed from the equilibrium of annealing and dissociating primers and targets (extension is T m (This is due to the rapid increase in temperature above 70 °C), and approximately 1-5% of the target copies will have primers. Therefore, by allowing the reaction longer time for annealing, it is possible to obtain approximately 100% target copies per cycle.
[0141] In various embodiments, the annealing temperature range is from 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, and 13°C at the lower end of the range to 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, 13°C, and 15°C at the upper end of the range, and does not exceed the melting temperature (e.g., empirically measured or calculated T ) of at least 25, 50, 60, 70, 75, 80, 90, 95, or 100% of the non-identical primers. m In various embodiments, the annealing temperature is between 1°C and 15°C (e.g., between 1°C and 10°C, between 1°C and 5°C, between 1°C and 3°C, between 3°C and 5°C, between 5°C and 10°C, between 5°C and 8°C, between 8°C and 10°C, between 10°C and 12°C, or between 12°C and 15°C, inclusive) and is at least 25; 50; 75; 100; 300; 500; 75 0; 1,000; 2,000; 5,000; 7,500; 10,000; 15,000; 19,000; 20,000; 25,000; 27,000; 28,000; 30,000; 40,000; 50,000; 75,000; 100,000; or the melting temperature (e.g., experimentally measured or calculated T mIn various embodiments, the annealing temperature is 1 to 15°C (e.g., 1°C to 10°C, 1°C to 5°C, 1°C to 3°C, 3°C to 5°C, 3°C to 8°C, 5°C to 10°C, 5°C to 8°C, 8°C to 10°C, 10°C to 12°C, or 12°C to 15°C, inclusive) and is higher than the melting temperatures (e.g., empirically measured or calculated T ) of at least 25%, 50%, 60%, 70%, 75%, 80%, 90%, 95%, or all of the non-identical primers. m ), and the length of the annealing step (per PCR cycle) is 5 to 180 minutes, for example, 15 to 120 minutes, 15 to 60 minutes, 15 to 45 minutes, or 20 to 60 minutes (inclusive).
[0142] In addition to thermal cycling during PCR, isothermal amplification is recognized as a means for amplifying nucleic acid molecules. For any of the PCR methods disclosed herein, those skilled in the art will understand how to adapt the method for use with this method. For example, in some embodiments, the reaction mixture can include tagged nucleic acid molecules, a pool of primers, nucleotide triphosphates, magnesium, and an isothermal polymerase. Several isothermal polymerases are available for performing isothermal amplification. These include Bst DNA polymerase, full length; Bst DNA polymerase, large fragment; Bst 2.0 DNA polymerase; Bst 2.0 armStart DNA polymerase; and Bst 3.0 DNA polymerase (all available from New England Biolabs). The polymerase used can depend on the method of isothermal amplification. Several types of isothermal amplification are available, including recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), nicking enzyme amplification reaction (NEAR), and template walking.
[0143] Sequencing of tagged nucleic acid molecules In some embodiments, the sequence of tagged nucleic acid molecules is directly determined by methods known in the art, particularly high-throughput sequencing.More typically, the sequence of tagged nucleic acid molecules is determined after one or more rounds of amplification carried out during sample preparation for high-throughput sequencing.Such amplification typically includes library preparation, clonal amplification, and amplification to add additional sequences or functions, such as sample barcodes, to sample nucleic acid molecules.During sample preparation for high-throughput sequencing, tagged nucleic acid molecules are typically clonally amplified on one or more solid supports.These monoclonal or substantially monoclonal colonies are then subjected to sequencing reactions.Furthermore, sample preparation for next-generation sequencing can typically include targeted amplification reactions after library preparation and before clonal amplification.Such targeted amplification can be multiplex amplification reactions.
[0144] In any of the embodiments disclosed herein, the methods and compositions can be used to identify amplification errors versus true sequence variations in a sample nucleic acid molecule. The present disclosure can further distinguish between possible sources of amplification errors and further identify the most likely true sequence of the initial sample nucleic acid molecule.
[0145] In some embodiments of the method provided herein, the sequence of at least a part, and in some embodiments, the whole sequence of at least one tagged nucleic acid molecule is determined.The method for determining the sequence of nucleic acid molecule is known in the art.Any sequencing method known in the art, such as Sanger sequencing, pyro-sequencing, reversible dye terminator sequencing, ligation sequencing or hybridization sequencing, can be used for this sequencing. In some embodiments, high-throughput next-generation (massively parallel) sequencing technologies can be used, including, but not limited to, Solexa (Illumina), Genome Analyzer IIx (Illumina), MiSeq (Illumina), HiSeq (Illumina), 454 (Roche), SOLiD (Life Technologies), Ion Torrent (Life Technologies, Carlsbad, CA), GS FLX+ (Roche), True Single Molecule Sequencing platform (Helicos), electron microscopy sequencing (Halcyon Molecular), or other sequencing methods can be used to sequence the tagged nucleic acid molecules generated by the methods provided herein. In some embodiments, any high-throughput, massively parallel sequencing method can be used, and those skilled in the art will understand how to adjust the disclosed methods to achieve appropriate MIT binding. Thus, high-throughput reactions, such as sequencing by synthesis or sequencing by ligation, can be used. Additionally, the sequencer can detect signals generated during the sequencing reaction, which can be fluorescent signals or ions such as hydrogen ions. All of these methods physically convert the genetic data stored in a sample of DNA into a set of genetic data, which is typically stored in a memory device until it is processed.
[0146] Identification of sample nucleic acid molecules The step of determining the sequence of the tagged nucleic acid molecule includes determining the sequence of at least a portion of the sample nucleic acid molecule, sample nucleic acid segment, or target locus, and the sequence of the tag (including the sequence of the MIT) that remains attached to the sample nucleic acid segment. In some embodiments, copies of tagged nucleic acid molecules derived from the same initial tagged nucleic acid molecule can be identified by comparing the MIT sequences attached to the tagged nucleic acid molecule. Copies derived from the same initial tagged nucleic acid molecule will have the same MIT attached to the same position relative to the sample nucleic acid segment. In some embodiments, the fragment-specific insert ends are mapped to specific locations within the genome of the organism, and these mapped locations or the sequences of the fragment-specific insert ends themselves, as discussed herein, are used together with the sequence of the MIT to identify the original tagged nucleic acid molecule from which the copies originate. In some embodiments, tagged nucleic acid molecules comprising complementary MIT and complementary nucleic acid segment sequences, i.e., tagged nucleic acid molecules derived from the same nucleic acid molecule and representing the plus and minus strands of the sample nucleic acid molecule, are identified and paired. In some embodiments, paired MIT families are used to verify differences in the original sequence. Any change in sequence must be present in all copies of the tagged nucleic acid molecule derived from the sample nucleic acid molecule. This information provides further confidence that the sequences of tagged nucleic acid molecules derived from the plus and minus strands of the sample represent differences in the sequence of the sample nucleic acid molecule and are not changes introduced during sample preparation or base calling errors during sequencing.
[0147] In some embodiments, two main types of tagged nucleic acid molecules useful for further analysis are generated: tagged nucleic acid molecules with identical binding MITs at the same position and substantially the same sample nucleic acid segment sequence, and tagged nucleic acid molecules with different binding MITs and substantially the same sample nucleic acid segment sequence. As discussed in detail herein, tagged nucleic acid molecules with identical binding MITs at the same position and substantially the same sample nucleic acid segment sequence can be used to identify amplification errors, and tagged nucleic acid molecules with at least one difference between binding MITs and substantially the same sample nucleic acid segment sequence can be used to identify true sequence variants.
[0148] After MIT binding, amplification errors can be identified by comparing the sequences of tagged nucleic acid molecules with identical MITs at the same relative positions with substantially the same sample nucleic acid sequence. If both strands of the initial sample nucleic acid molecule are tagged with the same one or more MITs, it is possible to identify paired MIT nucleic acid segment families with complementary MIT and nucleic acid segment sequences. These paired MIT nucleic acid segment families can be used to increase the confidence that sequence variations existed in both strands of the sample nucleic acid molecule. If tagged nucleic acid molecules derived from the sample nucleic acid molecule show differences in their sequences, this indicates either a mismatch in the sample nucleic acid molecule or an error was introduced during amplification or base calling. Sequences from paired MIT nucleic acid segment families with sequence differences would typically be discarded before further analysis. However, these paired MIT nucleic acid segment families with sequence differences can be used to identify mismatches in the sample nucleic acid molecule.
[0149] Amplification errors that introduce one or more changes in the sequence of a nucleic acid segment will not be present in all copies derived from the initial tagged nucleic acid molecule. If an error is introduced in the first round of amplification, up to 25% of copies derived from both strands of the initial tagged nucleic acid molecule will have an error in the sequence of the nucleic acid segment. If amplification proceeds at full efficiency, the proportion of copies with a particular error will be halved with each round of amplification; i.e., if an error is introduced in the second round, 12.5% of copies derived from the initial tagged nucleic acid molecule will have an error, and if an error is introduced during the third round of amplification, 6.25% of copies derived from the initial tagged nucleic acid molecule will have an error. This knowledge can be used to identify or estimate when an amplification error occurred; in embodiments in which multiple amplifications occur after MIT binding, this includes the introduction of an amplification error at that stage. In any of the embodiments disclosed herein, if an amplification error is present in a sample nucleic acid segment, the most likely sequence of the initial sample nucleic acid molecule can be determined using the methods detailed herein. For example, the most likely sequence can be determined from a pool of copies of the initial tagged nucleic acid molecule as the most common sequence. In some embodiments, prior probabilities can be used to determine the most likely sequence, such as, for example, known mutation rates at particular chromosomal sites in normal or diseased cells, or population frequencies of particular single nucleotide polymorphisms.
[0150] The likelihood of having the same amplification error in two or more tagged nucleic acid molecules with different MITs and substantially the same nucleic acid segment sequence is very low, and therefore identical sequence variants on tagged nucleic acid molecules with substantially the same sequence and identical MITs in the same relative positions are considered to be derived from the same molecule and not arose independently.
[0151] True sequence variations present in the sample nucleic acid segment can be identified because all copies derived from one initial tagged nucleic acid molecule have the same sequence at the mutation position, and at least one pool of copies of tagged nucleic acid molecules having substantially the same sample nucleic acid segment sequence and MIT difference have a different sequence at the same mutation position, where the MIT difference can be either at least one different combined MIT from the set of MITs or a different relative position of the same MIT.
[0152] In any of the embodiments disclosed herein, a sequence difference can be referred to as an amplification error if the proportion of copies derived from the same initial tagged nucleic acid molecule with sequence variation is less than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. In certain embodiments, copies are said to be derived from the same initial tagged nucleic acid molecule if the bound MITs are identical and the relative positions are identical, and if the sample nucleic acid segment sequences are substantially the same. In any of the embodiments disclosed herein, a sequence variation can be referred to as a true sequence variant in an initial tagged nucleic acid molecule if the sequence differs in at least two tagged nucleic acid molecules having substantially the same sample nucleic acid segment, and the pools of copies derived from each of the at least two tagged nucleic acid molecules having substantially the same sample nucleic acid segment are at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% identical within each pool, and each pool is identified by having at least one different MIT and / or MIT at a different position relative to the sample nucleic acid segment.
[0153] In some embodiments, the sequence of the tagged nucleic acid molecule is used to identify 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% of the lower end of the range. From the sample nucleic acid molecules, up to 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100% of the sample nucleic acid molecules can be identified at the upper end of the range.
[0154] In some embodiments, for each sample nucleic acid molecule, the method is used to detect an amplification error of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, 100, 250, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000 at the lower end of the range to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, Amplification errors of 16, 17, 18, 19, 20, 25, 50, 75, 100, 250, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, and up to 1,000,000 can be identified.In some embodiments, for each sample nucleic acid molecule, the method is used to select from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 50, 75, 100, 250, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, and 100,000 true sequence variants at the lower end of the range in the sample nucleic acid molecule to 2, 3, 4, 5, 6, 7, 8, 9 true sequence variants at the upper end of the range in the sample nucleic acid molecule. , 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 75, 100, 250, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000 , 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, and up to 1,000,000 true sequence variants can be identified.
[0155] Other uses of the embodiments disclosed herein will be apparent to those skilled in the art who understand how to adapt the methods. For example, the methods can be used to measure amplification bias, particularly changes in amplification bias of specific nucleic acid molecules after the introduction of amplification errors. The methods can also be used to characterize the mutation rate of polymerases. By dividing the sample and barcoding the reaction mixture, it is possible to simultaneously characterize the mutation rates of different polymerases.
[0156] MIT kit Any of the components used in the various embodiments disclosed herein can be assembled into a kit. The kit can include a container containing any of the MIT sets disclosed herein. The MIT can range from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, and 15 nucleotides in length at the lower end of the range to 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, and 30 nucleotides in length at the upper end of the range. The MIT can be a double-stranded nucleic acid adapter. These adapters can further include a portion of a Y adapter nucleic acid molecule having a base-paired double-stranded polynucleotide segment and at least one non-base-paired single-stranded polynucleotide segment. These Y adapters can contain identical sequences other than the MIT sequence. The double-stranded polynucleotide segment of the Y adaptor can be from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, and 25 nucleotides in length at the lower end of the range, to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, and 100 nucleotides in length at the higher end of the range. The single-stranded polynucleotide segment of the Y adaptor can be from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, and 25 nucleotides in length at the lower end of the range, to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, and 100 nucleotides in length at the higher end of the range.
[0157] In any of the embodiments disclosed herein, the MIT can be part of a polynucleotide segment comprising a universal primer binding sequence. In some embodiments, the MIT can be located 5' to the universal primer binding sequence. In some embodiments, the MIT can be positioned within the universal primer binding sequence such that the MIT sequence forms a non-base-paired loop when the polynucleotide segment binds to DNA. In any of the embodiments disclosed herein, the kit can include a set of sample-specific primers designed to bind to a sample nucleic acid molecule, nucleic acid segment, or internal sequence of a target locus. In some embodiments, the MIT can be part of a polynucleotide further comprising a sample-specific primer sequence. In these embodiments, the MIT can be positioned 5' to the sample-specific primer sequence or within the sample-specific primer sequence such that the MIT sequence forms a non-base-paired loop when the polynucleotide segment binds to DNA. In some embodiments, the set of sample-specific primers can include a forward and reverse primer for each target locus. In some embodiments, the set of sample-specific primers can be forward or reverse primers, and the set of universal primers can be used as reverse or forward primers, respectively.
[0158] In any of the embodiments disclosed herein, the kit can include single-stranded oligonucleotides on one or more immobilization substrates. In some embodiments, the single-stranded oligonucleotides on one or more immobilization substrates can be used to enrich a sample for specific sequences by performing hybrid capture and removing unbound nucleic acid molecules. In any of the embodiments disclosed herein, the kit can include a container containing a cell lysis buffer, a tube for performing cell lysis, and / or a tube for purifying DNA from a sample. In some embodiments, the cell lysis buffer, one and / or more tubes can be designed for a specific type of cell or sample, such as circulating cell-free fetal DNA and circulating cell-free tumor DNA found in blood samples.
[0159] Any of the kits disclosed herein can include an amplification reaction mixture comprising any of the following: a reaction buffer, dNTPs, and a polymerase. In some embodiments, the kit can include a ligation buffer and a ligase. In any of the embodiments disclosed herein, the kit can also include a means for clonally amplifying tagged nucleic acid molecules onto one or more solid supports. One of skill in the art will understand which components to include in a kit to enable use of such a kit for the various methods herein.
[0160] Determining the copy number of one or more chromosomes or chromosome segments of interest In some embodiments, the methods provided herein for identifying individual sample nucleic acid molecules using MIT can be used as part of a method for determining the copy number of one or more chromosomes or chromosomal segments of interest in a sample. As evidenced by the mathematical evidence provided in Example 3, significant cost and sample savings can be achieved by using a method including MIT to identify individual sample nucleic acid molecules as part of a method for determining the copy number of one or more chromosomes or chromosomal segments of interest in a sample. For example, based on the reduced noise and improved accuracy achieved using MIT to identify individual sample nucleic acid molecules shown in Example 1, acceptable and reliable results can be obtained using as little as 100 μl of plasma. Furthermore, acceptable and reliable results can be achieved with as few as 1,780,000 sequencing reads. Thus, two important limitations of current methods, namely sample volume and cost, can be overcome.
[0161] The present disclosure is useful, among other areas, in determining the copy number of one or more chromosomes or chromosomal segments of interest in a sample, as disclosed herein. Methods for determining the number of a chromosome or chromosome segment of interest that can be adapted for use in the methods of the present disclosure include, for example, U.S. patent application Ser. No. 13 / 499,086, filed March 29, 2012; U.S. patent application Ser. No. 14 / 692,703, filed April 21, 2015; U.S. patent application Ser. No. 14 / 877,925, filed October 7, 2015; U.S. patent application Ser. No. 14 / 918,544, filed October 20, 2015; "Noninvasive Prenatal Detection and Selective Analysis of Cell-Free DNA Obtained from Maternal Blood: Assessment of Trisomy 21 and Trisomy 18" (Sparks et al. April 2012. American Journal of Obstetrics and Gynecology. 206(4):319.e1-9); "Detection of Clonal and Subclonal Copy Number Variants in Cell-Free DNA from Breast Cancer Patients Using a Massively Multiplexed PCR Method" (Kirkizlar et al. October 2012). 2015. Translation Oncology. 8(5):407-416), each of which is incorporated herein by reference in its entirety.
[0162] Using MIT, a smaller sample volume of blood or a fraction thereof may be required to obtain acceptable and reliable results. In some embodiments, the blood sample may be a maternal blood sample for use in non-invasive prenatal testing. This can reduce the impact on the patient and reduce the cost of sample preparation. In any of the embodiments disclosed herein, the sample volume may range from 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.125, 0.15, 0.175, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, and 0.5 ml at the lower end of the range to 0.5 ml at the lower end of the range. On the upper end, it can be 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.125, 0.15, 0.175, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 3.5, 4, 4.5, and up to 5 ml. In some embodiments, the sample volume can be from 0.1, 0.125, 0.15, 0.175, 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, and 0.5 ml at the lower end of the range to 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 2.5, and 3 ml at the upper end of the range.
[0163] In any of the embodiments disclosed herein, the sample can be a maternal blood sample containing circulating cell-free DNA from the fetus and the fetus's mother. In some embodiments, these samples are used to perform non-invasive prenatal testing. In other embodiments, the sample can be a blood sample from a person suffering from or suspected of suffering from cancer. In some embodiments, the circulating cell-free DNA can include DNA fragments ranging from 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, and 150 nucleotides in length at the lower end of the range to 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, and 200 nucleotides in length at the upper end of the range.
[0164] In some embodiments, the length of any one or more chromosomal segments of interest is at the lower end of a range of 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 50,000, 60,000, 70,000, 80,000 , 90,000, and 100,000 nucleotides in length, at the upper end of the range: 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 3 00,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, 10,000,000, 15,000,000, 20,000,000, It can be up to 5,000,000, 30,000,000, 40,000,000, 50,000,000, 60,000,000, 70,000,000, 80,000,000, 90,000,000, 100,000,000, 125,000,000, 150,000,000, 175,000,000, 200,000,000, 250,000,000, and 300,000,000 nucleotides in length.
[0165] In one aspect, the disclosure features a method for determining the copy number of one or more chromosomes or chromosome segments of interest in a sample. In some embodiments, the method for determining the copy number of one or more chromosomes or chromosome segments of interest in a sample of blood or a fraction thereof includes: forming a reaction mixture of sample nucleic acid molecules and a set of molecular beacon tags (MITs) to generate a population of tagged nucleic acid molecules, wherein at least some of the sample nucleic acid molecules comprise one or more target loci of a plurality of target loci on a chromosome or chromosome segment of interest; amplifying the population of tagged nucleic acid molecules to generate a library of tagged nucleic acid molecules; determining the sequence of the bound MIT of a tagged nucleic acid molecule in the library of tagged nucleic acid molecules and the sequence of at least a portion of the sample nucleic acid segment to determine the identity of the sample nucleic acid molecule that gave rise to the tagged nucleic acid; Using the determined identities, measuring the amount of DNA for each target locus by counting the number of sample nucleic acid molecules containing each target locus; and determining, on a computer, the copy number of one or more chromosomes or chromosome segments of interest using the amount of DNA at each target locus in the sample nucleic acid molecule, wherein the number of target loci and the volume of the sample provide an effective amount of all target loci to achieve the desired sensitivity and specificity for the copy number determination. All target loci T L can be defined as the product of the total number of sample nucleic acid molecules spanning each target locus in the sample, C, and the number of target loci in the sample, L, where T L = C × L. Effective dose E Acan be defined as the amount required to obtain a particular number of total target loci for a target sensitivity and specificity. In some embodiments, the number of total target loci ranges from 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, and 100,000 total target loci at the lower end of the range to 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, and 100,000 total target loci at the upper end of the range. The total number of target loci may be up to 1,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, and 10,000,000 total target loci. The effective amount can take into account the sample preparation efficiency and the proportion of DNA in the mixed sample, for example, the proportion of fetal DNA in a maternal blood sample. Tables 1 and 3 in Example 3 show the total number of sequencing reads that are equivalent to all target loci required to achieve the target sensitivity and specificity for different methods of the present disclosure. In some embodiments, the total number of sample nucleic acid molecules in the sample nucleic acid molecule population is greater than the diversity of MITs in the MIT set. In further embodiments, the sample contains a mixture of two genetically different genomes. For example, the mixture can be a blood or plasma sample containing circulating cell-free tumor DNA and normal DNA, or maternal DNA and fetal DNA.
[0166] Example 3 herein provides a table identifying the total number of sequencing reads or total target loci required to achieve a particular level of specificity and sensitivity at different percentage mixtures ("proportion of G2 in the sample"). This can be, for example, the proportion of cancerous DNA versus normal DNA, or the proportion of fetal DNA versus maternal DNA. Total target loci are determined by multiplying the number of target loci for a chromosome or chromosome segment by the number of haploid copies of the target loci provided by the sample volume. For example, as shown in Example 3, to achieve 99% sensitivity and specificity in 4% fetal DNA or circulating cell-free DNA using a non-allelic method, 110,414 total target loci are required. This can be achieved using 0.5 ml of plasma, a plurality of at least 1,000 loci, and a sample preparation method that retains at least 25% of the initial total target loci using a set of at least 32 MITs. Thus, in this example, an effective amount is at least 1,000 loci and at least 0.5 ml of plasma.
[0167] In some embodiments, determining the copy number of one or more chromosomes or chromosome segments of interest can include comparing the amount of DNA at multiple target loci with the amount of DNA at multiple disomic loci on one or more chromosomes or chromosome segments that are predicted to be disomic. The amount of DNA at multiple disomic loci can be determined in the same manner as for multiple target loci, that is, by determining the sequence of the binding MIT of the tagged nucleic acid molecules in a library of tagged nucleic acid molecules and the sequence of at least a portion of the sample nucleic acid segment, using the determined sequence to determine the identity of the sample nucleic acid molecule that produced the tagged nucleic acid molecule, and using the determined identity to count the number of sample nucleic acid molecules that contain each target locus, thereby measuring the amount of DNA at each target locus. In some embodiments, the multiple disomic loci on one or more chromosomes or chromosome segments that are predicted to be disomic can be SNP loci.
[0168] In any of the embodiments disclosed herein, the number of loci in the plurality of target loci can range from 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, and 5,000 loci at the lower end of the range to 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, and 5,000 loci at the upper end of the range. The number of target loci can be up to 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, and 100,000 loci. In some embodiments, the number of target loci is at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 loci. In any of the embodiments disclosed herein, the number of loci in the plurality of disomic loci can range from 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, and 5,000 loci at the lower end of the range to 50, 60, 70, 80, 90 , 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, and up to 100,000 loci. In some embodiments, the number of disomic loci is at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 loci.
[0169] In various embodiments, a set of hypotheses regarding the copy number of one or more chromosomes or chromosomal segments of interest can be generated, and the measured DNA amount can be compared to the expected DNA amount based on each particular hypothesis. In the context of the present disclosure, a hypothesis can refer to the copy number of a chromosome or chromosomal segment of interest. It can also refer to possible ploidy states. It can also refer to possible allelic states or allelic imbalance. In some embodiments, a set of hypotheses can be designed so that one hypothesis from the set corresponds to the actual genetic state of a given individual. In some embodiments, a set of hypotheses can be designed so that all possible genetic states can be explained by at least one hypothesis from the set. In some embodiments of the present disclosure, the method can determine which hypothesis corresponds to the actual genetic state of the individual in question. In some embodiments, the set of hypotheses can include a hypothesis of fetal proportion in addition to the possible genetic states. In some embodiments, the set of hypotheses can include a hypothesis of average allelic imbalance in addition to the possible genetic states.
[0170] In some embodiments, a joint distribution model can be used to determine the relative probability of each hypothesis. A joint distribution model is a model that defines the probability of an event defined in terms of multiple random variables, given multiple random variables defined in the same probability space, where the probabilities of the variables are linked. In some embodiments, a degenerate case in which the probabilities of the variables are not linked can be used. In various embodiments of the present disclosure, determining the copy number of one or more chromosomes or chromosome segments of interest in a sample also includes combining the relative probability of each ploidy hypothesis determined using the joint distribution model with the relative probability of each ploidy hypothesis calculated using statistical methods derived from the group consisting of read count analysis, heterozygosity rate comparison, genotype signal probabilities normalized for specific parental status, and combinations thereof. In various embodiments, the joint distribution can combine the relative probability of each ploidy hypothesis with the relative probability of each fetal fraction hypothesis. In some embodiments of the present disclosure, determining the relative probability of each hypothesis can utilize the estimated proportion of DNA in the sample. In various embodiments, the joint distribution can combine the relative probability of each ploidy hypothesis with the relative probability of each allelic imbalance hypothesis. In some embodiments, determining the copy number of one or more chromosomes or chromosome segments of interest involves selecting the most probable hypothesis, which is performed using maximum likelihood estimation or maximum a posteriori techniques.
[0171] Maximum likelihood and maximum a posteriori estimation Most methods known in the art for detecting the presence or absence of a biological phenomenon or medical condition involve the use of single-hypothesis rejection tests, in which a metric value correlated with the condition is measured; if the metric value is on one side of a threshold, the condition is present; if the metric value is on the other side of the threshold, the condition is absent. In single-hypothesis rejection tests, when making a decision between null and alternative hypotheses, only the null hypothesis is examined. Without considering alternative distributions, it is not possible to estimate the likelihood of each hypothesis given the observed data, and therefore to calculate the confidence in the call. Thus, single-hypothesis rejection tests yield a "yes" or "no" answer without any emotion associated with the particular case.
[0172] In some embodiments, the methods disclosed herein can detect the presence or absence of a biological phenomenon or medical condition using maximum likelihood methods. This is a substantial improvement over methods that use single hypothesis rejection, because the threshold for calling the absence or presence of a condition can be appropriately adjusted for each case. This is particularly important for diagnostic techniques aimed at determining the presence or absence of aneuploidy in a pregnant fetus from genetic data obtained from a mixture of fetal and maternal DNA present in the free-floating DNA found in maternal plasma. This is because as the proportion of fetal DNA in the plasma-derived fraction changes, the optimal threshold for calling aneuploidy versus euploidy changes. As the fetal proportion decreases, the distribution of data related to aneuploidy becomes increasingly similar to the distribution of data related to euploidy.
[0173] Maximum likelihood estimation uses the distributions associated with each hypothesis to estimate the likelihood of the adjusted data based on each hypothesis. These conditional probabilities can then be converted into hypothesis calls and confidence levels. Similarly, maximum a posteriori estimation uses the same conditional probabilities as maximum likelihood estimation, but also incorporates population prior probabilities when selecting the best hypothesis and determining its confidence level. Thus, using maximum likelihood estimation (MLE) techniques or the closely related maximum a posteriori (MAP) techniques offers two advantages: first, it increases the likelihood of a correct call, and it also allows for the calculation of confidence levels for each call.
[0174] Exemplary Methods for Determining the Number of Sample Nucleic Acid Molecules Disclosed herein is a method for determining the number of DNA molecules in a sample by generating tagged nucleic acid molecules from each sample nucleic acid molecule by incorporating two MITs. Disclosed herein is a procedure for achieving the above objectives, followed by single molecule or clonal sequencing.
[0175] As detailed herein, this approach involves generating tagged nucleic acid molecules such that most or all of the tagged nucleic acid molecules from each locus have different combinations of MITs and can be identified by sequencing the MITs using clonal or single molecule sequencing. This identification can optionally use the mapped position of the nucleic acid segment. Each combination of MIT and nucleic acid segment represents a different sample nucleic acid molecule. This information can be used to determine the number of individual sample nucleic acid molecules in the original sample for each locus.
[0176] This method can be used in any application requiring quantitative assessment of the number of sample nucleic acid molecules. Furthermore, the number of individual nucleic acid molecules from one or more target loci can be related to the number of individual nucleic acid molecules from one or more disomic loci to determine relative copy number, copy number variation, allele distribution, allele ratio, allelic imbalance, or average allele imbalance. Alternatively, the copy numbers detected from various targets can be modeled by a distribution to identify the most likely copy number of the target locus. Applications include, but are not limited to, the detection of insertions and deletions, such as those found in carriers of Duchenne muscular dystrophy; quantification of chromosomal deletions or duplications, such as those observed in copy number variants; determination of chromosome copy number in samples from live individuals; and determination of chromosome copy number in samples from unborn children, such as embryos or fetuses.
[0177] This method can be combined with the simultaneous evaluation of the mutations contained in the determined sequence.This can be used to determine the number of sample nucleic acid molecules that are each allele in the original sample.This copy number method can be combined with the evaluation of SNP or other sequence mutations to determine the copy number of the target chromosome or chromosome segment from born or unborn individuals; the identification and quantification of copies from loci that have short sequence mutations but can be amplified by PCR from multiple target loci, such as in carrier detection of spinal muscular atrophy; and the determination of the copy number of nucleic acid molecules from different sources from samples that consist of a mixture of different individuals, such as in the detection of fetal aneuploidy from free-floating DNA obtained from maternal plasma.
[0178] In any of the embodiments disclosed herein, the method can include one or more of the following steps: (1) joining a Y adapter nucleic acid molecule having an MIT to a population of sample nucleic acid molecules by ligation; (2) performing one or more rounds of amplification; (3) using hybrid capture to enrich for target loci; and (4) measuring the amplified PCR products by multiple methods, such as clonal sequencing, to a sufficient number of bases to span the sequence.
[0179] In any of the embodiments disclosed herein, the method for a single target locus can include one or more of the following steps: (1) designing a standard pair of oligomers for amplification of a specific locus; (2) adding a specific sequence of bases with no or minimal complementarity to the target locus or genome to the 5' end of both target-specific PCR primers during synthesis. This sequence, called the tail, is a known sequence used for subsequent amplification, followed by an MIT. As a result, after synthesis, the tailed PCR primer pool will consist of a collection of oligomers that begin with the known sequence, followed by an MIT, followed by the target-specific sequence; (3) performing one round of amplification (denaturation, annealing, and extension) using only the tail oligomers; and (4) adding an exonuclease to the reaction, effectively terminating the PCR reaction, and incubating the reaction at an appropriate temperature to remove and extend any forward-leaning single-stranded oligos that did not anneal to the template to form double-stranded products. (5) Incubating the reaction at high temperature to denature the exonuclease and eliminate its activity, (6) Adding to the reaction a new oligonucleotide complementary to the tail of the oligomer used in the first reaction along with other target-specific oligomers to allow PCR amplification of the products generated in the first round of PCR, (7) Continuing amplification to generate enough product for downstream clonal sequencing, and (8) Measuring the amplified PCR products by a number of methods, for example, clonal sequencing, to a sufficient number of bases to span the sequence.
[0180] In some embodiments, the design and generation of primers with MIT can be summarized as follows: a primer with MIT consists of a sequence that is not complementary to the target sequence, followed by a region with MIT, followed by a target-specific sequence. The sequence 5' of the MIT can be used for subsequent PCR amplification and may contain a sequence useful for converting amplicons into a library for sequencing. In some embodiments, DNA can be measured by a sequencing method in which the sequence data represents the sequence of a single molecule. This can include methods in which single molecules are directly sequenced, or methods in which single molecules are amplified to form clones that can be detected by a sequencing instrument, but which are still single molecules, and are referred to herein as clonal sequencing.
[0181] In some embodiments, the method of the present disclosure comprises targeting multiple loci, either in parallel or not.Primers for different target loci can be independently prepared and mixed to produce multiplex PCR pools.In some embodiments, the original sample can be divided into subpools, and each subpool can target different loci, then recombined and sequenced.In some embodiments, after tagging and several amplification cycles, the pool can be subdivided to ensure efficient targeting of all targets, and then divided, and the smaller primer sets in the subdivided pools can be used to continue amplification, thereby improving subsequent amplification.
[0182] For example, imagine a heterozygous SNP in an individual's genome and a mixture of DNA from an individual where 10 sample nucleic acid molecules of each allele are present in the original DNA sample. After MIT incorporation and amplification, there may be 100,000 tagged nucleic acid molecules corresponding to that locus. Due to stochastic processes, the DNA ratio may be anywhere from 1:2 to 2:1, but because each sample nucleic acid molecule is tagged with MIT, it may be possible to determine that the DNA in the amplified pool originates from exactly 10 sample nucleic acid molecules from each allele. This method would therefore provide a more accurate measure of the relative abundance of each allele than methods that do not use this approach. For methods where it is desirable to minimize the relative amount of allele bias, this method would provide more accurate data.
[0183] The association of sequenced fragments with target loci can be achieved in several ways.In some embodiments, a sufficient length of sequence corresponding to MIT and a sufficient number of unique bases corresponding to target sequence is obtained from targeting fragments, allowing unambiguous identification of target loci.In other embodiments, the MIT primer containing MIT can also contain a locus-specific barcode (locus barcode) that identifies the target it is associated with.This locus barcode will be the same for all MIT primers for each individual target locus, and therefore the same for all resulting amplicons, but will be different from all other loci.In some embodiments, the tagging method disclosed herein can be combined with a one-sided nesting protocol.
[0184] One example of an application in which MIT may be particularly useful for determining copy number is non-invasive prenatal aneuploidy diagnosis, where the amount of DNA at one or more target loci can be used to help determine the copy number of one or more chromosomes or chromosome segments of interest in a fetus. In this context, it is desirable to amplify the DNA present in the initial sample while maintaining the relative abundance of various alleles. In some situations, particularly when very small amounts of DNA are present—for example, less than 5,000 copies of the genome, less than 1,000 copies of the genome, less than 500 copies of the genome, and less than 100 copies of the genome—a phenomenon known as bottlenecking can occur. This occurs when a small number of copies of a given allele are present in the initial sample, and amplification bias can result in an amplified pool of DNA with a significantly different ratio of alleles than in the initial DNA mixture. By using MIT on each DNA strand prior to standard PCR amplification, it is possible to eliminate n-1 copies of DNA from a set of n identically tagged nucleic acid molecules in a library derived from the same sample nucleic acid molecule. In this way, any allele bias or amplification bias can be eliminated from further analysis. In various embodiments of the present disclosure, the methods can be performed on fetuses at 4-5 weeks gestation, 5-6 weeks gestation, 6-7 weeks gestation, 7-8 weeks gestation, 8-9 weeks gestation, 9-10 weeks gestation, 10-12 weeks gestation, 12-14 weeks gestation, 14-20 weeks gestation, 20-40 weeks gestation, first trimester, second trimester, third trimester, or any combination thereof.
[0185] Another application where MIT is particularly useful for determining copy number or average allele imbalance is non-invasive cancer diagnosis, where the amount of genetic material at one locus or multiple loci can be used to help determine copy number variation or average allele imbalance. Allelic imbalance for aneuploidy determination, such as copy number variation determination, refers to the difference between the frequencies of alleles at a locus. This is an estimate of the copy number difference between homologs. Allelic imbalance can result from the complete loss of an allele or an increase in the copy number of one allele relative to the other allele. Allelic imbalance can be detected by measuring the ratio of one allele to the other allele in body fluids or cells from individuals who are constitutively heterozygous at a given locus. (Mei et al., Genome Res, 10:1126-37 (2000)). For a dimorphic SNP with alleles arbitrarily designated "A" and "B," the allelic ratio of the A allele is nA / (nA+nB), where nA and nB are the number of sequencing reads for alleles A and B, respectively. Allelic imbalance is the difference between the A and B allele ratios for a locus that is heterozygous in the germline. This definition is similar to that of SNVs, where the proportion of abnormal DNA is typically measured using the variant allele frequency, i.e., nm / (nm+nr), where nm and nr are the number of sequence reads for the variant and reference alleles, respectively. Thus, the proportion of abnormal DNA for CNVs can be measured by the average allelic imbalance (AAI), defined as |(H1-H2)| / (H1+H2), where Hi is the average copy number of homolog i in the sample and Hi / (H1+H2) is the fractional abundance, or homolog ratio, of homolog i. The maximum homolog ratio is the homolog ratio of the more abundant homolog.
[0186] Accurate measurement of allele distribution in a sample Current sequencing approaches can be used to estimate the distribution of alleles in a sample. One such method involves randomly sampling sequences from pooled DNA, known as shotgun sequencing. The proportion of a particular allele in sequencing data is typically very low and can be determined by simple statistics. The human genome contains approximately 3 billion base pairs. Therefore, if the sequencing method used produces 100-bp reads, a particular allele will be measured approximately once in every 30 million sequence reads.
[0187] In some embodiments, the method of the present disclosure is used to determine the presence or absence of two or more different haplotypes comprising the same set of loci in a DNA sample from the measured allele distribution of loci from chromosomes.The different haplotypes can be two different homologous chromosomes from one source, three different homologous chromosomes from one source, three different homologous haplotypes in a sample comprising a mixture of two genetically different genomes (wherein one haplotype is shared between genetically different genomes), three or four haplotypes in a sample comprising a mixture of two genetically different genomes (wherein one or two haplotypes are shared between genetically different genomes), or other combinations.The alleles that are polymorphic between haplotypes tend to be more informative, but any alleles where both genetically different genomes are not homozygous for the same allele can provide useful information through the measured allele distribution, beyond the information obtained from simple read count analysis.
[0188] However, shotgun sequencing of such samples is highly inefficient because it results in many sequence reads from loci that are not polymorphic between different haplotypes in the sample, or reads from unrelated chromosomes, and therefore provides no information about the proportion of the target haplotype. Disclosed herein is a method for specifically targeting and / or preferentially enriching segments of DNA in a sample that are more likely to be polymorphic in the genome, thereby increasing the yield of allele information obtained by sequencing. For the measured allele distribution in an enriched sample to truly represent the actual amount present in the target individual, it is important that there is little or no preferential enrichment of one allele compared to other alleles at a given locus in the target segment. Current methods known in the art for targeting polymorphic alleles are designed to reliably detect at least some of any alleles present. However, these methods were not designed to measure an unbiased allele distribution of polymorphic alleles present in the original mixture. It is difficult to predict which specific target enrichment method will produce an enriched sample, and the measured allele distribution will more accurately represent the allele distribution present in the original unamplified sample than other methods. Theoretically, many enrichment methods are anticipated to achieve this goal, but current amplification, targeting, and other preferential enrichment methods suffer from significant stochastic bias. One embodiment of the method disclosed herein allows multiple alleles found in a mixture of DNA corresponding to a given locus in the genome to be amplified or preferentially enriched so that each allele is approximately equally enriched. In other words, this method can increase the relative amount of alleles present in the entire mixture, while the ratio between the alleles corresponding to each locus remains the same as they were in the original DNA mixture. In some reported methods, preferential enrichment of loci can result in allele biases of more than 1%, 2%, 5%, or even 10%.This preferential enrichment is a capture bias, or amplification bias, that can be small with each cycle when using a hybrid capture approach, but can become significant when compounded over 20, 30, or 40 cycles. For purposes of this disclosure, the ratio remaining essentially the same means that the ratio of alleles in the original mixture divided by the ratio of alleles in the resulting mixture is 0.95-1.05, 0.98-1.02, 0.99-1.0, 0.995-1.005, 0.998-1.002, 0.999-1.001, or 0.9999-1.0001. Note that the allele ratio calculations presented herein cannot be used in determining the ploidy state of a target individual and can only be used as a metric for measuring allele bias. Because the methods disclosed herein can be used to specifically count the number of sample nucleic acid molecules, MIT can be used to remove errors due to capture bias, amplification bias, and allele bias.
[0189] In some embodiments, once a mixture is preferentially enriched for a set of target loci, it can be sequenced using either previous, current, or next-generation sequencing instruments, as discussed in more detail herein. The ratio can be assessed by sequencing through specific alleles within the chromosome or chromosomal segment of interest. These sequencing reads can be analyzed and counted according to allele type and the ratio of different alleles determined accordingly. For variants with lengths of one to several bases, allele detection is performed by sequencing, and it is essential that the sequencing read spans the allele in question to assess the allelic composition of the captured molecule. The total number of captured nucleic acid molecules measured for a genotype can be increased by extending the length of the sequencing read. Complete sequencing of all tagged nucleic acid molecules would ensure the collection of the maximum amount of data available in the enriched pool. However, sequencing is currently expensive, and methods that can measure allele distributions using fewer sequence reads would be of great value. Furthermore, there are technical limitations on the maximum read length, and as the read length increases, so too does the accuracy limit. The most useful alleles are one to a few bases in length, but theoretically any allele shorter than the length of a sequencing read can be used. Larger variants, such as segmental copy number variants, can often be detected by the collection of these small variants because the entire collection of SNPs within a segment overlaps. Variants larger than a few bases, such as STRs, require special consideration, and targeted approaches may or may not be successful.
[0190] There are several targeting approaches that can be used to specifically isolate and enrich one or more variant positions in a genome. Typically, these rely on utilizing invariant sequences adjacent to the variant sequence. Other researchers have reported on targeting in the context of sequencing when the substrate is maternal plasma (see, for example, Liao et al., Clin. Chem. 2011; 57(1): pp. 92-101). However, these approaches use targeting probes that target exons and do not focus on targeting polymorphic loci in the genome. In various embodiments, the methods of the present disclosure involve using targeting probes that focus exclusively or nearly exclusively on polymorphic loci. In some embodiments, the methods of the present disclosure involve using targeting probes that focus exclusively or nearly exclusively on SNPs. In some embodiments of the present disclosure, the targeted polymorphic sites consist of at least 10% SNPs, at least 20% SNPs, at least 30% SNPs, at least 40% SNPs, at least 50% SNPs, at least 60% SNPs, at least 70% SNPs, at least 80% SNPs, at least 90% SNPs, at least 95% SNPs, at least 98% SNPs, at least 99% SNPs, at least 99.9% SNPs, or entirely SNPs.
[0191] In some embodiments, the disclosed methods can be used to determine genotypes (the base composition of DNA at specific loci) and the relative proportions of those genotypes from a mixture of DNA molecules, where these DNA molecules can be derived from one or more genetically distinct genomes. In some embodiments, the disclosed methods can be used to determine genotypes at a set of polymorphic loci and the relative proportions of the amounts of different alleles present at those loci. In some embodiments, the polymorphic loci can consist entirely of SNPs. In some embodiments, the polymorphic loci can include SNPs, single tandem repeats, and other polymorphisms. In some embodiments, the disclosed methods can be used to determine the relative distribution of alleles at a set of polymorphic loci in a DNA mixture, where the DNA mixture includes DNA from an individual and a tumor growing within the individual.
[0192] In some embodiments, the mixture of DNA molecules can be derived from DNA extracted from multiple cells of a single individual. In some embodiments, the original collection of cells from which the DNA is derived can contain a mixture of diploid or haploid cells of the same or different genotypes if the individual is mosaic (germline or somatic). In some embodiments, the mixture of nucleic acid molecules can also be derived from DNA extracted from a single cell. In some embodiments, the mixture of DNA molecules can also be derived from DNA extracted from a mixture of two or more cells from the same individual or different individuals. In some embodiments, the mixture of DNA molecules can be derived from cell-free DNA such as that present in plasma. In some embodiments, when tumor DNA is present in plasma, such as in pregnancy where fetal DNA has been shown to be present in the mixture or in cancer, the biological material can be a mixture of DNA from one or more individuals. In some embodiments, the biological material can be derived from a mixture of cells found in maternal blood, where some of the cells are of fetal origin. In some embodiments, the biological material can be cells from the blood of a pregnant woman, which is enriched for fetal cells.
[0193] The algorithm used to determine the copy number of one or more chromosomes or chromosome segments of interest can calculate the predicted allele distribution for a target locus for a large number of possible fetal ploidy states and for various fetal cfDNA fractions, taking into account parental genotype and crossover frequency data (such as data from the HapMap database). Unlike allele ratio-based methods, this also accounts for linkage disequilibrium and can use non-Gaussian data models to describe the predicted distribution of allele measurements at SNPs given observed platform characteristics and amplification bias. The algorithm then compares the various predicted allele distributions with the actual allele distribution measured in the sample and calculates the likelihood of each hypothesis (monosomy, dinosomy, or trinosomy, of which there are multiple hypotheses based on various possible crossover analyses) based on the sequencing data. The algorithm sums the likelihoods of the individual monosomy, dinosomy, or trinosomy hypotheses and calls the hypothesis with the greatest total likelihood the copy number and fetal fraction. Similar algorithms can be used to determine the average allele imbalance in a sample, and those skilled in the art will understand how to modify the method.
[0194] The following examples are presented to provide those of ordinary skill in the art with a complete disclosure and description of how to use the embodiments provided herein, and are not intended to limit the scope of the disclosure or represent that the following examples are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for. Unless otherwise specified, parts are parts by volume and temperatures are in degrees Celsius. It is understood that variations in the methods described can be made without changing the basic aspects that the examples are intended to illustrate. [Example]
[0195] Example 1 Exemplary Workflow for Identifying Sample Nucleic Acid Molecules An example of a method for identifying sample nucleic acid molecules after amplification in a high-throughput sequencing workflow is provided herein. Non-limiting exemplary amplicon structures generated using such methods are shown in Figure 3. A set of nucleic acid sources is prepared by isolating nucleic acids from natural sources. For example, circulating cell-free DNA can be isolated from a sample of blood or a fraction thereof from a target patient using known methods. Some of the sample nucleic acids in the blood may contain one or more target sites. The sample nucleic acid molecules are treated to remove all overhangs in a blunt-end repair reaction using Klenow large fragment, and polynucleotide kinase is used to ensure that all 5' ends are phosphorylated. A 3' adenosine residue is added to the blunt-end-repaired sample nucleic acid molecules using Klenow fragment (exo) to increase ligation efficiency. A set of 206 MITs, each 6 nucleotides in length and with at least two base differences from all other MITs, is designed to be included in a double-stranded polynucleotide sequence adjacent to the 3' T overhangs of a standard high-throughput sequencing Y adapter, as illustrated in Figure 1. Next, a set of Y adaptors, each containing a different MIT, is ligated to both ends of each sample nucleic acid molecule using a ligase in a ligation reaction to generate a population of tagged nucleic acid molecules. For the ligation reaction, 10,000 sample nucleic acid molecules are tagged with a library of 206 MIT-containing Y adaptors. The resulting population of tagged nucleic acid molecules contains Y adaptors with MITs ligated to both ends of the sample nucleic acid molecules, as shown in Figure 1, such that the MITs are ligated to the ends of the sample nucleic acid segments, also referred to as inserts of the tagged nucleic acid molecules.
[0196] A library of tagged nucleic acid molecules is then prepared by amplifying the population of tagged nucleic acid molecules using a universal primer that binds to the primer binding site on the Y adapter. A target enrichment step is then performed to isolate and amplify tagged nucleic acid molecules containing sample nucleic acid segments with the target SNP. Target enrichment can be performed using a one-sided PCR reaction or hybrid capture. Either of these target enrichment reactions can be a multiplex reaction using a population of primers (one-sided PCR) or probes (hybrid capture) specific to the sample nucleic acid segments containing the target SNP. One or more additional PCR reactions are then performed using universal primers containing different barcode sequences for each patient sample and clonal amplification and sequencing primer binding sequences (R-tag and F-tag in Figure 3). The structures of the resulting amplified tagged nucleic acid molecules are shown schematically in Figure 3.
[0197] The amplified tagged nucleic acid molecules are then clonally amplified on a solid support using a universal sequence added during a single amplification reaction. The sequences of the clonally amplified tagged nucleic acid molecules are then determined using a high-throughput sequencing device such as an Illumina sequencer. For tagged nucleic acid molecules enriched using single-sided PCR, the MIT (i.e., insert) on the right side of the sample nucleic acid segment is the first base read by one of the sequencing reads. For tagged nucleic acid molecules enriched using hybrid capture, one MIT remains on each side of the sample nucleic acid segment (i.e., insert), and the first base of the first linked MIT on one end of the sample nucleic acid segment is the first base read in the first read, and the second linked MIT on the other end of the sample nucleic acid segment is the first base read in the second read. The resulting sequencing reads are then analyzed. Using the sequences of the fragment-specific insert ends, the location of each end of the nucleic acid segment is mapped to a specific location within the organism's genome, and these locations can be used in combination with the MIT to identify each tagged nucleic acid molecule. This information is then analyzed using commercially available software packages that are programmed to distinguish true sequence differences in the sample nucleic acid molecules from errors introduced in any of the sample preparation amplification reactions.
[0198] Example 2 Reducing the error rate using MIT on sample nucleic acid molecules An example is provided herein demonstrating the reduced error rate provided by using MIT to identify amplification errors in a high-throughput sequencing sample preparation workflow. In each experiment, 10,000 input copies of the human genome (10,000 copies × (3,000,000,000 bp / genome) / (150 bp / nucleic acid molecule) = 2 × 10 ) were used in 58 μl (final concentration 5.75 nM). 11 Total sample nucleic acid molecules (containing 2 x 10 11Three experiments were performed using two independent DNA samples containing all sample nucleic acid molecules to generate libraries of tagged nucleic acid molecules with MITs at the 5' and 3' ends as disclosed herein. For these experiments, a set of 196 MITs was used, ranging in concentration from 0.5 to 2 μM, resulting in a ratio of approximately 85:1 to approximately 350:1 between the total number of MITs in the reaction mixture and the total number of sample nucleic acid molecules in the reaction mixture. As shown, 2 × 10 11 For a sample with a total sample nucleic acid molecule of 196 MITs alone or approximately 40,000 combinations of two MITs were used.
[0199] In each experiment, libraries were enriched for tagged nucleic acid molecules containing TP53 exons by hybrid capture using a commercially available kit. The enriched libraries were then amplified by PCR using a universal primer capable of binding to the universal primer binding sequence previously incorporated into the tagged nucleic acid molecules. The universal primer contained a barcode sequence that was different for each sample and an additional sequence that enabled sequencing on an Illumina HiSeq 2500. Samples from each experiment were then pooled and paired-end sequenced on the HiSeq 2500 in fast mode for 150 cycles, with forward and reverse reads, respectively.
[0200] Sequencing data were demultiplexed using commercially available software. From each sequencing read, data on bases corresponding to the length of the MIT+T overhang (a total of 7 nucleotides in these experiments) was trimmed from the start of the read and recorded. The remaining trimmed reads were then combined and mapped to the human genome. The fragment end position for each read was recorded. All reads with at least one base covering the target locus (TP53 exon) were considered on-target. The average depth of the reads was calculated at the base-by-base level across the target locus. The average error rate (expressed as a percentage) was calculated by counting all base calls across the target locus that did not correspond to the reference genome (GRCh37) and dividing these by the total base calls across the target locus. Next, for each base position in the target locus, the sequencing data were grouped into MIT families, where each MIT family shared an identical MIT at the same relative position to the analyzed base position, as well as the same fragment end position and the same sequencing direction (positive or negative relative to the human genome). Each of these families represented a group of molecules that were likely to be clonal amplifications of the same sample nucleic acid molecule that entered the MIT library preparation process. Each sample nucleic acid molecule that entered the MIT library preparation process would have generated two families, one mapped to each of the positive and negative genomic orientations. Next, the two MIT families, one in the positive orientation and the other in the negative orientation, were used to generate paired MIT nucleic acid segment families, where each family contained a complementary MIT at the same relative position relative to the analyzed base position and complementary fragment end position. These paired MIT families represented a group of sequenced molecules that were more likely to be clonal amplifications of the same sample nucleic acid molecule that entered the MIT library preparation process. Next, the average error rate (expressed as a percentage) was calculated by counting all base calls within all paired MIT nucleic acid segment families across the target locus that did not correspond to the reference genome (GRCh37) and dividing these by the total base calls within all paired MIT families across the target locus.
[0201] Figure 4 shows the results of three experiments. Each sample contained 33 ng of DNA, representing 10,000 input copies of the haploid human genome. Sequencing data from these experiments yielded 4.4 million to 10.7 million mapped reads per sample and 3.0 million to 7.8 million on-target reads per sample. The ratio of on-target reads to mapped reads ranged from 68% to 74%. The average depth of reads across the target loci ranged from approximately 98,000 to approximately 244,000 read depths. When all data were included, the average error rate ranged from 0.15% to 0.26%. The average error rate calculated using data from only the paired MIT nucleic acid segment family ranged from 0.0036% to 0.0067%. The average mean error rate of the two samples in each experiment and the error rate of the paired MIT nucleic acid segment family demonstrate a dramatic reduction in error rate when using the paired MIT nucleic acid segment family (Figure 5). The residual errors observed here are likely due to single nucleotide polymorphisms in the samples because single nucleotide polymorphism positions were not excluded. The error rates of the paired MIT nucleic acid segment families were 23-73 times lower than their original error rates. In particular, experiments B and C, which had higher original error rates compared to experiment A, experienced a greater reduction in error rate when calculated using paired MIT families. These results demonstrate the usefulness of MIT for error removal.
[0202] Example 3 Mathematical analysis demonstrating low sample volume for copy number determination using MIT This example provides an analysis of the number of target loci and plasma sample volume to provide an effective amount of all target loci to achieve the desired sensitivity and specificity for copy number determination using MIT. In a sample containing a mixture of two genomes, G1 and G2, the copy number of a chromosome or chromosome segment of interest can be determined for one of the genomes. G1 and G2 can have various copy numbers of the chromosomes of interest, such as two copies of each chromosome in one set of chromosomes and one copy in another set. Assume that G2 has one or more reference chromosomes or chromosome segments on its genome with known copy numbers (typically one or more chromosomes or chromosome segments expected to be disomic) and one or more chromosomes or chromosome segments of interest on its genome with unknown copy numbers (although the possible copy numbers are assumed to be known). The copy number of G2 for a chromosome or chromosome segment of interest whose true copy number is unknown can be estimated (if the set of possible copy numbers is known). Note that the copy numbers of G1 on both the reference chromosome or chromosome segment and the chromosome or chromosome segment of interest are known. The measurement technique is modeled as capturing a nucleic acid molecule and identifying whether it belongs to one or more reference chromosomes or chromosome segments, or one or more chromosomes or chromosome segments of interest, where there is the possibility of error.
[0203] Assuming that the sample contains a finite number of nucleic acid molecules, sample nucleic acid molecules can be sampled until an accurate estimate of the number of nucleic acid molecules in the sample belonging to one or more reference chromosomes or chromosome segments and one or more chromosomes or chromosome segments of interest is obtained. Using the estimate of the proportion of G2 in the sample, test statistics for different copy number hypotheses of G2 in one or more chromosomes or chromosome segments of interest can be calculated as shown below.
[0204] Method 1: Quantitative non-allelic method In this method, the number of sample nucleic acid molecules is compared for one or more reference chromosomes or chromosome segments and one or more chromosomes or chromosome segments of interest. Once a tagged nucleic acid molecule is sequenced, it is assumed that there is an equal probability of sequencing a tagged nucleic acid molecule from one or more reference chromosomes or chromosome segments and one or more chromosomes or chromosome segments of interest. This probability is denoted by p, where p=0.5. An example of a test statistic that can be used is the number of nucleic acid molecules (n) from one or more chromosomes or chromosome segments of interest. t ) to the total number of nucleic acid molecules observed (n).
[0205] T=n t / n
[0206] For n>20, the distribution of T can be approximated by a normal distribution with variance (p(1-p)) / n=0.25 / n for p=0.5. The mean of the distribution depends on the copy number hypothesis of G2 being tested, and obtaining more observations (i.e., decreasing the variance) can improve the precision of the results. This allows the creation of an estimator that achieves a particular sensitivity and specificity.
[0207] Assume G2 represents 4% of the sample mixture (and G1 is 96% of the mixture). Assume further that G1 has two copies of each locus on both the reference chromosome or chromosome segment and the chromosome or chromosome segment of interest. Assume also that G2 has two copies of each locus in one or more reference chromosomes or chromosome segments. Consider two hypotheses: H2, where G2 has two copies of each locus in the chromosome or chromosome segment of interest, and H3, where G2 has three copies of each locus in the chromosome or chromosome segment of interest. As above, a normal distribution can be used to estimate the distribution of the above test statistic. Because the copy numbers of both G1 and G2 are identical on both the reference chromosome or chromosome segment and the chromosome or chromosome segment of interest, the mean of the test statistic for H2 is 0.5. The mean of the test statistic for H3 is:
[0208] ((1-4%) / 2+3 / 4×4%) / (1 / 2+1 / 2×(1-4%)+3 / 4×4%)=0.50495
[0209] mean μ and variance σ 2 To represent the normal distribution with 2 ) using the usual notation. Thus, the distribution of the test statistic for the two hypotheses is:
[0210] H2:N(0.5, 0.25 / n)
[0211] H3:N(0.50495, 0.25 / n)
[0212] Using this information, we can calculate the n needed to achieve a particular sensitivity and specificity. Suppose we want a sensitivity and specificity of 99%, then given a normal distribution X with mean 0 and variance 1, Prob(X<-2.326)=1%. Therefore,
[0213] ((0.5-0.505) / 2) / (0.5 / √n)<-2.326
[0214] Solving for n > 220,827, we therefore require approximately 110,414 observations for each chromosome or chromosome segment. See Table 1 for the number of observations required for one or more reference chromosomes or chromosome segments and one or more target chromosomes or chromosome segments, respectively, for a range of mixture proportions and target sensitivity and specificity. [Table 1]
[0215] Method 2: Using allelic ratios Similar to the quantitative approach described in Method 1, molecular-based methods can be used to determine heterozygosity rates at known SNPs. In this approach, for a SNP on one or more chromosomes or chromosomal segments of interest, which can take on allele values of A or B, the test statistic is the observed proportion of the reference allele. In particular, for a given SNP, let A and B denote the observed number of molecules with the A and B alleles, respectively. In this way, the heterozygosity rate can be defined.
[0216] H=AA+B
[0217] and the number of SNP molecules is
[0218] N=A+B.
[0219] Let A1 and A2 denote the number of A alleles in genomes G1 and G2, respectively, at a SNP of interest. Similarly, B1 and B2 denote the number of B alleles in genomes G1 and G2, respectively, at a SNP of interest. The distribution of A is a binomial distribution, whose parameters are functions of A1, A2, B1, B2, and N. Suppose A1 and B1 are known, and we wish to estimate A2 and B2. To do this, we calculate the probability of the observed heterozygosity rate H for all possible values of A2 and B2, and then use Bayes' rule to calculate the probabilities of A2 and B2 from the observed H. For example, suppose G2 is 4% of the sample mixture (hence, G1 is 96% of the mixture). Further, suppose G1 has two copies of each locus in the reference chromosome or chromosome segment and the chromosome or chromosome segment of interest. We consider two hypotheses for G2: having two or three copies. These two hypotheses are denoted as H2 (G2 has two copies) and H3 (G2 has three copies), respectively. Under these assumptions, the values of the binomial parameters p and A1, A2, B1, and B2 for each hypothesis are calculated as follows:
[0220] p=(0.96×A1+0.04×A2) / (0.96×A1+0.04×A2+0.96×B1+0.04×B2).
[0221] This gives the following values for p (Table 2): [Table 2]
[0222] We further know that A is distributed as bino(pN) and H has a normal distribution with mean p and variance p(1-p) / N. As the number of nucleic acid molecules increases, the variance of the distribution decreases, and various hypotheses can be more easily distinguished. For example, suppose (A1=1, B1=1) and we want to distinguish H2 from H3. For simplicity, we reduce this problem to distinguishing between (A2=1, B2=1) and (A2=2, B1=1). The model developed above can be used to calculate the minimum number of nucleic acid molecules required to achieve a certain specificity and sensitivity (Table 3). [Table 3]
[0223] Practical implications Using the methods analyzed above and the efficiency of sample and library preparation, it is possible to calculate the amount of sample required to obtain a specific number of unique sequencing reads for a particular sensitivity and specificity. An exemplary workflow would be: Sample Collection → Sample Preparation → Library Preparation → Hybrid Capture → Barcoding → Sequencing. Based on this workflow, and subject to some assumptions regarding the efficiency of each step, it is possible to work backward to determine sample requirements. In this example, the barcoding step is assumed to have no significant impact. When N unique sequencing reads from a chromosome or chromosome segment are required, the preferred approach is to exhaustively sequence the nucleic acid molecule. Results based on the "Coupon Collector's Problem" (see, e.g., Dawkins, Brian (1991), "Siobhan's problem: the coupon collector revisited", The American Statistician, 45 (1): 76-82) can be used as a guideline for how many sequence reads are needed to have a particular probability of sequencing all nucleic acid molecules. See the table below. For example, if there are 1,000 unique tagged nucleic acid molecules to be sequenced, a read depth of approximately 12x is required to have a 99% probability of observing all nucleic acid molecules. This estimate assumes that each sequence read is equally likely to be any of the 1,000 tagged nucleic acid molecules. If this is not the case, the calculated factor of 12 can be replaced with an empirically determined value. During the library preparation and hybrid capture steps, some of the sample nucleic acid molecules present in the blood vessel are lost. Assuming that 75% of the molecules are lost during these processes (i.e., 25% of the sample nucleic acid molecules are retained), more nucleic acid molecules are needed in the original sample to ensure that enough tagged nucleic acid molecules remain for barcoding. Here, the binomial distribution can be used to estimate the number of nucleic acid molecules in the sample required to have a certain probability of having a specific number of nucleic acid molecules after the library and hybrid capture steps.
[0224] Based on the above reasoning, using Method 1, for a sensitivity and specificity of 1% in a mixture with 4% G2, approximately 110,000 sequencing reads are required for both the reference chromosome or chromosome segment and the chromosome or chromosome segment of interest (Table 1). If the combination of the library preparation and hybrid capture steps has an overall efficiency of 25%, more than 110,000 starting copies are required in the sample. Using a simple binomial model, at least 443,000 sample nucleic acid molecules are required to ensure a greater than 99% probability of having at least 110,000 nucleic acid molecules available for barcoding and subsequent sequencing. Assuming library preparation starts with 443,000 nucleic acid molecules, the expected number of sample nucleic acid molecules after the library preparation and hybrid capture steps would be in the range of 110,000 to 111,400 molecules. To ensure measurement of all original molecules, a larger number, i.e., 111,400 nucleic acid molecules, can be used for further calculations. Due to variability in measuring nucleic acid molecules, substantially more measurements are required to have a high probability of measuring all 111,400 nucleic acid molecules. For example, to sequence all tagged nucleic acid molecules with a 99% probability, 16 times as many nucleic acid molecules must be sequenced. Therefore, approximately 1,780,000 reads are required for each chromosome or chromosome segment. This estimate assumes that each sequence read is equally likely to be any one of the 111,400 tagged nucleic acid molecules. If this is not the case, the calculated factor of 16 can be replaced with an empirically measured one.
[0225] As mentioned above, a total sample of approximately 443,000 nucleic acid molecules is required to achieve the aforementioned performance. The required 111,400 sequencing reads can be achieved by measuring multiple loci in each chromosome or chromosome segment. For example, if nucleic acid molecules are measured at 1,000 different loci, an average of approximately 112 unique nucleic acid molecules from each locus are required for sequencing, resulting in an average of approximately 443 unique nucleic acid molecules in the starting sample. If the underlying sample type is a human plasma sample, this contains 1,200–1,800 single haploid copies of the genome per ml of plasma. Furthermore, on average, a 1 ml blood sample contains approximately 0.5 ml of plasma. Therefore, taking these constraints into account, 1 ml of blood (0.5 ml of plasma and 600–900 unique nucleic acid molecules from each locus) should be sufficient to determine the copy number of a chromosome or chromosome segment of interest.
[0226] Here, MIT can be used to count individual sample nucleic acid molecules and reduce the variance associated with other quantitative methods. To simplify the counting of individual sample nucleic acid molecules, each sample nucleic acid molecule from a locus (i.e., each of the 443 nucleic acid molecules) must have a different combination of bound MITs. Assuming that each nucleic acid molecule has two types of MITs bound to it, the number of possible combinations of bound MITs is N 2 where N is the number of MITs in the set. There are approximately 443 copies of each locus, so N 2 must be greater than 443. Some margin is useful, so N 2 = 1,000, N would be approximately 32. The precise start and end genomic coordinates of the nucleic acid segment can also be used in combination with the sequence of the MIT to identify the sample nucleic acid molecule.
[0227] Those skilled in the art may devise numerous modifications and other embodiments within the scope and spirit of the present disclosure. Indeed, those skilled in the art may make variations in the described materials, methods, drawings, experiments, examples, and embodiments without changing the basic aspects of the present disclosure. Any disclosed embodiment may be used in combination with other disclosed embodiments. All headings herein are for the convenience of the reader and do not limit the disclosure in any way.
Claims
1. (a) extracting cell-free DNA (cfDNA) from a biological sample of a subject with or suspected of having cancer; (b) producing a non-natural composition of concentrated DNA, said production comprising: (i) tagging at least one end of the extracted cfDNA, or a DNA fragment derived from the extracted cfDNA, with at least one adapter to produce adapted DNA, wherein the at least one adapter comprises a molecular beacon tag (MIT); (ii) performing universal amplification on the adapted DNA to produce amplified adapted DNA; and (iii) selectively enriching at least a portion of the amplified adapted DNA to produce enriched DNA, wherein the selective enrichment comprises performing one-sided PCR using a universal primer and a plurality of target-specific primers to amplify at least a portion of the amplified adapted DNA, or capturing at least a portion of the amplified adapted DNA using a plurality of hybrid capture probes; and (c) performing an analysis of the concentrated DNA, the analysis comprising: (i) performing massively parallel sequencing on the enriched DNA to obtain sequence reads, wherein the sequence reads include determining the sequences of the MIT and the extracted cfDNA or DNA fragments derived from the extracted cfDNA; (ii) grouping the sequence reads based on a set of sequence features; and (iii) using the grouped sequence reads to identify one or more cancer mutations in the biological sample of the subject; A method comprising:
2. 2. The method of claim 1, wherein the set of sequence features for grouping the sequence reads comprises: the same target locus; the same MIT at the same relative position to the target locus; and, optionally, the same fragment end positions of the extracted cfDNA or DNA fragments derived from the extracted cfDNA.
3. 3. The method of claim 1 or 2, wherein the biological sample is a blood, plasma, serum, or urine sample.
4. The method according to any one of claims 1 to 3, wherein the biological sample is a plasma sample.
5. 5. The method of any one of claims 1 to 4, wherein the extracted cfDNA or DNA fragments derived from the extracted cfDNA are tagged with 50 to 1000 different MITs, each MIT being 3 to 8 nucleotides in length, and the sequences of the different MITs differ from each other by at least 2 nucleotides.
6. 6. The method of any one of claims 1 to 5, wherein the extracted cfDNA or DNA fragments derived from the extracted cfDNA are tagged with 100 to 500 different MITs, each MIT being 4 to 8 nucleotides in length, and the sequences of the different MITs differ from each other by at least 2 nucleotides.
7. The method according to any one of claims 1 to 6, wherein the adapter further comprises a universal priming sequence, and the MIT is located more internally than the universal priming sequence in the adapted DNA.
8. 8. The method of any one of claims 1 to 7, wherein the adapters further comprise a universal priming sequence, and step (b)(ii) comprises performing universal amplification using the universal priming sequence to produce amplified adapted DNA.
9. 9. The method of claim 1, wherein the grouping comprises grouping sequence reads that have the same target locus; a pair of identical MITs that are in the same relative position with respect to the target locus; and the same start and end genomic coordinates when the extracted cfDNA or a DNA fragment derived from the extracted cfDNA is mapped to a reference genome.
10. 10. The method of claim 9, wherein the grouping further comprises pairing a first family of grouped sequence reads having a positive genomic orientation with a second family of grouped complementary sequence reads having a negative genomic orientation, wherein the first family and the second family comprise complementary MITs at the same relative positions.
11. 3. The method of claim 2, wherein grouping the sequence reads using the MIT and the fragment end positions reduces the error rate of identifying the cancer mutations, and the error rate is calculated by counting the number of all base calls across the target loci that do not correspond to a reference genome and dividing that number by the total number of base calls across the target loci.
12. 10. The method of claim 9, wherein grouping the sequence reads using the MIT pair and the start and end genomic coordinates when the extracted cfDNA or a DNA fragment derived from the extracted cfDNA is mapped to a reference genome reduces the error rate of identifying the cancer mutation, and the error rate is calculated by counting the number of all base calls across the target loci that do not correspond to a reference genome and dividing that number by the total number of base calls across the target loci.
13. 11. The method of claim 10, wherein grouping the sequence reads using the MIT pair and the start and end genomic coordinates when the extracted cfDNA or a DNA fragment derived from the extracted cfDNA is mapped to a reference genome reduces the error rate of identifying the cancer mutation, and the error rate is calculated by counting the number of all base calls that do not correspond to a reference genome in all paired families across the target loci and dividing that number by the total number of base calls in all paired families across the target loci.
14. 14. The method of any one of claims 1 to 13, wherein step (b)(iii) comprises selectively enriching between 50 and 5000 target loci.
15. 15. The method of any one of claims 1 to 14, wherein step (b)(iii) comprises selectively enriching between 100 and 2500 target loci.
16. 16. The method of any one of claims 1 to 15, wherein step (b)(iii) comprises performing one-sided PCR using a universal primer and a plurality of target-specific primers to amplify at least a portion of the amplified adapted DNA.
17. 16. The method of any one of claims 1 to 15, wherein step (b)(iii) comprises capturing at least a portion of the amplified adapted DNA using a plurality of hybrid capture probes.
18. 18. The method of any one of claims 1 to 17, wherein the cancer mutation comprises a single nucleotide variant, an insertion, a deletion, or a copy number mutation.
19. 19. The method of any one of claims 1 to 18, wherein the enriched DNA is further amplified to introduce sample-specific barcodes, and the enriched DNA from multiple samples is pooled and sequenced together.
20. 20. The method of any one of claims 1-19, further comprising estimating the fraction of cancer DNA in the extracted cfDNA based on the sequence reads.