Target enrichment by unidirectional dual-probe primer extension

The unidirectional dual-probe primer extension method addresses the inefficiencies of current nucleic acid sequencing methods by enabling rapid and cost-effective enrichment of target nucleic acids, including unknown structural polymorphisms, with improved on-target yield and reduced sequencing waste.

JP7804630B2Active Publication Date: 2026-01-22F HOFFMANN LA ROCHE & CO AG +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023176172
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-12-21
Filing Date
2023-10-11
Publication Date
2026-01-22
Estimated Expiration
2038-12-19

AI Technical Summary

Technical Problem

Current nucleic acid sequencing methods for enriching target regions are either too complex and time-consuming or unable to accommodate unknown structural polymorphisms, leading to inefficiencies and high costs.

Method used

A method involving unidirectional dual-probe primer extension, where oligonucleotides are hybridized and extended with polymerases to create primer extension complexes, which are then captured and amplified using complementary adaptors and moieties, allowing for rapid and simple enrichment of target nucleic acids.

Benefits of technology

This approach provides a rapid, simple, and efficient method for enriching target nucleic acids, reducing turnaround time and material costs while accommodating unknown structural polymorphisms, with high on-target yield and reduced sequencing waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804630000007
    Figure 0007804630000007
  • Figure 0007804630000008
    Figure 0007804630000008
  • Figure 0007804630000009
    Figure 0007804630000009
Patent Text Reader

Abstract

To provide a method for enrichment of at least one target nucleic acid in a library of nucleic acids.SOLUTION: A first oligonucleotide is hybridized to a target nucleic acid in library of nucleic acids having first and second adapters. The hybridized first oligonucleotide is extended with a first polymerase, thereby producing a first primer extension complex including the target nucleic acid and the extended first oligonucleotide. The first primer extension complex is captured, enriched relative to the library of nucleic acids, and a second oligonucleotide is hybridized to the target nucleic acid. The hybridized second oligonucleotide is extended with a second polymerase, thereby producing a second primer extension complex including the target nucleic acid and the extended second oligonucleotide, and further liberating the extended first oligonucleotide from the first primer extension complex.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] This disclosure relates generally to enriching nucleic acid targets in a sample, and more particularly to enriching targets for nucleic acid sequencing, including high-throughput sequencing. [Background technology]

[0002]

[0002] The present invention belongs to a class of technologies that allow users to focus on regions of interest within nucleic acids to be sequenced. This reduces the costs associated with sequencing reactions and subsequent data analysis. Currently, there are three general types of technologies for selectively capturing regions of interest within nucleic acids present in a sample. The first is hybridization capture, in which the region of interest is captured by hybridization of probes that can selectively bind to a capture surface. This capture allows for the removal of non-target nucleic acids, followed by the release and recovery of the captured target molecules. Advantages of this type of technology include the ability to capture exome-sized regions and regions containing unknown structural polymorphisms. Disadvantages include long, complex protocols that tend to take more than eight hours to complete. The complexity is primarily due to the need to prepare randomly fragmented shotgun libraries prior to hybridization. The hybridization step alone can take up to three days to complete. Examples of this type of technology include the SECAP EZ Target Enrichment System (ROCHE) and the SURESELECT Target Enrichment System (AGILENT).

[0003] Another method of target enrichment is dual-target primer-based amplification. This method uses two probes at the boundaries of the target to enrich for the region of interest. This method tends to be completed in less than 8 hours and is simpler than hybridization capture methods. However, dual-primer-based techniques cannot enrich for sequences with unknown structural polymorphisms. The most established dual-primer approach is multiplex polymerase chain reaction (PCR). This is a very simple, single-step process, but it can only amplify a few dozen targets per reaction tube. Other, more recent technologies are now available, including the TRUSEQ Amplicon Sequencing Kit (ILLUMINA) and the ION TORRENT AMPLISEQ Sequencing Kit (LIFE TECHNOLOGIES), which can amplify hundreds to thousands of targets in a single reaction tube and require only a few processing steps.

[0004]

[0004] The third technique is single-target primer-based amplification. In this method, targets are enriched by amplifying a region defined by a single target primer and a terminally ligated universal primer. Similar to hybridization-based approaches, these techniques require the generation of a randomly fragmented shotgun library prior to selective hybridization of target oligonucleotides. However, instead of using these oligonucleotides to capture targets and wash away non-target molecules, an amplification step is employed to selectively amplify the region between the randomly generated termini and the target-specific oligonucleotides. The advantage of this technique is that, unlike dual-primer techniques, it allows for the detection of sequences with unknown structural polymorphisms. It is also faster and simpler than hybridization-based techniques. However, this type of technique is slower and more complex than dual-primer-based approaches. Examples of this type of technique are ARCHER's Anchor Multiplex PCR (ARCHER DX) and OVATION Target Enrichment System (NUGEN). Summary of the Invention [Problem to be solved by the invention]

[0005]

[0005] There remains an unmet need for a rapid and simple method of target enrichment that also accommodates unknown structural polymorphisms in the target sequence. [Means for solving the problem]

[0006] According to one embodiment, the present disclosure provides a method for enriching at least one target nucleic acid in a library of nucleic acids. The method includes hybridizing a first oligonucleotide to a target nucleic acid in the library of nucleic acids. Each nucleic acid in the library of nucleic acids has a first end comprising a first adaptor and a second end comprising a second adaptor. The method further includes extending the hybridized first oligonucleotide with a first polymerase, thereby generating a first primer extension complex comprising the target nucleic acid and the extended first oligonucleotide. The method further includes capturing the first primer extension complex, enriching the first primer extension complex with respect to the library of nucleic acids, hybridizing a second oligonucleotide to the target nucleic acid, and extending the hybridized second oligonucleotide with a second polymerase, thereby generating a second primer extension complex comprising the target nucleic acid and the extended second oligonucleotide, thereby releasing the extended first oligonucleotide from the first primer extension complex. The method further includes amplifying the target nucleic acid with a third polymerase, a first amplification primer, and a second amplification primer, wherein the first amplification primer has a 3' end complementary to the first adapter and the second amplification primer has a 3' end complementary to the second adapter.

[0007] In one embodiment, the method further comprises sequencing the amplified target nucleic acid.

[0008] In another embodiment, the first oligonucleotide comprises a capture moiety.

[0009] In another embodiment, capturing the first primer extension complex comprises capturing a capture moiety on a solid support.

[0008]

[0010] In another embodiment, the capture moiety is biotin and the solid support comprises streptavidin.

[0011] In another embodiment, the first oligonucleotide is bound to a solid support before hybridizing to the target nucleic acid, and hybridizing the first oligonucleotide to the target nucleic acid and extending the hybridized first oligonucleotide with a polymerase thereby captures the first primer extension complex on the solid support.

[0009]

[0012] In another embodiment, capturing the first primer extension complex is performed after extending the hybridized first oligonucleotide.

[0013] In another aspect, the method further comprises incorporating at least one modified nucleotide into at least one of the extended first oligonucleotide in the first primer extension complex and the extended second oligonucleotide in the second primer extension complex.

[0010]

[0014] In another embodiment, the modified nucleotides are selected from dUTP and nucleotides having a capture moiety.

[0015] In another embodiment, the method further comprises incorporating at least one modified nucleotide into the extended first oligonucleotide in the first primer extension complex, wherein the at least one modified nucleotide comprises a capture moiety.

[0011]

[0016] In another embodiment, capturing the first primer extension complex comprises capturing a capture moiety on a solid support.

[0017] In another aspect, the method further comprises incorporating at least one uracil into at least one of the extended first oligonucleotide of the first primer extension complex and the extended second oligonucleotide of the second primer extension complex, thereby forming a uracil-containing oligonucleotide product.

[0012]

[0018] In another embodiment, the method further comprises digesting the uracil-containing oligonucleotide product.

[0019] In another embodiment, digestion of uracil-containing oligonucleotide products is accomplished with at least one of uracil DNA glycosylase and DNA glycosylase lyase.

[0013]

[0020] In another embodiment, the DNA glycosylase lyase is selected from endonuclease IV and endonuclease VIII.

[0021] In another embodiment, the method further comprises contacting the library of nucleic acids with a blocking oligonucleotide.

[0014]

[0022] In another embodiment, the blocking oligonucleotide is at least partially complementary to at least one of the first adaptor and the second adaptor.

[0023] In another embodiment, the blocking oligonucleotide is a universal blocking oligonucleotide.

[0015]

[0024] In another embodiment, the first adaptor and the second adaptor have the same nucleic acid sequence.

[0025] In another embodiment, the first adaptor and the second adaptor have different nucleic acid sequences.

[0016]

[0026] In another embodiment, the first adaptor and the second adaptor are forked adaptors.

[0027] In another embodiment, the first adaptor and the second adaptor comprise at least one uracil.

[0017]

[0028] In another embodiment, at least one of the first polymerase and the second polymerase is a uracil-incompatible polymerase.

[0029] In another embodiment, the third polymerase is a uracil-compatible polymerase.

[0018]

[0030] In another embodiment, the second oligonucleotide hybridizes to the target nucleic acid at a position 5' relative to the first oligonucleotide.

[0031] In another embodiment, the third polymerase is a uracil-incompatible polymerase.

[0019]

[0032] In another aspect, at least one of the first adaptor, the second adaptor, the first amplification primer, and the second amplification primer comprises at least one of a unique identifier (UID) sequence, a molecular identifier (MID) sequence.

[0020]

[0033] According to another embodiment, the present disclosure provides a kit for enriching at least one target nucleic acid in a library of nucleic acids. The kit includes a first oligonucleotide complementary to a target nucleic acid in the library of nucleic acids, each nucleic acid having a first end comprising a first adaptor and a second end comprising a second adaptor. The kit further includes a second oligonucleotide complementary to the target nucleic acid, a first amplification primer, and a second amplification primer. The first oligonucleotide includes a capture moiety, the second oligonucleotide hybridizes to the target nucleic acid at a position 5' relative to the first oligonucleotide, the first amplification primer has a 3' end complementary to the first adaptor, and the second amplification primer has a 3' end complementary to the second adaptor.

[0021]

[0034] According to another embodiment, the present disclosure provides a kit for enriching at least one target nucleic acid in a nucleic acid library. The kit includes a first oligonucleotide complementary to a target nucleic acid in the nucleic acid library, each nucleic acid in the nucleic acid library having a first end comprising a first adaptor and a second end comprising a second adaptor. The kit further includes a modified nucleotide having a capture moiety, a second oligonucleotide complementary to the target nucleic acid, a first amplification primer, and a second amplification primer. The second oligonucleotide hybridizes to the target nucleic acid at a position 5' relative to the first oligonucleotide, the first amplification primer has a 3' end complementary to the first adaptor, and the second amplification primer has a 3' end complementary to the second adaptor.

[0022]

[0035] In one embodiment, the kit further comprises at least one of a uracil nucleotide, a uracil-compatible polymerase, a uracil-incompatible polymerase, and a blocking oligonucleotide.

[0023]

[0036] According to another embodiment, the present disclosure provides a composition comprising a library of nucleic acids including at least one target nucleic acid. Each nucleic acid in the library of nucleic acids has a first end comprising a first adaptor, a second end comprising a second adaptor, and a region of interest intermediate the first adaptor and the second adaptor. The composition further comprises an extended first oligonucleotide hybridized to the region of interest of the target nucleic acid. The extended first oligonucleotide comprises at least one capture moiety. The composition further comprises a solid support bound to the at least one capture moiety, a second oligonucleotide hybridized to the target nucleic acid at a position 5' relative to the first extended oligonucleotide, and a polymerase associated with the 3' end of the second oligonucleotide.

[0024]

[0037] In one embodiment, the composition further comprises a blocking oligo hybridized to each of the first adaptor and the second adaptor.

[0038] In another embodiment, at least one capture moiety is located at the 5' end of the extended first oligonucleotide.

[0025]

[0039] In another embodiment, at least one capture moiety is incorporated into the extension portion of the extended first oligonucleotide.

[0040] In another embodiment, the extended first oligonucleotide further comprises at least one uracil and at least one thymine.

[0026]

[0041] In another embodiment, the polymerase is a uracil-incompatible polymerase.

[0042] In another embodiment, at least one of the first adaptor and the second adaptor comprises at least one uracil and at least one thymine.

[0027]

[0043] In another aspect, releasing the extended first oligonucleotide from the first primer extension complex is accomplished using an enzyme having an activity selected from strand displacement activity, 5' to 3' exonuclease activity, and flap endonuclease activity.

[0028]

[0044] The foregoing and other aspects and advantages of the present invention will become apparent from the following description. In the description, reference is made to the accompanying drawings, which form a part hereof and which show, by way of example, preferred embodiments of the invention. Such embodiments do not, however, necessarily represent the full scope of the invention, and reference is therefore made to the claims and this specification for interpreting the scope of the invention. [Brief explanation of the drawings]

[0029] [Figure 1]

[0045] FIG. 1 shows a schematic flow diagram illustrating an embodiment of a method for enriching at least one target nucleic acid in a library of nucleic acids according to the present disclosure. [Figure 2]

[0046] 1 shows a schematic representation of a first embodiment of a method for enriching at least one target nucleic acid in a library of nucleic acids according to the present disclosure. In the illustrated embodiment, a first oligonucleotide comprises a capture moiety for liquid-phase capture of the target nucleic acid. [Figure 3]

[0047] 1 shows a schematic representation of a further second embodiment of a method for enriching at least one target nucleic acid in a library of nucleic acids according to the present disclosure. In the illustrated embodiment, a first oligonucleotide is bound to a solid support for in situ capture of the target nucleic acid. [Figure 4]

[0048] 1 shows a schematic representation of a further third embodiment of a method for enriching at least one target nucleic acid in a library of nucleic acids according to the present disclosure, in which one or more capture moieties are incorporated during extension of a first oligonucleotide hybridized to the target nucleic acid, thereby allowing capture of a complex comprising the target nucleic acid and the extended first oligonucleotide on a solid support. [Figure 5]

[0049] FIG. 1 shows a schematic illustration of multiple nucleic acids in a library molecule showing intermolecular adaptor-adaptor hybridization. [Figure 6]

[0050] Figure 6A shows fluorescent output traces from an electrophoretic DNA analyzer of a library of nucleic acids derived from human genomic DNA that was adapted to common adapter end sequences using a commercially available library preparation kit. Data were collected following standard PCR amplification of 1 μL of a 10 ng library and 1 μL of a 100 ng library for 5 and 12 cycles, respectively. The nucleic acid libraries were sampled and enriched for target nucleic acids according to the present disclosure.

[0051] FIG. 6B shows a fluorescence-based size analysis of the library of nucleic acids of FIG. 6A after primer extension target enrichment and amplification according to the present disclosure. [Figure 7]

[0052] Figure 7 shows bar graphs illustrating high-level sequencing metrics for the library of enriched nucleic acids from Figure 6B. Over 99% of the sequencing reads mapped to sequences known to be present in the library, and approximately half of the sequencing reads mapped to the enriched target nucleic acids. The fold-80 base penalty for the library was 1.4 and 1.5 for first oligonucleotide primer annealing temperatures of 60°C and 65°C, respectively. Within each cluster of three bars, data are shown for the percentage of trimmed reads mapped (left), the percentage of padded target nucleic acid bases (center), and the percentage of on-target, non-redundant reads mapped (right). DETAILED DESCRIPTION OF THE INVENTION

[0030] I. Definition

[0053] In this application, unless otherwise clear from the context, (i) the term "a" may be understood to mean "at least one"; (ii) the term "or" may be understood to mean "and / or"; (iii) the term "comprising" may be understood to encompass the listed component or step, whether presented by itself or together with one or more additional components or steps; and (iv) the terms "about" and "approximately" may be understood to allow for standard variations understood by one of ordinary skill in the art; and (v) when ranges are specified, endpoints are included.

[0031]

[0054] Adapter: As used herein, "adapter" refers to a nucleotide sequence that can be added to another sequence to impart additional properties to that sequence. Adapters can be single-stranded or double-stranded, or can have both single-stranded and double-stranded portions.

[0032]

[0055] Approximately: As used herein, the term "approximately" or "about," when applied to one or more reference values, refers to a value similar to the stated reference value. In certain embodiments, the term "approximately" or "about" refers to a range of values ​​that is within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater or less) of the stated reference value (except where such number exceeds 100% of possible values), unless otherwise stated or apparent from the context.

[0033]

[0056] Associated with: As the term is used herein, two events or entities are "associated" with one another if the presence, level, and / or form of one correlates with the other. For example, a particular entity (e.g., a polypeptide, gene signature, metabolite, etc.) is considered to be associated with a particular disease, disorder, or condition if its presence, level, and / or form correlates with the frequency of occurrence and / or susceptibility of the disease, disorder, or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically "associated" with one another if they interact, directly or indirectly, such that they are in physical proximity to one another and / or remain in physical proximity to one another. In some embodiments, two or more entities that are physically associated with one another are covalently bound to one another; in some embodiments, two or more entities that are physically associated with one another are not covalently bound to one another, but are non-covalently associated, e.g., by hydrogen bonding, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.

[0034]

[0057] Barcode: As used herein, "barcode" refers to a nucleotide sequence that confers identity to a molecule. A barcode can confer a unique identity to an individual molecule (and its copies). Such a barcode is a unique ID (UID). A barcode can confer identity to an entire population of molecules (and their copies) from the same source (e.g., a patient). Such a barcode is a multiplex ID (MID).

[0035]

[0058] Biological sample: As used herein, the term "biological sample" typically refers to a sample obtained or derived from a biological source of interest (e.g., a tissue or organism or cell culture), as described herein. In some embodiments, the source of interest comprises or consists of an organism, such as an animal or a human. In some embodiments, the biological sample comprises or consists of a biological tissue or bodily fluid. In some embodiments, the biological sample can be or comprise bone marrow; blood; blood cells; ascites; tissue or fine needle biopsy sample; cell-containing bodily fluid; free-floating nucleic acid; sputum; saliva; urine; cerebrospinal fluid, ascites; pleural effusion; feces; lymph; gynecological fluid; skin swab; vaginal swab; oral swab; nasal swab; washings or lavage fluids, such as ductal washings or bronchoalveolar lavage; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; other bodily fluids, secretions, and / or excretions; and / or cells therefrom, etc. In some embodiments, a biological sample comprises or consists of cells obtained from an individual. In some embodiments, the obtained cells are or comprise cells from the individual from whom the sample was obtained. In some embodiments, a sample is a "primary sample" obtained directly from a source of interest by any suitable means. For example, in some embodiments, a primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, stool, etc.), etc. In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components and / or adding one or more agents), e.g., filtration using a semi-permeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of specific components, etc.

[0036]

[0059] Blocking oligonucleotide: An oligonucleotide that is complementary to another nucleic acid present in the reaction mixture and can hybridize to such nucleic acid to prevent undesired hybridization of such nucleic acid. Such another nucleic acid can be a synthetic nucleic acid, for example, a primer or an adapter. When a primer or adapter is incorporated into a library nucleic acid molecule, the undesired hybridization that is to be prevented can occur. The blocking oligonucleotide does not need to be perfectly complementary to the nucleic acid to be protected from undesired hybridization, but must form a sufficiently stable hybrid to prevent the undesired event from occurring. For that purpose, the blocking oligonucleotide can be a universal base or a T m It may contain modified bases.

[0037]

[0060] Comprising: A composition or method described herein as "comprising" one or more specified elements or steps is open-ended, meaning that the specified elements or steps are required, although other elements or steps may be added within the composition or method. It should be understood that a composition or method described as "comprising" (or "comprising") one or more specified elements or steps also describes a corresponding, more limited composition or method that "consists essentially of" (or "consisting essentially of") the same specified elements or steps, meaning that the composition or method includes the specified required elements or steps, and may also include additional elements or steps that do not materially affect the basic and novel characteristics of the composition or method. It should also be understood that any composition or method described herein as "comprising" or "consisting essentially of" one or more specified elements or steps also describes a corresponding, more limited, closed-ended composition or method that "consists of" (or "consisting of") the specified elements or steps, excluding other unspecified elements or steps. In any composition or method disclosed herein, known or disclosed equivalents of any named essential element or step may be substituted for that element or step.

[0038]

[0061] Designed: As used herein, the term "designed" refers to (i) an agent whose structure is selected or chosen by the human hand; (ii) an agent produced by a process involving the human hand; and / or (iii) an agent that is distinct from natural substances and other known agents.

[0039]

[0062] Determining: Those skilled in the art who have read this specification will recognize that "determining" can be utilized or accomplished by using any of a variety of techniques available to those skilled in the art, including, for example, the specific techniques explicitly cited herein. In some embodiments, determining involves manipulation of a physical sample. In some embodiments, determining involves consideration and / or manipulation of data or information, for example, using a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or material from a source. In some embodiments, determining involves comparing one or more characteristics of the sample or entity to a comparable reference.

[0040]

[0063] Identity: As used herein, the term "identity" refers to the overall relatedness between polymer molecules, e.g., between nucleic acid molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. In some embodiments, polymer molecules are considered to be "substantially identical" to one another if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. Calculation of the percent identity of two nucleic acid or polypeptide sequences can be performed, for example, by aligning the two sequences for optimal comparison purposes (e.g., for optimal alignment, gaps can be introduced into one or both of the first and second sequences, and non-identical sequences can be disregarded for comparison purposes). In certain embodiments, the length of the sequences aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100% of the length of the reference sequence. Nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., nucleotide or amino acid) as the corresponding position in the second sequence, the molecules are identical at that position. The percent identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps that need to be introduced for optimal alignment of the two sequences and the length of each gap. Sequence comparison and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For example, percent identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989, 4:11-17), incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons performed with the ALIGN program use a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. The percent identity between two nucleotide sequences can alternatively be determined using the GAP program in the GCG software package using the NWSgapdna.CMP matrix.

[0041]

[0064] Ligation site: As used herein, a "ligation site" is a portion of a nucleic acid molecule (other than the blunt ends of a double-stranded molecule) that facilitates ligation. "Compatible ligation sites" present in two molecules allow for preferential ligation of the two molecules to one another.

[0042]

[0065] Sample: As used herein, the term "sample" refers to a substance that is or contains a composition of interest for qualitative and / or quantitative evaluation. In some embodiments, the sample is a biological sample (i.e., comes from a living organism (e.g., a cell or organism)). In some embodiments, the sample is derived from a geological, aquatic, astronomical, or agricultural source. In some embodiments, the source of interest comprises or consists of a living organism, such as an animal or a human. In some embodiments, a sample for forensic analysis is or comprises biological tissue, biological fluids, organic or non-organic matter, such as clothing, dirt, plastic, water, etc. In some embodiments, an agricultural sample comprises or consists of organic matter, such as leaves, petals, bark, wood, seeds, plants, fruit, etc.

[0043]

[0066] Single-stranded ligation: As used herein, "single-stranded ligation" is a ligation procedure that begins with at least one single-stranded substrate and typically includes one or more double-stranded or partially double-stranded adapters.

[0044]

[0067] Solid Support: As used herein, "solid support" refers to any solid material capable of interacting with a capture moiety. A solid support can be a solution-phase support (e.g., glass beads, magnetic beads, or other similar particles) that can be suspended in solution, or a solid-phase support (e.g., silicon wafer, glass slide, etc.). Examples of solution-phase supports include superparamagnetic spherical polymer particles such as DYNABEADS magnetic beads from INVITROGEN, or magnetic glass particles such as those described in U.S. Patents 656,568, 6,274,386, 7,371,830, 6,870,047, 6,255,477, 6,746,874, and 6,258,531.

[0045]

[0068] Substantially: As used herein, the term "substantially" refers to a qualitative state of exhibiting the whole or nearly whole extent or degree of a desired quality or characteristic. Those skilled in the art of biology will understand that biological and chemical phenomena rarely proceed to completion and / or completely, or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.

[0046]

[0069] Synthetic: As used herein, the term "synthetic" means produced by the hand of man, either by having a structure that is not found in nature, or by being associated with one or more other components that are not associated with nature, or by not being associated with one or more other components that are associated with nature, and therefore in a form that is not found in nature.

[0047]

[0070] Universal Primer: As used herein, "universal primer" and "universal priming site" refer to a primer and priming site that do not occur naturally in a target sequence. Typically, the universal priming site is present in an adapter or target-specific primer. A universal primer binds to the universal priming site and can direct primer extension therefrom.

[0048]

[0071] Variant: As used herein, the term "variant" refers to an entity that exhibits significant structural identity with a reference entity, but that differs structurally from the reference entity in the presence or level of one or more chemical moieties when compared to the reference entity. In many embodiments, a variant also differs functionally from the reference entity. Generally, whether a particular entity is properly considered to be a "variant" of a reference entity is based on the degree of structural identity with the reference entity. As will be recognized by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. A variant, by definition, is a distinct chemical entity that shares one or more such characteristic structural elements. To cite a few examples, small molecules have a characteristic core structural element (e.g., a macrocyclic core) and / or one or more characteristic pendant moieties, such that variants of small molecules share the core structural element and the characteristic pendant moieties but differ in other pendant moieties within the core and / or the type of bond present (single vs. double bond, E vs. Z, etc.); polypeptides may have characteristic sequence elements comprised of multiple amino acids that have designated positions relative to each other in linear or three-dimensional space and / or that contribute to a particular biological function; and nucleic acids may have characteristic sequence elements comprised of multiple nucleotide residues that have designated positions relative to each other in linear or three-dimensional space. For example, a variant polypeptide can differ from a reference polypeptide as a result of one or more differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc.) covalently attached to the polypeptide backbone. In some embodiments, the variant polypeptide exhibits an overall sequence identity with the reference polypeptide that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. Alternatively, or in addition, in some embodiments, the variant polypeptide does not share at least one characteristic sequence element with the reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, the variant polypeptide shares one or more biological activities of the reference polypeptide.In some embodiments, a variant polypeptide lacks one or more biological activities of a reference polypeptide. In some embodiments, a variant polypeptide exhibits a reduced level of one or more biological activities compared to the reference polypeptide. In many embodiments, a polypeptide of interest is considered a "variant" of a parent or reference polypeptide if the polypeptide of interest is identical to the parent amino acid sequence except for minor sequence changes at specific positions. Typically, less than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, or 2% of the residues in the variant are substituted compared to the parent. In some embodiments, a variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue compared to the parent. Often, a variant has a very small (e.g., less than 5, 4, 3, 2, or 1) number of substituted functional residues (i.e., residues responsible for a particular biological activity). Furthermore, mutations typically have no more than 5, 4, 3, 2, or 1 additions or deletions compared to the parent, although often there are no additions or deletions. Furthermore, any additions or deletions will typically be less than about 25, 20, 19, 18, 17, 16, 15, 14, 13, 10, 9, 8, 7, or 6 residues, and generally less than about 5, 4, 3, or 2 residues. In some embodiments, a mutation may also have one or more functional defects and / or may otherwise be considered a "mutant." In some embodiments, the parent or reference polypeptide is one found in nature. As will be understood by those skilled in the art, multiple mutations of a polypeptide of interest may be commonly found in nature, particularly when a particular polypeptide of interest is a polypeptide of an infectious agent.

[0049] II. Detailed Description of Specific Embodiments

[0072] In many nucleic acid enrichment techniques, it can be useful to first provide a shotgun library of nucleic acids, whereby longer nucleic acid sequences derived from a sample are subdivided into smaller fragments compatible with short-read sequencing techniques (i.e., approximately 50-500 nucleotides). To prepare a shotgun library, high-molecular-weight nucleic acid strands (typically cDNA or genomic DNA) are sheared into random fragments, optionally modified by ligation of common terminal sequences (i.e., adapters), and size-selected for downstream processing and analysis. For example, it can be useful to selectively capture a subset of nucleic acids in a shotgun library.

[0050]

[0073] Currently, two general categories of capture technologies exist: hybridization-based capture and amplification-based capture. Hybridization-based capture methods have the advantage of recovering the entire original shotgun library fragments, rather than replicating and recovering only a subset of the original library fragments. However, the on-target yield associated with hybridization-based capture is generally lower than that of amplification-based methods. In particular, a low on-target yield requires sequencing of off-target capture products, resulting in wasted sequencing capacity. Furthermore, the workflow associated with hybridization-based capture methods can be long and complex compared to amplification-based approaches. In contrast, amplification-based approaches, such as anchored multiplex PCR, offer the advantages of a simpler workflow, shorter turnaround time, and higher on-target yield compared to hybridization-based methods, but they also have several disadvantages. For example, target-specific primer sequences incorporated into the amplified library fragments result in wasted sequencing capacity. Furthermore, the library fragments are not necessarily representative of the original shotgun library because the template is necessarily truncated at the target-specific primer binding site. Thus, there remains an unmet need for a rapid and simple method of target enrichment that also accommodates unknown structural polymorphisms in the target sequence.

[0051]

[0074] These and other challenges can be overcome in the disclosed method for target enrichment by unidirectional dual-probe primer extension. In one aspect, the disclosure describes both a general approach for enrichment based on unidirectional dual-probe primer extension and improvements therefor. To this end, the disclosure provides a combination of primer extension and hybridization-based capture on a solid support for enrichment of one or more target nucleic acids from a library of target nucleic acids. The disclosure further provides an overall workflow that has many of the aforementioned advantages of anchor multiplex amplification-based enrichment methods and hybridization capture methods, but without many of the aforementioned disadvantages. Compared to many existing hybridization-based and anchor multiplex amplification-based capture methods, advantages of the disclosed kits, compositions, and methods include recovery of library molecules from whole shotgun molecules, a simple workflow (e.g., fewer overall steps and less hands-on time), reduced turnaround time, higher on-target rates, and lower overall material costs.

[0052]

[0075] In one embodiment, the present invention is a method for enriching at least one target nucleic acid in a library of nucleic acids. The method can include hybridizing a first oligonucleotide to a target nucleic acid in the library of nucleic acids. Each nucleic acid in the library of nucleic acids can have a first end comprising a first adaptor and a second end comprising a second adaptor. The method can further include extending the hybridized first oligonucleotide with a first polymerase, thereby generating a first primer extension complex comprising the target nucleic acid and the extended first oligonucleotide.

[0053]

[0076] In one embodiment, the method can further include capturing the first primer extension complex and enriching the first primer extension complex relative to a library of nucleic acids. In another embodiment, the method can include hybridizing a second oligonucleotide to the target nucleic acid and extending the hybridized second oligonucleotide with a second polymerase, thereby generating a second primer extension complex comprising the target nucleic acid and the extended second oligonucleotide, thereby releasing the extended first oligonucleotide from the first primer extension complex. The method can further include amplifying the target nucleic acid using a third polymerase, a first amplification primer, and a second amplification primer. The first amplification primer has a 3' end complementary to the first adapter and the second amplification primer has a 3' end complementary to the second adapter.

[0054]

[0077] The first, second, and third polymerases may be any suitable polymerase. One example of a polymerase is a Taq or Taq-derived polymerase (e.g., KAPA 2G polymerase from KAPA BIOSYSTEMS). Another example of a polymerase is a B-family DNA polymerase (e.g., KAPA HIFI polymerase from KAPA BIOSYSTEMS).

[0055]

[0078] In another embodiment, the present disclosure provides a kit for enriching at least one target nucleic acid in a nucleic acid library. The kit may include a first oligonucleotide complementary to the target nucleic acid in the nucleic acid library. Each nucleic acid in the nucleic acid library has a first end comprising a first adaptor and a second end comprising a second adaptor. The kit may further include a second oligonucleotide complementary to the target nucleic acid, a first amplification primer, and a second amplification primer. The first oligonucleotide may include a capture moiety. The second oligonucleotide is capable of hybridizing to the target nucleic acid at a 5' position relative to the first oligonucleotide. The first amplification primer has a 3' end complementary to the first adaptor, and the second amplification primer has a 3' end complementary to the second adaptor.

[0056]

[0079] In another embodiment, the present disclosure provides a kit for enriching at least one target nucleic acid in a nucleic acid library. The kit can include a first oligonucleotide complementary to the target nucleic acid in the nucleic acid library. Each nucleic acid in the nucleic acid library can have a first end comprising a first adaptor and a second end comprising a second adaptor. The kit can further include a modified nucleotide having a capture moiety, a second oligonucleotide complementary to the target nucleic acid, a first amplification primer, and a second amplification primer. The second oligonucleotide hybridizes to the target nucleic acid at a position 5' relative to the first oligonucleotide, the first amplification primer has a 3' end complementary to the first adaptor, and the second amplification primer has a 3' end complementary to the second adaptor.

[0057]

[0080] In yet another embodiment, the present disclosure provides a composition comprising a library of nucleic acids comprising at least one target nucleic acid. Each nucleic acid in the library of nucleic acids has a first end comprising a first adaptor, a second end comprising a second adaptor, and a region of interest intermediate the first adaptor and the second adaptor. The composition further comprises an extended first oligonucleotide hybridized to the region of interest of the target nucleic acid. The extended first oligonucleotide comprises at least one capture moiety. The composition further comprises a solid support bound to the at least one capture moiety, a second oligonucleotide hybridized to the target nucleic acid at a position 5' relative to the first extended oligonucleotide, and a polymerase associated with the 3' end of the second oligonucleotide.

[0058]

[0081] The methods of the present invention can be used as part of a sequencing protocol, including high-throughput single molecule sequencing protocols. The methods of the present invention generate a library of target nucleic acids to be sequenced. The target nucleic acids in the library may incorporate barcodes for molecular and sample identification.

[0059]

[0082] The present invention includes at least one linear primer extension step using target-specific primers. This linear extension step offers several advantages over the exponential amplification methods currently used in the art. Each target nucleic acid is characterized by a unique synthesis rate that depends on the rate at which the target-specific primer anneals and the rate at which the polymerase can read through the specific target sequence. Differences in extension and synthesis rates create biases that can result in small differences in a single round of synthesis. However, these small differences become exponentially amplified during PCR. The resulting gaps are called PCR biases. These biases can obscure any differences in the initial amounts of each sequence in the sample, preventing any quantitative analysis.

[0060]

[0083] The present invention limits the extension of target-specific primers (including gene-specific primers and degenerate primers that happen to be specific for binding sites within the genome) to a single step. All exponential amplification is performed using universal primers that are free from template-dependent bias or less biased than the target-specific primers.

[0061]

[0084] 1, a method 100 for target enrichment by unidirectional dual-probe primer extension includes preparing nucleic acid library fragments 102. In one aspect, the nucleic acid library fragments can be prepared from any nucleic acid source that includes one or more target nucleic acids. Generally, the target nucleic acids include regions or sequences of interest, and method 100 enables preferential enrichment of one or more target nucleic acids over non-target nucleic acids in a nucleic acid library for detection and analysis downstream of those regions or sequences of interest.

[0062]

[0085] Referring now to step 102, the nucleic acid is optionally fragmented and adapters are ligated to each end of the nucleic acid. Exemplary methods for preparing a library of nucleic acid fragments for use with the present disclosure include transposon-mediated fragmentation and labeling, mechanical shearing, enzymatic digestion, overhang (e.g., T / A) or blunt-end ligation, template-switching-mediated adapter ligation, and the like, and combinations thereof. Ultimately, the product of step 102, preparing nucleic acid library fragments, can result in a library of nucleic acids, wherein each nucleic acid in the library of nucleic acids has a first end comprising a first adapter and a second end comprising a second adapter. Notably, the first and second adapters can be the same or different and can take a variety of forms, including, but not limited to, forked or Y-shaped adapters having complementary and non-complementary portions, blunt-end adapters, overhang adapters, hairpin adapters, and the like, and combinations thereof. Generally, at least some of the aforementioned adapters are double-stranded; however, other adapter configurations may also be used in preparing libraries of nucleic acid fragments according to the present disclosure. Additionally, in the case of hairpin adapters, it may be useful to include a blocking element (e.g., a 3' dideoxynucleotide or a phosphate group) to prevent self-priming events.

[0063]

[0086] The next step 104 of method 100 can involve hybridizing a first oligonucleotide primer to a target nucleic acid present in the library of nucleic acids, thereby forming an unextended first primer-target complex. In one embodiment, the first oligonucleotide primer is a target-specific primer having a defined sequence that is complementary to the sequence of the target nucleic acid. An example of a target-specific primer is a gene-specific primer designed to hybridize to or near (e.g., upstream or 5' of) a gene of interest (e.g., cDNA, genomic DNA). The target nucleic acid can be RNA, DNA, or a combination thereof. The first oligonucleotide primer can be an oligonucleotide primer composed of ribonucleic acid, deoxyribonucleic acid, modified nucleic acid (e.g., biotinylated, locked nucleic acid, inosine, sila base, etc.), or other nucleic acid analogs known in the art.

[0064]

[0087] In various embodiments of the present disclosure, the first oligonucleotide primer may include one or more modified bases, a capture moiety, or a combination thereof. When the first oligonucleotide primer includes a capture moiety, the first oligonucleotide primer may be attached to a solid support or free in solution (i.e., not bound to or otherwise attached to a solid support) prior to step 104, in which the first oligonucleotide primer is hybridized to the target nucleic acid. In embodiments in which the first oligonucleotide primer including a capture moiety is not attached to a solid support via a capture moiety, step 104 can be performed in solution. In embodiments in which the first oligonucleotide primer including a capture moiety is attached to a solid support via a capture moiety, step 104 can be performed in situ. In particular, in the case of an in situ reaction, the resulting unextended primer-target complex is attached to the solid support. By separating the solution from the solid support to which the primer-target complex is bound, any remaining non-target nucleic acid in the solution or target nucleic acid not annealed to the first oligonucleotide primer can be removed.

[0065]

[0088] The next step 106 of method 100 involves performing a first primer extension reaction. In one aspect, step 106 involves extending the hybridized first oligonucleotide primer with a first polymerase. Following hybridization of the first oligonucleotide primer to the target nucleic acid template in step 104, the first oligonucleotide primer is extended by the first polymerase, thereby generating a first primer extension product or complex comprising a 3' region of the extended first oligonucleotide primer that comprises the reverse complement of at least a portion of the target nucleic acid template. As described herein, the hybridization and extension reactions are optionally performed simultaneously; however, in other embodiments, the hybridization and extension reactions may be performed separately (e.g., sequentially) and separated by a wash step that removes unannealed, uncaptured target nucleic acid from the reaction mixture. Additionally, step 104 may further include terminating the primer extension reaction to control the length of the extended first oligonucleotide primer. In particular, the length of the extended first oligonucleotide primer product can be actively controlled by techniques such as inactivating the polymerase added in step 104, or passively controlled by driving the reaction to completion, such as through the consumption of a limited amount of reactant, or by controlling / selecting the size of nucleic acid fragments in the library of nucleic acids in step 102 of method 100.

[0066]

[0089] Method 100 further includes step 108 of capturing the first primer extension complex. Capturing the first primer extension complex can be accomplished in various ways as disclosed herein and can be accomplished before, simultaneously with, or after either step 104 or step 106 of method 100. As described above, the first oligonucleotide can include a capture moiety that can be used to capture the first oligonucleotide primer onto a solid support before, during, or after step 104 or step 106 of method 100. In another example, extension of the first oligonucleotide primer following hybridization to the target nucleic acid includes incorporation of one or more modified nucleotides. The modified nucleotides can include a capture moiety, or downstream modification of the modified nucleotides can be configured to allow the capture moiety to be attached to or otherwise incorporated into the extended portion of the first primer extension complex. Thus, the first primer extension complex can be captured during or after step 106 by a capture moiety associated with one or more modified nucleotides. The choice of whether the target nucleic acid, the annealed primer-target complex, or the target-extended primer complex is captured further determines whether steps 104 and 106 of the method are performed in solution or in situ.

[0067]

[0090] The next step 110 of method 100 can involve enrichment of the first primer extension complexes. In one aspect, step 110 involves one or more purification and enrichment steps for recovery of the first primer extension complexes from non-target nucleic acids in the library and other molecules, such as unused reaction components (e.g., nucleotides, primer molecules, ATP, etc.), enzymes, buffers, etc. In some embodiments, step 110 involves enzymatic digestion, size-exclusion-based purification, affinity-based purification, etc., or a combination thereof. Notably, enrichment of the first primer extension products can be measured as enrichment relative to the entire library of nucleic acids. In one aspect, enrichment involves increasing the concentration of the target nucleic acid through depletion (i.e., removal) of other members of the library of nucleic acids that are not the target nucleic acid.

[0068]

[0091] The next step 112 of method 100 can involve hybridizing a second oligonucleotide primer to a target nucleic acid present in the library of nucleic acids. In one embodiment, the second oligonucleotide primer is a target-specific primer that binds to a region of interest within the target nucleic acid (rather than hybridizing to or being complementary to one or both of the first and second adapters). In another embodiment, the target nucleic acid is part of a first primer extension complex during step 112. For example, the second oligonucleotide primer can hybridize to the target nucleic acid at a position 5′ (i.e., upstream) relative to the extended first oligonucleotide primer in the first primer extension complex. In this case, the resulting unextended second primer-target complex includes the first extended oligonucleotide primer, the target nucleic acid hybridized to the first extended oligonucleotide primer, and the second (unextended) oligonucleotide primer. If the first primer extension product is attached to a solid support during step 112, the unextended second primer-target complex is similarly attached to the solid support. In other embodiments (e.g., after removing non-target nucleic acids from the reaction mixture), the first primer extension product is released from the solid support and in solution, allowing in-solution hybridization of the second oligonucleotide primer in step 112.

[0069]

[0092] The next step 114 of method 100 involves performing a second primer extension reaction. Following hybridization of the second oligonucleotide primer to the target nucleic acid template in step 112, the second oligonucleotide primer is coupled to a second polymer The second oligonucleotide primer is extended by a second polymerase, thereby generating a second primer extension product or complex containing the target nucleic acid. The extended second oligonucleotide primer comprises a 3' region that contains the reverse complement of at least a portion of the target nucleic acid template. In one embodiment, extension of the second oligonucleotide primer with a second polymerase releases the extended first oligonucleotide primer from the complex with the target nucleic acid. Releasing the extended first oligonucleotide from the first primer extension complex can involve one or more strand displacements (e.g., by a polymerase) or digestion (e.g., by a nuclease). For example, release of the extended first oligonucleotide can be achieved with an enzyme having at least one of strand displacement activity, 5' to 3' exonuclease activity, and flap endonuclease activity.

[0070]

[0093] As described herein, steps 112 and 114 are optionally performed simultaneously, although in other embodiments, steps 112 and 114 are performed separately (e.g., sequentially). Additionally, step 114 can further include terminating the primer extension reaction to control the length of the extended second oligonucleotide primer. In particular, the length of the extended second oligonucleotide primer product can be actively controlled, such as by inactivating the polymerase added in step 114, or passively controlled by driving the reaction to completion, such as through the consumption of a limited amount of reactant, or by controlling / selecting the size of nucleic acid fragments in a library of nucleic acids.

[0071]

[0094] If the extended first primer comprises one or more capture moieties attached to a solid support, release of the extended first oligonucleotide in step 114 results in a second primer extension complex that is free in solution rather than attached to a solid support. Thus, as described in step 110 of method 100, following step 114, one or more purification techniques can be implemented to recover the complex comprising the unbound second extension product or target nucleic acid from the first extended oligonucleotide primer attached to the support, the second polymerase, other reaction components, etc., and combinations thereof.

[0072]

[0095] Method 100 further includes an amplification step 116. Step 116 can include linear or exponential amplification (e.g., PCR). Generally, step 116 includes amplifying the target nucleic acid using a third polymerase, a first amplification primer, and a second amplification primer. In one embodiment, the first and second amplification primers are designed to be complementary to the sequences of the adapters incorporated into the target nucleic acid in the library of nucleic acids in step 102. For example, the first amplification primer can have a 3' end complementary to the first adapter, and the second amplification primer can have a 3' end complementary to the second adapter. However, the primers for amplification can include any sequence present in the target nucleic acid to be amplified (e.g., gene / target-specific primers, universal primers, etc.) and can support the synthesis of one or both strands (i.e., both the upper and lower strands of the double-stranded nucleic acid corresponding to the template for the amplification reaction).

[0073]

[0096] In some embodiments, step 116 allows for selective amplification of a target nucleic acid from a library of nucleic acids, rather than amplifying either the first or second extended oligonucleotide primer derived from the target nucleic acid. In one example, a uracil-compatible polymerase and dUTP are included in one or both of the extension reactions performed in step 106 and step 114. The extended oligonucleotide primer resulting from the reaction contains at least one uracil nucleotide, but the target nucleic acid template can be a DNA template without uracil nucleotides. A uracil-incompatible polymerase is then included in step 116 for amplification of the target nucleic acid. The uracil-incompatible polymerase is capable of amplifying uracil. However, a uracil-incompatible polymerase will be unable to replicate an extended oligonucleotide primer containing uracil. Alternatively, or in addition, uracil-containing products can be selectively digested or otherwise degraded, thereby leaving only the original molecule from the library of nucleic acids.

[0074]

[0097] After the step of amplification 116, method 100 can include a step 118 of analyzing the amplified target nucleic acid. Step 116 can include any method for determining the nucleic acid sequence of one or more products of method 100. Step 116 can further include sequence alignment, identification of sequence variations, measurement of unique primer extension products, etc., or a combination thereof.

[0075]

[0098] In addition to the elements of the present disclosure outlined in method 100, it may be useful to take into account several additional considerations when implementing the kits, compositions, and methods described herein. In one aspect, the primer hybridization step is mediated by a target-specific region of the primer. In some embodiments, the target-specific region can hybridize to a region of the gene located in an exon, intron, or untranslated portion of the gene, or a non-transcribed portion of the gene (e.g., a promoter or enhancer). In some embodiments, the gene is a protein-coding gene, while in other embodiments, the gene is not a protein-coding gene, such as an RNA-coding gene or a pseudogene. In still other embodiments, the target-specific region is located in an intergenic region. For mRNA or cDNA targets, the primer may include an oligo-dT sequence.

[0076]

[0099] Instead of a pre-designed target-specific region, the primer may contain a degenerate sequence (i.e., a string of randomly incorporated nucleotides). Such a primer can also find a binding site in the genome and act as a target-specific primer for that binding site. In particular, fully degenerate primers, in which each nucleotide position is degenerate, may not be useful for targeted enrichment. However, partially degenerate primers, in which only a portion of the nucleotide positions are degenerate, may be useful for use according to the present disclosure. For example, a primer with partial degeneracy at a single nucleotide position may be useful for capturing a target sequence containing one or more single nucleotide polymorphisms (SNPs).

[0077]

[0100] In addition to the target-specific region, the primer may contain additional sequences. In some embodiments, these sequences are located at the 5' end of the target-specific region. In other embodiments, these sequences may be included elsewhere in the primer, as long as the target-specific region is capable of hybridizing to the target and driving the primer extension reaction, as described below. The additional sequences in the primer may include one or more barcode sequences, such as a unique molecular identification sequence (UID) or a multiplex sample identification sequence (MID). The barcode sequence may be present as a single sequence or as two or more sequences.

[0078]

[0101] In some embodiments, the additional sequence comprises a sequence that facilitates ligation to the 5' end of the primer. The primer may also comprise a universal ligation sequence that allows for adapter ligation, as described in the next section.

[0079]

[0102] In some embodiments, the additional sequence comprises one or more binding sites for one or more universal amplification primers.

[0103] The primer extension step is carried out by a nucleic acid polymerase. Depending on the type of nucleic acid being analyzed, the polymerase may be a DNA-dependent DNA polymerase ("DNA polymerase") or an RNA-dependent DNA polymerase ("reverse transcriptase").

[0080]

[0104] In some embodiments, it is desirable to control the length of the nucleic acid strand synthesized in the primer extension reaction. As described below, this strand length determines the length of the nucleic acid that is subjected to subsequent steps of the method and any downstream applications. The extension reaction can be stopped by any method known in the art. For example, the reaction may be physically stopped by a temperature shift or the addition of a polymerase inhibitor. In some embodiments, the reaction is stopped by placing the reaction on ice. In other embodiments, the reaction is stopped by increasing the temperature to inactivate a thermolabile polymerase. In still other embodiments, the reaction is stopped by the addition of a chelating agent such as EDTA, which can sequester a critical cofactor of the enzyme, or another chemical or biological compound that can reversibly or irreversibly inactivate the enzyme.

[0081]

[0105] Alternative methods for controlling the length of primer extension products include limiting critical components (e.g., dNTPs) to directly limit extension length, or adding Mg to slow the rate of extension and improve the ability to control the extension termination point. 2+ The purpose of the present invention is to limit the amount of primer extension to a minimum, thereby starving the extension reaction. Those skilled in the art can experimentally or theoretically determine the appropriate amounts of the critical components that allow limited primer extension to produce products primarily of the desired length.

[0082]

[0106] Another method for controlling the length of primer extension products is the addition of terminator nucleotides, including reversible terminator nucleotides. Those skilled in the art can experimentally or theoretically determine the appropriate ratio of terminator and non-terminator nucleotides that allows limited primer extension to produce products of primarily the desired length. Examples of terminator nucleotides include dideoxynucleotides, 2'-phosphate nucleotides as described in U.S. Pat. No. 8,163,487 to Gelfand et al., 3'-O-blocked reversible terminators, and 3'-unblocked reversible terminators as described, for example, in U.S. Pat. App. Pub. No. 2014 / 0242579 to Zhuo et al. and Guo, J. et al., "Four-color DNA sequencing with 3'-O-modified nucleotide reversible terminators and chemically cleavable fluorescent dideoxynucleotides," PNAS 2008, 105(27), 9145-9150. Another method for controlling the length of primer extension products is to add a limited amount of uracil (dUTP) to the primer extension reaction. The uracil-containing DNA is then treated with uracil N-DNA glycosylase to generate abasic sites. DNA with abasic sites can be decomposed by heat treatment, optionally with the addition of alkali to improve the efficiency of decomposition, as described in U.S. Pat. No. 8,669,061 to Gupta et al. Those skilled in the art can experimentally or theoretically determine the appropriate ratio of dUTP to dTTP in an extension reaction that includes limited dUTP to primarily produce products of the desired length upon endonuclease treatment.

[0083]

[0107] In some embodiments, the length of the extension product is essentially limited by the length of the input nucleic acid. For example, cell-free DNA present in maternal plasma is less than 200 bp in length, with the majority being 166 bp in length. Yu, S.C.Y. et al., Size-based molecular diagnostics using plasma DNA for Noninvasive prenatal testing, PNAS USA 2014;111(23):8583-8. The median length of cell-free DNA found in plasma from healthy individuals and cancer patients is approximately 185-200 bp. Giacona, MB et al., Cell-free DNA in human blood plasma: length measurements in patients with pancreatic cancer and healthy controls, Pancreas 1998;17(1):89-97. Poorly preserved or chemically processed samples may contain chemically or physically degraded nucleic acids. For example, formalin-fixed, paraffin-embedded tissue (FFPET) typically produces nucleic acids with an average length of 150 bp.

[0084]

[0108] In some embodiments, the methods of the present invention include one or more purification steps after primer extension with a DNA polymerase or reverse transcriptase. Purification removes unused primer molecules and template molecules used to create the primer extension products. In some embodiments, all nucleic acid fragments other than the template nucleic acid and the extended primer are removed by exonuclease digestion. In such embodiments, the primer used in primer extension may have a 5'-end modification that renders the primer and any extension products resistant to exonuclease digestion. Examples of such modifications include phosphorothioate linkages. In other embodiments, the RNA template can be removed by enzymatic treatment that leaves DNA behind, such as RNase digestion, including RNase H digestion. In yet other embodiments, the primers and large template DNA are separated from the extension products by size exclusion methods, such as gel electrophoresis, chromatography, isotachophoresis, or epitachophoresis.

[0085]

[0109] In some embodiments, purification is by affinity binding. In a variation of this embodiment, the affinity is for a specific target sequence (sequence capture). In other embodiments, the primer includes an affinity tag. Any affinity tag known in the art can be used, such as biotin or an antibody or antigen for which a specific antibody exists. The affinity partner of the affinity tag can be present in solution, e.g., on a solution-phase solid support such as suspended particles or beads, or can be attached to a solid-phase support. During affinity purification, unbound components of the reaction mixture are washed away. In some embodiments, additional steps are taken to remove unused primers. In some embodiments, affinity capture alters the charge of the primer extension product. For example, the inclusion of one or more biotinylated nucleotides and the binding of streptavidin to it creates a charge change in the nascent nucleic acid strand. The altered charge can be used to separate the nascent strand (primer extension product) by isotachophoresis or epitachophoresis.

[0086]

[0110] Notably, the disclosed methods do not require a ligation step (e.g., adding a common sequence to the extended first or second oligonucleotide primer). However, in some embodiments, the invention includes a ligation step. For example, a homopolymer tail can be added to the 3' end of the nucleic acid. In this embodiment, the homopolymer can serve as a binding site for the reverse complement homopolymer (similar to a polyA tail using a polyT primer for mRNA). Ligation appends one or more adapter sequences to the primer extension product generated in the previous step. The adapter sequence provides one or more universal priming sites (for amplification or sequencing) and, optionally, one or more barcodes. The exact mode of adapter ligation is not important, as long as the adapter is attached to the primer extension product and subsequent steps, described below, are possible.

[0087]

[0111] In some embodiments described above, the method includes a target-specific primer that includes a universal priming sequence ("priming site") and generates a primer extension product with a single priming site. In such embodiments, only one additional priming sequence ("priming site") needs to be provided to allow exponential amplification. In other embodiments, the target-specific primer does not include a universal priming site. In such embodiments, two priming sites need to be provided to allow exponential amplification. The adapter with the universal priming site may be added by any single-stranded ligation method available in the art. good.

[0088]

[0112] An example of a single-stranded ligation method can be used in embodiments in which the extended primer contains a universal ligation site. In such embodiments, an adapter having a double-stranded region and a single-stranded overhang complementary to the universal ligation site of the primer can be annealed and ligated. Annealing the single-stranded 3' overhang of the adapter to the universal ligation site at the 5' end of the primer generates a nick double-stranded region in the strand containing the primer extension product. The two strands can be ligated at the nick by DNA ligase or another enzyme or non-enzymatic reagent capable of catalyzing the reaction between the 5'-phosphate of the primer extension product and the 3'-OH of the adapter. By joining the adapter, ligation provides a universal priming site at one end of the primer extension product.

[0089]

[0113] Another example of a single-stranded ligation method can be used to add universal priming sites to opposite ends of primer extension products (or to both ends of extension products in embodiments where the extended primer does not contain a universal ligation site). In this embodiment, one or both ends of the primer extension products to be ligated do not have a universal ligation site. Furthermore, in some embodiments, at least one end of the primer extension products to be ligated has an unknown sequence (e.g., due to a random termination event or unknown sequence variation). In such embodiments, a sequence-independent single-stranded ligation method is used. An exemplary method is described in U.S. Patent Application Publication No. 20140193860. Essentially, this method uses a population of adapters whose single-stranded 3' overhangs have random sequences, e.g., random hexamer sequences, instead of universal ligation sites. In some embodiments of the method, the adapters also have hairpin structures. Another example is the method enabled by the ACCEL-NGS 1S DNA Library Kit (Swift Biosciences, Ann Arbor, Mich.).

[0090]

[0114] The ligation step of the method utilizes a ligase or another enzyme with similar activity, or a non-enzymatic reagent. The ligase can be, for example, a DNA or RNA ligase of viral or bacterial origin, such as T4 or E. coli ligase, or the thermostable ligases Afu, Taq, Tfl, or Tth. In some embodiments, alternative enzymes, such as topoisomerases, can be used. Additionally, non-enzymatic reagents can be used to form a phosphodiester bond between the 5'-phosphate of the primer extension product and the 3'-OH of the adapter, as described and referenced in U.S. Patent Application Publication No. 2014 / 0193860 to Bevilacqua et al.

[0091]

[0115] In some embodiments of the method, the first ligation of the adapter is followed by an optional primer extension. The ligated adapter has a free 3' end that can be extended to generate a double-stranded nucleic acid. The opposite end of the adapter then becomes suitable for blunt-end ligation of another adapter. This double-stranded end of the molecule can be ligated to the double-stranded adapter by any ligase or other enzymatic or non-enzymatic means, avoiding the need for a single-stranded ligation procedure. The double-stranded adapter sequence provides one or more universal priming sites (for amplification or sequencing) and, optionally, one or more barcodes.

[0092]

[0116] In some embodiments, the methods of the present invention include one or more purification steps after the ligation step. Purification removes unused adapter molecules. Adapters and large ligation products are separated from the extension products by size exclusion methods, such as gel electrophoresis, chromatography, or isotachophoresis.

[0093]

[0117] In some embodiments, purification is by affinity binding. In a variation of this embodiment, the affinity is for a specific target sequence (sequence capture). In other embodiments, the adapter comprises an affinity tag. Any affinity tag known in the art can be used (e.g., biotin or an antibody or antigen for which a specific antibody exists). The affinity partner of the affinity tag can be associated with a solution-phase support (e.g., on a suspended particle or bead) or can be attached to a solid support. During affinity purification, unbound components of the reaction mixture are washed away. In some embodiments, additional steps are taken to remove unused adapters.

[0094]

[0118] In some embodiments, the present invention includes an amplification step. This step can include linear or exponential amplification (e.g., PCR). Primers for amplification can include any sequence present in the nucleic acid to be amplified and can support the synthesis of one or both strands. Amplification can be isothermal or involve thermocycling.

[0095]

[0119] In some embodiments, amplification is exponential and involves PCR. It is desirable to reduce PCR amplification bias. When one or more gene-specific primers are used to reduce bias, a limited number of amplification cycles (e.g., about 10 or fewer cycles) are involved in the method. In other variations of these embodiments, a universal primer is used to synthesize both strands. The universal primer sequence may be part of the original extension primer of one or both of the ligated adapters. One or two universal primers can be used. The extension primer and one or both adapters described above can be engineered to have the same primer binding site. In that embodiment, a single universal primer can be used to synthesize both strands. In other embodiments, the extension primer (or adapter) on one side of the amplified molecule and the adapter on the other side contain different universal primer binding sites. A universal primer may be paired with another universal primer (of the same or different sequence). In other embodiments, a universal primer may be paired with a gene-specific primer. Because sequence bias is reduced in PCR using universal primers, the number of amplification cycles need not be limited to the same extent as in PCR using gene-specific primers. The number of amplification cycles when universal primers are used can be low, but can also be as high as about 20, 30 or more.

[0096]

[0120] The present invention involves the use of molecular barcodes. Barcodes typically consist of 4-36 nucleotides. In some embodiments, barcodes are designed to have melting temperatures of 10°C or less with respect to each other. Barcodes can be designed to form minimal cross-hybridizing sets, i.e., combinations of sequences that form as few stable hybrids as possible with each other under the desired reaction conditions. The design, placement, and use of barcodes for sequence identification and enumeration are known in the art. See, e.g., U.S. Patent Nos. 7,393,665, 8,168,385, 8,481,292, 8,685,678, and 8,722,368.

[0097]

[0121] Barcodes can be used to identify each nucleic acid molecule in a sample and its progeny (i.e., the set of nucleic acid molecules produced using the original nucleic acid molecule). Such barcodes are "unique IDs" (UIDs).

[0098]

[0122] Barcodes can also be used to identify the sample from which the analyzed nucleic acid molecule originates. Such a barcode is a "multiplex sample ID" ("MID"). All molecules from the same sample share the same MID.

[0099]

[0123] Barcodes contain a unique sequence of nucleotides characteristic of each barcode. In some embodiments, the sequence of the barcode is pre-designed. In other embodiments, the barcode sequence is random. All or some of the nucleotides in a barcode can be random. Random sequences and random nucleotide bases within a known sequence are called "degenerate sequences" and "degenerate bases," respectively. In some embodiments, a molecule contains two or more barcodes: one for molecular identification (UID) and one for sample identification (MID). Sometimes, a UID or MID each contains several barcodes, which, taken together, allow for the identification of the molecule or sample.

[0100]

[0124] In some embodiments, the number of UIDs in a reaction can exceed the number of molecules being labeled. In some embodiments, one or more barcodes are used to group or bin sequences. For example, in some embodiments, one or more UIDs are used to group or bin sequences, and sequences in each bin contain the same UID, i.e., are amplicons derived from a single target molecule. In some embodiments, UIDs are used to align sequences. In other embodiments, target-specific regions are used to align sequences. In some embodiments of the present invention, UIDs are introduced into the initial primer extension event, while sample barcodes (MIDs) are introduced into the ligated adapters.

[0101]

[0125] After ligation, the nucleic acid product can be sequenced. Sequencing can be performed by any method known in the art. High-throughput single-molecule sequencing is particularly advantageous. Examples of such technologies include the 454 LIFE SCIENCES GS FLX platform (454 LIFE SCIENCES), the ILLUMINA HISEQ platform (ILLUMINA), the ION TORRENT platform (LIFE TECHNOLOGIES), the PACIFIC BIOSCIENCES platform utilizing SMRT sequencing technology (PACIFIC BIOSCIENCES), and any other currently existing or future single-molecule sequencing technology with or without sequencing by synthesis. In variations of these embodiments, sequencing utilizes universal primer sites present in one or both adapter sequences or one or both primer sequences. In yet other variations of these embodiments, gene-specific primers are used for sequencing. However, it should be noted that universal primers are associated with reduced sequencing bias compared to gene-specific primers.

[0102]

[0126] In some embodiments, the sequencing step comprises sequence alignment. In some embodiments, alignment is used to determine a consensus sequence from multiple sequences, e.g., multiple sequences with the same unique molecular ID (UID). In some embodiments, alignment is used to identify sequence variations, such as single nucleotide variations (SNVs). In some embodiments, a consensus sequence is determined from multiple sequences that all have the same UID. In other embodiments, UIDs are used to eliminate artifacts, i.e., variations present in the progeny of a single molecule (characterized by a particular UID). Such artifacts resulting from PCR or sequencing errors can be eliminated using UIDs.

[0103]

[0127] In some embodiments, the number of each sequence in a sample can be quantified by quantifying the relative number of sequences with each UID among a population with the same multiplex sample ID (MID). Each UID represents a single molecule in the original sample, and counting the different UIDs associated with each sequence variant allows the proportion of each sequence variant in the original sample in which all molecules share the same MID to be determined. One skilled in the art can identify consensus sequences. One can determine the number of sequence reads required for determination. In some embodiments, the relevant number is the reads per UID ("sequencing depth") that is required for accurate quantitative results. In some embodiments, the desired depth is 5-50 reads per UID.

[0104]

[0128] The samples used in the methods of the present invention include any individual (e.g., human, patient) or environmental sample containing nucleic acid. Polynucleotides can be extracted from the sample, or the sample can be directly subjected to the methods of the present invention. The starting sample can also be extracted or isolated nucleic acid, DNA, or RNA. The sample can constitute any tissue or body fluid obtained from an organism. For example, the sample can be a tumor biopsy or a blood or plasma sample. In some embodiments, the sample is a formalin-fixed, paraffin-embedded (FFPE) sample. The sample can contain nucleic acid from one or more sources, for example, one or more patients. In some embodiments, the tissue can be infected with a pathogen and therefore contain host and pathogen nucleic acid.

[0105]

[0129] Methods for DNA extraction are well known in the art. See J. Sambrook et al., "Molecular Cloning: A Laboratory Manual," 1989, 2nd ed., Cold Spring Harbor Laboratory Press: New York, NY. A variety of kits for extracting nucleic acids (DNA or RNA) from biological samples are commercially available (e.g., BD BIOSCIENCES CLONTECH (Palo Alto, Calif.), EPICENTRE TECHNOLOGIES (Madison, Wisc.); GENTRA SYSTEMS, INC. (Minneapolis, Minn.); and QIAGEN, INC. (Valencia, Calif.), AMBION, INC. (Austin, Tex.); BIORAD LABORATORIES (Hercules, Calif.); etc.).

[0106]

[0130] In some embodiments, the starting sample used in the methods of the invention is a library, such as a genomic or expression library, containing a plurality of polynucleotides. In other embodiments, the library is generated by the methods of the invention. When the starting material is a biological sample, the methods generate an amplified library, or a collection of amplicons representative of diversity or sequences. The library can be stored and used multiple times for further amplification or sequencing of the nucleic acids within the library.

[0107]

[0131] According to one embodiment of the present disclosure, a method for primer extension target enrichment can include in-solution primer-mediated capture of target nucleic acids. Referring now to Figures 2A-2E, a library of nucleic acids includes target nucleic acids 200 including a region of interest (ROI) 202 (Figure 2A). The target nucleic acid 200 further includes a first end including a first adaptor 204 and a second end including a second adaptor 206. In Figures 2A-2E, the target nucleic acid 200 is depicted as a single-stranded nucleic acid with the first adaptor 204 located at the 3' end (i.e., first end) of the target nucleic acid 200 and the second adaptor 206 located at the 5' end (i.e., second end) of the target nucleic acid 200. A first oligonucleotide 208 hybridizes to the target nucleic acid 200 in the library of nucleic acids. The first oligonucleotide 208 includes a 3' target-specific region 210 complementary to the target nucleic acid and a capture portion 212. In the illustrated embodiment, the target-specific region 210 is complementary to the ROI 202 .

[0108]

[0132] As shown in FIG. 2B, the hybridized first oligonucleotide 208 is extended with a first polymerase (not shown), thereby generating a first primer extension complex 214 comprising the target nucleic acid 200 and an extended first oligonucleotide 216 (the extended portion of the extended first oligonucleotide 216 is shown with a dashed line). The first primer extension complex 214 is captured on a solid support 218. The solid support can be a solution-phase support (e.g., a bead or another similar particle) or a solid-phase support (e.g., a silicon wafer, a glass slide, etc.). For example, the magnetic glass particles and devices using same described in U.S. Patent Nos. 656,568, 6,274,386, 7,371,830, 6,870,047, 6,255,477, 6,746,874, and 6,258,531 can be used. 2B, first primer extension complexes 214 are captured on a solid support via capture moieties 212. Following capture, first primer extension complexes 214 are enriched against a library of nucleic acids.

[0109]

[0133] 2C, a second oligonucleotide 220 is hybridized to the target nucleic acid 200. The second oligonucleotide 220 is complementary to the target nucleic acid 200 and hybridizes to the target nucleic acid 200 at a position 5′ relative to the target-specific region 210 of the first oligonucleotide 208. In the illustrated embodiment, the second oligonucleotide 220 is complementary to and hybridizes to the target nucleic acid 200 at a position just outside of the ROI 202. However, it will be appreciated that the first oligonucleotide 208 and the second oligonucleotide 220 can be designed to hybridize at any defined position along the length of the target nucleic acid 200, such that the second oligonucleotide 220 hybridizes to the target nucleic acid 200 at a position 5′ relative to the target-specific region 210 of the first oligonucleotide 208. From FIG. 2C, it can be seen that both the first extended oligonucleotide 216 (attached to the solid support 218) and the second oligonucleotide 220 are hybridized to the target nucleic acid 200.

[0110]

[0134] 2D , the hybridized second oligonucleotide 220 is extended with a second polymerase (not shown), thereby generating a second primer extension complex 222 comprising the target nucleic acid 200 and an extended second oligonucleotide 224 (the extended portion of the extended second oligonucleotide 224 is shown in dashed lines). In one embodiment, the extension of the hybridized second oligonucleotide 220 releases the extended first oligonucleotide 216 from the first primer extension complex 214. In another embodiment, the extended first oligonucleotide 216 (including the first oligonucleotide primer 208) remains attached to the solid support 218.

[0111]

[0135] 2E, target nucleic acid 200 is amplified with a third polymerase (not shown), a first amplification primer 226, and a second amplification primer 228. First amplification primer 226 includes a 3' end complementary to first adaptor 204, and second amplification primer 228 includes a 3' end complementary to second adaptor 206.

[0112]

[0136] According to another embodiment of the present disclosure, a method for primer extension target enrichment can include in situ primer-mediated capture of a target nucleic acid. Referring to Figures 3A and 3B, a library of nucleic acids includes a target nucleic acid 300 that includes a region of interest (ROI) 302 (Figure 3A). The target nucleic acid 300 further includes a first end including a first adaptor 304 and a second end including a second adaptor 306. In Figures 3A and 3B, the target nucleic acid 300 is illustrated as a single-stranded nucleic acid, with the first adaptor 304 located at the 3' end (i.e., the first end) of the target nucleic acid 300 and the second adaptor 306 located at the 5' end (i.e., the second end) of the target nucleic acid 300. A first oligonucleotide 308 is hybridized to the target nucleic acid 300 in the library of nucleic acids. The first oligonucleotide 308 includes a 3' target-specific region 310 that is complementary to the target nucleic acid 300 and a capture portion 312. In the illustrated embodiment, the target-specific region 310 is complementary to the ROI 302 .

[0113]

[0137] Compared to the embodiment shown in FIGS. 2A-2E, the first oligonucleotide 308 The first oligonucleotide 308 is captured on a solid support 318 prior to or simultaneously with hybridization of the first oligonucleotide 308 to the target nucleic acid 300. The solid support 318 may be a solution-phase support (e.g., beads or other similar particles) or a solid-phase support (e.g., a silicon wafer, glass slide, etc.). In the embodiment shown in FIG. 3A, the first oligonucleotide 308 is captured on the solid support 318 via the capture moiety 312. Referring to FIG. 2B, the hybridized first oligonucleotide 308 is extended with a first polymerase (not shown), thereby generating a first primer extension complex 314 comprising the target nucleic acid 300 and the extended first oligonucleotide 316 (the extended portion of the extended first oligonucleotide 316 is shown with a dashed line). Notably, the first primer extension complex 314 is captured on the solid support 318, allowing for enrichment of the target nucleic acid 300 relative to the library of nucleic acids. A second primer hybridization and extension reaction can then be performed as illustrated in Figures 2C and 2D, followed by an amplification step as illustrated in Figure 2E.

[0114]

[0138] According to yet another embodiment of the present disclosure, a method for primer extension target enrichment may include extension-mediated capture of a target nucleic acid. Referring to Figures 4A-4D, a library of nucleic acids includes a target nucleic acid 400 including a region of interest (ROI) 402 (Figure 4A). The target nucleic acid 400 further includes a first end including a first adaptor 404 and a second end including a second adaptor 406. In Figures 4A-4D, the target nucleic acid 400 is illustrated as a single-stranded nucleic acid, with the first adaptor 404 located at the 3' end (i.e., the first end) of the target nucleic acid 400 and the second adaptor 406 located at the 5' end (i.e., the second end) of the target nucleic acid 400. A first oligonucleotide 408 is hybridized to the target nucleic acid 400 in the library of nucleic acids. The first oligonucleotide 408 is complementary to the target nucleic acid 400. In particular, first oligonucleotide 408 does not necessarily include a capture portion, as compared to first oligonucleotide 208 of Figure 2A, which includes capture portion 212. In the embodiment shown in Figure 4A, first oligonucleotide 408 is complementary to a portion of ROI 402.

[0115]

[0139] As shown in FIG. 4B, the hybridized first oligonucleotide 408 is extended with a first polymerase (not shown), thereby generating a first primer extension complex 414 comprising the target nucleic acid 400 and an extended first oligonucleotide 416 (the extended portion of the extended first oligonucleotide 416 is shown in dashed lines). According to the embodiment illustrated in FIGS. 4A-4D, extension of the first oligonucleotide 408 is performed in the presence of one or more modified nucleic acids 412. Each modified nucleic acid includes a capture moiety 412a or can be modified to add a capture moiety 412a simultaneously with or after extension of the first oligonucleotide 416. Incorporation of one or more modified nucleic acids 412 including a capture moiety 412a enables extension-mediated capture of the target nucleic acid 400 on a solid support 418. Solid support 418 can be a solution phase support (e.g., beads or other similar particles) or a solid phase support (e.g., a silicon wafer, glass slide, etc.). In the embodiment illustrated in Figure 4B, first primer extension complexes 414 are captured on solid support 418 via modified nucleic acids 412 that include capture moieties 412a. Following capture, first primer extension complexes 414 are enriched against a library of nucleic acids.

[0116]

[0140] 4C , a second oligonucleotide 420 is hybridized to the target nucleic acid 400. The second oligonucleotide 420 is complementary to the target nucleic acid 400 and hybridizes to the target nucleic acid 400 at a position 5′ relative to the first oligonucleotide 408. In the illustrated embodiment, the second oligonucleotide 420 is complementary to and hybridizes to the target nucleic acid 400 at a position just inside the ROI 402. However, it will be appreciated that the first oligonucleotide 408 and the second oligonucleotide 420 can be designed to hybridize at any defined position along the length of the target nucleic acid 400, such that the second oligonucleotide 420 hybridizes to the target nucleic acid 400 at a position 5′ relative to the target-specific region 410 of the first oligonucleotide 408.

[0117]

[0141] 4C , the hybridized second oligonucleotide 420 is extended with a second polymerase (not shown), thereby generating a second primer extension complex 422 comprising the target nucleic acid 400 and an extended second oligonucleotide 424 (the extended portion of the extended second oligonucleotide 424 is shown with a dashed line). Prior to extension of the second oligonucleotide 420 with the second polymerase, the first extended oligonucleotide 416 (attached to a solid support 418) and the second oligonucleotide 420 are each hybridized to the target nucleic acid 400. Extension of the hybridized second oligonucleotide 420 releases the extended first oligonucleotide 416 from the first primer extension complex 414. In another embodiment, the extended first oligonucleotide 416 (comprising the first oligonucleotide primer 408 and the modified nucleic acid 412) remains attached to the solid support 418.

[0118]

[0142] 4D, target nucleic acid 400 is amplified with a third polymerase (not shown), a first amplification primer 426, and a second amplification primer 428. First amplification primer 426 includes a 3' end complementary to first adaptor 404, and second amplification primer 428 includes a 3' end complementary to second adaptor 406.

[0119]

[0143] In one embodiment, the target and non-target nucleic acids in the nucleic acid library may exhibit intermolecular interactions that result in a daisy-chain structure. As shown in FIG. 5, target nucleic acid 200 (see also FIG. 2A) includes an ROI, a first adaptor 204, and a second adaptor 206. A first oligonucleotide 208 is hybridized to target nucleic acid 200. First oligonucleotide 208 includes a 3' target-specific region 210 and a capture portion 212. The nucleic acid library may further include one or more non-target nucleic acids, including a first non-target nucleic acid 500 and a second non-target nucleic acid 500'. Similar to target nucleic acid 200, first non-target nucleic acid 500 and second non-target nucleic acid 500' each include a first end including first adaptor 504 and 504', respectively, and a second end including second adaptor 506 and 506', respectively. In one embodiment, first adaptor 204 is at least partially complementary to first adaptor 504, and second adaptor 506 is at least partially complementary to second adaptor 506'. Thus, as illustrated in Figure 5, target nucleic acid 200 can form a daisy chain with non-target nucleic acid 500 and non-target nucleic acid 500'.

[0120]

[0144] In various situations, it may be useful to minimize or eliminate the formation of daisy chain structures. For example, capture of target nucleic acid 200 by hybridization and extension of first oligonucleotide 208 may result in capture of non-target nucleic acid 500 and non-target nucleic acid 500 via association, thereby reducing the specificity of the capture and enrichment method. To reduce intermolecular interactions between the adaptor ends of target and non-target nucleic acids in a library of nucleic acids, a blocking oligonucleotide can be hybridized to the adaptor end sequence.

[0121]

[0145] To help reduce off-target hybridization, the blocking oligonucleotides have sequences complementary to the adapters (e.g., first adapter 204 and second adapter 206) and preferentially hybridize to these adapter sequences. Blocking oligos can be used in both single-plex and multiplex formats. When multiplexing is desired, different sample index sequences can be incorporated into the adapters. However, this requires the use of matching blocking oligonucleotides. When using multiple sample indexes (e.g., 24, 96, etc.), one possibility is to use a single "universal" blocking oligonucleotide. The universal blocking oligonucleotide has a unique sequence containing non-natural nucleotides that can bind to multiple different sample index sequences. As a result, only a single blocking oligonucleotide is added to the nucleic acid sample. Alternatively (or additionally), the single universal blocking oligonucleotide can be a mixture of oligonucleotides that collectively constitute the universal blocking oligonucleotide composition.

[0122]

[0146] In one embodiment, the universal blocking oligonucleotide comprises a non-specific region adjacent to the first and second specific regions. The non-specific region comprises, for example, a series of inosines that align with the sample index sequence when the universal blocking oligonucleotide hybridizes to the target adapter sequence. The specific region of the universal blocking oligonucleotide is complementary to the constant portion of the adapter sequence and is located between the T of the blocking oligonucleotide adapter duplex. m one or more melting temperatures (T m ) containing modified bases. m Examples of modified base substitutions are shown in Table 1.

[0123]

[0147]

[0124] [Table 1]

[0125]

[0148] In another embodiment, unamplified nucleic acid libraries prepared with two different adapter sequences may be processed without a blocking oligonucleotide, provided that the adapter ends do not hybridize to each other. Adapter types suitable for this approach include forked and Y-shaped adapters. [Example]

[0126] Example 1 Primer extension target enrichment by primer-mediated capture in solution (PETE-Cap)

[0149] Primer extension target enrichment with in-solution primer-mediated capture was performed according to the following protocol. Duplicate nucleic acid libraries were prepared from 10 ng and 100 ng of NA12878 human genomic DNA (CORIELL) using the KAPA HYPERPLUS Library Preparation Kit according to the manufacturer's instructions, including a 0.8X post-ligation cleanup step (Figure 6A). Target nucleic acids in the nucleic acid libraries were then enriched by primer extension target enrichment with in-solution primer-mediated capture according to the embodiment illustrated in Figure 2. Primers complementary to the plus or minus strand of the target nucleic acid were designed for the same exon of each gene of interest (i.e., target). The first (inner) oligonucleotide primer was 20–25 nucleotides long, and the second (outer) oligonucleotide primer was 50–60 nucleotides long. The additional length of the second oligonucleotide primer (compared to the first) was due to the inclusion of a 5' non-complementary tail sequence. In particular, the total length of the second oligonucleotide primer can be reduced by omitting the 5' non-complementary tail sequence.

[0127]

[0150] The first oligonucleotide (internal) primer hybridization and extension reactions were set up according to Table 2. The nucleic acid library consisted of unamplified products prepared with the KAPA HYPERPLUS Library Preparation Kit described above. A 0.8X post-ligation cleanup step was included in the reaction, and the total amount of nucleic acid library recovered after elution was determined; the final concentration of the nucleic acid library was not determined (nd). The mastermix consisted of a custom KAPA 2G polymerase PCR master mix. The primer mixture consisted of a set of 377 first oligonucleotide target-specific internal primers present at equimolar concentrations. Notably, each first oligonucleotide target-specific internal primer contained a 5' biotin capture moiety.

[0128]

[0151]

[0129] [Table 2]

[0130]

[0152] A first oligonucleotide primer was hybridized to a target nucleic acid in a library of nucleic acids and extended with a polymerase for a total of about 1 hour according to the thermal profile in Table 3. Notably, the protocol in Table 3 omits the use of thermal cycling.

[0131]

[0153]

[0132] [Table 3]

[0133]

[0154] Following hybridization and extension with a biotinylated first oligonucleotide primer, samples were transferred to DYNABEADS MYONE Streptavidin T1 capture beads (THERMO FISHER SCIENTIFIC) at a 1:1 ratio. The capture beads were mixed with 1× binding and wash buffer and resuspended in 2× binding and wash buffer before adding to the DNA sample. The composition of the binding and wash buffer is listed in Table 4.

[0134]

[0155]

[0135] [Table 4]

[0136]

[0156] The samples were incubated with 50 μL of MYONE capture beads for 10 minutes at room temperature on an automated sample rotator. Once the biotinylated DNA bound to the beads, the samples were placed on a magnet for 3 minutes to capture the beads, and the supernatant was removed and discarded. The beads were washed twice, once with 1X binding and wash buffer as described in Table 3 and once with 10 mM Tris-HCl, pH 8.0, to remove non-biotinylated DNA. The beads were then resuspended in 20 μL of 10 mM Tris-Cl, pH 8.0.

[0137]

[0157] The resuspended beads were added to the second oligonucleotide (outer) primer hybridization reaction mixture according to Table 5.

[0158]

[0138] [Table 5]

[0139]

[0159] The reaction mixtures listed in Table 5 were incubated at 55° C. for 165 minutes to enhance the specificity of target capture and allow the second oligonucleotide primer to hybridize to the target nucleic acid in the library of nucleic acids.

[0140]

[0160] Samples were then washed and eluted as previously described (i.e., a single wash with 1× binding and wash buffer and a single wash with 10 mM Tris-HCl, followed by resuspension in 20 μL of 10 mM Tris-HCl).

[0141]

[0161] The resuspended beads were added to a second extension reaction, resulting in the extension of the second oligonucleotide primer and the release of the target nucleic acid molecule into solution. The composition of the second extension reaction is listed in Table 6.

[0142]

[0162]

[0143] [Table 6]

[0144]

[0163] Following the second extension reaction, the sample was incubated at 50°C for 2 minutes, then placed directly on a magnet on ice for 1 minute. The supernatant was removed from the sample (without disturbing the beads) and added to an equal volume of KAPA PURE BEADS capture beads (KAPA BIOSYSTEMS). A 1X cleanup was performed, and the sample was eluted with 15 μL of 10 mM Tris-Cl, pH 8.0.

[0145]

[0164] The next step in the target enrichment protocol was amplification and cleanup with KAPA PURE BEAD capture beads (KAPA BIOSYSTEMS) according to the manufacturer's instructions for the KAPA HYPERPLUS Library Preparation Kit. The final product was eluted with 25 μL of Tris-HCl. The enriched target nucleic acid was then amplified and purified using the KAPA HYPERPLUS Library Preparation Kit (KAPA BIOSYSTEMS) according to the manufacturer's instructions (Figure 6B). The enriched and amplified library was sequenced on a MINISEQ DNA sequencer (ILLUMINA) using the mid-output kit with 2x150 bp reads, a loading concentration of 1.6 pM, and 1% PhiX DNA. The resulting sequencing data was processed using a pipeline developed for analysis of SEQCAP EZ Target Enrichment System (ROCHE) data to assess the degree of target enrichment (Figure 7).

[0146]

[0165] The described features, structures, or characteristics of the invention may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are set forth to provide a thorough understanding of system embodiments. However, one skilled in the art will recognize that both the system and method may be practiced without one or more specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations have not been shown or described in detail to avoid obscuring aspects of the invention. Accordingly, the foregoing description is intended to be illustrative, and not to limit the scope of the inventive concepts.

Claims

1. 1. A method for amplifying at least one target nucleic acid in a library of nucleic acids, comprising the steps of: (i) hybridizing a first oligonucleotide to at least one target nucleic acid in a library of nucleic acids, wherein the at least one target nucleic acid in the library of nucleic acids has a 3' end comprising a first adaptor and a 5' end comprising a second adaptor, wherein the first oligonucleotide comprises a sequence complementary to a region of interest (ROI) of the target nucleic acid located between the first adaptor and the second adaptor; (ii) extending the hybridized first oligonucleotide with a first polymerase, thereby generating a first primer extension complex comprising the target nucleic acid and the extended first oligonucleotide; (iii) capturing the first primer extension complex; (iv) enriching the first primer extension complexes, thereby enriching at least one target nucleic acid in the library of nucleic acids relative to at least one or more non-target nucleic acids in the library of nucleic acids; (v) hybridizing a second oligonucleotide to the target nucleic acid, wherein the second oligonucleotide is complementary to a sequence of the target nucleic acid located between the first and second adaptors, and the second oligonucleotide hybridizes to at least one target nucleic acid 5′ to the first oligonucleotide but outside the region of interest (ROI); (vi) extending the hybridized second oligonucleotide with a second polymerase, thereby generating a second primer extension complex comprising the target nucleic acid and the extended second oligonucleotide, wherein extension of the hybridized second oligonucleotide releases the extended first oligonucleotide from the first primer extension complex; and (vii) amplifying the target nucleic acid with a third polymerase, a first amplification primer, and a second amplification primer, wherein the first amplification primer is complementary to the first adapter at the 3' end of the target nucleic acid and the second amplification primer is complementary to the second adapter at the 5' end of the target nucleic acid; wherein amplifying the target nucleic acid produces a copy of the target nucleic acid, wherein the copy of the target nucleic acid comprises the entire sequence of the target nucleic acid, and wherein the copy of the target nucleic acid is longer than the extended first oligonucleotide and the extended second oligonucleotide; A method comprising:

2. The method of claim 1 , further comprising sequencing the amplified target nucleic acid.

3. The method of claim 1 , wherein the first oligonucleotide comprises a capture moiety.

4. 2. The method of claim 1, wherein the first oligonucleotide is bound to a solid support before hybridizing the first oligonucleotide to the target nucleic acid, and hybridizing the first oligonucleotide to the target nucleic acid and extending the hybridized first oligonucleotide with a polymerase thereby captures the first primer extension complex on the solid support.

5. 2. The method of claim 1, further comprising incorporating at least one modified nucleotide into at least one of the extended first oligonucleotide in the first primer extension complex and the extended second oligonucleotide in the second primer extension complex.

6. 2. The method of claim 1, further comprising incorporating at least one modified nucleotide into the extended first oligonucleotide in the first primer extension complex, wherein the at least one modified nucleotide comprises a capture moiety.

7. an extended first oligonucleotide in the first primer extension complex, and the extended second oligonucleotide in the second primer extension complex further comprising incorporating at least one uracil into at least one of 10. The method of claim 1, thereby forming a uracil-containing oligonucleotide product.

8. The method of claim 1 , further comprising contacting the library of nucleic acids with a blocking oligonucleotide.

9. The method of claim 1 , wherein the first adaptor and the second adaptor are forked adaptors.

10. 2. The method of claim 1, wherein the first adaptor and the second adaptor comprise at least one uracil.

11. 2. The method of claim 1, wherein at least one of the first adaptor, the second adaptor, the first amplification primer, and the second amplification primer comprises at least one of a unique identifier (UID) sequence, a molecular identifier (MID) sequence.

Citation Information

Patent Citations

  • Solid phase amplification process

    EP0672173B1

  • Methods for modifying DNA for microarray analysis

    US20050123956A1

  • Solid phase nucleic acid target capture and replication using strand displacing polymerases

    WO2016149837A1

  • Target enrichment by single probe primer extension

    WO2017021449A1

  • Elimination of primer-primer interactions during primer extension

    WO2017144457A1