Nucleic acid enrichment and detection
By designing a method of cleavage of mismatch sites of probes and nucleic acid molecules, low-frequency variant nucleic acid molecules are selectively enriched, solving the problem of enrichment difficulties in the prior art, and achieving efficient and low-cost nucleic acid detection and enrichment.
Patent Information
- Application Number
- CN202380082692.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-10-19
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively enrich low-frequency variant nucleic acid molecules during hybrid capture enrichment, resulting in high sequencing costs and difficult to detect, and conventional methods reduce sensitivity or increase complexity.
By designing probes to selectively modify and hybridize with nucleic acid molecules, cleavage using mismatch sites, selectively enrich or deplete target nucleic acid molecules, combined with the combination of different probes and solid support capture, efficient enrichment of variant molecules is achieved.
It improves the detection sensitivity and specificity of low-frequency variant nucleic acid molecules, reduces sequencing costs and complexity, and is suitable for efficient enrichment and detection of complex samples.
Smart Images

Figure CN120303410A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of U.S. Provisional Patent Application Serial No. 63 / 380,105, filed on October 19, 2022, the disclosure of which is hereby incorporated by reference in its entirety. Technical field
[0003] Provided herein are systems and methods for enriching and detecting nucleic acid molecules. In particular, provided herein are systems and methods for selectively enriching desired nucleic acid molecules in a sample containing undesired nucleic acid molecules. Background of the invention
[0005] Targeted detection of low - frequency variants in a wild - type molecular library has important clinical implications for early cancer detection, monitoring of cancer progression, targeting of cancer therapies, non - invasive prenatal testing, monitoring of T - cell populations targeting specific (neo) antigens, and early warning of organ transplant rejection. The combination of hybridization capture and next - generation sequencing (NGS) is the most commonly used method for targeted and multiplex detection of low - frequency variants, but has several sub - optimal features. Generally, hybridization capture enriches the target regions of interest but not variant molecules. This results in the vast majority of sequencing reads originating from wild - type rather than variant molecules. This causes waste, increases costs, and makes it difficult to detect rare variants from a background with errors from storage, library preparation, and sequencing. This results in insufficient specificity of NGS for routine detection of variants below ∼0.1%. Improved library preparation methods (such as duplex sequencing) can improve specificity and obtain accurate sequencing data from single molecules. Unfortunately, methods like duplex sequencing also reduce sensitivity (e.g., by modifying the library preparation method to avoid end - repair and by deliberately imposing a molecular bottleneck). The systems and methods described herein allow for the enrichment of variant molecules located within the target regions of interest, using, for example, modified hybridization capture methods based on selective modification (e.g., digestion) of undesired or desired molecules. This both reduces the number of sequencing reads required and opens the door to multiplex detection of low - frequency variant molecules below the current detection limit of NGS. Summary of the invention
[0007] Provided herein are systems and methods for enriching and detecting nucleic acid molecules. In particular, provided herein are systems and methods for selectively enriching desired nucleic acid molecules in a sample containing undesired nucleic acid molecules.
[0008] For example, in some embodiments, the following methods are provided herein, including contacting a sample (e.g., containing two or more different nucleic acid molecules) with a probe and a reagent, thereby selectively modifying (e.g., cleaving) a probe hybridized to a first nucleic acid molecule relative to a probe hybridized to a different second nucleic acid molecule. In some embodiments, the method further includes enriching or depleting the first nucleic acid molecule relative to the second nucleic acid molecule.
[0009] For example, in some embodiments, the method includes enriching or depleting a first nucleic acid molecule in a sample comprising a mixture of nucleic acid molecules by contacting the first nucleic acid molecule with a probe that is differentially complementary to a target region on the first nucleic acid molecule relative to other nucleic acid molecules in the sample; performing a cleavage reaction; and enriching or depleting the first nucleic acid molecule. In some embodiments, the probe is more complementary to the target region of the first nucleic acid molecule than it is to the corresponding target region of a second nucleic acid in the sample. In some embodiments, the probe is less complementary to the target region of the first nucleic acid molecule than it is to the corresponding target region of a second nucleic acid in the sample. In some embodiments, the difference between the first nucleic acid and the second nucleic acid is a sequence variation (e.g., point mutation, deletion, insertion, polynucleotide change, fusion, etc.). In some embodiments, the probe comprises a sequence that is perfectly complementary to the target region of the more complementary sequence and has one or more mismatches with the sequence variation found in the corresponding target region of the less complementary sequence. In some embodiments, the probe comprises one or more mismatches with the target regions of both the first and second nucleic acid molecules, but comprises more mismatches with the target region of the less complementary nucleic acid. In particular, the probe is designed such that when the probe hybridizes to the first nucleic acid relative to the second nucleic acid, the cleavage of the probe is different, thereby allowing selective enrichment or depletion of the first nucleic acid relative to the second nucleic acid.
[0010] For example, in some embodiments, the following methods are provided herein for increasing or decreasing the ratio of a first nucleic acid sequence to a second nucleic acid sequence in a sample, including: a) exposing a sample comprising the first and second nucleic acid sequences to a probe that is differentially complementary to the first and second nucleic acid sequences; b) modifying the probe (e.g., performing a cleavage reaction); and c) enriching or depleting the first nucleic acid sequence relative to the second nucleic acid sequence.
[0011] The probe can be equipped with one or more components that are removed before or during cleavage. For example, the probe can be provided with non-complementary flaps or other blocking groups (also known as capping groups) at its 3' or 5' end to prevent digestion or polymerase reactions until the blocking group is removed. The blocking group can be removed by any suitable mechanism (such as enzymatic cleavage, chemical reaction, temperature switching, physical cleavage, etc.). Any suitable blocking group can be used, including but not limited to phosphorothioate bonds, use of modified bases (such as 2'-O-methyl, 2'-fluoro, etc.), inverted or dideoxynucleotides, phosphorylation, inclusion of spacers, etc. The probe sequence that provides differential cleavage products when hybridized to different nucleic acid molecules can be located at any suitable position within the initial probe. For example, the mismatch sequence can be located at the 3' terminal base at the 3' end of the probe. The mismatch sequence can be located internally within the probe at the 3' end. The mismatch sequence can be located at the center of the probe. The mismatch sequence can be located within the 5' half of the probe or at the 5' end of the probe (such as the 5' terminal base).
[0012] The 5' or 3' end of the probe can contain a region (such as a 5' or 3' tail) that serves as an identifier (such as a sample identifier). Such sequences are used to selectively pull down captured molecules from a specific sample or from a specific region of a mixed sample. Such identifiers are particularly useful in multiplex reactions, where multiple different targets are reacted in the same sample or the same reaction vessel.
[0013] In some embodiments, the method employs a combination of probes in the same sample preparation, where some probes are designed to enrich or deplete sequences containing variants (e.g., as described above), while some probes are designed to simply capture regions of interest (e.g., using any known hybridization / capture technique or method). For example, in some embodiments, such a combination method is used to analyze microsatellite instability (MSI) or copy number variants in the same sequencing run as somatic variants: the somatic variants are enriched, while the genes / regions for MSI / CNV analysis are captured by standard methods. A potential problem is that standard hybridization / capture (hyb / cap) probes may be undesirably modified by the enzymes and / or reagents used in the enrichment / depletion methodology. To prevent this, in some embodiments, standard hybridization / capture oligonucleotides are made resistant to the modifying enzymes / reagents. For example, standard hybridization / capture oligonucleotides can be modified by using RNA instead of DNA, using modified bases or backbone modifications, or by introducing regions of intentional mismatches in the probe. In some cases, it may be beneficial to perform probe hybridization of the standard and enrichment / depletion probes simultaneously and then separate them. In some embodiments, this is accomplished by using different attachment chemistries on different probe types, or by different 5' or 3' sequences on the probes, which can be differentially captured onto a solid support via hybridization to complementary "adapter" oligomers, which themselves include an attachment moiety (or are pre-linked to the solid support).
[0014] The analytes / sequences to which the methods of the invention can be applied are those nucleic acids, such as naturally occurring or synthetic DNA or RNA molecules, which include the target polynucleotide sequence(s) being sought. In some embodiments, the analytes / sequences will typically be present in an aqueous solution containing them and other biological materials, and in some embodiments, the analytes / sequences will be present along with other background nucleic acid molecules that are not of interest for the testing purpose. In some embodiments, the analytes / sequences are present in low amounts relative to these other nucleic acid components. In some embodiments, for example, when the analyte is derived from a biological sample containing cellular material, some or all of these other nucleic acids and foreign biological materials are removed using sample preparation techniques such as filtration, centrifugation, chromatography, or electrophoresis before performing the enrichment or depletion methods described herein. In some embodiments, the DNA or RNA molecules are included in a mixture containing a sequencing library. The sequencing library can be derived from and / or include one or both of single-stranded or double-stranded molecules, and can include DNA and / or RNA. The library sequences can include modifications such as adapters, unique molecular identifiers (UMIs), primer binding sequences, etc., and can be prepared by any suitable method (e.g., tagged fragmentation).
[0015] The compositions and methods of the present invention can be used with any type of sample, including but not limited to environmental (e.g., water, soil, air, etc.) samples and biological samples. Biological samples can be from any source, including plants, animals, infectious disease agents, etc. Suitably, in some embodiments, the analyte / sequence is derived from a biological sample taken from a mammalian subject (especially a human patient), such as blood, plasma, sputum, urine, skin, biopsy, or surgical resection. In some embodiments, the biological sample is lysed in order to release the analyte / sequence by disrupting any cells present. In other embodiments, the analyte / sequence may already be present in the sample itself in a free form; for example, cell-free DNA circulating in blood or plasma. The compositions and methods of the present invention are particularly useful for sample types that have been historically challenging, which may have low allele fractions of analytes of interest. Such samples include blood, urine, samples collected by cytobrush (such as esophageal samples), samples derived from bronchoalveolar lavage (BAL), pleural fluid, and cerebrospinal fluid (CSF).
[0016] In some embodiments, the sample is a pooled sample. Pooled samples involve mixing multiple samples together in one batch, and testing the pooled sample in that batch. This method increases the number of individual samples that can be tested with a more limited amount of resources. Pooled samples of interest include but are not limited to donated blood samples, agricultural samples, food samples, sperm samples, and biological samples tested for the presence of infectious disease agents (such as SARS-CoV-2, HIV, HCV, etc.). In some embodiments, the pooled sample is an environmentally collected sample (e.g., wastewater sample), which, due to its nature of generation, has a pooled sample from multiple different sources. Although pooling of samples can reduce the allele fraction of variants due to dilution of the samples with each other, it can significantly increase the efficiency of screening. Because the techniques provided herein enable detection at very low allele fractions, it is particularly suitable for analyzing pooled samples. In some embodiments, a portion of each initial sample is pooled without using barcodes or other complex preparation steps, and the pooled sample is tested. If a positive result is obtained, the remaining portions of the unpooled samples can be tested individually.
[0017] The present disclosure also provides compositions (e.g., reagents, kits, reaction mixtures, instruments, software) for use with the methods described herein. For example, in some embodiments, the present disclosure provides a composition comprising one or more reagents that are necessary, sufficient, or useful for performing the methods described herein. For example, in some embodiments, the composition comprises: one or more probes comprising a sequence that is complementary to a known first sequence and different from a known second sequence (e.g., perfectly complementary to the known first sequence but not perfectly complementary to the known second sequence); one or more reagents that selectively modify probes hybridized to a desired nucleic acid relative to probes hybridized to a non-desired nucleic acid; and / or one or more reagents that cleave or digest the probes. In some embodiments, the composition further comprises a target nucleic acid isolation component for isolating the target nucleic acid molecule. In some embodiments, the composition comprises one or more solid supports. In some embodiments, the composition comprises one or more buffers. In some embodiments, the solid support is a bead (e.g., a magnetic or paramagnetic bead). In some embodiments, the composition further comprises one or more epigenetic modification-sensitive or -dependent restriction enzymes. In some embodiments, the composition further comprises one or more restriction endonucleases. In some embodiments, the composition further comprises one or more transposomes. In some embodiments, the composition comprises a Cas protein (e.g., Cas9). In some embodiments, the composition further comprises one or more transposases. In some embodiments, the composition further comprises one or more ligases. In some embodiments, the composition further comprises one or more blocking oligonucleotides. In some embodiments, the composition further comprises reagents for performing an amplification (e.g., PCR), sequencing (e.g., next-generation sequencing), or detection reaction. In some embodiments, the composition further comprises one or more molecular beacon probes. In some embodiments, the one or more molecular beacon probes are fluorescently labeled. In some embodiments, the composition further comprises components for transcribing RNA into cDNA.
[0018] In some embodiments, the composition is a reaction mixture that includes the reaction of any of the methods described herein at a particular time point. In some embodiments, the reaction mixture includes a probe / nucleic acid hybridization complex of the methods described herein. In some embodiments, the reaction mixture includes a captured nucleic acid molecule of the methods described herein. In some embodiments, the reaction mixture includes a region that contains a desired target nucleic acid concentration that is higher or lower than the desired target nucleic acid concentration present in the sample undergoing the digestion reaction. For example, in some embodiments, provided herein is a reaction mixture that includes: a sample; a reagent for modifying a probe that hybridizes to a target nucleic acid; a first nucleic acid molecule from the sample that hybridizes to a probe having a sequence, wherein the discrimination region of the probe is complementary to the first nucleic acid molecule; and a second nucleic acid molecule from the sample that hybridizes to the probe having the sequence, wherein the discrimination region of the probe is not perfectly complementary to the second nucleic acid molecule.
[0019] Also provided herein is the use of the composition (e.g., the use of a kit, the use of a reaction mixture, the use of a reagent, the use of an instrument, the use of software). For example, provided herein is the use of the composition for enriching or depleting a target nucleic acid in a sample.
[0020] In some embodiments, provided herein are devices and instruments that can be used in the methods described herein. In some embodiments, the devices and instruments can be used to collect a sample and dispense it into a reaction vessel. In some embodiments, the devices and instruments provide a reaction chamber for performing the method. In some embodiments, the devices and instruments provide multiple partitions or regions (e.g., wells, channels, etc.) for containing the reaction and / or for separating the enriched desired target nucleic acid or depleting the desired target nucleic acid. In some embodiments, the devices and instruments can be used to amplify or sequence nucleic acid molecules. In some embodiments, the devices and instruments can be used to detect nucleic acid molecules. In some embodiments, the devices and instruments can be used to receive or transmit information from a user. For example, the devices and instruments can include a user interface for receiving user instructions and a display for visually presenting the results to the user.
[0021] In some embodiments, provided herein is a computing device. The computing device can be used to control an instrument or device to facilitate the methods described herein. In some embodiments, the computing device collects, analyzes, and reports data. In some embodiments, the computing device includes one or more processors that run a computer program. In some embodiments, the computing device includes a non-transitory computer-readable medium (e.g., software) that contains instructions for guiding the processor to perform one or more computational steps.
[0022] In some embodiments, provided herein are methods that include enriching or depleting a first nucleic acid molecule in a sample comprising a mixture of nucleic acid molecules by contacting the sample with a probe that is differentially complementary to a target region of the first nucleic acid molecule relative to a second nucleic acid molecule in the sample; optionally, activating the probe hybridized to the first nucleic acid molecule by selectively modifying the probe hybridized to the first nucleic acid molecule relative to the probe hybridized to the second nucleic acid molecule; selectively digesting the probe hybridized to the first or the second nucleic acid molecule (relative to the other); and enriching or depleting the first nucleic acid molecule. In some embodiments, the first and second nucleic acid molecules comprise end-repaired nucleic acid molecules. In some embodiments, the first and second nucleic acid molecules are polyadenylated nucleic acid molecules. In some embodiments, the first and second nucleic acid molecules comprise one or more linker sequences.
[0023] In some embodiments, the first and second nucleic acid molecules are amplified. In some embodiments, the probe is activated by contacting with a cleavage agent. In some embodiments, the cleavage agent is selected from the group consisting of: restriction endonucleases, flap endonucleases, mismatch repair enzymes, RNases, Cas proteins, argonaute family enzymes, DNA-glycosylase formamidopyrimidine (Fpg), apurinic / apyrimidinic (AP) endonuclease (APE 1), and chemical cleavage agents.
[0024] In some embodiments, digestion includes contacting the probe with an exonuclease or an endonuclease. In some embodiments, the exonuclease is a 3' to 5' exonuclease. In some embodiments, the exonuclease is a 5' to 3' exonuclease.
[0025] In some embodiments, the probe comprises a binding moiety at its 3' or 5' end, or internally, the binding moiety optionally comprising biotin or a sequence to which a linker can hybridize, wherein the linker molecule is modified to bind to a solid support.
[0026] In some embodiments, the method further comprises the step of capturing the probe on a surface prior to digestion. In some embodiments, the surface comprises beads.
[0027] In some embodiments, the method further comprises the step of differentially releasing the first nucleic acid molecule or the second nucleic acid molecule from the probe. In some embodiments, release includes raising the temperature. In some embodiments, release includes changing the pH value. In some embodiments, release includes changing the salt concentration. In some embodiments, release includes a digestion step. In some embodiments, release occurs via melting after a cleavage event without otherwise changing the reaction conditions.
[0028] In some embodiments, the method further comprises the step of detecting the first or second nucleic acid molecule. In some embodiments, detecting comprises sequencing the first or second nucleic acid molecule.
[0029] In some embodiments, the probe comprises a base that is complementary to a position in the first nucleic acid molecule and mismatched to the corresponding position in the second nucleic acid molecule. In some embodiments, digestion comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests complementary strands over non-complementary strands (e.g., an enzyme that stalls at the mismatch position). In some embodiments, digestion comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests non-complementary strands over complementary strands.
[0030] In some embodiments, the probe comprises a blocking group at the 3'-end, 5'-end or internally. In some embodiments, the probe is activated using a cleavage agent that removes the blocking group from the probe hybridized to the first nucleic acid molecule but not from the probe hybridized to the second nucleic acid molecule. In some embodiments, digestion comprises contacting the probe with a nuclease that cleaves the probe lacking the blocking group but not the probe having the blocking group. In some embodiments, the probe is activated using a cleavage agent that removes the nucleic acid fragment comprising the blocking group from the probe hybridized to the first nucleic acid molecule but not from the probe hybridized to the second nucleic acid molecule. In some embodiments, digestion comprises contacting the nucleic acid fragment with a polymerase under conditions such that the fragment is extended and the probe hybridized to the first nucleic acid molecule is displaced or digested by using a polymerase having 5'-3' exonuclease activity, optionally using an upstream primer.
[0031] In some embodiments, the activation and digestion do not include pyrophosphorolysis.
[0032] In some embodiments, the method further comprises contacting the sample with a capture probe that hybridizes to the second nucleic acid molecule or to another nucleic acid molecule in the sample that is not the first nucleic acid molecule or the second nucleic acid molecule.
[0033] In some embodiments, the method includes enriching molecules having a specific fragmentation profile. In some embodiments, capture includes capturing nucleic acid fragments with probes having sequence identity to fragmentation breakpoints and adapter sequences. In some embodiments, capture can include capturing molecules having specific 5' and 3' breakpoints. In some embodiments, capture can include sequential hybridization and capture of each breakpoint. In other embodiments, capture can include hybridization and capture using multiple probes to capture molecules having specific 5' and 3' fragmentation breakpoints. The fragmentation pattern contains information on the nucleosome organization, chromatin structure, gene expression, and nuclease content of the originating tissue, resulting in a characteristic signature in the form of fragment size, nucleotide motifs at the fragment ends, single-stranded jagged ends, and genomic locations of the fragmentation endpoints. For example, fragmentation breakpoints can be used to identify the likely originating tissue of molecules in cell-free DNA. In oncology, this can be used to identify specific types or subtypes of cancer, particularly in screening tests.
[0034] The present disclosure further provides kits comprising reagents sufficient or useful for performing the above-described methods on a sample comprising the first and second nucleic acid molecules. Brief Description of the Drawings
[0036] Figure 1 Exemplary target enrichment and detection methods are shown.
[0037] Figure 2 Exemplary target enrichment and detection methods are shown.
[0038] Figure 3 Exemplary target enrichment and detection methods are shown.
[0039] Figure 4 Exemplary target enrichment and detection methods are shown.
[0040] Figure 5 Exemplary target enrichment and detection methods are shown. Detailed Description of the Invention
[0042] In some embodiments, the systems and methods employ one or more probes that hybridize at least in part to desired and / or undesired sequences. The probes are configured such that they are differentially cleaved or degraded relative to undesired sequences in the presence of desired sequence(s). As used herein, the term "target molecule" refers to a nucleic acid molecule in a sample that hybridizes to a probe. In some embodiments, the probe hybridizes to the target molecule at its 3'-end or 5'-end, or internally. In some embodiments, the probe or target molecule is attached to a solid surface. In some embodiments, the probe comprises DNA. In some embodiments, the probe comprises RNA. In some embodiments, the probe comprises a mixture of DNA and RNA. In some embodiments, the probe comprises one or more synthetic bases (e.g., locked nucleic acid (LNA)). In some embodiments, the probe comprises synthetic backbone modifications.
[0043] In some embodiments, the systems and methods cleave or digest the probe in the probe / target hybridization complex. In some embodiments, the cleavage / digestion mechanism employed is selected to be strand-specific (e.g., for the probe). In some embodiments, the cleavage / digestion mechanism is not strand-selective. In some embodiments, the probe or target is protected by a modification, e.g., a modification to a primer or adaptor used to generate a modified nucleic acid (e.g., 3' / 5'-end blocking modifications, internal base / backbone modifications, etc.). In some embodiments, the probe or target is protected by introducing a modification during amplification (e.g., phosphorothioate modification of the backbone).
[0044] In some embodiments, an activation step is employed to prepare the probe for the cleavage / digestion reaction. In some embodiments, the probe is made non-digestible via 3', 5' or internal modifications. In some embodiments, the modification is a mismatch relative to the target. In some embodiments, the mismatch is at one end of the probe (e.g., the 3'-end or 5'-end). In some embodiments, the mismatch is internal. In some embodiments, the mismatch comprises two or more bases. In some embodiments, the mismatch region provides a flap or bubble when the probe hybridizes to the target. In some embodiments, the probe is prepared with one or more modifications (e.g., backbone modifications such as phosphorothioate (PTO), base modifications, linker / adaptor, etc.). In some embodiments, the modification is the inclusion of an RNA base in an otherwise DNA probe. In some embodiments, the modification includes cyclization of the probe.
[0045] In some embodiments, such probes are selectively "activated" based on their differential hybridization to different targets (e.g., alleles), allowing the probe to subsequently be digested. In some embodiments, the activation step provides a nick or gap in the probe.
[0046] In some embodiments, activation is carried out using one or more restriction endonucleases. In some embodiments, the restriction endonuclease is a methylation / modification-specific endonuclease. In some embodiments, to avoid cleavage of both strands by the cleavage enzyme, one strand may contain modified bases or backbone modifications (such as phosphorothioate linkages) to block digestion. In some embodiments, activation is carried out using Flap endonuclease. In some embodiments, activation is carried out using mismatch repair enzymes or mismatch-specific endonucleases (such as CelI, S1, T7E1, Surveyor). In some embodiments, activation is carried out using RNase enzymes (such as RNaseH2). In some embodiments, activation is carried out using the CRISPR / Cas system (e.g., including the use of Cas enzymes or mutants that only cleave one strand). In some embodiments, activation is carried out using argonaute family enzymes (such as Ago1, Ago2, Ago3, Ago4, Hili, Hiwi, Hiwi2, Hiwi3, aubergine, PIWI, Ago5, Ago6, Ago7, Ago8, Ago9, Ago10, Alg-1, Alg-2, etc.). In some embodiments, activation is carried out using chemical cleavage of mismatches (CCM) (e.g., hydroxylamine + potassium permanganate, followed by piperidine).
[0047] In some embodiments, activation is carried out using DNA repair enzymes. For example, in some embodiments, the method involves hybridization of a probe that is designed to create a deliberate G:A mismatch at a desired cleavage position. The DNA repair enzyme MutY glycosylase recognizes the mismatch structure and selectively removes the mismatched A from the duplex to create an abasic site in the target strand. Addition of an AP-endonuclease such as endonuclease IV then cleaves the backbone, splitting the DNA strand into two fragments.
[0048] In some embodiments, digestion of the probe is selective based on differential hybridization to the target. In some embodiments, digestion of the probe employs a general digestion method that is selective for the activated probe. In some embodiments, digestion is in the 3'-5' direction. In some embodiments, digestion is in the 5'-3' direction. In some embodiments, digestion is via internal cleavage of the probe. In some embodiments, digestion does not use pyrophosphorolysis (where inorganic phosphate is the attacking group for cleavage).
[0049] In some embodiments where 3'-5' digestion is required, a 3'-5' exonuclease is used. Such exonucleases include, but are not limited to, Exonuclease I (ExoI) (e.g., E. coli ExoI), Exonuclease T (ExoT), Exonuclease VII (ExoVII), Exonuclease III (ExoIII), and DNS.
[0050] In some embodiments where 3'-5' digestion is required, a polymerase with 3'-5' exonuclease activity is used (e.g., a proofreading polymerase such as PHUSION (Thermo Fisher), Q5 (New England Biolabs), KAPAHIFI (Roche), KOD DNA polymerase).
[0051] In some embodiments where 5'-3' digestion is required, a 5'-3' exonuclease is used (e.g., Lambda exonuclease, RecJf, T7 exonuclease, Exonuclease V (ExoV), Exonuclease VIII (ExoVIII), T5 exonuclease).
[0052] In some embodiments where 5'-3' digestion is required, a DNA polymerase with 5'-3' exonuclease activity is used. In some such embodiments, a 5' (or internal) blocked probe can act as a blocking oligonucleotide to prevent primer extension against the target. Selective cleavage of the probe based on differential hybridization results in the melting of the blocked 5' end, allowing upstream primer extension, during which the remaining hybridized probe is digested (or displaced via strand displacement) in the 5'-3' direction, thereby releasing the target molecule. In similar embodiments, the 5' segment of the cleaved probe itself can serve as a primer.
[0053] For internal cleavage, any of the above "activation" methods can be used as an independent digestion method. In some embodiments, the fragment generated by cleaving the probe has a lower melting temperature (T m ) than the intact probe, thereby allowing selective denaturation from the target by increasing the temperature, pH, etc.
[0054] In some embodiments, the systems and methods use selective digestion as a means of enriching or depleting nucleic acid sequences. In some embodiments, the digestion reaction relies on complementarity between hybridized strands and only digests strands with specific characteristics (e.g., a specific sequence or structure). The reaction selectively shortens or digests certain sequences with a particular characteristic, leaving sequences without that characteristic undigested or less digested. The reaction can be carried out such that molecules hybridized to the sequences that are shortened or more digested are recovered and analyzed. Alternatively, the reaction can be carried out such that sequences that are less shortened or more digested are analyzed. Alternatively, the reaction can be carried out in order to analyze sequences that are not digested at all.
[0055] The enrichment or depletion can be repeated one or more times to further enrich the sample for sequences of interest. For example, in some embodiments, after completion of the first round, a second round of enrichment or depletion is carried out using the same reagents as in the first round. In some embodiments, an amplification reaction can be used between the first and second rounds of enrichment or depletion. In other embodiments, different probes that are selective for different sequences of the target nucleic acid to be enriched or removed are used. This method is particularly suitable, for example, when the sequences to be enriched or depleted differ from the sequences to be eliminated by at least two base positions. For example, a target nucleic acid containing two polymorphisms relative to the wild type can be subjected to a first round of enrichment or depletion based on the first polymorphism and a second round of enrichment or depletion based on the second polymorphism. Exponential enrichment or depletion can be achieved by using multiple rounds. In some embodiments, prior to enrichment or depletion, the target nucleic acid is modified to produce a synthetic sequence (e.g., addition of a polymorphism) such that the synthetic sequence is targeted for enrichment or depletion relative to the sequence without the synthetic sequence. In some embodiments, by using multiple probes, two or more nucleic acids can be enriched or depleted in any given reaction round. In some embodiments, in one or more rounds of the reaction, a particular nucleic acid can be enriched while a second nucleic acid is depleted.
[0056] In one aspect of the invention, a method of altering the ratio of a first nucleic acid sequence to a second nucleic acid sequence in a sample is provided, wherein the sample comprises at least the first and second sequences, the method comprising the steps of: a) introducing a sample comprising one or more nucleic acid analytes into a first reaction mixture comprising: i) probes that are complementarily different to the first and second sequences (e.g., the 3'-end or other region of the probe is perfectly complementary to one of the first or second sequences but not perfectly complementary to the other); ii) an optional activating enzyme or reagent; and iii) a nuclease that selectively cleaves or digests the probe hybridized to the first sequence relative to the second sequence; and b) separating the first sequence or its cleavage or digestion product from the second sequence or its cleavage or digestion product.
[0057] In some embodiments, any better (e.g., perfect) annealing probe complex can be isolated by using reaction conditions that favor better (e.g., perfect) annealing of probe sequence complexes over worse (e.g., imperfect) annealing of probe sequence complexes. This can take the form of a change in the temperature of the reaction mixture and / or a change in the pH of the reaction mixture and / or a change in the salinity of the reaction mixture. In some embodiments, chemical reagents (e.g., dimethyl sulfoxide (DMSO), formamide, etc.) are used to denature the nucleic acid to facilitate enrichment or depletion. In some embodiments, nucleic acid molecules (displacing oligonucleotides) and enzymes or proteins are used to separate hybridized nucleic acid molecules.
[0058] In some embodiments, two probes are used, one for the forward strand of the double-stranded target nucleic acid and one for the reverse strand. In some embodiments, it is beneficial to design the probes such that they do not hybridize to each other in a manner that interferes with the desired reaction.
[0059] In some embodiments, the probes are captured on a solid support before or after exposure to a sample containing the nucleic acid molecules to be enriched or depleted. In some embodiments, the probe contains a biotin moiety that is captured by the corresponding streptavidin moiety on the solid support. In some embodiments, the probe hybridizes to an adapter molecule that itself contains a modification such as a biotin moiety through which they are captured on the solid support. In some embodiments, the solid support is a bead. In some embodiments, the bead is a magnetic or paramagnetic bead.
[0060] The techniques are not limited to using capture to partition or enrich nucleic acid molecules of interest. Molecules can be partitioned or enriched, for example, based on differences in size, charge, or shape or other physical or chemical properties. In some embodiments, a moiety is added to the nucleic acid of interest (e.g., via click chemistry modification), whereby the added moiety confers a selectable distinguishing feature to the nucleic acid of interest.
[0061] In some embodiments, one or more wash steps are performed between one or more steps. In some embodiments, the wash and hybridization steps are performed at an elevated temperature of 25-95 °C.
[0062] In some embodiments, the sample includes a library of adapter-tagged nucleic acid analytes.
[0063] In some embodiments, sequences are identified by an amplification reaction (such as polymerase chain reaction (PCR), nucleic acid sequence-based amplification (NASBA), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), strand displacement amplification (SDA), rolling circle amplification (RCA), loop-mediated isothermal amplification (LAMP), recombinase polymerase amplification (RPA), helicase-dependent amplification (HDA), nicking and extension amplification reaction (NEAR), etc.). In some embodiments, sequences are identified by microarray analysis.
[0064] In some embodiments, sequences are identified by sequencing. In some embodiments, sequences are identified by next-generation sequencing (NGS) (such as, bridge amplification sequencing (Illumina), SMRT sequencing (PacBio), ion torrent sequencing, nanopore sequencing, pyrosequencing, etc.).
[0065] The compositions and methods of the present invention can be used for any type of sample, including but not limited to environmental (such as water, soil, air, etc.) samples and biological samples. Biological samples can be from any source, including plants, animals, infectious disease agents, etc. Suitably, in some embodiments, the analyte / sequence is derived from a biological sample taken from a mammalian subject (especially a human patient), such as blood, plasma, sputum, urine, skin, biopsy or surgical resection. In some embodiments, the biological sample will be lysed in order to release the analyte / sequence by disrupting any cells present. In other embodiments, the analyte / sequence may already be present in the sample itself in a free form; for example, cell-free DNA circulating in blood or plasma. The compositions and methods of the present invention are particularly useful for sample types that have been historically challenging, which may have a low allele fraction of the analyte of interest. Such samples include blood, urine, samples collected by cell sponge (such as esophageal samples), samples from bronchoalveolar lavage (BAL), pleural fluid, and cerebrospinal fluid (CSF).
[0066] In some embodiments, the sample is a pooled sample. A pooled sample involves mixing multiple samples together in a single batch, and testing the pooled sample in that batch. This approach increases the number of individual samples that can be tested with a more limited amount of resources. Pooled samples of interest include, but are not limited to, donated blood samples, agricultural samples, food samples, sperm samples, and biological samples tested for the presence of infectious disease agents such as SARS-CoV-2, HIV, HCV, etc. In some embodiments, the pooled sample is an environmentally collected sample (e.g., wastewater sample) that, due to its nature of generation, has a pooled sample from multiple different sources. While pooling of samples can reduce the allele fraction of variants due to dilution of the samples with each other, it can significantly increase the efficiency of screening. Because the techniques provided herein enable detection at very low allele fractions, it is particularly suitable for analyzing pooled samples. In some embodiments, a portion of each initial sample is pooled without using barcodes or other complex preparation steps, and the pooled sample is tested. If a positive result is obtained, the remaining portions of the unpooled samples can be tested individually.
[0067] The target nucleic acid to be analyzed can be any target nucleic acid of interest. In some embodiments, the analysis is for research, diagnostic, or therapeutic purposes. In some embodiments, when the sample is from a human subject, the purpose can be for the analysis, detection, treatment, or selection of treatment for one or more diseases or conditions. The techniques are particularly useful for analyzing multifactorial diseases and disorders, including but not limited to cancer (such as bladder cancer, breast cancer, cervical cancer, colorectal cancer, ovarian cancer, uterine cancer, vaginal cancer, vulvar cancer, head and neck cancer, eye cancer, brain cancer, kidney cancer, liver cancer, lung cancer, lymphoma, mesothelioma, myeloma, prostate cancer, skin cancer, thyroid cancer, pancreatic cancer, bone cancer, esophageal cancer, gallbladder cancer, gastric cancer, testicular cancer, anal cancer, rectal cancer, oral cancer, salivary gland cancer, sarcoma, and thyroid cancer), autoimmune diseases, asthma, ciliopathies, cleft palate, diabetes, heart disease, hypertension, inflammatory bowel disease, intellectual disability, mood disorders, obesity, infertility, and refractive errors.
[0068] In some embodiments, the nucleic acid being analyzed is a methylated sequence or derived from a methylated sequence. The techniques are used to distinguish the methylation status at any particular or multiple positions in the target nucleic acid. The methylated sequence can first be modified using chemical treatment (e.g., oxidation, reduction, bisulfite treatment) or by exposure to methylation-dependent restriction enzymes or any other suitable method, and then the modified sequence is enriched and / or identified.
[0069] In some embodiments, the targeted region of the RNA present in the sample is transcribed into DNA.
[0070] Some embodiments of the technology employ solid surfaces. Any solid surface compatible with nucleic acid capture can be used directly or indirectly. Solid supports include, but are not limited to, materials made of glass, metal, gel, plastic (such as polystyrene), ceramic, and filter paper, etc. In some embodiments, the solid support is beads (such as paramagnetic beads), microparticles, nanoparticles, columns, slides, etc. In some embodiments, the solid support is modified to include one or more of the following functional groups: amine, carboxylic acid, sulfonic acid, trimethylamine, and / or epoxide to facilitate reaction or coupling or conjugation with biomolecules. In some embodiments, streptavidin is attached to the solid surface. In some embodiments, the solid support is a dextran-modified surface. In some embodiments, the solid support is a polyethylene glycol (PEG) or PEG-modified surface. In some embodiments, the solid support is a polyvinylpyrrolidone (PVP) or PVP-modified surface. In some embodiments, the solid support is a polysaccharide or polysaccharide-modified surface. In some embodiments, the polysaccharide is selected from one or more of dextran, ficoll, glycogen, gum arabic, xanthan gum, carrageenan, amylose, agar, amylopectin, xylan, and / or β-glucan. In some embodiments, the solid support is a chemical resin or chemical resin-modified surface. In some embodiments, the chemical resin or chemical resin-modified surface is selected from one or more of the following resins: isocyanate, glycerol, piperidinyl-methyl, poly DMAP (polymer-bound dimethyl 4-aminopyridine), DIPAM (diisopropylaminomethyl, aminomethyl, polystyrene aldehyde, tris(2-aminomethyl)amine, morpholinyl-methyl, BOBA (3-benzyloxybenzaldehyde), triphenylphosphine, or benzylthio-methyl).
[0071] In some embodiments, a capture moiety is used to associate a probe or target nucleic acid with the solid surface. In some embodiments, the capture moiety is covalently attached to the solid support via a chemically cleavable linker, such as a disulfide, allyl, or azide-masked hemiaminal ether linker. In some embodiments, the capture moiety is covalently attached to the solid support via an amide or phosphorothioate bond. In some embodiments, a probe or target can be incorporated into a modified moiety to introduce a capture moiety. In some embodiments, the reaction includes reacting an alkyne-labeled oligonucleotide with an azide-biotin conjugate.
[0072] In some embodiments, the capture moiety comprises an oligonucleotide sequence and the solid surface comprises a complementary oligonucleotide sequence. In some embodiments, the oligonucleotide sequence comprises one or more modified bases and / or other such modifications known to those skilled in the art to alter the melting temperature. In some embodiments, the presence of one or more modified bases and / or other such modifications known to those skilled in the art results in a decrease in the melting temperature. In some embodiments, the presence of one or more modified bases and / or other such modifications known to those skilled in the art results in an increase in the melting temperature. In some embodiments, the length of the complementary sequence is 10, 20, 30, 40, 50, 100, 150 to 200 bases. In some embodiments, the length of the complementary sequence is 10, 20, 30, 40, 50 to 100 bases. In some embodiments, the length of the complementary sequence is 10-20, 10-30, 10-40 and 10-50 bases. In some embodiments, the length of the complementary sequence is 10-20, 10-30 and 10-40 bases. In some embodiments, the length of the complementary sequence is 10-20 and 10-30 bases. In some embodiments, the length of the complementary sequence is 10-20 bases.
[0073] In some embodiments, the capture moiety comprises a chemical modification and is attached to the solid support via an interaction between the chemical modification and the solid support. In some embodiments, the chemical modification is biotin and the solid support further comprises streptavidin. In some embodiments, the captured oligonucleotide sequence is released from the solid support. In some embodiments, the captured oligonucleotide sequence is released from the solid support by chemical denaturation. In some embodiments, chemical denaturation is achieved by using an appropriate concentration of base. In some embodiments, 0.1 M NaOH can be used. In some embodiments, the oligonucleotide sequence is released from the solid support by cleaving a chemical linker, wherein tris(2-carboxyethyl)phosphine (TCEP) or dithiothreitol (DTT) is added for a disulfide bond linker; a palladium complex or an allyl linker; or TCEP is added for an azide masked hemiaminal ether linker. In some embodiments, the oligonucleotide sequence is released from the solid support by removing non-canonical bases from the oligonucleotide sequence and cleaving at the resulting abasic site. In some embodiments, the non-canonical base is uracil, which is removed by uracil-DNA glycosylase. In another embodiment, the non-canonical base is 8-oxoguanine, which is removed by formamidopyrimidine DNA glycosylase (Fpg).
[0074] In some embodiments, the capture portion is an oligonucleotide region and is released by heating the reaction mixture. In some embodiments, the reaction mixture is heated to 37°C - 100°C for 1 - 20 minutes. In some embodiments, the reaction mixture is heated for 1 - 15 minutes. In some embodiments, the reaction mixture is heated for 1 - 10 minutes. In some embodiments, the reaction mixture is heated for 1 - 5 minutes. In some embodiments, the reaction mixture is heated for 5 minutes. In some embodiments, the reaction mixture is heated to 37°C - 85°C. In some embodiments, the reaction mixture is heated to 37°C - 75°C. In some embodiments, the reaction mixture is heated to 37°C - 65°C. In some embodiments, the reaction mixture is heated to 37°C - 55°C. In some embodiments, the reaction mixture is heated to 37°C - 45°C.
[0075] Those skilled in the art will understand that the temperature at which the reaction mixture is heated to release the complementary oligonucleotide region depends on many factors, including the length of the region.
[0076] In some embodiments, release is achieved by cleaving one or more oligonucleotide sequences. Such cleavage can be achieved by any of the means described previously or subsequently or any means known to those skilled in the art. In some embodiments, the oligonucleotide sequence is chemically cleaved. In some embodiments, the oligonucleotide sequence is enzymatically cleaved. In some embodiments, the oligonucleotide sequence is cleaved by a restriction enzyme. In some embodiments, the oligonucleotide sequence is cleaved by a restriction enzyme sensitive or dependent on epigenetic modification. In some embodiments, the oligonucleotide sequence is cleaved by a methylation-sensitive or -dependent restriction enzyme. In some embodiments, the oligonucleotide sequence is cleaved by a hydroxymethylation-sensitive or -dependent restriction enzyme.
[0077] In some embodiments, the sequence is enzymatically or chemically transformed to detect its methylation status before or after enriching the variant or wild-type sequence. Those skilled in the art will understand that the term "enrichment" refers to the selective separation of a target sequence from a mixture of target and non-target sequences as described previously or subsequently.
[0078] In some embodiments, the oligonucleotide sequence contains a photocleavable linker and the oligonucleotide sequence is released from the solid support by cleaving the linker (e.g., by UV light). Such modification can be a chemical backbone modification.
[0079] In some embodiments, the sample nucleic acid is fragmented prior to applying the methods disclosed herein. Those skilled in the art will understand that there are a variety of techniques available for fragmenting DNA. Such methods include sonication, needle shearing, nebulization, point-sink shearing, via a pressure cell (French press), and enzymatic methods. In some embodiments, fragmentation is achieved by sonication. In some embodiments, a (Denville, NJ) device may be used. In some embodiments, fragmentation is achieved by acoustic shearing. In some embodiments, a Covaris instrument (Woburn, MA) may be used. In some embodiments, fragmentation is achieved by nebulization. Nebulization forces the DNA through small holes in a nebulizer unit, which results in the formation of a fine mist for collection. The fragment size is determined by the gas pressure used to push the DNA through the nebulizer, the rate at which the DNA solution passes through the holes, the viscosity of the solution, and the temperature. In some embodiments, fragmentation is achieved by hydrodynamic shearing. In some embodiments, a Hydroshear from Digilab (Marlborough, MA) may be used. In some embodiments, fragmentation is achieved by point-sink shearing. In some embodiments, fragmentation is achieved by needle shearing. In some embodiments, fragmentation is achieved via use of a French press. In some embodiments, fragmentation is achieved by enzymatic fragmentation. In some embodiments, fragmentation is achieved by restriction endonuclease digestion. In some embodiments, the fragmentation is transposase-mediated fragmentation. In some embodiments, fragmentation is achieved by Cas9. In some embodiments, fragmentation is achieved by Cas9 as described in US10577644, the entire contents of which are incorporated herein by reference. In some embodiments, one or more different fragmentation techniques may be used. In some embodiments, one or more of the same or different fragmentation techniques may be used at one or more different points in the method.
[0080] In some embodiments, fragmentation of the sequence and adapter tagging occur at the same time or in the same step of the method. One such example is the Illumina Nextera DNA Library Prep kit.
[0081] Those skilled in the art will understand that there are a variety of techniques available for preparing adapter-tagged sequences / libraries.
[0082] In some embodiments, after fragmentation, the ends of the nucleic acid may be trimmed and polyadenylated, and then ligated to one or more adapters.
[0083] In some embodiments, after fragmentation, the ends of the nucleic acid can be trimmed and ligated to adapters in a blunt-end ligation reaction.
[0084] In some embodiments, after fragmentation, an adapter is ligated to single-stranded DNA.
[0085] In some embodiments, after fragmentation, terminal transferase is used to add non-templated bases to the 3' end of the fragment, thereby providing a site for priming to prepare double-stranded fragments.
[0086] In some embodiments, a topoisomerase can be used in place of DNA ligase.
[0087] In some embodiments, TOPO cloning can be used to add an adapter to fragmented DNA.
[0088] In some embodiments, after fragmentation, a transposase can be used to add an adapter sequence to the nucleic acid.
[0089] In some embodiments, after fragmentation, standard transposons can be used, but then modified using oligonucleotide replacement to create a Y-shaped adapter.
[0090] In some embodiments, where the sample is an adapter-tagged library, blocking oligonucleotides are used to prevent cross-hybridization of library molecules (so-called "daisy chaining").
[0091] In some embodiments, the sample contains one or more blocking oligonucleotides.
[0092] Any sequencing method can be used to analyze the target nucleic acid molecule. In some embodiments, the sequencing is Maxam-Gilbert sequencing. In some embodiments, the sequencing is Sanger sequencing. In some embodiments, the sequencing is shotgun sequencing. In some embodiments, the sequencing is single molecule real-time sequencing. In some embodiments, the sequencing is ion semiconductor sequencing. In some embodiments, the sequencing is pyrosequencing. In some embodiments, the sequencing is sequencing by synthesis. In some embodiments, the sequencing is combinatorial probe anchor synthesis (cPAS). In some embodiments, the sequencing is sequencing by ligation. In some embodiments, the sequencing is nanopore sequencing. In some embodiments, the sequencing is GenapSys sequencing. In some embodiments, the sequencing is next-generation sequencing (NGS).
[0093] In some embodiments, a method of screening a patient is provided, comprising detecting the presence or absence of one or more specific nucleic acid sequences in a sample derived from the patient using any of the previously or subsequently described embodiments of the method.
[0094] Those skilled in the art will understand that such screening will be useful for monitoring patients undergoing treatment for one or more conditions, the treatment status of which can be determined by the levels of one or more nucleic acid sequences in a patient sample.
[0095] For example, the treatment status of a patient undergoing treatment for one or more cancers can be determined by the levels of one or more nucleic acid sequences in their blood and / or the presence and / or absence of one or more specific variants. High levels of circulating tumor nucleic acid sequences and / or the presence and / or absence of one or more specific variants can be used to infer whether a particular treatment has the desired effect. Accordingly, a method for monitoring the success or failure of a particular treatment is provided, wherein such success can be inferred by the presence or absence and / or the respective levels of specific nucleic acid sequences in a sample derived from the patient.
[0096] In some embodiments, a method for monitoring patients in remission to detect any recurrence of disease is provided.
[0097] In some embodiments, a method for screening ostensibly healthy individuals to detect the presence of one or more disease states, including but not limited to cancer, is provided.
[0098] In some embodiments, a method for detecting the presence and / or absence of one or more genetic markers in a patient diagnosed with one or more disease states and using the presence and / or absence of one or more markers to determine which treatment the patient should receive is provided.
[0099] In some embodiments, a method for diagnosing and / or monitoring one or more cancers in a patient is provided, comprising detecting the presence or absence of one or more specific nucleic acid sequences in a sample derived from the patient using any of the previously or subsequently described embodiments of the method.
[0100] Those skilled in the art will understand that one or more specific nucleic acid sequences can be individual-specific (identified from a tissue biopsy or surgical resection, e.g., by identification methods such as sequencing), and in such cases, an individual patient-specific panel can be used.
[0101] Those skilled in the art will further understand that in some embodiments, the panel will cover known hotspots in the human genome, i.e., regions that are recurrently mutated in a given cancer type.
[0102] Those skilled in the art will further understand that in some embodiments, the panel will cover target regions or the entire human exome.
[0103] In some embodiments, a method of non-invasive prenatal testing (NIPT) is provided, including using any of the previously or subsequently described embodiments of the method to detect the presence or absence of one or more specific nucleic acid sequences in a sample derived from a patient, where the patient is a pregnant patient. In some embodiments, the sample is plasma and / or serum of the blood of a pregnant patient. In some embodiments, the methods provided herein are employed to enrich and / or quantify the fetal fraction of the sample, which uses a set of common SNPs associated with such a sample.
[0104] In some embodiments, a method of treating a patient is provided, including the steps of:
[0105] - performing any of the previously or subsequently described embodiments of the invention to detect the presence or absence of one or more specific nucleic acid sequences in a sample derived from the patient;
[0106] - making one or more treatment decisions based on the presence or absence of the sequence.
[0107] In some embodiments, the treatment decision is to initiate a specific treatment. In some embodiments, the treatment decision is to stop a specific treatment. In some embodiments, the treatment decision is to increase the dose of a specific treatment. In some embodiments, the treatment decision is to decrease the dose of a specific treatment. In some embodiments, the treatment decision is to increase the frequency of administration of a specific treatment. In some embodiments, the treatment decision is to decrease the frequency of administration of a specific treatment. In some embodiments, the treatment decision is to add an additional drug to an existing treatment regimen. In some embodiments, the treatment decision is to remove a drug from an existing treatment regimen.
[0108] In some embodiments, kits are provided that include one or more or all of the components necessary, sufficient, or useful for performing the methods described herein. For example, in some embodiments, the kit includes one or more probes, solid supports, enzymes, blocking oligonucleotides, buffers, detergents (e.g., sodium dodecyl sulfate (SDS), TWEEN20, etc.), adapters, crowding agents (e.g., polyethylene glycol (PEG), polyvinyl alcohol (PVA), dextran sulfate, etc.), solvents (e.g., formamide, ethylene carbonate, etc.), additives that hybridize to repetitive sequences (COT-1 DNA, salmon sperm DNA, oligonucleotides that block ribosomal RNA, etc.), capture moieties (e.g., biotin; e.g., as part of a probe), metal ions, blocking oligonucleotides, sequencing reagents, amplification reagents (isothermal amplification reagents; exponential amplification reagents (e.g., thermostable polymerase, primers, dNTPs, buffers, labeled detection probes)), transcription reagents, instructions for use, software, instruments, positive controls, negative controls, etc. One or more containers may separately house one or more of the components.
[0109] In some embodiments, kits are provided that include a plurality of probes, as described previously or subsequently.
[0110] In some embodiments, kits are provided that include from 1 to 1,000,000 individual probes. In some embodiments, kits are provided that include from 1 to 100,000 individual probes. In some embodiments, kits are provided that include from 1 to 10,000 individual probes. In some embodiments, kits are provided that include from 1 to 1,000 individual probes.
[0111] In some embodiments, bioinformatics methods are used to analyze sequencing data. In some embodiments, the presence or absence of specific variants is called. In other embodiments, data from multiple variants are combined to derive a probability estimate of the presence or absence of a specific target nucleic acid.
[0112] For example, Illumina sequencing data analysis includes using tools such as bcl2fastq to convert BCL files to FASTQ format and demultiplex. In some embodiments, the sequencing reads include molecular identifiers. In such cases, the molecular identifiers can be extracted from the sequencing reads, appended to the FASTQ headers, and the sequencing reads trimmed. In some embodiments, barcodes with non-canonical bases (not A, C, G, or T) can be filtered. The resulting reads can then be aligned using tools such as bwa mem, with the -C option to append the barcode sequence to the alignment. The alignment can then be sorted using biobambam2, bamsormadup, and bammarkduplicatesopt, to mark duplicate reads by coordinate sorting and annotate the reads using read coordinates, paired-read coordinates, and optical duplication auxiliary tags. Reads that are not marked as properly paired or are marked as optical duplicates, supplementary, QC failed, unmapped, or secondary alignments can be filtered. Each read can then be tagged with an auxiliary tag that consists of the reference name, sorted read, and matching fragmentation breakpoint, forward and reverse read barcodes, and read strand.
[0113] In some embodiments, a variant calling algorithm is used to analyze sequencing data, which uses different subsets of read flags and tags, including read and paired-read coordinates, optical duplicate flags, UMI sequences, MID sequences, alignment scores, secondary alignment scores, etc. In this case, the analysis of the sequencing data compares the probabilities of the observed data under two models. The first is a null model that specifies the distribution of sequencing artifacts. The second is a model that allows for true variants. In this case, if the probability under the alternative model exceeds the probability of the null model, the variant is called. In some embodiments, a set of pre-characterized samples can help model the error distribution of the first model.
[0114] In some embodiments, auxiliary tags can be used to identify reads that may be derived from the same input molecule and / or the same strand of the same input molecule. In some embodiments, the consensus base quality score can be derived from reads sharing the same auxiliary tag.
[0115] In some embodiments, an artificial intelligence algorithm such as a convolutional neural network is used to identify variants.
[0116] In some embodiments, the sequencing data can be further filtered to remove artifacts. Example filters include the number of mismatches present on a given sequencing read; alignment scores and second-best alignment scores; base quality scores or consensus base quality scores; the minimum number of reads covering a given variant locus; the position of the variant in the sequencing read; whether the read has been 5'-trimmed; whether the read is an incorrect pair; whether the read contains indels; and the variant allele fraction of a given variant. In some embodiments, regions of the genome containing common SNPs or regions prone to alignment artifacts are filtered out. Many other filters are known to those skilled in the art.
[0117] In some embodiments, control samples are sequenced to filter out variants. For example, DNA from oral epithelium or other tissue sources can be sequenced to remove germline variants. In another embodiment, buffy coat or white blood cell DNA can be sequenced to filter out somatic mutations arising from clonal hematopoiesis.
[0118] The compositions, methods, and kits of the present invention can be used in a wide range of applications and settings. In some embodiments, they are used in any methodology that desires to detect sequences in a sample. In some embodiments, they are used in any method that desires to detect a few (e.g., rare) sequences in a complex sample. In addition to the above-exemplified uses, many other illustrative uses are provided below.
[0119] In some embodiments, the compositions, methods, and kits are used for the analysis and treatment of infectious diseases. The technology is particularly valuable for detecting low-frequency mutations that may be present in a sample. For example, the technology can be used to detect low-frequency mutations associated with treatment-resistant (such as antibiotic-resistant, antiviral-resistant, etc.) infectious diseases (such as HIV, tuberculosis, etc.). The technology can also be used to selectively pull down bacterial or viral DNA or RNA for sequencing.
[0120] As described above, the technology is particularly suitable for analyzing and / or enriching analytes in complex samples. Microbiome analysis is an area of increasing research and clinical interest, where the technology can be used to provide a much higher specificity of selection for the desired bacterial DNA for sequencing or other analysis.
[0121] The technology can also be used for high-throughput analysis and multiplex analysis of many different samples. These benefits can be used in a wide variety of genotyping applications, including forensic analysis, paternity testing, disease analysis (such as cancer, infectious diseases), drug susceptibility testing, agriculture, and food testing (such as assisting in selective breeding, identifying trace contaminants, etc.).
[0122] The technology can be used for error correction of synthetic nucleic acids (such as DNA). Synthetic nucleic acids are used for research, diagnostic, and clinical indications. It is often important to avoid or minimize the use of nucleic acid molecules with unexpected or undesired sequences. The technology can be used to identify and isolate desired molecules from undesired ones.
[0123] Nucleic acid editing is becoming an important method in research, synthetic biology, and clinical applications. For example, CRISPR / CAS editing of nucleic acids and related methods are becoming important methods. Many of these editing techniques produce a mixed population of molecules, including the expected edited products, unedited products, and unintended edited products. The technology provided herein helps to identify, select, and isolate the expected edited products.
[0124] The technology can also be used for environmental monitoring. In addition to agricultural uses, the technology is particularly suitable for analyzing environmental samples that may contain trace amounts of analytes of interest. Such samples include, but are not limited to, the analysis of native and invasive biological species, early detection of invasive species, air and water pollution, and ancient DNA analysis. Sample types include, but are not limited to, soil, water, snow, feces, mucus, gametes, shed skin, cadavers, hair, and air.
[0125] The technology can be used to separate a desired subset of nucleic acids from other subsets in a specific sample. For example, the technology is used for the isolation and analysis of chloroplast and mitochondrial genomes.
[0126] This technology can be used for cell line screening of engineered and natural cells. Such cell lines include, but are not limited to, cell cultures (primary and immortalized), stem cells (embryonic, induced pluripotent, dedifferentiated, etc.), differentiated cells intended for cell therapy, ex vivo modified cells for research or clinical applications (such as CAR T cells), and genetically engineered cells.
[0127] This technology can be used to remove damaged or other unwanted nucleic acids from intact or desired nucleic acids. For example, this technology can be used to remove damaged DNA from a sample prior to methylation analysis.
[0128] This technology can be used for pre-implantation screening of cells (such as embryos, eggs, sperm), liposomes, exosomes, nucleic acid carriers (such as gene therapy vectors), etc. prior to their administration to a subject. In some embodiments, this technology can be applied to the culture medium for screening without disturbing living cells.
[0129] This technology can be used for drug toxicity screening. This technology is particularly suitable for identifying DNA damage, mutation generation, methylation changes, etc. that may be associated with the use of a specific drug.
[0130] This technology can be used for fragmentation profiling of nucleic acids. For example, probes spanning or aligned with breakpoints that associate a specific sequence with relevant associated information (such as origin tissue, association with diseases such as cancer, etc.) can be used.
[0131] This technology can be used for any application that requires reducing nucleic acid complexity. For example, this technology can be used to reduce the complexity of the whole genome. In some such embodiments, a step of using restriction enzyme digestion or other nucleic acid fragmentation methods is followed by a step of pulling out only the cleaved molecules using probes that match known end sequences.
[0132] This technology can be used to evaluate microsatellite instability (MSI). Enrich and / or identify target nucleic acid molecules in a sample that differ in the presence, number, or nature of repetitive nucleotides (such as GT / CA repeats). MSI is associated with a variety of diseases and conditions, including but not limited to colon cancer, gastric cancer, endometrial cancer, ovarian cancer, cholangiocarcinoma, urothelial cancer, brain cancer, and skin cancer.
[0133] This technology can also be used to evaluate tumor mutational burden (TMB). TMB has become a predictive biomarker for immune checkpoint therapy and is used for other purposes. Currently, next-generation whole exome sequencing is used to evaluate TMB or to evaluate genomic panels that provide sequences of gene subsets. A more sensitive TMB evaluation with significantly lower cost and burden can be performed using the technology provided herein.
[0134] This technique can be used for haplotype analysis. Genomic information reported as haplotypes rather than genotypes is becoming increasingly important for personalized medicine as well as a wide range of research applications. Haplotypes, which are more specific than less complex variants such as single nucleotide variants, can also be used for prognosis and diagnosis, tumor analysis, and tissue typing for transplantation. Currently, sequencing is the most common form of molecular haplotype analysis. The error rate of sequencing technology is an obstacle to obtaining accurate information. The techniques provided herein allow for efficient and highly accurate haplotype analysis.
[0135] In some embodiments, assay components are designed to avoid specific polymorphisms (e.g., SNPs). It is desirable to avoid unnecessary remnants of off-target molecules or insufficient recovery of on-target molecules. In some embodiments, probes are designed not to hybridize to regions containing known polymorphisms (e.g., SNPs). In some embodiments, probes and / or capture are configured to target a target sequence strand that does not contain the polymorphism to be avoided. In some embodiments, multiple probes are used for each target, each probe targeting a different allele (e.g., SNP allele). In some embodiments, a universal base (e.g., inosine) is added to the probe corresponding to a known polymorphism site (e.g., SNP site) and digested (or not digested) as if there were a sequence match at the polymorphism position, regardless of whether the position contains a polymorphic sequence or a wild-type sequence.
[0136] The ability of this technique to enrich any desired or sequence of interest enables this technique to enhance existing nucleic acid methodologies. For example, many nucleic acid sequencing methods encounter difficulties when there are repetitive sequence regions in the target nucleic acid. The techniques provided herein allow for the removal of repetitive regions, making such sequencing reactions more accurate and efficient. Examples
[0137] Example 1
[0138] Double-strand specific (mismatch-blocked) 3'-5' exonuclease digestion
[0139] The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. In the first stage, probe hybridization is performed on the target sequence. The probe sequence perfectly matches the mutated target sequence and has at least one mismatch with the wild-type target sequence. The probe has a modification at its 5' end, allowing attachment to a surface. A double-stranded specific DNA exonuclease digests the probe in the 3'-5' direction. Digestion of the probe annealed to the WT molecule stops due to the presence of the mismatch. The probe annealed to the mutated target is digested as long as it is sufficiently complementary to the mutated target, or until the reaction is stopped (e.g., by adding a protease). The mutated target is released (e.g., into the supernatant), but the wild-type molecule remains annealed to the partially digested probe, which is attached to a solid support (e.g., beads). Thus, the mutated target is readily separable from the wild-type target, allowing enrichment of the mutated target. Figure 1 A representative schematic of the process is shown.
[0140] Example 2
[0141] 3'-5' exonuclease digestion after restriction enzyme activation
[0142] The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. In the first stage, probe hybridization is performed on the target sequence. Digestion of the probe sequence at its 3' end is blocked by oligonucleotide modification or the presence of a mismatch (or other methods described above). The probe sequence perfectly matches the mutated target sequence and has at least one mismatch with the wild-type target sequence. In some embodiments, the probe is modified at its 5' end, allowing attachment to a surface. In the next stage, an enzyme (e.g., a restriction enzyme) selectively nicks the probe that is perfectly annealed to the mutated target and is unable to nick the probe annealed to the wild-type molecule. The 3' end block is removed from the probe, allowing 3'-5' digestion of the probe and subsequent release of the mutated target (e.g., into the supernatant). Figure 2 A representative schematic of the process is shown.
[0143] Example 3
[0144] 3'-5' exonuclease digestion after mismatch-specific endonuclease activation
[0145] The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. In the first stage, probe hybridization is performed on the target sequence. Digestion of the probe sequence at its 3'-end is blocked by oligonucleotide modification or the presence of a mismatch (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe has a modification at the 5'-end to allow attachment to a solid surface. In the next stage, the mismatch-specific endonuclease selectively nicks the probe at the mismatch site but cannot nick the probe annealed to the wild-type molecule. The 3'-end block is removed from the probe, allowing 3'-5' digestion of the probe and subsequent release of the mutant target (e.g., into the supernatant). Figure 3 A representative schematic of the process is shown in
[0146] Example 4
[0147] After activation of the mismatch-specific endonuclease, 5'-3' exonuclease digestion
[0148] The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. In the first stage, probe hybridization is performed on the target sequence. Digestion of the probe sequence at its 5'-end is blocked by oligonucleotide modification or the presence of a mismatch (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe is modified at the 3'-end to allow attachment to a solid surface. In the next stage, the mismatch-specific endonuclease selectively nicks the probe at the mismatch site but cannot nick the probe annealed to the wild-type molecule. The 5'-end block is removed from the probe, allowing 5'-3' digestion of the probe and subsequent release of the mutant target (e.g., into the supernatant). Figure 4 A representative schematic of the process is shown in
[0149] Example 5
[0150] After activation of the mismatch-specific endonuclease, displacement
[0151] The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. In the first stage, probe hybridization is performed on the target sequence. Digestion of the probe sequence at its 5' end is blocked by oligonucleotide modification or the presence of a mismatch (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe has a modification at its 3' end to allow attachment to a surface. In the next stage, the enzyme selectively nicks the probe at the mismatch site and it cannot nick the probe annealed to the wild-type molecule. By nicking, the probe is split into two oligonucleotides: the 5'-end blocked / closed oligonucleotide and the 3'-end bead-attached oligonucleotide. The 5'-end blocked oligonucleotide serves as a primer for extension by strand displacement DNA polymerase. By displacing the 3'-end bead-attached oligonucleotide, the mutant target sequence is released into the supernatant. Figure 5 A representative schematic of the process is shown.
Claims
1. A method, comprising: Enriching or depleting a first nucleic acid molecule in a sample containing a mixture of nucleic acid molecules, which is carried out by contacting the sample with a probe that is differentially complementary to a target region of the first nucleic acid molecule relative to a second nucleic acid molecule in the sample; optionally, activating the probe hybridized to the first nucleic acid molecule by selectively modifying the probe hybridized to the first nucleic acid molecule relative to the probe hybridized to the second nucleic acid molecule; selectively digesting the probe hybridized to the first or the second nucleic acid molecule relative to the other; and, enriching or depleting the first nucleic acid molecule.
2. The method of claim 1, wherein the first and second nucleic acid molecules comprise end-repaired nucleic acid molecules.
3. The method of claim 1 or 2, wherein the first and second nucleic acid molecules are poly(A)-tailed nucleic acid molecules.
4. The method of any one of claims 1 to 3, wherein the first and second nucleic acid molecules comprise a tag or adaptor sequence.
5. The method of any one of claims 1 to 4, wherein the first and second nucleic acid molecules are amplified.
6. The method of any one of claims 1 to 5, wherein the probe is activated by contacting with a cleavage agent.
7. The method of claim 6, wherein the cleavage agent is selected from the group consisting of: restriction endonucleases, flap endonucleases, mismatch repair enzymes, ribonucleases, Cas proteins, argonaute family enzymes, DNA-glycosylases for formamidopyrimidine, depurinating / apyrimidinic (AP) endonucleases, and chemical cleavage agents.
8. The method of any one of claims 1 to 7, wherein the digestion comprises contacting the probe with an exonuclease or endonuclease.
9. The method of claim 8, wherein the exonuclease is a 3'-to-5' exonuclease.
10. The method of claim 8, wherein the exonuclease is a 5'-to-3' exonuclease.
11. The method of any one of claims 1 to 10, wherein the probe comprises a binding moiety at its 3' or 5' end, the binding moiety optionally being biotin or a sequence to which an adaptor molecule can hybridize, wherein the adaptor molecule is modified to bind to a solid support.
12. The method of any one of claims 1 to 11, further comprising the step of capturing the probe on a surface before or after the digestion.
13. The method of claim 12, wherein the surface comprises beads.
14. The method of claim 12, further comprising the step of differentially releasing the first nucleic acid molecule or the second nucleic acid molecule from the probe.
15. The method of claim 14, wherein the release comprises raising the temperature.
16. The method of claim 14, wherein the release comprises changing the pH value.
17. The method of claim 14, wherein the release comprises changing the salt concentration.
18. The method of claim 14, wherein the release comprises the digestion.
19. The method of any one of claims 1 to 18, further comprising the step of detecting the first or second nucleic acid molecule.
20. The method of claim 19, wherein the detection comprises sequencing the first or second nucleic acid molecule.
21. The method of any one of claims 1 to 20, wherein the probe comprises a base that is complementary to a position in the first nucleic acid molecule and mismatched to the corresponding position in the second nucleic acid molecule.
22. The method of claim 21, wherein the digestion comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests the complementary strand relative to the non-complementary strand.
23. The method of claim 21, wherein the digestion comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests the non-complementary strand relative to the complementary strand.
24. The method of any one of claims 1 to 23, wherein the probe comprises a blocking group at the 3'-end, 5'-end or internally.
25. The method of claim 24, wherein the probe is activated using a cleavage agent that removes the blocking group from the probe hybridized to the first nucleic acid molecule, but not from the probe hybridized to the second nucleic acid molecule.
26. The method of claim 25, wherein the digestion comprises contacting the probe with a nuclease that cleaves the probe lacking the blocking group, but not the probe having the blocking group.
27. The method of claim 24, wherein the probe is activated using a cleavage agent that removes a nucleic acid fragment containing the blocking group from the probe hybridized to the first nucleic acid molecule, but not from the probe hybridized to the second nucleic acid molecule.
28. The method of claim 27, wherein the digestion comprises contacting the nucleic acid fragment with a polymerase under conditions such that the fragment is extended and the probe hybridized to the first nucleic acid molecule is displaced or digested by using a polymerase having 5'-3' exonuclease activity, optionally using an upstream primer.
29. The method of any one of claims 1 to 28, wherein the activation and digestion do not include pyrophosphorolysis.
30. The method of any one of claims 1 to 29, further comprising contacting the sample with a capture probe that hybridizes to the second nucleic acid molecule or to another nucleic acid molecule in the sample that is not the first nucleic acid molecule or the second nucleic acid molecule.
31. The method of any one of claims 1 to 30, wherein the first and / or second nucleic acid molecule is a single-stranded nucleic acid molecule.
32. The method of any one of claims 1 to 31, wherein the first and / or second nucleic acid contains one or more methylated nucleotides.
33. A kit comprising reagents sufficient to perform the method of any one of claims 1 to 32 on a sample comprising the first and second nucleic acid molecules.
Citation Information
Patent Citations
Method for fragmenting genomic DNA using CAS9
US10577644B2