Nucleic Acid Concentration and Detection

The method of using probes with differential complementarity and selective cleavage enriches desired nucleic acid molecules, addressing inefficiencies in detecting low-frequency variants by improving specificity and sensitivity in nucleic acid sequencing.

JP2025536281APending Publication Date: 2025-11-05BIOFIDELITY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025521258
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-19
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Current methods for targeted detection of low-frequency variants in nucleic acid sequencing are inefficient and costly due to the enrichment of wild-type molecules, leading to background errors and reduced sensitivity, making it challenging to detect rare variants below approximately 0.1%.

Method used

A method involving the use of probes with differential complementarity to target and non-target nucleic acid molecules, followed by selective cleavage and enrichment or depletion of desired nucleic acids using enzymes and reagents, allowing for improved specificity and sensitivity in detecting low-frequency variants.

Benefits of technology

Enhances the detection of low-frequency variants by reducing the number of sequencing reads required and improving the specificity of next-generation sequencing, enabling accurate sequencing data from a single molecule.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536281000001_ABST
    Figure 2025536281000001_ABST
Patent Text Reader

Abstract

Systems and methods for the enrichment and detection of nucleic acid molecules are provided. In particular, systems and methods are provided for selectively enriching desired nucleic acid molecules from a sample containing undesired nucleic acid molecules. This method uses probes that differ in complementarity between desired and undesired nucleic acid molecules, e.g., wild-type and mutant forms. The probes may be bound to a solid phase, for example, via biotin. The probes bound to the undesired nucleic acid molecules are cleaved, releasing the undesired nucleic acid molecules and enriching the desired nucleic acid molecules by binding to the solid phase. Furthermore, this method can also utilize differential release from the probes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 380,105, filed October 19, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0002] Provided herein are systems and methods for the enrichment and detection of nucleic acid molecules. In particular, provided herein are systems and methods for selectively enriching desired nucleic acid molecules in a sample containing undesired nucleic acid molecules. [Background technology]

[0003] Targeted detection of low-frequency variants in a pool of wild-type molecules is clinically important for early cancer detection, cancer progression monitoring, cancer therapy targeting, noninvasive prenatal testing, monitoring T cell populations targeting specific (neo)antigens, and early warning of transplant organ rejection. Hybridization capture combined with next-generation sequencing (NGS) is the most commonly applied method for targeted, multiplexed detection of low-frequency variants, but it has several suboptimal characteristics. Generally, hybridization capture enriches for target regions of interest but not for variant molecules. As a result, the majority of sequencing reads are derived from wild-type rather than variant molecules. This is wasteful, costly, and makes detecting rare variants challenging due to background errors from storage, library preparation, and sequencing. As a result, NGS lacks specificity for routine detection of variants below approximately 0.1%. By modifying the library preparation method (e.g., double sequencing), it is possible to improve specificity and obtain accurate sequencing data from a single molecule. Unfortunately, methods such as double sequencing also reduce sensitivity (e.g., by modifying the library preparation method to avoid end repair and by intentionally imposing a molecular bottleneck). The systems and methods described herein allow for enrichment of variant molecules located in the target region of interest, for example, by using a modified hybridization capture method based on selective modification (e.g., digestion) of undesired or desired molecules. This not only reduces the number of required sequencing reads, but also paves the way for multiplexed detection of low-frequency variant molecules that are below the detection limit of current NGS. Summary of the Invention [Problem to be solved by the invention]

[0004] Provided herein are systems and methods for the enrichment and detection of nucleic acid molecules. In particular, provided herein are systems and methods for selectively enriching desired nucleic acid molecules in a sample containing undesired nucleic acid molecules. [Means for solving the problem]

[0005] For example, in some embodiments, methods are provided herein that include contacting a sample (e.g., containing two or more different nucleic acid molecules) with a probe and a reagent for selectively modifying (e.g., cleaving) a probe hybridized to a first nucleic acid molecule relative to a probe hybridized to a second, different nucleic acid molecule. In some embodiments, the method further includes enriching or depleting the first nucleic acid molecule relative to the second nucleic acid molecule.

[0006] For example, in some embodiments, the method includes enriching or depleting a first nucleic acid molecule in a sample containing a mixture of nucleic acid molecules by contacting the first nucleic acid molecule with a probe that has a different complementarity to a target region on the first nucleic acid molecule than other nucleic acid molecules in the sample; performing a cleavage reaction; and enriching or depleting the first nucleic acid molecule. In some embodiments, the probe has a higher complementarity to the target region of the first nucleic acid molecule than the corresponding target region of a second nucleic acid in the sample. In some embodiments, the probe has a lower complementarity to the target region of the first nucleic acid molecule than the corresponding target region of the second nucleic acid in the sample. In some embodiments, the first and second nucleic acids differ by sequence variation (e.g., point mutation, deletion, insertion, multi-nucleotide change, fusion, etc.). In some embodiments, the probe comprises a sequence that is fully complementary to the target region of the highly complementary sequence and has one or more mismatches with the sequence variation found in the corresponding target region of the less complementary sequence. In some embodiments, the probe contains one or more mismatches with the target region of both the first and second nucleic acid molecules, but contains more mismatches with the target region of the less complementary nucleic acid. In particular, the probe is designed so that when the probe hybridizes to a first nucleic acid, cleavage of the probe is different from that of a second nucleic acid, allowing for selective enrichment or depletion of the first nucleic acid relative to the second nucleic acid.

[0007] For example, in some embodiments, provided herein are methods for increasing or decreasing the ratio of a first nucleic acid sequence to a second nucleic acid sequence in a sample, comprising: a) exposing a sample containing first and second nucleic acid sequences to probes that differ in complementarity to the first and second nucleic acid sequences; b) modifying the probes (e.g., performing a cleavage reaction); and c) enriching or depleting the first nucleic acid sequence relative to the second nucleic acid sequence.

[0008] A probe may comprise one or more components that are removed before or during cleavage. For example, a probe may comprise a non-complementary flap or other blocking group at the 3' or 5' end, which prevents digestion or polymerase reaction until the blocking group is removed. The blocking group can be removed by any suitable mechanism (e.g., enzymatic cleavage, chemical reaction, temperature change, physical cleavage, etc.). Any suitable blocking group can be used, including, but not limited to, the use of phosphorothioate linkages, modified bases (e.g., 2'-O-methyl, 2'-fluoro, etc.), inverted or dideoxynucleotides, phosphorylation, and spacers. The sequence of the probe that provides differential cleavage products when hybridized to distinct nucleic acid molecules can be located in any suitable position in the initial probe. For example, a mismatch sequence can be located at the 3'-terminal base of the 3' end of the probe. The mismatch sequence can be located within the 3' end of the probe. The mismatch sequence can be located in the center of the probe. The mismatch sequence can be located within the 5' half of the probe or at the 5' end of the probe (eg, at the 5' terminal base).

[0009] The 5' or 3' end of the probe may include a region (e.g., a 5' or 3' tail) that serves as an identifier (e.g., a sample identifier). Such a sequence is useful, for example, for selectively pulling down captured molecules of a specific sample or specific region from a mixed sample. Such an identifier may be particularly useful in multiplex reactions in which multiple different targets are reacted in the same sample or the same reaction vessel.

[0010] In some embodiments, methods use a combination of probes in the same sample preparation, some of which are designed to enrich or deplete variant-containing sequences (e.g., as described above), and some of which are designed simply to capture regions of interest (e.g., using any known hybridization / capture technology or approach). For example, in some embodiments, such combination methods are useful when MSI or copy number variants are analyzed in the same sequencing run as somatic variants: somatic variants are enriched, while genes / regions used for MSI / CNV analysis are captured by standard methods. One potential problem is that standard hyb / cap probes may be undesirably modified by the enzymes and / or reagents used in the enrichment / depletion methodology. To avoid this, in some embodiments, standard hybridization / capture oligonucleotides are made resistant to the modifying enzymes / reagents. For example, standard hybridization / capture oligonucleotides can be modified by using RNA rather than DNA, by using modified bases or backbone modifications, or by introducing intentionally mismatched regions into the probe. There are several scenarios in which it may be beneficial to perform simultaneous probe hybridization of the standard probe and the enrichment / depletion probe, followed by their subsequent separation. In some embodiments, this is done by using different attachment chemistries for the different probe types, or by different 5' or 3' sequences on the probes that can be differentially captured on the solid support via hybridization to complementary "linker" oligos that contain the attachment moiety (or that are pre-linked to the solid support).

[0011] The analytes / sequences to which the methods of the present invention can be applied are nucleic acids, such as naturally occurring or synthetic DNA or RNA molecules, that contain the target polynucleotide sequence being sought. In some embodiments, the analytes / sequences are typically present in an aqueous solution containing the analyte / sequence and other biological materials; in some embodiments, the analyte / sequence is present along with other background nucleic acid molecules that are not of interest for testing. In some embodiments, the abundance of the analyte / sequence is low relative to these other nucleic acid components. In some embodiments, for example, when the analyte is derived from a biological specimen containing cellular material, some or all of these other nucleic acids and extraneous biological materials are removed using sample preparation techniques, such as filtration, centrifugation, chromatography, or electrophoresis, before performing the enrichment or depletion methods described herein. In some embodiments, the DNA or RNA molecules are included in a mixture that contains a sequencing library. The sequencing library may be derived from and / or contain either or both single-stranded or double-stranded molecules, and may contain DNA and / or RNA. The library sequences may include modifications, such as adapters, unique molecular indexes (UMIs), or primer binding sequences, and may be prepared by any suitable process (e.g., tagged fragmentation).

[0012] The compositions and methods of the present invention can be used with any type of sample, including, but not limited to, environmental (e.g., water, soil, air, etc.) samples and biological samples. Biological samples can be from any source, including plants, animals, and infectious disease pathogens. Suitably, in some embodiments, the analytes / sequences are derived from biological samples collected from mammalian subjects (particularly human patients), such as blood, plasma, sputum, urine, skin, biopsy, or surgical resection samples. In some embodiments, the biological sample is subjected to lysis to release the analytes / sequences by disrupting any cells present. In other embodiments, the analytes / sequences may already be present in the sample itself in a free form (e.g., cell-free DNA circulating in blood or plasma). The compositions and methods of the present invention are particularly useful with previously challenging sample types where the allele frequency of the analyte of interest may be low. Such samples include blood, urine, cytosponge-collected samples (e.g., esophageal samples), bronchoalveolar lavage (BAL)-derived samples, pleural fluid, and cerebrospinal fluid (CSF).

[0013] In some embodiments, the sample is a pooled sample. Pooled samples involve mixing multiple samples together into a batch, which is then tested as a pooled collection. This approach increases the number of individual samples that can be tested using a more limited amount of resources. Pooled samples of interest include, but are not limited to, donated blood samples, agricultural samples, food samples, sperm samples, and biological samples that are tested for the presence of infectious disease pathogens (e.g., SARS-CoV-2, HIV, HCV, etc.). In some embodiments, the pooled sample is an environmentally collected sample (e.g., a wastewater sample) that, by its nature of production, has been pooled from multiple different sources. While pooling samples can reduce the allele frequency of variants due to dilution of the samples, it can significantly increase the efficiency of screening. The techniques provided herein are particularly well-suited for analyzing pooled samples, as they enable detection at very low allele frequencies. In some embodiments, portions of each initial sample are pooled without the use of barcoding or other complex preparation steps, and the pooled sample is tested. If a positive result is obtained, the remainder of the non-pooled samples can be tested individually.

[0014] Also provided herein are compositions (e.g., reagents, kits, reaction mixtures, instruments, software) useful in the methods described herein. For example, in some embodiments, compositions are provided herein that include one or more reagents necessary, sufficient, or beneficial for performing the methods described herein. For example, in some embodiments, the compositions include one or more probes that include sequences that are differentially complementary to a known first sequence and a known second sequence (e.g., perfectly complementary to the known first sequence but imperfectly complementary to the known second sequence); one or more reagents that selectively modify probes hybridized to desired nucleic acids relative to probes hybridized to non-desired nucleic acids; and / or one or more agents that cleave or digest the probes. In some embodiments, the compositions further include a target nucleic acid isolation component that separates target nucleic acid molecules. In some embodiments, the compositions include one or more solid supports. In some embodiments, the compositions include one or more buffers. In some embodiments, the solid supports are beads (e.g., magnetic or paramagnetic beads). In some embodiments, the composition further comprises one or more epigenetic modification-sensitive or -dependent restriction enzymes. In some embodiments, the composition further comprises one or more restriction endonucleases. In some embodiments, the composition further comprises one or more transposomes. In some embodiments, the composition comprises a Cas protein (e.g., Cas9). In some embodiments, the composition further comprises one or more transposases. In some embodiments, the composition further comprises one or more ligases. In some embodiments, the composition further comprises one or more blocking oligonucleotides. In some embodiments, the composition further comprises reagents for performing an amplification (e.g., PCR), sequencing (e.g., next-generation sequencing), or detection reaction. In some embodiments, the composition further comprises one or more molecular beacon probes. In some embodiments, the one or more molecular beacon probes are fluorescently labeled. In some embodiments, the composition further comprises components for transcribing RNA into cDNA.

[0015] In some embodiments, the configuration is a reaction mixture containing a reaction solution at a particular time point of any of the methods described herein. In some embodiments, the reaction mixture contains a probe / nucleic acid hybridization complex of the methods described herein. In some embodiments, the reaction mixture contains a captured nucleic acid molecule of the methods described herein. In some embodiments, the reaction mixture contains a region containing a concentration of a desired target nucleic acid that is higher or lower than the concentration of the desired target nucleic acid present in the sample subjected to the digestion reaction. For example, in some embodiments, a reaction mixture is provided herein that contains a sample; a reagent for modifying a probe hybridized to a target nucleic acid; a first nucleic acid molecule from the sample, the first nucleic acid molecule hybridized to a probe having a sequence complementary to that of the first nucleic acid molecule at a discriminator region of the probe; and a second nucleic acid molecule from the sample, the second nucleic acid molecule hybridized to a probe having a sequence that is not completely complementary to that of the second nucleic acid molecule at a discriminator region of the probe.

[0016] Uses of the compositions (e.g., use of kits, use of reaction mixtures, use of reagents, use of instruments, use of software) are also provided herein, e.g., use of the compositions to enrich or deplete target nucleic acids in a sample.

[0017] In some embodiments, provided herein are devices and instruments useful in the methods described herein. In some embodiments, the devices and instruments are useful for collecting and distributing samples into reaction vessels. In some embodiments, the devices and instruments provide reaction chambers for carrying out the methods. In some embodiments, the devices and instruments provide multiple zones or regions (e.g., wells, channels, etc.) for containing reaction solutions and / or for isolating enriched desired target nucleic acids or depleting desired target nucleic acids. In some embodiments, the devices and instruments are useful for amplifying or sequencing nucleic acid molecules. In some embodiments, the devices and instruments are useful for detecting nucleic acid molecules. In some embodiments, the devices and instruments are useful for receiving and communicating information from a user. For example, the devices and instruments may include a user interface for receiving user instructions and a display for visually presenting results to a user.

[0018] In some embodiments, a computing device is provided herein. The computing device is useful for controlling an instrument or device for facilitating the methods described herein. In some embodiments, the computing device collects, analyzes, and reports data. In some embodiments, the computing device comprises one or more processors that execute computer programs. In some embodiments, the computing device comprises a non-transitory computer-readable medium (e.g., software) that includes instructions that direct the processor to perform one or more computing steps.

[0019] In some embodiments, a method is provided herein for enriching or depleting a first nucleic acid molecule in a sample containing a mixture of nucleic acid molecules, the method comprising: contacting the sample with a probe whose complementarity to a target region of the first nucleic acid molecule is different from that of a second nucleic acid molecule in the sample; optionally activating the probe hybridized to the first nucleic acid molecule by selectively modifying the probe hybridized to the first nucleic acid molecule relative to the probe hybridized to the second nucleic acid molecule; selectively digesting the probe hybridized to the first nucleic acid molecule or the second nucleic acid molecule relative to other probes; and enriching or depleting the first nucleic acid molecule. In some embodiments, the first and second nucleic acid molecules comprise end-repaired nucleic acid molecules. In some embodiments, the first and second nucleic acid molecules are A-tailed nucleic acid molecules. In some embodiments, the first and second nucleic acid molecules comprise one or more adapter sequences.

[0020] In some embodiments, the first and second nucleic acid molecules are amplified. In some embodiments, the probe is activated by contact with a cleavage agent. In some embodiments, the cleavage agent is selected from the group consisting of restriction endonucleases, flap endonucleases, mismatch repair enzymes, RNases, Cas proteins, Argonaute family enzymes, DNA-formamidopyrimidine glycosylases (Fpg), apurinic / apyrimidinic (AP) endonucleases (APE1), and chemical cleavage agents.

[0021] In some embodiments, digesting comprises contacting the probe with an exonuclease or endonuclease. In some embodiments, the exonuclease is a 3' to 5' exonuclease. In some embodiments, the exonuclease is a 5' to 3' exonuclease.

[0022] In some embodiments, the probe comprises a binding moiety at its 3' or 5' end or internally, which is optionally biotin or a sequence to which a linker molecule can hybridize, and the linker molecule is modified to bind to a solid support.

[0023] In some embodiments, the method further comprises capturing the probes on a surface prior to digesting. In some embodiments, the surface comprises beads.

[0024] In some embodiments, the method further comprises differentially releasing the first nucleic acid molecule or the second nucleic acid molecule from the probe. In some embodiments, releasing comprises increasing the temperature. In some embodiments, releasing comprises changing the pH. In some embodiments, releasing comprises changing the salt concentration. In some embodiments, releasing comprises digesting. In some embodiments, releasing occurs by melting after a cleavage event without changing other reaction conditions.

[0025] In some embodiments, the method further comprises detecting the first or second nucleic acid molecule, hi some embodiments, detecting comprises sequencing the first or second nucleic acid molecule.

[0026] In some embodiments, the probe comprises a base that is complementary to a position in the first nucleic acid molecule and mismatches a corresponding position in the second nucleic acid molecule. In some embodiments, digesting comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests complementary strands over non-complementary strands (e.g., an enzyme that terminates at mismatched positions). In some embodiments, digesting comprises contacting the probe hybridized to the first and second nucleic acid molecules with an enzyme that preferentially digests non-complementary strands over complementary strands.

[0027] In some embodiments, the probe comprises a 3'-terminal, 5'-terminal, or internal blocking group. In some embodiments, probe activation is performed using a cleaving agent that removes the blocking group from a probe hybridized to the first nucleic acid molecule but not from a probe hybridized to the second nucleic acid molecule. In some embodiments, digesting comprises contacting the probe with a nuclease that cleaves a probe lacking a blocking group but not a probe with a blocking group. In some embodiments, probe activation is performed using a cleaving agent that removes a nucleic acid fragment comprising the blocking group from a probe hybridized to the first nucleic acid molecule but not from a probe hybridized to the second nucleic acid molecule. In some embodiments, digesting comprises contacting the nucleic acid fragment with a polymerase under conditions such that the fragment is extended and the probe hybridized to the first nucleic acid molecule is displaced or digested using a polymerase with 5' to 3' exonuclease activity, optionally with an upstream primer.

[0028] In some embodiments, the activating and digesting does not include pyrophosphorolysis.

[0029] In some embodiments, the method further comprises contacting the sample with a capture probe, wherein the capture probe hybridizes to the second nucleic acid molecule or to another nucleic acid molecule in the sample that is not the first nucleic acid molecule or the second nucleic acid molecule.

[0030] In some embodiments, the method includes enriching for molecules with a specific fragmentation profile. In some embodiments, capturing includes capturing nucleic acid fragments using probes with sequence identity to the fragmentation breakpoints and adapter sequences. In some embodiments, capturing can include capturing molecules with specific 5' and 3' breakpoints. In some embodiments, capturing can include sequential hybridization and capture of each breakpoint. In other embodiments, capturing can include hybridization and capture using multiple probes to capture molecules with specific 5' and 3' fragmentation breakpoints. The fragmentation pattern contains information about the nucleosome organization, chromatin structure, gene expression, and nuclease content of the tissue of origin, resulting in a characteristic signature in the form of fragment size, nucleotide motifs at the fragment ends, single-stranded jagged ends, and genomic location of the fragmentation endpoint. For example, fragmentation breakpoints can be used to identify the likely tissue of origin of molecules in cell-free DNA. In oncology, this can be used to identify specific types or subtypes of cancer, particularly in screening tests.

[0031] Further provided herein are kits that contain reagents sufficient, necessary, or beneficial for carrying out the above-described methods on a sample containing said first and second nucleic acid molecules. [Brief explanation of the drawings]

[0032] [Figure 1] FIG. 1 shows an exemplary target enrichment and detection method. [Figure 2] FIG. 1 shows an exemplary target enrichment and detection method. [Figure 3] FIG. 1 shows an exemplary target enrichment and detection method. [Figure 4] FIG. 1 shows an exemplary target enrichment and detection method. [Figure 5] FIG. 1 shows an exemplary target enrichment and detection method. DETAILED DESCRIPTION OF THE INVENTION

[0033] In some embodiments, the systems and methods employ one or more probes that at least partially hybridize to desired and / or undesired sequences. The probes are configured to be differentially cleaved or degraded in the presence of desired sequences relative to undesired sequences. As used herein, the term "target molecule" refers to a nucleic acid molecule in a sample that hybridizes to a probe. In some embodiments, the probe hybridizes to the target molecule at its 3' or 5' end, or internally. In some embodiments, either the probe or the target molecule is attached to a solid surface. In some embodiments, the probe comprises DNA. In some embodiments, the probe comprises RNA. In some embodiments, the probe comprises a mixture of DNA and RNA. In some embodiments, the probe comprises one or more synthetic bases (e.g., locked nucleic acid (LNA)). In some embodiments, the probe comprises a synthetic backbone modification.

[0034] In some embodiments, the systems and methods cleave or digest the probe in the probe / target hybridization complex. In some embodiments, the cleavage / digestion mechanism used is selective for one strand (e.g., the probe). In some embodiments, the cleavage / digestion mechanism is not strand-selective. In some embodiments, the probe or target is protected, for example, by modifications of the primers or adapters used to create the modified nucleic acid (e.g., 3' / 5' end blocking modifications, internal base / backbone modifications, etc.). In some embodiments, the probe or target is protected by introducing modifications during amplification (e.g., phosphorothioate modifications of the backbone).

[0035] In some embodiments, an activation step is used to render the probe susceptible to cleavage / digestion reactions. In some embodiments, the probe is rendered non-digestible by a 3', 5', or internal modification. In some embodiments, the modification is a mismatch to the target. In some embodiments, the mismatch is at the end of the probe (e.g., the 3' or 5' end). In some embodiments, the mismatch is internal. In some embodiments, the mismatch comprises two or more bases. In some embodiments, the mismatch region provides a flap or bubble when the probe hybridizes to the target. In some embodiments, the probe is prepared with one or more modifications (e.g., backbone modifications, e.g., phosphorothioates (PTO), base modifications, linkers, etc.). In some embodiments, the modification is the inclusion of RNA bases in an otherwise DNA probe. In some embodiments, the modification comprises circularizing the probe.

[0036] In some embodiments, such probes are selectively "activated" based on differential hybridization to different targets (e.g., alleles), after which the probes are allowed to be digested. In some embodiments, the activation step introduces a nick or gap into the probe.

[0037] In some embodiments, activation is performed using one or more restriction endonucleases. In some embodiments, the restriction endonuclease is a methylation / modification-specific endonuclease. In some embodiments, to prevent both strands from being cleaved by the cleavage enzyme, one strand may contain modified bases or backbone modifications (e.g., phosphorothioate linkages) to block digestion. In some embodiments, activation is performed using a flap endonuclease. In some embodiments, activation is performed using a mismatch repair enzyme or a mismatch-specific endonuclease (e.g., CeI, S1, T7E1, Surveyor). In some embodiments, activation is performed using an RNase enzyme (e.g., RNase H2). In some embodiments, activation is performed using a CRISPR / Cas system (e.g., including the use of a Cas enzyme or a mutant that cleaves only one strand). In some embodiments, activation is achieved using an Argonaute family enzyme (e.g., Ago1, Ago2, Ago3, Ago4, Hili, Hiwi, Hiwi2, Hiwi3, aubergine, PIWI, Ago5, Ago6, Ago7, Ago8, Ago9, Ago10, Alg-1, Alg-2, etc.). In some embodiments, activation is achieved using chemical cleavage of mismatches (CCM) (e.g., hydroxylamine plus potassium permanganate followed by piperidine).

[0038] In some embodiments, activation is achieved using a DNA repair enzyme. For example, in some embodiments, the method involves hybridization of a probe designed to create a deliberate G:A mismatch at the desired cleavage site. The DNA repair enzyme MutY glycosylase recognizes the mismatch structure and selectively removes the mismatched A from the duplex, creating an abasic site in the target strand. When an AP-endonuclease, such as endonuclease IV, is added, it cleaves the backbone and splits the DNA strand into two fragments.

[0039] In some embodiments, digestion of probes is selective based on differential hybridization to targets. In some embodiments, digestion of probes uses a general digestion method that is selective for activated probes. In some embodiments, the direction of digestion is 3' to 5'. In some embodiments, the direction of digestion is 5' to 3'. In some embodiments, digestion is by internal cleavage of the probe. In some embodiments, digestion does not use pyrophosphorolysis (cleavage in which inorganic phosphate is the attacking group).

[0040] In some embodiments where 3' to 5' digestion is desired, a 3' to 5' exonuclease is used, including, but not limited to, exonuclease I (ExoI) (e.g., E. coli ExoI), exonuclease T (ExoT), exonuclease VII (ExoVII), exonuclease III (ExoIII), and DNS.

[0041] In some embodiments where 3' to 5' digestion is desired, a polymerase with 3' to 5' exonuclease activity is used (e.g., a proofreading polymerase such as PHUSION (Thermo Fisher), Q5 (New England Biolabs), KAPA HIFI (Roche), KOD DNA polymerase).

[0042] In some embodiments where 5' to 3' digestion is desired, a 5' to 3' exonuclease is used (e.g., lambda exonuclease, RecJf, T7 exonuclease, exonuclease V (ExoV), exonuclease VIII (ExoVIII), T5 exonuclease).

[0043] In some embodiments where 5' to 3' digestion is desired, a DNA polymerase with 5' to 3' exonuclease activity is used. In some such embodiments, a 5' (or internally) blocked probe can act as a blocking oligonucleotide, preventing primer extension to the target. Selective cleavage of the probe based on differential hybridization results in melting of the blocked 5' end, allowing extension of the upstream primer while the remaining hybridized probe is digested in the 5' to 3' direction (or displaced by strand displacement), liberating the target molecule. In similar embodiments, the 5' section of the cleaved probe can itself act as a primer.

[0044] For internal cleavage, any of the "activation" methods described above can be used as the sole digestion method. In some embodiments, the fragments obtained from the cleaved probe have a melting temperature (T m ) is lower than that of the intact probe, allowing selective denaturation from the target by increasing temperature, pH, etc.

[0045] In some embodiments, the systems and methods use selective digestion as a method for enriching or depleting nucleic acid sequences. In some embodiments, the digestion reaction relies on complementarity between hybridized strands and digests only strands with specific properties (e.g., specific sequences or structures). This reaction selectively shortens or digests certain sequences with specific properties, while sequences that do not have these properties remain undigested or less digested. This reaction can be performed to recover and analyze molecules hybridized to the shortened or more digested sequences. Alternatively, this reaction can be performed to analyze more shortened or more digested sequences. Alternatively, this reaction can be performed to analyze undigested sequences.

[0046] A sample can be enriched or depleted one or more times to further enrich for sequences of interest. For example, in some embodiments, a second round of enrichment or depletion is performed after the first round is completed using the same reagents as the first round. In some embodiments, an amplification reaction can be used between the first and second rounds of enrichment or depletion. In other embodiments, different probes selective for different sequences of the target nucleic acid to be enriched or removed are used. This approach is particularly well-suited when the sequence to be enriched or depleted differs from the sequence to be removed at least two base positions. For example, a target nucleic acid containing two polymorphisms relative to the wild type can undergo a first round of enrichment or depletion based on the first polymorphism and a second round of enrichment or depletion based on the second polymorphism. Exponential enrichment or depletion can be achieved by using multiple rounds. In some embodiments, the target nucleic acid is modified (e.g., polymorphisms are added) to generate synthetic sequences prior to enrichment or depletion, so that the synthetic sequences are subject to enrichment or depletion relative to sequences that do not contain the synthetic sequences. In some embodiments, two or more nucleic acids can be enriched or depleted using multiple probes in any given round of reaction, hi some embodiments, a particular nucleic acid can be enriched while a second nucleic acid is depleted in one or more rounds of reaction.

[0047] In one aspect of the present invention, a method is provided for altering the ratio of a first nucleic acid sequence to a second nucleic acid sequence in a sample comprising at least a first and a second sequence, the method comprising: a) introducing the sample containing one or more nucleic acid analytes into a first reaction mixture comprising: i) probes that are differentially complementary to the first and second sequences (e.g., the 3' end or other region of the probe is perfectly complementary to one of the first or second sequences but imperfectly complementary to the other); ii) optionally, an activating enzyme or reagent; and iii) a nuclease that selectively cleaves or digests probes hybridized to the first sequence relative to the second sequence; and b) separating the first sequence, or its cleavage or digestion products, from the second sequence, or its cleavage or digestion products.

[0048] In some embodiments, separation of any probe complexes with better (e.g., complete) annealing is achieved using reaction conditions that favor probe-sequence complexes with better (e.g., complete) annealing over probe-sequence complexes with poorer (e.g., incomplete) annealing. This can take the form of changing the temperature of the reaction mixture, changing the pH of the reaction mixture, and / or changing the salt concentration of the reaction mixture. In some embodiments, chemical agents (e.g., dimethyl sulfoxide (DMSO), formamide, etc.) are used to denature nucleic acids to facilitate concentration or depletion. In some embodiments, nucleic acid molecules (displaced oligonucleotides) and enzymes or proteins are used to separate hybridized nucleic acid molecules.

[0049] In some embodiments, two probes are used, one for the forward strand and one for the reverse strand of a double-stranded target nucleic acid, and in some embodiments, it is beneficial to design the probes so that they do not hybridize to each other in a manner that would interfere with the desired reaction.

[0050] In some embodiments, the probe is captured on a solid support before or after exposure to a sample containing the nucleic acid molecules to be enriched or depleted. In some embodiments, the probe contains a biotin moiety that is captured by a corresponding streptavidin moiety on the solid support. In some embodiments, the probe hybridizes to a linker molecule that itself contains a modification, such as a biotin moiety, and is captured on the solid support by this modification. In some embodiments, the solid support is a bead. In some embodiments, the bead is a magnetic or paramagnetic bead.

[0051] The above techniques are not limited to the use of capture to separate or concentrate nucleic acid molecules of interest. Molecules can be separated or concentrated based on differences in, for example, size, charge, or shape, or other physical or chemical properties. In some embodiments, a moiety is added to a nucleic acid of interest (e.g., by click chemistry modification) such that the added moiety confers a selectable distinguishing characteristic to the nucleic acid of interest.

[0052] In some embodiments, one or more wash steps are performed between one or more steps, hi some embodiments, the wash and hybridization steps are performed at elevated temperatures between 25°C and 95°C.

[0053] In some embodiments, the sample comprises an adaptor-tagged library of nucleic acid analytes.

[0054] In some embodiments, the sequences are identified by an amplification reaction (e.g., polymerase chain reaction (PCR), nucleic acid sequence-based amplification (NASBA), multiple displacement amplification (MDA), transcription-mediated amplification (TMA), strand displacement amplification (SDA), rolling circle amplification (RCA), loop-mediated isothermal amplification (LAMP), recombinase polymerase amplification (RPA), helicase-dependent amplification (HDA), and nicking and extension amplification reaction (NEAR)). In some embodiments, the sequences are identified by microarray analysis.

[0055] In some embodiments, the sequence is identified by sequencing. In some embodiments, the sequence is identified by next-generation sequencing (NGS) (e.g., bridge amplification sequencing (Illumina), SMRT sequencing (PacBio), Ion Torrent sequencing, nanopore sequencing, and pyrosequencing, etc.).

[0056] The compositions and methods of the present invention can be used with any type of sample, including, but not limited to, environmental (e.g., water, soil, air, etc.) samples and biological samples. Biological samples can be from any source, including plants, animals, and infectious disease pathogens. Suitably, in some embodiments, the analytes / sequences are derived from biological samples collected from mammalian subjects (particularly human patients), such as blood, plasma, sputum, urine, skin, biopsy, or surgical resection samples. In some embodiments, the biological sample is subjected to lysis to release the analytes / sequences by disrupting any cells present. In other embodiments, the analytes / sequences may already be present in the sample itself in a free form (e.g., cell-free DNA circulating in blood or plasma). The compositions and methods of the present invention are particularly useful with previously challenging sample types where the allele frequency of the analyte of interest may be low. Such samples include blood, urine, cytosponge-collected samples (e.g., esophageal samples), bronchoalveolar lavage (BAL)-derived samples, pleural fluid, and cerebrospinal fluid (CSF).

[0057] In some embodiments, the sample is a pooled sample. Pooled samples involve mixing multiple samples together into a batch, which is then tested as a pooled collection. This approach increases the number of individual samples that can be tested using a more limited amount of resources. Pooled samples of interest include, but are not limited to, donated blood samples, agricultural samples, food samples, sperm samples, and biological samples that are tested for the presence of infectious disease pathogens (e.g., SARS-CoV-2, HIV, HCV, etc.). In some embodiments, the pooled sample is an environmentally collected sample (e.g., a wastewater sample) that, by its nature of production, has been pooled from multiple different sources. While pooling samples can reduce the allele frequency of variants due to dilution of the samples, it can significantly increase the efficiency of screening. The techniques provided herein are particularly well-suited for analyzing pooled samples, as they enable detection at very low allele frequencies. In some embodiments, portions of each initial sample are pooled without the use of barcoding or other complex preparation steps, and the pooled sample is tested. If a positive result is obtained, the remainder of the non-pooled samples can be tested individually.

[0058] The target nucleic acid to be analyzed can be any target nucleic acid of interest.In some embodiments, the analysis is for research, diagnostic or therapeutic purposes.In some embodiments, when the sample is from a human subject, the purpose can be for the analysis, detection, treatment or treatment selection of one or more diseases or pathologies. The techniques are particularly useful in the analysis of multifactorial diseases and disorders, including, but not limited to, cancer (e.g., bladder cancer, breast cancer, cervical cancer, colon cancer, ovarian cancer, uterine cancer, vaginal cancer, vulvar cancer, head and neck cancer, eye cancer, brain cancer, kidney cancer, liver cancer, lung cancer, lymphoma, mesothelioma, myeloma, prostate cancer, skin cancer, thyroid cancer, pancreatic cancer, bone cancer, esophageal cancer, gallbladder cancer, stomach cancer, testicular cancer, anal cancer, rectal cancer, oral cancer, salivary gland cancer, sarcoma, and thyroid cancer), autoimmune diseases, asthma, cilia-related diseases, cleft palate, diabetes, heart disease, high blood pressure, inflammatory bowel disease, intellectual disability, mood disorders, obesity, infertility, and refractive errors.

[0059] In some embodiments, the nucleic acid being analyzed is a methylated sequence or is derived from a methylated sequence. The techniques described above are used to distinguish the methylation state at any particular position or positions in a target nucleic acid. The methylated sequence can first be modified by using chemical treatment (e.g., oxidation, reduction, bisulfite treatment), or by exposure to a methylation-dependent restriction enzyme, or any other suitable approach, followed by enrichment and / or identification of the modified sequence.

[0060] In some embodiments, the targeted region of RNA present in the sample is transcribed into DNA.

[0061] Some embodiments of the above techniques use solid surfaces. Any solid surface compatible with nucleic acid capture can be used, directly or indirectly. Solid supports include, but are not limited to, those made of glass, metal, gel, plastic (e.g., polystyrene), ceramic, and filter paper, among others. In some embodiments, the solid support is a bead (e.g., paramagnetic bead), microparticle, nanoparticle, column, or slide. In some embodiments, the solid support is modified to include one or more of the following functional groups to facilitate reaction with, or coupling or conjugation to, a biomolecule: an amine group, a carboxylic acid group, a sulfonic acid group, a trimethylamine group, and / or an epoxide group. In some embodiments, streptavidin is attached to the solid surface. In some embodiments, the solid support is a dextran-modified surface. In some embodiments, the solid support is a polyethylene glycol (PEG) or PEG-modified surface. In some embodiments, the solid support is a polyvinylpyrrolidone (PVP) or PVP-modified surface. In some embodiments, the solid support is a polysaccharide or a polysaccharide-modified surface. In some embodiments, the polysaccharide is selected from one or more of dextran, ficoll, glycogen, gum arabic, xanthan gum, carrageenan, amylose, agar, amylopectin, xylan, and / or β-glucan. In some embodiments, the solid support is a chemical resin or a chemical resin-modified surface. In some embodiments, the chemical resin or chemical resin-modified surface is selected from one or more of the following resins: isocyanate, glycerol, piperidino-methyl, polyDMAP (polymer-bound dimethyl 4-aminopyridine), DIPAM (diisopropylaminomethyl), aminomethyl, polystyrene aldehyde, tris(2-aminomethyl)amine, morpholino-methyl, BOBA (3-benzyloxybenzaldehyde), triphenyl-phosphine, or benzylthio-methyl.

[0062] In some embodiments, a capture moiety is used to associate a probe or target nucleic acid with a solid surface. In some embodiments, the capture moiety is covalently attached to a solid support via a chemically cleavable linker, such as a disulfide, allyl, or azide-masked hemiaminal ether linker. In some embodiments, the capture moiety is covalently attached to a solid support via an amide or phosphorothioate bond. In some embodiments, the probe or target can incorporate a modified moiety to introduce the capture moiety. In some embodiments, the reaction comprises reacting an alkyne-labeled oligonucleotide with an azidobiotin conjugate.

[0063] In some embodiments, the capture moiety comprises an oligonucleotide sequence and the solid surface comprises a complementary oligonucleotide sequence. In some embodiments, the oligonucleotide sequence comprises one or more modified bases and / or other such modifications known to those of skill in the art to alter the melting temperature. In some embodiments, the presence of one or more modified bases and / or other such modifications known to those of skill in the art results in a lower melting temperature. In some embodiments, the presence of one or more modified bases and / or other such modifications known to those of skill in the art results in an increased melting temperature. In some embodiments, the length of the complementary sequence is between 10, 20, 30, 40, 50, 100, 150, and 200 bases. In some embodiments, the length of the complementary sequence is between 10, 20, 30, 40, 50, and 100 bases. In some embodiments, the length of the complementary sequence is between 10-20, 10-30, 10-40, and 10-50 bases. In some embodiments, the length of the complementary sequence is between 10-20, 10-30, and 10-40 bases. In some embodiments, the length of the complementary sequence is between 10-20 and 10-30 bases. In some embodiments, the length of the complementary sequence is between 10-20 bases.

[0064] In some embodiments, the capture moiety comprises a chemical modification and is attached to the solid support via an interaction between the chemical modification and the solid support. In some embodiments, the chemical modification is biotin, and the solid support further comprises streptavidin. In some embodiments, the captured oligonucleotide sequence is released from the solid support. In some embodiments, the captured oligonucleotide sequence is released from the solid support by chemical denaturation. In some embodiments, chemical denaturation is achieved by using an appropriate concentration of base. In some embodiments, 0.1 M NaOH can be used. In some embodiments, the oligonucleotide sequence is released from the solid support by cleavage of the chemical linker via the addition of tris(2-carboxyethyl)phosphine (TCEP) or dithiothreitol (DTT) in the case of a disulfide linker; via the addition of a palladium complex or an allyl linker; or via the addition of TCEP in the case of an azide-masked hemiaminal ether linker. In some embodiments, the oligonucleotide sequence is released from the solid support by removing a non-standard base from the oligonucleotide sequence and cleaving at the resulting abasic site. In some embodiments, the non-standard base is uracil, which is removed by uracil DNA glycosylase. In alternative embodiments, the non-standard base is 8-oxoguanine, which is removed by formamidopyrimidine DNA glycosylase (Fpg).

[0065] In some embodiments, the capture moiety is an oligonucleotide region and release is achieved by heating the reaction mixture. In some embodiments, the reaction mixture is heated to 37°C to 100°C for 1 to 20 minutes. In some embodiments, the reaction mixture is heated for 1 to 15 minutes. In some embodiments, the reaction mixture is heated for 1 to 10 minutes. In some embodiments, the reaction mixture is heated for 1 to 5 minutes. In some embodiments, the reaction mixture is heated for 5 minutes. In some embodiments, the reaction mixture is heated to 37°C to 85°C. In some embodiments, the reaction mixture is heated to 37°C to 75°C. In some embodiments, the reaction mixture is heated to 37°C to 65°C. In some embodiments, the reaction mixture is heated to 37°C to 55°C. In some embodiments, the reaction mixture is heated to 37°C to 45°C.

[0066] Those skilled in the art will appreciate that the temperature to which the reaction mixture is heated to release the complementary oligonucleotide regions will depend on many factors, including the length of the regions.

[0067] In some embodiments, liberation is achieved by cleavage of one or more oligonucleotide sequences. This cleavage can be achieved by any of the means described above or below, or by any means known to those of skill in the art. In some embodiments, the oligonucleotide sequence is cleaved chemically. In some embodiments, the oligonucleotide sequence is cleaved enzymatically. In some embodiments, the oligonucleotide sequence is cleaved by a restriction enzyme. In some embodiments, the oligonucleotide sequence is cleaved by an epigenetic modification-sensitive or -dependent restriction enzyme. In some embodiments, the oligonucleotide sequence is cleaved by a methylation-sensitive or -dependent restriction enzyme. In some embodiments, the oligonucleotide sequence is cleaved by a hydroxymethylation-sensitive or -dependent restriction enzyme.

[0068] In some embodiments, the sequences are enzymatically or chemically converted before or after enrichment for variant or wild-type sequences to allow for detection of methylation status. Those skilled in the art will understand that the term "enrichment" refers to the selective isolation of target sequences from a mixture of target and non-target sequences as described above or below.

[0069] In some embodiments, the oligonucleotide sequence comprises a photocleavable linker, and the oligonucleotide sequence is released from the solid support by cleavage of the linker (e.g., by UV light). The modification may be a chemical backbone modification.

[0070] In some embodiments, the sample nucleic acid is fragmented prior to applying the methods disclosed herein. Those skilled in the art will appreciate that there are multiple techniques that can be used to fragment DNA. These include sonication, needle shearing, nebulization, point-sink shearing, passage through a pressure cell (French press), and enzymatic methods. In some embodiments, fragmentation is achieved by sonication. In some embodiments, a Bioruptor® (Denville, NJ) device can be used. In some embodiments, fragmentation is achieved by acoustic shearing. In some embodiments, a Covaris® instrument (Woburn, MA) can be used. In some embodiments, fragmentation is achieved by nebulization. Nebulization involves forcing DNA through small holes in a nebulizer unit, resulting in the formation of a fine mist that is collected. Fragment size is determined by the pressure of the gas used to force the DNA through the nebulizer, the rate at which the DNA solution passes through the holes, the viscosity of the solution, and the temperature. In some embodiments, fragmentation is achieved by hydrodynamic shearing. In some embodiments, Hydroshear from Digilab (Marlborough, MA) can be used. In some embodiments, fragmentation is achieved by point sink shearing. In some embodiments, fragmentation is achieved by needle shearing. In some embodiments, fragmentation is achieved by use of a French press. In some embodiments, fragmentation is achieved by enzymatic fragmentation. In some embodiments, fragmentation is achieved by restriction endonuclease digestion. In some embodiments, fragmentation is transposome-mediated fragmentation. In some embodiments, fragmentation is achieved by Cas9. In some embodiments, fragmentation is achieved by Cas9, as described in US10577644, which is incorporated herein by reference in its entirety. In some embodiments, one or more different fragmentation techniques can be used. In some embodiments, one or more of the same or different fragmentation techniques can be used at one or more different points in the above methods.

[0071] In some embodiments, sequence fragmentation and adapter tagging occur at the same time or step in the above methods, one such example being the Nextera DNA Library Prep Kit by Illumina.

[0072] Those skilled in the art will appreciate that there are multiple techniques that can be used to prepare adaptor-tagged sequences / libraries.

[0073] In some embodiments, following fragmentation, the nucleic acids may be ligated to one or more adapters after polishing and A-tailing the ends.

[0074] In some embodiments, following fragmentation, the nucleic acids can be end-polished and ligated to adapters in a blunt-end ligation reaction.

[0075] In some embodiments, following fragmentation, adaptors are ligated to the single-stranded DNA.

[0076] In some embodiments, following fragmentation, a terminal transferase enzyme is used to add non-templated bases to the 3' ends of the fragments, providing priming sites for making the fragments double-stranded.

[0077] In some embodiments, a topoisomerase can be used in place of DNA ligase.

[0078] In some embodiments, TOPO cloning can be used to add adapters to the fragmented DNA.

[0079] In some embodiments, following fragmentation, a transposase can be used to add adapter sequences to the nucleic acid.

[0080] In some embodiments, following fragmentation, a standard transposon can be used, but then modified to create a Y-shaped adapter using oligonucleotide substitutions.

[0081] In some embodiments, when the sample is an adaptor-tagged library, blocking oligonucleotides are used to prevent cross-hybridization of library molecules (so-called "daisy-chaining").

[0082] In some embodiments, the sample comprises one or more blocking oligonucleotides.

[0083] Any sequencing methodology can be used to analyze the target nucleic acid molecule. In some embodiments, the sequencing is Maxam-Gilbert sequencing. In some embodiments, the sequencing is Sanger sequencing. In some embodiments, the sequencing is shotgun sequencing. In some embodiments, the sequencing is single-molecule real-time sequencing. In some embodiments, the sequencing is ion semiconductor sequencing. In some embodiments, the sequencing is pyrosequencing. In some embodiments, the sequencing is sequencing-by-synthesis. In some embodiments, the sequencing is combinatorial probe-anchor synthesis (cPAS). In some embodiments, the sequencing is sequencing-by-ligation. In some embodiments, the sequencing is nanopore sequencing. In some embodiments, the sequencing is GenapSys sequencing. In some embodiments, the sequencing is next-generation sequencing (NGS).

[0084] In some embodiments, methods are provided for screening patients, comprising detecting the presence or absence of one or more specific nucleic acid sequences in a sample from the patient using an embodiment of the method described either above or below.

[0085] Those skilled in the art will appreciate that such screening is useful for monitoring patients undergoing treatment for one or more disease conditions, where the treatment status can be ascertained by the level of one or more nucleic acid sequences in the patient's sample.

[0086] For example, the treatment status of a patient undergoing treatment for one or more cancers can be confirmed by the level of one or more nucleic acid sequences in blood and / or the presence and / or absence of one or more specific variants.High levels of circulating tumor nucleic acid sequences and / or the presence and / or absence of one or more specific variants can be used to predict whether a specific treatment has a desired effect.Therefore, provided is a method for monitoring whether a specific treatment is successful, and this success can be predicted by the presence or absence of specific nucleic acid sequences and / or their respective levels in patient-derived samples.

[0087] In some embodiments, methods are provided for monitoring patients in remission to detect disease recurrence.

[0088] In some embodiments, methods are provided for screening apparently healthy individuals to detect the presence of one or more disease states, including but not limited to cancer.

[0089] In some embodiments, methods are provided for detecting the presence and / or absence of one or more genetic markers in a patient diagnosed with one or more disease states, and using the presence and / or absence of the one or more markers to determine which treatment the patient should receive.

[0090] In some embodiments, methods are provided for diagnosing and / or monitoring one or more cancers in a patient, comprising detecting the presence or absence of one or more specific nucleic acid sequences in a sample from the patient using an embodiment of the method described either above or below.

[0091] Those skilled in the art will understand that one or more particular nucleic acid sequences may be specific to an individual (e.g., identified from a tissue biopsy or surgical resection by an identification method such as sequencing), and that in such cases, a panel specific to the individual patient may be used.

[0092] Those skilled in the art will further appreciate that in some embodiments, the panel covers known hotspot regions of the human genome, i.e., regions that are recurrently mutated in a given cancer type.

[0093] Those skilled in the art will further understand that in some embodiments, the panel covers targeted regions or the entire human exome.

[0094] In some embodiments, methods for non-invasive prenatal testing (NIPT) are provided, comprising detecting the presence or absence of one or more specific nucleic acid sequences in a sample from a patient using an embodiment of the method described above or below, wherein the patient is a pregnant patient. In some embodiments, the sample is plasma and / or serum from the pregnant patient's blood. In some embodiments, the methods provided herein are used to enrich and / or quantify the fetal fraction of a sample using a panel of common SNPs associated with such sample.

[0095] In some embodiments, there is provided a method of treating a patient, comprising: - carrying out any of the embodiments of the invention described above or below to detect the presence or absence of one or more specific nucleic acid sequences in a sample from a patient; - making one or more treatment decisions based on the presence or absence of said sequence.

[0096] In some embodiments, the treatment decision is to initiate a particular treatment. In some embodiments, the treatment decision is to discontinue a particular treatment. In some embodiments, the treatment decision is to increase the dose of a particular therapeutic agent. In some embodiments, the treatment decision is to decrease the dose of a particular therapeutic agent. In some embodiments, the treatment decision is to increase the frequency of administration of a particular therapeutic agent. In some embodiments, the treatment decision is to decrease the frequency of administration of a particular therapeutic agent. In some embodiments, the treatment decision is to add an additional drug to an existing treatment regimen. In some embodiments, the treatment decision is to remove a drug from an existing treatment regimen.

[0097] In some embodiments, a kit is provided that includes one or more or all of the components necessary, sufficient, or useful for carrying out the methods described herein. For example, in some embodiments, the kit includes one or more probes, solid supports, enzymes, blocking oligonucleotides, buffers, detergents (e.g., sodium dodecyl sulfate (SDS), TWEEN 20, etc.), adapters, crowding agents (e.g., polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), dextran sulfate, etc.), solvents (e.g., formamide, ethylene carbonate, etc.), additives that hybridize to repeat sequences (e.g., COT-1 DNA, salmon sperm DNA, and oligonucleotides that block ribosomal RNA), capture moieties (e.g., biotin; e.g., as part of a probe), metal ions, blocking oligonucleotides, sequencing reagents, amplification reagents (e.g., isothermal amplification reagents; exponential amplification reagents (e.g., thermostable polymerase, primers, dNTPs, buffers, labeled detection probes)), transcription reagents, instructions, software, instruments, positive controls, and negative controls, etc. The one or more containers may individually house one or more components.

[0098] In some embodiments, a kit is provided that includes a plurality of the probes described above or below.

[0099] In some embodiments, kits are provided that include between 1 and 1,000,000 distinct probes. In some embodiments, kits are provided that include between 1 and 100,000 distinct probes. In some embodiments, kits are provided that include between 1 and 10,000 distinct probes. In some embodiments, kits are provided that include between 1 and 1,000 distinct probes.

[0100] In some embodiments, bioinformatic approaches are used to analyze sequencing data. In some embodiments, the presence or absence of specific variants is assessed. In other embodiments, data from multiple variants is combined to derive a probabilistic estimate for the presence or absence of a particular target nucleic acid.

[0101] For example, analysis of Illumina sequencing data involves converting and demultiplexing BCL files to FASTQ format using tools such as bcl2fastq. In some embodiments, sequencing reads include molecular identifiers. In this case, the molecular identifiers can be extracted from the sequencing reads, added to the FASTQ header, and the sequencing reads clipped. In some embodiments, barcodes with non-standard bases (not A, C, G, or T) can be filtered. The resulting reads can then be aligned using a tool, such as bwamem, with the -C option to add the barcode sequence to the alignment. The alignment can then be sorted by coordinate, duplicate reads can be marked, and reads can be annotated with read coordinates, mate coordinates, and optical duplicate auxiliary tags using biobambam2, bamsormadup, and bammarkduplicatesopt. Reads can be filtered if they are not marked as a proper pair or if they are marked as optical duplicates, complementary, QC failed, unmapped, or secondary alignments. Each read can then be marked with an auxiliary tag comprising a reference name, sorted read and mate fragmentation breakpoints, forward and reverse read barcodes, and read strand.

[0102] In some embodiments, sequencing data is analyzed using a variant calling algorithm that uses various subsets of read flags and tags (including read and mate coordinates, optical overlap flags, UMI sequences, MID sequences, alignment scores, and secondary alignment scores). In this case, the analysis of sequencing data compares the probability of data observations under two models. The first model is a null model that specifies the distribution of sequencing artifacts. The second model is a model that allows for true variants. In this case, a variant is called when the probability under the alternative model exceeds the probability under the null model. In some embodiments, a panel of precharacterized samples can help model the error distribution of the first model.

[0103] In some embodiments, auxiliary tags can be used to identify reads that are likely to originate from the same input molecule and / or the same strand of the same input molecule, hi some embodiments, a consensus base quality score can be derived from reads that share the same auxiliary tag.

[0104] In some embodiments, variants are identified using artificial intelligence algorithms, for example, convolutional neural networks.

[0105] In some embodiments, sequencing data can be further filtered to remove artifacts.Exemplary filters include: the number of mismatches present on a given sequencing read; alignment score and suboptimal alignment score; base quality score or consensus base quality score; the minimum number of reads covering a given variant site; the position of the variant in the sequencing read; whether the read is 5'-clipped; whether the read is improperly paired; whether the read contains indels; and the variant allele frequency of a given variant.In some embodiments, the region of genome that contains common SNPs or is prone to alignment artifacts is filtered.Many other filters are known to those skilled in the art.

[0106] In some embodiments, control samples are sequenced to filter out variants. For example, DNA from buccal epithelium or other tissue sources can be sequenced to remove germline variants. In another embodiment, buffy coat or white blood cell DNA can be sequenced to filter out somatic mutations derived from clonal hematopoiesis.

[0107] The compositions, methods, and kits of the present invention are useful in a wide range of applications and settings. In some embodiments, they are useful in any methodology in which it is desirable to detect a sequence in a sample. In some embodiments, they are useful in any methodology in which it is desirable to detect low-abundance (e.g., rare) sequences in a complex sample. In addition to the exemplary uses described above, a number of additional exemplary uses are presented below.

[0108] In some embodiments, the compositions, methods, and kits are useful in the analysis and treatment of infectious diseases. The techniques are particularly valuable for detecting low-frequency mutations that may be present in a sample. For example, the techniques are useful for detecting low-frequency mutations associated with treatment-resistant (e.g., antibiotic-resistant, antiviral-resistant, etc.) infectious diseases (e.g., HIV, tuberculosis, etc.). The techniques are further useful for selective pull-down of bacterial or viral DNA or RNA for sequencing.

[0109] As noted above, the techniques are particularly well suited for the analysis and / or enrichment of analytes in complex samples. One area of ​​growing research and clinical interest is microbiome analysis, where the techniques are useful for selecting desired bacterial DNA with increased specificity for sequencing or other analyses.

[0110] The techniques are also useful for high-throughput and multiplexed analysis of many different samples. These advantages make them useful in a wide variety of genotyping applications, including forensic analysis, paternity / maternity testing, disease analysis (e.g., cancer, infectious diseases), drug susceptibility testing, and agricultural and food testing (e.g., to aid in selective breeding, to identify trace contaminants, etc.).

[0111] The above techniques are useful for error correction of synthetic nucleic acids (e.g., DNA). Synthetic nucleic acids are used in research, diagnostics, and clinical applications. In many cases, it is important to avoid or minimize the use of nucleic acid molecules with unintended or undesired sequences. The above techniques are useful for distinguishing and isolating desired molecules from undesired molecules.

[0112] Nucleic acid editing has emerged as an important process in research, synthetic biology, and clinical applications. For example, CRISPR / CAS editing of nucleic acids and related processes have emerged as important processes. Many of these editing methods result in a mixed population of molecules, including intended edited products, unedited products, and unintended edited products. The techniques provided herein facilitate the identification, selection, and isolation of intended edited products.

[0113] The technique is also useful in environmental monitoring. In addition to agricultural applications, the technique is particularly well suited for analyzing environmental samples that may contain trace amounts of analytes of interest. Such samples include, but are not limited to, analysis of native and invasive organisms, early detection of invasive species, air and water contamination, and ancient DNA analysis. Sample types include, but are not limited to, soil, water, snow, feces, mucus, gametes, shed skin, cadavers, hair, and air.

[0114] Such techniques are useful for isolating desired subsets of nucleic acids from a particular sample from other subsets, for example, they are useful in the isolation and analysis of chloroplast and mitochondrial genomes.

[0115] The above techniques are useful for cell line screening of engineered and natural cells, including, but not limited to, cell cultures (primary and immortalized), stem cells (embryonic, induced pluripotent, dedifferentiated, etc.), differentiated cells intended for cell therapy, ex vivo modified cells for research or clinical applications (e.g., CAR T cells), and genetically engineered cells.

[0116] The techniques are useful for separating and removing damaged or other unwanted nucleic acids from undamaged or desired nucleic acids, for example, they can be used to remove damaged DNA from a sample prior to methylation analysis.

[0117] The techniques are useful for preimplantation screening of cells (e.g., embryos, eggs, sperm), liposomes, exosomes, nucleic acid vectors (e.g., gene therapy vectors), etc., prior to administration to a subject. In some embodiments, the techniques can be applied to culture media to screen live cells without disturbing them.

[0118] The technique is useful in drug toxicity screening and is particularly well suited to identifying DNA damage, mutations, and methylation changes that may be associated with the use of a particular drug.

[0119] The above techniques are useful for nucleic acid fragmentome analysis, for example, by using probes that lie on or match breakpoints that associate specific sequences with relevant correlation information (e.g., tissue of origin, association with disease such as cancer, etc.).

[0120] The techniques can be used in any application where reducing nucleic acid complexity is desirable. For example, the techniques can be used to reduce the complexity of a whole genome. In some such embodiments, a restriction enzyme digestion or other nucleic acid fragmentation step is used, followed by a step of removing only the cleaved molecules using probes that match the known end sequences.

[0121] The above-mentioned technology is useful for evaluating microsatellite instability (MSI).The target nucleic acid molecules that differ in the presence, number, or nature of repeat nucleotides (such as GT / CA repeats) are enriched and / or identified in the sample.MSI is associated with many diseases and pathologies, including but not limited to colon cancer, gastric cancer, endometrial cancer, ovarian cancer, hepatobiliary cancer, urinary tract cancer, brain cancer, and skin cancer.

[0122] The above technology is also useful for assessing tumor mutation burden (TMB). TMB has emerged as a predictive biomarker for immune checkpoint therapy, among other applications. Currently, next-generation whole genome sequencing is employed to assess TMB, or a gene panel that provides the sequences of a subset of genes is evaluated. The use of the technology provided herein enables TMB assessment with greater sensitivity and significantly reduced cost and burden.

[0123] The above technology is useful for haplotype analysis. Genomic information reported as haplotypes rather than genotypes is becoming increasingly important in personalized medicine and a wide variety of research applications. Haplotypes are more specific than less complex variants, such as single nucleotide variants, and are applied in prognosis and diagnosis, tumor analysis, and tissue typing for transplantation. Currently, sequencing is the most common form of molecular haplotype analysis. The error rate of sequencing technology is a barrier to obtaining accurate information. The technology provided herein enables efficient and highly accurate haplotype analysis.

[0124] In some embodiments, assay components are designed to avoid specific polymorphisms (e.g., SNPs). This may be preferable to avoid unwanted carryover of off-target molecules or under-recovery of on-target molecules. In some embodiments, probes are designed not to hybridize to regions containing known polymorphisms (e.g., SNPs). In some embodiments, probes are designed and / or capture configured to target strands of the target sequence that do not contain the polymorphism to be avoided. In some embodiments, multiple probes are used for each target, with each probe targeting a different allele (e.g., SNP allele). In some embodiments, a universal base (e.g., inosine) is added to the probe corresponding to the known polymorphic site (e.g., SNP site), which is digested (or not digested) as if the sequence were a match at the polymorphic position, regardless of whether the position contains the polymorphic sequence or the wild-type sequence.

[0125] The above-mentioned technique can enrich any desired sequence or sequence of interest, so that the above-mentioned technique can enhance existing nucleic acid methodology.For example, many nucleic acid sequencing approaches are difficult to sequence when there are repetitive sequence regions in target nucleic acid.The technique provided herein can remove repetitive regions, making such sequencing reactions more accurate and more efficient. Example [Example]

[0126] Double-strand-specific (mismatch-blocked) 3' to 5' exonuclease digestion The extracted DNA undergoes standard library preparation steps: end repair, A-tailing, and adapter ligation. The first step involves probe hybridization to the target sequence. The probe sequence perfectly matches the mutant target sequence and has at least one mismatch with the wild-type target sequence. The probe has a modification at the 5' end that allows it to be attached to a surface. A double-strand-specific DNA exonuclease digests the probe in the 3'→5' direction. Digestion of probes annealed to WT molecules is halted by the presence of a mismatch. Probes annealed to mutant targets are digested as long as they are sufficiently complementary to the mutant target or until the reaction is stopped (e.g., by the addition of proteinase). The mutant target is released (e.g., in the supernatant), while the wild-type molecule remains annealed to the partially digested probe (attached to a solid support (e.g., beads)). Thus, the mutant target can be easily separated from the wild-type target, allowing for enrichment of the mutant target. A representative schematic of this process is shown in Figure 1. [Example]

[0127] Restriction enzyme activation followed by 3'→5' exonuclease digestion The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. The first step involves probe hybridization to the target sequence. The probe sequence is blocked from digestion at the 3' end by oligonucleotide modifications or the presence of mismatches (or other methods described above). The probe sequence perfectly matches the mutant target sequence and has at least one mismatch with the wild-type target sequence. In some embodiments, the probe is modified at the 5' end to allow attachment to a surface. In the next step, due to enzyme (e.g., restriction enzyme) selectivity, probes perfectly annealed to the mutant target are nicked, while probes annealed to the wild-type molecule are unable to be nicked. The 3' end block of the probe is removed from the mutant target, allowing for 3'→5' digestion of the probe and subsequent release (e.g., into the supernatant). A representative schematic of this process is shown in Figure 2. [Example]

[0128] Mismatch-specific endonuclease activation followed by 3'→5' exonuclease digestion The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. The first step involves probe hybridization to the target sequence. The probe sequence is blocked from digestion at its 3' end by oligonucleotide modifications or the presence of mismatches (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe has a modification at its 5' end, allowing it to be attached to a solid surface. In the next step, mismatch-specific endonuclease selectivity allows the probe to be nicked at the mismatch site, whereas probes annealed to wild-type molecules cannot be nicked. For mutant targets, the 3' end block of the probe is removed, allowing for 3'→5' digestion of the probe and subsequent release (e.g., into the supernatant). A representative schematic of this process is shown in Figure 3. [Example]

[0129] Mismatch-specific endonuclease activation followed by 5'→3' exonuclease digestion The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. The first step involves probe hybridization to the target sequence. The probe sequence is blocked from digestion at the 5' end by oligonucleotide modifications or the presence of mismatches (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe is modified at the 3' end to allow attachment to a solid surface. In the next step, mismatch-specific endonuclease selectivity allows the probe to be nicked at the mismatch site, whereas probes annealed to wild-type molecules cannot be nicked. For mutant targets, the 5' end block of the probe is removed, allowing for 5'→3' digestion of the probe and subsequent release (e.g., into the supernatant). A representative schematic of this process is shown in Figure 4. [Example]

[0130] Mismatch-specific endonuclease activation and subsequent displacement The extracted DNA undergoes standard library preparation: end repair, A-tailing, and adapter ligation. The first step involves probe hybridization to the target sequence. The probe sequence is blocked from digestion at the 5' end by oligonucleotide modifications or the presence of mismatches (or other methods described above). The probe sequence perfectly matches the wild-type target sequence and has at least one mismatch with the mutant target sequence. In some embodiments, the probe has a modification at the 3' end that allows it to be attached to a surface. In the next step, enzyme selectivity nicks the probe at the mismatch site, while probes annealed to wild-type molecules cannot be nicked. The probe is nicked into two oligonucleotides: a 5'-blocked oligonucleotide and a 3'-bead-attached oligonucleotide. The 5'-blocked oligonucleotide serves as a primer for extension using a strand-displacing DNA polymerase. The mutant target sequence is released into the supernatant by displacement of the 3'-bead-attached oligonucleotide. A representative schematic of this process is shown in Figure 5.

Claims

1. 1. A method comprising enriching or depleting a first nucleic acid molecule in a sample comprising a mixture of nucleic acid molecules, wherein the enriching or depleting comprises: contacting the sample with a probe whose complementarity to a target region of a first nucleic acid molecule differs from its complementarity to a second nucleic acid molecule in the sample; optionally activating the probe hybridized to the first nucleic acid molecule by selectively modifying the probe hybridized to the first nucleic acid molecule relative to the probe hybridized to the second nucleic acid molecule; selectively digesting the probe hybridized to the first nucleic acid molecule or the second nucleic acid molecule relative to other probes; and enriching or depleting the first nucleic acid molecule. The method is based on the above.

2. 10. The method of claim 1, wherein the first and second nucleic acid molecules comprise end-repaired nucleic acid molecules.

3. 3. The method of claim 1 or 2, wherein the first and second nucleic acid molecules are A-tailed nucleic acid molecules.

4. The method of claim 1 , wherein the first and second nucleic acid molecules comprise a tag or adapter sequence.

5. The method of claim 1 , wherein the first and second nucleic acid molecules are amplified.

6. The method of any one of claims 1 to 5, wherein the probe is activated by contact with a cleaving agent.

7. 7. The method of claim 6, wherein the cleavage agent is selected from the group consisting of restriction endonucleases, flap endonucleases, mismatch repair enzymes, RNases, Cas proteins, Argonaute family enzymes, DNA-formamidopyrimidine glycosylases, apurinic / apyrimidinic (AP) endonucleases, and chemical cleavage agents.

8. 8. The method of claim 1, wherein the digesting comprises contacting the probe with an exonuclease or endonuclease.

9. 9. The method of claim 8, wherein the exonuclease is a 3' to 5' exonuclease.

10. The method of claim 8, wherein the exonuclease is a 5' to 3' exonuclease.

11. 11. The method of claim 1, wherein the probe comprises a binding moiety at its 3' or 5' end, the binding moiety optionally being biotin or a sequence to which a linker molecule can hybridize, and the linker molecule is modified to bind to a solid support.

12. The method of claim 1 , further comprising capturing the probe on a surface before or after the digesting.

13. The method of claim 12 , wherein the surface comprises beads.

14. 13. The method of claim 12, further comprising differentially releasing the first nucleic acid molecule or the second nucleic acid molecule from the probe.

15. 15. The method of claim 14, wherein the liberating comprises increasing the temperature.

16. 15. The method of claim 14, wherein the liberating comprises changing the pH.

17. 15. The method of claim 14, wherein the releasing comprises changing the salt concentration.

18. 15. The method of claim 14, wherein said liberating comprises said digesting.

19. 19. The method of any preceding claim, further comprising detecting the first or second nucleic acid molecule.

20. 20. The method of claim 19, wherein said detecting comprises sequencing said first or second nucleic acid molecule.

21. 21. The method of any of claims 1 to 20, wherein the probe comprises a base that is complementary to a position in the first nucleic acid molecule and mismatches with a corresponding position in the second nucleic acid molecule.

22. 22. The method of claim 21, wherein said digesting comprises contacting probes hybridized to said first and second nucleic acid molecules with an enzyme that preferentially digests complementary strands over non-complementary strands.

23. 22. The method of claim 21, wherein said digesting comprises contacting probes hybridized to said first and second nucleic acid molecules with an enzyme that preferentially digests non-complementary strands over complementary strands.

24. 24. The method of any preceding claim, wherein the probe comprises a 3'-terminal, 5'-terminal or internal blocking group.

25. 25. The method of claim 24, wherein activation of the probe is performed using a cleaving agent that removes the blocking group from the probe hybridized to the first nucleic acid molecule but not from the probe hybridized to the second nucleic acid molecule.

26. 26. The method of claim 25, wherein said digesting comprises contacting the probe with a nuclease that cleaves a probe lacking a blocking group but does not cleave a probe with a blocking group.

27. 25. The method of claim 24, wherein activation of the probe is performed using a cleaving agent that removes the nucleic acid fragment containing the blocking group from the probe hybridized to the first nucleic acid molecule but not from the probe hybridized to the second nucleic acid molecule.

28. 28. The method of claim 27, wherein said digesting comprises contacting said nucleic acid fragment with a polymerase under conditions such that the fragment is extended and the probe hybridized to said first nucleic acid molecule is displaced or digested using a polymerase having 5' to 3' exonuclease activity, optionally with an upstream primer.

29. 29. The method of any one of claims 1 to 28, wherein said activating and said digesting do not involve pyrophosphorolysis.

30. 30. The method of any of claims 1 to 29, further comprising contacting the sample with a capture probe, wherein the capture probe hybridizes to the second nucleic acid molecule or to another nucleic acid molecule in the sample that is not the first nucleic acid molecule or the second nucleic acid molecule.

31. 31. The method of any one of claims 1 to 30, wherein the first and / or second nucleic acid molecule is a single-stranded nucleic acid molecule.

32. 32. The method of any preceding claim, wherein the first and / or second nucleic acid comprises one or more methylated nucleotides.

33. 33. A kit comprising sufficient reagents to perform the method of any of claims 1 to 32 on a sample containing said first and second nucleic acid molecules.