Linked target capture

The method of ligated target capture with universal and target-specific probes addresses the need for evaluating genome editing tools, offering precise integration rate assessment and non-specific detection, thereby improving the efficacy of CRISPR, TALENs, and ZFNs.

JP2025078774APending Publication Date: 2025-05-20NCAN GENOMICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025034743
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-10
Filing Date
2025-03-05
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Existing genome editing tools lack efficient methods for evaluating the integration rate and specificity of inserted sequences, necessitating improved techniques for assessing the performance of systems like CRISPR, TALENs, and ZFNs.

Method used

A method utilizing ligated target capture technology with universal and target-specific probes, combined with droplet-based methods, to enrich and quantify double-strand tag sequences, providing comprehensive assessment of on-target and off-target integration rates.

Benefits of technology

Enables precise evaluation of genome editing systems by measuring integration rates and identifying non-specific incorporation, enhancing the understanding and efficiency of genome editing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078774000005
    Figure 2025078774000005
  • Figure 2025078774000006
    Figure 2025078774000006
  • Figure 2025078774000007
    Figure 2025078774000007
Patent Text Reader

Abstract

To provide assessment of their efficiency and specificity including assessment of integration rates for inserted sequences.SOLUTION: The invention generally relates to using linked target capture probes to evaluate genome editing efficiency and specificity. Because multiple binding steps are required, specificity is improved over traditional single binding target capture techniques.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 859,486, filed June 10, 2019, the contents of which are incorporated herein by reference in their entirety.

[0002] The present invention relates generally to the capture, amplification, and sequencing of nucleic acids. [Background technology]

[0003] The advent of more powerful and user-friendly genome editing tools has opened a new world of possibilities for treating genetic disorders, eradicating disease, improving yields / resistance, and other potential benefits of modifying organisms. Systems including clustered regularly interspaced short palindromic repeats (CRISPR) and related enzymes, meganucleases, transcription activator effector-like nucleases (TALENs), and zinc finger nucleases allow for the introduction of double-stranded breaks in DNA at specific target sequences, which can allow for targeted mutations, including the insertion of desired sequences at the breakpoint.

[0004] To prove the validity of these tools and to promote their acceptance for general use, their efficiency and specificity need to be evaluated, including evaluation of the integration rate of the inserted sequences. Analysis of nonspecific cleavage and insertion will also be important. Summary of the Invention

[0005] The present invention provides a method for evaluating the integration rate and non-specific effect of any of the aforementioned genome editing tools.The inserted double-stranded tag sequence can be enriched and quantified to evaluate the success rate.The combination of monitoring non-specific (off-target) integration and quantifying on-target integration provides a powerful tool for evaluating genome editing systems.

[0006] In certain embodiments, the present invention provides a method of ligated target capture technology using probes that target double-stranded tags inserted using various genome editing tools. Target capture to detect double-stranded breaks can be performed in solution or using droplet-based methods. Ligated target capture probes are used that include a universal primer and a target-specific probe, and the reaction occurs under conditions that require the target-specific probe to bind to allow the binding of the universal primer. After the tag sequence is incorporated using the genome editing method to be analyzed, a duplex adapter with a universal priming site can be ligated to the end of the modified DNA. The target-specific probe can be complementary to the tag sequence, the genomic DNA sequence adjacent to the double-stranded break, or both. This heterogeneously-integrated DNA eNrichment, or HIDN-Seq process described herein, allows for enrichment of tag sequences or tags and adjacent sequences, providing data on the rate of incorporation as well as identifying non-specific incorporation and providing a comprehensive assessment of DNA editing performance. Enrichment of tag sequences allows measurement of all integration sites, including undesired non-specific sites, using probes designed only against tag sequences, whereas enrichment of desired integration sites allows measurement of integration rates at specific sites, using probes designed against expected genomic DNA integration sites.

[0007] The need for multiple binding steps improves specificity over conventional single binding target capture techniques. After binding of the ligated probes, the bound universal primer is extended using a strand displacing polymerase to generate a copy of the target strand. This can be amplified using PCR with the universal primer. The ligated capture probes can be used for both senses of DNA where higher specificity and dual information is required. Multiple linker types are possible, as described below. As with the solution-based target capture method of the present invention, we provide a droplet-based method that allows users to perform target capture for DNA integration analysis in droplets, rather than being limited to multiplex PCR in droplets.

[0008] Barcodes containing duplex unique molecular identifiers (UMIs) can be used to tag the amplified or enriched sequences, retaining sense information along with the starting molecule information of the double-stranded DNA being analyzed. Thus, sequencing results can be attributed to individual starting molecules for assessment of precise incorporation rates. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 shows an exemplary method for ligated target capture of double-stranded nucleic acids. [Diagram 2] FIG. 2 illustrates a method for amplifying linked target capture nucleic acids. [Diagram 3] 3A and 3B show steps of the droplet-based target capture method of the present invention. [Figure 4] FIG. 4 shows the incorporation of an exemplary tag sequence and the induced double-stranded DNA break. [Diagram 5] FIG. 5 shows an exemplary off-target discovery workflow using HIDN-Seq and ligated target capture probes specific for tag sequences. [Figure 6] FIG. 6 shows an exemplary off-target and adjacent discovery workflow using HIDN-Seq and ligated target capture probes specific for tag sequences and genomic DNA regions adjacent to the breakpoint. [Figure 7] FIG. 7 shows an exemplary combined workflow using HIDN-Seq and concatenated target capture probe sets specific for the tag sequence and genomic DNA regions adjacent to the breakpoint. [Figure 8] FIG. 8 shows an exemplary combined workflow performed in a single tube using HIDN-Seq and concatenated target capture probe sets specific for the tag sequence and genomic DNA regions adjacent to the breakpoint. [Figure 9] FIG. 9 shows an exemplary workflow using barcoding PCR and HIDN-Seq with quantification and sequencing. [Figure 10] FIG. 10 shows the experimental outline of Example 1. [Figure 11] FIG. 11 shows the number and percentage of SI, S2, and S3 clusters that contain the intended tag sequence in zero, one, or both reads in Example 1. [Figure 12] FIG. 12 shows the UMI coverage of the whole genome plotted against the number of bases in the genome, and the minimum UMI coverage for the SI, S2, and S3 groups of Example 1. [Figure 13] FIG. 13 shows the on-target fractions of Example 2 as determined by HIDN-Seq for spiked samples. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] The present invention relates generally to a method for targeted capture and analysis of double-strand breaks in DNA, particularly for analysis of the efficiency and specificity of genome editing systems. A linked target capture technique is used, in which a linked target capture probe is used that includes a universal primer and a target-specific probe, and a reaction occurs under conditions that require the target-specific probe to bind to allow binding of the universal primer. A universal priming site can be ligated to the end of a post-edited (e.g., cleavage and sequence insertion) fragment of genomic DNA. The target-specific portion of the linked target capture probe can then be designed to be specific to the target breakpoint of the DNA, the inserted tag sequence, or a combination of the two. By enriching the tag sequence alone or with the target site, information can be obtained about the integration rate and non-specific integration. That information is essential for evaluating existing and future technologies in the burgeoning field of genome editing. Linked target capture, as well as related amplification and sequencing techniques using linking molecules, are contemplated herein as described in U.S. Patent Publication No. 20190106729, which is incorporated herein by reference. The tag sequence may be specifically designed for evaluation or may be a functional sequence intended for use in genome modification. Target-specific probes targeting the tag sequence can be designed to bind to any sequence (evaluation specific tag or genomic DNA insert) to evaluate the general performance of the genome editing technology or to evaluate the performance of a specific modification using a specific insert.

[0011] The systems and methods described herein can be used to analyze any such technology, including those that rely on CRISPR-associated (Cas) endonucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or RNA-guided engineered nucleases (RGENs). Programmable nucleases and their uses are described, for example, in Zhang F, Wen Y, Guo X (2014). "CRISPR / Cas9 for genome editing: progress, implications and challenges".

[0012] Human Molecular Genetics. 23 (Rl): R40-6. doi:10.1093 / hmg / ddul25; Ledford H (March 2016). “CRISPR: gene editing is just the beginning”. Nature. 531 (7593): 156-9.

[0013] doi: 10.1038 / 531156a; Hsu PD, Lander ES, Zhang F (June 2014). “Development and applications of CRISPR-Cas9 for genome engineering”. Cell. 157 (6): 1262-78.

[0014] doi:10.1016 / j.cell.2014.05.010; Boch J (February 2011). “TALEs of genome targeting”. Nature Biotechnology. 29 (2): 135-6. doi:10.1038 / nbt.l767; Wood AJ, Lo TW, Zeitler B, Pickle CS, Ralston EJ, Lee AH, Amora R, Miller JC, Leung E, Meng ”Genome engineering with zinc-finger Genetics Society of America. 188 (4): 773-782.

[0015] doi:10.1534 / genetics.111.131433; Urnov, FD, Rebar, EJ, Holmes, MC, Zhang, HS, & Gregory, PD (2010). "Genome Editing with Engineered Zinc Finger Nucleases". Nature Reviews Genetics. 11 (9): 636-646. doi:10.1038 / nrg2842, the contents of each of which are incorporated herein by reference.

[0016] Existing techniques for identifying double-strand breaks and evaluating genome editing tools are described in U.S. Patent Nos. 9,822,407 and 9,850,484, which are incorporated herein by reference. And the back-end sequencing and analysis techniques described herein may be used with the linked target capture methods described herein for analysis of double-strand breaks and insertion efficiency.

[0017] An exemplary double-strand break and tag insertion is shown in FIG. 4. Any of the methods discussed (e.g., CRISPR-Cas RNA-guided nuclease (RGN), TALEN (transcription activator-like effector nuclease), and ZFN (zinc finger nuclease) can be used to introduce a double-strand break. After cleavage, the designed tag sequence can be incorporated as shown in FIG. 4. Tag incorporation can be achieved by methods such as those described in Tsai SQ, Zheng Z., Nguyen NT, Fiebers M., Topkar VV, et al. (2015) GUIDE-seq enables genome wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33: 187-197, which is incorporated herein by reference. Tag modifications (such as 5' phosphates at both ends and two phosphorothioate bonds) can be used to increase the rate of tag incorporation.

[0018] Assuming incomplete cleavage and integration, some target fragments will not be cleaved or will successfully integrate the tag sequence, some will successfully integrate the tag sequence, and some will integrate the tag sequence at an off-target site. The proportion of these outcomes can then be determined using the linked target capture (FTC) technique described herein.

[0019] In certain embodiments, ligated target capture probes are used that have a tag-specific probe ligated to a universal primer, as shown in FIG. 5. Prior to probe binding and amplification, adapters containing universal priming sites are ligated to the sample fragments, thereby providing a target site for the universal primer. The ligated adapters may contain unique molecular identifiers (UMIs) or other barcode sequences, which can later be used to discriminate the original molecule from which the sequence was ultimately derived. Such information can be used to determine consensus sequences for individual molecules, providing more accurate quantification of cleavage and on-target and off-target incorporation rates. Barcodes can be included in the stem portion of the y adapter or in the non-complementary portion of the y adapter to retain sense-specific tag information. Similarly, the universal priming site can be located in the stem or y portion of the adapter. In certain embodiments, the stem location is preferred to place the target site of the ligated target capture probe closer together for improved function. In such embodiments, despite the loss of sense-specific tag information, the error reduction benefits are still achieved, as discussed in U.S. patent application Ser. No. 16 / 239,100, which is incorporated herein by reference.

[0020] In the case of off-target detection by tag enrichment, the target-specific probe preferentially binds to the inserted sequence. Using the ligated target capture technique described below, amplification occurs only if both ligated probes bind relatively close to each other along the fragment. Because the ligated probe can contain another universal PCR priming sequence (different from the site of the ligated adapter), after several cycles of amplification with the ligated probe, the sample is indexed and a more robust amplification using conventional universal PCR primers can be used to create a sequencing library. The tag-specific ligated probe should capture and amplify any tag sequence along with the adjacent genomic DNA sequence between the tag and the ligated universal priming site. Thus, through sequencing and subsequent analysis, the comparative number of tags integrated at the correct site, tags integrated off-target, and tags not integrated can be assessed, thereby providing an assessment of the specificity and efficiency of the cleavage and integration techniques used in promising genome editing tools.

[0021] Off-target discovery can be combined with enrichment of flanking sequences, as shown in Figure 6. A ligated target capture and amplification technique is performed similarly to that shown in Figure 5, but with different probe-dependent primers (PDPs). Both PDPs contain a universal primer complementary to the ligated adapter sequence, but the target-specific probes preferentially bind to different targets. One target-specific probe binds to a tag-specific sequence, while the other binds to a portion of the genomic DNA adjacent to the integration site of interest. The resulting sequence capture should exclude unintegrated tags and non-specific integrations, and capture only correctly integrated targets. In certain embodiments, mismatches with the target-specific probes can be tolerated, thereby capturing integration errors that may be off-target by a few nucleotides or otherwise cause unintended changes at the breakpoint.

[0022] PDPs containing target-specific probes that target both sides of the genomic DNA adjacent to the breakpoint can also be used. This allows the capture of all genomic fragments containing the breakpoint of interest. The captured molecules should contain genomic DNA with successfully integrated tag sequences, as well as genomic DNA that was not cut or was not integrated and repaired. Thus, double-strand breaks and sequence integration efficiencies can be evaluated for the genome editing tool being tested.

[0023] As shown in Figure 7, these methods can be combined, where adapters are ligated to the fragments and probes are used that target both ends of the tag insert and the genomic DNA sequence flanking the breakpoint. Pre-amplification can be used in such an assay to measure integration rates. Such an assay provides off-target and on-target integration rates simultaneously, providing a complete genome editing performance assessment in a single assay. As shown in Figure 8, the combined assay can be performed in a single tube to reduce workflow complexity.

[0024] An exemplary method using back-end analysis is shown in FIG. 9. After target cleavage and incorporation of the tag sequence as shown in FIG. 4, an adaptor is ligated to the end of the tag-inserted genomic DNA. The adaptor contains a priming site and an optional barcode. Pre-amplification using a primer specific for the adaptor can optionally be used. Target capture is performed using ligated target capture probes as described with reference to FIGS. 5-8. Barcoding PCR is used, followed by DNA quantification and sequencing. Sequence analysis can optionally be used to then determine the consensus sequence of each uniquely-identified molecule. The raw sequencing data or collapsed reads can then be analyzed to determine the relative amount of unmodified genomic DNA at the target cleavage point, unincorporated tag sequence, on-target integration, and / or off-target integration, depending on the ligated target capture probe used. Any sequencing technology can be used, as well as known sequence analysis / comparison techniques or software.

[0025] Ligated target capture methods may involve solution-based capture of genomic regions of interest for targeted DNA sequencing. Figures 1 and 2 show an exemplary method of solution-based target capture. A universal priming site and any barcode (which may be sense-specific) are ligated to the extracted DNA. The ligated DNA product is then denatured and bound to a ligated target capture probe that includes a universal primer ligated to a target-specific probe. Target capture is performed at a temperature where the universal primer cannot bind alone unless it is at a high local concentration due to the binding of the target probe. A strand-displacing polymerase (e.g., Taq, BST, phi29, SD) is then used to extend the ligated probe bound to the target. As indicated by the black diamonds in Figures 1 and 2, the target probe is blocked from extension, so that extension occurs only along the bound universal primer, copying the bound target nucleic acid strand that remains ligated to the target primer. Several ligated PCR extension cycles can then be used to amplify the target sequence. PCR can then be performed using a universal primer that corresponds to the universal priming site from the ligated target capture probe to amplify one or both strands of the target nucleic acid. This PCR step can be performed in the same reaction without the need for a purification step. The amplified target sequence can then be sequenced as described above. Gaps are permitted, but no gaps are required between the ligated capture probes when used in opposite orientations. Capture probes can be created using a universal 5'-linker by attaching a universal primer to a pre-made capture probe. The capture probes can be attached by click chemistry or other means as described below.

[0026] In some embodiments, nucleic acid can be fragmented or broken into smaller nucleic acid fragments. The shorter fragments achieved before adapter ligation can help shorten the distance that the ligated probe needs to span, thereby increasing binding and enrichment efficiency. Nucleic acid, including genomic nucleic acid, can be fragmented using any of a variety of methods, such as mechanical fragmentation, chemical fragmentation, and enzymatic fragmentation. Methods of nucleic acid fragmentation are known in the art and include, but are not limited to, DNase digestion, sonication, mechanical shearing, and the like (J. Sambrook et al, "Molecular Cloning: A Laboratory Manual", 1989, 2nd Ed., Cold Spring Harbour Laboratory Press: New York, NY; P. Tijssen, "Hybridization with Nucleic Acid Probes- Laboratory Techniques in Biochemistry and Molecular Biology (Parts I and II)", 1993, Elsevier; CP Ordahl et al, Nucleic Acids Res., 1976, 3: 2985-2999; PJ Oefner et al, Nucleic Acids Res., 1996, 24: 3879-3889; YR Thorstenson et al, Genome Res., 1998, 8: 848-855). US Patent Publication No. 2005 / 0112590 provides a general overview of various methods of fragmentation known in the art.

[0027] The probe-dependent primers used in the target capture techniques described herein can link the 5' end of a target-specific DNA probe (e.g., complementary to a portion of the tag insertion sequence or an adjacent portion of the genomic DNA sequence at the breakpoint) to the 5' end of a universal primer. The DNA probe may contain an inverted dT, a C3 spacer, or other blocking moiety at its 3' end to prevent extension of the DNA probe, which supports extension of the subsequently bound universal primer, brought into proximity with the target nucleic acid fragment by the DNA probe that binds to a complementary target sequence within the fragment. Primers and probes can be synthesized separately and linked using the techniques described below.

[0028] Although target-specific sequences are preferred for the linked target capture probes, in certain embodiments, the 5' end of the universal primer (with any barcode, as described below) can be attached to the 5' end of a probe molecule, which may consist of any protein, nucleic acid, or other molecule that exhibits binding affinity to a specific target sequence or target feature in a nucleic acid. The probe molecule may be a DNA or RNA binding probe, and can be synthesized or isolated separately from the primer (e.g., the universal primer) before being linked together using, for example, click chemistry, biotin / streptavidin binding, or derivatives such as dual biotin and streptavidin, PEG, immuno-PCR chemistry, such as gold nanoparticles, chemical cross-linking or fusion proteins, or direct binding of proteins / antibodies to the DNA primer sequence. Linking methods are described in more detail below.

[0029] Exemplary DNA or RNA binding probes may include DNA or RNA probes for targeting specific DNA or RNA sequences. Zinc finger domains, TAL effectors, or other sequence-specific binding proteins can be modified and linked to universal adaptors or primers to create probe-dependent primers or adaptors that target specific DNA or RNA sequences. Methyl-CpG binding domains (MBDs) or antibodies (used in methylated DNA immunoprecipitation) can be linked to adaptors or primers that target methylated sequences. For use in the present system and method, the target-specific probes only need to preferentially bind to the desired portion of the incorporated tag or the breakpoint adjacent to the genomic DNA sequence. In certain embodiments, the tag may include a feature (e.g., a methylated sequence) that can be targeted using a specific probe.

[0030] Probe-dependent primers can be made by linking a universal primer and a target-specific probe with a link modification. Probes can be directly synthesized with the link modification. If this is not possible, such as for array-synthesized probes, the linker modification can be added by PCR. Probes can be synthesized in arrays on silicon chips and then amplified, rather than made in bulk by column-based synthesis. Array-based probes containing target sequencing and universal priming sites can be amplified by universal primers containing link modifications. Array-based oligos can be converted to linked target capture probes by adding a 5' linker modification, for example, by post-synthesis PCR. A 3' blocker can be substituted for the end of the frayed primer. After amplification, the modified probes can be ligated to universal primers and used as probe-dependent primers.

[0031] In certain embodiments, the linking molecule may be a streptavidin molecule, and the linked fragments may comprise biotinylated nucleic acids. In embodiments where linked primers are used to create linked nucleic acid fragments by amplification, the primers may be biotinylated and linked together on a streptavidin molecule. For example, four fragments may be linked with a tetrameric streptavidin. For example, four or more molecules may be linked by the formation of concatemers. In certain methods of the invention, two or more nucleic acid fragments may be linked via click chemistry. See Kolb, et al, Click Chemistry: Diverse Chemical Function from a Few Good Reactions, Angew Chem Int Ed Engl. 2001 Jun 1;40(11):2004-2021, incorporated herein by reference.

[0032] For example, binding molecules, and some known nanoparticles, can link multiple fragments and / or DNA-binding proteins, including hundreds or thousands of fragments, within a single binding molecule. One example of a binding nanoparticle may be a multivalent DNA-gold nanoparticle, which comprises colloidal gold modified with synthetic DNA sequences capped with thiols on their surface. See Mirkin, et al, 1996, A DNA-based method for rationally assembling nanoparticles into macroscopic materials, Nature, 382:607-609, incorporated herein by reference. The surface DNA sequence may be complementary to the desired template molecule sequence or may include a universal primer.

[0033] Linking molecules can also be useful for separating nucleic acid fragments. In a preferred embodiment, the fragments are oriented to prevent binding between them. The linker controls the spatial separation and orientation of the fragments, thereby avoiding and preventing folding and binding between the fragments.

[0034] In some embodiments, the linker may be polyethylene glycol (PEG) or modified PEG, for example, modified PEG such as DBCO-PEG4 or PEG-11 can be used to link two adaptors or nucleic acids. In another example, N-hydroxysuccinimide (NHS) modified PEG is used to link two adaptors. See, for example, Schlingman, et ak, Colloids and Surfaces B: Biointerfaces 83 (2011) 91-95. Any oligonucleotide or other molecule can be used to link adaptors or nucleic acids.

[0035] In some embodiments, an aptamer is used to bind the two probes. Aptamers can be designed to bind to various molecular targets, such as primers, proteins, or nucleic acids. Aptamers can be designed or selected by the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) method. An aptamer is a nucleic acid polymer that specifically binds to a target molecule. Like all nucleic acids, a particular nucleic acid ligand, or aptamer, can be described by a linear sequence of nucleotides (A, U, T, C, and G), usually 15-40 nucleotides long. In some preferred embodiments, the aptamer may include an inverted base or a modified base. In some embodiments, the aptamer or modified aptamer includes at least one inverted or modified base.

[0036] It is understood that the linker may be composed of an inverted base or may include at least one inverted base. Inverted or modified bases can be obtained through any commercial entity. Inverted or modified bases have been developed and are commercially available. Inverted or modified bases can be incorporated into other molecules. For example, 2-aminopurine can be substituted in oligonucleotides. 2-Aminopurine is a fluorescent base useful as a probe for monitoring DNA structure and dynamics. 2,6-diaminopurine (2-amino-dA) is a modified base that, when base-paired with dT, can form three hydrogen bonds and increase the Tm of short oligos. 5-bromodeoxyuridine is a photoreactive halogenated base that can be incorporated into oligonucleotides and crosslinked to DNA, RNA, or proteins upon exposure to UV light. Other examples of inverted or modified bases include deoxyuridine (dU), inverted dT, dideoxycytidine (ddC), 5-methyldeoxycytidine, or 2'-deoxyinosine (dI). It should be understood that any reversed or modified base can be used in ligating the template nucleic acid.

[0037] In a preferred embodiment, the linker comprises a molecule for linking two primers or two nucleic acid fragments. The linker may be a single molecule or multiple molecules. The linker may comprise some reversed or modified bases, or completely reversed or modified bases. The linker may comprise both Watson-Crick bases and reversed or modified bases.

[0038] It should be understood that any spacer or linking molecule can be used in the present invention. In some embodiments, the linker or spacer molecule can be a lipid or an oligosaccharide, or an oligosaccharide and a lipid. See US Patent No. 5,122,450. In this example, the molecule is preferably a lipid molecule, more preferably a glyceride or phosphatide having at least two hydrophobic polyalkylene chains.

[0039] The linker may be composed of any number of adapters, primers, and copies of the fragment. The linker may include two identical arms, where each arm is composed of a binding molecule, an amplification primer, a sequencing primer, an adapter, and a fragment. The linker may link any number of arms, such as three or four arms. It should be understood that in some aspects of the invention, the nucleic acid template is linked by a spacer molecule. The linker in the present invention may be any molecule or method for linking two fragments or primers. In some embodiments, polyethylene glycol or modified PEG such as DBCO-PEG4 or PEG-11 is used. In some embodiments, the linker is a lipid or a carbohydrate. In some embodiments, a protein can be attached to the adapter or the nucleic acid. In some embodiments, an oligosaccharide links the primer or the nucleic acid. In some embodiments, an aptamer links the primer or the nucleic acid. When the fragments are linked, the copies are oriented to be in phase, preventing binding between them.

[0040] In certain embodiments, the linker may be an antibody. The antibody may be a monomer, dimer, or pentamer. It is understood that any antibody for binding two primers or nucleic acids may be used. For example, it is known in the art that nucleosides can be made immunogenic by binding to proteins. See Void, BS (1979), Nucl Acids Res 7, 193-204. Furthermore, antibodies can be prepared to bind modified nucleic acids. See Biochemical Education, Vol. 12, Issue 3.

[0041] The linker may remain attached to the complex during amplification. In some embodiments, the linker is removed prior to amplification. In some embodiments, the linker is attached to a binding molecule, which is then attached to an amplification primer. When the linker is removed, the binding molecule or the binding primer is exposed. The exposed binding molecule also attaches to the solid support, forming an arch. The linker can be removed by any method known in the art, including washing with a solvent, applying heat, changing the pH, washing with a detergent or surfactant, etc.

[0042] The method of the present invention includes a droplet-based target capture, optionally using a universal link primer, to capture the double-stranded molecules. A droplet-based method using linked target capture probes, as described in US Patent Publication No. 20190106729, and as described therein and shown in Figures 1 and 2. A universal primer and any barcode (which may be sense-specific) are ligated to extracted DNA (e.g., cell-free DNA). An emulsion is created as described above using a duplex template molecule and a target capture probe that includes a universal primer ligated to a target-specific probe. As described above, target capture is performed at a temperature where the universal primer cannot bind alone unless the local concentration is high and the capture probe is prevented from extending itself by the binding of the target probe, but includes a universal priming site such that the universal primer and linked universal primer contained in the emulsion can be used to amplify the target nucleic acid and generate a linked double-stranded molecule that includes both the sense and antisense strands of the target nucleic acid. The universal linker can be omitted to perform target capture only. The emulsion is then broken and the unligated templates are enzymatically digested, leaving only ligated double-stranded molecules to seed clusters or be sequenced as above.

[0043] Figures 3A and 3B provide additional details of the droplet-based target capture method of the present invention. Step 0 in Figure 3A shows a duplex template molecule with a universal priming site and any barcodes linked to it are loaded into a droplet with a linked universal primer and target capture probe. The template DNA is denatured within the droplet, and the target capture probe binds to the denatured template strand at a temperature where the universal primer will not bind alone unless the target probe is also bound. The universal primer then binds only to the captured target. Extension by the strand-displacing polymerase occurs only on the captured target. Next, moving to Figure 3B, extension cycles are performed (e.g., 4-6 cycles) until the desired target capture probe and primer are exhausted. The resulting extension products are amplified using a universal link primer to generate a linked duplex molecule with a strand-specific barcode. As with the solution-based method, no gap is required between the linked capture probes if they are in opposite orientations. The linked capture probes can be used in one or both orientations if the universal linker is omitted and only target capture is performed. Conventional polymerases can be mixed with strand-displacing polymerases within the droplets to carry out the various extension and amplification steps of the method. EXAMPLES

[0044] Example 1: HIDN-Seq of Cas9 cell lines Using the Cas9 cell line, we added an insert to one group of cells as a control, adding the insert along with a guide RNA targeting the desired insertion breakpoint. HIDN-Seq as described above was then performed on the DNA of both cell groups. For the experimental (gRNA + insert) group, we waited for ligated target capture as described above, without PCR amplification after adapter ligation (i.e., directly from ligation to ligated target capture amplification, as shown in Figure 9). An overview of the experiment is shown in Figure 10. Here, SI represents HIDN-Seq performed on control DNA (no gRNA added), S2 represents HIDN-Seq using ligated target capture as described above, and S3 represents HIDN-Seq with ligated target capture as described, but without PCR amplification after adapter ligation. For all three samples, approximately 1 million clusters were sequenced. The results are shown in Figures 11 and 12.

[0045] Figure 11 shows the number and percentage of S1, S2, and S3 clusters that contain tag sequences in zero, one, or both reads. As shown, over 99% of the clusters for each sample contain at least one read with the expected tag sequence (within an edit distance of 4), meaning that there were essentially no wasted reads.

[0046] Figure 12 shows the UMI coverage across the genome, plotted against the number of bases in the genome and the minimum UMI coverage for the S1, S2, and S3 groups. The SI (tags only) group had much lower coverage, with a maximum coverage of less than 20. This indicates that only a few cut sites occurred at low integration rates, as would be expected in the absence of gRNA. However, the S2 and S3 groups show much higher coverage in certain regions, suggesting significant integration at multiple sites.

[0047] The sequencing results of off-target sites for the S2 and S3 groups are shown in the table below. The SI (tag only) group did not match the gRNA, but the S2 and S3 groups had gRNA sequences found in each of the top 50 coverage regions. The top 20 for each are shown in the table. Target sequences are underlined.

[0048] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]

[0049] The double-stranded tag sequences used in the experiments were as follows: BG Tag vl sequence (SEQ ID NO: 22): / 5Phos / C*A*GTGTTTAATTGAGTTGTCATATGTTAATAACGGTATCA*G*C BG Tag vl sequence (reverse complement, SEQ ID NO: 23): / 5Phos / G*C*TGATACCGTTATTAACATATGACAACTCAATTAAACAC*T*G

[0050] The forward probe (Tm=69.1° C.) sequence (SEQ ID NO: 24) was as follows: CA+GT+GTTTA+ATTGAGTTGTCATATGTTAATAACGG

[0051] The reverse probe (Tm=69.3° C.) sequence (SEQ ID NO: 25) was as follows: G+CT+GATACCGTTATTAACATATGACAACTCA

[0052] The tag sequences were selected to have a melting temperature high enough to allow binding of forward and reverse ligated target capture probes. The probe sequences were selected to have high specificity for the tag sequence, but low overlap temperature (e.g., below 60°C). Locked nucleic acids (LNA's, indicated with a "+" before the LNA base) were used to achieve the desired probe melting temperature.

[0053] Example 2: Tag Enrichment Genomic DNA containing the tag sequence was spiked into genomic DNA in various amounts and samples were subjected to HIDN-Seq using forward and reverse probes with tag-specific probes (see Figure 5). As shown in Figure 13, the percentage of sequence reads containing the tag sequence is greater than 99.8% at both 1E5 and 1E6 tag spike levels. Because ligated target capture amplifies the entire insert from the ligated adapter, the flanking sequences of the tag were recovered.

[0054] Incorporation by Reference References and citations to other documents such as patents, patent applications, patent publications, journals, books, articles, web content, etc. have been made throughout this disclosure. All such documents are incorporated herein by reference in their entirety for all purposes.

[0055] Equivalent The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics, and the foregoing embodiments are therefore to be considered in all respects illustrative and not limiting of the invention described herein.

Claims

1. 1. A method for detecting double stranded DNA insertions, comprising: ligating a universal priming site onto a plurality of double-stranded nucleic acid fragments, one or more of which comprises a tag sequence inserted into an insertion site, wherein the tag sequence comprises a segment of double-stranded DNA; denaturing the plurality of linked double-stranded nucleic acid fragments to generate single-stranded nucleic acid fragments that comprise universal priming sites; exposing the single-stranded nucleic acid fragments to a plurality of linked capture probes including a target probe comprising a segment of single-stranded DNA complementary to at least a portion of a sequence adjacent to the one or more tag sequences and the 3' or 5' side of the insertion site, wherein the target probe hybridizes to a target nucleic acid sequence and a linked single-stranded DNA universal primer hybridizes to the universal priming site; extending the ligated single-stranded DNA universal primer to create a copy of the insertion site or the tag region; and Sequencing the copies to determine the presence of the tag sequence at the insertion site.

2. The method of claim 1, wherein the sequences adjacent to the 3' or 5' side of the insertion site do not span the insertion site.

3. 2. The method of claim 1, wherein the sequence proximal to the 3' or 5' side of the insertion site is within 150 nucleotides of the insertion site.

4. 2. The method of claim 1, wherein the plurality of linked capture probes comprises a target probe having affinity for at least a portion of the tag sequence and a target probe having affinity for at least a portion of the sequence adjacent to the 3' or 5' side of the insertion site.

5. 2. The method of claim 1, further comprising inserting the tag sequence into the insertion site using a genome editing tool.

6. 6. The method of claim 5, wherein the genome editing tool is selected from the group consisting of clustered regularly interspaced short palindromic repeats (CRISPR) and related enzymes, meganucleases, transcription activator effector-like nucleases (TALENs), and zinc finger nucleases.

7. 6. The method of claim 5, further comprising comparing an amount of sequence comprising the tag sequence at the insertion site to an amount of sequence comprising the insertion site without the tag sequence inserted to determine an integration rate of the genome editing tool.

8. 6. The method of claim 5, further comprising comparing an amount of sequence comprising the tag sequence at the insertion site to an amount of sequence comprising the tag sequence inserted off-target at the insertion site to determine an off-target integration rate of the genome editing tool.

9. 2. The method of claim 1, wherein the melting temperature between the tag sequence and the probe sequence is sufficient to allow binding of the linked capture probes.

10. 10. The method of claim 1, wherein the ligating step further comprises ligating unique barcodes onto the plurality of double-stranded nucleic acid fragments.

11. 11. The method of claim 10, wherein the unique barcode is sense specific.

12. 10. The method of claim 1, further comprising linking the target probe and the universal primer together using a linking molecule.

13. 13. The method of claim 12, wherein the target probe and the universal primer are linked together using click chemistry.

14. 10. The method of claim 1, further comprising repeating the exposing and extending steps to amplify the genomic region of interest prior to the sequencing step.

15. 15. The method of claim 1 or 14, further comprising amplifying the genomic region of interest using an unlinked universal primer prior to the sequencing step.

16. 15. The method of claim 1 or 14, further comprising amplifying the genomic region of interest using PCR amplification and a universal primer complementary to the universal priming site.

17. The method of claim 1, wherein the double-stranded nucleic acid fragment is sheared prior to ligation.

Citation Information

Patent Citations

  • Genome-wide, unbiased identification of DSBs by sequencing (GUIDE-Seq)

    JP2017519508A

  • Ligated duplex target capture

    JP2019514357A

  • Linked target acquisition

    JP7646575B2

  • Linked ligation

    WO2018104908A2