Methods for selective amplification for efficient rearrangement detection
The method addresses the inefficiencies of current genomic rearrangement detection by combining tagmentation and negative selection with amplification techniques, enhancing sensitivity and reducing background noise, thus improving the detection of genomic rearrangements.
Patent Information
- Application Number
- JP2025528917
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-18
- Filing Date
- 2023-11-17
- Publication Date
- 2025-11-28
AI Technical Summary
Current methods for detecting genomic rearrangements after gene editing, such as CRISPR/Cas9, require significant amounts of genomic DNA input, are inefficient in DNA shearing, and suffer from high background noise due to low sensitivity and lack of enrichment steps.
A method combining tagmentation with negative selection and amplification techniques using sequence tags to reduce DNA input requirements, eliminate background noise, and enhance sensitivity by identifying individual molecular recombination events.
The method significantly reduces sample processing time and improves the sensitivity of genomic rearrangement detection by utilizing tagmentation and amplification steps with sequence tags, enabling high detection sensitivity and cost-effectiveness.
Smart Images

Figure 2025538503000041 
Figure 2025538503000042 
Figure 2025538503000043
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 384,299, filed November 18, 2022, the entire contents of which are incorporated by reference herein, including any drawings.
[0002] Field The present disclosure relates to methods for quantitatively detecting genomic rearrangements. [Background technology]
[0003] background In the case of gene editing therapy involving nucleases such as CRISPR / Cas9, cleavage of genomic DNA in living cells results in double-strand breaks (DSBs), which have the potential to rejoin during DSB repair to generate genomic rearrangements. Typically, such rearrangements occur between a target site and a frequently edited off-target site, but they can also occur between two off-target sites, or between a target or off-target site and a fragile site or any other genomic locus where chromosome breakage or nicking can occur.
[0004] Several molecular methods, including combinatorial multiplex PCR using AMP, HTGTS, UDiTaS, CAST-seq, and rhAmpseq, have been applied to enable the detection of rearrangements in cells or mammalian tissues after gene editing. However, most of these methods use low-efficiency DNA shearing and adapter ligation techniques and require significant amounts of genomic DNA as input. One method, UDiTaS, uses a more efficient tagmentation technique, but this method lacks an enrichment step for rearrangements and therefore suffers from low sensitivity in rearrangement detection due to a high background of unrearranged molecules.
[0005] The disclosure provided herein provides a method that uses a combination of tagmentation and negative selection to reduce DNA input requirements, eliminate time-consuming sample processing steps, reduce background caused by spurious products, and improve the sensitivity of quantitative rearrangement detection. Summary of the Invention
[0006] A brief summary The present disclosure generally relates to the development of a method for quantitatively detecting genome rearrangements.In particular, as described in more detail below, some embodiments of the present disclosure provide a method involving tagmentation combined with amplification technology to reduce sample processing time, reduce background, and increase the sensitivity of quantitative rearrangement detection.The above summary is merely illustrative and is not intended to be limiting in any way.In addition to the exemplary embodiments and features described herein, further aspects, embodiments, objects and features of the present disclosure will be fully apparent from the drawings, detailed description and claims.
[0007] The present disclosure includes, among other things, a method for detecting genome-wide rearrangements in a nucleic acid genome. The method involves (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons containing a first sequence tag at the 5' end of the transposon under conditions such that the transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagmented into a plurality of nucleic acid fragments, each of which contains a sequence tag at the 5' end of a partially double-stranded nucleic acid fragment. In (b), a first amplification reaction is performed to amplify the rearranged, tagmented nucleic acid fragments to produce primary amplification products containing nucleotide sequences that include the rearranged target sequence using (i) a first target-specific primer containing a nucleotide sequence complementary to a region adjacent to the target site, (ii) a second primer containing a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment, and (iii) a blocking oligonucleotide containing a nucleotide sequence complementary to a region distal to the target site relative to the first target-specific primer. In (c), the primary amplification products are then sequenced.
[0008] In some embodiments, the sequence tag comprises a non-interacting sequence orthogonal to the genome, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof. In some embodiments, the sequence tag comprises a non-interacting sequence orthogonal to the genome, a UMI, and a transposase recognition site. In some embodiments, the sequence tag comprises a non-interacting sequence orthogonal to the genome, a UMI, a transposase recognition site, and an index sequence.
[0009] In one embodiment, the blocking oligonucleotide comprises an absent or blocked 3' OH, a spacer, an inverted nucleotide, or other modification to block extension of the 3' end. In one embodiment, the sequence tag comprises uracil.
[0010] In one embodiment, the blocking nucleotides comprise one or more phosphorothioate linkages, spacers, or other modifications at the 3' and 5' ends to block exonuclease digestion at the 3' and 5' ends, LNA, BNA, PNA, RNA, DNA, modified nucleic acids, or combinations thereof.
[0011] In one embodiment, the first and second primers comprise second and third sequence tags.
[0012] In one embodiment, the method disclosed herein further comprises, prior to (c), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (b) and fourth and fifth sequence tags at the 5' ends of the nested primers to produce secondary amplification products comprising the rearranged target sequence and additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0013] In one embodiment, the third and / or fourth primer comprises a barcode sequence.
[0014] In one embodiment, the fifth and sixth primers comprise a sequencing tag and / or index sequence.
[0015] In one embodiment, the method disclosed herein further comprises, prior to (c), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (b), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce secondary amplification products comprising the rearranged target sequence and one or two additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction, to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0016] In one embodiment, the fourth and fifth primers comprise a sequencing tag and / or index sequence.
[0017] In one embodiment, (b) further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
[0018] In one embodiment, the method disclosed herein further comprises, prior to (b), contacting the plurality of tagmented nucleic acid fragments with ddNTPs or other 3'-modified dNTPs and a DNA polymerase or terminal deoxynucleotidyl transferase to block all extendable 3' ends.
[0019] The present disclosure also encompasses a method for detecting genome-wide rearrangements in nucleic acid genomes. The method includes: (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons, each comprising a first sequence tag at the 5' end of the transposon, under conditions such that the transposons are inserted into the genomic DNA sample, and the genomic DNA sample is tagged into a plurality of nucleic acid fragments, each comprising a sequence tag at the 5' end of the partially double-stranded nucleic acid fragment; (b) contacting the plurality of tagmented nucleic acid fragments with a sequence-specific cleavage reagent; (c) performing a first amplification reaction to amplify the rearranged tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region adjacent to the target site, and (ii) a second primer comprising a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment; (d) sequencing the primary amplification products.
[0020] In one embodiment, the sequence-specific cleavage reagent is an enzymatic reagent or a chemical cleavage agent. In one embodiment, the enzymatic cleavage agent comprises a CRISPR Cas / gRNA complex.
[0021] In one embodiment, the sequence tag comprises a non-interacting sequence separate from the genome, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof. In some embodiments, the sequence tag comprises a non-interacting sequence separate from the genome, a transposase recognition site, and a UMI. In some embodiments, the sequence tag comprises a non-interacting sequence separate from the genome, a transposase recognition site, a UMI, and an index sequence.
[0022] In one embodiment, the first and second primers comprise second and third sequence tags.
[0023] In one embodiment, the method further comprises, prior to (d), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (c) and additional sequence tags at the 5' ends of the nested primers to produce secondary amplification products comprising the rearranged target sequence and the additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0024] In one embodiment, the third and / or fourth primer comprises a barcode sequence.
[0025] In one embodiment, the fifth and sixth primers include a sequencing tag and / or index sequence.
[0026] In one embodiment, the method further comprises, prior to (d), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (c), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce secondary amplification products comprising the rearranged target sequence and one or two additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction, to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0027] In one embodiment, the second and / or third primer comprises a barcode sequence.
[0028] In one embodiment, the fourth and fifth primers comprise a sequencing tag and / or index sequence.
[0029] In one embodiment, the method further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
[0030] The present disclosure also relates to a method for detecting genome-wide rearrangements in a nucleic acid genome, comprising: (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons, each comprising a first sequence tag at the 5' end of the transposon, wherein the sequence tag comprises an RNA promoter sequence, under conditions such that the transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagged into a plurality of nucleic acid fragments, each comprising a sequence tag at the 5' end of the partially double-stranded nucleic acid fragment; (b) transcribing the plurality of tagmented nucleic acid fragments into RNA; (c) contacting the RNA with a target sequence-specific DNA oligonucleotide probe and RNase H; (d) performing a reverse transcription amplification reaction using (i) a first primer comprising a nucleotide sequence complementary to a portion of the sequence tag located at the 3' or 5' end of the fragment and (ii) a second primer comprising a nucleotide sequence complementary to a region adjacent to the target site to produce a primary amplification product; and (e) sequencing the primary amplification product.
[0031] In one embodiment, the sequence tag comprises a non-interacting sequence separate from the genome, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof. In some embodiments, the sequence tag comprises a non-interacting sequence separate from the genome, a transposase recognition site, and a UMI. In some embodiments, the sequence tag comprises a non-interacting sequence separate from the genome, a transposase recognition site, a UMI, and an index sequence.
[0032] In one embodiment, the first and second primers comprise a sequence tag.
[0033] In one embodiment, the method further comprises, prior to (e), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (d) and an additional sequence tag to produce secondary amplification products comprising the rearranged target sequence and the additional sequence tag. In one embodiment, the method further comprises performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0034] In one embodiment, the third and / or fourth primer comprises a barcode sequence.
[0035] In one embodiment, the fifth and sixth primers comprise an adapter sequence and / or an index sequence.
[0036] In one embodiment, the method further comprises, prior to (e), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (d), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce secondary amplification products comprising the rearranged target sequence and one or two additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction, to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0037] In one embodiment, the second and / or third primer comprises a barcode sequence.
[0038] In one embodiment, the fourth and fifth primers comprise an adapter sequence and / or an index sequence.
[0039] In one embodiment, the method includes where (d) further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site. [Brief explanation of the drawings]
[0040] BRIEF DESCRIPTION OF THE DRAWINGS The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings.
[0041] [Figure 1-1] FIG. 1 illustrates the workflow of an exemplary method of the present disclosure. [Figure 1-2] Same as above.
[0042] [Figure 2-1] FIG. 2 illustrates the workflow of an exemplary method of the present disclosure. [Figure 2-2] Same as above.
[0043] [Figure 3-1] FIG. 3 illustrates the workflow of an exemplary method of the present disclosure. [Figure 3-2] Same as above.
[0044] [Figure 4-1] FIG. 4 illustrates the workflow of an exemplary method of the present disclosure. [Figure 4-2] Same as above.
[0045] [Figure 5-1] FIG. 5 illustrates the workflow of an exemplary method of the present disclosure. [Figure 5-2] Same as above.
[0046] [Figure 6-1]FIG. 6 illustrates the workflow of an exemplary method of the present disclosure. [Figure 6-2] Same as above.
[0047] [Figure 7-1] FIG. 7 illustrates the workflow of an exemplary method of the present disclosure. [Figure 7-2] Same as above. [Figure 7-3] Same as above.
[0048] [Figure 8-1] FIG. 8 illustrates the workflow of an exemplary method of the present disclosure. [Figure 8-2] Same as above. [Figure 8-3] Same as above. [Figure 8-4] Same as above.
[0049] [Figure 9-1] FIG. 9 illustrates the workflow of an exemplary method of the present disclosure. [Figure 9-2] Same as above. [Figure 9-3] Same as above.
[0050] [Figure 10] Figures 10A and 10B show the blocking effect of the blocker oligo.
[0051] [Figure 11] 11A and 11B show the number of split reads identified in Example 2 of an exemplary method of the present disclosure.
[0052] [Figure 12] 12A and 12B show validation of interchromosomal translocations identified in genomic DNA used in all methods of this disclosure.
[0053] [Figure 13] 13A and 13B show the number of split reads identified in Example 4 of an exemplary method of the present disclosure.
[0054] [Figure 14] FIG. 14 shows the number of split reads identified in Example 6 of an exemplary method of the present disclosure.
[0055] [Figure 15] 15A and 15B show the number of split reads identified in Example 8 of an exemplary method of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0056] Detailed Description The present disclosure generally relates to new techniques for detecting genome rearrangements associated with genome editing at specific target sites. These methods address the problems of current detection technologies, such as the need for a significant amount of input genomic DNA, low sensitivity, and high background. Provided herein are, among other things, improved methods for detecting rearranged nucleic acid molecules, involving tagmentation and negative selection techniques, which enable high detection sensitivity and cost-effectiveness compared to currently available techniques for detecting genome rearrangements. For example, some methods of the present disclosure involve tagmentation and subsequent amplification steps using sequence tags, which allow for the identification of individual molecular recombination events. Other methods of the present disclosure involve tagmentation, sequence-specific cleavage, and subsequent amplification steps using sequence tags. The present disclosure also provides methods involving tagmentation, RNA transcription, and subsequent reverse transcription using sequencing tags.
[0057] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0058] While various features of the present disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, while the present disclosure may be described herein for clarity in the context of separate embodiments, the present disclosure may also be implemented in a single embodiment.
[0059] definition The singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. "A and / or B" is used herein to encompass all of the following alternatives: "A," "B," "A or B," and "A and B."
[0060] As used herein, "barcode" refers to one or more known nucleotide sequences that are used to identify the nucleic acid to which the barcode is associated. In some embodiments, the barcode sequence allows for multiplexing of products from different target sites in separate reactions or different samples.
[0061] As used herein, "UMI" (unique molecular identifier) refers to one or more randomized or semi-randomized nucleotide sequences that are used to group and de-duplicate all products derived from a single starting molecule after sequencing.
[0062] As used herein, "cleavage" refers to the breaking of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two separate single-strand cleavage events. DNA cleavage can result in the production of either blunt ends or cohesive ends.
[0063] As used herein, the term "detecting" a nucleic acid molecule or fragment thereof refers to determining the presence of a nucleic acid molecule, typically when the nucleic acid molecule or fragment thereof is completely or partially separated from other components of a sample or composition.
[0064] As used herein, the term "nuclease" refers to a polypeptide capable of cleaving phosphodiester bonds between nucleotide subunits of nucleic acids.
[0065] As used herein, the terms "nucleic acid," "nucleic acid molecule," or "polynucleotide" are used interchangeably herein. They refer to a polymer of deoxyribonucleotides or ribonucleotides in either single-stranded or double-stranded form, and unless otherwise specified, include known analogs of natural nucleotides that can function in a manner similar to naturally occurring nucleotides. These terms also include nucleic acid-like structures with synthetic backbones, as well as amplification products. Both DNA and RNA are polynucleotides. The polymers may be composed of natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine and 2-thiocytidine), chemically modified bases (2'-O,4'-C-methylene bridges / locked nucleic acids), biologically modified bases (e.g., methylated bases), intercalating bases, spacers, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0066] As used herein, the term "oligonucleotide" refers to a series of nucleotides or their analogs. Oligonucleotides may be obtained by several methods, including, for example, chemical synthesis, restriction enzyme digestion, or PCR. As will be understood by those of skill in the art, the length (i.e., number of nucleotides) of an oligonucleotide can often vary widely, depending on the intended function or use of the oligonucleotide. Generally, an oligonucleotide contains between about 5 and about 300 nucleotides, e.g., between about 15 and about 200 nucleotides, between about 15 and about 100 nucleotides, or between about 15 and about 50 nucleotides. Throughout this specification, whenever an oligonucleotide is represented by a sequence of letters (selected from the four base letters A, C, G, and T, representing adenosine, cytidine, guanosine, and thymidine, respectively), the nucleotides are represented in 5' to 3' order from left to right. In certain embodiments, the sequence of the oligonucleotide includes one or more degenerate residues as described herein.
[0067] As used herein, the terms "amplify," "amplified," or "amplifying" when used in reference to a nucleic acid or nucleic acid reaction refers to an in vitro method of making copies of a particular nucleic acid, such as a target nucleic acid or a tagged nucleic acid produced by the methods described herein.
[0068] As used herein, "primer" refers to a nucleic acid that has a sequence that is complementary and specific to a target or template nucleic acid, such as a known sequence in DNA.This means that they must be sufficiently complementary to their respective strands to hybridize to form desired hybridization products, and then be extendable by DNA polymerase.In some cases, primers have exact complementarity to target nucleic acid or template nucleic acid.However, in many situations, exact complementarity is impossible or unlikely, and there may be one or more mismatches that do not prevent hybridization or the formation of primer extension products using DNA polymerase.
[0069] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range, and any other stated or intervening value in that stated range, is encompassed within the disclosure unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, which are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed within the disclosure.
[0070] All ranges disclosed herein also encompass any and all possible subranges and combinations of subranges. Any recited range can be recognized as fully descriptive and permitting division of the same range into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range described herein can be readily divided into a lower third, middle third, and upper third, etc. Additionally, as will be understood by those skilled in the art, all language, such as "up to," "at least," "greater than," and "less than," encompasses the recited numbers and refers to a range that can be subsequently divided into subranges as described above. Finally, as will be understood by those skilled in the art, a range encompasses each individual member. Thus, for example, a group having 1 to 3 items refers to a group having 1, 2, or 3 items. Similarly, a group having 1 to 5 items refers to a group having 1, 2, 3, 4, or 5 items, etc.
[0071] It is understood that certain features of the present disclosure, which are described for clarity in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the present disclosure, which are described for brevity in the context of a single embodiment, may also be provided separately or in any subcombination. All combinations of the embodiments related to the present disclosure are specifically embraced by the present disclosure and are disclosed herein as if each and every combination were individually and expressly disclosed. Furthermore, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein as if each and every such subcombination were individually and expressly disclosed herein.
[0072] Although features of the present disclosure may be described in the context of a single embodiment, the features may be provided separately or in any suitable combination. Conversely, although the present disclosure may be described herein in the context of separate embodiments for clarity, the present disclosure may be implemented in a single embodiment. Any published patent applications and any other published references, documents, manuscripts, and scientific literature cited herein are incorporated by reference for any purpose. In the case of conflict, the present specification, including definitions, will prevail. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0073] Methods of the present disclosure Method I As described in more detail below, one embodiment of the present disclosure provides a method for detecting genome-wide rearrangements in a nucleic acid genome. The method involves (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons containing a first sequence tag at the 5' end of the transposon under conditions such that the transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagged into a plurality of nucleic acid fragments, each of which contains a sequence tag at the 5' end of a partially double-stranded nucleic acid fragment. A first amplification reaction is performed to amplify the rearranged, tagmented nucleic acid fragments using (i) a first target-specific primer containing a nucleotide sequence complementary to a region adjacent to the target site, (ii) a second primer containing a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment, and (iii) a blocking oligonucleotide containing a nucleotide sequence complementary to a region distal to the target site relative to the first target-specific primer, to produce a primary amplification product containing a nucleotide sequence containing the rearranged target sequence. The primary amplification product is then sequenced.
[0074] Genomic DNA sample The cells or tissues used in the methods provided herein can be of any eukaryotic cell type, including, but not limited to, human cells, non-human primate cells, mammalian cell types, vertebrate cell types, yeast, plant cells. These cells can include, for example, primary cells and / or tissues, cells or tissues that have been cultured for at least some period of time, or a combination of primary and cultured cells and / or tissues.
[0075] In some embodiments, the methods described herein are performed on genomic DNA from a single cell. For example, the genomic DNA from a single cell can be amplified before carrying out the methods described herein. Whole genome amplification methods are known in the art. Various protocols and / or commercially available kits can be used. Examples of commercially available kits include, but are not limited to, the REPLI-g Single Cell Kit from QIAGEN, the GENOMEPLEX® Single Cell Whole Genome Amplification Kit from Sigma Aldrich, the Ampli1™ WGA Kit from Silicon Biosystems, and the illustra Single Cell GenomiPhi DNA Amplification Kit from GE Healthcare Life Sciences.
[0076] Alternatively, or additionally, the methods disclosed herein can be used with genomic DNA samples from eukaryotic cells and / or tissues, or from prokaryotic cells. For example, the methods of the present disclosure can be performed using genomic DNA from isolates from microorganisms and / or patients (e.g., patients receiving antibiotics). In some embodiments, genomic DNA from a microbial community and / or one or more microbiota is used, for example, for metagenomic mining. (See, e.g., Delmont et al., "Metagenomic mining for microbiologists," ISME J. 2011 December;5(12):1837-43.)
[0077] Genomic DNA can be prepared using any of a variety of suitable methods, including, for example, the specific manipulation of cells and / or tissues described herein. Exemplary, non-limiting manipulations include contacting cells and / or tissues with nucleases (e.g., site-specific nucleases and / or RNA-guided nucleases) or genome editing systems containing such nucleases.
[0078] site-specific nucleases As described above, genomic DNA samples are obtained from cells or tissues that have been contacted with a site-specific nuclease. Generally, nucleases are site-specific in that they are known or predicted to cleave only at a specific sequence or set of sequences, referred to herein as the "target site" of the nuclease.
[0079] In the methods disclosed herein, the step of contacting with the nuclease is generally carried out under conditions favoring cleavage by the nuclease. That is, even if a given candidate target site or variant target site may not actually be cleaved by the nuclease, the incubation conditions are such that the nuclease will cleave at least a significant portion (e.g., at least 1%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%) of the template containing the known target site. For known and generally well-characterized nucleases, such conditions are generally known in the art and / or can be easily discovered or optimized. For newly discovered nucleases, such conditions can generally be approximated using information on related nucleases that are better characterized (e.g., homologs and orthologs).
[0080] In some embodiments, the nuclease is an endonuclease. In some embodiments, the nuclease is a site-specific endonuclease (e.g., a restriction endonuclease, a meganuclease, a transcription activator-like effector nuclease (TALEN), a zinc finger nuclease, etc.).
[0081] In some embodiments, the site specificity of site-specific nuclease is conferred by auxiliary molecule.For example, CRISPR-associated (Cas) nuclease is guided to specific site by " guide RNA " or gRNA as described herein.In some embodiments, nuclease is RNA-guided nuclease.In some embodiments, nuclease is CRISPR-associated nuclease.
[0082] In some embodiments, the nuclease is a homolog or ortholog of a previously known nuclease, eg, a newly discovered homolog or ortholog.
[0083] In some embodiments, the nuclease is a base editor.
[0084] RNA-guided nucleases The RNA-guided nuclease of the present disclosure includes, but is not limited to, naturally occurring class 2 CRISPR nucleases such as Cas9 and Cas12a, and other nucleases derived or obtained therefrom.In functional terms, an RNA-guided nuclease is defined as a nuclease that (a) interacts with (e.g., complexes with) gRNA; and (b) together with gRNA, associates with and optionally cleaves or modifies the target region of DNA, including (i) a sequence complementary to the targeting domain of gRNA, and optionally (ii) an additional sequence called "protospacer adjacent motif" or "PAM", which will be described in more detail below.RNA-guided nucleases can be broadly defined by their PAM specificity and cleavage activity, even though there may be variations between individual RNA-guided nucleases that share the same PAM specificity or cleavage activity. Those skilled in the art will understand that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using any suitable RNA-guided nuclease with a particular PAM specificity and / or cleavage activity. For this reason, unless otherwise specified, the term RNA-guided nuclease should be understood as a general term and is not limited to any particular type (e.g., Cas9 vs. Cas12a), species (e.g., S. pyogenes vs. S. aureus), or mutation (e.g., full-length vs. truncated or split; naturally occurring vs. engineered PAM specificity, etc.) of RNA-guided nuclease.
[0085] The PAM sequence takes its name from its sequence relationship to a "protospacer" sequence (or "spacer") that is complementary to the gRNA targeting domain. Together with the protospacer sequence, the PAM sequence defines the target region or sequence for a particular RNA-guided nuclease / gRNA combination.
[0086] Various RNA-guided nucleases may require different sequence relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence 3' of the visualized protospacer relative to the guide RNA targeting domain. Meanwhile, Cas12a generally recognizes the PAM sequence 5' of the protospacer.
[0087] In addition to recognizing specific sequence orientations of the PAM and protospacer, RNA-guided nucleases can also recognize specific PAM sequences. For example, S. aureus Cas9 recognizes the NNGRRT or NNGRRV PAM sequence, with the N residue immediately 3' to the region recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence. F. novicida Cas12a recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-guided nucleases, and strategies for identifying novel PAM sequences are described in Shmakov et al., 2015, Molecular Cell 60, 385-397, November 5, 2015. It should also be noted that an engineered RNA-guided nuclease may have a PAM specificity that differs from that of the reference molecule (e.g., in the case of an engineered RNA-guided nuclease, the reference molecule may be the naturally occurring variant from which the RNA-guided nuclease is derived, or the naturally occurring variant with the greatest amino acid sequence homology to the engineered RNA-guided nuclease).
[0088] In addition to their PAM specificity, RNA-guided nucleases can be characterized by their DNA cleavage activity: naturally occurring RNA-guided nucleases typically form DSBs in target nucleic acids, but engineered variants have been produced that generate only SSBs (see above) Ran & Hsu, et al., Cell 154(6), 1380-1389, Sep. 12, 2013 ("Ran"), incorporated herein by reference), or that do not cleave at all.
[0089] Cas9 The crystal structures of S. pyogenes Cas9 (Jinek et al., Science 343(6176),1247997,2014 ("Jinek 2014") and S. aureus Cas9 complexed with single-molecule guide RNA and target DNA have been determined (Nishimasu 2014; Anders et al., Nature. 2014 Sep. 25;513(7519):569-73 ("Anders 2014"); and Nishimasu 2015).
[0090] Naturally occurring Cas9 proteins contain two lobes: a recognition (REC) lobe and a nuclease (NUC) lobe; each of these contains specific structural and / or functional domains. The REC lobe contains an arginine-rich bridge helix (BH) domain and at least one REC domain (e.g., a REC1 domain and, optionally, a REC2 domain). The REC lobe does not share structural similarity with other known proteins, indicating that it is a unique functional domain. Without wishing to be bound by any theory, mutational analysis suggests specific functional roles for the BH and REC domains: the BH domain appears to play a role in gRNA:DNA recognition, while the REC domain is thought to interact with the repeat:anti-repeat duplex of the gRNA and mediate the formation of the Cas9 / gRNA complex.
[0091] The NUC lobe contains a RuvC domain, an HNH domain, and a PAM-interacting (PI) domain. The RuvC domain shares structural similarity with members of the retroviral integrase superfamily and cleaves the non-complementary (i.e., bottom) strand of the target nucleic acid. It can be formed from two or more split RuvC motifs (e.g., RuvCI, RuvCII, and RuvCIII in S. pyogenes and S. aureus). Meanwhile, the HNH domain is structurally similar to the HNN endonuclease motif and cleaves the complementary (i.e., top) strand of the target nucleic acid. As its name suggests, the PI domain contributes to PAM specificity.
[0092] While specific functions of Cas9 have been linked (but not necessarily fully determined) to specific domains described above, these and other functions may be mediated or influenced by other Cas9 domains or by multiple domains on either lobe. For example, as described in Nishimasu 2014, in S. pyogenes Cas9, the repeat:antirepeat duplex of the gRNA falls into a groove between the REC lobe and the NUC lobe, and nucleotides of the duplex interact with amino acids in the BH, PI, and REC domains. Some nucleotides in the first stem-loop structure also interact with amino acids in multiple domains (PI, BH, and REC1), as do some nucleotides in the second and third stem-loops (RuvC and PI domains).
[0093] Cas12a Crystal structure of Acidaminococcus species. Cas12a in complex with crRNA and a double-stranded (ds) DNA target containing a TTTN PAM sequence has been elucidated by Yamano et al. (Cell. 2016 May 5;165(4):949-962 ("Yamano"), incorporated herein by reference). Cas12a, like Cas9, has two lobes: the REC (recognition) lobe and the NUC (nuclease) lobe. The REC lobe contains the REC1 and REC2 domains, which lack similarity to any known protein structure. Meanwhile, the NUC lobe contains three RuvC domains (RuvC-I, -II, and -III) and a BH domain. However, in contrast to Cas9, the Cas12a REC lobe lacks the HNH domain and also contains other domains that lack similarity to known protein structures: a structurally unique PI domain, three Wedge (WED) domains (WED-I, -II, and -III), and a nuclease (Nuc) domain.
[0094] Although Cas9 and Cas12a share similarities in structure and function, it should be understood that certain Cas12a activities are mediated by structural domains that are not similar to either Cas9 domain. For example, cleavage of the complementary strand of target DNA appears to be mediated by the Nuc domain, which is sequence- and spatially distinct from the HNH domain of Cas9. Additionally, the non-targeting portion (handle) of Cas12a gRNA adopts a pseudo-knot structure, rather than the stem-loop structure formed by the repeat:anti-repeat duplex in Cas9 gRNA.
[0095] Base Editor Engineered base editors, such as Cas9 nickases, have been developed that provide a base-modifying enzyme domain (e.g., a deaminase) or reverse transcriptase along with a modified CRISPR-associated targeting domain (Gaudelli et al., Nature, 2017, 24644; Komor et al., Nature, 2016, 533:420-424; Yang et al., Nat Commun. 2016;7:13330; Anzalone, Nature. 2019,.doi:10.1038 / s41586-019-1711-4). Such base editors and prime editors typically introduce a single nick in one strand of the target site, then introduce one or more base modifications that resolve into base edits after DNA replication. Such enzymes have generally been thought not to induce double-strand breaks and resulting rearrangements. However, because single-stranded nicks can actually result in double-stranded breaks as a result of replisome disassembly and DNA cleavage (Vrtis et al., Mol Cell. 2021;81:1309-1318.e6.), it has been shown that homologous recombination and large-scale rearrangements can occur due to DNA cleavage at the site where DNA nicking occurs. Thus, the methods of the present invention are also applicable to detecting rearrangements induced by base editors, prime editors, and other site-specific gene editors that introduce single-stranded nicks.
[0096] Nucleic acid encoding an RNA-guided nuclease The present specification provides nucleic acids encoding RNA-guided nucleases, such as Cas9, Cas12a, or functional fragments thereof. Exemplary nucleic acids encoding RNA-guided nucleases have been previously described (see, for example, Cong et al., Science. 2013 Feb. 15; 339 (6121): 819-23 ("Cong 2013"); Wang et al., PLoS One. 2013 Dec. 31; 8 (12): e85650 ("Wang 2013"); Mali 2013; Jinek 2012).
[0097] In some cases, the nucleic acid encoding RNA-guided nuclease can be a synthetic nucleic acid sequence.For example, synthetic nucleic acid molecule can be chemically modified.In certain embodiments, the mRNA encoding RNA-guided nuclease has one or more (for example, all) of the following characteristics: can be capped; can be polyadenylated; and can be substituted with 5-methylcytidine and / or pseudouridine.
[0098] The synthetic nucleic acid sequence can also be codon-optimized, for example, at least one non-common or less common codon is replaced with a common codon. For example, the synthetic nucleic acid can direct the synthesis of an optimized messenger mRNA, for example, optimized for expression in a mammalian expression system as described herein. An example of a codon-optimized Cas9 coding sequence is provided in International Publication No. 2016 / 073990 ("Cotta-Ramusino").
[0099] Additionally or alternatively, the nucleic acid encoding the RNA-guided nuclease may contain a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.
[0100] Guide RNA (gRNA) molecules The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates the specific association (or "targeting") of an RNA-guided nuclease, such as Cas9 or Cas12a, to a target sequence, such as a genomic or episomal sequence within a cell. gRNAs can be unimolecular (comprising a single RNA molecule, or called chimeric) or modular (comprising more than one, typically two separate RNA molecules, such as crRNA and tracrRNA, which are usually associated with each other, e.g., by duplexing). gRNAs and their component parts are described throughout the literature, for example, in Briner et al. (Molecular Cell 56(2), 333-339, October 23, 2014 ("Briner") (incorporated by reference), and Cotta-Ramusino. In bacteria and archaea, type II CRISPR systems generally include an RNA-guided nuclease protein such as Cas9, a CRISPR RNA (crRNA) that includes a 5' region complementary to an exogenous sequence, and a trans-activating crRNA (tracrRNA) that includes a 5' region that is complementary to and forms a duplex with the 3' region of the crRNA. Without intending to be bound by any theory, it is believed that this duplex promotes the formation of the Cas9 / gRNA complex and is required for its activity.When Type II CRISPR systems were adapted for use in gene editing, it was discovered that the crRNA and tracrRNA could be linked into a single unimolecular or chimeric guide RNA by a four-nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that bridges, in one non-limiting example, the complementary region of the crRNA (at its 3' end) and the complementary region of the tracrRNA (at its 5' end) (Mali et al. Science. 2013 Feb. 15;339(6121):823-826 ("Mali 2013"); Jiang et al. Nat Biotechnol. 2013 March;31(3):233-239 ("Jiang"); and Jinek et al., 2012 Science August 17;337(6096):816-821 ("Jinek 2012"), all of which are incorporated herein by reference.
[0101] Guide RNAs, whether single-molecule or modular, include a "targeting domain" that is fully or partially complementary to a target domain within a target sequence, such as a DNA sequence in the genome of a cell where editing is desired. Targeting domains are referred to by various names in the literature, including, but not limited to, "guide sequence" (Hsu et al., Nat Biotechnol. 2013 September;31(9):827-832, ("Hsu"), incorporated herein by reference), "complementarity region" (Cotta-Ramusino), "spacer" (Briner), and generally "crRNA" (Jiang). Regardless of the name given to them, targeting domains are typically 10-30 nucleotides in length, and in certain embodiments, 16-24 nucleotides in length (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides in length), and are located at or near the 5' end in the case of Cas9 gRNAs and at or near the 3' end in the case of Cas12a gRNAs.
[0102] In addition to the targeting domain, gRNAs typically (but not necessarily, as described below) contain multiple domains that can affect the formation or activity of the gRNA / Cas9 complex. For example, as described above, the duplex structure formed by the first and second complementary domains of the gRNA (also called the repeat:anti-repeat duplex) can interact with the recognition (REC) lobe of Cas9 and mediate the formation of the Cas9 / gRNA complex. (Nishimasu et al., Cell 156, 935-949, February 27, 2014 ("Nishimasu 2014") and Nishimasu et al., Cell 162, 1113-1126, August 27, 2015 ("Nishimasu 2015"), both of which are incorporated herein by reference.) It should be noted that the first and / or second complementary domains may contain one or more polyA tracks that can be recognized by RNA polymerase as termination signals. Therefore, the sequences of the first and second complementary domains are optionally modified to eliminate these tracks and facilitate complete in vitro transcription of the gRNA, for example, by using AG swaps or AU swaps as described in Briner. These and other similar modifications to the first and second complementary domains are within the scope of the present disclosure.
[0103] Along with the first and second complementarity domains, Cas9 gRNAs typically contain two or more additional duplexed regions that are involved in nuclease activity in vivo, but not necessarily in vitro (Nishimasu 2015). The first stem-loop 1 near the 3' portion of the second complementarity domain is variously referred to as the "proximal domain" (Cotta-Ramusino), "stem-loop 1" (Nishimasu 2014 and 2015), and "nexus" (Briner). One or more additional stem-loop structures are generally present near the 3' end of the gRNA, the number of which varies by species: S. pyogenes gRNAs typically contain two 3' stem-loops (for a total of four stem-loop structures encompassing repeat:anti-repeat duplexes), while S. aureus and other species have only one (for a total of three stem-loop structures). A description of conserved stem-loop structures (and gRNA structures more generally) organized by species is provided in Briner.
[0104] While the above description focuses on gRNAs for use with Cas9, it should be understood that other RNA-guided nucleases have been discovered or invented (or may be discovered in the future) that utilize gRNAs that differ in some respects from those described thus far. For example, Cas12a ("CRISPR from Prevotella and Franciscella 1") is a recently discovered RNA-guided nuclease that does not require a tracrRNA to function. (Zetsche et al., 2015, Cell 163, 759-771 Oct. 22, 2015 ("Zetsche I"), incorporated herein by reference). gRNAs for use with the Cas12a genome editing system generally include a targeting domain and a complementarity domain (alternatively referred to as a "handle"). It should also be noted that in gRNAs for use with Cas12a, the targeting domain is typically located at or near the 3' end (the handle is located at or near the 5' end of the Cas12a gRNA), rather than the 5' end as described above in connection with the Cas9 gRNA.
[0105] However, those skilled in the art will understand that while structural differences may exist between gRNAs from different prokaryotic species or between Cas12a and Cas9 gRNAs, the principles by which gRNAs function are generally consistent. Because of this operational consistency, gRNAs can be broadly defined by their targeting domain sequences, and those skilled in the art will understand that a given targeting domain sequence can be incorporated into any suitable gRNA, including monomolecular or chimeric gRNAs, or gRNAs incorporating one or more chemical and / or sequence modifications (substitutions, additional nucleotides, truncations, etc.). Therefore, for the sake of economy of presentation in this disclosure, gRNAs will be described only in terms of their targeting domain sequences.
[0106] More generally, those skilled in the art will understand that some aspects of the present disclosure relate to systems, methods and compositions that can be implemented using multiple RNA-guided nucleases.For this reason, unless otherwise specified, the term gRNA should be understood to encompass any suitable gRNA that can be used with any RNA-guided nuclease, and not just the gRNA that is compatible with a specific species of Cas9 or Cas12a.For example, in certain embodiments, the term gRNA can encompass the gRNA that is used with any RNA-guided nuclease that occurs in class 2 CRISPR system, for example, type II or type V or CRISPR system, or the RNA-guided nuclease that is derived from or compatible with it.
[0107] Transposons and Tagmentation In the first step, genomic DNA samples obtained from cells or tissues exposed to site-specific nucleases are contacted with multiple transposons to introduce known DNA sequences into the ends of genomic DNA fragments. For example, Tn5 transposase catalyzes strand transfer via nucleophilic attack on target DNA by the activated 3-OH group at the transposon end, leaving a 9-bp gap at the target site (Vaezeslami et al., J Bacteriol. 2007;189:7436-7441). A transposome is a complex of the transposase enzyme and DNA containing transposon end sequences (also known as "transposase recognition sequences" or "mosaic ends" (MEs)). When a double-stranded oligonucleotide bearing a transposon recognition sequence at one end is provided, transposition results in the introduction of free DNA ends into the target molecule (e.g., "tagmentation"; Adey et al., Genome Biol. 2010;11:R119). This system can be adapted using a hyperactive transposase enzyme and a modified DNA oligonucleotide (sequence tag) containing an ME to introduce tags into both strands of the tagmented product with functional DNA molecules (e.g., primer binding sites). Any transposase enzyme with tagmentation activity can be used, for example, any transposase enzyme that can mediate strand transfer and incorporation of an oligonucleotide (e.g., a tag) at the end of the tagmented DNA. In some embodiments, the transposase is any transposase capable of conservative transposition. In some embodiments, the transposase is a cut-and-paste transposase. Other types of transposases are known in the art and are within the scope of this disclosure. For example, suitable transposase enzymes include, but are not limited to, Tn5, Tn5059, Mos-1, HyperMu™, Hermes, Tn7, or any functional variant or derivative of the transposase enzymes listed above.
[0108] For example, Tn5 transposase can be produced as a purified protein monomer. Tn5 transposase is also commercially available (e.g., manufacturer Illumina, Illumina.com, catalog number 15027865, TD Tagment DNA Buffer catalog number 15027866). Subsequently, a desired oligonucleotide can be loaded, for example, a ssDNA oligonucleotide containing an ME for Tn5 recognition and an additional functional sequence (e.g., a primer binding site, e.g., a UMI) can be annealed to form a dsDNA mosaic end oligonucleotide (MEDS) that is recognized by Tn5 during dimer assembly (e.g., transposome dimerization). In some embodiments, a hyperactive Tn5 transposase can be loaded with a tag (e.g., an oligonucleotide of interest) that can simultaneously fragment and tag the genome with a desired sequence.
[0109] As described above, transposons include a first sequence tag at the 5' end of the transposon. As used herein, the first sequence tag refers to a non-target nucleic acid component, generally DNA, that provides a means for addressing the nucleic acid fragment to which it is attached. For example, in some embodiments, the sequence tag is or includes a nucleotide sequence that allows for identification, recognition, and / or molecular or biochemical manipulation of the DNA to which the tag is attached (e.g., by providing a site for annealing an oligonucleotide, e.g., a primer for DNA polymerase extension, or an oligonucleotide for a capture or ligation reaction). In some embodiments, the sequencing tag is used to generate templates for next-generation sequencing on a specific sequencing platform (e.g., for an Illumina sequencing platform; for a Thermo Fisher Ion Torrent sequencing platform; for a Pacific Biosciences Sequel sequencing platform; for a MGI sequencing platform; or for any other sequencing platform). In some embodiments, the sequencing tag is a full-length Illumina forward (i5) adapter. In some embodiments, the sequencing tag is a full-length Illumina reverse (i7) adapter. In some embodiments, the sequence tag comprises a unique sequence (e.g., a sequence that is distinct from the genome and does not interact with it), a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof. In some embodiments, the sequence tag comprises a sequence that is distinct from the genome and does not interact with it, a unique molecular identifier (UMI), and a transposase recognition site. In some embodiments, the sequence tag comprises a sequence that is distinct from the genome and does not interact with it, a unique molecular identifier (UMI), an index sequence, and a transposase recognition site.
[0110] In some embodiments, the UMI is a randomly generated sequence. In some embodiments, the UMI is between 8 and 20 nucleotides in length, e.g., between 10 and 16 nucleotides in length, e.g., 10, 11, 12, 13, 14, 15, and 16 nucleotides in length. The production and use of UMIs in various contexts is known in the art.
[0111] Tagmentation is the common first step of all methods described herein. In some embodiments, the step of fragmenting the genomic DNA in the cells of the biological sample comprises contacting the biological sample containing genomic DNA with a transposase enzyme (e.g., a transposome, e.g., a reaction mixture (e.g., a solution) containing a transposase) under any suitable conditions. In some embodiments, transposomes are assembled by annealing transposon ends containing oligonucleotides at about 95, 96, 97, 98, or 99°C for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes, followed by about 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 cycles of about 80, 85, 90, 95, or 99°C for 1 minute, decreasing the temperature by 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, or 2°C per cycle. The transposase and annealed oligonucleotides can then be incubated at 15, 20, 25, 30, 35, or 40°C for 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, or 120 minutes. In some embodiments, such suitable conditions result in fragmentation (e.g., tagmentation) of genomic DNA of cells present in the biological sample. Typical conditions depend on the transposase enzyme used and can be determined using routine methods known in the art. Therefore, suitable conditions can be conditions (e.g., buffer, salt, concentration, pH, temperature, time conditions) under which the transposase enzyme is functional, for example, conditions under which the transposase enzyme exhibits transposase activity, particularly tagmentation activity, in the biological sample.In some embodiments, the tagmentation reaction comprises 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, or 500 ng of genomic DNA at about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, or 500 ng of genomic DNA at about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, , 5, 6, 7, or 8 μL of transposomes and incubating at 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, or 120 minutes ....
[0112] In one non-limiting example, the reaction mixture comprises a transposase enzyme in a buffer solution (eg, Tris-acetate) having a pH of about 6.5 to about 8.5, such as about 7.0 to about 8.0, for example about 7.5. Additionally or alternatively, the reaction mixture can be used at any suitable temperature based on the optimum temperature of the transposase enzyme, e.g., for Tn5, about 10° to about 55° C., e.g., about 10° to about 54° C., about 11° to about 53° C., about 12° to about 52° C., about 13° to about 51° C., about 14° to about 50° C., about 15° to about 49° C., about 16° to about 48° C., about 17° to about 47° C., e.g., about 10°, about 12°, about 15°, about 18°, about 20°, about 22°, about 25°, about 28° C., about 30°, about 33°, about 35° C., or about 37° C., preferably about 30° to about 40° C., e.g., about 37° C. In some embodiments, the transposase enzyme can be contacted with the biological sample for about 10 minutes to about 1 hour. In some embodiments, the transposase enzyme can be contacted with the biological sample for about 20, about 30, about 40, or about 50 minutes. In some embodiments, the transposase enzyme can be contacted with the biological sample for about 1 hour to about 4 hours.
[0113] Briefly, in some embodiments, Tn5 tagmentation uses staggered transposon ends with an optional 3' extension block (e.g., a 3' dideoxynucleotide, inverted nucleotide, or amino modifier) and a 5' sequence tag. As described above, in some embodiments, the sequence tag has a UMI, which typically contains a random 12-nucleotide sequence that differs between each transposon end. UMI allows products generated from a single starting genome fragment to be identified, grouped, and analyzed together after multiple amplification steps. This allows individual molecular recombination events to be identified and their frequency of occurrence in a population quantified relative to other unique events. It also allows the identification of variants associated with each haplotype in the starting genome. The use of tagmentation also avoids the inefficient and time-consuming steps of DNA shearing, end repair, and adapter ligation, reducing input DNA requirements and speeding up the process.
[0114] In the methods described herein, when multiple transposons are inserted into a genomic DNA sample, the genomic DNA sample is tagmented into multiple nucleic acid fragments, each containing a sequence tag at the 5' end of a partially double-stranded nucleic acid fragment.
[0115] amplification In the methods described herein, after tagmentation, a first amplification reaction is performed to amplify the rearranged, tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region adjacent to the target site, (ii) a second primer comprising a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment, and (iii) a blocking oligonucleotide comprising a nucleotide sequence complementary to a region distal to the target site relative to the first target-specific primer to produce a primary amplification product comprising a nucleotide sequence that includes the rearranged target sequence.
[0116] Numerous methods for amplifying nucleic acids are known in the art, including polymerase chain reaction (PCR), ligase chain reaction, strand displacement amplification reaction, rolling circle amplification reaction, transcription-mediated amplification methods such as NASBA (e.g., U.S. Pat. No. 5,409,818), and loop-mediated amplification methods (e.g., "LAMP" amplification using loop-forming sequences, e.g., as described in U.S. Pat. No. 6,410,278). The nucleic acid being amplified can be DNA, which comprises, consists of, or is derived from DNA or RNA, or a mixture of DNA and RNA, including modified DNA and / or RNA. The product resulting from amplification of a nucleic acid molecule(s) (i.e., "amplification product") can be a mixture of either DNA or RNA, or both DNA and RNA nucleosides or nucleotides, or can contain modified DNA or RNA nucleosides or nucleotides, regardless of whether the starting nucleic acid is DNA, RNA, or both. "Copy" does not necessarily imply perfect sequence complementarity or identity to the target sequence. For example, the copies may include nucleotide analogs such as deoxyinosine or deoxyuridine, intentional sequence changes (e.g., sequence changes introduced via primers containing sequences that are hybridizable to, but not complementary to, the target sequence), and / or sequence errors that occur during amplification. In some embodiments, the first amplification reaction is PCR.
[0117] The amplification protocol may, for example, include one amplification round or multiple amplification rounds. For example, a first amplification round may be followed by a second amplification round, with or without one or more processing steps (e.g., purification, concentration, etc.) between the two amplification rounds. In some embodiments, additional amplification rounds may be used, with or without one or more processing steps between them.
[0118] The number of amplification cycles can vary depending on the embodiment. For example, each amplification round can include at least 4, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 38, or 40 cycles. In embodiments involving more than one amplification round, the number of amplification cycles used in each round can be the same or different. When two amplification rounds are used and the number of amplification cycles in the two rounds is different, the first round can include more or fewer cycles than the second round. As an example, the first amplification round can include 12 cycles, and the second amplification round can include 15 cycles. As another example, the first amplification round can include 10 cycles, and the second amplification round can include 12 cycles.
[0119] The temperature of the amplification reaction varies depending on the step of the reaction. In some embodiments, initial denaturation can be carried out at about 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99°C for about 1, 2, 3, 4, or 5 minutes, and final extension can be carried out at about 67, 68, 69, 70, 71, 72, 73, 74, or 75°C for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes.
[0120] Blocking Oligonucleotides As described above, a first amplification reaction is performed to amplify the rearranged, tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region adjacent to the target site, (ii) a second primer comprising a nucleotide sequence identical to a portion of a sequence tag located at the 5' end of the fragment, and (iii) a blocking oligonucleotide comprising a nucleotide sequence complementary to a region distal to the target site relative to the first target-specific primer.
[0121] As used herein, a blocking oligonucleotide is an engineered single-stranded nucleic acid sequence. A blocking oligonucleotide can be composed of single-stranded DNA, RNA, peptide nucleic acid (PNA), locked nucleic acid (LNA), bridged nucleic acid (BNA), and / or other modified nucleotides. In some embodiments, it is a DNA oligonucleotide containing multiple modified bases. A blocking oligonucleotide can contain 10 to 100 nucleotides, but is preferably between 15 and 50 nucleotides in length.
[0122] The blocking oligonucleotide has strong avidity for the template DNA but cannot be extended by the polymerase used for amplification due to an absent or blocked 3' OH and modifications at or near the 3' end that confer resistance to any 3'>5' exonuclease activity of the polymerase (e.g., multiple phosphorothioate linkages at the 3' end, spacers, inverted nucleotides, LNA or BNA residues). The blocking oligonucleotide must be complementary to the same strand as the first target-specific primer and hybridize distal to the gene editor target site relative to the position of the first target-specific primer. In one embodiment, the blocking oligonucleotide contains an absent or blocked 3' OH, a spacer, an inverted nucleotide, or other modifications to block extension of the 3' end.
[0123] In one embodiment, the blocking oligonucleotide comprises one or more phosphorothioate linkages, spacers, or other modifications at the 3' and 5' ends to block polymerase extension and exonuclease digestion, and a backbone of DNA, RNA, LNA, BNA, PNA, modified nucleic acid, or a combination thereof.
[0124] In some embodiments, the blocking nucleotide comprises a nucleotide that increases duplex stability and melting temperature compared to native DNA. This allows the blocking oligonucleotide to hybridize at a higher temperature than the first and second primers (thus allowing hybridization to the template before the primers can hybridize during amplification). High duplex stability also reduces the likelihood of the blocking oligonucleotide dissociating and polymerase read-through during primer extension.
[0125] In some embodiments, the blocking oligonucleotide comprises a chemical or photocrosslinking moiety that allows efficient covalent binding to the hybridized DNA strand and prevents polymerase extension and dissociation. Exemplary chemical or photocrosslinking moieties include, but are not limited to, psoralens, click chemistry, and 3-cyanovinylcarbazole.
[0126] Sequencing The use of a blocking oligonucleotide, a first target-specific primer, and a second primer in an amplification reaction produces a primary amplification product containing a nucleotide sequence that includes the rearranged target sequence. The primary amplification product is then sequenced.
[0127] As used herein, "sequencing" encompasses any method for determining the sequence of a nucleic acid. Any sequencing method, including chain terminator (Sanger) sequencing and dye terminator sequencing, can be used in this method. In a preferred embodiment, next-generation sequencing (NGS), a high-throughput sequencing technology that performs thousands or millions of sequencing reactions in parallel, is used. While different NGS platforms use various assay chemistries, they all generate sequence data from multiple sequencing reactions run simultaneously on multiple templates. Typically, sequence data is collected using a scanner and then bioinformatically assembled and analyzed. Thus, sequencing reactions are performed, read, assembled, and analyzed in parallel. Exemplary techniques, systems, or technologies can be implemented using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR and droplet digital PCR (ddPCR), quantitative PCR, real-time PCR, multiplex PCR, PCR-based singleplex methods, emulsion PCR), and / or isothermal amplification. Non-limiting examples of nucleic acid sequencing methods include Maxam-Gilbert sequencing and chain termination, de novo sequencing methods including shotgun sequencing and bridge PCR, next generation methods including polony sequencing, 454 pyrosequencing, Illumina sequencing, Ion Torrent semiconductor sequencing, SMRT® sequencing, and Oxford Nanopore Technology sequencing.
[0128] In some embodiments, the primary reaction products can be purified to remove primers and quantified. In some embodiments, the sample can be split, and symmetric primer sets are used to evaluate both sides of the target site for rearrangements in separate reactions (this step and subsequent steps). The primary reaction products can also be used to detect rearrangements at other genomic loci, such as candidate or true off-target sites predicted by in silico algorithms, or detected by other experimental methods.
[0129] Additional Amplification As mentioned above, the first amplification round can be followed by a second amplification round.Therefore, in some embodiments, the second amplification comprises nested PCR and the introduction of second and third sequence tags.In some embodiments, the second amplification comprises hemi-nested PCR.
[0130] In embodiments in which a second amplification round is used, the methods disclosed herein can further include, prior to step (c), performing a second amplification reaction using third and fourth nested primers. The third and fourth primers contain nucleotide sequences complementary to the sequences proximal to the inner ends of the amplification products of (b) and fourth and fifth sequence tags at the 5' ends of the nested primers. The second amplification reaction produces secondary amplification products containing the rearranged target sequence and additional sequence tags. This nested amplification step enriches for the desired product because truncated extension products and spurious fragments resulting from mispriming in the first amplification step cannot serve as amplification templates.
[0131] In some embodiments in which a second amplification round is used, the methods disclosed herein can further include, prior to (c), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (b), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce a secondary amplification product comprising the rearranged target sequence and one or two additional sequence tags.
[0132] For the second amplification step, the amplification protocol can include, for example, one amplification round or multiple amplification rounds. For example, the first amplification round can be followed by a second amplification round, with or without one or more processing steps (e.g., purification, concentration, etc.) between the two amplification rounds. In some embodiments, additional amplification rounds can be used, with or without one or more processing steps between them.
[0133] The number of amplification cycles can vary depending on the embodiment. For example, each amplification round can include at least 4, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 38, or 40 cycles. In embodiments involving more than one amplification round, the number of amplification cycles used in each round can be the same or different. When two amplification rounds are used and the number of amplification cycles in the two rounds is different, the first round can include more or fewer cycles than the second round. As an example, the first amplification round can include 12 cycles, and the second amplification round can include 15 cycles. As another example, the first amplification round can include 10 cycles, and the second amplification round can include 12 cycles.
[0134] The temperature of the amplification reaction varies depending on the step of the reaction. In some embodiments, initial denaturation can be carried out at about 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99°C for about 1, 2, 3, 4, or 5 minutes, and final extension can be carried out at about 67, 68, 69, 70, 71, 72, 73, 74, or 75°C for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes.
[0135] In one embodiment, the method further comprises performing a third amplification reaction using fifth and sixth primers, or hemi-nested PCR followed by the fourth and fifth primers, that comprise nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products that contain the rearranged target sequence. In some embodiments, the third amplification reaction is hemi-nested PCR.
[0136] For the third amplification step, the amplification protocol can include, for example, one amplification round or multiple amplification rounds. For example, the first amplification round can be followed by a second amplification round, with or without one or more processing steps (e.g., purification, concentration, etc.) between the two amplification rounds. In some embodiments, additional amplification rounds can be used, with or without one or more processing steps between them.
[0137] The number of amplification cycles can vary depending on the embodiment. For example, each amplification round can include at least 4, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 38, or 40 cycles. In embodiments involving more than one amplification round, the number of amplification cycles used in each round can be the same or different. When two amplification rounds are used and the number of amplification cycles in the two rounds is different, the first round can include more or fewer cycles than the second round. As an example, the first amplification round can include 12 cycles, and the second amplification round can include 15 cycles. As another example, the first amplification round can include 10 cycles, and the second amplification round can include 12 cycles.
[0138] The temperature of the amplification reaction varies depending on the step of the reaction. In some embodiments, initial denaturation can be carried out at about 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99°C for about 1, 2, 3, 4, or 5 minutes, and final extension can be carried out at about 67, 68, 69, 70, 71, 72, 73, 74, or 75°C for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes.
[0139] In one embodiment, the third and / or fourth primer comprises a barcode sequence.
[0140] Barcode sequence A "barcode" is a molecular label or identifier that conveys or can convey information about a sample, a sequencing read, or a group of samples or sequencing reads (e.g., information about the analytes, templates, beads, microwells, primers, reads, and / or capture probes in the sample). A barcode can be part of an analyte or can be independent of the analyte. A barcode can be attached to an analyte. A particular barcode is unique relative to other barcodes and can include error correction features (e.g., Hamming codes) that allow for correct identification of each barcode in a mixture even in the presence of sequencing errors.
[0141] As used herein, a barcode comprises a custom polynucleotide of defined sequence introduced by primer extension (e.g., during PCR) or adapter ligation. Barcodes can allow for identification and / or quantification of groups of sequencing reads that share the same barcode. Barcodes can be 8-50 or more bases long, but are typically between 10-25 bases long.
[0142] Barcodes can be used to spatially resolve molecular components found in a biological sample, for example, at single-cell resolution (e.g., a barcode can be or can encompass a "spatial barcode"). In some embodiments, a barcode encompasses two or more sub-barcodes that can function together as a single barcode. For example, a polynucleotide barcode can encompass two or more polynucleotide sequences (e.g., sub-barcodes) separated by one or more non-barcode sequences.
[0143] Index sequences are a class of barcodes used by standard sequencing analysis software to distinguish and group sets of reads originating from different samples after sequencing.
[0144] Unique Molecular Identifiers Unique molecular identifiers, or "UMIs," comprise short stretches of random nucleotides (e.g., NNNNNNNN (SEQ ID NO: 90)) or semi-random nucleotides (e.g., NWYYRRV (SEQ ID NO: 91)) incorporated into custom primers or adapters. UMIs are typically between 6 and 25 bases long and are commonly introduced by tagmentation or adapter ligation, but can also be introduced by a single round of primer extension. UMIs serve to enable the identification and grouping of extension or amplification products (e.g., PCR products) derived from a single starting genomic DNA molecule. Such grouping allows for the counting of unique molecules in the starting sample represented by multiple reads after amplification and sequencing (e.g., counting unique recombination events). UMIs can also be used to distinguish between SNVs and sequencing errors in amplified molecules.
[0145] An individual sequencing read may contain multiple barcodes, indexes, and UMIs in addition to the sequence being characterized.
[0146] In one embodiment, the fifth and sixth primers comprise a sequencing tag and / or index sequence.
[0147] In one embodiment, (b) further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
[0148] An embodiment of the method described herein is illustrated in Figures 1 and 6 (Method IA). In some embodiments, the sequence tag comprises a non-interacting sequence distinct from the genome, a UMI, and a transposase recognition site (Figure 1). In some embodiments, the sequence tag comprises a non-interacting sequence distinct from the genome, a UMI, an index sequence, and a transposase recognition site (Figure 6). As shown in Figures 1 and 6, after tagmentation, the sample can be subjected to a three-step PCR amplification procedure to enrich for recombined fragments (steps 2-4). Each of these steps is performed in the presence of blocking oligonucleotides, which prevent products from native (unrearranged) sequences from undergoing exponential amplification. The amplification procedure may include a three-step or touchdown PCR cycling protocol to ensure that the blocking oligonucleotides hybridize before the PCR primers are extended.
[0149] In the first PCR step (step 2), one DNA primer (labeled "1" in Figures 1 and 6) matches the 5' transposon end tag (e.g., sequencing tag), and the other (labeled "2" in Figures 1 and 6) is complementary to a region adjacent to the gene editor target site being investigated as a possible site of rearrangement. In the presence of a blocker (labeled "0"), the DNA polymerase extension product derived from the target-specific primer (2) at the native gene editor target site (right, bottom of Figures 1 and 6) is blocked from extension. This prevents exponential PCR amplification from occurring (note that if the optional 3' end blocking step is included, the 3' end of the tagmented fragment cannot be extended). If a rearrangement occurs such that the sequence complementary to the blocking oligonucleotide is replaced by another sequence derived from elsewhere in the genome, the blocker cannot hybridize, and PCR products derived from the rearranged fragment are productively generated (left side of Figures 1 and 6). After amplification, the PCR product is purified to remove primers and quantified. Typically, the sample is split and symmetric primer sets are used to evaluate both sides of the target site for rearrangements in separate reactions (this step and subsequent steps). Further aliquots of the PCR product can be used to detect rearrangements at other genomic loci, such as off-target sites predicted by in silico algorithms, or detected by other experimental methods.
[0150] The second PCR step (step 3) uses a pair of nested primers (labeled "3" and "4" in Figure 1) complementary to sequences proximal to the inner ends of the PCR product generated in step 2. This nested PCR step enriches for the desired product because truncated extension products and spurious fragments resulting from mispriming at other loci in the first PCR step cannot serve as amplification templates. The nested primers contain sequence tags that can be used as priming sites for the third PCR step (step 4) and may also contain barcodes, allowing for multiplexing of products from different target sites in separate reactions or different samples. Alternatively, if the first sequencing tag encompasses an index primer (Figure 6), the second PCR step involves a hemi-nested PCR step using a pair of primers (labeled "3" and "4" in Figure 6) in which primer "4" is complementary to a sequence proximal to the inner end of the PCR product generated in step 2 and primer "3" matches the 5' transposon end tag (e.g., sequencing tag). Again, in the presence of a blocking oligonucleotide (Blocker 0), only PCR products derived from the rearranged target molecule can be further amplified with nested or hemi-nested primers. Fragments derived from the native, unrearranged target site cannot be amplified (only a single truncated extension product can be generated in each PCR cycle). After amplification, the PCR product is purified to remove primers and quantified.
[0151] In the third PCR step (step 4), primers (labeled "5" and "6") complementary to the tags introduced by PCR in the second PCR step (step 3) are used to further enrich for the rearranged molecules. These primers also provide sequences required for NGS sequencing, such as adapter sequences for cluster generation on an Illumina flow cell, UMIs, and index sequences for sample multiplexing. After amplification, the PCR products are purified to remove primers and quantified. The resulting library is then subjected to NGS sequencing, for example, on an Illumina platform.
[0152] Variation of Method I In one embodiment, the methods disclosed herein further include, prior to (b), contacting the plurality of tagmented nucleic acid fragments with ddNTPs or other 3'-modified dNTPs and a DNA polymerase (e.g., Klenow fragment or terminal deoxynucleotidyl transferase) to block all extendable 3' ends immediately after tagmentation.
[0153] The tagmentation and amplification reactions can be carried out as described above for Method I.
[0154] An exemplary embodiment of the method described herein is shown in Figure 2 (Method IB). As shown in Figures 2 and 7, after tagmentation, the 3' ends of the short ME fragments and the 3' ends of the 9-bp gap genomic DNA are extended and blocked by DNA polymerase or ddNTPs using terminal deoxynucleotidyl transferase (TdT). This prevents any possible extension of the ME or genomic fragment ends to generate priming sites for primers 1 and 3 during PCR, as well as preventing end extension by hybridization with repeat sequences or short sequences that share homology with the 3' fragment ends. Other steps are the same as those described above for Figure 1.
[0155] In one embodiment, the sequence tag in the above method comprises uracil.
[0156] According to this embodiment, an exemplary embodiment of the method described herein is shown in Figures 3 and 8 (Method IC). As shown in Figures 3 and 8, the sequence tags used to assemble transposomes carry uracil. Unlike other current protocols, there is no strand displacement step after tagmentation. Instead, the DNA is purified, and then a PCR reaction is set up using Primer 2 and a uracil-resistant DNA polymerase (e.g., Phusion U Hot Start DNA Polymerase, catalog number F555S, Thermo Fisher Scientific). Extension from Primer 2 using unrearranged DNA as a template is blocked by Blocker 0, but Blocker 0 cannot block extension using rearranged DNA as a template. This Primer 2 extension step can be performed multiple times to generate multiple copies of the Primer 2-extended strand. Alternatively, isothermal amplification methods, such as recombinase polymerase amplification (RPA) and strand invasion-based amplification (SIBA), can be used to produce more copies. The purified DNA is then treated with USER (Uracil-Specific Excision Reagent) enzyme (a mixture of uracil DNA glycosylase (UDG) and DNA glycosylase-lysase endonuclease VIII) at the uracil residues synthetically incorporated into the adapters. In some embodiments, USER enzyme cleavage can be performed by mixing 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 μL of USER enzyme with the product from the previous step and incubating at 16, 20, 25, 30, 35, 40, 45, 50, 55, or 60 minutes at 16, 20, 25, 30, 35, 40, 45, 50, 55, or 60 minutes. Subsequent steps in Figures 3 and 8 are similar to those in Figure 1 from step 2 onwards.
[0157] Method II The present disclosure also encompasses another method for detecting genome-wide rearrangements in a nucleic acid genome. The method includes (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons containing a first sequence tag at the 5' end of the transposon under conditions in which the plurality of transposons are inserted into the genomic DNA sample, and the genomic DNA sample is tagged into a plurality of nucleic acid fragments, each containing a sequence tag at the 5' end of a partially double-stranded nucleic acid fragment. In some embodiments, the sequence tag comprises a non-interacting sequence distinct from the genome, a UMI, and a transposase recognition site (Figure 5). In some embodiments, the sequence tag comprises a non-interacting sequence distinct from the genome, a UMI, an index sequence, and a transposase recognition site (Figure 9). The method then includes (b) contacting the plurality of tagmented nucleic acid fragments with a sequence-specific cleavage reagent; (c) performing a first amplification reaction to amplify the rearranged tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region flanking the target site, and (ii) a second primer comprising a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment; and (d) sequencing the primary amplification products.
[0158] Methods and embodiments of step (a) are described above for Method I. However, unlike the above methods utilizing blocking oligonucleotides, methods according to this aspect of the disclosure involve contacting a plurality of tagmented nucleic acid fragments with a sequence-specific cleavage reagent.
[0159] Sequence-specific cleavage agents In one embodiment, the sequence-specific cleavage reagent is an enzymatic reagent or a chemical cleavage agent.
[0160] In the methods disclosed herein, the step of contacting with the nuclease is generally carried out under conditions favorable for cleavage by the nuclease. That is, even if a given candidate target site or variant target site may not actually be cleaved by the nuclease, the incubation conditions are such that the nuclease will cleave at least a significant portion (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%) of the template containing the known target site. For known and generally well-characterized nucleases, such conditions are generally known in the art and / or can be easily discovered or optimized. For newly discovered nucleases, such conditions can generally be approximated using information on related nucleases that are better characterized (e.g., homologs and orthologs).
[0161] In some embodiments, the nuclease is an endonuclease. In some embodiments, the nuclease is a site-specific endonuclease (e.g., a restriction endonuclease, a meganuclease, a transcription activator-like effector nuclease (TALEN), a zinc finger nuclease, etc.).
[0162] In some embodiments, the site specificity of site-specific nuclease is conferred by auxiliary molecule.For example, CRISPR-associated (Cas) nuclease is guided to specific site by " guide RNA " or gRNA as described herein.In some embodiments, nuclease is RNA-guided nuclease.In some embodiments, nuclease is CRISPR-associated nuclease.
[0163] In some embodiments, the nuclease is a homolog or ortholog of a previously known nuclease, eg, a newly discovered homolog or ortholog.
[0164] RNA-guided nucleases according to the present disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases such as Cas9 and Cas12a, as well as other nucleases derived or obtained therefrom. In functional terms, RNA-guided nucleases are defined as nucleases that (a) interact with (e.g., complex with) gRNA; and (b) together with gRNA, associate with and optionally cleave or modify a target region of DNA that includes (i) a sequence complementary to the targeting domain of gRNA, and optionally (ii) an additional sequence called a "protospacer adjacent motif" or "PAM," which will be described in more detail below.
[0165] In one embodiment, the enzymatic cleavage agent comprises a CRISPR Cas / gRNA complex.
[0166] In one embodiment, a guide RNA homologous to a blocking oligonucleotide locus can be used in conjunction with a Cas nuclease (e.g., SpCas9) to cleave tagmented genomic DNA, selectively preventing amplification of the native, unrearranged fragment and allowing the rearranged locus (uncleaved) to be amplified in exactly the same way as the blocking oligonucleotide procedure.
[0167] In one embodiment, the sequence-specific cleavage reagent is a chemical cleavage agent.Chemical cleavage can include any method that utilizes non-nucleic acid and non-enzymatic chemical reagents to facilitate / achieve the cleavage of one or both strands of a double-stranded nucleic acid molecule.Optionally, one or both strands of a double-stranded nucleic acid molecule can include one or more non-nucleotide chemical moieties and / or non-natural nucleotides and / or non-natural backbone bonds to enable chemical cleavage reaction.
[0168] In one embodiment, a chemical cleavage agent can be linked to the oligonucleotide. Exemplary chemical cleavage agents linked to oligonucleotides that can be useful in the methods herein include, but are not limited to, chelated metal ions (e.g., Ag, Cu, Fe).
[0169] In one embodiment, the sequence tag comprises a sequence that is separate from the genome and does not interact with it, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof, as described above for Method I.
[0170] In one embodiment, the first and second primers comprise second and third sequence tags, as described above for Method I.
[0171] As described above for Method I, the first amplification round can be followed by a second amplification round. Thus, in some embodiments, the second amplification comprises nested PCR and the introduction of second and third sequence tags. In some embodiments, the second amplification comprises hemi-nested PCR.
[0172] In one embodiment, when a second amplification round is used, the method disclosed herein can further include, prior to (d), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (c) and additional sequence tags at the 5' ends of the nested primers to produce secondary amplification products comprising the rearranged target sequence and the additional sequence tags. In one embodiment, the method further includes performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0173] In one embodiment in which a second amplification round is used, the method disclosed herein can further comprise, prior to (d), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (c), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce secondary amplification products comprising the rearranged target sequence and one or two additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction, to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0174] In one embodiment, the products of the first amplification reaction are subjected to repeated cleavage reactions before proceeding with the second and / or third amplification reactions, respectively.
[0175] In one embodiment, the third and / or fourth primer includes a barcode sequence as described above for Method I.
[0176] In one embodiment, the fifth and sixth primers include a sequencing tag and / or index sequence as described above for Method I.
[0177] In one embodiment, the method further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
[0178] Exemplary embodiments of the above methods described herein are shown in Figures 4 and 9. As shown in Figures 4 and 9, instead of using a blocking oligo as in Method I, Method II uses sequence-specific cleavage using CRISPR / Cas, TALEN, ZFN, or similar sequence-directed enzymatic cleavage reagents. In some embodiments, the cleavage reaction can be carried out by mixing the guide RNA with Cas9 nuclease at 15, 20, 25, 30, 35, 40, or 45°C for 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120 minutes.
[0179] Another alternative is the use of a chemical cleavage reagent linked to an oligonucleotide. This selectively prevents amplification of the native, unrearranged fragment, while allowing the rearranged locus (uncleaved) to be amplified in exactly the same way as the blocking oligo procedure. After the cleavage reaction is complete, all 3' ends are blocked by treatment with ddNTPs and terminal deoxynucleotidyl transferase (TdT), Klenow exo-, or another suitable enzyme (TdT has the advantage of being able to incorporate 3'-blocked nucleotides at staggered or blunt ends). End blocking reduces the generation of interfering by-products in subsequent PCR reactions.
[0180] Method III The present disclosure also relates to a method for detecting genome-wide rearrangements in a nucleic acid genome, comprising: (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons, each comprising a first sequence tag at the 5' end of the transposon, wherein the sequence tag comprises an RNA promoter sequence, under conditions such that the transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagged into a plurality of nucleic acid fragments, each comprising a sequence tag at the 5' end of the partially double-stranded nucleic acid fragment; (b) transcribing the plurality of tagmented nucleic acid fragments into RNA; (c) contacting the RNA with a target sequence-specific DNA oligonucleotide probe and RNase H; (d) performing a reverse transcription amplification reaction using (i) a first primer comprising a nucleotide sequence complementary to a portion of the sequence tag located at the 3' or 5' end of the fragment and (ii) a second primer comprising a nucleotide sequence complementary to a region adjacent to the target site to produce a primary amplification product; and (e) sequencing the primary amplification product.
[0181] RNA promoter sequence RNA promoter sequences useful in the methods described herein are sequences capable of binding RNA polymerase and containing a transcription initiation site. Thus, promoter sequences typically encompass between about 15 and about 250 nucleotides, preferably between about 25 and about 60 nucleotides, from a naturally occurring RNA polymerase promoter, a consensus promoter sequence (Alberts et al., in Molecular Biology of the Cell, 2nd Ed., Garland, NY (1989)), or a modified version thereof. Exemplary RNA promoters are the T3, T7, and SP6 phage promoter / polymerase systems.
[0182] Unlike previous methods described herein, this aspect of the disclosure includes the steps of (b) transcribing the plurality of tagmented nucleic acid fragments into RNA; (c) contacting the RNA with a target sequence-specific DNA oligonucleotide probe and RNase H; and (d) performing a reverse transcription amplification reaction. The steps before and after this method are similar to those described for the other methods above.
[0183] In one embodiment, the sequence tag comprises a sequence that is separate from and does not interact with the genome, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof, as described above.
[0184] In one embodiment, the first and second primers comprise sequence tags as described above.
[0185] An exemplary embodiment of the method described herein is shown in Figure 5. As shown in Figure 5, the specific endoribonuclease activity of RNase H (ribonuclease H) specifically hydrolyzes phosphodiester bonds in RNA when hybridized to DNA. In this method, a T7 promoter (or other in vitro transcription promoter) is incorporated into the transposon adapter. After tagmentation, the DNA is transcribed into RNA in vitro. In some embodiments, the in vitro transcription reaction can be performed using the MAXIscript™ T7 Transcription Kit by incubating at 25, 30, 35, 40, or 45°C for 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, or 600 minutes.
[0186] A sequence-specific DNA probe is then hybridized to the RNA and treated with RNase H, resulting in specific cleavage of the RNA region hybridized with the DNA probe. Next, one-step reverse transcription PCR is performed using sequence-specific primer 2. In some embodiments, primer 2 is designed specifically for the sequence to the right of the potential rearrangement site, and a DNA probe is designed specifically for the sequence to the left of the same site. Thus, unrearranged templates are cleaved by RNase H and cannot be amplified by primers 1 and 2, while rearranged templates are not cleaved by the DNA probe and can be exponentially amplified. In some embodiments, for DNA probe hybridization and RNase H treatment, two separate reactions are performed for each target: one reaction is for detecting genomic rearrangements on the left side of the target, and the other reaction is for detecting genomic rearrangements on the right side of the target. In some embodiments, the reaction is denatured at about 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99° C. for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, RNase H is added to the reaction and incubated at about 16, 20, 25, 30, 35, 40, or 45° C. for about 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, or 300 minutes. In some embodiments, RNase H is inactivated for about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, or 240 minutes at 60, 65, 70, 75, 80, 85, or 90° C. After first strand synthesis, the following steps are the same as in FIG. 4 from step 3 onwards.
[0187] In one embodiment, the method can further comprise, prior to (e), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (d) and an additional sequence tag to produce secondary amplification products comprising the rearranged target sequence and the additional sequence tag. In one embodiment, the method further comprises performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0188] In one embodiment, the method can further comprise, prior to (e), performing a second hemi-nested amplification reaction using a second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (d), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce secondary amplification products comprising the rearranged target sequence and one or two additional sequence tags. In one embodiment, the method further comprises performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction, to produce further enriched tertiary amplification products comprising the rearranged target sequence.
[0189] In one embodiment, the third and / or fourth primers include a barcode sequence as described above.
[0190] In one embodiment, the fifth and sixth primers comprise an adapter sequence and / or an index sequence as described above.
[0191] In one embodiment, the method includes where (d) further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
[0192] Data analysis For the methods described herein, data analysis can be performed using paired-end sequencing reads from Illumina sequencing that are merged using paired-end read merging software such as PEAR (Zhang et al, Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al, PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). Merged reads can be trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads can be aligned to a human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie 2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100).Aligned BAM files can be converted to BED files using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34). For methods using UMIs, software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer ( / github.com / smangul1 / UMI-Reducer), or custom scripts can be used to collapse UMIs and remove redundant sequencing reads.
[0193] To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site can be identified and counted. Statistical tests can be applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) can be used to facilitate the identification and analysis of rearrangements on a genome-wide scale.
[0194] Systems and Kits Also provided herein are systems and kits that include the necessary reagents to practice the methods described herein, as well as written instructions for making and using the same.
[0195] Any of the above systems and kits can further include one or more additional reagents.
[0196] In some embodiments, the system or kit can further include instructions for practicing the method using the components of the kit. The instructions for practicing the method are generally recorded on a suitable recording medium. For example, the instructions can be printed on a substrate such as paper or plastic. The instructions can be present in the kit as a package insert, on a label on the kit's container or a component thereof (i.e., associated with the package or subpackage), etc. The instructions can be present as an electronic storage data file present on a suitable computer-readable storage medium, such as a CD-ROM, diskette, flash drive, etc. In some cases, the actual instructions are not present in the kit, but a means for obtaining the instructions from a remote source (e.g., via the Internet) can be provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions can be recorded on a suitable substrate. [Example]
[0197] Example While certain alternative forms of the present disclosure have been disclosed, it is to be understood that various modifications and combinations are possible and are contemplated within the true spirit and scope of the appended claims. Accordingly, there is no intention to be limited to the precise summary and disclosure presented herein.
[0198] Example 1: Exemplary Method of Method IA Exemplary oligonucleotides for use in the present methods are listed in Table 1. Table 1 - Exemplary oligonucleotide sequences. [Table 1-1] [Table 1-2] [Table 1-3]
[0199] Exemplary oligonucleotide sets are shown in Tables 2 and 3 below. For example, for EMX1 or CCR5#1, one reaction is performed to detect rearrangements to the left of the on-target cleavage site, and a second reaction is performed to detect rearrangements to the right of the on-target cleavage site. For each reaction, a target-specific primer and blocker set should be used. For the left rearrangement of EMX1, oligos 03 / 05 / 07 / 13 are used, and for the right rearrangement of EMX1, oligos 14 / 15 / 16 / 16 are used. Similarly, for the left rearrangement of CCR5#1, oligos 18 / 19 / 20 / 21 are used, and for the right rearrangement of CCR5#1, oligos 22 / 23 / 24 / 25 are used. Table 2. Target-specific primers and blockers [Table 2] Table 3. Transposome oligonucleotides [Table 3]
[0200] Exemplary sgRNA target sequences are shown in Table 4 below. Table 4 [Table 4]
[0201] Additional reagents include Tn5 transposase, e.g., Robust Tn5 transposase (catalog number EMQZ1422, Creative Biogene), DNA polymerase PCR master mix, e.g., Platinum™ SuperFi II PCR Master Mix (ThermoFisher, catalog number 12368010), DNA polymerase with PCR buffer, e.g., Platinum™ SuperFi II DNA polymerase (ThermoFisher, catalog number 12361010), dNTPs, e.g., dNTP mix (10 mM each) (ThermoFisher, catalog number R0194), DNA purification kit: column kit, e.g., DNA Clean&Concentrator-5 (Zymo Research, catalog number D4004), magnetic beads, e.g., AMPure XP reagent (Beckman Coulter, catalog number A63880), Illumina sequencing kit, e.g., NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Cat. No. 20050264), MiSeq Reagent Kit v3 (150 cycles) (Cat. No. MS-102-3001), MiSeq Reagent Kit v2 (300 cycles) (Cat. No. MS-102-2002), PhiX Control v3 (Illumina, Cat. No. FC-110-3001), and STE buffer (10 mM Tris-HCl (pH 8.0), 1 mM EDTA, 0.1 M NaCl).
[0202] In the first step, the Tn5 transposome is assembled. Transposon end-containing oligos (e.g., oligo-01 and oligo-02 for the non-UMI version, or oligo-26 and oligo-02 for the UMI version) are annealed in STE buffer using the following program: 95–99°C for 1–10 min, followed by 1–99°C for 1 min for 40–500 cycles (decrease the temperature by 0.2–2°C per cycle until a temperature of 1–12°C is reached). The transposome complex is assembled in TPS buffer according to the following procedure, as recommended by the manufacturer (Creative Biogene, catalog number EMQZ1422). Mix the reagents thoroughly and incubate at 15–40°C for 5–120 min. Component Volume Tn5 transposase 1-10 μL Adapter 0.5~8μL 2μL of 10x TPS buffer Add sterilized water up to 20 μL
[0203] The workflow for this example is shown in Figure 1. A tagmentation reaction is performed in which genomic DNA extracted from the gene-edited sample or control (using standard methods) is tagmented with transposome complexes assembled according to the procedure. An exemplary reaction mixture is shown below: Component Volume Genomic DNA variation (5-500ng) 6μL of 5×LM buffer Transposome 0.1~8μL Add sterilized water up to 30 μL
[0204] Mix the reaction components thoroughly and incubate for 5-120 min at 30-65 °C. The tagmented DNA is purified using column-based purification.
[0205] The tagmented DNA from the previous step was used in two separate PCR reactions for each target: one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target. Each PCR was then performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR was performed using a three-step PCR cycling protocol with tagmented DNA, primer 1 (oligo-04, common to all targets and rearrangements on both sides), primer 2 (e.g., oligo-05 for EMX1 left rearrangement, oligo-15 for EMX1 right rearrangement, oligo-19 for CCR5#1 left rearrangement, or oligo-23 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement). The PCR was performed using II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the following program: initial denaturation at 90-99°C for 1-5 min, 20-40 cycles of four-step denaturation-annealing-amplification (90-99°C for 5-60 s, 75-92°C for 5-60 s, 55-75°C for 5-60 s, 67-75°C for 5-60 s), and a final extension at 67-75°C for 1-10 min and a hold time of 4-12°C. PCR products were purified using column-based purification.
[0206] Next, perform nested PCR using the barcode tags. PCR is performed with the PCR product from the previous step, primer 3 (oligo-06, common to all targets and both sides of the rearrangement), primer 4 (e.g., oligo-07 for EMX1 left rearrangement, oligo-16 for EMX1 right rearrangement, oligo-20 for CCR5#1 left rearrangement, or oligo-24 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0207] Next, tag PCR is performed using the sequencing primers. PCR is performed with the PCR product from the previous step, primer 5 (e.g., oligo-08, an example of an Illumina index primer from New England Biolabs, NEB#E7603A), primer 6 (e.g., oligo-09, an example of an Illumina index primer from New England Biolabs, NEB#E7611A), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi PCR. II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0208] Next, perform next-generation sequencing using the PCR products from the previous step, with optional ≥0–95% PhiX (Illumina, catalog number FC-110-3001) for sequencing using an Illumina platform (e.g., Illumina MiSeq or NextSeq 2000).
[0209] Paired-end sequencing reads from Illumina sequencing are merged using paired-end read merging software such as PEAR (Zhang et al., Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al., PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). The merged reads are then trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads are aligned to the human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100). The aligned BAM file is converted to a BED file using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34).For methods using UMIs, collapse UMIs to remove redundant sequencing reads using software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer (github.com / smangul1 / UMI-Reducer), or custom scripts. To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site are identified and counted. Statistical tests are then applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) facilitate genome-wide identification and analysis of rearrangements.
[0210] Example 2: Exemplary Method of Method Ia Exemplary oligonucleotides for use in the present methods are listed in Table 5. Table 5 - Exemplary oligonucleotide sequences. [Table 5-1] [Table 5-2] [Table 5-3]
[0211] Exemplary oligonucleotide sets used are shown in Tables 6 and 7 below. For example, for CCR5#1, one reaction was performed to detect rearrangements to the left of the on-target cleavage site, and a second reaction was performed to detect rearrangements to the right of the on-target cleavage site. For each reaction, a target-specific primer and blocker set was used. Oligos 04A / 05A / 19 / 20 were used for the left side of CCR5#1, and oligos 04A / 09A / 10A / 23 / 24 were used for the right side of CCR5#1. Table 6. Target-specific primers and blockers [Table 6] Table 7. Transposome oligonucleotides [Table 7]
[0212] Exemplary sgRNA target sequences used are shown in Table 8 below. Table 8 [Table 8]
[0213] Additional reagents included unloaded Tn5 transposase (Diagenode, C01070010, Diagenode), 2X Tagmentation buffer (Diagenode, C01019043), Platinum™ SuperFi II PCR Master Mix (ThermoFisher, Cat. No. 12368010), dNTP mix (10 mM each) (ThermoFisher, Cat. No. R0194), DNA purification kit: DNA Clean & Concentrator-5 (Zymo Research, Cat. No. D4004), AMPure XP Reagent (Beckman Coulter, Cat. No. A63880), glycerol (MilliporeSigma, 356350500ML), 10% SDS (Fisher, BP2436200), 5 M NaCl (Invitrogen, AM9760G), Tris HCl 1M, pH 8.0 (Invitrogen, 15568-025), genomic DNA (extracted from edited and unedited 293T cells), NEBNext library Quant Kit for Illumina (NEB, E7630), Illumina NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Illumina, catalog number 20050264), and PhiX Control v3 (Illumina, catalog number FC-110-3001).
[0214] Prior to library preparation, we tested the ability of LNA blockers to efficiently block amplification of unintended targets, including wild-type and small indel edits. We incorporated LNA blockers (Oligo-05-A and Oligo-09-A) into PCR to amplify amplicons using wild-type genomic DNA. The PCR mixture was prepared as follows: 50 ng of gDNA, 25 μL of 2x PCR master mix, 2.5 μL of Primer F (Oligo-19, 10 μM), 2.5 μL of Primer R (Oligo-23, 10 μM), different concentrations of blockers, and ddH2O added up to 50 μL. PCR was performed using the following program: initial denaturation at 98°C for 30 seconds, 35 cycles of a four-step denaturation-annealing-amplification (98°C for 10 seconds, 80°C for 10 seconds, 60°C for 10 seconds, and 72°C for 5 seconds), and a final extension at 72°C for 5 minutes, followed by a 4°C hold. PCR products were directly profiled on a 2% EX agarose gel. As shown in Figure 10A, when the concentration of the blocker reached 50% of that of the PCR primers, both were able to efficiently block amplicon amplification. To further optimize the blocking effect of the blocker oligo, different temperatures were tested, as shown in Figure 10B. All tested temperatures between 77°C and 82.4°C achieved complete elimination of PCR products. 78°C was used for the four-step denaturation-annealing-amplification (98°C for 10 seconds, 78°C for 10 seconds, 60°C for 10 seconds, and 72°C for 5 seconds) for library preparation.
[0215] In the first step, transposon-end-containing oligos (e.g., oligo-01A-501 and oligo-03A, or oligo-01A-502 and oligo-03A) were resuspended in annealing buffer (40 mM Tris-HCl (pH 8.0), 50 mM NaCl) to a stock concentration of 100 μM. In a PCR tube, 10 μl of oligo-01 and 10 μl of oligo-03 were mixed, vortexed, and placed in a thermocycler with the following program: 95°C for 5 min, cool to 65°C (-0.1°C / s), 65°C for 5 min, cool to 25°C (-0.1°C / s), 25°C for 5 min, and hold at 4°C.
[0216] Transposome complexes were assembled according to the following procedure, as recommended by the manufacturer. Reagents were mixed in a PCR tube: 10 μL Tn5 transposase (2 μg / ul), 10 μL annealed adapters. The reagents were mixed thoroughly and incubated at 23°C for 30 minutes. 10 μL glycerol was added and mixed. The assembled transposome complexes were stored at -20°C.
[0217] The workflow for this example is shown in Figure 6. A tagmentation reaction was performed, and genomic DNA extracted from gene-edited samples or controls (using standard methods) was tagmented with the assembled transposome complex using the following procedure. The reaction mixture contained 100 ng of genomic DNA (unedited control or edited sample), 20 μL of 2x tagmentation buffer, 200 ng of Tn5 transposase, and HO added up to 40 μL. The reaction mixture was mixed thoroughly and incubated at 55°C for 15 minutes. 10 μL of 0.2% SDS was added. Tn5 was then inactivated at 70°C for 10 minutes. The tagmented DNA was purified using a Zymo column according to the manufacturer's instructions and eluted in 21 μL.
[0218] The tagmented DNA from the previous step was used in two separate PCR reactions for each target (one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target). Each PCR was then performed using a common primer annealed to the Tn5 adapter region, a target-specific blocker, and a target-specific primer. SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) was used with primer 1 (oligo-04A, common to all targets and both rearrangements), primer 2 (e.g., oligo-19 for the left rearrangement of CCR5#1 or oligo-23 for the right rearrangement of CCR5#1), and blocker 0 (e.g., oligo-05A for the left rearrangement of CCR5#1 or oligo-09A for the right rearrangement of CCR5#1). The PCR mixture was prepared as follows: 25 μL of 2x PCR master mix, 10 μL of tagmented DNA, and 1.25 μL of primer 1 (10 μM), 2.5 μL of primer 2 (10 μM), 10 μL of blocker 0 (10 μM), and HO (added to 50 μL). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of a four-step denaturation-annealing-amplification sequence (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, and 72°C for 60 s), followed by a final extension at 72°C for 5 min and a 4°C hold. The PCR product was purified using AMpure XP beads (1x) and eluted in 12 μL.
[0219] Nested PCR was then performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog no. 12368010) with the PCR product from the previous step, primer 3 (oligo-04A, common to all targets and both rearrangements), primer 4 (e.g., oligo-20 for the CCR5#1 left-sided rearrangement, or oligo-24 for the CCR5#1 right-sided rearrangement), and blocker 0 (e.g., oligo-05A for the CCR5#1 left-sided rearrangement, or oligo-09A for the CCR5#1 right-sided rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 1.25 μL of primer 3 (10 μM), 2.5 μL of primer 4 (10 μM), 10 μL of blocker 0 (10 μM), and HO up to 50 μL. PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 12 μL.
[0220] Next, tag PCR was performed using the sequencing primers. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the PCR product from the previous step, primer 5 (oligo-04A, common to all target and bilateral rearrangements), primer 6 (e.g., any oligo-13A or Illumina index primer from New England Biolabs, NEB#E7611A), and blocker 0 (e.g., oligo-05A for CCR5#1 left-sided rearrangement or oligo-09A for CCR5#1 right-sided rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 5 μL of primer 5 (10 μM), 5 μL of primer 6 (10 μM), and 5 μL of blocker 0 (20 μM). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 52 μL.
[0221] 50 μL of PCR product was transferred to a new tube for library quantification using the NEBNext Library Quantitation Kit for Illumina (NEB, E7630) according to the manufacturer's instructions. Libraries were normalized and pooled for sequencing.
[0222] Next-generation sequencing was then performed using the PCR products from the previous step, and sequencing was performed using an Illumina NextSeq 2000 with 20% PhiX (Illumina, catalog no. FC-110-3001).
[0223] Paired-end sequencing reads from Illumina sequencing were processed using a customized NGS analysis pipeline. Sequence reads were first checked for overall quality using FastQC, followed by integration of FastQC results using MultiQC. Reads were then trimmed of adapter and primer sequences using Cutadapt (DOI: 10.14806 / ej.17.1.200), followed by trimming of low-quality bases using Trimmomatic (Bolger et al., Bioinformatics. 2014;30(15):2114-20). Adapter removal was confirmed by running the FastQC and MultQC steps on the trimmed read fastq file. The resulting reads were aligned to the human genome reference sequence (hg38) using the BWA-MEM (Li and Durbin, Bioinformatics, 2009;25:1754-1760) tool with default parameters. The generated Sequence Alignment Map (SAM) files were sorted by coordinate and converted to Binary Alignment Map (BAM) files using PicardTools (http: / / broadinstitute.github.io / picard). The BAM files were then filtered to remove low-quality reads and retain reads with a quality score ≥ 30 using Samtools. The resulting high-quality BAM files were then checked using the CollectInsertSizeMatrix and CollectAlignmentSummaryMatrics modules from PicardTools, followed by the generation of FastQC and MultiQC reports. Reads within the sorted high-quality read BAM were grouped based on UMI using UMI-Tools with the pairs option, saved as grouped BAMs, and subsequently BAM-indexed using Samtools.The grouped BAM files were then de-duplicated using UMI-Tools with the pairs option to remove PCR duplicate reads. The FGSV tool was then used to discover structural variant pileups by searching for split-read mappings and read pairs that map across breakpoints in the BAM files using the FGSV SVPileup module. The AggregateSvPileup module in FGSV then aggregates information across nearby pileups to call structural variants. Only pileups containing at least one breakpoint on target were considered true hits, and high-confidence hits with at least 10 split reads were considered. IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) was used to facilitate the identification and analysis of genome-wide rearrangements.
[0224] To verify the DNA rearrangements identified by the above experiments, we performed PCR using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog no. 12368010) with the gene-edited genomic DNA, a primer for the interchromosomal rearrangement event CCR5#1C13 (oligo-60, common to both rearrangements), a primer for the CCR5 side (e.g., oligo-20 for the CCR5#1 left rearrangement, or oligo-23 for the CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-05A for the CCR5#1 left rearrangement, or oligo-09A for the CCR5#1 right rearrangement). The PCR mixture was prepared as follows: 100 ng gDNA, 25 µL of 2x PCR Master Mix, 5 µL of primer CCR5#1C13 (10 µM), 5 µL of primer CCR5 side (10 µM), and 5 µL of blocker 0 (20 µM). PCR was performed using the following program: initial denaturation at 98°C for 30 seconds, 35 cycles of four-step denaturation-annealing-amplification (98°C for 10 seconds, 78°C for 10 seconds, 60°C for 10 seconds, 72°C for 60 seconds), and a final extension at 72°C for 5 minutes and a hold at 4°C. PCR products were profiled directly on a 2% EX agarose gel. The genomic DNA of all experiments performed in this application was from the same batch; therefore, the PCR method is the same for all methods, and the results are applicable to all methods.
[0225] Table 9 below summarizes the DNA rearrangement events detected in control and edited cells for Example 2. To eliminate noise introduced by PCR during library preparation for next-generation sequencing, we set the threshold for a high-confidence hit at 10 split reads, with one breakpoint located at the on-target site. In control cells, 0 hits passed the high-confidence filter for left-side events and 18 hits were enriched for right-side events in Table 9. However, the number of split reads supporting these hits was significantly lower compared to the events in edited cells in Figures 11A-B. In edited samples, a total of 63 and 68 hits were captured for left-side and right-side events, respectively. These data support the generation or increase of hits captured in edited cells by CRISPR / Cas9 gene editing. Table 9. Summary of DNA rearrangement events detected for Example 2 [Table 9]
[0226] The DNA rearrangements identified in Example 2 in CCR5-edited cells are shown in Table 10 below. The table includes the genomic coordinates of the DNA breakpoints between the DNA rearrangements, the number of unique split reads based on UMI-tools, the reads extracted to indicate the split reads, and notes to further explain the potential mechanism of this DNA rearrangement. It is noteworthy that DNA rearrangements between CCR5 and its homologous gene CCR2, as well as an interchromosomal translocation between CCR5 on chromosome 3 and RNF17 / CENPJ on chromosome 13, were also identified by CAST-seq (Turchiano et al., 2021). PCR was also used to verify the presence of fusion DNA between chr3 / chr13 and the acentric chromosome (Figure 12A-B). Therefore, the SAFER detection method IA can efficiently capture both intrachromosomal and interchromosomal rearrangements. Table 10. Examples of split reads for DNA rearrangement events. [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4] [Table 10-5] [Table 10-6]
[0227] Example 3: Exemplary Method of Method IB Exemplary oligonucleotides are listed in Table 1.
[0228] Additional reagents include Tn5 transposase, e.g., Robust Tn5 transposase (catalog number EMQZ1422, Creative Biogene), DNA polymerase PCR master mix, e.g., Platinum™ SuperFi II PCR Master Mix (ThermoFisher, catalog number 12368010), DNA polymerase with PCR buffer, e.g., Platinum™ SuperFi II DNA polymerase (ThermoFisher, catalog number 12361010), dNTPs, e.g., dNTP mix (10 mM each) (ThermoFisher, catalog number R0194), ddNTPs, e.g., dideoxynucleoside triphosphate set (MilliporeSigma, catalog number 03732738001), DNA purification kit: column kit, e.g., DNA Clean&Concentrator-5 (Zymo Research, catalog number D4004), magnetic beads, e.g., AMPure These include XP Reagents (Beckman, Catalog No. A63880), Illumina sequencing kits such as NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Cat. No. 20050264), MiSeq Reagent Kit v3 (150 cycles) (Cat. No. MS-102-3001), MiSeq Reagent Kit v2 (300 cycles) (Cat. No. MS-102-2002), PhiX Control v3 (Illumina, Catalog No. FC-110-3001), and STE buffer (10 mM Tris-HCl (pH 8.0), 1 mM EDTA, 0.1 M NaCl).
[0229] In the first step, the Tn5 transposome is assembled. Transposon end-containing oligos (e.g., oligo-01 and oligo-02 for the non-UMI version, or oligo-26 and oligo-02 for the UMI version) are annealed in STE buffer using the following program: 90–99°C for 1–10 min, followed by 40–500 cycles of 90–99°C for 1 min (decrease the temperature 0.2–2°C per cycle until a temperature of 1–12°C is reached). The transposome complex is assembled in TPS buffer according to the following procedure, as recommended by the manufacturer (Creative Biogene, catalog number EMQZ1422). Mix the reagents thoroughly and incubate at 15–40°C for 5–120 min. Component Volume Tn5 transposase 1-10 μL Adapter 0.5~8μL 2μL of 10x TPS buffer Add sterilized water up to 20 μL
[0230] The workflow for this example is shown in Figure 2. A tagmentation reaction is performed in which genomic DNA extracted from the gene-edited sample or control (using standard methods) is tagmented with transposome complexes assembled according to the procedure: Component Volume Genomic DNA variation (5-500ng) 6μL of 5×LM buffer Transposome 0.1~8μL Add sterilized water up to 30 μL Mix the reaction components thoroughly and incubate for 5-120 min at 30-65 °C. The tagmented DNA is purified using column-based purification.
[0231] The 3' ends of the DNA molecules are then blocked using ddNTPs and DNA polymerase SuperFi II (ThermoFisher, catalog number 12361010). The products are purified using column-based purification.
[0232] The tagmented DNA from the previous step was used in two separate PCR reactions for each target: one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target. Each PCR was then performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR was performed using a three-step PCR cycling protocol with tagmented DNA, primer 1 (oligo-04, common to all targets and rearrangements on both sides), primer 2 (e.g., oligo-05 for EMX1 left rearrangement, oligo-15 for EMX1 right rearrangement, oligo-19 for CCR5#1 left rearrangement, or oligo-23 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement). The PCR was performed using II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the following program: initial denaturation at 90-99°C for 1-5 min, 20-40 cycles of four-step denaturation-annealing-amplification (90-99°C for 5-60 s, 75-92°C for 5-60 s, 55-75°C for 5-60 s, 67-75°C for 5-60 s), and a final extension at 67-75°C for 1-10 min and a hold time of 4-12°C. PCR products were purified using column-based purification.
[0233] Next, perform nested PCR using the barcode tags. PCR is performed with the PCR product from the previous step, primer 3 (oligo-06, common to all targets and both sides of the rearrangement), primer 4 (e.g., oligo-07 for EMX1 left rearrangement, oligo-16 for EMX1 right rearrangement, oligo-20 for CCR5#1 left rearrangement, or oligo-24 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0234] Next, tag PCR is performed using the sequencing primers. PCR is performed with the PCR product from the previous step, primer 5 (e.g., oligo-08, an example of an Illumina index primer from New England Bio Labs, NEB#E7603A), primer 6 (e.g., oligo-09, an example of an Illumina index primer from New England Bio Labs, NEB#E7611A), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi PCR. II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0235] Next, perform next-generation sequencing using the PCR products from the previous step, with optional ≥0–95% PhiX (Illumina, catalog number FC-110-3001) for sequencing using an Illumina platform (e.g., Illumina MiSeq or NextSeq 2000).
[0236] Paired-end sequencing reads from Illumina sequencing are merged using paired-end read merging software such as PEAR (Zhang et al., Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al., PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). The merged reads are then trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads are aligned to the human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100). The aligned BAM file is converted to a BED file using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34).For methods using UMIs, collapse UMIs to remove redundant sequencing reads using software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer (github.com / smangul1 / UMI-Reducer), or custom scripts.
[0237] To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site are identified and counted. Statistical tests are then applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) facilitate genome-wide identification and analysis of rearrangements.
[0238] Example 4: Exemplary Method of Method Ib Exemplary oligonucleotides used are listed in Table 5.
[0239] Additional reagents included unloaded Tn5 transposase (Diagenode, C01070010, Diagenode), 2X Tagmentation buffer (Diagenode, C01019043), Platinum™ SuperFi II PCR Master Mix (ThermoFisher, Cat. No. 12368010), dNTP mix (10 mM each) (ThermoFisher, Cat. No. R0194), DNA purification kit: DNA Clean & Concentrator-5 (Zymo Research, Cat. No. D4004), AMPure XP Reagent (Beckman Coulter, Cat. No. A63880), glycerol (MilliporeSigma, 356350500ML), 10% SDS (Fisher, BP2436200), 5 M NaCl (Invitrogen, AM9760G), Tris HCl The reagents included 1M, pH 8.0 (Invitrogen, 15568-025), genomic DNA (extracted from edited and unedited 293T cells), ddNTPs, dideoxynucleoside triphosphate set (MilliporeSigma, catalog number 03732738001), NEBNext library Quant Kit for Illumina (NEB, E7630), terminal transferase (TdT, NEB, M0315S), USER (NEB#M5508), Illumina NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Illumina, catalog number 20050264), and PhiX Control v3 (Illumina, catalog number FC-110-3001).
[0240] In the first step, transposon-end-containing oligos (e.g., oligo-01A-501 and oligo-03A, or oligo-01A-502 and oligo-03A) were resuspended in annealing buffer (40 mM Tris-HCl (pH 8.0), 50 mM NaCl) to a stock concentration of 100 μM. In a PCR tube, 10 μl of oligo-01A and 10 μl of oligo-03A were mixed, vortexed, and placed in a thermocycler with the following program: 95°C for 5 min, cool to 65°C (-0.1°C / s), 65°C for 5 min, cool to 25°C (-0.1°C / s), 25°C for 5 min, and hold at 4°C.
[0241] Transposome complexes were assembled according to the following procedure, as recommended by the manufacturer. Reagents were mixed in a PCR tube: 10 μL Tn5 transposase (2 μg / ul), 10 μL annealed adapters. The reagents were mixed thoroughly and incubated at 23°C for 30 minutes. 10 μL glycerol was added and mixed. The assembled transposome complexes were stored at -20°C.
[0242] The workflow for this example is shown in Figure 7. Tagmentation reactions were performed using genomic DNA extracted from the gene-edited sample or control (using standard methods) along with assembled transposome complexes using the following procedure: 100 ng of genomic DNA from the unedited control or edited sample, 20 μL of 2x tagmentation buffer, 200 ng of loaded Tn5 transposase, and HO added to 40 μL. The reaction mixture was mixed thoroughly and incubated at 55°C for 15 minutes. 10 μL of 0.2% SDS was added. Tn5 was then inactivated at 70°C for 10 minutes. The tagmented DNA was purified using a Zymo column according to the manufacturer's instructions and eluted in 11 μL.
[0243] The 3' ends of the DNA molecules were then blocked with ddNTPs and terminal transferase (TdT, NEB, M0315S) using the following procedure: 10 μL tagmented DNA, 5 μL 10× TdT buffer, 2 μL (40 U) TdT, 5 μL CoCl solution, 1 μL ddNTPs (10 mM), and add up to 50 μL. The reaction was incubated at 37°C for 2 hours, 70°C for 10 minutes, and cooled to 4°C. The DNA was then purified using a Zymo column according to the manufacturer's instructions and eluted in 21 μL.
[0244] The tagmented and 3'-blocked DNA from the previous step was used in two separate PCR reactions for each target (one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target). Each PCR was then performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the tagmented DNA, primer 1 (oligo-04A, common to all targets and both rearrangements), primer 2 (e.g., oligo-19 for the left rearrangement of CCR5#1 or oligo-23 for the right rearrangement of CCR5#1), and blocker 0 (e.g., oligo-05A for the left rearrangement of CCR5#1 or oligo-09A for the right rearrangement of CCR5#1) using a four-step PCR cycling protocol. The PCR mixture was prepared as follows: 25 μL of 2x PCR master mix, 10 μL of tagmented DNA, and 1.25 μL of primer 1 (10 μM), 2.5 μL of primer 2 (10 μM), 10 μL of blocker 0 (10 μM), and HO (added to 50 μL). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of a four-step denaturation-annealing-amplification sequence (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, and 72°C for 60 s), followed by a final extension at 72°C for 5 min and a 4°C hold. The PCR product was purified using AMpure XP beads (1x) and eluted in 12 μL.
[0245] Next, nested PCR was performed using the barcode tags. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the PCR product from the previous step, primer 3 (oligo-04A, common to all targets and both rearrangements), primer 4 (e.g., oligo-20 for the CCR5#1 left-side rearrangement or oligo-24 for the CCR5#1 right-side rearrangement), and blocker 0 (e.g., oligo-05A for the CCR5#1 left-side rearrangement or oligo-09A for the CCR5#1 right-side rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 1.25 μL of primer 3 (10 μM), 2.5 μL of primer 4 (10 μM), 10 μL of blocker 0 (10 μM), and HO added to 50 μL. PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 12 μL.
[0246] Next, tag PCR was performed using the sequencing primers. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the PCR product from the previous step, primer 1 (oligo-04A, common to all target and bilateral rearrangements), primer 6 (e.g., any oligo-13A or Illumina index primer from New England Biolabs, NEB#E7611A), and blocker 0 (e.g., oligo-05A for CCR5#1 left-sided rearrangement or oligo-09A for CCR5#1 right-sided rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 5 μL of primer 5 (10 μM), 5 μL of primer 6 (10 μM), and 5 μL of blocker 0 (20 μM). PCR was performed using the following program: initial denaturation at 98°C for 30 seconds, 15 cycles of a four-step denaturation-annealing-amplification (98°C for 10 seconds, 78°C for 10 seconds, 60°C for 10 seconds, 72°C for 10 seconds), and a final extension at 72°C for 5 minutes and a hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 52 μL.
[0247] 50 μL of the supernatant was transferred to a new tube for library quantification using the NEBNext Library Quant Kit for Illumina (NEB, E7630) according to the manufacturer's instructions. Libraries were normalized and pooled for sequencing.
[0248] Next, next-generation sequencing was performed using the PCR products from the previous step, and sequencing was performed using an Illumina platform (NextSeq 2000) with 20% PhiX (Illumina, catalog number FC-110-3001).
[0249] Paired-end sequencing reads from Illumina sequencing were processed using a customized NGS analysis pipeline. Sequence reads were first checked for overall quality using FastQC, followed by integration of FastQC results using MultiQC. Reads were then trimmed of adapter and primer sequences using Cutadapt (DOI: 10.14806 / ej.17.1.200), followed by trimming of low-quality bases using Trimmomatic (Bolger et al., Bioinformatics. 2014;30(15):2114-20). Adapter removal was confirmed by running the FastQC and MultQC steps on the trimmed read fastq file. The resulting reads were aligned to the human genome reference sequence (hg38) using the BWA-MEM (Li and Durbin, Bioinformatics, 2009;25:1754-1760) tool with default parameters. The generated Sequence Alignment Map (SAM) files were sorted by coordinate and converted to Binary Alignment Map (BAM) files using PicardTools (http: / / broadinstitute.github.io / picard). The BAM files were then filtered to remove low-quality reads and retain reads with a quality score ≥ 30 using Samtools. The resulting high-quality BAM files were then checked using the CollectInsertSizeMatrix and CollectAlignmentSummaryMatrics modules from PicardTools, followed by the generation of FastQC and MultiQC reports. Reads within the sorted high-quality read BAM were grouped based on UMI using UMI-Tools with the pairs option, saved as grouped BAMs, and subsequently BAM-indexed using Samtools.The grouped BAM files were then de-duplicated using UMI-Tools with the pairs option to remove PCR duplicate reads. The FGSV tool was then used to discover structural variant pileups by searching for split-read mappings and read pairs that map across breakpoints in the BAM files using the FGSV SVPileup module. The AggregateSvPileup module in FGSV then aggregates information across nearby pileups to call structural variants. Only pileups containing at least one breakpoint on target were considered true hits, and high-confidence hits with at least 10 split reads were considered. IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) was used to facilitate the identification and analysis of genome-wide rearrangements.
[0250] The DNA rearrangement events detected for Example 4 in control and edited cells are shown in Table 11 below. The criteria used to remove noise were the same as in Method IA. As shown in Table 11, high-confidence hits of 36 and 37 for left- and right-sided DNA rearrangement events, respectively, passed the filter in edited cells, while only 0 and 10 hits were captured for left-sided and right-sided events, respectively, in control cells. However, in addition to more hits, the number of split reads supporting these hits was also much higher in edited cells, as shown in Figures 13A-B. Table 11. Summary of DNA rearrangements according to Example 4 in control and CCR5-edited cells. [Table 11]
[0251] The DNA rearrangements identified in Example 4 in CCR5-edited cells are shown in Table 12 below. Table 12 includes the genomic coordinates of the DNA breakpoints between the DNA rearrangements, the number of unique split reads based on the BMI tool, the reads extracted to indicate the split reads, and notes to further explain the potential mechanism of this DNA rearrangement. It is noteworthy that DNA rearrangements between CCR5 and its homologous gene CCR2, as well as an interchromosomal translocation between CCR5 on chromosome 3 and RNF17 / CENPJ on chromosome 13, were also identified by CAST-seq (Turchiano et al., 2021). PCR was also used to verify the presence of fusion DNA between chr3 / chr13 and the acentric chromosome (Figures 13A-B). Therefore, the SAFER detection method IB can efficiently capture both intrachromosomal and interchromosomal rearrangements. Table 12. Examples of DNA rearrangements identified by Example 4 in CCR5-edited cells. [Table 12-1] [Table 12-2] [Table 12-3] [Table 12-4]
[0252] Example 5: Exemplary Method of Method Ic Exemplary oligonucleotides are listed in Table 1.
[0253] Additional reagents include Tn5 transposase, e.g., Robust Tn5 transposase (catalog number EMQZ1422, Creative Biogene), DNA polymerase PCR master mix, e.g., Platinum™ SuperFi II PCR Master Mix (ThermoFisher, catalog number 12368010), DNA polymerase with PCR buffer, e.g., Platinum™ SuperFi II DNA polymerase (ThermoFisher, catalog number 12361010), dNTPs, e.g., dNTP mix (10 mM each) (ThermoFisher, catalog number R0194), DNA purification kit: column kit, e.g., DNA Clean&Concentrator-5 (Zymo Research, catalog number D4004), magnetic beads, e.g., AMPure XP reagent (Beckman, catalog number A63880), Illumina sequencing kit, e.g., NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Cat. No. 20050264), MiSeq Reagent Kit v3 (150 cycles) (Cat. No. MS-102-3001), MiSeq Reagent Kit v2 (300 cycles) (Cat. No. MS-102-2002), PhiX Control v3 (Illumina, Cat. No. FC-110-3001), USER® Enzyme (New England Biolabs, Cat. No. M5505S), and STE buffer (10 mM Tris-HCl (pH 8.0), 1 mM EDTA, 0.1 M NaCl).
[0254] In the first step, the Tn5 transposome is assembled. Transposon end-containing oligos (e.g., oligo-10 and oligo-02 for the non-UMI version, or oligo-27 and oligo-02 for the UMI version) are annealed in STE buffer using the following program: 90–99°C for 1–10 min, followed by 40–500 cycles of 90–99°C for 1 min (decrease the temperature by 0.2–2°C per cycle until a temperature of 1–12°C is reached). The transposome complex is assembled in TPS buffer according to the following procedure, as recommended by the manufacturer (Creative Biogene, catalog number EMQZ1422). Mix the reagents thoroughly and incubate at 15–40°C for 5–120 min. Component Volume Tn5 transposase 1-10 μL Adapter 0.5~8μL 2μL of 10x TPS buffer Add sterilized water up to 20 μL
[0255] The workflow for this example is shown in Figure 3. A tagmentation reaction is performed in which genomic DNA extracted from the gene-edited sample or control (using standard methods) is tagmented with the transposome complex assembled according to the procedure. An exemplary reaction mixture is shown below: Component Volume Genomic DNA variation (5-500ng) 6μL of 5×LM buffer Transposome 0.1~8μL Add sterilized water up to 30 μL
[0256] Mix the reaction components thoroughly and incubate for 5-120 min at 30-65 °C. The tagmented DNA is purified using column-based purification.
[0257] Next, perform the first new strand synthesis. PCR reactions are set up with the fragmented DNA from the previous step, primer 2 (e.g., oligo-05 for EMX1 left rearrangement, oligo-15 for EMX1 right rearrangement, oligo-19 for CCR5#1 left rearrangement, or oligo-23 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement), and then PCR is performed using SuperFix. II Polymerase PCR Master Mix (ThermoFisher, catalog no. 12368010) is performed with the following program: denaturation at 90–99°C for 1–10 min, blocker annealing at 75–92°C for 5–120 s, primer annealing at 55–75°C for 1 min, primer extension at 67–75°C for 1–10 min, and a hold at 4–12°C. Products are purified using column-based purification.
[0258] Next, USER enzyme cleavage is performed. Briefly, USER enzyme cleavage is performed as follows: Add 0.5-10 μL of USER enzyme to the product from the previous step, mix, and incubate at 16-45°C for 5-60 minutes. Purify the product using a DNA purification kit.
[0259] The USER-treated DNA from the previous step was used in two separate PCR reactions for each target: one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target. Each PCR was then performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR was performed using a three-step PCR cycling protocol with tagmented DNA, primer 1 (oligo-04, common to all targets and rearrangements on both sides), primer 2 (e.g., oligo-05 for EMX1 left rearrangements, oligo-15 for EMX1 right rearrangements, oligo-19 for CCR5#1 left rearrangements, or oligo-23 for CCR5#1 right rearrangements), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangements, oligo-14 for EMX1 right rearrangements, oligo-18 for CCR5#1 left rearrangements, or oligo-22 for CCR5#1 right rearrangements). The PCR was performed using II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the following program: initial denaturation at 90-99°C for 1-5 min, 20-40 cycles of four-step denaturation-annealing-amplification (90-99°C for 5-60 s, 75-92°C for 5-60 s, 55-75°C for 5-60 s, 67-75°C for 5-60 s), and a final extension at 67-75°C for 1-10 min and a hold time of 4-12°C. PCR products were purified using column-based purification.
[0260] Next, perform nested PCR using the barcode tags. PCR is performed with the PCR product from the previous step, primer 3 (oligo-06, common to all targets and both sides of the rearrangement), primer 4 (e.g., oligo-07 for EMX1 left rearrangement, oligo-16 for EMX1 right rearrangement, oligo-20 for CCR5#1 left rearrangement, or oligo-24 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0261] Next, tag PCR is performed using the sequencing primers. PCR is performed with the PCR product from the previous step, primer 5 (e.g., oligo-08, an example of an Illumina index primer from New England Bio Labs, NEB#E7603A), primer 6 (e.g., oligo-09, an example of an Illumina index primer from New England Bio Labs, NEB#E7611A), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi PCR. II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0262] Next, perform next-generation sequencing using the PCR products from the previous step, with optional ≥0–95% PhiX (Illumina, catalog number FC-110-3001) for sequencing using an Illumina platform (e.g., Illumina MiSeq or NextSeq 2000).
[0263] Paired-end sequencing reads from Illumina sequencing are merged using paired-end read merging software such as PEAR (Zhang et al., Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al., PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). The merged reads are then trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads are aligned to the human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100). The aligned BAM file is converted to a BED file using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34).For methods using UMIs, collapse UMIs to remove redundant sequencing reads using software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer (github.com / smangul1 / UMI-Reducer), or custom scripts. To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site are identified and counted. Statistical tests are then applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) facilitate genome-wide identification and analysis of rearrangements.
[0264] Example 6: Exemplary Method of Method Ic Exemplary oligonucleotides used are listed in Table 5.
[0265] Additional reagents included unloaded Tn5 transposase (Diagenode, C01070010, Diagenode), 2X Tagmentation buffer (Diagenode, C01019043), Platinum™ SuperFi II PCR Master Mix (ThermoFisher, Cat. No. 12368010), dNTP mix (10 mM each) (ThermoFisher, Cat. No. R0194), DNA purification kit: DNA Clean & Concentrator-5 (Zymo Research, Cat. No. D4004), AMPure XP Reagent (Beckman Coulter, Cat. No. A63880), glycerol (MilliporeSigma, 356350500ML), 10% SDS (Fisher, BP2436200), 5 M NaCl (Invitrogen, AM9760G), Tris HCl The reagents included 1M, pH 8.0 (Invitrogen, 15568-025), genomic DNA (extracted from edited and unedited 293T cells), ddNTPs, dideoxynucleoside triphosphate set (MilliporeSigma, catalog number 03732738001), NEBNext library Quant Kit for Illumina (NEB, E7630), terminal transferase (TdT, NEB, M0315S), USER (NEB#M5508), Phusion U hot-start DNA polymerase (Thermofisher, F555S), Illumina NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Illumina, catalog number 20050264), and PhiX Control v3 (Illumina, catalog number FC-110-3001).
[0266] In the first step, transposon end-containing oligos (e.g., oligo-02A and oligo-03A) were resuspended in annealing buffer (40 mM Tris-HCl (pH 8.0), 50 mM NaCl) to a stock concentration of 100 μM. In a PCR tube, 10 μl of oligo-01A and 10 μl of oligo-03A were mixed, vortexed, and placed in a thermocycler with the following program: 95°C for 5 min, cool to 65°C (-0.1°C / s), 65°C for 5 min, cool to 25°C (-0.1°C / s), 25°C for 5 min, and hold at 4°C.
[0267] Transposome complexes were assembled using the following procedure, as recommended by the manufacturer. Reagents were mixed in a PCR tube: 10 μL Tn5 transposase (2 μg / ul), 10 μL annealed adapters. The reagents were mixed thoroughly and incubated at 23°C for 30 minutes. 10 μL glycerol was added and mixed. Assembled transposome complexes were stored at -20°C.
[0268] The workflow for this example is shown in Figure 8. Tagmentation reactions were performed using genomic DNA extracted from gene-edited samples or controls (using standard methods) as follows: Add 100 ng of genomic DNA (unedited control or edited sample), 20 μL of 2x tagmentation buffer, 200 ng of Tn5 transposase, and HO up to 40 μL. The reaction mixture was mixed thoroughly and incubated at 55°C for 15 minutes. Add 10 μL of 0.2% SDS. Tn5 was then inactivated at 70°C for 10 minutes. The tagmented DNA was purified using a Zymo column according to the manufacturer's instructions and eluted in 11 μL.
[0269] The 3' ends of the DNA molecules were then blocked with ddNTPs and terminal transferase (TdT, NEB, M0315S) using the following procedure: 10 μL of tagmented DNA, 5 μL of 10× TdT buffer, 2 μL (40 U) TdT, 5 μL of CoCl solution, 1 μL of ddNTPs (10 mM), and HO added to 50 μL. The reaction was incubated at 37°C for 2 hours, 70°C for 10 minutes, and cooled to 4°C. The DNA was then purified using a Zymo column according to the manufacturer's instructions and eluted in 21 μL.
[0270] First-strand synthesis was then performed for each target using the DNA from the previous step in two separate PCR reactions (one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target). Each PCR was then performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR reactions were set up using Phusion U Hot Start DNA Polymerase (Thermofisher, F555S) with the fragmented and 3'-blocked DNA from the previous step, primer 2 (e.g., oligo-19 for the left rearrangement of CCR5#1, or oligo-23 for the right rearrangement of CCR5#1), and blocker 0 (e.g., oligo-05A for the left rearrangement of CCR5#1, or oligo-09A for the right rearrangement of CCR5#1). The reaction mixture was prepared as follows: 10 μL of blocked DNA, 10 μL of 5X Phusion buffer, 0.5 μL of Phusion U polymerase, 1 μL of dNTPs (10 mM), 2.5 μL of Primer 2 (10 μM), 10 μL of Blocker 0 (10 μM), and HO (to 50 μL). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 1 or 20 cycles of a four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for 60 s), and a final extension at 72°C for 5 min and a 4°C hold. The PCR product was purified using AMpure XP beads (1X) and eluted in 12 μL.
[0271] USER enzyme cleavage was then performed. Briefly, USER enzyme cleavage was performed using the following procedure: 10 μL purified DNA, 5 μL (10X) rCutsmart buffer, 2 μL USER, and HO added to 50 μL. The mixture was incubated at 37°C for 30 minutes. DNA was purified using a Zymo column and eluted in 11 μl for PCR.
[0272] The user-treated DNA from the previous step was used for the first PCR amplification. Each PCR was performed using a common primer annealed to the adapter region, a target-specific blocker, and a target-specific primer. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with tagmented DNA, primer 1 (oligo-04A, common to all targets and both rearrangements), primer 2 (e.g., oligo-06A for the CCR5#1 left rearrangement or oligo-10A for the CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-05A for the CCR5#1 left rearrangement or oligo-09A for the CCR5#1 right rearrangement) using a four-step PCR cycling protocol. The PCR mixture was prepared as follows: 25 μL of 2x PCR master mix, 10 μL of tagmented DNA, and 1.25 μL of primer 1 (10 μM), 2.5 μL of primer 2 (10 μM), 10 μL of blocker 0 (10 μM), and HO (added to 50 μL). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of a four-step denaturation-annealing-amplification sequence (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, and 72°C for 60 s), followed by a final extension at 72°C for 5 min and a 4°C hold. The PCR product was purified using AMpure XP beads (1x) and eluted in 12 μL.
[0273] Next, nested PCR was performed using the barcode tags. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the PCR product from the previous step: primer 3 (oligo-04A, common to all targets and both rearrangements), primer 4 (e.g., oligo-20 for the CCR5#1 left-side rearrangement or oligo-24 for the CCR5#1 right-side rearrangement), and blocker 0 (e.g., oligo-05A for the CCR5#1 left-side rearrangement or oligo-09A for the CCR5#1 right-side rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 1.25 μL of primer 3 (oligo-04, 10 μM), 2.5 μL of primer 4 (10 μM), 10 μL of blocker 0 (10 μM), and HO added to a volume of 50 μL. PCR was performed using the following program: initial denaturation at 98°C for 30 seconds, 15 cycles of a four-step denaturation-annealing-amplification (98°C for 10 seconds, 78°C for 10 seconds, 60°C for 10 seconds, 72°C for 10 seconds), and a final extension at 72°C for 5 minutes and a hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 12 μL.
[0274] Next, tag PCR was performed using the sequencing primers. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the product from the previous step, primer 5 (oligo-04A, common to all target and bilateral rearrangements), primer 6 (e.g., any oligo-13A or Illumina index primer from New England Biolabs, NEB#E7611A), and blocker 0 (e.g., oligo-05A for CCR5#1 left-sided rearrangement or oligo-09A for CCR5#1 right-sided rearrangement). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 5 μL of primer 5 (10 μM), 5 μL of primer 6 (10 μM), and 5 μL of blocker 0 (20 μM). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 52 μL.
[0275] 50 μL of the supernatant was transferred to a new tube for library quantification using the NEBNext Library Quant Kit for Illumina (NEB, E7630) according to the manufacturer's instructions. Libraries were normalized and pooled for sequencing.
[0276] Next-generation sequencing was then performed using the PCR products from the previous step, and the sequences were sequenced using an Illumina platform (NextSeq 2000) at 20% PhiX (Illumina, catalog number FC-110-3001).
[0277] Paired-end sequencing reads from Illumina sequencing were processed using a customized NGS analysis pipeline. Sequence reads were first checked for overall quality using FastQC, followed by integration of FastQC results using MultiQC. Reads were then trimmed of adapter and primer sequences using Cutadapt (DOI: 10.14806 / ej.17.1.200), followed by trimming of low-quality bases using Trimmomatic (Bolger et al., Bioinformatics. 2014;30(15):2114-20). Adapter removal was confirmed by running the FastQC and MultQC steps on the trimmed read fastq file. The resulting reads were aligned to the human genome reference sequence (hg38) using the BWA-MEM (Li and Durbin, Bioinformatics, 2009;25:1754-1760) tool with default parameters. The generated Sequence Alignment Map (SAM) files were sorted by coordinate and converted to Binary Alignment Map (BAM) files using PicardTools (http: / / broadinstitute.github.io / picard). The BAM files were then filtered to remove low-quality reads and retain reads with a quality score ≥ 30 using Samtools. The resulting high-quality BAM files were then checked using the CollectInsertSizeMatrix and CollectAlignmentSummaryMatrics modules from PicardTools, followed by the generation of FastQC and MultiQC reports. Reads within the sorted high-quality read BAM were grouped based on UMI using UMI-Tools with the pairs option, saved as grouped BAMs, and subsequently BAM-indexed using Samtools.The grouped BAM files were then de-duplicated using UMI-Tools with the pairs option to remove PCR duplicate reads. The FGSV tool was then used to discover structural variant pileups by searching for split-read mappings and read pairs that map across breakpoints in the BAM files using the FGSV SVPileup module. The AggregateSvPileup module in FGSV then aggregates information across nearby pileups to call structural variants. Only pileups containing at least one breakpoint on target were considered true hits, and high-confidence hits with at least 10 split reads were considered. IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) was used to facilitate the identification and analysis of genome-wide rearrangements.
[0278] The DNA rearrangements from Example 6 in control and CCR5-edited cells are shown in Table 13 below. The criteria used to remove noise were the same as in Method IA. In this example, only the left-sided event library was performed. As shown in Table 13, 38 and 57 hits passed the filter with high confidence in the control and edited cells, respectively. However, the number of unique splits supporting these hits was still significantly higher in the edited cells (Figure 14). Table 13. Summary of DNA rearrangements according to Example 6 in control and CCR5-edited cells. [Table 13]
[0279] The DNA rearrangements identified in Example 6 in CCR5-edited cells are shown in Table 14 below. The table includes the genomic coordinates of the DNA breakpoints between the DNA rearrangements, the number of unique split reads based on UMI-tools, the reads extracted to indicate the split reads, and notes to further explain the potential mechanism of this DNA rearrangement. It is noteworthy that an interchromosomal translocation between CCR5 on chromosome 3 and RNF17 / CENPJ on chromosome 13 was also identified by CAST-seq (Turchiano et al., 2021). PCR was also used to verify the presence of fusion DNA between chr3 / chr13 and the acentric chromosome (Figure 13A-B). Therefore, the SAFER detection method can efficiently capture both intrachromosomal and interchromosomal rearrangements. Table 14. Examples of DNA rearrangements identified by Example 6 in CCR5-edited cells. [Table 14-1] [Table 14-2]
[0280] Example 7: Exemplary Method of Method II Exemplary oligonucleotides are listed in Table 1.
[0281] Additional reagents include Tn5 transposase, e.g., Robust Tn5 transposase (catalog number EMQZ1422, Creative Biogene), DNA polymerase PCR master mix, e.g., Platinum™ SuperFi II PCR Master Mix (ThermoFisher, catalog number 12368010), DNA polymerase with PCR buffer, e.g., Platinum™ SuperFi II DNA polymerase (ThermoFisher, catalog number 12361010), dNTPs, e.g., dNTP mix (10 mM each) (ThermoFisher, catalog number R0194), DNA purification kit: column kit, e.g., DNA Clean&Concentrator-5 (Zymo Research, catalog number D4004), magnetic beads, e.g., AMPure XP reagent (Beckman, catalog number A63880), Illumina sequencing kit, e.g., NextSeq 1000 / 2000 P1 Reagents (300 cycles) (catalog number 20050264), MiSeq Reagent Kit v3 (150 cycles) (catalog number MS-102-3001), MiSeq Reagent Kit v2 (300 cycles) (catalog number MS-102-2002), PhiX Control v3 (Illumina, catalog number FC-110-3001), sequence-specific nuclease and ribonucleoprotein assembly buffer, such as Cas9 nuclease, S. pyogenes (New England Biolabs, catalog number M0386S), EnGen® Lba Cas12a (Cpf1) (New England Biolabs, catalog number M0653S), SpRY, guide RNA specific for the target region, STE buffer (10 mM Tris-HCl (pH 8.0), 1 mM EDTA, 0.1 M NaCl). Exemplary guide RNA sequences are shown in Table 15 below. Table 15. [Table 15-1] [Table 15-2]
[0282] In the first step, the Tn5 transposome is assembled. Transposon end-containing oligos (e.g., oligo-01 and oligo-11 for the non-UMI version, or oligo-26 and oligo-11 for the UMI version) are annealed in STE buffer using the following program: 90–99°C for 1–10 min, followed by 40–500 cycles of 90–99°C for 1 min (decrease the temperature by 0.2–2°C per cycle until a temperature of 1–12°C is reached). The transposome complex is assembled in TPS buffer according to the following procedure, as recommended by the manufacturer (Creative Biogene, catalog number EMQZ1422). Mix the reagents thoroughly and incubate at 15–40°C for 5–120 min. Component Volume Tn5 transposase 1-10 μL Adapter 0.5~8μL 2μL of 10x TPS buffer Add sterilized water up to 20 μL
[0283] The workflow for this example is shown in Figure 4. A tagmentation reaction is performed in which genomic DNA extracted from the gene-edited sample or control (using standard methods) is tagmented with transposome complexes assembled according to the procedure. An exemplary reaction mixture is shown below: Component Volume Genomic DNA variation (5-500ng) 6μL of 5×LM buffer Transposome 0.1~8μL Add sterilized water up to 30 μL
[0284] Mix the reaction components thoroughly and incubate for 5-120 min at 30-65 °C. The tagmented DNA is purified using column-based purification.
[0285] Two separate ribonucleoprotein (RNP) assembly reactions and corresponding in vitro cleavage (IVC) reactions are then performed for each target: one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target. CRISPR / Cas9 nuclease, S. pyogenes (New England Biolabs, catalog number M0386S), and sgRNA (Table 5) are assembled into RNPs as follows. NEBuffer r3.1 3 μL 300nM sgRNA 0.6~15μL (final 6~150nM) 1μM Cas9 nuclease 0.2~5μL (6~150nM final) Add sterilized water up to 30 μL The reaction mixture is mixed thoroughly and incubated for 1-120 min at 15-45°C. RNP is added to the tagmented DNA at an RNP:DNA ratio of 0.01:1 to 50:1. The cleavage reaction is incubated for 5-120 min at 15-45°C. The product is purified using column-based purification.
[0286] If necessary, the 3' ends of the DNA molecules are then blocked using ddNTPs and DNA polymerase SuperFi II (ThermoFisher, catalog number 12361010). The products are purified using column-based purification.
[0287] Next, perform PCR using the IVC DNA from the previous step with a common primer annealed to the adapter region and a target-specific primer. PCR is performed using a three-step PCR cycling protocol with IVC DNA, primer 1 (oligo-04, common to all targets and both-sided rearrangements), and primer 2 (e.g., oligo-05 for EMX1 left-side rearrangements, oligo-15 for EMX1 right-side rearrangements, oligo-19 for CCR5#1 left-side rearrangements, or oligo-23 for CCR5#1 right-side rearrangements) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog no. 12368010) with the following program: initial denaturation at 90–99 °C for 1–5 min, 20–40 cycles of four steps of denaturation-annealing-amplification (90–99 °C for 5–60 s, 75–92 °C for 5–60 s, 55–75 °C for 5–60 s, and 67–75 °C for 5–60 s), and a final extension at 67–75 °C for 1–10 min and a hold time of 4–12 °C. The PCR product is purified using column-based purification.
[0288] Next, perform nested PCR using the barcode tags. PCR is performed with the PCR product from the previous step, primer 3 (oligo-06, common to all targets and both sides of the rearrangement), primer 4 (e.g., oligo-07 for EMX1 left rearrangement, oligo-16 for EMX1 right rearrangement, oligo-20 for CCR5#1 left rearrangement, or oligo-24 for CCR5#1 right rearrangement), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0289] Next, tag PCR is performed using the sequencing primers. PCR is performed with the PCR product from the previous step, primer 5 (e.g., oligo-08, an example of an Illumina index primer from New England Bio Labs, NEB#E7603A), primer 6 (e.g., oligo-09, an example of an Illumina index primer from New England Bio Labs, NEB#E7611A), and blocker 0 (e.g., oligo-03 for EMX1 left rearrangement, oligo-14 for EMX1 right rearrangement, oligo-18 for CCR5#1 left rearrangement, or oligo-22 for CCR5#1 right rearrangement) using SuperFi PCR. II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) is used, with the following program: initial denaturation at 90–99°C for 1–5 min, 4–40 cycles of four-step denaturation–annealing–amplification (90–99°C for 5–60 s, 75–92°C for 5–60 s, 55–75°C for 5–60 s, 67–75°C for 5–60 s), and a final extension at 67–75°C for 1–10 min and a hold at 4–12°C. PCR products are purified using column-based purification.
[0290] Next, perform next-generation sequencing using the PCR products from the previous step, with optional ≥0–95% PhiX (Illumina, catalog number FC-110-3001) for sequencing using an Illumina platform (e.g., Illumina MiSeq or NextSeq 2000).
[0291] Paired-end sequencing reads from Illumina sequencing are merged using paired-end read merging software such as PEAR (Zhang et al., Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al., PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). The merged reads are then trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads are aligned to the human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100). The aligned BAM file is converted to a BED file using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34).For methods using UMIs, collapse UMIs to remove redundant sequencing reads using software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer (github.com / smangul1 / UMI-Reducer), or custom scripts. To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site are identified and counted. Statistical tests are then applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) facilitate genome-wide identification and analysis of rearrangements.
[0292] Example 8: Exemplary Method of Method II Exemplary oligonucleotides used are listed in Table 5.
[0293] Additional reagents included unloaded Tn5 transposase (Diagenode, C01070010, Diagenode), 2X Tagmentation buffer (Diagenode, C01019043), Platinum™ SuperFi II PCR Master Mix (ThermoFisher, Cat. No. 12368010), dNTP mix (10 mM each) (ThermoFisher, Cat. No. R0194), DNA purification kit: DNA Clean & Concentrator-5 (Zymo Research, Cat. No. D4004), AMPure XP Reagent (Beckman Coulter, Cat. No. A63880), glycerol (MilliporeSigma, 356350500ML), 10% SDS (Fisher, BP2436200), 5 M NaCl (Invitrogen, AM9760G), Tris HCl 1M, pH 8.0 (Invitrogen, 15568-025), genomic DNA (extracted from edited and unedited 293T cells), ddNTPs, dideoxynucleoside triphosphate set (MilliporeSigma, catalog number 03732738001), NEBNext library Quant Kit for Illumina (NEB, E7630), terminal transferase (TdT, NEB, M0315S), Alt-R™ SpHiFi Cas9 nuclease V3 (IDT, 1081060), NEB Buffer r3.1 (NEB, B6003S), Illumina NextSeq 1000 / 2000 P1 Reagents (300 cycles) (Illumina, catalog number 20050264), PhiX Control v3 (Illumina, catalog number FC-110-3001). Table 16. Exemplary guide RNA sequences are shown below. [Table 16]
[0294] In the first step, transposon-end-containing oligos (e.g., oligo-01A and oligo-03A, or oligo-02A and oligo-03A) were resuspended in annealing buffer (40 mM Tris-HCl (pH 8.0), 50 mM NaCl) to a stock concentration of 100 μM. In a PCR tube, 10 μl of oligo-01 and 10 μl of oligo-03 were mixed, vortexed, and placed in a thermocycler with the following program: 95°C for 5 min, cool to 65°C (-0.1°C / s), 65°C for 5 min, cool to 25°C (-0.1°C / s), 25°C for 5 min, and hold at 4°C. Transposome complexes were assembled according to the following procedure, as recommended by the manufacturer. Reagents were mixed in a PCR tube: 10 μL Tn5 transposase (2 μg / ul), 10 μL annealed adapters. The reagents were mixed thoroughly and incubated at 23°C for 30 minutes. 10 μL glycerol was added and mixed. The assembled transposome complexes were stored at -20°C.
[0295] The workflow for this example is shown in Figure 9. Tagmentation reactions were performed using genomic DNA extracted (using standard methods) from gene-edited samples or controls using the following procedure: 100 ng of genomic DNA from an unedited control or edited sample was combined with 20 μL of 2x tagmentation buffer, 200 ng of loaded Tn5 transposase, and HO was added to a volume of 40 μL. The reaction mixture was mixed thoroughly and incubated at 55°C for 15 minutes. 10 μL of 0.2% SDS was added, and then the Tn5 was inactivated at 70°C for 10 minutes. The tagmented DNA was purified using a Zymo column according to the manufacturer's instructions and eluted in 21 μL.
[0296] Two separate ribonucleoprotein (RNP) assembly reactions and corresponding in vitro cleavage (IVC) reactions were then performed for each target (one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target). CRISPR / Cas9 nuclease, Alt-R™ SpHiFi Cas9 nuclease V3 (IDT, 1081060), and CCR5 sgRNA (e.g., CCR5_L for left-side rearrangements, or CCR5_R for right-side rearrangements) (Table 5A) were assembled into RNPs as follows: 32.5 μL HO, 5 μL (10X) NEBuffer r3.1, 1.5 μL CCR5 sgRNA (10 μM), 0.8 μL HiFi Cas9 nuclease (diluted to 6.2 μM). The RNP reactions were thoroughly mixed and incubated at room temperature for 10 minutes.
[0297] The tagmented DNA was cleaved in vitro using the Cas9 RNP assembled in the previous step. 10 μL of DNA was added to the RNP solution and incubated at 37°C for 2 hours. The product was purified using a Zymo column according to the manufacturer's instructions and eluted in 11 μl.
[0298] The 3' ends of the DNA molecules were then blocked with ddNTPs and terminal transferase (TdT, NEB, M0315S) using the following procedure: 10 μL tagmented DNA, 5 μL 10× TdT buffer, 2 μL (40 U) TdT, 5 μL CoCl solution, 1 μL ddNTPs (10 mM), and added to 50 μL. The reaction was incubated at 37°C for 2 hours, 70°C for 10 minutes, and cooled to 4°C. The DNA was then purified using a Zymo column according to the manufacturer's instructions and eluted in 11 μL.
[0299] Next, PCR was performed using the blocked IVC DNA from the previous step with a common primer annealed to the adapter region and a target-specific primer. PCR was performed using a four-step PCR cycling protocol with SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) containing tagmented DNA, primer 1 (oligo-04A, common to all targets and both rearrangements), and primer 2 (e.g., oligo-06A for the left-side rearrangement of CCR5#1, or oligo-10A for the right-side rearrangement of CCR5#1). The PCR mixture was prepared as follows: 25 µL of 2x PCR Master Mix, 10 µL of tagmented DNA, and 1.25 µL of primer 1 (10 µM), 2.5 µL of primer 2 (10 µM), 10 µL of blocker 0 (10 µM), and HO up to 50 µL. PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for 60 s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads (1×) and eluted in 12 μL.
[0300] Next, nested PCR was performed using the barcode tags. PCR was performed using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the PCR product from the previous step, primer 3 (oligo-04A, common to all targets and both rearrangements), and primer 4 (e.g., oligo-20 for the left-side rearrangement of CCR5#1, or oligo-24 for the right-side rearrangement of CCR5#1). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 1.25 μL of primer 3 (10 μM), 2.5 μL of primer 4 (10 μM), 10 μL of blocker 0 (10 μM), and HO (added to 50 μL). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 12 μL.
[0301] Next, tag PCR was performed using the sequencing primers. PCR was performed using the PCR product from the previous step, primer 1 (oligo-04A, common to all targets and both rearrangements), (e.g., any oligo-13A or Illumina index primer from New England Biolabs, NEB#E7611A), and SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010). The PCR mixture was prepared as follows: 25 μL of 2x PCR Master Mix, 10 μL of PCR product, 5 μL of Primer 5 (10 μM), 5 μL of Primer 6 (10 μM), and 5 μL of Blocker 0 (20 μM). PCR was performed using the following program: initial denaturation at 98°C for 30 s, 15 cycles of four-step denaturation-annealing-amplification (98°C for 10 s, 78°C for 10 s, 60°C for 10 s, 72°C for s), and a final extension at 72°C for 5 min and hold at 4°C. PCR products were purified using AMpure XP beads and eluted in 52 μL.
[0302] 50 μL of the supernatant was transferred to a new tube for library quantification using the NEBNext Library Quant Kit for Illumina (NEB, E7630) according to the manufacturer's instructions. Libraries were normalized and pooled for sequencing.
[0303] Next, next-generation sequencing was performed using the PCR products from the previous step, and sequencing was performed using an Illumina platform (NextSeq 2000) with 20% PhiX (Illumina, catalog number FC-110-3001).
[0304] Paired-end sequencing reads from Illumina sequencing were processed using a customized NGS analysis pipeline. Sequence reads were first checked for overall quality using FastQC, followed by integration of FastQC results using MultiQC. Reads were then trimmed of adapter and primer sequences using Cutadapt (DOI: 10.14806 / ej.17.1.200), followed by trimming of low-quality bases using Trimmomatic (Bolger et al., Bioinformatics. 2014;30(15):2114-20). Adapter removal was confirmed by running the FastQC and MultQC steps on the trimmed read fastq file. The resulting reads were aligned to the human genome reference sequence (hg38) using the BWA-MEM (Li and Durbin, Bioinformatics, 2009;25:1754-1760) tool with default parameters. The generated Sequence Alignment Map (SAM) files were sorted by coordinate and converted to Binary Alignment Map (BAM) files using PicardTools (http: / / broadinstitute.github.io / picard). The BAM files were then filtered to remove low-quality reads and retain reads with a quality score ≥ 30 using Samtools. The resulting high-quality BAM files were then checked using the CollectInsertSizeMatrix and CollectAlignmentSummaryMatrics modules from PicardTools, followed by the generation of FastQC and MultiQC reports. Reads within the sorted high-quality read BAM were grouped based on UMI using UMI-Tools with the pairs option, saved as grouped BAMs, and subsequently BAM-indexed using Samtools.The grouped BAM files were then de-duplicated using UMI-Tools with the pairs option to remove PCR duplicate reads. The FGSV tool was then used to discover structural variant pileups by searching for split-read mappings and read pairs that map across breakpoints in the BAM files using the FGSV SVPileup module. The AggregateSvPileup module in FGSV then aggregates information across nearby pileups to call structural variants. Only pileups containing at least one breakpoint on target were considered true hits, and high-confidence hits with at least 10 split reads were considered. IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) was used to facilitate the identification and analysis of genome-wide rearrangements.
[0305] The DNA rearrangements detected by Example 8 in control and CCR5-edited cells are shown in Table 17 below. The criteria used to remove noise are the same as those in Method IA. As shown in Table 17, high-confidence 27 and 30 hits for left- and right-sided DNA rearrangement events, respectively, passed the filter in edited cells, while only 0 and 12 hits were captured for left-sided and right-sided events, respectively, in control cells. However, in addition to more hits, the number of split reads supporting these hits was also much higher in edited cells in Figure 15. Table 17. Summary of DNA rearrangements according to Example 8 in control and CCR5-edited cells. [Table 17] Comparison of split reads between control and CCR5 edited cells shown in Figure 15.
[0306] The DNA rearrangements identified in Example 8 in CCR5-edited cells are shown in Table 18 below. The table includes the genomic coordinates of the DNA breakpoints between the DNA rearrangements, the number of unique split reads based on UMI-tools, the reads extracted to indicate the split reads, and notes to further explain the potential mechanism of this DNA rearrangement. It is noteworthy that DNA rearrangements between CCR5 and its homologous gene CCR2, as well as an interchromosomal translocation between CCR5 on chromosome 3 and RNF17 / CENPJ on chromosome 13, were also identified by CAST-seq (Turchiano et al., 2021). PCR was also used to verify the presence of fusion DNA between chr3 / chr13 and the acentric chromosome (Figure 13A-B). Therefore, the SAFER detection method II can efficiently capture both intrachromosomal and interchromosomal rearrangements. Table 18. Summary of DNA rearrangements in CCR5-edited cells according to Example 8. [Table 18-1] [Table 18-2] [Table 18-3]
[0307] Example 9: Exemplary Method of Method III Exemplary oligonucleotides are listed in Table 1.
[0308] Additional reagents include Tn5 transposase, e.g., Robust Tn5 transposase (catalog number EMQZ1422, Creative Biogene), DNA polymerase PCR master mix, e.g., Platinum™ SuperFi II PCR Master Mix (ThermoFisher, catalog number 12368010), DNA polymerase with PCR buffer, e.g., Q5 High-Fidelity DNA Polymerase (New England Biolabs, catalog number M0491S), dNTPs, e.g., dNTP mix (10 mM each) (ThermoFisher, catalog number R0194), DNA purification kit: column kit, e.g., DNA Clean&Concentrator-5 (Zymo Research, catalog number D4004), magnetic beads, e.g., AMPure XP Reagent (Beckman, catalog number A63880), RNA purification kit, e.g., RNA Clean&Concentrator-5 (Zymo Research, catalog number D4004). Research, catalog number R1013), Illumina sequencing kits such as NextSeq 1000 / 2000 P1 Reagents (300 cycles) (catalog number 20050264), MiSeq Reagent Kit v3 (150 cycles) (catalog number MS-102-3001), MiSeq Reagent Kit v2 (300 cycles) (catalog number MS-102-2002), PhiX Control v3 (Illumina, catalog number FC-110-3001), strand-displacing polymerases such as phi29 DNA polymerase (New England Biolabs, catalog number M0269S) or Bst 3.These include DNA polymerase (New England Biolabs, catalog number M0374S), DNase I (RNase-free) (New England Biolabs, catalog number M0303S), RNase H (New England Biolabs, catalog number M0297S), in vitro transcription kits such as the MAXIscript™ T7 Transcription Kit (ThermoFisher, catalog number AM1312), reverse transcription kits such as the SuperScript IV First-Strand Synthesis System (ThermoFisher, catalog number 18091050), STE buffer (10 mM Tris-HCl (pH 8.0), 1 mM EDTA, 0.1 M NaCl), and magnesium chloride (MgCl2) solution (New England Biolabs, catalog number B9021S).
[0309] In the first step, the Tn5 transposome is assembled. Transposon end-containing oligos (e.g., oligo-01 and oligo-02 for the non-UMI version, or oligo-26 and oligo-02 for the UMI version) are annealed in STE buffer using the following program: 90–99°C for 1–10 min, followed by 40–500 cycles of 90–99°C for 1 min (decrease the temperature 0.2–2°C per cycle until a temperature of 1–12°C is reached). The transposome complex is assembled in TPS buffer according to the following procedure, as recommended by the manufacturer (Creative Biogene, catalog number EMQZ1422). Mix the reagents thoroughly and incubate at 15–40°C for 5–120 min. Component Volume Tn5 transposase 1-10 μL Adapter 0.5~8μL 2μL of 10x TPS buffer Add sterilized water up to 20 μL
[0310] The workflow for this example is shown in Figure 5. A tagmentation reaction is performed in which genomic DNA extracted from the gene-edited sample or control (using standard methods) is tagmented with the transposome complex assembled according to the procedure. An exemplary reaction mixture is shown below: Component Volume Genomic DNA variation (5-500ng) 6μL of 5×LM buffer Transposome 0.1~8μL Add sterilized water up to 30 μL
[0311] The reaction components are mixed thoroughly and incubated at 30-65°C for 5-120 minutes. The transposase is then removed in the presence of 0.1-10 mM EDTA at 60-80°C for 5-120 minutes. The ends of the transposed fragment are filled in and extended with Q5 High-Fidelity DNA polymerase (New England Biolabs, catalog number M0491S) in the presence of 0.2-4 mM MgCl2 and 20-1000 μM dNTPs for 5-600 seconds at 60-75°C without supplementing the specific Q5 reaction buffer. The product is purified using column-based purification.
[0312] Then, perform in vitro transcription using the MAXIscript™ T7 Transcription Kit (ThermoFisher, catalog number AM1312) and the tagmented DNA from the previous step by incubating at 25-45 °C for 1-600 min. Perform the next step immediately.
[0313] Add TURBO DNase from the MAXIscript™ T7 Transcription Kit (ThermoFisher, Catalog No. AM1312) to the previous reaction mixture and incubate for 1-240 minutes at 20-45° C. Purify the RNA product using an RNA purification kit, such as RNA Clean & Concentrator-5 (Zymo Research, Catalog No. R1013).
[0314] DNA probe hybridization and subsequent RNase H treatment are then performed in two separate reactions for each target: one reaction to detect genomic rearrangements on the left side of the target, and the other reaction to detect genomic rearrangements on the right side of the target. The probe (e.g., oligo-13 for EMX1 left rearrangements, oligo-17 for EMX1 right rearrangements, oligo-21 for CCR5#1 left rearrangements, or oligo-25 for CCR5#1 right rearrangements) is added to the purified RNA product from the previous step in RNase H reaction buffer. The mixture is denatured at 85-99°C for 1-10 minutes and cooled to 4-28°C. RNase H is added to the reaction and incubated at 16-45°C for 1-300 minutes. RNase H is then inactivated at 60-90°C for 5-240 minutes. The product is purified using an RNA purification kit, such as RNA Clean & Concentrator-5 (Zymo Research, Cat. No. R1013).
[0315] Reverse transcription is then performed. Briefly, RNase H-treated RNA is reverse transcribed into DNA using target-specific primers (e.g., oligo-05 for the left-side rearrangement of EMX1, oligo-15 for the right-side rearrangement of EMX1, oligo-19 for the left-side rearrangement of CCR5#1, or oligo-23 for the right-side rearrangement of CCR5#1) and the SuperScript IV First-Strand Synthesis System (ThermoFisher, catalog number 18091050). First, the RNA, dNTPs, and primers are mixed and heated to 65°C for 1-20 min, then incubated at 1-12°C for 10-600 s for primer annealing. Next, 5x SSIV buffer, DTT, ribonuclease inhibitors, and SuperScript IV reverse transcriptase are added to the mixture, and the combined reaction mixture is incubated at 40-60°C for 2-120 min. The reaction is then inactivated by incubating at 60-90°C for 5-240 min. To remove RNA, add 0.1-10 µL of RNase H to the reaction mixture, mix, and incubate at 20-55 °C for 5-120 min. The product is purified using column-based purification.
[0316] Use the first-strand cDNA synthesized from the previous step in the first PCR with common primer 1 (oligo-04, common to all targets and both-side rearrangements), target-specific primer 2 (e.g., oligo-05 for EMX1 left-side rearrangements, oligo-15 for EMX1 right-side rearrangements, oligo-19 for CCR5#1 left-side rearrangements, or oligo-23 for CCR5#1 right-side rearrangements), and SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog no. 12368010) with the following program: initial denaturation at 90-99 °C for 1-5 min, 20-40 cycles of three-step denaturation-annealing-amplification (90-99 °C for 5-60 s, 55-75 °C for 5-60 s, 67-75 °C for 5-60 s), and a final extension at 67-75 °C for 1-10 min and a hold time of 4-12 °C. The PCR product is purified using column-based purification.
[0317] Next, perform nested PCR using the barcode tags. PCR is performed using the PCR product from the previous step, primer 3 (oligo-06, common to all targets and both sides of the rearrangement), and primer 4 (e.g., oligo-07 for EMX1 left-side rearrangement, oligo-16 for EMX1 right-side rearrangement, oligo-20 for CCR5#1 left-side rearrangement, or oligo-24 for CCR5#1 right-side rearrangement) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the following program: initial denaturation at 90-99 °C for 1-5 min, three-step denaturation-annealing-amplification (90-99 °C for 5-60 s, 55-75 °C for 5-60 s, 67-75 °C for 5-60 s) for 4-40 cycles, and a final extension at 67-75 °C for 1-10 min and a hold time of 4-12 °C. The PCR product is purified using column-based purification.
[0318] Next, tag PCR is performed using the sequencing primers. PCR is performed with the PCR product from the previous step, primer 5 (e.g., oligo-08, an example of an Illumina index primer from New England Bio Labs, NEB#E7603A), and primer 6 (e.g., oligo-09, an example of an Illumina index primer from New England Bio Labs, NEB#E7611A) using SuperFi II Polymerase PCR Master Mix (ThermoFisher, catalog number 12368010) with the following program: initial denaturation at 90-99°C for 1-5 min, three-step denaturation-annealing-amplification (90-99°C for 5-60 s, 55-75°C for 5-60 s, 67-75°C for 5-60 s) for 4-40 cycles, and a final extension at 67-75°C for 1-10 min and a 4-12°C hold. PCR products are purified using column-based purification.
[0319] Next, perform next-generation sequencing using the PCR products from the previous step, with optional ≥0–95% PhiX (Illumina, catalog number FC-110-3001) for sequencing using an Illumina platform (e.g., Illumina MiSeq or NextSeq 2000).
[0320] Paired-end sequencing reads from Illumina sequencing are merged using paired-end read merging software such as PEAR (Zhang et al., Bioinformatics. 2014;30:614-620), FLASH (Magoc and Salzberg, Bioinformatics. 2011;27:2957-2963), BBMerge (Bushnell et al., PLoS One. 2017;12:e0185056), or NGmerge (Gaspar, Bioinformatics. 2018;19:536). The merged reads are then trimmed and filtered using software such as BBmap (https: / / jgi.doe.gov / data-and-tools / software-tools / bbtools / ), samtools (Danecek et al., Gigascience. 2021;10.doi:10.1093 / gigascience / giab008), or custom scripts to remove Illumina adapter sequences, low-quality reads, reads containing target flanking sequences on the side of the target site where the blocking oligonucleotide or cleavage reagent binds, and reads that do not contain target flanking sequences on the side of the target site where the target-specific primer hybridizes. Selected reads are aligned to the human reference genome, e.g., GRCh37 (hg19), GRCh38 (hg38), or a telomere-to-telomere assembly (e.g., T2T-CHM13v2.0), using alignment software such as Bowtie2 (Langmead and Salzberg, Nat Methods. 2012;9:357-359), BWA (Li and Durbin, Bioinformatics. 2009;25:1754-1760), or Minimap2 (Li, Bioinformatics. 2018;34:3094-3100). The aligned BAM file is converted to a BED file using BEDTools (Quinlan, Bioinformatics. 2014;47:11.12.1-34).For methods using UMIs, collapse UMIs to remove redundant sequencing reads using software such as Gencore (Chen et al., Bioinformatics. 2019;20:606), UMI-tools (Smith et al., Genome Res. 2017;27:491-499), UMI-Reducer (github.com / smangul1 / UMI-Reducer), or custom scripts.
[0321] To quantify the number of rearrangements between the target site and other genomic loci, reads with candidate translocation breakpoints within an appropriate window (e.g., 1, 3, 5, 10, 20, 40, 60, 80, 100, 200, or more bases) adjacent to the target site are identified and counted. Statistical tests are then applied to establish confidence scores for each rearrangement group at each distal rearranged locus (e.g., within a defined window of 50, 100, 200, 500, 1000, 2000, 3000, or more bases at the distal rearranged locus). Visualization tools such as IGV (Thorvaldsdottir et al., Brief Bioinform. 2013;14:178-192) or Circos (Krzywinski et al., Genome Res. 2009;19:1639-1645) facilitate genome-wide identification and analysis of rearrangements.
[0322] Any reference cited herein is not admitted to constitute prior art. The discussion of a reference states what its author asserts, and the applicant reserves the right to challenge the accuracy and pertinence of the cited reference. Although many sources of information, including scientific journal articles, patent documents, and textbooks, are referenced herein, it should be clearly understood that this reference does not constitute an admission that any of these documents form part of the common general knowledge in the art.
[0323] The general method discussion provided herein is for illustrative purposes only. Other alternative methods and substitutes will be apparent to those skilled in the art upon review of this disclosure and are to be encompassed within the spirit and scope of this application.
[0324] Throughout this specification, reference is made to various patents, patent applications and other types of publications (e.g., journal articles, electronic database entries, etc.) The disclosures of all patents, patent applications and other publications cited herein are incorporated herein by reference in their entirety for all purposes. Unofficial sequence listing [Table 19-1] [Table 19-2] [Table 19-3] [Table 19-4] [Table 19-5] [Table 19-6]
Claims
1. 1. A method for detecting genome-wide rearrangements in a nucleic acid genome, comprising: (a) contacting a genomic DNA sample obtained from the cell or tissue contacted with the site-specific nuclease with a plurality of transposons comprising a first sequence tag at the 5' end of the transposon under conditions such that the plurality of transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagmented into a plurality of partially double-stranded nucleic acid fragments, each of the plurality of nucleic acid fragments comprising the sequence tag at the 5' end of the nucleic acid fragment; (b) performing a first amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region adjacent to a target site, (ii) a second primer comprising a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment, and (iii) a blocking oligonucleotide comprising a nucleotide sequence complementary to a region distal to the target site relative to the first target-specific primer to produce a primary amplification product comprising a nucleotide sequence that includes the rearranged target sequence; (c) sequencing the primary amplification products.
2. 2. The method of claim 1, wherein the sequence tag comprises a sequence that is separate from the genome and does not interact with the genome, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof.
3. 2. The method of claim 1, wherein the blocking oligonucleotide comprises an absent or blocked 3' OH, a spacer, an inverted nucleotide, or other modification to block extension of the 3' end.
4. 2. The method of claim 1, wherein the blocking nucleotide comprises one or more phosphorothioate bonds, spacers, or other modifications at the 3' and 5' ends to block exonuclease digestion at the 3' and 5' ends, LNA, BNA, PNA, RNA, DNA, modified nucleic acid, or a combination thereof.
5. 2. The method of claim 1, wherein the first and second primers comprise second and third sequence tags.
6. (c) prior to (c), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (b) and fourth and fifth sequence tags at the 5' ends of the nested primers to produce secondary amplification products comprising the rearranged target sequence and additional sequence tags; The method of claim 1 further comprising:
7. performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; The method of claim 6 further comprising:
8. The method of claim 6 , wherein the third and / or fourth primer comprises a barcode sequence.
9. The method of claim 7, wherein the fifth and sixth primers comprise a sequencing tag and / or an index sequence.
10. (c) prior to (c), performing a second hemi-nested amplification reaction using the second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (b), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce a secondary amplification product comprising the rearranged target sequence and one or two additional sequence tags; The method of claim 1 further comprising:
11. performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; The method of claim 10 further comprising:
12. 12. The method of claim 11 , wherein the fourth and fifth primers comprise a sequencing tag and / or an index sequence.
13. 10. The method of Claim 1, wherein (b) further comprises performing a separate amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
14. (b), prior to contacting the plurality of tagmented nucleic acid fragments with ddNTPs or other 3'-modified dNTPs and a DNA polymerase or terminal deoxynucleotidyl transferase to block all extendable 3' ends; The method of claim 1 further comprising:
15. The method of claim 2 , wherein the sequence tag comprises uracil.
16. 1. A method for detecting genome-wide rearrangements in a nucleic acid genome, comprising: (a) contacting a genomic DNA sample obtained from the cell or tissue contacted with the site-specific nuclease with a plurality of transposons comprising a first sequence tag at the 5' end of the transposon under conditions such that the plurality of transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagmented into a plurality of partially double-stranded nucleic acid fragments, each of the plurality of nucleic acid fragments comprising the sequence tag at the 5' end of the nucleic acid fragment; (b) contacting the plurality of tagmented nucleic acid fragments with a sequence-specific cleavage reagent; (c) performing a first amplification reaction to amplify the rearranged, tagmented nucleic acid fragments using (i) a first target-specific primer comprising a nucleotide sequence complementary to a region flanking a target site, and (ii) a second primer comprising a nucleotide sequence identical to a portion of the sequence tag located at the 5' end of the fragment; (d) sequencing the primary amplification products.
17. 17. The method of claim 16, wherein the sequence-specific cleavage reagent is an enzymatic reagent or a chemical cleavage agent.
18. 18. The method of claim 17, wherein the enzymatic cleavage agent comprises a CRISPR Cas / gRNA complex.
19. (b) followed by contacting the plurality of tagmented nucleic acid fragments with ddNTPs or other 3'-modified dNTPs and a DNA polymerase or deoxynucleotidyl transferase to block all extendable 3' ends; 17. The method of claim 16, further comprising:
20. 17. The method of claim 16, wherein the sequence tag comprises a sequence that is separate from the genome and does not interact with it, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof.
21. 17. The method of claim 16, wherein the first and second primers comprise second and third sequence tags.
22. (d) prior to (d), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (c) and additional sequence tags at the 5' ends of the nested primers to produce secondary amplification products comprising the rearranged target sequence and additional sequence tags; 17. The method of claim 16, further comprising:
23. performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; 23. The method of claim 22, further comprising:
24. 23. The method of claim 22, wherein the third and / or fourth primer comprises a barcode sequence.
25. 23. The method of claim 22, wherein the fifth and sixth primers comprise a sequencing tag and / or an index sequence.
26. (d) prior to (d), performing a second hemi-nested amplification reaction using the second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (c), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce a secondary amplification product comprising the rearranged target sequence and one or two additional sequence tags; 17. The method of claim 16, further comprising:
27. performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; 27. The method of claim 26, further comprising:
28. 27. The method of claim 26, wherein the second and / or third primer comprises a barcode sequence.
29. 28. The method of claim 27, wherein the fourth and fifth primers comprise a sequencing tag and / or an index sequence.
30. 17. The method of Claim 16, wherein (b) further comprises performing a separate amplification reaction to amplify the rearranged tagmented nucleic acid fragments using a second target-specific primer opposite the target site.
31. 1. A method for detecting genome-wide rearrangements in a nucleic acid genome, comprising: (a) contacting a genomic DNA sample obtained from a cell or tissue contacted with a site-specific nuclease with a plurality of transposons comprising a first sequence tag at a 5' end of the transposon, wherein the sequence tag comprises an RNA promoter sequence, under conditions such that the plurality of transposons are inserted into the genomic DNA sample and the genomic DNA sample is tagmented into a plurality of partially double-stranded nucleic acid fragments, each of the plurality of nucleic acid fragments comprising the sequence tag at a 5' end of the nucleic acid fragment; (b) transcribing the plurality of tagmented nucleic acid fragments into RNA; (c) contacting the RNA with a target sequence-specific DNA oligonucleotide probe and RNase H; (d) performing a reverse transcription amplification reaction using (i) a first primer comprising a nucleotide sequence complementary to a portion of the sequence tag located at the 3' or 5' end of the fragment and (ii) a second primer comprising a nucleotide sequence complementary to a region adjacent to a target site to produce a primary amplification product; (e) sequencing the primary amplification products.
32. 32. The method of claim 31 , wherein the sequence tag comprises a sequence that is separate from the genome and does not interact with it, a unique molecular identifier (UMI), a transposase recognition site, an index sequence, or a combination thereof.
33. 32. The method of claim 31 , wherein the first and second primers comprise sequence tags.
34. (e) prior to (e), performing a second amplification reaction using third and fourth nested primers comprising nucleotide sequences complementary to sequences proximal to the inner ends of the amplification products of (d) and an additional sequence tag to produce secondary amplification products comprising the rearranged target sequence and the additional sequence tag; 32. The method of claim 31 , further comprising:
35. performing a third amplification reaction using fifth and sixth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; 35. The method of claim 34, further comprising:
36. 35. The method of claim 34, wherein the third and / or fourth primer comprises a barcode sequence.
37. 36. The method of claim 35, wherein the fifth and sixth primers comprise an adapter sequence and / or an index sequence.
38. (e) prior to (e), performing a second hemi-nested amplification reaction using the second primer and a third nested primer, wherein the third primer comprises a nucleotide sequence complementary to a sequence proximal to the target-specific primer end of the amplification product of (d), and optionally fourth and fifth sequence tags at the 5' ends of the second and / or third primers, to produce a secondary amplification product comprising the rearranged target sequence and one or two additional sequence tags; 32. The method of claim 31 , further comprising:
39. performing a third amplification reaction using fourth and fifth primers comprising nucleotide sequences complementary to the amplification products of the second amplification reaction to produce further enriched tertiary amplification products comprising rearranged target sequences; 39. The method of claim 38, further comprising:
40. 39. The method of claim 38, wherein the second and / or third primer comprises a barcode sequence.
41. 40. The method of claim 39, wherein the fourth and fifth primers comprise an adapter sequence and / or an index sequence.
42. 32. The method of Claim 31, wherein (d) further comprises performing a separate amplification reaction to amplify said rearranged tagmented nucleic acid fragments using a second target-specific primer opposite said target site.