Methods for isolating double-strand breaks
INDUCE-seq avoids PCR amplification and directly measures DNA double-strand breaks using oligonucleotide characteristics, solving the signal distortion problem in existing technologies and achieving accurate detection and pattern identification of DSBs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV COLLEGE CARDIFF CONSULTANTS LTD
- Filing Date
- 2021-08-20
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies introduce significant biases during the PCR amplification stage when measuring DNA double-strand breaks (DSBs) in nucleic acid samples, leading to signal distortion and making it impossible to accurately determine the composition of the original DSBs.
The INDUCE-seq method was used to react nucleic acid samples with specific oligonucleotides under ligation conditions, avoiding PCR amplification and directly measuring DSB. The 5' and 3' binding characteristics of oligonucleotides were used to hybridize with sequencing primers to separate and sequence DSB, thereby enhancing the signal-to-noise ratio.
It enables direct measurement and pattern identification of DSB, and can simultaneously detect DSB caused by physiological and induced factors, thus improving the accuracy of measurement and signal-to-noise ratio.
Smart Images

Figure CN122484255A_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 202180072261.2 entitled "Method for Separating Double-Strand Breaks", filed on August 20, 2021. Technical Field
[0002] This invention relates to a method for determining the number and nature of DNA double-strand breaks (DSBs) in a nucleic acid sample, ideally genomic DNA (gDNA); a kit for carrying out the aforementioned method, comprising at least a plurality of oligonucleotides for ligating the nucleic acid sample; and oligonucleotides for use in the kit and the method. Background of the Invention Nucleic acid breaks (where one or more strands of the nucleic acid backbone are severed) are particularly harmful to cells and organisms because they can lead to genome rearrangements. Such breaks include double-strand breaks (DSBs) commonly found in DNA, which are the most dangerous of all DNA damage because they directly impair genome stability if not repaired. Besides causing cell death and the inability to properly repair them, DSBs can contribute to cancer by forming structural genomic alterations, including deletions, insertions, DNA translocations, and mitotic recombination events in somatic cells. In healthy cells, it is estimated that up to 50 endogenous DSBs are formed per cell during the cell cycle. These low-level physiological breaks occur occasionally as a result of normal cellular processes such as DNA replication, transcription, and chromatin circularization, or at higher levels due to repeated breaks programmed by the cell to facilitate processes during meiosis, such as V(D)J recombination. V(D)J recombination is a defining characteristic of the adaptive immune system. It is a unique mechanism of genetic recombination that occurs in developing lymphocytes during the early stages of B cell and T cell maturation. It involves somatic cell recombination and results in a highly diverse lineage of antibodies / immunoglobulins and T cell receptors (TCRs) found in B cells and T cells, respectively.
[0003] In addition to DSBs obtained through normal cellular processes, a variety of exogenous physical and chemical agents, such as ionizing radiation, chemotherapeutic drugs, and more recently, CRISPR genome editing technology, are also effective inducers of strand breaks (especially genomic DSBs).
[0004] Understanding the processes that generate and repair breaks in nucleic acids (especially the genome) is crucial for genomic medicine. From identifying the causes and cures of cancer to the safe development of genome editing technologies, precise and accurate measurements of the frequency, location, and causes of DNA double-strand breaks in the genome are paramount. However, simultaneously measuring the genome-wide landscape of endogenous and exogenous / induced DSBs in cells is challenging, primarily due to the vast range of rare, sporadic break events relative to frequently induced recurrence events.
[0005] The advent of next-generation sequencing (NGS) has spurred the development of various methods for detecting and measuring DNA sequences associated with nucleic acid strand breaks (especially DSBs) at the genome scale. These methods can be broadly categorized into three types: i) Indirect break markers that use proteins as substitutes for breakage (e.g., ƴH2AX ChIP-seq, DISCOVER-seq); ii) Indirect markers of the repaired fracture (e.g., GUIDE-seq, HTGTS), and iii) Direct marking of unrepaired broken ends in cells (e.g., BLESS, DSBCapture, END-seq, BLISS).
[0006] Despite numerous incremental technical improvements in each iteration of these methods, they all suffer from a common fundamental flaw: reliance on standard DNA library preparation required for NGS, following PCR labeling and enrichment of DSBs in the sample prior to sequencing. The PCR amplification stage of this method introduces significant bias into the DNA sequencing library, resulting in an indirect and distorted representation of the original pattern of DSBs present in the sample. This is a well-known phenomenon, making it impossible to quantify the original DSB composition in the sample. For many NGS applications, such as whole-genome / exome sequencing, this may not represent a significant problem. However, for quantitative, genome-wide measurements of specific features such as DSBs, PCR amplification bias introduces high levels of noise into systems where the signal (DSB) is already very low.
[0007] To overcome this drawback, we designed a novel DNA library preparation protocol that avoids the need for PCR amplification of fragmented sequences before DSB detection and further enhances the DSB signal. By improving the signal-to-noise ratio of DSB detection, we obtained a direct measurement of genomic fragments in the sample, where one sequence read is equivalent to a fragment present in the sample, ideally a DSB. Invention Overview In this paper, we describe our novel method, called INDUCE-seq, for the direct measurement of nucleic acid breaks (typically DSB). Furthermore, we demonstrate the ability of this method to identify break patterns that can characterize breaks formed by a variety of different physiological and induced causes.
[0008] INDUCE-seq can thus simultaneously detect the presence of low-level sporadic double-strand breaks caused by physiological transcription and DNA replication, as well as higher-level recurrent breaks induced by restriction enzymes or genome editing nucleases (such as CRISPR-Cas9). INDUCE-seq can therefore be used to determine the origins of DNA break formation and their repair mechanisms, as well as for the safe development of genome editing.
[0009] In one aspect of the invention, a method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, the method comprising: i) Under ligation conditions, a sample of nucleic acid suspected of containing DSB is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and comprises a 3' binding feature enabling the oligonucleotide to ligate to the second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. ii) Fragment the nucleic acid in the sample into fragments; iii) Exposing the fragment to a second pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the fragmented nucleic acid and a hybridization site (RD2 SP) thereto which a second sequencing primer can bind; and a second longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the fragmented nucleic acid, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features; iv) Denature the fragment to provide single-stranded nucleic acid; v) Divide the strand of part iv) into two groups: Group A, those fragments that have a first hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end; and Group B, those fragments that have not had a hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and have not had a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end; and vi) Sequencing of the strands in group A using primers that bind to the first and / or second hybridization sites, wherein each sequence is equivalent to a DSB break, and further wherein the number and nature of the base pair deletions can be determined by comparing each sequence with the genome of the species from which the sample was collected.
[0010] The second pair of oligonucleotides containing the 5' binding feature may not contain the binding sequence used to separate the fragmented nucleic acid.
[0011] The oligonucleotides in parts i) and iii) are interchangeable, so that after fragmentation in step ii), the nucleic acid is first exposed to the oligonucleotides in part iii) and then subsequently exposed to the oligonucleotides in part i).
[0012] Nucleic acid samples can be gDNA.
[0013] The 5' and / or 3' binding features may include one of the following: a phosphate group; a triphosphate "T-tail," preferably a deoxythymidine triphosphate "T-tail"; a triphosphate "A-tail," preferably a deoxyadenosine triphosphate "A-tail"; at least one random N nucleotide and multiple N nucleotides. The 5' and / or 3' protective features may include features providing resistance to any one or more of the following: phosphorylation activity, phosphatase activity, terminal transferase activity, nucleic acid hybridization, endonuclease activity, exonuclease activity, ligase activity, polymerase activity, and protein binding. Protective features may include a phosphorothioate linkage, a dideoxynucleotide or covalent block, phosphoramide, or a C3 spacer phosphoramide (3SpC3).
[0014] The 5' binding feature of the first oligonucleotide in part i) may be a phosphate group, and the 3' binding feature of the second oligonucleotide in part i) may be a triphosphate tail. The first and second oligonucleotides in part i) may also include index features, which are specific nucleotide sequences that enable the identification of the source of the pooled sample.
[0015] The first oligonucleotide of part i), read from 5' to 3', may include a 5' binding feature and subsequently, optionally, a protective feature, a hybridization site (RD1 SP) to which sequencing primers can bind, an index sequence, a binding sequence for separating the DSB from the DSB pool, and a 3' binding and / or protective feature. The binding feature may be a phosphate group. The second oligonucleotide of part i), read from 3' to 5', may include a 3' binding feature and subsequently, optionally, a hybridization sequence (RD1 SP) to which sequencing primers can bind, an index sequence, a binding sequence for separating the DSB from the DSB pool, and a 5' binding and / or protective feature. The 3' binding feature may include a 3' deoxythymidine triphosphate "T-tail" and a phosphate thioester linker.
[0016] The first and / or second oligonucleotide of the first oligonucleotide pair in part i) may contain two different terminal protective features. The 5' binding feature of the first oligonucleotide of the second oligonucleotide pair in part iii) may be a phosphate group, and the 3' binding feature of the second oligonucleotide in part iii) may be a triphosphate tail.
[0017] The first and second oligonucleotides of part iii) may also include index features, which are specific nucleotide sequences that enable determination of the source of the pooled sample.
[0018] The second oligonucleotide of part iii), read from 5' to 3', may include a 5' binding feature and subsequently optional protective features, optional additional sequences for bridging amplification, an index sequence, a hybridization site (RD2 SP) to which sequencing primers can bind, and a 3' binding and / or protective feature.
[0019] The first or second oligonucleotide of the second oligonucleotide pair in part iii) may contain two different terminal protective features.
[0020] The oligonucleotide pair of part i) may comprise a first oligonucleotide having SEQ ID NO. 1 and a second oligonucleotide having SEQ ID NO. 2; or an oligonucleotide sharing at least 80% identity or homology with SEQ ID NO. 1 or 2. The second pair of oligonucleotides of part iii) may comprise a first oligonucleotide having SEQ ID NO. 3 and a second oligonucleotide having SEQ ID NO. 4; or an oligonucleotide sharing at least 80% identity or homology with SEQ ID NO. 3 or 4. The second oligonucleotide of the second pair of oligonucleotides of part iii) may comprise any of the following sequences: SEQ ID NO. 4-28; or an oligonucleotide sharing at least 80% identity or homology with one of SEQ ID NO. 4-28.
[0021] The sample can be a mammal or a human. The ligation in part i) can occur in situ or in vitro using a cell or tissue sample. Prior to step i), the sample may be exposed to a permeabilizing agent. Prior to step i), the sample may be exposed to at least one reagent for arginine tail repair.
[0022] Part i) may also include extracting gDNA from the sample prior to proceeding with subsequent steps.
[0023] The method may further include, after part ii) and / or part iv), removing fragments whose size is less than about 100 bp or less than about 150 bp, and / or retaining fragments whose size is greater than about 150 bp.
[0024] The separation of part v) may involve using the binding sequence provided by the oligonucleotide of part i) to bind the partner body, and thus separate the group A strand of part iv) from any other strand.
[0025] The complementary binding strand of the binding sequence provided by the oligonucleotide of part i) can be anchored to the substrate, and the nucleic acid single strand flows through or through the anchored complementary binding strand.
[0026] Part vi) may involve bridging amplification, in which the single strand separated in part v) is cloned and amplified on a substrate to which the binding sequence of the first oligonucleotide in part i) and the additional sequence of the second oligonucleotide in part iii) are anchored.
[0027] Before proceeding with the claimed method, a sample containing or suspected of containing single-chain fractures may be joined or broken to ensure that the single-chain fracture is converted into a double-chain fracture.
[0028] In another aspect of the invention, a kit for identifying DNA double-strand breaks (DSBs) in gDNA samples is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to link to the first strand of the DSB, a hybridization site (RD1SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, complementary to the first oligonucleotide of the first pair, and comprising a 3' binding feature for binding to the second strand of the DSB; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to link to the first strand of the DSB, and a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second longer oligonucleotide, which is partially complementary to the first oligonucleotide of the second pair, and comprises a 3' binding feature for binding to the second strand of the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features.
[0029] In another aspect of the invention, a sample preparation kit for identifying DSB in a gDNA sample is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a strand enabling the oligonucleotide to be linked to a double-stranded nucleic acid and comprises a 5' binding feature according to the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31); and a second oligonucleotide complementary to the first oligonucleotide of the first pair; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides does not contain a sequence of more than 5, 10, 15 or 20 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and the second oligonucleotide comprises one strand of a double-stranded nucleic acid such that the oligonucleotide can be linked to one strand of the double-stranded nucleic acid, and comprises a 3' binding feature according to the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32); and wherein one or both of the oligonucleotides respectively comprise 3' and / or 5' protective features.
[0030] In another aspect of the invention, a sample preparation kit for identifying DSB in a gDNA sample is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a strand enabling the oligonucleotide to be linked to a double-stranded nucleic acid and comprises a 5' binding feature according to the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and a second oligonucleotide complementary to the first oligonucleotide of the first pair; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides does not contain a sequence of more than 5, 10, or 15 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31); and the second oligonucleotide comprises one strand of a double-stranded nucleic acid such that the oligonucleotide can be linked to it, and comprises a 3' binding feature according to the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34); and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively.
[0031] The first oligonucleotide of the first and / or second oligonucleotide pair may contain a hybridization site to which the first sequencing primer can bind. The first and second oligonucleotides of portions i) and / or ii) may contain index features, which are specific nucleotide sequences that enable the determination of the source of the pooled sample.
[0032] The kit may further contain at least one primer that binds to a first and / or second hybridization site for sequencing purposes. The kit may further contain fragmenting and / or denaturing agents, respectively, for fragmenting and / or denaturing nucleic acids into fragments and / or single strands.
[0033] In another aspect of the invention, a double-stranded adapter is provided for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, comprising: a first oligonucleotide chain including a 5' binding feature enabling the oligonucleotide to link to a first strand of the DSB, a hybridization site (RD1 SP) to which sequencing primers can bind, and a binding sequence for separating the DSB from a pool of DSBs; and a second oligonucleotide chain complementary to the first oligonucleotide and including a 3' binding feature for binding to a second strand of the DSB; wherein one or both of the oligonucleotides include 3' and / or 5' protective features.
[0034] In another aspect of the invention, a double-stranded linker is provided for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, comprising: a first oligonucleotide chain including a 5' binding feature that enables the oligonucleotide to link to a first strand of the DSB and a hybridization site (RD2 SP) thereto which sequencing primers can bind; and a second, longer oligonucleotide chain that is partially complementary to the first oligonucleotide chain and includes a 3' binding feature for binding to a second strand of the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides include 3' and / or 5' protective features.
[0035] The adaptor may contain first and second oligonucleotides containing index features, which are specific nucleotide sequences that enable determination of the source of the pooled sample.
[0036] In another aspect of the invention, a double-stranded linker for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, comprising: a first oligonucleotide chain that does not contain a sequence of more than 5, 10, 15, or 20 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and a second oligonucleotide comprising a 3' binding feature that enables the oligonucleotide to be linked to one strand of the double-stranded nucleic acid and comprises a 3' binding feature according to the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32); and wherein one or both of the oligonucleotides respectively comprise 3' and / or 5' protective features.
[0037] In another aspect of the invention, a double-stranded linker for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, comprising: a first oligonucleotide chain that does not contain a sequence of more than 5, 10, or 15 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31); and a second oligonucleotide comprising a strand that enables the oligonucleotide to be linked to a double-stranded nucleic acid, and comprising a 3' binding feature according to the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34); and wherein one or both of said oligonucleotides respectively comprise 3' and / or 5' protective features.
[0038] In another aspect of the invention, a sample preparation method for identifying DNA double-strand breaks (DSBs) in nucleic acid samples is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing immobilized primers, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor comprises an oligonucleotide at the 3' end of one strand of a DSB capable of ligating to the DSB and said oligonucleotide comprises a sequence capable of binding to primers immobilized on a substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor contains an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide does not contain a sequence capable of binding to primers immobilized on a substrate by hybridization.
[0039] The oligonucleotide in step d) may contain the same sequence as the region of the second primer. The substrate may contain the first and second immobilized primers; the oligonucleotide in step b) may contain a sequence capable of binding to the first immobilized primer by hybridization; and the oligonucleotide in step d) may contain the same sequence as the region of the second immobilized primer.
[0040] In one embodiment, step b) is: exposing a plurality of nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of the DSB, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and the first oligonucleotide is ligable to the 3' end and comprises a sequence capable of binding to primers immobilized on the substrate by hybridization; and step d) is: Multiple nucleic acids are exposed to a second adaptor pair under conditions that facilitate ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break, but not to a first oligonucleotide of the first adaptor pair, wherein the second adaptor contains oligonucleotides that are complementary to the first and second portions, and the first oligonucleotide is ligated to the 5' end and contains the same sequence as the region of the second primer, and the second oligonucleotide does not contain a sequence complementary to the sequence that is the same as the region of the second primer.
[0041] In one implementation of the method, the second connector pair includes: An oligonucleotide containing a first oligonucleotide of the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32), and a second oligonucleotide containing a sequence of more than 5, 10, 15, or 20 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or a sequence that does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); or The first oligonucleotide contains the sequence according to AATGATACGGCGACCACCGA (SEQ ID NO: 34), and the second oligonucleotide contains a sequence of more than 5, 10, or 15 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or a sequence of all 20 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31).
[0042] The first and / or second oligonucleotides of the first adaptor pair may contain 3' and / or 5' protective features; and / or wherein the first and / or second oligonucleotides of the second adaptor pair contain 3' and / or 5' protective features. In one embodiment, the second adaptor cannot be linked to the first adaptor due to the presence of a 3' modification of the first adaptor.
[0043] In one embodiment, the oligonucleotide that can be linked to the second adaptor at the 5' end contains the same sequence as the immobilizing primer with 5, 10, 15, 20, 21, 24 or more bases.
[0044] The method may further include denaturing multiple nucleic acids to form multiple single-stranded nucleic acids. The method may further include contacting the multiple nucleic acids with a substrate containing the fixed primers under conditions suitable for hybridization of the immobilized primers and complementary nucleic acids. The method may further include obtaining the sequence information of any nucleic acid hybridized to the substrate.
[0045] In one implementation of the method, the sample containing multiple nucleic acids is gDNA.
[0046] In some embodiments, the steps are performed in the order of a), b), c), and then d). In other embodiments, the steps are performed in the order of a), c), d), and then b); wherein the sample is exposed to conditions that can cause or are suspected of causing DSB between steps d) and b).
[0047] In some implementations, the sample is exposed to conditions that can induce DSB at target features in the nucleic acid sample.
[0048] In another aspect of the invention, a sample preparation method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing a fixed first primer, the method comprising: 1) Provide samples containing multiple nucleic acids; 2) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide at the 3' end of a strand capable of ligating to a DSB and said oligonucleotide contains a sequence capable of hybridizing with a second primer; 3) Fragmenting multiple nucleic acids; 4) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of ligating to the 5' end of one strand at a fragmentation-induced break but not to the first adaptor, and said oligonucleotide comprises the same sequence as the region of the immobilized first primer; and 5) Under conditions suitable for primer extension, contact multiple nucleic acids with the second primer.
[0049] In one implementation, step 2) is: exposing a plurality of nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of the DSB, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and the first oligonucleotide is ligable to the 3' end and comprises a sequence capable of hybridizing with a second primer; and step 4) is: Multiple nucleic acids are exposed to a second adaptor pair under conditions that facilitate ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break but not to a first oligonucleotide of the first adaptor pair, wherein the second adaptor contains an oligonucleotide complementary to the first and second portions, and the first oligonucleotide is ligable to the 5' end and contains the same sequence as the region of the immobilized first primer, and the second oligonucleotide does not contain a sequence complementary to the sequence of the same region as the immobilized first primer.
[0050] The second adaptor pair may contain a first oligonucleotide with the sequence according to AACCCACTACGCCTCCGCTTTCC (SEQ ID NO: 40); and a sequence of more than 5, 10, 15, or 20 bases that does not contain the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36), or a second oligonucleotide that does not contain all 22 bases of the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36).
[0051] The first and / or second oligonucleotides of the first adaptor pair may contain 3' and / or 5' protective features; and / or wherein the first and / or second oligonucleotides of the second adaptor pair contain 3' and / or 5' protective features. In one embodiment, the second adaptor cannot be linked to the first adaptor due to the presence of a 3' modification of the first adaptor.
[0052] In one embodiment, the oligonucleotide that can be linked to the second adaptor at the 5' end contains the same sequence as the immobilizing primer with 5, 10, 15, 20, 21, 24 or more bases.
[0053] The method may further include denaturing multiple nucleic acids to form multiple single-stranded nucleic acids. The method may further include contacting the multiple nucleic acids with a substrate containing a fixed first primer under conditions suitable for hybridization of the immobilized first primer with a complementary nucleic acid. The method may further include obtaining sequence information of any nucleic acid hybridized to the substrate. The sample included may be gDNA.
[0054] In one embodiment, the steps are performed in the order of 1), 2), 3), 4), and then 5). In another embodiment, the steps are performed in the order of 1), 3), 4), 2), and then 5); wherein the sample is exposed to conditions that can cause or are suspected of causing DSB between steps 4) and 2).
[0055] In one implementation, the sample is exposed to conditions that can induce DSB at a target feature in the nucleic acid sample.
[0056] In another aspect of the invention, a sample preparation method for identifying a target feature in a nucleic acid sample is provided, wherein the preparation includes modifying a nucleic acid associated with the target feature to be suitable for binding to a substrate containing immobilized primers, the method comprising: a) Provide a sample containing multiple nucleic acids, expose the multiple nucleic acids to conditions that can cleave at least one strand of the nucleic acid at the target feature, and denature the multiple nucleic acids into single-stranded nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor comprises an oligonucleotide at the 3' end of a strand capable of ligating to a cleavage site and said oligonucleotide comprises a sequence capable of binding to a primer immobilized on a substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor contains an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide does not contain a sequence capable of binding to primers immobilized on a substrate by hybridization.
[0057] The target feature can be any feature that can be specifically cleaved. The target feature can be a cyclobutane pyrimidine dimer (CPD), 8-oxoguanine, or a debasing site.
[0058] The oligonucleotide in step d) may contain the same sequence as the region of the second primer. The substrate may contain the first and second immobilized primers; the oligonucleotide in step b) may contain a sequence capable of binding to the first immobilized primer by hybridization; and the oligonucleotide in step d) may contain the same sequence as the region of the second immobilized primer.
[0059] In one embodiment, step b) is: exposing a plurality of nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of a cleavage site, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and the first oligonucleotide is ligable to the 3' end and comprises a sequence capable of binding to a primer immobilized on a substrate via hybridization; and wherein step d) is: Multiple nucleic acids are exposed to a second adaptor pair under conditions that facilitate ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break but not to a first oligonucleotide of the first adaptor pair, wherein the second adaptor contains an oligonucleotide complementary to the first and second portions, and the first oligonucleotide is ligable to the 5' end and contains the same sequence as the region of the second primer, and the second oligonucleotide does not contain a sequence complementary to the sequence identical to the region of the second primer.
[0060] In one implementation, the second connector pair includes: An oligonucleotide containing a first oligonucleotide of the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32), and a second oligonucleotide containing a sequence of more than 5, 10, 15, or 20 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or a sequence that does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); or The first oligonucleotide contains the sequence according to AATGATACGGCGACCACCGA (SEQ ID NO: 34), and the second oligonucleotide contains a sequence of more than 5, 10, or 15 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or a sequence of all 20 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31).
[0061] The first and / or second oligonucleotides of the first adaptor pair may contain 3' and / or 5' protective features; and / or wherein the first and / or second oligonucleotides of the second adaptor pair may contain 3' and / or 5' protective features. In one embodiment, the second adaptor cannot be linked to the first adaptor due to the presence of a 3' modification of the first adaptor.
[0062] The oligonucleotide that can be linked to the second adaptor at the 5' end may contain the same sequence as the immobilizing primer, consisting of 5, 10, 15, 20, 21, 24 or more bases.
[0063] In one embodiment, the method further includes denaturing multiple nucleic acids to form multiple single-stranded nucleic acids. In one embodiment, the method further includes contacting the multiple nucleic acids with a substrate containing the immobilized primers under conditions suitable for hybridization of the immobilized primers and complementary nucleic acids. In one embodiment, the method further includes obtaining sequence information of any nucleic acid hybridized to the substrate.
[0064] In one embodiment, the sample containing multiple nucleic acids is gDNA. In one embodiment, the steps are performed in the order of a), b), c), and then d). In another embodiment, the steps are performed in the order of c), d), a), and then b).
[0065] In another aspect of the invention, a double-stranded linker for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, is provided, comprising: a first oligonucleotide chain that does not contain a sequence of more than 5, 10, 15, or 20 bases of the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36), or does not contain all 22 bases of the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36); and a second oligonucleotide comprising a 3' binding feature that enables the oligonucleotide to be linked to one strand of the double-stranded nucleic acid and comprises a 3' binding feature according to the sequence AACCCACTACGCCTCCGCTTTCC (SEQ ID NO: 40); and wherein one or both of said oligonucleotides respectively comprise 3' and / or 5' protective features. Attached Figure Description
[0066] Figure 1 • Overview of INDUCE-seq. In situ break labeling in immobilized and permeabilized cells is achieved by ligating a full-length chemically modified P5 sequencing adaptor to pre-prepared DSB ends. Genomic DNA is then extracted, fragmented, end-prepared, and ligated using a chemically modified semi-functional P7 adaptor. The resulting DNA library contains a mixture of functional DSB-tagged fragments (P5:P7) and non-functional genomic DNA fragments (P7:P7). Subsequent sequencing of the INDUCE-seq library enriches the DNA-tagged fragments and eliminates all other non-functional DNA. Because INDUCE-seq library preparation is PCR-free, each sequencing read obtained is equivalent to a single labeled DSB end. Figure 2• Detailed schematic diagram of the flow cell enrichment adapter design. (a) Structure of the complete adapter-ligated dsDNA fragments used for sequencing. The 3' P5 and P7 adapters hybridize with the flow cell. Sequencing primers bind to the Read 1 sequencing primer (RD1 SP) and Read 2 sequencing primer (RD2 SP) sequences during the first and second sequencing reads. The index allows differentiation of different samples in the library pool. (b) Structure of DNA fragments present in the INDUCE-seq library. Only DSB-ligated fragments consist of all the adapter components required for sequencing. (c) Loading the INDUCE-seq library into the sequencing flow cell will enrich the DSB-ligated fragments via hybridization at the 3' end of the P5 adapter sequence. No other fragments can interact with the flow cell and will be removed; Figure 3• INDUCE-seq demonstrates unprecedented sensitivity and dynamic range compared to alternative DSB sequencing technologies. (a) INDUCE-seq simultaneously detects highly recurrent induced DSBs and single endogenous DSBs at high resolution. After in situ digestion with the restriction endonuclease HindIII, the genome browser view of INDUCE-seq reads is mapped to a 10mb portion of the genome from HEK293T cells. (Top) At high levels of observation (10mb, 0-1000 reads), highly recurrent enzyme-induced breaks represent the vast majority of reads. (Bottom) Zooming in (pink highlight, 500kb, 0-20 reads) shows low levels of single endogenous breaks (green highlight) in untreated samples and in recurrent HindIII-induced breaks. (b) Mapping of INDUCE-seq reads at HindIII target sites demonstrates the precision of single nucleotide break mapping. The figure includes TACTCAAGCTTACCCCTA (SEQ ID NO: 35) and GGGGGGTAAGCTTGAGTA (SEQ ID NO: 43). (c) Quantification of breaks measured per cell in HindIII-treated and control samples. INDUCE-seq quantitatively detects breaks per cell across three orders of magnitude between samples. (d and e) Comparison between INDUCE-seq and DSBCapture in detecting restriction sites cleaved in vitro by the enzymes HindIII and EcoRV. (d) Mapping a larger proportion of sequenced and genome-aligned reads to restriction sites using INDUCE-seq. (e) Using 1 / 800 of the cells, INDUCE-seq identified a HindIII restriction site proportion (92.7%) similar to that identified by DSBCapture (93.7%). (f) Dynamic range of induced DSB detection using INDUCE-seq. In addition to the breaks identified at the HindIII target sequence (AAGCTT), multiple 1bp and 2bp off-target sites were identified. The number of induced break events measured by INDUCE-seq spanned eight orders of magnitude, from ~150 million breaks identified at the HindIII target site to 5 breaks identified at the rarest off-target sites. (g) Comparison between INDUCE-seq, DSBCapture, and BLISS in detecting AsiSI-induced breaks in live DiVA cells. The number of sequenced reads (top) is compared to the number of AsiSI sites identified in each experiment (bottom). INDUCE-seq sensitively detected the largest number of AsiSI sites using 1 / 40th the reads of DSBCapture and 1 / 23rd the reads of BLISS; Figure 4 • INDUCE-seq sensitively detects and quantifies CRISPR / Cas9-induced on-target and off-target DSBs. (a) Off-target sequences and number of fragments of EMX1 sgRNA identified using INDUCE-seq. This figure includes GAGTCCCGAGCAGAAGAAGAANGG (SEQ ID NO: 44). (b) INDUCE-seq shows the kinetics of EMX1-induced DSB formation in cell populations. Quantification of the number of fragments detected per million reads per sample shows high on-target and off-target Cas9 activity occurring immediately after nuclear transfection. (c) Comparison of off-targets identified by INDUCE-seq with established in vitro methods CIRCLE-seq and DiGenome-seq, and cell-based methods GUIDE-seq, BLISS, and HTGTS. INDUCE-seq detected many off-targets that were previously only detectable by in vitro methods. Significantly more off-target sites were identified compared to any current cell-based method. INDUCE-seq also identified several off-targets that could not be detected by any other method. (d) Amplicon sequencing was performed to measure the frequency of off-target insertions and deletions identified by INDUCE-seq. Amplicon sequencing was able to identify only 4 of the 60 off-targets identified by INDUCE-seq, and was limited by a background insertion / deletion false detection rate of 0.1%. These findings are consistent with previous studies that measured the frequency of EMX1 off-target insertions and deletions 48 hours after EMX1 RNP nuclear transfection; Figure 5 • Comparison of INDUCE-seq with current DSB mapping workflows. (a) Overview of the INDUCE-seq workflow. Sequencing of INDUCE-seq libraries produces quantitative output, where one read is equivalent to one break. (b) Overview of DSBCapture, BLISS, and END-seq workflows. Sequencing after standard library construction produces output, where one read is not equivalent to a single DSB; Figure 6 • Comparison between the number of sequenced reads and the number of DSBs defined for INDUCE-seq and BLISS NGS libraries. (a) Scatter plot of the number of sequenced INDUCE-seq reads and the number of breaks defined by a single INDUCE-seq experiment. (b) Scatter plot showing the number of BLISS reads sequenced from a single BLISS experiment with the correct read 1 (R1) barcode prefix and the number of breaks defined after removing duplicates using UMI correction; Figure 7Genome browser view of DSB hotspots in HEK293 cells. (a) 11kb view of ch17 DSB hotspot. Purple arrows indicate DSB ends labeled on the right (+ strand), while blue arrows indicate DSB ends labeled on the left (- strand). Recurrent DSBs are evenly distributed throughout the hotspot region. (b) 5kb view of chr11 DSB hotspot. Recurrent DSBs can be detected at different locations on the positive and negative strands; Figure 8 A schematic diagram of the off-target detection process used in INDUCE-seq. Figure 8 The sequence in the image is SEQ ID NO: 44; Figure 9 • CRISPR off-target detection using INDUCE-seq is highly reproducible. (a) Comparison between the number of EMX1 off-targets detected in r1 and r2 time-history experiments. (b) Scatter plot showing the number of breakages found at CRISPR off-target sites identified in two independent experiments; Figure 10 • Venn diagrams showing the cross-cutting off-target results identified by INDUCE-seq, CIRCLE-seq, GUIDE-seq, and BLISS. (a and b) Overlap of individual samples from 0 to 30 hours, calculated from independent experiments r1 (a) and r2 (b). (c and d) Combined overlap of all time points from sets r1 (c) and r2 (d). (e) Overlap calculated between methods when all INDUCE-seq samples are combined; Figure 11 • CRISPR-induced DSB patterns at target and off-target sites correlate with editing results. EMX1 shows a 180bp coverage trajectory across the target (a) and top-ranking off-target sites (sequences shown from top to bottom are SEQ ID NO: 45-65), OT-1 (sequences shown from top to bottom are SEQ ID NO: 66-87) (b), and OT-2 (sequences shown from top to bottom are SEQ ID NO: 88-105) (c). Close-up views of a 40bp region around each target site show a unique 1bp protruding cut pattern, rather than the usual Cas9-induced blunt-end DSB. The corresponding insertion / deletion spectra at each site show the mutation locations associated with the observed breakpoints; Figure 12 An exemplary embodiment of an exemplary chemical modification is shown.
[0067] Figure 13 An exemplary embodiment in which a semi-functional adaptor is linked before a fully functional adaptor. This figure illustrates an embodiment for detecting artificially induced DSBs, for example, for detecting off-target DSBs induced by nucleases.
[0068] Figure 14 An exemplary workflow for applying the method of the present invention to bead sequencing technology.
[0069] Figure 15 Exemplary adapters for use with microbead sequencing technology.
[0070] Figure 16 An exemplary implementation for detecting DNA damage affecting one strand. This implementation involves the ligation of a semi-functional adaptor preceding a fully functional adaptor.
[0071] Figure 17 An exemplary implementation of in vitro detection of CRISPR base editing-cytosine base editor (CBE) is shown.
[0072] Figure 18 An exemplary implementation of in vitro detection of CRISPR base editing-adenine base editor (ABE) is shown.
[0073] Table 1. Example connecting subsequences of preferred embodiments of the present invention. Invention Details In one aspect of the invention, a method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, the method comprising: i) Under ligation conditions, a sample of nucleic acid suspected of containing DSB is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and comprises a 3' binding feature enabling the oligonucleotide to ligate to the second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. ii) Fragment the nucleic acid sample (ideally gDNA) into fragments; iii) Under ligation conditions, the fragment is exposed to a second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the fragment (ideally at a binding site remote from the first pair of oligonucleotides), and a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second, longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the fragment or the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. iv) Denature the fragment to provide a single nucleic acid; v) Divide the strand of part iv) into two groups: Group A, those fragments that have a first hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end; and Group B, those fragments that have not had a hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and have not had a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end; and vi) Sequencing of the strands in group A using primers that bind to the first and / or second hybridization sites, wherein each sequence is equivalent to a break, typically a DSB. Optionally further, the number and nature of the base pair deletions can be determined by comparing each sequence with the genome of the species from which the sample was collected.
[0074] The steps of the method of the present invention do not need to be performed at one time or in one location. For example, in one aspect, the method of the present invention allows for the preparation of nucleic acid samples that have been labeled in a manner that allows for the selection and isolation of DSBs. Thus, in one aspect of the present invention, a sample preparation method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, comprising: i) Under ligation conditions, a sample of nucleic acid suspected of containing DSB is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and comprises a 3' binding feature enabling the oligonucleotide to ligate to the second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. ii) Fragmenting the nucleic acid sample (ideally gDNA) into fragments; and iii) Exposing the fragment to a second pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to bind to the first strand of the fragment (optionally at a binding site remote from the first pair of oligonucleotides), and a hybridization site (RD2 SP) thereto which a second sequencing primer can bind; and a second longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the fragment, a sequence complementary to the hybridization site, and an additional sequence, optionally a binding sequence for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features.
[0075] Alternatively, the method of the present invention can be performed until the strands are separated into groups A and B, and can be sequenced separately. Thus, this method allows for the preparation of nucleic acid samples in which any DSB has been selected and isolated. Therefore, in one embodiment, a sample preparation method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, comprising: i) Under ligation conditions, a sample of nucleic acid suspected of containing DSB is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and comprises a 3' binding feature enabling the oligonucleotide to ligate to the second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. ii) Fragment the nucleic acid sample (ideally gDNA) into fragments; iii) Exposing the fragment to a second pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the fragment (optionally at a binding site remote from the first pair of oligonucleotides), and a hybridization site (RD2 SP) thereto which a second sequencing primer can bind; and a second longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the fragment, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. iv) Denature the fragment to provide a single nucleic acid; and v) Divide the strand of part iv) into two groups: group A, those fragments that have a first hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end, and group B, those fragments that have not a hybridization site and binding sequence provided by the oligonucleotide of part i) at the first end and have not a second hybridization site and additional sequence provided by the oligonucleotide of part iii) at the other end.
[0076] In one aspect of the invention, a method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, comprising: i) Exposing a sample of nucleic acid suspected of containing DSB to a first pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to a first strand of the DSB, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide that is complementary to the first oligonucleotide of the first pair; wherein one or both of the oligonucleotides comprises a 3' and / or a 5' protective feature. ii) Fragmenting nucleic acid samples (e.g., gDNA) into fragments; iii) Exposing the fragment to a second pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the second pair of oligonucleotides is complementary to the second oligonucleotide of the second pair; and a second oligonucleotide comprising a 3' binding feature for binding to the second strand of the fragment, and optionally a binding sequence for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features; iv) Denature the fragment to provide a single nucleic acid; v) Divide the strand of part iv) into two groups: group A, those fragments that have a binding sequence provided by the oligonucleotide of part i) attached to the first end and optionally a binding sequence provided by the oligonucleotide of part iii) for bridging amplification attached to the other end; and group B, those fragments that have not a binding sequence provided by the oligonucleotide of part i) attached to the first end and optionally have not a binding sequence provided by the oligonucleotide of part iii) for bridging amplification attached to the other end; and vi) Sequencing of the strands in group A, wherein each sequence is equivalent to a break, typically a DSB, and optionally further, wherein the number and nature of the base pair deletions can be determined by comparing each sequence with the genome of the species from which the sample was collected. As described above, the method may include steps i), ii), and iii); or i), ii), iii), iv), and v); or all of these steps.
[0077] In a particular implementation, the second pair of oligonucleotides containing the 5' binding feature does not contain a binding sequence for separating the fragmented nucleic acid.
[0078] In an embodiment of the invention, the oligonucleotides of portion i) and portion iii) can be interchanged, thereby exposing the nucleic acid first to the oligonucleotide of portion iii) and subsequently to the oligonucleotide of portion i).
[0079] In one embodiment of the invention, the oligonucleotides of portion i) and portion iii) are interchangeable, thereby first exposing the nucleic acid to the oligonucleotide of portion iii), and subsequently exposing the nucleic acid fragment to the oligonucleotide of portion i) after fragmentation.
[0080] Implementations involving the interchange of parts i) and iii) are particularly relevant to methods for detecting induced DSB. In such implementations, the sample may be fragmented, and step iii) may then be performed. Subsequently, the sample may be processed to potentially induce DSB, and after induction, step i) may be performed. Implementations in which DSB is introduced are further discussed herein, including implementations for detecting off-target effects of nucleases, etc. Figure 13 , 17 Figures 1 and 18 illustrate specific implementation schemes.
[0081] In the method of the present invention, either of the oligonucleotide pairs may be or may be referred to as an adaptor.
[0082] In this article, the term "adaptor" (or "adaptor head") refers to a linker in genetic engineering, and it is an oligonucleotide that can be attached to the end of another DNA molecule. Adaptors can be double-stranded. Double-stranded adaptors can be synthesized with blunt ends on both ends, or with one end being a sticky end and the other a blunt end. Adaptors can be short and can be chemically synthesized.
[0083] The first pair of oligonucleotides can be configured to allow binding to the substrate via hybridization with an oligonucleotide immobilized on the substrate, wherein the immobilized oligonucleotide is oriented such that the 5' end is proximal to the immobilization site and the 3' end is distal to the immobilization site. For example, the first pair of oligonucleotides (i.e., oligonucleotides containing the 5' binding feature) at the 3' end of a strand attached to the DSB can be at least partially complementary to the immobilized oligonucleotide. The degree of complementarity allows binding to the immobilized oligonucleotide via hybridization. The second pair of oligonucleotides can be configured to disallow binding to the substrate via hybridization with an oligonucleotide immobilized on the substrate, wherein the immobilized oligonucleotide is oriented such that the 5' end is proximal to the immobilization site and the 3' end is distal to the immobilization site. For example, the second pair of oligonucleotides (i.e., oligonucleotides containing the 5' binding feature) at the 3' end of a strand attached to the fragmentation site can be insufficiently complementary to any immobilized oligonucleotide to prevent binding via hybridization. In some embodiments, for example to allow subsequent bridging amplification, the second pair of oligonucleotides (i.e., oligonucleotides containing the 3' binding feature) binding to the 5' end of one strand of the fragmentation site may be at least partially identical in sequence to the immobilized oligonucleotides. In another embodiment, for example to allow microbead emulsion amplification, other arrangements may be employed (see, for example, see...). Figure 14 ).
[0084] In one embodiment, a sample preparation method for identifying DNA double-strand breaks (DSBs) in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing immobilized oligonucleotides, the method comprising: i) Under ligation conditions, a sample of nucleic acid suspected of containing DSB is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to ligate to a first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool, wherein the binding sequence is at least partially complementary to the immobilized oligonucleotide; and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and comprises a 3' binding feature enabling the oligonucleotide to ligate to a second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. ii) Fragmenting the nucleic acid sample (ideally gDNA) into fragments; and iii) Exposing the fragment to a second pair of oligonucleotides under ligation conditions, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the fragment (ideally at a binding site remote from the first pair of oligonucleotides), a hybridization site (RD2 SP) to which a second sequencing primer can bind, and wherein the first oligonucleotide does not contain a sequence capable of hybridizing with a fixed oligonucleotide; and a second, longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the fragment, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features.
[0085] DSB-associated nucleic acids are nucleic acids located on one side of the DSB. Therefore, sequencing of DSB-associated nucleic acids allows for the identification of the DSB, for example, its location within the genome.
[0086] In one embodiment, the method of the present invention is designed for use with Illumina P5 and P7 adaptors. Thus, in a particular embodiment, the second oligonucleotide pair does not contain a sequence of more than 5, 10, 15, or 20 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30). In another embodiment, the second oligonucleotide pair does not contain a sequence of more than 5, 10, or 15 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31).
[0087] Therefore, in a particular implementation, step iii) is: Under ligation conditions, the fragment is exposed to a second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the fragment (ideally at a binding site distant from the first pair of oligonucleotides), and a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second, longer oligonucleotide that is partially complementary to the first oligonucleotide of the second pair and comprises a 3' binding feature for binding to the second strand of the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features; and wherein the second pair of oligonucleotides does not contain a sequence of more than 5, 10, 15, 20, or 24 bases as specified in SEQ ID NO: 30 and / or does not contain a sequence of more than 5, 10, 15, or 20 bases as specified in SEQ ID NO: 31.
[0088] In a preferred method of the present invention, the 5' or 3' binding feature of the oligonucleotide pair comprises one of the following: a phosphate group; a triphosphate "T tail", preferably a deoxythymidine triphosphate "T tail"; a triphosphate "A tail", preferably a deoxyadenosine triphosphate "A tail"; at least one random N nucleotide, preferably multiple N nucleotides, or any other known binding group that allows the adaptor to be linked to the DSB.
[0089] The nucleic acid sample can be any DNA sample capable of containing a DSB. In a preferred embodiment of the invention, the nucleic acid sample is gDNA.
[0090] Ideally, the 5' binding feature of the first oligonucleotide in part i) is a phosphate group, while the 3' binding feature of the second oligonucleotide in part i) is a triphosphate tail.
[0091] In a preferred embodiment of the invention, the 5' and / or 3' protective features of the first pair of oligonucleotides include features providing resistance to one or more of the following: phosphorylation activity, phosphatase activity, terminal transferase activity, nucleic acid hybridization, endonuclease activity, exonuclease activity, ligase activity, polymerase activity, and protein binding. This can be achieved by any means known to those skilled in the art, such as, but not limited to, phosphate thioester linkages, phosphoramide spacers, phosphate groups, 2'-O-methyl groups, reverse deoxy or dideoxy T modifications, locked nucleic acid bases, dideoxy nucleotides, etc. Examples of the activities provided by these features are shown in Table 2.
[0092] Preferably, the first oligonucleotide of portion i) comprises a 3' protective feature that provides resistance to exonuclease activity, such as a phosphate thioester linker. Additionally or alternatively, the 3' protective feature also provides resistance to ligase activity and / or polymerase activity (ideally 5'>3' polymerase activity), and is, for example, a dideoxynucleotide or physical block (ideally in phosphoramide form, particularly C3 spacer phosphoramide (3SpC3)), or any other protective feature known to those skilled in the art, such as those in Table 2 that provide resistance to exonuclease activity and / or ligase activity and / or polymerase activity.
[0093] More preferably, the first and second oligonucleotides in portion i) further comprise an index feature, which is a specific nucleotide sequence (e.g., GATCT) that enables the determination of the source of the assembled sequencing library; in other words, it enables demultiplexing of the assembled sequencing library. Ideally, the index feature is located between the hybridization site and the binding sequence.
[0094] Most preferably, the first oligonucleotide in portion i), read from 5' to 3', includes a 5' binding feature and subsequently, optionally, a protective feature (ideally, the binding feature is a phosphate group), a hybridization site (RD1 SP) to which a first sequencing primer can bind, an index sequence, a binding sequence for separating the DSB from the DSB pool, and a 3' binding and / or protective feature. Preferably, the protective feature provides resistance to any one or more of the following: exonuclease activity, ligase activity, and / or polymerase activity.
[0095] Most preferably, the second oligonucleotide in portion i), read from 3' to 5', includes a 3' binding feature and subsequently, optionally, a protective feature, a hybridization sequence (RD1 SP) to which a first sequencing primer can bind, an index sequence, a binding sequence for separating the DSB from the DSB pool, and a 5' binding and / or protective feature. Preferably, the protective feature provides resistance to any one or more of the following: exonuclease activity, ligase activity, and / or polymerase activity. Ideally, the binding feature is 3'.
[0096] In a preferred embodiment of the invention, the first oligonucleotide of portion i) is one of a first oligonucleotide pair, and the second oligonucleotide of the first oligonucleotide pair is complementary to the first oligonucleotide and includes 5' and 3' protective features, preferably providing resistance to any one or more of the following: exonuclease activity, ligase activity, and / or polymerase activity. Thus, in some embodiments, the oligonucleotide used to implement the invention lacks a 5' or 3' phosphate group.
[0097] In a preferred embodiment of the invention, the first oligonucleotide of portion i) contains a hybridization site ( / sequence) for sequencing the ligated DSBs and a binding sequence for separating the DSBs from the DSB pool. However, those skilled in the art will understand that the second oligonucleotide of portion i) may contain a hybridization site ( / sequence) for sequencing the ligated DSBs and a binding sequence for separating the DSBs from the DSB pool. This is because both sequencing and separation will be determined by the nature of the primers used for sequencing and the nature of the oligonucleotides used for separation. Most typically, the orientation of all oligonucleotides, i.e., the hybridization sites (RD1 SP and RD2 SP) to which at least the first and / or second sequencing primers can bind, the binding sequence for separating the DSBs from the DSB pool, and optionally the orientation of additional sequences for achieving bridging amplification, such that when separation is performed in portion v), only the A group strands can be extracted using the binding sequence for separating the DSBs from the DSB pool. Furthermore, this orientation also ensures that when sequencing is performed in portion vi), only the A group strands can be bridged (when a bridging amplification sequence is present).
[0098] Ideally, the first or second oligonucleotide of part i) of the first oligonucleotide pair contains two different terminal protective features.
[0099] More preferably, the second oligonucleotide of the first oligonucleotide pair in part i) contains a 3'-deoxythymidine triphosphate "T-tail" to provide a substrate for linking to an "A+ tail" DNA fragment, and also ideally contains a phosphate thioester linker, thereby conferring resistance to exonuclease activity.
[0100] In a further preferred method of the invention, the 5' binding feature of the first oligonucleotide of the second oligonucleotide pair in part iii) is a phosphate group, and the 3' binding feature of the second oligonucleotide of the second oligonucleotide pair in part iii) is a triphosphate tail.
[0101] In a preferred embodiment of the invention, the 5' and / or 3' protective features of the second pair of oligonucleotides include features providing resistance to one or more of the following: phosphorylation activity, phosphatase activity, terminal transferase activity, nucleic acid hybridization, endonuclease activity, exonuclease activity, ligase activity, polymerase activity, and protein binding. This can be achieved by any means known to those skilled in the art, such as, but not limited to, phosphate thioester linkages, phosphoramide spacers, phosphate groups, 2'-O-methyl groups, reverse deoxy or dideoxy T modifications, locked nucleic acid bases, dideoxy nucleotides, etc. Examples of the activities provided by these features are shown in Table 2.
[0102] Preferably, the first oligonucleotide of portion iii) further comprises a 3' protective feature, such as a phosphoramide spacer, providing resistance to exonuclease activity. Additionally or alternatively, the 3' protective feature also provides resistance to ligase activity and / or polymerase activity (ideally 5'>3' polymerase activity), and is, for example, a dideoxynucleotide or physical block (ideally in phosphoramide form, particularly C3 spacer phosphoramide (3SpC3)), or any other protective feature known to those skilled in the art, such as those in Table 2 that provide resistance to exonuclease activity and / or ligase activity and / or polymerase activity.
[0103] More preferably, the first and second oligonucleotides of portion iii) further comprise an index feature, which is a specific nucleotide sequence (e.g., GATCT) that enables the determination of the source of the assembled sequencing library; in other words, it enables the splitting of the assembled sequencing library. Ideally, the index feature is located between the hybridization site and the additional sequence (if present).
[0104] Ideally, the first or second oligonucleotide of the second oligonucleotide pair in part iii) contains two different terminal protective features.
[0105] Most preferably, the second oligonucleotide in portion iii), read from 5' to 3', includes a 5' binding and / or protective feature, an optional additional sequence for bridging amplification, an index sequence, a hybridization site (RD2 SP) to which sequencing primers can bind, and a 3' binding and / or protective feature. Preferably, one or two protective features provide resistance to any one or more of the following: exonuclease activity, ligase activity, and / or polymerase activity.
[0106] In a preferred embodiment of the invention, the second oligonucleotide of portion iii) is complementary to the first oligonucleotide of the oligonucleotide pair and includes 5' and 3' protective features, preferably providing resistance to any one or more of the following: exonuclease activity, ligase activity, and / or polymerase activity. Thus, in some embodiments, the oligonucleotide used to implement the invention lacks a 5' or 3' phosphate group.
[0107] More preferably, the second oligonucleotide of part iii) contains a 3'-deoxythymidine triphosphate "T-tail" to provide a substrate for linking to "A-tail" DNA fragments, ideally with a thiophosphate linker, thereby conferring resistance to exonuclease activity.
[0108] In a preferred embodiment of the invention, the ligation in part i) occurs in situ or in vitro using a cell or tissue suspension, and thus occurs within intact cells. More preferably, the oligonucleotide to be ligated is brought into access to the DSB site by chemical, electronic, mechanical, or physiological permeation of the cell or tissue, thereby facilitating the ligation via its terminal binding characteristics (e.g., phosphate). Preferably, permeation of the cell is performed by incubation in a lysis buffer. More preferably, the DSB site undergoes arginine tail repair prior to ligation with the oligonucleotide.
[0109] As will be understood, linking the first pair of oligonucleotides / adaptors to the DSB prior to further processing ensures that the identified DSB is a real event rather than the result of subsequent processing steps (and therefore represents an artifact of processing).
[0110] In a further preferred embodiment of the invention, part i) further includes extracting gDNA from the cells using any conventional means (such as extraction buffer) prior to performing subsequent steps.
[0111] In a preferred embodiment of the invention, step ii) comprises fragmenting the gDNA into smaller fragments by any means known in the art, such as sonication or tagging.
[0112] In a further preferred embodiment of the invention, the method further includes, after portions ii) and / or iv), the optional step of removing fragments smaller than about 100 bp, more preferably smaller than about 150 bp, and retaining fragments larger than about 150 bp. Those skilled in the art will understand that this step advantageously removes any oligonucleotide chains / dimers that may have formed and would otherwise subsequently contribute to sequence artifacts. Ideally, this can be done using conventional means (such as using a Bioruptor sonicator) and size selection using SPRI microbeads (GC Biotech, CNGS-0005) to remove fragments <150 bp. Alternatively, fragments within a preferred range can be selected, such as, but not limited to, 150-1000 bp, ideally 200-800 bp, more ideally 250-750 bp, and most preferably 300-500 bp.
[0113] In a further preferred embodiment of the invention, the separation of portion v) involves using the binding sequence provided by the oligonucleotide of portion i) to bind the partner, and thus separate the group A strand of portion iv) from any other strand. Typically, the complementary binding strand of the binding sequence provided by the oligonucleotide of portion i) is anchored to the substrate, and the nucleic acid single strand flows through or through the anchored complementary binding strand.
[0114] In a further preferred method of the invention, part vi) involves bridging amplification, wherein the single strand isolated in part v) is cloned and amplified on a substrate to which the oligonucleotide / binding sites of the first oligonucleotide binding sequence in part i) and the additional sequence of the second oligonucleotide in part iii) are anchored. In this way, the single strand of group A fragments can be bound to both ends of the substrate to facilitate bridging amplification. More preferably, the sequencing can be performed using synthetic sequencing with labeled nucleotides, each of which emits a characteristic signal, which is read as the sequence extends to provide readout of sequence information. Typically, multiple strands are sequenced in parallel. If desired, the index sequence can be sequenced separately from the DSB, thereby providing an indication of the nucleic acid origin before sequencing the DSB. Alternatively, both indices can be sequenced and thus read together.
[0115] In one particular embodiment, the first pair of oligonucleotides in part i) comprises a first oligonucleotide containing the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31) and a second oligonucleotide containing the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34).
[0116] In another embodiment that may be combined with the subject matter of the foregoing paragraphs, the second pair of oligonucleotides comprises a first oligonucleotide containing a sequence of more than 5, 10, 15, or 20 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or a sequence of all 24 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), and a second oligonucleotide containing a sequence according to CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32).
[0117] In one particular embodiment, the first pair of oligonucleotides in part i) comprises a first oligonucleotide containing the sequence according to ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30) and a second oligonucleotide containing the sequence according to CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32).
[0118] In another embodiment that may be combined with the subject matter of the foregoing paragraphs, the second pair of oligonucleotides may comprise a sequence of more than 5, 10, or 15 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or a first oligonucleotide that does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), and a second oligonucleotide that comprises the sequence according to AATGATACGGCGACCACCGA (SEQ ID NO: 34).
[0119] SEQ ID NO: 30, 31, 32 and 34 may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2 or 1 modification, such as substitution, deletion or insertion. In one embodiment, the modification is substitution.
[0120] In a further preferred embodiment of the invention, the first pair of oligonucleotides in portion i) comprises a first oligonucleotide having SEQ ID NO. 1 and a second oligonucleotide having SEQ ID NO. 2, or sharing at least 80% identity or homology with them, and more preferably, in an enhanced priority order, oligonucleotides sharing 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identity or homology with them. The oligonucleotides according to SEQ ID NO: 1 and SEQ ID NO: 2 may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution. Homology as used herein may be referred to as similarity.
[0121] In a further preferred embodiment of the invention, the first oligonucleotide pair of portion i) comprises a first oligonucleotide of the sequence GATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT[index]TCGGTGGTCGCCGTATCATTC, or the first oligonucleotide comprises 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. The first oligonucleotide pair of portion i) may further comprise a second oligonucleotide of the sequence AATGATACGGCGACCACCGA[index]ACACTCTTTCCCTACACGACGCTCTTCCGATCT, or the second oligonucleotide comprises 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion.
[0122] In a further preferred embodiment of the invention, the first oligonucleotide pair of portion i) comprises the sequence GATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT (SEQ ID NO: 37), an index, and TCGGTGGTCGCCGTATCATTC (SEQ ID NO: 38). The first oligonucleotide pair of portion i) may further comprise the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34), an index, and ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 39). The index may be any base (n) and may be, for example, 5 to 15 or 6 to 10 base pairs in length. SEQ ID NOs: 34, 37, 38, and 39 may contain 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0123] In a further embodiment of the invention, the first oligonucleotide pair of portion i) comprises a first oligonucleotide containing SEQ ID NO. 3, an index, and the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30). The first oligonucleotide pair of portion i) may further comprise a second oligonucleotide containing, in a 5' to 3' sequence, the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32), an index, and the sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 33). The index may be any base (n) and may be, for example, 5 to 15 or 6 to 10 base pairs in length. SEQ ID NOs: 3, 30, 32, and 33 may contain 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0124] In a further preferred embodiment of the invention, the second pair of oligonucleotides in portion iii) comprises a first oligonucleotide having SEQ ID NO. 3 and a second oligonucleotide having the sequence CAAGCAGAAGACGGCATACGAGAT[index]GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT or sharing at least 80% identity or homology with it, and more preferably, in an enhanced priority order, sharing 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identity or homology with it. The oligonucleotide according to SEQ ID NO: 3 may contain 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2 or 1 modification, such as substitution, deletion or insertion. The oligonucleotide according to CAAGCAGAAGACGGCATACGAGAT[index]GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0125] In a further preferred embodiment of the invention, the second oligonucleotide pair of portion iii) comprises a first oligonucleotide having SEQ ID NO. 3, and a second oligonucleotide having the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32) in a 5' to 3' sequence, an index, and the sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 33). The index can be any base (n) and can be, for example, 5 to 15 or 6 to 10 base pairs in length. The oligonucleotide may comprise a sequence having 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identity or homology with any of SEQ ID NO: 3, 32, or 33. SEQ ID NO: 3, 32, or 33 may contain 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0126] In a further preferred embodiment of the invention, the second oligonucleotide of the second pair of oligonucleotides in portion iii) comprises any one of the following sequences SEQ ID NO: 4-28, 30-34, 37-39, or 41-42, or shares at least 80% identity or homology with them, and more preferably, in an enhanced priority order, shares 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% identity or homology with them. SEQ ID NO: 4-28, 30-34, 37-39, or 41-42 may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0127] In a preferred embodiment, the sample is a mammalian sample, ideally a human sample.
[0128] According to another aspect of the present invention, a kit ideally suited for identifying DNA double-strand breaks (DSBs) in gDNA samples is provided, comprising: i) A first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to link to the first strand of the DSB, a hybridization site (RD1SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide, complementary to the first oligonucleotide of the first pair, and comprising a 3' binding feature for binding to the second strand of the DSB; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to be linked to the first strand of the DSB, and a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second longer oligonucleotide, which is partially complementary to the first oligonucleotide of the second pair, and comprises a 3' binding feature for binding to the second strand of the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively.
[0129] In a preferred kit of the present invention, the kit further comprises at least one primer that binds to a first and / or second hybridization site for sequencing purposes.
[0130] More preferably, the kit further comprises fragmenting agents and / or denaturing agents for fragmenting and / or denaturing nucleic acids into fragments and / or single strands, respectively.
[0131] In another aspect of the invention, a kit suitable for identifying DNA double-strand breaks (DSBs) in gDNA samples is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature enabling the oligonucleotide to be linked to a first strand of the DSB and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide complementary to the first oligonucleotide of the first pair; and wherein one or both of the oligonucleotides respectively comprise 3' and / or 5' protective features; and ii) a second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides is partially complementary to the second oligonucleotide of the second pair; and a second oligonucleotide comprising a 3' binding feature for binding to the second strand of the DSB and a binding sequence for achieving bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively.
[0132] In one aspect of the invention, a kit suitable for sample preparation for identifying DSBs in gDNA samples is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a strand enabling the oligonucleotide to be linked to a double-stranded nucleic acid and comprises a 5' binding feature according to the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31); and a second oligonucleotide complementary to the first oligonucleotide of the first pair; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides does not contain a sequence of more than 5, 10, 15 or 20 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and the second oligonucleotide comprises one strand of a double-stranded nucleic acid such that the oligonucleotide can be linked to one strand of the double-stranded nucleic acid, and comprises a 3' binding feature according to the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32); and wherein one or both of the oligonucleotides respectively comprise 3' and / or 5' protective features.
[0133] In another aspect of the invention, a kit suitable for sample preparation for identifying DSBs in gDNA samples is provided, comprising: i) a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a strand enabling the oligonucleotide to be linked to a double-stranded nucleic acid and comprises a 5' binding feature according to the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and a second oligonucleotide, which is complementary to the first oligonucleotide of the first pair and optionally comprises a 3' binding feature; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively; and ii) A second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides does not contain a sequence of more than 5, 10, or 15 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), and optionally contains a 5' binding feature; and the second oligonucleotide comprises one strand of a double-stranded nucleic acid such that the oligonucleotide can be linked to one strand of the double-stranded nucleic acid, and contains a 3' binding feature according to the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34); and wherein one or both of the oligonucleotides contain 3' and / or 5' protective features, respectively.
[0134] SEQ ID NO: 30, 31, 32 and 34 may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2 or 1 modification, such as substitution, deletion or insertion. In one embodiment, the modification is substitution.
[0135] The second oligonucleotide in part i) may contain a 3' binding feature, and the first oligonucleotide in part ii) may contain a 5' binding feature. The first oligonucleotide of the first and second pairs of oligonucleotides may include a hybridization site (RD1 SP) to which the sequencing primer can bind.
[0136] According to another aspect of the invention, a double-stranded linker is provided for ideally identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, comprising: A first oligonucleotide chain comprising a 5' binding feature enabling the oligonucleotide to link to a first strand of the DSB, a hybridization site (RD1 SP) to which sequencing primers can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide chain complementary to the first oligonucleotide and comprising a 3' binding feature for binding to a second strand of the DSB; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively.
[0137] According to another aspect of the invention, a double-stranded linker is provided for ideally identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, comprising: A first oligonucleotide chain comprising a 5' binding feature enabling the oligonucleotide to link to a first strand of the DSB and a hybridization site (RD2 SP) to which sequencing primers can bind; and a second, longer oligonucleotide chain partially complementary to the first oligonucleotide chain and comprising a 3' binding feature for binding to a second strand of the DSB, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features, respectively.
[0138] According to another aspect of the invention, a double-stranded linker suitable for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, is provided, comprising: The first oligonucleotide chain does not contain a sequence of more than 5, 10, 15, or 20 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or does not contain all 24 bases of the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30); and the second oligonucleotide comprises a 3' binding feature that enables the oligonucleotide to be linked to a strand of a double-stranded nucleic acid and comprises a 3' binding feature according to the sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32); and one or both of the oligonucleotides respectively comprise 3' and / or 5' protective features. The first oligonucleotide may comprise a 5' binding feature.
[0139] According to another aspect of the invention, a double-stranded linker suitable for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, is provided, comprising: The first oligonucleotide chain does not contain a sequence of more than 5, 10, or 15 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or does not contain all 20 bases of the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31); and the second oligonucleotide comprises one strand enabling the oligonucleotide to be linked to a double-stranded nucleic acid, and includes a 3' binding feature according to the sequence AATGATACGGCGACCACCGA (SEQ ID NO: 34); and one or both of the oligonucleotides respectively include 3' and / or 5' protective features. The first oligonucleotide may include a 5' binding feature.
[0140] In one aspect of the invention, a sample preparation method for identifying DSBs in nucleic acid samples is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing immobilized primers, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor comprises an oligonucleotide at the 3' end of one strand of a DSB capable of ligating to the DSB and said oligonucleotide comprises a sequence capable of binding to primers immobilized on a substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of attaching to the 5' end of one strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide does not contain a sequence capable of binding to a primer immobilized on a substrate via hybridization. The oligonucleotide of the second adaptor capable of attaching to the 5' end of one strand at a fragmentation-induced break may contain the same sequence as the region of the second primer.
[0141] In some embodiments, the nucleic acids in a sample containing multiple nucleic acids are double-stranded during steps b) to d). In other embodiments, the nucleic acids may be single-stranded during steps b) to d). In still further embodiments, the nucleic acids may be double-stranded for some steps and single-stranded for others (see, for example, [link to relevant documentation]). Figure 16 In embodiments involving the ligation of a double-stranded adaptor to a single-stranded nucleic acid, the adaptor may comprise a "splint oligo." The splint oligo may comprise random nucleotides, such as 6-8 random nucleotides, located at the 3' end of the oligonucleotide of the adaptor that is not ligated to the nucleic acid sample. The splint oligo facilitates the ligation process.
[0142] In one embodiment, a sample preparation method for identifying DSBs in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for substrate binding containing a first immobilized primer and suitable for amplification including the use of the first immobilized primer and a second primer, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide at the 3' end of one strand capable of ligating to a DSB, and said oligonucleotide contains a sequence capable of binding to a primer immobilized on a substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor contains an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide contains the same sequence as the region of the second primer.
[0143] Nucleic acid samples suitable for amplification, including those using a first fixed primer and a second primer, are samples capable of being amplified using primers with the same sequence as the first fixed primer and primers with the same sequence as the second primer. Amplification suitability can be determined in solution.
[0144] In some implementations, the second primer is not immobilized on the substrate. For example, the substrate may be microbeads, and amplification may be performed via microbead emulsion amplification.
[0145] In other embodiments, the second primer is immobilized on a substrate. For example, the substrate may be a flow cell, and amplification may be performed via bridge amplification.
[0146] In some implementations, the adaptor is a single-stranded oligonucleotide.
[0147] In other embodiments, the adaptor comprises at least partially complementary first and second oligonucleotides. In such embodiments, the first adaptor pair is capable of attaching to at least the 3' end of one strand of the DSB, and the first adaptor pair comprises at least partially complementary first and second oligonucleotides, wherein the first oligonucleotide is attachable to the 3' end and comprises a sequence capable of binding to a primer immobilized on the substrate via hybridization. Furthermore, in such embodiments, the second adaptor pair is capable of attaching to at least the 5' end of one strand at a fragmentation-induced break, but not to the first oligonucleotide of the first adaptor pair, wherein the second adaptor comprises first and second partially complementary oligonucleotides, wherein the first oligonucleotide is attachable to the 5' end and comprises a sequence identical to the region of the second primer, and the second oligonucleotide does not contain a sequence complementary to the sequence identical to the region of the second primer.
[0148] Therefore, in one embodiment, a sample preparation method for identifying DNA DSB in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing a first immobilized primer and suitable for amplification using the first immobilized primer and a second primer, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing multiple nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of the DSB, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and the first oligonucleotide is ligated to the 3' end and comprises a sequence capable of binding to a primer immobilized on the substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor pair under conditions conducive to ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break, but not to the first oligonucleotide of the first adaptor pair, wherein the second adaptor contains oligonucleotides complementary to the first and second portions, wherein the first oligonucleotide is ligable to the 5' end and contains the same sequence as the region of the second primer, and the second oligonucleotide does not contain a sequence complementary to the sequence identical to the region of the second primer.
[0149] In some embodiments, the substrate comprises a first immobilized primer and a second immobilized primer. In some embodiments, the immobilized primer may be adapted to function as primers during bridging amplification.
[0150] Thus, in one embodiment, a sample preparation method for identifying DNA DSB in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing first and second immobilization primers, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide at the 3' end of a strand capable of ligating to a DSB and the oligonucleotide contains a sequence capable of binding to a first immobilized primer by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor contains an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide contains the same sequence as the region of the second immobilization primer.
[0151] In other embodiments, the adaptor is an adaptor pair comprising at least partially complementary first and second oligonucleotides. Thus, a sample preparation method for identifying DNA DSBs in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate comprising first and second immobilization primers, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing a plurality of nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of the DSB, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and wherein the first oligonucleotide is ligable to the 3' end and comprises a sequence capable of binding to a first immobilized primer by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor pair under conditions conducive to ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break but not to a first oligonucleotide of the first adaptor pair, wherein the second adaptor contains oligonucleotides complementary to the first and second portions, and the oligonucleotide ligated to the 5' end contains the same sequence as the region of the second immobilized primer, and wherein another oligonucleotide does not contain a sequence complementary to the sequence of the region of the second immobilized primer.
[0152] As described herein, in some embodiments, the nucleic acids in the sample can remain as double-stranded molecules during steps a) to d). Therefore, in a particular embodiment, the method may further include: Multiple double-stranded nucleic acids are denatured to form multiple single-stranded nucleic acids. This may be "step e)" in some implementations.
[0153] In one particular implementation, the method may further include: Under conditions suitable for hybridization of the immobilized primers and complementary nucleic acids, multiple single-stranded nucleic acids are contacted with a substrate containing the immobilized primers. This may be referred to as "step f)" in some implementations.
[0154] Step i) of the method disclosed herein also applies to step b). Step ii) of the method disclosed herein also applies to step c). Step iii) of the method disclosed herein also applies to step d). Step iv) of the method disclosed herein also applies to step e). Step v) of the method disclosed herein also applies to step f).
[0155] The substrate can be a solid surface, such as a flow cell, microbeads, glass slide, or membrane surface. In particular, the substrate can be a flow cell. The substrate can be a patterned or unpatterned flow cell. The substrate can comprise glass, quartz, silica, metal, ceramic, or plastic. The substrate surface can comprise a polyacrylamide matrix or coating.
[0156] As used herein, the term "flow cell" is intended to have the general meaning in the art, particularly in the field of sequencing by synthesis. Exemplary flow cells include, but are not limited to, those used in nucleic acid sequencing devices, such as flow cells of the Genome Analyzer®, MiSeq®, NextSeq®, HiSeq®, or NovaSeq® platforms commercially available from Illumina, Inc. (San Diego, Calif.); or flow cells of the SOLiD™ or Ion Torrent™ sequencing platforms commercially available from Life Technologies (Carlsbad, Calif.). Exemplary flow cells and methods of their manufacture and use are also described, for example, in WO2014 / 142841A1; U.S. Patent Application Publication No. 2010 / 0111768 A1 and U.S. Patent No. 8,951,781.
[0157] The substrate may contain immobilized primers, such as two types of primers that can together act as forward and reverse primers for bridging amplification. Immobilization to the substrate means that the primers remain bound to the substrate even under conditions that would denature the double-stranded nucleic acid. For example, the primers may be covalently bound to the substrate. The primers are oriented such that the 5' end is proximal to the immobilization site, and the 3' end is distal to the immobilization site. This arrangement is standard in the art.
[0158] The steps can be performed in the following order: step c), step d), and then step b). This order is particularly relevant to embodiments where the DSB is potentially induced in the sample, in contrast to embodiments for detecting pre-existing DSBs. Referring to the statement of the characteristic steps of the invention defined in Roman numerals, the steps can be performed in the following order: step ii), step iii), and then step i). In such embodiments, the DSB can be induced after the connection of the second connector and before the connection of the first connector. See, for example, [link to relevant documentation]. Figure 13 Induction of DSB may include exposing a sample to conditions that can cause or are suspected of causing DSB.
[0159] In an embodiment where step d) is performed prior to step b), the connector in step d) may include 3' and / or 5' protective features to prevent the connectors of step b) from connecting to those of step d). These protective features prevent the first connector from connecting to the second connector. In such embodiments, the second connector remains unable to connect to the first connector because the connection occurred before the first connector was present.
[0160] In some embodiments, the steps are performed in the order of steps b), c), and d), wherein step b) is performed in cells or in situ. In other embodiments, the steps are performed in the order of steps c), d), and subsequently step b), wherein step c) is performed in vitro after the nucleic acid sample is isolated.
[0161] Therefore, in one embodiment, a sample preparation method for identifying DSBs in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing immobilized primers, the method comprising: a) Provide a sample containing multiple nucleic acids; c) Fragmenting multiple nucleic acids; d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break and said oligonucleotide does not contain a sequence capable of binding to a primer immobilized on a substrate by hybridization; optionally containing a sequence identical to the region of the second primer; and b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide capable of ligating to the 3' end of one strand of a DSB but not to a second adaptor, and the oligonucleotide contains a sequence capable of binding to primers immobilized on a substrate by hybridization.
[0162] In step a), the sample containing multiple double-stranded nucleic acids can be any DNA sample capable of containing DSBs, such as gDNA.
[0163] For any of the methods disclosed herein, the sample may contain DSB or may have been treated in a manner that may or does introduce DSB.
[0164] For example, the methods disclosed herein can be used to detect DNA damage or alterations limited to single strands. Thus, the methods disclosed herein can be used to detect and / or quantify target features in nucleic acid samples. For instance, damage in one strand of double-stranded DNA can be enzymatically converted to DSB, which can then be detected using the methods disclosed herein.
[0165] In some instances, the damage can be a single-strand break. In other instances, the damage is a base change in a nucleic acid sample. For example, the methods disclosed herein can be used to detect CRISPR / Cas-induced base editing, such as cytosine base editors or adenosine base editors. Such editing can be converted to DSB and detected as disclosed herein. Figure 17 and Figure 18 An exemplary implementation is shown.
[0166] Thus, in one embodiment, the method includes providing a sample containing a plurality of nucleic acids, said nucleic acids containing or suspected of containing a target feature or damage; and exposing the sample to conditions capable of converting the target feature or damage into a DSB.
[0167] For such implementations, the method may follow the order of: transforming the damage into DSB, b), c), and d); or it may follow the order of: c), d), transforming the damage into DSB, and b).
[0168] In other embodiments, the sample has been treated with a nuclease, such as a transcription activator-like effector nuclease (TALEN), CRISPR / Cas endonuclease, zinc finger nuclease, a wide range of nucleases, or any restriction endonuclease, and the method can be used to detect off-target effects. The sample can be treated with an agent (such as a potential therapeutic agent) to determine whether the agent can induce DSB or off-target DSB.
[0169] The methods disclosed herein can be used to identify sites of protein-nucleic acid binding. For example, a sample can be contacted with a protein binder (such as an antibody) specific to a target protein, where the target protein may bind to DNA. The DNA can be a sample that has already been contacted with the target protein. The protein binder can associate directly or indirectly with a nuclease, thus forming a DSB at any site of binding to the target protein. Any DSB can then be detected by the methods disclosed herein. The method can be a nuclease-targeted cleavage and run (CUT&RUN) method. Thus, the method may include: step a); contacting the sample with the target protein; contacting the sample with a nuclease capable of directly or indirectly associating with the target protein to form a DSB in the protein-bound nucleic acid; steps b), c), and subsequently d). This sequence is particularly useful for embodiments in which the method is performed in cells up to step b). Alternatively, the method may include: step a); step c); step d); contacting the sample with the target protein; contacting the sample with a nuclease capable of directly or indirectly associating with the target protein to form a DSB in the protein-bound nucleic acid; and subsequently step b). This sequence is particularly useful for in vitro embodiments.
[0170] In other embodiments, DSB can be intentionally induced at known or target sites. For example, targeting nucleases, such as CRISPRR / Cas or TALEN, can be used to induce DSB in a target sequence, making the method of the present invention usable for isolating the target sequence for further analysis. The target sequence can be a specific gene, a whole exome, or a specific locus in the genome. Thus, in one embodiment, the method includes contacting a sample with a nuclease capable of targeting and inducing DSB in a target sequence.
[0171] The method of this invention can be used to detect viral or bacterial insertion events or DNA damage caused by viral or bacterial insertion events. These events can be measured using unique genetic sequences associated with the bacteria and viruses that will be inserted into the genome. The method of this invention can therefore reveal the sites of these foreign DNA insertions (which may occur via DSB intermediate structures).
[0172] In other embodiments, the methods disclosed herein can be used to assess risks associated with therapeutic agents, such as gene therapy. For example, the methods disclosed herein can be applied to samples taken from a patient, where the sample has been treated with a therapeutic agent, to determine the risk, nature, or frequency of off-target effects in that patient. Thus, the methods disclosed herein can be used for personalized medicine by providing patient-specific patterns and frequencies of DSBs induced by the agent of interest. Thus, in one embodiment, the method may include obtaining a sample from a subject and exposing the sample to an agent, such as a therapeutic agent. The method may include determining the nature and / or frequency of any damage or DSBs in the sample after exposure. The order of the steps of the invention may be any as described herein.
[0173] The methods disclosed in this paper can be used to detect contamination of samples by any reagent that can cause DSB.
[0174] The methods disclosed in this paper can be used to measure the stability of artificially assembled or synthesized genomes.
[0175] The methods disclosed herein can be used for next-generation risk assessment (NGRA) in genetic toxicology. NGRA is defined as an exposure-oriented, hypothesis-driven risk assessment approach that integrates new technology methods (NAM) to ensure safety without the need for animal testing. Disseminated basis of sex (DSB) exposure is a direct measure of genotoxic exposure and can be quantified using the methods of this invention. Therefore, the methods of this invention can be used for risk assessment requiring quantification of genotoxic exposure.
[0176] In summary, samples of nucleic acids suspected of containing DSB may contain the DSB due to naturally occurring DNA damage, treatment with a potential DSB initiator, intentional induction of DSB at the target site, or for any other reason.
[0177] In step b), conditions are set such that the first adaptor can be linked to the DSB. The first adaptor comprises an oligonucleotide at the 3' end of one strand of the DSB and may also comprise another oligonucleotide at the 5' end of one strand of the DSB. Thus, the first adaptor is covalently linked to the DSB. The linking can be direct or indirect. For example, it can be linked to another nucleotide introduced at the DSB. Alternatively, an oligonucleotide or oligonucleotide pair can be linked to the DSB, and the first adaptor can be linked to the one or more oligonucleotides. The linking method can be any linking method disclosed herein. The two oligonucleotides of the first adaptor pair can be completely complementary.
[0178] The first and / or second oligonucleotide of the first adaptor pair may contain 3' and / or 5' protective features. These protective features may be any protective features disclosed herein, particularly any protective features disclosed in conjunction with the first pair of oligonucleotides discussed in relation to step i) of the method disclosed herein.
[0179] The first and / or second oligonucleotide of the first adaptor pair may contain any features disclosed in connection with step i) of the method disclosed herein. In particular, the first adaptor pair may or may not contain the sequence disclosed in connection with the first oligonucleotide pair discussed in step i).
[0180] The first adaptor comprises an oligonucleotide that can be attached to the 3' end and contains a sequence capable of binding to the immobilized primer via hybridization. This means that when the attached first adaptor pair denatures into a single strand, one strand is complementary to the primer immobilized on the substrate. The adaptor includes a complementary sequence of sufficient length to achieve binding that is not released during washing or polymerization steps. The length of the complementary region can be 5, 10, 15, 20, 21, 24 or more bases. Alternatively, the complementary region may include 5, 10, 15, 20, 21, 24 or more complementary bases.
[0181] Step c) may include any fragmentation method described herein.
[0182] In step d), conditions are set such that the nucleic acid can be ligated to a second adaptor. The second adaptor comprises an oligonucleotide capable of ligating to the 5' end at the fragmentation site and may also comprise an oligonucleotide capable of ligating to the 3' end at the fragmentation site. The ligation method may be any method disclosed herein.
[0183] The second adaptor cannot connect to the first adaptor. In one particular embodiment, this prevention is due to the inclusion of a 3' protective feature on the first adaptor. In related embodiments, this prevention is due to the inclusion of 5' and / or 3' protective features on the first adaptor pair. Specifically, the first oligonucleotide of the first adaptor (i.e., the oligonucleotide capable of binding to the 3' end) may include a 3' protective feature (such as a spacer C3 3' chain terminator) to prevent the adaptor from connecting to that chain (see [link to relevant documentation]). Figure 12 ).
[0184] The first and / or second oligonucleotide of the second adaptor pair may contain 3' and / or 5' protective features. These protective features may be any protective features disclosed herein, particularly any protective features disclosed in conjunction with the second pair of oligonucleotides discussed in step iii) of the method disclosed herein.
[0185] In the actual implementation prior to step b), the first and / or second oligonucleotides of the second adaptor pair may contain 3' and / or 5' protective features, and at least one protective feature prevents the adaptors of step b) from being attached to those of step d). For example, the oligonucleotides of the second adaptor capable of binding to the 3' end may contain 3' protective features, such as a spacer C3 3' chain terminator.
[0186] The first and / or second oligonucleotide of the second adaptor pair may contain any features disclosed in connection with the second oligonucleotide pair discussed in step iii) of the method disclosed herein. In particular, the second adaptor pair may or may not include the sequence disclosed in connection with the second oligonucleotide pair discussed in step iii).
[0187] The second connector pair can be introduced via tagging. Therefore, the fragmentation and joining of the second connector can be done in the same step. For example, steps c) and d) can be combined.
[0188] The second adaptor comprises an oligonucleotide that can be linked to the 5' end, and said oligonucleotide contains at least a portion of the sequence identical to that of the primer immobilized to the substrate. Therefore, when this oligonucleotide acts as a template during polymerization, the new chain will include a sequence capable of hybridizing with said primer. Thus, the lengths of the relevant regions of the adaptor sequence and the relevant regions of the primer should be sufficient to satisfy this function. The length of the identical region can be 5, 10, 15, 20, 21, 24, or more bases. The other oligonucleotide in the second adaptor does not include a sequence complementary to this so-called identical region. Therefore, the second adaptor cannot bind to the substrate via hybridization.
[0189] Following step f), the method may further include contacting any hybridized nucleic acid with a polymerase to synthesize a nucleic acid, which is a chain of nucleotides complementary to the hybridized nucleic acid, under conditions suitable for immobilizing primer extension. The newly formed nucleic acid can then be amplified. In some embodiments, the primers used for amplification are also immobilized on a substrate and may be adapted, for example, for bridging amplification. This process is known in the art and forms clonal clusters of nucleic acids. In other embodiments, such as for embodiments where the substrate is microbeads, the primers used for implication may be in solution. The amplified nucleic acid can then be sequenced in a conventional manner, for example, by sequencing-by-synthesis. The adaptor may contain binding sites for sequencing primers to facilitate this process. The adaptor may also contain indexes disclosed herein.
[0190] Therefore, in one implementation, the method may further include: g) Obtain the sequence information of any nucleic acids that hybridized with the substrate in step f).
[0191] The method including step g) can be referred to as a method for identifying DNA DSB in nucleic acid samples.
[0192] The features disclosed in step vi) of the method disclosed herein also apply to step g).
[0193] In any embodiment of a product or method comprising an oligonucleotide containing AATGATACGGCGACCACCGA (SEQ ID NO: 34) or a variant thereof, the oligonucleotide may comprise AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO: 41), or a variant defined in SEQ ID NO: 34. In any embodiment of a product or method comprising an oligonucleotide containing SEQ ID NO: 31 or containing the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31) of no more than 5, 10, 15, or 20 bases, the oligonucleotide may comprise SEQ ID NO: 38 or 42, or may comprise the sequence GTGTAGATCTCGGTGGTCGCCGTATCATT (SEQ ID NO: 42) of no more than 5, 10, 15, 20, 25, or 29 bases, or the sequence TCGGTGGTCGCCGTATCATTC (SEQ ID NO: 38) of no more than 5, 10, 15, or 19 bases.
[0194] Figure 16Embodiments of identifying target features (such as damage in a single strand of double-stranded DNA) using the method of the present invention are disclosed. The target feature can be any feature capable of being specifically cleaved. For example, any feature that would cause cleavage of a single strand of double-stranded DNA at the site of the target feature. The target feature can be, for example, a cyclobutanepyrimidine dimer (CPD), 8-oxoguanine, an abase site, or any combination thereof. In an embodiment, the DNA strand containing the target feature can be cleaved, and the double-stranded sample can be denatured to produce a 3' end to which an adaptor can connect.
[0195] Therefore, in one aspect of the invention, a sample preparation method for identifying a target feature in a nucleic acid sample is provided, wherein the preparation includes modifying a nucleic acid associated with the target feature to be suitable for binding to a substrate containing immobilized primers, the method comprising: a) Provide a sample containing multiple nucleic acids, expose the multiple nucleic acids to conditions that can cleave at least one strand of the nucleic acid at a target feature, and denature the multiple nucleic acids into single-stranded nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor comprises an oligonucleotide at the 3' end of a strand capable of ligating to a cleavage site and said oligonucleotide comprises a sequence capable of binding to a primer immobilized on a substrate by hybridization; c) Fragmenting multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of attaching to the 5' end of one strand at a fragmentation-induced break but not to the first adaptor, and the oligonucleotide does not contain a sequence capable of binding to a primer immobilized on a substrate via hybridization. The oligonucleotide of the second adaptor capable of attaching to the 5' end of one strand at a fragmentation-induced break may contain the same sequence as the region of the second primer.
[0196] All the features disclosed in the methods for identifying DNA DSBs are also relevant to the methods used to identify the target features. Specifically, steps b), c), and d) of the sample preparation method for identifying the target features in a nucleic acid sample can be identical to steps b), c), and d) disclosed in the sample preparation method for identifying DNA DSBs in a nucleic acid sample. The adaptor can be the same as that disclosed for identifying DSBs, and the method can be used to prepare libraries suitable for hybridization with any substrate disclosed herein. For embodiments including a double-stranded adaptor ligated to a single-stranded nucleic acid, the splice oligonucleotides disclosed herein can be included.
[0197] In some implementations, the steps may be performed in the following order: c), d), a), and then b). This is Figure 16The order shown in the text. Therefore, in one embodiment, a sample preparation method for identifying a target feature in a nucleic acid sample is provided, wherein the preparation includes modifying the nucleic acid associated with the target feature to be suitable for binding to a substrate containing immobilized primers, the method comprising: c) Fragmenting samples containing multiple nucleic acids; and d) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of ligating to the 5' end of a strand at a fragmentation-induced break and said oligonucleotide does not contain a sequence capable of binding to a primer immobilized on a substrate by hybridization; optionally, it contains a sequence identical to the region of the second primer; a) Exposing multiple nucleic acids to conditions that cleave at least one strand of the nucleic acid at the target feature, and denaturing the multiple nucleic acids into single-stranded nucleic acids; b) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide capable of ligating to the 3' end of a strand that is capable of ligating to a cleavage site but not to a second adaptor, and the oligonucleotide contains a sequence capable of binding to a primer immobilized on a substrate by hybridization.
[0198] All downstream steps and characteristics disclosed in the methods for identifying DNA DSBs are also relevant to methods for identifying target characteristics. Specifically, methods may include denaturing the sample into single-stranded nucleic acids, for example, denaturing any double-stranded connectives. Methods may further include contacting multiple nucleic acids with a substrate containing the fixed primers under conditions suitable for hybridization of the immobilized primers and complementary nucleic acids. Furthermore, methods may further include obtaining sequence information of any nucleic acids hybridized to the substrate.
[0199] In another embodiment, a sample preparation method for identifying a target feature in a nucleic acid sample is provided, wherein the preparation includes modifying the nucleic acid associated with the target feature to be suitable for binding to a substrate containing immobilized primers, the method comprising: α) Expose a sample containing multiple nucleic acids to conditions that can cleave at least one strand of the nucleic acid at the target feature, and denature the multiple nucleic acids into single-stranded nucleic acids; β) Under ligation conditions, the sample is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide that is complementary to the first oligonucleotide of the first pair; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features. γ) Fragmenting the nucleic acid of the sample into fragments; and δ) Under ligation conditions, the fragment is exposed to a second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides contains a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second, longer oligonucleotide, which is partially complementary to the first oligonucleotide of the second pair and contains a 3' binding feature for binding to the second strand of the fragmented nucleic acid, a sequence complementary to the hybridization site, and an additional sequence, optionally a binding sequence for bridging amplification; and wherein one or both of the oligonucleotides contain 3' and / or 5' protective features.
[0200] In some embodiments, the order of the steps may be step γ), step δ), step α), and then step β). Thus, in another embodiment, a sample preparation method for identifying a target feature in a nucleic acid sample is provided, wherein the preparation includes modifying the nucleic acid associated with the target feature to be suitable for binding to a substrate containing immobilized primers, the method comprising: γ) Fragmenting nucleic acid samples into nucleic acid fragments; δ) Under ligation conditions, the fragment is exposed to a second pair of oligonucleotides, wherein the first oligonucleotide of the second pair of oligonucleotides contains a hybridization site (RD2 SP) to which a second sequencing primer can bind; and a second, longer oligonucleotide, which is partially complementary to the first oligonucleotide of the second pair and contains a 3' binding feature for binding to the second strand of the fragmented nucleic acid, a sequence complementary to the hybridization site, and additional sequences, optionally binding sequences for bridging amplification; and wherein one or both of the oligonucleotides contain 3' and / or 5' protective features; α) Exposing the sample to conditions capable of cleaving at least one strand of nucleic acid at the target feature, and denaturing multiple nucleic acids into single-stranded nucleic acids; and β) Under ligation conditions, the sample is exposed to a first pair of oligonucleotides, wherein the first oligonucleotide of the first pair of oligonucleotides comprises a 5' binding feature that enables the oligonucleotide to ligate to the first strand of the DSB, a hybridization site (RD1 SP) to which a first sequencing primer can bind, and a binding sequence for separating the DSB from the DSB pool; and a second oligonucleotide that is complementary to the first oligonucleotide of the first pair; wherein one or both of the oligonucleotides comprise 3' and / or 5' protective features.
[0201] The method may further include: ε) denatures the fragment to provide single-stranded nucleic acid.
[0202] The method may further include: ζ) divides the chain of part ε) into two groups: group A, those fragments that have a first hybridization site and binding sequence provided by oligonucleotides of part β) at the first end and a second hybridization site and additional sequence provided by oligonucleotides of part δ) at the other end, and group B, those fragments that have not a hybridization site and binding sequence provided by oligonucleotides of part β) at the first end and a second hybridization site and additional sequence provided by oligonucleotides of part δ) at the other end.
[0203] The method may further include: η) Sequencing of the group A strands using primers that bind to the first and / or second hybridization sites.
[0204] In another aspect of the invention, the method can be adapted to be particularly suitable for use with microbead-based systems, such as IonTorrent sequencing. Figure 14 and 15 The example illustrates a specific implementation scheme.
[0205] Therefore, in one aspect of the invention, a sample preparation method for identifying DNA DSB in a nucleic acid sample is provided, wherein the preparation includes modifying DSB-associated nucleic acids to be suitable for binding to a substrate containing a fixed first primer, the method comprising: 1) Provide samples containing multiple nucleic acids; 2) Exposing multiple nucleic acids to a first adaptor under conditions conducive to ligation, wherein the first adaptor contains an oligonucleotide at the 3' end of a strand capable of ligating to a DSB and said oligonucleotide contains a sequence capable of hybridizing with a second primer; 3) Fragmenting multiple nucleic acids; 4) Exposing multiple nucleic acids to a second adaptor under conditions conducive to ligation, wherein the second adaptor comprises an oligonucleotide capable of ligating to the 5' end of one strand at a fragmentation-induced break but not to the first adaptor, and said oligonucleotide comprises the same sequence as the region of the immobilized first primer; and 5) Under conditions suitable for primer extension, contact multiple nucleic acids with the second primer.
[0206] A sample containing multiple nucleic acids can be any sample disclosed herein. In particular, a sample may be suspected of containing DSB, may be treated with an agent that can induce suspected DSB, or may contain or be suspected of containing target features that can be converted into DSB. The sample may include target features that can be specifically cleaved.
[0207] The sample can be any DNA sample that can contain DSB, such as gDNA.
[0208] In one particular implementation, step 2) is: Multiple nucleic acids are exposed to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair is capable of ligating to at least the 3' end of one strand of a DSB, and wherein the first adaptor pair comprises at least partially complementary first and second oligonucleotides, and the first oligonucleotide is ligable to the 3' end and comprises a sequence capable of hybridizing with a second primer; and Step 4) is: Multiple nucleic acids are exposed to a second adaptor pair under conditions that facilitate ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break but not to a first oligonucleotide of the first adaptor pair, wherein the second adaptor contains an oligonucleotide complementary to the first and second portions, and the first oligonucleotide is ligable to the 5' end and contains the same sequence as the region of the immobilized first primer, and the second oligonucleotide does not contain a sequence complementary to the sequence of the same region as the immobilized first primer.
[0209] In a particular embodiment, the oligonucleotide that can be linked to the second adaptor pair at the 5' end comprises 5, 10, 15, 20, or all 23 bases of the sequence AACCCACTACGCCTCCGCTTTCC (SEQ ID NO: 40). Another oligonucleotide has a sequence containing more than 5, 10, 15, or 20 bases that does not contain the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36), or a sequence containing all 22 bases that does not contain the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36).
[0210] The oligonucleotides according to SEQ ID NO: 36 and SEQ ID NO: 40 may contain 1 to 12, 1 to 10, 1 to 8, 1 to 5, 1 to 3, 2, or 1 modification, such as substitution, deletion, or insertion. In one embodiment, the modification is substitution.
[0211] The first and / or second oligonucleotide of the first adaptor pair may contain 3' and / or 5' protective features. These protective features may be any protective features disclosed herein, particularly any protective features disclosed in conjunction with the first pair of oligonucleotides discussed in relation to step i) of the method disclosed herein.
[0212] The first and / or second oligonucleotide of the second adaptor pair may contain 3' and / or 5' protective features. These protective features may be any protective features disclosed herein, particularly any protective features disclosed in conjunction with the second pair of oligonucleotides discussed in step iii) of the method disclosed herein.
[0213] In certain implementations, the second connector cannot be connected to the first connector due to the presence of the 3' modification of the first connector. For example, the C3 3' chain terminator.
[0214] The oligonucleotide that can be linked to the 5' end of the second adaptor contains the same sequence as the region of the second primer, such that when the complementary strand is generated, the complementary strand contains the region to which the primer can bind by hybridization. This sequence, which is the same as the region of the second primer, can be identical to 5, 10, 15, 20, 21, 24 or more bases of the immobilizing primer.
[0215] The method may further include denaturing multiple nucleic acids to form multiple single-stranded nucleic acids. This step may be characterized by any feature disclosed herein in conjunction with other embodiments of the invention.
[0216] The method may further include contacting multiple nucleic acids with a substrate containing the immobilized first primer under conditions suitable for hybridization of the immobilized first primer with complementary nucleic acids. For example, a sample of nucleic acids with ligated adaptors can subsequently bind to a substrate (such as microbeads containing the immobilized primer). The primers can be immobilized onto the beads such that the 5' end is proximal to the immobilization site and the 3' end is distal to the immobilization site.
[0217] Conventional techniques can be used to enable a substrate (such as microbeads) to display multiple copies of nucleic acids with the same sequence.
[0218] The method may further include obtaining sequence information of any nucleic acids hybridized to the substrate. The adaptor may contain binding sites for sequencing primers to assist the process. The adaptor may contain an index sequence to assist the process.
[0219] According to another aspect of the invention, a double-stranded linker suitable for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, such as gDNA samples, is provided, comprising: Contains a first oligonucleotide with the sequence according to AACCCACTACGCCTCCGCTTTCC (SEQ ID NO: 40); and A second oligonucleotide containing more than 5, 10, 15, or 20 bases of the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36), or a sequence containing all 22 bases of the sequence GGAAAGCGGAGGCGTAGTGGTT (SEQ ID NO: 36); wherein one or both of the oligonucleotides contain 3' and / or 5' protective features, respectively.
[0220] In the following claims and the foregoing description of the invention, unless the context requires otherwise due to explicit language or necessary meaning, the word “comprises” or variations such as “comprises” or “comprising” are used in an inclusive sense, that is, specifying the presence of the features described in the various embodiments of the invention, but not excluding the presence or addition of other features.
[0221] All references cited in this specification, including any patents or patent applications, are hereby incorporated by reference. No acknowledgment is made that any reference constitutes prior art. Furthermore, no acknowledgment is made that any prior art constitutes part of common general knowledge in the art.
[0222] Preferred features of each aspect of the invention may be as described in conjunction with any other aspect.
[0223] Any feature disclosed in the statements of inventive feature steps i), ii), iii), etc., may be combined with any feature disclosed in the statements of inventive feature steps a), b), c), etc.
[0224] Other features of the invention will become apparent from the following embodiments. Generally, the invention extends to any new one or any new combination of features disclosed in this specification (including the appended claims and drawings). Thus, features, integers, properties, compounds, or chemical portions described in connection with a particular aspect, embodiment, or example of the invention should be understood to be applicable to any other aspect, embodiment, or example described herein, unless incompatible therewith.
[0225] Furthermore, unless otherwise stated, any feature disclosed herein may be replaced by an alternative feature used for the same or similar purpose.
[0226] The invention will now be described by way of example only, with particular reference to the following figures, in which... Example
[0227] method Cell culture and processing HEK293, HEK293T, and U2OS DIvA cells were cultured at 37˚C in 5% CO2 in DMEM supplemented with 10% FBS (Life Technologies). Lonza 4D-Nucleofector X units with pulse-coding CM-130 were used at 224 pmol ribonucleoprotein (RNP) / 3.5 × 10⁻⁶. 5 HEK293 cells were nuclearly transfected. Cells were harvested at 0, 7, 12, 24, and 30 hours post-transfection for INDUCE-seq processing. To stimulate AsiSI-dependent DSB induction, DIvA cells were treated with 300 nM 4OHT (Sigma, H7904) for 4 hours.
[0228] The EMX1-targeting guide RNA (GAGTCCGAGCAGAAGAAGAA; SEQ ID NO: 29) was synthesized as a full-length, unmodified sgRNA oligonucleotide (Synthego). The Cas9 protein was produced internally (AstraZeneca) and contained an N-terminal 6xHN tag.
[0229] Cells were loaded at a rate of ~1×10 5Cells were seeded at a density of / wells into 96-well plates pre-coated with poly-D-lysine (GreinerBio-One, 655940) and crosslinked in 4% paraformaldehyde (PFA) (Pierce, 28908) at room temperature (rt) for 10 min. Cells were washed in 1×PBS to remove formaldehyde and stored at 4°C for up to 30 days. The INDUCE-seq method began with cell permeabilization. Between incubation steps, cells were washed in 1×PBS at room temperature. Cells were permeabilized by incubating for one hour at room temperature in lysis buffer 1 (10 mM Tris-HCl pH 8, 10 mM NaCl, 1 mM EDTA, 0.2% Triton X-100, pH 8, at 4°C), followed by incubation at 37°C in lysis buffer 2 (10 mM Tris-HCl, 150 mM NaCl, 1 mM EDTA, 0.3% SDS, pH 8, at 25°C) for one hour. Permeabilized cells were washed three times in 1×CutSmart® buffer (NEB, B7204S) and then blunted at the ends for one hour at room temperature using a final volume of 50 μL with 100 µg / mL BSA using the NEB Rapid End-Bending Kit (E1201L). Cells were then washed three times in 1×CutSmart® buffer and A-tailed at 37°C for 30 minutes using a final volume of 50 μL with the NEBNext® dA-tailing module (NEB, E6053L) to add a single dATP to the 3' end of the double-stranded DNA. A-tailed cells were washed three times in 1×CutSmart® buffer and then incubated for 5 minutes at room temperature in 1×T4 DNA ligase buffer (NEB, B0202S). A-tailed ends were labeled using T4 DNA ligase (NEB, M0202M) + 0.4 μM P5 adaptor at a final volume of 50 μL for 16–20 hours at 16°C. Following ligation, excess P5 adaptor was removed by washing cells 10 times at room temperature in washing buffer (10 mM Tris-HCl, 2 M NaCl, 2 mM EDTA, 0.5% Triton X-100, pH 8, at 25°C), with each wash incubated for 2 minutes. Cells were washed once in PBS and then once in nuclease-free H2O (IDT, 11-05-01-04). Genomic DNA was extracted by incubating cells in a final volume of 100 µL at room temperature for 5 minutes in DNA extraction buffer (10 mM Tris-HCl, 100 mM NaCl, 50 mM EDTA, 1.0% SDS, pH 8, at 25°C) + 1 mg / mL proteinase K (Invitrogen, AM2584).Cell lysates were transferred to 1.5 mL Eppendorf RNA / DNA LoBind tubes (Fisher Scientific, 13-698-792) and incubated at 65°C for 1 hour with shaking at 800 rpm. DNA was purified using Genomic DNA Clean & Concentrator™-10 (Zymo Research, D4010) and eluted with 100 μL of elution buffer. DNA yield was assessed using 1 μL of sample and the Qubit DNA HS kit (Invitrogen, Q32854) before library preparation. Genomic DNA was fragmented to 300–500 bp using a Bioruptor sonicator, and size selection was performed using SPRI beads (GC Biotech, CNGS-0005) to remove fragments <150 bp. End repair was performed on the fragmented and size-selected DNA using the NEBNext® Ultra™ II DNA module (NEB, E7546L). Using the NEBNext® Ultra™ II ligation module (NEB, E7595L), following the manufacturer's instructions, 7.5 μM of modified semi-functional P7 adaptor was used, omitting the USER enzyme addition. Fragmented and end-repaired DNA was directly added to the ligation reaction. The ligated sequencing library was purified using SPRI microbeads. The library was purified twice more using SPRI microbeads, and size selection was performed to remove fragments <200 bp to remove residual adaptor DNA. The final clean library was quantified by qPCR using the KAPA Library Quantification Kit (Roche, 07960255001) from the Illumina® platform. The sample was collected and concentrated to the required volume for sequencing using SpeedVac. Sequencing was performed on an Illumina NextSeq 550 using a 1×75 bp high-capacity flow cell.
[0230] All modified INUCE-seq adapter oligonucleotides were purchased from IDT. Single-stranded oligonucleotides were annealed to a final concentration of 10 μM in nuclease-free double-stranded buffer (IDT, 11-01-03-01) by heating to 95°C for 5 minutes and then slowly cooling to 25°C using a thermal cycler. Figure 2 The document provides an overview of the structure of the adaptor oligonucleotide.
[0231] In situ DSB induction with HindIII INDUCE-seq preliminary experiments were performed in HEK293T cells using the restriction enzyme HindIII-HF® (NEB, R3104S) to induce DSB in situ. The procedure was identical to that described in the full INDUCE-seq method, with DSB induction added before end-point blunting. After cell permeabilization, DSB was induced with 50 U HindIII-HF® in 1×CutSmart® buffer at a final volume of 50 μL. Digestion was performed at 37°C for 18 hours.
[0232] The split FASTQ files were obtained and, using default settings, were trimmed using Galore! (Krueger 2015) to remove the 3' adaptor sequences from the reads. Sequencing data quality was assessed using FastQC (Andrews 2010). After read alignment with the human reference genome (GRCh37 / hg19) using BWA-mem (Li and Durbin 2009), alignments mapped to low scores (MAPQ < 30) were removed using SAMtools (Li et al., 2009), and a custom AWK script was used to filter soft-splitting reads to ensure accurate DSB allocation. The resulting BAM files were converted to BED files using the bedtools bam2bed function (Quinlan and Hall, 2010), and a list of read coordinates was then filtered using poorly mapped regions, chromosome ends, and incomplete reference genome contigs to remove these features from the data. The DSB position was specified as the first 5' nucleotide upstream of the read relative to strand orientation, and the output was a "breakends" BED file. Optical repeats are carefully removed while preserving true recurrent DSB events. By maintaining the read ID, a custom AWK script is used to filter out optical repeats using the X and Y position information of the flow cell. The final output is a BED file containing a quantified list of individual nucleotide break locations.
[0233] The locations of HindIII target sites in hg19 were first predicted on a computer using the tool SeqKit locate (Shen et al., 2016), allowing for a maximum mismatch of 2 bp with the HindIII target sequence AAGCT. The number of breaks overlapping these predicted sites was calculated using bedtoolsintersect. To compare with the DSBCaptrureEcoRV experiment (Lensing et al., 2016), the same coverage threshold of ≥5 breaks / sites was used to define each HindIII-induced break site.
[0234] The locations of AsiSI target sites were calculated in the same manner as HindIII, but mismatches were not allowed, and the sequence GCGATCGC was used. Since DIvA cells are female, sites present on the Y chromosome were removed, leaving 1211 sites for chr1-X. To rigorously calculate true AsiSI-induced breaks, 8 bp AsiSI sites were reduced to 1 bp genomic spacers at the predicted break locations. This reduced each 8 bp genomic spacer to two 1 bp spaces; at position 6 on the positive strand and position 3 on the negative strand. The direct overlap between the 1 bp break ends and the predicted AsiSI break sites was then calculated using bedtoolsintersect. For each overlap, a matching strand orientation was required to be considered a true AsiSI-induced break site.
[0235] First, two sets of potential off-target sites for EMX1 in hg19 were predicted using a command-line version of Cas-OFFinder (Bae et al., 2014). The first set allowed up to 6 mismatches in the spacer and standard PAM, while the second set allowed up to 7 mismatches. Next, the predicted sequences from both sets were filtered based on the number of mismatches in a seed region defined as 12 nucleotides proximal to the PAM. Each set was filtered for up to 2, 3, 4, and 5 mismatches in the seed, generating a set of eight files with different mismatch filtering parameters. To constrain CRISPR-induced DSBs, each 23 bp predicted site was first reduced to a 2 bp spacer flanking the expected CRISPR breakpoint and 3 bp upstream of the PAM. The overlap between these 2 bp expected breakpoints and the 1 bp breakpoint in INDUCE-seq was then calculated using bedtools intersecT (Quinlan and Hall, 2010), returning a set of DSBs identified at the expected CRISPR breakpoints. Finally, DSBs overlapping with CRISPR sites were filtered based on the number of site mismatches and the number of breaks detected at the sites. Sites with mismatches >n needed to have more than one DSB overlap to be retained as true off-target sites. Each set of break overlaps was filtered using mismatch values >2, >3, >4, and >5, resulting in a total of 32 filtering conditions and off-target datasets for each INDUCE-seq sample.
[0236] Calculate the overlap between CRISPR off-target detection methods EMX1 off-target sites were compared with alternative methods CIRCLE-seq, Digenome-seq, GUIDE-seq, BLISS, and HTGTS. Genomic spacer files were generated for each corresponding off-target detection method. The overlap of EMX1 off-targets detected by each method was calculated using bedtoolsintersecT (Quinlan and Hall, 2010).
[0237] Amplicon sequencing validation of mutation results Amplicon sequencing DNA libraries were prepared using a custom-designed rhAmpSeq RNase-H-dependent primer set (IDT) located on the off-target flanking regions identified by INDUCE-seq on the EMX1. Multiplex PCR was performed according to the manufacturer's instructions using rhAmpSeq HotStart Master Mix 1, the custom primer mix, and 10 ng of genomic DNA. PCR products were purified using SPRI microbeads and incorporated into Illumina sequencing P5 and P7 index sequences using a second multiplex PCR with rhAmpSeq HotStart Master Mix 2. The resulting sequencing libraries were pooled and sequenced using an Illumina NextSeq 550 Mid output flow cell with 2 x 150 bp chemistry. Editing results at the target and off-target regions were determined using CRISPResso software (Pinello et al., 2016) v2.0.32 with the following parameters: CRISPRessoPooled-q30-ignore_substitutions--max_paired_end_reads_overlap 151. Use CRISPROSTOCompare to compare insertion and missing frequencies.
[0238] result The fragmentation measurement was achieved via INDUCE-seq through a two-stage, PCR-free library preparation process. Figure 1 and 5Phase 1 consists of in-situ labeling of the end-prepared DSBs by ligating a full-length, chemically modified P5 adaptor to the DSB ends. In Phase 2, the extracted, fragmented, and end-prepared gDNA is ligated using a second chemically modified semi-functional P7 adaptor. The resulting DSB-labeled DNA fragments containing both the P5 and semi-functional P7 adaptors can interact with an Illumina flow cell and are subsequently sequenced using single-end sequencing. DNA fragments without the P5 adaptor remain non-functional because they lack the sequence required for hybridization with the flow cell. This method allows for the enrichment of functionally labeled DSB sequences and the elimination of all other genomic DNA fragments that would otherwise contribute to system noise. Avoidance of break-amplification produces sequencing outputs where a single sequencing read is equivalent to a single labeled DSB end (compare). Figure 5 (a and 5b). This innovation enables the direct measurement and quantification of DSB through sequencing, representing a significant advance in the accurate measurement of DSB in the genome.
[0239] Following in-situ breakpoint labeling, currently available DSB detection methods, such as BLISS, DSBCapture, and END-seq, all employ enrichment protocols to separate the DSB-labeled DNA fragments from the remaining genomic DNA. This is followed by PCR-based library preparation and sequencing. Figure 5 b). This amplification-based approach produces sequencing output where a single read is not equivalent to a single break, requiring PCR error correction schemes (such as unique molecular identifiers (UMIs)) to attempt to quantify DSBs. Figure 5 b and Figure 6 (b)(Yan et al., 2017). Importantly, the new INDUCE-Seq library preparation and adaptor combination are compatible with any publicly available in-situ DSB labeling scheme.
[0240] To demonstrate the characteristics of the INDUCE-seq method, we first examined how INDUCE-seq measures genome-wide DSBs after inducing defined DSBs in fixed and permeabilized cells using a high-fidelity HindIII restriction endonuclease. This method has previously been used in detection methods such as BLISS, END-seq, and DSBCapture. Figure 3 As shown in Figure a, INDUCE-seq detected hundreds of millions of highly recurrent HindIII-induced DSBs simultaneously from the same sample, in addition to detecting hundreds of thousands of low-level endogenous DSBs, without requiring any form of error correction. This, for the first time, enables the precise quantification and characterization of endogenous DSBs in the genome. Figure 7An example of endogenous DSB detection using INDUCE-seq is shown. In summary, these observations demonstrate the significantly wider dynamic range and sensitivity achievable with INDUCE-seq. Figure 3 b shows the expected HindIII cleavage pattern detected by INDUCE-seq, where two semi-overlapping symmetrical segments of the forward and reverse sequencing reads are mapped to known HindIII cleavage sites on the forward and reverse strands. This confirms that INDUCE-seq can be used to accurately measure DSB end structure at single nucleotide resolution. We measured a significant increase in breakage per cell after treatment with HindIII, from fewer than 10 intrinsic breaks in untreated cells to more than 3000 induced breaks per treated cell. This demonstrates that INDUCE-seq can quantitatively measure breakage per cell across three orders of magnitude. Figure 3 c). Compared to an equivalent experiment using the DSBCaptrure method to detect EcoRV-induced breaks, we found a larger proportion of the summed and aligned INDUCE-seq reads mapped to restriction sites. Significantly, 96.7% of the aligned reads mapped to restriction sites, indicating a 25% improvement in break detection fidelity compared to DSBCapture. Figure 3 d). Importantly, INDUCE-seq used 1 / 800th of the cells used in the DSBCapture experiment and identified a HindIII restriction site proportion (92.7%) similar to that identified by DSBCapture (93.7%). Figure 3 e). In addition to the target HindIII site containing the sequence AAGCTT, we identified a large number of DSBs at various HindIII off-target sequences differing by one or two mismatched bases. Figure 3f). The total number of HindIII-induced DSBs measured by INDUCE-seq ranged from ~150,000,000 at the target site to only five DSBs at the lowest-ranked off-target site. Therefore, INDUCE-seq quantitatively detected breaks across eight orders of magnitude, demonstrating a significantly enhanced dynamic range for break detection compared to other methods. Enhanced sensitivity was also observed when performing similar experiments measuring DSBs induced at AsiSI sites in live DiVA cells. Despite sequencing 1 / 40th of the reads for the corresponding DSBcapture experiment and 1 / 23rd of the reads for BLISS, INDUCE-seq detected the presence of breaks at 230 AsiSI sites. This represents an increase beyond the 214 sites detected by BLISS and the 121 sites detected by DSBCapture. This demonstrates that INDUCE-seq is significantly more sensitive, efficient, and therefore more cost-effective than any current break detection method. Figure 3 g).
[0241] We have already established the characteristics of break detection using INDUCE-seq, and next, we applied it to detect CRISPR / Cas9-induced on-target and off-target DSBs in the genome. This analysis is significant for the safety profile analysis in the development of CRISPR-based therapies. Following RNP nuclear transfection of HEK293 cells with widely characterized EMX1 sgRNA, off-target DSBs were measured at 0, 7, 12, 24, and 30 hours post-transfection, as defined by our customized data analysis pipeline. Figure 4 a, Figure 8 and Figure 9 ). Figure 3 a also shows the number of breaks detected at each of the 60 on-target and off-target sites (mismatches ranging from 2 bp to 6 bp compared to the target site) on EMX1.
[0242] This experiment reveals the dynamics of fracture-induced fracture guided by EMX1, providing insights into the mechanism of the CRISPR / Cas9 editing process. Figure 4 As shown in b, most of the on-target and off-target activities were observed immediately after nuclear transfection and in the early stages of the process. These results demonstrate the rapidity of editing via this sgRNA. When compared with existing technologies, we found that INDUCE-seq significantly outperformed alternative cell-based methods GUIDE-seq, BLISS, and HTGTS, and captured several sites previously identified using only in vitro off-target discovery methods CIRCLE-seq and DiGenome-seq. Importantly, INDUCE-seq also detected novel off-target breakpoints that had not been detected by any other method previously. Figure 4 c and Figure 10 ).
[0243] Finally, using DNA from the same samples, we measured the editing results at these target and off-target breakpoints identified by INDUCE-seq using amplicon sequencing. Amplicon sequencing can only detect insertions and deletions with a background error detection rate higher than 0.1%. Therefore, evidence of editing was detected at the target sites, and evidence of editing was detected at only four of the 60 off-target sites identified throughout the time course. Figure 4 d). This observation is consistent with the same previous study that identified five off-target effects with an insertion / deletion frequency >0.1% 48 hours after nuclear transfection of HEK293 cells with EMX1 RNP. Figure 4 d, far right column, 48 hours). This data demonstrates that INDUCE-seq can detect and quantify CRISPR-induced off-target DSBs with significantly higher sensitivity than amplicon sequencing for detecting insertions and deletions. This observation underscores the need for more sensitive methods to detect edit outcomes in assessing the safety of genome editing. Of particular interest is the careful examination of break patterns at the target site and the first two off-target sites, along with subsequent insertion / deletion profiles, revealing sgRNA-specific cleavage patterns reflected in and associated with the edit outcome. Figure 11 This increases the likelihood of using DSB patterns observed at CRISPR-induced breakpoints to predict the final editing results at both target and off-target sites.
[0244] discuss We have developed a novel PCR-free method for preparing DNA libraries for next-generation sequencing of DSBs in the genome. This enables, for the first time, the direct measurement of broken ends in cells. Our method overcomes the poor signal-to-noise ratio (SNR) of DSBs associated with PCR amplification typically used in standard NGS library preparation. The novel INDUCE-seq adaptor design essentially allows the sequencing flow cell to be used for enriching labeled DSB sequences without amplifying them. An improvement in SNR is achieved by filtering out noisy broken ends generated during sample preparation that are irrelevant to the true physiological DSBs found in cells. We demonstrate the characterization of INDUCE-seq in measuring genomic DSBs in a range of different applications. We reveal its ability to simultaneously and sensitively detect both low-level endogenous and high-level restriction enzyme-induced breaks, which was previously impossible without complex and expensive error correction methods with their own limitations and drawbacks. We compare our results with currently available break detection methods to demonstrate how it improves upon these, not only in terms of accuracy and sensitivity, but also in terms of simplicity, scalability, ease of use, and cost-effectiveness. These are fundamental characteristics of assays that can be used to assess the safety profile of synthetic guidance for CRISPR genome editing. We demonstrate how INDUCE-seq performs compared to several current DSB assays for detecting on-target and off-target editing via EMX1sgRNA. We reveal that, in addition to detecting numerous off-target sites cumulatively measured by five other methods, INDUCE-seq also identifies a significant number of novel off-target sites that current methods cannot detect. We suggest that INDUCE-seq may be a crucial method for safety profiling and designing synthetic guide RNAs for the future development of genome editing as a therapeutic modality. These features include genome-wide mutations, single-strand breaks and gaps, as well as other types of DNA damage that can translate into breaks. The development of INDUCE-seq and its derivative assays has significant implications for a range of diverse biomedical applications.
[0245] Table 1. * Phosphothiophosphate linkages resist exonuclease activity Resistance of / 3SpC3 / C3 spacer phosphoramide (covalent block) to exonuclease, ligase, and 5'>3' polymerase activities. -P 5' phosphate groups promote 5' linkage. INDEX Illumina sequencing index sequences can split pooled sequencing libraries. *T The 3'-deoxythymidine triphosphate "T-tail" with a thiophosphate linkage provides a substrate for ligation to "A+ tail" DNA fragments. Table 2 References Andrews, S. (2010). FastQC: a quality control tool for highthroughput sequence data. Available at:http: / / www.bioinformatics.babraham.ac.uk / projects / fastqc [Accessed: 2019]. Bae, S., Park, J., and Kim, J. S. (2014). Cas-OFFinder: a fast andversatile algorithm that searches forpotential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30(10):1473-1475. Krueger, F. (2015). Trim Galore: A wrapper tool around Cutadapt andFastQC to consistently apply quality and adapter trimming to FastQ files.Available at: http: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / [Accessed: 2019]. Lensing, S. V., Marsico, G., Hansel-Hertsch, R., Lam, E. Y.,Tannahill, D., and Balasubramanian, S. (2016). DSBCapture: in situ capture andsequencing of DNA breaks. Nature Methods 13(10):855-+. Li, H. and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25(14):1754-1760. Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N.,. . . Genome Project Data, P. (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics 25(16):2078-2079. Pinello, L., Canver, M. C., Hoban, M. D., Orkin, S. H., Kohn, D. B., Bauer, D. E. and Yuan, G. C. (2016). Analyzing CRISPR genome-editing experiments with CRISPResso. Nature Biotechnology 34(7):695-697. Quinlan, A. R. and Hall, I. M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26(6):841-842. Shen, W., Le, S., Li, Y. and Hu, F. Q. (2016). SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA / Q File Manipulation. Plos One 11(10). Yan, W. X., Mirzazadeh, R., Garnerone, S., Scott, D., Schneider, M.W., Kallas, T., . . . Crosetto, N. (2017). BLISS is a versatile andquantitativemethod for genome-wide profiling of DNA double-strand breaks.Nature Communications 8.
Claims
1. A sample preparation method for identifying DNA double-strand breaks (DSBs) in nucleic acid samples, wherein the preparation includes modifying DSB-associated nucleic acids to enable them to bind to a substrate containing at least a first immobilized primer, the method comprising: a) Provide a sample containing multiple nucleic acids; b) Exposing the plurality of nucleic acids to a first adaptor pair under conditions conducive to ligation, wherein the first adaptor pair comprises a first oligonucleotide and a second oligonucleotide that are at least partially complementary to each other, and the first oligonucleotide is ligable to the 3' end of one strand of the DSB and comprises a sequence capable of binding to the first immobilizing primer by hybridization. c) Fragment the multiple nucleic acids; and d) Exposing the plurality of nucleic acids to a second adaptor pair under conditions conducive to ligation, wherein the second adaptor pair is capable of ligating to at least the 5' end of a strand at a fragmentation-induced break, but not to the first oligonucleotide of the first adaptor pair, wherein the second adaptor pair comprises a first oligonucleotide and a second oligonucleotide that are partially complementary to each other, and the first oligonucleotide is ligable to the 5' end and contains the same sequence as the region of the second primer, and the second oligonucleotide does not contain a sequence complementary to the sequence that is the same as the region of the second primer.
2. The method of claim 1, wherein the method includes modifying the DSB-related nucleic acid to enable it to amplify, wherein the amplification includes using the first immobilization primer and the second primer.
3. The method of claim 1, wherein the second primer is a second fixed primer contained in the substrate.
4. The method of claim 3, wherein the method includes modifying the DSB-related nucleic acid to enable it to amplify, wherein the amplification includes using the first immobilization primer and the second immobilization primer.
5. The method of claim 4, wherein the amplification is bridge amplification.
6. The method of claim 1, wherein the second adaptor pair comprises a first oligonucleotide containing a sequence according to CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO: 32), and a second oligonucleotide containing a sequence of more than 5, 10, 15, or 20 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30), or a sequence of all 24 bases that does not contain the sequence ATCTCGTATGCCGTCTTCTGCTTG (SEQ ID NO: 30).
7. The method of claim 1, wherein the second adaptor pair comprises a first oligonucleotide containing the sequence according to AATGATACGGCGACCACCGA (SEQ ID NO: 34), and a second oligonucleotide containing a sequence of more than 5, 10, or 15 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31), or a sequence of all 20 bases that does not contain the sequence TCGGTGGTCGCCGTATCATT (SEQ ID NO: 31).
8. The method of claim 1, wherein the first oligonucleotide of the first adaptor pair comprises a 3' protective feature; and / or wherein the second oligonucleotide of the second adaptor pair comprises a 3' protective feature.
9. The method of claim 1, wherein the second connector pair cannot be connected to the first connector pair due to the presence of the 3' modification of the first connector pair.
10. The method of claim 1, wherein the oligonucleotide of the second adaptor pair that can be linked to the 5' end contains the same sequence as 5, 10, 15, 20, 21, 24 or more bases as the second primer.
11. The method of claim 1, further comprising denaturing the plurality of nucleic acids to form a plurality of single-stranded nucleic acids.
12. The method of claim 1, further comprising contacting the plurality of nucleic acids with a substrate containing the immobilized primers under conditions suitable for hybridization of the immobilized primers and complementary nucleic acids.
13. The method of claim 12, further comprising obtaining sequence information of any nucleic acid hybridized with the substrate.
14. The method of claim 1, wherein the sample comprising a plurality of nucleic acids is genomic DNA (gDNA).
15. The method of claim 1, wherein the steps are performed in the order of a), b), c) and subsequently d).
16. The method of claim 15, wherein prior to step b), the sample is exposed to conditions capable of inducing DSB at a characteristic site in the nucleic acid sample.
17. The method of claim 1, wherein the steps are performed in the order of a), c), d) and then b); wherein the sample is exposed to conditions that can cause or are suspected of causing DSB between steps d) and b).
18. The method of claim 17, wherein the sample is exposed to conditions capable of inducing DSB at a characteristic site in the nucleic acid sample.
19. The method of claim 1, wherein the connection in step b) occurs in situ or in vitro using a cell or tissue sample.
20. The method of claim 1, wherein the first adaptor pair includes a first hybridization site to which a first sequencing primer can bind and / or the second adaptor pair includes a second hybridization site to which a second sequencing primer can bind.