Methods and kits for detecting base editor editing sites
By introducing marker molecules and enriching them during base editing and sequencing, the accuracy and cost issues of detecting off-target effects of base editors in existing technologies have been resolved, achieving high-sensitivity detection at the whole genome level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2022-05-20
- Publication Date
- 2026-05-12
AI Technical Summary
Existing whole-genome sequencing technologies cannot effectively and unbiasedly assess the off-target effects of base editors, especially in the real environment of living cells, leading to inaccurate test results and high costs.
By introducing marker molecules at the single-strand breaks generated during base editing, and enriching and sequencing the marker products using specific binding molecules, highly sensitive detection of base editor editing sites and off-target effects can be achieved.
This invention provides a method for detecting base editor editing sites and off-target effects with high sensitivity and non-bias at the whole genome level. It is applicable to various base editing tools and improves the accuracy and cost-effectiveness of detection.
Smart Images

Figure CN115386623B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gene editing (particularly base editing) technology. Specifically, this application relates to a method for detecting the sites of nucleic acid editing by a base editor (e.g., a single-base editor or a dual-base editor), and a kit for performing said method. This application also relates to a method for detecting the editing efficiency or off-target effects of nucleic acid editing by a base editor (e.g., a single-base editor or a dual-base editor). Background Technology
[0002] In 2016, David Liu et al. fused the rat rAPOBEC1 protein with the nCas9 (D10A) protein based on the CRISPR / Cas9 system to develop a cytosine base editor (CBE) (Komor, et al. Nature 533, 420-424, doi:10.1038 / nature17946(2016)). The editing principle of the design is as follows: First, nCas9, which has lost some nucleic acid cleavage activity, can still be guided by sgRNA, which drives rAPOBEC1 linked to nCas9 to the target site; then, sgRNA forms an R-loop structure with the DNA sequence of the target gene, so that the non-target strand of DNA in the single-stranded state can be bound by APOBEC1, deaminated cytosine (C) in a certain range on the strand into uracil (U); finally, these uracils can complete the conversion of uracil to thymine through the subsequent DNA replication process, thereby ultimately achieving the C-to-T base conversion. Subsequently, various new CBE editing systems with varying degrees of optimization in terms of editing efficiency, active editing window, and editable sequence range have been developed, such as YE1-BE and BE4max (Kim, YB et al. Nature biotechnology 35, 371-376, doi:10.1038 / nbt.3803 (2017); Suzuki, K. et al. Nature 540, 144-149, doi:10.1038 / nature20565 (2016)).
[0003] Furthermore, in 2020, David Liu et al. reported an RNA-free mitochondrial cytosine base editor, DdCBE (DddA-derived CBE), which achieved a major breakthrough in mitochondrial gene editing (Mok, BY et al. Nature 583,631-+, doi:10.1038 / s41586-020-2477-4(2020)). Previously, due to the presence of the mitochondrial double membrane, the introduction of sgRNA into mitochondria still faced significant challenges, severely limiting the application of CRISPR / Cas9-based CBE tools in mitochondrial gene editing. Compared to CRISPR / Cas9-based CBE tools, the main changes in DdCBE include the following two points: First, it uses TALE protein instead of sgRNA to recognize the target DNA strand, avoiding the problem that sgRNA has difficulty entering mitochondria; second, it uses a newly discovered double-stranded DNA deaminase, DddA, instead of APOBEC to deaminate dC on the double-stranded DNA at the target site into dU, ultimately achieving the base conversion from dC to dT.
[0004] In summary, various cytosine base editing systems targeting the cell nucleus or mitochondria already exist, and their capabilities are constantly being expanded. However, their core principle remains the same: at the targeted editing site, cytosine (C) is deaminated to uracil (U); finally, these uracils can complete the conversion from uracil (U) to thymine (T) through subsequent DNA replication, thereby ultimately achieving a C-to-T base conversion.
[0005] Since David Liu developed the cytosine base editor (Komor et al., 2016) in 2016, the adenine base editor (ABE) (Gaudelli et al., 2017) followed in 2017. The main editing principle of this technology is as follows: Cas9, guided by sgRNA, reaches the target editing site, opens the DNA double helix to form an R-loop structure, and then the adenine deaminase fused with Cas9 deaminates the adenine within the editing window to form inosine (I). During repair and replication, inosine is read as G by DNA polymerase, ultimately resulting in the conversion of adenine (A) to guanine (G). After several years of development, the ABEmax system is currently widely used. This system is based on the initial ABE version and has undergone a series of improvements, including mutation screening, codon optimization, and the introduction of nuclear localization signals, resulting in continuously improved editing efficiency at the target site. In 2020, David Liu and Jennifer A. Doudna reported a new version of ABE with higher activity, named ABE8e (Richter et al., 2020). ABE8e retains only one TadA element from ABEmax and has undergone multiple mutations, which not only improves the enzyme's in vitro activity (Lapinaite et al., 2020) but also greatly enhances the editing efficiency of the target site in cells.
[0006] Similarly, similar to the CBE editing system, various ABE editing systems have been developed. Their core principle is to deaminate adenine into hypoxanthine at the targeted editing site. Then, these hypoxanthines can complete the conversion from hypoxanthine to guanine through the subsequent DNA replication process, thereby ultimately achieving the base conversion from adenine (A) to guanine (G) (A-to-G).
[0007] In addition, in 2020, four research groups successively developed the adenine and cytosine dual-base editing system (ACBE) (Grunewald et al., 2020; Li et al., 2020; Sakata et al., 2020; Zhang et al., 2020). The basic principle is to combine the previously developed ABE and CBE technologies to achieve simultaneous editing of adenine and cytosine within the same target editing window.
[0008] Ideally, gene-editing tools should only edit the target site. However, in reality, both ZFN / TALEN and CRISPR / Cas systems have consistently been found to have off-target risks. Off-target effects occur when the gene-editing tool unnecessarily edits at non-target locations. Once an off-target event occurs, it can damage the gene sequence or chromosome structure at that location, disrupt genome stability and normal cell function, potentially leading to various serious side effects and even inducing cancer. Therefore, off-target effects are a major drawback of gene-editing technology for applications with high safety requirements (such as clinical treatment applications). If base editors are to be applied in practice, their off-target effects must be thoroughly, comprehensively, and accurately assessed beforehand.
[0009] Theoretically, the simplest and most direct way to detect the off-target effects of base editors is to directly detect single nucleotide mutations (SNVs) generated by base editors through whole genome sequencing (WGS). However, WGS is known to have many inherent limitations: First, the genome naturally contains many single nucleotide variations (SNVs), and the DNA replication process and subsequent high-throughput sequencing also introduce considerable random errors. These factors create a genomic background that affects detection accuracy, resulting in extremely low sensitivity for WGS in detecting SNVs. Second, when using high-throughput sequencing technology to perform WGS sequencing on the whole genome, the coverage of the sequencing reads is highly uneven, often requiring a huge amount of data to obtain sufficient information for whole-genome evaluation. Therefore, conventional WGS cannot effectively detect the off-target effects of base editors at the whole-genome level.
[0010] Another approach is to first use software prediction (such as Cas-OFFinder) to identify potential off-target sites, or to select sites where base editing tools might cause off-target editing from the identification results of the CRISPR / Cas9 nuclease system using GUIDE-seq, and then obtain the accurate editing frequency of these sites through targeted deep sequencing. GUIDE-seq is a technique that detects off-target sites by tracking double-stranded breaks (DSBs) generated during the editing process of a nuclease system. This technique is not suitable for gene editing technologies that produce almost no DSBs (such as various base editors). While predicting locations first and then performing single-point deep sequencing can quickly identify and compare the off-target risks of different base editing tools to a certain extent, the results are not based on a comprehensive consideration at the whole genome level, and the conclusions may vary greatly depending on the selected sites.
[0011] Currently, there are two main technologies used to comprehensively evaluate the off-target effects of base editing systems: one is detection technology based on in vitro incubation, such as Digenome-seq; the other is technology based on SNP detection, such as GOTI.
[0012] In 2017, Jin-Soo Kim's team from South Korea made some modifications to the CBE system based on their existing Digenome-seq technology, achieving in vitro detection of off-target effects of the system at the whole genome level (Kim, D. et al. Nature biotechnology 35, 475-480, doi:10.1038 / nbt.3852(2017)). The detection principle is as follows: First, the genomic DNA incubated with BE3ΔUGI (BE3 with UGI removed) is treated with the UDG enzyme to generate a single-strand break at the location of dU (for CBE), or the edited strand is cut with the endonuclease Endo V, which recognizes dI, to create a nick (for ABE), forming a DSB together with the single-strand break formed by nCas9 cleavage; then, the editing site information is obtained by capturing characteristic reads in subsequent high-throughput sequencing results.
[0013] In 2019, Yang Hui's team reported an off-target detection technology called GOTI (genome-wide off-target analysis by two-cell embryo injection) (Zuo, E. et al. Science 364, 289-292, doi:10.1126 / science.aav9973(2019)). The core of this technology lies in the use of two-cell embryo injection. In the two-cell stage of mouse embryos, a gene-editing system carrying a red fluorescent signal is injected into one of the cells. After the embryo develops to a sufficient number of cells, the entire embryo is digested into multiple single cells, and flow cytometry is used to separate the edited and unedited cell progeny. Theoretically, red fluorescent positive and negative cells originate from the same fertilized egg and should therefore have the same genomic background. Subsequent whole-genome sequencing (WGS) comparing these two groups of cells can reveal the differences caused by gene editing, thus providing off-target information.
[0014] Of the existing whole-genome sequencing technologies, Digenome-seq is an in vitro detection technique. Off-target editing behavior is theoretically influenced by the actual chromatin state and local protein concentration within living cells, thus this technology cannot effectively reflect the true off-target situation in vivo. On the other hand, while technologies like GOTI employ a two-cell embryo injection strategy to minimize the influence of genomic background factors such as SNVs, they still cannot avoid the DNA replication error background caused by single-cell amplification. Furthermore, this method involves embryo manipulation, has limited universality, and is technically challenging and time-consuming. In addition, this method still relies on whole-genome sequencing analysis; achieving sufficient data coverage for all embryo samples involved in the experiment inevitably incurs high sequencing costs, making it unsuitable for high-throughput screening and evaluation. More importantly, the conclusions of these two methods regarding the DNA off-target effects of base editing tools are almost completely contradictory. For example, Kim's team found that CBE has high specificity, causing only a limited number of Cas-dependent off-targets, while Yang Hui's team identified only a large number of non-Cas-dependent off-targets. As is well known, the understanding of off-target effects largely determines the direction of subsequent optimization of base editors. There is clearly a need in this field for a better, more comprehensive off-target detection technology that is free from detection bias.
[0015] Therefore, there is an urgent need to develop a new, sensitive, unbiased, and cost-effective detection technology for comprehensive evaluation of the off-target effects of base editing systems at the whole-genome level. Summary of the Invention
[0016] Based on in-depth research, the inventors of this application have developed a novel method for detecting the editing sites, editing efficiency, or off-target effects of nucleic acids edited by base editors (e.g., single-base editors or dual-base editors). This method can capture base editing intermediates generated in living cells during the editing process of various base editors (e.g., single-base editors or dual-base editors) and effectively label and enrich the editing sites. Therefore, this method is universally applicable to the detection of editing sites of various base editing tools, can evaluate their editing efficiency or off-target effects, and can achieve high-sensitivity detection at the whole-genome level.
[0017] Therefore, in one aspect, this application provides a method for detecting the editing site, editing efficiency, or off-target effects of a base editor (e.g., a single-base editor or a dual-base editor) on a target nucleic acid, comprising the following steps:
[0018] (1) Provides an editing product of a base editor for editing a target nucleic acid, comprising a base editing intermediate, the base editing intermediate comprising a first nucleic acid strand and a second nucleic acid strand; wherein the first nucleic acid strand comprises edited bases generated by the base editor for editing the target nucleic acid;
[0019] (2) In the first nucleic acid chain, a single-strand break is generated in the segment containing the edited base (e.g., in the segment from 10 nt upstream to 10 nt downstream of the edited base);
[0020] (3) Introduce a nucleotide labeled with a first labeling molecule at or downstream of the single-strand break to produce a labeled product containing the first labeling molecule;
[0021] (4) Separate or enrich the labeled product; for example, use a first binding molecule that can specifically recognize and bind to the first labeled molecule to separate or enrich the labeled product;
[0022] (5) Determine the sequence of the labeled product;
[0023] Thus, the editing site, editing efficiency, or off-target effect of the base editor on the target nucleic acid can be determined.
[0024] The method of this application can be used to detect the editing sites, editing efficiency, or off-target effects of various base editors on target nucleic acids. In some preferred embodiments, the base editor is a single-base editor or a dual-base editor. In some preferred embodiments, the base editor is selected from cytosine single-base editors, adenine single-base editors, and adenine and cytosine dual-base editors.
[0025] The method of this application is not limited to the target nucleic acid being edited. In some preferred embodiments, the target nucleic acid is a genomic nucleic acid. In some preferred embodiments, the target nucleic acid is a mitochondrial nucleic acid.
[0026] In some preferred embodiments, the editing product of step (1) is the product of the base editor editing the target nucleic acid outside the cell, inside the cell, or in an organelle (e.g., the nucleus or mitochondria).
[0027] In some preferred embodiments, the method further includes, prior to step (1), the following step: contacting the base editor with the target nucleic acid under conditions that allow the base editor to edit the target nucleic acid, thereby generating the edited product. The conditions allowing the base editor to edit the target nucleic acid can be any conditions suitable for the base editor to exert its editing activity.
[0028] In some preferred embodiments, the base editor is brought into contact with the target nucleic acid, either extracellularly, intracellularly, or within an organelle (e.g., the nucleus or mitochondria), under conditions that allow the base editor to edit the target nucleic acid, thereby generating the edited product.
[0029] For example, the method further includes the following steps before step (1): introducing the base editor into a cell or organelle, so that the base editor contacts the target nucleic acid in the cell or organelle and performs base editing, thereby generating an edited product; or, introducing a nucleic acid molecule encoding the base editor into a cell or organelle and expressing the base editor, so that the base editor contacts the target nucleic acid in the cell or organelle and performs base editing, thereby generating an edited product.
[0030] In some preferred embodiments, in step (1), the base-edited target nucleic acid is extracted or isolated from the cell or organelle and optionally fragmented to obtain the edited product.
[0031] The fragmentation can be performed using any method suitable for nucleic acid fragmentation, such as sonication or random enzymatic digestion. In some embodiments, the edited product may or may not contain dangling ends. In some preferred embodiments, the fragmentation (e.g., fragmentation using a nuclease) produces a nucleic acid fragment containing dangling ends (e.g., sticky ends). In such embodiments, optionally, the nucleic acid fragment containing dangling ends is end-repaired to produce a nucleic acid fragment with blunt ends, which can be used as an edited product for the next step. For example, the end-repair may include the flattening of 5' dangling ends (e.g., by nucleic acid polymerization) and / or the removal of 3' dangling ends. In some preferred embodiments, the end-repair includes the flattening of 5' dangling ends (e.g., by nucleic acid polymerization).
[0032] In some preferred embodiments, the second nucleic acid strand has not undergone base editing or does not contain edited bases.
[0033] However, it is readily understood that due to off-target effects, base editors may perform base editing at multiple editing sites (including targeted and off-target sites). For example, a base editor may edit both nucleic acid strands of genomic DNA or organelle DNA (e.g., mitochondrial DNA). Therefore, in some cases, the second nucleic acid strand may potentially have undergone base editing and may contain edited bases. Thus, in some embodiments, the second nucleic acid strand has undergone base editing and / or contains edited bases.
[0034] In some preferred embodiments, the edited base is selected from uracil or hypoxanthine.
[0035] In some preferred embodiments, in step (2), a single-strand break is created at the location of the edited base or upstream of it (e.g., within 10nt, 9nt, 8nt, 7nt, 6nt, 5nt, 4nt, 3nt, 2nt, or 1nt upstream) or downstream of it (e.g., within 10nt, 9nt, 8nt, 7nt, 6nt, 5nt, 4nt, 3nt, 2nt, or 1nt downstream).
[0036] In some preferred embodiments, prior to step (2), the method further includes the step of repairing single-strand breaks (SSBs) (e.g., endogenous single-strand breaks) that may be present in the edited product. For example, prior to step (2), the method further includes using a nucleic acid polymerase, nucleotides (e.g., unlabeled nucleotides; e.g., unlabeled dNTPs), and nucleic acid ligases (e.g., DNA ligases) to repair SSBs (e.g., endogenous SSBs) that may be present in the edited product.
[0037] For example, prior to step (2), the method further includes: (i) incubating the edited product with a nucleic acid polymerase (e.g., DNA polymerase) and nucleotide molecules (preferably unlabeled dNTPs) under conditions that allow nucleic acid polymerization; and (ii) ligating the gap in the product of step (i) using a nucleic acid ligase (e.g., DNA ligase). In some preferred embodiments, the nucleic acid polymerase (e.g., DNA polymerase) has strand displacement activity.
[0038] Without being theoretically limited, it is advantageous to repair SSBs prior to step (2). For example, SSB repair can eliminate gaps that may exist in the edited product, including endogenously present SSBs and SSBs that may be introduced by nucleic acid manipulation (e.g., nucleic acid fragmentation). This avoids the introduction of nucleotides labeled with the first marker molecule at or downstream of these pre-existing SSBs in subsequent steps, thus preventing interference with the detection results from these pre-existing SSBs.
[0039] In some preferred embodiments, in step (2), a single-strand break is created in the first nucleic acid strand using a nuclease (e.g., nuclease V, nuclease VIII, or AP nuclease).
[0040] In some preferred embodiments, the nucleotide labeled with the first labeling molecule is selected from uracil deoxyribonucleotides labeled with the first labeling molecule (e.g., dUTP labeled with the first labeling molecule), cytosine deoxyribonucleotides labeled with the first labeling molecule (e.g., dCTP labeled with the first labeling molecule), thymine deoxyribonucleotides labeled with the first labeling molecule (e.g., dTTP labeled with the first labeling molecule), adenine deoxyribonucleotides labeled with the first labeling molecule (e.g., dATP labeled with the first labeling molecule), guanine deoxyribonucleotides labeled with the first labeling molecule (e.g., dGTP labeled with the first labeling molecule), or any combination thereof.
[0041] In some preferred embodiments, the nucleotide labeled with the first labeling molecule is a uracil deoxyribonucleotide labeled with the first labeling molecule (e.g., dUTP labeled with the first labeling molecule) or a guanine deoxyribonucleotide labeled with the first labeling molecule (e.g., dGTP labeled with the first labeling molecule).
[0042] In some preferred embodiments, the first labeled molecule and the first binding molecule constitute a molecular pair capable of specific interaction (e.g., specific binding). Such molecular pairs capable of specific interaction (e.g., specific binding) are well known to those skilled in the art, for example, biotin or its functional variants—avidin or its functional variants (e.g., biotin-avidin, biotin-streptavidin), antigen / hapten-antibody, enzymes and cofactors, receptor-ligands, molecular pairs capable of click chemistry (e.g., alkynyl-azido compounds), etc. In some preferred embodiments, the first labeled molecule is biotin or its functional variant, and the first binding molecule is avidin or its functional variant; or, the first labeled molecule is a hapten or antigen, and the first binding molecule is an antibody specifically against said hapten or antigen; or, the first labeled molecule is an alkynyl-containing group (e.g., ethynyl), and the first binding molecule is an azido compound capable of click chemistry with said alkynyl group (e.g., ethynyl). For example, the nucleotide labeled by the first labeling molecule is a nucleotide containing an acetylenic group (e.g., 5-Ethynyl-dUTP), and the first binding molecule is an azide compound (e.g., azide-modified magnetic beads) that can undergo a click chemical reaction with the acetylenic group.
[0043] In some preferred embodiments, the linking of the first marker molecule to the nucleotide is reversible or irreversible.
[0044] In some preferred embodiments, the linking of the first marker molecule to the nucleotide in the nucleotide labeled with the first marker molecule is reversible. In such embodiments, after step (4), the method may further include the step of removing the first marker molecule from the labeled product. In some cases, the removal of the first marker molecule is advantageous, for example, as it can avoid adverse effects on subsequent amplification and / or sequencing steps.
[0045] In some preferred embodiments, the linking of the first marker molecule to the nucleotide is irreversible. In such embodiments, preferably, the presence of the first marker molecule does not adversely affect the amplification and / or sequencing of the labeled product. For example, in some preferred embodiments, the labeled product generated in step (3) is capable of nucleic acid amplification. For example, the labeled product is capable of nucleic acid amplification by a nucleic acid polymerase (e.g., a high-fidelity or low-fidelity nucleic acid polymerase).
[0046] In some preferred embodiments, the nucleotide labeled with the first marker molecule is introduced into the single-strand break site or downstream thereof via a nucleic acid polymerization reaction, thereby producing a labeled product containing the first marker molecule. For example, in step (3), a nucleic acid polymerase (e.g., a nucleic acid polymerase with chain displacement activity) is used to introduce the nucleotide labeled with the first marker molecule into the single-strand break site or downstream thereof. For example, in step (3), under conditions allowing nucleic acid polymerization, the first nucleic acid chain is incubated with the nucleic acid polymerase and the nucleotide labeled with the first marker molecule; wherein the nucleic acid polymerase initiates an extension reaction at the single-strand break site using the second nucleic acid chain as a template, and incorporates the nucleotide labeled with the first marker molecule into the single-strand break site or downstream thereof.
[0047] In some preferred embodiments, step (3) further includes using a nucleic acid ligase (e.g., DNA ligase) to ligate the gap in the labeled product containing the first labeled molecule.
[0048] In some preferred embodiments, in step (3), a nucleotide labeled with a second labeling molecule is introduced at or downstream of the single-strand break, thereby producing a labeled product containing the first and second labeling molecules.
[0049] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is a nucleotide molecule that is capable of base pairing with different nucleotides under different conditions (e.g., before and after treatment). For example, the nucleotide labeled with the second labeling molecule is capable of base pairing with the first nucleotide before treatment and with the second nucleotide after treatment.
[0050] In some preferred embodiments, the nucleotide molecule containing the second label is selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4 C(N4-acetylcytosine deoxyribonucleotide).
[0051] In some preferred embodiments, the nucleotide molecule containing the second label is a modified cytosine deoxyribonucleotide that, before treatment, is capable of base pairing with the first nucleotide (e.g., guanine deoxyribonucleotide), and after treatment, is capable of base pairing with the second nucleotide (e.g., adenine deoxyribonucleotide). In some preferred embodiments, the nucleotide molecule containing the second label is selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4 C(N4-acetylcytosine deoxyribonucleotide).
[0052] For example, the nucleotide labeled with the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide. 5-aldehyde cytosine deoxyribonucleotide is capable of base pairing with guanine deoxyribonucleotide before treatment with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or azidoindenidine), and after treatment with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane or 2-methylpyridineborane), or azidoindenidine), is capable of base pairing with adenine deoxyribonucleotide (see, for example, Liu, Y. et al. Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution. Nature Biotechnology). 37,424-429, doi:10.1038 / s41587-019-0041-2(2019).; Patent document WO2015043493A1, the full text of which is incorporated herein by reference).
[0053] For example, the nucleotide labeled with the second labeling molecule is 5-carboxycytosine deoxyribonucleotide. 5-Carboxycytosine deoxyribonucleotide is capable of base pairing with guanine deoxyribonucleotide before treatment with a compound (e.g., a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane)), and after treatment with a compound (e.g., a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane)) (see, for example, Liu, Y. et al. Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution. Nature biotechnology 37, 424-429, doi:10.1038 / s41587-019-0041-2 (2019), the full text of which is incorporated herein by reference).
[0054] For example, the nucleotide labeled with the second labeling molecule is 5-hydroxymethylcytosine deoxyribonucleotide. 5-hydroxymethylcytosine deoxyribonucleotide can be converted to 5-aldehyde cytosine deoxyribonucleotide under the catalysis of an oxidant (e.g., potassium ruthenate) or an oxidase (e.g., TET (ten-eleven translocation) protein). 5-aldehyde cytosine deoxyribonucleotide is capable of base pairing with guanine deoxyribonucleotide before treatment with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or azidoindenedione), and is capable of base pairing with adenine deoxyribonucleotide after treatment with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane or 2-methylpyridineborane), or azidoindenedione).
[0055] For example, the nucleotide labeled by the second labeling molecule is N4-acetylcytosine deoxyribonucleotide (dac 4 C). N4-acetylcytosine deoxyribonucleotides are able to form base pairings with guanine deoxyribonucleotides before treatment with a compound (e.g., sodium cyanoborohydride), and after treatment with a compound (e.g., sodium cyanoborohydride), they are able to form base pairings with adenine deoxyribonucleotides (see, for example, Nature 583, 638-643 (2020), DOI:10.1038 / s41586-020-2418-2, the full text of which is incorporated herein by reference).
[0056] In some preferred embodiments, the nucleotides labeled with a first marker molecule and the nucleotides labeled with a second marker molecule are introduced at or downstream of the single-strand break via a nucleic acid polymerization reaction, thereby generating a labeled product containing the first and second marker molecules. For example, in step (3), under conditions allowing nucleic acid polymerization, the first nucleic acid chain is incubated with a nucleic acid polymerase (e.g., a nucleic acid polymerase with chain displacement activity), the nucleotides labeled with the first marker molecule, and the nucleotides labeled with the second marker molecule; wherein the nucleic acid polymerase initiates an extension reaction at the single-strand break using the second nucleic acid chain as a template, and incorporates the nucleotides labeled with the first and second marker molecules into the single-strand break or downstream of it. In some preferred embodiments, step (3) further includes the step of using a ligase to ligate the gap in the labeled product containing the first and second marker molecules.
[0057] It is understood that the nucleotides labeled with the first labeling molecule and the nucleotides labeled with the second labeling molecule can be introduced in the same nucleic acid polymerization reaction or in different nucleic acid polymerization reactions, as long as a labeled product containing the first labeling molecule and the second labeling molecule can be produced.
[0058] In some embodiments, the use or incorporation of a nucleotide labeled with a second labeling molecule is advantageous. It is readily understood that a nucleotide labeled with a second labeling molecule can be incorporated into a labeled product via nucleic acid polymerization through base complementarity pairing. In this case, the nucleotide labeled with the second labeling molecule (e.g., 5-aldehyde cytosine deoxyribonucleotide) is incorporated into the labeled product through its complementary pairing ability with a first base (e.g., guanine deoxyribonucleotide). Subsequently, the labeled product can be treated (e.g., with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or azidoindenidine)) whereby the nucleotide labeled with the second labeling molecule in the labeled product will be modified or altered, and base complementarity pairing will occur with a second base (e.g., adenine deoxyribonucleotide). Therefore, when the processed labeled product is sequenced, the nucleotide at the incorporation site of the second-labeled nucleotide will pair with the second base and be read as the complementary base of the second base (not the complementary base of the first base) in the sequencing results. In other words, in the sequencing results of the processed labeled product, a base mutation signal (e.g., a C-to-T mutation signal) will be generated at the site of incorporation of the second-labeled nucleotide, from the complementary base of the first base to the complementary base of the second base. By detecting this base mutation signal, the incorporation site of the second-labeled nucleotide can be determined, and the adjacent edited bases can be precisely located. Furthermore, through nucleic acid polymerization, one or more second-labeled nucleotides can be incorporated into the labeled product, thereby detecting one or more base mutation signals in the sequencing results of the processed labeled product. This amplifies the base mutation signal and improves the detection sensitivity.
[0059] Therefore, in embodiments using nucleotides labeled with a second labeling molecule, preferably, after step (3), the labeled product is treated to alter the base complementarity of the nucleotides labeled with the second labeling molecule contained therein.
[0060] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is a modified cytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is treated to alter the base pairing ability of the modified cytosine deoxyribonucleotide it contains (e.g., to pair it with adenine deoxyribonucleotide instead of guanine deoxyribonucleotide).
[0061] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is a 5-aldehyde cytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is treated with a compound (e.g., malononitrile, a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane), or indanedione) to alter the base-pairing ability of the 5-aldehyde cytosine deoxyribonucleotide it contains.
[0062] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is 5-carboxycytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is treated with a compound (e.g., a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane)) to alter the base-pairing ability of the 5-carboxycytosine deoxyribonucleotide it contains.
[0063] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is 5-hydroxymethylcytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is first treated with an oxidant (e.g., potassium ruthenate) or an oxidase (e.g., TET protein), and then with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or indanedione) to alter the base pairing ability of the 5-hydroxymethylcytosine deoxyribonucleotide it contains.
[0064] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is N4-acetylcytosine deoxyribonucleotide (dac 4 C). In such embodiments, after step (3), the labeled product is treated with a compound (e.g., sodium cyanoborohydride) to alter the base-complementary pairing ability of the N4-acetylcytosine deoxyribonucleotide it contains.
[0065] Preferably, the processing step of the labeled product is performed before the labeled product is sequenced, for example, before step (4) or before step (5).
[0066] In some cases, nucleotides labeled with a second marker molecule (e.g., 5-aldehyde cytosine deoxyribonucleotide, 5-hydroxymethyl cytosine deoxyribonucleotide) may be naturally occurring nucleotides within the cell. To avoid the adverse effects of such naturally occurring nucleotides labeled with a second marker molecule (e.g., causing false positive signals), the nucleotides that may be present in the editing product during the editing process can be protected (e.g., by protecting endogenous 5-aldehyde cytosine deoxyribonucleotide with ethyl hydroxylamine, or by protecting endogenous 5-hydroxymethyl cytosine deoxyribonucleotide with glycosylation catalyzed by β-glucosyltransferase (βGT)) before step (3) (e.g., before step (2)) to prevent changes in their base complementarity.
[0067] Therefore, in some embodiments that use nucleotides labeled with a second marker molecule (e.g., 5-aldehyde cytosine deoxyribonucleotide, 5-hydroxymethyl cytosine deoxyribonucleotide), the nucleotides labeled with the second marker molecule that may be present in the editing product are protected before step (3) (e.g., before step (2)).
[0068] For example, in some embodiments, the nucleotide labeled by the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide. In such embodiments, preferably, endogenous 5-aldehyde cytosine deoxyribonucleotide is protected with ethyl hydroxylamine before step (3) (e.g., before step (2)).
[0069] For example, in some embodiments, the nucleotide labeled by the second marker molecule is 5-hydroxymethylcytosine deoxyribonucleotide. In such embodiments, preferably, before step (3) (e.g., before step (2)), the endogenous 5-hydroxymethylcytosine deoxyribonucleotide is protected by a βGT-catalyzed glycosylation reaction (see Cell, 18 Apr 2013, 153(3):678-691, DOI:10.1016 / j.cell.2013.04.001, the full text of which is incorporated herein by reference).
[0070] In some cases, the nucleotides labeled with the second marker molecule (e.g., 5-carboxycytosine deoxyribonucleotide, N4-acetylcytosine deoxyribonucleotide) are not naturally occurring nucleotides in the cell, or although they are naturally occurring nucleotides in the cell, their content is extremely low. In such cases, there is no need to perform nucleotide protection treatment on the edited product before step (3).
[0071] Therefore, in some embodiments that use nucleotides labeled with a second marker molecule (e.g., 5-carboxycytosine deoxyribonucleotide, N4-acetylcytosine deoxyribonucleotide), the edited product is not nucleotide protected prior to step (3).
[0072] In some preferred embodiments, in step (2), a single-strand break is created at the location of the edited base; and in step (3), the nucleotide labeled with the first labeling molecule and the nucleotide labeled with the second labeling molecule are introduced at and downstream of the single-strand break to generate a labeled product containing the first labeling molecule and the second labeling molecule.
[0073] In some preferred embodiments, in step (2), a single-strand break is generated downstream of the edited base; and in step (3), the nucleotide labeled with the first labeling molecule is introduced at or downstream of the single-strand break, and optionally, a nucleotide labeled with a second labeling molecule is introduced, thereby generating a labeled product containing the first labeling molecule and optionally the second labeling molecule.
[0074] In some preferred embodiments, in step (4), the labeled product is separated or enriched using a first binding molecule attached to a solid support. Various suitable solid supports can be used to support the first binding molecule. For example, the solid support may be selected from magnetic beads, agarose beads, or chips.
[0075] In some preferred embodiments, prior to step (5), the method further includes: amplifying the labeled products isolated or enriched in step (4); and / or constructing a sequencing library from the labeled products isolated or enriched in step (4).
[0076] In some preferred embodiments, in step (4), the nucleic acid single strands containing the first label and / or the second label in the labeled product are separated or enriched. For example, in some embodiments, the labeled product may be subjected to a melting treatment (e.g., alkaline treatment), and then a first binding molecule capable of specifically recognizing and binding to the first label molecule may be used to separate or enrich the nucleic acid single strands containing the first label and / or the second label in the labeled product. In some preferred embodiments, the melting treatment (e.g., alkaline treatment) is performed while the first label molecule and the first binding molecule remain bound.
[0077] In some preferred embodiments, prior to step (5), the labeled product isolated or enriched in step (4) is amplified using a nucleic acid polymerase (e.g., a low-fidelity nucleic acid polymerase and / or a high-fidelity nucleic acid polymerase). For example, in some preferred embodiments, the amplification step includes:
[0078] Polymerase chain reaction was performed using a low-fidelity nucleic acid polymerase for up to 5 cycles (e.g., up to 1, 2, 3, 4, 5); and,
[0079] At least three (e.g., at least three, at least five, at least ten, at least twenty, at least thirty, at least forty) polymerase chain reactions were performed using a high-fidelity nucleic acid polymerase.
[0080] It is understood that various suitable methods can be used to construct sequencing libraries from the labeled products isolated or enriched in step (4). Such methods for constructing sequencing libraries are not limited. For example, sequencing libraries with corresponding characteristics can be constructed depending on the sequencing method used. For example, oligonucleotide adapters for sequencing or amplification can be added to the ends of the labeled products as needed for sequencing. In some embodiments, a dA tail can be added to the 3' end of the labeled product, which can be used to ligate with oligonucleotide adapters containing dT tails.
[0081] In some preferred embodiments, in step (5), the sequence of the labeled product is determined by sequencing (e.g., second-generation sequencing or third-generation sequencing), hybridization or mass spectrometry.
[0082] In some preferred embodiments, the method further includes aligning the sequence determined in step (5) with a reference sequence to determine the editing site, editing efficiency, or off-target effect of the base editor on the target nucleic acid.
[0083] In some preferred embodiments, the reference sequence is the target nucleic acid sequence before base editing. For example, the target nucleic acid sequence before base editing can be obtained from a database or by sequencing methods.
[0084] Cytosine base editor and its evaluation
[0085] In a preferred embodiment, the base editor is a cytosine base editor (e.g., a nuclear cytosine base editor, an organelle cytosine base editor). In some preferred embodiments, the cytosine base editor is a cytosine base editor capable of editing cytosine to uracil. For a detailed description of cytosine base editors, see, for example, Andrew V. Anzalone, et al. Nature biotechnology 38(7), 824-844, doi:10.1038 / s41587-020-0561-9 (2020), the full text of which is incorporated herein by reference. In some preferred embodiments, the base editor is a cytosine base editor capable of editing nuclear nucleic acids or a cytosine base editor capable of editing mitochondrial nucleic acids.
[0086] In some preferred embodiments, the edited base is uracil.
[0087] In some preferred embodiments, the base editing intermediate is a nucleic acid molecule (e.g., a DNA molecule) containing uracil.
[0088] In some preferred embodiments, the nucleotide molecule containing the second label is a modified cytosine deoxyribonucleotide that, before treatment, is capable of base pairing with the first nucleotide (e.g., guanine deoxyribonucleotide), and after treatment, is capable of base pairing with the second nucleotide (e.g., adenine deoxyribonucleotide). In some preferred embodiments, the nucleotide molecule containing the second label is selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4 C(N4-acetylcytosine deoxyribonucleotide).
[0089] In some preferred embodiments, in step (2), an AP site-specific endonuclease (e.g., AP endonuclease) is used to create a single-strand break at the location of the edited base in the first nucleic acid strand; and in step (3), the nucleotide labeled with the first marker molecule and the nucleotide labeled with the second marker molecule are introduced at and downstream of the single-strand break to produce a labeled product containing the first and second marker molecules. Subsequently, steps (4) to (5) can be performed as previously described to determine the editing site, editing efficiency, or off-target effects of the cytosine base editor on the target nucleic acid.
[0090] In some preferred embodiments, prior to step (2), the method further includes forming an AP site at the location of the edited base in the first nucleic acid strand.
[0091] For example, in some preferred embodiments, prior to step (2), the method further includes incubating the edited product with UDG (uracil-DNA glycosylase). UDG specifically recognizes uracil nucleotides in the nucleic acid chain and specifically removes uracil from the nucleotides, thereby forming AP sites (depurine / depyrimidine sites) in the nucleic acid chain. Therefore, incubation of the edited product with UDG can convert the edited base (uracil) in the first nucleic acid chain into AP sites.
[0092] In some preferred embodiments, prior to the incubation with UDG step, the method further includes a step of repairing any AP sites that may be present in the edited product.
[0093] In some preferred embodiments, the AP site repair step includes:
[0094] (a) The AP endonuclease is incubated with the editing product, where an AP site may be present, under conditions that allow the AP endonuclease to exert its cleavage activity.
[0095] (b) Under conditions that allow nucleic acid polymerization, the product of step (a) is incubated with a nucleic acid polymerase (e.g., DNA polymerase) and nucleotide molecules (e.g., nucleotide molecules that do not contain a first label or a second label; e.g., dNTPs that do not contain labels);
[0096] (c) Incubate the product of step (b) with a nucleic acid ligase (e.g., DNA ligase) under conditions that allow the nucleic acid ligase to exert its ligation activity.
[0097] Thus, AP sites that may exist in the edited product are repaired.
[0098] It is readily understood that in step (a), the AP endonuclease enables the edited product to create a single-strand break at a possible AP site. In step (b), the nucleic acid polymerase initiates an extension reaction at the single-strand break using a second nucleic acid strand as a template, repairing the single-strand break generated in step (a). In step (c), a nucleic acid ligase (e.g., DNA ligase) ligates the gap in the product of step (b). In some preferred embodiments, the nucleic acid polymerase (e.g., DNA polymerase) in step (b) has strand displacement activity.
[0099] Without theoretical limitations, it is advantageous to repair AP sites prior to step (2). For example, repairing AP sites can eliminate any AP sites that may be present in the edited product. This avoids the introduction of nucleotides labeled with the first and second marker molecules at or downstream of these pre-existing AP sites in subsequent steps, thus preventing interference from these pre-existing AP sites with the detection results.
[0100] In some preferred embodiments, after step (3), the labeled product is treated to alter the base-pairing ability of the nucleotides it contains that are labeled with the second labeling molecule. In some preferred embodiments, the nucleotides labeled with the second labeling molecule are modified cytosine deoxyribonucleotides. In such embodiments, after step (3), the labeled product is treated to alter the base-pairing ability of the modified cytosine deoxyribonucleotides it contains (e.g., to pair with adenine deoxyribonucleotides instead of guanine deoxyribonucleotides).
[0101] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is a 5-aldehyde cytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is treated with a compound (e.g., malononitrile, a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane), or indanedione) to alter the base-pairing ability of the 5-aldehyde cytosine deoxyribonucleotide it contains.
[0102] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is 5-carboxycytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is treated with a compound (e.g., a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane)) to alter the base-pairing ability of the 5-carboxycytosine deoxyribonucleotide it contains.
[0103] In some preferred embodiments, the nucleotide labeled by the second labeling molecule is 5-hydroxymethylcytosine deoxyribonucleotide. In such embodiments, after step (3), the labeled product is first treated with an oxidant (e.g., potassium ruthenate) or an oxidase (e.g., TET protein), and then with a compound (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or indanedione) to alter the base pairing ability of the 5-hydroxymethylcytosine deoxyribonucleotide it contains.
[0104] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is N4-acetylcytosine deoxyribonucleotide (dac 4 C). In such embodiments, after step (3), the labeled product is treated with a compound (e.g., sodium cyanoborohydride) to alter the base-complementary pairing ability of the N4-acetylcytosine deoxyribonucleotide it contains.
[0105] Preferably, the processing step of the labeled product is performed before the labeled product is sequenced, for example, before step (4) or before step (5).
[0106] In some embodiments, prior to step (3) (e.g., prior to step (2)), the nucleotides labeled with a second marker molecule that may be present in the editing product are protected. For example, prior to step (3) (e.g., prior to step (2)), endogenous 5-aldehyde cytosine deoxyribonucleotides may be protected with ethyl hydroxylamine, or endogenous 5-hydroxymethyl cytosine deoxyribonucleotides may be protected with βGT-catalyzed glycosylation.
[0107] For example, in some embodiments that use nucleotides labeled with a second marker molecule (e.g., 5-aldehyde cytosine deoxyribonucleotide, 5-hydroxymethyl cytosine deoxyribonucleotide), the nucleotides labeled with the second marker molecule that may be present in the editing product are protected before step (3) (e.g., before step (2)).
[0108] For example, in some embodiments, the nucleotide labeled by the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide. In such embodiments, preferably, endogenous 5-aldehyde cytosine deoxyribonucleotide is protected with ethyl hydroxylamine before step (3) (e.g., before step (2)).
[0109] For example, in some embodiments, the nucleotide labeled by the second labeling molecule is 5-hydroxymethylcytosine deoxyribonucleotide. In such embodiments, preferably, the endogenous 5-hydroxymethylcytosine deoxyribonucleotide is protected by a βGT-catalyzed glycosylation reaction prior to step (3) (e.g., prior to step (2)).
[0110] In some embodiments that use nucleotides labeled with a second marker molecule (e.g., 5-carboxycytosine deoxyribonucleotide, N4-acetylcytosine deoxyribonucleotide), the edited product is not nucleotide protected prior to step (3).
[0111] Adenine base editor and its evaluation
[0112] In a preferred embodiment, the base editor is an adenine base editor. In some preferred embodiments, the adenine base editor is an adenine base editor capable of editing adenine to hypoxanthine, such as adenine base editors ABE7.10, ABEmax, and ABE8e. A detailed description of adenine base editors can be found, for example, Andrew V. Anzalone, et al. Nature biotechnology 38(7), 824-844, doi:10.1038 / s41587-020-0561-9 (2020), the full text of which is incorporated herein by reference.
[0113] In some preferred embodiments, the edited base is hypoxanthine.
[0114] In some preferred embodiments, the base editing intermediate is a nucleic acid molecule (e.g., a DNA molecule) containing hypoxanthine.
[0115] In some preferred embodiments, in step (2), a hypoxanthine site-specific endonuclease (e.g., endonuclease V, or endonuclease VIII) is used to create a single-strand break at or downstream of the site of the edited base in the first nucleic acid strand; and in step (3), the nucleotide labeled with the first marker molecule is introduced at and downstream of the single-strand break, and optionally, a nucleotide labeled with a second marker molecule is introduced to produce a labeled product containing the first marker molecule and optionally the second marker molecule. Subsequently, steps (4) to (5) can be performed as previously described to determine the editing site, editing efficiency, or off-target effects of the adenine base editor on the target nucleic acid.
[0116] In some preferred embodiments, in step (2), a single-strand break is created downstream of the edited base in the first nucleic acid strand using endonuclease V; or, a single-strand break is created at the position of the edited base in the first nucleic acid strand using endonuclease VIII.
[0117] In such embodiments, hypoxanthine in the labeled product is read as guanine (G) during sequencing, thereby generating an A-to-G base mutation signal in the sequencing results of the labeled product. By detecting this base mutation signal, the edited base can be precisely located. Therefore, in such embodiments, the use of a nucleotide labeled with a second marker molecule is not necessary. Thus, in some exemplary embodiments, in step (3), no nucleotide labeled with a second marker molecule is introduced at or downstream of the single-strand break.
[0118] However, it is readily understood that the base mutation signal can be further amplified using nucleotides labeled with a second labeling molecule, thereby improving the sensitivity of detection. Therefore, in some exemplary embodiments, in step (3), a nucleotide labeled with a second labeling molecule is introduced at or downstream of the single-strand break.
[0119] It is also readily understood that the detailed description above of nucleotides labeled with a second label molecule also applies here. For example, in some preferred embodiments, the second-labeled nucleotide molecule is selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4 C(N4-acetylcytosine deoxyribonucleotide).
[0120] Furthermore, as described above, in embodiments using nucleotides labeled with a second labeling molecule, preferably, after step (3), the labeled product is treated to alter the base complementarity of the nucleotides labeled with the second labeling molecule contained therein; and / or, before step (3) (e.g., before step (2), the nucleotides labeled with the second labeling molecule that may be present in the edited product are protected. For details regarding the treatment and protection of nucleotides labeled with the second labeling molecule, please refer to the detailed description above.
[0121] Dibase Editor and its Evaluation
[0122] In a preferred embodiment, the base editor is a dibase editor.
[0123] In some preferred embodiments, the base editor is a base editor capable of editing cytosine to uracil and adenine to hypoxanthine.
[0124] In some preferred embodiments, the edited base is hypoxanthine and / or uracil.
[0125] In some preferred embodiments, the base editing intermediate is a nucleic acid molecule (e.g., a DNA molecule) containing hypoxanthine and / or uracil.
[0126] It is easy to understand that the editing products of target nucleic acids edited by dual-base editors (such as adenine and cytosine dual-base editors) also contain the same editing bases as those generated by editing target nucleic acids by single-base editors (such as cytosine base editors and adenine base editors). Therefore, the descriptions above regarding cytosine base editors and adenine base editors and their evaluations also apply to adenine and cytosine dual-base editors.
[0127] In some preferred embodiments, the scheme described above for cytosine base editors is used to detect the editing sites, editing efficiency, or off-target effects of a dual-base editor (e.g., an adenine-cytosine dual-base editor) on target nucleic acids. For example, the scheme can be used to detect the editing sites, editing efficiency, or off-target effects of cytosine in target nucleic acids edited by a dual-base editor (e.g., an adenine-cytosine dual-base editor).
[0128] In some preferred embodiments, the scheme described above for an adenine base editor is used to detect the editing sites, editing efficiency, or off-target effects of a dual-base editor (e.g., an adenine-cytosine dual-base editor) on the target nucleic acid. For example, the scheme can be used to detect the editing sites, editing efficiency, or off-target effects of adenine in the target nucleic acid edited by a dual-base editor (e.g., an adenine-cytosine dual-base editor).
[0129] In one aspect, this application also provides a kit comprising an enzyme or combination of enzymes capable of generating single-strand breaks within a segment containing edited bases, comprising a nucleotide molecule labeled with a first labeling molecule and a first binding molecule capable of specifically recognizing and binding to the first labeling molecule; wherein the endonuclease or combination thereof is capable of specifically recognizing the base-editing intermediate containing the edited bases and is capable of generating phosphodiester bond breaks within a segment from 10 nt upstream (e.g., 10 nt, 9 nt, 8 nt, 7 nt, 6 nt, 5 nt, 4 nt, 3 nt, 2 nt, 1 nt) to 10 nt downstream (e.g., 10 nt, 9 nt, 8 nt, 7 nt, 6 nt, 5 nt, 4 nt, 3 nt, 2 nt, 1 nt) of the edited bases.
[0130] In some preferred embodiments, the enzyme or combination of enzymes capable of producing single-strand breaks within the segment containing the edited bases is endonuclease V or endonuclease VIII.
[0131] In some preferred embodiments, the enzyme or combination of enzymes capable of producing single-strand breaks within the segment containing the edited bases is a combination of UDG enzyme and AP endonuclease.
[0132] In some preferred embodiments, the kit further comprises nucleotide molecules labeled with a second labeling molecule, said nucleotide being capable of base-complementary pairing with different nucleotides under different conditions (e.g., before and after treatment). In some preferred embodiments, the nucleotide molecules labeled with the second labeling molecule are selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4C(N4-acetylcytosine deoxyribonucleotide).
[0133] In some preferred embodiments, the nucleotide molecule containing the second label is a modified cytosine deoxyribonucleotide that, before treatment, is capable of base pairing with the first nucleotide (e.g., guanine deoxyribonucleotide), and after treatment, is capable of base pairing with the second nucleotide (e.g., adenine deoxyribonucleotide). In some preferred embodiments, the nucleotide molecule containing the second label is selected from d5fC (5-aldehyde cytosine deoxyribonucleotide), d5caC (5-carboxycytosine deoxyribonucleotide), d5hmC (5-hydroxymethylcytosine deoxyribonucleotide), and dac 4 C(N4-acetylcytosine deoxyribonucleotide).
[0134] In some preferred embodiments, the kit further comprises a reagent for protecting the nucleotide molecules labeled with the second labeling molecule (e.g., ethyl hydroxylamine, reagents required for βGT-catalyzed glycosylation reactions (e.g., β-glucosyltransferase, glucosyl compounds), or any combination thereof), and / or a reagent for treating the nucleotide molecules labeled with the second labeling molecule to alter their base pairing ability (e.g., malononitrile, indane, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), potassium ruthenate, TET protein, sodium cyanoborohydride, or any combination thereof).
[0135] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide. In such embodiments, the kit may also contain a reagent (e.g., ethyl hydroxylamine) for protecting the nucleotide molecule labeled with the second labeling molecule, and / or a reagent (e.g., malononitrile, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), or indendione) for treating the nucleotide molecule labeled with the second labeling molecule to alter its base pairing ability.
[0136] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is 5-hydroxymethylcytosine deoxyribonucleotide. In such embodiments, the kit may also contain a reagent protecting the nucleotide molecule labeled with the second labeling molecule (e.g., a reagent required for βGT-catalyzed glycosylation (e.g., β-glucosyltransferase, a glucosyl compound)), and / or a reagent treating the nucleotide molecule labeled with the second labeling molecule to alter its base pairing ability (e.g., potassium ruthenate or TET protein, and malononitrile or borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane) or azidoindenide).
[0137] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is 5-carboxycytosine deoxyribonucleotide. In such embodiments, the kit may also include a reagent (e.g., a borane compound (e.g., a pyridineborane compound, such as pyridineborane or 2-methylpyridineborane)) for treating the nucleotide molecule labeled with the second labeling molecule to alter its base pairing ability.
[0138] In some preferred embodiments, the nucleotide labeled with the second labeling molecule is N4-acetylcytosine deoxyribonucleotide. In such embodiments, the kit may also include a reagent (e.g., sodium cyanoborohydride) for treating the nucleotide molecule labeled with the second labeling molecule to alter its base pairing ability.
[0139] In some preferred embodiments, the kit further comprises a nucleic acid polymerase (e.g., a nucleic acid polymerase with strand displacement activity), a nucleic acid ligase (e.g., DNA ligase), unlabeled nucleotide molecules, a reagent for protecting nucleotide molecules labeled with a second labeling molecule (e.g., ethyl hydroxylamine, a reagent required for βGT-catalyzed glycosylation (e.g., β-glucosyltransferase, a glucosyl compound), or any combination thereof), a reagent for treating nucleotide molecules labeled with a second labeling molecule to alter their base pairing ability (e.g., malononitrile, indanedione azide, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), potassium ruthenate, TET protein, sodium cyanoborohydride, or any combination thereof), or any combination thereof.
[0140] It is readily understood that the kit is used to implement the methods of this application. Therefore, the detailed descriptions above of base editors (e.g., single-base editors and double-base editors), first labeling molecules, first binding molecules, nucleotide molecules labeled with the first labeling molecule, second labeling molecules, nucleotide molecules labeled with the second labeling molecule, nucleic acid polymerases, nucleic acid ligases, UDG enzymes, AP endonucleases, endonucleases V or VIII, etc., also apply here.
[0141] In some preferred embodiments, the kit is used to detect the editing site, editing efficiency, or off-target effects of a base editor (e.g., a single-base editor or a dual-base editor) editing a target nucleic acid.
[0142] In some preferred embodiments, the kit is used to detect the editing site, editing efficiency, or off-target effects of a cytosine base editor editing a target nucleic acid. In some preferred embodiments, the kit comprises: a UDG enzyme, an AP endonuclease, a nucleotide molecule labeled with a first labeling molecule, a first binding molecule, and a nucleotide molecule labeled with a second labeling molecule (e.g., d5fC, d5caC, d5hmC, or dac). 4C); optionally also comprising, nucleic acid polymerase, nucleic acid ligase, unlabeled nucleotide molecule, reagent for protecting nucleotide molecules labeled with a second labeling molecule (e.g., ethylhydroxylamine, reagent required for βGT-catalyzed glycosylation (e.g., β-glucosyltransferase, glucosyl compound), or any combination thereof), reagent for treating nucleotide molecules labeled with a second labeling molecule to alter their base pairing ability (e.g., malononitrile, azidoindenidine, borane compound (e.g., pyridineborane compound, such as pyridineborane or 2-methylpyridineborane), potassium ruthenate, TET protein, sodium cyanoborohydride, or any combination thereof), or any combination thereof.
[0143] In some preferred embodiments, the kit is used to detect the editing site, editing efficiency, or off-target effects of an adenine base editor on target nucleic acids. In some preferred embodiments, the kit includes an endonuclease V or VIII, a nucleotide molecule labeled with a first labeling molecule, and a first binding molecule; optionally, it also includes a nucleic acid polymerase, a nucleic acid ligase, and a nucleotide molecule labeled with a second labeling molecule (e.g., d5fC, d5caC, d5hmC, or dac). 4 C) Unlabeled nucleotide molecules, reagents protecting nucleotide molecules labeled with a second labeling molecule (e.g., ethylhydroxylamine, reagents required for βGT-catalyzed glycosylation reactions (e.g., β-glucosyltransferase, glucosyl compounds), or any combination thereof), reagents that treat nucleotide molecules labeled with a second labeling molecule to alter their base pairing ability (e.g., malononitrile, azidoindenidine, borane compounds (e.g., pyridineborane compounds, such as pyridineborane or 2-methylpyridineborane), potassium ruthenate, TET protein, sodium cyanoborohydride, or any combination thereof), or any combination thereof.
[0144] In some preferred embodiments, the kit is used to detect the editing site, editing efficiency, or off-target effects of a dual-base editor (e.g., an adenine-cytosine dual-base editor) on the target nucleic acid. In some preferred embodiments, the kit comprises: UDG enzyme, AP endonuclease, endonuclease V or VIII, a nucleotide molecule labeled with a first labeling molecule, a first binding molecule, and a nucleotide molecule labeled with a second labeling molecule (e.g., d5fC, d5caC, d5hmC, or dac). 4C); optionally also comprising, nucleic acid polymerase, nucleic acid ligase, unlabeled nucleotide molecule, reagent for protecting nucleotide molecules labeled with a second labeling molecule (e.g., ethylhydroxylamine, reagent required for βGT-catalyzed glycosylation (e.g., β-glucosyltransferase, glucosyl compound), or any combination thereof), reagent for treating nucleotide molecules labeled with a second labeling molecule to alter their base pairing ability (e.g., malononitrile, azidoindenidine, borane compound (e.g., pyridineborane compound, such as pyridineborane or 2-methylpyridineborane), potassium ruthenate, TET protein, sodium cyanoborohydride, or any combination thereof), or any combination thereof.
[0145] Terminology Definition
[0146] In this application, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the nucleic acid chemistry laboratory procedures used herein are all standard procedures widely used in the relevant fields. To better understand this invention, definitions and explanations of related terms are provided below. Unless specifically defined or described differently elsewhere herein, the terms and descriptions relating to this invention below should be understood according to the definitions given below.
[0147] When the terms “for example,” “such as,” “like,” “including,” “contains,” or variations thereof are used herein, these terms will not be considered restrictive terms but will be interpreted as meaning “but not limited to” or “not limited to.”
[0148] Unless otherwise specified herein or clearly contradicted by the context, the terms “an” and “a kind” as well as “the” and similar designations shall be interpreted to cover both the singular and the plural in the context of describing the invention (especially in the context of the following claims).
[0149] As used herein, the term "base editor" refers to a reagent comprising a polypeptide capable of editing or modifying bases (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA or RNA). In some embodiments, the base editor is a single-base editor or a two-base editor.
[0150] In some embodiments, the base editor is a single-base editor capable of editing a single base within a nucleic acid molecule (e.g., a DNA molecule); for example, it is capable of deaminating a single base within a nucleic acid molecule (e.g., a DNA molecule). In some embodiments, the single-base editor is capable of deaminating adenine (A) in DNA. In some embodiments, the single-base editor is capable of deaminating cytosine (C) in DNA. In some embodiments, the single-base editor comprises adenosine deaminase and a nucleic acid-programmable DNA-binding protein (napDNAbp), for example, a fusion protein comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) fused with adenosine deaminase. In some embodiments, the single-base editor comprises cytidine deaminase and a nucleic acid-programmable DNA-binding protein (napDNAbp), for example, a fusion protein comprising napDNAbp fused with cytidine deaminase. In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 protein, such as Cas9 Nickase (nCaS9), which can only cleave one strand of a nucleic acid duplex, or Cas9 (dCaS9), which has no nuclease activity.
[0151] In some embodiments, the single-base editor comprises adenosine deaminase and a Cas9 protein, for example, a Cas9 protein fused with adenosine deaminase. In some embodiments, the single-base editor comprises cytidine deaminase and a Cas9 protein, for example, a Cas9 protein fused with cytidine deaminase. In some embodiments, the single-base editor comprises adenosine deaminase and nCaS9, for example, an nCaS9 fused with adenosine deaminase. In some embodiments, the single-base editor comprises cytidine deaminase and nCaS9, for example, an nCaS9 fused with cytidine deaminase. In some embodiments, the single-base editor comprises adenosine deaminase and dCaS9, for example, dCaS9 fused with adenosine deaminase. In some embodiments, the single-base editor comprises cytidine deaminase and dCaS9, for example, dCaS9 fused with cytidine deaminase.
[0152] In some embodiments, the base editor is a two-base editor capable of editing two bases within a nucleic acid molecule (e.g., a DNA molecule); for example, it can deaminate two bases within a nucleic acid molecule (e.g., a DNA molecule). In some embodiments, the two-base editor can deaminate adenine (A) and cytosine (C) in DNA. In some preferred embodiments, the two-base editor can deaminate adenine (A) and cytosine (C) located within the same editing window in DNA. In some embodiments, the two-base editor comprises adenosine deaminase, cytidine deaminase, and a nucleic acid programmable DNA-binding protein (napDNAbp). In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is a Cas9 protein, such as Cas9 Nickase (nCaS9), which can only cleave one strand of a nucleic acid duplex, or Cas9 (dCaS9), which has no nuclease activity. In some embodiments, the two-base editor comprises adenosine deaminase, cytidine deaminase, and Cas9 protein. In some embodiments, the dual-base editor comprises adenosine deaminase, cytidine deaminase, and Cas9Nickase (nCaS9). In some embodiments, the dual-base editor comprises adenosine deaminase, cytidine deaminase, and nuclease-free Cas9 (dCaS9). In some embodiments, the dual-base editor is a complex or fusion protein comprising adenosine deaminase, cytidine deaminase, and nap DNAbp.
[0153] It is readily understood that the dual-base editor may comprise one or more (e.g., one or two) nucleic acid-programmable DNA-binding proteins (napDNAbp). In some embodiments, the dual-base editor comprises two napDNAbps, each independently fused to adenosine deaminase and cytidine deaminase. In some embodiments, the dual-base editor comprises one napDNAbp, simultaneously fused to both adenosine deaminase and cytidine deaminase. In some embodiments, the dual-base editor is a combination of two single-base editors.
[0154] In some embodiments, the base editor is fused to a base excision repair inhibitor (e.g., a UGI or DISN domain). In some embodiments, the fusion protein comprises an nCas9 fused to a deaminase and a base excision repair inhibitor, such as a UGI or DISN domain. In some embodiments, the base excision repair inhibitor, such as a UGI or DISN domain, is provided in the system but is not fused to the Cas9 protein (or dCas9, nCas9). It is important to emphasize that the terms "fused with" or "fused to" herein include fusion or linkage between proteins (or their functional domains) with or without a linker. In some embodiments, the "linker" is a peptide linker. In some embodiments, the "linker" is a non-peptide linker.
[0155] In some embodiments, the deaminase included in the base editor and the nucleic acid-programmable DNA-binding protein are structurally independent of each other; that is, the deaminase included in the base editor and the nucleic acid-programmable DNA-binding protein are not fused or linked through a linker. In some embodiments, the deaminase included in the base editor and the nucleic acid-programmable DNA-binding protein are non-covalently linked or bound.
[0156] It is readily understood that the deaminase can be a specific deaminase of a glycoside of any base formation or a combination thereof (e.g., adenosine deaminase, cytidine deaminase).
[0157] In some embodiments, the programmable DNA-binding protein may be selected from TALEs, ZFs, Casx, Casy, Cpf1, C2c1, C2c2, C2c3, Argonaute proteins, or derivatives thereof. In some embodiments, the programmable DNA-binding protein does not have nuclease activity. In some embodiments, the programmable DNA-binding protein can only cleave one strand of the nucleic acid double helix. In some embodiments, the programmable DNA-binding protein does not have the activity of forming a double-strand break in the nucleic acid.
[0158] In some implementations, the base editor is a cytosine base editor, such as the cytosine base editor BE3, the upgraded cytosine base editor BE4max, the mitochondrial cytosine base editor DdCBE, and various CBE editing systems. For a description of various cytosine base editors, see, for example, Andrew V. Anzalone, et al. Nature biotechnology 38(7), 824-844, doi:10.1038 / s41587-020-0561-9 (2020), the full text of which is incorporated herein by reference.
[0159] In some embodiments, the base editor is an adenine base editor, such as the adenine base editor ABE7.10, the adenine base editor ABEmax, and the adenine base editor ABE8e, as well as various ABE editing systems. For a detailed description of various adenine base editors, see, for example, Andrew V. Anzalone, et al. Nature biotechnology 38(7), 824-844, doi:10.1038 / s41587-020-0561-9 (2020), the full text of which is incorporated herein by reference.
[0160] In some implementations, the base editor is a base editor capable of editing both adenine and cytosine, such as ACBE.
[0161] As used herein, the term "base editing intermediate" refers to a product of a base editor (e.g., a single-base editor or a dual-base editor) editing a target nucleic acid, containing edited bases generated by the base editor editing the target nucleic acid. The target nucleic acid can be derived from any living organism (e.g., eukaryotic cells, prokaryotic cells, viruses, and viroids) or a non-living organism (e.g., a nucleic acid molecular library). In some embodiments, the base editing intermediate is a direct product of the base editor editing the target nucleic acid. In some embodiments, the base editing intermediate is a product obtained by enrichment and / or nucleic acid fragmentation of the direct product of the base editor editing the target nucleic acid. In some embodiments, the edited base is a base (e.g., uracil, hypoxanthine) modified by a corresponding active element (e.g., cytidine deaminase, adenosine deaminase) in the base editor. Generally, the bases before and after modification / editing have different base-pairing capabilities (i.e., they can pair complementaryly with different bases). For example, cytosine in nucleic acids is converted into uracil by cytidine deaminase in a base editor. Uracil pairs complementaryly with adenine, not guanine. Similarly, adenine in nucleic acids is converted into hypoxanthine by adenine deaminase in a base editor. Hypoxanthine pairs complementaryly with cytosine, not thymine.
[0162] As used herein, the term "borane compound" refers to a borane compound that can be used to treat the second-labeled nucleotides of this application to alter their base pairing ability. In particular, pyridineborane compounds include pyridineborane and its derivatives. Non-limiting examples of said pyridineborane compounds are pyridineborane and 2-methylpyridineborane (see, for example, Liu, Y. et al. Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution. Nature biotechnology 37, 424-429, doi:10.1038 / s41587-019-0041-2 (2019), the entire text of which is incorporated herein by reference).
[0163] As used herein, the term “upstream” describes the relative position of two nucleic acid sequences (or two nucleic acid molecules) and has a meaning commonly understood by those skilled in the art. For example, stating that “one nucleic acid sequence is upstream of another nucleic acid sequence” means that, when aligned in a 5’ to 3’ direction, the former is located further forward (i.e., closer to the 5’ end) than the latter. As used herein, the term “downstream” has the opposite meaning to “upstream”.
[0164] As used herein, the term "first labeled molecule" refers to a molecule capable of specifically forming an interacting molecule pair with a first binding molecule. According to the method of this application, the specific binding of the first binding molecule to the first labeled molecule can be used to enrich the labeled product containing the first labeled molecule. In some embodiments, the first labeled molecule binds to the first binding molecule reversibly or irreversibly. In some preferred embodiments, the first labeled molecule binds to the first binding molecule reversibly.
[0165] As used herein, the term "nucleotide labeled with a first labeling molecule" refers to a nucleotide molecule containing a group in the first labeling molecule capable of specifically forming an interacting molecule pair with a first binding molecule. In some preferred embodiments, the nucleotide labeled with the first labeling molecule refers to a mononucleotide molecule, such as dUTP, dATP, dTTP, dCTP, or dGTP labeled with the first labeling molecule, or any combination thereof.
[0166] In some embodiments, the labeled nucleotide molecule is reversibly or irreversibly linked to the first labeling molecule. In some embodiments, the ribose, base, or phosphate moiety of the labeled nucleotide molecule is reversibly or irreversibly linked to the first labeling molecule. In some preferred embodiments, the labeled nucleotide molecule is reversibly linked to the first labeling molecule. It should be noted that in some cases, the nucleotide molecule labeled with the first labeling molecule does not contain the complete structure of the first labeling molecule, but contains groups in the first labeling molecule that can specifically form interacting molecular pairs with the first binding molecule.
[0167] As used herein, the term "second marker molecule" refers to a molecule that can modify a base in a nucleotide molecule to produce a modified base that can pair complementaryly with different bases under different conditions (e.g., before and after treatment).
[0168] As used herein, the term "nucleotide labeled with a second labeling molecule" refers to a nucleotide molecule that can perform base complementary pairing with different nucleotides under different conditions (e.g., before and after treatment). In some preferred embodiments, the nucleotide labeled with a second labeling molecule refers to a single nucleotide molecule.
[0169] As used herein, a nucleic acid polymerase with "strand displacement activity" refers to a nucleic acid polymerase that, during the extension of a new nucleic acid strand, encounters a downstream nucleic acid strand complementary to the template strand and can continue the extension reaction while degrading (rather than stripping) the complementary nucleic acid strand. In some preferred embodiments, the nucleic acid polymerase with "strand displacement activity" also has 5' to 3' exonuclease activity.
[0170] As used herein, "high-fidelity nucleic acid polymerase" refers to a nucleic acid polymerase that, during nucleic acid amplification, introduces erroneous nucleotides (i.e., has a lower error rate) than wild-type Taq polymerases (e.g., Taq polymerases with sequences such as UniProt Acession: P19821.1). For example, Start High-Fidelity DNA Polymerase.
[0171] As used herein, "low-fidelity nucleic acid polymerase" refers to a nucleic acid polymerase that, during nucleic acid amplification, has a higher probability (i.e., error rate) of introducing incorrect nucleotides than wild-type Taq polymerases (e.g., Taq polymerases with sequences such as UniProt Acession: P19821.1). For example, MightyAmp DNA Polymerase.
[0172] As used herein, unless the context clearly indicates otherwise, the term "nucleotide" as used herein preferably refers to nucleoside triphosphates, such as deoxyribonucleoside triphosphates.
[0173] Beneficial effects
[0174] This application provides a novel method for detecting the site, efficiency, or off-target effects of nucleic acid editing by a base editor (e.g., a cytosine base editor, an adenine base editor, or an adenine and cytosine dual base editor), which has one or more beneficial technical effects selected from the following:
[0175] (1) The method of the present invention can capture base editing intermediates (e.g., nucleic acids containing uracil or hypoxanthine) generated by base editing tools in living cells, and thus can obtain site information of actual base editing events.
[0176] (2) The method of the present invention can effectively label and enrich the editing site, thereby making it very easy to distinguish it from gene background such as SNV and sequencing errors.
[0177] (3) In existing technologies, when detecting base editing sites using whole-genome sequencing, the coverage of the sequencing reads across the entire genome is highly uneven, requiring a huge amount of data to obtain sufficient information to evaluate the editing sites throughout the genome. The method of this invention overcomes this difficulty and can obtain strong detection signals at the whole-genome level with a relatively small amount of data.
[0178] (4) The method of this invention is not biased towards various base editing tools (e.g., CBE, ABE). As mentioned above, various optimized base editing tools have been developed to meet practical needs. Since the method of this invention can capture base editing intermediates (e.g., nucleic acids containing uracil or hypoxanthine) generated in various base editing processes, the method of this invention is universally applicable to the detection of editing sites of various base editing tools and can evaluate their editing efficiency or off-target effects.
[0179] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0180] Figure 1 An exemplary scheme 1 for detecting the edit site of a base editor using the method of the present invention is shown, wherein the base editor is a cytosine base editor.
[0181] The first step involves extracting nucleic acids (e.g., genomic DNA or mitochondrial DNA) edited by a cytosine base editor, containing a base-editing intermediate (e.g., DNA containing uracil). This intermediate is a product of the cytosine base editor editing the target nucleic acid and comprises a first nucleic acid strand and a second nucleic acid strand. The first nucleic acid strand contains edited bases (e.g., uracil) generated by the cytosine base editor editing the target nucleic acid. The nucleic acids are then fragmented using methods such as sonication to form, for example, nucleic acid fragments of approximately 300 bp. The fragmented genomic DNA fragments are then trimmed to blunt ends using an end-repair process. In some exemplary embodiments, the end-repair process includes the removal of 3' overhangs and the filling of 5' overhangs. In some preferred embodiments, the end-repair process can be performed using a nucleic acid polymerase containing 3' to 5' exonuclease activity.
[0182] The second step involves incorporating, via in vitro BER (Base Excision Repair Pathway) labeling, the location of the edited base (e.g., uracil) in the base editing intermediate and downstream thereof, with a nucleotide (e.g., uracil deoxyribonucleotide) labeled by a first labeling molecule (e.g., biotin) and a nucleotide (e.g., 5-aldehyde cytosine deoxyribonucleotide) labeled by a second labeling molecule. In some exemplary embodiments, the BER labeling method includes: using UDG (uracil-DNA glycosylase) to specifically recognize and excise uracil on the edited product generated by the cytosine base editor editing the target nucleic acid, generating an AP site; using an AP endonuclease to excise the debasement site, generating a single-strand gap; using a DNA polymerase containing strand displacement activity to perform a DNA strand displacement reaction from the generated single-strand gap along the 5' to 3' direction; and using a DNA ligase to ligate the single-strand cut in the DNA strand displacement reaction product. In the DNA strand replacement reaction system, at least one nucleotide substrate (e.g., biotin-uracil ribonucleotide) labeled with a first marker molecule (e.g., biotin) is used instead of a conventional nucleotide substrate (e.g., thymine deoxyribonucleotide). In some preferred embodiments, the DNA strand replacement reaction system further includes at least one nucleotide substrate (e.g., 5-aldehyde cytosine deoxyribonucleotide) labeled with a second marker molecule instead of a conventional nucleotide substrate (e.g., cytosine deoxyribonucleotide). The incorporation of the nucleotide labeled with the first marker molecule (e.g., biotin-uracil deoxyribonucleotide) allows for subsequent enrichment of the nucleic acid fragment containing the first marker molecule using a first binding molecule (e.g., streptavidin), wherein the first binding molecule specifically interacts with the first marker molecule. The nucleotide labeled with the second marker molecule can form complementary base pairs with different nucleotides under different conditions (e.g., before and after treatment). For example, the nucleotide labeled by the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide (d5fC); it can form a base pairing with guanine deoxyribonucleotide before treatment with a compound (e.g., malononitrile, or azidoindone), and can form a base pairing with adenine deoxyribonucleotide after treatment with a compound (e.g., malononitrile, or azidoindone). Thus, the labeled product containing d5fC can generate a C-to-T mutation signal at the d5fC incorporation site through a subsequent chemical reaction, thereby achieving precise localization of the edited base (e.g., uracil).
[0183] In some preferred embodiments, to avoid false positive signals that may arise from DNA damage or modification (e.g., SSB or AP sites) introduced during endogenous or nucleic acid manipulation, the method further includes nucleic acid repair treatment of the edited product before proceeding to the second step. In some exemplary embodiments, the treatment includes: removing the AP site with an AP endonuclease to create a single-strand gap; performing a DNA strand displacement reaction with a DNA polymerase starting from the created single-strand gap or a possible SSB gap in the nucleic acid strand along the 5' to 3' direction; and ligating the gap in the strand displacement reaction product with a DNA ligase. In some preferred embodiments, the DNA polymerase has strand displacement activity.
[0184] In some preferred embodiments, to avoid the adverse effects of endogenous nucleotides labeled with a second marker molecule (e.g., endogenous 5-aldehyde cytosine deoxyribonucleotides), the method further includes protecting the nucleotides labeled with the second marker molecule that may be present in the editing product before proceeding to the second step. For example, ethyl hydroxylamine (EtONH2) may be used to protect the 5-aldehyde cytosine deoxyribonucleotides that may be present in the editing product before proceeding to the second step to prevent them from reacting with compounds (e.g., malononitrile, or indanedione) and forming false positive base switching signals.
[0185] The third step involves treating the nucleic acid containing the nucleotides labeled with the second marker molecule generated in the previous step to alter the base-pairing ability of the labeled nucleotides. In some preferred embodiments, the labeled nucleotide is 5-aldehyde cytosine deoxyribonucleotide. As described above, the 5-aldehyde cytosine deoxyribonucleotide treated with a compound (e.g., malononitrile or azidoindenidine) will pair with adenine deoxyribonucleotide during subsequent DNA replication, thereby generating a C-to-T mutation signal at the position of the 5-aldehyde cytosine deoxyribonucleotide in the sequencing results of the amplified product of the treated nucleic acid.
[0186] The fourth step involves enriching DNA fragments containing a first labeled molecule (e.g., biotin) using a solid support (e.g., magnetic beads) coupled with a first binding molecule (e.g., streptavidin); these fragments, optionally amplified and / or used for library construction, can then be used for high-throughput sequencing. Based on the sequencing results, the location information of the editing sites in the base-editing intermediates generated after the cytosine base editor edits the target nucleic acid can be analyzed.
[0187] In some preferred embodiments, prior to amplification and / or library construction of the enriched DNA fragments, the DNA fragments enriched on a solid support (e.g., magnetic beads) may be treated (e.g., alkali treatment) to remove the complementary strand of the nucleic acid single strand containing the first marker molecule (e.g., biotin).
[0188] In some exemplary embodiments, oligonucleotide adapters are ligated to the ends of enriched DNA fragments via an adapter ligation reaction prior to treatment with an alkali (e.g., NaOH) to remove the complementary strand of the nucleic acid single strand containing a first labeling molecule (e.g., biotin), to facilitate amplification or sequencing of the DNA fragments. In some preferred embodiments, a dA tail is added to the 3' end of the DNA fragment, which can be used for ligation to oligonucleotide adapters containing dT tails.
[0189] Figure 2 A schematic diagram (a) of different pattern sequences used in the method of Embodiment 1 of the present invention is shown, as well as the enrichment results (b) of the method of Embodiment 1 of the present invention on different pattern sequences.
[0190] Figure 3 The high-throughput sequencing signals generated on the pattern sequence by the method of Embodiment 1 of the present invention are shown. (a) High-throughput sequencing results of the pattern sequence containing dU:dG base pairs. Gray dashed lines indicate the positions of dU:dG base pairs, and red blocks represent C-to-T mutation signals; (b) Statistical calculation results of the proportion of C-to-T mutations at different positions on the pattern sequence based on high-throughput sequencing data. Gray dashed lines indicate the positions of dU:dA base pairs, solid red dots indicate the positions of continuous C-to-T mutation signals, and hollow dots indicate the positions of C signals below the background level.
[0191] Figure 4The following diagram illustrates the signals generated on genomic DNA by the method of Example 1 of this invention. (a) Signals generated at on-target sites. The upper half indicates the signals generated at the EMX1 on-target site in samples obtained from the HEK293T cell line using different editing components and treatment methods according to the method of this invention. The lower half indicates the signals generated at the VEGFA_site_2 on-target site in samples obtained from the HEK293T cell line using different editing components and treatment methods according to the method of this invention. In the sample names, "IN" indicates the input sample, "NT" indicates the sample transfected with BE4max and non-target sgRNA, "rep1" indicates repeat 1, and "rep2" indicates repeat 2; the green "A" is equivalent to indicating C-to-T signals on the non-target strand; (b) Statistics of continuous C-to-T mutation signals generated at the whole genome level. The left half counts the distance of generated mutation signals, and the right half counts the number of generated mutations; (c) Signal at a certain off-target site in the VEGFA_site_2 sample. The red blocks indicate "C-to-T" mutations on the non-target strand, the red inverted triangles indicate the actual location edited by CBE, the black inverted triangles indicate "G-to-T" SNVs, and the brown shading indicates pRBS, i.e., the putative sgRNA binding site; (d) Comparison of the signal of this invention (left) and WGS signal (right) within a 4kb range before and after pRBS (dark blue) or random site (light green).
[0192] Figure 5 This diagram shows the composition of the plasmids used in the comparative experiments of deleting different components of the CBE system.
[0193] Figure 6The results of the detection of non-Cas-dependent off-targets are shown. (a) Examples of signals from non-Cas-dependent off-target sites in different samples. The red "T" in the (-)sgRNA sample indicates the C-to-T signal generated by the method of this invention, which was not observed in other samples; (b) The number of non-Cas-dependent off-target sites identified in different samples; (c) The intersection of non-Cas-dependent off-target sites identified in each All and (-)sgRNA sample; (d) Sequence motif analysis results at such non-Cas-dependent off-target sites in different samples. The 10 bp neighboring sequences on both sides of each site (referencing the hg38 genome) were extracted and sequenced using WebLogo software; (e) The non-Cas-dependent off-target sites identified by the method of this invention are enriched in transcriptionally active regions of the genome; (f) The non-Cas-dependent off-target sites identified by this invention are more concentrated in regions of highly expressed genes. All P-values were calculated using a one-sided Student's t-test.
[0194] Figure 7 The results of Cas-dependent off-target detection are shown. (a) Signal examples of Cas-dependent off-target sites in different samples. In the enlarged IGV (Integrative Genomics Viewer) image on the right, the green blocks represent "G-to-A" mutations, equivalent to "C-to-T" mutations on the non-target strand; (b) Cas-dependent off-target sites identified in the two biological replicate samples "VEGFA_site_2-ALL". Under very strict bioinformatics analysis identification rules (cufoff), 384 sites were identified as duplicates (orange dots; including on-target), but the signal intensity of the remaining rep-only sites (blue dots) was not low in both samples; (c) Comparison of the signals of all Cas-dependent off-target sites at the whole genome level in different samples. The signal of naturally occurring endogenous dU modification (gray dots) remained basically unchanged in the diagonal position, while the signal intensity of on-target sites (red dots) and Cas-dependent off-target sites (orange dots) varied with the removed components.
[0195] Figure 8 The figure shows a comparison between the signal intensity detected by the method of Embodiment 1 of the present invention and the results of site-directed deep sequencing. ρ is the Spearman correlation coefficient. Note: The figures show validation data for Cas-dependent off-target sites.
[0196] Figure 9Two examples of Cas-dependent off-target detection by the method of the present invention are shown, verified by site-directed deep sequencing. (a) Actual editing efficiency at the off-target site “VEGFA_site_2pRBS-237” in different samples; (b) Actual editing efficiency at the off-target site “VEGFA_site_2pRBS-67” in different samples.
[0197] Figure 10 The distribution of the “EMX1,” “VEGFA_site_2,” and “HEK293 site_4” sgRNA targeted editing sites and Cas-dependent off-target editing sites on various chromosomes, detected at the whole-genome level using the method of this invention, is shown. Targeted editing sites and Cas-dependent off-target editing sites are indicated by red squares and blue circles, respectively.
[0198] Figure 11 The Venn diagram shows a comparison of the method of Embodiment 1 of the present invention with Cas-dependent off-target sites detected by GUIDE-seq(a) and Digenome-seq(b).
[0199] Figure 12 The results of the re-evaluation of the specificity of the CBE optimization tool YE1-BE4max using the method of this invention are shown. (a) Comparison of detection signals of all Cas-dependent off-target sites at the genome-wide level in the YE1-BE4max (vertical axis) and WT-BE4max (horizontal axis) samples; (b) Editing efficiency of YE1-BE4max and WT-BE4max at different sites. Red triangles indicate locations with a large number of remaining off-target edits.
[0200] Figure 13 This image shows the Cas-dependent off-target effects at the LbCpf1-BE level on the “RUNX1” and “DYRK1A” sites detected by the method of Example 1 of this invention at the whole-genome level. The horizontal and vertical axes represent the signal intensities identified by this invention in two biological replicate samples.
[0201] Figure 14 Examples of TALE-dependent off-target effects (a) and non-TALE-dependent off-target effects (b) caused by the CRISPR-free DdCBE tool, detected using the method of Example 1 of this invention, are shown. The top image is an enlarged IGV (Integrative Genomics Viewer) image, with red blocks representing "C-to-T" mutations and green blocks representing "G-to-A" mutations, equivalent to "C-to-T" mutations on the complementary strand; the middle image shows mCherry, a negative control sample; the bottom image shows sequencing results verifying the off-target sites detected by the method of this invention using site-directed deep sequencing.
[0202] Figure 15 An exemplary embodiment 2 is shown, which uses the method of the present invention to detect the editing site of a base editor, wherein the base editor is an adenine base editor.
[0203] First, the first step involves extracting nucleic acid (e.g., genomic DNA) edited by an adenine base editor, containing a base editing intermediate (e.g., DNA containing hypoxanthine). This intermediate is a product of the adenine base editor editing the target nucleic acid and comprises a first nucleic acid strand and a second nucleic acid strand. The first nucleic acid strand contains edited bases (e.g., hypoxanthine) generated by the adenine base editor editing the target nucleic acid. The nucleic acid is then broken down using methods such as sonication to form, for example, nucleic acid fragments of approximately 300 bp. The broken genomic DNA fragments are then trimmed to blunt ends using an end-repair process. In some exemplary embodiments, the end-repair process includes the removal of 3' overhangs and the filling of 5' overhangs. In some preferred embodiments, the end-repair process can be performed using a nucleic acid polymerase containing 3' to 5' exonuclease activity.
[0204] The second step involves incorporating a nucleotide (e.g., uracil deoxyribonucleotide) labeled with a first labeling molecule (e.g., biotin) downstream of the position of the edited base (e.g., hypoxanthine) in the base editing intermediate using an in vitro labeling method. In some exemplary embodiments, the labeling experiment includes: specifically recognizing hypoxanthine in the base editing intermediate using the endonuclease Endo V and cleaving the second phosphodiester bond at the 3' end of the hypoxanthine deoxyribonucleotide to form a single-stranded gap; performing a DNA strand replacement reaction from the generated single-stranded gap along the 5' to 3' direction using a DNA polymerase containing strand replacement activity; and ligating the single-stranded cut in the DNA strand replacement reaction product using a DNA ligase. In the DNA strand replacement reaction system, at least one nucleotide substrate (e.g., biotin-uracil deoxyribonucleotide) labeled with a first labeling molecule (e.g., biotin) is used instead of a conventional nucleotide acid substrate (e.g., thymine deoxyribonucleotide). The incorporation of a nucleotide labeled with a first marker molecule (e.g., biotin-uracil deoxyribonucleotide) allows for the subsequent enrichment of DNA fragments containing the first marker molecule using the first binding molecule (e.g., streptavidin). The edited base (e.g., hypoxanthine) contained in the base editing intermediate pairs complementaryly with cytosine during subsequent DNA replication and sequencing. Consequently, an A-to-G mutation signal is generated at the hypoxanthine position in the sequencing results of the labeled product. Therefore, by detecting the presence of the mutation signal, the precise location of the edited base (e.g., hypoxanthine) can be achieved.
[0205] In some preferred embodiments, to avoid false positive signals that may arise from DNA damage (e.g., SSB) introduced during endogenous or nucleic acid manipulation, the method further includes nucleic acid repair treatment of the edited product before proceeding to the second step. In some exemplary embodiments, the treatment includes: performing a DNA strand displacement reaction with a DNA polymerase starting at the SSB gap along the 5' to 3' direction; and ligating the gap in the strand displacement reaction product with a DNA ligase. In some preferred embodiments, the DNA polymerase has strand displacement activity.
[0206] The third step involves enriching DNA fragments containing a first labeled molecule (e.g., biotin) using a solid support (e.g., magnetic beads) coupled with a first binding molecule (e.g., streptavidin); these fragments, optionally amplified and / or used for library construction, can then be used for high-throughput sequencing. Based on the sequencing results, the location of the editing site in the base-editing intermediates (e.g., DNA containing hypoxanthine) generated after the adenine base editor edits the target nucleic acid can be analyzed.
[0207] In some preferred embodiments, prior to amplification and / or library construction of the enriched DNA fragments, the DNA fragments enriched on a solid support (e.g., magnetic beads) may be treated (e.g., alkali treatment) to remove the complementary strand of the nucleic acid single strand containing the first marker molecule (e.g., biotin).
[0208] In some exemplary embodiments, oligonucleotide adapters are ligated to the ends of enriched DNA fragments via an adapter ligation reaction prior to treatment with an alkali (e.g., NaOH) to remove the complementary strand of the nucleic acid single strand containing a first labeling molecule (e.g., biotin), to facilitate amplification or sequencing of the DNA fragments. In some preferred embodiments, a dA tail is added to the 3' end of the DNA fragment, which can be used for ligation to oligonucleotide adapters containing dT tails.
[0209] Figure 16 The enrichment results of the method of Embodiment 2 of the present invention on different pattern sequences are shown.
[0210] Figure 17 This shows the high-throughput sequencing results of ABE at the target site of HEK293_site_4sgRNA (HEK4) for each sample group. The shaded areas indicate the sequence positions of the on-target sites, where "G" represents the A-to-G mutation signal.
[0211] Figure 18 The high-throughput sequencing results of ABE at an off-target site (4) in HEK4 are shown for each sample group. The shaded areas indicate the sequence locations where sgRNA may bind, with "G" indicating an A-to-G mutation signal.
[0212] Figure 19 This shows the site-specific depth sequencing validation results of ABE at the off-target site of HEK4. The first two rows of sequences are the on-target sequence and the off-target sequence, respectively; the last six rows represent the proportion of A, G, C, T bases and the percentage of insertions and deletions.
[0213] Figure 20 High-throughput sequencing results of HEK4 sgRNA at the targeted editing sites in the ABE, ABE8e, and ACBE systems are shown. Orange G represents A-to-G mutation signals; red T represents C-to-T mutation signals.
[0214] Figure 21 High-throughput sequencing results of HEK4 sgRNA at off-target sites (4) in the ABE, ABE8e, and ACBE systems are shown. Orange G represents A-to-G mutation signals; red T represents C-to-T mutation signals.
[0215] Figure 22 The results show high-throughput sequencing of the ABE, ABE8e, and ACBE systems at ABE8e-only off-target sites. The blue C represents the T-to-C mutation signal, which is also the A-to-G mutation signal on its complementary strand.
[0216] Figure 23 The diagram shows the characterization results of the spike-in sequence after replacing the malononitrile labeling step in this invention with other 5fC labeling methods (pyridineborane labeling reaction or 2-methylpyridineborane labeling reaction). (The diagram shows the characterization results of the spike-in sequence after replacing the malononitrile labeling step in this invention with other 5fC labeling methods (pyridineborane labeling reaction or 2-methylpyridineborane labeling reaction). Figure 23 a) qPCR enrichment results of different pattern sequences (AP:dA, dU:dA, or dU:dG) after replacing the chemical labeling method with pyridineborane or 2-methylpyridineborane; Figure 23 (b) Sanger sequencing results of the dU:dG base pair pattern sequence after the replacement with chemical labeling methods such as pyridineborane (pyridineborane or 2-methylpyridineborane). The red arrows indicate the C-to-T mutation signal induced by the chemical labeling.
[0217] Figure 24 The results show the qPCR enrichment of different pattern sequences (Nick, AP:dA, dU:dA, or dU:dG) after replacing Biotin-dU with Biotin-dG.
[0218] Sequence information
[0219] Information about the sequence involved in this invention is provided in Table 1 below.
[0220] Table 1
[0221]
[0222]
[0223]
[0224] Note: The symbol “^” represents the Nick site; N = A, T, G, or C; the symbol “P” represents phosphorylation modification; “AMN” represents C7 Aminolinker blocking. Detailed Implementation
[0225] The invention will now be described with reference to the following embodiments, which are intended to illustrate the invention (and not limit it).
[0226] Unless otherwise specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the embodiments are described by way of example and are not intended to limit the scope of protection claimed in this application.
[0227] Example 1: Detection of CBE Editing Sites
[0228] Experimental methods:
[0229] 1. DNA fragmentation
[0230] Genomic DNA was extracted from live cells transfected with the CBE system, specifically HEK293T (ATCC, catalog number: CRL-11268) or MCF7 (ATCC, catalog number: HTB-22). For the CBE system transfection method, please refer to (Xiao Wang, et al. Nature biotechnology 36, 946-949, doi:10.1038 / nbt.4198(2018)). For the genomic DNA extraction method, please refer to the kit instructions (Kangwei Century, catalog number: CW2298M).
[0231] The extracted genomic DNA was broken down into fragments of approximately 300 bp in length using a Covaris ME220 ultrasonic disruptor, and then recovered using a DNA Clean & Concentrator-5 Kit (purchased from VISTECH, catalog number: DC2005).
[0232] 2. DNA fragment end repair
[0233] The fragmented DNA from step 1 above will have some nicks and overhangs. If these are not repaired, they will be labeled with biotin in subsequent labeling reactions, resulting in false positives. Therefore, this step uses the NEB end repair module (catalog number: E6050) and E. coli DNA ligase (purchased from NEB, catalog number: M0205) to repair the genomic DNA damage that may have been caused during the fragmentation process.
[0234] Prepare the reaction system according to Table 2:
[0235] Table 2: End-of-phase remediation reaction system
[0236]
[0237] After mixing the above reaction system on ice, it was reacted at 20°C for 30 min, and then recovered with 2.0×AMPure XP beads (purchased from Beckman Coulter, catalog number: NC9933872) and eluted with ddH2O.
[0238] 3. EtONH2 protection
[0239] The end-repaired DNA fragment prepared in step 2 was incubated in 80 μL of 100 mM MES buffer (pH 5.0) containing 10 mM EtONH2 at 37°C for 6 h. This protected the naturally occurring d5fC modification in cells, preventing false positives from the subsequent reaction with malononitrile. The DNA was then recovered using the DNA Clean & Concentrator-5 Kit.
[0240] 4.Add dA tail
[0241] Add a dA to the 3' end of each DNA fragment obtained in step 3 to facilitate subsequent ligation of sequencing adapters using the A / T complementation rule.
[0242] Prepare the reaction system according to Table 3:
[0243] Table 3: Reaction systems with added dA tails
[0244]
[0245] The above reaction system was mixed on ice and reacted at 37°C for 30 min. It was then recovered with 2.0×AMPure XP beads and eluted with ddH2O.
[0246] 5. DNA damage repair
[0247] The purpose of this step is to repair and remove DNA modifications or damage that may produce false positive signals, such as naturally occurring AP sites, SSB, and Nick, before dU labeling.
[0248] Prepare the reaction system according to Table 4:
[0249] Table 4: Damage Repair Response System
[0250] Components Total system (50 μL) DNA prepared in step 4 38μL (~2.7ug) NEBuffer 3.0 (purchased from NEB, part number: B7003S) 5μL <![CDATA[50 mM of NAD + > 1μL 2.5mM dNTPs 1μL Endo IV (purchased from NEB, item number: M0304) 2μL Bst full-length polymerase (purchased from NEB, item number: M0328) 1μL Taq DNA ligase (purchased from NEB, product number: M0208) 2μL
[0251] After mixing the above reaction system, the reaction was first carried out at 37°C for 60 min, and then at 45°C for 60 min. The mixture was recovered using 2.0×AMPureXP beads and eluted with ddH2O.
[0252] 6. In vitro BER marker assay
[0253] 0.5 μL of the DNA obtained in step 5 above was added to 0.5 μL of ddH2O as input, and the remaining sample was labeled according to the following steps.
[0254] Prepare the reaction system according to Table 5:
[0255] Table 5: In vitro labeling reaction system
[0256] Components Total system (50 μL) DNA prepared in step 5 37μL (~2.5ug) NEBuffer 3.0 5μL <![CDATA[50 mM of NAD + > 1μL 5μM dATP / dGTP / Biotin-dUTP / 20μM d5fCTP 2μL UDG (purchased from NEB, item number: M0280) 1μL Endo IV 1.5μL Bst full-length polymerase 0.8μL Taq DNA ligase 1.7μL
[0257] After mixing the above reaction system, it was placed at 37℃ for 40 min and recovered with 2.0×AMPure XP beads and eluted with ddH2O.
[0258] 7. Malononitrile reaction
[0259] The DNA recovered in step 6 was placed in 50 mM Tris-HCl (pH 7.0) containing 75 mM Malononitrile and reacted in a mixer at 37°C and 800 rpm for 20 h. It was then recovered again using 2×AMPure XP beads and eluted with ddH2O.
[0260] 8. Fragment enrichment
[0261] Each pull-down sample corresponded to 10 μL of Streptavidin C1 beads (purchased from Invitrogen, catalog number: 65002). Sufficient beads were washed three times with 1×B&W buffer (5 mM Tris-HCl (pH 7.5), 1 M NaCl, 0.5 mM EDTA, 0.05% Tween-20), then resuspended in 40 μL of 2×B&W buffer. An equal volume of the sample DNA treated in step 7 above was added, and the mixture was incubated at room temperature for 1 hour with rotation. The beads were then washed three times with 1×B&W buffer, followed by one wash with 10 mM Tris-HCl (pH 8.0), each time with rotation at room temperature for 5 minutes. Finally, the Tris-HCl solution was aspirated from the magnetic rack, and the remaining magnetic beads (approximately 1 μL) containing the DNA fragment were used for the adapter ligation reaction.
[0262] 9. Connecting connectors
[0263] 1) Dilute the adapter stock solution (30 μM) to 1.5 μM with 10 mM Tris-HCl on ice. The Y-type adapter used was obtained by annealing two single-stranded sequences, wherein the 5' end of the forward single strand is phosphorylated and the 3' end is blocked by a C7 Aminolinker, the sequence of which is shown in SEQ ID NO:7, and the reverse single-stranded sequence is shown in SEQ ID NO:8.
[0264] 2) Use The Quick Ligation Module (purchased from NEB, item number: E6056) performs a connector connection reaction on the Input sample (aqueous solution) retained in step 6 and the PD sample (connected to the magnetic bead) obtained in step 8 above.
[0265] Prepare the reaction system according to Table 6:
[0266] Table 6: Joint Connection Reaction System
[0267] Components Total system (25 μL) <![CDATA[ddH2O]]> 14μL NEB Quick Ligation Buffer 5μL 1.5μM Y-type adapter 2.5μL Quick T4 DNA Ligase 2.5μL PD or Input Sample DNA 1μL
[0268] For the adapter ligation reaction of PD samples: After mixing the above reaction system, place it at about 20°C for 1 hour by rotation (to avoid magnetic bead sedimentation), then add 50 μl of 1×B&W buffer and continue to rotate and incubate at room temperature for 1 hour (to allow the small amount of DNA fragments that detached during the ligation process to rebind to the magnetic beads), and then proceed to the next step of the reaction.
[0269] For the adapter ligation reaction of the input sample: after mixing the above reaction system, place it in a PCR instrument at 20℃ for 40 min, and use 1×AMPure XP beads for recovery and retention to remove unligated adapters.
[0270] 10. NaOH treatment
[0271] The PD samples on the magnetic beads obtained in step 9 above were washed three times with 1×B&W buffer, and then once with 1×SSC buffer. Each time, the magnetic beads were gently inverted to swirl and then rotated at room temperature for 5 min. The supernatant was then removed, and the remaining magnetic beads were resuspended in 20 μl of 0.15M NaOH solution and incubated at room temperature for 10 min. Then, the beads were washed once each with 1×SSC buffer and 10 mM Tris-HCl (pH 8.0). Finally, the magnetic beads were treated with ddH2O at 95°C for 3 min to elute the DNA library from the magnetic beads for the next step of PCR amplification.
[0272] 11. Library Augmentation
[0273] 1) Because the amplification process of high-fidelity DNA polymerase is easily interrupted by Biotin-dU and d5fC labeled with malononitrile, the library was first amplified using MightyAmp DNA Polymerase (purchased from TaKaRa, catalog number: R076A), which has slightly lower fidelity.
[0274] Prepare the reaction system according to Table 7:
[0275] Table 7: MightyAmp amplification system
[0276]
[0277]
[0278] After mixing the above reaction mixture, PCR was performed. The program was as follows: 98℃ for 30s; 98℃ for 10s, 65℃ for 90s (2 cycles); 72℃ for 5min. The DNA after reaction was recovered using a DNA Clean & Concentrator-5 Kit (VISTECH).
[0279] 2) Use high-fidelity DNA polymerase for subsequent amplification to ensure a low overall sequencing noise background.
[0280] Prepare the reaction system according to Table 8:
[0281] Table 8: High-fidelity amplification system
[0282]
[0283] After mixing the above reaction system, perform the PCR reaction. The program is as follows: 98℃ for 30s; 98℃ for 10s, 65℃ for 90s (8-9 cycles for PD samples; 6-7 cycles for Input samples); 72℃ for 5min. Recover the PCR products with 0.9×AMPure XP beads and elute with ddH2O.
[0284] 12. Document Quality Inspection
[0285] Library concentration was determined using a Qubit2.0 precision spectrophotometer;
[0286] The distribution of library fragments was examined using the Fragment Analyzer 12 fully automated capillary electrophoresis system;
[0287] The pattern sequence was relatively quantified and enrichment fold was calculated using qPCR. Primers used for qPCR are shown in SEQ ID NOs:11-22. Data processing employed a 2... -△△Ct The enrichment factor is the fold change in the relative amount of spike-in DNA molecules containing a specific type of modification in the PD sample (with the control pattern sequence as a reference) compared to the corresponding input sample. Based on this fold change, the enrichment status of this batch of experiments can be evaluated.
[0288] The pattern sequence was amplified by full-length PCR, and the resulting PCR product was sequenced by Sanger sequencing. The labeling status of this batch of experiments can be evaluated by the sequencing results.
[0289] Finally, the obtained library was delivered to the Illumina Hiseq X-ten platform for paired-end sequencing (150bp reads).
[0290] Sequencing data processing and analysis:
[0291] 1. Replying and filtering of data in this invention
[0292] After the data is processed, the sequencing reads in the FASTQ file of the sequencing results are first removed using the cutadapt (version 1.18) software. The specific command parameters are: cutadapt --times 1-e 0.1-O3 --quality-cutoff 25-m 50. After removing the adapters, considering that the sequencing results of this invention may contain C to T mutations, the sequencing reads with the adapters removed are first pasted back to the reference genome (version hg38) using the Bismark (version 0.22.3) software. Sequencing reads that fail to align or have an alignment quality MAQP lower than 20 are extracted again and then re-aligned using BWA MEM (version 0.7.17). Finally, the sequencing data after the two alignments are merged are screened again. Only alignment results with an alignment quality MAQ greater than 20, i.e., an alignment error rate of less than 1%, are retained for downstream analysis. Next, the high-quality alignment results are deduplicated using the PicardMarkDuplicates command (version 1.9). This step primarily aims to remove molecular redundancy generated during library construction due to amplification. After these steps, the genome re-alignment results (BAM format file) suitable for downstream analysis are obtained.
[0293] 2. Preliminary identification of the signal of this invention
[0294] The BAM file was converted to an mpileup file using the `samtools mpileup-q 20-Q 20` command (version 1.9). Then, the `parse-mpileup` and `bmat2pmat` commands from a self-developed software tool (see, for example, https: / / github.com / menghaowei / Detect-seq) were used to generate a pmat file. Next, the `pmat-merge` command was used to scan and organize all tandem C-to-T mutation signals across the entire genome and record them as mpmat format files. Finally, the `mpmat-select` command was used for selection to obtain the preliminary sequencing signals for this invention.
[0295] 3. Identification of the enriched signal in this invention
[0296] After obtaining the initial sequencing signals for this invention, enrichment detection of these candidate regions is required. First, the `find-significant-mpmat` command in the software tool is used to perform statistical tests on the candidate regions. The results of the statistical tests are then corrected using the BH method to obtain the false discovery rate (FDR). Regions with an FDR less than 0.01, a normalized enrichment fold greater than 2 compared to the control group, fewer than 3 reads with mutation signals in the control group samples, and at least 5 sequencing reads with mutation signals in the treatment group samples are considered the final identified regions for this invention.
[0297] 4. Removal of endogenous deoxyuracil sites
[0298] In the enrichment detection, the experimental group and the control group were respectively set as samples transfected only with empty vector plasmid and subjected to the enrichment library preparation procedure described in this method, and samples not treated with the enrichment library preparation procedure described in this method. This allows for the acquisition of endogenous deoxyuridine location information. To ensure a lower false negative rate for this identification method, a relatively lenient threshold is used in this step: FDR less than 0.05, and the normalized enrichment fold of the experimental group compared to the control group greater than 1.5.
[0299] 5. Alignment of off-target site gene sequences with sgRNA sequences
[0300] Within the enriched signal regions identified in the above steps after removing endogenous dU, the binding sites of sgRNA / crRNA can be inferred through sequence alignment. These inferred sgRNA / crRNA binding sites are called pRBS (putative sgRNA / crRNA binding sites). A modified semi-global alignment method was used for sequence alignment between sgRNA / crRNA and the enriched signal regions. For sgRNA, the PAM sequence (NAG / NGG) is first searched within the region. Then, for the found PAM sites, a 30-nt sequence in the 5' direction of the PAM is extracted and semi-global double sequence alignment is performed with the sgRNA. The best reported result is the pRBS. For crRNA, the PAM sequence (TTTV, V = A / C / G) is first searched within the region. Then, for the found PAM sites, a 30-nt sequence in the 3' direction of the PAM is extracted and semi-global double sequence alignment is performed with the crRNA. The best reported result is the pRBS of the crRNA. In the above process, if no PAM is found within the region, a semi-global alignment of the sgRNA / crRNA with the sequence in that region is performed directly. The optimal alignment result is the pRBS of the sgRNA / crRNA. The alignment parameters used in this step are: match +5; mismatch -4; open gap -24; gap extension -8. The alignment procedure for this step is included in the mpmat-to-art command in the Detect-seq software toolbox.
[0301] Experimental results:
[0302] 1. Specific labeling and enrichment of pattern sequences containing dU
[0303] To demonstrate the specificity and efficiency of the method of the present invention, Figure 2 The pattern sequences and control sequences (SEQ ID NOs: 1-6) containing different modified bases, as shown in figure a, were incorporated into the fragmented genomic DNA, and library construction was performed according to the experimental method described above. Finally, the changes in the proportions of different pattern sequences in the samples before and after pull-down were calculated and compared using quantitative real-time PCR (all were relative to the control sequence without any modifications (the control pattern sequence shown in SEQ ID NO: 1)). The enrichment folds of different pattern sequences in the samples before and after pull-down were also calculated. The enrichment folds are shown below. Figure 2As shown in Figure b, for pattern sequences containing single dU:dA and dU:dG base pairs, the method provided by this invention can enrich them by approximately 60-fold and 30-fold, respectively; while for pattern sequences containing AP sites and d5fC, almost no enrichment is achieved. This demonstrates that the method provided by this invention can specifically enrich DNA fragments containing dU.
[0304] On the other hand, according to the design principle, this invention has a certain probability of continuously incorporating multiple d5fCTPs at the 3' end of the dU position, thereby generating continuous C-to-T mutations, thus achieving signal amplification and facilitating detection. This is based on the results of Sanger sequencing and high-throughput sequencing (…). Figure 3 We did indeed observe continuous C-to-T mutation signals in the pattern sequences containing dU, indicating that the strategy of introducing C-to-T mutation signals through chemical reactions in the process of this invention can indeed achieve the labeling of dU positions.
[0305] In summary, by capturing this highly distinctive C-to-T mutation signal, very sensitive and accurate dU detection can be achieved.
[0306] 2. Specific detection signals are generated at the CBE editing site.
[0307] In human HEK293T and MCF7 cell lines, several representative sgRNAs were selected to test the off-target effects of the method provided in this invention on the highly efficient CBE tool BE4max. The method for transfecting cells with the CBE4max editing system is described in (Xiao Wang, et al. Nature biotechnology 36, 946-949, doi:10.1038 / nbt.4198(2018)). The representative sgRNAs were “VEGFA_site_2” (SEQ ID NO:23) and “HEK293 site_4” (SEQ ID NO:24), which are known to have very low specificity in vivo; “EMX1” (SEQ ID NO:25), which has moderate specificity; “RNF2” (SEQ ID NO:26), which has not been reported to have off-target sites; and “RUNX1” (SEQ ID NO:27), which has been studied less previously.
[0308] Test results as follows Figure 4 As shown, by Figure 4As can be seen from this, the method of the present invention creates a very obvious read enrichment peak at the corresponding on-target editing site. Upon further amplification, a distinct and characteristic continuous C-to-T mutation signal can be observed. Furthermore, these enriched mutation signals were not observed in the negative control NT samples (i.e., samples transfected with BE4max and non-target sgRNA), indicating that the present invention has very good detection specificity. By comparing with the on-target editing results of these sgRNAs in previous studies, we found that the C with the strongest C-to-T mutation signal is usually the cytosine position with the highest actual editing efficiency. Moreover, it is likely because the polymerase nick translation reaction in the present invention can incorporate multiple d5fCTPs at once, so even if only one or two Cs are edited, a distinct continuous C-to-T mutation signal will be generated. Figure 4 As can be seen from b, generally 2-6 consecutive C-to-T mutations will be generated in the 4-9bp region after the edited C.
[0309] In addition, with Figure 4 Taking c as an example, it is clear that the characteristic signal of continuous C-to-T mutations generated by this invention can be easily distinguished from SNVs. Furthermore, at the whole-genome level, the signal generated by the method of this invention, with the same amount of data, is far stronger than that of conventional WGS sequencing, and is much easier to distinguish from sequencing background errors, with lower requirements for sequencing coverage. Figure 4 d).
[0310] In summary, the above observations demonstrate that the signal features generated by the method of the present invention can greatly enhance the detection signal at the editing site, thereby significantly improving the detection sensitivity of the present invention and reducing the detection cost.
[0311] 3. Assessment of Cas-dependent and non-Cas-dependent off-target effects caused by CBE
[0312] By conducting deletion comparison experiments on different components of the CBE system, the off-target site properties detected by this invention at the whole-genome level and their possible generation mechanisms can be verified. Specifically, during cell transfection, we removed the APOBEC1, UGI, and sgRNA portions of the BE4max system, respectively. After removal, the plasmid structures are as follows: Figure 5 As shown, Vector samples transfected only with mCherry plasmid were used as negative control samples, and the genomic DNA of these samples after transfection was detected using the method of the present invention.
[0313] Detection results of non-Cas-dependent off-target effects are as follows Figure 6As shown, it exhibits three distinct characteristics: 1) The gene location of the signal has almost no similarity to the sgRNA sequence ( Figure 6 a) ; 2) The signal strength is usually very low, mostly just above the background level ( Figure 6 a) ; 3) More likely to appear in transcriptionally active regions ( Figure 6 e). These characteristics are consistent with previously reported non-Cas-dependent off-target behaviors. More importantly, further analysis of these off-target sites revealed that: when all components of the CBE system were present, a large number of such off-target sites were found, exhibiting a very distinct "TC" motif; even after removing the sgRNA component, the number of such sites remained high, and the motif persisted; however, after deleting the APOBEC1 component, the number of such sites dropped to background levels, and the motif disappeared. Figure 6 (bd). APOBEC1 is known to have a natural substrate-binding preference for the "TC" motif. These experimental data and characteristics suggest that such off-target sites are not dependent on the Cas system, but only on APOBEC1, and should be off-target edits randomly generated by APOBEC1 overexpression.
[0314] Detection results of Cas-dependent off-target effects are as follows Figure 7 As shown, it exhibits the following characteristics: 1) The signal intensity is much stronger than that of non-Cas-dependent off-target events. At some sites, signal intensity comparable to that at the on-target site can even be observed. Figure 7 a) This indicates that the editing efficiency of such off-target sites would be much higher; 2) The signal is repeatedly and stably generated in biological repeat groups ( Figure 7 b); 3) Gene sequences with some similarity to sgRNA can usually be found in the genomic region where the signal is located. Component deletion comparison experiments showed that: compared to the All samples with all elements present, the signal intensity of such sites in the (-)sgRNA and (-)APOBEC samples all decreased below the background level, while the signal intensity in the (-)UGI samples varied in degree; while the signal intensity of dU modification sites present in the cell was almost completely unaffected by component deletion. Figure 7 c). These experimental data indicate that these off-target sites are generated simultaneously by sgRNA and APOBEC, and should indeed be classic Cas-dependent off-target sites. Furthermore, the number of Cas-dependent off-target sites identified in this invention varies depending on the specificity of the sgRNA: for example, under the same bioinformatics analysis identification rule (cufoff), for the known poorly specific "VEGFA_site_2", this invention identified a total of 511 such off-target sites ( Figure 7b); however, for “RNF2”, which is known to have excellent specificity, this invention did not detect such off-target sites.
[0315] 4. Validation results of off-target sites
[0316] To verify the authenticity of the detection results of the method of this invention, targeted deep sequencing technology was used to measure the actual editing efficiency at the off-target sites identified by this invention. Targeted deep sequencing technology involves performing targeted PCR amplification on the target site and then performing high-throughput sequencing on the PCR products. This allows the sequencing depth of at least tens of thousands of reads to be covered at the tested genomic site, thus obtaining a very precise editing efficiency for this site.
[0317] The results of site-specific deep sequencing verification of the sites detected by the method of this invention are as follows: Figure 8 As shown in the figure, among the 151 randomly selected sites of the present invention with signal intensities ranging from low to high, 50 / 50 "EMX1" sites, 51 / 51 "VEGFA_site_2" sites, 43 / 43 "HEK293site_4" sites, and 7 / 7 "RUNX1" sites were all successfully verified by deep sequencing, achieving a true positive rate of nearly 100%. Furthermore, even when the actual editing efficiency was relatively low, the corresponding signal intensity of the present invention was already very high, further demonstrating that the present invention indeed possesses very high detection sensitivity.
[0318] Furthermore, site-specific deep sequencing verified that the Cas-dependent off-target effects identified by the method of this invention (selecting more than 20 sites) are indeed dependent on sgRNA production. Figure 9 The image shows the deep sequencing signals at two sites in the sample groups with and without sgRNA. Figure 9 The results show that the two off-target sites are indeed dependent on sgRNA production. In summary, the above data proves the high reliability of the method of the present invention.
[0319] Figure 10 The distribution of “EMX1”, “VEGFA_site_2” and “HEK293 site_4” sgRNA targeted editing sites and Cas-dependent off-target editing sites on each chromosome is shown at the whole-genome level using the method of the present invention.
[0320] 5. Comparison of detection results between the method of this invention (Detect-seq) and other related methods
[0321] GUIDE-seq is a well-known off-target detection technique in the field of gene editing, primarily used to detect Cas-dependent off-target effects caused by the CRISPR / Cas9 nuclease system. Given that CBE tools are also constructed based on inactivated or partially inactivated Cas9 proteins, some researchers directly assess the off-target effects of the CBE system using sites identified by GUIDE-seq. However, in reality, even using the same sgRNA, genome-wide off-target effects caused by the CBE system are quite different from those caused by Cas9 nucleases (Kim, D. et al. Nature biotechnology 35, 475-480, doi:10.1038 / nbt.3852(2017).).
[0322] A comparison of the method of this invention with the detection results of GUIDE-seq is as follows: Figure 11 As shown in Figure a, for “VEGFA_site_2” and “EMX1”, the method of this invention detected most of the Cas-dependent off-target sites in the GUIDE-seq results; for “HEK293 site_4”, the method of this invention detected about half of the sites in the GUIDE-seq results; the method of this invention discovered a large number of off-target sites that were not reported by GUIDE-seq. The results of site-specific deep sequencing validation at randomly selected sites showed that, compared to GUIDE-seq, the 41 new off-target sites detected by the method of this invention were indeed real off-target sites, while the 15 / 17 sites not reported by the method of this invention but reported by GUIDE-seq did not experience CBE editing events in live cells; and the 37 off-target sites identified by both methods were successfully validated.
[0323] A comparison of the method of this invention with the Digenome-seq detection results of the CBE system developed by Kim et al. is as follows: Figure 11 As shown in b, Digenome-seq is essentially an in vitro off-target detection technique based on WGS. Similar to the results compared with conventional WGS, the signal values exhibited by this invention at off-target sites are significantly higher than those of Digenome-seq under the same sequencing volume. The method of this invention detected most of the Cas-dependent off-target sites reported by Digenome-seq, but discovered a much larger number of new off-target sites (…). Figure 11 b). The results of site-specific deep sequencing verification at random sites showed that 10 / 15 sites not reported in this invention but reported by Digenome-seq did not experience CBE editing events in living cells; all 18 off-target sites identified by both methods were successfully verified.
[0324] The above results also indicate that the true positive rate reported by this invention is close to 100%, while the true negative rate is approximately 80%. It is worth mentioning that, upon further careful examination of the detection results of the method of this invention, detection signals of varying degrees can actually be observed at the seven real off-target sites that were not successfully reported, but these may not have been reported because they did not reach the bioinformatics analysis cutoff.
[0325] 6. Assessment of off-target effects of the optimized CBE tool
[0326] Recently, many improved CBE tools have been reported in this field that have performed well in reducing DNA or RNA off-target effects. Among them, YE1-BE4max has been reported by several independent studies as the best CBE version overall (Doman, et al. Nature biotechnology 38, 620-628, doi:10.1038 / s41587-020-0414-6(2020); Zuo, E. et al. NatMethods 17, 600-604, doi:10.1038 / s41592-020-0832-x(2020)).
[0327] The method of this invention can detect that YE1-BE4max does indeed reduce most of the off-target signal levels caused by WT-BE4max. However, taking "EMX1" sgRNA as an example, among the 48 Cas-dependent off-target sites identified from the WT-BE4max sample, there are still 4, 3, and a dozen or so sites that retain high, medium, and low intensity detection signals in YE1-BE4max. Figure 12 a).
[0328] Validation results from site-specific deep sequencing show that, with similar on-target site editing efficiencies, YE1-BE4max indeed produces virtually no editing results at sites reported negative in this invention (e.g., the "EMX1 pRBS_1" site). However, at the three strong signal sites identified in this invention ("EMX1 pRBS_4", "EMX1 pRBS_3", and "EMX1 pRBS_2"), YE1-BE4max still exhibits a very high off-target editing ratio (up to nearly half the on-target editing efficiency), and at one of these sites (the "EMX1 pRBS_2" site), it shows no reduction compared to WT-BE4max. Therefore, using this invention to evaluate the overall off-target effect of the new optimized tool has high reliability. Similarly, other optimized versions of CBE tools (such as the CBE system built using APOBEC3A) can also be comprehensively evaluated for off-target effects using this invention.
[0329] Furthermore, these data also illustrate that previous methods of assessing off-target effects of CBE tools by randomly selecting only a subset of sites identified by GUIDE-seq were insufficient, and the conclusions drawn could vary depending on the selected sites. This invention provides an assessment platform based on a comprehensive, genome-wide assessment, offering a basis for the optimization and comparison of CBE tools.
[0330] 7. Detection of off-target effects of CBE tools built on other CRISPR systems
[0331] Given the same APOBEC deamination editing principle, CBE tools built on other CRISPR systems, such as Cpf1(Cas12a)-BE, can also use the method of this invention for off-target evaluation. Figure 13 This study demonstrates 949 and 240 Cas-dependent off-target sites induced by LbCpf1-BE at the whole-genome level in “RUNX1” (SEQ ID NO:37) and “DYRK1A” (SEQ ID NO:38) crRNA using the method of this invention. Similarly, site-directed deep sequencing confirmed that 18 / 18 of these were actual off-target editing sites.
[0332] 8. Off-target detection of the CRISPR-free DdCBE tool
[0333] HEK293T cells were transfected with the DdCBE system, which targets different DNA sites in mitochondria. The transfection method is described in (Mok, BY et al. Nature 583,631-+, doi:10.1038 / s41586-020-2477-4(2020)). Three days later, the genome was extracted to detect the editing efficiency at the mitochondrial target sites. Sanger sequencing results showed that the editing efficiency was between 35% and 55%. Since the deaminase DddA in the DdCBE system converts dC on double-stranded DNA into dU, the method of this invention can also be used to detect the intermediate product dU, thereby assessing the off-target effects caused by DdCBE.
[0334] Although DdCBE is a mitochondrial DNA cytosine editing tool, Detect-seq results show that each DdCBE generates hundreds of off-target edits in the cell nucleus. Based on the characteristics and causes of off-target signals, they can be divided into two main categories: TALE-dependent off-targets and non-TALE-dependent off-targets. In this invention, 36 off-target sites were randomly selected for verification. Site-specific deep sequencing results confirmed that all 36 sites indeed exhibited a certain off-target editing rate, with some sites showing an off-target efficiency as high as 8%, indicating that Detect-seq can indeed be used to detect off-target effects caused by DdCBE. Figure 14The illustrations show sequencing signal maps of TALE-dependent and non-TALE-dependent off-targets detected by the method of the present invention, as well as sequencing results that were validated using site-directed depth sequencing.
[0335] Example 2: ABE Edit Site Detection
[0336] Experimental methods:
[0337] 1. DNA fragmentation
[0338] Genomic DNA was extracted from live HEK293T cells (purchased from ATCC, catalog number: CRL-11268) transfected using the ABE system. The method for transfecting cells using the ABE system is described in (Xiao Wang, et al. Nature biotechnology 36, 946-949, doi:10.1038 / nbt.4198(2018)), and the method for extracting genomic DNA from the cells is described in the kit instructions (purchased from Kangwei Century, catalog number: CW2298M).
[0339] The extracted genomic DNA was broken down into fragments of approximately 300 bp in length using a Covaris ME220 ultrasonic disruptor, and then recovered using a DNA Clean & Concentrator-5 Kit.
[0340] 2. DNA fragment end repair
[0341] This step uses the NEB end repair module and E. coli DNA ligase to smooth out some nicks and overhangs in fragmented DNA, as well as repair genomic DNA damage that may have been caused during the breakup process.
[0342] Prepare the reaction system according to Table 9:
[0343] Table 9: End-of-phase remediation reaction system
[0344]
[0345] After mixing the above reaction system on ice, it was reacted at 20°C for 30 min, then recovered with 2.0×AMPure XP beads and eluted with 40 μL ddH2O.
[0346] 3.Add dA tail
[0347] Add a dA to the 3' end of each of the DNA fragments obtained in step 2 to facilitate subsequent ligation to sequencing adapters using the A / T complementation rule. The experimental procedure is the same as in Example 1.
[0348] 4. DNA damage repair
[0349] Prepare the reaction system according to Table 10:
[0350] Table 10: Damage Repair Reaction System
[0351] Components Total system (50 μL) DNA prepared in step 3 40 μL (~3.3 μg) NEBuffer 3.0 5μL <![CDATA[50 mM of NAD + > 1μL 2.5mM dNTPs 1μL Bst full-length polymerase 1μL Taq DNA ligase 2μL
[0352] After mixing the above reaction system, the reaction was first carried out at 37℃ for 60 min, and then at 45℃ for 60 min. The mixture was recovered using 2.0×AMPureXP beads, eluted with 17 μL of ddH2O, and 1 μL of the sample was used as input for subsequent library construction.
[0353] 5. dI recognition
[0354] The purpose of this step is to break the second phosphodiester bond at the 3' end of dI, thereby creating a nick for subsequent labeling.
[0355] Prepare the reaction system according to Table 11:
[0356] Table 11: Incision Formation Reaction System
[0357] Components Total system (20 μL) DNA prepared in step 4 16 μL (~3 μg) NEBuffer 4 2μL Endonuclease V (purchased from NEB, item number: M0305) 2μL
[0358] After mixing the above reaction system, react at 37°C for 80 min, then purify with twice the volume of XP beads, and finally elute with 43 μL of water.
[0359] 6. Biotin-marker
[0360] The purpose of this step is to add biotin-tagged dUTP at the location where detection is required.
[0361] Prepare the reaction system according to Table 12:
[0362] Table 12: Biotin-labeled reaction system
[0363] Components Total system (50 μL) DNA prepared in step 5 42 μL (~2.7 μg) NEBuffer 3 5μL 100mM dATP 0.5μL 100mM dCTP 0.5μL 100mM dGTP 0.5μL 5μM Biotin-16-AA-2'-dUTP 0.5μL Full length Bst DNA polymerase 1μL
[0364] After mixing the above reaction system thoroughly, react at 37°C for 40 min. After the reaction is complete, add 1 μL of 50 mM NAD to the tube. + Add 2 μL of Taq DNA ligase and incubate in a PCR instrument at 37°C for 40 min. After the reaction, purify with 2×XP beads and finally wash with 41 μL of water.
[0365] 7. Fragment enrichment
[0366] Each pull-down sample corresponds to 10 μL of Streptavidin C1 beads. Sufficient beads were washed three times with 1×B&W buffer (5 mM Tris-HCl (pH 7.5), 1 M NaCl, 0.5 mM EDTA, 0.05% Tween-20), then resuspended in 40 μL of 2×B&W buffer. An equal volume of the sample DNA treated in step 6 above was added, and the mixture was incubated at room temperature for 1 hour. The beads were then washed three times with 1×B&W buffer, followed by one wash with 10 mM Tris-HCl (pH 8.0), each time incubated at room temperature for 5 minutes. Finally, the Tris-HCl solution was aspirated from the magnetic rack, and the remaining beads containing the DNA fragments were used for the adapter ligation reaction.
[0367] 8. Connecting connectors
[0368] 1) Dilute the adapter stock solution (30 μM) to 1.5 μM with 10 mM Tris-HCl on ice. The Y-type adapter used was obtained by annealing two single-stranded sequences, wherein the forward single-stranded sequence has a phosphorylation modification at the 5' end, and its sequence is shown in SEQ ID NO:7, and the reverse single-stranded sequence is shown in SEQ ID NO:8.
[0369] 2) Use The Quick Ligation Module performs a connector bonding reaction on the Input sample (aqueous solution) retained in step 4 and the PD sample (attached to the magnetic bead) obtained in step 7 above.
[0370] Prepare the reaction system according to Table 13:
[0371] Table 13: Joint Connection Reaction System
[0372] Components Total system (25 μL) <![CDATA[ddH2O]]> 14μL NEB Quick Ligation Buffer 5μL 1.5μM Y-type adapter 2.5μL Quick T4 DNA Ligase 2.5μL PD or Input Sample DNA 1μL
[0373] For the adapter ligation reaction of PD samples: After mixing the above reaction system, place it at about 20°C for 1 hour by rotation (to avoid magnetic bead sedimentation), then add 50 μl of 1×B&W buffer and continue to rotate and incubate at room temperature for 1 hour (to allow the small amount of DNA fragments that detached during the ligation process to rebind to the magnetic beads), and then proceed to the next step of the reaction.
[0374] For the adapter ligation reaction of the input sample: after mixing the above reaction system, place it in a PCR instrument at 20°C for 1 hour, and use 1×AMPure XP beads for recovery and retention to remove unligated adapters.
[0375] 9. Cleaning and purification process
[0376] The sample (PD sample) ligated to the beads after step 8 was washed three times with 1 mL 1×BW, then washed once with 200 μL EB (10 mM Tris-HCl), and finally eluted the DNA library in the PD sample with 25 μL ddH2O in a shaker at 95℃ and 1200 rpm.
[0377] 10. Library amplification
[0378] The experimental procedure is the same as in Example 1.
[0379] 11. Document Quality Inspection
[0380] Library concentration was determined using a Qubit2.0 precision spectrophotometer;
[0381] The distribution of library fragments was examined using the Fragment Analyzer 12 fully automated capillary electrophoresis system;
[0382] The pattern sequence was relatively quantified and enrichment folds were calculated using qPCR. Primers used for qPCR are shown in SEQ ID NOs:11-12,31-36. Data processing employed a 2... -△△Ct The enrichment factor is the fold change in the relative amount of spike-in DNA molecules containing a specific type of modification in the PD sample (with the control pattern sequence as a reference) compared to the corresponding input sample. Based on this fold change, the enrichment status of this batch of experiments can be evaluated.
[0383] The pattern sequence was amplified by full-length PCR, and the resulting PCR product was sequenced by Sanger sequencing. The labeling status of this batch of experiments can be evaluated by the sequencing results.
[0384] Finally, the obtained library was delivered to the Illumina Hiseq X-ten platform for paired-end sequencing (150bp reads).
[0385] Sequencing data processing and analysis:
[0386] 1. Replying and filtering of data in this invention
[0387] After the data is processed, the sequencing reads in the FASTQ file of the sequencing results are first removed using the cutadapt (version 1.18) software. The specific command parameters are: cutadapt --times 1-e 0.1-O3 --quality-cutoff 25-m 50. The sequencing reads after adapter removal are then back-aligned to the reference genome (version hg38) using BWA MEM (version 0.7.17). Alignment results with a MAPQ greater than 20 (i.e., an alignment error rate below 1%) are retained for downstream analysis. Subsequently, the Picard MarkDuplicates command (version 1.9) is used to remove duplicates from the selected high-quality alignments. This step primarily aims to remove molecular redundancy generated during library construction due to amplification. After these steps, the genome back-aligned results (BAM format file) suitable for downstream analysis are obtained.
[0388] 2. Preliminary identification of the signal of this invention
[0389] After obtaining the filtered BAM file, the `samtools mpileup-q 20-Q20` command (version 1.9) was first used to convert the BAM file into an mpileup file. Then, the `parse-mpileup` and `bmat2pmat` commands from the aforementioned software tools were used to generate a pmat file. Next, the `pmat-merge` command from the same software tools was used to scan, organize, and record all tandem C-to-T mutation signals across the entire genome into mpmat format files. Finally, the `mpmat-select` command from the same software tools was used for filtering to obtain the preliminary sequencing signals for this invention.
[0390] 3. Identification of the enriched signal in this invention
[0391] After obtaining the initial sequencing signals for this invention, enrichment detection of these candidate regions is required. First, the `find-significant-mpmat` command in the software tool is used to perform statistical tests on the candidate regions. The results of the statistical tests are then corrected using the BH method to obtain the false discovery rate (FDR). Regions with an FDR less than 0.01, a normalized enrichment fold greater than 2 compared to the control group, fewer than 3 reads with mutation signals in the control group samples, and at least 5 sequencing reads with mutation signals in the treatment group samples are considered the final identified regions for this invention.
[0392] 4. Comparison of off-target site gene sequences with sgRNA sequences
[0393] Within the enriched signal regions identified in the above steps, the binding site of sgRNA can be inferred through sequence alignment. The inferred sgRNA binding site is called the pRBS (putative sgRNA binding site). A modified semi-global alignment method is used when aligning the sgRNA with the enriched signal regions. First, a PAM sequence (NAG / NGG) is searched within the enriched regions. Then, for each found PAM site, a 30-nt sequence in the 5' direction of the PAM is extracted and semi-global aligned with the sgRNA. The best result reported during this alignment is the pRBS. If no PAM is found within the region, the sgRNA is directly semi-global aligned with the sequence in that region, and the best result is the pRBS for that sgRNA. The alignment parameters used in this step are: match +5; mismatch -4; open gap -24; gap extension -8. The alignment procedure for this step is included in the mpmat-to-art command within the Detect-seq software toolbox.
[0394] Experimental results:
[0395] 1. Specific labeling and enrichment of pattern sequences containing dI
[0396] To demonstrate the specificity and efficiency of the method of this invention, pattern sequences containing different modified bases and control sequences (SEQ ID NOs: 1, 28-30) were incorporated into the library construction samples. Finally, the changes in the proportions of different pattern sequences in the samples before and after pull-down were calculated and compared using qPCR (all were relative to the control sequence without any modifications (the control pattern sequence shown in SEQ ID NO: 1)), and the enrichment folds of different pattern sequences in the samples before and after pull-down were calculated. The enrichment folds are shown below. Figure 16 As shown in the figure, for pattern sequences containing single dI:dC and dI:dT base pairs, the method of the present invention can enrich them by about 220 times and about 50 times or more, respectively, while pattern sequences containing only Nick are almost not enriched at all. This proves that the method of the present invention can specifically and efficiently enrich DNA fragments containing dI.
[0397] 2. Enrichment of DNA containing ABE actual editing sites
[0398] Genomic DNA was extracted from HEK293T cells transfected with ABEmax. The method for transfecting cells with ABEmax is described in (Xiao Wang, et al. Nature biotechnology 36, 946-949, doi:10.1038 / nbt.4198(2018)). A second-generation sequencing library was constructed using the method of this invention, and then a series of supporting bioinformatics analyses were performed to obtain information on the ABEmax editing sites at the whole genome level. Figure 17 The high-throughput sequencing results of ABE at the HEK293_site_4 (abbreviated as HEK4) (SEQ ID NO:24) target site are shown. As can be seen from the figure, no mutation signal was detected in the negative control vector sample, while the experimental group sample all-PD showed an A-to-G mutation signal, where the mutation location is the editing site; moreover, compared with the vector sample, the number of reads containing mutations in the all-PD sample was significantly increased, which also indicates that enrichment did occur at this point.
[0399] Figure 18 The high-throughput sequencing results for one of the off-target sites are shown in the figure. It can be seen from the figure that there is no mutation signal in the vector sample, while the all-PD sample contains A-to-G mutation information, which is the off-target signal.
[0400] 3. Validation results of off-target sites detected by the method of the present invention
[0401] Figure 19 The figure shows the verification results of one of the off-target sites detected by the method of the present invention through site-specific deep sequencing. As can be seen from the figure, the off-target editing rate of this site is as high as 10.82%. Furthermore, the comparison between the on-target sequence and the off-target sequence in the figure shows that they are very similar, suggesting that the off-target effect here is CAS-dependent.
[0402] 4. Assessment of off-target effects of various ABE systems
[0403] In addition to the ABEmax system, the invention can be used to identify off-target sites for the two novel tools, ABE8e and ACBE, as well as other base editing systems based on adenine deaminase that may be developed in the future.
[0404] Figure 20-22This image shows high-throughput sequencing results at on-target and off-target sites when the method of this invention is applied to the off-target detection of two novel tools, ABE8e (Richter et al., 2020) and ACBE (Grunewald et al., 2020; Li et al., 2020; Sakata et al., 2020; Zhang et al., 2020). For on-target sites, from... Figure 20 It can be observed that all three systems have corresponding A-to-G mutation signals within the sgRNA binding region. Among them, the signal of ABE8e is stronger than that of ABE, and in addition to the A-to-G mutation signal, ACBE also has a C-to-T mutation signal.
[0405] For off-target sites, such as the aforementioned off-target 4 sites, off-target signals were detected in all three systems, only with different signal intensities. Figure 21 In addition to the off-target sites common to all three systems, this invention also detected off-target sites unique to ABE8e. For example... Figure 22 As shown, off-target signals were detected only in samples transfected with the ABE8e system at this location, while no corresponding off-target signals were detected in the other two samples. Previous literature reported that ABE8e has much higher activity than ABE, and the off-target signals detected by this invention are indeed much more numerous than those detected by ABE8e, which to some extent demonstrates the reliability of this invention.
[0406] Example 3
[0407] The inventors of this application replaced step 7 (malononitrile labeling step) of the experimental method in Example 1 with other 5fC labeling methods, which can also induce C to T mutation signals at d5fC without affecting the enrichment results, and can ultimately achieve labeling at the dU position.
[0408] Taking chemical labeling methods such as pyridine borane as an example, the inventors replaced malononitrile in Example 1 with pyridine borane or 2-methylpyridine borane for the reaction (other experimental steps are described in Example 1). The characterization results of the spike in pattern sequence after treatment by the method of this invention are as follows: Figure 23 As shown. Figure 23 The results showed that: 1) Pattern sequences containing single dU:dA (SEQ ID NO:2) and dU:dG (SEQ ID NO:5) base pairs were enriched by approximately 60-fold and 20-fold, respectively, while the pattern sequence containing AP sites (SEQ ID NO:4) showed almost no enrichment. Figure 23a) ; 2) Through Sanger sequencing results, continuous C-to-T mutation signals were observed in the dU-containing pattern sequences. Figure 23 b). The above results demonstrate that the present invention can also introduce continuous C-to-T mutation signals using other similar chemical reactions without affecting the enrichment results, and ultimately achieve dU position labeling. It should be noted that compared to the malononitrile labeling method, the proportion of C-to-T mutation signals generated by the pyridineborane labeling method is lower ( Figure 23 b).
[0409] Example 4
[0410] The Biotin-dU marker molecules in Examples 1 and 2 can also be replaced with other marker molecules with enrichment effects. For example, after the inventors of this application replaced Biotin-dU in Example 1 with Biotin-dG, the pattern sequences containing single dU:dA (SEQ ID NO:3) and dU:dG (SEQ ID NO:5) base pairs were enriched by approximately 30-fold and 20-fold, respectively, while the pattern sequences containing AP sites (SEQ ID NO:4) and Nick (SEQ ID NO:30) showed almost no enrichment. Figure 24 This result demonstrates that, by switching to Biotin-dG, this invention also specifically enriches DNA fragments containing dU.
[0411] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the teachings disclosed, and all such changes are within the scope of protection of the invention. The full scope of the invention is given by the appended claims and any equivalents thereof. SEQUENCE LISTING <110> Beijing University <120> Methods and kits for detecting base editor editing sites <130> IDC220153 <150> CN202110551156.9 <151> 2021-05-20 <160> 38 <170> PatentIn version 3.5 <210> 1 <211> 137 <212> DNA <213> Artificial Sequence <220> <223> Control mode sequence <400> 1 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatgtc agcatgacgt 60 tatctggcgg tggagatggc tccgtgtggc agagctgaaa gaggagcttg atgacacgta 120 atgcttgcgt ggcaaac 137 <210> 2 <211> 137 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-1 pattern sequence <220> <221> misc_feature <222> (48)..(48) <223> n is uracil deoxyribonucleotide. <400> 2 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatntc agcatgacgg 60 cggtaagcac gaactcaggc tccgtgtggc agagctgaaa gaggagcttg atgacacggg 120 aaataccgtg gtgtggc 137 <210> 3 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-2 pattern sequence <220> <221> misc_feature <222> (48)..(48) <223> n is uracil deoxyribonucleotide. <400> 3 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatntc agcatgacgc 60 atgagtgccc tcagcagtag ctccgtgtgg cagagctgaa agaggagctt gatgacacgt 120 ccaaccttta ggagccatg 139 <210> 4 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> AP:dA pattern sequence <220> <221> misc_feature <222> (48)..(48) <223> n represents the AP site. <400> 4 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatntc agcatgacgc 60 atgagtgccc tcagcagtag ctccgtgtgg cagagctgaa agaggagctt gatgacacgt 120 ccaaccttta ggagccatg 139 <210> 5 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> dU:dG mode sequence <220> <221> misc_feature <222> (48)..(48) <223> n is uracil deoxyribonucleotide. <400> 5 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatntc agcatgacgg 60 cggctggagc ggtaattttg ctccgtgtgg cagagctgaa agaggagctt gatgacacgt 120 aatgacgttg ccagccagt 139 <210> 6 <211> 145 <212> DNA <213> Artificial Sequence <220> <223> d5fC:dG mode sequence <220> <221> misc_feature <222> (100) <223> n is 5-aldehyde cytosine deoxyribonucleotide. <400> 6 catgagtgcc ctcagcagta agtaactgac cagatctctc gtgcctcttg aggctactga 60 gttatccaac ctttaggagc catgcatcga tagcatccgn cacaggcagt gaggctactg 120 agtcatgcac gcagaaagaa atagc 145 <210> 7 <211> 31 <212> DNA <213> Artificial Sequence <220> <223> Y-type adapter forward single-stranded sequence <400> 7 gatcggaaga gcacacgtct gaactccagt c 31 <210> 8 <211> 33 <212> DNA <213> Artificial Sequence <220> <223> Y-type adapter reverse single-stranded sequence <400> 8 acactctttc cctacacgac gctcttccga tct 33 <210> 9 <211> 58 <212> DNA <213> Artificial Sequence <220> <223> Universal Primer sequence <400> 9 aatgatacgg cgaccaccga gatctacact ctttccctac acgacgctct tccgatct 58 <210> 10 <211> 64 <212> DNA <213> Artificial Sequence <220> <223> Index Primer sequence <220> <221> misc_feature <222> (25) (30) <223> n are each independently selected from guanine deoxyribonucleotide, adenine deoxyribonucleotide, thymine pyridine deoxyribonucleotide or cytosine deoxyribonucleotide <400> 10 caagcagaag acggcatacg agatnnnnnn gtgactggag ttcagacgtg tgctcttccg 60 atct 64 <210> 11 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> Control pattern sequence qPCR primer-1 <400> 11 ttatctggcg gtggagatg 19 <210> 12 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> Control pattern sequence qPCR primer-2 <400> 12 gtttgccacg caagcatta 19 <210> 13 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-1 pattern sequence qPCR primer-1 <400> 13 gcggtaagca cgaactcag 19 <210> 14 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-1 pattern sequence qPCR primer-2 <400> 14 gccacaccac ggtatttcc 19 <210> 15 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-2 pattern sequence qPCR primer-1 <400> 15 catgagtgcc ctcagcagta 20 <210> 16 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dU:dA-2 pattern sequence qPCR primer-2 <400> 16 catggctcct aaaggttgga 20 <210> 17 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> AP:dA pattern sequence qPCR primer-1 <400> 17 catgagtgcc ctcagcagta 20 <210> 18 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> AP:dA pattern sequence qPCR primer-2 <400> 18 catggctcct aaaggttgga 20 <210> 19 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dU:dG pattern sequence qPCR primer-1 <400> 19 gcggctggag cggtaatttt 20 <210> 20 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dU:dG pattern sequence qPCR primer-2 <400> 20 actggctggc aacgtcatta 20 <210> twenty one <211> 20 <212> DNA <213> Artificial Sequence <220> <223> d5fC:dG pattern sequence qPCR primer-1 <400> twenty one catgagtgcc ctcagcagta 20 <210> twenty two <211> 20 <212> DNA <213> Artificial Sequence <220> <223> d5fC:dG pattern sequence qPCR primer-2 <400> twenty two catggctcct aaaggttgga 20 <210> twenty three <211> 20 <212> DNA <213> Artificial Sequence <220> <223> VEGFA_site_2 sgRNA target site sequence <400> twenty three gaccccctcc accccgcctc 20 <210> twenty four <211> 20 <212> DNA <213> Artificial Sequence <220> <223> HEK293 site_4 sgRNA target site sequence <400> twenty four ggcactgcgg ctggaggtgg 20 <210> 25 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> EMX1 sgRNA target site sequence <400> 25 gagtccgagc agaagaagaa 20 <210> 26 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> RNF2 sgRNA target site sequence <400> 26 gtcatcttag tcattacctg 20 <210> 27 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> RUNX1 sgRNA target site sequence <400> 27 tcccctctgc tggatacctc 20 <210> 28 <211> 137 <212> DNA <213> Artificial Sequence <220> <223> dI:dC mode sequence <220> <221> misc_feature <222> (52)..(52) <223> n represents inosine deoxyribonucleotide. <400> 28 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatttc ancatgacga 60 atgtggatgc cgcagttggc tccgtgtggc agagctgaaa gaggagcttg atgacacgca 120 accgggacat cacggat 137 <210> 29 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> dI:dT mode sequence <220> <221> misc_feature <222> (40)..(40) <223> n represents inosine deoxyribonucleotide. <400> 29 aactgattgc ccgtctccgc tcgctgggtg aacaactgan ccgtgatgtc agcatgacgc 60 tacgcaaact ggctgtcaag ctccgtgtgg cagagctgaa agaggagctt gatgacacgt 120 catggacgct acctcacag 139 <210> 30 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> Nick pattern sequence <400> 30 aactgattgc ccgtctccgc tcgctgggtg aacaactgaa ccgtgatgtc agcatgacga 60 ggccaacata catgccttcg ctccgtgtgg cagagctgaa agaggagctt gatgacacgg 120 aatggcagag tcaaggagc 139 <210> 31 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> dI:dC pattern sequence qPCR primer-1 <400> 31 aatgtggatg ccgcagttg 19 <210> 32 <211> 19 <212> DNA <213> Artificial Sequence <220> <223> dI:dC pattern sequence qPCR primer-2 <400> 32 atccgtgatg tcccggttg 19 <210> 33 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dI:dT pattern sequence qPCR primer-1 <400> 33 ctacgcaaac tggctgtcaa 20 <210> 34 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> dI:dT pattern sequence qPCR primer-2 <400> 34 ctgtgaggta gcgtccatga 20 <210> 35 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> Nick pattern sequence qPCR primer-1 <400> 35 aggccaacat acatgccttc 20 <210> 36 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> Nick pattern sequence qPCR primer-2 <400> 36 gctccttgac tctgccattc 20 <210> 37 <211> twenty three <212> DNA <213> Artificial Sequence <220> <223> RUNX1 crRNA target site sequence <400> 37 ttctcccctc tgctggatac ctc 23 <210> 38 <211> twenty three <212> DNA <213> Artificial Sequence <220> <223> DYRK1A crRNA target site sequence <400> 38 gaagcacatc aaggacattc taa 23
Claims
1. A method for detecting, for non-diagnostic purposes, the editing site, editing efficiency, or off-target effects of a base editor-edited target nucleic acid, comprising the following steps: (1) Providing an editing product for editing a target nucleic acid using a base editor, comprising a base editing intermediate, wherein the base editing intermediate comprises a first nucleic acid strand and a second nucleic acid strand; wherein, The first nucleic acid strand contains edited bases generated by the base editor editing the target nucleic acid; (2) In the first nucleic acid chain, a single-strand break is generated in the segment containing the edited base using a nuclease; (3) Using a nucleic acid polymerase with chain displacement activity, a nucleotide labeled with a first labeling molecule and a nucleotide labeled with a second labeling molecule are introduced at or downstream of the single-strand break to produce a labeled product containing the first labeling molecule and the second labeling molecule; (4) Use a first binding molecule that can specifically recognize and bind to the first labeled molecule to separate or enrich the labeled product; (5) Determine the sequence of the labeled product; Thus, the editing site, editing efficiency, or off-target effects of the base editor on the target nucleic acid can be determined; Wherein, the nucleotide labeled by the second labeling molecule is 5-aldehyde cytosine deoxyribonucleotide; and the method further comprises: Prior to step (3), nucleotides that may be labeled with a second marker molecule in the edited product are protected; and, After step (3), before determining the sequence of the labeled product, the labeled product is treated with a compound to alter the base-complementary pairing ability of the nucleotides labeled by the second labeling molecule contained therein, the compound being selected from malononitriles or borane compounds.
2. The method of claim 1, wherein, In step (2), a single-strand break is generated in the first nucleic acid chain within a 10nt upstream to 10nt downstream segment of the edited base.
3. The method of claim 1, wherein, The base editor is either a single-base editor or a double-base editor.
4. The method of claim 1, wherein, The base editor is a cytosine base editor, an adenine base editor, or an adenine and cytosine dual-base editor.
5. The method of claim 1, wherein, The target nucleic acid is either genomic nucleic acid or mitochondrial nucleic acid.
6. The method of claim 1, wherein, The edited product is the product of the base editor editing the target nucleic acid outside or inside the cell.
7. The method of claim 1, wherein, The edited product is the result of the base editor editing the target nucleic acid within the organelle.
8. The method of claim 1, wherein, The method further includes the following step before step (1): under the condition that the base editor is allowed to edit the target nucleic acid, the base editor is brought into contact with the target nucleic acid to generate the edited product.
9. The method of claim 1, wherein, The method further includes the following step before step (1): under conditions that allow the base editor to edit the target nucleic acid, the base editor is brought into contact with the target nucleic acid, either outside the cell or inside the cell, thereby generating an edited product.
10. The method of claim 1, wherein, The method further includes the following step before step (1): under conditions that allow the base editor to edit the target nucleic acid, the base editor is brought into contact with the target nucleic acid within an organelle to generate an edited product.
11. The method of claim 1, wherein, The method further includes the following steps before step (1): introducing the base editor into the cell, so that the base editor contacts the target nucleic acid in the cell and performs base editing, thereby generating an edited product; or, introducing the nucleic acid molecule encoding the base editor into the cell and making it express the base editor, so that the base editor contacts the target nucleic acid in the cell and performs base editing, thereby generating an edited product.
12. The method of claim 1, wherein, The method further includes the following steps before step (1): introducing the base editor into the organelle, so that the base editor contacts the target nucleic acid in the organelle and performs base editing, thereby generating an edited product; or, introducing the nucleic acid molecule encoding the base editor into the organelle and making it express the base editor, so that the base editor contacts the target nucleic acid in the organelle and performs base editing, thereby generating an edited product.
13. The method of claim 11, wherein, In step (1), the base-edited target nucleic acid is extracted or isolated from the cell and optionally fragmented to obtain the edited product.
14. The method of claim 12, wherein, In step (1), the base-edited target nucleic acid is extracted or isolated from the organelle and optionally fragmented to obtain the edited product.
15. The method of claim 11, wherein, In step (1), the base-edited target nucleic acid is extracted or isolated from the cell and then fragmented and repaired to obtain the edited product.
16. The method of claim 12, wherein, In step (1), the base-edited target nucleic acid is extracted or isolated from the organelle and then fragmented and repaired to obtain the edited product.
17. The method of claim 15 or 16, wherein, The end repair refers to the filling of the 5' end overhang and / or the excision of the 3' end overhang.
18. The method of claim 1, wherein, The second nucleic acid strand has not undergone base editing or does not contain edited bases.
19. The method of claim 1, wherein, The edited bases are selected from uracil or hypoxanthine.
20. The method of claim 1, wherein, In step (2), a single-strand break is created at the location of the edited base or upstream or downstream of it.
21. The method of claim 1, wherein, Before performing step (2), the method further includes the step of repairing single-chain breaks (SSBs) that may exist in the edited product.
22. The method of claim 1, wherein, Before performing step (2), the method further includes using a nucleic acid polymerase, nucleotides, and a nucleic acid ligase to repair any single-strand breaks (SSBs) that may exist in the edited product.
23. The method of claim 22, wherein, The nucleotides used to repair SSBs that may be present in the edited product are unlabeled nucleotides.
24. The method of claim 21 or 22, wherein, The single-chain break (SSB) is an endogenous single-chain break.
25. The method of claim 1, wherein, The endonuclease is selected from endonuclease V, endonuclease VIII, or AP endonuclease.
26. The method of claim 1, wherein, The nucleotide labeled with the first labeling molecule is selected from uracil deoxyribonucleotide labeled with the first labeling molecule, cytosine deoxyribonucleotide labeled with the first labeling molecule, thymine deoxyribonucleotide labeled with the first labeling molecule, adenine deoxyribonucleotide labeled with the first labeling molecule, guanine deoxyribonucleotide labeled with the first labeling molecule, or any combination thereof.
27. The method of claim 1, wherein, The first labeling molecule is biotin or a functional variant thereof, and the first binding molecule is avidin or a functional variant thereof; or, the first labeling molecule is a hapten or antigen, and the first binding molecule is an antibody specifically against the hapten or antigen; or, the first labeling molecule contains an alkyne group, and the first binding molecule is an azide compound capable of undergoing a click chemical reaction with the alkyne group.
28. The method of claim 1, wherein, The nucleotide labeled by the first labeling molecule is a nucleotide containing an acetylene group, and the first binding molecule is an azide compound capable of undergoing a click chemical reaction with the acetylene group.
29. The method of claim 28, wherein, The nucleotide containing the acetylene group is 5-Ethynyl-dUTP, and / or the azide compound capable of undergoing a click chemical reaction with the acetylene group is an azide-modified magnetic bead.
30. The method of claim 1, wherein, The borane compound is a pyridineborane compound.
31. The method of claim 30, wherein, The pyridineborane compounds are pyridineborane or 2-methylpyridineborane.
32. The method of claim 1, wherein, The processing steps for the labeled product are performed before step (4) or before step (5).
33. The method of claim 1, wherein, The protection includes: protecting endogenous 5-aldehyde cytosine deoxyribonucleotides with ethyl hydroxylamine.
34. The method of claim 1 or 33, wherein, The protection is performed before step (2).
35. The method of claim 1, wherein, In step (4), the labeled product is separated or enriched using a first binding molecule attached to a solid support.
36. The method of claim 1, wherein, Before performing step (5), the method further includes: amplifying the labeled products isolated or enriched in step (4); and / or constructing a sequencing library from the labeled products isolated or enriched in step (4).
37. The method of claim 1, wherein, In step (5), the sequence of the labeled product is determined by sequencing, hybridization or mass spectrometry.
38. The method of claim 1, wherein, The method further includes comparing the sequence determined in step (5) with a reference sequence to determine the editing site, editing efficiency or off-target effect of the base editor on the target nucleic acid.
39. The method of claim 38, wherein, The reference sequence is the target nucleic acid sequence before base editing.
40. The method of claim 39, wherein, The target nucleic acid sequence before base editing is obtained from a database or through sequencing methods.
41. The method of claim 1, wherein, The base editor is a cytosine base editor.
42. The method of claim 41, wherein, The cytosine base editor is either a nuclear cytosine base editor or an organelle cytosine base editor.
43. The method of claim 41, wherein, The cytosine base editor is a cytosine base editor capable of editing cytosine into uracil.
44. The method of claim 41, wherein, The base editor is either a cytosine base editor capable of editing nuclear nucleic acids or a cytosine base editor capable of editing mitochondrial nucleic acids.
45. The method of claim 41, wherein, The edited base is uracil.
46. The method of claim 41, wherein, The base editing intermediate is a DNA molecule containing uracil.
47. The method of claim 41, wherein, The nucleotide molecule containing the second label is 5-aldehyde cytosine deoxyribonucleotide, which is capable of base pairing with guanine deoxyribonucleotide before treatment and with adenine deoxyribonucleotide after treatment.
48. The method of claim 41, wherein, In step (2), an AP site-specific endonuclease is used to generate a single-strand break at the location of the edited base in the first nucleic acid strand; and in step (3), the nucleotide labeled with the first labeling molecule and the nucleotide labeled with the second labeling molecule are introduced at and downstream of the single-strand break to generate a labeled product containing the first labeling molecule and the second labeling molecule.
49. The method of claim 48, wherein, The AP site-specific endonuclease is the AP endonuclease.
50. The method of claim 48, wherein, Before performing step (2), the method further includes the step of forming an AP site at the location of the base edited in the first nucleic acid strand.
51. The method of claim 48, wherein, Before performing step (2), the method further includes the step of incubating the edited product with UDG (uracil-DNA glycosylation enzyme).
52. The method of claim 51, wherein, Prior to the incubation with UDG step, the method further includes the step of repairing any AP sites that may be present in the edited product.
53. The method of claim 52, wherein, The AP site repair steps include: (a) The AP endonuclease is incubated with the editing product, where an AP site may be present, under conditions that allow the AP endonuclease to exert its cleavage activity; (b) Under conditions that allow nucleic acid polymerization, the product of step (a) is incubated with nucleic acid polymerase and nucleotide molecules; (c) Under conditions that allow the nuclease to exert its ligation activity, incubate the product of step (b) with the nuclease. Thus, AP sites that may exist in the edited product are repaired.
54. The method of claim 53, in step (b), wherein the product of step (a) is incubated with a nucleic acid polymerase and nucleotide molecules not containing a first label or a second label under conditions that allow nucleic acid polymerization.
55. The method of claim 1, wherein, The base editor is an adenine base editor.
56. The method of claim 55, wherein, The adenine base editor is an adenine base editor capable of editing adenine into hypoxanthine.
57. The method of claim 55, wherein, The edited base is hypoxanthine.
58. The method of claim 55, wherein, The base editing intermediate is a DNA molecule containing hypoxanthine.
59. The method of claim 55, wherein, In step (2), a hypoxanthine site-specific endonuclease is used to generate a single-strand break at or downstream of the position of the edited base in the first nucleic acid strand; and in step (3), the nucleotide labeled with the first labeling molecule and the nucleotide labeled with the second labeling molecule are introduced at and downstream of the single-strand break to generate a labeled product containing the first labeling molecule and the second labeling molecule.
60. The method of claim 59, wherein, The hypoxanthine site-specific endonuclease is endonuclease V or endonuclease VIII.
61. The method of claim 1, wherein, The base editor is a two-base editor.
62. The method of claim 61, wherein, The base editor is a base editor capable of editing cytosine to uracil and adenine to hypoxanthine.
63. The method of claim 61, wherein, The edited bases are hypoxanthine and / or uracil.
64. The method of claim 61, wherein, The base editing intermediate is a DNA molecule containing hypoxanthine and / or uracil.