Improved genome editing method

CN120813698APending Publication Date: 2025-10-17PEKING UNIV INST OF ADVANCED AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014660.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-22
Filing Date
2024-09-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing genome editing technologies have problems of low accuracy and low efficiency when accurately inserting or replacing exogenous sequences, especially low HDR efficiency.

Method used

The donor DNA fragment with a sticky end and a site-directed nuclease are used to cleave the target sequence of the genomic DNA, and the sticky end and the sticky end of the genomic DNA are paired complementarily to achieve the integration of the donor DNA fragment. This system is called the mortise and tenon connection system (MT system).

Benefits of technology

It significantly improves the integration efficiency and accuracy of exogenous DNA fragments at designated sites, and improves the universality and efficiency of polynucleotide knock-in and replacement compared with traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000033_0000
    Figure 00000033_0000
  • Figure 00000033_0001
    Figure 00000033_0001
  • Figure 00000034_0000
    Figure 00000034_0000
Patent Text Reader

Abstract

Relates to the field of gene engineering. The invention relates to an improved genome editing method, which can be used for accurately inserting an exogenous sequence into a genome at a fixed point.
Need to check novelty before this filing date? Find Prior Art

Description

Improved genome editing methods Technical Field

[0001] The present invention relates to the field of genetic engineering. Specifically, the present invention relates to an improved genome editing method that can precisely insert exogenous sequences into a genome.

[0002] Background of the Invention

[0003] In recent years, with the continuous development of gene editing technology, a large number of gene editing tools have been developed, improved, and applied. After genome editing technology introduces double-strand breaks at a specific site in the genome, the editing results and purposes mainly include base / fragment deletions, insertions, and substitutions. The strategies for achieving this goal mainly include the following four methods: 1) using the organism's non-homologous end joining (NHEJ) mechanism to introduce random mutations; 2) using the organism's non-homologous end joining mechanism to delete a sequence; 3) using the organism's non-homologous end joining mechanism to insert a sequence; and 4) using homologous recombination to insert, delete, or replace a sequence. Targeted DNA insertion or replacement holds great potential for gene therapy, breeding improvement, and theoretical research. Currently, a variety of gene editing tools that rely on double-strand breaks are capable of inserting or replacing DNA fragments, such as using the homologous recombination HDR repair mechanism for gene editing. However, HDR efficiency is relatively low. In 2020, Zhu Jiankang et al. used SpCas9 to cut the recipient DNA and integrate the modified donor DNA fragment into the recipient plant genome, but this method has disadvantages such as uncertain direction of the knock-in fragment and low accuracy.

[0004] Therefore, this field still needs methods that can insert or replace specified exogenous sequences into double-strand break gaps in the genome, especially accurate, efficient, and universal polynucleotide knock-in and replacement gene editing systems.

[0005] Summary of the Invention

[0006] The present invention provides a method for inserting or replacing a given DNA into a double-strand break in a genome. The method utilizes a donor DNA fragment with sticky ends as a tenon structure, and a site-directed nuclease to cut a targeted region of a genomic DNA target sequence to produce sticky ends that are not fully complementary to themselves, which serve as a tenon structure. The donor DNA sticky ends complementarily pair with the genomic DNA sticky ends, thereby integrating the donor DNA fragment into the cut site. Therefore, this system is named the mortise-tenon joint system (MT system for short).

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1. Schematic diagram of the MT system's operation. (A) Schematic diagram of the operation with non-complementary 5' sticky ends at both ends of the cut. (B) Schematic diagram of the operation with non-complementary 5' sticky ends and 3' blunt ends at both ends of the cut. (C) Schematic diagram of the operation with non-complementary 3' sticky ends and 5' blunt ends at both ends of the cut. (D) Schematic diagram of the operation with non-complementary 3' sticky ends at both ends of the cut. (E) Schematic diagram of the operation with non-complementary 5' sticky ends and 3' sticky ends at both ends of the cut. (F) Schematic diagram of the operation with non-complementary 3' sticky ends and 5' sticky ends at both ends of the cut.

[0009] Figure 2. Schematic diagram showing the principle of site-directed knock-in of target fragments using DNA glycosylase or DNA glycosylase plus deaminase, nuclease, AP lyase (apurinic and apyrimidinic site lyase) and donor DNA fragments.

[0010] Figure 3. Schematic diagrams illustrating three methods for generating 5' sticky ends and the site-specific knock-in of a donor DNA fragment into a target fragment. (A): Schematic diagram illustrating the site-specific knock-in of a target fragment using cytosine deaminase and thymine glycosylase, nuclease, AP lyase (apurinic apyrimidinic site lyase), and a donor DNA fragment; (B): Schematic diagram illustrating the site-specific knock-in of a target fragment using adenine deaminase and hypoxanthine glycosylase, nuclease, AP lyase, and a donor DNA fragment; (C): Schematic diagram illustrating the site-specific knock-in of a target fragment using DNA glycosylase, nuclease, AP lyase, and a donor DNA fragment.

[0011] Figure 4. Schematic diagrams showing three types of site-specific knock-in of target fragments using specific nucleases plus exonucleases and donor DNA fragments. (A): Schematic diagram showing the site-specific knock-in of target fragments using TALEN plus exonucleases and donor DNA fragments; (B): Schematic diagram showing the site-specific knock-in of target fragments using ZFN plus exonucleases and donor DNA fragments; (C): Schematic diagram showing the site-specific knock-in of target fragments using CRISPR nucleases plus exonucleases and donor DNA fragments.

[0012] Figure 5. Targeted knock-in genomic PCR and sequencing results in the examples. (A) PCR detection results for the targeted knock-in of the translation enhancer at the rice SLR1 gene locus. (B) Sequencing results for the targeted knock-in of the translation enhancer at the rice SLR1 gene locus. (C) PCR detection results for the targeted knock-in of the miRNA396 recognition site at the rice NRT1.1B gene locus. (D) Sequencing results for the targeted knock-in of the miRNA396 recognition site at the rice NRT1.1B gene locus.

[0013] Figure 6. Shows the efficiency and accuracy of insertion of sequences of different lengths at different target sites using the MT1, MT2, and Cas9 systems: precise insertion efficiency, precise insertion efficiency only at the 5' end, precise insertion efficiency only at the 3' end, imprecise insertion efficiency at both the 5' and 3' ends, and reverse insertion rate. (A): 21 bp donor DNA fragment inserted into the NRT1.1B target site; (B): 66 bp donor DNA fragment inserted into the ACTIN1 target site; (C): 21 bp donor DNA fragment inserted into the GRF1 target site; (D): 21 bp donor DNA fragment inserted into the IPA1 target site; (E): 64 bp donor DNA fragment inserted into the SOS1 target site; (F): 64 bp donor DNA fragment inserted into the SLR1 target site.

[0014] Figure 7. Phenotypes of genome-targeted knock-in. (A) Western blot of transgenic seedlings with a Flag tag targeted to the C-terminus of the rice ACTIN1 gene. (B) Phenotype of transgenic seedlings with a translation enhancer targeted to the SLR1 gene.

[0015] FIG8 shows a comparative analysis of the knock-in efficiency of SpCas9 and MT systems at different target sites.

[0016] Figure 9. Illustrated showing the creation of sticky ends for insertion of foreign sequences through dual-target targeting.

[0017] Figure 10 shows the precise site-directed insertion of foreign sequences using the adenine deaminase-based ABE-MT system.

[0018] Figure 11. Insertion of a donor DNA fragment formed by multi-fragment splicing by the MT system.

[0019] Detailed Description of the Invention

[0020] 1. Definition

[0021] In the present invention, unless otherwise indicated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. In addition, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are terms and routine procedures widely used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are more fully described in the following literature: Sambrook, J., Fritsch, EF and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). At the same time, in order to better understand the present invention, definitions and explanations of relevant terms are provided below.

[0022] As used herein, the term "and / or" encompasses all combinations of items connected by the term, and should be treated as if each combination had been individually listed herein. For example, "A and / or B" encompasses "A," "A and B," and "B." For example, "A, B, and / or C" encompasses "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."

[0023] When the word "comprising" is used herein to describe a sequence of a protein or nucleic acid, the protein or nucleic acid may be composed of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present invention. In addition, it is clear to those skilled in the art that the methionine encoded by the start codon at the N-terminus of the polypeptide may be retained in certain practical situations (for example, when expressed in a specific expression system), but it does not substantially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of this application, although it may not contain a methionine encoded by a start codon at the N-terminus, a sequence containing the methionine is also covered, and accordingly, its encoding nucleotide sequence may also contain a start codon; and vice versa.

[0024] As used herein, a "genome editing system" refers to a combination of components required for genome editing of a cell's genome. The individual components of the system may exist independently or in any combination as a composition.

[0025] "Genome" as used herein encompasses not only the chromosomal DNA present in the nucleus of a cell, but also the organelle DNA present in subcellular components of the cell (eg, mitochondria, plastids).

[0026] As used herein, "organism" includes any organism suitable for genome editing, preferably a eukaryotic organism. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis.

[0027] "Genetically modified organism" or "genetically modified cell" refers to an organism or cell that contains an exogenous polynucleotide or a modified gene or expression control sequence within its genome. For example, the exogenous polynucleotide is capable of stably integrating into the genome of the organism or cell and being inherited through successive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. A modified gene or expression control sequence is one that contains single or multiple deoxynucleotide substitutions, deletions, and additions within the genome of the organism or cell.

[0028] "Exogenous" with respect to a sequence refers to a sequence that is from a foreign species, or, if from the same species, has been significantly altered from its native form in composition and / or locus through deliberate human intervention.

[0029] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single-stranded or double-stranded polymers of RNA or DNA that optionally contain synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single-letter designations as follows: "A" for adenosine or deoxyadenosine (RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide. Although nucleotide sequences herein may be presented as DNA sequences (including T), when reference is made to RNA, one skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U).

[0030] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residues is an artificial chemical analog of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" may also include modified forms including, but not limited to, glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.

[0031] Sequence "identity" has a meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the entire length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are many methods to measure the identity between two polynucleotides or polypeptides, the term "identity" is well known to those of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).

[0032] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. Generally, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).

[0033] As used herein, an "expression construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to produce mRNA or functional RNA) and / or translation of RNA into a precursor or mature protein.

[0034] The "expression construct" of the present invention can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be an RNA (such as mRNA) that can be translated.

[0035] An "expression construct" of the present invention may comprise regulatory sequences and a nucleotide sequence of interest from different sources, or regulatory sequences and a nucleotide sequence of interest from the same source but arranged in a manner different from that normally found in nature.

[0036] "Regulatory sequence" and "regulatory element" are used interchangeably to refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence and that influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.

[0037] "Promoter" refers to a nucleic acid fragment that is capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter that is capable of controlling the transcription of a gene in a cell, whether or not it is derived from the cell. A promoter can be a constitutive promoter, a tissue-specific promoter, a developmentally regulated promoter, or an inducible promoter.

[0038] "Constitutive promoter" refers to a promoter that will generally cause a gene to be expressed in most cell types under most circumstances. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed primarily, but not necessarily exclusively, in one tissue or organ, and may also be expressed in one specific cell or cell type. "Developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (environmental, hormonal, chemical signals, etc.).

[0039] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. Examples of pol I promoters include the chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, the cytomegalovirus immediate early (CMV) promoter, the Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and the simian virus 40 (SV40) immediate early promoter. Examples of pol III promoters include the U6 and H1 promoters. Inducible promoters such as the metallothionein promoter can be used. Other examples of promoters include the T7 phage promoter, the T3 phage promoter, the β-galactosidase promoter, and the Sp6 phage promoter. When used in plants, the promoter can be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, or the rice actin promoter.

[0040] As used herein, the term "operably linked" refers to the connection of a regulatory element (e.g., but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (e.g., a coding sequence or an open reading frame) such that transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory element. Techniques for operably linking regulatory element regions to nucleic acid molecules are known in the art.

[0041] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism refers to transforming an organism cell with the nucleic acid or protein so that the nucleic acid or protein can function in the cell. "Transformation" as used herein includes stable transformation and transient transformation.

[0042] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into a genome, resulting in the stable inheritance of the exogenous gene. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof.

[0043] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell where it functions without the foreign gene being stably inherited. In transient transformation, the foreign nucleic acid sequence does not integrate into the genome.

[0044] 2. Improved Genome Editing Methods

[0045] In one aspect, the present invention provides a genome editing method for site-specific insertion of an exogenous nucleotide sequence into a cell genome, the method comprising:

[0046] (a) generating a DNA double-strand break at a target region in the genome of the cell, the DNA double-strand break resulting in at least one sticky end of genomic DNA; and

[0047] (b) providing a donor DNA fragment comprising an exogenous nucleotide sequence to be inserted to the cell, wherein the donor DNA fragment comprises a sticky end complementary to at least one sticky end of the genomic DNA.

[0048] The present inventors surprisingly discovered that the site-specific integration of the donor DNA fragment into the target region of the genome can be significantly improved by complementary pairing of the cohesive end in the donor DNA fragment with at least one cohesive end generated by the DNA double-strand break.

[0049] In some embodiments, the DNA double-strand break results in a sticky end and a blunt end of the genomic DNA.

[0050] In some embodiments, the DNA double-strand break results in two sticky ends of genomic DNA that are not fully complementary. Self-complementary pairing of the two sticky ends of the genomic DNA may reduce the efficiency of complementary pairing between the sticky ends of the donor DNA fragment and the sticky ends of the genomic DNA, thereby reducing the insertion efficiency. Therefore, it is desirable to reduce the self-complementary pairing of the two sticky ends of the genomic DNA. In some embodiments, the two sticky ends of the genomic DNA cannot form a stable or effective duplex through complementary hybridization. In some embodiments, when complementary alignment is achieved, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, or no more than 5% of the nucleotides of any sticky end of the genomic DNA can pair with nucleotides of the other sticky end of the genomic DNA. In some embodiments, when complementary alignment is achieved, the two sticky ends of the genomic DNA can form no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, no more than 5, no more than 4, no more than 3, no more than 2, or no more than 1 base pairs (e.g., consecutive base pairs).

[0051] In some embodiments, the DNA double-strand break results in a 3' cohesive end and a blunt end of the genomic DNA, and the donor DNA fragment comprises a 3' cohesive end complementary to the 3' cohesive end of the genomic DNA and a blunt end.

[0052] As used herein, "3' sticky end" refers to a single-stranded region with a free 3' end present at the end of a double-stranded DNA; "5' sticky end" refers to a single-stranded region with a free 5' end present at the end of a double-stranded DNA; and "blunt end" refers to a double-stranded DNA end without a single-stranded region.

[0053] In some embodiments, the DNA double-strand break results in a 5' sticky end and a blunt end of the genomic DNA, and the donor DNA fragment comprises a 5' sticky end complementary to the 5' sticky end of the genomic DNA and a blunt end.

[0054] In some embodiments, the DNA double-strand break results in two 3' sticky ends of the genomic DNA that cannot complementarily pair, and the donor DNA fragment comprises two 3' sticky ends that complementarily pair with the two 3' sticky ends of the genomic DNA, respectively.

[0055] In some embodiments, the DNA double-strand break results in two 5' sticky ends of the genomic DNA that cannot complementarily pair, and the donor DNA fragment comprises two 5' sticky ends that complementarily pair with the two 5' sticky ends of the genomic DNA, respectively.

[0056] In some embodiments, the DNA double-strand break results in a 5' sticky end and a 3' sticky end of the genomic DNA, and the donor DNA fragment comprises a 5' sticky end that is complementary to the 5' sticky end of the genomic DNA and a 3' sticky end that is complementary to the 3' sticky end of the genomic DNA.

[0057] In some embodiments, the sticky ends of the genomic DNA are complementary to the corresponding sticky ends of the donor DNA fragments. In some embodiments, there are no more than 5, no more than 4, no more than 3, no more than 2, or no more than 1 mismatch between the sticky ends of the genomic DNA and the corresponding sticky ends of the donor DNA fragments. In some embodiments, the sticky ends of the genomic DNA and the corresponding sticky ends of the donor DNA fragments are perfectly complementary.

[0058] By utilizing two cohesive ends of genomic DNA that cannot completely complement each other or utilizing one cohesive end and one blunt end, the exogenous nucleotide sequence can be inserted in a selected direction (directionally).

[0059] In some embodiments, the sticky ends are about 2 to about 50 nucleotides in length, preferably 8 to 9 nucleotides in length. In some embodiments, the sticky ends of the genomic DNA and the corresponding sticky ends of the donor DNA fragments are substantially the same length (e.g., differing by no more than 5, no more than 4, no more than 3, no more than 2, no more than 1 nucleotide in length), preferably having the same length.

[0060] In some embodiments, the method generates the DNA double-strand break at a target region in the genome of the cell by introducing into the cell a sequence-specific nuclease and / or an expression construct comprising a nucleotide sequence encoding a sequence-specific nuclease.

[0061] The sequence-specific nuclease can be any sequence-specific nuclease known in the art. For example, the sequence-specific nuclease can be a ZFN, a TALEN, a restriction endonuclease or a CRISPR nuclease, preferably a CRISPR nuclease.

[0062] In some embodiments, the method comprises introducing into the cell:

[0063] i) sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases, and optionally uracil-DNA glycosylases (UDG), and / or expression constructs comprising nucleotide sequences encoding said sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases, and optionally uracil-DNA glycosylases (UDG);

[0064] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0065] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method further comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the CRISPR nuclease to a target region in the cell genome.

[0066] Cytosine deaminase deaminates cytosine from deoxycytidine to produce uracil, which is then glycosylated by uracil glycosylase to form an AP site. Subsequently, AP lyase excises the AP site, creating a single-stranded nick that can interact with the double-stranded nick created by sequence-specific nucleases, such as CRISPR nucleases, to produce a 5' sticky end. See Figure 3A.

[0067] In some embodiments, the sequence-specific nucleases, such as CRISPR nuclease, cytosine deaminase, AP lyase, and optionally uracil-DNA glycosylase (UDG), can form one or more fusion proteins or form a complex within the cell.

[0068] Thus, in some embodiments, the method comprises introducing into the cell:

[0069] i) a fusion protein and / or an expression construct comprising a nucleotide sequence encoding a fusion protein;

[0070] wherein the fusion protein comprises a cytosine deaminase, a sequence-specific nuclease such as a CRISPR nuclease, an AP lyase, and optionally a uracil-DNA glycosylase (UDG), wherein the sequence-specific nuclease is capable of targeting the fusion protein to a target region in the genome of the cell,

[0071] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0072] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method further comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the fusion protein and the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs or can be the same expression construct. In some embodiments, the fusion protein is an isolated fusion protein and / or the guide RNA is an isolated guide RNA.

[0073] In some embodiments, the method comprises introducing into the cell:

[0074] i) a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein; and

[0075] ii) a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein;

[0076] wherein the first fusion protein comprises cytosine deaminase, a sequence-specific nuclease such as a CRISPR nuclease, and optionally uracil-DNA glycosylase (UDG), and the second fusion protein comprises AP lyase, wherein the sequence-specific nuclease is capable of targeting the first fusion protein to a target region in the genome of the cell,

[0077] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0078] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method further comprises introducing into the cell iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the first fusion protein, the expression construct comprising the nucleotide sequence encoding the second fusion protein and / or the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs, or any two or all of them are the same expression construct. In some embodiments, the first fusion protein is an isolated fusion protein, the second polypeptide is an isolated fusion protein and / or the guide RNA is an isolated RNA.

[0079] In some embodiments, the sequence-specific nuclease, cytosine deaminase, AP lyase, and optionally uracil-DNA glycosylase (UDG) can form a complex in cells via a protein recruitment system.

[0080] In some embodiments, the method comprises introducing into the cell:

[0081] i) sequence-specific nucleases such as CRISPR nuclease, adenine deaminase, AP lyase and hypoxanthine glycosylase, and / or expression constructs comprising nucleotide sequences encoding said CRISPR nuclease, adenine deaminase, AP lyase and hypoxanthine glycosylase,

[0082] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0083] Adenine deaminase deaminates adenine from deoxyadenosine to produce hypoxanthine, which is then glycosylated by hypoxanthine glycosylases, such as N-methylpurine DNA glycosylase (MPG), to form an AP site. AP lyase excises the AP site, creating a single-stranded nick that can interact with the double-stranded nick created by sequence-specific nucleases, such as CRISPR nucleases, to generate a 5' sticky end. See Figure 3B.

[0084] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the CRISPR nuclease to a target region in the genome of the cell.

[0085] In some embodiments, the sequence-specific nucleases, such as CRISPR nuclease, adenine deaminase, AP lyase, and hypoxanthine glycosylase, can form one or more fusion proteins or form a complex within the cell.

[0086] Thus, in some embodiments, the method comprises introducing into the cell:

[0087] i) a fusion protein and / or an expression construct comprising a nucleotide sequence encoding a fusion protein;

[0088] wherein the fusion protein comprises adenine deaminase, a sequence-specific nuclease such as CRISPR nuclease, AP lyase, and hypoxanthine glycosylase, wherein the sequence-specific nuclease is capable of targeting the fusion protein to a target region in the cell genome,

[0089] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0090] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the fusion protein and the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs or can be the same expression construct. In some embodiments, the fusion protein is an isolated fusion protein and / or the guide RNA is an isolated guide RNA.

[0091] In some embodiments, the method comprises introducing into the cell:

[0092] i) a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein; and

[0093] ii) a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein;

[0094] wherein the first fusion protein comprises adenine deaminase, a sequence-specific nuclease such as a CRISPR nuclease, and hypoxanthine glycosylase, and the second fusion protein comprises AP lyase, wherein the sequence-specific nuclease is capable of targeting the first fusion protein to a target region in the cell genome,

[0095] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0096] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the process comprises introducing into the cell iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the first fusion protein, the expression construct comprising the nucleotide sequence encoding the second fusion protein, and / or the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs, or any two or all of them can be the same expression construct. In some embodiments, the first fusion protein is an isolated fusion protein, the second polypeptide is an isolated fusion protein, and / or the guide RNA is an isolated RNA.

[0097] In some embodiments, the sequence-specific nucleases, such as CRISPR nuclease, adenine deaminase, AP lyase, and hypoxanthine glycosylase, can form a complex in cells through a protein recruitment system.

[0098] In some embodiments, the method comprises introducing into the cell:

[0099] i) sequence-specific nucleases such as CRISPR nucleases, DNA glycosylases, and AP lyases, and / or expression constructs comprising nucleotide sequences encoding said CRISPR nucleases, DNA glycosylases, and AP lyases;

[0100] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0101] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the CRISPR nuclease to a target region in the genome of the cell.

[0102] The DNA glycosylase may be guanine glycosylase, thymine glycosylase, adenine glycosylase, or cytosine glycosylase.

[0103] Guanine on DNA can be glycosylated and removed by guanine DNA glycosylases (such as N-methylpurine DNA glycosylase (MPG), 8-oxoguanine DNA glycosylase (OGG1), and mutY DNA glycosylase (MUTYH)) to form AP sites. Similarly, thymine glycosylase TDG can mediate the removal of thymine to produce AP sites, adenine glycosylases (such as MUTYH) can mediate the removal of adenine to produce AP sites, and cytosine glycosylases (such as 5-methylcytosine glycosylase ROS1) can mediate the removal of cytosine to produce AP sites. Subsequently, the single-stranded gap generated by the removal of the AP site by AP lyase can cooperate with the double-stranded gap generated by sequence-specific nucleases such as CRISPR nuclease to produce 5' sticky ends. See Figure 3C.

[0104] In some embodiments, sequence-specific nucleases such as CRISPR nucleases, DNA glycosylases, and AP lyases can form one or more fusion proteins or form a complex within a cell.

[0105] Thus, in some embodiments, the method comprises introducing into the cell:

[0106] i) a fusion protein and / or an expression construct comprising a nucleotide sequence encoding a fusion protein;

[0107] wherein the fusion protein comprises a DNA glycosylase, a sequence-specific nuclease such as a CRISPR nuclease and an AP lyase, wherein the sequence-specific nuclease is capable of targeting the fusion protein to a target region in the cell genome,

[0108] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0109] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the fusion protein and the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs or can be the same expression construct. In some embodiments, the fusion protein is an isolated fusion protein and / or the guide RNA is an isolated guide RNA.

[0110] In some embodiments, the method comprises introducing into the cell:

[0111] i) a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein; and

[0112] ii) a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein;

[0113] wherein the first fusion protein comprises a DNA glycosylase and a sequence-specific nuclease such as a CRISPR nuclease, and the second fusion protein comprises an AP lyase, wherein the sequence-specific nuclease is capable of targeting the first fusion protein to a target region in the cell genome,

[0114] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0115] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method further comprises introducing into the cell iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the fusion protein to a target region in the cell genome. In some embodiments, the expression construct comprising the nucleotide sequence encoding the first fusion protein, the expression construct comprising the nucleotide sequence encoding the second fusion protein and / or the expression construct comprising the nucleotide sequence encoding the guide RNA can be different expression constructs, or any two or all of them are the same expression construct. In some embodiments, the first fusion protein is an isolated fusion protein, the second polypeptide is an isolated fusion protein and / or the guide RNA is an isolated RNA.

[0116] In some embodiments, the sequence-specific nucleases, such as CRISPR nucleases, DNA glycosylases, and AP lyases, can form a complex in cells via a protein recruitment system.

[0117] In some embodiments, the method comprises introducing into the cell:

[0118] i) sequence-specific nucleases such as CRISPR nucleases and exonucleases, and / or expression constructs comprising nucleotide sequences encoding sequence-specific nucleases such as CRISPR nucleases and exonucleases,

[0119] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0120] Sequence-specific nucleases, such as CRISPR nucleases, create double-stranded DNA breaks at the target region in the cell's genome, while exonucleases can cleave the DNA strands in either direction from the double-stranded break to the 3' terminal nucleotide (3'→5' exonuclease) or the 5' terminal nucleotide (5'→3' exonuclease), thereby forming 5' sticky ends or 3' sticky ends. See Figure 4.

[0121] Sequence-specific nucleases include, but are not limited to, ZFN, Talen, or CRISPR nuclease, preferably CRISPR nuclease. The exonuclease may be a 3'→5' exonuclease or a 5'→3' exonuclease.

[0122] In some embodiments, the sequence-specific nuclease is a CRISPR nuclease. In some embodiments, the method further comprises introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the CRISPR nuclease to a target region in the cell genome.

[0123] In some embodiments, the sequence-specific nuclease, such as a CRISPR nuclease, and the endonuclease can form a complex in the cell through a protein recruitment system; or the sequence-specific nuclease, such as a CRISPR nuclease, and the endonuclease form a fusion protein.

[0124] "ZFN (zinc finger nuclease)" refers to a sequence-specific nuclease based on one or more zinc finger desmin domains (ZFPs). A "zinc finger desmin domain (ZFP)" typically contains 3-6 individual zinc finger repeats, each of which can recognize a unique sequence of, for example, 3 bp. By combining different zinc finger repeats, different genomic sequences can be targeted. ZFNs can be obtained by fusing ZFPs with, for example, FolkI nucleases. Those skilled in the art can easily prepare ZFNs targeting specific sequences.

[0125] "TALEN (transcription activator-like effector nuclease)" refers to sequence-specific nucleases based on one or more "transcription activator-like effector domains". "Transcription activator-like effector domains" are the DNA binding domains of transcription activator-like effectors (TALEs). TALEs can be engineered to bind to almost any desired DNA sequence. TALENs can be obtained by fusing transcription activator-like effector domains to, for example, FolkI nuclease. Those skilled in the art can easily prepare TALENs targeting specific sequences.

[0126] As used herein, the term "CRISPR nuclease" generally refers to a nuclease present in a naturally occurring CRISPR system, as well as a modified form thereof, a variant thereof, or a catalytically active fragment thereof. CRISPR nucleases can recognize, bind to, and / or cut a target nucleic acid structure by interacting with a guide RNA. The term encompasses any nuclease or functional variant thereof that is capable of achieving gene editing in a cell based on a CRISPR system. In some embodiments, the functional variant retains its double-strand cleavage activity, i.e., the ability to form a double-strand break (DSB) in the target sequence.

[0127] The CRISPR nuclease used in the present invention can be selected from, for example, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, GSU0054, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx11, Csx16, CsaX, Csx 3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, C2c3, C2c8, C2c10, Cas12a, Cas12a2, Cas12b (also known as C2c1), Cas12c, Cas12c1, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12f, Cas12k, Cas12m, Cas12n, Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas13m.3, Cas13m.6, Cas14, Casφ, Casλ, TnpB protein, or functional variants of these nucleases.

[0128] In some embodiments, the CRISPR nuclease includes a Cas9 nuclease or a variant thereof. The Cas9 nuclease can be a Cas9 nuclease from a different species, such as spCas9 from Streptococcus pyogenes (S. pyogenes). The Cas9 nuclease variant can, for example, include a highly specific variant of the Cas9 nuclease, such as the Cas9 nuclease variants eSpCas9 (1.0) (K810A / K1003A / R1060A) and eSpCas9 (1.1) (K848A / K1003A / R1060A) by Feng Zhang et al., and the Cas9 nuclease variant SpCas9-HF1 (N497A / R661A / Q695A / Q926A) developed by J. Keith Joung et al. In some specific embodiments, the Cas9 nuclease has the amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the Cas9 nuclease comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO: 1, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 1.

[0129] In some embodiments, the CRISPR nuclease may further comprise a Cpf1 nuclease or a variant thereof, such as a highly specific variant. The Cpf1 nuclease may be a Cpf1 nuclease from a different species, such as a Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6, and Lachnospiraceae bacterium ND2006.

[0130] As used herein, the term "cytosine deaminase" refers to a deaminase that can accept DNA, such as single-stranded DNA, as a substrate and catalyze the deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively. Examples of cytosine deaminases include, but are not limited to, APOBEC1 deaminase, activation-induced cytidine deaminase (AID), APOBEC3G, CDA1, human APOBEC3A deaminase, and truncated APOBEC3B deaminase.

[0131] In some embodiments, the cytosine deaminase is human APOBEC3A deaminase, for example, whose amino acid sequence is shown in SEQ ID NO: 2. In some specific embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 2, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 2.

[0132] In some preferred embodiments, the cytosine deaminase is a truncated APOBEC3B deaminase, for example, whose amino acid sequence is shown in SEQ ID NO: 3. In some specific embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 3, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 3.

[0133] As used herein, uracil-DNA glycosylase (UDG) or uracil-N-glycosylase (UNG) refers to an enzyme that can recognize a U base and remove the N-glycosidic bond of the base to form an apurinic or apyrimidinic site. The UDG can be of different origins, for example, from Escherichia coli. In some embodiments, UDG has the amino acid sequence shown in SEQ ID NO: 4. In some embodiments, UDG comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 3, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 4.

[0134] "AP lyase," "AP lyase," AP endonuclease, and "apurinic pyrimidine lyase" are used interchangeably herein to refer to an enzyme that recognizes and cleaves apurinic or apyrimidinic sites on nucleic acids. The AP lyase can be derived from various sources, such as from Escherichia coli. In some embodiments, the AP lyase has the amino acid sequence set forth in SEQ ID NO:5. In some embodiments, the AP lyase comprises an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:5, or has one or more conservative amino acid substitutions relative to SEQ ID NO:5.

[0135] Examples of "adenine deaminases" that can be used herein include, but are not limited to, Escherichia coli adenine deaminase TadA (tRNA adenosine deaminase (TadA)), comprising the amino acid sequence of SEQ ID NO: 6, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 6, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 6.

[0136] Examples of "hypoxanthine glycosylase" that can be used herein include, but are not limited to, N-methylpurine DNA glycosylase (MPG), which comprises the amino acid sequence shown in SEQ ID NO: 7, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 7, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 7.

[0137] Examples of "guanine glycosylases" that can be used herein include, but are not limited to, N-methylpurine DNA glycosylase (MPG) comprising the amino acid sequence of SEQ ID NO:7, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:7, or having one or more conservative amino acid substitutions relative to SEQ ID NO:7; 8-oxoguanine DNA glycosylase (OGG1) comprising the amino acid sequence of SEQ ID NO:8, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:8, or having one or more conservative amino acid substitutions relative to SEQ ID NO:8; mutY DNA glycosylase (MUTYH) comprising the amino acid sequence of SEQ ID NO:9, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:9; NO:9 has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO:9, or has one or more conservative amino acid substitutions relative to SEQ ID NO:9.

[0138] Examples of "thymine glycosylase" that can be used herein include, but are not limited to, thymine glycosylase TDG, comprising the amino acid sequence shown in SEQ ID NO: 10, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 10, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 10.

[0139] Examples of "adenine glycosylases" that can be used herein include, but are not limited to, mutY DNA glycosylase (MUTYH), comprising the amino acid sequence of SEQ ID NO: 9, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 9, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 9.

[0140] Examples of "cytosine glycosylases" that can be used herein include, but are not limited to, 5-methylcytosine glycosylase ROS1, comprising the amino acid sequence shown in SEQ ID NO: 11, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 11, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 11.

[0141] Examples of "5'→3' exonucleases" that can be used herein include, but are not limited to: a lambda exonuclease comprising the amino acid sequence of SEQ ID NO: 12, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 12, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 12; a T5 exonuclease comprising the amino acid sequence of SEQ ID NO: 13, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 13, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 13; and a recE exonuclease comprising the amino acid sequence of SEQ ID NO: 14, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 14. NO:14 has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the sequence of SEQ ID NO:14, or has one or more conservative amino acid substitutions relative to SEQ ID NO:14.

[0142] Examples of "3'→5' exonucleases" that can be used herein include, but are not limited to: exonuclease II comprising the amino acid sequence of SEQ ID NO: 15, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 15, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 15; and exonuclease III comprising the amino acid sequence of SEQ ID NO: 16, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 16, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 16.

[0143] In some embodiments of the invention, the cytosine deaminase or adenine deaminase or DNA glycosylase is fused to the N-terminus of the CRISPR nuclease.

[0144] In some embodiments of the present invention, the parts of the fusion protein are directly linked to each other.

[0145] In some embodiments of the present invention, the parts of the fusion protein are connected by a linker. The linker can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids in length, without secondary or higher structure. For example, the linker can be a flexible linker, such as GGGGS, GS, GAP, (GGGGS) x 3, GGS, and (GGS) x 7.

[0146] In some embodiments of the present invention, the various enzymes of the present invention and / or fusion proteins also include a nuclear localization sequence (NLS). Generally speaking, one or more NLS in the enzyme and / or fusion protein should have enough intensity, so that in the nucleus of the cell, the enzyme and / or fusion protein are driven to accumulate with the amount that can realize its function. Generally speaking, the intensity of nuclear localization activity is determined by the number, position, one or more specific NLS used or the combination of these factors of NLS in the enzyme and / or fusion protein. In addition, according to the DNA position required for editing, the enzyme of the present invention and / or fusion protein can also include other localization sequences, such as cytoplasmic localization sequence, chloroplast localization sequence, mitochondrial localization sequence etc.

[0147] As used herein, "gRNA" and "guide RNA" are used interchangeably and refer to an RNA molecule that can form a complex with a CRISPR nuclease and can target the complex to a target sequence due to a certain complementarity with the target sequence. For example, for the Cas9 nuclease, its gRNA is generally composed of crRNA and tracrRNA molecules that are partially complementary to form a complex, wherein the crRNA comprises a sequence that has sufficient homology to the target sequence so as to hybridize with the complementary sequence of the target sequence and guide the CRISPR complex (Cas9+crRNA+tracrRNA) to specifically bind to the target sequence sequence. However, it is known in the art that single guide RNA (sgRNA) can be designed, which includes the features of crRNA and tracrRNA at the same time. And in the case of the Cpf1 nuclease, gRNA is generally composed of mature crRNA molecules only. It is within the capabilities of those skilled in the art to design a suitable gRNA based on the CRISPR nuclease used and the target sequence to be edited.

[0148] In some embodiments, the sgRNA for Cas9 comprises the scaffold sequence shown in SEQ ID NO:17.

[0149] As used herein, a "target sequence" is a sequence that is complementary or identical (depending on the different CRISPR nucleases) to a guide sequence of approximately 20 nucleotides contained in a guide RNA. The guide RNA targets the target sequence by base pairing with the target sequence or its complementary strand. For the Cas9 nuclease, the target sequence it recognizes generally needs to contain a PAM sequence at the 3' end, such as 5'-NGG-3'. The preferred target sequence length for the Cas9 nuclease is 20 nucleotides. The target region described herein may contain one or more of the target sequences.

[0150] In the present invention, CRISPR nucleases such as Cas9 can form DNA double-strand breaks (DSBs) at specific positions in the target sequence (e.g., between -3 and -4 positions upstream of the PAM sequence). They can also guide cytosine deaminase or adenine deaminase or DNA glycosylase to act on the chain where the DNA target sequence is located, and form one or more apurinic or apyrimidinic sites (ap sites) upstream of the DSB, which can be recognized and removed by AP lyase. This process will cause the genomic DNA to form a 5' overhanging end, i.e., a 5' sticky end, and a blunt end. The composition and length of the 5' sticky end can be determined according to the composition of the target sequence and the possible AP site (determined based on the cytosine deaminase or adenine deaminase or DNA glycosylase used), and accordingly, a suitable donor DNA fragment can be designed.

[0151] In some embodiments, the target sequence comprises one or more C nucleotides between positions -1 and -17 upstream of the PAM sequence.

[0152] The DNA double-strand break resulting in at least one sticky end of the genomic DNA in the present invention can be generated by various methods.

[0153] For example, in some embodiments, the method comprises introducing into the cell:

[0154] i) a CRISPR nickase, and / or an expression construct comprising a nucleotide sequence encoding said CRISPR nickase;

[0155] ii) four guide RNAs and / or an expression construct comprising a nucleotide sequence encoding said four guide RNAs,

[0156] Two of the guide RNAs target different target sequences on the sense strand of genomic DNA, and the other two guide RNAs target different target sequences on the antisense strand of genomic DNA.

[0157] The DNA double-strand break is thereby generated at the target region in the genome of the cell.

[0158] In some embodiments, the CRISPR nickase is a Cas9 nickase. In some embodiments, the Cas9 nickase is derived from SpCas9 of Streptococcus pyogenes (S. pyogenes) and comprises at least the amino acid substitution H840A relative to wild-type SpCas9. In some embodiments, the Cas9 nickase is capable of forming a nick between the -3 nucleotide (the first nucleotide at the 5' end of the PAM sequence is the +1 position) and the -4 nucleotide of the PAM of the target sequence.

[0159] By selecting four appropriate target sequences and four appropriate guide RNAs in the target region, two sticky ends that cannot be completely complementary can be obtained on the genomic DNA.

[0160] In some embodiments, the donor DNA fragments are linear DNA fragments. In some embodiments, the phosphodiester bonds between one or more (e.g., 2, 3, 4, or 5) nucleotides at the 5' and / or 3' ends of each strand of the donor DNA fragments comprise thiolation modifications. In some embodiments, the 5' end of each strand of the donor DNA fragments comprises a phosphorylation modification. In some embodiments, the donor DNA fragments are artificially synthesized.

[0161] In some embodiments, the length of the exogenous nucleotide sequence to be inserted can be about 1 bp to about 1000 bp or longer, for example, about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, about 600 bp, about 700 bp, about 1000 bp or longer, or any value therebetween.

[0162] In some embodiments, the "donor DNA fragment containing the exogenous nucleotide sequence to be inserted" can be composed of multiple (e.g., 2, 3, 4 or more) DNA fragments, each of which contains at least one sticky end and can be connected to each other to form the "donor DNA fragment containing the exogenous nucleotide sequence to be inserted". For example, the introduction of the "donor DNA fragment containing the exogenous nucleotide sequence to be inserted" can be achieved by introducing the multiple DNA fragments, each of which contains at least one sticky end and can be connected to each other. Specific examples can be seen in Figure 11A. By using the multi-fragment splicing method, the problem of difficulty in preparing long-fragment donor DNA can be avoided, and the precise insertion efficiency is higher. The length of each of the multiple DNA fragments can be about 1 bp to about 1000 bp or longer, for example, about 10 bp, about 20 bp, about 30 bp, about 40 bp, about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp or longer, or any value therebetween. Preferably, the length of each of the multiple DNA fragments can be about 50 bp. In some embodiments, a DNA ligase such as T4 DNA ligase or an expression construct encoding a DNA ligase such as T4 DNA ligase may also be introduced. The DNA ligase can facilitate the ligation between DNA fragments (such as the plurality of DNA fragments).

[0163] In some embodiments, the exogenous nucleotide sequence is associated with a trait of the organism, such that insertion of the exogenous nucleotide sequence results in the organism having an altered (preferably improved) trait relative to a wild-type organism.

[0164] In some embodiments, the exogenous nucleotide sequence is used to replace an endogenous nucleotide sequence in the cell. In some embodiments, the exogenous nucleotide sequence comprises one or more nucleotide substitutions, deletions and / or additions relative to the endogenous nucleotide sequence to be replaced.

[0165] For example, peptide tags such as Flag tags, His tags, etc. can be accurately inserted into the coding sequence of the protein of interest in accordance with the reading frame to mark the protein of interest. Alternatively, an enhancer sequence or miRNA sequence can be inserted into the expression control region of the gene to improve or regulate gene expression. In some specific embodiments, for example, referring to the examples of the present application, a 66bp 3*Flag tag is fused to the nitrogen end or carbon end of the protein for labeling the protein, a 54bp translation enhancer ADHE or a 100bp translation enhancer GAGA is inserted into the gene promoter region to increase protein expression, or a 21bp miRNA binding site is inserted to regulate gene expression.

[0166] In the methods of the present invention, the donor DNA fragment, the enzyme, the fusion protein, and / or the expression construct can be introduced into cells by methods known in the art. Available introduction methods include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, and other viruses), gene gun technique, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, and the like.

[0167] Cells that can be used in the methods of the present invention can be derived from animals, plants, or microorganisms. For example, the animals can be vertebrates and invertebrates, such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats, chickens, ducks, geese, zebrafish, and the like. The plants include monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis. The microorganisms can be eukaryotic or prokaryotic, such as Escherichia coli, Agrobacterium, cyanobacteria, and yeast.

[0168] In some embodiments, the methods of the present invention are performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.

[0169] In other embodiments, the methods of the present invention can also be performed in vivo. For example, the cells are cells in an organism, and the donor DNA fragment, enzyme, fusion protein, and / or expression construct of the present invention can be introduced into the cells in vivo, for example, by viral or Agrobacterium-mediated methods.

[0170] 3. Test Kit

[0171] In another aspect, the present invention provides a kit for use in the above method of the present invention, the kit comprising

[0172] 1) an agent capable of generating a DNA double-strand break at a target site in the genome of a cell, the DNA double-strand break resulting in at least one sticky end of the genomic DNA;

[0173] 2) A donor DNA fragment comprising an exogenous nucleotide sequence to be inserted, wherein the donor DNA fragment comprises a sticky end complementary to at least one sticky end of the genomic DNA.

[0174] The DNA double-strand break resulting in at least one sticky end of the genomic DNA is as defined above. The donor DNA fragment is as defined above.

[0175] In some embodiments, the "donor DNA fragment comprising an exogenous nucleotide sequence to be inserted" can be a plurality of (e.g., 2, 3, 4 or more) DNA fragments, each of which comprises at least one sticky end and can be ligated to each other to form the donor DNA fragment.

[0176] In some embodiments, the agent is a sequence-specific nuclease as defined above and / or an expression construct comprising a nucleotide sequence encoding a sequence-specific nuclease.

[0177] In some embodiments, the agent is as defined above

[0178] i) sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases, and optionally uracil-DNA glycosylases (UDG), and / or expression constructs comprising nucleotide sequences encoding said sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases, and optionally uracil-DNA glycosylases (UDG); or

[0179] Sequence-specific nucleases such as CRISPR nuclease, adenine deaminase, AP lyase and hypoxanthine glycosylase, and / or expression constructs comprising nucleotide sequences encoding said sequence-specific nucleases such as CRISPR nuclease, adenine deaminase, AP lyase and hypoxanthine glycosylase; or

[0180] Sequence-specific nucleases such as CRISPR nucleases, DNA glycosylases, and AP lyases, and / or expression constructs comprising nucleotide sequences encoding such sequence-specific nucleases such as CRISPR nucleases, DNA glycosylases, and AP lyases;

[0181] and / or

[0182] If the sequence-specific nuclease is a CRISPR nuclease, ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding a guide RNA.

[0183] In some embodiments, the agent is i) a fusion protein and / or an expression construct comprising a nucleotide sequence encoding a fusion protein as defined above; and

[0184] If the sequence-specific nuclease is a CRISPR nuclease, ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding a guide RNA.

[0185] In some embodiments, the agent is as defined above

[0186] i) a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein;

[0187] ii) a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein; and

[0188] If the sequence-specific nuclease is a CRISPR nuclease, iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding a guide RNA.

[0189] In some embodiments, the agent is as defined above

[0190] i) sequence-specific nucleases such as CRISPR nucleases and exonucleases, and / or expression constructs comprising nucleotide sequences encoding sequence-specific nucleases such as CRISPR nucleases and exonucleases,

[0191] If the sequence-specific nuclease is a CRISPR nuclease, ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding a guide RNA.

[0192] In some embodiments, the agent is as defined above

[0193] i) a CRISPR nickase, and / or an expression construct comprising a nucleotide sequence encoding said CRISPR nickase;

[0194] ii) four guide RNAs and / or an expression construct comprising a nucleotide sequence encoding said four guide RNAs,

[0195] Two of the guide RNAs target different target sequences on the sense strand of genomic DNA, and the other two guide RNAs target different target sequences on the antisense strand of genomic DNA.

[0196] The kit may also include a label indicating the intended use and / or instructions for use of the contents of the kit. The term label includes any written or recorded material provided on or with the kit or otherwise associated with the kit. Example

[0197] The present invention will be further described in detail below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the present invention. The experimental methods in the following examples, for which detailed conditions are not specified, are generally performed according to conventional conditions such as those described in "Molecular Cloning Laboratory Manual" by Sambrook.J et al. (translated by Huang Peitang et al., Beijing: Science Press, 2002), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are calculated by weight. The experimental materials and reagents used in the following examples can be obtained from commercial sources unless otherwise specified.

[0198] Example 1: Knock-in of a translation enhancer fragment into the 5'UTR of the rice SLR1 gene

[0199] The MT1 system is constructed by comprising a polypeptide comprising cytosine deaminase APOBEC3A, a CRISPR nuclease Cas9, an AP lyase, and optionally a uracil-DNA glycosylase (UDG) and a sticky-end donor DNA; the MT2 system is constructed by comprising a polypeptide comprising cytosine deaminase APOBEC3Bctd, a CRISPR nuclease Cas9, an AP lyase, and optionally a uracil-DNA glycosylase (UDG) and a sticky-end donor DNA.

[0200] Using an in vitro synthesized DNA fragment as the donor DNA, a peptide containing CRISPR nuclease, cytosine deaminase, AP lyase, and optionally uracil-DNA glycosylase (UDG) was used to create cohesive ends. The translational enhancer fragment was then introduced into the 5'UTR region of the rice SLR1 gene by homologous ligation. PCR, electrophoresis, and DNA sequencing demonstrated efficient integration of the fragment into the target site. The detailed protocol is as follows.

[0201] In vitro synthesis of single-stranded oligonucleotide fragments:

[0202] Table 1

[0203] Where 5'P represents 5'-terminal phosphorylation, and * represents interbase phosphorothioate modification. The synthesized single-stranded oligonucleotide fragments were dissolved in water to 100 μM, then diluted to 10 μM in annealing buffer (10 mM Tris-Cl, 0.1 mM EDTA, 50 mM NaCl, pH 8.0). Annealing was performed using a PCR instrument to form double-stranded donor DNA (Figure 6).

[0204] The experimental group (with sticky ends) was formed by annealing SLR1-DOR-MT-F and SLR1-DOR-MT-R to form the SLR1-DOR-MT donor fragment; the control group (without sticky ends) was formed by annealing SLR1-DOR-Cas9-F and SLR1-DOR-MT-R to form the SLR1-DOR-Cas9 donor fragment.

[0205] A gRNA was designed targeting the 5'UTR region of the rice DRO1 gene, and the gRNA guide sequence was constructed into rice MT1 and MT2 vectors and CRISPR / Cas9 plasmids. The plasmids were prepared in vitro.

[0206] The partial sequence of the 5'UTR of the rice SLR1 gene is as follows:

[0207] The gRNA target sequence is underlined (reverse complement, PAM in italics), and the bold and framed bases between CC and GA are the MT cleavage sites. The bold and framed bases between CC are the Cas9 cleavage sites in the control group.

[0208] The expected sequence is as follows:

[0209] Expected sticky ends are underlined.

[0210] The MT plasmid and donor fragment SLR1-DOR-MT, the CRISPR / Cas9 plasmid and donor fragment SLR1-DOR-Cas9, and gold powder were mixed according to the optimized transduction system in Table 2. Rice callus pretreated with hypertonic medium for 4 hours was transformed using the Bio-Rad PDS-1000 tabletop gene gun operating manual. Using hygromycin as a selection marker, positive resistant calli were obtained after conventional tissue culture screening and further differentiated to obtain stably transformed plants.

[0211] Table 2

[0212] DNA was extracted from individual stable transgenic plants obtained through screening. PCR amplification was performed using primers SLR1-F: 5'-gctactactagttgcttgcctc-3' (SEQ ID NO: 23) and DOR-R2: 5'-ATAGATCGAAAGCTCAGGTTTTC-3' (SEQ ID NO: 24). As shown in Figure 5, when the donor DNA was inserted into the target site in the forward orientation, PCR amplified a 202 bp target fragment. Sequencing results showing precise insertion of the donor fragment are shown in Figure 5. Sequencing efficiency statistics showed that the MT1 system achieved a precise insertion efficiency of 23.5%, 4.1 times that of the SpCas9 system, which relies on NHEJ repair, and the MT2 system achieved a precise insertion efficiency of 33.6%, 5.8 times that of the SpCas9 system, which relies on NHEJ repair (Figure 8). Transgenic plants with the translation enhancer knocked into the 5' UTR of the rice SLR1 gene were significantly shorter than wild-type plants (Figure 7), indicating that the translation enhancer was successfully functional.

[0213] Example 2: Knocking in the miRNA396 recognition site fragment in the coding region of the rice NRT1.1B gene

[0214] Using an in vitro-synthesized DNA fragment containing the miRNA396 recognition site as donor DNA, a peptide containing CRISPR nuclease, cytosine deaminase, AP lyase, and optionally uracil-DNA glycosylase (UDG) was used to cleave the rice NRT1.1B gene coding region to create sticky ends. The miRNA396 recognition site fragment was then introduced into the rice NRT1.1B gene coding region using the sticky ends. PCR, electrophoresis, and DNA sequencing demonstrated that the fragment was efficiently integrated into the target site. The specific procedure is as follows.

[0215] In vitro synthesis of single-stranded oligonucleotide fragments:

[0216] Table 3

[0217]

[0218] Where, 5'P represents 5'-terminal phosphorylation, and * represents interbase phosphorothioate modification. The synthesized single-stranded oligonucleotide fragments were dissolved in water to 100 μM, then diluted to 10 μM in annealing buffer (10 mM Tris-Cl, 0.1 mM EDTA, 50 mM NaCl, pH 8.0). Annealing was performed using a PCR instrument to form double-stranded donor DNA (Figure 6). NRT1.1-DOR-MT served as the experimental group, and NRT1.1B-DOR-Cas9 served as the control group.

[0219] The experimental group (with sticky ends) was formed by annealing NRT1.1B-DOR-MT-F and NRT1.1B-DOR-MT-R to form the NRT1.1-DOR-MT donor fragment; the control group (without sticky ends) was formed by annealing NRT1.1B-DOR-Cas9-F and NRT1.1B-DOR-MT-R to form the NRT1.1B-DOR-Cas9 donor fragment.

[0220] A gRNA was designed targeting the 5'UTR region of the rice DRO1 gene, and the gRNA guide sequence was constructed into the rice MT1 vector, MT2 vector, and CRISPR / Cas9 plasmid. The plasmid was prepared in vitro.

[0221] The target coding region sequence of the rice NRT1.1B gene is as follows:

[0222] The gRNA target sequence is underlined (PAM in italics), and the bold bases between the TC and AT are the MT cleavage sites. The bold and boxed bases between the AT are the Cas9 cleavage sites in the control group.

[0223] According to the transformation method of Example 1, rice callus was transformed to obtain stable transgenic plants. DNA was extracted from individual plants and PCR amplification was performed using primers NRT1.1B-F: 5'-agctaggagtagagaacgagacatatac (SEQ ID NO: 29) and NRT1.1B-DOR-R2: 5'-ATCCACAGGCTTTCTTGAACG (SEQ ID NO: 30). As shown in Figure 5, when the donor DNA was knocked into the target site in the forward direction, a 342bp target fragment could be amplified by PCR. The sequencing results of the precise insertion of the donor fragment are shown in Figure 5. The efficiency statistics showed that the precise knock-in efficiency of the MT1 system was 25%, which was 6.3 times that of the SpCas9 system that relies on NHEJ repair, and the precise knock-in efficiency of the MT2 system was 28%, which was 7 times that of the SpCas9 system that relies on NHEJ repair (Figure 8).

[0224] Example 3: Comparison of the efficiency of MT1, MT2, and Cas9 systems for inserting sequences of different lengths

[0225] Using methods similar to those in Examples 1 or 2, the efficiency and accuracy of inserting sequences of varying lengths at different target sites using the MT1, MT2, and Cas9 systems were compared: a 21bp donor DNA fragment with a miRNA396 recognition site was inserted into the NRT1.1B target site; a 66bp donor DNA fragment with a 3×Flag tag was inserted into the ACTIN1 target site; a 21bp donor DNA fragment with a miRNA156 recognition site was inserted into the GRF1 target site; a 21bp donor DNA fragment with a miRNA396 recognition site was inserted into the IPA1 target site; a 64bp donor DNA fragment with a translation enhancer was inserted into the SOS1 target site; and a 64bp donor DNA fragment with a translation enhancer was inserted into the SLR1 target site. Based on the different sticky ends that can be generated by the IPA1 target site, different donor DNA fragments with corresponding sticky ends were designed. The target site and donor DNA fragment sequences are shown in Figure 6.

[0226] The results, shown in Figure 6, demonstrate that, compared to the Cas9 system, the MT1 and MT2 systems can efficiently and precisely insert DNA fragments of varying lengths at different target sites. As shown by the immunoblotting results in Figure 7, transgenic plants containing a 66 bp 3× Flag tag donor DNA fragment inserted into the C-terminal target site of the ACTIN1 gene successfully expressed the Flag tag.

[0227] Example 4: Creating sticky ends by dual-target targeting to insert exogenous sequences

[0228] Using the MT system, we targeted two targets to create sticky ends in genomic DNA, inserting a 21bp donor DNA fragment into the NRT1.1B and D53 genes. Based on the potential sticky ends generated by the target sites, different donor DNA fragments with corresponding sticky ends were designed. The target site and donor DNA fragment sequences are shown in Figure 9.

[0229] The results are shown in Figure 9, indicating that compared with the Cas9 system, the MT system can efficiently and accurately insert DNA fragments of different lengths through a dual-target strategy.

[0230] Example 5: Precise insertion of exogenous sequences using the adenine deaminase-based ABE-MT system

[0231] An ABE-MT system was constructed, consisting of adenine deaminase TadA, the CRISPR nuclease Cas9, AP lyase, and hypoxanthine glycosylase polypeptides, and a cohesive-ended donor DNA. Similar to the previous examples, the efficiency and accuracy of the ABE-MT system and Cas9 in inserting sequences of varying lengths at different target sites were compared. A 64-bp exogenous sequence was inserted into the SOS1 gene, and a 21-bp exogenous sequence was inserted into the NRT1.1B gene. The target site and donor DNA fragment sequences are shown in Figure 10.

[0232] The results are shown in Figure 10. The ABE-MT system can accurately insert DNA fragments of different lengths, but its efficiency is not as good as that of the Cas9 system.

[0233] Example 6: Insertion of a donor DNA fragment formed by splicing multiple fragments using the MT system

[0234] Longer donor DNA fragments are difficult to prepare efficiently due to low synthesis and annealing efficiencies. This example investigates the feasibility of using a MT system to create multiple DNA fragments that can be end-ligated to form donor DNA fragments. The specific design is shown in Figure 11A.

[0235] A 110 bp exogenous sequence was inserted into the NRT1.1B gene using a single DNA donor fragment through the MT system and Cas9 system. At the same time, a 110 bp exogenous sequence was inserted into the NRT1.1B gene using the MT system using two DNA fragments split into 53 bp and 57 bp that could be connected by sticky ends (with or without the introduction of T4 ligase).

[0236] The results are shown in FIG11B . The insertion efficiency of the spliced ​​two short fragments was significantly higher than that of a single long donor fragment. The introduction of T4 ligase further increased the insertion efficiency.

[0237] Sequence information involved in this application

[0238] References

[0239] 1. Wang, S. et al. Precise, predictable multi-nucleotide deletions in rice and wheat using APOBEC–Cas9. Nature Biotechnology 38, 1460-1465 (2020).

[0240] 2.Tong,H.et al.Programmable A-to-Y base editing by fusing an adenine base editor with an N-methylpurine DNA glycosylase.Nature Biotechnology 41,1080-1084(2023).

[0241] 3.Tong,H.et al.Programmabledeaminase-free base editors for G-to-Y conversion by engineeredglycosylase.National Science Review 10(2023).

[0242] 4.Cortellino,S.et al.Thymine DNA Glycosylase Is Essential for Active DNA Demethylation by LinkedDeamination-Base Excision Repair.Cel 146,67-79(2011).

[0243] 5.Nakamura,T.et al.Structure of the mammalian adenine DNA glycosylase MUTYH:insights into thebase excision repair pathway and cancer.Nucleic Acids Research 49,7154-7163(2021).

[0244] 6.Ponferrada-Marín,M.I.,Roldán-Arjona,T.&Ariza,R.R.ROS1 5-methylcytosine DNA glycosylase isa slow-turnover catalyst that initiates DNA demethylation in a distributive fashion.Nucleic AcidsResearch 37,4264-4274(2009).

[0245] 7.Chen,L.et al.Re-engineering the adenine deaminase TadA-8e for efficient and specific CRISPR-based cytosine base editing.Nature Biotechnology 41,663-672(2022).

Claims

1. A genome editing method for inserting an exogenous nucleotide sequence at a specific site in a cell genome, the method comprising: (a) generating a DNA double-strand break at a target region in the genome of the cell, the DNA double-strand break resulting in at least one sticky end of the genomic DNA; and (b) providing a donor DNA fragment comprising an exogenous nucleotide sequence to be inserted to the cell, wherein the donor DNA fragment comprises a sticky end that complementarily pairs with at least one sticky end of the genomic DNA.

2. The method of claim 1, wherein The DNA double-strand break results in a 3' sticky end and a blunt end of the genomic DNA, and the donor DNA fragment comprises a 3' sticky end complementary to the 3' sticky end of the genomic DNA and a blunt end; or The DNA double-strand break results in a 5' sticky end and a blunt end of the genomic DNA, and the donor DNA fragment comprises a 5' sticky end complementary to the 5' sticky end of the genomic DNA and a blunt end.

3. The method of claim 1, wherein the DNA double-strand break results in two sticky ends of the genomic DNA, and the two sticky ends cannot be completely complementary to each other. For example, when complementary alignment is achieved, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, or no more than 5% of the nucleotides at any sticky end of the genomic DNA can be paired with the nucleotides at the other sticky end of the genomic DNA.

4. The method of claim 3, wherein The DNA double-strand break results in two 3' sticky ends of the genomic DNA that cannot be completely complementary to each other, and the donor DNA fragment comprises two 3' sticky ends that are respectively complementary to the two 3' sticky ends of the genomic DNA; The DNA double-strand break results in two 5' sticky ends of the genomic DNA that cannot be complementary to each other, and the donor DNA fragment comprises two 5' sticky ends that are complementary to the two 5' sticky ends of the genomic DNA, respectively; or The DNA double-strand break results in a 5' sticky end and a 3' sticky end of the genomic DNA, and the donor DNA fragment comprises a 5' sticky end complementary to the 5' sticky end of the genomic DNA and a 3' sticky end complementary to the 3' sticky end of the genomic DNA.

5. The method of any one of claims 1 to 4, wherein there are no more than 5, no more than 4, no more than 3, no more than 2, or no more than 1 mismatch between the sticky ends of the genomic DNA and the corresponding sticky ends of the donor DNA fragments. Preferably, the sticky ends of the genomic DNA and the corresponding sticky ends of the donor DNA fragments are perfectly complementary.

6. The method of any one of claims 1 to 5, wherein the sticky ends are about 2 to about 50 nucleotides in length, preferably 8 or 9 nucleotides in length.

7. The method of any one of claims 1 to 6, wherein the method generates the DNA double-strand break at a target region in the genome of the cell by introducing into the cell a sequence-specific nuclease and / or an expression construct comprising a nucleotide sequence encoding a sequence-specific nuclease.

8. The method of any one of claims 1 to 7, wherein a) The method comprises introducing into the cell: sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases and optionally uracil-DNA glycosylases (UDG), and / or expression constructs comprising nucleotide sequences encoding said sequence-specific nucleases such as CRISPR nucleases, cytosine deaminases, AP lyases and optionally uracil-DNA glycosylases (UDG), thereby generating the DNA double-strand break at the target region in the genome of the cell; or, b) the method comprises introducing into the cell: Fusion proteins and / or expression constructs comprising nucleotide sequences encoding fusion proteins; wherein the fusion protein comprises cytosine deaminase, a sequence-specific nuclease such as a CRISPR nuclease, an AP lyase, and optionally a uracil-DNA glycosylase (UDG), thereby generating the DNA double-strand break at the target region in the genome of the cell; or c) the method comprises introducing into the cell: a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein; and a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein; wherein the first fusion protein comprises cytosine deaminase, a sequence-specific nuclease such as a CRISPR nuclease, and optionally a uracil-DNA glycosylase (UDG), and the second fusion protein comprises an AP lyase, The DNA double-strand break is thereby generated at the target region in the genome of the cell.

9. The method of claim 8, wherein the cytosine deaminase is selected from the group consisting of APOBEC1 deaminase, activation-induced cytidine deaminase (AID), APOBEC3G, CDA1, human APOBEC3A deaminase, truncated APOBEC3B deaminase.

10. The method of claim 9, wherein the cytosine deaminase is a human APOBEC3A deaminase, e.g., comprising the amino acid sequence of SEQ ID NO:2, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:2, or having one or more conservative amino acid substitutions relative to SEQ ID NO:2; or the cytosine deaminase is a truncated APOBEC3B deaminase, e.g., comprising the amino acid sequence of SEQ ID NO:3, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:3, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

3.

11. The method of any one of claims 8-10, wherein the UDG comprises the amino acid sequence of SEQ ID NO:4, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:4, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

4.

12. The method of any one of claims 1 to 7, wherein a) The method comprises introducing into the cell: Sequence-specific nucleases such as CRISPR nucleases, adenine deaminase, AP lyase and hypoxanthine glycosylase, and / or sequences encoding the CRISPR nucleases, adenine deaminase, AP lyase and hypoxanthine glycosylase An expression construct of a nucleotide sequence of an enzyme, thereby generating the DNA double-strand break at the target region in the genome of the cell; or b) the method comprises introducing into the cell: Fusion proteins and / or expression constructs comprising nucleotide sequences encoding fusion proteins; wherein the fusion protein comprises adenine deaminase, sequence-specific nuclease such as CRISPR nuclease, AP lyase and hypoxanthine glycosylase, thereby generating the DNA double-strand break at the target region in the genome of the cell; or c) the method comprises introducing into the cell: a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein; and a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein; wherein the first fusion protein comprises adenine deaminase, a sequence-specific nuclease such as CRISPR nuclease and hypoxanthine glycosylase, and the second fusion protein comprises AP lyase, The DNA double-strand break is thereby generated at the target region in the genome of the cell.

13. The method of claim 12, wherein the adenine deaminase is Escherichia coli adenine deaminase TadA, for example, comprising the amino acid sequence shown in SEQ ID NO:6, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:6, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

6.

14. The method of claim 13, wherein the hypoxanthine glycosylase is N-methylpurine DNA glycosylase (MPG), for example, comprising the amino acid sequence shown in SEQ ID NO:7, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:7, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

7.

15. The method of any one of claims 1 to 7, wherein a) The method comprises introducing into the cell: sequence-specific nucleases such as CRISPR nucleases, DNA glycosylases, and AP lyases, and / or expression constructs comprising nucleotide sequences encoding said CRISPR nucleases, DNA glycosylases, and AP lyases, thereby generating the DNA double-strand break at the target region in the genome of the cell; or b) the method comprises introducing into the cell: Fusion proteins and / or expression constructs comprising nucleotide sequences encoding fusion proteins; wherein the fusion protein comprises a DNA glycosylase, a sequence-specific nuclease such as a CRISPR nuclease and an AP lyase, thereby generating the DNA double-strand break at the target region in the genome of the cell; or c) the method comprises introducing into the cell: a first fusion protein and / or an expression construct comprising a nucleotide sequence encoding the first fusion protein, and a second fusion protein and / or an expression construct comprising a nucleotide sequence encoding the second fusion protein, wherein the first fusion protein comprises a DNA glycosylase and a sequence-specific nuclease such as a CRISPR nuclease, and the second fusion protein comprises an AP lyase, The DNA double-strand break is thereby generated at the target region in the genome of the cell.

16. The method of claim 15, wherein the DNA glycosylase is guanine glycosylase, thymine glycosylase, adenine glycosylase, or cytosine glycosylase.

17. The method of claim 16, wherein The guanine glycosylase is selected from: N-methylpurine DNA glycosylase (MPG), which comprises the amino acid sequence of SEQ ID NO:7, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:7, or having one or more conservative amino acid substitutions relative to SEQ ID NO:7; 8-oxoguanine DNA glycosylase (OGG1), which comprises the amino acid sequence of SEQ ID NO:8, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:8, or having one or more conservative amino acid substitutions relative to SEQ ID NO:8; mutY DNA glycosylase (MUTYH), which comprises the amino acid sequence of SEQ ID NO:9, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:9, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

9. NO:9 having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or has one or more conservative amino acid substitutions relative to SEQ ID NO:9; The thymine glycosylase is thymine glycosylase TDG, which comprises the amino acid sequence of SEQ ID NO: 10, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO: 10, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 10; The adenine glycosylase is a mutY DNA glycosylase (MUTYH) comprising the amino acid sequence of SEQ ID NO:9, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:9, or having one or more conservative amino acid substitutions relative to SEQ ID NO:9; or The cytosine glycosylase is 5-methylcytosine glycosylase ROS1, which comprises the amino acid sequence shown in SEQ ID NO:11, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO:11, or has one or more conservative amino acid substitutions relative to SEQ ID NO:

11.

18. The method of any one of claims 8-17, wherein the AP lyase comprises the amino acid sequence of SEQ ID NO:5, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:5, or having one or more conservative amino acid substitutions relative to SEQ ID NO:

5.

19. The method of any one of claims 1 to 7, wherein a) The method comprises introducing into the cell: sequence-specific nucleases such as CRISPR nucleases and exonucleases, and / or expression constructs comprising nucleotide sequences encoding sequence-specific nucleases such as CRISPR nucleases and exonucleases, thereby generating the DNA double-strand break at the target region in the genome of the cell; or b) the method comprises introducing into the cell: A fusion protein, and / or an expression construct comprising a nucleotide sequence encoding the fusion protein, wherein the fusion protein comprises a sequence-specific nuclease such as a CRISPR nuclease and an exonuclease, The DNA double-strand break is thereby generated at the target region in the genome of the cell.

20. The method of claim 19, wherein the exonuclease is a 3'→5' exonuclease or a 5'→3' exonuclease.

21. The method of claim 20, wherein The "5'→3' exonuclease" is selected from: a lambda exonuclease comprising the amino acid sequence of SEQ ID NO: 12, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO: 12, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 12; a T5 exonuclease comprising the amino acid sequence of SEQ ID NO: 13, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO: 13, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 13; a recE exonuclease comprising the amino acid sequence of SEQ ID NO: 14, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO: 14, or having one or more conservative amino acid substitutions relative to SEQ ID NO: 14; NO:14 has an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or has one or more conservative amino acid substitutions relative to SEQ ID NO:14; or The "3'→5' exonuclease" is selected from: exonuclease II, which comprises the amino acid sequence shown in SEQ ID NO:15, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:15, or has one or more conservative amino acid substitutions relative to SEQ ID NO:15; exonuclease III, which comprises the amino acid sequence shown in SEQ ID NO:16, or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:16, or has one or more conservative amino acid substitutions relative to SEQ ID NO:

16.

22. The method of any one of claims 7-21, wherein the sequence-specific nuclease is selected from a ZFN, a TALEN, a restriction endonuclease, or a CRISPR nuclease.

23. The method of claim 22, wherein the sequence-specific nuclease is a CRISPR nuclease, the method further comprising introducing into the cell ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA is capable of targeting the CRISPR nuclease or a fusion protein comprising the CRISPR nuclease to a target region in the cell genome.

24. The method of any one of claims 22-23, wherein the CRISPR nuclease is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, GSU0054, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx11 、Csx16、CsaX、Csx3、Csx1、Csx15、Csf1、Csf2、Csf3、Csf4、C2c3、C2c8、C2c10、Cas12a、Cas12a2、Cas12b (also known as C2c1), Cas12c、Cas12c1、Cas12e、Cas12g、Cas12h、Cas12i、Cas12j、Cas12f、Cas12k、Cas12m、Cas12n、Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas13m.3, Cas13m.6, Cas14, Casφ, Casλ, TnpB, Preferably, the CRISPR nuclease is a Cas9 nuclease, for example, the Cas9 nuclease has the amino acid sequence shown in SEQ ID NO:1, or comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:1, or has one or more conservative amino acid substitutions relative to SEQ ID NO:

1.

25. The method of any one of claims 1-7, wherein the method comprises introducing into the cell: i) a CRISPR nickase, and / or an expression construct comprising a nucleotide sequence encoding the CRISPR nickase; ii) 4 guide RNAs and / or an expression construct comprising a nucleotide sequence encoding the 4 guide RNAs, Two of the guide RNAs target different target sequences on the sense strand of genomic DNA, and the other two guide RNAs target different target sequences on the antisense strand of genomic DNA. The DNA double-strand break is thereby generated at the target region in the genome of the cell.

26. The method of claim 25, wherein the CRISPR nickase is a Cas9 nickase, such as the Cas9 nickase derived from SpCas9 of Streptococcus pyogenes (S. pyogenes) and comprising at least the amino acid substitution H840A relative to wild-type SpCas9.

27. The method of any one of claims 1-26, wherein the donor DNA fragments are linear DNA fragments.

28. The method of any one of claims 1 to 27, wherein the phosphodiester bonds between one or more nucleotides at the 5' and / or 3' ends of each strand of the donor DNA fragment comprise thioate modifications.

29. The method of any one of claims 1-28, wherein the 5' end of each strand of the donor DNA fragment comprises a phosphorylation modification.

30. The method according to any one of claims 1 to 28, wherein the length of the exogenous nucleotide sequence to be inserted may be about 1 bp to about 1000 bp or longer.

31. The method according to any one of claims 1 to 30, wherein the donor DNA fragment comprising the exogenous nucleotide sequence to be inserted comprises a plurality of DNA fragments, each of the plurality of DNA fragments comprising at least one sticky end and capable of being ligated to each other to form the donor DNA fragment comprising the exogenous nucleotide sequence to be inserted.

32. The method of any one of claims 1 to 31, wherein the exogenous nucleotide sequence is associated with a trait of the cell or an organism derived therefrom, whereby insertion of the exogenous nucleotide sequence results in the cell or an organism derived therefrom having an altered (preferably improved) trait relative to a wild-type cell or an organism derived therefrom.

33. The method of any one of claims 1 to 32, wherein the exogenous nucleotide sequence is used to replace an endogenous nucleotide sequence in the cell.

34. The method of claim 33, wherein the exogenous nucleotide sequence comprises one or more nucleotide substitutions, deletions and / or additions relative to the endogenous nucleotide sequence to be replaced.

35. The method of any one of claims 1-34, wherein the cell is from an animal, a plant or a microorganism, for example, the animal is a vertebrate and an invertebrate, such as a human, a mouse, a rat, a monkey, a dog, a pig, a sheep, a cow, a cat, a chicken, a duck, a goose, a zebrafish; the plant includes a monocotyledonous plant and a dicotyledonous plant, for example, rice, corn, wheat, sorghum, barley, soybean, peanut, Arabidopsis; the microorganism is a eukaryotic microorganism or a prokaryotic microorganism, for example, Escherichia coli, Agrobacterium, cyanobacteria, yeast.