Improved gene editing system

By combining CRISPR nuclease, cytosine deaminase, AP lyase and uracil-DNA glycosylase, combined with guide RNA, the problem of difficulty in achieving short genome fragment deletion in the prior art is solved, and efficient, accurate and predictable gene editing effects are achieved.

CN119932089APending Publication Date: 2025-05-06SUZHOU QI BIODESIGN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510113127.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-05-07
Filing Date
2020-05-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing gene editing technologies are difficult to achieve efficient, accurate and predictable short fragment deletion of eukaryotic cell genomes.

Method used

By combining CRISPR nuclease, cytosine deaminase, AP lyase and uracil-DNA glycosylase, guide RNA guides to target the cell genome, achieving precise deletion from double-strand break sites within the target sequence to specific C nucleotide sites.

Benefits of technology

Efficient, accurate and predictable short fragment deletion are achieved, significantly improving the frequency and accuracy of polynucleotide deletion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119932089A_ABST
    Figure CN119932089A_ABST
Patent Text Reader

Abstract

The invention provides a gene editing system for editing a target gene in the genome of a cell. The gene editing system comprises CRISPR nuclease, cytosine deaminase, AP lyase, guide RNA and optional uracil-DNA glycosylase. The invention also provides a method for producing a genetically modified cell, and a kit comprising the gene editing system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application with application number 202080034110.3, application date May 7, 2020, and invention name “Improved Gene Editing System”. Technical Field

[0002] The present invention relates to the field of genetic engineering. Specifically, the present invention relates to an improved gene editing system. More specifically, the present invention relates to a gene editing system capable of accurately editing the genome of eukaryotic cells, especially predictable and accurate polynucleotide deletion. Background of the Invention

[0004] In recent years, with the continuous development of genome editing technology, a large number of gene editing tools have been developed, improved and applied, including gene knockout tools mediated by SpCas9 to single-base editing tools mediated by nCas9 (D10A) fused with cytosine deaminase, etc. Under the guidance of guide RNA, SpCas9 binds and cuts double-stranded DNA to form double-strand breaks (DSBs), which often introduce insertions and / or deletions of different fragment lengths during the repair process of the body, but these insertions and / or deletions are random, imprecise and unpredictable (Wang et al., 2014; Zhang et al., 2016). et al. (2017) used Cas9 fused to 3' repair exonuclease 2 (Trex2), which significantly increased the frequency of deletion mutations and the deletion fragments were longer, but the mutation type was still imprecise and unpredictable; using a pair of sgRNAs for targeted deletion can obtain specific long-fragment deletions, but it will also produce inversions, small fragments InDel, etc., which also greatly reduces the efficiency of the former ( et al., 2017). In order to obtain precise fragment deletion, Wolfs et al. (2016) fused Cas9 with TevI nuclease, which recognizes the restriction site and cuts the double-stranded DNA. The cut and the DSB cut by Cas9 together form a 33-36bp deletion, but the efficiency of the system is low due to the restriction of the restriction site. So far, no tool has been developed that can perform efficient, precise and predictable short fragment deletion within the protospacer sequence.

[0005] Therefore, there is still a need in the art for a gene editing system that can accurately edit the genome of eukaryotic cells, especially predictable and accurate polynucleotide deletions. Brief description of the invention

[0007] In one aspect, the present invention provides a gene editing system for editing a target sequence in a cell genome, comprising:

[0008] i) a first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide;

[0009] ii) a second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide; and

[0010] iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA,

[0011] The first polypeptide comprises a CRISPR nuclease, a cytosine deaminase, and optionally a uracil-DNA glycosylase (UDG), and the second polypeptide comprises an AP lyase, wherein the guide RNA is capable of targeting the first polypeptide to a target sequence in the genome of a cell.

[0012] In one aspect, the present invention provides a gene editing system for editing a target sequence in a cell genome, comprising:

[0013] i) a polypeptide and / or an expression construct comprising a nucleotide sequence encoding the polypeptide; and

[0014] ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA,

[0015] The polypeptide comprises a CRISPR nuclease, a cytosine deaminase, an AP lyase, and optionally a uracil-DNA glycosylase (UDG), wherein the guide RNA is capable of targeting the polypeptide to a target sequence in the genome of a cell.

[0016] In one aspect, the invention provides a method for producing a genetically modified cell, comprising introducing the gene editing system of the invention into a cell.

[0017] In one aspect, the invention provides a kit comprising the gene editing system of the invention, and instructions for use. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 .Shows the working mode of the ACD system.

[0020] Figure 2 .Shows a comparative analysis of InDel production efficiency between SpCas9 and ACD systems at different target sites.

[0021] Figure 3 The graph shows the types and efficiency of deletion mutations formed by the ACD system at the sgF3HT4 site.

[0022] Figure 4 Shown are the types and efficiency of Deletion mutations formed by the ACD system at the sgLART4 site.

[0023] Figure 5 Shown are the types and efficiency of deletion mutations formed by the ACD system at the sgMYBT2 site.

[0024] Figure 6 Shown are the types and efficiency of Deletion mutations formed by the ACD system at the sgPMKT1 site.

[0025] Figure 7 The graph shows the types and efficiency of deletion mutations formed by the ACD system at the sgVRN1T1 site.

[0026] Figure 8 The figure shows the types and efficiency of deletion mutations formed by the ACD system at the sgGS6T2 site.

[0027] Fig. 9 .Shows the differences in deamination activity and deamination window of different cytosine deaminases.

[0028] Fig.10 . Schematic diagram showing the vector construction of two different types of AFID systems.

[0029] Fig.11 .Shows the deletion efficiency of Cas9, AFID-3, and eAFID-3 at different endogenous target sites in rice.

[0030] Fig.12 .Shows the deletion efficiency of Cas9, AFID-3, and eAFID-3 at different endogenous targets in wheat.

[0031] Fig.13 .Shows the types and proportions of deletion mutations of AFID-3 and eAFID-3 at endogenous target sites in rice.

[0032] Fig.14 .Shows the types and proportions of deletion mutations of AFID-3 and eAFID-3 at endogenous target sites in wheat.

[0033] Fig.15 . Shows the preference of AFID-3 and eAFID-3 for the cytosine base that predicts the start of a deletion.

[0034] Fig.16 .Shows the mutation types and their proportions of Cas9, AFID-3, and eAFID-3 producing the desired predictable in-frame deletions at the miR396h binding site of the rice OsGRF1 gene and the miR156 binding site of the OsIPA1 gene, respectively.

[0035] Fig.17 .Shows a schematic diagram of the construction of AFID-3 vector for rice Agrobacterium infection.

[0036] Fig.18 .Shows the types of regenerated plant mutants produced by Cas9 and AFID-3 on the rice OsCDC48 gene. DETAILED DESCRIPTION OF THE INVENTION

[0038] I. Definition

[0039] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. In addition, the protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology related terms and laboratory operation procedures used herein are terms and routine procedures widely used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are more fully described in the following documents: Sambrook, J., Fritsch, EF and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). At the same time, in order to better understand the present invention, the definitions and explanations of the relevant terms are provided below.

[0040] As used herein, the term "and / or" encompasses all combinations of items connected by the term, and each combination should be considered to have been listed separately herein. For example, "A and / or B" encompasses "A," "A and B," and "B." For example, "A, B, and / or C" encompasses "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."

[0041] When the term "comprising" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present invention. In addition, it is clear to those skilled in the art that the methionine encoded by the start codon at the N-terminus of the polypeptide may be retained in certain practical situations (for example, when expressed in a specific expression system), but it does not substantially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of this application, although it may not contain a methionine encoded by a start codon at the N-terminus, a sequence containing the methionine is also covered at this time, and accordingly, its encoding nucleotide sequence may also contain a start codon; and vice versa.

[0042] "Genome" as used herein encompasses not only the chromosomal DNA present in the cell nucleus, but also the organellar DNA present in subcellular components of the cell (eg, mitochondria, plastids).

[0043] As used herein, "organism" includes any organism suitable for genome editing, preferably a eukaryotic organism. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants including monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, etc.

[0044] "Genetically modified organism" or "genetically modified cell" means an organism or cell that contains an exogenous polynucleotide or a modified gene or expression control sequence in its genome. For example, the exogenous polynucleotide can be stably integrated into the genome of the organism or cell and inherited for consecutive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. The modified gene or expression control sequence is a sequence in the genome of the organism or cell that contains single or multiple deoxynucleotide substitutions, deletions and additions.

[0045] "Exogenous" with respect to a sequence refers to a sequence that is from a foreign species, or, if from the same species, a sequence that has been significantly altered in composition and / or locus from its native form through deliberate human intervention.

[0046] "Polynucleotide", "nucleic acid sequence", "nucleotide sequence" or "nucleic acid fragment" are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers that optionally may contain synthetic, non-natural or altered nucleotide bases. Nucleotides are referred to by their single letter names as follows: "A" is adenosine or deoxyadenosine (RNA or DNA, respectively), "C" represents cytidine or deoxycytidine, "G" represents guanosine or deoxyguanosine, "U" represents uridine, "T" represents deoxythymidine, "R" represents purine (A or G), "Y" represents pyrimidine (C or T), "K" represents G or T, "H" represents A or C or T, "I" represents inosine, and "N" represents any nucleotide.

[0047] "Polypeptide", "peptide", and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide", "peptide", "amino acid sequence" and "protein" may also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma carboxylation of glutamic acid residues, hydroxylation and ADP-ribosylation.

[0048] Sequence "identity" has a meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are many methods to measure the identity between two polynucleotides or polypeptides, the term "identity" is well known to those of skill (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).

[0049] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art, and generally can be carried out without changing the biological activity of the resulting molecule. Generally, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially change biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).

[0050] As used herein, "expression construct" refers to a vector such as a recombinant vector suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (such as transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.

[0051] The "expression construct" of the present invention can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be a translatable RNA (such as mRNA).

[0052] An "expression construct" of the present invention may comprise regulatory sequences and a nucleotide sequence of interest from different sources, or regulatory sequences and a nucleotide sequence of interest from the same source but arranged in a manner different from that normally found in nature.

[0053] "Regulatory sequence" and "regulatory element" are used interchangeably and refer to nucleotide sequences located upstream (5' non-coding sequence), in the middle or downstream (3' non-coding sequence) of a coding sequence and affecting the transcription, RNA processing or stability or translation of the relevant coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns and polyadenylation recognition sequences.

[0054] "Promoter" refers to a nucleic acid fragment that can control the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter that can control the transcription of a gene in a cell, whether or not it is derived from the cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally regulated promoter or an inducible promoter.

[0055] "Constitutive promoter" refers to a promoter that will generally cause a gene to be expressed in most cell types under most circumstances. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed primarily, but not necessarily exclusively, in one tissue or organ, and may also be expressed in one specific cell or cell type. "Developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. "Inducible promoter" selectively expresses an operably linked DNA sequence in response to endogenous or exogenous stimuli (environmental, hormonal, chemical signals, etc.).

[0056] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II or pol III promoters. Examples of pol I promoters include chicken RNA pol I promoters. Examples of pol II promoters include, but are not limited to, cytomegalovirus immediate early (CMV) promoters, Rous sarcoma virus long terminal repeat (RSV-LTR) promoters, and simian virus 40 (SV40) immediate early promoters. Examples of pol III promoters include U6 and H1 promoters. Inducible promoters such as metallothionein promoters can be used. Other examples of promoters include T7 phage promoters, T3 phage promoters, β-galactosidase promoters, and Sp6 phage promoters. When used for plants, the promoter can be a cauliflower mosaic virus 35S promoter, a corn Ubi-1 promoter, a wheat U6 promoter, a rice U3 promoter, a corn U3 promoter, a rice actin promoter.

[0057] As used herein, the term "operably linked" refers to the connection of a regulatory element (e.g., but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (e.g., a coding sequence or an open reading frame) such that transcription of the nucleotide sequence is controlled and regulated by the transcription regulatory element. Techniques for operably linking a regulatory element region to a nucleic acid molecule are known in the art.

[0058] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism refers to transforming an organism cell with the nucleic acid or protein so that the nucleic acid or protein can function in the cell. "Transformation" as used in the present invention includes stable transformation and transient transformation.

[0059] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into the genome, resulting in stable inheritance of the exogenous gene. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof.

[0060] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform its function without the foreign gene being stably inherited. In transient transformation, the foreign nucleic acid sequence is not integrated into the genome.

[0061] 2. Improved gene editing system

[0062] The inventors surprisingly found that by targeting CRISPR nuclease to a target sequence in the cell genome through a guide RNA to form a double-strand break (DSB), and by converting C in the target sequence or its complementary sequence into U through a cytosine deaminase fused to the CRISPR nuclease, and then through the combined action of endogenous or exogenous uracil-DNA glycosylase (UDG) and AP lyase, precise deletion from the DSB site in the target sequence to the C nucleotide site can be achieved.

[0063] Therefore, the present invention provides a gene editing system for editing a target sequence in a cell genome, comprising:

[0064] i) a first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide;

[0065] ii) a second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide; and

[0066] iii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA,

[0067] Wherein the first polypeptide comprises cytosine deaminase, CRISPR nuclease and optionally uracil-DNA glycosylase (UDG), and the second polypeptide comprises AP lyase, wherein the guide RNA is capable of targeting the first polypeptide to a target sequence in the cell genome. In some embodiments, the expression construct comprising a nucleotide sequence encoding the first polypeptide, the expression construct encoding the second polypeptide, and / or the expression construct comprising a nucleotide sequence encoding the guide RNA can be different expression constructs, or any two or all of them are the same expression construct. In some embodiments, the first polypeptide is an isolated polypeptide, the second polypeptide is an isolated polypeptide, and / or the guide RNA is an isolated RNA.

[0068] As used herein, "gene editing system" refers to a combination of components required for gene editing of a genome in a cell. The individual components of the system, such as polypeptides, gRNAs, etc., may exist independently of each other, or may exist in any combination as a composition.

[0069] In some embodiments, the gene editing system comprises at least an expression construct comprising a nucleotide sequence encoding the first polypeptide, a nucleotide sequence encoding a self-cleaving peptide, and a nucleotide sequence encoding the second polypeptide connected in frame. In some embodiments, the nucleotide sequence encoding the first polypeptide, the nucleotide sequence encoding the self-cleaving peptide, and the nucleotide sequence encoding the second polypeptide are arranged in a 5' to 3' direction.

[0070] As used herein, "self-cleaving peptide" means a peptide that can achieve self-cleavage in a cell. For example, the self-cleaving peptide may include a protease recognition site, thereby being recognized and specifically cleaved by a protease in the cell.

[0071] Alternatively, the self-cleaving peptide can be a 2A polypeptide. 2A polypeptides are a class of short peptides from viruses, and their self-cleavage occurs during translation. When two different target polypeptides are expressed in the same reading frame using a 2A polypeptide, two target polypeptides are generated at almost a 1:1 ratio. Commonly used 2A polypeptides can be P2A from porcine techovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitisA virus, and F2A from foot-and-mouth disease virus. Among them, P2A has the highest cleavage efficiency and is therefore preferred. A variety of functional variants of these 2A polypeptides are also known in the art, and these variants can also be used in the present invention. In some embodiments, the self-cleaving peptide is P2A shown in SEQ ID NO:9.

[0072] In some embodiments, the gene editing system comprises at least an expression construct comprising a nucleotide sequence encoding the amino acid sequence shown in SEQ ID NO:10 or SEQ ID NO:11.

[0073] In another aspect, the present invention also provides a gene editing system for editing a target sequence in a cell genome, comprising:

[0074] i) a polypeptide and / or an expression construct comprising a nucleotide sequence encoding the polypeptide; and

[0075] ii) a guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA,

[0076] Wherein the polypeptide comprises cytosine deaminase, CRISPR nuclease, AP lyase and optionally uracil-DNA glycosylase (UDG), wherein the guide RNA is capable of targeting the polypeptide to a target sequence in the cell genome. In some embodiments, the expression construct comprising a nucleotide sequence encoding a polypeptide and the expression construct comprising a nucleotide sequence encoding a guide RNA may be different expression constructs or may be the same expression construct. In some embodiments, the polypeptide is an isolated polypeptide and / or the guide RNA is an isolated RNA. In some embodiments, the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 10 or SEQ ID NO: 11.

[0077] As used herein, the term "CRISPR nuclease" generally refers to a nuclease present in a naturally occurring CRISPR system, as well as a modified form thereof, a variant thereof, or a catalytically active fragment thereof. CRISPR nucleases can recognize, bind to and / or cut a target nucleic acid structure by interacting with a guide RNA. The term encompasses any nuclease or a functional variant thereof that is capable of achieving gene editing in a cell based on a CRISPR system. In some embodiments, the functional variant retains its double-strand cleavage activity, i.e., the ability to form a double-strand break (DSB) in a target sequence.

[0078] The CRISPR nuclease used in the gene editing system of the present invention can be selected from, for example, Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Cas9, Csn2, Cas4, Cpf1, C2c1, C2c3 or C2c2 proteins, or functional variants of these nucleases.

[0079] In some embodiments, the CRISPR nuclease includes a Cas9 nuclease or a variant thereof. The Cas9 nuclease can be a Cas9 nuclease from a different species, such as spCas9 from Streptococcus pyogenes (S. pyogenes). The Cas9 nuclease variant can, for example, include a highly specific variant of the Cas9 nuclease, such as the Cas9 nuclease variants eSpCas9 (1.0) (K810A / K1003A / R1060A) and eSpCas9 (1.1) (K848A / K1003A / R1060A) of Feng Zhang et al., and the Cas9 nuclease variant SpCas9-HF1 (N497A / R661A / Q695A / Q926A) developed by J. Keith Joung et al. In some specific embodiments, the CRISPR nuclease has an amino acid sequence shown in SEQ ID NO: 1. In some embodiments, the CRISPR nuclease comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:1, or having one or more conservative amino acid substitutions relative to SEQ ID NO:1.

[0080] In some embodiments, the CRISPR nuclease may also include a Cpf1 nuclease or a variant thereof, such as a high-specificity variant. The Cpf1 nuclease may be a Cpf1 nuclease from a different species, such as a Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6, and Lachnospiraceae bacterium ND2006.

[0081] As used herein, the "cytosine deaminase" refers to a deaminase that can accept single-stranded DNA as a substrate and can catalyze the deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively. Examples of cytosine deaminases include, but are not limited to, for example, APOBEC1 deaminase, activation-induced cytidine deaminase (AID), APOBEC3G, CDA1, human APOBEC3A deaminase, and truncated APOBEC3B deaminase. In some embodiments, the cytosine deaminase is human APOBEC3A deaminase, for example, its amino acid sequence is shown in SEQ ID NO: 2. In some specific embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO: 2, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 2. In some embodiments, the cytosine deaminase is a truncated APOBEC3B deaminase (APOBEC3Bctd), for example, whose amino acid sequence is shown in SEQ ID NO: 7. In some specific embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO: 7, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 7.

[0082] As used herein, uracil-DNA glycosylase (UDG) or uracil-N-glycosylase (UNG) refers to an enzyme that can recognize the U base and remove the N-glycosidic bond of the base to form an apurinic or apyrimidinic site. The UDG may have different sources, such as from Escherichia coli. In some specific embodiments, UDG has the amino acid sequence shown in SEQ ID NO: 3. In some specific embodiments, UDG comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO: 3, or has one or more conservative amino acid substitutions relative to SEQ ID NO: 3.

[0083] "AP lyase", "AP lyase", AP endonuclease and "apurinic pyrimidine lyase" are used interchangeably herein and refer to an enzyme that can recognize apurinic or apyrimidinic sites on nucleic acids and cleave nucleic acids. The AP lyase can be from different sources, such as from Escherichia coli. In some specific embodiments, the AP lyase has the amino acid sequence shown in SEQ ID NO:4. In some specific embodiments, the AP lyase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:4, or has one or more conservative amino acid substitutions relative to SEQ ID NO:4.

[0084] As used herein, "gRNA" and "guide RNA" are used interchangeably and refer to RNA molecules that can form a complex with a CRISPR nuclease and can target the complex to a target sequence due to a certain complementarity with the target sequence. For example, in a gene editing system based on Cas9, gRNA is generally composed of crRNA and tracrRNA molecules that are partially complementary to form a complex, wherein crRNA contains a sequence that is sufficiently complementary to the target sequence so as to hybridize with the target sequence and guide the CRISPR complex (Cas9+crRNA+tracrRNA) to specifically bind to the target sequence sequence. However, it is known in the art that a single guide RNA (sgRNA) can be designed, which simultaneously contains the features of crRNA and tracrRNA. In a genome editing system based on Cpf1, gRNA is generally composed of only mature crRNA molecules, wherein the sequence contained in crRNA has sufficient homology with the target sequence so as to hybridize with the complementary sequence of the target sequence and guide the complex (Cpf1+crRNA) to specifically bind to the target sequence sequence. It is within the capabilities of those skilled in the art to design a suitable gRNA based on the CRISPR nuclease used and the target sequence to be edited.

[0085] As used herein, a "target sequence" is a sequence that is complementary or identical (depending on the different CRISPR nucleases) to a guide sequence of about 20 nucleotides contained in a guide RNA. The guide RNA targets the target sequence by base pairing with the target sequence or its complementary strand.

[0086] In some embodiments of the present invention, the gene editing results in the deletion of one or more nucleotides in the target sequence, preferably resulting in the deletion of multiple consecutive nucleotides in the target sequence. The type and length of the deletion depends on the double-strand break (DSB) position caused by the CRISPR nuclease and the number and position of cytosine (C) bases present in the target sequence or its complementary sequence. In some embodiments, the length of the deletion does not exceed the length of the target sequence. For example, the deletion can be about 1-17 nucleotides, such as 10-17 nucleotides, such as 10, 11, 12, 13, 14, 15, 16, 17 nucleotides.

[0087] In some embodiments of the invention, the cytosine deaminase is fused to the N-terminus of the CRISPR nuclease.

[0088] In some embodiments of the invention, the cytosine deaminase, the CRISPR nuclease, the UDG and / or the AP lyase are directly linked to each other.

[0089] In some embodiments of the present invention, the cytosine deaminase, the CRISPR nuclease, the UDG and / or the AP lyase are connected by a linker. The linker can be a non-functional amino acid sequence with a length of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids and no secondary structure or above. For example, the linker can be a flexible linker, such as GGGGS, GS, GAP, (GGGGS) x 3, GGS and (GGS) x 7, etc. In some embodiments, the linker comprises the amino acid sequence shown in SEQ ID NO: 8.

[0090] In some embodiments of the present invention, the polypeptide of the present invention further comprises a nuclear localization sequence (NLS). In general, one or more NLS in the polypeptide should have sufficient strength to drive the polypeptide to accumulate in the nucleus of the cell in an amount that can achieve its gene editing function. In general, the intensity of nuclear localization activity is determined by the number, position, one or more specific NLS used, or a combination of these factors of the NLS in the polypeptide.

[0091] In some embodiments of the present invention, the NLS of the polypeptide of the present invention may be located at the N-terminus and / or the C-terminus. In some embodiments of the present invention, the NLS of the polypeptide of the present invention may be located between the cytosine deaminase, the CRISPR nuclease, the UDG and / or the AP lyase. In some embodiments, the polypeptide comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the polypeptide comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the N-terminus. In some embodiments, the polypeptide comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as one or more NLSs at the N-terminus and one or more NLSs at the C-terminus. When there is more than one NLS, each can be selected to be independent of the other NLSs.

[0092] In general, NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the surface of the protein, but other types of NLS are also known. Non-limiting examples of NLS include: KKRKV (nucleotide sequence 5'-AAGAAGAGAAAGGTC-3'), PKKKRKV (nucleotide sequence 5'-CCCAAGAAGAAGAGGAAGGTG-3' or CCAAAGAAGAAGAGGAAGGTT), or SGGSPKKKRKV (nucleotide sequence 5'-TCGGGGGGGAGCCCAAAGAAGAAGCGGAAGGTG-3').

[0093] In addition, depending on the DNA location to be edited, the polypeptide of the present invention may also include other localization sequences, such as cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.

[0094] In some specific embodiments of the present invention, the first polypeptide comprises the amino acid sequence shown in SEQ ID NO: 5. In some specific embodiments of the present invention, the second polypeptide comprises the amino acid sequence shown in SEQ ID NO: 6.

[0095] In order to obtain efficient expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the polypeptide is codon-optimized for the organism from which the cell to be gene-edited comes.

[0096] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with codons that are more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit specific preferences for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the efficiency of translation of messenger RNA (mRNA), which in turn is believed to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons that are most frequently used for peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adapted for use in different ways. See, Nakamura et al., 2001. Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0097] The organisms from which the cells that can be gene-edited by the system of the present invention come are preferably eukaryotic organisms, including but not limited to mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants including monocotyledonous and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, and the like.

[0098] 3. Methods for modifying target sequences in cell genomes

[0099] In another aspect, the present invention provides a method for modifying a target sequence in a cell genome, comprising introducing the gene editing system of the present invention into the cell.

[0100] In some embodiments, the modification results in the deletion of one or more nucleotides, preferably multiple consecutive nucleotides in the target sequence. In the present invention, the type and length of the deletion caused by the modification depends on the double-strand break (DSB) position caused by the CRISPR nuclease and the number and position of cytosine (C) bases present in the target sequence or its complementary sequence. In some embodiments, the deletion is located within the target sequence. In some embodiments, the modification does not include insertion and / or substitution mutations.

[0101] In another aspect, the present invention also provides a method for producing genetically modified cells, comprising introducing the gene editing system of the present invention into the cells.

[0102] In another aspect, the present invention also provides a genetically modified organism comprising a genetically modified cell produced by the method of the present invention or a progeny cell thereof.

[0103] In the present invention, the target sequence to be modified can be located at any position in the genome, for example, in a functional gene such as a protein coding gene, or, for example, in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the gene function or modification of gene expression. The modification in the cell target sequence can be detected by T7EI, PCR / RE or sequencing methods.

[0104] In the method of the present invention, the gene editing system can be introduced into cells by various methods well known to those skilled in the art.

[0105] Methods that can be used to introduce the gene editing system of the present invention into cells include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun technique, PEG-mediated protoplast transformation, and soil Agrobacterium-mediated transformation.

[0106] Cells that can be gene-edited by the methods of the present invention can come from, for example, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; plants, including monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, and the like.

[0107] In some embodiments, the methods of the invention are performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.

[0108] In other embodiments, the method of the present invention can also be performed in vivo. For example, the cell is a cell in an organism, and the system of the present invention can be introduced into the cell in vivo by, for example, a virus or Agrobacterium-mediated method.

[0109] IV. Test Kit

[0110] The present invention also includes a kit for the method of the present invention, the kit comprising the gene editing system of the present invention, and instructions for use. The kit generally includes a label indicating the intended use and / or method of use of the contents of the kit. The term label includes any written or recorded material provided on or with the kit or otherwise provided with the kit. Example

[0111] Materials and Methods

[0112] 1. Vector construction

[0113] In order to construct pA3A-Cas9-UDG and pJIT163-Ubi-AP vectors, UDG and AP lyase sequences from Escherichia coli were obtained from NCBI (accession numbers AMB53293.1 and WP_115209270.1, respectively), and rice codon optimization and gene synthesis were carried out at Suzhou GeneWeizhi Company. Finally, the APOBEC3A, Cas9 and UDG fusion protein gene fragments and the APlyase gene fragment were introduced into the pJIT163 vector backbone, respectively, to obtain pA3A-SpCas9-UDG and pJIT163-Ubi-AP vectors.

[0114] In addition, APOBEC3A was fused to the N-terminus of Cas9 using an XTEN linker, UDG was fused to the C-terminus of Cas9, and AP lyase was fused to the C-terminus of UDG using a self-cleaving 2A polypeptide (P2A), and finally the fusion protein gene fragments were introduced into the pJIT163 vector backbone to construct the transient transformation vector AFID-3. Then, APOBEC3Bctd (APOBC3B sequence from humans (accession number NM_004900.5), and the C-terminal functional catalytic domain of APOBEC3B was truncated (APOBEC3Bctd)) was used to replace APOBEC3 in AFID-3 to construct eAFID-3. In addition, the fusion gene fragment with APOBEC3A was integrated into the pHUE411 backbone together with the sgRNA expression component using the Gibson method to construct a stable transformation vector pH-AFID-3 for Agrobacterium infection-mediated rice genetic transformation.

[0115] The gene coding sequences of key enzyme genes in the synthesis of wheat seed coat pigment (flavanone-3-hydroxylase gene, TaF3H-A1 / B1 / D1; colorless anthocyanin reductase gene TaLAR-A1 / B1 / D1) and their regulatory genes (TaMYB10-A1 / B1 / D1), plasma membrane kinase related to plant disease resistance (TaPMK-A1 / B1 / D1), vernalization response-related genes (TaVRN1-A1 / B1 / D1), and gibberellin-stimulated regulatory factor genes related to growth and development (TaGASR6-A1 / B1 / D1) were obtained. The targeting site sequences (sgF3HT4, sgLART4, sgMYBT2, sgPMKT1, sgVRN1T1 and sgGS6T2, see Table 1 for specific sequences) were compiled, sgRNA targeting site primers were synthesized, annealed and ligated into the pTaU6-sgRNA vector using T4 ligase to obtain pTaU6-sgF3HT4, pTaU6-sgLART4, pTaU6-sgMYBT2, pTaU6-sgPMKT1, pTaU6-sgVRN1T1 and pTaU6-sgGS6T2 vectors, respectively.

[0116] Table 1. sgRNA targeting primers

[0117]

[0118] Nine endogenous target sites were selected from seven rice genes (OsAAT, OsACC, OsCDC48, OsNRT1.1B, OsPDS, OsGRF1, and OsSPL14 / OsIPA1) for the construction of pOsU3-sgRNA vectors, and four endogenous target sites were selected from four wheat genes (TaF3H, TaGASR6, TaMYB10, and TamiR396) for the construction of pTaU6-sgRNA vectors. All target site sequences are shown in Table 2. sgRNA target site primers were synthesized, annealed, and ligated to the sgRNA vector using T4 ligase.

[0119] Table 2. sgRNA targeting sites and sequences

[0120]

[0121] PAM sequence is shown in bold

[0122] 2. Protoplast isolation and transformation (4 biological replicates)

[0123] 2.1 Rice or wheat seedling culture

[0124] The Zhonghua 11 rice seeds were first rinsed with 75% ethanol for 1 minute, then treated with 4% sodium hypochlorite for 30 minutes, washed with sterile water for more than 5 times, and cultured on M6 medium for 3-4 weeks at 26°C in the dark.

[0125] The wheat seeds were potted and planted in a culture room, and cultured for about 1-2 weeks (about 10 days) at a temperature of 25±2°C, a light intensity of 1000Lx, and a light intensity of 14-16h / d.

[0126] 2.2 Protoplast isolation

[0127] (1) Take young leaves of rice or wheat, cut the middle part into 0.5-1 mm strips with a blade, put them into 0.6 M Mannitol solution and keep them away from light for 10 min, then filter them with a filter, put them into 50 mL of enzymatic solution (filtered with 0.45 μm filter membrane), evacuate (pressure about 15 KPa) for 30 min, take them out and place them on a shaker (10 rpm) for enzymatic hydrolysis at room temperature for 5 h; (2) Add 30-50 mL of W5 to dilute the enzymatic hydrolysis product, filter the enzymatic hydrolysis solution with a 75 μm nylon filter membrane and place it in a round-bottom centrifuge tube (50 mL); (3) Centrifuge at 23°C, 100 g (rcf), up 3 times and down 3 times, and discard the supernatant; (4) Gently suspend with 10 mL of W5 and place on ice for 30 min; the protoplasts gradually settle and the supernatant is discarded; (5) Add an appropriate amount of MMG to suspend, place on ice and wait for transformation.

[0128] 2.3 Protoplast transformation

[0129] (1) Add 10 μg of each desired transformation vector to a 2 mL centrifuge tube and mix well. Use a sharpened pipette tip to draw up 200 μL of protoplasts and flick to mix well. Immediately add 250 μL of PEG4000 solution and flick to mix well. Induce transformation at room temperature in the dark for 20-30 min. (2) Add 800 μL of W5 (room temperature) and gently invert to mix. Centrifuge at 100 g (rcf), up 3 times and down 3 times for 3 minutes and discard the supernatant. (3) Add 1 mL of W5 and gently invert to mix well. Gently transfer to a 6-well plate that has been pre-added with 1 mL of W5. Wrap the 6-well plate with tin foil and culture at 23°C in the dark for 48 h.

[0130] 3. Protoplast DNA extraction and amplicon sequencing analysis

[0131] 3.1 Protoplast DNA extraction

[0132] The protoplasts were collected in 2 mL centrifuge tubes, and the protoplast DNA (~30 μL) was extracted using the CTAB method. The concentration (30-60 ng / μL) was determined using a NanoDrop ultra-micro spectrophotometer and stored at -20°C.

[0133] 3.2 Amplicon sequencing analysis

[0134] (1) Use genomic universal primers to perform PCR amplification on the protoplast DNA template. The 20μL amplification system contains 4μL 5×Fastpfu buffer, 1.6μL dNTPs (2.5mM), 0.4μL Forward primer (10μM), 0.4μL Reverseprimer (10μM), 0.4μL FastPfu polymerase (2.5U / μL), and 2μL DNA template (~60ng). Amplification conditions: 95℃ pre-denaturation for 5min; 95℃ denaturation for 30s, 50-64℃ annealing for 30s, 72℃ extension for 30s, 35 cycles; 72℃ full extension for 5min, 12℃ storage;

[0135] (2) The above amplification product was diluted 10 times, and 1 μL was taken as the template for the second round of PCR amplification. The amplification primer was a sequencing primer containing a barcode. The 50 μL amplification system contained 10 μL 5×Fastpfu buffer, 4 μL dNTPs (2.5 mM), 1 μL Forward primer (10 μM), 1 μL Reverse primer (10 μM), 1 μL FastPfu polymerase (2.5 U / μL), and 1 μL DNA template. The amplification conditions were the same as above, and the number of amplification cycles was 38 cycles.

[0136] (3) The PCR products were separated by 2% agarose gel electrophoresis, and the target fragments were recovered by gel extraction using AxyPrepTM DNA Gel Extraction kit. The recovered products were quantitatively analyzed using NanoDrop ultra-micro spectrophotometer; 100 ng of the recovered products were mixed and sent to Genewise Biotechnology Co., Ltd. for amplicon sequencing library construction and amplicon sequencing analysis.

[0137] (4) After sequencing is completed, the raw data are split according to the sequencing primers. The sgRNA sequence and its flanking sequences are used as reference sequences, and the WT is used as a control to compare and analyze the type and efficiency of gene editing at different gene targeting sites in four repeated experiments.

[0138] Example 1: Construction of a gene editing system for precise short fragment deletion (ACD)

[0139] The single-base editing system was established in 2016 (Komor et al., 2016; Ma et al., 2016; Nishida et al., 2016). The system uses nCas9 (D10A) to guide cytosine deaminase to act on the non-complementary strand of the DNA target site and deaminate cytosine (C) in a specific region into uracil (U). Uracil (U) will be replaced by thymine (T) during DNA replication, thereby achieving precise single-base replacement of C-to-T. In the repair process of plant and animal organisms, uracil-DNA glycosylase (UDG) will preferentially recognize the U base and remove the N-glycosidic bond of the base to form an apurinic or apyrimidinic site (AP site), and then the U base will be repaired to the original C base through base excision repair under the action of AP lyase (AP lyase). Therefore, uracil-DNA glycosylase inhibitor (UGI) is often introduced into the single-base editing system to improve the C-to-T editing efficiency.

[0140] The inventors surprisingly found that replacing the nCas9 of the fusion protein in the single-base editing system with wild-type Cas9 enables the fusion protein to regain the ability to break the double-stranded DNA, and at the same time replacing UGI with UDG to recognize the U base and remove its glycosidic bond to form an AP site, which is then recognized by AP lyase and the glycosylated U base is removed, ultimately achieving efficient, accurate and predictable short-fragment deletion in cells. The inventors thus constructed an efficient, accurate and predictable short-fragment deletion system (APOBEC3A Coupled Deletion, ACD) composed of Cas9, APOBEC3A, UDG and AP lyase, wherein Cas9 mediates the generation of DSB at the DNA target site, while APOBEC3A, UDG and AP lyase mediate the generation of multiple gaps at the non-complementary chain C base upstream of the DSB, thereby causing the deletion of single-stranded DNA fragments on the non-complementary chain, and ultimately forming a short-fragment deletion of the DNA double strand under the action of the body's DNA repair ( Figure 1Without wishing to be bound by any theory, APOBEC3A can efficiently mediate the C-to-U replacement of the non-targeted strand upstream of DSB, while UDG and AP lyase mediate the formation of a gap at the U base, resulting in the loss of single-stranded DNA fragments on the non-targeted strand. At this time, the targeted strand forms a 5' overhanging end, which will first be recognized and removed by the Artemis-DNA-PK complex during the body's non-homologous end repair process, and then form a short-fragment missing DNA double-strand under the action of a ligation complex composed of DNA ligase IV, XRCC4, XRCC4-like factor (XLF) and its paralogs (PAXX) (Chang et al. 2017).

[0141] The efficiency of insertion and deletion produced by SpCas9 and ACD at the targeted editing sites of sgF3HT4, sgLART4, sgMYBT2, sgPMKT1, sgVRN1T1 and sgGS6T2 was compared and analyzed. The results showed that compared with SpCas9, the insertion mutation rate of the ACD system was significantly reduced, while the deletion mutation rate was significantly increased, and the deletion mutation rate was 1.5-23.6 times that of SpCas9, which fully demonstrated the high efficiency of the ACD system ( Figure 2 ).

[0142] Example 2: Analysis of Deletion Types Generated by the ACD System

[0143] Sequence analysis of the Deletion mutations generated by the ACD system at different target sites ( Figure 3-8 ), except for a few types, most mutation types are in line with expectations, all of which are deletions between the APOBEC3A action base (NGG (PAM) corresponds to the C base; CCN (PAM) corresponds to the G base) and the Cas9 cutting site. However, due to the asymmetry of Cas9 when cutting double strands, Cas9 will cut between the 3rd and 4th positions or the 4th and 5th positions near the PAM end. In addition, the bases where APOBEC3A acts on the non-targeted chain will use the targeted chain as a template to form 1-2 bases that pair with the complementary chain during the repair process. Therefore, 1-2 bases that complement the targeted chain may also be introduced.

[0144] The ACD system has a very low efficiency in producing insertions, but a very high efficiency in producing deletions, and deletions only occur within the 20-bp pre-spacer sequence. In these targeted sites, most deletions are 10 to 17 nt in length, and different types of deletions can be stably detected in more than three biological repeat experiments, which is something that SpCas9 and other tools cannot do, which fully demonstrates the accuracy and predictability of the ACD system.

[0145] Example 3: Construction of AFID (APOBEC-Cas9 Fusion-Induced Deletion) system

[0146] The present invention selects human APOBEC3A with high deamination activity and wide deamination window to construct the AFID-3 system, and selects an APOBEC3Bctd with higher deamination activity and narrow window to replace APOBEC3A to construct the eAFID-3 system ( Fig. 9 and Fig.10 ). A comparative analysis of the deletion efficiency of Cas9, AFID-3 and eAFID-3 at endogenous gene targets in rice and wheat showed that the efficiency of deletion mutations produced by Cas9, AFID-3 and eAFID-3 increased significantly, with their average deletion mutation rates being 2.2 and 2.6 times that of Cas9, respectively, which fully demonstrates the high efficiency of the AFID system.

[0147] Example 4: Analysis of mutation types generated by the AFID system

[0148] The mutation types and proportions generated by AFID-3 and eAFID-3 at different endogenous target sites were analyzed. The results showed that the length of the deletion fragment mainly depends on the position of the deamination C nucleotide and its deamination activity. At target sites with strong deamination activity, deletion mutations are the main mutation type; but at target sites with weak deamination activity, a certain proportion of insertion mutations will appear. The mutation types with a larger proportion are predictable multi-nucleotide deletion mutations (from the C nucleotide where the deaminase acts to the Cas9 cleavage site (Cas9 cuts the double strands asymmetric, resulting in the Cas9 cleavage site appearing between the 3rd and 4th or 4th and 5th positions near the PAM end) Fig.13 , 14 ). In addition, it was also found that during the NHEJ repair process, templated insertion of C nucleotides occurred at the deaminated C nucleotides ( Fig.13 , 14 ), which is mainly due to the fact that during the excision process of the 5' protruding end on the targeted chain, the DNA polymerase can easily use this 5' protruding end as a template to perform base repair on the non-targeted chain.

[0149] In order to detect the preference of AFID-3 and eAFID-3 for the C base at the start of the predictable deletion, the proportion of deletion mutations from the AC, TC, CC and GC motifs to DSBs at different targets was counted. The results showed that AFID-3 could mediate predictable deletion mutations from the AC, TC, CC and GC motifs to DSBs. Compared with AFID-3, eAFID-3 showed an enhanced preference for the TC base, and most of its predictable deletion mutations were deletion mutations from the TC motif to DSBs ( Fig.15 ). In addition, the types and proportions of mutations produced by Cas9, AFID-3 and eAFID-3 at the miR396h binding site of the rice OsGRF1 gene and the miR156 binding site of the OsIPA1 gene were analyzed. The results showed that Cas9 was almost unable to produce such predictable in-frame deletion mutations; while both AFID-3 and eAFID-3 could produce such predictable deletion mutations, but the proportion produced by eAFID-3 was significantly higher than that of AFID-3 ( Fig.16 ). This also fully demonstrates the accuracy and predictability of the AFID system.

[0150] Example 5: AFID system mediates predictable polynucleotide deletion mutations in plants

[0151] To determine whether the AFID system can mediate predictable multinucleotide deletion mutations in plants, two targets (TamiR396 and TaGASR6) were selected in wheat, and Cas9 and AFID-3 were delivered to immature embryos of wheat by gene gun bombardment. Three targets (OsCDC48-T2, OsSPL14, and OsPDS) were selected in rice to construct the corresponding pH-Cas9 and pH-AFID-3 Agrobacterium vectors ( Fig.17 ) and transformed rice callus using Agrobacterium infection. The results showed that among the tested targets, Cas9 did not produce predictable polynucleotide deletion mutants, and its mutation types were mainly 1-bp insertion and 1-3bp deletion; while AFID-3 produced mostly polynucleotide deletion mutants, of which the predictable proportion accounted for 25.0-55.5% (Table 3, Fig.18 ). It can be seen that the AFID system can mediate predictable multi-nucleotide deletion mutations in plants.

[0152] Table 3 Statistics of predictable deletion plant mutants generated by AFID-3

[0153]

[0154] Sequence Listing

[0155] SEQ ID NO: 1SpCas9

[0156] KDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAETRL

[0157] KRTARRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAY

[0158] HEKYPTIYHLRKKLVDSTDKADLRLIYLALAMHIKFRGHFLIEGDLNPDNSDVDKLFIQLVQT

[0159] YNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF

[0160] DLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSA

[0161] SMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKM

[0162] DGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIP

[0163] YYVGPLARGNSRFAWMTRKSEEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKH

[0164] SLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIE

[0165] CFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKT

[0166] YAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD

[0167] SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0168] ARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVD

[0169] QELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLL

[0170] NAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLI

[0171] REVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVGTALIKKYPKLESEFVYGDY

[0172] KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFKTEITLANGEIRKRPLIETNGETGEIVWDKG

[0173] RDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVA

[0174] YSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLF

[0175] ELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHY

[0176] LDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0177] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD

[0178] SEQ ID NO:2APOBEC3A

[0179] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKN

[0180] LLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRI

[0181] FAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQA

[0182] LSGRLRAILQNQGNSGSETPGTSESATPES

[0183] SEQ ID NO:3UDG

[0184] ANELTWHDVLAEEKQQPYFLNTLQTVASERQSGVTIYPPQKDVFNAFRFTELGDVKVVILGQ

[0185] DPYHGPGQAHGLAFSVRPGIAIPPSLLNMYKELENTIPGFTRPNHGYLESWARQGVLLLNTVL

[0186] TVRAGQAHSHASLGWETFTDKVISLINQHREGVVFLLWGSHAQKKGAIIDKQRHHVLKAPHP

[0187] SPLSAHRGFFGCNHFVLANQWLEQRGETPIDWMPVLPAESE

[0188] SEQ ID NO:4AP lyase

[0189] MPEGPEIRRAADNLEAAIKGKPLTDVWFAFPQLKPYQSQLIGQHVTHVETRGKALLTHFSNDL

[0190] TLYSHNQLYGVWRVVDTGEEPQTTRVLRVKLQTADKTILLYSASDIEMLTPEQLTTHPFLQRVG

[0191] PDVLDPNLTPEVVKERLLSPRFRNRQFAGLLLDQAFLAGLGNYLRVEILWQVGLTGNHKAKD

[0192] LNAAQLDALAHALLEIPRFSYATRGQVDENKHHGALFRFKVFHRDGEPCERCGSIIEKTTLSSR

[0193] PFYWCPGCQH

[0194] SEQ ID NO:5 Exemplary First Polypeptide

[0195] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKN

[0196] LLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRI

[0197] FAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQA

[0198] LSGRLRAILQNQGNSGSETPGTSESATPESLKDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFK

[0199] VLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSF

[0200] FHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM

[0201] IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLI

[0202] AQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYAD

[0203] LFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFD

[0204] QSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLG

[0205] ELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEITITPWNFEEV

[0206] VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGE

[0207] QKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFL

[0208] DNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIR

[0209] DKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKK

[0210] GILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKE

[0211] HPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRS

[0212] DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQL

[0213] VETRQITKHVAQILDSRMNTKYDENDKLIVEKVITLSKLVSDFRKDFQFYKVREINNYHHA

[0214] HDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0215] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIL

[0216] PKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSF

[0217] EKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLY

[0218] LASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK

[0219] PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLG

[0220] GDKRPAATKKAGQAKKKKTRDSGGSANELTWHDVLAEEKQQPYFLNTLQTVASERQSGVTI

[0221] YPPQKDVFNAFRFTELGDVKVVILGQDPYHGPGQAHGLAFSVRPGIAIPPSLLNMYKELENTIP

[0222] GFTRPNHGYLESWARQGVLLLNTVLTVRAGQAHSHASLGWETFTDKVISLINQHREGVVFLL

[0223] WGSHAQKKGAIIDKQRHHVLKAPHPSPLSAHRGFFGCNHFVLANQWLEQRGETPIDWMPVL

[0224] PAESEPKKKRKV

[0225] SEQ ID NO:6 Exemplary Second Polypeptide

[0226] MPEGPEIRRAADNLEAAIKGKPLTDVWFAFPQLKPYQSQLIGQHVTHVETRGKALLTHFSNDL

[0227] <h2 style=";text-align:left;direction:ltr">TLYSHNQLYGVWRVVDTGEEPQTTRVLRVKLQTADKTILLYSASDIEMLTPEQLTTHPFLQRVG<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0228] <h2 style=";text-align:left;direction:ltr"> PDVLDPNLTPEVVKERLLSPRFRNRQFAGLLLDQAFLAGLGNYLRVEILWQVGLTGNHKAKD<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0229] <h2 style=";text-align:left;direction:ltr"> LNAAQLDALAHALLEIPRFSYATRGQVDENKHHGALFRFKVFHRDGEPCERCGSIIEKTTLSSR<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0230] <h2 style=";text-align:left;direction:ltr"> PFYWCPGCQHPKKKRKV<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0231] <h2 style=";text-align:left;direction:ltr"> SEQ ID NO:7APOBEC3Bctd<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0232] <h2 style=";text-align:left;direction:ltr"> MEILRYLMDPDTFTFNNNDPLVLRRRQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQNQGN<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0233] <h2 style=";text-align:left;direction:ltr"> SEQ ID NO:8 XTEN linker<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0234] <h2 style=";text-align:left;direction:ltr"> SGSETPGTSESATPES<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0235] <h2 style=";text-align:left;direction:ltr"> SEQ ID NO:9 P2A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0236] <h2 style=";text-align:left;direction:ltr"> GSGATNFSLLKQAGDVEENPGPPE<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0237] <h2 style=";text-align:left;direction:ltr"> SEQ ID NO:10 AFID-3<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0238]

[0239] SEQ ID NO:11 eAFID-3

[0240]

[0241] GDKRPAATKKAGQAKKKKTRDSGGSANELTWHDVLAEEKQQPYFLNTLQTVASERQSGVTI

[0242] YPPQKDVFNAFRFTELGDVKVVILGQDPYHGPGQAHGLAFSVRPGIAIPPSLLNMYKELENTIP

[0243] GFTRPNHGYLESWARQGVLLLNTVLTVRAGQAHSHASLGWETFTDKVISLINQHREGVVFLL

[0244] WGSHAQKKGAIIDKQRHHVLKAPHPSPLSAHRGFFGCNHFVLANQWLEQRGETPIDWMPVL

[0245] PAESEPKKKRKVSAGSGATNFSLLKQAGDVEENPGPPEGPEIRRAADNLEAAIKGKPLTDVWF

[0246] AFPQLKPYQSQLIGQHVTHVETRGKALLTHFSNDLTLYSHNQLYGVWRVVDTGEEPQTTRVL

[0247] RVKLQTADKTILLYSASDIEMLTPEQLTTHPFLQRVGPDVLDPNLTPEVVKERLLSPRFRNRQFA

[0248] GLLLDQAFLAGLGNYLRVEILWQVGLTGNHKAKDLNAAQLDALAHALLEIPRFSYATRGQVD

[0249] ENKHHGALFRFKVFHRDGEPCERCGSIIEKTTLSSRPFYWCPGCQHPKKKRKV。

Claims

1. A gene editing system for editing a target sequence in a cell genome, comprising: i) a polypeptide, and / or an expression construct comprising a nucleotide sequence encoding the polypeptide; and ii) a guide RNA, and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, The polypeptides include cytosine deaminase, CRISPR nuclease, uracil-DNA glycosylase and AP lyase, and the enzymes are connected in sequence. wherein the guide RNA is capable of targeting the polypeptide to a target sequence in the cell genome, wherein the CRISPR nuclease is a wild-type Cas9 nuclease, The wild-type Cas9 nuclease has double-stranded DNA cleavage activity. The cytosine deaminase is the C-terminal functional catalytic domain of APOBEC3B deaminase, The amino acid sequence of the C-terminal functional catalytic domain of APOBEC3B deaminase is SEQ ID NO: 7, The amino acid sequence of the uracil-DNA glycosylase is SEQ ID NO: 3, The amino acid sequence of the AP lyase is SEQ ID NO:

4.

2. The gene editing system of claim 1, wherein the wild-type Cas9 nuclease is spCas9.

3. The gene editing system of claim 1, wherein the amino acid sequence of the polypeptide is SEQ ID NO:

11.

4. A method for producing a genetically modified plant cell, comprising introducing the gene editing system of any one of claims 1-3 into the plant cell, wherein the genetic modification is a deletion of 2-18 consecutive nucleotides in the target sequence.

5. The method of claim 4, wherein the genetic modification is a deletion of 10-17 consecutive nucleotides in the target sequence.

6. The method of claim 4 or 5, wherein the plants include monocots and dicots.

7. The method of claim 6, wherein the plant is selected from the group consisting of rice, corn, wheat, sorghum, barley, soybean, peanut, and Arabidopsis.

8. A kit comprising the gene editing system of any one of claims 1-3.