A system and method for editing a nucleic acid

By using a functional complex of H840A-mutated spCas9 with MLV-RT fusion protein and PegRNA, the problem of inaccurate large-fragment gene editing in existing technologies has been solved, achieving efficient and precise gene knock-in and replacement.

CN113913405BActive Publication Date: 2025-11-07INST OF ZOOLOGY CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110780360.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-10
Filing Date
2021-07-09
Publication Date
2025-11-07
Estimated Expiration
2041-07-09

AI Technical Summary

Technical Problem

Existing gene editing technologies are difficult to efficiently perform site-specific knock-in and replacement of large exogenous gene fragments, especially insertions and replacements larger than 1Kbp. Furthermore, existing methods suffer from inaccurate editing and high costs.

Method used

By employing a fusion protein consisting of spCas9 with an H840A mutation and the reverse transcriptase MLV-RT, along with PegRNA, a functional complex is formed to cleave at the genome target site and the reverse transcriptase is used to reverse transcribe and extend the edited sequence at the cleavage site, achieving efficient gene knock-in and replacement.

Benefits of technology

It enables efficient and precise gene knock-in and replacement, especially for the insertion and replacement of large exogenous gene fragments, avoiding additional base deletions or insertions and improving the accuracy of editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113913405B_ABST
    Figure CN113913405B_ABST
Patent Text Reader

Abstract

The present application relates to systems and kits for editing nucleic acids and uses thereof, and methods of editing nucleic acids. The systems, kits, and methods of the present application can be used to cleave a double-stranded target nucleic acid and form overhangs at its ends, and can be used to insert a target nucleic acid into a nucleic acid molecule of interest (e.g., genomic DNA) or to replace a nucleotide fragment in a nucleic acid molecule of interest (e.g., genomic DNA) with a target nucleic acid.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of genetic engineering and molecular biology. In particular, the present application relates to systems and kits for editing nucleic acids and uses thereof, as well as methods of editing nucleic acids. The systems, kits and methods of the present application can be used to break a double-stranded target nucleic acid and form overhangs at its ends, in particular at the ends of the break, and can be used to insert a target nucleic acid in a nucleic acid molecule of interest, such as genomic DNA, or to replace a fragment of nucleotides in a nucleic acid molecule of interest, such as genomic DNA, with a target nucleic acid. BACKGROUND

[0002] Gene editing technology is a hot field of biomedical research, and has broad application prospects in clinical treatment of genetic diseases, construction of animal models, genetic breeding of crops, etc. Gene editing technology includes deletion, addition and replacement of single nucleotides or a segment of DNA sequence at specific sites in the genome. Site-directed knock-in of exogenous genes can be achieved by homologous recombination (HDR): introduction of a 500-3000 bp homologous arm on both sides of the exogenous gene can achieve precise site-directed integration of the exogenous gene, but the efficiency is extremely low, only about 0.01%. By artificially constructing nucleases such as ZFN (zinc-finger nucleases), TALEN (transcription activator-like effector nucleases) or CRISPR / Cas9 (clustered regularly interspaced short palindromic repeats / CRISPR-associated protein-9 nuclease), cleavage at the targeted site in the genome generates a double-strand break (DSB), which can promote site-directed knock-in of exogenous genes mediated by homologous recombination. However, due to the fact that most mammalian cells mainly rely on NHEJ (non-homologous end joining) for DSB repair, the efficiency of site-directed knock-in based on nucleases and homologous recombination is still very low, generally about 1%. In addition, since homologous recombination only occurs in the S / G2 phase of the cell cycle, most somatic cells in the terminal differentiation stage cannot achieve site-directed integration of exogenous genes by the above methods.

[0003] The linear single-stranded DNA donor can also achieve site-specific integration of exogenous DNA fragments. The single-stranded DNA donor has a 30-50 nt homologous arm at each end. After the nuclease cuts the specific site of the genome, the single-stranded DNA is integrated into the DSB site by SDSA (synthesis-dependent strand annealing), thereby achieving site-specific integration of the genome. Linear single-stranded DNA is more efficient than HDR, but less accurate: additional base insertion and deletion often occur at the 5' end of the single-stranded DNA. In addition, the chemical synthesis of long linear DNA single strands is costly and difficult to obtain. Therefore, this method is not suitable for site-specific knock-in of large exogenous genes (more than 1 Kb). In addition, when the inserted fragment is more than 1 Kb, the integration efficiency will also be significantly reduced.

[0004] NHEJ-based site-specific knock-in, such as HITI (Homology-independent target integration) technology, does not rely on homologous arms at both ends of the exogenous gene. In this method, the nuclease cuts the specific site of the genome and also cuts the donor vector. Then the linearized exogenous gene DNA fragment is inserted into the broken site of the genome through the NHEJ DNA repair pathway. NHEJ-based site-specific knock-in has no directionality, and the position of the junction is often inaccurate, which can easily produce additional base insertion or deletion. The MMEJ-based site-specific knock-in method introduces micro-homologous arms at both ends of the exogenous gene based on NHEJ, but the efficiency is still very low.

[0005] Prime Editing is a new type of gene editing method. This method uses a fusion protein composed of spCas9 (nCas9) with a H840A mutation and reverse transcriptase MLV-RT (Murine Leukemia Virus-Reverse Transcriptase), and a PegRNA (Prime editing guide RNA) modified from gRNA (guide RNA), which can realize the conversion / transversion of any single base or the deletion, addition and replacement of small fragments of DNA. The PegRNA is generated by introducing a PBS (Prime binding site) sequence and a template sequence at the 3' end of the gRNA, wherein the template sequence contains an editing sequence and a homologous sequence of a genomic DSB site. In this method, the complex formed by nCas9 and PegRNA binds to the genomic target site and cuts the PAM strand, then the PBS sequence on the PegRNA pairs with the 3' end released from the upstream of the PAM strand, and then the MLV-RT extends the editing sequence and the homologous sequence at the 3' end of the PAM strand cut with the template sequence of the PegRNA as a template. Subsequently, through the process of DNA single strand replacement and mismatch repair, repair can be completed at the cut site and the editing sequence can be integrated into the target site. Since H840A nCas9 only cuts one strand of double-stranded DNA (i.e. the PAM strand), it does not induce NHEJ caused by DSB, so this method is less likely to introduce additional base deletions or insertions, and has high editing accuracy. However, due to the length limitation of the template sequence on the PegRNA, the length of the editable sequence is limited, and Prime Editing is only suitable for the deletion or knock-in of base sequences less than 100 bp.

[0006] Therefore, it is of great importance to establish a method capable of efficiently performing gene site-directed knock-in and replacement, especially a method capable of efficiently performing insertion and replacement of large fragments (more than 1 Kbp) of exogenous genes, for expanding the application of gene editing technology in production and medical treatment. SUMMARY

[0007] In the present application, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. And the nucleic acid chemical laboratory operation steps used herein are conventional steps widely used in the corresponding field. At the same time, in order to better understand the present application, the definitions and explanations of related terms are provided as follows.

[0008] The term "Cas protein" or "Cas nuclease" is an RNA-guided nuclease. Cas proteins are also known as casnl nucleases or CRISPR-associated nucleases. CRISPR (clustered regularly interspaced short palindromic repeats) is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain repeats and spacers, where the spacers are sequences complementary to mobile genetic elements that can target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA also requires the involvement of a trans-encoded small RNA (tracrRNA). Thus, in nature, type II CRISPR systems require a Cas protein and two RNAs for cleavage of DNA. However, by engineering, the crRNA and tracrRNA can be incorporated into a single guide RNA (abbreviated as "sgRNA" or "gNRA"). See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.

[0009] As used herein, the term "complementary" means that two nucleic acid sequences are capable of forming hydrogen bonds between each other according to the base pairing rules (Waston-Crick rules) and thereby form a duplex. In the present application, the term "complementary" includes "substantially complementary" and "perfectly complementary". As used herein, the term "perfectly complementary" means that every base in one nucleic acid sequence is capable of pairing with a base in another nucleic acid sequence without the presence of mismatches or gaps. As used herein, the term "substantially complementary" means that a substantial number of bases in one nucleic acid sequence are capable of pairing with bases in another nucleic acid sequence, which allows the presence of mismatches or gaps (e.g., one or several nucleotides of mismatches or gaps). Generally, under conditions that allow nucleic acid hybridization, annealing, or amplification, two nucleic acid sequences that are "complementary" (e.g., substantially complementary or perfectly complementary) will selectively / specifically hybridize or anneal and form a duplex.

[0010] As used herein, the term "DNA polymerase" refers to an enzyme that is capable of synthesizing one nucleic acid strand (e.g., a DNA strand or an RNA strand) using another nucleic acid strand (e.g., a DNA strand or an RNA strand) as a template. In the present application, a DNA polymerase can be a DNA-dependent DNA polymerase (i.e., an enzyme that is capable of synthesizing a complementary DNA strand using a DNA strand as a template) or an RNA-dependent DNA polymerase (i.e., an enzyme that is capable of synthesizing a complementary DNA strand using an RNA strand as a template). In certain embodiments, a DNA polymerase used in the present application is an RNA-dependent DNA polymerase, such as a reverse transcriptase.

[0011] As used herein, the term "reverse transcriptase (RT)" refers to an enzyme that is capable of synthesizing a complementary DNA strand using an RNA strand as a template. Reverse transcriptases of the present application include, but are not limited to, reverse transcriptases from retroviruses or other viruses or bacteria, as well as DNA polymerases having reverse transcription activity, such as TTH DNA polymerase, Taq DNA polymerase, TNE DNA polymerase, TMA DNA polymerase, and the like. Reverse transcriptases from retroviruses include, but are not limited to, reverse transcriptases from Moloney murine leukemia virus (M-MLV), human immunodeficiency virus (HIV), avian sarcoma-leukosis virus (ASLV), Rous sarcoma virus (RSV), avian myeloblastosis virus (AMV), avian erythroblastosis virus helper virus, avian myelocytomatosis virus MC29 helper virus, avian reticuloendotheliosis virus helper virus, avian sarcoma virus UR2 helper virus, avian sarcoma virus Y73 helper virus, Rous-associated virus, and myeloblastosis-associated virus (MAV). Specific examples of reverse transcriptases can also be found in, for example, U.S. Patent Application 2002 / 0198944 (which is incorporated herein by reference in its entirety). In addition, reverse transcriptases of the present application include, but are not limited to, any form, such as, for example, naturally occurring reverse transcriptases, naturally occurring mutant reverse transcriptases, engineered mutant reverse transcriptases, or other variants (e.g., truncated variants that retain reverse transcription activity).

[0012] As used herein, the terms "hybridization" and "annealing" mean the process by which complementary single-stranded nucleic acid molecules form a double-stranded nucleic acid. In the present application, "hybridization" and "annealing" have the same meaning and are used interchangeably. Typically, two nucleic acid sequences that are completely complementary or substantially complementary can hybridize or anneal. The degree of complementarity required for two nucleic acid sequences to hybridize or anneal depends on the hybridization conditions, particularly the temperature.

[0013] As used herein, "conditions that allow nucleic acid hybridization" has the meaning generally understood by those skilled in the art, and can be determined by routine methods. For example, two nucleic acid molecules having complementary sequences can hybridize under suitable hybridization conditions. Such hybridization conditions can involve factors such as temperature, pH value, composition of the hybridization buffer, and ionic strength, etc. of the hybridization buffer, and can be determined according to the length and GC content of the two complementary nucleic acid molecules. For example, when the length of the two complementary nucleic acid molecules is relatively short and / or the GC content is relatively low, low stringent hybridization conditions can be employed. When the length of the two complementary nucleic acid molecules is relatively long and / or the GC content is relatively high, high stringent hybridization conditions can be employed. Such hybridization conditions are well known to those skilled in the art, and can be found in, for example, Joseph Sambrook, et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2001); and M.L.M. Anderson, Nucleic Acid Hybridization, Springer- Verlag New York Inc. N.Y. (1999). In the present application, "hybridization" and "annealing" have the same meaning, and can be used interchangeably. Accordingly, the expressions "conditions that allow nucleic acid hybridization" and "conditions that allow nucleic acid annealing" also have the same meaning, and can be used interchangeably.

[0014] As used herein, the term "upstream" is used to describe the relative positional relationship of two nucleic acid sequences (or two nucleic acid molecules), and has the meaning generally understood by those skilled in the art. For example, the expression "one nucleic acid sequence is upstream of another nucleic acid sequence" means that, when arranged in the 5' to 3' direction, the former is located at a more forward position (i.e., a position closer to the 5' end) compared to the latter. As used herein, the term "downstream" has the opposite meaning of "upstream".

[0015] As used herein, the term "linker" refers to a chemical entity used to connect two entity elements (e.g., two nucleic acids or two polypeptides). For example, a linker used to connect two polypeptides can be a peptide linker (e.g., a linker comprising a plurality of amino acid residues); a linker used to connect two nucleic acids can be a nucleic acid linker (e.g., a linker comprising a plurality of nucleotides).

[0016] As used herein, the term "guide sequence" is a targeting sequence comprised by a guide RNA. In certain instances, a guide sequence is a polynucleotide sequence that has sufficient complementarity to a target sequence such that it is capable of hybridizing to the target sequence and directing specific binding of a CRISPR / Cas complex to the target sequence. In certain embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Methods of determining the complementarity of two nucleic acid sequences are within the ability of one of ordinary skill in the art. For example, there are published and commercially available alignment algorithms and programs such as, but not limited to, ClustalW, Smith-Waterman in matlab, Bowtie, Geneious, Biopython, and SeqMan.

[0017] As used herein, the term "scaffold sequence" is a sequence recognized and bound by a Cas protein in a guide RNA. In certain instances, a scaffold sequence can comprise or consist of a repeat sequence of a CRISPR.

[0018] As used herein, the term "functional complex" refers to a complex formed by the binding of a guide RNA (gRNA) to a Cas protein that is capable of recognizing and cleaving a polynucleotide to which the guide RNA is directed.

[0019] As used herein, the term "target nucleic acid" or "target sequence" is a polynucleotide to which a guide sequence is directed, e.g., a sequence that has complementarity to the guide sequence. Perfect complementarity of a guide sequence to a target sequence is not required, so long as there is sufficient complementarity to cause hybridization and to promote binding of a CRISPR / Cas complex. A target sequence can comprise any polynucleotide, such as DNA or RNA. In certain instances, the target sequence is located in the nucleus or cytoplasm of a cell. In certain instances, the target sequence can be located within an organelle of a eukaryotic cell, such as a mitochondrion or a chloroplast.

[0020] In the present disclosure, the expression "target sequence" or "target nucleic acid" can be any endogenous or exogenous polynucleotide to a cell (e.g., a eukaryotic cell). For example, the target nucleic acid can be a polynucleotide (e.g., genomic DNA) present in the nucleus of a eukaryotic cell or a polynucleotide (e.g., vector DNA) introduced into a cell exogenously. For example, the target nucleic acid can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). In some cases, the target nucleic acid or target sequence comprises or is adjacent to a protospacer adjacent motif (PAM). The exact sequence and length of the PAM depends on the Cas protein used. Typically, the PAM is a sequence of 2-5 base pairs adjacent to a protospacer in a CRISPR cluster. Those skilled in the art are able to identify PAM sequences for use with a given Cas protein.

[0021] As used herein, the term "vector" refers to a nucleic acid vehicle into which a polynucleotide can be inserted. When the vector is capable of mediating expression of the inserted polynucleotide, the vector is referred to as an expression vector. Vectors can be introduced into host cells by transformation, transduction, or transfection and allow the carried genetic material elements to be expressed in the host cells. Vectors are well known to those skilled in the art and include, but are not limited to, plasmids; phagemids; cosmids; nanolipid particles; exosomes; artificial chromosomes, such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs); bacteriophages, such as lambda phage or M13 phage; and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpesviruses (e.g., herpes simplex viruses), poxviruses, baculoviruses, papillomaviruses, papova viruses (e.g., SV40). A vector can contain multiple elements that control expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, a vector can contain a replication origin. Those skilled in the art will appreciate that the design of an expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression desired, etc. When a vector carries foreign DNA to be integrated into a host genome, and non-protein expression elements associated with the integration of the foreign DNA, the vector is referred to as a donor vector. Foreign DNA includes, but is not limited to, entire genes or gene fragments, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and protein coding sequences. Non-protein expression elements associated with the integration of the foreign DNA include, but are not limited to, homologous sequences of the insertion site, targeted cleavage sequences for tool enzymes, etc. Adeno-associated viral vectors include, but are not limited to, adeno-associated viruses of different serotypes such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV-DJ, and other engineered serotypes of adeno-associated viruses.

[0022] In the present application, the "intein" refers to a kind of internal protein element that can mediate post-translational protein splicing. The intein is located in the middle of the polypeptide sequence, is removed after processing, and catalyzes the ligation of the proteins on both ends into a mature protein molecule. The "intein splitting system" is a system for efficiently splitting and splicing larger protein molecules using inteins. The intein can be separated into an N-terminal segment and a C-terminal segment. The target protein is split into an N-terminal segment and a C-terminal segment, which are connected to the N-terminal segment and the C-terminal segment of the intein, respectively, to form a fusion protein. Only when the N-terminal segment and the C-terminal segment of the fusion protein meet, the intein in the split precursor protein undergoes protein splicing removal, the N-terminal segment and the C-terminal segment of the target protein are spliced, and then a functional target protein is formed. The intein suitable for use in the present application is derived from, but not limited to, DnaE DNA polymerase of Synechocystis sp. PCC6803 and Nostoc punctiforme PCC73102 (Npu).

[0023] As used herein, the term "host cell" refers to a cell that can be used for introducing a vector, including but not limited to prokaryotic cells such as E. coli or Bacillus subtilis, fungal cells such as yeast cells or Aspergillus, insect cells such as S2 Drosophila cells or Sf9, or animal cells such as fibroblast cells, CHO cells, COS cells, NSO cells, HeLa cells, BHK cells, HEK 293 cells or human cells.

[0024] In a first aspect, the present application provides a system or kit comprising the following four components:

[0025] (1) a first Cas protein or a nucleic acid molecule A1 containing a nucleotide sequence encoding the first Cas protein, wherein the first Cas protein is capable of cleaving or breaking a first double-stranded target nucleic acid;

[0026] (2) a template-dependent first DNA polymerase or a nucleic acid molecule B1 containing a nucleotide sequence encoding the first DNA polymerase;

[0027] (3) a first gRNA or a nucleic acid molecule C1 containing a nucleotide sequence encoding the first gRNA, wherein the first gRNA is capable of binding to the first Cas protein and forming a first functional complex; the first functional complex is capable of breaking both strands of the first double-stranded target nucleic acid to form a broken target nucleic acid fragment;

[0028] (4) a first tag primer or a nucleic acid molecule D1 containing a nucleotide sequence encoding the first tag primer, wherein the first tag primer contains a first tag sequence and a first target-binding sequence, the first tag sequence is located upstream or 5' of the first target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the first target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment to form a double-stranded structure, and the first tag sequence does not bind to the target nucleic acid fragment, being in a free, single-stranded state.

[0029] In certain embodiments, the first Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologs or modified versions thereof.

[0030] In certain embodiments, the first Cas protein is capable of cleaving a first double-stranded target nucleic acid and generating a sticky end or a blunt end.

[0031] In certain embodiments, the first Cas protein is a Cas9 protein, such as a Cas9 protein of S. pyogenes (spCas9).

[0032] In certain embodiments, the first Cas protein has an amino acid sequence set forth in SEQ ID NO: 1.

[0033] The sequences and structures of various Cas proteins are well known to those of skill in the art. Currently, a variety of Cas9 proteins and homologues thereof have been reported in a variety of species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Other suitable Cas9 proteins will be apparent to those of skill in the art based on the disclosure herein, for example, the Cas9 proteins disclosed in Chylinski, Rhun, and Charpentier. The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems. (2013) RNA Biology 10:5, 726-737 (the entire contents of which are incorporated herein by reference).

[0034] In some embodiments, the Cas9 is from a species of Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquatus I (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Streptococcus pyogenes (NCBI Ref: NC_017053.1).

[0035] In certain embodiments, the first DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0036] In certain embodiments, the first DNA polymerase is an RNA-dependent DNA polymerase.

[0037] In certain embodiments, the first DNA polymerase is a reverse transcriptase, for example, a reverse transcriptase listed above, for example, the reverse transcriptase of Moloney murine leukemia virus.

[0038] In certain embodiments, the first DNA polymerase has an amino acid sequence set forth in SEQ ID NO: 4.

[0039] In certain embodiments, the first Cas protein is linked to the first DNA polymerase.

[0040] In certain embodiments, the first Cas protein is covalently linked to the first DNA polymerase, with or without a linker.

[0041] In certain embodiments, the linker is a peptide linker, e.g., a flexible peptide linker; for example, the linker has an amino acid sequence set forth in SEQ ID NO: 51.

[0042] In certain embodiments, the first Cas protein is fused to the first DNA polymerase, with or without a peptide linker, to form a first fusion protein.

[0043] In certain embodiments, the first Cas protein is linked or fused to the N-terminus of the first DNA polymerase, optionally with a linker; or, the first Cas protein is linked or fused to the C-terminus of the first DNA polymerase, optionally with a linker.

[0044] In certain embodiments, the first fusion protein has an amino acid sequence set forth in SEQ ID NO: 52.

[0045] In some embodiments, the linker is a peptide linker. In some embodiments, the peptide linker is 5-200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length.

[0046] In certain embodiments, the first fusion protein or the first cas protein can be split into two parts by an intein split system. It is readily understood that the intein split system can split at any amino acid position of the first fusion protein or the first cas protein. For example, in certain embodiments, the intein split system splits internally in the first cas protein. Thus, in certain embodiments, the first cas protein is split into an N-terminal segment and a C-terminal segment. For example, the N-terminal segment and the C-terminal segment of the first cas protein can be fused to the N-terminal segment and the C-terminal segment of an intein, respectively, or to the C-terminal segment and the N-terminal segment of an intein, respectively, and both are capable of reconstituting into an active first cas protein in a cell. In certain embodiments, the N-terminal segment and the C-terminal segment of the first cas protein are each not active in an isolated state, but are capable of reconstituting into an active first cas protein in a cell. Accordingly, in certain embodiments, the nucleic acid molecule Al can be split into two parts, each of which comprises a nucleotide sequence encoding the N-terminal segment and the C-terminal segment of the first cas protein. Further, it is readily understood that in the first fusion protein, the first DNA polymerase can be fused to the N-terminal segment or the C-terminal segment of the first cas protein. In certain embodiments, the first DNA polymerase is fused to the C-terminal segment of the first cas protein.

[0047] In certain embodiments, the first gRNA contains a first guide sequence, and, under conditions permissive for nucleic acid hybridization or annealing, the first guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a first double-stranded target nucleic acid.

[0048] In certain embodiments, the first guide sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0049] In certain embodiments, the first gRNA further contains a first scaffold sequence, which is capable of being recognized and bound by the first Cas protein, thereby forming a first functional complex.

[0050] In certain embodiments, the first scaffold sequence is at least 20 nt in length, e.g., 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0051] In certain embodiments, the first guide sequence is upstream or 5' to the first scaffold sequence.

[0052] In certain embodiments, the first functional complex is capable of cleaving both strands of the first double-stranded target nucleic acid upon binding of the first guide sequence to the first double-stranded target nucleic acid.

[0053] In certain embodiments, the first target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment under conditions that allow nucleic acid hybridization or annealing, and the 3' end is formed as a result of the first functional complex cleaving the first double-stranded target nucleic acid.

[0054] In certain embodiments, the first target-binding sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0055] In certain embodiments, the first tag sequence is at least 4 nt in length, e.g., 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0056] In certain embodiments, the first DNA polymerase is capable of extending the 3' end of the nucleic acid strand with the first tag primer as a template upon hybridization or annealing of the first target-binding sequence to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment. In certain embodiments, the extension forms a first overhang.

[0057] In certain embodiments, the first tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0058] In certain embodiments, the first tag primer is a single-stranded ribonucleic acid and the first DNA polymerase is an RNA-dependent DNA polymerase; or, the first tag primer is a single-stranded deoxyribonucleic acid and the first DNA polymerase is a DNA-dependent DNA polymerase.

[0059] In certain embodiments, the nucleic acid strand to which the first guide sequence binds is different from the nucleic acid strand to which the first target-binding sequence binds. In certain embodiments, the nucleic acid strand to which the first guide sequence binds is the opposite strand of the nucleic acid strand to which the first target-binding sequence binds.

[0060] In certain embodiments, the first tag primer is linked to the first gRNA.

[0061] In certain embodiments, the first tag primer is covalently linked to the first gRNA, with or without a linker.

[0062] In certain embodiments, the first tag primer is linked to the 3' end of the first gRNA, with or without a linker.

[0063] In certain embodiments, the linker is a nucleic acid linker (e.g., a ribonucleic acid linker or a deoxyribonucleic acid linker).

[0064] In certain embodiments, the first tag primer is a single-stranded ribonucleic acid, and it is linked to the 3' end of the first gRNA, with or without a ribonucleic acid linker, to form a first PegRNA.

[0065] In certain embodiments, the nucleic acid molecule Al is capable of expressing the first Cas protein in a cell. In certain embodiments, the nucleic acid molecule Bl is capable of expressing the first DNA polymerase in a cell. In certain embodiments, the nucleic acid molecule Cl is capable of transcribing the first gRNA in a cell. In certain embodiments, the nucleic acid molecule Dl is capable of transcribing the first tag primer in a cell.

[0066] In certain embodiments, the nucleic acid molecule Al is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule Al is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the first Cas protein.

[0067] In certain embodiments, the nucleic acid molecule Bl is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule Bl is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the first DNA polymerase.

[0068] In certain embodiments, the nucleic acid molecule Cl is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule Cl is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the first gRNA.

[0069] In certain embodiments, the nucleic acid molecule Dl is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule Dl is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the first tag primer.

[0070] In certain embodiments, the nucleic acid molecule A1 and the nucleic acid molecule B1 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule A1 and the nucleic acid molecule B1 are capable of expressing the first Cas protein and the first DNA polymerase separately, or a first fusion protein containing the first Cas protein and the first DNA polymerase, in a cell.

[0071] In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are comprised in the same expression vector (e.g., eukaryotic expression vector); in certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in a cell.

[0072] In certain embodiments, two, three or four of the nucleic acid molecules A1, B1, C1 and D1 are comprised in the same expression vector (e.g., eukaryotic expression vector).

[0073] In certain embodiments, the system or kit comprises:

[0074] (M1-1) a first fusion protein containing the first Cas protein and the first DNA polymerase, or, a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and the first DNA polymerase separately, or a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separately; and,

[0075] (M2) a first PegRNA containing the first gRNA and the first tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA.

[0076] In certain embodiments, the system or kit further comprises:

[0077] (5) a second gRNA or a nucleic acid molecule C2 containing a nucleotide sequence encoding the second gRNA, wherein the second gRNA is capable of binding to a second Cas protein and forming a second functional complex; the second functional complex is capable of cleaving two strands of a second double-stranded target nucleic acid to form a cleaved target nucleic acid fragment.

[0078] In certain embodiments, the second Cas protein is the same as or different from the first Cas protein. In certain embodiments, the second Cas protein is the same as the first Cas protein.

[0079] In certain embodiments, the second gRNA comprises a second guide sequence, and, under conditions that allow nucleic acid hybridization or annealing, the second guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a second double-stranded target nucleic acid.

[0080] In certain embodiments, the second functional complex cleaves both strands of the second double-stranded target nucleic acid upon binding of the second guide sequence to the second double-stranded target nucleic acid.

[0081] In certain embodiments, the second guide sequence is different from the first guide sequence.

[0082] In certain embodiments, the second double-stranded target nucleic acid is the same or different from the first double-stranded target nucleic acid.

[0083] In certain embodiments, the second double-stranded target nucleic acid is the same as the first double-stranded target nucleic acid, and the second functional complex cleaves the same double-stranded target nucleic acid at a different location than the first functional complex.

[0084] In certain embodiments, the second functional complex cleaves the same double-stranded target nucleic acid as the first functional complex, and the nucleic acid strand to which the first guide sequence binds is different from the nucleic acid strand to which the second guide sequence binds; in certain embodiments, the nucleic acid strand to which the first guide sequence binds is the opposite strand of the nucleic acid strand to which the second guide sequence binds.

[0085] In certain embodiments, the second guide sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0086] In certain embodiments, the second gRNA further comprises a second scaffold sequence that is capable of being recognized and bound by the second Cas protein, thereby forming a second functional complex.

[0087] In certain embodiments, the second scaffold sequence is at least 20 nt in length, e.g., 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0088] In certain embodiments, the second scaffold sequence is the same or different from the first scaffold sequence; in certain embodiments, the second scaffold sequence is the same as the first scaffold sequence.

[0089] In certain embodiments, the second guide sequence is upstream or 5' of the second scaffold sequence.

[0090] In some embodiments, the nucleic acid molecule C2 is capable of transcribing the second gRNA in a cell.

[0091] In some embodiments, the nucleic acid molecule C2 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule C2 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the second gRNA.

[0092] In some embodiments, the second Cas protein is different from the first Cas protein; and the system or kit further comprises:

[0093] (6) the second Cas protein or a nucleic acid molecule A2 containing a nucleotide sequence encoding the second Cas protein, wherein the second Cas protein is capable of cleaving or nicking a second double-stranded target nucleic acid.

[0094] In some embodiments, the second Cas protein is capable of nicking a second double-stranded target nucleic acid and generating a sticky end or a blunt end.

[0095] In some embodiments, the second Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologues thereof or modified versions thereof.

[0096] In some embodiments, the second Cas protein is a Cas9 protein, for example, a Cas9 protein of S. pyogenes (spCas9).

[0097] In some embodiments, the second Cas protein has an amino acid sequence set forth in SEQ ID NO: 1.

[0098] In some embodiments, the nucleic acid molecule A2 is capable of expressing the second Cas protein in a cell.

[0099] In some embodiments, the nucleic acid molecule A2 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule A2 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the second Cas protein.

[0100] In some embodiments, the system or kit further comprises:

[0101] (7) a second tag primer or a nucleic acid molecule D2 containing a nucleotide sequence encoding the second tag primer, wherein the second tag primer contains a second tag sequence and a second target-binding sequence, the second tag sequence is located upstream or 5’ of the second target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the second target-binding sequence is capable of hybridizing or annealing to the 3’ end of one nucleic acid strand of the fragmented target nucleic acid fragment, forming a double-stranded structure, and the second tag sequence does not bind to the target nucleic acid fragment, being in a free, single-stranded state.

[0102] In some embodiments, under conditions permitting nucleic acid hybridization or annealing, the second target-binding sequence is capable of hybridizing or annealing to the 3’ end of one nucleic acid strand of the fragmented target nucleic acid fragment, and the 3’ end is formed as a result of the second functional complex fragmenting the second double-stranded target nucleic acid.

[0103] In some embodiments, the second target-binding sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0104] In some embodiments, the second target-binding sequence is different from the first target-binding sequence. In some embodiments, the nucleic acid strand to which the second target-binding sequence binds is different from the nucleic acid strand to which the first target-binding sequence binds. In some embodiments, the nucleic acid strand to which the second target-binding sequence binds is the opposite strand of the nucleic acid strand to which the first target-binding sequence binds.

[0105] In certain embodiments, the second tag sequence is at least 4 nt in length, e.g., 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0106] In certain embodiments, the second tag sequence is the same as or different from the first tag sequence. In certain embodiments, the second tag sequence is different from the first tag sequence.

[0107] In certain embodiments, upon hybridization or annealing of the second target-binding sequence to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment, a second DNA polymerase is capable of extending the 3' end of the nucleic acid strand using the second tag primer as a template. In certain embodiments, the extension forms a second overhang.

[0108] In certain embodiments, the second DNA polymerase is the same as or different from the first DNA polymerase. In certain embodiments, the second DNA polymerase is the same as the first DNA polymerase.

[0109] In certain embodiments, the second tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0110] In certain embodiments, the second tag primer is a single-stranded ribonucleic acid and the second DNA polymerase is an RNA-dependent DNA polymerase; or, the second tag primer is a single-stranded deoxyribonucleic acid and the second DNA polymerase is a DNA-dependent DNA polymerase.

[0111] In certain embodiments, the nucleic acid strand to which the second guide sequence binds is different from the nucleic acid strand to which the second target-binding sequence binds. In certain embodiments, the nucleic acid strand to which the second guide sequence binds is the opposite strand of the nucleic acid strand to which the second target-binding sequence binds.

[0112] In certain embodiments, the second guide sequence binds to the same nucleic acid strand as the first target-binding sequence, and the binding position of the second guide sequence is upstream or 5' of the binding position of the first target-binding sequence.

[0113] In certain embodiments, the first guide sequence binds to the same nucleic acid strand as the second target-binding sequence, and the binding position of the first guide sequence is upstream or 5' of the binding position of the second target-binding sequence.

[0114] In certain embodiments, the first and second overhangs are comprised on the same target nucleic acid fragment and are located on opposite nucleic acid strands from each other.

[0115] In certain embodiments, the nucleic acid molecule D2 is capable of transcribing the second tag primer in a cell.

[0116] In certain embodiments, the nucleic acid molecule D2 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule D2 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the second tag primer.

[0117] In certain embodiments, the second DNA polymerase is different from the first DNA polymerase; and the system or kit further comprises:

[0118] (8) the second DNA polymerase or a nucleic acid molecule B2 containing a nucleotide sequence encoding the second DNA polymerase.

[0119] In certain embodiments, the second DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0120] In certain embodiments, the second DNA polymerase is an RNA-dependent DNA polymerase.

[0121] In certain embodiments, the second DNA polymerase is a reverse transcriptase, such as the reverse transcriptases listed above, such as the reverse transcriptase of Moloney murine leukemia virus.

[0122] In certain embodiments, the second DNA polymerase has the amino acid sequence set forth in SEQ ID NO: 4.

[0123] In certain embodiments, the nucleic acid molecule B2 is capable of expressing the second DNA polymerase in a cell.

[0124] In certain embodiments, the nucleic acid molecule B2 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule B2 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the second DNA polymerase.

[0125] In certain embodiments, the second tag primer is linked to the second gRNA.

[0126] In certain embodiments, the second tag primer is covalently linked to the second gRNA with or without a linker.

[0127] In certain embodiments, the second tag primer is optionally linked to the 3’ end of the second gRNA with a linker.

[0128] In some embodiments, the linker is a nucleic acid linker (e.g., a ribonucleic acid linker or a deoxyribonucleic acid linker).

[0129] In some embodiments, the second tag primer is a single-stranded ribonucleic acid, and it is connected to the 3' end of the second gRNA with or without a ribonucleic acid linker, forming a second PegRNA.

[0130] In some embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., a eukaryotic expression vector); in some embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of transcribing a second PegRNA containing the second gRNA and the second tag primer in a cell.

[0131] In some embodiments, the system or kit comprises: a second PegRNA containing the second gRNA and the second tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA.

[0132] In some embodiments, the second Cas protein is separate from or linked to the second DNA polymerase.

[0133] In some embodiments, the second Cas protein is covalently linked to the second DNA polymerase with or without a linker.

[0134] In some embodiments, the linker is a peptide linker, e.g., a flexible peptide linker; for example, the linker has the amino acid sequence set forth in SEQ ID NO: 51.

[0135] In some embodiments, the second Cas protein is fused to the second DNA polymerase with or without a peptide linker, forming a second fusion protein.

[0136] In some embodiments, the second Cas protein is optionally linked or fused to the N-terminus of the second DNA polymerase with a linker; or the second Cas protein is optionally linked or fused to the C-terminus of the second DNA polymerase with a linker.

[0137] In some embodiments, the second fusion protein has the amino acid sequence set forth in SEQ ID NO: 52.

[0138] In certain embodiments, the second fusion protein or the second cas protein can be split into two parts by an intein split system. It is readily understood that the intein split system can split at any amino acid position of the second fusion protein or the second cas protein. For example, in certain embodiments, the intein split system splits internally in the second cas protein. Thus, in certain embodiments, the second cas protein is split into an N-terminal segment and a C-terminal segment. For example, the N-terminal segment and the C-terminal segment of the second cas protein can be fused to the N-terminal segment and the C-terminal segment of an intein, respectively, or to the C-terminal segment and the N-terminal segment of an intein, respectively, and both can reconstitute into an active second cas protein in a cell. In certain embodiments, the N-terminal segment and the C-terminal segment of the second cas protein are not active in an isolated state, but can reconstitute into an active second cas protein in a cell. Accordingly, in certain embodiments, the nucleic acid molecule Al can be split into two parts, which respectively comprise nucleotide sequences encoding the N-terminal segment and the C-terminal segment of the second cas protein. Furthermore, it is readily understood that in the second fusion protein, the second DNA polymerase can be fused to the N-terminal segment or the C-terminal segment of the second cas protein. In certain embodiments, the second DNA polymerase is fused to the C-terminal segment of the second cas protein.

[0139] In certain embodiments, the nucleic acid molecule A2 and the nucleic acid molecule B2 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule A2 and the nucleic acid molecule B2 are capable of expressing the second Cas protein and the second DNA polymerase, separately, or a second fusion protein comprising the second Cas protein and the second DNA polymerase, in a cell.

[0140] In certain embodiments, the system or the kit comprises a second fusion protein comprising the second Cas protein and the second DNA polymerase, or a nucleic acid molecule comprising a nucleotide sequence encoding the second fusion protein. Alternatively, the second Cas protein and the second DNA polymerase, separately, or nucleic acid molecules capable of expressing the second Cas protein and the second DNA polymerase, separately.

[0141] In certain embodiments, the first and second Cas proteins are the same Cas protein, the first and second DNA polymerases are the same DNA polymerase; and the system or the kit comprises:

[0142] (M1-1) a first fusion protein comprising the first Cas protein and the first DNA polymerase, or, a nucleic acid molecule comprising a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and the first DNA polymerase separated, or, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separated;

[0143] (M2) a first PegRNA comprising the first gRNA and the first tag primer, or, a nucleic acid molecule comprising a nucleotide sequence encoding the first PegRNA;

[0144] (M3) a second PegRNA comprising the second gRNA and the second tag primer, or, a nucleic acid molecule comprising a nucleotide sequence encoding the second PegRNA.

[0145] In certain embodiments, the system or kit further comprises a nucleic acid vector.

[0146] In certain embodiments, the nucleic acid vector is double-stranded.

[0147] In certain embodiments, the nucleic acid vector is a circular double-stranded vector.

[0148] In certain embodiments, the nucleic acid vector comprises a first guide-binding sequence capable of hybridizing or annealing to the first guide sequence (e.g., a complement of the first guide sequence), and / or, a second guide-binding sequence capable of hybridizing or annealing to the second guide sequence (e.g., a complement of the second guide sequence). In certain embodiments, the nucleic acid vector further comprises a restriction enzyme site between the first guide-binding sequence and the second guide-binding sequence.

[0149] In certain embodiments, the first guide-binding sequence and the second guide-binding sequence are located on opposite strands of the nucleic acid vector.

[0150] In certain embodiments, the nucleic acid vector further comprises a first PAM sequence recognized by the first Cas protein, and / or, a second PAM sequence recognized by the second Cas protein.

[0151] In certain embodiments, the first functional complex is capable of binding to and cleaving the nucleic acid vector through the first guide-binding sequence and the first PAM sequence; and / or, the second functional complex is capable of binding to and cleaving the nucleic acid vector through the second guide-binding sequence and the second PAM sequence.

[0152] In certain embodiments, the nucleic acid vector further comprises a gene of interest.

[0153] In certain embodiments, the gene of interest is located between the first guide binding sequence and the second guide binding sequence.

[0154] In certain embodiments, the first functional complex and the second functional complex cleave the nucleic acid vector, resulting in a nucleic acid fragment containing the gene of interest.

[0155] In certain embodiments, under conditions permitting nucleic acid hybridization or annealing, the first tag primer is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the nucleic acid fragment via the first target binding sequence, forming a double-stranded structure, and the first tag sequence of the first tag primer is in a free state; in certain embodiments, the nucleic acid strand to which the first target binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand containing the first guide binding sequence.

[0156] In certain embodiments, under conditions permitting nucleic acid hybridization or annealing, the second tag primer is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the nucleic acid fragment via the second target binding sequence, forming a double-stranded structure, and the second tag sequence of the second tag primer is in a free state; in certain embodiments, the nucleic acid strand to which the second target binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand containing the second guide binding sequence.

[0157] In certain embodiments, the nucleic acid strand to which the first target binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand to which the second target binding sequence hybridizes or anneals.

[0158] In certain embodiments, the nucleic acid vector further comprises a first target sequence; wherein, under conditions permitting nucleic acid hybridization or annealing, the first tag primer is capable of hybridizing or annealing to the first target sequence via the first target binding sequence, forming a double-stranded structure, and the first tag sequence of the first tag primer is in a free state. In certain embodiments, the first target sequence is located between the first guide binding sequence and the second guide binding sequence. In certain embodiments, the first target sequence is located on the opposite strand of the first guide binding sequence. In certain embodiments, after the first functional complex cleaves the nucleic acid vector, the nucleic acid strand containing the first target sequence is capable of extension with the first tag primer annealed to the first target sequence as a template (in certain embodiments, forming a first overhang). In certain embodiments, the site at which the first functional complex cleaves the nucleic acid vector is located 3' to the first target sequence or a 3' portion thereof. In certain embodiments, the first target sequence is located 3' to the nucleic acid strand of the nucleic acid fragment containing the gene of interest.

[0159] and / or,

[0160] The nucleic acid carrier further comprises a second target sequence; wherein the second tag primer is capable of hybridizing or annealing to the second target sequence via the second target binding sequence to form a double-stranded structure under conditions permitting nucleic acid hybridization or annealing, and the second tag sequence of the second tag primer is in a free state; in certain embodiments, the second target sequence is located between the first guide binding sequence and the second guide binding sequence. In certain embodiments, the second target sequence is located on the opposite strand of the second guide binding sequence. In certain embodiments, after the second functional complex cleaves the nucleic acid carrier, the nucleic acid strand containing the second target sequence is capable of extension with the second tag primer annealed to the second target sequence as a template (in certain embodiments, forming a second overhang). In certain embodiments, the site at which the second functional complex cleaves the nucleic acid carrier is located at the 3' end or 3' portion of the second target sequence; in certain embodiments, the second target sequence is located at the 3' end of the nucleic acid strand containing the gene of interest.

[0161] In certain embodiments, the nucleic acid strand containing the first target sequence is located on the opposite strand of the nucleic acid strand containing the second target sequence.

[0162] In certain embodiments, the nucleic acid carrier further comprises a restriction enzyme site between the first target sequence and the second target sequence.

[0163] In certain embodiments, the nucleic acid carrier further comprises a gene of interest between the first target sequence and the second target sequence.

[0164] In certain embodiments, the system or kit further comprises:

[0165] (9) a third gRNA or a nucleic acid molecule C3 containing a nucleotide sequence encoding the third gRNA, wherein the third gRNA is capable of binding to a third Cas protein to form a third functional complex; the third functional complex is capable of cleaving both strands of a third double-stranded target nucleic acid to form cleaved nucleotide fragments a1 and a2.

[0166] In certain embodiments, the third Cas protein is the same as or different from the first Cas protein or the second Cas protein; in certain embodiments, the first, second, and third Cas proteins are the same Cas protein.

[0167] In certain embodiments, the third gRNA contains a third guide sequence, and the third guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a third double-stranded target nucleic acid under conditions permitting nucleic acid hybridization or annealing.

[0168] In certain embodiments, the third functional complex cleaves both strands of the third double- stranded target nucleic acid upon binding of the third guide sequence to the third double- stranded target nucleic acid.

[0169] In certain embodiments, the third guide sequence is the same as or different from the first guide sequence or the second guide sequence. In certain embodiments, the first, second, and third guide sequences are different from each other.

[0170] In certain embodiments, the third double- stranded target nucleic acid is the same as or different from the first double- stranded target nucleic acid or the second double- stranded target nucleic acid. In certain embodiments, the second double- stranded target nucleic acid is the same as the first double- stranded target nucleic acid, and the third double- stranded target nucleic acid is different from the first and second double- stranded target nucleic acids. In certain embodiments, the third double- stranded target nucleic acid is genomic DNA.

[0171] In certain embodiments, the third guide sequence is at least 5 nt in length, such as 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0172] In certain embodiments, the third gRNA further comprises a third scaffold sequence that is recognized and bound by the third Cas protein to form a third functional complex.

[0173] In certain embodiments, the third scaffold sequence is at least 20 nt in length, such as 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0174] In certain embodiments, the third scaffold sequence is the same as or different from the first scaffold sequence or the second scaffold sequence. In certain embodiments, the first, second, and third scaffold sequences are the same.

[0175] In certain embodiments, the third guide sequence is upstream or 5' of the third scaffold sequence.

[0176] In certain embodiments, the nucleic acid molecule C3 is capable of transcribing the third gRNA in a cell.

[0177] In certain embodiments, the nucleic acid molecule C3 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule C3 is an expression vector (e.g., a eukaryotic expression vector) comprising a nucleotide sequence encoding the third gRNA.

[0178] In certain embodiments, the third functional complex is capable of cleaving both strands of a third double-stranded target nucleic acid to form cleaved nucleotide fragments al and a2.

[0179] In certain embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to cleaved nucleotide fragment al under conditions that allow nucleic acid hybridization or annealing. In certain embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to cleaved nucleotide fragment al at a terminus formed by the third functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to a 3’ end or 3’ portion of one nucleic acid strand of cleaved nucleotide fragment al, and the 3’ end or 3’ portion is formed by the third functional complex cleaving the third double-stranded target nucleic acid.

[0180] In certain embodiments, the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to a 3’ portion of one nucleic acid strand of cleaved nucleotide fragment al, and there is a first spacer region between the 3’ portion of the nucleotide fragment al and the cleaved terminus formed by the third double-stranded target nucleic acid.

[0181] In certain embodiments, the first spacer region is 1 nt to 200 nt in length, e.g., 1-10 nt, 10-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, or 100-200 nt.

[0182] In certain embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a2 under conditions that allow nucleic acid hybridization or annealing. In certain embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a2 at a terminus formed by the third functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to a 3’ end or 3’ portion of one nucleic acid strand of cleaved nucleotide fragment a2, and the 3’ end or 3’ portion is formed by the third functional complex cleaving the third double-stranded target nucleic acid.

[0183] In certain embodiments, the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to a 3’ portion of one nucleic acid strand of cleaved nucleotide fragment a2, and there is a second spacer region between the 3’ portion of the nucleotide fragment a2 and the cleaved terminus formed by the third double-stranded target nucleic acid.

[0184] In certain embodiments, the second spacer region has a length of 1 nt-200 nt, such as 1-10 nt, 10-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, or 100-200 nt.

[0185] In certain embodiments, the third Cas protein is different from the first Cas protein or the second Cas protein; and the system or kit further comprises:

[0186] (10) the third Cas protein or a nucleic acid molecule A3 containing a nucleotide sequence encoding the third Cas protein, wherein the third Cas protein is capable of cleaving or nicking a third double-stranded target nucleic acid.

[0187] In certain embodiments, the third Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologues or modified versions thereof.

[0188] In certain embodiments, the third Cas protein is capable of nicking a third double-stranded target nucleic acid and generating a sticky end or a blunt end.

[0189] In certain embodiments, the third Cas protein is a Cas9 protein, such as a Cas9 protein of S. pyogenes (spCas9).

[0190] In certain embodiments, the third Cas protein has an amino acid sequence as set forth in SEQ ID NO: 1.

[0191] In some embodiments, the nucleic acid molecule A3 is capable of expressing the third Cas protein in a cell.

[0192] In some embodiments, the nucleic acid molecule A3 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule A3 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the third Cas protein.

[0193] In some embodiments, the first, second and third Cas proteins are the same Cas protein, the first and second DNA polymerases are the same DNA polymerase; and, the system or kit comprises:

[0194] (M1-1) a first fusion protein containing the first Cas protein and the first DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and the first DNA polymerase are separated, or a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase are separated;

[0195] (M2) a first PegRNA containing the first gRNA and the first tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA;

[0196] (M3) a second PegRNA containing the second gRNA and the second tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA;

[0197] (M4) the third gRNA or a nucleic acid molecule containing a nucleotide sequence encoding the third gRNA.

[0198] In some embodiments, the system or kit further comprises: a nucleic acid vector as defined in the foregoing.

[0199] In some embodiments, the system or kit further comprises:

[0200] (11) a third tag primer or a nucleic acid molecule D3 containing a nucleotide sequence encoding the third tag primer, wherein the third tag primer contains a third tag sequence and a third target binding sequence, the third tag sequence is located upstream or 5' end of the third target binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the third target binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the broken nucleotide fragment a1 or a2, forming a double-stranded structure, and the third tag sequence is not bound to the nucleotide fragment a1 or a2, being in a free single-stranded state.

[0201] In certain embodiments, the third target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the cleaved nucleotide fragment al or a2 under conditions that allow nucleic acid hybridization or annealing, and the 3' end is formed as a result of cleavage of a third double-stranded target nucleic acid by the third functional complex.

[0202] In certain embodiments, the nucleic acid strand to which the third target-binding sequence binds is different from the nucleic acid strand to which the third guide sequence binds; in certain embodiments, the nucleic acid strand to which the third target-binding sequence binds is the opposite strand of the nucleic acid strand to which the third guide sequence binds.

[0203] In certain embodiments, the third target-binding sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0204] In certain embodiments, the third target-binding sequence is different from the first or second target-binding sequence.

[0205] In certain embodiments, upon hybridization or annealing of the third target-binding sequence to the 3' end of one nucleic acid strand of the cleaved nucleotide fragment al or a2, a third DNA polymerase is capable of extending the 3' end of the nucleic acid strand using the third tag primer as a template; in certain embodiments, the extension forms a third overhang.

[0206] In certain embodiments, the third tag sequence is at least 4 nt in length, e.g., 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0207] In certain embodiments, the third tag sequence is the same as or different from the first or second tag sequence. In certain embodiments, the third tag sequence is different from the first or second tag sequence.

[0208] In certain embodiments, the complement of the third tag sequence or the third overhang is capable of hybridizing or annealing to a nucleic acid strand containing the first overhang or the second overhang under conditions that allow nucleic acid hybridization or annealing; in certain embodiments, the complement of the third tag sequence or the third overhang hybridizes or anneals to the first overhang or the second overhang or nucleotides upstream thereof.

[0209] In certain embodiments, the third DNA polymerase is the same as or different from the first or second DNA polymerase; in certain embodiments, the first, second, and third DNA polymerases are the same DNA polymerase.

[0210] In certain embodiments, the third tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0211] In certain embodiments, the third tag primer is a single-stranded ribonucleic acid and the third DNA polymerase is an RNA-dependent DNA polymerase; or, the third tag primer is a single-stranded deoxyribonucleic acid and the third DNA polymerase is a DNA-dependent DNA polymerase.

[0212] In certain embodiments, the nucleic acid molecule D3 is capable of transcribing the third tag primer in a cell.

[0213] In certain embodiments, the nucleic acid molecule D3 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule D3 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the third tag primer.

[0214] In certain embodiments, the third DNA polymerase is different from the first or second DNA polymerase; and the system or kit further comprises:

[0215] (12) the third DNA polymerase or a nucleic acid molecule B3 containing a nucleotide sequence encoding the third DNA polymerase.

[0216] In certain embodiments, the third DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0217] In certain embodiments, the third DNA polymerase is an RNA-dependent DNA polymerase.

[0218] In certain embodiments, the third DNA polymerase is a reverse transcriptase, such as the reverse transcriptases listed above, such as the reverse transcriptase of Moloney murine leukemia virus.

[0219] In certain embodiments, the third DNA polymerase has the amino acid sequence set forth in SEQ ID NO: 4.

[0220] In certain embodiments, the nucleic acid molecule B3 is capable of expressing the third DNA polymerase in a cell.

[0221] In some embodiments, the nucleic acid molecule B3 is contained in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule B3 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the third DNA polymerase.

[0222] In some implementations, the third tag primer is linked to the third gRNA.

[0223] In some implementations, the third tag primer is covalently linked to the third gRNA via a adapter or without a adapter.

[0224] In some implementations, the third tag primer is optionally linked to the 3' end of the third gRNA via a adapter.

[0225] In some embodiments, the adapter is a nucleic acid adapter (e.g., a ribonucleic acid adapter or a deoxyribonucleic acid adapter).

[0226] In some embodiments, the third tag primer is a single-stranded ribonucleic acid (RNA), and it is attached to the 3' end of the third gRNA via or without a ribonucleic acid adapter to form a third PegRNA.

[0227] In some embodiments, nucleic acid molecules C3 and D3 are contained in the same expression vector (e.g., a eukaryotic expression vector). In some embodiments, nucleic acid molecules C3 and D3 are capable of being transcribed in cells into a third PegRNA containing the third gRNA and the third tag primer.

[0228] In some embodiments, the system or kit comprises: a third PegRNA containing the third gRNA and the third tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the third PegRNA.

[0229] In some embodiments, the third Cas protein is either separate from or linked to the third DNA polymerase.

[0230] In some embodiments, the third Cas protein is covalently linked to the third DNA polymerase, either via a adapter or without a adapter.

[0231] In some embodiments, the adapter is a peptide adapter, such as a flexible peptide adapter; for example, the adapter has the amino acid sequence shown in SEQ ID NO:51.

[0232] In some embodiments, the third Cas protein is fused to the third DNA polymerase via a peptide linker or without a peptide linker to form a third fusion protein.

[0233] In some embodiments, the third Cas protein is optionally linked or fused to the N-terminus of the third DNA polymerase via a linker; or, the third Cas protein is optionally linked or fused to the C-terminus of the third DNA polymerase via a linker.

[0234] In some embodiments, the third fusion protein has an amino acid sequence as set forth in SEQ ID NO: 52.

[0235] In some embodiments, the third fusion protein or the third cas protein can be split into two parts by an intein split system. It is readily understood that the intein split system can split at any amino acid position of the third fusion protein or the third cas protein. For example, in some embodiments, the intein split system splits internally in the third cas protein. Accordingly, in some embodiments, the third cas protein is split into an N-terminal segment and a C-terminal segment. For example, the N-terminal segment and the C-terminal segment of the third cas protein can be fused to the N-terminal segment and the C-terminal segment of an intein, respectively (or to the C-terminal segment and the N-terminal segment of an intein, respectively), and both are capable of reconstituting into an active third cas protein in a cell. In some embodiments, the N-terminal segment and the C-terminal segment of the third cas protein are each not active in an isolated state, but are capable of reconstituting into an active third cas protein in a cell. Accordingly, in some embodiments, the nucleic acid molecule A1 can be split into two parts, which respectively comprise nucleotide sequences encoding the N-terminal segment and the C-terminal segment of the third cas protein. Further, it is readily understood that in the third fusion protein, the third DNA polymerase can be fused to the N-terminal segment or the C-terminal segment of the third cas protein. In some embodiments, the third DNA polymerase is fused to the C-terminal segment of the third cas protein.

[0236] In some embodiments, the nucleic acid molecule A3 and the nucleic acid molecule B3 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In some embodiments, the nucleic acid molecule A3 and the nucleic acid molecule B3 are capable of expressing the third Cas protein and the third DNA polymerase separately, or the third fusion protein containing the third Cas protein and the third DNA polymerase, in a cell.

[0237] In some embodiments, the system or the kit comprises a third fusion protein containing the third Cas protein and the third DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the third fusion protein. Alternatively, the third Cas protein and the third DNA polymerase separately, or a nucleic acid molecule capable of expressing the third Cas protein and the third DNA polymerase separately.

[0238] In certain embodiments, the first, second, and third Cas proteins are the same Cas protein, the first, second, and third DNA polymerases are the same DNA polymerase; and, the system or kit comprises:

[0239] (M1-1) a first fusion protein containing the first Cas protein and the first DNA polymerase, or, a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and the first DNA polymerase separated, or, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separated;

[0240] (M2) a first PegRNA containing the first gRNA and the first tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA;

[0241] (M3) a second PegRNA containing the second gRNA and the second tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA;

[0242] (M4) a third PegRNA containing the third gRNA and the third tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the third PegRNA.

[0243] In certain embodiments, the system or kit further comprises: a nucleic acid vector as defined in the foregoing.

[0244] In certain embodiments, the system or kit further comprises:

[0245] (13) a fourth gRNA or a nucleic acid molecule C4 containing a nucleotide sequence encoding the fourth gRNA, wherein the fourth gRNA is capable of binding to a fourth Cas protein, and forming a fourth functional complex; the fourth functional complex is capable of cleaving two strands of a fourth double-stranded target nucleic acid, forming cleaved target nucleic acid fragments b1 and b2.

[0246] In certain embodiments, the fourth Cas protein is the same as or different from the first, second, or third Cas protein; in certain embodiments, the first, second, third, and fourth Cas proteins are the same Cas protein.

[0247] In certain embodiments, the fourth gRNA contains a fourth guide sequence, and, under conditions permitting nucleic acid hybridization or annealing, the fourth guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a fourth double-stranded target nucleic acid.

[0248] In certain embodiments, the fourth functional complex cleaves both strands of the fourth double-stranded target nucleic acid upon binding of the fourth guide sequence to the fourth double-stranded target nucleic acid.

[0249] In certain embodiments, the fourth guide sequence is the same as or different from the first, second, or third guide sequence. In certain embodiments, the first, second, third, and fourth guide sequences are different from each other.

[0250] In certain embodiments, the fourth double-stranded target nucleic acid is the same as or different from the first, second, or third double-stranded target nucleic acid. In certain embodiments, the second double-stranded target nucleic acid is the same as the first double-stranded target nucleic acid, and the fourth double-stranded target nucleic acid is the same as the third double-stranded target nucleic acid but different from the first or second double-stranded target nucleic acid. In certain embodiments, the fourth functional complex cleaves the same double-stranded target nucleic acid as the third functional complex, but at a different location.

[0251] In certain embodiments, the fourth functional complex cleaves the same double-stranded target nucleic acid as the third functional complex, and the nucleic acid strand to which the fourth guide sequence binds is different from the nucleic acid strand to which the third guide sequence binds. In certain embodiments, the nucleic acid strand to which the fourth guide sequence binds is the opposite strand of the nucleic acid strand to which the third guide sequence binds.

[0252] In certain embodiments, the fourth double-stranded target nucleic acid is genomic DNA.

[0253] In certain embodiments, the fourth guide sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0254] In certain embodiments, the fourth gRNA further comprises a fourth scaffold sequence that is recognized and bound by the fourth Cas protein to form the fourth functional complex.

[0255] In certain embodiments, the fourth scaffold sequence is at least 20 nt in length, e.g., 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0256] In certain embodiments, the fourth scaffold sequence is the same as or different from the first, second, or third scaffold sequence. In certain embodiments, the first, second, third, and fourth scaffold sequences are the same.

[0257] In some embodiments, the fourth guide sequence is located upstream or 5' of the fourth scaffold sequence.

[0258] In some embodiments, the nucleic acid molecule C4 is capable of transcribing the fourth gRNA in a cell.

[0259] In some embodiments, the nucleic acid molecule C4 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule C4 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the fourth gRNA.

[0260] In some embodiments, the fourth double-stranded target nucleic acid is identical to the third double-stranded target nucleic acid, and the third and fourth functional complexes cleave the same double-stranded target nucleic acid at different locations to form cleaved nucleotide fragments a1, a2 and a3; wherein, prior to cleavage, nucleotide fragments a1, a2 and a3 are arranged in sequence in the same double-stranded target nucleic acid (i.e., nucleotide fragment a1 is connected to nucleotide fragment a3 via nucleotide fragment a2); in some embodiments, the third and fourth functional complexes cause the separation of nucleotide fragments a1 and a2, respectively, and the separation of nucleotide fragments a2 and a3.

[0261] In some embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a1 under conditions that allow nucleic acid hybridization or annealing; in some embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a1 at the end formed by the cleavage of the third double-stranded target nucleic acid by the third functional complex. In some embodiments, the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of cleaved nucleotide fragment a1, and the 3' end or 3' portion is formed as a result of the cleavage of the third double-stranded target nucleic acid by the third functional complex.

[0262] In some embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a3 under conditions that allow nucleic acid hybridization or annealing; in some embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to cleaved nucleotide fragment a3 at the end formed by the cleavage of the third double-stranded target nucleic acid by the fourth functional complex. In some embodiments, the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of cleaved nucleotide fragment a3, and the 3' end or 3' portion is formed as a result of the cleavage of the third double-stranded target nucleic acid by the fourth functional complex.

[0263] In certain embodiments, the fourth Cas protein is different from the first, second, or third Cas protein; and the system or kit further comprises:

[0264] (14) the fourth Cas protein or a nucleic acid molecule A4 containing a nucleotide sequence encoding the fourth Cas protein, wherein the fourth Cas protein is capable of cleaving or nicking a fourth double-stranded target nucleic acid.

[0265] In certain embodiments, the fourth Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologues thereof or modified versions thereof.

[0266] In certain embodiments, the fourth Cas protein is capable of nicking a fourth double-stranded target nucleic acid and generating a sticky end or a blunt end.

[0267] In certain embodiments, the fourth Cas protein is a Cas9 protein, such as a Cas9 protein of S. pyogenes (spCas9).

[0268] In certain embodiments, the fourth Cas protein has an amino acid sequence as set forth in SEQ ID NO: 1.

[0269] In certain embodiments, the nucleic acid molecule A4 is capable of expressing the fourth Cas protein in a cell.

[0270] In certain embodiments, the nucleic acid molecule A4 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule A4 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the fourth Cas protein.

[0271] In certain embodiments, the first, second, third, and fourth Cas proteins are the same Cas protein, the first and second DNA polymerases (and optionally the third DNA polymerase) are the same DNA polymerase; and, the system or kit comprises:

[0272] (M1-1) a first fusion protein containing the first Cas protein and the first DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and first DNA polymerase are separated, or a nucleic acid molecule capable of expressing the first Cas protein and first DNA polymerase are separated;

[0273] (M2) a first PegRNA containing the first gRNA and first tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA;

[0274] (M3) a second PegRNA containing the second gRNA and second tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA;

[0275] (M4) the third gRNA or a nucleic acid molecule containing a nucleotide sequence encoding the third gRNA; or, a third PegRNA containing the third gRNA and third tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the third PegRNA;

[0276] (M5) the fourth gRNA or a nucleic acid molecule containing a nucleotide sequence encoding the fourth gRNA.

[0277] In certain embodiments, the system or kit further comprises: a nucleic acid vector as defined in the preceding.

[0278] In certain embodiments, the system or kit further comprises:

[0279] (15) a fourth tag primer or a nucleic acid molecule D4 comprising a nucleotide sequence encoding the fourth tag primer, wherein the fourth tag primer comprises a fourth tag sequence and a fourth target-binding sequence, the fourth tag sequence is located upstream or 5' to the fourth target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the fourth target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment b1 or b2, forming a double-stranded structure, and the fourth tag sequence does not bind to the target nucleic acid fragment b1 or b2, being in a free, single-stranded state.

[0280] In certain embodiments, under conditions permitting nucleic acid hybridization or annealing, the fourth target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment b1 or b2, and the 3' end is formed as a result of the fourth functional complex fragmenting the fourth double-stranded target nucleic acid.

[0281] In certain embodiments, the nucleic acid strand to which the fourth target-binding sequence binds is different from the nucleic acid strand to which the fourth guide sequence binds. In certain embodiments, the nucleic acid strand to which the fourth target-binding sequence binds is the opposite strand of the nucleic acid strand to which the fourth guide sequence binds.

[0282] In certain embodiments, the fourth target-binding sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0283] In certain embodiments, the fourth target-binding sequence is different from the first, second, or third target-binding sequence. In certain embodiments, the nucleic acid strand to which the fourth target-binding sequence binds is different from the nucleic acid strand to which the third target-binding sequence binds. In certain embodiments, the nucleic acid strand to which the fourth target-binding sequence binds is the opposite strand of the nucleic acid strand to which the third target-binding sequence binds.

[0284] In certain embodiments, after the fourth target-binding sequence hybridizes or anneals to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment b1 or b2, a fourth DNA polymerase is capable of extending the 3' end of the nucleic acid strand using the fourth tag primer as a template. In certain embodiments, the extension forms a fourth overhang.

[0285] In certain embodiments, the fourth tag sequence is at least 4 nt in length, e.g., 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0286] In certain embodiments, the fourth tag sequence is the same as or different from the first, second, or third tag sequence. In certain embodiments, the fourth tag sequence is different from the first, second, or third tag sequence.

[0287] In certain embodiments, the fourth DNA polymerase is the same as or different from the first, second, or third DNA polymerase. In certain embodiments, the first, second, third, and fourth DNA polymerases are the same DNA polymerase.

[0288] In certain embodiments, the fourth tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0289] In certain embodiments, the fourth tag primer is a single-stranded ribonucleic acid and the fourth DNA polymerase is an RNA-dependent DNA polymerase. Alternatively, the fourth tag primer is a single-stranded deoxyribonucleic acid and the fourth DNA polymerase is a DNA-dependent DNA polymerase.

[0290] In certain embodiments, the fourth guide sequence binds to the same nucleic acid strand as the third target-binding sequence, and the binding position of the third target-binding sequence is located upstream or 5' of the binding position of the fourth guide sequence.

[0291] In certain embodiments, the third guide sequence binds to the same nucleic acid strand as the fourth target-binding sequence, and the binding position of the fourth target-binding sequence is located upstream or 5' of the binding position of the third guide sequence.

[0292] In certain embodiments, the third overhang and the fourth overhang are comprised on different target nucleic acid fragments, and in certain embodiments, are located on opposite nucleic acid strands from each other.

[0293] In certain embodiments, the complement of the fourth tag sequence or the fourth overhang is capable of hybridizing or annealing to the nucleic acid strand containing the first overhang or the second overhang under conditions that allow nucleic acid hybridization or annealing. In certain embodiments, the complement of the fourth tag sequence or the fourth overhang is capable of hybridizing or annealing to the first overhang or the second overhang or the nucleotide sequence upstream thereof.

[0294] In certain embodiments, the complement of the third tag sequence or the third overhang is capable of hybridizing or annealing to the first overhang or the nucleotide sequence upstream thereof, and the complement of the fourth tag sequence or the fourth overhang is capable of hybridizing or annealing to the second overhang or the nucleotide sequence upstream thereof, under conditions permissive for nucleic acid hybridization or annealing; or, the complement of the third tag sequence or the third overhang is capable of hybridizing or annealing to the second overhang or the nucleotide sequence upstream thereof, and the complement of the fourth tag sequence or the fourth overhang is capable of hybridizing or annealing to the first overhang or the nucleotide sequence upstream thereof, under conditions permissive for nucleic acid hybridization or annealing.

[0295] In certain embodiments, the nucleic acid molecule D4 is capable of transcribing the fourth tag primer in a cell.

[0296] In certain embodiments, the nucleic acid molecule D4 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule D4 is an expression vector (e.g., a eukaryotic expression vector) containing a nucleotide sequence encoding the fourth tag primer.

[0297] In certain embodiments, the fourth DNA polymerase is different from the first, second, or third DNA polymerase; and the system or kit further comprises:

[0298] (16) the fourth DNA polymerase or a nucleic acid molecule B4 containing a nucleotide sequence encoding the fourth DNA polymerase.

[0299] In certain embodiments, the fourth DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0300] In certain embodiments, the fourth DNA polymerase is an RNA-dependent DNA polymerase.

[0301] In certain embodiments, the fourth DNA polymerase is a reverse transcriptase, such as the reverse transcriptases listed above, such as the reverse transcriptase of Moloney murine leukemia virus.

[0302] In certain embodiments, the fourth DNA polymerase has the amino acid sequence set forth in SEQ ID NO: 4.

[0303] In certain embodiments, the nucleic acid molecule B4 is capable of expressing the fourth DNA polymerase in a cell;

[0304] In certain embodiments, the nucleic acid molecule B4 is comprised in an expression vector (e.g., a eukaryotic expression vector), or the nucleic acid molecule B4 is an expression vector (e.g., a eukaryotic expression vector) comprising a nucleotide sequence encoding the fourth DNA polymerase.

[0305] In certain embodiments, the fourth tag primer is linked to the fourth gRNA.

[0306] In certain embodiments, the fourth tag primer is covalently linked to the fourth gRNA, with or without a linker.

[0307] In certain embodiments, the fourth tag primer is linked to the 3' end of the fourth gRNA, optionally with a linker.

[0308] In certain embodiments, the linker is a nucleic acid linker (e.g., a ribonucleic acid linker or a deoxyribonucleic acid linker).

[0309] In certain embodiments, the fourth tag primer is a single-stranded ribonucleic acid, and it is linked to the 3' end of the fourth gRNA, with or without a ribonucleic acid linker, to form a fourth PegRNA.

[0310] In certain embodiments, the nucleic acid molecule C4 and the nucleic acid molecule D4 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C4 and the nucleic acid molecule D4 are capable of being transcribed in a cell to produce a fourth PegRNA comprising the fourth gRNA and the fourth tag primer.

[0311] In certain embodiments, the system or kit comprises: a fourth PegRNA comprising the fourth gRNA and the fourth tag primer, or a nucleic acid molecule comprising a nucleotide sequence encoding the fourth PegRNA.

[0312] In certain embodiments, the fourth Cas protein is separate from or linked to the fourth DNA polymerase.

[0313] In certain embodiments, the fourth Cas protein is covalently linked to the fourth DNA polymerase, with or without a linker.

[0314] In certain embodiments, the linker is a peptide linker, e.g., a flexible peptide linker; for example, the linker has an amino acid sequence set forth in SEQ ID NO: 51.

[0315] In certain embodiments, the fourth Cas protein is fused to the fourth DNA polymerase, with or without a peptide linker, to form a fourth fusion protein.

[0316] In some embodiments, the fourth Cas protein is optionally linked or fused to the N-terminus of the fourth DNA polymerase via a linker; or, the fourth Cas protein is optionally linked or fused to the C-terminus of the fourth DNA polymerase via a linker.

[0317] In some embodiments, the fourth fusion protein has an amino acid sequence as set forth in SEQ ID NO: 52.

[0318] In some embodiments, the fourth fusion protein or the fourth cas protein can be split into two parts by an intein split system. It is readily understood that the intein split system can split at any amino acid position of the fourth fusion protein or the fourth cas protein. For example, in some embodiments, the intein split system splits internally in the fourth cas protein. Accordingly, in some embodiments, the fourth cas protein is split into an N-terminal segment and a C-terminal segment. For example, the N-terminal segment and the C-terminal segment of the fourth cas protein can be fused to the N-terminal segment and the C-terminal segment of an intein, respectively (or to the C-terminal segment and the N-terminal segment of an intein, respectively), and both are capable of reconstituting into an active fourth cas protein in a cell. In some embodiments, the N-terminal segment and the C-terminal segment of the fourth cas protein are each not active in an isolated state, but are capable of reconstituting into an active fourth cas protein in a cell. Accordingly, in some embodiments, the nucleic acid molecule A1 can be split into two parts, which respectively comprise a nucleotide sequence encoding the N-terminal segment and the C-terminal segment of the fourth cas protein. Further, it is readily understood that in the fourth fusion protein, the fourth DNA polymerase can be fused to the N-terminal segment or the C-terminal segment of the fourth cas protein. In some embodiments, the fourth DNA polymerase is fused to the C-terminal segment of the fourth cas protein.

[0319] In some embodiments, the nucleic acid molecule A4 and the nucleic acid molecule B4 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In some embodiments, the nucleic acid molecule A4 and the nucleic acid molecule B4 are capable of expressing the isolated fourth Cas protein and the fourth DNA polymerase, or the fourth fusion protein containing the fourth Cas protein and the fourth DNA polymerase, in a cell.

[0320] In some embodiments, the system or the kit comprises a fourth fusion protein containing the fourth Cas protein and the fourth DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the fourth fusion protein. Alternatively, the isolated fourth Cas protein and the fourth DNA polymerase, or a nucleic acid molecule capable of expressing the isolated fourth Cas protein and the fourth DNA polymerase.

[0321] In certain embodiments, the first, second, third and fourth Cas proteins are the same Cas protein, the first, second, third and fourth DNA polymerases are the same DNA polymerase; and, the system or kit comprises:

[0322] (M1-1) a first fusion protein containing the first Cas protein and the first DNA polymerase, or, a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein; or, (M1-2) the first Cas protein and the first DNA polymerase separated, or, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separated;

[0323] (M2) a first PegRNA containing the first gRNA and the first tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA;

[0324] (M3) a second PegRNA containing the second gRNA and the second tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA;

[0325] (M4) a third PegRNA containing the third gRNA and the third tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the third PegRNA;

[0326] (M5) a fourth PegRNA containing the fourth gRNA and the fourth tag primer, or, a nucleic acid molecule containing a nucleotide sequence encoding the fourth PegRNA.

[0327] In certain embodiments, the system or kit further comprises: a nucleic acid vector as defined in the foregoing.

[0328] In certain embodiments, the fourth double-stranded target nucleic acid is the same as the third double-stranded target nucleic acid, and the third and fourth functional complexes cleave the same double-stranded target nucleic acid at different locations to form cleaved nucleotide fragments a1, a2 and a3. Wherein, in the same double-stranded target nucleic acid, the nucleotide fragment a1 is connected to the nucleotide fragment a3 through the nucleotide fragment a2.

[0329] In certain embodiments, the third and fourth functional complexes respectively cause the separation of the nucleotide fragments a1 and a2 and the separation of the nucleotide fragments a2 and a3.

[0330] In certain embodiments, the nucleotide fragment a1 has a third overhang formed by extension with the third tag primer as a template; and, the nucleotide fragment a3 has a fourth overhang formed by extension with the fourth tag primer as a template.

[0331] In certain embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment al under conditions that allow nucleic acid hybridization or annealing. In certain embodiments, the first tag sequence or its complement or the first overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment al at a terminus formed by the third functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the fragmented nucleotide fragment al, and the 3' end or 3' portion is formed as a result of the third functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the first overhang is capable of hybridizing or annealing to a third overhang of the fragmented nucleotide fragment al or a nucleotide sequence upstream thereof.

[0332] In certain embodiments, the complement of the third tag sequence or the third overhang is capable of hybridizing or annealing to the first overhang or a nucleotide sequence upstream thereof under conditions that allow nucleic acid hybridization or annealing.

[0333] In certain embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment a3 under conditions that allow nucleic acid hybridization or annealing. In certain embodiments, the second tag sequence or its complement or the second overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment a3 at a terminus formed by the fourth functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the fragmented nucleotide fragment a3, and the 3' end or 3' portion is formed as a result of the fourth functional complex cleaving the third double-stranded target nucleic acid. In certain embodiments, the second overhang is capable of hybridizing or annealing to a fourth overhang of the fragmented nucleotide fragment a3 or a nucleotide sequence upstream thereof.

[0334] In certain embodiments, the complement of the fourth tag sequence or the fourth overhang is capable of hybridizing or annealing to the second overhang or a nucleotide sequence upstream thereof under conditions that allow nucleic acid hybridization or annealing.

[0335] In certain embodiments, the kit further comprises an additional component.

[0336] In certain embodiments, the additional component comprises one or more selected from the group consisting of:

[0337] (1) one or more (e.g., 2, 3, 4, 5, 10, 15, 20, or more) additional gRNAs or nucleic acid molecules containing nucleotide sequences encoding the additional gRNAs, wherein the additional gRNAs are capable of binding to a Cas protein and forming a functional complex. In certain embodiments, the functional complex is capable of cleaving or nicking both strands of a double-stranded target nucleic acid.

[0338] (2) one or more (e.g., 2, 3, 4, 5, 10, 15, 20, or more) additional Cas proteins or nucleic acid molecules containing nucleotide sequences encoding the additional Cas proteins. In certain embodiments, the Cas proteins are capable of cleaving or nicking a double-stranded target nucleic acid.

[0339] (3) one or more (e.g., 2, 3, 4, 5, 10, 15, 20, or more) additional tag primers or nucleic acid molecules containing nucleotide sequences encoding the additional tag primers, wherein the additional tag primers contain a tag sequence and a target-binding sequence, and the tag sequence is located upstream or 5' of the target-binding sequence. In certain embodiments, the target-binding sequence is capable of hybridizing or annealing to the 3' end of one of the nucleic acid strands of the nicks of the fragmented target nucleic acid fragments to form a double-stranded structure under conditions that allow nucleic acid hybridization or annealing, and the tag sequence does not bind to the target nucleic acid fragments and is in a free, single-stranded state.

[0340] (4) one or more (e.g., 2, 3, 4, 5, 10, 15, 20, or more) additional DNA polymerases or nucleic acid molecules containing nucleotide sequences encoding the additional DNA polymerases. In certain embodiments, the additional DNA polymerases are selected from the group consisting of DNA-dependent DNA polymerases and RNA-dependent DNA polymerases. In certain embodiments, the additional DNA polymerases are RNA-dependent DNA polymerases, such as reverse transcriptases.

[0341] In a second aspect, the present application provides a fusion protein comprising a Cas protein and a template-dependent DNA polymerase, wherein the Cas protein is capable of nicking a double-stranded target nucleic acid.

[0342] In certain embodiments, the Cas protein is capable of nicking a double-stranded target nucleic acid and generating sticky ends or blunt ends.

[0343] In certain embodiments, the Cas protein is selected from, but not limited to, Cas9 protein, Cas12a protein, cas12b protein, cas12c protein, cas12d protein, cas12e protein, cas12f protein, cas12g protein, cas12h protein, cas12i protein, cas14 protein, Cas13a protein, Cas1 protein, Cas1B protein, Cas2 protein, Cas3 protein, Cas4 protein, Cas5 protein, Cas6 protein, Cas7 protein, Cas8 protein, Cas10 protein, Csy1 protein, Csy2 protein, Csy3 protein, Cse1 protein, Cse2 protein, Csc1 protein, Csc2 protein, Csa5 protein, Csn2 protein, Csm2 protein, Csm3 protein, Csm4 protein, Csm5 protein, Csm6 protein, Cmr1 protein, Cmr3 protein, Cmr4 protein, Cmr5 protein, Cmr6 protein, Csb1 protein, Csb2 protein, Csb3 protein, Csx17 protein, Csx14 protein, Csx10 protein, Csx16 protein, CsaX protein, Csx3 protein, Csx1 protein, Csx15 protein, Csf1 protein, Csf2 protein, Csf3 protein, Csf4 protein, and homologues thereof or modified versions thereof.

[0344] In certain embodiments, the Cas protein is a Cas9 protein, for example, a Cas9 protein of S. pyogenes (spCas9).

[0345] In certain embodiments, the Cas protein has an amino acid sequence as set forth in SEQ ID NO: 1.

[0346] In certain embodiments, the DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0347] In certain embodiments, the DNA polymerase is an RNA-dependent DNA polymerase.

[0348] In certain embodiments, the DNA polymerase is a reverse transcriptase, for example, a reverse transcriptase listed above, for example, a reverse transcriptase of Moloney murine leukemia virus.

[0349] In certain embodiments, the DNA polymerase has an amino acid sequence as set forth in SEQ ID NO: 4.

[0350] In certain embodiments, the Cas protein is covalently linked to the DNA polymerase via a linker or without a linker.

[0351] In certain embodiments, the linker is a peptide linker, e.g., a flexible peptide linker; for example, the linker has an amino acid sequence set forth in SEQ ID NO: 51.

[0352] In certain embodiments, the Cas protein is linked or fused to the N-terminus of the DNA polymerase, optionally through a linker; or, the Cas protein is linked or fused to the C-terminus of the DNA polymerase, optionally through a linker.

[0353] In certain embodiments, the fusion protein has an amino acid sequence set forth in SEQ ID NO: 52.

[0354] In a third aspect, the present application provides a nucleic acid molecule comprising a polynucleotide encoding a fusion protein as previously described.

[0355] In a fourth aspect, the present application provides a vector comprising a nucleic acid molecule as previously described.

[0356] In certain embodiments, the vector is an expression vector.

[0357] In certain embodiments, the vector is a eukaryotic expression vector.

[0358] In a fifth aspect, the present application provides a host cell comprising a nucleic acid molecule as previously described or a vector as previously described.

[0359] In certain embodiments, the host cell is a prokaryotic cell, e.g., an E. coli cell; or the host cell is a eukaryotic cell, e.g., a yeast cell, a fungal cell, a plant cell, an animal cell.

[0360] In certain embodiments, the host cell is a mammalian cell, e.g., a human cell.

[0361] In a fifth aspect, the present application provides a method of producing a fusion protein as previously described, comprising, (1) culturing a host cell as previously described under conditions permitting protein expression; and (2) isolating the fusion protein expressed by the host cell.

[0362] In a sixth aspect, the present application provides a complex comprising a first Cas protein and a first template-dependent DNA polymerase, wherein the first Cas protein has the ability to cleave a double-stranded target nucleic acid, and the first Cas protein is complexed with the first DNA polymerase, covalently or non-covalently.

[0363] In certain embodiments, the first Cas protein is capable of cleaving a double-stranded target nucleic acid and generating a sticky end or a blunt end.

[0364] In certain embodiments, the first Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologs thereof or modified versions thereof.

[0365] In certain embodiments, the first Cas protein is a Cas9 protein, such as a Cas9 protein of S. pyogenes (spCas9).

[0366] In certain embodiments, the first Cas protein has the amino acid sequence set forth in SEQ ID NO: 1.

[0367] In certain embodiments, the first DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0368] In certain embodiments, the first DNA polymerase is an RNA-dependent DNA polymerase.

[0369] In certain embodiments, the first DNA polymerase is a reverse transcriptase, such as the reverse transcriptases listed above, such as the reverse transcriptase of Moloney murine leukemia virus.

[0370] In certain embodiments, the first DNA polymerase has the amino acid sequence set forth in SEQ ID NO: 4.

[0371] In certain embodiments, the first Cas protein is covalently linked to the first DNA polymerase via a linker or not.

[0372] In some embodiments, the linker is a peptide linker, e.g., a flexible peptide linker; for example, the linker has an amino acid sequence set forth in SEQ ID NO: 51.

[0373] In some embodiments, the first Cas protein is fused to the first DNA polymerase through a peptide linker or without a peptide linker, forming a first fusion protein.

[0374] In some embodiments, the first Cas protein is optionally linked or fused to the N-terminus of the first DNA polymerase through a linker; or, the first Cas protein is optionally linked or fused to the C-terminus of the first DNA polymerase through a linker.

[0375] In some embodiments, the first fusion protein has an amino acid sequence set forth in SEQ ID NO: 52.

[0376] In some embodiments, the complex further comprises a first gRNA.

[0377] In some embodiments, the first gRNA is capable of binding to the first Cas protein and forming a first functional unit; the first functional unit is capable of binding to a double- stranded target nucleic acid and cleaving both strands of the double-stranded target nucleic acid to form a cleaved target nucleic acid fragment.

[0378] In some embodiments, the first gRNA contains a first guide sequence, and, under conditions permissive for nucleic acid hybridization or annealing, the first guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a double-stranded target nucleic acid.

[0379] In some embodiments, the first guide sequence has a length of at least 5 nt, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0380] In some embodiments, the first gRNA further contains a first scaffold sequence, which is capable of being recognized and bound by the first Cas protein, thereby forming a first functional unit.

[0381] In some embodiments, the first scaffold sequence has a length of at least 20 nt, e.g., 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0382] In some embodiments, the first guide sequence is located upstream or 5’ to the first scaffold sequence.

[0383] In certain embodiments, the complex or the first functional unit is capable of cleaving both strands of the double-stranded target nucleic acid to form a cleaved target nucleic acid fragment after the first guide sequence binds to the double-stranded target nucleic acid.

[0384] In certain embodiments, the complex further comprises a double-stranded target nucleic acid,

[0385] In certain embodiments, the double-stranded target nucleic acid contains a first PAM sequence recognized by the first Cas protein and a first guide binding sequence capable of hybridizing or annealing to the first guide sequence, whereby the first functional unit binds to the double-stranded target nucleic acid through the first guide binding sequence and the first PAM sequence.

[0386] In certain embodiments, the complex further comprises a first tag primer hybridized or annealed to the double-stranded target nucleic acid; wherein the first tag primer contains a first target binding sequence capable of hybridizing or annealing to the double-stranded target nucleic acid.

[0387] In certain embodiments, the tag primer contains a first tag sequence and a first target binding sequence, the first tag sequence is located upstream or 5’ of the first target binding sequence; and the first target binding sequence is capable of hybridizing or annealing to the double-stranded target nucleic acid under conditions permitting nucleic acid hybridization or annealing. In certain embodiments, the first target binding sequence is capable of hybridizing or annealing to the double-stranded target nucleic acid at the position cleaved by the first functional unit; in certain embodiments, the first target binding sequence is capable of hybridizing or annealing to the 3’ end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure. In certain embodiments, the 3’ end is formed as a result of the first functional unit cleaving the double-stranded target nucleic acid; in certain embodiments, the first tag sequence is not bound to the target nucleic acid fragment and is in a free, single-stranded state.

[0388] In certain embodiments, the first target binding sequence is at least 5 nt in length, such as 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0389] In certain embodiments, the first tag sequence is at least 4 nt in length, such as 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0390] In certain embodiments, the first tag primer binds to the fragmented target nucleic acid fragment via the first target-binding sequence; in certain embodiments, the first DNA polymerase binds to the fragmented target nucleic acid fragment and the first tag primer.

[0391] In certain embodiments, the first tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0392] In certain embodiments, the first tag primer is a single-stranded ribonucleic acid and the first DNA polymerase is an RNA-dependent DNA polymerase; or, the first tag primer is a single-stranded deoxyribonucleic acid and the first DNA polymerase is a DNA-dependent DNA polymerase.

[0393] In certain embodiments, the fragmented target nucleic acid fragment is extended by the first DNA polymerase using the first tag primer as a template, forming a first overhang.

[0394] In certain embodiments, the first gRNA-bound nucleic acid strand is different from the first tag primer-bound nucleic acid strand. In certain embodiments, the first gRNA-bound nucleic acid strand is the opposite strand of the first tag primer-bound nucleic acid strand.

[0395] In certain embodiments, the first tag primer is linked to the first gRNA.

[0396] In certain embodiments, the first tag primer is covalently linked to the first gRNA via or without a linker.

[0397] In certain embodiments, the first tag primer is linked to the 3' end of the first gRNA, optionally via a linker.

[0398] In certain embodiments, the linker is a nucleic acid linker (e.g., a ribonucleic acid linker or a deoxyribonucleic acid linker).

[0399] In certain embodiments, the first tag primer is a single-stranded ribonucleic acid and it is linked to the 3' end of the first gRNA via or without a ribonucleic acid linker, forming a first PegRNA.

[0400] In certain embodiments, the complex further comprises a second Cas protein and a second gRNA, wherein the second Cas protein has the ability to cleave a double-stranded target nucleic acid, and the second gRNA is capable of binding to the second Cas protein and forms a second functional unit; the second functional unit is capable of binding to a double-stranded target nucleic acid and cleaving both strands of the double-stranded target nucleic acid, forming a fragmented target nucleic acid fragment.

[0401] In certain embodiments, the second Cas protein is the same as or different from the first Cas protein. In certain embodiments, the second Cas protein is the same as the first Cas protein.

[0402] In certain embodiments, the second Cas protein is capable of cleaving a double- stranded target nucleic acid and generating a sticky end or a blunt end.

[0403] In certain embodiments, the second Cas protein is selected from, but not limited to, a Cas9 protein, a Cas12a protein, a cas12b protein, a cas12c protein, a cas12d protein, a cas12e protein, a cas12f protein, a cas12g protein, a cas12h protein, a cas12i protein, a cas14 protein, a Cas13a protein, a Cas1 protein, a Cas1B protein, a Cas2 protein, a Cas3 protein, a Cas4 protein, a Cas5 protein, a Cas6 protein, a Cas7 protein, a Cas8 protein, a Cas10 protein, a Csy1 protein, a Csy2 protein, a Csy3 protein, a Cse1 protein, a Cse2 protein, a Csc1 protein, a Csc2 protein, a Csa5 protein, a Csn2 protein, a Csm2 protein, a Csm3 protein, a Csm4 protein, a Csm5 protein, a Csm6 protein, a Cmr1 protein, a Cmr3 protein, a Cmr4 protein, a Cmr5 protein, a Cmr6 protein, a Csb1 protein, a Csb2 protein, a Csb3 protein, a Csx17 protein, a Csx14 protein, a Csx10 protein, a Csx16 protein, a CsaX protein, a Csx3 protein, a Csx1 protein, a Csx15 protein, a Csf1 protein, a Csf2 protein, a Csf3 protein, a Csf4 protein, and homologs thereof or modified versions thereof.

[0404] In certain embodiments, the second Cas protein is a Cas9 protein, such as a Cas9 protein of S. pyogenes (spCas9).

[0405] In certain embodiments, the second Cas protein has an amino acid sequence set forth in SEQ ID NO: 1.

[0406] In certain embodiments, the second gRNA contains a second guide sequence, and, under conditions permissive for nucleic acid hybridization or annealing, the second guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a double-stranded target nucleic acid.

[0407] In certain embodiments, the second guide sequence is different from the first guide sequence; in certain embodiments, the nucleic acid strand bound by the first guide sequence is different from the nucleic acid strand bound by the second guide sequence. In certain embodiments, the nucleic acid strand bound by the first guide sequence is the opposite strand of the nucleic acid strand bound by the second guide sequence.

[0408] In certain embodiments, the second guide sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0409] In certain embodiments, the second gRNA further comprises a second scaffold sequence that is capable of being recognized and bound by the second Cas protein, thereby forming a second functional unit.

[0410] In certain embodiments, the second scaffold sequence is the same as or different from the first scaffold sequence. In certain embodiments, the second scaffold sequence is the same as the first scaffold sequence.

[0411] In certain embodiments, the second scaffold sequence is at least 20 nt in length, e.g., 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0412] In certain embodiments, the second guide sequence is located upstream or 5’ of the second scaffold sequence.

[0413] In certain embodiments, the double-stranded target nucleic acid comprises a second PAM sequence recognized by the second Cas protein and a second guide binding sequence capable of hybridizing or annealing to the second guide sequence, whereby the second functional unit binds the double-stranded target nucleic acid through the second guide binding sequence and the second PAM sequence.

[0414] In certain embodiments, the complex further comprises a second template-dependent DNA polymerase covalently or non-covalently complexed with the second Cas protein.

[0415] In certain embodiments, the second DNA polymerase is selected from, but not limited to, a DNA-dependent DNA polymerase and an RNA-dependent DNA polymerase.

[0416] In certain embodiments, the second DNA polymerase is an RNA-dependent DNA polymerase.

[0417] In certain embodiments, the second DNA polymerase is a reverse transcriptase, such as the reverse transcriptases listed above, for example, the reverse transcriptase of Moloney murine leukemia virus.

[0418] In certain embodiments, the second DNA polymerase has an amino acid sequence as set forth in SEQ ID NO: 4.

[0419] In certain embodiments, the second DNA polymerase is the same as or different from the first DNA polymerase. In certain embodiments, the second DNA polymerase is the same as the first DNA polymerase.

[0420] In certain embodiments, the second Cas protein is covalently linked to the second DNA polymerase via a linker or not.

[0421] In certain embodiments, the linker is a peptide linker, such as a flexible peptide linker; for example, the linker has an amino acid sequence as set forth in SEQ ID NO: 51.

[0422] In certain embodiments, the second Cas protein is fused to the second DNA polymerase via a peptide linker or not, forming a second fusion protein.

[0423] In certain embodiments, the second Cas protein is optionally linked or fused to the N-terminus of the second DNA polymerase via a linker; or, the second Cas protein is optionally linked or fused to the C-terminus of the second DNA polymerase via a linker.

[0424] In certain embodiments, the second fusion protein has an amino acid sequence as set forth in SEQ ID NO: 52.

[0425] In certain embodiments, the complex further comprises a second tag primer hybridized or annealed to the double-stranded target nucleic acid; wherein the second tag primer contains a second target-binding sequence that is capable of hybridizing or annealing to the double-stranded target nucleic acid.

[0426] In certain embodiments, the tag primer contains a second tag sequence and a second target-binding sequence, the second tag sequence is located upstream or 5' of the second target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the second target-binding sequence is capable of hybridizing or annealing to the double-stranded target nucleic acid. In certain embodiments, the second target-binding sequence is capable of hybridizing or annealing to the double-stranded target nucleic acid at the location where the double-stranded target nucleic acid is cleaved by the second functional unit; in certain embodiments, the second target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment, forming a double-stranded structure. In certain embodiments, the 3' end is formed as a result of the cleavage of the double-stranded target nucleic acid by the second functional unit. In certain embodiments, the second tag sequence is not bound to the target nucleic acid fragment, in a free, single-stranded state.

[0427] In certain embodiments, the second target-binding sequence is at least 5 nt in length, e.g., 5-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0428] In certain embodiments, the second target-binding sequence is different from the first target-binding sequence. In certain embodiments, the nucleic acid strand to which the second target-binding sequence binds is different from the nucleic acid strand to which the first target-binding sequence binds. In certain embodiments, the nucleic acid strand to which the second target-binding sequence binds is the opposite strand from the nucleic acid strand to which the first target-binding sequence binds.

[0429] In certain embodiments, the second tag sequence is at least 4 nt in length, e.g., 4-10 nt, 10-15 nt, 15-20 nt, 20-25 nt, 25-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, 100-200 nt, or longer.

[0430] In certain embodiments, the second tag sequence is the same as or different from the first tag sequence. In certain embodiments, the second tag sequence is different from the first tag sequence.

[0431] In certain embodiments, the second tag primer binds to the cleaved target nucleic acid fragment via the second target-binding sequence. In certain embodiments, the second DNA polymerase binds to the cleaved target nucleic acid fragment and the second tag primer.

[0432] In certain embodiments, the second tag primer is a single-stranded deoxyribonucleic acid or a single-stranded ribonucleic acid.

[0433] In certain embodiments, the second tag primer is a single-stranded ribonucleic acid, and the second DNA polymerase is an RNA-dependent DNA polymerase; or, the second tag primer is a single-stranded deoxyribonucleic acid, and the second DNA polymerase is a DNA-dependent DNA polymerase.

[0434] In certain embodiments, the fragmented target nucleic acid fragments are extended by the second DNA polymerase using the second tag primer as a template, forming second overhangs.

[0435] In certain embodiments, the second gRNA-bound nucleic acid strand is different from the second tag primer-bound nucleic acid strand; in certain embodiments, the second gRNA-bound nucleic acid strand is the opposite strand of the second tag primer-bound nucleic acid strand.

[0436] In certain embodiments, the second tag primer is linked to the second gRNA.

[0437] In certain embodiments, the second tag primer is covalently linked to the second gRNA, with or without a linker.

[0438] In certain embodiments, the second tag primer is linked to the 3' end of the second gRNA, optionally with a linker.

[0439] In certain embodiments, the linker is a nucleic acid linker (e.g., a ribonucleic acid linker or a deoxyribonucleic acid linker).

[0440] In certain embodiments, the second tag primer is a single-stranded ribonucleic acid, and it is linked to the 3' end of the second gRNA, with or without a ribonucleic acid linker, forming a second PegRNA.

[0441] In certain embodiments, the first and second functional units bind to a double-stranded target nucleic acid in a predetermined positional relationship.

[0442] In certain embodiments, the second guide sequence binds to the same nucleic acid strand as the first target-binding sequence; and / or, the first guide sequence binds to the same nucleic acid strand as the second target-binding sequence.

[0443] In certain embodiments, the binding position of the second guide sequence is upstream or 5' to the binding position of the first target-binding sequence; and / or, the binding position of the first guide sequence is upstream or 5' to the binding position of the second target-binding sequence.

[0444] In certain embodiments, the binding site of the second guide sequence is downstream or 3’ of the binding site of the first target binding sequence; and / or, the binding site of the first guide sequence is downstream or 3’ of the binding site of the second target binding sequence.

[0445] In certain embodiments, the double-stranded target nucleic acid is selected from, but not limited to, genomic DNA and nucleic acid vector DNA.

[0446] In a seventh aspect, the present application provides a method for cleaving a double-stranded target nucleic acid and adding an overhang at its 3’ end, wherein the method comprises using a system or a kit as previously described.

[0447] In certain embodiments, the method comprises the following steps:

[0448] i. providing a double-stranded target nucleic acid; and

[0449] providing the first Cas protein, the first gRNA, the first DNA polymerase and the first tag primer;

[0450] ii contacting the double-stranded target nucleic acid with the first Cas protein, the first gRNA, the first DNA polymerase and the first tag primer.

[0451] In certain embodiments, in step ii:

[0452] the first Cas protein and the first gRNA form a first functional complex, and the first functional complex binds to and cleaves the double-stranded target nucleic acid to form a cleaved target nucleic acid fragment; and,

[0453] the first tag primer hybridizes or anneals to the 3’ end of one nucleic acid strand of the cleaved target nucleic acid fragment via the first target binding sequence; and,

[0454] the first DNA polymerase extends the cleaved target nucleic acid fragment using the first tag primer annealed to the cleaved target nucleic acid fragment as a template to form a first overhang.

[0455] In certain embodiments, the method is performed in a cell.

[0456] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule A1, the first DNA polymerase or nucleic acid molecule B1, the first gRNA or nucleic acid molecule C1, and the first tag primer or nucleic acid molecule D1 are delivered into a cell to provide the first Cas protein, the first gRNA, the first DNA polymerase and the first tag primer in the cell.

[0457] In certain embodiments, in step i, the nucleic acid molecule A1, the nucleic acid molecule B1, the first gRNA or nucleic acid molecule C1, and the first tag primer or nucleic acid molecule D1 are delivered into the cell to provide the first Cas protein, first gRNA, first DNA polymerase, and first tag primer within the cell.

[0458] In certain embodiments, in step i, the nucleic acid molecule A1, B1, C1, and D1 are delivered into the cell to provide the first Cas protein, first gRNA, first DNA polymerase, and first tag primer within the cell.

[0459] In certain embodiments, the nucleic acid molecule A1 and nucleic acid molecule B1 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule A1 and nucleic acid molecule B1 are capable of expressing the first Cas protein and the first DNA polymerase separately, or a first fusion protein containing the first Cas protein and the first DNA polymerase, in the cell. In certain embodiments, in step i, the nucleic acid molecules capable of expressing the first Cas protein and the first DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, are delivered into the cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase within the cell.

[0460] In certain embodiments, the nucleic acid molecule C1 and nucleic acid molecule D1 are comprised in the same expression vector (e.g., eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C1 and nucleic acid molecule D1 are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in the cell. In certain embodiments, in step i, the first PegRNA is delivered into the cell to provide the first gRNA and the first tag primer within the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into the cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer within the cell.

[0461] In certain embodiments, in step i, the nucleic acid molecules capable of expressing the first Cas protein and the first DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, and a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA are delivered into the cell and transcribed and expressed in the cell, thereby providing the first Cas protein, first gRNA, first DNA polymerase, and first tag primer within the cell.

[0462] In certain embodiments, in step i, the double-stranded target nucleic acid or a nucleic acid molecule T containing the double-stranded target nucleic acid is delivered into a cell to provide the double-stranded target nucleic acid within the cell.

[0463] In certain embodiments, the first Cas protein, first gRNA, first DNA polymerase, or first tag primer are as defined in the foregoing.

[0464] In certain embodiments, the double-stranded target nucleic acid or nucleic acid molecule T contains a first PAM sequence recognized by the first Cas protein. In certain embodiments, in step ii, the first functional complex binds to the double-stranded target nucleic acid or nucleic acid molecule T through the first PAM sequence and the first gRNA and cleaves it.

[0465] In an eighth aspect, the present application provides a method for cleaving a double-stranded target nucleic acid into a target nucleic acid fragment and adding overhangs to the two 3’ ends of the target nucleic acid fragment, respectively, wherein the method comprises using a system or kit as described in the foregoing; wherein the first double-stranded target nucleic acid is identical to the second double-stranded target nucleic acid.

[0466] In certain embodiments, the method comprises the following steps:

[0467] i. providing a double-stranded target nucleic acid; and

[0468] providing the first Cas protein, first gRNA, first DNA polymerase, first tag primer, the second Cas protein, second gRNA, second DNA polymerase, and second tag primer;

[0469] ii. contacting the double-stranded target nucleic acid with the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, and second tag primer.

[0470] In certain embodiments, in step ii:

[0471] the first Cas protein and first gRNA bind together to form a first functional complex, and the second Cas protein and second gRNA bind together to form a second functional complex; and, the first and second functional complexes bind to and cleave the double-stranded target nucleic acid to form a target nucleic acid fragment F1; and,

[0472] the first tag primer hybridizes or anneals to the 3’ end of one nucleic acid strand of the target nucleic acid fragment F1 through the first target-binding sequence; and, the second tag primer hybridizes or anneals to the 3’ end of the other nucleic acid strand of the target nucleic acid fragment F1 through the second target-binding sequence; and,

[0473] the first DNA polymerase and the second DNA polymerase extend the target nucleic acid fragment F1 using the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang.

[0474] In certain embodiments, the method is performed in a cell.

[0475] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule A1, the first DNA polymerase or nucleic acid molecule B1, the first gRNA or nucleic acid molecule C1, the first tag primer or nucleic acid molecule D1, the second Cas protein or nucleic acid molecule A2, the second DNA polymerase or nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, and the second tag primer or nucleic acid molecule D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, and second tag primer in the cell.

[0476] In certain embodiments, in step i, the nucleic acid molecule A1, the nucleic acid molecule B1, the first gRNA or nucleic acid molecule C1, the first tag primer or nucleic acid molecule D1, the nucleic acid molecule A2, the nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, and the second tag primer or nucleic acid molecule D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, and second tag primer in the cell.

[0477] In certain embodiments, in step i, the nucleic acid molecules A1, B1, C1, D1, A2, B2, C2, and D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, and second tag primer in the cell.

[0478] In certain embodiments, the nucleic acid molecule A1 and the nucleic acid molecule B1 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule A1 and the nucleic acid molecule B1 are capable of expressing, in a cell, the first Cas protein and the first DNA polymerase separately, or a first fusion protein containing the first Cas protein and the first DNA polymerase. In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, is delivered into the cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase in the cell.

[0479] In certain embodiments, the nucleic acid molecule A2 and the nucleic acid molecule B2 are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule A2 and the nucleic acid molecule B2 are capable of expressing, in a cell, the second Cas protein and the second DNA polymerase separately, or a second fusion protein containing the second Cas protein and the second DNA polymerase. In certain embodiments, in step i, a nucleic acid molecule capable of expressing the second Cas protein and the second DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the second fusion protein, is delivered into the cell and expressed in the cell to provide the second Cas protein and the second DNA polymerase in the cell.

[0480] In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are comprised in the same expression vector (e.g., eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are capable of transcribing, in a cell, a first PegRNA containing the first gRNA and the first tag primer. In certain embodiments, in step i, the first PegRNA is delivered into the cell to provide the first gRNA and the first tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into the cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer in the cell.

[0481] In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of being transcribed in a cell to produce a second PegRNA comprising the second gRNA and the second tag primer. In certain embodiments, in step i, the second PegRNA is delivered into the cell to provide the second gRNA and the second tag primer in the cell, or a nucleic acid molecule comprising a nucleotide sequence encoding the second PegRNA is delivered into the cell and the second PegRNA is transcribed in the cell to provide the second gRNA and the second tag primer in the cell.

[0482] In certain embodiments, in step i, the double-stranded target nucleic acid or the nucleic acid molecule T comprising the double-stranded target nucleic acid is delivered into the cell to provide the double-stranded target nucleic acid in the cell.

[0483] In certain embodiments, the first Cas protein, the first gRNA, the first DNA polymerase, or the first tag primer are as defined in the foregoing.

[0484] In certain embodiments, the second Cas protein, the second gRNA, the second DNA polymerase, or the second tag primer are as defined in the foregoing.

[0485] In certain embodiments, the double-stranded target nucleic acid or the nucleic acid molecule T comprises a first PAM sequence recognized by the first Cas protein and a second PAM sequence recognized by the second Cas protein. In certain embodiments, in step ii, the first functional complex binds to the double-stranded target nucleic acid or the nucleic acid molecule T through the first PAM sequence and the first gRNA and cleaves it; and, the second functional complex binds to the double-stranded target nucleic acid or the nucleic acid molecule T through the second PAM sequence and the second gRNA and cleaves it.

[0486] In certain embodiments, the second Cas protein is the same as the first Cas protein, and the second DNA polymerase is the same as the first DNA polymerase; wherein the first Cas protein forms a first functional complex with the first gRNA and a second functional complex with the second gRNA, respectively, and the first DNA polymerase extends the target nucleic acid fragment F1 with the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, respectively, to form a target nucleic acid fragment F2 with a first overhang and a second overhang.

[0487] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, and the second tag primer or nucleic acid molecule D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, and second tag primer within the cell.

[0488] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, and the second tag primer or nucleic acid molecule D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, and second tag primer within the cell.

[0489] In certain embodiments, in step i, the nucleic acid molecules Al, Bl, Cl, Dl, C2, and D2 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, and second tag primer within the cell.

[0490] In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are capable of expressing the first Cas protein and the first DNA polymerase separately or a first fusion protein containing the first Cas protein and the first DNA polymerase in a cell. In certain embodiments, in step i, nucleic acid molecules capable of expressing the first Cas protein and first DNA polymerase separately or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein are delivered into a cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase within the cell.

[0491] In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in a cell. In certain embodiments, in step i, the first PegRNA is delivered into the cell to provide the first gRNA and the first tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into the cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer in the cell.

[0492] In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of transcribing a second PegRNA containing the second gRNA and the second tag primer in a cell. In certain embodiments, in step i, the second PegRNA is delivered into the cell to provide the second gRNA and the second tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA is delivered into the cell and the second PegRNA is transcribed in the cell to provide the second gRNA and the second tag primer in the cell.

[0493] In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA, and a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA are delivered into the cell and transcribed and expressed in the cell, thereby providing the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second gRNA, and the second tag primer in the cell.

[0494] In a ninth aspect, the present application provides a method for inserting a target nucleic acid fragment into a nucleic acid molecule of interest; wherein the method comprises using a system or a kit as described previously; wherein the first double-stranded target nucleic acid and the second double-stranded target nucleic acid are the same for providing the target nucleic acid fragment; and the third double-stranded target nucleic acid is the nucleic acid molecule of interest.

[0495] In certain embodiments, the method comprises:

[0496] a. by the method as described in the foregoing, the first double-stranded target nucleic acid is fragmented into a target nucleic acid fragment Fl, and overhangs are added to both 3' ends of the target nucleic acid fragment Fl, respectively, to form a target nucleic acid fragment F2 having a first overhang and a second overhang;

[0497] b. the nucleic acid molecule of interest is fragmented with the third functional complex to form fragmented nucleotide fragments al and a2; and,

[0498] c. the nucleotide fragments al and a2 are ligated with the target nucleic acid fragment F2, thereby inserting the target nucleic acid fragment into the nucleic acid molecule of interest.

[0499] In certain embodiments, the method comprises the steps of:

[0500] i. providing a double-stranded target nucleic acid and a nucleic acid molecule of interest; and

[0501] providing the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second Cas protein, the second gRNA, the second DNA polymerase, the second tag primer, the third Cas protein, and the third gRNA;

[0502] ii. contacting the double-stranded target nucleic acid with the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second Cas protein, the second gRNA, the second DNA polymerase, and the second tag primer, and, contacting the nucleic acid molecule of interest with the third Cas protein and the third gRNA.

[0503] In certain embodiments, in step ii:

[0504] the first Cas protein and the first gRNA combine to form a first functional complex, the second Cas protein and the second gRNA combine to form a second functional complex, and the third Cas protein and the third gRNA combine to form a third functional complex; and,

[0505] the first and second functional complexes bind to and fragment the double-stranded target nucleic acid to form a target nucleic acid fragment Fl, and, the third functional complex binds to and fragments the nucleic acid molecule of interest to form fragmented nucleotide fragments al and a2; and,

[0506] the first tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the target nucleic acid fragment Fl via the first target-binding sequence; and, the second tag primer hybridizes or anneals to the 3' end of the other nucleic acid strand of the target nucleic acid fragment Fl via the second target-binding sequence; and,

[0507] the first DNA polymerase and the second DNA polymerase extend the target nucleic acid fragment F1 using the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first overhang and the second overhang are capable of hybridizing or annealing to the broken nucleotide fragment a1 and a2, respectively; and,

[0508] the target nucleic acid fragment F2 is hybridized or annealed to the nucleotide fragment a1 and a2 through the first overhang and the second overhang, respectively, and is further inserted or ligated between the nucleotide fragment a1 and a2, thereby inserting the target nucleic acid fragment into the nucleic acid molecule of interest.

[0509] In certain embodiments, the first overhang is capable of hybridizing or annealing to a 3’ end or a 3’ portion of one nucleic acid strand of the nucleotide fragment a1, and the 3’ end or the 3’ portion is formed due to the breaking of the nucleic acid molecule of interest by the third functional complex.

[0510] In certain embodiments, the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to a 3’ portion of one nucleic acid strand of the broken nucleotide fragment a1, and there is a first spacer region between the 3’ portion of the nucleotide fragment a1 and the broken end formed by the third double-stranded target nucleic acid.

[0511] In certain embodiments, the first spacer region has a length of 1 nt-200 nt, such as 1-10 nt, 10-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, or 100-200 nt.

[0512] In certain embodiments, the second overhang is capable of hybridizing or annealing to a 3’ end or a 3’ portion of one nucleic acid strand of the nucleotide fragment a2, and the 3’ end or the 3’ portion is formed due to the breaking of the nucleic acid molecule of interest by the third functional complex.

[0513] In certain embodiments, the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to a 3’ portion of one nucleic acid strand of the broken nucleotide fragment a2, and there is a second spacer region between the 3’ portion of the nucleotide fragment a2 and the broken end formed by the third double-stranded target nucleic acid.

[0514] In certain embodiments, the second spacer region has a length of 1 nt-200 nt, such as 1-10 nt, 10-20 nt, 20-30 nt, 30-40 nt, 40-50 nt, 50-100 nt, or 100-200 nt.

[0515] In certain embodiments, the method is performed in a cell.

[0516] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second Cas protein or nucleic acid molecule A2, the second DNA polymerase or nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third Cas protein or nucleic acid molecule A3, and the third gRNA or nucleic acid molecule C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, and third gRNA in the cell.

[0517] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the nucleic acid molecule A2, the nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the nucleic acid molecule A3, and the third gRNA or nucleic acid molecule C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, and third gRNA in the cell.

[0518] In certain embodiments, in step i, the nucleic acid molecules Al, Bl, Cl, Dl, A2, B2, C2, D2, A3, and C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, and third gRNA in the cell.

[0519] In certain embodiments, in step i, the double-stranded target nucleic acid or a nucleic acid molecule containing the double-stranded target nucleic acid T is delivered into a cell to provide the double-stranded target nucleic acid in the cell.

[0520] In certain embodiments, the double-stranded target nucleic acid or nucleic acid molecule T contains a first PAM sequence recognized by the first Cas protein and a second PAM sequence recognized by the second Cas protein. In certain embodiments, in step ii, the first functional complex binds to and cleaves the double-stranded target nucleic acid or nucleic acid molecule T through the first PAM sequence and the first gRNA; and, the second functional complex binds to and cleaves the double-stranded target nucleic acid or nucleic acid molecule T through the second PAM sequence and the second gRNA.

[0521] In certain embodiments, the nucleic acid molecule of interest contains a third PAM sequence recognized by a third Cas protein; in certain embodiments, in step ii, the third functional complex binds to and cleaves the nucleic acid molecule of interest through the third PAM sequence and the third gRNA.

[0522] In certain embodiments, the nucleic acid molecule of interest is genomic DNA of the cell.

[0523] In certain embodiments, the first Cas protein, the first gRNA, the first DNA polymerase, or the first tag primer are as defined in the foregoing.

[0524] In certain embodiments, the second Cas protein, the second gRNA, the second DNA polymerase, or the second tag primer are as defined in the foregoing.

[0525] In certain embodiments, the third Cas protein and the third gRNA are as defined in the foregoing.

[0526] In certain embodiments, the first, second, and third Cas proteins are the same Cas protein, and the second DNA polymerase is the same as the first DNA polymerase; wherein the first Cas protein forms a first, a second, and a third functional complex with the first, the second, and the third gRNA, respectively, and the first DNA polymerase extends the target nucleic acid fragment F1 with the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, respectively, to form a target nucleic acid fragment F2 with a first overhang and a second overhang.

[0527] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, and the third gRNA or nucleic acid molecule C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, and third gRNA within the cell.

[0528] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, and the third gRNA or nucleic acid molecule C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, and third gRNA within the cell.

[0529] In certain embodiments, in step i, the nucleic acid molecules Al, Bl, Cl, Dl, C2, D2, and C3 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, and third gRNA within the cell.

[0530] In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are capable of expressing the first Cas protein and the first DNA polymerase separately or a first fusion protein containing the first Cas protein and the first DNA polymerase in a cell. In certain embodiments, in step i, nucleic acid molecules capable of expressing the first Cas protein and first DNA polymerase separately or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein are delivered into a cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase within the cell.

[0531] In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in a cell. In certain embodiments, in step i, the first PegRNA is delivered into the cell to provide the first gRNA and the first tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into the cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer in the cell.

[0532] In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of transcribing a second PegRNA containing the second gRNA and the second tag primer in a cell. In certain embodiments, in step i, the second PegRNA is delivered into the cell to provide the second gRNA and the second tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA is delivered into the cell and the second PegRNA is transcribed in the cell to provide the second gRNA and the second tag primer in the cell;

[0533] In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA, a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA, and a nucleic acid molecule containing a nucleotide sequence encoding the third gRNA are delivered into the cell and transcribed and expressed in the cell, thereby providing the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second gRNA, the second tag primer, and the third gRNA in the cell.

[0534] In a tenth aspect, the present application provides a method for replacing a nucleotide fragment in a nucleic acid molecule of interest with a target nucleic acid fragment; wherein the method comprises using a system or a kit as described previously; wherein the first double-stranded target nucleic acid and the second double-stranded target nucleic acid are the same for providing the target nucleic acid fragment; and the third double-stranded target nucleic acid and the fourth double-stranded target nucleic acid are the same for the nucleic acid molecule of interest.

[0535] In certain embodiments, the method comprises:

[0536] a. cleaving the first double-stranded target nucleic acid into a target nucleic acid fragment Fl by a method as previously described, and adding overhangs to both 3' ends of the target nucleic acid fragment Fl, respectively, to form a target nucleic acid fragment F2 having a first overhang and a second overhang;

[0537] b. cleaving the nucleic acid molecule of interest with the third and fourth functional complexes to form cleaved nucleotide fragments al, a2 and a3; wherein, prior to cleavage, nucleotide fragments al, a2 and a3 are arranged in the nucleic acid molecule of interest in sequence (i.e., nucleotide fragment al is connected to nucleotide fragment a3 by nucleotide fragment a2); and,

[0538] c. ligating the nucleotide fragments al and a3 with the target nucleic acid fragment F2, thereby replacing nucleotide fragment a2 in the nucleic acid molecule of interest with the target nucleic acid fragment.

[0539] In certain embodiments, the method comprises the steps of:

[0540] i. providing a double-stranded target nucleic acid and a nucleic acid molecule of interest; and

[0541] providing the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second Cas protein, the second gRNA, the second DNA polymerase, the second tag primer, the third Cas protein, the third gRNA, the fourth Cas protein, and the fourth gRNA;

[0542] ii contacting the double-stranded target nucleic acid with the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second Cas protein, the second gRNA, the second DNA polymerase, and the second tag primer, and contacting the nucleic acid molecule of interest with the third Cas protein, the third gRNA, the fourth Cas protein, and the fourth gRNA.

[0543] In certain embodiments, in step ii:

[0544] the first Cas protein and the first gRNA combine to form a first functional complex, the second Cas protein and the second gRNA combine to form a second functional complex, the third Cas protein and the third gRNA combine to form a third functional complex, and the fourth Cas protein and the fourth gRNA combine to form a fourth functional complex; and,

[0545] the first and second functional complexes bind to and cleave the double-stranded target nucleic acid to form a target nucleic acid fragment Fl, and the third and fourth functional complexes bind to and cleave the nucleic acid molecule of interest to form cleaved nucleotide fragments al, a2 and a3; and,

[0546] the first tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the target nucleic acid fragment F1 via the first target-binding sequence; and, the second tag primer hybridizes or anneals to the 3' end of the other nucleic acid strand of the target nucleic acid fragment F1 via the second target-binding sequence; and,

[0547] the first and second DNA polymerases, respectively, extend the target nucleic acid fragment F1 using the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first and second overhangs are capable of hybridizing or annealing to the broken nucleotide fragments a1 and a3, respectively; and,

[0548] the target nucleic acid fragment F2 hybridizes or anneals to the nucleotide fragments a1 and a3 via the first and second overhangs, respectively, and in turn is ligated between the nucleotide fragments a1 and a3, thereby replacing the nucleotide fragment a2 in the nucleic acid molecule of interest with the target nucleic acid fragment;

[0549] In certain embodiments, the first overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a1, and the 3' end or 3' portion is formed as a result of the breaking of the nucleic acid molecule of interest by the third functional complex;

[0550] In certain embodiments, the second overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a3, and the 3' end or 3' portion is formed as a result of the breaking of the nucleic acid molecule of interest by the fourth functional complex.

[0551] In certain embodiments, the method is performed in a cell.

[0552] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second Cas protein or nucleic acid molecule A2, the second DNA polymerase or nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third Cas protein or nucleic acid molecule A3, the third gRNA or nucleic acid molecule C3, the fourth Cas protein or nucleic acid molecule A4, and the fourth gRNA or nucleic acid molecule C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, third gRNA, fourth Cas protein, and fourth gRNA within the cell.

[0553] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the nucleic acid molecule A2, the nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the nucleic acid molecule A3, the third gRNA or nucleic acid molecule C3, the nucleic acid molecule A4, and the fourth gRNA or nucleic acid molecule C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, third gRNA, fourth Cas protein, and fourth gRNA within the cell.

[0554] In certain embodiments, in step i, the nucleic acid molecule Al, Bl, Cl, Dl, A2, B2, C2, D2, A3, C3, A4, and C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second Cas protein, second gRNA, second DNA polymerase, second tag primer, third Cas protein, third gRNA, fourth Cas protein, and fourth gRNA within the cell.

[0555] In certain embodiments, in step i, the double-stranded target nucleic acid or a nucleic acid molecule containing the double-stranded target nucleic acid T is delivered into a cell to provide the double-stranded target nucleic acid within the cell.

[0556] In certain embodiments, the double-stranded target nucleic acid or nucleic acid molecule T contains a first PAM sequence recognized by a first Cas protein and a second PAM sequence recognized by a second Cas protein. In certain embodiments, in step ii, the first functional complex binds to the double-stranded target nucleic acid or nucleic acid molecule T through the first PAM sequence and the first gRNA and cleaves it; and, the second functional complex binds to the double-stranded target nucleic acid or nucleic acid molecule T through the second PAM sequence and the second gRNA and cleaves it.

[0557] In certain embodiments, the nucleic acid molecule of interest contains a third PAM sequence recognized by a third Cas protein and a fourth PAM sequence recognized by a fourth Cas protein. In certain embodiments, in step ii, the third functional complex binds to the nucleic acid molecule of interest through the third PAM sequence and the third gRNA and cleaves it; and, the fourth functional complex binds to the nucleic acid molecule of interest through the fourth PAM sequence and the fourth gRNA and cleaves it.

[0558] In certain embodiments, the nucleic acid molecule of interest is genomic DNA of the cell.

[0559] In certain embodiments, the first Cas protein, the first gRNA, the first DNA polymerase or the first tag primer are as defined in the foregoing.

[0560] In certain embodiments, the second Cas protein, the second gRNA, the second DNA polymerase or the second tag primer are as defined in the foregoing.

[0561] In certain embodiments, the third Cas protein and the third gRNA are as defined in the foregoing.

[0562] In certain embodiments, the fourth Cas protein and the fourth gRNA are as defined in the foregoing.

[0563] In certain embodiments, the first, second, third and fourth Cas proteins are the same Cas protein, and the second DNA polymerase is the same as the first DNA polymerase; wherein, the first Cas protein forms a first, a second, a third and a fourth functional complex with the first, the second, the third and the fourth gRNA, respectively, and the first DNA polymerase extends the target nucleic acid fragment F1 with the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, respectively, to form a target nucleic acid fragment F2 with a first overhang and a second overhang.

[0564] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third gRNA or nucleic acid molecule C3, and the fourth gRNA or nucleic acid molecule C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, third gRNA, and fourth gRNA within the cell.

[0565] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third gRNA or nucleic acid molecule C3, and the fourth gRNA or nucleic acid molecule C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, third gRNA, and fourth gRNA within the cell.

[0566] In certain embodiments, in step i, the nucleic acid molecule Al, Bl, Cl, Dl, C2, D2, C3, and C4 are delivered into a cell to provide the first Cas protein, first gRNA, first DNA polymerase, first tag primer, second gRNA, second tag primer, third gRNA, and fourth gRNA within the cell.

[0567] In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule Al and nucleic acid molecule Bl are capable of expressing the first Cas protein and the first DNA polymerase separately, or a first fusion protein containing the first Cas protein and the first DNA polymerase, in a cell; in certain embodiments, in step i, nucleic acid molecules capable of expressing the first Cas protein and first DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, are delivered into a cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase within the cell.

[0568] In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C1 and the nucleic acid molecule D1 are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in a cell. In certain embodiments, in step i, the first PegRNA is delivered into the cell to provide the first gRNA and the first tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into the cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer in the cell.

[0569] In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of transcribing a second PegRNA containing the second gRNA and the second tag primer in a cell. In certain embodiments, in step i, the second PegRNA is delivered into the cell to provide the second gRNA and the second tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA is delivered into the cell and the second PegRNA is transcribed in the cell to provide the second gRNA and the second tag primer in the cell.

[0570] In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA, a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA, a nucleic acid molecule containing a nucleotide sequence encoding the third gRNA, and a nucleic acid molecule containing a nucleotide sequence encoding the fourth gRNA are delivered into the cell and transcribed and expressed in the cell, thereby providing the first Cas protein, the first gRNA, the first DNA polymerase, the first tag primer, the second gRNA, the second tag primer, the third gRNA, and the fourth gRNA in the cell.

[0571] In certain embodiments, the method comprises the following steps:

[0572] i. providing a double-stranded target nucleic acid and a nucleic acid molecule of interest; and

[0573] providing the first, second, third, and fourth Cas proteins, the first, second, third, and fourth gRNAs, the first, second, third, and fourth DNA polymerases, and the first, second, third, and fourth tag primers;

[0574] ii contacting the double-stranded target nucleic acid with the first and second Cas proteins, the first and second gRNAs, the first and second DNA polymerases, the first and second tag primers, and, contacting the nucleic acid molecule of interest with the third and fourth Cas proteins, the third and fourth gRNAs, the third and fourth DNA polymerases, and the third and fourth tag primers.

[0575] In certain embodiments, in step ii:

[0576] the first Cas protein and the first gRNA combine to form a first functional complex, the second Cas protein and the second gRNA combine to form a second functional complex, the third Cas protein and the third gRNA combine to form a third functional complex, and the fourth Cas protein and the fourth gRNA combine to form a fourth functional complex; and,

[0577] the first and second functional complexes bind to and cleave the double-stranded target nucleic acid, forming a target nucleic acid fragment F1, and, the third and fourth functional complexes bind to and cleave the nucleic acid molecule of interest, forming cleaved nucleotide fragments a1, a2, and a3; and,

[0578] the first tag primer hybridizes or anneals to the 3’ end of one nucleic acid strand of the target nucleic acid fragment F1 via the first target-binding sequence; and, the second tag primer hybridizes or anneals to the 3’ end of the other nucleic acid strand of the target nucleic acid fragment F1 via the second target-binding sequence; and,

[0579] the first and second DNA polymerases extend the target nucleic acid fragment F1 using the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, forming a target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first and second overhangs are capable of hybridizing or annealing to the cleaved nucleotide fragments a1 and a3, respectively; and,

[0580] the third tag primer hybridizes or anneals to the 3’ end of one nucleic acid strand of the nucleotide fragment a1 via the third target-binding sequence, wherein the 3’ end is formed as a result of cleavage of the nucleic acid molecule of interest by the third functional complex; and, the fourth tag primer hybridizes or anneals to the 3’ end of one nucleic acid strand of the nucleotide fragment a3 via the fourth target-binding sequence, wherein the 3’ end is formed as a result of cleavage of the nucleic acid molecule of interest by the fourth functional complex; and,

[0581] the third DNA polymerase extends the nucleotide fragment a1 using the third tag primer annealed to the nucleotide fragment a1 as a template to form a nucleotide fragment a1 with a third overhang; and the fourth DNA polymerase extends the nucleotide fragment a3 using the fourth tag primer annealed to the nucleotide fragment a3 as a template to form a nucleotide fragment a3 with a fourth overhang; wherein the third and fourth overhangs are capable of hybridizing or annealing to the target nucleic acid fragment F2, respectively; and

[0582] The target nucleic acid fragment F2 hybridizes or anneals to the nucleotide fragments a1 and a3, respectively, via the first, second, third and fourth overhangs, and is ligated between the nucleotide fragments a1 and a3, thereby replacing the nucleotide fragment a2 in the nucleic acid molecule of interest with the target nucleic acid fragment.

[0583] In certain embodiments, the first overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a1, and the 3' end or 3' portion is formed as a result of the cleavage of the nucleic acid molecule of interest by the third functional complex. In certain embodiments, the first overhang is capable of hybridizing or annealing to the third overhang of the nucleotide fragment a1 or a nucleotide sequence upstream thereof.

[0584] In certain embodiments, the second overhang is capable of hybridizing or annealing to a 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a3, and the 3' end or 3' portion is formed as a result of the cleavage of the nucleic acid molecule of interest by the fourth functional complex. In certain embodiments, the second overhang is capable of hybridizing or annealing to the fourth overhang of the nucleotide fragment a3 or a nucleotide sequence upstream thereof.

[0585] In certain embodiments, the third overhang is capable of hybridizing or annealing to the first overhang or a nucleotide sequence upstream thereof.

[0586] In certain embodiments, the fourth overhang is capable of hybridizing or annealing to the second overhang or a nucleotide sequence upstream thereof.

[0587] In certain embodiments, the method is performed in a cell.

[0588] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second Cas protein or nucleic acid molecule A2, the second DNA polymerase or nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third Cas protein or nucleic acid molecule A3, the third DNA polymerase or nucleic acid molecule B3, the third gRNA or nucleic acid molecule C3, the third tag primer or nucleic acid molecule D3, the fourth Cas protein or nucleic acid molecule A4, the fourth DNA polymerase or nucleic acid molecule B4, the fourth gRNA or nucleic acid molecule C4, and the fourth tag primer or nucleic acid molecule D4 are delivered into a cell to provide the first, second, third, and fourth Cas proteins, the first, second, third, and fourth gRNAs, the first, second, third, and fourth DNA polymerases, and the first, second, third, and fourth tag primers within the cell.

[0589] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the nucleic acid molecule A2, the nucleic acid molecule B2, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the nucleic acid molecule A3, the nucleic acid molecule B3, the third gRNA or nucleic acid molecule C3, the third tag primer or nucleic acid molecule D3, the nucleic acid molecule A4, the nucleic acid molecule B4, the fourth gRNA or nucleic acid molecule C4, and the fourth tag primer or nucleic acid molecule D4 are delivered into a cell to provide the first, second, third, and fourth Cas proteins, the first, second, third, and fourth gRNAs, the first, second, third, and fourth DNA polymerases, and the first, second, third, and fourth tag primers within the cell.

[0590] In certain embodiments, in step i, the nucleic acid molecules Al, Bl, Cl, Dl, A2, B2, C2, D2, A3, B3, C3, D3, A4, B4, C4, D4 are delivered into a cell to provide the first, second, third, and fourth Cas proteins, the first, second, third, and fourth gRNAs, the first, second, third, and fourth DNA polymerases, and the first, second, third, and fourth tag primers within the cell.

[0591] In certain embodiments, in step i, the double-stranded target nucleic acid or a nucleic acid molecule containing the double-stranded target nucleic acid T is delivered into a cell to provide the double-stranded target nucleic acid within the cell.

[0592] In certain embodiments, the double-stranded target nucleic acid or nucleic acid molecule T contains a first PAM sequence recognized by a first Cas protein and a second PAM sequence recognized by a second Cas protein; in certain embodiments, in step ii, the first functional complex binds to and cleaves the double-stranded target nucleic acid or nucleic acid molecule T through the first PAM sequence and the first gRNA; and, the second functional complex binds to and cleaves the double-stranded target nucleic acid or nucleic acid molecule T through the second PAM sequence and the second gRNA.

[0593] In certain embodiments, the nucleic acid molecule of interest contains a third PAM sequence recognized by a third Cas protein and a fourth PAM sequence recognized by a fourth Cas protein. In certain embodiments, in step ii, the third functional complex binds to and cleaves the nucleic acid molecule of interest through the third PAM sequence and the third gRNA. And, the fourth functional complex binds to and cleaves the nucleic acid molecule of interest through the fourth PAM sequence and the fourth gRNA.

[0594] In certain embodiments, the nucleic acid molecule of interest is genomic DNA of the cell.

[0595] In certain embodiments, the first Cas protein, first gRNA, first DNA polymerase, or first tag primer are as defined in the foregoing.

[0596] In certain embodiments, the second Cas protein, second gRNA, second DNA polymerase, or second tag primer are as defined in the foregoing.

[0597] In certain embodiments, the third Cas protein, third gRNA, third DNA polymerase, or third tag primer are as defined in the foregoing.

[0598] In certain embodiments, the fourth Cas protein, fourth gRNA, fourth DNA polymerase, or fourth tag primer are as defined in the foregoing.

[0599] In certain embodiments, the first, second, third, and fourth Cas proteins are the same Cas protein, and the first, second, third, and fourth DNA polymerases are the same DNA polymerase; wherein the first Cas protein forms a first, second, third, and fourth functional complex with the first, second, third, and fourth gRNAs, respectively; and the first DNA polymerase extends the target nucleic acid fragment F1 with the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; and the first DNA polymerase extends the nucleotide fragments al and a3 with the third and fourth tag primers as templates, to form third and fourth overhangs.

[0600] In certain embodiments, in step i, the first Cas protein or nucleic acid molecule Al, the first DNA polymerase or nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third gRNA or nucleic acid molecule C3, the third tag primer or nucleic acid molecule D3, the fourth gRNA or nucleic acid molecule C4, and the fourth tag primer or nucleic acid molecule D4 are delivered into a cell to provide the first Cas protein, first DNA polymerase, first, second, third, and fourth gRNAs, and first, second, third, and fourth tag primers within the cell.

[0601] In certain embodiments, in step i, the nucleic acid molecule Al, the nucleic acid molecule Bl, the first gRNA or nucleic acid molecule Cl, the first tag primer or nucleic acid molecule Dl, the second gRNA or nucleic acid molecule C2, the second tag primer or nucleic acid molecule D2, the third gRNA or nucleic acid molecule C3, the third tag primer or nucleic acid molecule D3, the fourth gRNA or nucleic acid molecule C4, and the fourth tag primer or nucleic acid molecule D4 are delivered into a cell to provide the first Cas protein, first DNA polymerase, first, second, third, and fourth gRNAs, and first, second, third, and fourth tag primers within the cell.

[0602] In certain embodiments, in step i, the nucleic acid molecules Al, Bl, Cl, Dl, C2, D2, C3, D3, C4, and D4 are delivered into a cell to provide the first Cas protein, first DNA polymerase, first, second, third, and fourth gRNAs, and first, second, third, and fourth tag primers within the cell.

[0603] In certain embodiments, the nucleic acid molecule Al and the nucleic acid molecule Bl are comprised in the same or different expression vectors (e.g., eukaryotic expression vectors). In certain embodiments, the nucleic acid molecule Al and the nucleic acid molecule Bl are capable of expressing the first Cas protein and the first DNA polymerase separately, or a first fusion protein containing the first Cas protein and the first DNA polymerase, in a cell. In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase separately, or a nucleic acid molecule containing a nucleotide sequence encoding the first fusion protein, is delivered into a cell and expressed in the cell to provide the first Cas protein and the first DNA polymerase in the cell.

[0604] In certain embodiments, the nucleic acid molecule Cl and the nucleic acid molecule Dl are comprised in the same expression vector (e.g., eukaryotic expression vector). In certain embodiments, the nucleic acid molecule Cl and the nucleic acid molecule Dl are capable of transcribing a first PegRNA containing the first gRNA and the first tag primer in a cell. In certain embodiments, in step i, the first PegRNA is delivered into a cell to provide the first gRNA and the first tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA is delivered into a cell and the first PegRNA is transcribed in the cell to provide the first gRNA and the first tag primer in the cell.

[0605] In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are comprised in the same expression vector (e.g., eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C2 and the nucleic acid molecule D2 are capable of transcribing a second PegRNA containing the second gRNA and the second tag primer in a cell. In certain embodiments, in step i, the second PegRNA is delivered into a cell to provide the second gRNA and the second tag primer in the cell, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA is delivered into a cell and the second PegRNA is transcribed in the cell to provide the second gRNA and the second tag primer in the cell.

[0606] In certain embodiments, the nucleic acid molecule C3 and the nucleic acid molecule D3 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C3 and the nucleic acid molecule D3 are capable of being transcribed in a cell to yield a third PegRNA comprising the third gRNA and the third tag primer. In certain embodiments, in step i, the third PegRNA is delivered into the cell to provide the third gRNA and the third tag primer within the cell, or a nucleic acid molecule comprising a nucleotide sequence encoding the third PegRNA is delivered into the cell and the third PegRNA is transcribed in the cell to provide the third gRNA and the third tag primer within the cell.

[0607] In certain embodiments, the nucleic acid molecule C4 and the nucleic acid molecule D4 are comprised in the same expression vector (e.g., a eukaryotic expression vector). In certain embodiments, the nucleic acid molecule C4 and the nucleic acid molecule D4 are capable of being transcribed in a cell to yield a fourth PegRNA comprising the fourth gRNA and the fourth tag primer. In certain embodiments, in step i, the fourth PegRNA is delivered into the cell to provide the fourth gRNA and the fourth tag primer within the cell, or a nucleic acid molecule comprising a nucleotide sequence encoding the fourth PegRNA is delivered into the cell and the fourth PegRNA is transcribed in the cell to provide the fourth gRNA and the fourth tag primer within the cell.

[0608] In certain embodiments, in step i, a nucleic acid molecule capable of expressing the first Cas protein and the first DNA polymerase or a nucleic acid molecule comprising a nucleotide sequence encoding the first fusion protein, a nucleic acid molecule comprising a nucleotide sequence encoding the first PegRNA, a nucleic acid molecule comprising a nucleotide sequence encoding the second PegRNA, a nucleic acid molecule comprising a nucleotide sequence encoding the third PegRNA, and a nucleic acid molecule comprising a nucleotide sequence encoding the fourth PegRNA are delivered into the cell and the first fusion protein is expressed and the first, second, third, and fourth PegRNAs are transcribed in the cell, thereby providing the first Cas protein, the first DNA polymerase, the first, second, third, and fourth gRNAs, and the first, second, third, and fourth tag primers within the cell.

[0609] Advantages of the invention

[0610] Compared with the prior art, the nucleic acid editing system, kit and method provided by the application can break double-stranded nucleic acid and extend / add one or two overhangs of arbitrary base sequences at the end (3' end) thereof. On this basis, the system, kit and method of the application can realize efficient and accurate insertion and replacement of exogenous nucleic acid (especially large fragment exogenous nucleic acid).

[0611] Embodiments of the application will be described in detail below with reference to the accompanying drawings and examples, but those skilled in the art will understand that the following drawings and examples are only used to illustrate the application, and are not a limitation on the scope of the application. According to the following detailed description of the preferred embodiments and the accompanying drawings, various objects and advantages of the application will become apparent to those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS

[0612] Figure 1 A schematic diagram showing the principle of the method of the application mediating insertion of an exogenous gene into a genome. In which, the black double solid line represents the genomic sequence or the backbone sequence of the donor vector; the blue double solid line represents the exogenous gene to be inserted; the orange and green single solid lines in the gray dashed line circle represent the first overhang and the second overhang of extension, and the orange and green double solid lines in the gray dashed line circle represent the first homologous sequence and the second homologous sequence complementary to the first overhang and the second overhang at the end of the genomic break; the black solid triangle indicates the genomic specific site spacerX, which can be recognized and cleaved by Cas9-MLV-RT / GeneX-gRNA (a complex formed by Cas9 protein, reverse transcriptase (MLV-RT) and GeneX-gRNA); the black hollow triangle indicates the sapcerA site upstream of the exogenous gene on the donor vector, which can be recognized and cleaved by Cas9-MLV-RT / spacerA-pegRNA (a complex formed by Cas9 protein, reverse transcriptase (MLV-RT) and spacerA-pegRNA); the gray hollow triangle indicates the sapcerK site downstream of the exogenous gene on the donor vector, which can be recognized and cleaved by Cas9-MLV-RT / spacerK-pegRNA (a complex formed by Cas9 protein, reverse transcriptase (MLV-RT) and spacerK-pegRNA); and spacerA and sapcerK are located on opposite nucleic acid strands of each other.

[0613] The sapcerA site is recognized and cleaved by Cas9-MLV-RT / spacerA-pegRNA, wherein the spacerA-pegRNA contains a first target binding sequence and a first tag sequence (which is complementary to one strand of the first homologous sequence), the first target binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the first tag sequence does not bind to the target nucleic acid fragment and is in a free single-stranded state. Therefore, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand using the first tag primer as a template, forming a first overhang (i.e., the orange single solid line, which can be, for example, 35 nt in length). Similarly, the spacerK site is recognized and cleaved by Cas9-MLV-RT / spacerK-pegRNA, wherein the spacerK-pegRNA contains a second target binding sequence and a second tag sequence (which is complementary to one strand of the second homologous sequence), the second target binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the second tag sequence does not bind to the target nucleic acid fragment and is in a free single-stranded state. Therefore, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand using the second tag primer as a template, forming a second overhang (i.e., the green single solid line, which can be, for example, 35 nt in length). Through double cleavage, the exogenous gene fragment is cleaved from the vector and has a first overhang and a second overhang added at both ends.

[0614] In addition, the spacerX site is recognized and cleaved by Cas9-MLV-RT / GeneX-gRNA to form a broken genome; and the two ends at the broken site respectively contain a first homologous sequence (complementary to the first overhang) and a second homologous sequence (complementary to the second overhang). Thus, the exogenous gene fragment with the first overhang and the second overhang can be integrated into the broken site of the genome through interchain annealing, achieving site-specific insertion of the exogenous gene.

[0615] Figure 2A schematic diagram showing the principle of the method of the present application mediating the replacement of a specific nucleotide fragment of the genome with an exogenous gene. In which, the black double solid line represents the sequence of the genome or the backbone sequence of the donor vector; the black double dashed line represents the genomic fragment to be replaced; the blue double solid line represents the exogenous inserted gene to be replaced; the same color solid line in the gray dashed line circle represents the homologous sequence capable of complementing each other, the orange single solid line represents the first overhang extended, the green single solid line represents the second overhang extended, the red single solid line represents the third overhang extended, and the purple single solid line represents the fourth overhang extended; the black solid triangle indicates the cleavage site RC-PegRNA upstream of the fragment to be replaced on the genome, which can be recognized, cleaved and extended by Cas9-MLV-RT / RC-PegRNA (a complex formed by Cas9 protein, reverse transcriptase (MLV-RT) and RC-PegRNA); the gray solid triangle indicates the cleavage site RT-PegRNA downstream of the fragment to be replaced on the genome, which can be recognized, cleaved and extended by Cas9-MLV-RT / RT-PegRNA (a complex formed by Cas9 protein, reverse transcriptase (MLV-RT) and RT-PegRNA); the black hollow triangle represents the cleavage site RC-pegA upstream of the exogenous inserted fragment on the donor vector, which can be recognized, cleaved and extended by Cas9-MLV-RT / RC-pegA; the gray hollow triangle represents the cleavage site RT-pegK downstream of the exogenous inserted fragment on the donor vector, which can be recognized, cleaved and extended by Cas9-MLV-RT / RT-pegK; the RC-pegA and RT-pegK sites are located on opposite nucleic acid strands of each other; and the RC-PegRNA and RT-PegRNA sites are located on opposite nucleic acid strands of each other.

[0616] The site RC-pegA is recognized and cleaved by Cas9-MLV-RT / RC-PegRNA, wherein the RC-PegRNA contains a first target binding sequence and a first tag sequence (which is complementary to one strand of the first homologous sequence), the first target binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the first tag sequence is not bound to the target nucleic acid fragment and is in a free single-stranded state. Therefore, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand with the first tag primer as a template to form a first overhang (i.e. orange single solid line), and the first overhang can be complementary to the nucleotide sequence (orange double solid line) upstream of the RC-PegRNA site.

[0617] Similarly, site RT-pegK is recognized and cleaved by Cas9-MLV-RT / RT-pegK, where RT-pegK contains a second target-binding sequence and a second tag sequence (which is complementary to one strand of a second homologous sequence), the second target-binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the second tag sequence does not bind to the target nucleic acid fragment, being in a free single-stranded state. Thus, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand using the second tag primer as a template to form a second overhang (i.e., the green single solid line), and the second overhang can be complementary to the nucleotide sequence downstream of the RT-PegRNA site (green double solid line).

[0618] By double cleavage, the exogenous gene fragment is cleaved from the vector, and the first overhang and the second overhang are added at both ends.

[0619] Site RC-PegRNA is recognized and cleaved by Cas9-MLV-RT / RC-PegRNA, where RC-PegRNA contains a third target-binding sequence and a third tag sequence (which is complementary to one strand of a third homologous sequence), the third target-binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the third tag sequence does not bind to the target nucleic acid fragment, being in a free single-stranded state. Thus, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand using the third tag primer as a template to form a third overhang (i.e., the red single solid line), and the third overhang can be complementary to the nucleotide sequence downstream of the RC-pegA site (red double solid line).

[0620] Site RT-PegRNA is recognized and cleaved by Cas9-MLV-RT / RT-PegRNA, where RT-PegRNA contains a fourth target-binding sequence and a fourth tag sequence (which is complementary to one strand of a fourth homologous sequence), the fourth target-binding sequence hybridizes to the 3' end of one nucleic acid strand of the cleaved target nucleic acid fragment to form a double-stranded structure, and the fourth tag sequence does not bind to the target nucleic acid fragment, being in a free single-stranded state. Thus, the reverse transcriptase (MLV-RT) can extend the 3' end of the nucleic acid strand using the fourth tag primer as a template to form a fourth overhang (i.e., the purple single solid line), and the fourth overhang can be complementary to the nucleotide sequence upstream of the RT-pegK site (purple double solid line).

[0621] By double cleavage, the fragment to be replaced is excised from the genome, and the third overhang and the fourth overhang are added at both ends of the broken genome.

[0622] Thus, the exogenous gene fragment with the first and second overhangs can be inserted into the broken genome with the third and fourth overhangs by interchain annealing, thereby achieving the replacement of a specific nucleotide fragment on the genome.

[0623] Figure 3 A flowchart showing the site-directed knock-in of an exogenous gene (IRES-EGFP) into the GAPDH gene of a human cell genome.

[0624] Figure 4 A flowchart showing the site-directed knock-in of an exogenous gene (IRES-EGFP) into the GAPDH gene of a human cell genome. Figure 4 A shows the ratio of EGFP positive cells produced by different methods analyzed by flow cytometric fluorescence sorting technique (FACS). Figure 4 B is the result of PCR identification of the nucleotide sequence at the junction of the reporter gene IRES-EGFP (5' end and 3' end) and genomic DNA. Figure 4 C is the result of Sanger sequencing of the nucleotide sequence at the junction of the reporter gene IRES-EGFP (5' end and 3' end) and genomic DNA.

[0625] Figure 5 A flowchart showing the site-directed knock-in of an exogenous gene (IRES-EGFP) into the ACTB gene of a human cell genome.

[0626] Figure 6 A flowchart showing the site-directed knock-in of an exogenous gene (IRES-EGFP) into the ACTB gene of a human cell genome. Figure 6 A shows the ratio of EGFP positive cells produced by different methods analyzed by flow cytometric fluorescence sorting technique (FACS). Figure 6 B is the result of PCR identification of the nucleotide sequence at the junction of the reporter gene IRES-EGFP (5' end and 3' end) and genomic DNA. Figure 6 C is the result of Sanger sequencing of the nucleotide sequence at the junction of the reporter gene IRES-EGFP (5' end and 3' end) and genomic DNA.

[0627] Figure 7 A flowchart showing the site-directed replacement of a nucleotide fragment in the GAPDH gene of a human cell genome with a target nucleic acid fragment containing an exogenous gene (T2A-EGFP).

[0628] Figure 8 A flowchart showing the site-directed replacement of a nucleotide fragment in the GAPDH gene of a human cell genome with a target nucleic acid fragment containing an exogenous gene (T2A-EGFP).

[0629] Figure 9 A shows a schematic diagram of the results of using the HDR method and the method of the present application (EPTI) to knock in an exogenous gene (T2A-EGFP) at the stop codon of the ACTB gene in 293T cells, so as to express T2A-EGFP in fusion with the ACTB gene.

[0630] Figure 9 A shows a schematic diagram of the results of using the HDR method and the method of the present application (EPTI) to knock in an exogenous gene (T2A-EGFP) at the stop codon of the ACTB gene in 293T cells, so as to express T2A-EGFP in fusion with the ACTB gene.

[0631] Figure 9 B shows a sequence schematic diagram of EPTI-mediated knock-in of an exogenous gene (T2A-EGFP) at the stop codon of the ACTB gene.

[0632] In the first sequence list, the sequence represents the human ACTB gene sequence; the blue sequence represents the protein coding sequence of the ACTB gene (the blue sequence is also a homologous sequence); the black triangle represents the targeted cleavage site of the genome (sgACTB2); wherein, TAG is the stop codon, and the targeted cleavage site of the genome is between the "T" base and the "A" base of the stop codon.

[0633] The second sequence list represents the sequence of the donor vector, wherein the black triangle represents the targeted cleavage site upstream of the exogenous gene on the donor vector. According to the above description, when the site is recognized and cleaved by the complex of Cas9-MLV-RT and pegRNA, the pegRNA extends the 3' end of the nucleic acid strand upstream of the exogenous gene to form an overhang sequence.

[0634] The third sequence list represents the sequence after the overhang sequence upstream of the exogenous gene and the homologous sequence of the ACTB gene anneal; due to the interstrand annealing, the spacer sequence (for example, the base "T") between the homologous sequence of the ACTB gene and the break site forms a free base;

[0635] In the fourth sequence list, the blue sequence is the homologous sequence of the ACTB gene; the blue sequence in the gray box is the overhang sequence formed upstream of the exogenous gene; the red sequence in the gray box is the sequence of the donor vector;

[0636] The fourth sequence list represents the sequence after the exogenous gene (T2A-EGFP) is knocked in at the stop codon of the ACTB gene, wherein the free "T" base is removed, realizing the continuity of the open reading frame and the fusion expression of the protein.

[0637] Figure 9C shows the efficiency comparison of HDR and EPTI-mediated site-directed knock-in of a foreign gene (T2A-EGFP) before the stop codon of ACTB gene.

[0638] Figure 9 D is the result of PCR identification of the nucleotide sequence at the junction of the reporter gene EGFP (5' end and 3' end) and genomic DNA.

[0639] Sequence information

[0640] Information of part of the sequences involved in the present application is provided in Table 1 below.

[0641] Table 1: Description of sequences

[0642]

[0643]

[0644]

[0645]

[0646]

[0647]

[0648]

[0649] DETAILED DESCRIPTION

[0650] The present application will now be described with reference to the following examples, which are intended to illustrate the present application (but not to limit the present application).

[0651] Unless specifically indicated otherwise, the experiments and methods described in the examples were performed essentially according to conventional methods well known in the art and described in various references. For example, the general techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, among others, used in the present application can be found in Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd Ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F. M. Ausubel et al. eds., (1987)); the series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (M. J. MacPherson, B. D. Hames, and G. R. Taylor eds., (1995)), and ANIMAL CELL CULTURE (R. I. Freshney ed., (1987)).

[0652] In addition, unless otherwise specified, the examples were performed under conventional conditions or under conditions recommended by the manufacturer. The reagents or instruments used, when not specified by the manufacturer, were all conventional products available on the market. The person skilled in the art knows that the examples describe the application by way of example and are not intended to limit the scope of the application as claimed. All the publications and other references mentioned herein are incorporated by reference in their entirety.

[0653] Example 1. Site-directed insertion of a foreign gene into the human GAPDH gene using the EPTI system

[0654] To verify the effect of EPTI system on site-specific insertion of exogenous gene into genome, the following experiment was designed: using EPTI system to knock-in reporter gene IRES-EGFP into 3'UTR region of human GAPDH gene, and using HITI system as control. The principle of EPTI system for site-specific insertion of exogenous gene is shown in Figure 1 The specific process of site-specific knock-in of exogenous gene (IRES-EGFP) into human cell genome GAPDH gene using EPTI system is shown in Figure 3

[0655] ​GAPDH gene is located in chromosome 12, encoding glycerolaldehyde-3-phosphate dehydrogenase, which is an important housekeeping gene and has high expression abundance in 293T cells. The reporter gene is knocked into the 3'UTR region of GAPDH, which can be transcribed together with the GAPDH gene, and the IRES sequence in it can recruit ribosomes, so that EGFP can be expressed. The fluorescence signal of EGFP can be directly observed and quantified by fluorescence microscope, and the cells expressing EGFP can be captured and quantified by flow cytometry.

[0656] The pCAG-Cas9-mCherry plasmid (which can express Cas9 protein (SEQ ID NO: 1) and mCherry protein (SEQ ID NO: 2)) and pUC19-U6-gRNA (which can transcribe gRNA (SEQ ID NO: 3) lacking a guide sequence) used in this example were obtained from the Li Wei group of the Institute of Animal Sciences, Chinese Academy of Sciences.

[0657] The nucleotide fragment encoding MLV-TR (SEQ ID NO: 4) was amplified from the pCMV-PE2 (#132775) plasmid from addgene company, and the partial nucleotide fragment encoding Cas9 and the nucleotide fragment encoding mCherry were amplified from the pCAG-Cas9-mCherry plasmid. The above-mentioned amplified nucleotide fragments were connected to the pCAG-Cas9-mCherry plasmid digested by AscI / BsrGI by In-fusion cloning technology, to obtain the pCAG-Cas9-MLV RT-mCherry plasmid, which can express Cas9 protein, MLV-TR protein and mCherry protein.

[0658] The primers Gapdh-gRNA-F (SEQ ID NO: 5) and Gapdh-gRNA-R (SEQ ID NO: 6) were annealed and ligated to the pUC19-U6-gRNA plasmid digested by BsaI with T4 ligase, to obtain the pUC19-U6-Gapdh-gRNA plasmid, which can transcribe Gapdh-gRNA (SEQ ID NO: 7) to guide Cas9 protein to target the 3'URT region of human GAPDH site.

[0659] Primers spacerA-gRNA-F (SEQ ID NO: 17) and spacerA-gRNA-R (SEQ ID NO: 18) were annealed and ligated by T4 ligase to the pUC19-U6-gRNA plasmid vector digested with Bsal enzyme to obtain the pUC19-U6-spacerA-gRNA plasmid, which can transcribe spacerA-gRNA (SEQ ID NO: 19) to guide the Cas9 protein to target the spacer A sequence (SEQ ID NO: 49).

[0660] Primers spacerK-gRNA-F (SEQ ID NO: 20) and spacerK-gRNA-R (SEQ ID NO: 21) were annealed and ligated by T4 ligase to the pUC19-U6-gRNA plasmid vector digested with Bsal enzyme to obtain the pUC19-U6-spacerK-gRNA plasmid, which can transcribe spacerK-gRNA (SEQ ID NO: 22) to guide the Cas9 protein to target the spacer K sequence (SEQ ID NO: 50).

[0661] Primers Gapdh-pegA-F (SEQ ID NO: 23) and Gapdh-pegA-R (SEQ ID NO: 24) were subjected to overlap extension PCR, and the obtained fragment was recycled and ligated by In-fusion cloning technology to the pUC19-U6-spacerA-gRNA plasmid vector digested with Hindlll enzyme to obtain the pUC19-U6-Gapdh-pegA plasmid, which can transcribe Gapdh-pegA (SEQ ID NO: 25) to guide the Cas9 protein to target the spacer A sequence (SEQ ID NO: 49).

[0662] Primers Gapdh-pegK-F (SEQ ID NO: 26) and Gapdh-pegK-R (SEQ ID NO: 27) were subjected to overlap extension PCR, and the obtained fragment was recycled and ligated by In-fusion cloning technology to the pUC19-U6-spacerK-gRNA plasmid vector digested with Hindlll enzyme to obtain the pUC19-U6-Gapdh-pegK plasmid, which can transcribe Gapdh-pegK (SEQ ID NO: 28) to guide the Cas9 protein to target the spacer K sequence (SEQ ID NO: 50).

[0663] The reporter IRES-EGFP (SEQ ID NO: 47) was synthesized by Jierui Company and ligated to the EcoRV-digested pGH vector (provided by Jierui Company) by T4 ligase as the donor vector. The reporter gene has spacer A-gRNA / Gapdh-pegA and spacer K-gRNA / Gapdh-pegK recognition and cleavage sites spacer A and spacer K (sequences are SEQ ID NO: 49 and SEQ ID NO: 50, respectively) on both sides.

[0664] In the implementation of the EPTI system, pCAG-Cas9-MLV RT-mCherry, Gapdh-gRNA, Gapdh-pegA, Gapdh-pegK together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent from Invitrogen Company. In the negative control group of the EPTI system, pCAG-Cas9-MLV RT-mCherry, Gapdh-pegA, Gapdh-pegK together with the donor vector were transfected into 293T cells. The 293T cell line was from the ATCC cell library. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were cultured for 5 days, after which the ratio of EGFP-positive cells was analyzed by flow cytometry. In the implementation of the HITI system, pCAG-Cas9-mCherry, Gapdh-gRNA, spacer A-gRNA, spacer K-gRNA together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent from Invitrogen Company. In the negative control group of the HITI system, pCAG-Cas9-mCherry, spacer A-gRNA, spacer K-gRNA together with the donor vector were transfected into 293T cells. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were cultured for 5 days, after which the ratio of EGFP-positive cells was analyzed by flow cytometry. Comparing the proportion of EGFP-positive cells in the two different systems can reflect the difference in the efficiency of site-directed knock-in of the GAPDH gene in the two systems. The exogenous gene integration efficiency results are shown in FIG. 2A and FIG. 2B. Figure 4 The results show that the ratio of EGFP-positive cells in the HITI system is less than 5%, while in the EPTI system, the ratio of EGFP-positive cells is more than 60%. And in the negative control group without Gapdh-gRNA components, no EGFP-positive cells were detected in both the HITI system and the EPTI system, indicating that EGFP-positive cells can reflect the specific integration of the exogenous gene IRES-EGFP on the donor vector at the human GAPDH target site.

[0665] Genomic DNA of EGFP positive cells was extracted, and then the nucleotide sequences at the junctions of the reporter gene IRES-EGFP (5' end and 3' end) and the genomic DNA were identified by PCR and Sanger sequencing analysis using primers GAPDH-P1 (SEQ ID NO: 68) / GAPDH-P2 (SEQ ID NO: 69), GAPDH-P3 (SEQ ID NO: 70) / GAPDH-P4 (SEQ ID NO: 71), respectively.

[0666] The PCR identification results are shown in Figure 4 B. The results show that the primers GAPDH-P1 / GAPDH-P2 and GAPDH-P3 / GAPDH-P4 used can amplify two fragments with expected sizes using the genomic DNA of EGFP positive cells as a template. The Sanger sequencing analysis results are shown in Figure 4 C. The results show that the EPTI method of the present application can efficiently, site-specifically and accurately mediate the ligation of an exogenous gene to the broken ends of the genomic DNA. In summary, the EPTI system described in the present application can greatly improve the site-specific integration efficiency of an exogenous gene.

[0667] Example 2. Site-directed insertion of a foreign gene into the human ACTB gene using the EPTI system

[0668] To verify the effect of the EPTI system on site-specific insertion of an exogenous gene into the genome, the following experiment was designed: the reporter gene IRES-EGFP (SEQ ID NO: 47) was site-specifically knocked into the 3' UTR region of the human ACTB gene using the EPTI system, and the HITI system was used as a control. The specific procedure for site-specifically knocking an exogenous gene (IRES-EGFP) into the ACTB gene of the human cell genome is shown in Figure 5 The synthesis and cleavage steps of the reporter gene were the same as in Example 1.

[0669] Primers Actb-gRNA-F (SEQ ID NO: 8) and Actb-gRNA-R (SEQ ID NO: 9) were annealed and ligated to the pUC19-U6-gRNA plasmid digested with BsaI enzyme using T4 ligase to obtain the pUC19-U6-Actb-gRNA plasmid, which can transcribe Actb-gRNA (SEQ ID NO: 10) to guide the Cas9 protein to target the 3' UTR region of the human ACTB site.

[0670] The primers Actb-pegA-F (SEQ ID NO: 29) and Actb-pegA-R (SEQ ID NO: 30) were subjected to overlap extension PCR, and the obtained fragment was recovered and ligated to the pUC19-U6-spacerA-gRNA plasmid vector digested with Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-Actb-pegA plasmid, which can transcribe Actb-pegA (SEQ ID NO: 31) to guide the Cas9 protein to target the spacer A sequence (SEQ ID NO: 49).

[0671] The primers Actb-pegK-F (SEQ ID NO: 32) and Actb-pegK-R (SEQ ID NO: 33) were subjected to overlap extension PCR, and the obtained fragment was recovered and ligated to the pUC19-U6-spacerK-gRNA plasmid vector digested with Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-Actb-pegK plasmid, which can transcribe Actb-pegK (SEQ ID NO: 34) to guide the Cas9 protein to target the spacer K sequence (SEQ ID NO: 50).

[0672] In the implementation of the EPTI system, pCAG-Cas9-MLV RT-mCherry, Actb-gRNA, Actb-pegA, Actb-pegK together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent of Invitrogen. In the negative control group of the EPTI system, pCAG-Cas9-MLV RT-mCherry, Actb-pegA, Actb-pegK together with the donor vector were transfected into 293T cells. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were further cultured for 5 days. The ratio of EGFP-positive cells was analyzed by flow cytometry. In the implementation of the HITI system, pCAG-Cas9-mCherry, Actb-gRNA, spacerA-gRNA, spacerK-gRNA together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent of Invitrogen. In the negative control group of the HITI system, pCAG-Cas9-mCherry, spacerA-gRNA, spacerK-gRNA together with the donor vector were transfected into 293T cells. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were further cultured for 5 days. The ratio of EGFP-positive cells was analyzed by flow cytometry. By comparing the ratio of EGFP-positive cells in the two different systems, the difference in the efficiency of the two systems in the site-directed knock-in of the ACTB gene can be reflected. The results of the comparison of the integration efficiency of exogenous genes by different methods are shown in FIG. 6. Figure 6 As shown in FIG. 6, the results show that the ratio of EGFP-positive cells in the HITI system is below 8%, while the ratio of EGFP-positive cells in the EPTI system is more than 50%. And in the negative control group without the Actb-gRNA component, no EGFP-positive cells were detected in both the HITI system and the EPTI system, indicating that EGFP-positive cells can reflect the specific integration of the exogenous gene IRES-EGFP on the donor vector at the human ACTB target site.

[0673] The genomic DNA of the EGFP-positive cells was extracted, and then the nucleotide sequences at the junctions of the reporter gene IRES-EGFP (5' end and 3' end) and the genomic DNA were identified by PCR and analyzed by Sanger sequencing using primers ACTB-P1 (SEQ ID NO: 72) / ACTB-P2 (SEQ ID NO: 73), ACTB-P3 (SEQ ID NO: 74) / ACTB-P4 (SEQ ID NO: 75), respectively.

[0674] The PCR identification results are shown in FIG. 7.Figure 6 The results show that the primers ACTB-P1 / ACTB-P2, ACTB-P3 / ACTB-P4 used can amplify two fragments with expected sizes from the genomic DNA of EGFP positive cells. The results of Sanger sequencing analysis are shown in FIG. 3B. The results show that the EPTI method of the present application can efficiently, site-specifically and accurately mediate the ligation of the exogenous gene to the broken end of the genomic DNA. In summary, the EPTI system described in the present application can greatly improve the site-specific integration efficiency of the exogenous gene. Figure 6 The results show that the primers ACTB-P1 / ACTB-P2, ACTB-P3 / ACTB-P4 used can amplify two fragments with expected sizes from the genomic DNA of EGFP positive cells. The results of Sanger sequencing analysis are shown in FIG. 3B. The results show that the EPTI method of the present application can efficiently, site-specifically and accurately mediate the ligation of the exogenous gene to the broken end of the genomic DNA. In summary, the EPTI system described in the present application can greatly improve the site-specific integration efficiency of the exogenous gene.

[0675] Example 3. Site-directed replacement of a nucleotide fragment in the human GAPDH gene with a foreign gene using the EPTI system

[0676] To verify the effect of the EPTI system on site-specific replacement of the exogenous gene into the genome, the following experiment was designed: using the EPTI system to site-specifically replace the DNA fragment containing the reporter gene T2A-EGFP (SEQ ID NO: 48) into the DNA sequence of the human GAPDH gene, and using the HITI system as a control. The schematic diagram of the principle of the EPTI method mediating the replacement of the exogenous gene into the specific nucleotide fragment of the genome is shown in FIG. 4A, and the specific process of site-specific replacement of the target nucleic acid fragment containing the exogenous gene (T2A-EGFP) into the nucleotide fragment of the GAPDH gene of the human cell genome is shown in FIG. 4B. Figure 2 The results show that the primers ACTB-P1 / ACTB-P2, ACTB-P3 / ACTB-P4 used can amplify two fragments with expected sizes from the genomic DNA of EGFP positive cells. The results of Sanger sequencing analysis are shown in FIG. 3B. The results show that the EPTI method of the present application can efficiently, site-specifically and accurately mediate the ligation of the exogenous gene to the broken end of the genomic DNA. In summary, the EPTI system described in the present application can greatly improve the site-specific integration efficiency of the exogenous gene. Figure 7 The results show that the primers ACTB-P1 / ACTB-P2, ACTB-P3 / ACTB-P4 used can amplify two fragments with expected sizes from the genomic DNA of EGFP positive cells. The results of Sanger sequencing analysis are shown in FIG. 3B. The results show that the EPTI method of the present application can efficiently, site-specifically and accurately mediate the ligation of the exogenous gene to the broken end of the genomic DNA. In summary, the EPTI system described in the present application can greatly improve the site-specific integration efficiency of the exogenous gene.

[0677] The primers GapdhRC-gRNA-F (SEQ ID NO: 11) and GapdhRC-gRNA-R (SEQ ID NO: 12) were annealed and ligated to the pUC19-U6-gRNA plasmid digested with Bsal enzyme using T4 ligase to obtain the pUC19-U6-GapdhRC-gRNA plasmid, which can transcribe GapdhRC-gRNA (SEQ ID NO: 13) to guide the Cas9 protein to target the intron region between the 4th and 5th exons of the human GAPDH gene.

[0678] The primers GapdhRT-gRNA-F (SEQ ID NO: 14) and GapdhRT-gRNA-R (SEQ ID NO: 15) were annealed and ligated to the pUC19-U6-gRNA plasmid digested with Bsal enzyme using T4 ligase to obtain the pUC19-U6-GapdhRT-gRNA plasmid, which can transcribe GapdhRT-gRNA (SEQ ID NO: 16) to guide the Cas9 protein to target the downstream region of the human GAPDH gene.

[0679] The primers GapdhRC-pegRNA-F (SEQ ID NO: 35) and GapdhRC-pegRNA-F (SEQ ID NO: 36) were subjected to overlap extension PCR, and the obtained fragment was recovered and connected to the pUC19-U6-GapdhRC-gRNA plasmid vector digested by Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-GapdhRC-pegRNA plasmid, which can transcribe GapdhRC-pegRNA (SEQ ID NO: 37) to guide the Cas9 protein to target the intron region between the 4th and 5th exons of the human GAPDH gene.

[0680] The primers GapdhRT-pegRNA-F (SEQ ID NO: 38) and GapdhRT-pegRNA-F (SEQ ID NO: 39) were subjected to overlap extension PCR, and the obtained fragment was recovered and connected to the pUC19-U6-GapdhRT-gRNA plasmid vector digested by Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-GapdhRT-pegRNA plasmid, which can transcribe GapdhRT-pegRNA (SEQ ID NO: 40) to guide the Cas9 protein to target the downstream region of the human GAPDH gene.

[0681] The primers GapdhRC-pegA-F (SEQ ID NO: 41) and GapdhRC-pegA-R (SEQ ID NO: 42) were subjected to overlap extension PCR, and the obtained fragment was recovered and connected to the pUC19-U6-spacerA-gRNA plasmid vector digested by Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-GapdhRC-pegA plasmid, which can transcribe GapdhRC-pegA (SEQ ID NO: 43) to guide the Cas9 protein to target the spacerA sequence (SEQ ID NO: 49).

[0682] The primers GapdhRT-pegK-F (SEQ ID NO: 44) and GapdhRT-pegK-R (SEQ ID NO: 45) were subjected to overlap extension PCR, and the obtained fragment was recovered and connected to the pUC19-U6-spacerK-gRNA plasmid vector digested by Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-GapdhRT-pegK plasmid, which can transcribe GapdhRT-pegK (SEQ ID NO: 46) to guide the Cas9 protein to target the spacerK sequence (SEQ ID NO: 50).

[0683] The exogenous gene fragment containing the reporter gene T2A-EGFP (SEQ ID NO: 48) was synthesized by Jiery Company and ligated to the EcoRV-digested pGH vector (provided by Jiery Company) by T4 ligase as the donor vector. The reporter gene has the recognition and cleavage sites spacerA and spacerK (sequences are SEQ ID NO: 49 and SEQ ID NO: 50, respectively) of spacerA-gRNA / Gapdh-pegA and spacerK-gRNA / Gapdh-pegK on both sides.

[0684] In the implementation of the EPTI system, pCAG-Cas9-MLV RT-mCherry, GapdhRC-pegRNA, GapdhRT-pegRNA, GapdhRC-pegA, GapdhRT-pegK, together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent of Invitrogen Company. In the negative control group of the EPTI system, pCAG-Cas9-MLV RT-mCherry, GapdhRC-pegA, GapdhRT-pegK, together with the donor vector were transfected into 293T cells. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were cultured for 5 days. The ratio of EGFP-positive cells was analyzed by flow cytometry. In the implementation of the HITI system, pCAG-Cas9-mCherry, GapdhRC-gRNA, GapdhRT-gRNA, spacerA-gRNA, spacerK-gRNA, together with the donor vector were transfected into 293T cells using Lipofectamine 3000 liposome transfection reagent of Invitrogen Company. In the negative control group of the HITI system, pCAG-Cas9-mCherry, spacerA-gRNA, spacerK-gRNA, together with the donor vector were transfected into 293T cells. The mCherry-positive cells were sorted by flow cytometry 24 hours after transfection, and the sorted cells were cultured for 5 days. The ratio of EGFP-positive cells was analyzed by flow cytometry. By comparing the proportion of EGFP-positive cells in the two different systems, the difference in the efficiency of site-directed replacement of the GAPDH gene in the two systems can be reflected.

[0685] The experimental results are as follows Figure 8As shown, the proportion of EGFP-positive cells in the HITI system is around 3%, while in the EPTI system, the proportion is higher than 30%. Furthermore, in the negative control group where the sgRNA / pegRNA component targeting the GAPDH gene was removed, neither the HITI nor EPTI systems detected any EGFP-positive cells, indicating that EGFP-positive cells reflect site-specific substitution of the exogenous gene fragment integrated with T2A-EGFP on the donor vector at the human GAPDH target site with the genomic fragment. In summary, the EPTI system described in this invention can significantly improve the efficiency of site-specific substitution of exogenous genes at the human GAPDH gene site.

[0686] Example 4. Site-directed insertion of a foreign gene into the human ACTB gene using the EPTI system ACTB-T2A-EGFP fusion protein

[0687] To further verify the effectiveness of the EPTI system in precisely inserting exogenous genes into the genome, this embodiment designed the following experiment: The reporter gene T2A-EGFP (SEQ ID NO:82) was knocked into the ACTB gene of 293T cells before the stop codon using the EPTI system, with the HDR system as a control. The procedure for precisely knocking in the exogenous gene (T2A-EGFP) into the ACTB gene of the 293T cell genome is as follows: Figure 9 As shown in A and 9B.

[0688] like Figure 9 As shown in Figure B, the targeted cleavage site of the ACTB gene is not between the last coding codon and the stop codon, but between the bases "T" and "A" of the stop codon. To ensure the foreign gene fragment (T2A-EGFP) can be inserted precisely after the last coding codon of the ACTB gene for fusion expression, the "T" base of the stop codon needs to be removed during genome editing. For this purpose, a tag sequence in the pegRNA is designed such that the overhang (first overhang) on ​​the resulting foreign gene fragment is complementary to the upstream of the targeted cleavage site of the ACTB gene; that is, there is a spacer sequence between the genomic sequence targeted by the first overhang and the targeted cleavage site of the ACTB gene (in this embodiment, the spacer sequence is the "T" base to be removed). With this design, after the first overhang undergoes interstrand annealing with its targeted genomic sequence strand, the spacer sequence becomes free and is cleaved; thus, the foreign gene fragment (T2A-EGFP) can be precisely linked after the last coding codon of the ACTB gene.

[0689] The pCAG-Cas9-mCherry plasmid and pCAG-Cas9-MLV RT-mCherry used in this embodiment are the same as those in Example 1.

[0690] The primers sgACTB2-F (SEQ ID NO: 53) and sgACTB2-R (SEQ ID NO: 54) were annealed and ligated with T4 ligase to the pUC19-U6-gRNA plasmid digested with Bsal enzyme to obtain the pUC19-U6-sgACTB2 plasmid, which can transcribe sgACTB2 (SEQ ID NO: 55) to guide Cas9 protein to target and cut the specific site of the ACTB gene of 293T cells (i.e., to cut between the base "T" and the base "A" of the termination codon of the ACTB gene).

[0691] The primers ACTB2-sgL-F (SEQ ID NO: 56) and ACTB2-sgL-R (SEQ ID NO: 57) were annealed and ligated with T4 ligase to the pUC19-U6-gRNA plasmid vector digested with Bsal enzyme to obtain the pUC19-U6-ACTB2-sgL plasmid, which can transcribe ACTB2-sgL (SEQ ID NO: 58).

[0692] The primers ACTB2-sgR-F (SEQ ID NO: 59) and ACTB2-sgR-R (SEQ ID NO: 60) were annealed and ligated with T4 ligase to the pUC19-U6-gRNA plasmid vector digested with Bsal enzyme to obtain the pUC19-U6-ACTB2-sgR plasmid, which can transcribe ACTB2-sgR (SEQ ID NO: 61).

[0693] The primers ACTB2-pegL-F (SEQ ID NO: 62) and ACTB2-pegL-R (SEQ ID NO: 63) were subjected to overlap extension PCR, and the obtained fragment was ligated to the pUC19-U6-ACTB2-sgL plasmid vector digested with Hindlll enzyme by In-fusion cloning technology to obtain the pUC19-U6-ACTB2-pegL plasmid, which can transcribe ACTB2-pegL (SEQ ID NO: 64) to guide Cas9 protein to target the spacer L sequence (SEQ ID NO: 80).

[0694] The primers ACTB2-pegR-F (SEQ ID NO: 65) and ACTB2-pegR-R (SEQ ID NO: 66) were subjected to overlap extension PCR, and the obtained fragment was recovered and ligated to the pUC19-U6-ACTB-pegR plasmid vector digested with Hind III enzyme by In-fusion cloning technology to obtain the pUC19-U6-ACTB-pegR plasmid, which can transcribe ACTB2-pegR (SEQ ID NO: 67) to guide the Cas9 protein to target the spacer R sequence (SEQ ID NO: 81).

[0695] The reporter gene T2A-EGFP (SEQ ID NO: 82) was synthesized by Jiery Company and ligated to the pGH vector (provided by Jiery Company) digested with EcoRV enzyme by T4 ligase as a donor vector. For the donor vector of the EPTI system, the reporter gene T2A-EGFP has the spacer L sequence (SEQ ID NO: 80) on both sides, which is reverse to the T2A-EGFP gene; and the spacer R sequence (SEQ ID NO: 81), which is the same direction as the T2A-EGFP gene. For the donor vector of the HDR system, the reporter gene T2A-EGFP has the left homologous arm ACTB2 LHA (SEQ ID NO: 83) and the right homologous arm ACTB2 RHA (SEQ ID NO: 84) on both sides.

[0696] In the experiment using the EPTI system, pCAG-Cas9-MLV RT-mCherry, U6-sgACTB2, U6-ACTB2-pegL, U6-ACTB2-pegR together with the donor vector were transfected into 293T cells using the Lipofectamine 3000 liposome transfection reagent of Invitrogen. In the negative control group of the EPTI system, pCAG-Cas9-MLV RT-mCherry, U6-ACTB2-pegL, U6-ACTB2-pegR together with the donor vector were transfected into 293T cells. In the experiment using the HDR system, pCAG-Cas9-mCherry, U6-sgACTB2 together with the HDR donor vector were transfected into 293T cells using the Lipofectamine 3000 liposome transfection reagent of Invitrogen. In the negative control group of the HDR system, pCAG-Cas9-mCherry together with the HDR donor vector were transfected into 293T cells. The 293T cell line was from the ATCC cell bank. After 24 hours of transfection, the mCherry-positive cells, i.e. the successfully transfected cells, were sorted by flow cytometry. The sorted cells were further cultured for 5 days, and then the ratio of EGFP-positive cells was analyzed by flow cytometry. Comparison of the ratio of EGFP-positive cells in the two different systems can reflect the difference in the efficiency of site-directed knock-in of the ACTB gene in different systems. The comparison results of the efficiency of inserting exogenous genes using different methods are shown in FIG. 8. Figure 9 C. The results show that in the case of using the EPTI system, the ratio of EGFP-positive cells is about 30%, which is significantly higher than the method using the HDR system.

[0697] The EGFP-positive cells after 5 days of culture were subjected to genomic DNA extraction, and then the nucleotide sequences at the junction of the reporter gene EGFP (5' end and 3' end) and the genomic DNA were identified by PCR using primers ACTB2-P1 (SEQ ID NO: 76) / ACTB2-P2 (SEQ ID NO: 77), ACTB2-P3 (SEQ ID NO: 78) / ACTB2-P4 (SEQ ID NO: 79), respectively.

[0698] The PCR identification results are shown in FIG. 9. Figure 9The results show that the primers pair ACTB2-P1 / ACTB2-P2, ACTB2-P3 / ACTB2-P4 used can amplify two fragments with expected size from the genomic DNA of EGFP positive cells. These results indicate that the EPTI system of the present application can site-specifically and accurately insert the exogenous gene (T2A-EGFP) into the specific position of ACTB gene of 293T cells. In addition, these results also indicate that the EPTI system of the present application can still greatly improve the site-specific integration efficiency of the exogenous gene into the target genome in the case that there is a gap between the genomic sequence targeted by the binding of the overhang and the site of the genome targeted for cleavage.

[0699] While the specific embodiments of the application have been described in detail, those skilled in the art will appreciate that various modifications and alterations can be made to the details of the application according to the teachings of all the teachings without departing from the scope of the application. The entire scope of the application is given by the following claims and any equivalents thereof. SEQUENCE LISTING <110> Institute of Zoology, Chinese Academy of Sciences, Beijing Stem Cell and Regenerative Medicine Research Institute <120> A system and method for editing nucleic acid <130> IDC210247 <150> 202010663076.8 <151> 10 July 2020 <160> 84 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> artificial <220> <223> Cas9 protein <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser lie Lys Lys Asn Leu lie 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg lie Cys 65 70 75 80 Tyr Leu Gin Glu lie Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro lie Phe Gly Asn lie Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr lie Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu lie Tyr Leu Ala Leu Ala His 145 150 155 160 Met lie Lys Phe Arg Gly His Phe Leu lie Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe lie Gin Leu Val Gin Thr Tyr 180 185 190 Asn Gin Leu Phe Glu Glu Asn Pro lie Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala lie Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu lie Ala Gin Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu lie Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gin Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gin lie Gly Asp Gin Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala lie Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu lie Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met lie Lys Arg Tyr Asp Glu His His Gin Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gin Gin Leu Pro Glu Lys Tyr Lys Glu lie Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val lie Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro lie Arg Glu Gin Ala Glu Asn 1295 1300 1305 lie lie His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr lie Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu lie His Gin Ser lie Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg lie Asp Leu Ser Gin Leu Gly Gly Asp 1355 1360 1365 <210> 2 <211> 236 <212> PRT <213> artificial <220> <223> mCherry <400> 2 Met Val Ser Lys Gly Glu Glu Asp Asn Met Ala lie lie Lys Glu Phe 1 5 10 15 Met Arg Phe Lys Val His Met Glu Gly Ser Val Asn Gly His Glu Phe 20 25 30 Glu lie Glu Gly Glu Gly Glu Gly Arg Pro Tyr Glu Gly Thr Gin Thr 35 40 45 Ala Lys Leu Lys Val Thr Lys Gly Gly Pro Leu Pro Phe Ala Trp Asp 50 55 60 lie Leu Ser Pro Gin Phe Met Tyr Gly Ser Lys Ala Tyr Val Lys His 65 70 75 80 Pro Ala Asp lie Pro Asp Tyr Leu Lys Leu Ser Phe Pro Glu Gly Phe 85 90 95 Lys Trp Glu Arg Val Met Asn Phe Glu Asp Gly Gly Val Val Thr Val 100 105 110 Thr Gin Asp Ser Ser Leu Gin Asp Gly Glu Phe lie Tyr Lys Val Lys 115 120 125 Leu Arg Gly Thr Asn Phe Pro Ser Asp Gly Pro Val Met Gin Lys Lys 130 135 140 Thr Met Gly Trp Glu Ala Ser Ser Glu Arg Met Tyr Pro Glu Asp Gly 145 150 155 160 Ala Leu Lys Gly Glu lie Lys Gin Arg Leu Lys Leu Lys Asp Gly Gly 165 170 175 His Tyr Asp Ala Glu Val Lys Thr Thr Tyr Lys Ala Lys Lys Pro Val 180 185 190 Gln Leu Pro Gly Ala Tyr Asn Val Asn Ile Lys Leu Asp Ile Thr Ser 195 200 205 His Asn Glu Asp Tyr Thr Ile Val Glu Gln Tyr Glu Arg Ala Glu Gly 210 215 220 Arg His Ser Thr Gly Gly Met Asp Glu Leu Tyr Lys 225 230 235 <210> 3 <211> 76 <212> DNA <213> artificial <220> <223> gRNA lacking a guide sequence <400> 3 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 60 ggcaccgagt cggtgc 76 <210> 4 <211> 722 <212> PRT <213> artificial <220> <223> MLV RT <400> 4 Met Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser 1 5 10 15 Gly Gly Ser Ser Gly Gly Ser Ser Thr Leu Asn Ile Glu Asp Glu Tyr 20 25 30 Arg Leu His Glu Thr Ser Lys Glu Pro Asp Val Ser Leu Gly Ser Thr 35 40 45 Trp Leu Ser Asp Phe Pro Gln Ala Trp Ala Glu Thr Gly Gly Met Gly 50 55 60 Leu Ala Val Arg Gln Ala Pro Leu Ile Ile Pro Leu Lys Ala Thr Ser 65 70 75 80 Thr Pro Val Ser Ile Lys Gln Tyr Pro Met Ser Gln Glu Ala Arg Leu 85 90 95 Gly Ile Lys Pro His Ile Gln Arg Leu Leu Asp Gln Gly Ile Leu Val 100 105 110 Pro Cys Gln Ser Pro Trp Asn Thr Pro Leu Leu Pro Val Lys Lys Pro 115 120 125 Gly Thr Asn Asp Tyr Arg Pro Val Gln Asp Leu Arg Glu Val Asn Lys 130 135 140 Arg Val Glu Asp Ile His Pro Thr Val Pro Asn Pro Tyr Asn Leu Leu 145 150 155 160 Ser Gly Leu Pro Pro Ser His Gln Trp Tyr Thr Val Leu Asp Leu Lys 165 170 175 Asp Ala Phe Phe Cys Leu Arg Leu His Pro Thr Ser Gln Pro Leu Phe 180 185 190 Ala Phe Glu Trp Arg Asp Pro Glu Met Gly Ile Ser Gly Gln Leu Thr 195 200 205 Trp Thr Arg Leu Pro Gin Gly Phe Lys Asn Ser Pro Thr Leu Phe Asn 210 215 220 Glu Ala Leu His Arg Asp Leu Ala Asp Phe Arg Ile Gin His Pro Asp 225 230 235 240 Leu Ile Leu Leu Gin Tyr Val Asp Asp Leu Leu Leu Ala Ala Thr Ser 245 250 255 Glu Leu Asp Cys Gin Gin Gly Thr Arg Ala Leu Leu Gin Thr Leu Gly 260 265 270 Asn Leu Gly Tyr Arg Ala Ser Ala Lys Lys Ala Gin Ile Cys Gin Lys 275 280 285 Gln Val Lys Tyr Leu Gly Tyr Leu Leu Lys Gin Gly Gin Arg Trp Leu 290 295 300 Thr Gin Ala Arg Lys Gin Thr Val Met Gly Gin Pro Thr Pro Lys Thr 305 310 315 320 Pro Arg Gin Leu Arg Glu Phe Leu Gly Lys Ala Gly Phe Cys Arg Leu 325 330 335 Phe Ile Pro Gly Phe Ala Gin Met Ala Ala Pro Leu Tyr Pro Leu Thr 340 345 350 Lys Pro Gly Thr Leu Phe Asn Trp Gly Pro Asp Gin Gin Lys Ala Tyr 355 360 365 Gln Glu Ile Lys Gln Ala Leu Leu Thr Ala Pro Ala Leu Gly Leu Pro 370 375 380 Asp Leu Thr Lys Pro Phe Glu Leu Phe Val Asp Glu Lys Gln Gly Tyr 385 390 395 400 Ala Lys Gly Val Leu Thr Gln Lys Leu Gly Pro Trp Arg Arg Pro Val 405 410 415 Ala Tyr Leu Ser Lys Lys Leu Asp Pro Val Ala Ala Gly Trp Pro Pro 420 425 430 Cys Leu Arg Met Val Ala Ala Ile Ala Val Leu Thr Lys Asp Ala Gly 435 440 445 Lys Leu Thr Met Gly Gln Pro Leu Val Ile Leu Ala Pro His Ala Val 450 455 460 Glu Ala Leu Val Lys Gln Pro Pro Asp Arg Trp Leu Ser Asn Ala Arg 465 470 475 480 Met Thr His Tyr Gln Ala Leu Leu Leu Asp Thr Asp Arg Val Gln Phe 485 490 495 Gly Pro Val Val Ala Leu Asn Pro Ala Thr Leu Leu Pro Leu Pro Glu 500 505 510 Glu Gly Leu Gln His Asn Cys Leu Asp Ile Leu Ala Glu Ala His Gly 515 520 525 Thr Arg Pro Asp Leu Thr Asp Gln Pro Leu Pro Asp Ala Asp His Thr 530 535 540 Trp Tyr Thr Asp Gly Ser Ser Leu Leu Gln Glu Gly Gln Arg Lys Ala 545 550 555 560 Gly Ala Ala Val Thr Thr Glu Thr Glu Val Ile Trp Ala Lys Ala Leu 565 570 575 Pro Ala Gly Thr Ser Ala Gln Arg Ala Glu Leu Ile Ala Leu Thr Gln 580 585 590 Ala Leu Lys Met Ala Glu Gly Lys Lys Leu Asn Val Tyr Thr Asp Ser 595 600 605 Arg Tyr Ala Phe Ala Thr Ala His Ile His Gly Glu Ile Tyr Arg Arg 610 615 620 Arg Gly Trp Leu Thr Ser Glu Gly Lys Glu Ile Lys Asn Lys Asp Glu 625 630 635 640 Ile Leu Ala Leu Leu Lys Ala Leu Phe Leu Pro Lys Arg Leu Ser Ile 645 650 655 Ile His Cys Pro Gly His Gln Lys Gly His Ser Ala Glu Ala Arg Gly 660 665 670 Asn Arg Met Ala Asp Gln Ala Ala Arg Lys Ala Ala Ile Thr Glu Thr 675 680 685 Pro Asp Thr Ser Thr Leu Leu Ile Glu Asn Ser Ser Pro Ser Gly Gly 690 695 700 Ser Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Pro Lys Lys Lys Arg 705 710 715 720 Lys Val <210> 5 <211> 24 <212> DNA <213> artificial <220> <223> Gapdh-gRNA-F <400> 5 ccggagagag agaccctcac tgct 24 <210> 6 <211> 24 <212> DNA <213> artificial <220> <223> Gapdh-gRNA-R <400> 6 aaacagcagt gagggtctct ctct 24 <210> 7 <211> 96 <212> DNA <213> artificial <220> <223> Gapdh-gRNA <400> 7 agagagagac cctcactgct gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 8 <211> 24 <212> DNA <213> artificial <220> <223> Actb‐gRNA‐F <400> 8 ccggatcccc caaagttcac aatg 24 <210> 9 <211> 24 <212> DNA <213> artificial <220> <223> Actb‐gRNA‐R <400> 9 aaaccattgt gaactttggg ggat 24 <210> 10 <211> 96 <212> DNA <213> artificial <220> <223> Actb‐gRNA <400> 10 atcccccaaa gttcacaatg gttttagagc tagaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 11 <211> 24 <212> DNA <213> artificial <220> <223> GapdhRC‐gRNA‐F <400> 11 ccggtagcgt tgacccgacc ccaa 24 <210> 12 <211> 24 <212> DNA <213> artificial <220> <223> GapdhRC-gRNA-R <400> 12 aaacttgggg tcgggtcaac gcta 24 <210> 13 <211> 96 <212> DNA <213> artificial <220> <223> GapdhRC-gRNA <400> 13 tagcgttgac ccgaccccaa gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 14 <211> 24 <212> DNA <213> artificial <220> <223> GapdhRT-gRNA-F <400> 14 ccgggtaagc acacgtgcaa agtg 24 <210> 15 <211> 24 <212> DNA <213> artificial <220> <223> GapdhRT-gRNA-R <400> 15 aaaccacttt gcacgtgtgc ttac 24 <210> 16 <211> 96 <212> DNA <213> artificial <220> <223> GapdhRT-gRNA <400> 16 gtaagcacac gtgcaaagtg gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 17 <211> 24 <212> DNA <213> artificial <220> <223> spacerA-gRNA-F <400> 17 ccgggagatc gagtgccgca tcac 24 <210> 18 <211> 24 <212> DNA <213> artificial <220> <223> spacerA-gRNA-R <400> 18 aaacgtgatg cggcactcga tctc 24 <210> 19 <211> 96 <212> DNA <213> artificial <220> <223> spacerA-gRNA <400> 19 gagatcgagt gccgcatcac gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 20 <211> 24 <212> DNA <213> artificial <220> <223> spacerK-gRNA-F <400> 20 ccgggtcgccc tcgaacttcacct 24 <210> 21 <211> 24 <212> DNA <213> artificial <220> <223> spacerK-gRNA-R <400> 21 aaacaggtga agttcgaggg cgac 24 <210> 22 <211> 96 <212> DNA <213> artificial <220> <223> spacerK-gRNA <400> 22 gtcgccctcg aacttcacct gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 23 <211> 65 <212> DNA <213> artificial <220> <223> Gapdh-pegA-F <400> 23 aaagtggcac cgagtcggtg cagcaagagc acaagaggaa gagagagacc ctcactatgc 60 ggcac 65 <210> 24 <211> 62 <212> DNA <213> artificial <220> <223> Gapdh-pegA-R <400> 24 acagctatga ccatgattac gccaagctta aaaaaaatcg agtgccgcat agtgagggtc 60 tc 62 <210> 25 <211> 144 <212> DNA <213> artificial <220> <223> Gapdh-pegA <400> 25 gagatcgagt gccgcatcac gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcagca agagcacaag aggaagagag 120 agaccctcac tatgcggcac tcga 144 <210> 26 <211> 56 <212> DNA <213> artificial <220> <223> Gapdh-pegK-F <400> 26 aaagtggcac cgagtcggtg ctggtggggg actgagtgtg gcagggactc cccagc 56 <210> 27 <211> 60 <212> DNA <213> artificial <220> <223> Gapdh-pegK-R <400> 27 gaccatgatt acgccaagct taaaaaaaac cctcgaactt cagctgggga gtccctgcca 60 <210> 28 <211> 144 <212> DNA <213> artificial <220> <223> Gapdh-pegK <400> 28 gtcgccctcg aacttcacct gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctggt gggggactga gtgtggcagg 120 gactccccag ctgaagttcg aggg 144 <210> 29 <211> 62 <212> DNA <213> artificial <220> <223> Actb-pegA-F <400> 29 aaagtggcac cgagtcggtg ccagtcggtt ggagcgagca tcccccaaag ttcacaatgc 60 gg 62 <210> 30 <211> 66 <212> DNA <213> artificial <220> <223> Actb-pegA-R <400> 30 acagctatga ccatgattac gccaagctta aaaaaaatcg agtgccgcat tgtgaacttt 60 ggggga 66 <210> 31 <211> 144 <212> DNA <213> artificial <220> <223> Actb-pegA <400> 31 gagatcgagt gccgcatcac gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgccagt cggttggagc gagcatcccc 120 caaagttcac aatgcggcac tcga 144 <210> 32 <211> 66 <212> DNA <213> artificial <220> <223> Actb-pegK-F <400> 32 aaagtggcac cgagtcggtg caaacaacaa tgtgcaatca aagtcctcgg ccacattgaa 60 gttcga 66 <210> 33 <211> 63 <212> DNA <213> artificial <220> <223> Actb-pegK-R <400> 33 acagctatga ccatgattac gccaagctta aaaaaaaccc tcgaacttca atgtggccga 60 gga 63 <210> 34 <211> 144 <212> DNA <213> artificial <220> <223> Actb-pegK <400> 34 gtcgccctcg aacttcacct gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcaaac aacaatgtgc aatcaaagtc 120 ctcggccaca ttgaagttcg aggg 144 <210> 35 <211> 60 <212> DNA <213> artificial <220> <223> GapdhRC-pegRNA-F <400> 35 aagtggcacc gagtcggtgc tttacagcct ggcctttgga gatcgagtgc cgcatgggtc 60 <210> 36 <211> 65 <212> DNA <213> artificial <220> <223> GapdhRC-pegRNA-R <400> 36 aacagctatg accatgatta cgccaagctt aaaaaaaatt gacccgaccc atgcggcact 60 cgatc 65 <210> 37 <211> 143 <212> DNA <213> artificial <220> <223> GapdhRC-pegRNA <400> 37 tagcgttgac ccgaccccaa gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcttta cagcctggcc tttggagatc 120 gagtgccgca tgggtcgggt caa 143 <210> 38 <211> 61 <212> DNA <213> artificial <220> <223> GapdhRT‐pegRNA‐F <400> 38 aagtggcacc gagtcggtgc gttactcccg ggcctcacgt cgccctcgaa cttcatttgc 60 to 61 <210> 39 <211> 68 <212> DNA <213> artificial <220> <223> GapdhRT‐pegRNA‐R <400> 39 aacagctatg accatgatta cgccaagctt aaaaaaagc acacgtgcaa atgaagttcg 60 agggcgac 68 <210> 40 <211> 144 <212> DNA <213> artificial <220> <223> GapdhRT‐pegRNA <400> 40 gtaagcacac gtgcaaagtg gttttagagc tagaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgtta ctcccgggcc tcacgtcgcc 120 ctcgaacttc atttgcacgt gtgc 144 <210> 41 <211> 59 <212> DNA <213> artificial <220> <223> GapdhRC‐pegA‐F <400> 41 aagtggcacc gagtcggtgc gcccttcccc tgccagccta gcgttgaccc gacccatgc 59 <210> 42 <211> 67 <212> DNA <213> artificial <220> <223> GapdhRC‐pegA‐R <400> 42 acagctatga ccatgattac gccaagctta aaaaaaatcg agtgccgcat gggtcgggtc 60 aacgcta 67 <210> 43 <211> 144 <212> DNA <213> artificial <220> <223> GapdhRC‐pegA <400> 43 gagatcgagt gccgcatcac gttttagagc tagaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgccc ttcccctgcc agcctagcgt 120 tgacccgacc catgcggcac tcga 144 <210> 44 <211> 60 <212> DNA <213> artificial <220> <223> GapdhRT-pegK-F <400> 44 aagtggcacc gagtcggtgc tacttttgtc tccactaggt aagcacacgt gcaaatgaag 60 <210> 45 <211> 71 <212> DNA <213> artificial <220> <223> GapdhRT-pegK-R <400> 45 aacagctatg accatgatta cgccaagctt aaaaaaaacc ctcgaacttc atttgcacgt 60 gtgcttacct a 71 <210> 46 <211> 144 <212> DNA <213> artificial <220> <223> GapdhRT-pegK <400> 46 gtcgccctcg aacttcacct gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctact tttgtctcca ctaggtaagc 120 acacgtgcaa atgaagttcg aggg 144 <210> 47 <211> 1447 <212> DNA <213> artificial <220> <223> IRES-EGFP <400> 47 ccggtgatgc ggcactcgat ctcgaattcc ccctctccct cccccccccc taacgttact 60 ggccgaagcc gcttggaata aggccggtgt gcgtttgtct atatgttatt ttccaccata 120 ttgccgtctt ttggcaatgt gagggcccgg aaacctggcc ctgtcttctt gacgagcatt 180 cctaggggtc tttcccctct cgccaaagga atgcaaggtc tgttgaatgt cgtgaaggaa 240 gcagttcctc tggaagcttc ttgaagacaa acaacgtctg tagcgaccct ttgcaggcag 300 cggaaccccc cacctggcga caggtgcctc tgcggccaaa agccacgtgt ataagataca 360 cctgcaaagg cggcacaacc ccagtgccac gttgtgagtt ggatagttgt ggaaagagtc 420 aaatggctct cctcaagcgt attcaacaag gggctgaagg atgcccagaa ggtaccccat 480 tgtatgggat ctgatctggg gcctcggtgc acatgcttta catgtgttta gtcgaggtta 540 aaaaaacgtc taggcccccc gaaccacggg gacgtggttt tcctttgaaa aacacgatga 600 taatatggcc acaacgctcg gtttaaaaag cttctatgcc tgaataggtg accggaggtc 660 ggcacctttc ctttgcaatt actgacccta tgaatacagg atctatggtg agcaagggcg 720 aggagctgtt caccggggtg gtgcccatcc tggtcgagct ggacggcgac gtaaacggcc 780 acaagttcag cgtgtccggc gagggcgagg gcgatgccac ctacggcaag ctgaccctga 840 agttcatctg cactacgggg aaactgcccg tgccctggcc caccctcgtg accaccctga 900 cctacggcgt gcagtgcttc agccgctacc ccgaccacat gaagcagcac gacttcttca 960 agtccgccat gcccgaaggc tacgtccagg agcgcaccat cttcttcaag gacgacggca 1020 actacaagac ccgcgctgaa gtcaaattcg agggcgacac cctggtgaac cgcatcgagc 1080 tgaagggcat cgacttcaag gaggacggca acatcctggg gcacaagctg gagtacaact 1140 acaacagcca caacgtctat atcatggccg acaagcagaa gaacggcatc aaggtgaact 1200 tcaagatccg ccacaacatc gaggacggca gcgtgcagct cgccgaccac taccagcaga 1260 acacccccat cggcgacggc cccgtgctgc tgcccgacaa ccactacctg agcacccagt 1320 ccgccctgag caaagacccc aacgagaagc gcgatcacat ggtcctgctg gagttcgtga 1380 ccgccgccgg gatcactctc ggcatggacg agctgtacaa gtaagtcgcc ctcgaacttc 1440 acctcgg 1447 <210> 48 <211> 4259 <212> DNA <213> artificial <220> <223> Gene fragment containing T2A-EGFP <400> 48 ccggtgatgc ggcactcgat ctccaaaggc caggctgtaa atgtcaccgg gaggattggg 60 tgtctgggcg cctcggggaa cctgcccttc tccccattcc gtcttccgga aaccagatct 120 cccaccgcac cctggtctga ggttaaatat agctgctgac ctttctgtag ctgggggcct 180 gggctggggc tctctcccat cccttctccc cacacacatg cacttacctg tgctcccact 240 cctgatttct ggaaaagagc taggaaggac aggcaacttg gcaaatcaaa gccctgggac 300 tagggggtta aaatacagct tcccctcttc ccacccgccc cagtctctgt cccttttgta 360 ggagggactt agagaagggg tgggcttgcc ctgtccagtt aatttctgac ctttactcct 420 gccctttgag tttgatgatg ctgagtgtac aagcgttttc tccctaaagg gtgcagctga 480 gctaggcagc agcaagcatt cctggggtgg catagtgggg tggtgaatac catgtacaaa 540 gcttgtgccc agactgtggg tggcagtgcc ccacatggcc gcttctcctg gaagggcttc 600 gtatgactgg gggtgttggg cagccctgga gccttcagtt gcagccatgc cttaagccag 660 gccagcctgg cagggaagct caagggagat aaaattcaac ctcttgggcc ctcctggggg 720 taaggagatg ctgcattcgc cctcttaatg gggaggtggc ctagggctgc tcacatattc 780 tggaggagcc tcccctcctc atgccttctt gcctcttgtc tcttagattt ggtcgtattg 840 ggcgcctggt caccagggct gcttttaact ctggtaaagt ggatattgtt gccatcaatg 900 accccttcat tgacctcaac tacatggtga gtgctacatg gtgagcccca aagctggtgt 960 gggaggagcc acctggctga tgggcagccc cttcataccc tcacgtattc ccccaggttt 1020 acatgttcca atatgattcc acccatggca aattccatgg caccgtcaag gctgagaacg 1080 ggaagcttgt catcaatgga aatcccatca ccatcttcca ggagtgagtg gaagacagaa 1140 tggaagaaat gtgctttggg gaggcaacta ggatggtgtg gctcccttgg gtatatggta 1200 accttgtgtc cctcaatatg gtcctgtccc catctccccc ccacccccat aggcgagatc 1260 cctccaaaat caagtggggc gatgctggcg ctgagtacgt cgtggagtcc actggcgtct 1320 TCTGGAGCCTGGAGTGGGAGCCCTGGGGAGGCCTGGAGCCTGGAGCCTGGA 60 CAGCCCTGCA AAGGCAGGAC CCAGGTTCAT AACTGTCTGC TTCTCTGCTG 1380 TGCAGGGGGG AGCCAAAAGG GTCATCATCT CTGCCCCCTC TGCTGATGCC CCCATGTTCG 1740 TTCATGGGTG TGAACCATGA GAAGTATGAC AACAGCCTCA AGATCATCAG GTGAGGAAGG 1800 AGGGCCCgtGGAGAAGCGCCAGCCTGGCACCCCTATGGACACGCTCCCCCG ACTTGCGCC 1860 CCGCTCCCTC TTTCTTTGCA GCAATGCCTC CTGCACCACC AACTGCTTAG CACCCCTGGC 1920 CAAGGTcatc CATGACAAC TTTGGTATCG TGGAAggACTCATGGTAT GAGAGCTGGG 1980 TGGGACTGAG GCTCCCACCT TTCTCATCCA AGACTGGCTC CTCCCTGCCG GGCTGCgtG 2040 CAACCCTGGG GTTGGGGGTT CTGGGGACTG GCTTTCCCAT AATTTCTTTT C AAGGTGGGG 2100 AGGGAGGTAG AGGGGTGATG TGGGGAGTAC GCTGCAGGGC CTCACTCCTT TTGCAGACCA 2160 CAGTCCATGC ATCCTGCCAC CCCAGAAGAC TGTGGATGGC CCCTCCGGGA AACTGTGGC 2220 GTGATGGCCG CggGGCTCTC CAGAACATCA TCCCTGCCTC TACTGGCGCT GCCAAGGCTG 2280 tgggcaaggt catccctgag ctgaacggga agctcactgg catggccttc cgtgtcccca 2100 ctgccaacgt gtcagtggtg gacctgacct gccgtctaga aaaacctgcc aaatatgatg 2160 acatcaagaa ggtggtgaag caggcgtcgg agggccccct caagggcatc ctgggctaca 2220 ctgagcacca ggtggtctcc tctgacttca acagcgacac ccactcctcc acctttgacg 2280 ctggggctgg cattgccctc aacgaccact ttgtcaagct catttcctgg tatgtggctg 2340 gggccagaga ctggctctta aaaagtgcag ggtctggcgc cctctggtgg ctggctcaga 2400 aaaagggccc tgacaactct tttcatcttc taggtatgac aacgaatttg gctacagcaa 2460 cagggtggtg gacctcatgg cccacatggc ctccaaggag gagggcagag gaagtcttct 2520 aacatgcggt gacgtggagg agaatcccgg cccaatggtg agcaagggcg aggagctgtt 2580 caccggggtg gtgcccatcc tggtcgaact cgatggagat gtgaacggcc acaagttcag 2640 cgtgtccggc gagggcgagg gcgatgccac ctacggcaag ctgaccctga agttcatctg 2700 cactacgggg aaactgcccg tgccctggcc caccctcgtg accaccctga cctacggcgt 2760 gcagtgcttc agccgctacc ccgaccacat gaagcagcac gacttcttca agtccgccat 2820 gccagaggga tatgtgcaag agcgcaccat cttcttcaag gacgacggca actacaagac 2880 ccgcgctgaa gtcaaattcg agggcgacac cctggtgaac cgcatcgagc tgaagggcat 2940 cgacttcaag gaggacggca acatcctggg gcacaagctg gagtacaact acaacagcca 3000 caacgtctat atcatggccg acaagcagaa gaacggcatc aaggtgaact tcaagatccg 3060 ccacaacatc gaggacggca gcgtgcagct cgccgaccac taccagcaga acacccccat 3120 cggcgacggc cccgtgctgc tgcccgacaa ccactacctg agcacccagt ccgccctgag 3180 caaagacccc aacgagaagc gcgatcacat ggtcctgctg gagttcgtga ccgccgccgg 3240 gatcactctc ggcatggacg agctgtacaa gtaagacccc tggaccacca gccccagcaa 3300 gagcacaaga ggaagagaga gaccctcact gctggggagt ccctgccaca ctcagtcccc 3360 caccacactg aatctcccct cctcacagtt gccatgtaga ccccttgaag aggggagggg 3420 cctagggagc cgcaccttgt catgtaccat caataaagta ccctgtgctc aaccagttac 3480 ttgtcctgtc ttattctagg gtctggggca gaggggaggg aagctgggct tgtgtcaagg 3540 tgagacattc ttgctgggga gggacctggt atgttctcct cagactgagg gtagggcctc 3600 caaacagcct tgcttgcttc gagaaccatt tgcttcccgc tcagacgtct tgagtgctac 3660 aggaagctgg caccactact tcagagaaca aggccttttc ctctcctcgc tccagtccta 3720 ggctatctgc tgttggccaa acatggaaga agctattctg tgggcagccc cagggaggct 3780 gacaggtgga ggaagtcagg gctcgcactg ggctctgacg ctgactggtt agtggagctc 3840 agcctggagc tgagctgcag cgggcaattc cagcttggcc tccgcagctg tgaggtcttg 3900 agcacgtgct ctattgcttt ctgtgccctc gtgtcttatc tgaggacatc gtggccagcc 3960 cctaaggtct tcaagcagga ttcatctagg taaaccaagt acctaaaacc atgcccaagg 4020 cggtaaggac tatataatgt ttaaaaatcg gtaaaaatgc ccacctcgca tagttttgag 4080 gaagatgaac tgagatgtgt cagggtgact tatttccatc atcgtcctta ggggaacttg 4140 ggtaggggca aggcgtgtag ctgggaccta ggtccagacc cctggctctg ccactgaacg 4200 gctcagttgc tttgggcagt tactcccggg cctcacgtcg ccctcgaact tcacctcgg 4259 <210> 49 <211> 23 <212> DNA <213> artificial <220> <223> SpacerA <400> 49 gagatcgagt gccgcatcac cgg 23 <210> 50 <211> 23 <212> DNA <213> artificial <220> <223> SpacerK <400> 50 gtcgccctcg aacttcacct cgg 23 <210> 51 <211> 10 <212> PRT <213> artificial <220> <223> linker <400> 51 Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly 1 5 10 <210> 52 <211> 2099 <212> PRT <213> artificial <220> <223> fusion protein <400> 52 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val lie Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser lie Lys Lys Asn Leu lie 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg lie Cys 65 70 75 80 Tyr Leu Gin Glu lie Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro lie Phe Gly Asn lie Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr lie Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu lie Tyr Leu Ala Leu Ala His 145 150 155 160 Met lie Lys Phe Arg Gly His Phe Leu lie Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gin Leu Val Gin Thr Tyr 180 185 190 Asn Gin Leu Phe Gin Gin Asn Pro Ile Asn Ala Ser Gin Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Gin Asn 210 215 220 Leu Ile Ala Gin Leu Pro Gin Gin Lys Lys Asn Gin Leu Phe Gin Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gin Leu Thr Pro Asn Phe Gin Ser Asn Phe 245 250 255 Asp Leu Ala Gin Asp Ala Lys Leu Gin Leu Ser Lys Asp Thr Tyr Gin 260 265 270 Asp Asp Leu Gin Asn Leu Leu Ala Gin Ile Gin Asp Gin Tyr Gin Gin 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Gin Ala Ile Leu Leu Ser Gin 290 295 300 Ile Leu Arg Val Asn Thr Gin Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Gin His His Gin Gin Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gin Gin Leu Pro Glu Lys Tyr Lys Glu He Phe Phe 340 345 350 Asp Gin Ser Lys Asn Gly Tyr Ala Gly Tyr He Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe He Lys Pro He Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gin Arg Thr Phe Asp Asn Gly Ser He Pro His Gin He His Leu 405 410 415 Gly Glu Leu His Ala He Leu Arg Arg Gin Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys He Glu Lys He Leu Thr Phe Arg He 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr He Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gin Ser Phe He Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gin Leu Gin Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gin Glu Leu Asp He Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His He Val Pro Gin Ser Phe Leu Lys 835 840 845 Asp Asp Ser He Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gin Leu Leu Asn Ala Lys Leu He Thr Gin Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe He Lys Arg Gin Leu Val Glu Thr Arg Gin He Thr 915 920 925 Lys His Val Ala Gin He Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu He Arg Glu Val Lys Val He Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gin Phe Tyr Lys Val Arg 965 970 975 Glu lie Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu lie Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met lie Ala 1010 1015 1020 Lys Ser Glu Gin Glu lie Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn lie Met Asn Phe Phe Lys Thr Glu lie Thr Leu Ala 1040 1045 1050 Asn Gly Glu lie Arg Lys Arg Pro Leu lie Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu lie Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gin Val Asn lie Val Lys Lys Thr 1085 1090 1095 Glu Val Gin Thr Gly Gly Phe Ser Lys Glu Ser lie Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gin Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gin Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gin Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gin Leu Gly Gly Asp 1355 1360 1365 Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly 1370 1375 1380 Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly 1385 1390 1395 Gly Ser Ser Thr Leu Asn Ile Glu Asp Glu Tyr Arg Leu His Glu 1400 1405 1410 Thr Ser Lys Glu Pro Asp Val Ser Leu Gly Ser Thr Trp Leu Ser 1415 1420 1425 Asp Phe Pro Gln Ala Trp Ala Glu Thr Gly Gly Met Gly Leu Ala 1430 1435 1440 Val Arg Gln Ala Pro Leu Ile Ile Pro Leu Lys Ala Thr Ser Thr 1445 1450 1455 Pro Val Ser Ile Lys Gln Tyr Pro Met Ser Gln Glu Ala Arg Leu 1460 1465 1470 Gly Ile Lys Pro His Ile Gln Arg Leu Leu Asp Gln Gly Ile Leu 1475 1480 1485 Val Pro Cys Gln Ser Pro Trp Asn Thr Pro Leu Leu Pro Val Lys 1490 1495 1500 Lys Pro Gly Thr Asn Asp Tyr Arg Pro Val Gln Asp Leu Arg Glu 1505 1510 1515 Val Asn Lys Arg Val Glu Asp Ile His Pro Thr Val Pro Asn Pro 1520 1525 1530 Tyr Asn Leu Leu Ser Gly Leu Pro Pro Ser His Gln Trp Tyr Thr 1535 1540 1545 Val Leu Asp Leu Lys Asp Ala Phe Phe Cys Leu Arg Leu His Pro 1550 1555 1560 ​Thr Ser Gin Pro Leu Phe Ala Phe Glu Trp Arg Asp Pro Glu Met 1565 1570 1575 Gly lie Ser Gly Gin Leu Thr Trp Thr Arg Leu Pro Gin Gly Phe 1580 1585 1590 Lys Asn Ser Pro Thr Leu Phe Asn Glu Ala Leu His Arg Asp Leu 1595 1600 1605 Ala Asp Phe Arg lie Gin His Pro Asp Leu lie Leu Leu Gin Tyr 1610 1615 1620 Val Asp Asp Leu Leu Leu Ala Ala Thr Ser Glu Leu Asp Cys Gin 1625 1630 1635 Gln Gly Thr Arg Ala Leu Leu Gin Thr Leu Gly Asn Leu Gly Tyr 1640 1645 1650 Arg Ala Ser Ala Lys Lys Ala Gin lie Cys Gin Lys Gin Val Lys 1655 1660 1665 Tyr Leu Gly Tyr Leu Leu Lys Glu Gly Gin Arg Trp Leu Thr Glu 1670 1675 1680 Ala Arg Lys Glu Thr Val Met Gly Gin Pro Thr Pro Lys Thr Pro 1685 1690 1695 Arg Gin Leu Arg Glu Phe Leu Gly Lys Ala Gly Phe Cys Arg Leu 1700 1705 1710 Phe lie Pro Gly Phe Ala Glu Met Ala Ala Pro Leu Tyr Pro Leu 1715 1720 1725 Thr Lys Pro Gly Thr Leu Phe Asn Trp Gly Pro Asp Gln Gln Lys 1730 1735 1740 Ala Tyr Gin Glu lie Lys Gin Ala Leu Leu Thr Ala Pro Ala Leu 1745 1750 1755 Gly Leu Pro Asp Leu Thr Lys Pro Phe Glu Leu Phe Val Asp Glu 1760 1765 1770 Lys Gin Gly Tyr Ala Lys Gly Val Leu Thr Gin Lys Leu Gly Pro 1775 1780 1785 Trp Arg Arg Pro Val Ala Tyr Leu Ser Lys Lys Leu Asp Pro Val 1790 1795 1800 Ala Ala Gly Trp Pro Pro Cys Leu Arg Met Val Ala Ala lie Ala 1805 1810 1815 Val Leu Thr Lys Asp Ala Gly Lys Leu Thr Met Gly Gin Pro Leu 1820 1825 1830 Val lie Leu Ala Pro His Ala Val Glu Ala Leu Val Lys Gin Pro 1835 1840 1845 Pro Asp Arg Trp Leu Ser Asn Ala Arg Met Thr His Tyr Gin Ala 1850 1855 1860 Leu Leu Leu Asp Thr Asp Arg Val Gin Phe Gly Pro Val Val Ala 1865 1870 1875 Leu Asn Pro Ala Thr Leu Leu Pro Leu Pro Gin Gin Gly Leu Gin 1880 1885 1890 His Asn Cys Leu Asp He Leu Ala Glu Ala His Gly Thr Arg Pro 1895 1900 1905 Asp Leu Thr Asp Gin Pro Leu Pro Asp Ala Asp His Thr Trp Tyr 1910 1915 1920 Thr Asp Gly Ser Ser Leu Leu Gin Glu Gly Gin Arg Lys Ala Gly 1925 1930 1935 Ala Ala Val Thr Thr Glu Thr Glu Val He Trp Ala Lys Ala Leu 1940 1945 1950 Pro Ala Gly Thr Ser Ala Gin Arg Ala Glu Leu He Ala Leu Thr 1955 1960 1965 Gln Ala Leu Lys Met Ala Glu Gly Lys Lys Leu Asn Val Tyr Thr 1970 1975 1980 Asp Ser Arg Tyr Ala Phe Ala Thr Ala His He His Gly Glu He 1985 1990 1995 Tyr Arg Arg Arg Gly Trp Leu Thr Ser Glu Gly Lys Glu He Lys 2000 2005 2010 Asn Lys Asp Glu lie Leu Ala Leu Leu Lys Ala Leu Phe Leu Pro 2015 2020 2025 Lys Arg Leu Ser lie lie His Cys Pro Gly His Gin Lys Gly His 2030 2035 2040 Ser Ala Glu Ala Arg Gly Asn Arg Met Ala Asp Gin Ala Ala Arg 2045 2050 2055 Lys Ala Ala lie Thr Glu Thr Pro Asp Thr Ser Thr Leu Leu lie 2060 2065 2070 Glu Asn Ser Ser Pro Ser Gly Gly Ser Lys Arg Thr Ala Asp Gly 2075 2080 2085 Ser Glu Phe Glu Pro Lys Lys Lys Arg Lys Val 2090 2095 <210> 53 <211> 24 <212> DNA <213> artificial <220> <223> sgACTB2-F <400> 53 ccggccaccg caaatgcttc tagg 24 <210> 54 <211> 24 <212> DNA <213> artificial <220> <223> sgACTB2-R <400> 54 aaaccctaga agcatttgcg gtgg 24 <210> 55 <211> 96 <212> DNA <213> artificial <220> <223> sgACTB2 <400> 55 ccaccgcaaa tgcttctagg gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 56 <211> 24 <212> DNA <213> artificial <220> <223> ACTB2-sgL-F <400> 56 ccgggagctg gacggcgacg taaa 24 <210> 57 <211> 24 <212> DNA <213> artificial <220> <223> ACTB2-sgL-R <400> 57 aaactttacg tcgccgtcca gctc 24 <210> 58 <211> 96 <212> DNA <213> artificial <220> <223> ACTB2-sgL <400> 58 gagctggacg gcgacgtaaa gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 59 <211> 24 <212> DNA <213> artificial <220> <223> ACTB2-sgR-F <400> 59 ccggcatgcc cgaaggctac gtcc 24 <210> 60 <211> 24 <212> DNA <213> artificial <220> <223> ACTB2-sgR-R <400> 60 aaacggacgt agccttcggg catg 24 <210> 61 <211> 96 <212> DNA <213> artificial <220> <223> ACTB2-sgR <400> 61 catgcccgaa ggctacgtcc gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgc 96 <210> 62 <211> 60 <212> DNA <213> artificial <220> <223> ACTB2-pegL-F <400> 62 aaagtggcac cgagtcggtg cggcccctcc atcgtccacc gcaaatgctt cacgtcgccg 60 <210> 63 <211> 60 <212> DNA <213> artificial <220> <223> ACTB2-pegL-R <400> 63 aacagctatg accatgatta cgccaagctt aaaaaaaaga cggcgacgtg aagcatttgc 60 <210> 64 <211> 137 <212> DNA <213> artificial <220> <223> ACTB2-pegL <400> 64 gagctggacg gcgacgtaaa gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcggcc cctccatcgt ccaccgcaaa 120 tgcttcacgt cgccgtc 137 <210> 65 <211> 58 <212> DNA <213> artificial <220> <223> ACTB2-pegR-F <400> 65 aaagtggcac cgagtcggtg cggtgtaacg caactaagtc atagtccgcc tcgtagcc 58 <210> 66 <211> 63 <212> DNA <213> artificial <220> <223> ACTB2-pegR-R <400> 66 aacagctatg accatgatta cgccaagctt aaaaaaaacg aaggctacga ggcggactat 60 gac 63 <210> 67 <211> 137 <212> DNA <213> artificial <220> <223> ACTB2-pegR <400> 67 catgcccgaa ggctacgtcc gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcggtg taacgcaact aagtcatagt 120 ccgcctcgta gccttcg 137 <210> 68 <211> 18 <212> DNA <213> artificial <220> <223> GAPDH-P1 <400> 68 aaaagtgcag ggtctggc 18 <210> 69 <211> 20 <212> DNA <213> artificial <220> <223> GAPDH-P2 <400> 69 acaccggcct tattccaagc 20 <210> 70 <211> 22 <212> DNA <213> artificial <220> <223> GAPDH-P3 <400> 70 ccgaccacta ccagcagaac ac 22 <210> 71 <211> 25 <212> DNA <213> artificial <220> <223> GAPDH-P4 <400> 71 ccagacccta gaataagaca ggaca 25 <210> 72 <211> 22 <212> DNA <213> artificial <220> <223> ACTB-P1 <400> 72 gggagctgtc acatccaggg tc 22 <210> 73 <211> 22 <212> DNA <213> artificial <220> <223> ACTB-P2 <400> 73 aagacggcaa tatggtggaa aa 22 <210> 74 <211> 20 <212> DNA <213> artificial <220> <223> ACTB-P3 <400> 74 ctgcccgaca accactacct 20 <210> 75 <211> 22 <212> DNA <213> artificial <220> <223> ACTB-P4 <400> 75 CTAAGGCTGC TCAATGTCAA GG 22 <210> 76 <211> 22 <212> DNA <213> artificial <220> <223> ACTB2-P1 <400> 76 GGGAGCTGTC ACATCCAGGG TC 22 <210> 77 <211> 22 <212> DNA <213> artificial <220> <223> ACTB2-P2 <400> 77 CATCTCCATC GAGTTCGACC AG 22 <210> 78 <211> 20 <212> DNA <213> artificial <220> <223> ACTB2-P3 <400> 78 CTGCCCGACA ACCACTACCT 20 <210> 79 <211> 22 <212> DNA <213> artificial <220> <223> ACTB2-P4 <400> 79 CTAAGGCTGC TCAATGTCAA GG 22 <210> 80 <211> 23 <212> DNA <213> artificial <220> <223> spacer L <400> 80 gagctggacg gcgacgtaaa cgg 23 <210> 81 <211> 23 <212> DNA <213> artificial <220> <223> spacer R <400> 81 catgcccgaa ggctacgtcc agg 23 <210> 82 <211> 771 <212> DNA <213> artificial <220> <223> T2A‑EGFP <400> 82 gagggcagag gaagtcttct aacatgcggt gacgtggagg agaatcccgg cccagtgagc 60 aagggcgagg agctgttcac cggggtggtg cccatcctgg tcgaactcga tggagatgtg 120 aacggccaca agttcagcgt gtccggcgag ggcgagggcg atgccaccta cggcaagctg 180 accctgaagt tcatctgcac tacggggaaa ctgcccgtgc cctggcccac cctcgtgacc 240 accctgacct acggcgtgca gtgcttcagc cgctaccccg accacatgaa gcagcacgac 300 ttcttcaagt ccgccatgcc agagggatat gtgcaagagc gcaccatctt cttcaaggac 360 gacggcaact acaagacccg cgctgaagtc aaattcgagg gcgacaccct ggtgaaccgc 420 GAGAAGAAGA AGAAGAAGAA GAAGAAGAAG AAGAAGAAGA AGAAG 48 TACAACAGCC ACAACGTCTA TATCATGGCC GACAAGCAGA AGAECGGCAT CAAG 540 GTGAAC TTC A AGATCCGCCACAACATCGAG GACGGCAGCGTGCAGCTCGCC GACCAC TAC 600 CAGCAGAACACCCCCATCGGCACGGCCCCGTGCTGCTGCCGACAACCACTACCTGAGC 660 ACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGA G 720 TTCGTGACCGCCGCCGGGATC ACTCTCGGC ATGGACGAGCTGTACAAGTA A 771 <210> 83 <211> 800 <212> DNA <213> artificial <220> <223> ACTB2 LHA <400> 83 TGGACCTGGC TGGCCGGGAC CTGACTGACT ACCTCATGAA GATCCTCACC GAGCGC GGCT 60 ACAGCTTCAC CACCGGCCGA GC GGGAAATCGTGC GTGACATTAAGGAGAAGCTGTGCT 120 ACGTCGCCCT GGACTTCGAG CAAGAGATGG CCACGGCTGC TTCCAGCTCC TCCCTGGAGA 180 AGAGCTACGA GCTGCCTGAC GGCCAGGTCA TCACCATTGG CAATGAGCGG TTCCGCTGCC 240 AGAGCTACGA GCTGCCTGAC GGCCAGGTCA TCACCATTGG CAATGAGCGG TTCCGCTGCC 240ctgaggcact cttccagcct tccttcctgg gtgagtggag actgtctccc ggctctgcct 300 gacatgaggg ttacccctcg gggctgtgct gtggaagcta agtcctgccc tcatttccct 360 ctcaggcatg gagtcctgtg gcatccacga aactaccttc aactccatca tgaagtgtga 420 cgtggacatc cgcaaagacc tgtacgccaa cacagtgctg tctggcggca ccaccatgta 480 ccctggcatt gccgacagga tgcagaagga gatcactgcc ctggcaccca gcacaatgaa 540 gatcaaggtg ggtgtctttc ctgcctgagc tgacctgggc aggtcggctg tggggtcctg 600 tggtgtgtgg ggagctgtca catccagggt cctcactgcc tgtccccttc cctcctcaga 660 tcattgctcc tcctgagcgc aagtactccg tgtggatcgg cggctccatc ctggcctcgc 720 tgtccacctt ccagcagatg tggatcagca agcaggagta tgacgagtcc ggcccctcca 780 tcgtccaccg caaatgcttc 800 <210> 84 <211> 800 <212> DNA <213> artificial <220> <223> ACTB2 RHA <400> 84 aggcggacta tgacttagtt gcgttacacc ctttcttgac aaaacctaac ttgcgcagaa 60 aacaagatga gattggcatg gctttatttg ttttttttgt tttgttttgg tttttttttt 120 ttttttggct tgactcagga tttaaaaact ggaacggtga aggtgacagc agtcggttgg 180 agcgagcatc ccccaaagtt cacaatgtgg ccgaggactt tgattgcaca ttgttgtttt 240 tttaatagtc attccaaata tgagatgcgt tgttacagga agtcccttgc catcctaaaa 300 gccaccccac ttctctctaa ggagaatggc ccagtcctct cccaagtcca cacaggggag 360 gtgatagcat tgctttcgtg taaattatgt aatgcaaaat ttttttaatc ttcgccttaa 420 tactttttta ttttgtttta ttttgaatga tgagccttcg tgccccccct tccccctttt 480 ttgtccccca acttgagatg tatgaaggct tttggtctcc ctgggagtgg gtggaggcag 540 ccagggctta cctgtacact gacttgagac cagttgaata aaagtgcaca ccttaaaaat 600 gaggccaagt gtgactttgt ggtgtggctg ggttgggggc agcagagggt gaaccctgca 660 ggagggtgaa ccctgcaaaa gggtggggca gtgggggcca acttgtcctt acccagagtg 720 caggtgtgtg gagatccctc ctgccttgac attgagcagc cttagagggt gggggaggct 780 caggggtcag gtctctgttc 800

Claims

1. A system or kit for editing a nucleic acid, comprising: (1) a first Cas protein or a nucleic acid molecule A1 containing a nucleotide sequence encoding said first Cas protein, wherein the first Cas protein is capable of cleaving or breaking a first double-stranded target nucleic acid; (2) a first template-dependent reverse transcriptase or a nucleic acid molecule B1 containing a nucleotide sequence encoding the first reverse transcriptase; (3) a first PegRNA containing a first gRNA and a first tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA; wherein, the first gRNA is capable of binding to the first Cas protein and forming a first functional complex; the first functional complex is capable of breaking both strands of the first double-stranded target nucleic acid to form a fragmented target nucleic acid fragment; the first tag primer contains a first tag sequence and a first target-binding sequence, the first tag sequence is located upstream or 5' of the first target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the first target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment to form a double-stranded structure, and the first tag sequence is not bound to the target nucleic acid fragment and is in a free single-stranded state; (4) a second Cas protein or a nucleic acid molecule A2 containing a nucleotide sequence encoding the second Cas protein, wherein the second Cas protein is capable of cleaving or breaking a second double-stranded target nucleic acid; (5) a second PegRNA containing a second gRNA and a second tag primer, or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA; wherein, the second gRNA is capable of binding to the second Cas protein and forming a second functional complex; the second functional complex is capable of breaking both strands of the second double-stranded target nucleic acid to form a fragmented target nucleic acid fragment; the second tag primer contains a second tag sequence and a second target-binding sequence, the second tag sequence is located upstream or 5' of the second target-binding sequence; and, under conditions permitting nucleic acid hybridization or annealing, the second target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment to form a double-stranded structure, and the second tag sequence is not bound to the target nucleic acid fragment and is in a free single-stranded state; (6) a second reverse transcriptase or a nucleic acid molecule B2 containing a nucleotide sequence encoding the second reverse transcriptase; (7) a third Cas protein or a nucleic acid molecule A3 containing a nucleotide sequence encoding the third Cas protein, wherein the third Cas protein is capable of cleaving or breaking a third double-stranded target nucleic acid; (8) a third gRNA or a nucleic acid molecule C3 containing a nucleotide sequence encoding the third gRNA, wherein the third gRNA is capable of binding to the third Cas protein and forming a third functional complex; the third functional complex is capable of breaking both strands of the third double-stranded target nucleic acid to form fragmented nucleotide fragments a1 and a2; and the first reverse transcriptase is capable of extending the 3' end of the nucleic acid strand using the first tag primer as a template and forming a first overhang; and the second reverse transcriptase is capable of extending the 3' end of the nucleic acid strand using the second tag primer as a template and forming a second overhang; and the first overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment a1 and the second overhang is capable of hybridizing or annealing to the fragmented nucleotide fragment a2 under conditions that allow nucleic acid hybridization or annealing; wherein the second double-stranded target nucleic acid is identical to the first double-stranded target nucleic acid, the second functional complex and the first functional complex cleave the identical double-stranded target nucleic acid at different locations; and the first overhang and the second overhang are comprised on the same target nucleic acid fragment and are located on opposite nucleic acid strands from each other. wherein the third Cas protein, the second Cas protein and the first Cas protein are Cas9 proteins; and the second reverse transcriptase is the same as or different from the first reverse transcriptase.

2. The system or kit of claim 1, wherein, The third Cas protein, the second Cas protein and the first Cas protein have the amino acid sequence set forth in SEQ ID NO:

1.

3. The system or kit of claim 1, wherein, The first reverse transcriptase and the second reverse transcriptase are each independently a reverse transcriptase from Moloney murine leukemia virus, human immunodeficiency virus (HIV), Rous sarcoma virus (RSV), avian myeloblastosis virus (AMV), avian erythroblastosis virus helper virus, avian myelocytomatosis virus MC29 helper virus, avian reticuloendotheliosis virus helper virus, avian sarcoma virus UR2 helper virus, avian sarcoma virus Y73 helper virus, Rous-associated virus and myeloblastosis-associated virus (MAV).

4. The system or kit of claim 1, wherein, The first reverse transcriptase and / or the second reverse transcriptase have the amino acid sequence set forth in SEQ ID NO:

4.

5. The system or kit of claim 1, wherein, The first gRNA contains a first guide sequence, and the first guide sequence is capable of hybridizing or annealing to one nucleic acid strand of the first double-stranded target nucleic acid under conditions that allow nucleic acid hybridization or annealing.

6. The system or kit of claim 5, wherein, The first gRNA further contains a first scaffold sequence, which is capable of forming a first functional complex with the first Cas protein.

7. The system or kit of claim 6, wherein, The system or the kit has one or more selected from the following: (1) the first guide sequence is located upstream or 5' of the first scaffold sequence; (2) the first functional complex is capable of cleaving both strands of the first double-stranded target nucleic acid after the first guide sequence binds to the first double-stranded target nucleic acid.

8. The system or kit of claim 5, wherein, The system or the kit has one or more selected from the following: (1) the first target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment under conditions that allow nucleic acid hybridization or annealing, and the 3' end is formed as a result of the cleavage of the first double-stranded target nucleic acid by the first functional complex; (2) the first reverse transcriptase is capable of extending the 3' end of the nucleic acid strand using the first tag primer as a template after the first target-binding sequence hybridizes or anneals to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment; wherein the extension forms a first overhang; (3) the first tag primer is a single-stranded ribonucleic acid, and the first reverse transcriptase is an RNA-dependent reverse transcriptase; (4) the nucleic acid strand to which the first guide sequence binds is different from the nucleic acid strand to which the first target-binding sequence binds; (5) the nucleic acid strand to which the first guide sequence binds is the opposite strand of the nucleic acid strand to which the first target-binding sequence binds.

9. The system or kit of claim 1, wherein, The system or kit has one or more features selected from the following: (1) the first tag primer is optionally linked to the 3' end of the first gRNA via a linker; (2) the first tag primer is a single-stranded ribonucleic acid, and it is linked to the 3' end of the first gRNA with or without a ribonucleic acid linker to form a first PegRNA.

10. The system or kit of claim 1, having one or more features selected from the following: (1) the nucleic acid molecule Al is capable of expressing the first Cas protein in a cell; (2) the nucleic acid molecule Bl is capable of expressing the first reverse transcriptase in a cell.

11. The system or kit of claim 1, having one or more features selected from the following: (1) the nucleic acid molecule Al is comprised in an expression vector, or the nucleic acid molecule Al is an expression vector comprising a nucleotide sequence encoding the first Cas protein; (2) the nucleic acid molecule Bl is comprised in an expression vector, or the nucleic acid molecule Bl is an expression vector comprising a nucleotide sequence encoding the first reverse transcriptase.

12. The system or kit of claim 1, wherein the nucleic acid molecule Al and the nucleic acid molecule Bl are comprised in the same expression vector.

13. The system or kit of claim 1, wherein the nucleic acid molecule Al and the nucleic acid molecule Bl are comprised in different expression vectors.

14. The system or kit of claim 8, wherein, The second gRNA comprises a second guide sequence, and the second guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a second double-stranded target nucleic acid under conditions that allow nucleic acid hybridization or annealing.

15. The system or kit of claim 14, wherein, The system or kit has one or more features selected from the following: (1) the second functional complex cleaves both strands of the second double-stranded target nucleic acid after the second guide sequence binds to the second double-stranded target nucleic acid; (2) the second functional complex cleaves the same double-stranded target nucleic acid as the first functional complex, and the nucleic acid strand to which the first guide sequence binds is different from the nucleic acid strand to which the second guide sequence binds; (3) the nucleic acid strand to which the first guide sequence binds is the opposite strand of the nucleic acid strand to which the second guide sequence binds; (4) the second gRNA further comprises a second scaffold sequence that is capable of binding to the second Cas protein and forming the second functional complex.

16. The system or kit of claim 15, wherein, The system or kit has one or more features selected from the following: (1) the second guide sequence is different from the first guide sequence; (2) the second guide sequence is located upstream or 5' to the second scaffold sequence.

17. The system or kit of claim 1, having one or more features selected from the following: (1) the nucleic acid molecule A2 is capable of expressing the second Cas protein in a cell; (2) the nucleic acid molecule A2 is comprised in an expression vector, or the nucleic acid molecule A2 is an expression vector comprising a nucleotide sequence encoding the second Cas protein.

18. The system or kit of claim 15, wherein, The system or kit has one or more technical features selected from the following: (1) the second target-binding sequence is capable of hybridizing or annealing to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment under conditions that allow nucleic acid hybridization or annealing, and the 3' end is formed as a result of the second functional complex fragmenting the second double-stranded target nucleic acid; (2) the second target-binding sequence is different from the first target-binding sequence; (3) the nucleic acid strand to which the second target-binding sequence binds is different from the nucleic acid strand to which the first target-binding sequence binds; (4) the nucleic acid strand to which the second target-binding sequence binds is the opposite strand of the nucleic acid strand to which the first target-binding sequence binds; (5) the second tag sequence is different from the first tag sequence; (6) after the second target-binding sequence hybridizes or anneals to the 3' end of one nucleic acid strand of the fragmented target nucleic acid fragment, the second reverse transcriptase is capable of extending the 3' end of the nucleic acid strand using the second tag primer as a template; wherein the extension forms a second overhang; (7) the second tag primer is a single-stranded ribonucleic acid, and the second reverse transcriptase is an RNA-dependent reverse transcriptase; (8) the nucleic acid strand to which the second guide sequence binds is different from the nucleic acid strand to which the second target-binding sequence binds; (9) the nucleic acid strand to which the second guide sequence binds is the opposite strand of the nucleic acid strand to which the second target-binding sequence binds; (10) the second guide sequence binds to the same nucleic acid strand as the first target-binding sequence, and the binding position of the second guide sequence is upstream or 5' of the binding position of the first target-binding sequence; (11) the first guide sequence binds to the same nucleic acid strand as the second target-binding sequence, and the binding position of the first guide sequence is upstream or 5' of the binding position of the second target-binding sequence.

19. The system or kit of claim 1, wherein, The system or kit has one or more technical features selected from the following: (1) the nucleic acid molecule B2 is capable of expressing the second reverse transcriptase in a cell; (2) the nucleic acid molecule B2 is comprised in an expression vector, or the nucleic acid molecule B2 is an expression vector comprising a nucleotide sequence encoding the second reverse transcriptase.

20. The system or kit of claim 18 or 19, wherein, The second tag primer is a single-stranded ribonucleic acid, and it is connected to the 3' end of the second gRNA via a ribonucleic acid linker or without a ribonucleic acid linker, forming a second PegRNA.

21. The system or kit of claim 19, wherein, The system or kit has one or more technical features selected from the following: (1) the second Cas protein and the second reverse transcriptase are separate; (2) the nucleic acid molecule A2 and the nucleic acid molecule B2 are comprised in the same or different expression vectors.

22. The system or kit of claim 15, wherein, The system or kit further comprises a nucleic acid vector.

23. The system or kit of claim 22, wherein, The system or kit has any one technical feature selected from the following: (1) the nucleic acid vector is double-stranded; (2) the nucleic acid vector comprises a first guide-binding sequence capable of hybridizing or annealing to the first guide sequence, and a second guide-binding sequence capable of hybridizing or annealing to the second guide sequence; optionally, the nucleic acid vector further comprises a restriction enzyme site between the first guide-binding sequence and the second guide-binding sequence; (3) the nucleic acid vector further comprises a first PAM sequence recognized by the first Cas protein, and a second PAM sequence recognized by the second Cas protein; or, (4) the nucleic acid vector comprises a first guide-binding sequence capable of hybridizing or annealing to the first guide sequence, and a second guide-binding sequence capable of hybridizing or annealing to the second guide sequence; the nucleic acid vector further comprises a first PAM sequence recognized by the first Cas protein, and a second PAM sequence recognized by the second Cas protein; and the first functional complex is capable of binding to and cleaving the nucleic acid vector through the first guide-binding sequence and the first PAM sequence; and the second functional complex is capable of binding to and cleaving the nucleic acid vector through the second guide-binding sequence and the second PAM sequence.

24. The system or kit of claim 23, wherein, The system or kit has one or more technical features selected from the following: (1) the nucleic acid vector is a circular double-stranded vector; (2) the first guide-binding sequence and the second guide-binding sequence are located on opposite strands of the nucleic acid vector.

25. The system or kit of claim 23, wherein, The nucleic acid vector further comprises a gene of interest; wherein the nucleic acid vector further comprises a first target sequence; wherein under conditions permitting nucleic acid hybridization or annealing, the first tag primer is capable of hybridizing or annealing to the first target sequence through the first target-binding sequence to form a double-stranded structure, and the first tag sequence of the first tag primer is in a free state; and the nucleic acid vector further comprises a second target sequence; wherein under conditions permitting nucleic acid hybridization or annealing, the second tag primer is capable of hybridizing or annealing to the second target sequence through the second target-binding sequence to form a double-stranded structure, and the second tag sequence of the second tag primer is in a free state.

26. The system or kit of claim 25, wherein, The system or kit has one or more technical features selected from the following: (1) the first functional complex and the second functional complex cleave the nucleic acid vector, resulting in a nucleic acid fragment containing the gene of interest; (2) under conditions permitting nucleic acid hybridization or annealing, the first tag primer is capable of hybridizing or annealing to the 3' end of one of the nucleic acid strands of the nucleic acid fragment through the first target-binding sequence to form a double-stranded structure, and the first tag sequence of the first tag primer is in a free state; (3) the nucleic acid strand to which the first target-binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand containing the first guide-binding sequence; (4) under conditions permitting nucleic acid hybridization or annealing, the second tag primer is capable of hybridizing or annealing to the 3' end of one of the nucleic acid strands of the nucleic acid fragment through the second target-binding sequence to form a double-stranded structure, and the second tag sequence of the second tag primer is in a free state; (5) the nucleic acid strand to which the second target-binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand containing the second guide-binding sequence; (6) the nucleic acid strand to which the first target-binding sequence hybridizes or anneals is the opposite strand of the nucleic acid strand to which the second target-binding sequence hybridizes or anneals.

27. The system or kit of claim 26, wherein, The system or kit has one or more technical features selected from the following: (1) the first target sequence is located between the first guide-binding sequence and the second guide-binding sequence; (2) the first target sequence is located on the opposite strand of the first guide-binding sequence; (3) the site at which the first functional complex cleaves the nucleic acid vector is located at or in the 3' end or 3' portion of the first target sequence; (4) the second target sequence is located between the first guide-binding sequence and the second guide-binding sequence; (5) the second target sequence is located on the opposite strand of the second guide-binding sequence; (6) the site of the nucleic acid vector cleaved by the second functional complex is located at or in the 3' end or 3' portion of the second target sequence; (7) the nucleic acid strand containing the first target sequence is located on the opposite strand of the nucleic acid strand containing the second target sequence; (8) the nucleic acid vector further comprises a restriction enzyme site between the first target sequence and the second target sequence.

28. The system or kit of claim 15, wherein, The system or kit has one or more technical features selected from the following: (1) the third gRNA comprises a third guide sequence, and the third guide sequence is capable of hybridizing or annealing to one nucleic acid strand of a third double-stranded target nucleic acid under conditions permitting nucleic acid hybridization or annealing; (2) the second double-stranded target nucleic acid is identical to the first double-stranded target nucleic acid, and the third double-stranded target nucleic acid is different from the first and second double-stranded target nucleic acids; (3) the third gRNA further comprises a third scaffold sequence, which is capable of binding to the third Cas protein and forming a third functional complex; (4) the nucleic acid molecule C3 is capable of transcribing the third gRNA in a cell; (5) the nucleic acid molecule C3 is comprised in an expression vector, or the nucleic acid molecule C3 is an expression vector comprising a nucleotide sequence encoding the third gRNA.

29. The system or kit of claim 28, wherein, The system or kit has one or more technical features selected from the following: ( 1) the third guide sequence is located upstream or 5' of the third scaffold sequence; (2) the third double-stranded target nucleic acid is genomic DNA; (3) the third functional complex cleaves both strands of the third double-stranded target nucleic acid after the third guide sequence binds to the third double-stranded target nucleic acid; (4) the third guide sequence is identical to or different from the first guide sequence or the second guide sequence.

30. The system or kit of claim 1, wherein, The system or kit has one or more technical features selected from the following: (1) the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a1, and the 3' end or 3' portion is formed due to cleavage of the third double-stranded target nucleic acid by the third functional complex; (2) the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a2, and the 3' end or 3' portion is formed due to cleavage of the third double-stranded target nucleic acid by the third functional complex.

31. The system or kit of claim 30, wherein, The system or kit has one or more technical features selected from the following: (1) the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to the 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a1, and there is a first spacer region between the 3' portion of the nucleotide fragment a1 and the cleaved end formed by the third double-stranded target nucleic acid; (2) the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to the 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a2, and there is a second spacer region between the 3' portion of the cleaved nucleotide fragment a2 and the cleaved end of the third double-stranded target nucleic acid.

32. The system or kit of claim 31, wherein, The system or kit has one or more technical features selected from the following: (1) the nucleic acid molecule A3 is capable of expressing the third Cas protein in a cell; (2) the nucleic acid molecule A3 is comprised in an expression vector, or, the nucleic acid molecule A3 is an expression vector comprising a nucleotide sequence encoding the third Cas protein.

33. The system or kit of claim 29, wherein, The system or kit further comprises: a third PegRNA comprising a third gRNA and a third tag primer, and a third reverse transcriptase; a fourth Cas protein or a nucleic acid molecule A4 comprising a nucleotide sequence encoding the fourth Cas protein, and the fourth Cas protein is a Cas9 protein; a fourth PegRNA comprising a fourth gRNA and a fourth tag primer, and a fourth reverse transcriptase; wherein the fourth gRNA is capable of binding to the fourth Cas protein and forming a fourth functional complex; wherein the fourth Cas protein is capable of cleaving or cutting the fourth double-stranded target nucleic acid; the second double-stranded target nucleic acid is the same as the first double-stranded target nucleic acid, and the fourth double-stranded target nucleic acid is the same as the third double-stranded target nucleic acid but different from the first or second double-stranded target nucleic acid; and the third tag primer comprises a third tag sequence and a third target-binding sequence, the third tag sequence is located upstream or 5' of the third target-binding sequence; and the third target-binding sequence hybridizes or anneals to the 3' end of one nucleic acid strand of the cleaved nucleotide fragment a1 or a2 under conditions permitting nucleic acid hybridization or annealing, and the third reverse transcriptase extends the 3' end of the nucleic acid strand using the third tag primer as a template; wherein the extension forms a third overhang; and the fourth tag primer comprises a fourth tag sequence and a fourth target-binding sequence, the fourth tag sequence is located upstream or 5' of the fourth target-binding sequence; and the fourth target-binding sequence hybridizes or anneals to the 3' end of the other nucleic acid strand of the cleaved target nucleotide fragment a1 or a2 under conditions permitting nucleic acid hybridization or annealing, and the fourth reverse transcriptase extends the 3' end of the nucleic acid strand using the fourth tag primer as a template; wherein the extension forms a fourth overhang; and the third overhang is capable of hybridizing or annealing to the upstream nucleotide sequence of the first overhang, and the fourth overhang is capable of hybridizing or annealing to the upstream nucleotide sequence of the second overhang under conditions permitting nucleic acid hybridization or annealing; the first overhang is capable of hybridizing or annealing to the upstream nucleotide sequence of the third overhang, and the second overhang is capable of hybridizing or annealing to the upstream nucleotide sequence of the fourth overhang.

34. The system or kit of claim 33, wherein, The fourth functional complex cleaves the same double-stranded target nucleic acid as the third functional complex, and the nucleic acid strand to which the fourth guide sequence binds is different from the nucleic acid strand to which the third guide sequence binds.

35. A method for cleaving a double-stranded target nucleic acid into target nucleic acid fragments and adding overhangs to each of the two 3' ends of the target nucleic acid fragments, wherein, The method comprises using the system or kit of any one of claims 1-34; wherein the first double-stranded target nucleic acid is the same as the second double-stranded target nucleic acid.

36. The method of claim 35, wherein, The method comprises the following steps: i. providing a double-stranded target nucleic acid; and providing a first Cas protein, a first PegRNA, a first reverse transcriptase, a second Cas protein, a second PegRNA and a second reverse transcriptase; ii. contacting the double-stranded target nucleic acid with the first Cas protein, the first PegRNA, the first reverse transcriptase, the second Cas protein, the second PegRNA and the second reverse transcriptase.

37. The method of claim 36, wherein, In step ii: the first Cas protein and the first PegRNA form a first functional complex, and the second Cas protein and the second PegRNA form a second functional complex; and, the first and second functional complexes bind to and cleave the double-stranded target nucleic acid to form a target nucleic acid fragment F1; and, the first tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the target nucleic acid fragment F1 through the first target-binding sequence; and, the second tag primer hybridizes or anneals to the 3' end of the other nucleic acid strand of the target nucleic acid fragment F1 through the second target-binding sequence; and, the first reverse transcriptase and the second reverse transcriptase respectively extend the target nucleic acid fragment F1 with the first tag primer and the second tag primer annealed thereto as templates to form a target nucleic acid fragment F2 with a first overhang and a second overhang.

38. The method of claim 36, wherein, The method has one or more technical features selected from the following: (1) the method is performed in a cell; (2) in step i, the first Cas protein or nucleic acid molecule A1, the first reverse transcriptase or nucleic acid molecule B1, the first PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA, the second Cas protein or nucleic acid molecule A2, the second reverse transcriptase or nucleic acid molecule B2, and the second PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA are delivered into the cell to provide the first Cas protein, the first PegRNA, the first reverse transcriptase, the second Cas protein, the second PegRNA and the second reverse transcriptase in the cell; (3) the nucleic acid molecule A1 and the nucleic acid molecule B1 are comprised in the same or different expression vectors; (4) the nucleic acid molecule A2 and the nucleic acid molecule B2 are comprised in the same or different expression vectors; (5) in step i, the double-stranded target nucleic acid or a nucleic acid molecule T containing the double-stranded target nucleic acid is delivered into the cell to provide the double-stranded target nucleic acid in the cell; (6) the double-stranded target nucleic acid or the nucleic acid molecule T contains a first PAM sequence recognized by the first Cas protein and a second PAM sequence recognized by the second Cas protein; (7) in step ii, the first functional complex binds to and cleaves the double-stranded target nucleic acid or the nucleic acid molecule T through the first PAM sequence and the first gRNA; and, the second functional complex binds to and cleaves the double-stranded target nucleic acid or the nucleic acid molecule T through the second PAM sequence and the second gRNA.

39. The method of claim 38, wherein, the second Cas protein is the same as the first Cas protein, and the second reverse transcriptase is the same as the first reverse transcriptase; wherein the first Cas protein forms a first functional complex with the first and second PegRNAs, and the first reverse transcriptase extends the target nucleic acid fragment F1 using the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang.

40. A method for inserting a target nucleic acid fragment into a nucleic acid molecule of interest; wherein, The method comprises using the system or kit of any one of claims 1-32; wherein the first double-stranded target nucleic acid and the second double-stranded target nucleic acid are the same, for providing the target nucleic acid fragments; and the third double-stranded target nucleic acid is a nucleic acid molecule of interest.

41. The method of claim 40, wherein, The method comprises: a. fragmenting the first double-stranded target nucleic acid into a target nucleic acid fragment F1 and adding overhangs to both 3' ends of the target nucleic acid fragment F1 by the method of any one of claims 35-39, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; b. fragmenting the nucleic acid molecule of interest with a third functional complex to form fragmented nucleotide fragments a1 and a2; and, c. ligating the nucleotide fragments a1 and a2 with the target nucleic acid fragment F2, thereby inserting the target nucleic acid fragment into the nucleic acid molecule of interest.

42. The method of claim 40, wherein, The method comprises the following steps: i. providing a double-stranded target nucleic acid and a nucleic acid molecule of interest; and providing a first Cas protein, a first PegRNA, a first reverse transcriptase, a second Cas protein, a second PegRNA, a second reverse transcriptase, a third Cas protein, and a third gRNA; ii. contacting the double-stranded target nucleic acid with the first Cas protein, the first PegRNA, the first reverse transcriptase, the second Cas protein, the second PegRNA, and the second reverse transcriptase, and contacting the nucleic acid molecule of interest with the third Cas protein and the third gRNA.

43. The method of claim 42, wherein, In step ii: the first Cas protein and the first PegRNA combine to form a first functional complex, the second Cas protein and the second PegRNA combine to form a second functional complex, and the third Cas protein and the third gRNA combine to form a third functional complex; and, the first and second functional complexes bind to and fragment the double-stranded target nucleic acid to form a target nucleic acid fragment F1, and the third functional complex binds to and fragments the nucleic acid molecule of interest to form fragmented nucleotide fragments a1 and a2; and, the first tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the target nucleic acid fragment F1 via the first target-binding sequence; and the second tag primer hybridizes or anneals to the 3' end of the other nucleic acid strand of the target nucleic acid fragment F1 via the second target-binding sequence; and, the first reverse transcriptase and the second reverse transcriptase extend the target nucleic acid fragment F1 using the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first overhang and the second overhang are capable of hybridizing or annealing to the fragmented nucleotide fragments a1 and a2, respectively; and, the first reverse transcriptase and the second reverse transcriptase extend the target nucleic acid fragment F1 using the first and second tag primers annealed to the target nucleic acid fragment F1 as templates, to form a target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first overhang and the second overhang are capable of hybridizing or annealing to the fragmented nucleotide fragments a1 and a2, respectively; and, The target nucleic acid fragment F2 is hybridized or annealed to the first overhang and the second overhang of the nucleotide fragment a1 and a2, respectively, and is inserted or ligated between the nucleotide fragments a1 and a2, thereby inserting the target nucleic acid fragment into the nucleic acid molecule of interest.

44. The method of claim 43, wherein, The method has one or more features selected from the following: (1) the first overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a1, and the 3' end or 3' portion is formed due to the cleavage of the nucleic acid molecule of interest by the third functional complex; (2) the complement of the first tag sequence or the first overhang is capable of hybridizing or annealing to the 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a1, and there is a first spacer region between the 3' portion of the nucleotide fragment a1 and the cleaved end formed by the third double-stranded target nucleic acid; (3) the second overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment a2, and the 3' end or 3' portion is formed due to the cleavage of the nucleic acid molecule of interest by the third functional complex; (4) the complement of the second tag sequence or the second overhang is capable of hybridizing or annealing to the 3' portion of one nucleic acid strand of the cleaved nucleotide fragment a2, and there is a second spacer region between the 3' portion of the nucleotide fragment a2 and the cleaved end formed by the third double-stranded target nucleic acid; (5) the method is performed in a cell; (6) in step i, the double-stranded target nucleic acid or the nucleic acid molecule T containing the double-stranded target nucleic acid is delivered into the cell to provide the double-stranded target nucleic acid in the cell; (7) the double-stranded target nucleic acid or the nucleic acid molecule T contains a first PAM sequence recognized by the first Cas protein and a second PAM sequence recognized by the second Cas protein; (8) in step ii, the first functional complex binds to the double-stranded target nucleic acid or the nucleic acid molecule T through the first PAM sequence and the first PegRNA, and cleaves it; and the second functional complex binds to the double-stranded target nucleic acid or the nucleic acid molecule T through the second PAM sequence and the second PegRNA, and cleaves it; (9) the nucleic acid molecule of interest contains a third PAM sequence recognized by the third Cas protein; (10) in step ii, the third functional complex binds to the nucleic acid molecule of interest through the third PAM sequence and the third gRNA, and cleaves it; (11) the nucleic acid molecule of interest is the genomic DNA of the cell.

45. The method of claim 44, wherein, The first, second and third Cas proteins are the same Cas protein, and the second reverse transcriptase is the same as the first reverse transcriptase; wherein the first Cas protein forms the first, second and third functional complexes with the first, second and third gRNAs, respectively, and the first reverse transcriptase extends the target nucleic acid fragment F1 using the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates, respectively, to form the target nucleic acid fragment F2 with the first overhang and the second overhang.

46. A method for replacing a nucleotide segment in a nucleic acid molecule of interest with a target nucleic acid segment; wherein, The method comprises using the system or kit of claim 33 or 34; wherein the first double-stranded target nucleic acid and the second double-stranded target nucleic acid are the same, for providing the target nucleic acid fragment; and the third double-stranded target nucleic acid and the fourth double-stranded target nucleic acid are the same, for the nucleic acid molecule of interest.

47. The method of claim 46, wherein, The method comprises: a. cleaving the first double-stranded target nucleic acid into the target nucleic acid fragment F1 by the method of any one of claims 35-39, and adding overhangs to both 3' ends of the target nucleic acid fragment F1, respectively, to form the target nucleic acid fragment F2 having a first overhang and a second overhang; b. cleaving the nucleic acid molecule of interest with the third and fourth functional complexes to form the cleaved nucleotide fragments b1, b2 and b3; wherein, before cleavage, the nucleotide fragments b1, b2 and b3 are arranged in the nucleic acid molecule of interest in sequence; and, c. ligating the nucleotide fragments b1 and b3 with the target nucleic acid fragment F2, thereby replacing the nucleotide fragment b2 in the nucleic acid molecule of interest with the target nucleic acid fragment.

48. The method of claim 46 or 47, wherein, The method comprises the following steps: i. providing the double-stranded target nucleic acid and the nucleic acid molecule of interest; and providing the first Cas protein, the first PegRNA, the first reverse transcriptase, the second Cas protein, the second PegRNA, the second reverse transcriptase, the third Cas protein, the third PegRNA, the third reverse transcriptase, the fourth Cas protein, the fourth PegRNA and the fourth reverse transcriptase; ii contacting the double-stranded target nucleic acid with the first Cas protein, the first PegRNA, the first reverse transcriptase, the second Cas protein, the second PegRNA and the second reverse transcriptase, and contacting the nucleic acid molecule of interest with the third Cas protein, the third PegRNA, the third reverse transcriptase, the fourth Cas protein, the fourth reverse transcriptase and the fourth PegRNA.

49. The method of claim 48, wherein, In step ii: the first Cas protein and the first PegRNA combine to form a first functional complex, the second Cas protein and the second PegRNA combine to form a second functional complex, the third Cas protein and the third PegRNA combine to form a third functional complex, and the fourth Cas protein and the fourth PegRNA combine to form a fourth functional complex; and, the first and second functional complexes bind to and cleave the double-stranded target nucleic acid to form the target nucleic acid fragment F1, and the third and fourth functional complexes bind to and cleave the nucleic acid molecule of interest to form the cleaved nucleotide fragments b1, b2 and b3; and, the first tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the target nucleic acid fragment F1 via the first target binding sequence; and the second tag primer hybridizes or anneals to the 3' end of the other nucleic acid strand of the target nucleic acid fragment F1 via the second target binding sequence; and, the first reverse transcriptase and the second reverse transcriptase respectively extend the target nucleic acid fragment F1 using the first tag primer and the second tag primer annealed to the target nucleic acid fragment F1 as templates to form the target nucleic acid fragment F2 having a first overhang and a second overhang; wherein the first overhang and the second overhang are respectively capable of hybridizing or annealing to the cleaved nucleotide fragments b1 and b3; and, the third tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the nucleotide fragment b1 via the third target-binding sequence, wherein the 3' end is formed as a result of the third functional complex breaking the nucleic acid molecule of interest; and the fourth tag primer hybridizes or anneals to the 3' end of one nucleic acid strand of the nucleotide fragment b3 via the fourth target-binding sequence, wherein the 3' end is formed as a result of the fourth functional complex breaking the nucleic acid molecule of interest; and the third reverse transcriptase extends the nucleotide fragment b1 using the third tag primer annealed to the nucleotide fragment b1 as a template to form a nucleotide fragment b1 with a third overhang; and the fourth reverse transcriptase extends the nucleotide fragment b3 using the fourth tag primer annealed to the nucleotide fragment b3 as a template to form a nucleotide fragment b3 with a fourth overhang; wherein the third overhang and the fourth overhang are capable of hybridizing or annealing to the target nucleic acid fragment F2, respectively; and the target nucleic acid fragment F2 hybridizes or anneals to the nucleotide fragments b1 and b3 via the first, second, third and fourth overhangs, respectively, and is ligated between the nucleotide fragments b1 and b3, thereby replacing the nucleotide fragment b2 in the nucleic acid molecule of interest with the target nucleic acid fragment.

50. The method of claim 49, wherein, The method has one or more features selected from the group consisting of: (1) the first overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment b1, and the 3' end or 3' portion is formed as a result of the third functional complex breaking the nucleic acid molecule of interest; (2) the first overhang is capable of hybridizing or annealing to the third overhang of the nucleotide fragment b1 or a nucleotide sequence upstream thereof; (3) the second overhang is capable of hybridizing or annealing to the 3' end or 3' portion of one nucleic acid strand of the nucleotide fragment b3, and the 3' end or 3' portion is formed as a result of the fourth functional complex breaking the nucleic acid molecule of interest; (4) the second overhang is capable of hybridizing or annealing to the fourth overhang of the nucleotide fragment b3 or a nucleotide sequence upstream thereof; (5) the third overhang is capable of hybridizing or annealing to the first overhang or a nucleotide sequence upstream thereof; (6) the fourth overhang is capable of hybridizing or annealing to the second overhang or a nucleotide sequence upstream thereof; (7) the method is performed in a cell; (8) in step i, the first Cas protein or nucleic acid molecule A1, the first reverse transcriptase or nucleic acid molecule B1, the first PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the first PegRNA, the second Cas protein or nucleic acid molecule A2, the second reverse transcriptase or nucleic acid molecule B2, the second PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the second PegRNA, the third Cas protein or nucleic acid molecule A3, the third reverse transcriptase or nucleic acid molecule B3, the third PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the third PegRNA, the fourth Cas protein or nucleic acid molecule A4, the fourth reverse transcriptase or nucleic acid molecule B4, the fourth PegRNA or a nucleic acid molecule containing a nucleotide sequence encoding the fourth PegRNA are delivered into the cell to provide the first, second, third and fourth Cas proteins, the first, second, third and fourth PegRNAs, the first, second, third and fourth reverse transcriptases in the cell; (9) in step i, the double-stranded target nucleic acid or a nucleic acid molecule T containing the double-stranded target nucleic acid is delivered into the cell to provide the double-stranded target nucleic acid in the cell; (10) the double-stranded target nucleic acid or the nucleic acid molecule T contains a first PAM sequence recognized by the first Cas protein and a second PAM sequence recognized by the second Cas protein; (11) in step ii, the first functional complex binds to and cleaves the double-stranded target nucleic acid or the nucleic acid molecule T through the first PAM sequence and the first PegRNA; and the second functional complex binds to and cleaves the double-stranded target nucleic acid or the nucleic acid molecule T through the second PAM sequence and the second PegRNA; (12) the nucleic acid molecule of interest contains a third PAM sequence recognized by the third Cas protein and a fourth PAM sequence recognized by the fourth Cas protein; (13) in step ii, the third functional complex binds to and cleaves the nucleic acid molecule of interest through the third PAM sequence and the third PegRNA; and the fourth functional complex binds to and cleaves the nucleic acid molecule of interest through the fourth PAM sequence and the fourth PegRNA; (14) the nucleic acid molecule of interest is the genomic DNA of the cell.

Citation Information

Patent Citations

  • Method for distributing large files to multiple recipients

    US20020198944A1

  • APP modification via base editing using the crispr / CAS9 system

    WO2020124257A1