Composition and method for prime editing technique

By developing fusion proteins, combining V-type Cas proteins and reverse transcriptase, the problems of inefficient and limited application range of existing guided editing technologies are solved, and efficient and flexible gene editing is achieved.

WO2025130907A1PCT designated stage expired Publication Date: 2025-06-26SHANDONG SHUNFENG BIOTECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140239
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing guided editing technology has low editing efficiency and is limited to type II CRISPR/Cas protein, and its application range is limited.

Method used

Developed a fusion protein containing V-type Cas protein and reverse transcriptase for guided editing, extending the scope of application of guided editing.

Benefits of technology

It improves the efficiency and flexibility of guided editing, expands the scope of application, and enables efficient gene editing in non-dividing cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140239_26062025_PF_FP_ABST
    Figure CN2024140239_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of nucleic acid editing, in particular to the technical field of clustered regularly interspaced short palindromic repeats (CRISPRs). Specifically, provided in the present invention is a fusion protein. The fusion protein can be used for performing prime editing on a target nucleic acid. The fusion protein of the present invention expands a prime editing system, thereby enabling the system to be compatible with the Type V CRISPR system and allowing for precise editing.
Need to check novelty before this filing date? Find Prior Art

Description

A composition and method for guided editing technology

[0001] This application claims priority to Chinese patent application CN202311777306.3, filed on December 22, 2023. This application incorporates the entire text of the aforementioned Chinese patent application. Technical Field

[0002] The present invention relates to the field of gene editing, in particular to the field of guided editing technology. Specifically, the present invention relates to a fusion protein and its application, and in particular to a V-type Cas protein fused to a reverse transcriptase and its application. Background Art

[0003] CRISPR / Cas technology is a widely used gene editing technology. It uses RNA to specifically bind to target sequences on the genome and cut DNA to produce double-strand breaks. It uses biological non-homologous end joining (NHEJ) or homologous recombination (HDR) for site-directed gene editing. Traditional homology-directed repair (HDR) is mostly very inefficient, especially in non-dividing cells, and competitive non-homologous ends dominate, leading to insertion-deletion byproducts.

[0004] Prime editing technology enables base substitution or fragment insertion at specific sites. However, the editing efficiency of prime editing technology still needs to be improved. So far, prime editing systems are limited to type II CRISPR / Cas proteins, such as Cas9 from Streptococcus pyogenes and Staphylococcus aureus.

[0005] Therefore, the present invention develops a V-type Cas protein that can be used for guide editing, which can be combined with reverse transcriptase to form a fusion protein for guide editing, thereby expanding the application scope of guide editing. Summary of the Invention

[0006] In one aspect, the present invention provides a fusion protein comprising a Cas protein and a reverse transcriptase.

[0007] In one embodiment, the Cas protein is a V-type Cas protein or a variant thereof.

[0008] In a preferred embodiment, the Cas protein is selected from Cas9, Cas9n, dCas9, CasX, CasY, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Cas12a, Cas12b, Cas12c, Cas12e, Cas12d, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, Cas9-NG, LbCasl2a, enAsCasl2a, Cas9-KKH, circular displacement Cas9, Argonaute (Ago) domain, SmacCas9, Spy-macCas9, SpGas9-NRRH, SpaCas9-NRTH, SpaCas9-NRCH, Cas9-NG-CP1041, Cas9-NG-VRQR, dCas12i, nCas12i or Argonaute and variants thereof, preferably, the Cas protein is Cas12i.

[0009] In one embodiment, the amino acid sequence of the Cas protein has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to SEQ ID NO.1.

[0010] In one embodiment, the amino acid sequence of the Cas protein has one or more amino acid substitutions, deletions or additions compared to SEQ ID NO. 1, for example, 1-20 amino acid substitutions, deletions or additions, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amino acid substitutions, deletions or additions; and the biological function of the Cas protein is substantially retained.

[0011] The biological function of the Cas protein includes the biological function of the parent V-type Cas protein, for example, the activity of binding to the guide RNA, the endonuclease activity, or the activity of binding to and cutting a specific site of the target sequence under the guidance of the guide RNA (including but not limited to Cis cutting activity and Trans cutting activity).

[0012] In one embodiment, the amino acid sequence of the Cas protein is shown in SEQ ID NO.1.

[0013] In one embodiment, the reverse transcriptase is a naturally occurring reverse transcriptase sequence from a retrovirus or retrotransposon, or a variant thereof.

[0014] In some embodiments, the reverse transcriptase is selected from any one or more of M-MLV reverse transcriptase (Moloney murine leukemia virus reverse transcriptase) or AMV reverse transcriptase (avian myoblastosis virus reverse transcriptase).

[0015] In a preferred embodiment, the reverse transcriptase is M-MLV reverse transcriptase.

[0016] In some embodiments, the amino acid sequence of the reverse transcriptase has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity compared to SEQ ID NO.2.

[0017] In one embodiment, the amino acid sequence of the reverse transcriptase has one or more amino acid substitutions, deletions or additions compared to SEQ ID NO. 2, for example, 1-20 amino acid substitutions, deletions or additions, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 amino acid substitutions, deletions or additions; and the biological function of the reverse transcriptase is substantially retained.

[0018] In a preferred embodiment, the amino acid sequence of the reverse transcriptase is shown as SEQ ID NO.2.

[0019] It is clear to those skilled in the art that the structure of a protein can be changed without adversely affecting its activity and functionality. For example, one or more conservative amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Those skilled in the art are aware of examples and embodiments of conservative amino acid substitutions. Specifically, the amino acid residue can be replaced with another amino acid residue belonging to the same group as the site to be replaced, that is, a non-polar amino acid residue can be substituted for another non-polar amino acid residue, a polar uncharged amino acid residue can be substituted for another polar uncharged amino acid residue, a basic amino acid residue can be substituted for another basic amino acid residue, and an acidic amino acid residue can be substituted for another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, conservative substitutions in which one amino acid is replaced by another amino acid belonging to the same group fall within the scope of the present invention. Therefore, the Cas protein or reverse transcriptase of the present invention may contain one or more conservative substitutions in the amino acid sequence, and these conservative substitutions are preferably generated by substitution according to Table 1. In addition, the present invention also encompasses proteins that further comprise one or more other non-conservative substitutions, as long as the non-conservative substitutions do not significantly affect the desired functions and biological activities of the proteins of the present invention.

[0020] Conservative amino acid replacement can be carried out at one or more predicted non-essential amino acid residues.A "non-essential" amino acid residue is an amino acid residue that can be changed (deleted, substituted or replaced) without changing biological activity, while an "essential" amino acid residue is required for biological activity.A "conservative amino acid replacement" is a replacement in which an amino acid residue is replaced by an amino acid residue with a similar side chain.Amino acid replacement can be carried out in the non-conserved region of the above-mentioned engineered V-type Cas protein.In general, such replacement is not carried out on conserved amino acid residues, or on amino acid residues located within a conserved motif, where such residues are required for protein activity.However, it will be appreciated by those skilled in the art that functional variants can have fewer conservative or non-conservative changes in conserved regions.

[0021] Table 1

[0022] It is well known in the art that one or more amino acid residues can be changed (replaced, deleted, truncated or inserted) from the N and / or C terminus of a protein while still retaining its functional activity. Therefore, proteins in which one or more amino acid residues are changed from the N and / or C terminus of a Cas protein or reverse transcriptase while retaining its desired functional activity are also within the scope of the present invention. These changes may include changes introduced by modern molecular methods such as PCR, which includes PCR amplification of protein coding sequences by means of including amino acid coding sequences in oligonucleotides used in PCR amplification.

[0023] It will be appreciated that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such manipulations are generally known in the art. For example, amino acid sequence variants of the above-described proteins can be prepared by mutations in the DNA. Other forms of mutagenesis and / or directed evolution can also be accomplished, for example, using known mutagenesis, recombination, and / or shuffling methods, in combination with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.

[0024] Those skilled in the art will appreciate that these minor amino acid changes in the Cas protein or reverse transcriptase of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be changed, but the polypeptide may retain its activity. If the mutations present are not close to the catalytic domain, active site, or other functional domains, a smaller effect can be expected.

[0025] Those skilled in the art can identify the essential amino acids of the Cas protein or reverse transcriptase of the present invention according to methods known in the art, such as site-directed mutagenesis or protein evolution or analysis of bioinformatics systems. The catalytic domain, active site or other functional domain of the protein can also be determined by physical analysis of the structure, such as by the following techniques: such as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, combined with mutations of amino acids at putative key sites.

[0026] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine ​​(Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0027] Specific amino acid positions (numbers) within the proteins of the present invention are determined by aligning the amino acid sequence of the target protein with a reference amino acid sequence (e.g., SEQ ID No. 1) using standard sequence alignment tools, such as the Smith-Waterman algorithm or the CLUSTALW2 algorithm, wherein the sequences are considered aligned when the alignment score is the highest. The alignment score can be calculated according to the method described in Wilbur, WJ and Lipman, DJ (1983) Rapid similarity searches of nucleic acid and protein data banks. Proc. Natl. Acad. Sci. USA, 80:726-730. Preferably, the default parameters are used in the ClustalW2 (1.82) algorithm: protein gap open penalty = 10.0; protein gap extension penalty = 0.2; protein matrix = Gonnet; protein / DNA end gap = -1; protein / DNAGAPDIST = 4. Preferably, the AlignX program (part of the vectorNTI group) is used to determine the position of specific amino acids within the protein of the present invention by aligning the amino acid sequence of the protein with SEQ ID No. 1 using default parameters suitable for multiple alignment (gap opening penalty: 10.0 gap extension penalty 0.05).

[0028] In one embodiment, the fusion protein further comprises a linker connecting the Cas protein and the reverse transcriptase.

[0029] In the present invention, a linker can be used to connect any peptide or protein domain of the present invention. In certain embodiments, the linker is a polypeptide. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, a disulfide bond, a carbon-heteroatom bond, etc.). In certain embodiments, the linker is an amide-connected carbon-nitrogen bond. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In certain embodiments, the linker is a polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminocaproic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker can include a functionalized portion to facilitate attachment of a nucleophile (e.g., thiol, amino) from a peptide to the linker. Any electrophile can be used as a part of a linker.

[0030] In some embodiments, the linker can be a GS linker. In some embodiments, the linker can comprise the amino acid sequence (GGS)n, GS, SG, GSSG, S(GGS)n, SGGS, or (GGGGS)n, where n is an integer from 1-20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20). In some embodiments, the linker can comprise the amino acid sequence: SGGSGGSGGS. In some embodiments, the linker can comprise the amino acid sequence: SGSETPGTSESATPES, also known as an XTEN linker. In some embodiments, the linker can comprise the amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGS, also known as a GS-XTEN-GS linker.

[0031] In a preferred embodiment, the linker is an XTEN linker with the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO. 3).

[0032] In one embodiment, the fusion protein further comprises a modifying portion selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting portion, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), a nuclease domain (e.g., Fok1), and a domain having an activity selected from the following: nucleotide deaminase (e.g., adenosine deaminase or cytidine deaminase), methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof. The NLS sequence is well known to those skilled in the art, and examples thereof include, but are not limited to, the SV40 large T antigen, EGL-13, c-Myc, and TUS protein.

[0033] In one embodiment, the fusion protein of the present invention further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In other embodiments, both the N-terminus and the C-terminus of the fusion protein are connected to the NLS.

[0034] In some embodiments, the NLS is fused to the N-terminus of the Cas protein. In some embodiments, the NLS is fused to the C-terminus of the Cas protein. In some embodiments, the NLS is fused to the N-terminus of the reverse transcriptase. In some embodiments, the NLS is fused to the C-terminus of the reverse transcriptase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. Nuclear localization sequences (NLS) are known in the art and are obvious to technicians. In some embodiments, the sequence of the NLS comprises the amino acid sequence KRPAATKKAGQAKKKK (SEQ ID No.6), or PKKKRKV (SEQ ID No.7).

[0035] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can select other suitable epitope tags (for example, purification, detection or tracing).

[0036] The reporter gene sequence is well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0037] In one embodiment, the fusion protein of the present invention comprises a domain capable of binding to a DNA molecule or an intracellular molecule, such as maltose binding protein (MBP), the DNA binding domain (DBD) of Lex A, the DBD of GAL4, and the like.

[0038] In one embodiment, the fusion protein of the invention comprises a detectable label, such as a fluorescent dye, eg, FITC or DAPI.

[0039] In one embodiment, one end of the Cas protein is connected to the reverse transcriptase, and the other end can also be connected to a single-chain binding protein or a ubiquitin-like modification protein.

[0040] In one embodiment, the reverse transcriptase is connected to the N-terminus or C-terminus of the Cas protein; in a preferred embodiment, the reverse transcriptase is connected to the N-terminus of the Cas protein.

[0041] In one embodiment, the single-chain binding protein or ubiquitin-like modifying protein is connected to the N-terminus or C-terminus of the Cas protein.

[0042] In one embodiment, one end of the reverse transcriptase is connected to the Cas protein, and the other end can also be connected to a single-chain binding protein or a ubiquitin-like modification protein.

[0043] In one embodiment, the single-chain binding protein or ubiquitin-like modifying protein is linked to the N-terminus or C-terminus of the reverse transcriptase.

[0044] In one embodiment, the single-stranded binding protein is capable of stabilizing the ssDNA exposed during double-strand breaks, and the single-stranded binding protein is selected from Brex27, EcRecA (single-stranded binding protein from Escherichia coli), BsRecA (single-stranded binding protein from Bacillus subtilis) or T4SSB (single-stranded binding protein from T4 phage), etc. In a preferred embodiment, the single-stranded binding protein is Brex27, and preferably, the Brex amino acid sequence is shown in SEQ ID NO.5.

[0045] In one embodiment, the ubiquitin-like modified protein can promote accurate DNA repair. Preferably, the amino acid sequence of the ubiquitin-like modified protein is shown in SEQ ID NO.15.

[0046] In some embodiments, the connection is a direct connection, or a connection through a linker.

[0047] In a preferred embodiment, the amino acid sequence of the linker is shown as SEQ ID NO.3.

[0048] Protein-nucleic acid complexes / compositions

[0049] In one aspect, the present invention also provides a complex for guiding editing, the complex comprising

[0050] (i) a protein component selected from: the above-mentioned fusion protein comprising a Cas protein and a reverse transcriptase;

[0051] (ii) a nucleic acid component, which is a guide editing guide RNA (PEgRNA), wherein the PEgRNA comprises a guide RNA (gRNA) and a nucleic acid extension arm, wherein the nucleic acid extension arm comprises a primer binding site sequence (PBS) and a reverse transcription template sequence (RTT); the gRNA comprises a guide sequence and a backbone sequence; the guide sequence is capable of pairing with the target nucleic acid, and the backbone sequence is capable of interacting with the Cas protein;

[0052] The protein component and the nucleic acid component combine with each other to form a complex.

[0053] In a preferred embodiment, the backbone sequence is shown as SEQ ID NO.4.

[0054] In one embodiment, the complex or composition is non-naturally occurring or modified. In one embodiment, at least one component of the complex or composition is non-naturally occurring or modified. In one embodiment, the first component is non-naturally occurring or modified; and / or the second component is non-naturally occurring or modified.

[0055] In one embodiment, the reverse transcription template stores target site editing information, and the fusion protein is capable of targeting the target sequence when combined with the guide editing guide RNA.

[0056] In one embodiment, the target sequence comprises a target strand and a complementary non-target strand.

[0057] In one embodiment, the gRNA is capable of hybridizing to the complementary non-target strand (non-PAM strand).

[0058] In one embodiment, the nucleic acid extension arm is located at the 3' or 5' end of the guide RNA, or at an intramolecular position in the guide RNA, and the nucleic acid extension arm is DNA or RNA.

[0059] In one embodiment the nucleic acid extension arm further comprises a homology arm sequence.

[0060] In one embodiment, the homology arm is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides or at least 30 nucleotides. Preferably, the homology arm is 17 nucleotides.

[0061] In one embodiment, the nucleic acid extension arm is 5-200bp in length, preferably, the nucleic acid extension arm is 17-159bp in length, preferably, the nucleic acid extension arm is 20-150bp in length, preferably, the nucleic acid extension arm is 30-140bp in length, preferably, the nucleic acid extension arm is 40-135bp in length, preferably, the nucleic acid extension arm is 50-130bp, preferably, the nucleic acid extension arm is 60-120bp, preferably, the nucleic acid extension arm is 70-120bp, preferably, the nucleic acid extension arm is 80-115bp, preferably, the nucleic acid extension arm is 90-110bp, preferably, the nucleic acid extension arm is 100bp.

[0062] In one embodiment, the length of the PBS sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, or at least 15 nucleotides. Preferably, the length of the PBS sequence is 9-79 bp, preferably, 8-48 bp, preferably, 48 bp, and preferably, 40 bp.

[0063] In one embodiment, the length of the RTT sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, or at least 15 nucleotides, preferably, 8-80 bp, preferably, 52-72 bp, preferably, 36-72 bp.

[0064] In one embodiment, the PEgRNA further comprises at least one additional structure selected from the group consisting of a linker, a stem-loop, a hairpin, a toe-loop, an aptamer, or an RNA-protein recruitment domain.

[0065] In one embodiment, the PBS sequence and the guide sequence are paired with the same strand (non-PAM strand) of the target nucleic acid.

[0066] In one embodiment, the PBS sequence and the guide sequence are paired with the two strands of the target nucleic acid, respectively.

[0067] In one embodiment, the complementary homology of the RTT sequence to the target nucleic acid sequence is less than 100%, or less than 95%, or less than 90%, or less than 85%, or less than 80%, or less than 70%, or less than 60%, or less than 50%, or less than 40%, or less than 30%, or less than 20%, or less than 10%, or less than 5%.

[0068] In one embodiment, the RTT sequence is not complementary to the target nucleic acid sequence.

[0069] In one embodiment, the RTT sequence is complementary to the target nucleic acid sequence and has one or more nucleotide mispairing, i.e., at least 1%, at least 5%, at least 10%, at least 20%, at least 40%, at least 60%, at least 80%, at least 90%, at least 95% nucleotide mispairing.

[0070] In one embodiment, the PEgRNA includes gRNA, RTT sequence and PBS sequence from 5' to 3' direction.

[0071] In one embodiment, the PEgRNA includes gRNA, RTT sequence and PBS sequence from 3' to 5' direction.

[0072] In a preferred embodiment, the sequence of the PEgRNA is shown as SEQ ID NO.9, SEQ ID NO.10 or SEQ ID NO.11.

[0073] In one embodiment, the PEgRNA further comprises a termination signal.

[0074] In a preferred embodiment, one end of the PEgRNA is further connected to a DNA guanine quadruplex (G-quadruplex). Preferably, the nucleotide sequence of the G-quadruplex is shown in SEQ ID NO.8.

[0075] In one embodiment, the RTT can be used as a template sequence by a reverse transcriptase for synthesizing a corresponding single-stranded DNA flap having a 5' end, wherein the DNA flap is complementary to the strand of the endogenous target DNA sequence adjacent to the nicking site, and wherein the single-stranded DNA flap comprises nucleotide modifications encoded by the editing template.

[0076] In one embodiment, the single-stranded DNA flap displaces endogenous single-stranded DNA having a 3' end in the nicked target DNA sequence.

[0077] In one embodiment, cellular repair of the single-stranded DNA flap results in the installation of the nucleotide modification, thereby forming the desired product.

[0078] In one embodiment, the nucleotide modification is a single nucleotide substitution, deletion, insertion, or large-scale deletion, insertion, or replacement.

[0079] In one embodiment, the result of the gene editing is a single nucleotide substitution, deletion, insertion, or large fragment deletion, insertion, or replacement, etc.

[0080] In one embodiment, the nicking site of the complex is located in the protospacer (or spacer) of the PAM chain; in one embodiment, the position of the nucleotide modification or gene editing is the protospacer (or spacer) and the phosphodiester bond between any bases within 5bp, 10bp, 20bp, 30bp, 40bp, 50bp, 60bp, 70bp, 80bp, 90bp, 100bp, 150bp, 200bp, 250bp, or 300bp upstream and / or downstream of the protospacer (or spacer).

[0081] In one embodiment, the position of the nucleotide modification or gene editing is -3 to 17 (the first base of the original spacer of the PAM chain is site 1), or -3 to 23 (the first base of the original spacer of the PAM chain is site 1).

[0082] In a preferred embodiment, the substitution of the single nucleotide is a conversion or replacement, preferably, the substitution is selected from the following: G to T substitution, G to A substitution, G to C substitution, T to G substitution, T to A substitution, T to C substitution, C to G substitution, C to T substitution, C to A substitution, A to T substitution, A to G substitution, or A to C substitution; preferably, the conversion is selected from the following: conversion of a G:C base pair to a T:A base pair, conversion of a G:C base pair to an A:T base pair, G The following example converts a T:A base pair to a C:G base pair, a T:A base pair to a G:C base pair, a T:A base pair to an A:T base pair, a T:A base pair to a C:G base pair, a C:G base pair to a G:C base pair, a C:G base pair to a T:A base pair, a C:G base pair to an A:T base pair, an A:T base pair to a T:A base pair, an A:T base pair to a G:C base pair, or an A:T base pair to a C:G base pair.

[0083] In a preferred embodiment, the nucleotide modification is a nucleotide deletion, and the length of the deletion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0084] In a preferred embodiment, the nucleotide modification is a nucleotide insertion, and preferably, the length of the insertion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 43 nucleotides, at least 44 nucleotides, at least 45 In some embodiments, the insertion sequence is at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides. In a preferred embodiment, the insertion sequence is a polypeptide encoding sequence.

[0085] In a preferred embodiment, the nucleotide modification is a nucleotide substitution, preferably, the substituted nucleotides are at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0086] Nucleic Acids

[0087] In another aspect, the present invention provides an isolated polynucleotide comprising:

[0088] (a) a polynucleotide sequence encoding an engineered fusion protein or complex of the present invention;

[0089] Alternatively, a polynucleotide complementary to the polynucleotide described in (a).

[0090] In one embodiment, the nucleotide sequence is codon optimized for expression in prokaryotes. In one embodiment, the nucleotide sequence is codon optimized for expression in eukaryotic cells.

[0091] In one embodiment, the cell is an animal cell, eg, a mammalian cell.

[0092] In one embodiment, the cell is a human cell.

[0093] In one embodiment, the cell is a plant cell, such as a cell from a cultivated plant (such as cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.

[0094] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0095] Primer editing guide RNA (PEgRNA)

[0096] On the other hand, the present invention provides a PEgRNA comprising a guide RNA (gRNA) and a nucleic acid extension arm, wherein the nucleic acid extension arm comprises a DNA binding template, a DNA binding template primer binding site sequence (PBS) and a reverse transcription template sequence (RTT); the gRNA includes a guide sequence and a backbone sequence; the guide sequence is capable of pairing with the target nucleic acid, and the backbone sequence is capable of interacting with the Cas protein.

[0097] The principle of guide editing technology is that the protein binding sequence of the guide RNA programmable by nucleic acid can interact with the DNA binding protein (Cas protein) in the fusion protein of the present invention, thereby forming a complex between the DNA binding protein (Cas protein) and the guide RNA.

[0098] In one embodiment, the reverse transcription template stores target site editing information, and the fusion protein is capable of targeting the target sequence when combined with the guide editing guide RNA.

[0099] In one embodiment, the target sequence comprises a target strand and a complementary non-target strand.

[0100] In one embodiment, the gRNA hybridizes to the complementary non-target strand (non-PAM strand) to form an RNA-DNA hybrid and an R-loop.

[0101] In one embodiment, the nucleic acid extension arm is located at the 3' or 5' end of the guide RNA, or at an intramolecular position in the guide RNA, and the nucleic acid extension arm is DNA or RNA.

[0102] In a preferred embodiment, the backbone sequence is shown as SEQ ID NO.4, and the guide sequence is shown as SEQ ID NO.14.

[0103] carrier

[0104] The present invention also provides a vector comprising the fusion protein, complex, isolated nucleic acid molecule or polynucleotide as described above; preferably, it further comprises a regulatory element operably linked thereto.

[0105] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.

[0106] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.

[0107] In some embodiments, the vector included in the system is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.

[0108] In one embodiment, the fusion protein and PEgRNA are located in the same vector.

[0109] In one embodiment, the fusion protein and PEgRNA are located in different vectors.

[0110] host cells

[0111] The present invention also relates to an in vitro, ex vivo or in vivo cell or cell line or their progeny, wherein the cell or cell line or their progeny comprises: the fusion protein, nucleic acid molecule, protein-nucleic acid complex, vector, and delivery composition of the present invention.

[0112] In certain embodiments, the cell is a prokaryotic cell.

[0113] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, cattle, sheep, pig, dog, monkey, rabbit, rodent (such as rat or mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of poultry (such as chicken), fish or crustacean (such as clams, shrimp). In certain embodiments, the cell is a plant cell, such as a cell or cultivated plant or food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice that a monocot or dicot has, such as an algae, tree or production plant, fruit or vegetable (for example, trees such as citrus trees, nut trees; Solanaceae, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0114] In certain embodiments, the cell is a stem cell or a stem cell line.

[0115] In certain cases, the host cells of the invention comprise genetic or genomic modifications that are not present in their wild-type form.

[0116] Delivery and delivery compositions

[0117] The fusion proteins, PEgRNAs, nucleic acid molecules, vectors, systems, complexes, and compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipofection, nucleofection, microinjection, sonoporation, gene guns, calcium phosphate-mediated transfection, cationic transfection, lipofection, dendritic transfection, heat shock transfection, nucleofection, magnetofection, lipofection, puncture transfection, optical transfection, agent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial virions, and the like.

[0118] Therefore, in another aspect, the present invention provides a delivery composition comprising a delivery vector and one or more selected from the following: the fusion protein, PEgRNA, nucleic acid molecule, vector, system, complex and composition of the present invention.

[0119] In one embodiment, the delivery vehicle is a particle.

[0120] In one embodiment, the delivery vehicle is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses or adeno-associated viruses).

[0121] Gene Editing Methods and Applications

[0122] The fusion proteins, nucleic acids, complexes, CIRSPR / Cas systems, vector systems, delivery compositions, or host cells of the present invention can be used for any one or more of the following purposes: targeting and / or editing target nucleic acids; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids. In other embodiments, they can also be used to prepare reagents or kits for any one or more of the above purposes.

[0123] The present invention also provides the use of the above-mentioned fusion protein, nucleic acid, complex, CIRSPR / Cas system, vector system, delivery composition or host cell in gene editing, gene targeting or gene cleavage; or, use in the preparation of reagents or kits for gene editing, gene targeting or gene cleavage.

[0124] In one embodiment, the gene editing, gene targeting or gene cleavage is performed inside and / or outside the cell.

[0125] The present invention also provides a method for gene editing, gene targeting or gene cleavage, comprising contacting a target nucleic acid with the above-mentioned fusion protein, nucleic acid, complex, CIRSPR / Cas system, vector system, delivery composition or host cell. In one embodiment, the method is to edit the target nucleic acid, target the target nucleic acid or cleave the target nucleic acid inside or outside the cell.

[0126] The gene editing or editing of target nucleic acid includes modifying genes, knocking out genes, inserting genes, changing the expression of gene products, repairing mutations, and / or inserting polynucleotides, gene mutations.

[0127] The editing can be performed in prokaryotic cells and / or eukaryotic cells.

[0128] In certain embodiments, the cell is a prokaryotic cell.

[0129] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as a cell of a non-human primate, cattle, sheep, pig, dog, monkey, rabbit, rodent (such as rat or mouse). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as a cell of poultry (such as chicken), fish or crustacean (such as clams, shrimp). In certain embodiments, the cell is a plant cell, such as a cell or cultivated plant or food crop such as cassava, corn, sorghum, soybean, wheat, oat or rice that a monocot or dicot has, such as an algae, tree or production plant, fruit or vegetable (for example, trees such as citrus trees, nut trees; Solanaceae, cotton, tobacco, tomato, grape, coffee, cocoa, etc.).

[0130] In another aspect, the present invention provides the use of the above-mentioned fusion protein, nucleic acid, composition, CIRSPR / Cas system, vector system, delivery composition, or host cell in preparing a preparation or kit, wherein the preparation or kit is used for:

[0131] (i) gene or genome editing;

[0132] (ii) target nucleic acid detection and / or diagnosis;

[0133] (iii) editing a target sequence in a target locus to modify an organism;

[0134] (iv) treatment of disease;

[0135] (v) targeting target genes;

[0136] (vi) Cutting the target gene.

[0137] Preferably, the above-mentioned gene or genome editing is performed inside or outside the cell.

[0138] Preferably, the target nucleic acid detection and / or diagnosis is performed in vitro.

[0139] Preferably, the treatment of the disease is the treatment of a condition caused by a defect in the target sequence in the target locus.

[0140] Method for specifically modifying target nucleic acid

[0141] On the other hand, the present invention also provides a method for specifically modifying a target nucleic acid, the method comprising: contacting the target nucleic acid with the above-mentioned fusion protein, nucleic acid, composition, CIRSPR / Cas system, vector system or delivery composition.

[0142] The specific modification can occur in vivo or in vitro.

[0143] The specific modification can occur inside or outside the cell.

[0144] In some cases, the cell is selected from a prokaryotic cell or a eukaryotic cell, eg, an animal cell, a plant cell, or a microbial cell.

[0145] In one embodiment, the nucleotide modification is a single nucleotide substitution, deletion, insertion, or large-scale deletion and insertion, etc.

[0146] In a preferred embodiment, the substitution of the single nucleotide is a conversion or replacement, preferably, the substitution is selected from the following: G to T substitution, G to A substitution, G to C substitution, T to G substitution, T to A substitution, T to C substitution, C to G substitution, C to T substitution, C to A substitution, A to T substitution, A to G substitution, or A to C substitution; preferably, the conversion is selected from the following: conversion of a G:C base pair to a T:A base pair, conversion of a G:C base pair to an A:T base pair, G The following example converts a T:A base pair to a C:G base pair, a T:A base pair to a G:C base pair, a T:A base pair to an A:T base pair, a T:A base pair to a C:G base pair, a C:G base pair to a G:C base pair, a C:G base pair to a T:A base pair, a C:G base pair to an A:T base pair, an A:T base pair to a T:A base pair, an A:T base pair to a G:C base pair, or an A:T base pair to a C:G base pair.

[0147] In a preferred embodiment, the nucleotide modification is a nucleotide deletion, and the length of the deletion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0148] In a preferred embodiment, the nucleotide modification is a nucleotide insertion, and preferably, the length of the insertion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 43 nucleotides, at least 44 nucleotides, at least 45 At least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0149] In a preferred embodiment, the insertion is a sequence encoding a polypeptide.

[0150] On the other hand, the present invention also provides a method for introducing desired nucleotide changes in a double-stranded DNA sequence, the method comprising: contacting the double-stranded DNA sequence with a complex comprising the above-mentioned fusion protein and PEgRNA, wherein the fusion protein comprises a Cas protein and a reverse transcriptase, wherein the PEgRNA comprises a DNA synthesis template (i.e., a reverse transcription template sequence, RTT) containing the desired nucleotide change and a primer binding site; thereby generating a nick in the double-stranded DNA sequence, thereby generating a free single-stranded DNA having a 3' end of a targeted chain; thereby hybridizing the 3' end of the free single-stranded DNA with the primer binding site, thereby activating the reverse transcriptase to polymerize the DNA chain at the 3' end hybridized with the primer binding site using RTT as a template, thereby generating a single-stranded DNA comprising the desired nucleotide change and complementary to the RTT template; thereby utilizing the cell's endogenous DNA repair mechanism to replace the endogenous DNA chain adjacent to the cleavage site with the single-stranded DNA, thereby installing the desired nucleotide change in the double-stranded DNA sequence.

[0151] In one embodiment, the desired nucleotide change is installed in an editing window between -3 and +17 or between -3 and 23 of the protospacer (the first base of the protospacer of the PAM strand is position 1). In other embodiments, the desired nucleotide change is installed within the protospacer and in an editing window between -5 and +5, or between -10 and +10, or between -20 and +20, or between -30 and +30, or between -40 and +40, or between -50 and +50, or between -60 and +60, or between -70 and +70, or between -80 and +80, or between -90 and +90, or between -100 and +100, or between -200 and +200 of the protospacer.

[0152] In one embodiment, the desired nucleotide change is installed in an editing window of about -7 to +23 of the PAM sequence. In other embodiments, the desired nucleotide change is installed in an editing window of about -5 to +5 of the nicking site, or about -10 to +10 of the nicking site, or about -20 to +20 of the nicking site, or about -30 to +30 of the nicking site, or about -40 to +40 of the nicking site, or about -50 to +50 of the nicking site, or about -60 to +60 of the nicking site, or about -70 to +70 of the nicking site, or about -80 to +80 of the nicking site, or about -90 to +90 of the nicking site, or about -100 to +100 of the nicking site, or about -200 to +200 of the nicking site.

[0153] In one embodiment, the desired nucleotide change is a single nucleotide substitution, deletion, insertion, or large-scale deletion, insertion, or replacement.

[0154] In a preferred embodiment, the substitution of the single nucleotide is a conversion or replacement, preferably, the substitution is selected from the following: G to T substitution, G to A substitution, G to C substitution, T to G substitution, T to A substitution, T to C substitution, C to G substitution, C to T substitution, C to A substitution, A to T substitution, A to G substitution, or A to C substitution; preferably, the conversion is selected from the following: conversion of a G:C base pair to a T:A base pair, conversion of a G:C base pair to an A:T base pair, G The following example converts a T:A base pair to a C:G base pair, a T:A base pair to a G:C base pair, a T:A base pair to an A:T base pair, a T:A base pair to a C:G base pair, a C:G base pair to a G:C base pair, a C:G base pair to a T:A base pair, a C:G base pair to an A:T base pair, an A:T base pair to a T:A base pair, an A:T base pair to a G:C base pair, or an A:T base pair to a C:G base pair.

[0155] In a preferred embodiment, the desired nucleotide change is a nucleotide deletion, and the length of the deletion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 43 nucleotides, at least 44 nucleotides, at least 45 At least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0156] In a preferred embodiment, the desired nucleotide change is a nucleotide insertion, and preferably, the length of the insertion is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides. , at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, or at least 100 nucleotides.

[0157] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. DETAILED DESCRIPTION

[0158] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0159] As used herein, the "polynucleotide", "nucleotide sequence", "nucleic acid sequence", "nucleic acid molecule" and "nucleic acid" are used interchangeably and include DNA, RNA or hybrids thereof, which may be double-stranded or single-stranded.

[0160] The terms "homology" or "identity" are used to refer to the matching of sequences between two polypeptides or between two nucleic acids. When a position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of the two DNA molecules is occupied by adenine, or a position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. Between two sequences. Typically, comparison is made when the two sequences are aligned for maximum identity. Alignment methods are conventional methods known to those skilled in the art, such as the BLAST algorithm.

[0161] The term "complementary" or "complementary pairing" refers to the specific matching relationship between nucleic acid molecules based on the principle of complementary base pairing. For example, in DNA molecules, there are four bases: adenine (A), thymine (T), cytosine (C), and guanine (G). A and T have a complementary pairing relationship, and C and G have a complementary pairing relationship. Therefore, if a base on one strand of a DNA molecule is A, the corresponding base on its complementary strand is T; if a base on one strand of a DNA molecule is C, the corresponding base on its complementary strand is G.

[0162] The term "genetic engineering" refers to the technology of artificially modifying and utilizing the nucleotides that control the genetic information of organisms to obtain new genetic traits, new varieties, or new products. This includes all genetic modification technologies disclosed in the art, such as gene mutagenesis, transgenic technology, and gene editing. Genetic mutagenesis methods include, but are not limited to, physical mutagenesis (such as ultraviolet mutagenesis), chemical mutagenesis (such as acridine dyes), and biological mutagenesis (such as viral and bacteriophage mutagenesis).

[0163] The term "encode" refers to the inherent property of a particular nucleotide sequence in a polynucleotide, such as a gene, cDNA, or mRNA, to serve as a template for the synthesis of other polymers and macromolecules in a biological process having a defined nucleotide sequence (i.e., rRNA, tRNA, and mRNA) or a defined amino acid sequence and the resulting biological properties. Thus, a gene encodes a protein if transcription and translation of the mRNA corresponding to that gene produces the protein in a cell or other biological system.

[0164] The term "amino acid" refers to a carboxylic acid containing an amino group. Various proteins in living organisms are composed of 20 basic amino acids.

[0165] The terms "protein," "polypeptide," and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, including polymers in which one or more amino acid residues is a chemical analog of a naturally occurring amino acid residue. The proteins and polypeptides of the present invention can be produced recombinantly or by chemical synthesis.

[0166] In the present invention, amino acid residues can be represented by single letters or three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamine (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine ​​(Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), arginine (Arg, R).

[0167] The term "regulatory element," also known as a "regulatory element," as used herein, is intended to include promoters, terminator sequences, leader sequences, polyadenylation sequences, signal peptide coding regions, marker genes, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences), which are described in detail in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, CA (1990). In some cases, regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or special cell types (e.g., lymphocytes). In some cases, regulatory elements can also direct expression in a temporally dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell-type-specific. In some cases, the term "regulatory element" encompasses enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I ((Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequences between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0168] The term "promoter" has a meaning well known to those skilled in the art and refers to a non-coding nucleotide sequence located upstream of a gene that can initiate expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in a cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell essentially only when the cell is a cell of the tissue type corresponding to the promoter.

[0169] The term "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" proteins for import into the cell nucleus via nuclear transport. That is, proteins with an NLS are transported to the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the protein surface. Exemplary NLSs include, but are not limited to, NLSs from the SV40 large T antigen, EGL-13, c-Myc, and TUS proteins.

[0170] The term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0171] The nucleic acid sequence, nucleic acid construct or expression vector of the present invention can be introduced into the host cell by a variety of techniques, including transformation, transfection, transduction, viral infection, gene gun or Ti-plasmid-mediated gene delivery, as well as calcium phosphate transfection, DEAE-dextran-mediated transfection, lipofection or electroporation.

[0172] CRISPR system

[0173] As used herein, the terms "clustered regularly interspaced short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meaning commonly understood by those skilled in the art, which generally includes transcripts or other elements associated with the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of directing the activity of the Cas genes.

[0174] Cas proteins

[0175] Cas protein, or CRISPR-related protein refers to a nuclease suitable for the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) system. Preferably, the Cas protein is a CRISPR enzyme, and its types include but are not limited to: Cas9 protein, Cas12 protein, Cas13 protein, Cas14 protein, Csm1 protein, FDK1 protein. The Cas protein may have different structures depending on its source, such as SpCas9 from Streptococcus pyogenes and SaCas9 from Staphylococcus aureus; it may also be classified according to structural features (such as domains), such as the Cas12 family including Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12i, etc. The Cas protein may have double-stranded or single-stranded or no cutting activity. The Cas protein of the present invention may be a wild type or a mutant thereof, and the mutation type of the mutant includes amino acid replacement, substitution or deletion, and the mutant may or may not change the enzymatic activity of the Cas protein. For example, nCas9 refers to a Cas9 (H840A) mutant, which has single-stranded nucleic acid cleavage activity. As known to those skilled in the art, a variety of Cas proteins with nucleic acid cleavage activity have been reported in the prior art. The known protein or its modified variant can achieve the function of the present invention, and is herein incorporated by reference into the scope of protection.

[0176] gRNA

[0177] As used herein, the term "gRNA" or "CRISPR RNA" refers to a guide RNA (or guide RNA, guide RNA) suitable for the CRISPR system, which includes a guide sequence (or spacer sequence) and a backbone region, wherein the backbone region can interact with the CRISPR protein (or Cas protein), thereby forming a complex between the Cas protein and the gRNA and guiding the complex to bind to the target nucleic acid; its guide sequence (or spacer sequence) is complementary to the target nucleic acid sequence.

[0178] Prime Editing Technology

[0179] Prime Editing (guide editing) technology, such as patents CN113891936A, CN113891937A, CN114127285A, or CN114729365A, refers to a method for gene editing using a nucleic acid programmable DNA binding protein (such as Cas9), a polymerase (such as reverse transcriptase), and a guide editing guide RNA (PEgRNA). The PEgRNA includes a guide RNA (composed of a guide sequence and a backbone sequence) and an extension arm, which includes a DNA binding template (including an editing template and an RTT of a homology arm) and a primer binding site (PBS). The principle of guide editing technology is that a nucleic acid programmable DNA binding protein (such as Cas9) cuts a strand (non-target strand) of the target nucleic acid, and the target nucleic acid strand that produces the cut interacts with the extension arm of the pegRNA, initiating polymerization, and a polymerase (such as reverse transcriptase) synthesizes ssDNA containing the target sequence. The DNA chain containing the target sequence can hybridize with the endogenous target nucleic acid chain due to the inclusion of the homology arm sequence, and replaces the original DNA chain, and the target sequence is ultimately inserted into the target nucleic acid.

[0180] The sequence involved in the present invention is as follows:

[0181] The main advantages of the present invention are:

[0182] The present invention provides a fusion protein that can be used for guide editing, which realizes the combination of V-type Cas protein Cas12i and reverse transcriptase for guide editing for the first time, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0183] Figure 1. Schematic diagram of PE vector construction. Figure 1a is a schematic diagram of vector SF01-PE, wherein G-quadruplex is a guanine tetraplex (shown in SEQ ID NO.8), tgcrRNA is the PEgRNA (shown in SEQ ID NO.9, or SEQ ID NO.10, or SEQ ID NO.11); U6 is the promoter of PEgRNA (SEQ ID NO.12), CMV is the fusion protein promoter (shown in SEQ ID NO.13), RT is the reverse transcriptase (SEQ ID NO.2), and SF01 is the Cas protein (SEQ ID NO.1); Figure 1b is a schematic diagram of vector SF01-Brex, wherein G-quadruplex is a guanine tetraplex (shown in SEQ ID NO.8), tgcrRNA is the PEgRNA (shown in SEQ ID NO.9, or SEQ ID NO.10, or SEQ ID NO.11), U6 is the promoter of PEgRNA (SEQ ID NO.12), CMV is the fusion protein promoter (SEQ ID NO.13), RT is a reverse transcriptase (SEQ ID NO.2), SF01 is a Cas protein (SEQ ID NO.1), and Brex is a Brex27 single-stranded binding protein (SEQ ID NO.5); Figure 1c is a schematic diagram of the vector SF01-LinkerBrex, wherein G-quadruplex is a guanine quadruplex (shown in SEQ ID NO.8), tgcrRNA is the PEgRNA (shown in SEQ ID NO.9, or shown in SEQ ID NO.10, or shown in SEQ ID NO.11), U6 is the promoter of PEgRNA (SEQ ID NO.12), CMV is a fusion protein promoter (shown in SEQ ID NO.13), RT is a reverse transcriptase (SEQ ID NO.2), SF01 is a Cas protein (SEQ ID NO.1), Brex is a Brex27 single-stranded binding protein (SEQ ID NO.5), and Liker is a linker XTEN (SEQ ID NO.3) connecting the single-stranded binding protein and the Cas protein.

[0184] Figure 2. Precise replacement editing efficiency of the PE vector system.

[0185] Figure 3. Schematic diagram of the precise editing principle based on the PE technology of the present invention.

[0186] Figure 4. Schematic diagram of PE vector construction.

[0187] Figure 5. Precise editing efficiency and indel efficiency of PE vectors.

[0188] Figure 6. Schematic diagram of the editing position of the PE vector.

[0189] Figure 7. Efficiency of base substitution at different positions of PE vector.

[0190] Figure 8. Length range of substituted bases in PE vectors.

[0191] Figure 9. Position and length results of base deletion using PE vector.

[0192] Implementation Method

[0193] The following examples are only used to describe the present invention, but are not intended to limit the present invention. Unless otherwise specified, the experiment and method described in the embodiment are carried out substantially according to conventional methods well known in the art and described in various references. For example, the conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, " molecular cloning: laboratory manual " (MOLECULAR CLONING:A LABORATORY MANUAL), 2nd edition (1989); " current protocols in molecular biology " (FM Ausubel et al. edit, (1987)); " methods in enzymes " (METHODS IN ENZYMOLOGY) series (Academic Publishing Company): " PCR 2: practical methods " (PCR 2:A PRACTICAL APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), Harlow and Lane, eds. (1988) ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).

[0194] In addition, if specific conditions are not specified in the examples, the experiments were performed under conventional conditions or the conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments is not specified, they are all conventional products that can be obtained commercially. It is understood that the examples describe the present invention by way of example and are not intended to limit the scope of the present invention. All publications and other references mentioned herein are incorporated herein by reference in their entirety.

[0195] Example 1. Construction of PE carrier system

[0196] For the known Cas protein (Cas12f.4 in CN111757889B, referred to as Cas12i in this embodiment), the applicant predicted the key amino acid sites that may affect its biological function through bioinformatics, and mutated the amino acid sites to obtain a Cas mutant protein with improved editing activity, whose amino acid sequence is shown in SEQ ID NO.1. It is connected to the reverse transcriptase to form a fusion protein. The amino acid sequence of the reverse transcriptase is shown in SEQ ID NO.2. Three targets are selected on the DNMT1 gene, and the PEgRNA sequences are designed as shown in Table 2 below (the bold lowercase part is the backbone sequence of the gRNA, the italic part is the guide sequence of the gRNA, the underline is the RTT sequence, the bold uppercase nucleotides are the nucleotides replaced by the PE precise editing, the uppercase part is the PBS sequence, and the rest is the homology arm sequence of the nucleic acid extension arm).

[0197] Table 2. PEgRNA sequence information

[0198] The pcDNA3.3 vector was modified to carry the EGFP fluorescent protein and the PuroR resistance gene. The SV40 NLS-Cas-XX fusion protein was inserted via the restriction sites XbaI and PstI. The U6 promoter and PE gRNA sequence were inserted via the restriction site Mfe1. Expression of the SV40 NLS-Cas-XX-NLS-GFP fusion protein is driven by a CMV promoter. The Cas-XX-NLS protein and GFP protein are linked by a T2A linker. Expression of the puromycin resistance gene is driven by the EF-1α promoter.

[0199] The results of PE vector construction are shown in Figure 1 , where Figure 1a shows the case without the addition of Brex27 single-chain binding protein, Figure 1b shows the case where the Cas protein is directly linked to the Brex27 single-chain binding protein, and Figure 1b shows the case where the Cas protein is linked to the Brex27 single-chain binding protein via a linker.

[0200] As shown in Figure 3, the complex containing the fusion protein and PEgRNA contacts the double-stranded DNA, causing a cut in the double-stranded DNA sequence, exposing free single-stranded DNA at the 3' end of the targeted chain; the free single-stranded DNA at the 3' end can hybridize with the primer binding site, thereby activating the reverse transcriptase to polymerize the DNA chain using RTT as a template at the 3' end hybridized with the primer binding site, thereby generating a single-stranded DNA containing a nucleotide replacement fragment and complementary to the RTT template; the cell's endogenous DNA repair mechanism is used to replace the endogenous DNA chain adjacent to the cleavage site with single-stranded DNA, thereby producing precise editing in the double-stranded DNA sequence.

[0201] Example 2: Verification of Editing Activity of the PE Vector System

[0202] The editing activity of the PE vector system was verified in 293T cells. Each constructed vector was transferred into 293T cells. Plating: 293T cells were plated to a confluency of 70-80%, and the number of cells seeded in a 12-well plate was 8*10^4 cells / well. Transfection: 24 hours after plating, transfection was performed. 6.25μl Hieff Transfection was added to 100μl Opti-MEM. TM Liposome nucleic acid transfection reagent, mix well; add 2.5ug plasmid to 100μl opti-MEM and mix well. TM Mix the diluted plasmid with the liposome transfection reagent and incubate at room temperature for 20 minutes. Add the incubated mixture to the culture medium containing cells for transfection. 48 hours after transfection, digest the cells with trypsin-EDTA (0.05%) and sort the cells with GFP signal using flow cytometry (FACS).

[0203] DNA extraction, PCR amplification of the region surrounding the edited site, and hiTOM sequencing: Cells were collected after trypsin digestion, and genomic DNA was extracted using a cell / tissue genomic DNA extraction kit (Biotech). The genomic DNA was amplified from the region surrounding the target site. PCR products were subjected to hiTOM sequencing. Sequencing data were analyzed to count the number and proportion of sequences within 15 nt upstream and 10 nt downstream of the target site, and to calculate the probability of precise substitutions and large deletions in the sequence.

[0204] Design hiTOM sequencing primers for each target site:

[0205] The detection results of each target in 293T cells are shown in Figure 2. TGTC-ACAG is target 1. When precise editing occurs, TGTC at target 1 will be replaced by ACAG; ACCT-TGCA is target 2. When precise editing occurs, ACCT at target 2 will be replaced by TGCA; CGG-GCA is target 3. When precise editing occurs, CGG at target 3 will be replaced by GCA.

[0206] As shown in Figure 2, at different target sites, the constructed PE editing system has obvious editing activity. At target sites 1-3, the precise replacement editing efficiency of the SF01-PE vector system is significantly higher than that of the SF01-Brex vector system and the SF01-LinkerBrex vector system.

[0207] Example 3: Construction and activity verification of PE carrier system

[0208] The Cas mutant protein in Example 1 (Cas-SF01, whose amino acid sequence is shown in SEQ ID NO.1) was connected to the reverse transcriptase M-MLV to form a fusion protein (the amino acid sequence of the reverse transcriptase is shown in SEQ ID NO.2), the target 1 in Example 1 was selected on the DNMT1 gene, and the PEgRNA1 sequence in Example 1 was used for the experiment. The experimental method is the same as that in Example 1. Four PE vectors are constructed. The construction results are shown in Figure 4. Figures 4a and 4c are PE vectors without Ubv (ubiquitin-like modifier protein); the N-terminus of the Cas protein in Figure 4a is directly connected to the reverse transcriptase (MMLV-SF01 in Figure 5); the N-terminus of the Cas protein in Figure 4b is directly connected to the reverse transcriptase, and the C-terminus of the Cas protein is directly connected to Ubv (MMLV-SF01-Ubv in Figure 5); the C-terminus of the Cas protein in Figure 4c is directly connected to the reverse transcriptase (SF01-MMLV in Figure 5); the C-terminus of the Cas protein in Figure 4d is directly connected to the reverse transcriptase, and the C-terminus of the reverse transcriptase is directly connected to Ubv (SF01-MMLV-Ubv in Figure 5). In Figure 4, SF01 refers to the Cas mutant Cas-SF01, RT is the reverse transcriptase M-MLV, and Ubv is a ubiquitin-like modifier protein (shown in SEQ ID No. 15). The other elements are the same as in Figure 1. Among them, Ubv is a ubiquitin-like modifier protein that mainly acts on the 53BP1 protein, inhibiting its recruitment to the DNA double-strand break site, thereby promoting precise DNA repair.

[0209] The editing activity of the above-mentioned PE vector was verified using the method described in Example 2. The results are shown in Figure 5. For the PE vector without Ubv, the PE vector in which the N-terminus of the Cas-SF01 protein is connected to the reverse transcriptase MMLV (the vector shown in Figure 4a) has a better precision repair effect than the PE vector in which the C-terminus of the Cas-SF01 protein is connected to the reverse transcriptase MMLV (the vector shown in Figure 4c); fusing Ubv to the PE vector (the vectors shown in Figures 4b and 4d) can significantly improve the precision repair efficiency. In particular, the PE vector in which the N-terminus of the Cas protein is connected to the reverse transcriptase and the C-terminus of the Cas protein is connected to Ubv (the vector shown in Figure 4b) has the best precision repair effect. In Figure 5, SF01 refers to the Cas mutant protein Cas-SF01, MMLV is the reverse transcriptase, and Ubv is the ubiquitin-like modification protein.

[0210] Example 4: Editing Scope of the PE Vector System

[0211] Using the PE vector constructed in Example 1, as shown in Figure 1a, we selected target site 1 in Example 1 on the DNMT1 gene to verify the editing range of the PE vector. The experimental method was the same as in Example 1-2. As shown in Figure 6, the spacer was set to start at position 1, and the PAM position was between positions -3 and -1.

[0212] The results of the base substitution range are shown in Figure 7. The editing range of PE (MMLV-SF01) is from position -3 to 17, with the highest editing efficiency when the bases are replaced at positions 10-13, with an average efficiency exceeding 30%, significantly higher than the control group. In Figure 7, the horizontal axes P-3, 1, P2, 5, P6, 9, P10, 13, and P14, 17 refer to the start and end positions of the four replaced bases at the target site when editing occurs; MMLV-SF01 is the control group, with the replaced bases at positions 12-15.

[0213] The results of the replacement base length are shown in Figure 8. The range of base replacement lengths for PE (MMLV-SF01) can reach up to 24bp. Among them, the precision editing efficiency was the highest with a base replacement length of 8bp (P10-17 group), exceeding 30%, significantly higher than the control group. After the base replacement length exceeded 8bp, the replacement efficiency showed a downward trend as the replacement length increased. In Figure 8, the horizontal axes P10,17 (replacement base length of 8bp), P8,19 (replacement base length of 12bp), P6,21 (replacement base length of 16bp), P4,23 (replacement base length of 20bp), and P2,25 (replacement base length of 24bp) refer to the start and end positions of the replacement base at the target site when editing occurs; MMLV-SF01 is the control group, and its replacement base positions are the four bases 12-15.

[0214] The editing results of PE (MMLV-SF01) were designed to be base deletions. The range and length of the deleted bases are shown in Figure 9. The range of base deletions for PE (MMLV-SF01) is from -7 to 23, and the length of the deleted bases can reach 30 bp. Among them, the highest precision editing efficiency was achieved when the deleted bases were located from -1 to 23 and the deletion length was 24 bp (P-1,23 group). When the length of the deleted bases was ≤24 bp, the longer the length, the higher the precision editing efficiency. In Figure 9, the horizontal axes P15, 18 (deleted base length is 4 bp), P9, 16 (deleted base length is 8 bp), P10, 21 (deleted base length is 12 bp), P8, 23 (deleted base length is 16 bp), P-1, 23 (deleted base length is 24 bp), and P-7, 23 (deleted base length is 30 bp) refer to the start and end positions of the deleted bases at the target site when editing occurs; MMLV-SF01 is the control group, and its editing result is the replacement of the four bases at positions 12-15.

[0215] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.

Claims

1. A fusion protein comprising a Cas protein and a reverse transcriptase, characterized in that: The Cas protein is a V-type Cas protein or a variant thereof, preferably, the Cas protein is Cas12i, more preferably, the amino acid sequence of the Cas protein is as shown in SEQ ID NO.1, the reverse transcriptase is selected from M-MLV reverse transcriptase, AMV reverse transcriptase and others, preferably, the reverse transcriptase is M-MLV reverse transcriptase, preferably, the amino acid sequence of the reverse transcriptase is as shown in SEQ ID NO.2, preferably, the fusion protein comprises a linker connecting the Cas protein and the reverse transcriptase.

2. The fusion protein according to claim 1, characterized in that One end of the Cas protein is connected to the reverse transcriptase, and the other end can also be connected to the single-chain binding protein or ubiquitin-like modified protein; or, one end of the reverse transcriptase is connected to the Cas protein, and the other end can also be connected to the single-chain binding protein or ubiquitin-like modified protein; Preferably, the single-chain binding protein is selected from Brex27, EcRecA, BsRecA or T4SSB, and the amino acid sequence of the ubiquitin-like modified protein is shown in SEQ ID No. 15; more preferably, the single-chain binding protein is Brex27; Preferably, the reverse transcriptase is connected to the N-terminus or C-terminus of the Cas protein. Preferably, the reverse transcriptase is connected to the N-terminus of the Cas protein; Preferably, the connection may be a direct connection or a connection via a connector.

3. A complex for guide editing, characterized in that The compound comprises: (i) a protein component selected from the group consisting of: a fusion protein according to any one of claims 1 to 2; (ii) a nucleic acid component, which is a guide editing guide RNA (PEgRNA), wherein the PEgRNA comprises a guide RNA (gRNA) and a nucleic acid extension arm, wherein the nucleic acid extension arm comprises a primer binding site sequence (PBS) and a reverse transcription template sequence (RTT); preferably, the nucleic acid extension arm may further comprise a homology arm sequence, and the gRNA comprises a guide sequence and a backbone sequence; The protein component and the nucleic acid component are combined with each other to form a complex.

4. An isolated polynucleotide, characterized in that The polynucleotide is a polynucleotide sequence encoding the fusion protein according to any one of claims 1 to 2, or a polynucleotide sequence encoding the complex according to claim 3.

5. A carrier, characterized in that The vector comprises the polynucleotide according to claim 4 and a regulatory element operably linked thereto.

6. An engineered host cell, characterized in that The host cell comprises the fusion protein according to any one of claims 1-2, or the complex according to claim 3, or the polynucleotide according to claim 4, or the vector according to claim 5.

7. Use of the fusion protein of any one of claims 1-2, or the complex of claim 3, or the polynucleotide of claim 4, or the vector of claim 5, or the host cell of claim 6 in gene editing, gene targeting or gene cleavage; or, use in the preparation of a reagent or kit for gene editing, gene targeting or gene cleavage.

8. Use of the fusion protein according to any one of claims 1 to 2, or the complex according to claim 3, or the polynucleotide according to claim 4, or the vector according to claim 5, or the host cell according to claim 6 in preparing a preparation or a kit, wherein the preparation or the kit is used for: (i) gene or genome editing; (ii) target nucleic acid detection and / or diagnosis; (iii) editing a target sequence in a target locus to modify an organism; (iv) treatment of disease; (v) targeting target genes; (vi) Cutting the target gene.

9. A method for gene editing, gene targeting or gene cleavage, the method comprising: The target nucleic acid sequence is contacted with the fusion protein according to any one of claims 1 to 2, or the complex according to claim 3, or the polynucleotide according to claim 4, or the vector according to claim 5, or the host cell according to claim 6.

10. A method for introducing a desired nucleotide change in a double-stranded DNA sequence, characterized in that: The method comprises: contacting the double-stranded DNA sequence with a complex comprising the fusion protein of claim 1 or 2 and PEgRNA, wherein the fusion protein comprises a Cas protein and a reverse transcriptase, and the PEgRNA comprises a DNA synthesis template (reverse transcription template sequence, RTT) containing the desired nucleotide change and a primer binding site; thereby generating a nick in the double-stranded DNA sequence, thereby generating free single-stranded DNA having a 3' end of the targeted chain; thereby hybridizing the 3' end of the free single-stranded DNA with the primer binding site, thereby activating the reverse transcriptase to polymerize the DNA chain at the 3' end hybridized with the primer binding site using RTT as a template, thereby generating single-stranded DNA comprising the desired nucleotide change and complementary to the RTT sequence; thereby utilizing the cell's endogenous DNA repair mechanism to replace the endogenous DNA chain adjacent to the cleavage site with the single-stranded DNA, thereby installing the desired nucleotide change in the double-stranded DNA sequence.

Citation Information

Patent Citations

  • Novel CRISPR / Cas12f enzymes and systems

    CN113106081A

  • Engineered Cas12i nuclease as well as effect protein and application thereof

    CN113151215A

  • Methods and compositions for editing nucleotide sequences

    CN113891936A

  • Methods and compositions for editing nucleotide sequences

    CN114127285A

  • Methods and compositions for editing nucleotide sequences

    CN114729365A