Reverse prime editing system

By using a reverse-guided editing system, precise small-fragment DNA editing can be performed upstream of the target site using targeted strand nickase and circular RNA. This solves the problem that traditional guided editors can only edit downstream, enabling wider application and more efficient editing results.

WO2026045988A1PCT designated stage Publication Date: 2026-03-05INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2025/115543
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2025-08-19
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Traditional guided editors can only perform precise editing downstream of the target sequence cleavage site, and cannot achieve the desired editing upstream of the target site, which limits their application scope and potential.

Method used

Develop a reverse-guided editing system that utilizes target strand cleavage enzymes such as D10A-Cas9 and double-stranded endonucleases such as WT-Cas9, combined with pegRNA and circular RNA, to achieve precise editing of small DNA fragments upstream of the target site.

Benefits of technology

It has expanded the scope of editing, reduced editing byproducts, improved editing efficiency and safety, and broadened its application potential in areas such as disease treatment, crop improvement, and biological research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025115543_05032026_PF_FP_ABST
    Figure CN2025115543_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a reverse prime editing system, which is used for achieving site-specific and precise small-fragment DNA editing upstream of a target site. In the reverse prime editing system, a targeted strand nickase (such as D10A-Cas9) is used to cleave a nick on a targeted strand, thereby enabling precise small-fragment DNA base insertion, deletion and substitution upstream of the target site under the action of a reverse transcriptase.
Need to check novelty before this filing date? Find Prior Art

Description

Reverse-guided editing system Technical Field

[0001] This invention relates to the field of genetic engineering. Specifically, it relates to a reverse-guided editing system for achieving precise, site-specific small-fragment DNA editing upstream of a target site. More specifically, the reverse-guided editing system of this invention utilizes a target strand nickase (such as D10A-Cas9) or a double-stranded endonuclease (such as WT-Cas9) to create a nick on the target strand, and, under the action of reverse transcriptase, performs precise small-fragment DNA base insertion, deletion, and substitution upstream of the target site, as well as genetically modified but non-transgenic organisms and their offspring produced by said method. Background Technology

[0002] Precise and efficient genome editing has wide applications in disease treatment, crop improvement, and biological research. Guided editors are widely used in various life science fields because they can perform all single-base transformations of the four base types without causing double-strand breaks, fully covering the single-base editing function of base editors, and enabling precise insertion, deletion, and replacement of small DNA fragments. In guided editors, the reverse transcriptase M-MLV RT (Moloney-murine leukemia virus reverse transcriptase) uses the reverse transcriptase template (RTT) and the primer binding site (PBS) to generate the desired 3' end single-stranded DNA sequence along the 5'→3' direction. Therefore, guided editing can produce the desired precise edit downstream of the target sequence cleavage site. However, reverse transcriptase cannot perform reverse transcription in the 3'→5' direction using RNA as a template to produce the desired 5' end single-stranded DNA sequence. Therefore, traditional guided editors can only produce precise edits downstream of the target sequence cleavage site, not upstream, which greatly limits the application scope and potential of guided editing. Furthermore, Lee et al. (J.Lee,K.Lim,A.Kim,etal.Prime editing with genuine Cas9 nickases minimizes unwanted indels.Nat Commun.2023Mar30;14(1):1786.) found that the target strand nickase D10A-Cas9 has more precise single-strand cleavage activity than the non-target strand nickase H840A-Cas9. The traditional guide editor uses H840A for single-strand cleavage, which is one of the reasons why guide editors produce more byproducts. Summary of the Invention

[0003] The problem the invention aims to solve

[0004] Reverse-guided editing systems, which utilize targeted chain cleavage enzymes (such as D10A-Cas9) or double-stranded endonucleases (such as WT-Cas9 and various other Cas proteins) to perform editing upstream of the editing target site, thus opposing the direction of traditional guided editing, will be able to further increase the editing range and expand application potential.

[0005] Solution for solving the problem

[0006] This invention first utilizes pegRNA (prime editing guide RNA) and Cas9 to develop novel inverse prime editing systems (iPE), including nCas9-D10A-dependent iPE (nickase-dependent iPE, Figures 1a and 2a) and WTCas9-dependent nu-iPE (nuclease-dependent iPE, Figures 1b and 2b), for generating precise guided editing upstream of the target sequence cleavage site. Simultaneously, circular RNA with unwinding capability is used to replace pegRNA in expressing RTT-PBS sequences, and nCas9-D10A-dependent inverse prime editing systems (ciPE, nickase-dependent circular RNA-mediated iPE, Figures 1c and 2c) and WTCas9-dependent inverse prime editing systems (nu-ciPE, nuclease-dependent circular RNA-mediated iPE, Figures 1d and 2d) are developed. Finally, to further improve the editing efficiency of nickase-dependent reverse-guided editing systems, this invention develops a helicase-assisted ciPE reverse-guided editing system, hciPE (Figures 1e and 2e). iPE and the highly efficient ciPE reverse-guided editing system will enable more comprehensive and precise editing of all types of single-base transformations and small DNA fragments across a wider range of genomes in organisms such as animals, bacteria, dicotyledons, and monocotyledons.

[0007] This invention provides a reverse-guided editing system, comprising:

[0008] (i)a) An expression construct containing a CRSIPR nuclease and / or a nucleotide sequence encoding the CRSIPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase; or

[0009] b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRSIPR nuclease and / or a reverse transcriptase;

[0010] (ii) a guide RNA (pegRNA) targeting a target sequence in genomic DNA and / or an expression construct containing a nucleotide sequence encoding said guide RNA; and / or

[0011] (iii) An expression construct containing a circular RNA with a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or a nucleotide sequence encoding the circular RNA, and an expression construct containing a guide RNA (sgRNA) targeting a sequence in genomic DNA and / or a nucleotide sequence encoding the guide RNA;

[0012] Optionally, the CRSIPR nuclease is a target strand nickase or a double-stranded endonuclease.

[0013] In some embodiments, the CRSIPR nuclease is a Cas9 nuclease or a variant thereof.

[0014] In some embodiments, the Cas9 nuclease comprises the sequence shown in SEQ ID NO:14; and / or

[0015] The Cas9 nuclease variant is an nCas9 nuclease, which contains the sequence shown in SEQ ID NO:15.

[0016] In some embodiments, the reverse enzyme is M-MLV reverse transcriptase;

[0017] Preferably, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted, and it contains the sequence shown in SEQ ID NO:16.

[0018] In some embodiments, the reverse transcriptase and / or the guide editing fusion protein may also comprise one or more RNA aptamer-binding protein sequences.

[0019] In some embodiments, the RNA aptamer-binding protein is an MCP protein;

[0020] Optionally, the MCP protein comprises the sequence shown in SEQ ID NO:17.

[0021] In some embodiments, the guide editing fusion protein comprises a sequence selected from any of SEQ ID NO:1-3,7-11.

[0022] In some embodiments, it also includes a guide RNA (nicking sgRNA) targeting a non-target strand and / or an expression construct containing a nucleotide sequence encoding the guide RNA targeting a non-target strand.

[0023] In some embodiments, the guide RNA comprises a backbone sequence as shown in SEQ ID NO:4; and / or

[0024] The guide RNA that targets the non-target strand contains a backbone sequence as shown in SEQ ID NO:4.

[0025] In some implementations, the reverse transcription template (RTT) sequence and the primer binding site (PBS) sequence are directly linked in the circular RNA;

[0026] Preferably, the reverse transcription template (RTT) sequence is located at the 5' end of the primer binding site (PBS) sequence.

[0027] In some embodiments, the circular RNA comprises at least one RNA aptamer sequence, and optionally, the aptamer comprises MS2.

[0028] In some embodiments, the primer binding site sequence is complementary to at least a portion of the target sequence;

[0029] Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the target strand caused by the nick, particularly to the nucleotide sequence at the 3' end of the 3' free single strand.

[0030] In some embodiments, the expression construct containing the nucleotide sequence encoding the circular RNA comprises the following coding sequence: 5'-first ribozyme-first cyclic arm-RTT-PBS-second cyclic arm-second ribozyme-3';

[0031] Optionally, after transcription into RNA within the cell, the first and second ribozymes can self-cleave to produce 5'-first cyclic arm-RTT-PBS-second cyclic arm-3', and the first and second cyclic arms can connect with each other to form a circular RNA containing RTT and PBS.

[0032] In some embodiments, the coding sequence of the first ribozyme includes the sequence shown in SEQ ID NO:22, the coding sequence of the second ribozyme includes the sequence shown in SEQ ID NO:23, the first cyclic arm includes the nucleotide sequence shown in SEQ ID NO:24, and the second cyclic arm includes the nucleotide sequence shown in SEQ ID NO:25.

[0033] In some implementations, the nucleotide sequence encoding the circular RNA is expressed by a U6 promoter.

[0034] In some embodiments, the reverse-guided editing system further includes the MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor;

[0035] Optionally, the MLH1dn protein factor comprises the sequence shown in SEQ ID NO:26.

[0036] In some embodiments, the reverse-guided editing system further includes a helicase and / or an expression construct containing a nucleotide sequence encoding the helicase;

[0037] Optionally, the helicase comprises the sequence shown in SEQ ID NO:27;

[0038] Preferably, the helicase and reverse transcriptase form a fusion protein; more preferably, the fusion protein comprises the sequence shown in SEQ ID NO:12.

[0039] This invention provides the application of the reverse-guided editing system described above in either (A) or (B) below:

[0040] (A) Editing of the genome sequence of an organism or its cells;

[0041] (B) Products that are prepared by editing the genome sequence of an organism or biological cell.

[0042] The present invention provides a kit comprising the reverse bootstrapping editing system as described above.

[0043] The present invention provides a method for producing genetically modified cells, comprising introducing a reverse-guided editing system as described above into at least one of the cells, thereby resulting in modification of the genome sequence of the at least one cell.

[0044] In some embodiments, the cells are derived from microorganisms, animals, or plants;

[0045] The microorganisms include bacteria and fungi;

[0046] The animals include mammals and poultry;

[0047] The plants include monocotyledons and dicotyledons;

[0048] Optionally, the mammals include humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats;

[0049] Optionally, the poultry includes chickens, ducks, and geese;

[0050] Optionally, the monocotyledonous plants include rice, corn, wheat, sorghum, and barley, and the dicotyledonous plants include soybean, peanut, Arabidopsis thaliana, cotton, and rapeseed.

[0051] The effects of the invention

[0052] This invention constructs a reverse guided editing system based on target strand nickases such as D10A-Cas9, further reducing editing byproducts and improving editing safety. Simultaneously, circular RNA is introduced for stable expression of RTT and PBS sequences, resulting in low off-target rates and significantly improved editing efficiency, enhancing the application potential of the reverse guided editing system in disease treatment, crop improvement, and biological research. Attached Figure Description

[0053] Figures 1a to 1e: Schematic diagrams of different reverse-guided editing systems.

[0054] Figure 2: Schematic diagram of different reverse-guided editing system carriers.

[0055] Figure 3: Reverse guided editing in human cells using the iPE system with pegRNA.

[0056] Figure 4: The nu-iPE system, which utilizes pegRNA for reverse guided editing, demonstrates more efficient reverse guided editing than iPE in the human cell reporter system.

[0057] Figure 5: The nu-iPE system, which utilizes pegRNA for reverse guided editing, demonstrates more efficient reverse guided editing than iPE in human cells.

[0058] Figure 6: The ciPE system, which utilizes circular RNA for reverse guided editing, produces more efficient reverse guided editing in human cells than iPE while generating fewer InDels byproducts.

[0059] Figure 7: The nu-ciPE system, which utilizes circular RNA for reverse guided editing, also produces highly efficient reverse guided editing in human cells, but with more InDels byproducts.

[0060] Figure 8: Helicase-assisted hciPE produces higher reverse guided editing efficiency in human cells than ciPE.

[0061] Figure 9: Comparison of the reverse bootstrap editor hciPE with other current editors.

[0062] Figure 10: The reverse bootstrap editor exhibits low off-target activity. Detailed Implementation

[0063] Various exemplary embodiments, features, and aspects of the present invention will be described in detail below. The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0064] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art should understand that the present invention can be practiced without certain specific details. In other instances, methods, means, apparatus, and steps well known to those skilled in the art have not been described in detail in order to highlight the spirit of the present invention.

[0065] Unless otherwise stated, all units used in this specification are international standard units, and all numerical values ​​and ranges appearing in this invention should be understood to include systematic errors that are unavoidable in industrial production.

[0066] In this specification, the word "may" has two meanings: to perform a certain process and not to perform a certain process.

[0067] In this specification, references to "some specific / preferred embodiments," "other specific / preferred embodiments," "implementation," etc., refer to specific elements (e.g., features, structures, properties, and / or characteristics) related to that embodiment, which are included in at least one of the embodiments described herein and may or may not be present in other embodiments. Furthermore, it should be understood that these elements may be combined in any suitable manner in various embodiments.

[0068] As used in this article, “containing,” “having,” or “including” includes “containing,” “mainly composed of,” “substantially composed of,” and “composed of”; “mainly composed of,” “substantially composed of,” and “composed of” are subordinate concepts of “containing,” “having,” or “including.”

[0069] While the disclosure supports the definition of the term "or" as merely a substitute and "and / or", the term "or" in the claims means "and / or" unless expressly stated as merely a substitute or as mutually exclusive. In this specification, the term "and / or", when used to connect two or more options, should be understood to mean any one or any two or more of the options.

[0070] In this specification, "optional" and "optionally" mean that the events or circumstances described below may or may not occur, and the description includes both cases where the events or circumstances occur and cases where the events or circumstances do not occur.

[0071] In this specification, the range of values ​​referred to as "value A to value B" refers to the range including the endpoint values ​​A and B.

[0072] In this specification, the term "polynucleotide" refers to a polymer composed of nucleotides. Polynucleotides can be in the form of individual fragments or as a component of a larger nucleotide sequence structure, derived from a nucleotide sequence isolated at least once in number or concentration, and capable of being recognized, manipulated, and recovered using standard molecular biology methods (e.g., using cloning vectors). This also includes an RNA sequence (i.e., A, T, G, C) when a nucleotide sequence is represented by a DNA sequence (i.e., A, U, G, C), where "U" replaces "T". In other words, "polynucleotide" refers to a polymer of nucleotides removed from other nucleotides (individual fragments or entire fragments), or it can be a component or part of a larger nucleotide structure, such as an expression vector or a polycistronic sequence. Polynucleotides include DNA, RNA, and cDNA sequences.

[0073] In this specification, the term "genome" as used herein encompasses not only chromosomal DNA present in the cell nucleus, but also organelle DNA present in subcellular components of the cell, such as mitochondria and plastids.

[0074] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein and refer to amino acid polymers of any length. The polymer may be linear or branched, may contain modified amino acids, and may be separated by non-amino acid segments. The term also includes amino acid polymers that have been modified (e.g., through disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with labeled components).

[0075] In this specification, the term "expression construct" can refer to a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (such as mRNA), for example, RNA transcribed in vitro. For use in mammals such as humans, the expression construct is preferably a viral vector, such as adeno-associated virus (AAV), lentivirus, or adenovirus vector, with AAV vectors being the most preferred.

[0076] Prime editing (PE) systems are genome editing strategies consisting of a prime editor. Currently, conventional PEs use a Cas9 nickase (H840A) fused with an engineered Moloney murine leukemia virus (M-MLV) reverse transcriptase (RT). The RT is programmed with the desired prime editing guide RNA (pegRNA). The pegRNA contains a reverse transcription RNA template (RTT) sequence comprising sgRNA, a primer binding site (PBS), and the reverse transcriptase sequence, used to generate the desired edit at the target site.

[0077] As used herein, the terms “guide RNA,” “directing RNA,” and “sgRNA” are generally used interchangeably. A sgRNA comprises a directing sequence and a double-strand-forming region (e.g., the double-strand-forming region of crRNA, which may also be referred to as a crRNA repeat sequence). A sgRNA is a polynucleotide that can specifically target a target sequence and can form a complex with a nucleic acid programmable nucleotide binding domain protein (e.g., Cas9).

[0078] In this specification, "downstream of the target cleavage site" refers to the downstream of the CRSIPR nuclease cleavage site. In this invention, the cleavage site of the Cas9 protein is usually located three nucleotides upstream of the PAM sequence.

[0079] I. Guided Editing System

[0080] This invention provides a reverse-guided editing system, the system comprising:

[0081] (i)a) An expression construct containing a CRSIPR nuclease and / or a nucleotide sequence encoding the CRSIPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase; or

[0082] b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRSIPR nuclease and / or a reverse transcriptase;

[0083] (ii) a guide RNA (pegRNA) targeting a target sequence in genomic DNA and / or an expression construct containing a nucleotide sequence encoding said guide RNA; and / or

[0084] (iii) An expression construct containing a circular RNA with a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or a nucleotide sequence encoding the circular RNA, and an expression construct containing a guide RNA (sgRNA) targeting a sequence in genomic DNA and / or a nucleotide sequence encoding the guide RNA.

[0085] In some alternative implementations, the CRSIPR nuclease is a target strand nickase or a double-stranded endonuclease.

[0086] In this specification, the "reverse-guided editing system" refers to a combination of components required for reverse transcription-based genome editing within cells. The individual components of this system, such as CRISPR nucleases, reverse transcriptases, guided editing fusion proteins, gRNAs, circular RNAs, or their expression constructs, can exist independently or in any combination as a composition. The individual components of the system can be packaged independently or together into viruses, virus-like particles, virions, liposomes, vesicles, exosomes, liposome nanoparticles (LNPs), etc.

[0087] In some alternative embodiments, the CRSIPR nuclease is a CRSIPR nickase, meaning it forms a nick only on one strand of the double-stranded nucleic acid molecule and does not completely cleave the double-stranded nucleic acid. In some embodiments, the CRISPR nickase forms a nick only on the strand containing the target sequence in the double-stranded nucleic acid molecule (i.e., the target strand), and is also called a target strand nickase. The complementary strand of the strand containing the target sequence in the double-stranded nucleic acid molecule is called the non-target strand.

[0088] In this specification, the term "target sequence" refers to a sequence of approximately 20 nucleotides in length characterized by a PAM (pre-interstitial sequence adjacent motif) sequence flanking the genome at 5' or 3'. Typically, the PAM is necessary for the recognition of the target sequence by a complex formed by a CRISPR nuclease or a variant thereof with guide RNA.

[0089] In some specific implementations, the "target strand" or "target sequence" refers to a DNA strand in the double-stranded DNA that is complementary to the spacer sequence of the guide RNA. The nucleotide sequence of this strand is complementary to the spacer region of the guide RNA and can specifically bind to the guide RNA through base pairing, thereby guiding the Cas protein to target the region of the double-stranded DNA. This strand is another strand complementary to the DNA strand containing the PAM (protospacer adjacent motif) sequence.

[0090] In some embodiments, the Cas9 nuclease comprises the sequence shown in SEQ ID NO:14; the Cas9 nuclease variant is an nCas9 nuclease (named nCas9-D10A or D10A-Cas9 herein) which comprises the sequence shown in SEQ ID NO:15.

[0091] In some embodiments of the present invention, the reverse transcriptase is a viral reverse transcriptase, such as M-MLV reverse transcriptase derived from Moloney mouse leukemia virus. Further, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted. In some embodiments, the M-MLV reverse transcriptase with a mutated or deleted RNase H domain comprises the sequence shown in SEQ ID NO:16.

[0092] In some embodiments, the reverse transcriptase and / or the guided editing fusion protein of the present invention may further comprise one or more RNA aptamer-binding protein sequences. In some embodiments, the reverse transcriptase of the present invention further comprises one or more RNA aptamer-binding protein sequences. In some embodiments, the guided editing fusion protein of the present invention further comprises one or more RNA aptamer-binding protein sequences.

[0093] In this specification, "RNA aptamer-binding protein" refers to a protein capable of specifically binding to an "RNA aptamer" having a specific sequence and structure. Proteins containing one or more RNA aptamer-binding protein sequences can thus be recruited within the cell to nucleic acid molecules containing the corresponding RNA aptamer, and vice versa. Many such combinations of specifically binding "RNA aptamer-binding proteins" and "RNA aptamers" are known in the art.

[0094] In some embodiments, the RNA aptamer-binding protein is the MCP protein. Based on this, the present invention achieves in-situ recruitment of the MCP-RT fusion protein by attaching an MS2 aptamer and utilizing the interaction between the MS2 aptamer and the MCP protein. The MCP protein can specifically bind MS2 or multiple MS2 molecules in tandem. Exemplarily, the MCP comprises the sequence shown in SEQ ID NO:17. Exemplarily, the MS2 comprises the sequence shown in SEQ ID NO:18.

[0095] In this invention, when connecting various elements such as MCP, MS2, CRSIPR nuclease, and reverse transcriptase, they can be directly connected or connected through adapter sequences.

[0096] In some embodiments, the linker may be a non-functional amino acid sequence of one or more amino acids without secondary or higher-order structures. In some embodiments, the linker may be about 5 to 100 amino acids in length, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. For example, the linker may be a flexible linker. In some specific embodiments, the linker comprises the sequence shown in SEQ ID NO:19.

[0097] In some embodiments of the present invention, depending on the location of the DNA to be edited, the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein of the present invention may also include various localization sequences, such as nuclear localization sequences, cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.

[0098] In some exemplary embodiments of the present invention, the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein of the present invention may further comprise one or more nuclear localization sequences (NLS). Generally, one or more NLS in the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein should have sufficient strength to drive the accumulation of the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein in the nucleus of the cell to achieve its gene-editing function. Generally, the strength of nuclear localization activity is determined by the number, location, one or more specific NLS used, or a combination of these factors in the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein. Exemplarily, the NLS may comprise the sequence shown in SEQ ID NO:20 or SEQ ID NO:21.

[0099] In some embodiments, the guide editing fusion protein comprises a sequence selected from any of SEQ ID NO:1-3,7-11.

[0100] To achieve efficient expression in different organisms, in this invention, the nucleotide sequence encoding the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein can be codon-optimized for the organism whose genome is to be modified. "Codon optimization" refers to configuring the nucleotide sequence encoding the polypeptide to contain codons preferred by the host cell or organism to improve gene expression and translation efficiency in the host cell or organism.

[0101] The guide RNA (sgRNA) described in this invention refers to an sgRNA compatible with the CRISPR nuclease. The sgRNA typically comprises a scaffold sequence and a guide sequence (also called a seed sequence or spacer sequence). The guide sequence is configured to have sufficient sequence identity (preferably 100%) with the target sequence, thereby enabling sequence-specific targeting by binding to the complementary strand of the target sequence through base pairing. The scaffold sequence typically depends on the CRISPR nuclease used. In some alternative embodiments of this invention, the sgRNA may also target a non-target strand (i.e., the non-edited strand), referred to herein as nicking sgRNA. Nicking sgRNA is a specially designed single-guide RNA used to guide the Cas9 protein in the CRISPR-Cas9 system to generate a single-strand break (SSB) on the target DNA strand, rather than a double-strand break (DSB). Specifically, the nicking sgRNA targets a position approximately 50 base pairs from the PAM sequence in the non-target strand.

[0102] In some exemplary embodiments, the backbone sequence of the corresponding sgRNA for the Cas9 nuclease or variants thereof of the present invention comprises the sequence shown in SEQ ID NO:4.

[0103] In some embodiments, the pegRNA comprises sgRNA, a reverse transcription modulus (RTT) sequence, and a primer binding site (PBS) sequence (RTT-PBS sequence). Exemplarily, the pegRNA comprises a sequence as shown in SEQ ID NO:5 or SEQ ID NO:6.

[0104] In some implementations, the reverse transcription modulus (RTT) sequence and primer binding site (PBS) sequence can also be expressed via circular RNA, i.e., the circular RNA contains directly linked RTT and PBS sequences.

[0105] The RTT sequence is the reverse complementary sequence of the three bases from the 3' end of the target sequence and the subsequent continuous genomic sequence, in which the target mutation is introduced. This sequence serves as a reverse transcription template for reverse transcriptase, producing cDNA, which is then used as a repair template to repair the genomic DNA. The RTT sequence length can further be 8–34 bp.

[0106] Typically, the RTT sequence contains only the target mutation site (i.e. the site where the mutation is desired), and a mutant base is introduced at the target mutation site.

[0107] The PBS sequence (primer binding site sequence) is the reverse complementary sequence (1 ≤ n < 17) of the target sequence from the nth to the 17th base from the 5' end of the target sequence. Further, it is configured to be complementary to at least a portion of the target sequence (preferably perfectly paired with at least a portion of the target sequence). Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the target strand (TS) caused by cleavage (preferably perfectly paired with at least a portion of the 3' free single strand), particularly complementary to the nucleotide sequence at the 3' end of the 3' free single strand (preferably perfectly paired). When the 3' free single strand of the strand binds to the primer binding sequence through base pairing, the 3' free single strand can act as a primer, using the reverse transcription template sequence adjacent to the primer binding site sequence as a template. Reverse transcription is performed under the action of reverse transcriptase in the guide editing fusion protein or reverse transcriptase recruited by RNA aptamer binding protein, extending the DNA sequence corresponding to the reverse transcription template (RTT) sequence.

[0108] In some optional embodiments, the design methods or principles of the RTT sequence and the PBS sequence may refer to the design methods or principles of RTT sequences and PBS sequences of pegRNA in prime editing (PE) techniques reported in the prior art.

[0109] In other specific embodiments, the circular RNA further comprises one or more RNA aptamer sequences, such as the MS2 sequence. The one or more RNA aptamer sequences in the circular RNA can be used to recruit reverse transcriptases containing RNA aptamer-binding proteins in cells, or to recruit the circular RNA to the CRISPR nuclease-gRNA-genomic DNA complex in cells.

[0110] In some embodiments, the expression construct containing the nucleotide sequence encoding the circular RNA includes the following coding sequence: 5'-first ribozyme-first cyclic arm-RTT-PBS-second cyclic arm-second ribozyme-3'. Upon intracellular transcription into RNA, the first and second ribozymes self-cleave to generate "5'-first cyclic arm-RTT-PBS-second cyclic arm-3'", and the first and second cyclic arms can connect to form a circular RNA containing RTT and PBS. In some embodiments, one or more RNA aptamer sequences, such as MS2 sequences, may also be flanked by the RTT-PBS.

[0111] In some embodiments, the first ribozyme comprises the nucleotide sequence shown in SEQ ID NO:22, and the second ribozyme comprises the nucleotide sequence shown in SEQ ID NO:23. In some embodiments, the first cyclic arm comprises the nucleotide sequence shown in SEQ ID NO:24, and the second cyclic arm comprises the nucleotide sequence shown in SEQ ID NO:25.

[0112] In some optional embodiments, the reverse-guided editing system further includes the MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor. The MLH1dn protein factor is a dominant-negative mutant of MLH1; optionally, the MLH1dn protein factor comprises the sequence shown in SEQ ID NO:26.

[0113] In some alternative embodiments, the reverse editing system further comprises a helicase and / or an expression construct containing a nucleotide sequence encoding the helicase. In this invention, the helicase is used to untie double helices formed by two complementary DNA strands, double helices formed by DNA and complementary RNA, and other double strands containing DNA.

[0114] In some embodiments, the helicase is derived from porcine circovirus, Escherichia coli, or humans. Exemplarily, the helicase is derived from Escherichia coli and contains the sequence shown in SEQ ID NO:27.

[0115] II. Application of Reverse Editing Systems

[0116] This invention provides the application of the reverse-guided editing system of this invention in the following (A) or (B):

[0117] (A) Editing of the genome sequence of an organism or its cells;

[0118] (B) Products that are prepared by editing the genome sequence of an organism or biological cell.

[0119] In some specific embodiments, the cells are microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats, and poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis thaliana, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.

[0120] III. Methods for modifying target sequences in the cellular genome, methods for generating genetically modified cells, and genetically modified organisms.

[0121] This invention provides a method for producing genetically modified cells, the method comprising introducing the reverse-guided editing system of the invention into at least one cell, thereby resulting in modification of the genomic sequence of the at least one cell. The modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the modification includes one or more substitutions selected from the following: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or includes the deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide deletions; and / or includes the insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide insertions. The modification can be located in or near the target sequence of the sgRNA, for example, upstream of the target sequence.

[0122] In another aspect, the present invention provides a method for generating genetically modified cells, comprising introducing the guided editing system of the present invention into the cells.

[0123] In another aspect, the present invention also provides genetically modified organisms comprising genetically modified cells or their progeny cells produced by the method of the present invention.

[0124] In this invention, the modification can be located anywhere in the genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter or enhancer region, thereby achieving modification of gene function or gene expression. The modification in the cell genome sequence can be detected using T7EI, PCR / RE, or sequencing methods.

[0125] In this invention, the reverse guided editing system can be introduced into cells using various methods well known to those skilled in the art. For example, methods for introducing the guided editing system of this invention into cells include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, and other viruses), gene gun method, PEG-mediated protoplast transformation, and Agrobacterium-mediated transformation. Cells that can be gene-edited using the methods of this invention can be derived from, for example, microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.

[0126] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.

[0127] In other embodiments, the method of the present invention can also be performed in vivo. For example, the cells are cells within an organism, and the system of the present invention can be introduced into the cells in vivo via, for example, a viral or Agrobacterium-mediated method.

[0128] IV. Methods for producing genetically modified plants and plant breeding methods

[0129] This invention provides a method for producing genetically modified plants, comprising introducing the reverse-guided editing system of this invention into at least one of the plants, thereby resulting in modifications in the genome of the at least one plant. The modifications include substitutions, deletions, and / or additions of one or more nucleotides. For example, the modification includes one or more substitutions selected from the following: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or includes the deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide deletions; and / or includes the insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide insertions.

[0130] In some embodiments, the method further includes screening plants with desired modifications from the at least one plant.

[0131] In the method of this invention, the reverse guided editing system can be introduced into plants using various methods well known to those skilled in the art. Methods for introducing the guided editing system of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the guided editing system is introduced into plants via transient transformation.

[0132] In the method of this invention, genome modification can be achieved simply by introducing or generating relevant proteins and RNA in plant cells, and the modification can be stably inherited without the need for stable transformation of plants with exogenous polynucleotides encoding components of the reverse-guided editing system. This avoids the potential off-target effects of a stably existing (continuously generated) editing system and also avoids the integration of exogenous nucleotide sequences into the plant genome, thus providing higher biosafety.

[0133] In some preferred embodiments, the introduction is performed in the absence of selection pressure, thereby avoiding the integration of exogenous nucleotide sequences into the plant genome.

[0134] In some embodiments, the introduction includes converting the reverse-guided editing system of the present invention into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into complete plants. Preferably, the regeneration is performed without selection pressure, i.e., without using any selection agents targeting the selection genes carried on the expression vector during tissue culture. Not using selection agents can improve the regeneration efficiency of the plants, resulting in modified plants free of exogenous nucleotide sequences.

[0135] In other embodiments, the reverse-guided editing system of the present invention can be applied to specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.

[0136] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) are directly transformed into the plant. The proteins and / or RNA molecules can be guided to edit within plant cells and are subsequently degraded by the cells, avoiding the integration of exogenous nucleotide sequences into the plant genome.

[0137] Therefore, in some embodiments, using the methods of the present invention to genetically modify and breed plants can yield plants whose genomes are free of foreign polynucleotide integration, i.e., non-transgene-free modified plants.

[0138] In some embodiments of the invention, the modified genomic region is associated with plant traits such as agronomic traits, whereby the modification results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.

[0139] In some embodiments, the method further includes the step of screening plants with desired modifications and / or desired traits such as agronomic traits.

[0140] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring have the desired modification and / or desired traits such as agronomic traits.

[0141] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein said plants are obtained by the methods described above. In some embodiments, the genetically modified plants or their offspring or portions thereof are non-GMO. Preferably, the genetically modified plants or their offspring have the desired genetic modification and / or desired traits such as agronomic traits.

[0142] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant obtained by the method described above with a second plant that does not contain the modification, thereby introducing the modification into the second plant. Preferably, the genetically modified first plant has desired traits such as agronomic traits.

[0143] "Agronomic traits" specifically refer to measurable parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total nitrogen content of plants, nitrogen content of fruits, nitrogen content of seeds, nitrogen content of plant vegetative tissues, total free amino acid content of plants, free amino acid content of fruits, free amino acid content of seeds, free amino acid content of plant vegetative tissues, total protein content of plants, protein content of fruits, protein content of seeds, protein content of plant vegetative tissues, herbicide resistance and drought resistance, nitrogen uptake, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt tolerance, and tiller number, etc.

[0144] V. Therapeutic Uses

[0145] This invention also provides the application of the reverse-guided editing system of this invention in disease treatment.

[0146] By modifying disease-related genes using the reverse-guided editing system of this invention, it is possible to achieve upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the genomic modifications described in this invention can be located within the protein-coding region of the disease-related gene, or, for example, within gene expression regulatory regions such as promoter regions or enhancer regions, thereby enabling modifications to the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modifications to the disease-related gene itself (e.g., protein-coding regions), as well as modifications to its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).

[0147] "Disease-associated" genes are any genes that produce transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control groups. In cases where altered expression is associated with the onset and / or progression of the disease, it can be a gene expressed at abnormally high levels; it can also be a gene expressed at abnormally low levels. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked to one or more genes responsible for the etiology of the disease in disequilibrium. Such mutations or genetic variations are, for example, single nucleotide variants (SNVs). The transcribed or translated products can be known or unknown and can be at normal or abnormal levels.

[0148] Therefore, the present invention also provides a method for treating a disease in a subject in need, comprising delivering an effective amount of the guided editing system of the present invention to the subject to modify a gene associated with the disease. The present invention also provides the use of the guided editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need, wherein the guided editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need, comprising the guided editing system of the present invention and optionally a pharmaceutically acceptable vector, wherein the guided editing system is used to modify a gene associated with the disease.

[0149] Preferably, the "object" referred to in this invention is a mammal, such as a human.

[0150] In some implementations, the guided editing system described in this invention is used to introduce point mutations into nucleic acids.

[0151] In some embodiments, the guided editing system described herein is used to correct genetic defects, such as in correcting point mutations that result in loss of function in a gene product. In some embodiments, the genetic defect is associated with a disease or condition (e.g., lysosomal storage disease or metabolic disease, such as, for example, type 1 diabetes). In some embodiments, the methods provided herein can be used to introduce inactive point mutations into a gene or allele encoding a gene product associated with a disease or condition.

[0152] In some embodiments, the purpose of the schemes described in this invention is to treat diseases associated with or caused by point mutations, which can be corrected using the guided editing system provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neonatal disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.

[0153] In some embodiments, the purposes of the solutions described in this invention are for the treatment of mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological disorders, loss of motor control, muscle weakness and pain, gastrointestinal disorders and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, visual / hearing problems, lactic acidosis, developmental delay, and susceptibility to infection.

[0154] Examples of diseases described in this invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat amplification disorders, hearing disorders, gene-targeted therapy for non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, β-thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, and schizophrenia. Other diseases that can be treated by correcting point mutations or introducing inactive mutations into disease-related genes are known to those skilled in the art, and therefore this disclosure is not limited in this respect. In addition to the diseases exemplarily described in this invention, other related diseases can also be treated using the strategies and guided editing systems provided in this invention, and this application will be apparent to those skilled in the art. The diseases or targets to which this invention can be applied refer to the base editing systems listed in WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1 (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), and WO2021155065A1 (PCT / US2021 / 015580) for applicable diseases.

[0155] The administration of the guided editing system or pharmaceutical composition of the present invention can be tailored to the patient's or subject's weight and species. The frequency of administration is within medically or veterinary limits. It depends on conventional factors including the patient's or subject's age, sex, general health condition, other conditions, and the specific symptom or condition being addressed.

[0156] VI. Reagent Kit

[0157] The present invention also provides a kit for use with the methods of the present invention, the kit comprising the components of the reverse-guided editing system of the present invention. The kit may also contain reagents for introducing the reverse-guided editing system into an organism or somatic cells. The kit generally includes a label indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise accompanied by the kit.

[0158] Example

[0159] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.

[0160] Materials and Methods

[0161] 1. Carrier Construction

[0162] The vector was constructed using the existing guide editor vector skeleton in the laboratory. The main vectors constructed were:

[0163] iPEs:

[0164] (1) iPE2 contains iPE-nCas9-D10A and pegRNA.

[0165] The iPE-nCas9-D10A sequence is shown in SEQ ID NO:2.

[0166] The pegRNA sequence is shown in SEQ ID NO:5;

[0167] (2)iPE3 contains iPE-nCas9-D10A, pegRNA, and nicking sgRNA

[0168] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the rest of the sequence is the same as iPE2;

[0169] (3) iPE4 contains iPE-nCas9-D10A, pegRNA, and MLH1dn protein factor.

[0170] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPE2;

[0171] The iPE-nCas9-H840A is a direct replacement of the nCas9-D10A in the iPE2 with the nCas9-H840A.

[0172] (4) iPE5 contains iPE-nCas9-D10A, pegRNA, MLH1dn protein factor, and nicking sgRNA.

[0173] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPE2;

[0174] (5) iPEmax2 contains iPEmax-nCas9-D10A and epegRNA

[0175] The iPEmax-nCas9-D10A sequence is shown in SEQ ID NO:3.

[0176] The epegRNA sequence is shown in SEQ ID NO:6;

[0177] (6) iPEmax3 contains iPEmax-nCas9-D10A, pegRNA, and nicking sgRNA.

[0178] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the rest of the sequence is the same as iPEmax2;

[0179] (7) iPEmax4 contains iPEmax-nCas9-D10A, pegRNA, and MLH1dn protein factor.

[0180] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPEmax2.

[0181] (8) iPEmax5 contains iPEmax-nCas9-D10A, pegRNA, MLH1dn protein factor, and nicking sgRNA.

[0182] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPEmax2.

[0183] nu-iPEs:

[0184] (1) nu-iPE2 contains iPE-WT-Cas9 and pegRNA

[0185] The iPE-WT-Cas9 sequence is shown in SEQ ID NO:1.

[0186] The pegRNA sequence is shown in SEQ ID NO:5;

[0187] (2) nu-iPEmax2 contains iPEmax-WT-Cas9 and epegRNA

[0188] The iPEmax-WT-Cas9 sequence is shown in SEQ ID NO:29.

[0189] The epegRNA sequence is shown in SEQ ID NO:6;

[0190] (3) nu-iPEmax4 contains iPEmax-WT-Cas9, epeRNA, and MLH1dn protein factor.

[0191] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as nu-iPEmax2.

[0192] ciPEs:

[0193] (1) ciPE2 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, and sgRNA.

[0194] The nCas9-D10A sequence is shown in SEQ ID NO:10.

[0195] The reverse transcriptase M-MLV RTΔRNase H sequence is shown in SEQ ID NO:11.

[0196] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.

[0197] The sgRNA sequence is shown in SEQ ID NO:4.

[0198] (2) ciPE3 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and nicking sgRNA.

[0199] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the rest of the sequence is the same as ciPE2;

[0200] (3) ciPE4 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.

[0201] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as ciPE2.

[0202] (4) ciPE5 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, nicking sgRNA, and MLH1dn protein factor.

[0203] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as ciPE2.

[0204] The equivalent of ciPE-nCas9-H840A is to replace nCas9-D10A in ciPE2 with nCas9-H840A.

[0205] nu-ciPE:

[0206] (1) nu-ciPE2 contains WT-Cas9, reverse transcriptase M-MLV RTΔRNase H, circular RNA, and sgRNA.

[0207] The WT-Cas9 sequence is shown in SEQ ID NO:9.

[0208] The reverse transcriptase M-MLV RTΔRNase H sequence is shown in SEQ ID NO:11.

[0209] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.

[0210] The sgRNA sequence is shown in SEQ ID NO:4.

[0211] (2) nu-ciPE4 contains WT-Cas9, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.

[0212] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as nu-ciPE2.

[0213] hciPEs(ciPEs+Rep-X):

[0214] (1) hciPE2 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, and sgRNA.

[0215] The sequence of nCas9-D10A is shown in SEQ ID NO:10.

[0216] The sequence of MCP-Helicase-M-MLV RTΔRNase H is shown in SEQ ID NO:12.

[0217] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.

[0218] The sgRNA sequence is shown in SEQ ID NO:4.

[0219] (2) hciPE3 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and nicking sgRNA.

[0220] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the remaining sequences are the same as hciPE2;

[0221] (3) hciPE4 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.

[0222] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as hciPE2.

[0223] (4) hciPE5 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.

[0224] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as hciPE2.

[0225] For details on the editing systems split IPE2, split IPEmaxs△R (split IPEmax2△R, split IPEmax3△R, split IPEmax4△R, split IPEmax5△R), PAMless-PEs, twinPE, upPE, downPE, SpG-PEs (SpG-PE2, SpG-PE3, SpG-PE4, SpG-PE5), and SpRY-PEs (SpRY-PE2, SpRY-PE3, SpRY-PE4, SpRY-PE5), please refer to Liang, R., Wang, S., Cai, Y. et al. Circular RNA-mediated inverse prime editing in human cells. Nat Commun 16, 5057 (2025), which is incorporated herein by reference.

[0226] 2. Human cell culture and transformation

[0227] 2.1 Thawing cells

[0228] (1) Remove HEK293T, U2OS and K562 cell lines from liquid nitrogen and thaw them in a 37°C water bath for 2 minutes.

[0229] (2) Remove the cells from the cryopreservation tube and slowly add them to a 15 mL centrifuge tube containing 10 mL of culture medium.

[0230] (3) Centrifuge at 500×g for 3 minutes to precipitate cells and remove the supernatant.

[0231] (4) Add 10 mL of culture medium to the precipitate, transfer it to a T175 culture flask, and place it in a 37°C incubator containing 5% CO2.

[0232] (5) Replace the culture medium after 24 hours and continue culturing in an incubator containing 5% CO2 at 37°C.

[0233] 2.2 Cell Culture

[0234] (1) Carefully aspirate the cell culture medium from the T175 culture flask.

[0235] (2) Gently wash the cells with 10 mL of PBS solution.

[0236] (3) Treat the cells with 2 mL of trypsin: Place the T175 cell culture flask in a 37°C incubator for 2 minutes to allow the trypsin to fully dissociate the cells.

[0237] (4) Add 10 mL of cell culture medium to the T175 cell culture flask to resuspend the cells and inactivate trypsin.

[0238] (5) Use a 10mL pipette to blow the cell suspension up and down to obtain a single-cell suspension.

[0239] (6) After diluting the cells at a ratio of 1:10, continue to culture them in an incubator at 37°C containing 5% CO2.

[0240] 2.3 Cell Deployment

[0241] (1) Take 15 μL of the single cell suspension from step (5) in 2.1 and mix it with 15 μL of trypan blue. Let it stand at room temperature for 1-2 minutes.

[0242] (2) Open a new cell counter. Add 10 μL of well-mixed cell and trypan blue solution to each of the counting chambers A and B, insert the counting plate into the cell counter and read the count.

[0243] (3) Calculate the required dilution factor so that the final cell count is 45,000 cells per well in 250 μL of cell culture medium in a 48-well plate. Dilution factor = (A count + B count) / (2 × 4 × 45000).

[0244] (4) Dilute the cells according to the dilution factor and use an adjustable width multichannel pipette to dispense the cell suspension into 48-well plates, adding 250uL of cell suspension to each well.

[0245] (5) Place the 48-well plate containing cells in a 37°C incubator and incubate for 16 to 24 hours before using it for liposome conversion.

[0246] 2.4 Transfecting cells

[0247] (1) Before transfecting cells, observe the cell growth status under a microscope, preferably with 60% coverage.

[0248] (2) Mix 300 ng of protein expression plasmid and 100 ng of RNA expression plasmid in Opti-MEM to make a total volume of 12.5 μL.

[0249] (3) Add 1 μL of Lipofectamine 2000 and 10-20 ng of copGFP expression plasmid to Opti-MEM to make the total volume 12.5 μL.

[0250] (4) Mix (2) and (3) to a total volume of 25 μL and incubate for 5-15 minutes. Then add the solution to the cell culture in a 48-well plate. Incubate at 37°C for 48-72 hours.

[0251] 3. Fluorescence observation of the reporting system

[0252] Human cells cultured at 37°C for 48 hours were observed under a fluorescence microscope or a laser confocal microscope. The efficiency of different guide editors was compared based on the fluorescence, and the fluorescence images were saved.

[0253] In the following embodiments, unless otherwise specified or marked, HEK293T cells were used.

[0254] 4. Human cell lysis and amplicon sequencing analysis

[0255] 4.1 Lysing human cells

[0256] (1) After culturing the cells in a 37°C incubator for 72 hours, carefully aspirate the cell culture medium using an adjustable width multichannel pipette.

[0257] (2) Gently add 300uL of PBS buffer, shake gently to wash away dead cells, and then aspirate the PBS buffer.

[0258] (3) Add 200uL of lysis buffer (with proteinase K added) and treat at 55℃ for 30 minutes.

[0259] (4) Pipette 50 μL of cell lysis buffer into a 96-well plate and treat at 95°C for 5 minutes before use.

[0260] 4.2 Amplicon Miseq Sequencing Analysis

[0261] (1) PCR amplification of human cell lysate was performed using Miseq first-round primers.

[0262] The first round of 15μL amplification system consisted of: 7.5μL 2×Phanta Max Master Mix, 4.5μL ddH2O, 1μL forward primer (10μM), 1μL reverse primer (10μM), and 1μL cell lysis buffer.

[0263] The first round of 15μL amplification conditions were as follows: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 15 s, 50-60℃ annealing for 15 s, 72℃ extension for 20 s, 34 cycles; and 72℃ full extension for 5 min.

[0264] (2) The first-round amplification product was diluted 5 times, and 1 μL was used as the template for the second-round PCR amplification. The amplification primers were the Miseq second-round sequencing primers containing barcode.

[0265] The second round of amplification system consisted of 25 μL: 12.5 μL 2×Phanta Max Master Mix, 9.5 μL ddH2O, 1 μL forward primer (10 μM), 1 μL reverse primer (10 μM), and 1 μL DNA template.

[0266] Second round amplification conditions: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 15 s, 50-60℃ annealing for 15 s, 72℃ extension for 20 s, 10 cycles; 72℃ full extension for 5 min.

[0267] (3) The PCR products were detected by 2% agarose gel electrophoresis, and the target fragment was recovered by gel extraction using the AxyPrep DNA Gel Extraction kit. The recovered products were quantitatively analyzed using a NanoDrop ultra-micro spectrophotometer. 100 ng of the recovered products were mixed and sent to Beijing Qihe Biotechnology Co., Ltd. for amplicon sequencing analysis.

[0268] (4) After sequencing is completed, the editing type and editing efficiency of the products are compared and analyzed at different gene target sites in at least three repeated experiments.

[0269] 3. Target sites and target sequences

[0270] Example 1. An inverse prime editing system (iPE) developed using pegRNA and nCas9-D10A.

[0271] Reverse-guided editing occurs in human cells.

[0272] Traditional guide editors primarily utilize nCas9-H840A and the RNA-dependent 5'→3' reverse transcriptase M-MLV RT to generate guide editing downstream of the target sequence cleavage site. Since no reverse transcriptase capable of reverse transcription along the 3'→5' axis has been found in nature to work with nCas9-H840A, traditional guide editors cannot generate guide editing upstream of the cleavage site.

[0273] This invention first utilizes pegRNA and nCas9-D10A to develop the reverse-guided editing system iPE (Figures 1a and 2a). The invention initially developed the iPE2 editor, then developed the iPE3 editor by adding nicking sgRNA, the iPE4 editor by adding the MLH1dn protein factor, and finally the iPE5 editor by adding both. Testing in human HEK293T cells revealed low efficiency, producing editing efficiencies of 0.04%–1.06% at target sites HBB, HEXA, FANCF, and PDCD1 (Figure 3a). Simultaneously, the inventors constructed more efficient iPEmax2, iPEmax3, iPEmax4, and iPEmax5 editors using the iPEmax backbone, finding that they produced higher editing efficiencies than iPEs, reaching 8.6% at DMD sites (Figure 3b).

[0274] Example 2. The nu-iPE reverse-guided editing system, developed using pegRNA and WTcas9, generates reverse-guided editing in human cells HEK293T.

[0275] This invention further utilizes pegRNA and WTCas9 to develop a reverse guided editing system, nu-iPEs (Figures 1b and 2b), which was then tested in the HEK293T reporter system. In the reporter system, a one-base deletion and a two-base substitution mutation were generated on copGFP. Only when precise and efficient reverse guided editing occurred on the copGFP plasmid could the HEK293T reporter system recover fluorescence. Experimental results showed that nu-iPEs produced higher reverse guided editing efficiency in the reporter system than iPEs, resulting in brighter green fluorescence (Figures 4a-4c). Simultaneously, the inventors found the same results in HEK293T cells, where nu-iPEs produced higher reverse guided editing efficiency than iPEs (Figure 5). The inventors analyzed that traditional guided editing is based on nCas9-H840A. The PBS sequence binding site is within 20 bp of the target site, and the target sequence region is well unwound by nCas9-H840A, which is beneficial for reverse transcriptase to use the RTT-PBS sequence for reverse transcription, thereby generating the desired edit. However, the downstream position of the target sequence cleavage site is not unwound by Cas9, so the reverse guided editing constructed using nCas9-D10A is less efficient than traditional guided editors. In contrast, after DNA double-strand breaks are generated using WTCas9, the body's repair mechanisms further open the DNA double strand downstream of the target sequence cleavage site, which is beneficial for PBS binding and thus facilitates reverse guided editing.

[0276] Example 3. Using circular RNA and ciPE and nu-ciPE editors developed with nCas9-D10A and WTCas9 respectively, efficient reverse-guided editing was generated on the HEK293T target site in human cells.

[0277] The results in Examples 1 and 2 of this invention suggest that the inventors need to increase the unwinding of the DNA double strand downstream of the target sequence cleavage site to facilitate reverse-guided editing. Therefore, this invention innovatively introduces circular RNA (SEQ ID NO: 13, with the RTT and PBS sequences of pegRNA inserted at the MCS site of the polyclonal restriction enzyme) into the reverse-guided editing system. Because circular RNA itself has a certain unwinding ability, it can unwind downstream of the target sequence cleavage site and bind the carried RTT-PBS sequence downstream of the target sequence cleavage site, initiating M-MLV RTΔRNase H-mediated reverse transcription to generate the desired edited sequence. The inventors developed a ciPE editor (Figures 1c and 2c) and tested it in human cells HEK293T. They found that ciPEs produced higher reverse-guided editing efficiency than iPEs at human cell targets DMD, FANCF, HEK3, RNF2, PDCD1, CXCR4, HEXA, HEK4, BCL11A, and TRAC, with an editing efficiency of up to 24.7% at the HEK4 site (Figures 6a-6b).

[0278] The inventors also developed a reverse-guided editing system, nu-ciPE, using circular RNA (Figures 1d and 2d), based on WTCas9, because WTCas9 may better open the DNA double strand downstream of the target sequence cleavage site during repair, which is beneficial for reverse-guided editing. The inventors performed reverse-guided editing at six target sites in human HEK293T cells, and the results showed that the reverse-guided editors nu-ciPEs produced the expected precise editing upstream of the PAH target site, with an editing efficiency of approximately 19% (Figure 7).

[0279] Example 4. Helicase-assisted reverse guide editor hciPE produces more efficient reverse guide editing than ciPE in human cells HEK293T.

[0280] This invention further incorporates a 3'→5' DNA helicase to unwind downstream of the target sequence cleavage site, facilitating the binding of the PBS sequence carried by the circular RNA to the target sequence. Under the action of reverse transcriptase M-MLV RTΔRNase H, the desired DNA sequence for editing is generated, enabling reverse guided editing. This invention utilizes the previously optimized 3'→5' DNA helicase Rep-X to assist in reverse guided editing, developing the hciPE reverse guided editor (Figures 1e and 2e). Tested in human cells HEK293T, K562, and U2OS, the hciPE reverse guided editor demonstrated higher reverse guided editing efficiency than ciPEs at six target sites (HEK4, DMD, BCL11A, PSMB2, GFAP, and HEXA), reaching a maximum of 55.4% (Figures 8a-8d).

[0281] Example 5. The helicase-assisted reverse-guided editor hciPE produces efficient editing in regions such as disease sites that were previously difficult to edit.

[0282] This invention further compares the efficiency of the hciPE editor with existing editors PAMless-PE and twinPE, finding that PAMless-PEs, including SpRY-PE and SpG-PE, only achieve an average maximum editing efficiency of 4.5% for the three target sites of human HEK293T cells, while the hciPE editor achieves an average editing efficiency of up to 14.0%, significantly higher than PAMless-PEs (Figures 9a-9b). Similarly, for the disease targets BRCA1 and RPE65, PAMless-PEs and twinPE produce editing efficiencies of 5.7% and 2.9%, and 0.8% and 0.5%, respectively, while hciPE can produce editing efficiencies of 13.3% and 9.5% (Figures 9c-9d).

[0283] Example 6. The ciPE and hciPE editors produce lower off-target activity.

[0284] This invention further utilized Cas-OFFinder to predict off-target sites for GFAP, HEK4, and DMD targets in human HEK293T cells, identifying a total of 26 off-target sites. These off-target sites were then validated using ciPE and hciPE editors. At GFAP and DMD sites, only background levels of InDels products (<0.04%) were detected, and no off-target reverse-guided editing products were detected (Figures 10a-10b). However, among the nine off-target sites in HEK4, off-target site 3 generated 1.32%–6.32% InDels off-target activity and 0.05%–0.67% reverse-guided editing off-target activity. Given the low off-target activity of guided editing, the inventors further analyzed the reasons for the off-target activity of reverse guided editing at off-target site 3. They found that this site has 7 bases identical to the target sequence in the PBS region, which will greatly improve the binding efficiency of the PBS sequence to this off-target site, thus resulting in a certain off-target activity (Figures 10a-10b).

[0285] The sequences involved in this invention:

[0286] The sequence reporting system sequence shown in the attached figure:

[0287] Repaired sequence:

Claims

1. A reverse-guided editing system, comprising: (i)a) An expression construct containing a CRSIPR nuclease and / or a nucleotide sequence encoding the CRSIPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase; or b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRSIPR nuclease and / or a reverse transcriptase; (ii) a guide RNA (pegRNA) targeting a target sequence in genomic DNA and / or an expression construct containing a nucleotide sequence encoding said guide RNA; and / or (iii) An expression construct containing a circular RNA with a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or a nucleotide sequence encoding the circular RNA, and an expression construct containing a guide RNA (sgRNA) targeting a sequence in genomic DNA and / or a nucleotide sequence encoding the guide RNA; Optionally, the CRSIPR nuclease is a target strand nickase or a double-stranded endonuclease.

2. The reverse-guided editing system according to claim 1, wherein, The CRSIPR nuclease is a Cas9 nuclease or a variant thereof.

3. The reverse-guided editing system according to claim 2, wherein, The Cas9 nuclease comprises the sequence shown in SEQ ID NO:14; and / or The Cas9 nuclease variant is an nCas9 nuclease, which contains the sequence shown in SEQ ID NO:

15.

4. The reverse-guided editing system according to any one of claims 1 to 3, wherein, The reverse enzyme is M-MLV reverse transcriptase; Preferably, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted, and it contains the sequence shown in SEQ ID NO:

16.

5. The reverse-guided editing system according to any one of claims 1 to 4, wherein, The reverse transcriptase and / or the guided editing fusion protein may also contain one or more RNA aptamer-binding protein sequences.

6. The reverse-guided editing system according to claim 5, wherein, The RNA aptamer-binding protein is the MCP protein; Optionally, the MCP protein comprises the sequence shown in SEQ ID NO:

17.

7. The reverse-guided editing system according to any one of claims 1 to 6, wherein, The guided editing fusion protein comprises a sequence selected from any of the sequences shown in SEQ ID NO: 1-3, 7-11.

8. The reverse guided editing system according to any one of claims 1 to 7 further comprises a guide RNA (nicking sgRNA) targeting a non-target strand and / or an expression construct containing a nucleotide sequence encoding the guide RNA targeting a non-target strand.

9. The reverse-guided editing system according to any one of claims 1 to 8, wherein, The guide RNA comprises a backbone sequence as shown in SEQ ID NO:4; and / or The guide RNA that targets the non-target strand contains a backbone sequence as shown in SEQ ID NO:

4.

10. The reverse-guided editing system according to any one of claims 1 to 9, wherein, In circular RNA, the reverse transcription template (RTT) sequence and the primer binding site (PBS) sequence are directly linked; Preferably, the reverse transcription template (RTT) sequence is located at the 5' end of the primer binding site (PBS) sequence.

11. The reverse-guided editing system according to any one of claims 1 to 10, wherein, The circular RNA contains at least one RNA aptamer sequence, and optionally, the aptamer includes MS2.

12. The reverse-guided editing system according to any one of claims 1 to 11, wherein, The primer binding site (PBS) sequence is complementary to at least a portion of the target sequence; Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the target strand caused by the nick, particularly to the nucleotide sequence at the 3' end of the 3' free single strand.

13. The reverse-guided editing system according to any one of claims 1 to 12, wherein, The expression construct containing the nucleotide sequence encoding the circular RNA comprises the following coding sequence: 5'-first ribozyme-first cyclic arm-RTT-PBS-second cyclic arm-second ribozyme-3'; Optionally, after transcription into RNA within the cell, the first and second ribozymes can self-cleave to produce 5'-first cyclic arm-RTT-PBS-second cyclic arm-3', and the first and second cyclic arms can connect with each other to form a circular RNA containing RTT and PBS.

14. The reverse-guided editing system according to claim 13, wherein, The coding sequence of the first ribozyme includes the sequence shown in SEQ ID NO:22, the coding sequence of the second ribozyme includes the sequence shown in SEQ ID NO:23, the first cyclic arm includes the nucleotide sequence shown in SEQ ID NO:24, and the second cyclic arm includes the nucleotide sequence shown in SEQ ID NO:

25.

15. The reverse-guided editing system according to any one of claims 1 to 14, wherein, The nucleotide sequence encoding the circular RNA is expressed by the U6 promoter.

16. The reverse-guided editing system according to any one of claims 1 to 15, wherein, It also includes the MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor; Optionally, the MLH1dn protein factor comprises the sequence shown in SEQ ID NO:

26.

17. The reverse-guided editing system according to any one of claims 1 to 16, further comprising a helicase and / or an expression construct containing a nucleotide sequence encoding said helicase; Optionally, the helicase comprises the sequence shown in SEQ ID NO:27; Preferably, the helicase and reverse transcriptase form a fusion protein; more preferably, the fusion protein comprises the sequence shown in SEQ ID NO:

12.

18. The application of the reverse-guided editing system according to any one of claims 1 to 17 in either (A) or (B): (A) Editing of the genome sequence of an organism or its cells; (B) Products that are prepared by editing the genome sequence of an organism or biological cell.

19. A kit comprising the reverse boot editing system according to any one of claims 1 to 17.

20. A method of producing genetically modified cells, comprising introducing a reverse-guided editing system according to any one of claims 1 to 17 into at least one of the cells, thereby resulting in modification of the genome sequence of the at least one cell.

21. The method of claim 20, wherein the cells are derived from microorganisms, animals, or plants; The microorganisms include bacteria and fungi; The animals include mammals and poultry; The plants include monocotyledons and dicotyledons; Optionally, the mammals include humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; Optionally, the poultry includes chickens, ducks, and geese; Optionally, the monocotyledonous plants include rice, corn, wheat, sorghum, and barley, and the dicotyledonous plants include soybean, peanut, Arabidopsis thaliana, cotton, and rapeseed.

Citation Information

Patent Citations

  • Method for inserting exogenous sequence in genome at fixed point

    CN117126876A

  • Systems and methods for inserting and editing large nucleic acid fragments

    CN118043457A

  • Engineered ADAR recruitment RNA and methods of use thereof

    CN118202045A

  • Guide editing system based on circular RNA

    CN118995701A

  • Artificial DNA replisome and methods of use thereof

    WO2022178167A1

Cited By

  • Split-phase nicking enzyme mediated pilot editor and editing system and application of split-phase nicking enzyme mediated pilot editor

    CN122168603A