Reverse guidance editing system
By using a reverse-guided editing system, precise DNA editing can be performed upstream of the target site using targeted strand nickase and circular RNA. This solves the problem that traditional guided editors can only edit downstream, enabling wider application and more efficient editing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional guided editors can only perform precise editing downstream of the target sequence cleavage site, and cannot achieve the desired editing upstream of the target site, which limits their application scope and potential.
Develop a reverse-guided editing system that utilizes targeted strand nickases such as D10A-Cas9 and double-stranded endonucleases such as WT-Cas9, combined with pegRNA and circular RNA, to achieve precise DNA editing upstream of the target site. Small DNA fragments can be inserted, deleted, and replaced using iPE, nu-iPE, ciPE, and hciPE systems.
It expands the scope of editing, reduces editing byproducts, improves editing efficiency and safety, and enhances its application potential in disease treatment, crop improvement, and biological research.
Smart Images

Figure CN121628933A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of genetic engineering. Specifically, it relates to a reverse-guided editing system for achieving site-specific, precise small-fragment DNA editing upstream of a target site. More specifically, the reverse-guided editing system of this invention utilizes a target strand nickase (such as D10A-Cas9) or a double-stranded endonuclease (such as WT-Cas9) to make a nick on the target strand, and under the action of reverse transcriptase, performs precise small-fragment DNA base insertion, deletion, and substitution upstream of the target site, as well as genetically modified but non-transgenic organisms and their offspring produced by the method. Background Technology
[0002] Precise and efficient genome editing has wide applications in disease treatment, crop improvement, and biological research. Guided editors are widely used in various life science fields because they can perform all single-base transformations of the four base types without causing double-strand breaks, fully covering the single-base editing function of base editors, and enabling precise insertion, deletion, and replacement of small DNA fragments. In guided editors, the reverse transcriptase M-MLV RT (Moloney-murine leukemiavirus reverse transcriptase) uses a reverse transcriptase template (RTT) and a primer binding site (PBS) to generate the desired 3' end single-stranded DNA sequence along the 5'→3' direction. Therefore, guided editing can produce the desired precise edit downstream of the target sequence cleavage site. However, reverse transcriptase cannot perform reverse transcription in the 3'→5' direction using RNA as a template to produce the desired 5' end single-stranded DNA sequence. Therefore, traditional guided editors can only produce precise edits downstream of the target sequence cleavage site, not upstream, which greatly limits the application scope and potential of guided editing. Furthermore, Lee et al. (J.Lee,K.Lim,A.Kim,etal.Prime editing with genuine Cas9 nickases minimizes unwantedindels.Nat Commun.2023Mar30;14(1):1786.) found that the target strand nickase D10A-Cas9 has more precise single-strand cleavage activity than the non-target strand nickase H840A-Cas9. The traditional guide editor uses H840A for single-strand cleavage, which is one of the reasons why guide editors produce more byproducts. Summary of the Invention
[0003] The problem the invention aims to solve
[0004] Reverse-guided editing systems, which utilize targeted chain cleavage enzymes (such as D10A-Cas9) or double-stranded endonucleases (such as WT-Cas9 and various other Cas proteins) to perform editing upstream of the editing target site, thus opposing the direction of traditional guided editing, will be able to further increase the editing range and expand application potential.
[0005] Solution for solving the problem
[0006] This invention first utilizes pegRNA (prime editing guide RNA) and Cas9 to develop a novel inverse prime editing system, iPE (inverse prime editing system), including nCas9-D10A-dependent iPE (nickase-dependent iPE). Figure 1a and Figure 2 a) and nu-iPEs that depend on WTCas9 (nuclease-dependent iPE, nu-iPE, Figure 1b and Figure 2 (b) This method is used to generate precise guided editing upstream of the target sequence cleavage site. Simultaneously, circular RNA with unwinding capability is used to replace pegRNA in expressing RTT-PBS sequences, and an nCas9-D10A-dependent reverse guided editing system, ciPE (nickase-dependent circular RNA-mediated iPE), is developed. Figure 1c and Figure 2 c) and nu-ciPE (nuclease-dependent circular RNA-mediated iPE), a WTCas9-dependent reverse guided editing system. Figure 1d and Figure 2 d). Finally, to further improve the editing efficiency of nickase-dependent reverse-guided editing systems, this invention develops a helicase-assisted ciPE reverse-guided editing system, hciPE (helicase-assisted hciPE). Figure 1e and Figure 2 e). iPE and the highly efficient ciPE reverse-guided editing system will enable more comprehensive single-base transformations and precise, efficient editing of small DNA fragments across a wider range of genomes in organisms such as animals, bacteria, dicotyledons, and monocotyledons.
[0007] This invention provides a reverse-guided editing system, comprising:
[0008] (i)a) An expression construct containing a CRSIPR nuclease and / or a nucleotide sequence encoding the CRSIPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase; or
[0009] b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRSIPR nuclease and / or a reverse transcriptase;
[0010] (ii) a guide RNA (pegRNA) targeting a target sequence in genomic DNA and / or an expression construct containing a nucleotide sequence encoding said guide RNA; and / or
[0011] (iii) An expression construct containing a circular RNA with a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or a nucleotide sequence encoding the circular RNA, and an expression construct containing a guide RNA (sgRNA) targeting a sequence in genomic DNA and / or a nucleotide sequence encoding the guide RNA;
[0012] Optionally, the CRSIPR nuclease is a target strand nickase or a double-stranded endonuclease.
[0013] In some embodiments, the CRSIPR nuclease is a Cas9 nuclease or a variant thereof.
[0014] In some embodiments, the Cas9 nuclease comprises the sequence shown in SEQ ID NO:14; and / or
[0015] The Cas9 nuclease variant is an nCas9 nuclease, which contains the sequence shown in SEQ ID NO:15.
[0016] In some embodiments, the reverse enzyme is M-MLV reverse transcriptase;
[0017] Preferably, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted, and it contains the sequence shown in SEQ ID NO:16.
[0018] In some embodiments, the reverse transcriptase and / or the guide editing fusion protein may also comprise one or more RNA aptamer-binding protein sequences.
[0019] In some embodiments, the RNA aptamer-binding protein is an MCP protein;
[0020] Optionally, the MCP protein comprises the sequence shown in SEQ ID NO:17.
[0021] In some embodiments, the guide editing fusion protein comprises a sequence selected from any of SEQ ID NO:1-3,7-11.
[0022] In some embodiments, it also includes a guide RNA (nicking sgRNA) targeting a non-target strand and / or an expression construct containing a nucleotide sequence encoding the guide RNA targeting a non-target strand.
[0023] In some embodiments, the guide RNA comprises a backbone sequence as shown in SEQ ID NO:4; and / or
[0024] The guide RNA that targets the non-target strand contains a backbone sequence as shown in SEQ ID NO:4.
[0025] In some implementations, the reverse transcription template (RTT) sequence and the primer binding site (PBS) sequence are directly linked in the circular RNA;
[0026] Preferably, the reverse transcription template (RTT) sequence is located at the 5' end of the primer binding site (PBS) sequence.
[0027] In some embodiments, the circular RNA comprises at least one RNA aptamer sequence, and optionally, the aptamer comprises MS2.
[0028] In some embodiments, the primer binding site sequence is complementary to at least a portion of the target sequence;
[0029] Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the target strand caused by the nick, particularly to the nucleotide sequence at the 3' end of the 3' free single strand.
[0030] In some embodiments, the expression construct containing the nucleotide sequence encoding the circular RNA comprises the following coding sequence: 5'-first ribozyme-first cyclic arm-RTT-PBS-second cyclic arm-second ribozyme-3';
[0031] Optionally, after transcription into RNA within the cell, the first and second ribozymes can self-cleave to produce 5'-first cyclic arm-RTT-PBS-second cyclic arm-3', and the first and second cyclic arms can connect with each other to form a circular RNA containing RTT and PBS.
[0032] In some embodiments, the coding sequence of the first ribozyme includes the sequence shown in SEQ ID NO:22, the coding sequence of the second ribozyme includes the sequence shown in SEQ ID NO:23, the first cyclic arm includes the nucleotide sequence shown in SEQ ID NO:24, and the second cyclic arm includes the nucleotide sequence shown in SEQ ID NO:25.
[0033] In some implementations, the nucleotide sequence encoding the circular RNA is expressed by a U6 promoter.
[0034] In some embodiments, the reverse-guided editing system further includes the MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor;
[0035] Optionally, the MLH1dn protein factor comprises the sequence shown in SEQ ID NO:26.
[0036] In some embodiments, the reverse-guided editing system further includes a helicase and / or an expression construct containing a nucleotide sequence encoding the helicase;
[0037] Optionally, the helicase comprises the sequence shown in SEQ ID NO:27;
[0038] Preferably, the helicase and reverse transcriptase form a fusion protein; more preferably, the fusion protein comprises the sequence shown in SEQ ID NO:12.
[0039] This invention provides the application of the reverse-guided editing system described above in either (A) or (B) below:
[0040] (A) Editing of the genome sequence of an organism or its cells;
[0041] (B) Products that are prepared by editing the genome sequence of an organism or biological cell.
[0042] The present invention provides a kit comprising the reverse bootstrapping editing system as described above.
[0043] The present invention provides a method for producing genetically modified cells, comprising introducing a reverse-guided editing system as described above into at least one of the cells, thereby resulting in modification of the genome sequence of the at least one cell.
[0044] In some embodiments, the cells are derived from microorganisms, animals, or plants;
[0045] The microorganisms include bacteria and fungi;
[0046] The animals include mammals and poultry;
[0047] The plants include monocotyledons and dicotyledons;
[0048] Optionally, the mammals include humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats;
[0049] Optionally, the poultry includes chickens, ducks, and geese;
[0050] Optionally, the monocotyledonous plants include rice, corn, wheat, sorghum, and barley, and the dicotyledonous plants include soybean, peanut, Arabidopsis thaliana, cotton, and rapeseed.
[0051] The effects of the invention
[0052] This invention constructs a reverse guided editing system based on target strand nickases such as D10A-Cas9, further reducing editing byproducts and improving editing safety. Simultaneously, circular RNA is introduced for stable expression of RTT and PBS sequences, resulting in low off-target rates and significantly improved editing efficiency, enhancing the application potential of the reverse guided editing system in disease treatment, crop improvement, and biological research. Attached Figure Description
[0053] Figures 1a-1e : Schematic diagram of different reverse-guided editing systems.
[0054] Figure 2 : Schematic diagram of different reverse-guided editing system carriers.
[0055] Figure 3 The iPE system, a reverse-guided editing system based on pegRNA, generates reverse-guided editing in human cells.
[0056] Figure 4 The nu-iPE system, which utilizes pegRNA for reverse guided editing, demonstrates more efficient reverse guided editing than iPE in the human cell reporter system.
[0057] Figure 5 The nu-iPE system, which utilizes pegRNA for reverse guided editing, demonstrates more efficient reverse guided editing than iPE in human cells.
[0058] Figure 6 The circular RNA reverse-guided editing system ciPE produces more efficient reverse-guided editing in human cells than iPE while generating fewer InDels byproducts.
[0059] Figure 7 The nu-ciPE system, which utilizes circular RNA for reverse-guided editing, also produces highly efficient reverse-guided editing in human cells, but it produces more InDels byproducts.
[0060] Figure 8 Helicase-assisted hciPE produces higher reverse-guided editing efficiency in human cells than ciPE.
[0061] Figure 9 A comparison between the reverse bootstrap editor hciPE and other current editors.
[0062] Figure 10 The reverse bootstrap editor exhibits low off-target activity. Detailed Implementation
[0063] Various exemplary embodiments, features, and aspects of the present invention will be described in detail below. The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0064] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art should understand that the present invention can be practiced without certain specific details. In other instances, methods, means, apparatus, and steps well known to those skilled in the art have not been described in detail in order to highlight the spirit of the present invention.
[0065] Unless otherwise stated, all units used in this specification are international standard units, and all numerical values and ranges appearing in this invention should be understood to include systematic errors that are unavoidable in industrial production.
[0066] In this specification, the word "may" has two meanings: to perform a certain process and not to perform a certain process.
[0067] In this specification, references to "some specific / preferred embodiments," "other specific / preferred embodiments," "implementation," etc., refer to specific elements (e.g., features, structures, properties, and / or characteristics) related to that embodiment, which are included in at least one of the embodiments described herein and may or may not be present in other embodiments. Furthermore, it should be understood that these elements may be combined in any suitable manner in various embodiments.
[0068] As used in this article, “containing,” “having,” or “including” includes “containing,” “mainly composed of,” “substantially composed of,” and “composed of”; “mainly composed of,” “substantially composed of,” and “composed of” are subordinate concepts of “containing,” “having,” or “including.”
[0069] While the disclosure supports the definition of the term "or" as merely a substitute and "and / or", the term "or" in the claims means "and / or" unless expressly stated as merely a substitute or as mutually exclusive. In this specification, the term "and / or", when used to connect two or more options, should be understood to mean any one or any two or more of the options.
[0070] In this specification, "optional" and "optionally" mean that the events or circumstances described below may or may not occur, and the description includes both cases where the events or circumstances occur and cases where the events or circumstances do not occur.
[0071] In this specification, the range of values referred to as "value A to value B" refers to the range including the endpoint values A and B.
[0072] In this specification, the term "polynucleotide" refers to a polymer composed of nucleotides. Polynucleotides can be in the form of individual fragments or as a component of a larger nucleotide sequence structure, derived from a nucleotide sequence isolated at least once in number or concentration, and capable of being recognized, manipulated, and recovered using standard molecular biology methods (e.g., using cloning vectors). This also includes an RNA sequence (i.e., A, T, G, C) when a nucleotide sequence is represented by a DNA sequence (i.e., A, U, G, C), where "U" replaces "T". In other words, "polynucleotide" refers to a polymer of nucleotides removed from other nucleotides (individual fragments or entire fragments), or it can be a component or part of a larger nucleotide structure, such as an expression vector or a polycistronic sequence. Polynucleotides include DNA, RNA, and cDNA sequences.
[0073] In this specification, the term "genome" as used herein encompasses not only chromosomal DNA present in the cell nucleus, but also organelle DNA present in subcellular components of the cell, such as mitochondria and plastids.
[0074] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein and refer to amino acid polymers of any length. The polymer may be linear or branched, may contain modified amino acids, and may be separated by non-amino acid segments. The term also includes amino acid polymers that have been modified (e.g., through disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with labeled components).
[0075] In this specification, the term "expression construct" can refer to a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (such as mRNA), for example, RNA transcribed in vitro. For use in mammals such as humans, the expression construct is preferably a viral vector, such as adeno-associated virus (AAV), lentivirus, or adenovirus vector, with AAV vectors being the most preferred.
[0076] Prime editing (PE) systems are genome editing strategies consisting of a prime editor. Currently, conventional PEs use a Cas9 nickase (H840A) fused with an engineered Moloney murine leukemia virus (M-MLV) reverse transcriptase (RT). The RT is programmed with the desired prime editing guide RNA (pegRNA). The pegRNA contains a reverse transcription RNA template (RTT) sequence comprising sgRNA, a primer binding site (PBS), and the reverse transcriptase sequence, used to generate the desired edit at the target site.
[0077] As used herein, the terms “guide RNA,” “directing RNA,” and “sgRNA” are generally used interchangeably. A sgRNA comprises a directing sequence and a double-strand-forming region (e.g., the double-strand-forming region of crRNA, which may also be referred to as a crRNA repeat sequence). A sgRNA is a polynucleotide that can specifically target a target sequence and can form a complex with a nucleic acid programmable nucleotide binding domain protein (e.g., Cas9).
[0078] In this specification, "downstream of the target cleavage site" refers to the downstream of the CRSIPR nuclease cleavage site. In this invention, the cleavage site of the Cas9 protein is usually located three nucleotides upstream of the PAM sequence.
[0079] I. Guided Editing System
[0080] This invention provides a reverse-guided editing system, the system comprising:
[0081] (i)a) An expression construct containing a CRSIPR nuclease and / or a nucleotide sequence encoding the CRSIPR nuclease, and an expression construct containing a reverse transcriptase and / or a nucleotide sequence encoding the reverse transcriptase; or
[0082] b) A guide editing fusion protein and / or an expression construct containing a nucleotide sequence encoding the guide editing fusion protein, wherein the guide editing fusion protein comprises a CRSIPR nuclease and / or a reverse transcriptase;
[0083] (ii) a guide RNA (pegRNA) targeting a target sequence in genomic DNA and / or an expression construct containing a nucleotide sequence encoding said guide RNA; and / or
[0084] (iii) An expression construct containing a circular RNA with a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or a nucleotide sequence encoding the circular RNA, and an expression construct containing a guide RNA (sgRNA) targeting a sequence in genomic DNA and / or a nucleotide sequence encoding the guide RNA.
[0085] In some alternative implementations, the CRSIPR nuclease is a target strand nickase or a double-stranded endonuclease.
[0086] In this specification, the "reverse-guided editing system" refers to a combination of components required for reverse transcription-based genome editing within cells. The individual components of this system, such as CRISPR nucleases, reverse transcriptases, guided editing fusion proteins, gRNAs, circular RNAs, or their expression constructs, can exist independently or in any combination as a composition. The individual components of the system can be packaged independently or together into viruses, virus-like particles, virions, liposomes, vesicles, exosomes, liposome nanoparticles (LNPs), etc.
[0087] In some alternative embodiments, the CRSIPR nuclease is a CRSIPR nickase, meaning it forms a nick only on one strand of the double-stranded nucleic acid molecule and does not completely cleave the double-stranded nucleic acid. In some embodiments, the CRISPR nickase forms a nick only on the strand containing the target sequence in the double-stranded nucleic acid molecule (i.e., the target strand), and is also called a target strand nickase. The complementary strand of the strand containing the target sequence in the double-stranded nucleic acid molecule is called the non-target strand.
[0088] In this specification, the term "target sequence" refers to a sequence of approximately 20 nucleotides in length characterized by a PAM (pre-interstitial sequence adjacent motif) sequence flanking the genome at 5' or 3'. Typically, the PAM is necessary for the recognition of the target sequence by a complex formed by a CRISPR nuclease or a variant thereof with guide RNA.
[0089] In some specific implementations, the "target strand" or "target sequence" refers to a DNA strand in the double-stranded DNA that is complementary to the spacer sequence of the guide RNA. The nucleotide sequence of this strand is complementary to the spacer region of the guide RNA and can specifically bind to the guide RNA through base pairing, thereby guiding the Cas protein to target the region of the double-stranded DNA. This strand is another strand complementary to the DNA strand containing the PAM (protospacer adjacent motif) sequence.
[0090] In some embodiments, the Cas9 nuclease comprises the sequence shown in SEQ ID NO:14; the Cas9 nuclease variant is an nCas9 nuclease (named nCas9-D10A or D10A-Cas9 herein) which comprises the sequence shown in SEQ ID NO:15.
[0091] In some embodiments of the present invention, the reverse transcriptase is a viral reverse transcriptase, such as M-MLV reverse transcriptase derived from Moloney mouse leukemia virus. Further, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted. In some embodiments, the M-MLV reverse transcriptase with a mutated or deleted RNase H domain comprises the sequence shown in SEQ ID NO:16.
[0092] In some embodiments, the reverse transcriptase and / or the guided editing fusion protein of the present invention may further comprise one or more RNA aptamer-binding protein sequences. In some embodiments, the reverse transcriptase of the present invention further comprises one or more RNA aptamer-binding protein sequences. In some embodiments, the guided editing fusion protein of the present invention further comprises one or more RNA aptamer-binding protein sequences.
[0093] In this specification, "RNA aptamer-binding protein" refers to a protein capable of specifically binding to an "RNA aptamer" having a specific sequence and structure. Proteins containing one or more RNA aptamer-binding protein sequences can thus be recruited within the cell to nucleic acid molecules containing the corresponding RNA aptamer, and vice versa. Many such combinations of specifically binding "RNA aptamer-binding proteins" and "RNA aptamers" are known in the art.
[0094] In some embodiments, the RNA aptamer-binding protein is the MCP protein. Based on this, the present invention achieves in-situ recruitment of the MCP-RT fusion protein by attaching an MS2 aptamer and utilizing the interaction between the MS2 aptamer and the MCP protein. The MCP protein can specifically bind MS2 or multiple MS2 molecules in tandem. Exemplarily, the MCP comprises the sequence shown in SEQ ID NO:17. Exemplarily, the MS2 comprises the sequence shown in SEQ ID NO:18.
[0095] In this invention, when connecting various elements such as MCP, MS2, CRSIPR nuclease, and reverse transcriptase, they can be directly connected or connected through adapter sequences.
[0096] In some embodiments, the linker may be a non-functional amino acid sequence of one or more amino acids without secondary or higher-order structures. In some embodiments, the linker may be about 5 to 100 amino acids in length, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. For example, the linker may be a flexible linker. In some specific embodiments, the linker comprises the sequence shown in SEQ ID NO:19.
[0097] In some embodiments of the present invention, depending on the location of the DNA to be edited, the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein of the present invention may also include various localization sequences, such as nuclear localization sequences, cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.
[0098] In some exemplary embodiments of the present invention, the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein of the present invention may further comprise one or more nuclear localization sequences (NLS). Generally, one or more NLS in the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein should have sufficient strength to drive the accumulation of the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein in the nucleus of the cell to achieve its gene-editing function. Generally, the strength of nuclear localization activity is determined by the number, location, one or more specific NLS used, or a combination of these factors in the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein. Exemplarily, the NLS may comprise the sequence shown in SEQ ID NO:20 or SEQ ID NO:21.
[0099] In some embodiments, the guide editing fusion protein comprises a sequence selected from any of SEQ ID NO:1-3,7-11.
[0100] To achieve efficient expression in different organisms, in this invention, the nucleotide sequence encoding the CRISPR nuclease, reverse transcriptase, or guided editing fusion protein can be codon-optimized for the organism whose genome is to be modified. "Codon optimization" refers to configuring the nucleotide sequence encoding the polypeptide to contain codons preferred by the host cell or organism to improve gene expression and translation efficiency in the host cell or organism.
[0101] The guide RNA (sgRNA) described in this invention refers to an sgRNA compatible with the CRISPR nuclease. The sgRNA typically comprises a scaffold sequence and a guide sequence (also called a seed sequence or spacer sequence). The guide sequence is configured to have sufficient sequence identity (preferably 100%) with the target sequence, thereby enabling sequence-specific targeting by binding to the complementary strand of the target sequence through base pairing. The scaffold sequence typically depends on the CRISPR nuclease used. In some alternative embodiments of this invention, the sgRNA may also target a non-target strand (i.e., the non-edited strand), referred to herein as nickingsgRNA. Nicking sgRNA is a specially designed single-guide RNA used to guide the Cas9 protein in the CRISPR-Cas9 system to generate a single-strand break (SSB) on the target DNA strand, rather than a double-strand break (DSB). Specifically, the nickingsgRNA targets a position in the non-target strand approximately 50 base pairs from the PAM sequence.
[0102] In some exemplary embodiments, the backbone sequence of the corresponding sgRNA for the Cas9 nuclease or variants thereof of the present invention comprises the sequence shown in SEQ ID NO:4.
[0103] In some embodiments, the pegRNA comprises sgRNA, a reverse transcription modulus (RTT) sequence, and a primer binding site (PBS) sequence (RTT-PBS sequence). Exemplarily, the pegRNA comprises a sequence as shown in SEQ ID NO:5 or SEQ ID NO:6.
[0104] In some implementations, the reverse transcription modulus (RTT) sequence and primer binding site (PBS) sequence can also be expressed via circular RNA, i.e., the circular RNA contains directly linked RTT and PBS sequences.
[0105] The RTT sequence is the reverse complementary sequence of the three bases from the 3' end of the target sequence and the subsequent continuous genomic sequence, in which the target mutation is introduced. This sequence serves as a reverse transcription template for reverse transcriptase, producing cDNA, which is then used as a repair template to repair the genomic DNA. The RTT sequence length can further be 8–34 bp.
[0106] Typically, the RTT sequence contains only the target mutation site (i.e. the site where the mutation is desired), and a mutant base is introduced at the target mutation site.
[0107] The PBS sequence (primer binding site sequence) is the reverse complementary sequence (1 ≤ n < 17) of the target sequence from the nth to the 17th base from the 5' end of the target sequence. Further, it is configured to be complementary to at least a portion of the target sequence (preferably perfectly paired with at least a portion of the target sequence). Preferably, the primer binding site sequence is complementary to at least a portion of the 3' free single strand in the target strand (TS) caused by cleavage (preferably perfectly paired with at least a portion of the 3' free single strand), particularly complementary to the nucleotide sequence at the 3' end of the 3' free single strand (preferably perfectly paired). When the 3' free single strand of the strand binds to the primer binding sequence through base pairing, the 3' free single strand can act as a primer, using the reverse transcription template sequence adjacent to the primer binding site sequence as a template. Reverse transcription is performed under the action of reverse transcriptase in the guide editing fusion protein or reverse transcriptase recruited by RNA aptamer binding protein, extending the DNA sequence corresponding to the reverse transcription template (RTT) sequence.
[0108] In some optional embodiments, the design methods or principles of the RTT sequence and the PBS sequence may refer to the design methods or principles of RTT sequences and PBS sequences of pegRNA in prime editing (PE) techniques reported in the prior art.
[0109] In other specific embodiments, the circular RNA further comprises one or more RNA aptamer sequences, such as the MS2 sequence. The one or more RNA aptamer sequences in the circular RNA can be used to recruit reverse transcriptases containing RNA aptamer-binding proteins in cells, or to recruit the circular RNA to the CRISPR nuclease-gRNA-genomic DNA complex in cells.
[0110] In some embodiments, the expression construct containing the nucleotide sequence encoding the circular RNA includes the following coding sequence: 5'-first ribozyme-first cyclic arm-RTT-PBS-second cyclic arm-second ribozyme-3'. Upon intracellular transcription into RNA, the first and second ribozymes self-cleave to generate "5'-first cyclic arm-RTT-PBS-second cyclic arm-3'", and the first and second cyclic arms can connect to form a circular RNA containing RTT and PBS. In some embodiments, one or more RNA aptamer sequences, such as MS2 sequences, may also be flanked by the RTT-PBS.
[0111] In some embodiments, the first ribozyme comprises the nucleotide sequence shown in SEQ ID NO:22, and the second ribozyme comprises the nucleotide sequence shown in SEQ ID NO:23. In some embodiments, the first cyclic arm comprises the nucleotide sequence shown in SEQ ID NO:24, and the second cyclic arm comprises the nucleotide sequence shown in SEQ ID NO:25.
[0112] In some optional embodiments, the reverse-guided editing system further includes the MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor. The MLH1dn protein factor is a dominant-negative mutant of MLH1; optionally, the MLH1dn protein factor comprises the sequence shown in SEQ ID NO:26.
[0113] In some alternative embodiments, the reverse editing system further comprises a helicase and / or an expression construct containing a nucleotide sequence encoding the helicase. In this invention, the helicase is used to untie double helices formed by two complementary DNA strands, double helices formed by DNA and complementary RNA, and other double strands containing DNA.
[0114] In some embodiments, the helicase is derived from porcine circovirus, Escherichia coli, or humans. Exemplarily, the helicase is derived from Escherichia coli and contains the sequence shown in SEQ ID NO:27.
[0115] II. Application of Reverse Editing Systems
[0116] This invention provides the application of the reverse-guided editing system of this invention in the following (A) or (B):
[0117] (A) Editing of the genome sequence of an organism or its cells;
[0118] (B) Products that are prepared by editing the genome sequence of an organism or biological cell.
[0119] In some specific embodiments, the cells are microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats, and poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis thaliana, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.
[0120] III. Methods for modifying target sequences in the cell genome, methods for generating genetically modified cells, and genetically modified... organisms
[0121] This invention provides a method for producing genetically modified cells, the method comprising introducing the reverse-guided editing system of the invention into at least one cell, thereby resulting in modification of the genomic sequence of the at least one cell. The modification includes substitution, deletion, and / or addition of one or more nucleotides. For example, the modification includes one or more substitutions selected from the following: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or includes the deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide deletions; and / or includes the insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide insertions. The modification can be located in or near the target sequence of the sgRNA, for example, upstream of the target sequence.
[0122] In another aspect, the present invention provides a method for generating genetically modified cells, comprising introducing the guided editing system of the present invention into the cells.
[0123] In another aspect, the present invention also provides genetically modified organisms comprising genetically modified cells or their progeny cells produced by the method of the present invention.
[0124] In this invention, the modification can be located anywhere in the genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter or enhancer region, thereby achieving modification of gene function or gene expression. The modification in the cell genome sequence can be detected using T7EI, PCR / RE, or sequencing methods.
[0125] In this invention, the reverse guided editing system can be introduced into cells using various methods well known to those skilled in the art. For example, methods for introducing the guided editing system of this invention into cells include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, and other viruses), gene gun method, PEG-mediated protoplast transformation, and Agrobacterium-mediated transformation. Cells that can be gene-edited using the methods of this invention can be derived from, for example, microorganisms such as bacteria and fungi; animals, including mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants, including monocots and dicots, with monocots such as rice, corn, wheat, sorghum, and barley, and dicots such as soybeans, peanuts, Arabidopsis, rapeseed, and cotton. In some preferred embodiments, the cells are derived from humans.
[0126] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.
[0127] In other embodiments, the method of the present invention can also be performed in vivo. For example, the cells are cells within an organism, and the system of the present invention can be introduced into the cells in vivo via, for example, a viral or Agrobacterium-mediated method.
[0128] IV. Methods for producing genetically modified plants and plant breeding methods
[0129] This invention provides a method for producing genetically modified plants, comprising introducing the reverse-guided editing system of this invention into at least one of the plants, thereby resulting in modifications in the genome of the at least one plant. The modifications include substitutions, deletions, and / or additions of one or more nucleotides. For example, the modification includes one or more substitutions selected from the following: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or includes the deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide deletions; and / or includes the insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotide insertions.
[0130] In some embodiments, the method further includes screening plants with desired modifications from the at least one plant.
[0131] In the method of this invention, the reverse guided editing system can be introduced into plants using various methods well known to those skilled in the art. Methods for introducing the guided editing system of this invention into plants include, but are not limited to: gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the guided editing system is introduced into plants via transient transformation.
[0132] In the method of this invention, genome modification can be achieved simply by introducing or generating relevant proteins and RNA in plant cells, and the modification can be stably inherited without the need for stable transformation of plants with exogenous polynucleotides encoding components of the reverse-guided editing system. This avoids the potential off-target effects of a stably existing (continuously generated) editing system and also avoids the integration of exogenous nucleotide sequences into the plant genome, thus providing higher biosafety.
[0133] In some preferred embodiments, the introduction is performed in the absence of selection pressure, thereby avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0134] In some embodiments, the introduction includes converting the reverse-guided editing system of the present invention into isolated plant cells or tissues, and then regenerating the converted plant cells or tissues into complete plants. Preferably, the regeneration is performed without selection pressure, i.e., without using any selection agents targeting the selection genes carried on the expression vector during tissue culture. Not using selection agents can improve the regeneration efficiency of the plants, resulting in modified plants free of exogenous nucleotide sequences.
[0135] In other embodiments, the reverse-guided editing system of the present invention can be applied to specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young spikelets, or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.
[0136] In some embodiments of the present invention, in vitro expressed proteins and / or in vitro transcribed RNA molecules (e.g., the expression construct is an in vitro transcribed RNA molecule) are directly transformed into the plant. The proteins and / or RNA molecules can be guided to edit within plant cells and are subsequently degraded by the cells, avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0137] Therefore, in some embodiments, using the methods of the present invention to genetically modify and breed plants can yield plants whose genomes are free of foreign polynucleotide integration, i.e., non-transgene-free modified plants.
[0138] In some embodiments of the invention, the modified genomic region is associated with plant traits such as agronomic traits, whereby the modification results in the plant having altered (preferably improved) traits, such as agronomic traits, relative to the wild-type plant.
[0139] In some embodiments, the method further includes the step of screening plants with desired modifications and / or desired traits such as agronomic traits.
[0140] In some embodiments of the invention, the method further includes obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring have the desired modification and / or desired traits such as agronomic traits.
[0141] In another aspect, the present invention also provides genetically modified plants or their offspring or portions thereof, wherein said plants are obtained by the methods described above. In some embodiments, the genetically modified plants or their offspring or portions thereof are non-GMO. Preferably, the genetically modified plants or their offspring have the desired genetic modification and / or desired traits such as agronomic traits.
[0142] In another aspect, the present invention also provides a plant breeding method, comprising crossing a genetically modified first plant obtained by the method described above with a second plant that does not contain the modification, thereby introducing the modification into the second plant. Preferably, the genetically modified first plant has desired traits such as agronomic traits.
[0143] "Agronomic traits" specifically refer to measurable parameters of crop plants, including but not limited to: leaf greenness, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total nitrogen content of plants, nitrogen content of fruits, nitrogen content of seeds, nitrogen content of plant vegetative tissues, total free amino acid content of plants, free amino acid content of fruits, free amino acid content of seeds, free amino acid content of plant vegetative tissues, total protein content of plants, protein content of fruits, protein content of seeds, protein content of plant vegetative tissues, herbicide resistance and drought resistance, nitrogen uptake, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt tolerance, and tiller number, etc.
[0144] V. Therapeutic Uses
[0145] This invention also provides the application of the reverse-guided editing system of this invention in disease treatment.
[0146] By modifying disease-related genes using the reverse-guided editing system of this invention, it is possible to achieve upregulation, downregulation, inactivation, activation, or mutation correction of disease-related genes, thereby achieving disease prevention and / or treatment. For example, the genomic modifications described in this invention can be located within the protein-coding region of the disease-related gene, or, for example, within gene expression regulatory regions such as promoter regions or enhancer regions, thereby enabling modifications to the function or expression of the disease-related gene. Therefore, the modification of disease-related genes described herein includes modifications to the disease-related gene itself (e.g., protein-coding regions), as well as modifications to its expression regulatory regions (e.g., promoters, enhancers, introns, etc.).
[0147] "Disease-associated" genes are any genes that produce transcriptional or translational products at abnormal levels or in abnormal forms in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control groups. In cases where altered expression is associated with the onset and / or progression of the disease, it can be a gene expressed at abnormally high levels; it can also be a gene expressed at abnormally low levels. Disease-associated genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked to one or more genes responsible for the etiology of the disease in disequilibrium. Such mutations or genetic variations are, for example, single nucleotide variants (SNVs). The transcribed or translated products can be known or unknown and can be at normal or abnormal levels.
[0148] Therefore, the present invention also provides a method for treating a disease in a subject in need, comprising delivering an effective amount of the guided editing system of the present invention to the subject to modify a gene associated with the disease. The present invention also provides the use of the guided editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need, wherein the guided editing system is used to modify a gene associated with the disease. The present invention also provides a pharmaceutical composition for treating a disease in a subject in need, comprising the guided editing system of the present invention and optionally a pharmaceutically acceptable vector, wherein the guided editing system is used to modify a gene associated with the disease.
[0149] Preferably, the "object" referred to in this invention is a mammal, such as a human.
[0150] In some implementations, the guided editing system described in this invention is used to introduce point mutations into nucleic acids.
[0151] In some embodiments, the guided editing system described herein is used to correct genetic defects, such as in correcting point mutations that result in loss of function in a gene product. In some embodiments, the genetic defect is associated with a disease or condition (e.g., lysosomal storage disease or metabolic disease, such as, for example, type 1 diabetes). In some embodiments, the methods provided herein can be used to introduce inactive point mutations into a gene or allele encoding a gene product associated with a disease or condition.
[0152] In some embodiments, the purpose of the schemes described in this invention is to treat diseases associated with or caused by point mutations, which can be corrected using the guided editing system provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neonatal disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.
[0153] In some embodiments, the purposes of the solutions described in this invention are for the treatment of mitochondrial diseases or disorders. As used herein, "mitochondrial disease" refers to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological disorders, loss of motor control, muscle weakness and pain, gastrointestinal disorders and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, visual / hearing problems, lactic acidosis, developmental delay, and susceptibility to infection.
[0154] Examples of diseases described in this invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat amplification disorders, hearing disorders, gene-targeted therapy for non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, β-thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, and schizophrenia. Other diseases that can be treated by correcting point mutations or introducing inactive mutations into disease-related genes are known to those skilled in the art, and therefore this disclosure is not limited in this respect. In addition to the diseases exemplarily described in this invention, other related diseases can also be treated using the strategies and guided editing systems provided in this invention, and this application will be apparent to those skilled in the art. The diseases or targets to which this invention can be applied are related to the base editing systems listed in WO2015089465A1 (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), and WO2021155065A1 (PCT / US2021 / 015580).
[0155] The administration of the guided editing system or pharmaceutical composition of the present invention can be tailored to the patient's or subject's weight and species. The frequency of administration is within medically or veterinary limits. It depends on conventional factors including the patient's or subject's age, sex, general health condition, other conditions, and the specific symptom or condition being addressed.
[0156] VI. Reagent Kit
[0157] The present invention also provides a kit for use with the methods of the present invention, the kit comprising the components of the reverse-guided editing system of the present invention. The kit may also contain reagents for introducing the reverse-guided editing system into an organism or somatic cells. The kit generally includes a label indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise accompanied by the kit.
[0158] Example
[0159] The embodiments of the present invention will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of the invention. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer are followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products.
[0160] Materials and Methods
[0161] 1. Carrier Construction
[0162] The vector was constructed using the existing guide editor vector skeleton in the laboratory. The main vectors constructed were:
[0163] iPEs:
[0164] (1) iPE2 contains iPE-nCas9-D10A and pegRNA.
[0165] The iPE-nCas9-D10A sequence is shown in SEQ ID NO:2.
[0166] The pegRNA sequence is shown in SEQ ID NO:5;
[0167] (2)iPE3 contains iPE-nCas9-D10A, pegRNA, and nicking sgRNA
[0168] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the rest of the sequence is the same as iPE2;
[0169] (3) iPE4 contains iPE-nCas9-D10A, pegRNA, and MLH1dn protein factor.
[0170] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPE2;
[0171] The iPE-nCas9-H840A is a direct replacement of the nCas9-D10A in the iPE2 with the nCas9-H840A.
[0172] (4) iPE5 contains iPE-nCas9-D10A, pegRNA, MLH1dn protein factor, and nicking sgRNA.
[0173] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPE2;
[0174] (5) iPEmax2 contains iPEmax-nCas9-D10A and epegRNA
[0175] The iPEmax-nCas9-D10A sequence is shown in SEQ ID NO:3.
[0176] The epegRNA sequence is shown in SEQ ID NO:6;
[0177] (6) iPEmax3 contains iPEmax-nCas9-D10A, pegRNA, and nicking sgRNA.
[0178] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the rest of the sequence is the same as iPEmax2;
[0179] (7) iPEmax4 contains iPEmax-nCas9-D10A, pegRNA, and MLH1dn protein factor.
[0180] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPEmax2.
[0181] (8) iPEmax5 contains iPEmax-nCas9-D10A, pegRNA, MLH1dn protein factor, and nicking sgRNA.
[0182] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as iPEmax2.
[0183] nu-iPEs:
[0184] (1) nu-iPE2 contains iPE-WT-Cas9 and pegRNA
[0185] The iPE-WT-Cas9 sequence is shown in SEQ ID NO:1.
[0186] The pegRNA sequence is shown in SEQ ID NO:5;
[0187] (2) nu-iPEmax2 contains iPEmax-WT-Cas9 and epegRNA
[0188] The iPEmax-WT-Cas9 sequence is shown in SEQ ID NO:29.
[0189] The epegRNA sequence is shown in SEQ ID NO:6;
[0190] (3) nu-iPEmax4 contains iPEmax-WT-Cas9, epeRNA, and MLH1dn protein factor.
[0191] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as nu-iPEmax2.
[0192] ciPEs:
[0193] (1) ciPE2 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, and sgRNA.
[0194] The nCas9-D10A sequence is shown in SEQ ID NO:10.
[0195] The reverse transcriptase M-MLV RTΔRNase H sequence is shown in SEQ ID NO:11.
[0196] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.
[0197] The sgRNA sequence is shown in SEQ ID NO:4.
[0198] (2) ciPE3 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and nicking sgRNA. The nicking sgRNA sequence is shown in SEQ ID NO:4, and the remaining sequences are the same as ciPE2.
[0199] (3) ciPE4 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.
[0200] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as ciPE2.
[0201] (4) ciPE5 contains nCas9-D10A, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, nicking sgRNA, and MLH1dn protein factor.
[0202] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as ciPE2.
[0203] The equivalent of ciPE-nCas9-H840A is to replace nCas9-D10A in ciPE2 with nCas9-H840A.
[0204] nu-ciPE:
[0205] (1) nu-ciPE2 contains WT-Cas9, reverse transcriptase M-MLV RTΔRNase H, circular RNA, and sgRNA.
[0206] The WT-Cas9 sequence is shown in SEQ ID NO:9.
[0207] The reverse transcriptase M-MLV RTΔRNase H sequence is shown in SEQ ID NO:11.
[0208] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.
[0209] The sgRNA sequence is shown in SEQ ID NO:4.
[0210] (2) nu-ciPE4 contains WT-Cas9, reverse transcriptase M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.
[0211] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as nu-ciPE2.
[0212] hciPEs(ciPEs+Rep-X):
[0213] (1) hciPE2 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, and sgRNA.
[0214] The sequence of nCas9-D10A is shown in SEQ ID NO:10.
[0215] The sequence of MCP-Helicase-M-MLV RTΔRNase H is shown in SEQ ID NO:12.
[0216] The circular RNA (U6-5'+3'MS2-CirRNA) sequence is shown in SEQ ID NO:13.
[0217] The sgRNA sequence is shown in SEQ ID NO:4.
[0218] (2) hciPE3 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and nicking sgRNA.
[0219] The nicking sgRNA sequence is shown in SEQ ID NO:4, and the remaining sequences are the same as hciPE2;
[0220] (3) hciPE4 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.
[0221] The MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as hciPE2.
[0222] (4) hciPE5 contains nCas9-D10A, MCP-Helicase-M-MLV RTΔRNase H, circular RNA, sgRNA, and MLH1dn protein factor.
[0223] The nicking sgRNA sequence is shown in SEQ ID NO:4, the MLH1dn protein factor sequence is shown in SEQ ID NO:26, and the remaining sequences are the same as hciPE2.
[0224] For details on the editing systems split IPE2, split IPEmaxsΔR (split IPEmax2ΔR, split IPEmax3ΔR, split IPEmax4ΔR, split IPEmax5ΔR), PAMless-PEs, twinPE, upPE, downPE, SpG-PEs (SpG-PE2, SpG-PE3, SpG-PE4, SpG-PE5), and SpRY-PEs (SpRY-PE2, SpRY-PE3, SpRY-PE4, SpRY-PE5), please refer to Liang, R., Wang, S., Cai, Y. et al. Circular RNA-mediated inverse primeediting in human cells. Nat Commun 16, 5057 (2025), which is incorporated herein by reference.
[0225] 2. Human cell culture and transformation
[0226] 2.1 Thawing cells
[0227] (1) Remove HEK293T, U2OS and K562 cell lines from liquid nitrogen and thaw them in a 37°C water bath for 2 minutes.
[0228] (2) Remove the cells from the cryopreservation tube and slowly add them to a 15 mL centrifuge tube containing 10 mL of culture medium.
[0229] (3) Centrifuge at 500×g for 3 minutes to precipitate cells and remove the supernatant.
[0230] (4) Add 10 mL of culture medium to the precipitate, transfer it to a T175 culture flask, and place it in a 37°C incubator containing 5% CO2.
[0231] (5) Replace the culture medium after 24 hours and continue culturing in an incubator containing 5% CO2 at 37°C.
[0232] 2.2 Cell Culture
[0233] (1) Carefully aspirate the cell culture medium from the T175 culture flask.
[0234] (2) Gently wash the cells with 10 mL of PBS solution.
[0235] (3) Treat the cells with 2 mL of trypsin: Place the T175 cell culture flask in a 37°C incubator for 2 minutes to allow the trypsin to fully dissociate the cells.
[0236] (4) Add 10 mL of cell culture medium to the T175 cell culture flask to resuspend the cells and inactivate trypsin.
[0237] (5) Use a 10mL pipette to blow the cell suspension up and down to obtain a single cell suspension.
[0238] (6) After diluting the cells at a ratio of 1:10, continue to culture them in an incubator at 37°C containing 5% CO2.
[0239] 2.3 Cell Deployment
[0240] (1) Take 15 μL of the single cell suspension from step (5) in 2.1 and mix it with 15 μL of trypan blue. Let it stand at room temperature for 1-2 minutes.
[0241] (2) Open a new cell counter. Add 10 μL of well-mixed cell and trypan blue solution to each of the counting chambers A and B, insert the counting plate into the cell counter and read the count.
[0242] (3) Calculate the required dilution factor so that the final cell count is 45,000 cells per well in 250 μL of cell culture medium in a 48-well plate. Dilution factor = (A count + B count) / (2 × 4 × 45000).
[0243] (4) Dilute the cells according to the dilution factor and use an adjustable width multichannel pipette to dispense the cell suspension into 48-well plates, adding 250uL of cell suspension to each well.
[0244] (5) Place the 48-well plate containing cells in a 37°C incubator and incubate for 16 to 24 hours before using it for liposome conversion.
[0245] 2.4 Transfecting cells
[0246] (1) Before transfecting cells, observe the cell growth status under a microscope, preferably with 60% coverage.
[0247] (2) Mix 300 ng of protein expression plasmid and 100 ng of RNA expression plasmid in Opti-MEM to make a total volume of 12.5 μL.
[0248] (3) Add 1 μL of Lipofectamine 2000 and 10-20 ng of copGFP expression plasmid to Opti-MEM to make the total volume 12.5 μL.
[0249] (4) Mix (2) and (3) to a total volume of 25 μL and incubate for 5-15 minutes. Then add the solution to the cell culture of a 48-well plate. Incubate at 37°C for 48-72 hours.
[0250] 3. Fluorescence observation of the reporting system
[0251] Human cells cultured at 37°C for 48 hours were observed under a fluorescence microscope or a laser confocal microscope. The efficiency of different guide editors was compared based on the fluorescence, and the fluorescence images were saved.
[0252] In the following embodiments, unless otherwise specified or marked, HEK293T cells were used.
[0253] 4. Human cell lysis and amplicon sequencing analysis
[0254] 4.1 Lysing human cells
[0255] (1) After culturing the cells in a 37°C incubator for 72 hours, carefully aspirate the cell culture medium using an adjustable width multichannel pipette.
[0256] (2) Gently add 300uL of PBS buffer, shake gently to wash away dead cells, and then aspirate the PBS buffer.
[0257] (3) Add 200uL of lysis buffer (with proteinase K added) and treat at 55℃ for 30 minutes.
[0258] (4) Pipette 50 μL of cell lysis buffer into a 96-well plate and treat at 95°C for 5 minutes before use.
[0259] 4.2 Amplicon Miseq Sequencing Analysis
[0260] (1) PCR amplification of human cell lysate was performed using Miseq first-round primers.
[0261] The first round of 15μL amplification system consisted of: 7.5μL 2×Phanta Max Master Mix, 4.5μL ddH2O, 1μL forward primer (10μM), 1μL reverse primer (10μM), and 1μL cell lysis buffer.
[0262] The first round of 15μL amplification conditions were as follows: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 15 s, 50-60℃ annealing for 15 s, 72℃ extension for 20 s, 34 cycles; and 72℃ full extension for 5 min.
[0263] (2) The first-round amplification product was diluted 5 times, and 1 μL was used as the template for the second-round PCR amplification. The amplification primers were the Miseq second-round sequencing primers containing barcode.
[0264] The second round of amplification system consisted of 25 μL: 12.5 μL 2×Phanta Max Master Mix, 9.5 μL ddH2O, 1 μL forward primer (10 μM), 1 μL reverse primer (10 μM), and 1 μL DNA template.
[0265] Second round amplification conditions: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 15 s, 50-60℃ annealing for 15 s, 72℃ extension for 20 s, 10 cycles; 72℃ full extension for 5 min.
[0266] (3) The PCR products were detected by 2% agarose gel electrophoresis, and the target fragment was recovered by gel extraction using AxyPrep DNA Gel Extraction Kit. The recovered products were quantitatively analyzed using NanoDrop ultra-micro spectrophotometer. 100 ng of the recovered products were mixed and sent to Beijing Qihe Biotechnology Co., Ltd. for amplicon sequencing analysis.
[0267] (4) After sequencing is completed, the editing type and editing efficiency of the products are compared and analyzed at different gene target sites in at least three repeated experiments.
[0268] 3. Target sites and target sequences
[0269]
[0270]
[0271]
[0272]
[0273] Example 1. Inverse Prime Editing (iPE) system, developed using pegRNA and nCas9-D10A, generates reverse prime editing in human cells.
[0274] Traditional guide editors primarily utilize nCas9-H840A and the RNA-dependent 5'→3' reverse transcriptase M-MLV RT to generate guide editing downstream of the target sequence cleavage site. Since no reverse transcriptase capable of reverse transcription along the 3'→5' axis has been found in nature to work with nCas9-H840A, traditional guide editors cannot generate guide editing upstream of the cleavage site.
[0275] This invention first utilizes pegRNA and nCas9-D10A to develop the reverse guided editing system iPE ( Figure 1a and Figure 2 a). This invention first developed the iPE2 editor, then developed the iPE3 editor by adding nicking sgRNA, and the iPE4 editor by adding the MLH1dn protein factor. The iPE5 editor was developed by adding both simultaneously. Testing in human HEK293T cells revealed low efficiency, producing editing efficiencies of 0.04%–1.06% at target sites HBB, HEXA, FANCF, and PDCD1. Figure 3 a). Simultaneously, the inventors utilized the iPEmax skeleton to construct more efficient iPEmax2, iPEmax3, iPEmax4, and iPEmax5 editors, finding that they produced higher editing efficiency than iPEs, reaching 8.6% at DMD sites ( Figure 3 b).
[0276] Example 2. The nu-iPE reverse-guided editing system, developed using pegRNA and WTcas9, generates reverse-guided editing in human cells HEK293T.
[0277] This invention further utilizes pegRNA and WTCas9 to develop the reverse guided editing system nu-iPEs ( Figure 1b and Figure 2(b) Simultaneously, testing was conducted on the HEK293T reporter system in human cells. A one-base deletion and a two-base substitution mutation were generated on the copGFP plasmid in the reporter system. Only with precise and efficient reverse guided editing on the copGFP plasmid could the HEK293T reporter system in human cells regain fluorescence. Experimental results showed that nu-iPEs produced higher reverse guided editing efficiency in the reporter system than iPEs, resulting in brighter green fluorescence. Figure 4 a- Figure 4 c). Simultaneously, the inventors tested nu-iPEs in HEK293T cells and found the same results as with the reporter system: nu-iPEs produced higher reverse-guided editing efficiency in human HEK293T cells than iPEs. Figure 5 The inventors analyzed that traditional guided editing is based on nCas9-H840A. The PBS sequence binding site is located within 20 bp of the target site, and the target sequence region is well unwound by nCas9-H840A, which is beneficial for reverse transcriptase to use the RTT-PBS sequence for reverse transcription, thereby generating the desired edit. However, the downstream position of the target sequence cleavage site is not unwound by Cas9, so the reverse guided editing constructed using nCas9-D10A is less efficient than traditional guided editors. In contrast, after generating a DNA double-strand break using WTCas9, the body's repair mechanism further opens the DNA double strand downstream of the target sequence cleavage site, which is beneficial for PBS binding, thus facilitating reverse guided editing.
[0278] Example 3. Using circular RNA and ciPE and nu-ciPE editors developed with nCas9-D10A and WTCas9 respectively, efficient reverse-guided editing was generated on the HEK293T target site in human cells.
[0279] The results in Examples 1 and 2 of this invention suggest that the inventors need to increase the unwinding of DNA double strands downstream of the target sequence cleavage site to facilitate reverse-guided editing. Therefore, this invention innovatively introduces circular RNA (SEQ ID NO: 13, with the RTT and PBS sequences of pegRNA inserted at the MCS site of the polyclonal restriction enzyme) into the reverse-guided editing system. Because circular RNA itself has a certain unwinding ability, it can unwind downstream of the target sequence cleavage site and bind the carried RTT-PBS sequence downstream of the target sequence cleavage site, initiating M-MLV RTΔRNase H-mediated reverse transcription to generate the desired edited sequence. The inventors have developed a ciPE editor (… Figure 1c and Figure 2c) In human HEK293T cells, ciPEs were tested and found to produce higher reverse-guided editing efficiency than iPEs at human cell targets DMD, FANCF, HEK3, RNF2, PDCD1, CXCR4, HEXA, HEK4, BCL11A, and TRAC, with an editing efficiency of up to 24.7% at the HEK4 site. Figure 6 a- Figure 6 b).
[0280] The inventors also developed nu-ciPE, a reverse-guided editing system using circular RNA, based on WTCas9. Figure 1d and Figure 2 d) Because WTCas9 may better open the DNA double strand downstream of the target sequence cleavage site during repair, it is beneficial for reverse-guided editing. The inventors performed reverse-guided editing at six target sites in human HEK293T cells, and the results showed that the reverse-guided editor nu-ciPEs produced the expected precise editing upstream of the PAH target site, with an editing efficiency of approximately 19%. Figure 7 ).
[0281] Example 4. Helicase-assisted reverse guide editor hciPE produces more efficient reverse guide editing than ciPE in human cells HEK293T.
[0282] This invention further adds a 3'→5' DNA helicase to unwind downstream of the target sequence cleavage site, facilitating the binding of the PBS sequence carried by the circular RNA to the target sequence. Under the action of reverse transcriptase M-MLV RTΔRNase H, the desired DNA sequence is generated, thus enabling reverse guided editing. This invention selects the previously optimized 3'→5' DNA helicase Rep-X to assist in reverse guided editing and develops the hciPE reverse guided editor (…). Figure 1e and Figure 2 e). This invention was tested in human cells HEK293T, K562, and U2OS. The reverse-guided editor hciPEs demonstrated higher reverse-guided editing efficiency than ciPEs at six targets (HEK4, DMD, BCL11A, PSMB2, GFAP, and HEXA), with a maximum efficiency of 55.4%. Figure 8 a- Figure 8 d).
[0283] Example 5. The helicase-assisted reverse-guided editor hciPE produces efficient editing in regions such as disease sites that were previously difficult to edit.
[0284] This invention further compares the efficiency of the hciPE editor with existing editors PAMless-PE and twinPE, finding that PAMless-PEs, including SpRY-PE and SpG-PE, only achieve an average maximum editing efficiency of 4.5% at the three target sites of human HEK293T cells, while the hciPE editor can achieve an average editing efficiency of up to 14.0%, significantly higher than PAMless-PEs. Figure 9 a- Figure 9 b). Similarly, for the disease targets BRCA1 and RPE65, PAMless-PEs and twinPEs produced editing efficiencies of 5.7% and 2.9%, 0.8% and 0.5%, respectively, while hciPEs produced editing efficiencies of 13.3% and 9.5%. Figure 9 c- Figure 9 d).
[0285] Example 6. The ciPE and hciPE editors produce lower off-target activity.
[0286] This invention further utilized Cas-OFFinder to predict off-target sites for GFAP, HEK4, and DMD targets in human HEK293T cells, identifying a total of 26 off-target sites. These off-target sites were then validated using ciPE and hciPE editors. Only background-level InDels products (<0.04%) were detected at the GFAP and DMD sites; no off-target reverse-guided editing products were detected. Figure 10 a- Figure 10 (b) However, among the nine off-target sites of HEK4, off-target site 3 generated 1.32%–6.32% InDels off-target activity and 0.05%–0.67% reverse guided editing off-target activity. Given the low off-target activity of guided editing, the inventors further analyzed the reason for the reverse guided editing off-target activity at off-target site 3 and found that this site shares 7 bases with the target sequence in the PBS region, which greatly increases the binding efficiency of the PBS sequence to this off-target site, thus leading to a certain degree of off-target activity. Figure 10 a- Figure 10 b).
[0287] The sequences involved in this invention:
[0288] SEQ ID NO:1
[0289]
[0290] SEQ ID NO:2
[0291] SEQ ID NO:3
[0292]
[0293] SEQ ID NO:4 sgRNA or nicking sgRNA backbone sequence
[0294] GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUGGCACCGAGUC
[0295] GGUGC
[0296] SEQ ID NO:5
[0297]
[0298] SEQ ID NO:6
[0299]
[0300] SEQ ID NO:7
[0301]
[0302]
[0303] SEQ ID NO:8
[0304] SEQ ID NO:9
[0305]
[0306]
[0307] SEQ ID NO:10
[0308]
[0309] SEQ ID NO:11
[0310]
[0311] SEQ ID NO:12
[0312]
[0313] SEQ ID NO:13
[0314]
[0315] SEQ ID NO:14 WT-Cas9 nuclease
[0316] DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC
[0317] YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALA
[0318] HMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLF
[0319] GNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPL
[0320] SASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE
[0321] DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPW
[0322] NFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLF
[0323] KTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIE
[0324] ERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKA
[0325] QVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIK
[0326] ELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGK
[0327] SDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKY
[0328] DENDKLIREVKIVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK
[0329] MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTE
[0330] VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKN
[0331] PIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQK
[0332] QLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKR
[0333] YTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0334] SEQ ID NO: 15 D10A-Cas9, that is, nCas9-D10A
[0335] DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC
[0336] YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALA
[0337] HMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLF
[0338] GNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPL
[0339] SASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE
[0340] DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPW
[0341] NFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLF
[0342] KTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIE
[0343] ERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKA
[0344] QVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIK
[0345] ELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGK
[0346] SDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKY
[0347] DENDKLIREVKIVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK
[0348] MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTE
[0349] VQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKN
[0350] PIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQK
[0351] QLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKR
[0352] YTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0353] SEQ ID NO:16 Amino acid sequence of M-MLV reverse transcriptase (ΔRnaseH)
[0354] ELPGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPL
[0355] SEQ ID NO:17 Amino acid sequence of MCP
[0356] ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLN MELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY
[0357] SEQ ID NO:18MS2 sequence
[0358] GCACATGAGGATCACCCATGTGC
[0359] SEQ ID NO:19 32aa connector sequence
[0360] SGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0361] SEQ ID NO:20NLS sequence
[0362] MKRTADGSEFESPKKKRKV
[0363] SEQ ID NO:21NLS sequence
[0364] KRTADGSEFEPKKKRKV
[0365] SEQ ID NO:22 5'ribozyme coding sequence
[0366] GCCATCAGTCGCCGGTCCCAAGCCCGGATAAAATGGGAGGGGGCGGGAAACCGCCT
[0367] SEQ ID NO:23 3'ribozyme coding sequence
[0368] AACACTGCCAATGCCGGTCCCAAGCCCGGATAAAAGTGGAGGGTACAGTCCACGC
[0369] SEQ ID NO:24 5' loop arm sequence
[0370] AACCATGCCGACTGATGGCAGAACTAA
[0371] SEQ ID NO:25 3' loop arm sequence
[0372] AAATTAACACTGCCATCAGTCGGCGTGGACTGTAG
[0373] SEQ ID NO:26 MLH1 dn Protein Factor Sequence (Homo sapiens)
[0374] MSFVAGVIRRLDETVVNRIAAGEVIQRPANAIKEMIENCLDAKSTSIQVIVKEGGLKLIQIQDNGTGIRKEDLDIVCERFTTSKLQSFEDLASISTYGFRGEALASISHVAHVTITTKTADGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEEYGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDK TLAFKMNGYISNANYSVKKCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSNSSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSDKVYAHQMVRTDSREQKLDAFLQPLSKPLSSQPQAIVTEDKTDISSGRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVF
[0375] SEQ ID NO:27 Helicase Rep-X (Escherichia)
[0376] MRLNPGQQQAVEFVTGPLLVLAGAGSGKTRVITNKIAHLIRGSGYQARHIAAVTFTNKAAREMKERVGQTLGRKEARGLMISTFHTLGLDIIKREYAALGMKANFSLFDDTDQLALLKELTEGLIEDDKVLLQQLISTISNWKNDLKTPSQAAASAIGERDRIFAHVYGLYDAHLKACNVLDFDDLILLPTLLLQRNEEVRKRWQNKIRYLLVDEYQDTNTSQYELVKLLVGSRARFTVVGDDDQSIYSWRGARPQNLVLLSQDFPALKVIKLEQNYRSSGRILKAANILIANNPHVFEKRLFSELGYGAELKVLSANNEEHEAERVTGELIAHHFVNKTQYKDYAILYRGNHQSRVFEKFLMQNRIPYKISGGTSFFSIPEIKDLLAYLRVLTNPDDDCAFLRIVNTPKREIGPATLKKLGEWAMTRNKSMFTASFDMGLSQTLSGRGYEALTRFTHWLAEIQRLAEREPIAAVRDLIHGMDYESWLYETSPSPKAAEMRMKNVNQLFSWMTEMLEGSELDEPMTLTQVVTRFTLRDMMERGESEEELDQVQLMTLHASKGLEFPYVYMVGMEEGFLPHQSSIDEDNIDEERRLAYVGITRAQKELTFTLAKERRQYGELVIPEPSRFLLELPQDDLIWEQERKVVSAEERMQKGQSHLANLKAMMAAKRGK
[0377] SEQ ID NO:28 MLH1dn protein factor sequence
[0378] MSFVAGVIRRLDETVVNRIAAGEVIQRPANAIKEMIENCLDAKSTSIQVIVKEGGLKLIQIQDNGTGIRKEDLDIVCERFTTSKLQSFEDLASISTYGFRGEALASISHVAHVTITTKTADGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEEYGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDKTLAFKMNGYISNANYSVKKCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSNSSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSD KVYAHQMVRTDSREQKLDAFLQPLSKPLSSQPQAIVTEDKTDISSGRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVF
[0379] SEQ ID NO:29
[0380]
[0381]
Claims
1. A reverse prime editing system, comprising: (i)a) a CRISPR nuclease and / or an expression construct comprising a nucleotide sequence encoding the CRISPR nuclease, and a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase; or b) a prime editing fusion protein and / or an expression construct comprising a nucleotide sequence encoding the prime editing fusion protein, wherein the prime editing fusion protein comprises a CRISPR nuclease and / or a reverse transcriptase; (ii) a prime editing guide RNA (pegRNA) directed to a target sequence in genomic DNA and / or an expression construct comprising a nucleotide sequence encoding the prime editing guide RNA; and / or (iii) a circular RNA comprising a reverse transcription template (RTT) sequence and a primer binding site (PBS) sequence and / or an expression construct comprising a nucleotide sequence encoding the circular RNA, and a guide RNA (sgRNA) directed to a target sequence in genomic DNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA; Optionally, the CRISPR nuclease is a target strand nicking enzyme or a double-strand endonuclease. The CRISPR nuclease is a Cas9 nuclease or a variant thereof.
2. The prime editing system of claim 1, wherein, The Cas9 nuclease comprises a sequence as set forth in SEQ ID NO: 14; and / or 3. The prime editing system of claim 2, wherein, The Cas9 nuclease variant is an nCas9 nuclease comprising a sequence as set forth in SEQ ID NO:
15. The reverse transcriptase is an M-MLV reverse transcriptase; 4. The prime editing system of any one of claims 1-3, wherein, Preferably, the RNase H domain of the M-MLV reverse transcriptase is mutated or deleted, which comprises a sequence as set forth in SEQ ID NO:
16. The reverse transcriptase and / or the prime editing fusion protein can further comprise one or more RNA aptamer binding protein sequences.
5. The prime editing system of any one of claims 1-4, wherein, The RNA aptamer binding protein is an MCP protein; 6. The prime editing system of claim 5, wherein, Optionally, the MCP protein comprises a sequence as set forth in SEQ ID NO:
17. The prime editing fusion protein comprises a sequence selected from any one of SEQ ID NOs: 1-3, 7-11.
7. The prime editing system of any one of claims 1-6, wherein, 8.The reverse prime editing system of any one of claims 1-7, further comprising a guide RNA targeting a non-target strand (nicking sgRNA) and / or an expression construct comprising a nucleotide sequence encoding the guide RNA targeting a non-target strand. The guide RNA comprises a backbone sequence as set forth in SEQ ID NO: 4; and / or 9. The prime editing system of any one of claims 1-8, wherein, The guide RNA targeting a non-target strand comprises a backbone sequence as set forth in SEQ ID NO:
4. In the circular RNA, the reverse transcription template (RTT) sequence and the primer binding site (PBS) sequence are directly connected; 10. The prime editing system of any one of claims 1-9, wherein, Preferably, the reverse transcription template (RTT) sequence is located at the 5’ end of the primer binding site (PBS) sequence. The circular RNA comprises at least one RNA aptamer sequence, and optionally, the aptamer comprises MS2.
11. The prime editing system of any one of claims 1-10, wherein, The primer binding site (PBS) sequence is complementary to at least a portion of the target sequence.
12. The prime editing system of any one of claims 1-11, wherein, Preferably, the primer binding site sequence is complementary to at least a part of the 3' overhang single strand resulted from the nick in the target strand of the target sequence, in particular, to the nucleotide sequence of the 3' end of the 3' overhang single strand.
13. The prime editing system of any one of claims 1-12, wherein, The expression construct containing the nucleotide sequence encoding the circular RNA comprises a coding sequence of 5'-first ribozyme-first loop-forming arm-RTT-PBS-second loop-forming arm-second ribozyme-3'; Optionally, when transcribed into RNA in a cell, the first ribozyme and the second ribozyme are capable of self-cleavage to produce 5'-first loop-forming arm-RTT-PBS-second loop-forming arm-3', and the first loop-forming arm and the second loop-forming arm are capable of connecting to each other to form a circular RNA containing RTT and PBS.
14. The prime editing system of claim 13, wherein, The coding sequence of the first ribozyme comprises a sequence as set forth in SEQ ID NO: 22, the coding sequence of the second ribozyme comprises a sequence as set forth in SEQ ID NO: 23, the first loop-forming arm comprises a nucleotide sequence as set forth in SEQ ID NO: 24, and the second loop-forming arm comprises a nucleotide sequence as set forth in SEQ ID NO:
25.
15. The prime editing system of any one of claims 1-14, wherein, The nucleotide sequence encoding the circular RNA is driven by a U6 promoter.
16. The prime editing system of any one of claims 1-15, wherein, Further comprising an MLH1dn protein factor and / or an expression construct containing a nucleotide sequence encoding the MLH1dn protein factor; Optionally, the MLH1dn protein factor comprises a sequence as set forth in SEQ ID NO:
26.
17. The prime editing system of any one of claims 1-16, further comprising a helicase and / or an expression construct containing a nucleotide sequence encoding the helicase. Optionally, the helicase comprises a sequence as set forth in SEQ ID NO:
27. Preferably, the helicase forms a fusion protein with the reverse transcriptase; more preferably, the fusion protein comprises a sequence as set forth in SEQ ID NO:
12.
18. Use of the prime editing system of any one of claims 1-17 in (A) or (B) below: (A) editing of a genomic sequence of an organism or a biological cell; (B) preparing a product of editing of a genomic sequence of an organism or a biological cell.
19. A kit comprising the prime editing system of any one of claims 1-17.
20. A method of producing a genetically modified cell, comprising introducing the prime editing system of any one of claims 1-17 into at least one of the cells, thereby causing modification of a genomic sequence of the at least one cell.
21. The method of claim 20, wherein the cell is from a microorganism, an animal, or a plant. The microorganism includes bacteria and fungi. The animal includes mammals and poultry. The plant includes monocotyledonous plants and dicotyledonous plants. Optionally, the mammal includes humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats. Optionally, the poultry includes chickens, ducks, and geese. Optionally, the monocotyledonous plant includes rice, corn, wheat, sorghum, and barley, and the dicotyledonous plant includes soybeans, peanuts, Arabidopsis thaliana, cotton, and oilseed rape.
Citation Information
Patent Citations
Delivery, use and therapeutic applications of the crispr-CAS systems and compositions for HBV and viral diseases and disorders
WO2015089465A1
Novel crispr enzymes and systems
WO2016205711A1
Uses of adenosine base editors
WO2019079347A1
Methods and compositions for editing nucleotide sequences
WO2020191233A1
Base editors, compositions, and methods for modifying the mitochondrial genome
WO2021155065A1