Ligation of nucleic acid molecules and the use thereof for genome editing
The ligation of nucleic acid fragments with splint DNA under controlled conditions addresses the challenges of producing high-quality, longer nucleic acid molecules like pegRNA and epegRNA, enhancing stability and efficiency.
Patent Information
- Application Number
- PCT/CN2025/079468
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-02-27
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods for synthesizing longer nucleic acid molecules, such as pegRNA and epegRNA, face challenges with low quality and high cost, and existing techniques like solid-phase synthesis and in vitro transcription struggle to produce site-specific modifications and are inefficient in producing longer RNA fragments.
A method involving the ligation of short single-stranded nucleic acid fragments using single-stranded splint DNA fragments under controlled temperature conditions, followed by nucleic acid ligase, to form longer nucleic acid molecules, including pegRNA and epegRNA, with specific chemical modifications to enhance stability.
The method enables the production of high-quality, longer nucleic acid fragments with improved stability and editing efficiency, reducing production costs and overcoming the limitations of existing synthesis methods.
Smart Images

Figure PCTCN2025079468-FTAPPB-I100001 
Figure PCTCN2025079468-FTAPPB-I100002 
Figure PCTCN2025079468-FTAPPB-I100003
Abstract
Description
LIGATION OF NUCLEIC ACID MOLECULES AND THE USE THEREOF FOR GENOME EDITINGBACKGROUND
[0001] Nucleic acid molecules can be synthesized or transcribed in vitro. In vitro transcription, however, has very limited ability to make specifically modified nucleic acid molecules, such as those used in RNA vaccines and certain guide RNA molecules in genome editing.
[0002] CRISPR-based genome editing systems introduce targeted alternation of genome. CRISPR-Cas nucleases cause double-stranded break to facilitate small insertions or deletions (indels) or targeted insertion via non-homologous end joining or homology-directed repair, respectively. Applying catalytically impaired Cas nucleases together with various deaminases effectively convert C-to-T or A-to-G, referred as base editing. Fusion of Cas9 nickase with a reverse transcriptase (RT) created prime editor (PE) to generate small modifications around the nicked site without donor DNA. The prime editing guide RNA (pegRNA) contains a reverse transcription template (RTT) and primer binding sequence (PBS) at the 3’ end of sgRNA, resulting in at least approximately 125-145 nucleotides (nt) in length. Unlike sgRNA, which is protected by a Cas9 protein, the 3’ extension of pegRNA, including RTT and PBS, is susceptible to degradation by ribonucleases within the cell. Structured RNA motifs have been incorporated at the 3’ end of pegRNA to enhance its stability and prevent the degradation of PBS and RTT sequences. The evopreQ1 motif has been included into pegRNA to generate the engineered pegRNA (epegRNA) , and it has been shown to improve the editing efficiencies of PE by 3-4 folds in several cell lines tested.
[0003] Chemical modifications of sgRNA can improve its stability against degradation and greatly enhance its performance in cells and in vivo. Chemically modified sgRNA has been extensively used in research and therapeutic applications. The first FDA approved CRISPR-based therapy, CASGEVY, delivered chemically modified sgRNA and Cas9 protein as ribonucleoprotein (RNP) into hematopoietic stem and progenitor cells (HSPC) , as an ex vivo therapy for β-thalassemia and sickle cell disease. In vivo CRISPR-based therapy has entered late-stage clinical development to treat transthyretin (ATTR) amyloidosis with cardiomyopathy. Chemically modified sgRNA and Cas9 mRNA, which are encapsulated into lipid nanoparticles, have been shown to effectively edit ATTR gene in vivo in patients. Chemical modifications have been applied to pegRNA for RNP and mRNA delivery of PE.
[0004] In vitro RNA synthesis, such as by solid-phase synthesis, needs to be stretched to produce chemically modified pegRNA and epegRNA with sharply increased cost, and the quality of chemically synthesized RNA with such length is not satisfied. While RNP and RNA delivery of Cas9 / sgRNA usually surpass the plasmids delivery in terms of editing efficiency, the editing efficiencies of PE via RNP and RNA delivery is lower than expectation. It is believed that the low editing efficiency via RNP and RNA delivery of PE is in part due to the low quality of chemically synthesized pegRNA or epegRNA.
[0005] In vitro transcription (IVT) , on the other hand, employs RNA polymerase such as T7 and several others to generate RNA sequence up to 30 kb. Modified nucleotides including Pseudouridine (Ψ) , N1-methylpseudouridine (m1Ψ) and 5-methylcytidine (m5C) are introduced randomly or completely into mRNA to attenuate immunogenicity. The standard IVT however, cannot generate site-specific modifications, and it cannot tolerate modified nucleotides such as 2’ OH modifications of ribose for enhancing resistance to ribonucleases. Moreover, IVT by T7 enzyme usually produce sequences containing heterogeneous 5’a nd / or 3’ end products.
[0006] A hybrid solid-liquid phase transcription approach was combined with automated robotic platform to generate RNA with position-selective modification. Recently, a biocatalytic method was developed to generate oligonucleotides by combining polymerases and endonucleases in one pot. However, these two methods are suitable to generate modified short oligonucleotides in large quantity and high quality rather than to produce longer than 100 nt RNA.
[0007] Solid-phase chemical synthesis enables production of desired RNA sequences with position-selective modifications, to incorporate modified residues at the base and sugar phosphate backbone to enhance stability. The linear chemical synthesis of oligonucleotides uses nucleoside phosphoramidite as building blocks, and the synthesis relies on repeated rounds of several chemical reactions in order, extending one nucleotide each round. Although the efficiency of individual elongation cycle is high, the overall yield of synthesis drops sharply with longer sequences. Therefore, the current solid-phase synthesis generally is limited to preparation of RNA fragments with a length of approximately 100 nt only.
[0008] When longer RNA fragments are attempted, the yield of production steeply declines, and frequent failures occurs. Truncated sequences arise from incomplete coupling reactions, and they can block solid support. All these factors contribute to product impurities, some of which are difficult to be removed by purification.
[0009] Therefore, synthesis of pegRNAs (about 125-145 nt) is associated with low quality and high cost, while that of epegRNA (about 170-190 nt) is nearly impractical.SUMMARY
[0010] Provided are compositions and methods useful for ligating short fragments of single-stranded nucleic acid fragments, with the assistance of one or more single-stranded splint DNA fragment, to prepare much longer nucleic acid fragments, such as pegRNA molecules.
[0011] One embodiment of the present disclosure provides a method for ligating two or more single-stranded target nucleic acid fragments, comprising: (a) incubating, in a sample, the two or more single-stranded target nucleic acid fragments with one or more single-stranded splint DNA fragment (s) at a denaturing temperature; (b) reducing the temperature of the sample to allow the target nucleic acid fragments and the splint DNA fragment (s) to form an at least partially duplex molecule by virtue of their sequence complementarity, wherein there is no gap between each adjacent target nucleic acid fragment in the duplex molecule; (c) incubating the sample at a ligation temperature, in the presence of a nucleic acid ligase, to allow the ligase to ligate the target nucleic acid fragments in the duplex molecule to form a ligated nucleic acid fragment; and (d) repeating steps (a) - (c) for at least one time to form more ligated nucleic acid fragments.
[0012] In some embodiments, the target nucleic acid fragments are RNA fragments, RNA fragments with modified nucleic acids or RNA / DNA chimera fragments.
[0013] In some embodiments, the denaturing temperature is 55 ℃ to 98 ℃, preferably 60 ℃to 80 ℃, and more preferably 65 ℃ to 75 ℃.
[0014] In some embodiments, the ligation temperature is 20 ℃ to 42 ℃, preferably 32 ℃ to 39 ℃, more preferably 35 ℃ to 38 ℃.
[0015] In some embodiments, in the duplex molecule, one of the strands includes only the target nucleic acid fragments and the other strand includes only the splint DNA fragment (s) . In some embodiments, the duplex molecule includes a single splint DNA fragment. In some embodiments, the duplex molecule includes two or three target nucleic fragments. In some embodiments, the duplex molecule includes two or more splint DNA fragments.
[0016] In some embodiments, the metho further entails (e) digesting the splint DNA fragment (s) .
[0017] In some embodiments, at least one of the target nucleic fragments is chemically modified or comprises a non-natural nucleotide.
[0018] In some embodiments, at least one of the target nucleic fragments comprises a chemically modified ribose selected from the group consisting of 2’ -O-methyl (2’ -O-Me) , 2’ -Fluoro (2’ -F) , 2’ -deoxy-2’ -fluoro-beta-D-arabino-nucleic acid (2’ F-ANA) , 4’ -S, 4’ -SFANA, 2’-azido, UNA, 2’ -O-methoxy-ethyl (2’ -O-ME) , 2’ -O-Allyl, 2’ -O-Ethylamine, 2’ -O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, and N-MeO-aminooxy BNA.
[0019] In some embodiments, at least one of the target nucleic fragments comprises a phosphodiester linkage selected from the group consisting of Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’ , 5’ -amide, N3’ -phosphoramidate (NP) , Phosphodiester (PO) , or 2’ , 5’ -phosphodiester (2’ , 5’ -PO) . In some embodiments, the target nucleic acid fragments, following ligation, constitute a pegRNA.
[0020] In some embodiments, the target nucleic acid fragments, following ligation, have a total length of at least 150 nt, or at least 180 nt, 200 nt, 250 nt, or 300 nt.
[0021] In some embodiments, each target nucleic acid fragment and each splint DNA fragment have a molar ratio, in the sample, that is from 0.5: 1.5 to 1.5: 0.5, or preferably from 0.8: 1.2 to 1.2: 0.8.
[0022] In some embodiments, the nucleic acid ligase is T4 RNA ligase 2 (T4 Rnl2) .
[0023] In some embodiments, at each step (c) when it is repeated, an amount of 0.5 to 5 μg / μl of the nucleic acid ligase is added.
[0024] In some embodiments, the steps (a) - (c) are repeated once or twice in step (d) .
[0025] In some embodiments, in step (c) , the sample is incubated at the ligation temperature for 15 minutes to 3 hours, preferably 20 minutes to 90 minutes, more preferably 25 to 65 minutes.
[0026] In some embodiments, each target nucleic acid fragment has a length of 10 nt to 200 nt.
[0027] Also provided is a prime editing guide RNA (pegRNA) , comprising (a) a single guide RNA (sgRNA) , (b) a reverse transcriptase (RT) template, and (c) an RNA aptamer or an RNA pseudoknot motif, wherein the pegRNA is at least 150 nt in length, preferably at last 180 nt, 200 nt, 250 nt or 300 nt in length.
[0028] In some embodiments, the pegRNA is chemically modified or comprises a non-natural nucleotide.
[0029] In some embodiments, the pegRNA comprises a chemically modified ribose selected from the group consisting of 2’ -O-methyl (2’ -O-Me) , 2’ -Fluoro (2’ -F) , 2’ -deoxy-2’ -fluoro-beta-D-arabino-nucleic acid (2’ F-ANA) , 4’ -S, 4’ -SFANA, 2’ -azido, UNA, 2’ -O-methoxy-ethyl (2’ -O-ME) , 2’ -O-Allyl, 2’ -O-Ethylamine, 2’ -O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, and N-MeO-aminooxy BNA.
[0030] In some embodiments, the pegRNA comprises a phosphodiester linkage selected from the group consisting of Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’, 5’ -amide, N3’ -phosphoramidate (NP) , Phosphodiester (PO) , or 2’ , 5’ -phosphodiester (2’, 5’ -PO) .
[0031] In some embodiments, the RNA aptamer is selected from the group consisting of an MS2 aptamer, a tetrahymena thermophila group I intron, a theophylline aptamer, an ATP aptamer, a guanine quadruplex aptamer, a thrombin aptamer (TBA) , a vascular endothelial growth factor (VEGF) aptamer, an aptamer against HIV-1 reverse transcriptase, a hemin binding aptamer (HBA) , a riboswitch, a ribozyme.
[0032] In some embodiments, the RNA pseudoknot motif is selected from the group consisting of an H-type pseudoknot, a kissing-loop pseudoknot, a triple helix pseudoknot, a complex pseudoknot, a frame-shift pseudoknot, a riboswitch pseudoknot, a G-quadruplex pseudoknot, an intermolecular pseudoknot.
[0033] In some embodiments, the RNA pseudoknot motif is evopreQ1. In some embodiments, the pegRNA comprises two or more of the RNA aptamer or RNA pseudoknot motif.
[0034] In some embodiments, at least 20%of the 50 3’ end nucleotides are modified, optionally at least 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the 50 3’ end nucleotides are modified. In some embodiments, at least 5 of the 15 3’ nucleotides are modified, optionally at least 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of the 15 3’ nucleotides are modified. In some embodiments, the modification is a ribose modification selected from 2'-O-methyl (2'-O-Me) , 2’ -fluoro (2’ -F) , and the combination thereof, preferably wherein the modification is a ribose modification with 2'-O-Me.
[0035] In some embodiments, the (d) comprises at least a stem-loop, and wherein no more than 50%of the nucleotides in the stem-loop are modified, preferably wherein no more than 40%, 30%, 20%or 10%of the nucleotides in the stem-loop are modified.
[0036] In some embodiments, the sgRNA comprises a spacer and a scaffold comprising, from 5’ to 3’ , a first, a second, a third and a fourth stem-loops, and wherein the first and third stem-loops comprise modifications. In some embodiments, each of the first and third stem-loops comprises at least 2, 3, 4, 5, 6, 7, or 8 modified ribose’s selected from 2'-O-methyl (2'-O-Me) ribose, 2’ -fluoro (2’ -F) ribose, and the combination thereof, preferably wherein the modified ribose’s are 2'-O-methyl (2'-O-Me) ribose’s. In some embodiments, the fourth stem-loop comprises no more than 5, 4, 3, 2 or 1 modified ribose’s.
[0037] In some embodiments, the spacer comprises three consecutive phosphorothioate (PS) linkages at the 5’ end.
[0038] In a preferred embodiment, in the pegRNA, (i) the spacer comprises three consecutive phosphorothioate (PS) linkages at the 5’ end, (ii) the first and third stem-loops of the scaffold of the sgRNA each comprises at least 8 nucleotides with 2'-O-methyl (2'-O-Me) modifications, and (iii) at least 20%of the 50 3’ end nucleotides of the pegRNA are modified with 2'-O-methyl (2'-O-Me) .
[0039] In some embodiments, (i) the spacer comprises only three consecutive PS linkages at the 5’ end, (ii) the fourth stem-loop of the scaffold of the sgRNA comprises no more than 4 modified ribose’s, and (iii) the step-loop of (d) comprises no more than 40%modified ribose’s.BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG. 1a-h: Optimization of RNA ligation enables assemble of sgRNA. a. Overview of design scheme for sgRNA ligation. Acceptor RNA is 20 nt spacer sequence by chemical synthesis, and donor RNA is 82 nt scaffold sequence generated by IVT. b. Urea-PAGE of ligated sgRNA (for VEGFA locus) . c. Urea-PAGE analysis of ligated sgRNA. A mixture of 100 pmol acceptor RNA, 100 pmol donor RNA, and 100 pmol 40 nt splint DNA were annealed and ligated with 0.5 μl T4 RNA Ligase 2. The ligation reactions were performed at 25℃ or 37℃, respectively. d. (Left) Overview of design scheme for sgRNA ligation using various lengths of splint DNA (20, 40, or 59nt) . (Right) Urea-PAGE of ligation products. The reactions were performed at 37℃ following the conditions described above. e. Urea-PAGE of sgRNA ligation products at different doses of splint DNA. f. Urea-PAGE of sgRNA ligation products with different doses of T4 RNA Ligase 2. g. The 20+82 ligated RNA: sgRNA was ligated using 20 nt synthetic acceptor RNA and 82 nt IVT-generated donor RNA; *20+82 Ligated RNA: sgRNA was ligated using 20 nt synthetic acceptor RNA with 5’ modification and 82nt IVT donor RNA. h. In vitro cleavage of ligated sgRNA in TAE agarose gel. The molar ratio of spCas9 protein and sgRNA was 1: 1. The RNP was incubated at room temperature for 10 minutes, followed by the addition of the DNA template and incubation at 37℃ for 1 hour.
[0041] FIG. 2a-f: Ligated pegRNA mediates efficient prime editing. a. Overview of the initial design scheme for ligation of pegRNA. The acceptor RNA is a 32 nt 5’ end modified RNA, and it includes the spacer sequence and part of the sgRNA scaffold sequence. The donor RNA was produced by IVT, and it includes the rest of scaffold sequence and RTT-PBS. The acceptor RNA has a hydroxyl group at 3’ end, and the donor RNA has a phosphate group at 5’ end. T4 Rnl2 ligates facilitates the joining of RNA molecules in the presence of splint DNA. b. Ligation of pegRNA (for +5G to T mutation at VEGFA locus) . From left to right, Lane 1: marker; Lane 2: chemically synthesized 32 nt acceptor RNA; Lane 3: 105 nt donor RNA by IVT; Lane 4: full-length pegRNA by IVT as a control; Lane 5: ligated 137 nt pegRNA. The samples were run in 6%denaturing urea polyacrylamide gel electrophoresis (Urea-PAGE) . c-f. The efficiencies of RNP-mediated prime editing in HEK293T cells were determined by deep sequencing. Four different pegRNAs were used as +5 G to T mutation at FANCF (c) , 3 bp insertion at HEK3 (d) , +5 G to T mutation at VEGFA (e) and 3 bp deletion at VEGFA loci (f) . “IVT” indicates full length pegRNA generated by IVT; 32+100 / 96 / 105 / 102 indicates ligated pegRNA with 5’ end modified; 32+100 / 96 / 105 / 102 (HPLC) indicated ligated pegRNA with HPLC purification. For each electroporation, 140 pmol PE protein, 186 pmol pegRNA, and 62 pmol nicking sgRNA were used. Modification indicates three nucleosides harboring 2′-O-methyl modification and with three phosphorothioate linkages. Data and error bars represent the mean and standard deviation of three or more independent biological replicates. One-way ANOVA with Tukey’s multiple comparisons test was applied; NS, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0042] FIG. 3a-c: HPLC purification and measurement of ligated pegRNA. a. HPLC purification of ligated pegRNA (+5 G to T Mutation in VEGFA locus) . b. Analysis of pegRNA purity using area under curve (AUC) of each peak. c. Detection of purity by HPLC analysis of purified. mAU, milli-absorbance unit; time, the execution time of the program for (a) and (c) .
[0043] FIG. 4a-h: Comparison of different ligation strategies for pegRNA. a-c. Overview of design schemes to generate epegRNA with only 5’ end modification (a) , pegRNA with both 5’ and 3’ ends modifications (b) , and epegRNA with both 5’a nd 3’ ends modifications (c) . d-f. The efficiencies of RNP-mediated prime editing in HEK293T cells were determined by deep sequencing for 3 bp insertion at HEK3 locus (d) , VEGFA locus +5 G to T mutation (e) , and VEGFA locus 3 bp deletion (f) . *32+141 / 147 / 150evQ1: ligated epegRNA by a 32 nt synthetic acceptor RNA with 5’ modifications and IVT-generated donor RNA containing evopreQ1. *32+96 / 102 / 105*: ligated pegRNA by a 32 nt synthetic acceptor RNA with 5’ end modifications and synthetic donor RNA with 3’ end modifications. *91+82 / 88 / 91evQ1*: ligated epegRNA by a 91 nt synthetic acceptor RNA with 5’ end modifications and synthetic donor RNA with 3’ end modifications at evopreQ1 sequence. g. Urea-PAGE (6%) of ligated pegRNA / epegRNA (for 3 bp deletion at VEGFA locus) produced by three different ligation designs. M, marker. h. Comparison of RNP-mediated prime editing efficiencies of L-epegRNA and IVT-generated epegRNA. For each electroporation, 70 pmol PE protein, 90 pmol pegRNA, and 60 pmol nicking sgRNA were used. Modification indicates three nucleosides harboring 2′-O-methyl modification and with three phosphorothioate linkages. Data and error bars represent the mean and standard deviation of three or more independent biological replicates. Data analysis used One-way ANOVA with Tukey’s multiple comparisons test; NS, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0044] FIG. 5a-f: Dose optimization of RNP delivery. a-b. PE2 doses optimization. Doses of pegRNA and protein were increased or decreased in equal proportions. The abscissa represents the protein dose. c-d. Optimization of PE protein and pegRNA ratios. For each sample, 70 pmol PE protein, and 70, 140, and 280 pmol pegRNA were used respectively. e-f. Optimization of nickRNA dosages for PE3 system. For each sample, 70 pmol PE protein, 140 pmol pegRNA, and 10, 30, 60, and 100 pmol nickRNA were used, respectively. Dose optimizations were for editing at VEGFA locus using *32+150evQ1 pegRNA (a, c, e) and HEK3 locus using 32+158evQ1 pegRNA (b, d, f) , respectively. Data and error bars represent the mean and standard deviation of three independent biological replicates.
[0045] FIG. 6a-e: Production of different PE proteins and their prime editing efficiencies via RNP. a. Illustration of the protein expression vectors. b. SDS-PAGE of PE proteins after NI column purification: ΔRH refers to the RT enzyme of PE lacking RNase H domain. The arrow points to the target protein, M, marker. c. Yield of PE proteins after purification. μg / L: protein yield purified from 1 L bacterial solution. Data and error bars represent the mean and standard deviation from at least two independent biological replicates. Data were analyzed by two-tailed unpaired Student’s t-test; NS indicates no significance; *p < 0.05; **p < 0.01; ***p < 0.001. d-e. Prime editing efficiencies mediated by different PE proteins in HEK293T cells were determined by deep sequencing. +5 G to T mutation (c) and deletion of 3 bp (d) at VEGFA locus. For each sample, 70 pmol PE protein, 140 pmol L-pegRNA, and 60 pmol nicking sgRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used One-way ANOVA; NS, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0046] FIG. 7a-d: The editing efficiencies of PE4max and PE5max via RNP in K562 cells. a-b. Prime editing efficiencies of PE4max and PE5max via RNP delivery for insertion of 3 bp (a) and +1 T to A mutation (b) at HEK3 locus in K562 cells. For each sample, purified MLH1dn of 0 pmol, 17.5 pmol, 35 pmol, 70 pmol and 140 pmol were used, and 70 pmol PEmax ΔRH protein, 140 pmol L-epegRNA, and 60 pmol nicking sgRNA were used. c-d. Prime editing efficiencies of PE4max and PE5max via RNP delivery for insertion of 3 bp (c) and +1 T to A mutation (d) at HEK3 locus in K562 cells. For each sample, 800 ng PE expression plasmid, 200 ng pegRNA plasmid, and 83 ng nickRNA plasmid were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used One-way ANOVA; NS, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0047] FIG. 8a-k: Comparison of prime editing efficiencies via plasmid and L-epegRNA-mediated RNP delivery. a. Schematic representation of L-epegRNA-mediated RNP and plasmid delivery for prime editing. Plasmid and RNP were introduced into cells via electroporation. b-f. The efficiencies of prime editing for five endogenous loci of HEK3 (b) , RUNX1 (c) , DMNT1 (d) , VEGFA (e) , EMX1 (f) loci in HEK293T cells using L-epegRNA-mediated RNP delivery. g-k. Comparisons of prime editing efficiencies using plasmid and L-epegRNA-mediated RNP delivery in HEK293T (g) , K562 (h) , HUH7 (i) , HeLa (g) and U2OS cells (k) . For each sample of RNP delivery, 70 pmol PEmaxΔRH protein, 140 pmol L-epegRNA, and 60 pmol nicking sgRNA were used. Data and error bars represent the mean and standard deviation from at least three independent biological replicates. Data were analyzed by two-tailed unpaired Student’s t-test; NS indicates no significance; *p < 0.05; **p < 0.01; ***p < 0.001.
[0048] FIG. 9a-d: Prime editing efficiencies by L-epegRNA-mediated RNA delivery. a. Illustration of RNA delivery for prime editing. L-pegRNA and mRNA encoding PEmax were co-transfected into cells through electroporation. b-d. Comparison of prime editing efficiencies by L-epegRNA or IVT-generated epegRNA when they were co-transfected with mRNA encoding PEmax in Huh-7 (b) , HEK293T (b) , and K562 cells (d) . For each sample, 2 μg PEmax mRNA, 180pmol epegRNA, and 60 pmol nicking sgRNA were used.
[0049] FIG. 10a-d: Prime editing efficiencies using L-epegRNA in human primary T cells and hematopoietic stem cells. a-d. Prime editing efficiencies of L-epegRNA-mediated RNP in human T cells (a) and CD34+ HSPCs (c) , and RNA delivery in human T cells (b) and CD34+ HSPCs (d) were determined by deep sequencing. For RNP delivery, 70 pmol PEmax ΔRH protein, 140 pmol L-epegRNA, and 60 pmol nicking sgRNA were used for each sample. For RNA delivery, 2 μg mRNA encoding PEmax was co-transfected with 180 pmol L-epegRNA and 60 pmol nicking sgRNA.
[0050] FIG. 11a-e: Multi-fragments assembled L-epegRNA mediates efficient 17 bp insertion. a. Overview of epegRNA ligation design for insertion of 17 bp at HEK3 locus. (Top) Two-fragment ligation strategy using chemically synthesized 105 nt acceptor RNA and 105 nt donor RNA, to be ligated by splint DNA. (Bottom) Three-fragment ligation using chemically synthesized 76 nt, 54 nt, and 80 nt RNA ligated by a 84 nt splint DNA. b. (Left) Urea-PAGE of two-fragment ligation products; (Right) Urea-PAGE of three-fragment ligation products. M, marker. c-d. Comparison of prime editing efficiencies using different epegRNA for insertion of 17 bp at HEK3 locus in HEK293T cells. PE was delivered in RNP (c) and RNA format (d) , respectively. IVT-epegRNA indicates full-length epegRNA produced by IVT. L-epegRNA indicates epegRNA generated through three-fragment ligation. L-epegRNA (HPLC) indicates L-epegRNA that has been purified by HPLC. For each sample, 70 pmol PEmax protein, 140 pmol L-epegRNA and 60 pmol nicking sgRNA were used for RNP, and 2 μg mRNA encoding PEmax was co-transfected with 180 pmol L-epegRNA and 60 pmol nicking sgRNA for RNA delivery. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used One-way ANOVA with Tukey’s multiple comparisons test; ***p < 0.001. e. Comparison of chemically synthesized pegRNA, epegRNA and L-epegRNA.
[0051] FIG. 12a-c: Ligated pegDNA mediates efficient prime editing. a. Overview of design schemes to generate epegDNA with both 5’a nd 3’ ends modifications (Left) , pegDNA (MS2) with both 5’ and 3’ ends modifications (Right) . The DT (DNA template) and PBS regions are replaceable DNA segments. b. Urea-PAGE (6%) of RNA / DNA chimera ligation products (epegDNA: +6 G to T at PRNP locus) . M, marker. c. Prime editing efficiencies by L-epegDNA and L-pegDNA (MS2) mediated RNA delivery. For L-epegDNA, 2 μg PEmax mRNA, 180pmol L-epegDNA, and 60 pmol nicking sgRNA were used. For L-pegDNA (MS2) , 2 μg nCas9 (H840A) mRNA, 2 μg MCP-RT mRNA, 180pmol L-pegDNA (MS2) , and 60 pmol nicking sgRNA were used.
[0052] FIG. 13a-g: Modifications in the evopreQ1 regions of epegRNA enhance prime editing efficiency. a. Schematic representation of end-modified epegRNA. b. Schematic of modifications in the evopreQ1 region. "evopreQ1-EM" contains three phosphorothioate linkages and three nucleotides with 2'-O-methyl modifications at the 3'end. "evopreQ1-M1" includes three phosphorothioate linkages and additional nucleotides with 2'-O-methyl modifications at the 3'end and loop. "evopreQ1-M2" has the same modifications as "evopreQ1-M1" but without modification at the loop. Red represents 2'-O-methyl modifications; thick lines represent phosphorothioate linkages; green indicates unmodified nucleotides. c-e. Prime editing efficiencies of modified epegRNAs from (b) were assessed in HEK293T cells by amplicon sequencing. The mRNA encoding PEmax (c) , PE6c (d) , and PE6d (e) were used to induce a 3 bp deletion at the EMX1 locus and a +1 T to A mutation at the HEK3 locus. For EMX1, each electroporation used 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol of nicking sgRNA. For HEK3, each electroporation used 0.5 μg PE mRNA, 3 pmol epegRNA, and 1 pmol nicking sgRNA. f-g. Prime editing efficiencies of evopreQ1-region modified epegRNAs from (b) were evaluated in HEK293T cells via deep sequencing at high doses. PEmax mRNA was used to induce a 3 bp deletion at the EMX1 locus (f) and a +1 T to A mutation at the HEK3 locus (g) . Each electroporation included 2 μg PE mRNA, 90 pmol epegRNA, and 30 pmol nicking sgRNA. Data are presented as the mean ± standard deviation from three independent biological replicates. Statistical analysis was performed using one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0053] FIG. 14a-g: Modifications in the scaffold regions of epegRNA enhance prime editing efficiency. a. Schematic representation of end-modified epegRNA and intended regions for modification. b. Schematic of modifications in the scaffold region. "Scaffold-UM" represents no modifications; "Scaffold-M1" contains 40 nucleotides modified with 2'-O-methyl modifications across three loops; Scaffold-M2" has 20 nucleotides modified in two loops. Red indicates 2'-O-methyl modifications, and orange represents unmodified nucleotides. c-e. Prime editing efficiencies of modified epegRNAs from (b) were evaluated in HEK293T cells via amplicon sequencing. The mRNA encoding PEmax (c) , PE6c (d) , and PE6d (e) were used to induce a 3 bp deletion at the EMX1 locus and a +1 T to A mutation at the HEK3 locus. For EMX1, each electroporation used 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol of nicking sgRNA. For HEK3, each electroporation used 0.5 μg PE mRNA, 3 pmol epegRNA, and 1 pmol nicking sgRNA. f-g. Prime editing efficiencies of scaffold-region modified epegRNAs from (b) were evaluated in HEK293T cells via deep sequencing at high doses. PEmax mRNA was used to induce a 3 bp deletion at the EMX1 locus (f) and a +1 T to A mutation at the HEK3 locus (g) . Each electroporation included 2 μg PE mRNA, 90 pmol epegRNA, and 30 pmol nicking sgRNA. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0054] FIG. 15a-d: Editing efficiencies of spacer-modified epegRNAs. a. Schematic representation of modification strategies in the spacer region of epegRNA. "3M3S" , "2M2S" and "1M1S" indicate three, two, or one phosphorothioate linkages, and three, two, or one nucleotide with 2'-O-methyl modifications at the 5'end of the spacer, respectively. Red indicates 2'-O-methyl modifications, thick lines denote phosphorothioate linkages, and blue indicates unmodified nucleotides. b-c. Prime editing efficiencies of spacer-modified epegRNAs described in (a) were evaluated in HEK293T cells using amplicon sequencing. The mRNA encoding PEmax, PE6c, and PE6d were used to induce a 3 bp deletion at the EMX1 locus (b) and a +1 T to A mutation at the HEK3 locus (c) . For EMX1, each electroporation included 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol nicking sgRNA. For HEK3, each electroporation included 0.5 μg PE mRNA, 3 pmol epegRNA, and 1 pmol nicking sgRNA. d. Prime editing efficiencies of spacer-modified epegRNAs described in (a) were assessed in HEK293T cells through amplicon sequencing following RNP delivery. For each sample of RNP delivery, 70 pmol PEmax protein, 140 pmol epegRNA, and 60 pmol nicking sgRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0055] FIG. 16a-f: Modifications in the RTT regions of epegRNA generally enhance prime editing efficiency. a. Schematic representation of end-modified epegRNA and intended regions for modification. b. Schematic of modifications in the RTT region. "RTT-UM" indicates no modifications; "RTT-M1" represents complete 2'-O-methyl modifications; "RTT-M2" indicates 2'-O-methyl modifications for the half close to the scaffold; "RTT-M3" represents complete 2'-fluoro modifications; "RTT-M4" indicates 2'-fluoro modifications for the half close to the scaffold; "RTT-M5" represents every other nucleotide being modified with 2'-fluoro. Red represents 2'-O-methyl modifications; yellow indicates 2'-fluoro modifications; gray represents unmodified nucleotides. c-d. Prime editing efficiencies of modified epegRNAs from (b) were evaluated in HEK293T cells via amplicon sequencing. The mRNA encoding PEmax, PE6c, and PE6d were used for a 3 bp deletion at the EMX1 locus (c) and a +1 T to A mutation at the HEK3 locus (d) . For EMX1, each electroporation used 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol nicking sgRNA. For HEK3, each electroporation used 0.5 μg PE mRNA, 3 pmol epegRNA, and 1 pmol nicking sgRNA. e-f. Prime editing efficiencies in HEK293T cells were determined for RTT-region-modified epegRNAs using amplicon sequencing at high doses. The mRNA encoding PEmax was employed to induce a 3 bp deletion at the EMX1 locus (e) and a +1 T to A mutation at the HEK3 locus (f) . Each electroporation consisted of 2 μg PE mRNA, 90 pmol epegRNA, and 30 pmol nicking sgRNA. Data and error bars represent the mean and standard deviation of three independent biological replicates. Statistical analysis was performed using one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p<0.001.
[0056] FIG. 17a-f: Modifications in the PBS regions of epegRNA generally enhance prime editing efficiency. a. Schematic representation of end-modified epegRNA and intended regions for modification. b. Schematic of modifications in the PBS region. "PBS-UM" indicates no modifications; "PBS-M1" represents complete 2'-O-methyl modifications; "PBS-M2" indicates complete 2'-fluoro modifications; "PBS-M3" represents 2'-fluoro modifications for the half close to evopreQ1; "PBS-M4" represents every other nucleotide being modified with 2'-fluoro. c-d. Prime editing efficiencies of modified epegRNAs from (b) were evaluated in HEK293T cells via deep sequencing. The mRNA encoding PEmax, PE6c, and PE6d were used for a 3 bp deletion at the EMX1 locus (c) and a +1 T to A mutation at the HEK3 locus (d) . For EMX1, each electroporation used 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol nicking sgRNA. For HEK3, each electroporation used 0.5 μg PE mRNA, 3 pmol epegRNA, and 1 pmol nicking sgRNA. e-f. Prime editing efficiencies in HEK293T cells were determined for PBS-region-modified epegRNAs using amplicon sequencing at high doses. The mRNA encoding PEmax was employed to induce a 3 bp deletion at the EMX1 locus (e) and a +1 T to A mutation at the HEK3 locus (f) . Each electroporation consisted of 2 μg PE mRNA, 90 pmol epegRNA, and 30 pmol nicking sgRNA. Data and error bars represent the mean and standard deviation of three independent biological replicates. Statistical analysis was performed using one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p<0.001.
[0057] FIG. 18a-d: Combinations of modifications in epegRNA facilitate efficient prime editing. a. Schematic representation of combinations of modified epegRNA designs. "Scaffold-M2" or "Sca-M2" represents 20 nucleotides modified in two loops of the scaffold; "RTT-M3" represents complete 2'-fluoro modifications of the RTT region. "PBS-M3" indicates 2'-fluoro modifications for half of the PBS sequence close to evopreQ1; "evopreQ1-M2" or "evQ1-M2" includes three phosphorothioate linkages and 15 nucleotides with 2'-O-methyl modifications at the 3'end. Red indicates 2'-O-methyl modifications; yellow represents 2'-fluoro modifications; thick lines denote phosphorothioate linkages; and gray indicates unmodified nucleotides. b-d. Pairwise combinations of the modifications shown in (a) were performed, including at least one fixed region modification. Prime editing efficiencies of these combinations for epegRNAs were assessed in HEK293T cells via amplicon sequencing. The mRNA encoding PEmax (b) , PE6c (c) , and PE6d (d) were used for a 3 bp deletion at the EMX1 locus. For each electroporation, 2 μg PE mRNA, 10 pmol epegRNA, and 3.3 pmol nicking sgRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0058] FIG. 19a-d: Indel percentages relative to total editing for chemically modified epegRNAs. a-c. The percentage of indels relative to total editing for chemically modified epegRNAs was calculated based on the data presented in FIG. 18. Analysis was conducted for mRNA encoding PEmax (a) , PE6c (b) , or PE6d (c) , to induce a 3 bp deletion at the EMX1 locus. d. Prime editing efficiencies in HEK293T cells using chemically modified epegRNAs at moderate doses were assessed by amplicon sequencing. PEmax mRNA was used to induce a 3 bp deletion at the EMX1 locus. For each electroporation, 2 μg PE mRNA, 30 pmol epegRNA, and 10 pmol nicking sgRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0059] FIG. 20a-d: Prime editing efficiencies in Hepa1-6 cells. a-b. Prime editing efficiencies of different epegRNAs (a) and nicking sgRNAs (b) for the Pcsk9 locus in Hepa1-6 cells were determined using plasmid delivery. For each sample, 800 ng PEmax plasmid, 200 ng epegRNA plasmid, and 83 ng nicking sgRNA plasmid were used. c. Prime editing efficiencies mediated by mRNA encoding different versions of PE were determined by amplicon sequencing using the end-modified epegRNA. d. Prime editing efficiencies in Hepa1-6 cells using three pegRNA variants and mRNA encoding PE6c were determined by amplicon sequencing. For c and d, 100 ng PE mRNA, 150 ng epegRNA, and 50 ng nicking sgRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0060] FIG. 21a-g: Modified epegRNA enhances prime editing efficiency in vivo. a. Schematic representation of the experimental procedure for animal studies. A mixture of LNPs that encapsulate PE mRNA and modified epegRNA with nicking sgRNA were administered via intravenous tail vein injection. Seven days after injection, mouse liver tissues were collected for analysis. b. Prime editing efficiency at the Pcsk9 locus for a 4 bp insertion (insTTAC) in mice livers by mRNA encoding three different versions of PE. The end-modified pegRNA (EM-pegRNA) , with modifications of three nucleotides at both the 5'and 3'ends, was used. c. Prime editing efficiency at the Pcsk9 locus for a 4 bp insertion in mice livers by three pegRNA variants: end-modified pegRNA (EM-pegRNA) , end-modified epegRNA (EM-epegRNA) , and heavily modified epegRNA (HM-epegRNA) . PE6c mRNA was used. d. Prime editing efficiency at the Pcsk9 locus for a 4 bp insertion mice livers, with the addition of mRNA encoding MLH1dn or MLH1-SB. PE6c and MLH1dn or MLH1-SB mRNA were co-encapsulated in the same LNP. e. Prime editing efficiency at the Pcsk9 locus for a 4 bp insertion mice livers, with the addition of mRNA encoding Vpx or both Vpx and MLH1-SB. The mRNA encoding PE6c and Vpx or Vpx plus MLH1-SB were co-encapsulated in LNPs. f-g. Serum Pcsk9 (f) and cholesterol (g) levels in mice were determined. The HM-epegRNA and PE6c mRNA were used for d-f. Each data point represents an individual mouse. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; ns, not significant; *p < 0.05; **p < 0.01; ***p < 0.001.
[0061] FIG. 22a-c: The mRNA encoding different proteins for prime editing. a-b. Prime editing efficiencies at the Pcsk9 locus for a 4 bp insertion in Hepa1-6 cells were determined, with the addition of mRNA encoding MLH1dn or MLH1-SB (a) , or Vpx or Vpx with MLH1-SB (b) . The end-modified epegRNA and PE6c mRNA were used. For each sample, 100 ng PE mRNA, 150 ng epegRNA, 50 ng nicking sgRNA, and 34 ng MLH1dn / MLH1-SB / Vpx mRNA were used. Data and error bars represent the mean and standard deviation of three independent biological replicates. Data analysis used one-way ANOVA with Tukey's multiple comparisons test; **p < 0.01. c. Deep sequencing results of a 4 bp insertion at the Pcsk9 locus in mouse livers were obtained by amplicon sequencing. The HM-epegRNA and mRNA encoding PE6c, Vpx, and MLH1-SB were used, as shown in FIG. 4e.DETAILED DESCRIPTIONDefinitions
[0062] It is to be noted that the term “a” or “an” entity refers to one or more of that entity. for example, “an antibody, ” is understood to represent one or more antibodies. As such, the terms “a” (or “an” ) , “one or more, ” and “at least one” can be used interchangeably herein. Ligation of Nucleic Acid Fragments
[0063] Ligating relatively short RNA molecules (e.g., 100 nt each) can potentially produce much longer ones (e.g., 200 or 300 nt) than what a conventional chemical synthesis can accomplish. The current RNA ligation technology, however, is met with significant challenges. In an RNA ligation process, RNA ligases are applied to join RNA with 5′-phosphate and 3′-hydroxyl. The so-called “splint ligation” employes a splint DNA to hybridize to the 3’a nd 5’ ends of two RNA fragments and direct ligase activity in joining them. The efficiency of splint ligation varies but often remains low.
[0064] The instant disclosure has designed and tested a new splint ligation process, which can effectively generate chemically modified RNA molecules (e.g., pegRNAs and epegRNAs) with high yield and purity. The resulting RNA molecules are referred to as L-pegRNA and L-epegRNA, respectively. The potency of the L-epegRNA was examined in different cell lines and human primary cells via RNA delivery. Compared to epegRNA produced by IVT, L-epegRNA demonstrated up to 845.2-fold improved editing efficiency in prime editing. In addition, the L-epegRNA greatly improved the editing efficiency of PE in RNP format, enabling RNP delivery of PE surpass plasmids encoded PE in editing efficiency for most comparisons examined.
[0065] With the new splint ligation method, the inventors also generated chemically modified epegRNA with a relatively large insertion. The resulting 210 nt RNA was beyond the current limitation of chemical synthesis (e.g., solid-phase synthesis) . This epegRNA enabled approximately 60%efficiencies of 17 bp insertion via either RNP or RNA delivery. L-epegRNA dramatically boosted the efficiency of PE via RNP and RNA delivery, and the substantially reduced cost and efficient ligation process can facilitate L-epegRNA preparation for broad usage including therapeutic applications.
[0066] Prime editing (PE) has emerged as a promising tool for various applications, particularly in the field of therapeutics. Despite its potential, the editing efficiencies of PE through ribonucleoprotein (RNP) and RNA delivery are not optimal due to the challenge in synthesizing long pegRNA (>125 nt) via solid-phase synthesis. The instant technology, therefore, provides an efficient, rapid, and cost-effective way for generating chemically modified pegRNA (125-145 nt) and epegRNA (170-190 nt) for prime editing. The instant technology, therefore, paves the way for the use of PE in therapeutics and various other applications.
[0067] The instant technology is greatly improved over the conventional splint-mediated nucleic acid ligation technology. In the conventional splint-mediated RNA ligation technology, a splint DNA fragment that is partially complementary to the 3’ terminal portion of one RNA molecule and partially complementary to the 5’ terminal portion of another RNA molecule is used to hybridize to both and thus bring the two termini to proximity of each other. In the presence of an RNA ligase, the two termini are connected, resulting in a fusion of both RNA molecules.
[0068] The instant inventors made multiple modifications to the conventional process to make it practical to generate high quality and high purity long RNA molecules at low cost. One of such modifications is to repeat the denaturing-annealing process and another is to increase the ligation temperature from 25 ℃ to about 37 ℃. In the conventional RNA ligation technology, at 25 ℃, the ligation process takes overnight to complete, at very low ligation rates. Unexpectedly, when the ligation temperature was raised to 37 ℃, each ligation step can be shortened to just 30-60 minutes, with greatly higher ligation rates. When this process is repeated, even the total time with the 2-3 cycles is no longer than 3 hours. Also importantly, as demonstrated in the experimental examples, the resulting ligated RNA molecules had much higher functions.
[0069] Yet another discovery of the present disclosure is that the T4 RNA ligase 2 (T4 Rnl2) is particularly suitable for the improved process, in particular when three or more RNA fragments are ligated. In such a process, the use of T4 Rnl2 can minimize self-circularization of each RNA fragment. Certain other conditions of the ligation process have also been optimized including, for instance, fragment / splint ratios and ligase concentrations, without limitation.
[0070] Accordingly, one embodiment of the present disclosure provides a method for ligating two or more single-stranded target nucleic acid fragments. In one embodiment, the method entails (a) incubating, in a sample, the two or more single-stranded target nucleic acid fragments with one or more single-stranded splint DNA fragment (s) at a denaturing temperature. In some embodiments, subsequently, the method entails (b) reducing the temperature of the sample to allow the target nucleic acid fragments and the splint DNA fragment (s) to form an at least partially duplex molecule by virtue of their sequence complementarity, wherein there is no gap between each adjacent target nucleic acid fragment in the duplex molecule. In some embodiments, the method further entails (c) incubating the sample at a ligation temperature, in the presence of a nucleic acid ligase, to allow the ligase to ligate the target nucleic acid fragments in the duplex molecule to form a ligated nucleic acid fragment. In some embodiments, steps (a) - (c) are repeated for at least one more time to form more ligated nucleic acid fragments.
[0071] A relatively simple version of the process includes ligating two single-stranded target nucleic acid molecules (e.g., two RNA fragments) . With reference to FIG. 2a, one of the RNA fragments is referred to as “acceptor RNA” while the other is referred to as “donor RNA. ” The splint DNA (also single-stranded) does not have to very long, so long as it has sufficient complementarity with both RNA fragments. For instance, the splint DNA can have 15-nt sequence complementarity to the 3’ -terminal portion of the acceptor RNA and 15-nt sequence complementarity to the 5’ -terminal portion of the donor RNA.
[0072] As provided, once the at least partially duplex molecule if formed among the two RNA fragments and the splint DNA, no gap is left between the 3’ terminus of the acceptor RNA and the 5’ terminus of the donor RNA, such that these two ends can be ligated in the presence of a suitable nucleic acid ligase.
[0073] Nevertheless, it has been observed that relatively longer splint DNA can improve the ligation efficiency. In one embodiment, each splint DNA is at least 30 nt, or at least 40 nt, 50 nt, 60 nt, 70 nt, 80 nt, 90 nt, or 100 nt in length. (Perhaps longer splint DNA fragments are required)
[0074] In some embodiments, each target molecule (e.g., RNA fragment) has at least 10 nt sequence complementarity to a corresponding splint DNA. In some embodiments, each target molecule (e.g., RNA fragment) has at least 15 nt, 20 nt, 25 nt, or 30 nt sequence complementarity to a corresponding splint DNA. In some embodiments, each target molecule (e.g., RNA fragment) has no more than 100 nt, 90 nt, 80 nt, 70 nt, 60 nt, 50 nt, 40 nt, 30 nt, 25 nt, or 20 nt sequence complementarity to a corresponding splint DNA.
[0075] It is appreciated that the sequence complementarity does not need to be 100%sequence complementarity, so long as it is sufficient (e.g., 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%or 75%) to allow hybridization under the annealing conditions as described.
[0076] In some embodiments, each target nucleic acid fragment has a length of 10 nt to 150 nt. In some embodiments, the resulting, ligated target nucleic acid molecule has a length that is at least 50 nt in length, or at least 60 nt, 70 nt, 80 nt, 90 nt, 100 nt, 110 nt, 120 nt, 130 nt, 140 nt, 150 nt, 160 nt, 170 nt, 180 nt, 190 nt, 200 nt, 210 nt, 220 nt, 230 nt, 240 nt, 250 nt, 260 nt, 270 nt, 280 nt, 290 nt, 300 nt, 350 nt, 400 nt, or 500 nt.
[0077] In some embodiments, the sample includes three or more single-stranded target nucleic acid molecules (e.g., three or more RNA fragments) . One such example is illustrated in FIG. 11a, lower panel. In this example, a single splint DNA is used, which hybridizes to the 3’ terminal portion of the acceptor RNA, the 5’ termina portion of the donor RNA, and the entirety of a “bridge” RNA fragment in the middle that has 100%sequence complementarity to the splint DNA.
[0078] In some embodiments, two or more splint DNA fragments can be used, each is used to connect two or more single-stranded target nucleic acid molecules. Apparently, the splint DNA fragments themselves do not need to leave no gap between them, so long as they serve as “splints” between the single-stranded target nucleic acid molecules.
[0079] Therefore, when multiple single-stranded target nucleic acid molecules and multiple splint DNA fragments are used, the resulting duplex has two strands, one of which concatenates all of the single-stranded target nucleic acid molecules, with no gap between each adjacent ones, and the other strand includes the splint DNA fragments, which optionally have gaps between them.
[0080] In some embodiments, the ratio between the two types of nucleic acid molecules are also tweaked. In one embodiment, each target nucleic fragment and each splint DNA fragment have a molar ratio, in the sample, that is from 0.5: 1.5 to 1.5: 0.5, or preferably from 0.8: 1.2 to 1.2: 0.8.
[0081] In step (a) , the sample with the single-stranded target nucleic acid molecules (e.g., two RNA fragments) and the splint DNA fragment (s) is incubated at the denaturing temperature. The denaturing temperature, in some embodiments, is 55 ℃ to 98 ℃. In some embodiments, it is 60 ℃ to 80 ℃, or 65 ℃ to 75 ℃, or 66 ℃ to 74 ℃, or 67 ℃ to 73 ℃, or 68 ℃ to 72 ℃, or 69 ℃ to 71 ℃, or at 70 ℃. The denaturing temperature, in some embodiments, is 55 ℃ to 65 ℃, 60 ℃ to 70 ℃, 65 ℃ to 75 ℃, 70 ℃ to 80 ℃, 75 ℃ to 85 ℃, 80 ℃ to 90 ℃, or 85 ℃ to 95 ℃.
[0082] The denaturing step can take, for instance, from 10 seconds to 10 minutes. In some embodiments, the denaturing step takes 20 seconds to 8 minutes, 30 seconds to 5 minutes, 40 seconds to 4 minutes, 1 to 4 minutes, 2 to 5 minutes, 1 to 3 minutes, 1 to 2 minutes, 2 to 3 minutes, 3 to 4 minutes, 4 to 5 minutes, without limitation.
[0083] Upon completion of the denaturing step, the temperature of the sample is reduced to allow ligation of the nucleic acid fragments. The ending temperature after the reduction, in some embodiments, is the desired ligation temperature. However, the temperature can be reduced further to, e.g., room temperature, before it is raised to the ligation temperature (e.g., 37 ℃) .
[0084] In some embodiments, the reducing of temperature is at a pace that is about 0.01 ℃ / sto 0.5 ℃ / s. In some embodiments, the reducing of temperature is at a pace that is about 0.02 ℃ / sto 0.4 ℃ / s, 0.04 ℃ / sto 0.3 ℃ / s, 0.06 ℃ / sto 0.2 ℃ / s, 0.08 ℃ / sto 0.15 ℃ / s, or at about 0.1 ℃ / s.
[0085] The ligation temperature at the ligation step (c) , in some embodiments, is 20 ℃ to 42 ℃. As described above, however, a ligation temperature that is higher than the conventional one, 25 ℃, can increase the ligation efficiency. Accordingly, in some embodiments, the ligation temperature is 28 ℃ to 42 ℃, 30 ℃ to 41 ℃, 31 ℃ to 40 ℃, 32 ℃ to 39 ℃, 33 ℃ to 38 ℃, 34 ℃ to 38 ℃, 35 ℃ to 38 ℃, 36 ℃ to 38 ℃, or 36.5 ℃ to 37.5 ℃.
[0086] In some embodiments, the ligation step is for a time period that is 15 minutes to 3 hours. In some embodiments, the ligation step is for a time period that is 20 minutes to 2 hours, 20 minutes to 1.5 hours, 25 minutes to 75 minutes, or 25 to 65 minutes. In some embodiments, in the first iteration of steps (a) - (c) , the ligation time can be relatively longer, such as 45-100 minutes, 45-90 minutes, or 50-70 minutes. In some embodiments, in the subsequent iteration (s) of steps (a) - (c) , the ligation time can be relatively shorter, such as 10-50 minutes, 20-45 minutes, or 25-35 minutes.
[0087] The ligation step, in some embodiments, is carried out in the presence of a suitable nucleic acid ligase and other required reagents such as ATP and buffer. Nevertheless, the ligase and / or any other regents can be added at any step, e.g., before or during the denaturing step. In some embodiments, an additional amount of the nucleic acid ligase can be added after the denaturing temperature is reduced such as when these steps are repeated.
[0088] In some embodiments, the ligase is T4 RNA ligase 2 (T4 Rnl2) . In some embodiments, each time when step (c) is carried, a new dose of the ligase is added. In some embodiments, an amount of 0.5 to 5 μg of the ligase is added. In some embodiments, an amount of 1 to 4 μg of the ligase is added. In some embodiments, in the first iteration of step (c) , a relatively higher dose of the ligase is added, such as 2 to 4 μg, 2.5 to 3.5 μg, or 2.8 to 3.2 μg. In some embodiments, in the subsequent iterations of step (c) , a relatively lower dose of the ligase is added, such as 0.5 to 3 μg, 1 to 2 μg, or 1.4 to 1.6 μg, without limitation.
[0089] In some embodiments, upon completion of the ligation, the ligated target ligated single-stranded target nucleic acid molecule is purified, which entails, e.g., digesting and removing the splint DNA fragment (s) .
[0090] As noted, the steps (a) - (c) are repeated. In some embodiments, they are repeated once, twice, three time, or four times. In some embodiments, they are repeated once or twice. In some embodiments, they are repeated twice, for a total of three times.
[0091] As noted, the present technology is particular useful for preparing long nucleic acid molecules having chemically modification or including non-natural nucleotide (s) , which cannot be readily made through in vitro transcription or solid-phase synthesis.
[0092] In some embodiments, at least one of the single-stranded target nucleic acid fragments (e.g., RNA fragments) is chemically modified, such as on the uridine nucleosides. In some embodiments, the chemically modified uridine nucleosides are N1-methylpseudouridines. In some embodiments, the chemically modified nucleobase is selected from 5-formylcytidine (5fC) , 5-methylcytidine (5meC) , 5-methoxycytidine (5moC) , 5-hydroxycytidine (5hoC) , 5-hydroxymethylcytidine (5hmC) , 5-formyluridine (5fU) , 5-methyluridine (5-meU) , 5-methoxyuridine (5moU) , 5-carboxymethylesteruridine (5camU) , pseudouridine (Ψ) , N1-methylpseudouridine (me1Ψ) , N6-methyladenosine (me6A) , or thienoguanosine (thG) .
[0093] In some embodiments, the chemically modified ribose is selected from 2’ -O-methyl (2’-O-Me) , 2’ -Fluoro (2’ -F) , 2’ -deoxy-2’ -fluoro-beta-D-arabino-nucleic acid (2’ F-ANA) , 4’ -S, 4’ -SFANA, 2’ -azido, UNA, 2’ -O-methoxy-ethyl (2’ -O-ME) , 2’ -O-Allyl, 2’ -O-Ethylamine, 2’-O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, or N-MeO-aminooxy BNA.
[0094] In some embodiments, one or more phosphodiester linkage in the RNA is chemically modified. In some embodiments, the chemically modified phosphodiester linkage is selected from Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’ , 5’ -amide, N3’ -phosphoramidate (NP) , Phosphodiester (PO) , or 2’ , 5’ -phosphodiester (2’ , 5’ -PO) . pegRNA Molecules
[0095] The instant technology finally makes is possible to prepare long (e.g., longer than 100 nt, 120 nt, or 150 nt) RNA molecules with site-specific modifications, which is incapable to be produced by in vitro transcription. Accordingly, the instant technology can prepare prime editing guide RNA (pegRNA) having additional, unconventional elements.
[0096] Accordingly, one embodiment provides a prime editing guide RNA that includes (a) a single guide RNA (sgRNA) , (b) a reverse transcriptase (RT) template (RTT) and primer binding site (PBS) , and (c) an RNA aptamer or an RNA pseudoknot motif. In some embodiments, the pegRNA is at least 125 nt in length, preferably at last 150 nt, 180 nt, 200 nt, 250 nt or 300 nt in length. In some embodiments, the pegRNA includes at least two, three, or four of such RNA aptamers and / or RNA pseudoknot motifs.
[0097] RNA aptamers and RNA pseudoknot motifs are useful elements for RNA molecules, but it has been challenging to prepare long RNA (e.g., pegRNA) with such elements which includes site-specific modifications.
[0098] RNA aptamers are short RNA sequences that bind a specific target molecule, or family of target molecules. Examples, without limitation, include an MS2 aptamer, a tetrahymena thermophila group I intron, a theophylline aptamer, an ATP aptamer, a guanine quadruplex aptamer, a thrombin aptamer (TBA) , a vascular endothelial growth factor (VEGF) aptamer, an aptamer against HIV-1 reverse transcriptase, a hemin binding aptamer (HBA) , a riboswitch, a ribozyme.
[0099] The MS2 coat protein is an RNA binding protein. MS2 aptamers have been developed based on the MS2 coat protein’s RNA binding domain. These aptamers, referred to as MS2 aptamers, are RNA sequences that mimic the natural RNA binding sites recognized by the MS2 coat protein.
[0100] Tetrahymena thermophila Group I Introns can fold into a complex three-dimensional structure. Theophylline aptamers bind specifically to theophylline, a methylxanthine drug. An ATP aptamer binds to adenosine triphosphate (ATP) with high affinity. Guanine quadruplex aptamers fold into guanine quadruplex structures and can bind to various targets, including proteins, small molecules, and metal ions.
[0101] Thrombin aptamer (TBA) binds specifically to thrombin, a key enzyme in blood coagulation. Vascular endothelial growth factor (VEGF) aptamer specifically binds to VEGF, a signaling protein involved in angiogenesis. Aptamer against HIV-1 reverse transcriptase target viral proteins such as HIV-1 reverse transcriptase. Hemin binding aptamers (HBA) bind to hemin, an iron-containing molecule.
[0102] Riboswitches are RNA elements that bind specific small molecules, leading to changes in gene expression. They are found in the untranslated regions of mRNA and regulate gene expression in response to cellular metabolites.
[0103] Ribozymes are RNA molecules with catalytic activity. They can be engineered to have binding properties similar to aptamers, and some ribozymes have been utilized in biosensing applications.
[0104] Non-limiting examples of RNA pseudoknot motifs include an H-type pseudoknot, a kissing-loop pseudoknot, a triple helix pseudoknot, a complex pseudoknot, a frame-shift pseudoknot, a riboswitch pseudoknot, a G-quadruplex pseudoknot, an intermolecular pseudoknot.
[0105] RNA pseudoknots are structural motifs commonly found in RNA molecules. They occur when a single-stranded region of RNA, typically a loop, forms base pairs with a region of the RNA molecule outside of the loop, causing the RNA to fold back on itself. This results in a complex three-dimensional structure where a helix (stem) is interrupted by a loop, which is then followed by another helix that stacks on top of the first helix, creating a knot-like structure.
[0106] Pseudoknots are classified based on the number and arrangement of helices involved. The simplest type is the H-type pseudoknot, which consists of two helices (H1 and H2) connected by a loop, where the loop region of H1 base pairs with a region upstream or downstream of H2. There are also more complex pseudoknots with multiple helices and loops.
[0107] RNA pseudoknots play important roles in various biological processes, including translation, RNA catalysis, ribosomal frameshifting, and RNA-protein interactions. Their structural diversity and functional significance make them interesting targets for RNA engineering and design in synthetic biology and biotechnology applications.
[0108] H-type Pseudoknot is the simplest type of pseudoknot, consisting of two helices (H1 and H2) connected by a loop, where the loop of H1 base pairs with a region upstream or downstream of H2.
[0109] Kissing-loop Pseudoknot: In this motif, two hairpin loops form base pairs with each other, creating a pseudoknot structure. This type of pseudoknot is often involved in RNA-RNA interactions, such as in viral RNA genomes and riboswitches.
[0110] Triple Helix Pseudoknot: This pseudoknot consists of three helices, where the third helix forms base pairs with both the loop region of the first helix and the stem region of the second helix.
[0111] Complex Pseudoknots: These pseudoknots contain multiple helices and loops, often forming intricate three-dimensional structures. They are commonly found in ribosomal RNA, viral RNA genomes, and RNA enzymes (ribozymes) .
[0112] Frame-shift Pseudoknots: These pseudoknots play a crucial role in ribosomal frameshifting, a mechanism used by certain viruses to regulate the translation of their RNA genomes. The pseudoknot causes a shift in the reading frame during translation, leading to the production of alternative protein products.
[0113] Riboswitch Pseudoknots: Some riboswitches, which are RNA elements that regulate gene expression in response to small molecule binding, contain pseudoknot structures. These pseudoknots are involved in ligand binding and conformational changes that control gene expression.
[0114] G-quadruplex Pseudoknots: In this motif, a G-quadruplex structure (formed by guanine-rich sequences) interacts with a nearby single-stranded region to form a pseudoknot. These structures are involved in various cellular processes, including telomere maintenance and gene regulation.
[0115] Intermolecular Pseudoknots: While most pseudoknots involve interactions within a single RNA molecule, intermolecular pseudoknots occur when two separate RNA molecules interact to form a pseudoknot structure. These interactions can occur between different regions of viral RNA genomes or between RNA molecules and proteins.
[0116] Another example is evopreQ1, developed as useful for stabilizing RNA molecules, which is a modified prequeosine1-1 riboswitch aptamer (Nelson et al., Nat Biotechnol. 2022 Mar; 40 (3) : 402-410) .
[0117] In some embodiments, the pegRNA includes at least two, three, or four of such RNA aptamers and / or RNA pseudoknot motifs.
[0118] Through comprehensive engineering of modification patterns for epegRNAs, the instant inventors have identified a number of modification patterns to each region within the epegRNA, which significantly enhanced editing efficiencies both in vitro and in vivo as compared to conventionally end-modified epegRNAs.
[0119] A typical epegRNA includes, from the 5’ end to the 3’ end, a sgRNA (spacer +scaffold) , an RTT (reverse transcriptase template) , a PBS (primer binding site) , and an RNA aptamer or an RNA pseudoknot motif (e.g., evopreQ1) which includes a RNA motif and a 3’ single-stranded region.
[0120] In the pseudoknot motif, it was shown that “heavy” modification of the 3’ single-stranded region alone significantly increased editing efficiency (e.g., FIG. 13b, evopreQ1-M2) while additional modification to the stem-loop motif did not further increase it (e.g., evopreQ1-M1) . In evopreQ1-M2, all 15 3’ terminal ribose’s were subjected to 2'-O-methyl modifications. In addition, the three 3’ terminal linages were modified.
[0121] The phosphodiester linkage of an pegRNA can be chemically modified. In some embodiments, the chemically modified phosphodiester linkage is selected from Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’ , 5’ -amide, N3’ -phosphoramidate (NP) , Phosphodiester (PO) , or 2’ , 5’ -phosphodiester (2’ , 5’ -PO) . In some embodiments, the modified linages are phosphorothioate (PS) linkages.
[0122] In some embodiments, therefore, in an pegRNA of the present disclosure, the 3’ end portion is “heavily” modified. As used herein, the 3’ end portion of the pegRNA may include the 3’ (last) 10, 20, 30, 40, 50, 60, 70, or 80 nucleotides. In some embodiments, the 3’ end portion includes the entire RNA aptamer or RNA pseudoknot motif. In some embodiments, the heavy modification refers to modification of at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the nucleotides within the 3’s end portion. For instance, in embodiment, at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the nucleotides within the 20, 30, 40, 50, 60, 70 or 80 3’s end nucleotides, or of the entire RNA aptamer or RNA pseudoknot motif, are modified.
[0123] In a particular embodiment, at least 5 of the 15 3’ nucleotides are modified. In some embodiments, at least 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of the 15 3’ nucleotides are modified. In some embodiments, among these modifications, at least the three 3’ terminal nucleotides are modified.
[0124] The modifications, without limitation, may be ribose modification selected from 2’ -O-methyl (2’ -O-Me) , 2’ -Fluoro (2’ -F) , 2’ -deoxy-2’ -fluoro-beta-D-arabino-nucleic acid (2’ F-ANA) , 4’ -S, 4’ -SFANA, 2’ -azido, UNA, 2’ -O-methoxy-ethyl (2’ -O-ME) , 2’ -O-Allyl, 2’ -O-Ethylamine, 2’ -O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, or N-MeO-aminooxy BNA. In one example, the modification is ribose modification with 2’ -O-methyl (2’ -O-Me) . In another example, the modification is ribose modification with 2’ -Fluoro (2’ -F) . In a preferred embodiment, all of the modifications within the 15 3’ nucleotides are ribose modifications with 2’ -O-methyl (2’ -O-Me) .
[0125] In some embodiments, the one or more motif (e.g., stem-loop) within the RNA aptamer or RNA pseudoknot motif also includes “heavy” modifications (e.g., at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the nucleotides are modified) . In some other embodiments, the one or more motif (e.g., stem-loop) within the RNA aptamer or RNA pseudoknot motif can include no or low modifications. For instance, in one embodiment, no more than 50%of the nucleotides in at least one of the stem-loops are modified. In some embodiments, no more than 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%or 5%of the nucleotides in at least one of the stem-loops are modified. In some embodiments, no more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%or 5%of the linkages in at least one of the stem-loops are modified. In some embodiments, no more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%or 5%of the nucleotides in each of the stem-loops are modified. In some embodiments, no more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%or 5%of the linkages in each of the stem-loops are modified.
[0126] A typical sgRNA scaffold includes 4 stem-loops (e.g., based the sequence of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUU GAAAAAGUGGGACCGAGUCGGGUCC) . As shown in Example 2 and FIG. 14, “heavy” modification to the distant portion of the first stem-loop (numbered from 5’ to 3’ ) and to the third stem-loop significantly improved editing efficiency while those to the fourth stem-loop did not.
[0127] In some embodiments, therefore, in an pegRNA of the present disclosure, the scaffold of the sgRNA includes, from 5’ to 3’ , a first, a second, a third and a fourth stem-loops, and wherein the first and / or the third stem-loops include modifications. In some embodiments, the first stem-loop includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 modified ribose’s. In some embodiments, the third stem-loop includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 modified ribose’s.
[0128] In some embodiments, the fourth stem-loop includes no more than 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 modifications. In some embodiments, the fourth stem-loop includes no more than 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 modified ribose’s.
[0129] In some embodiments, the second stem-loop includes no more than 9, 8, 7, 6, 5, 4, 3, 2 or 1 modifications. In some embodiments, the second stem-loop includes no more than 9, 8, 7, 6, 5, 4, 3, 2 or 1 modified ribose’s.
[0130] The modifications, without limitation, may be ribose modification selected from 2’ -O-methyl (2’ -O-Me) , 2’ -Fluoro (2’ -F) , 2’ -deoxy-2’ -fluoro-beta-D-arabino-nucleic acid (2’ F-ANA) , 4’ -S, 4’ -SFANA, 2’ -azido, UNA, 2’ -O-methoxy-ethyl (2’ -O-ME) , 2’ -O-Allyl, 2’ -O-Ethylamine, 2’ -O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, or N-MeO-aminooxy BNA. In one example, the modification is ribose modification with 2’ -O-methyl (2’ -O-Me) . In another example, the modification is ribose modification with 2’ -Fluoro (2’ -F) . In a preferred embodiment, all of the modifications within the 15 3’ nucleotides are ribose modifications with 2’ -O-methyl (2’ -O-Me) .
[0131] In a preferred embodiment, the first stem-loop includes at least 8, 10 or 12 ribose modifications with 2’ -O-methyl (2’ -O-Me) . In a preferred embodiment, the third stem-loop includes at least 8, 10 or 12 ribose modifications with 2’ -O-methyl (2’ -O-Me) . In a preferred embodiment, the second and fourth stem-loops do not include modifications.
[0132] With respect to the 5’ spacer region, it was discovered the reduced 5’ end modification improved editing efficiency (FIG. 15d) . In some embodiments, therefore, the 5’ spacer includes no modification except on the three 3’ end nucleotides. In some embodiments, the 5’ spacer includes no modification except on the two 3’ end nucleotides. The two or three 3’ end modifications may be ribose modifications with 2’ -O-methyl (2’ -O-Me) and / or phosphorothioate (PS) linkages.
[0133] With respect to the RTT, it was discovered that heavy 2'-fluoro modifications led to significantly improved editing efficiencies (FIG. 16, RTT-M3) . In some embodiments, therefore, in an pegRNA of the present disclosure, the RTT includes modifications on at least 40%, 50%, 60%, 70%, 80%or 90%of the nucleotides. In some embodiments, such modifications are 2'-fluoro modifications.
[0134] Moreover, with respect to the PBS, 2'-fluoro modifications generally improved editing efficiencies and full 2'-fluoro modifications of the 3’ half of the PBS region achieved the best results (FIG. 17, PBS-M3) . In some embodiments, therefore, in an pegRNA of the present disclosure, the PBS includes modifications on at least 20%, 30%, 40%, 50%, 60%, 70%, 80%or 90%of the nucleotides. In some embodiments, the PBS includes modifications on at least 50%, 60%, 70%, 80%or 90%of the nucleotides within the 3’ half, and no more than 50%, 40%, 30%, 20%or 10%of the nucleotides within the 5’ half. In some embodiments, such modifications are 2'-fluoro modifications.
[0135] In an example embodiment, the preRNA includes a combination of the following modifications: (i) the spacer include three consecutive phosphorothioate (PS) linkages at the 5’end, (ii) the first and third stem-loops of the scaffold of the sgRNA each comprises at least 5, 6, 7, 8, 9, 10, 11 or 12 nucleotides with 2'-O-methyl (2'-O-Me) modifications, and (iii) at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 3’ end nucleotides of the pegRNA are modified with 2'-O-methyl (2'-O-Me) .
[0136] In some embodiments, (i) the spacer comprises exactly three consecutive PS linkages at the 5’ end, (ii) the third and / or fourth stem-loop of the scaffold of the sgRNA includes no more than 6, 5, 4, 3, 2, 1 or 0 modified ribose’s, and (iii) the step-loop of (d) comprises no more than 50%, 40%, 30%, 20%or 10%modified ribose’s. EXAMPLES Example 1. Ligated pegRNA Enables Efficient Prime Editing
[0137] Prime editing (PE) has emerged as a promising tool for various applications, particularly in the field of therapeutics. Despite its potential, the editing efficiencies of PE through ribonucleoprotein (RNP) and RNA delivery are not optimal due to the challenge in synthesizing long pegRNA (>125 nt) via solid-phase synthesis.
[0138] This example developed an efficient, rapid, and cost-effective way for generating chemically modified pegRNA (125-145 nt) and epegRNA (170-190 nt) . These RNAs, referred as L-pegRNA and L-epegRNA, were made using an optimized splint ligation approach, resulting in approximately 90%production efficiency. This example examined the potency of L-epegRNA in different cell lines and human primary cells. L-epegRNA greatly enhanced editing efficiencies, achieving up to several hundred-fold improvements, when utilizing RNA delivery of PE in comparison to epegRNA generated through in vitro transcription (IVT) . Notably, L-epegRNA also significantly enhanced the efficiency of PE in the RNP format, outperforming plasmid-encoded PE in the majorities of the comparisons.
[0139] Chemically modified epegRNA for large insertion exceeds the current limits of solid-phase synthesis. In this approach, a modified 210 nt L-epegRNA was successfully generated, enabling a 17 bp insertion at ~60%editing efficiency through both RNP and RNA delivery of PE. This technology, therefore, expands the capabilities of long RNA synthesis, providing a cost-effective and high-quality solution for obtaining pegRNA and epegRNA with desired chemical modifications. This approach paves the way for the use of PE in therapeutics and various other applications. Methods and Materials Mammalian cell culture
[0140] Human HEK293T, Huh7, Hela, U2OS, K562, and mouse Hepa1-6 cells were procured from ATCC. HEK293T, Huh7, and Hepa1-6 cells were cultured in DMEM (Thermo Fisher) . Hela and U2OS cells were maintained in DMEM / F12 (1: 1) medium (HyClone) , while K562 cells were cultured in RPMI-1640 medium with l-glutamine (Gibco) . Each culture medium was supplemented with 10%Fetal Bovine Serum (FBS) and 1%Penicillin / Streptomycin (P / S) . All cell types underwent incubation, maintenance, and culture at 37 ℃ with 5%CO2, and were routinely tested to ensure the absence of mycoplasma contamination. Plasmid construction
[0141] Plasmids for the mammalian expression of prime editors and other proteins were cloned using the -Basic Seamless Cloning and Assembly Kit (TransGen Biotech, CU201) . Plasmids designed for expressing pegRNAs were constructed utilizing a pGCL acceptor plasmid. The pegRNA sequences were obtained by PCR using Phanta Max Super-Fidelity DNA Polymerase (Vazyme, P505) . The U6 promoter-driven nicking gRNA mammalian expression plasmid was created using the gRNA cloning Vector (Addgene #41824) . Nicking gRNA cloning was executed by phosphorylation of oligonucleotides corresponding to spacer sequences with T4 PolyNucleotide Kinase (PNK) (New England Biolabs, M0201) , followed by annealing and ligation into BbsI-digested Gcl vector. Plasmids designed for expressing the prime editor protein were derived from pCMV-PEmax (Addgene #174820) and pCMV-PEmax-P2A-hMLH1dn (Addgene #174828) .
[0142] To obtain the PE2 protein, we constructed a pET28a-PE2-His vector that underwent Escherichia coli codon optimization. The SpCas9 (H840A) sequence was derived from pET28a-Cas9-His (Addgene #98158) and subjected to point mutations. The nuclear localization signal, linker, and M-MLV RT sequence were synthesized by GenScript and codon-optimized for Escherichia coli. For the PEmax protein expression vector (pET28a-PEmax-His) , mutations were introduced based on the pET28a-PE2-His vector, and the nuclear localization signal and linker were replaced with known techniques. The protein expression vectors for PE2ΔRH (pET28a-PE2 ΔRH-His) and PEmaxΔRH (pET28a-PEmax ΔRH-His) were generated by excluding the Rnase H domain from M-MLV ORFs. To express the MLH1dn protein, we generated an Escherichia coli codon-optimized MLH1dn sequence (GenScript) and cloned it into pET28a-His expression vector.
[0143] Plasmids used as templates for in vitro transcription were constructed using the pUC18 acceptor plasmid containing the T7 promoter and UTR sequences, with the target sequence obtained by PCR. Sequences for PEmax, PE6a, PE6c, PE6d, and MLH1dn were obtained from pCMV-PEmax (Addgene #174820) , pCMV-PE6a (Addgene #207851) , pCMV-PE6c (Addgene #207853) , pCMV-PE6d (Addgene #207854) , and pCMV-PEmax-P2A-hMLH1dn (Addgene #174828) , respectively. Vpx and MLH1-SB were synthesized by GenScript Corporation according to recent publications. Purification of PE proteins and T4 RNA Ligase2
[0144] Rosetta (DE3) competent cells (WEIDI, EC1010) and BL21 (DE3) competent cells (WEIDI, EC1010) were transformed with pET28a-PE2 or pET30C-gp24.1, respectively, following the manufacturer’s instructions. A 5 ml overnight culture grown from a single colony in LB with 50 μg / ml kanamycin was transferred into 0.5 liters of the same medium and cultured at 37 ℃. Once the OD600 of the culture reached 0.7–0.8, the expressions of PE and T4 RNA Ligase2 were induced with 1 mM isopropyl β-D-1-thiogalactopyranoside (IPTG) for 18 hours at 18 ℃, and for 3-6 hours at 37 ℃, respectively. Cells were collected and lysed in a lysis buffer using a high-pressure homogenizer (ATS, AH1500) . The supernatant was collected after centrifugation and then filtered with 0.45-micron filters. Affinity purification followed by a size exclusion chromatographic step was performed for protein purification. In brief, clarified lysate was loaded onto a HisTrap HP column (Cytiva) in the NGC Quest 10 Chromatography System (Biorad) . The column was pre-balanced in lysis buffer. Protein was eluted in buffer B1 using a gradient program. Different elution fractions were collected and then verified by SDS-PAGE to identify the target protein. The resulting protein was then loaded to HiLoad 16 / 60 Superdex 200pg column (Cytiva) in buffer B2.Eluted protein was concentrated by centrifugal filters (Millipore) and stored in buffer B2 at -80℃ with 10%glycerol. The other four proteins, PEΔRH, PEmax, PEmaxΔRH, MLH1dn, were also purified using the method described above.
[0145] The lysis buffer for purifying PE proteins and MLH1dn contains 20 mM Tris-HCl, 500 mM NaCl, pH 7.4, the buffer B1 consists of 20mM Tris-HCl, 500mM NaCl, 500mM Imidazole, pH7.4, and the buffer B2 consists of 20mM Tris-HCl, 200mM NaCl, pH7.4. The lysis buffer for purifying T4 RNA Ligase 2 contains 50mM Tris-HCl, 250mM NaCl, 10%(w / v) sucrose, pH 7.5, the buffer B1 consists of 50 mM Tris-HCl, 250mM NaCl, 10%(v / v) glycerol, 500mM imidazole, pH 8.0, and the buffer B2 consists of 10 mM Tris-HCl, 0.1mM EDTA, 0.1mM DTT, 10%glycerol, 90mM NaCl, pH 7.5. In vitro transcription of pegRNA
[0146] In vitro transcription (IVT) was performed using T7 RNA polymerase in a reaction mixture containing 5× T7 RNA Polymerase Reaction Buffer (0.2 M Tris-HCl pH 8.0, 0.125 M NaCl, 40 mM MgCl2, 10 mM Spermidine- (HCl) 3) , 50 mM DTT, pyrophosphatase, and 0.5-1 μg DNA in a total volume of 50 μl. The NTPs (GTP, ATP, CTP, UTP) were supplemented to a final concentration of 2 mM (Sigma, A2383 / U6625 / C1506 / G8877) 70. For transcripts necessitating a 5′-phosphate for subsequent ligation, the final concentration of GTP was adjusted to 2 mM, and GMP was added to achieve a final concentration of 16 mM. DNA template was then enzymatically digested by adding 4 U RQ1 RNase-Free DNase (Promega, M6101) at 37℃ for 15 minutes. The resulting products were purified using RNA Cleanup Kit (New England Biolabs, T2040) . The pegRNAs obtained by IVT were treated with CIP enzyme (New England Biolabs, M0525S) before being utilized for cell transfection. Nucleic acids synthesis
[0147] Chemically modified short 32nt RNA, nicking sgRNA, and unmodified splint DNA fragments were synthesized by GenScript and Sangon Biotech. Other chemically modified RNAs were synthesized by General Bio. RNA ligation
[0148] For RNA ligation, 200 pmol of acceptor RNA, 200 pmol of donor RNA, and equimolar amounts of splints DNA, together with 5× annealing buffer (50mM Tris-HCl pH 8.0, 100mM NaCl) , were added for a total volume of 30 μl. This mixture was incubated for 3 minutes at 70℃ and then gradually cooled to room temperature at a rate of 0.1℃ / s. Subsequently, ligation was performed by adding 4 μl of 10 × T4 RNA Ligase 2 buffer (500 mM Tris-HCl pH 7.5, 20 mM MgCl2, 10 mM DTT, 4 mM ATP) , 3 μl T4 RNA Ligase 2 (1 μg / μl) , and 3 μl nuclear-free water for incubating at 37℃ for 1 hour. The annealing and ligation process was repeated twice, with the addition of 1.5 μl T4 RNA ligase in the second and third ligation rounds, followed by incubation at 37℃ for 30 minutes. The DNA splint was digested with RNase-Free DNase (Promega) for 30 minutes at 37℃. The ligation products were purified using the RNA Cleanup Kit (New England Biolabs) . HPLC purification of ligated pegRNA
[0149] The HPLC purification of ligated pegRNA was conducted using the Agilent 1260 Infinity II system. The purification process employed a column matrix composed of alkylated non-porous polystyrene-divinylbenzene copolymer microspheres (Agilent PL1512-5802) . For HPLC separation, buffer containing 0.1M triethylammonium acetate (TEAA, pH 7.0) was used. The HPLC columns were initially equilibrated with a solvent consisting of 10%acetonitrile. Subsequently, the RNA sample was loaded onto the column and subjected to a linear gradient elution with acetonitrile, transitioning from 13%to 18%acetonitrile over a duration of 20–30 minutes, with a flow rate of 0.5ml / min. In vitro transcription of PE mRNA
[0150] The in vitro transcribed (IVT) of PE2, PEmax, nCas9 (H840A) and MCP-RT mRNA followed an established protocol using VSW3-RNA polymerase. In brief, the VSW3 promoter sequence was introduced into PE constructs to serve as the template for PCR, and the reverse primer introduced Poly A tail to the 3’ UTR of the resulting PCR product. The IVT reaction mixture (10 μL) comprised 40 mM Tris-HCl (pH 8.0) , 16 mM MgCl2, 5 mM DTT, 2 mM spermidine, 35 ng / μL of the purified PCR product as template, 1.5 U / μL of RNase inhibitor, 0.2 μM inorganic pyrophosphatase, 4 mM of each of the NTPs, and 0.15 μM of VSW-3 RNA polymerase, and the reaction was conducted at 25℃ for 12-16 hours. RNase-Free DNase (Promega) was employed to eliminate the DNA template, and the resulting mRNA was precipitated through lithium chloride. The mRNA was enzymatically capped using the Vaccinia Capping System (New England Biolab) and further purified using the Monarch RNA Cleanup Kit (New England Biolab) .
[0151] For the preparation of mRNA encoding PEmax, PE6c, PE6d, Vpx, MLH1dn, and MLH1-SB, linear DNA templates containing the T7 RNA polymerase promoter and a 110-base poly (A) tail were amplified by PCR. Each in vitro transcription reaction included 10 μL of 5× T7 RNA polymerase reaction buffer (0.2 M Tris-HCl, pH 8.0, 0.125 M NaCl, 40 mM MgCl2, 10 mM spermidine- (HCl) 3 (Sigma) , 4 μL of rNTP mix (Vazyme Biotech Co., Ltd, DD4108-PA, DD4106-PA, DD4107-PA, DD4114-PA) , 0.5-1 μg of DNA template, 4 μL of PPase (Sigma) , 1.25 μL of RNase inhibitor (Promega) , 6.25 μL of CAG Trimer (3’ -OMe) (Vazyme Biotech Co., Ltd, GMP4119PC) , 2 μL of T7 RNA polymerase, and nuclease-free water to a final volume of 50 μL. The reaction mixture was incubated at 37℃ for 1 hour, followed by DNase treatment to remove the DNA template. The resulting mRNA was purified using the RNA Cleanup Kit. Cell Transfection
[0152] All electroporation procedures were conducted using the Lonza 4D Nucleofector system in B1mix buffer following established protocols. Each procedure comprised 2e5 cells in 20 μL. RNPs were pre-incubation at room temperature for 15 minutes before electroporation. The programs of electroporation were used as following: EO-115 for HEK293T, EO-138 for Huh-7 and K562, DS-137 for U2OS, and DN-130 for Hela cells. The initial RNP electroporation applied 140 pmol PE protein, 186 pmol pegRNA, and 62 pmol of nicking sgRNA. The optimized RNP electroporation used 70 pmol PEmax protein, 140 pmol L-epegRNA, and 60 pmol nicking sgRNA. The mRNA electroporation included 0.5-2 μg of PEmax mRNA, 3-180 pmol epegRNA, and 1-60 pmol nicking sgRNA. For electroporation of plasmid, 800 ng PE plasmid, 200 ng pegRNA plasmid, and 83 ng nicking sgRNA plasmid were used for HEK293T, Huh-7 and K562, and the doses of plasmids were doubled for U2OS and Hela cells, followed previously established procedures. For mRNA transfection of Hepa1-6 cells, 100 ng of PE mRNA, 150 ng of chemically modified epegRNA, and 50 ng of nicking sgRNA were transfected using LipofectamineTM MessengerMAXTM (Thermo Fisher) with 7.5 × 104 cells per well in a 48-well plate. Isolation and electroporation of primary human T cells and CD34+ human HSPCs
[0153] Peripheral blood mononuclear cells (PBMCs) were purchased Biotechnologies. CD3+ T cells were isolated from PBMCs using the EasySep Human T Cell Isolation Kit (STEMCELL) and cultured in X-Vivo15 medium (Lonza) supplemented with 5%fetal bovine serum (FBS) (Gibco) , human IL-2 (50 ng / ml, Peprotech) , human IL-7 (10 ng / ml, Peprotech) , and 1%P / S. To activate T cells before electroporation, CD3 / CD28 Dynabeads (Thermo Fisher, 11131D) were added to the culture at a 1: 3 ratio and cultured for 3 days. After removing the beads, the cells were allowed to rest for 5-7 hours before electroporation, which was carried out using EO-138 program in B1mix buffer. The electroporation details are the same as described above for RNP and mRNA delivery. Three days after electroporation, cells were collected by centrifugation, and genomes were isolated.
[0154] Cryopreserved CD34+ HSPCs were purchased from Biotechnologies, and thawed according to the manufacturer’s instructions. CD34+ HSPCs were cultured in StemSpan SFEM (StemCell Technologies) supplemented with human stem cell factor, human thrombopoietin, and human Flt3 ligand (100 ng / ml for each, Peprotech) , with 1%P / S. CD34+ human HSPCs were electroporated with RNA or RNP 24 hours post-thaw using the Lonza 4D Nucleofector system with the CM-137 program in B1mix buffer. Each electroporation procedure consisted of 20 μl containing 2e5 CD34+ human HSPCs. The doses of RNA or RNP followed the details described above. Genome were collected three days after electroporation. Preparation of LNPs
[0155] LNPs were prepared using a microfluidic device as previously described. Lipid 2950 (Xiamen Sinopeg Biotech Co., Ltd) , DSPC (1, 2-distearoyl-sn-glycero-3-phosphocholine) (A.V. T. Pharmaceutical Co. Ltd) , cholesterol (A. V. T. Pharmaceutical Co. Ltd) , and DMG-PEG2000 (Xiamen Sinopeg Biotech Co., Ltd) were dissolved in ethanol at molar ratios of 50:10: 38.5: 1.5, respectively. RNA was dissolved in 6.25 mM acetate buffer, PH 5.0. The lipid and RNA solutions were mixed at a 1: 3 ratio using a microfluidic device at a flow rate of 12 mL / min. The resulting mixture was dialyzed against PBS to remove ethanol. LNPs were then concentrated, and the RNA encapsulation efficiency was determined using the Quant-iTTM RiboGreenTM RNA Assay (Thermo Fisher) . The prepared LNPs were stored at 4℃ until further use. LNP injection and subsequent tissue collection and analysis
[0156] During LNP preparation, mRNA, chemically modified pegRNA / epegRNA, and nicking sgRNA were packaged into separate LNPs. For LNPs containing MLH1dn, MLH1-SB, or Vpx mRNA, the mass ratio of PE mRNA to MLH1dn, MLH1-SB, or Vpx was 5: 1. For LNPs co-packaging PE, MLH1-SB, and Vpx, the mass ratio was 5: 1: 1. For pegRNA and nicking sgRNA, the mass ratio of pegRNA to nick-sgRNA was 3: 1. After quantifying the RNA in LNPs, mRNA-containing LNPs were mixed with pegRNA-containing LNPs at a mass ratio of 2: 1 and injected into 6-to 8-week-old female C57BL / 6J mice at a total dose of 3 mg / kg. Seven days post-injection, the mice were euthanized, and liver tissue and blood samples were collected. Genomic DNA from the liver was purified using a genomic DNA extraction kit (Magen) according to the manufacturer's instructions, followed by amplicon sequencing of the edited loci. The Pcsk9 concentrations in mouse blood samples were measured using the Mouse Pcsk9 ELISA Kit (Proteintech, KE10050) . All animal studies were approved by the Institutional Animal Care and Use Committee (IACUC) of Wuhan University. Genomic DNA extraction
[0157] The collected cells were centrifuged at 300g for 5 minutes, and the pellet was resuspended in 50 μl lysis buffer (10 mM Tris-HCl pH 8.0, 0.05%SDS) and 1 μl proteinase K (Thermo Fisher, EO0491) . The mixture was then incubated for 2 hours at 37℃, followed by an additional 30 minutes at 85℃ to inactivate the proteinase K. The editing loci were subsequently PCR amplified and prepared for amplicon sequencing (Illumina) . Deep sequencing analysis
[0158] PE target sites were amplified and sequenced using the Illumina sequencing platform (Novaseq 6000) . Genomic DNA samples were amplified with primers specific for Illumina adapters using Phanta Max Super-Fidelity DNA Polymerase (Vazyme) . The first round of PCR reactions was carried out with the following parameters: 95℃ for 3 min, followed by 23 PCR cycles, and a final 72 ℃ extension for 5min. In the second round of PCR reactions, 13 cycles were performed. The primers contained Illumina adapters and a 7-bp / 8-bp index. The PCR products were purified before sequencing in the Illumina sequencing platform.
[0159] Raw paired-end reads were merged using the fastp software to generate full-length reads. Reads with a mean quality score <30 or adaptor contamination were discarded using Trimmomatic. Alignment of reads to a reference sequence was performed using CRISPResso2. Based on the deep sequencing results, target editing was classified into two categories: intended edits and indels, which include undesired mutations. Statistical analysis
[0160] GraphPad Prism 9 software was applied to analyze the data. Two-tailed Student’s t-tests was used for comparisons of differences between two groups, and one-way analysis of variance (ANOVA) with Tukey’s multiple comparisons test was performed for multiple groups comparisons. Results Optimization of RNA ligation process
[0161] We first established a method to obtain full length sgRNA and pegRNA with defined chemical modifications using RNA ligation, and optimized the ligation process. We tested a panel of RNA ligases, and selected bacteriophage T4 RNA ligase 2 (T4 Rnl2) for RNA ligation.
[0162] We first used T4 Rnl2 to obtain chemically defined sgRNA, as the structure of sgRNA is simpler than pegRNA. We divided sgRNA into two parts, a 20 nt spacer region serving as acceptor RNA and an 82 nt scaffold region as donor RNA (FIG. 1a) . We chemically synthesized a 20 nt acceptor RNA with a hydroxyl group at the 3’ end, and used IVT to cost-effectively obtain the 82 nt scaffold RNA. To ensure a single phosphate modification at the 5’ end of donor RNA, GMP was introduced during IVT. The presence of a splint DNA facilitated RNA ligation (FIG. 1b) . We assessed the precision of ligation by RT-PCR and Sanger sequencing, and the results indicated this enzymatic ligation of two RNAs obtained a full length sgRNA.
[0163] To improve the efficiencies of ligation, we optimized ligation temperature, the length of splint DNA, the ratio of splint DNA to RNA, and dose of ligase. Our results indicate ligation at 37 degrees is better than previously described 25 degrees for ligation of small RNA (FIG. 1c) (Viollet, S et al., BMC Biotechnol 11, 72 (2011) ) . A longer length of splint DNA improved the ligation efficiencies (FIG. 1d) . Interestingly, optimal molar ratio of RNA to splint DNA is 1: 1, and a higher ratio of splint DNA adversely affected ligation (FIG. 1e) . Moderate doses of T4 Rnl2 were sufficient for ligation, as further increasing enzyme quantity did not improve ligation efficiency (FIG. 1f) . We also synthesized 3’ end modified and unmodified acceptor RNAs separately for sgRNA ligation, and both showed effective ligation, indicating modification away from ligation site did not affect the enzymatic reaction (FIG. 1g) . Finally, we used ligated sgRNA for in vitro cleavage of substrate DNA, and both 3’ end modified and unmodified ligated sgRNA showed effective cleavage of substrate DNA (FIG. 1h) . In summary, with optimized conditions of RNA ligation, a chemically defined, functional sgRNA can be efficiently obtained. Ligation of pegRNA
[0164] The pegRNA has additional RTT, PBS sequence, and optional protection sequence at the 3’ end of sgRNA. Based on the optimized ligation process described above, we synthesized two relatively short RNAs. The short RNA is within the range of high-quality solid phase synthesis, and the cost of synthesis is much more affordable than full pegRNA synthesis. We synthesized a 32-nt acceptor RNA with three nucleosides harboring 2′-O-methyl modification and with phosphorothioate linkages at the 5’ end, along with a donor RNA containing a 5’ phosphate generated through IVT (FIG. 2a) .
[0165] The ligation results demonstrated successful assembly of full-length pegRNA by 3 cycles of annealing and ligation process for a total of 2 hours at 37 degrees (FIG. 2b, methods section) . Subsequently, we used the same method to ligate different pegRNAs to introduce editing events at different loci in HEK293T cells. Our results revealed that ligated pegRNA with chemical modifications at the 5’ end exhibited 1.7 to 6.1 folds higher editing efficiency than pegRNA produced by IVT for PE2, and 1.4 to 3.9 folds higher for PE3 via RNP delivery (FIG. 2c-f) . Our initial donor RNA was produced by IVT to reduce the cost of synthesis to the greatest extent (FIG. 2b-f) . However, due to the indispensable nature of GTP in IVT, a significant portion of donor RNAs lacked the 5’ single phosphate, rendering them incapable of being ligated by T4 Rnl2. As a result, un-ligated RNA persisted in the ligation products (FIG. 2b) . Analysis by high-performance liquid chromatography (HPLC) indicated that approximately 30%of ligation products were un-ligated RNA (FIG. 3a-b) .
[0166] Subsequently, we subjected the ligation products to HPLC purification, resulting in substantial improvement in the purity of the ligated products (FIG. 3c) . HPLC-purified ligated pegRNAs revealed 2.5 to 10.5 folds and 2.3 to 5.4 folds higher editing efficiency than pegRNA produced by IVT for PE2 and PE3, respectively, via RNP delivery (FIG. 2c-f) . Notably, two out of four HPLC-purified pegRNAs did not exhibit significantly higher editing efficiencies than unpurified ligated pegRNA (FIG. 2c-f) . It is contemplated that these pegRNA may have reached the upper-limit of editing efficiency due to the lack of 3’ end modifications. Enhanced editing efficiency by stabilizing the 3’ end of ligated pegRNA
[0167] The 3’ extension of pegRNA, when left unshielded, is exposed in the cellular environment, making it more susceptible to degradation by cellular nucleases, and 3’ end truncated pegRNA would lose its PBS and RTT but retain the ability to bind Cas9, severely impeding the ability of PE. The evopreQ1 motif has been utilized to protect the 3’ extension of pegRNA to generate “epegRNA” , leading to increased stability and enhanced prime editing efficiency.
[0168] We contemplated that stabilizing the 3’ end of pegRNA via ligation could improve efficiency of PE. Therefore, we designed three different ligation strategies to allow 3’ end been protected (FIG. 4a-c) . The first one ligated synthetic 5’ -modified acceptor RNA with evopreQ1-containing donor RNA generated by IVT (referred as *32+X evQ1, X indicate the length of each fragment, and *indicates chemical modifications) (FIG. 4a) . The second strategy still used 32 nt synthetic 5’ -modified acceptor RNA, and used a chemically synthesized donor RNA. Due to the length limitation of chemically synthesized RNA, the donor RNA is without evopreQ1 (referred as *32+X*) (FIG. 4b) . To add evopreQ1 onto the ligated pegRNA and to avoid the penalty of chemically synthesizing long RNA sequences, we split acceptor and donor RNA to comparable lengths (refer as *91+X evQ1*) (FIG. 4c) . The donor RNA is 91 nt, containing most part of the sgRNA sequence, and the acceptor RNA is about 80 to 95 nt (FIG. 4c-g) . For the second and third designs, three nucleotides at 5’ end of the donor RNA and 3’ end of the acceptor RNA are modified (FIG. 4b-c) . We found that when the donor RNA was chemically synthesized, the efficiency of ligation reaction was higher than the reaction in which donor RNA was generated by IVT, resulting in 91%ligation efficiency for epegRNA, in part due to the predominant 5’ phosphate in chemically synthesized donor RNA (FIG. 4g) .
[0169] We compared the editing efficiencies of three designs. In all pegRNAs examined, the *91+X evQ1*showed the highest editing efficiencies for both PE2 and PE3 via RNP delivery, reached 43.4%to 69.2%for PE2 and 42.2%to 72.3%for PE3 (FIG. 4d-f) . We named the third design (*91+X evQ1*) as L-epegRNA, which is around 170-190 nt in length for PE-mediated point mutations, small deletions, and small insertions. It is worth noting that this length range of L-epegRNA is extremely difficult to be generated by chemical synthesis directly. Due to length limitation, solid-phase synthesis of pegRNA usually generates RNA sequences without evopreQ1, as the same sequence and modifications generated by the second design (FIG. 4b) . Moreover, due to the low quality of RNA product, solid-phase synthesized epegRNA did not substantially improve editing efficiencies in comparison to chemically synthesized pegRNA. The editing efficiencies of L-pegRNA are comparable or slightly higher than reported numbers using solid-phase synthesized pegRNA for the same sequences.
[0170] Our optimal ligation strategy of L-epegRNAs revealed 2.4 to 4.9 folds and 2.1 to 3.8 folds higher editing efficiencies than L-pegRNA for PE2 and PE3, respectively, via RNP delivery (FIG. 4d-f) . We then compared L-epegRNA with epegRNA generated by IVT. For five pegRNA sequences examined, L-epegRNA demonstrated average editing efficiencies of 44.7%for PE2 and 54.2%for PE3 in HEK293T cells via RNP delivery, 1.9 to 11.4 folds higher than epegRNA generated by IVT (FIG. 4h) . Optimization of ligated pegRNA-mediated RNP delivery
[0171] Due to the limitation of pegRNA synthesis, previous studies on RNP delivery of PE have been constrained from systematic optimization. Here we optimized various conditions that may affect RNP delivery of PE using ligated pegRNA. To ensure non-saturation editing, experiments were conducted using *32+X evQ1 ligated RNA. Initially, we optimized the dosages of various components in the RNP delivery system. For PE2 editing, we proportionally adjusted the dose of RNP, and the results demonstrated that 70 to 140 pmol RNP was sufficient to maintain editing efficiency (FIG. 5a-b) . Subsequently, we explored the impact of the ratio of PE protein to pegRNA on editing efficiency in PE2. We observed a slight improvement in editing efficiency with an increased ratio of pegRNA (FIG. 5c-d) . Finally, we examined the dosage of nick sgRNA in PE3, and identified 30-60 pmol nick sgRNA is optimal for editing (FIG. 5e-f) .
[0172] It was suggested that RNaseH domain could be dispensable for prime editing, and PEmax version often exhibited superior performance compared to classic PE. Therefore, we expressed and purified four versions of PE, as PE, PEΔRH, PEmax, and PEmaxΔRH (FIG. 6a-b) . Interestingly, we found that the removal of the RNaseH domain significantly increased yield of protein in Ecoli expression system (FIG. 6c) . Subsequently, we examined the editing efficiency of these PE proteins. Removal of RNaseH domain showed a trend of improved editing efficiency, and PEmaxΔRH had a slightly improved efficiency over PEΔRH for PE2 editing (FIG. 6d-e) . Therefore, we used PEmaxΔRH for most of studies below, except the case that long RTT sequence forming secondary structure.
[0173] Suppression of the Mismatch Repair (MMR) pathway via overexpression dominant negative MMR protein (MLH1dn) has been shown to improve efficiencies of PE in several cell lines including K562 cells. Therefore, we expressed and purified the MLH1dn protein to examine its effect in RNP delivery of PE. Surprisingly, the addition of purified MLH1dn protein at varying dosages (up to 140 pmol) for PE2 and PE3 did not result in increased editing efficiency (FIG. 7a-b) . As controls, when delivered as plasmids, PE4max (PE2max with overexpression of MLH1dn) showed superior editing to PE2max for both pegRNAs examined, and PE5max exhibited improved efficiencies than PE3max in one of two pegRNAs examined (FIG. 7c-d) . L-epegRNA-mediated RNP delivery is generally superior to plasmids
[0174] We examined the editing efficiencies using L-epegRNA under the optimized RNP delivery described above. Using the PEmaxΔRH protein, we examined three types of changes (point mutation, small insertion, and small deletion) at five endogenous sites. L-epegRNA can efficiently edit all loci with PE2 and PE3, exceeding 70%efficiencies for two loci (FIG. 8a-f) .
[0175] Furthermore, we compared RNP delivery using five different L-epegRNAs and corresponding plasmids in five different cell lines including HEK293T, K562, Huh-7, Hela and U2OS cells. L-epegRNA-mediated RNP delivery exhibited significantly higher editing efficiencies for most comparisons (22 out of 25 for PE2, 20 out of 25 for PE3) (FIG. 8g-k) . Strikingly, while PE via plasmid delivery exhibited extremely low or undetectable editing at several loci, L-epegRNA-mediated RNP delivery achieved excellent editing efficiencies. For example, at the RNUX1 locus (+5G to T) in Hela cells, plasmid delivery of PE2 and PE3 yielded average efficiencies of 0.14%and 0.53%, respectively. In contrast, L-epegRNA-mediated RNP delivery resulted in average efficiencies of 11.6%and 47.9%for PE2 and PE3, respectively (FIG. 8g) . A similar phenomenon was observed at the DMNT1 site (+5G to T) in U2OS cells, whereas plasmid delivery of PE3 yielded editing efficiencies 0.17%. In comparison, L-epegRNA-mediated RNP delivery resulted in 22.9%editing (FIG. 8k) . These data suggest L-pegRNA-mediated prime editing can be tested in loci that plasmid delivery of PE failed to generate efficient editing. L-epegRNA facilitates efficient editing via RNA delivery of PE
[0176] Besides RNP delivery, RNA delivery of PE can also be used for therapeutics and research. We explored the use of L-epegRNA in RNA delivery of PE platform in three cell lines (FIG. 9a) . In comparison to full-length epegRNA generated by IVT, L-epegRNA exhibited 56.6 to 845.2 folds, 9.8 to 19.9 folds, and 100.4 to 112.5 folds improvement for PE2 in HEK293T, K562 and Huh-7 cells, respectively (FIG. 9b-d) . Similarly, L-epegRNA exhibited 11.5 to 563.0 folds, 13.1 to 21.8 folds, and 79.0 to 521.4 folds improvement in comparison to IVT epegRNA for PE3 in these cells, respectively (FIG. 9b-d) . L-epegRNA achieves efficient prime editing in primary cells
[0177] Primary cells often exhibit poor tolerance to plasmids, and applications of PE in primary cells usually require the use of synthetic pegRNA for RNP or RNA delivery of PE. We employed L-epegRNA for both RNP and RNA delivery PE in human primary T cells and human CD34+ hematopoietic stem cells (HSPC) . L-epegRNA generated efficient editing in both cell types via RNP and RNA delivery, and the editing efficiency generated by RNA delivery seemed to be higher than that of RNP delivery, 4.7%to 34.6%for RNP and 14.7%to 71.4%for RNA delivery, respectively (FIG. 10a-d) . L-epegRNA achieves relatively large insertions by RNP and RNA delivery
[0178] The upper limit of chemical synthesis of RNA is approximately 200 nt, and the feasibility of synthesis depends on the complexity of the sequence. While epegRNA reached a size of 170 to 190 nt for point mutations or a few bp insertion / deletions, the size of epegRNA for longer insertion is over 200 nt. We aimed to ligate an epegRNA for insertion of 17 bp at HEK3 locus, and the length of epegRNA reached 210 nt. We designed two ligation strategies, namely, two-segment ligation (105 nt + 105 nt) and three-segment ligation (76 nt +54 nt + 80 nt) , respectively (FIG. 11a) . Our data indicates the ligation efficiencies of three-segment is higher than two-segment, likely attribute to the reduced accuracy of synthesizing the 5’ monophosphate of longer than 100 nt donor RNA or the complexity of RNA structure in the two-segment connection (FIG. 11b) . L-epegRNA generated by three-segment ligation without HPLC purification was tested, and the efficiencies of successful 17 bp insertion were 42.1%and 49.4%via RNP or RNA delivery, respectively (FIG. 11c-d) . In contrast, IVT-generated epegRNA induced 0.1 to 8.7%of 17-bp insertion (FIG. 11c-d) . We then applied HPLC to purify L-epegRNA. HPLC purification further boost the rate of 17-bp insertion by L-epegRNA to 57.8%and 59.7%for RNP and RNA delivery of PE, respectively (FIG. 11c-d) .
[0179] While directly chemical synthesis of pegRNA is difficult and expensive, generation of epegRNA by such way is extremely challenging. The feasibility depends on sequence complexity, and the cost is usually not affordable (FIG. 11e) . The performance of these chemical modified pegRNA and epegRNA is not satisfactory, due to impurity of direct chemical synthesis of long RNA despite HPLC purification (Nelson, J. W. et al. Nature Biotechnology 40, 402-410 (2021) ; Liu, B. et al. Nature Biotechnology (2023) doi. org / 10.1038 / s41587-023-01947-w. ) . In contrast, L-epegRNA in our study showed up to 11.4-fold and 845.2-fold higher efficiencies than IVT-generated epegRNA for RNP and RNA delivery of PE, respectively (FIG. 4h; FIG. 9b-d) . Moreover, the synthesis of L-pegRNA is straightforward, the purity of L-pegRNA is ensured, and the total cost is affordable (FIG. 11e) .
[0180] Primer editing by DNA polymerase requires DNA as a template for extension. DNA templates also showed higher editing efficiency in previous reports. Therefore, we chemically synthesized RNA / DNA chimeric strands to partially replace the RNA in the original RTT-PBS region with DNA. And we ligated the chimeric strand with RNA sequence to obtain L-epegDNA and L-pegDNA with MS2 loop. Both the L-epegDNA and L-pegDNA (MS2) can successfully induce editing of the target site. (FIG. 12a-c)
[0181] In this example, we optimized the conditions of RNA splint ligation and extended its usage for generating chemically modified sgRNA and pegRNA (FIG. 1&FIG. 2b, 4g) . The ligated pegRNA with the best design (L-epegRNA) demonstrated superior editing activity via RNP and RNA delivery of PE is efficient in both cell lines and human primary cells, promoting the therapeutical applications using PE system (FIG. 8-10) . For the same loci, our optimized RNP delivery efficiencies are 2.8-3.6 folds higher in HEK293T cells and 2.4-5.8 folds higher in human prime T cells than previously reported (FIG. 8b &10a) . Interestingly, L-epegRNA exhibited much lower indel formation, in comparison with published data (FIG. 10b) (Chen, P. J. et al. Cell 184, 5635-5652. e5629 (2021) ) . The instability of pegRNA may cause degradation of RTT and PBS region, making pegRNA to function as sgRNA and generate nick or double-stranded break. The reduced indel formation by L-epegRNA could attribute to its stability by both chemical modifications and evopreQ1 motif to protect PBS-RTT region.
[0182] Prior to our report, plasmid transfection reached much higher efficiencies than RNP delivery of PE. With the activity of L-pegRNA and optimization of various parameters for RNP delivery of PE, we demonstrated for the first time that efficiencies RNP delivery of PE can be superior to plasmid delivery for majority of comparisons examined (FIG. 8g-k) . L-epegRNA can be applied for ex vivo therapy and in vivo delivery in RNA format via encapsulation into non-viral vectors such as lipid nanoparticles.
[0183] It is extremely challenging to chemically synthesize ~200 nt RNA or beyond, as it is beyond the limit of solid-phase synthesis. Employing a three-segment ligation approach, we successfully generated a 210-nucleotide epegRNA with chemical modification, achieving high-efficiency of 17 bp insertion (approximately 60%editing) via both RNP and RNA delivery (FIG. 11c-d) . These results underscore the potential of L-epegRNA in facilitating the insertion of longer sequence using PE. L-epegRNA can be applied to insert a landing pad for integrases and recombinases, facilitating targeted insertion of large fragment for research and therapeutic applications. Furthermore, using three-segment ligation approach, it is possible to include additional RNA structures into chemically modified epegRNA. For example, MS2 can be incorporated into L-pegRNA to recruit effector proteins to the PE targeting site. This optimized splint ligation protocol can be applied to synthesis high quality long RNA / nucleic acids with chemical modifications. For example, it can be used for generating stable sgRNA with MS2 / PP7 motif for recruitment of effectors, and chemically modified oligonucleotides to bind endogenous adenosine deaminases acting on RNA (ADAR) enzymes.
[0184] In conclusion, this example demonstrates optimized RNA ligation can be applied to generate chemically modified epegRNA for efficient prime editing from point mutations to 17 bp insertions in human cells. L-epegRNA and the optimized RNP delivery of PE facilitate efficient prime editing that surpass commonly-used PE plasmids. Example 2. Modifications to epegRNA
[0185] This example designed and tested various modifications to epegRNA. Heavy modifications of evopreQ1 improve prime editing efficiency
[0186] An epegRNA typically includes five functional domains: a spacer for targeting, a scaffold for nCas9 binding, RTT and PBS for reverse transcription, and an evopreQ1 motif for protecting the 3'end from degradation (FIG. 13a) . The scaffold and evopreQ1 sequences are fixed, while the spacer, RTT, and PBS sequences vary depending on the target sites and editing types. Given the functional differences among these regions, each may require specific modifications, making it necessary to assess modification tolerance individually.
[0187] Without proper protection, the 3'end of pegRNA is exposed to the cellular environment and susceptible to nuclease degradation, which severely reduces editing efficiencies of prime editing. Although chemical modifications with 2'-O-methyl and phosphorothioate at three nucleotides on the 3’ end, combined with the evopreQ1 motif, can significantly enhance RNA stability and improve prime editing efficiency in cell culture, it remains unclear whether these chemical modifications are sufficient for both in vitro and in vivo applications. We envisioned that adding more chemical modifications to the 3'end of epegRNA could further boost prime editing efficiency. In a first approach, we increased the number of 2'-O-methyl modifications to enhance protection of the 3'end. In our first design (evopreQ1-M1) , we added 26 additional 2'-O-methyl modifications to the original evopreQ1-EM design (EM referred to the end modification of the last three nucleotides on pegRNAs) to protect both free 3’ end and the stem-loop of evopreQ1 (FIG. 13b) . In the second design (evopreQ1-M2) , we added 12 additional 2'-O-methyl modifications to the original design to protect the free 3’ but not the stem-loop of evopreQ1 (FIG. 13b) . These chemically modified epegRNAs were prepared using the ligation method we recently developed.
[0188] To evaluate the effects of these modifications, we tested these modified epegRNAs with three versions of PE mRNA (PEmax, PE6c, and PE6d) for a 3-base pair deletion at the EMX1 locus and a +1 T-to-Amutation at the HEK3 locus. We first examined the performance of these epegRNAs at low doses to expand the potential therapeutic window. The epegRNAs with either evopreQ1-M1 or evopreQ1-M2 modifications exhibited a 2-to 2.8-fold improvement in editing efficiency compared to classically end-modified epegRNAs, indicating that the additional 3'end modifications provided enhanced protection (FIG. 13c-e) . Notably, epegRNAs with evopreQ1-M1 modifications did not show further improvement in efficiency over those with evopreQ1-M2, suggesting that additional modifications to the evopreQ1 stem-loop may not be as important (FIG. 13c-e) .
[0189] Next, we examined the editing efficiencies under a high dose of epegRNA (90 pmol) , and the results mirrored those observed under low-dose conditions. The evopreQ1-M2 achieved editing efficiencies of 87.9%at the EMX1 locus, compared with 53%by evopreQ1-EM at the same dose (FIG. 13f) . Similarly, evopreQ1-M2 reached 56.7%at the HEK3 locus, compared to 38.6%with evopreQ1-EM (FIG. 13g) . Notably, the HEK3 locus contains a single nucleotide polymorphism (SNP) in one allele of triploid HEK293T cells, rendering approximately 30%of the sequence challenging to edit. Considering that evopreQ1-M2 achieved optimal editing efficiency with fewer modifications, we selected this design as the preferred option for subsequent studies. Modification of the scaffold region enhances efficiency
[0190] The scaffold region is a critical component of sgRNA, and its modification substantially impacts editing efficiency. Complete modification of the scaffold region has been shown to disrupt its function and drastically reduce editing efficiency. Furthermore, the scaffold region seems not tolerate 2'-fluoro modification well. Therefore, we introduced partial 2'-O-methyl modifications to the scaffold, referencing two strategies that have shown improved efficiency for sgRNA. In one design, we adopted a "highly modified" approach, which has demonstrated high efficiency for both Cas9 nuclease and base editing in vivo. Following this strategy, we modified three stem-loops of the scaffold with 2'-O-methyl modifications of 40 nucleotides, in addition to the end modification of epegRNA. We named this design as Scaffold-M1. In the second design, we used a structure-guided approach to modify the two stem-loops not protected by nCas9, resulting in 20 modified nucleotides. This modification in combination with the end modification was designated Scaffold-M2 (FIG. 14a-b) .
[0191] Compared to the unmodified scaffold (Scaffold-UM) , epegRNAs with Scaffold-M1 modifications exhibited lower editing efficiency across all three PE versions and both loci (FIG. 14c-d) . In contrast, Scaffold-M2 modifications increased editing efficiency by 1.3-to 1.7-fold compared to epegRNAs with unmodified scaffold, indicating that appropriate scaffold modification can enhance prime editing efficiency (FIG. 14c-d) . We further examined these epegRNAs at a higher dose, and found a similar trend as observed under low-dose conditions. Scaffold-M2 epegRNA achieved editing efficiencies of 76.6%at the EMX1 locus and 50.2%at the HEK3 locus (FIG. 14f-g) . Given the improved performance of Scaffold-M2 modifications, we selected this design for further analysis. Modification of the spacer region
[0192] The sequence of the spacer region varies depending on the target site, and excessive modification of the spacer region in crRNA reduces Cas9 editing efficiency. In the PE system, extension of the spacer's 5'end can impair the editing efficiency of epegRNAs, suggesting that the precision of the 5'end is crucial for optimal efficiency. Therefore, we hypothesized that excessive modifications at 5'end might negatively impact efficiency. The commonly adopted modification strategy involves adding three 2'-O-methyl modifications and three phosphorothioate linkages at the 5'end of the spacer (referred to as 3M3S) . However, a previous report suggested that reducing 5'end modifications could improve Cas9 editing efficiency. Therefore, we reduced the number of modifications to create two additional designs: 2M2S, containing two 2'-O-methyl modified nucleotides and two phosphorothioate linkages, and 1M1S, containing one 2'-O-methyl modified nucleotide and one phosphorothioate linkage (FIG. 15a) . We introduced these epegRNAs into cells, and the results indicated that reducing modifications in the spacer region did not improve editing efficiency (FIG. 15b-c) . Consequently, we decided to continue using the classic 3M3S modification for 5'end modification in RNA delivery. However, in RNP delivery, reducing 5'end modifications in the spacer region could improve editing efficiency (FIG. 15d) . Modification of RTT region enhances editing efficiency for PEmax and PE6c
[0193] The sequence of the RTT region varies depending on the target site and editing type (FIG. 16a) . Previous attempts to modify the RTT region of pegRNA have generally been ineffective in increasing editing efficiencies, typically using 2'-O-methyl, phosphorothioate linkages, their combinations, or DNA substitutions. Here we excluded phosphorothioate linkages for RTT modification due to their potential for creating stereoisomeric complexity, and focused on testing 2'-O-methyl and 2'-fluoro modifications. We generated several modification patterns for RTT, including a comprehensive 2'-O-methyl modification to the RTT region, referred to as RTT-M1; half of the RTT region near the scaffold modified with 2'-O-methyl, referred to as RTT-M2; a comprehensive 2'-fluoro modification, referred to as RTT-M3; half of the RTT region modified with 2'-fluoro modification, referred to as RTT-M4; and an alternating 2'-fluoro modification, referred to as RTT-M5 (FIG. 16b) .
[0194] Compared with RTT unmodified epegRNAs (RTT-UM) , RTT-M1 reduced efficiency, in particular for PEmax, while half 2'-O-methyl modification decreased efficiency to a less extent (FIG. 16c-d) . In contrast, full 2'-fluoro modification significantly improved editing efficiencies for PEmax and PE6c (FIG. 16c-d) . For PE6d, the RTT modifications did not lead to increased efficiency, suggesting that mutations in the RNase H domain and RT enzyme might affect tolerance to these modifications (FIG. 16c-d) . Reducing the number of 2'-fluoro modified nucleotides (RTT-M4 and RTT-M5) did not improve efficiencies (FIG. 16c-d) . We then examined the effect of RTT modification with a high dose of epegRNAs, and observed that 2'-fluoro modified significantly improved editing efficiency for one of the two loci (FIG. 16e-f) . These data collectively indicate that 2'-fluoro modification of RTT can be tolerated by PE and improved efficiency for PEmax and PE6c. Therefore, RTT-M3 was selected for further analysis. Modification of PBS enhances editing efficiency
[0195] The sequence of the PBS region also varies depending on the target site (FIG. 17a) . Since PBS is located at the 3'end of pegRNA, for the study investigating the effect of PBS modification in pegRNA, it could be difficult to determine whether the editing efficiency improvements by PBS modification were due to improved stability of PBS or increased protection of the 3'end. Here we investigated the PBS in epegRNA to reduce the probability of 3’ end protection (FIG. 17a) . We designed four modification patterns for PBS: full 2'-O-methyl modification of PBS (PBS-M1) , full 2'-fluoro modification (PBS-M2) , 2'-fluoro modifications on half of the nucleotides close to evopreQ1 (PBS-M3) , and 2'-fluoro modifications applied to alternating nucleotides (PBS-M4) (FIG. 17b) .
[0196] Full 2'-O-methyl modification of PBS decreased editing efficiency at the EMX1 locus, across all three PE versions (FIG. 17c) . Interestingly, this modification pattern maintained efficiency for the HEK3 locus with PEmax and PE6c mRNA, and provide a slight but significant increase for PE6d (FIG. 17d) . In contrast, 2'-fluoro modifications in PBS generally improved editing efficiencies, with up to approximately a 2-fold increase (FIG. 17c-d) . At a high dose of epegRNA, different 2'-fluoro modifications in PBS also significantly improved editing efficiencies, and PBS-M3 achieved an efficiency of 82.5%at the EMX1 locus and 60.1%at the HEK3 locus (FIG. 17e-f) . Considering that PBS-M3 more consistently improved editing efficiencies than other PBS modification patterns, we selected it for further studies. Combined modifications of different epegRNA regions further enhance editing efficiency
[0197] While individual modifications of different regions improved prime editing efficiency, we next examined whether combining these modifications could further improve editing efficiency. The combinations included modifications of the fixed regions Scaffold-M2 and evopreQ1-M2, along with the variable regions RTT-M3 and PBS-M3 (FIG. 18a) .
[0198] For PEmax, combinations of fixed region modifications with PBS but not RTT modifications, improved editing efficiency compared with single fixed region modification (FIG. 18b) . The combination of Scaffold-M2 and evopreQ1-M2 resulted in the highest editing efficiency, averaging 49.8%, compared with 15.7%by end-modified epegRNA (FIG. 18b) . For PE6c, combinations of fix region modifications with PBS did not further improve efficiency compared with single fixed region modifications, while reduced efficiencies were observed with combination involving RTT (FIG. 18c) . For PE6d, combinations involving RTT and PBS regions all decreased efficiency (FIG. 18d) . For both PE6c and PE6d, the combination of Scaffold-M2 and evopreQ1-M2 showed the highest efficiency, reaching 52%and 49.3%, respectively with a low dose of epegRNA, compared with 19.6%and 14.9%by end-modified epegRNA (FIG. 18c-d) .
[0199] We analyzed the percentage of indels within the total editing products, and found that 2’-fluoro modifications in the RTT, but not PBS, significantly increased this ratio, indicating an increased proportion of indels after RTT modification for all three PE versions (FIG. 19a-c) . The ratio for the Scaffold-M2 and evopreQ1-M2 combination was comparable to that of the-end modified control, indicating that this combination does not lead to an increased proportion of indel formation (FIG. 19a-c) .
[0200] Notably, the editing efficiencies induced by the Scaffold-M2 and evopreQ1-M2 combination-modified epegRNA at a low dose of 10 pmol (approximately 50%) was comparable to those of the end-modified epegRNA at a 90 pmol dose (52.7%, FIG. 13f) , indicating that heavily modified epegRNA can significantly reduce RNA usage while achieving comparable efficiency. We then examined the combined modifications at a medium dose. The pattern of increase is consistent with the low dose of pegRNA, and the combination of Scaffold-M2 and evopreQ1-M2 showed the highest efficiency, reaching 71.7%with a 30 pmol dose (FIG. 19d) . Overall, epegRNA with the Scaffold-M2 and evopreQ1-M2 combination demonstrated consistent high efficiency across multiple PE versions. Efficient In Vivo prime editing via non-viral delivery
[0201] To evaluate the editing efficiency of modified epegRNA in vivo, we selected the mouse Proprotein Convertase Subtilisin / Kexin Type 9 (Pcsk9) gene as the target. Pcsk9 is highly expressed in the liver, and previous studies using LNP delivery and Cas9 nuclease or base editors have demonstrated efficient editing of this gene. However, LNP delivery of PE showed much lower than expected efficiency in previous studies. We first screened epegRNAs targeting the Pcsk9 locus and selected an epegRNA that introduced a TTAC four-base insertion with high editing efficiency in Hepa1-6 cells, a murine hepatoma cell line (FIG. 20a) . This insertion creates a premature stop codon, leading to a disrupted reading frame. We then screened a few nicking sgRNAs for this epegRNA, and selected the most efficient nicking sgRNA for in vivo experiments (FIG. 20b) .
[0202] We tested several PE versions in Hepa1-6 cells for editing efficiency at the Pcsk9 locus. The editing efficiencies by mRNA encoding PEmax were significantly higher than those of PE6c and PE6d in these cells (FIG. 20c) . We encapsulated the mRNA encoding a prime editor effector and pegRNA with the nicking sgRNA into lipid nanoparticles (LNP) , and injected the LNP intravenously into mice, collecting liver tissues 7 days later (FIG. 21a) . The pegRNA without evopreQ1 and the nicking sgRNA were end-modified as previously described. The LNP-mediated delivery of pegRNA / nicking sgRNA with PEmax mRNA showed 5.5%prime editing efficiency in bulk liver, comparable to previously described results. Surprisingly, PE6c and PE6d mRNA showed significantly higher editing efficiencies in mouse liver, opposite to the results in Hepa1-6 cells (FIG. 20c, FIG. 21b) . Specifically, PE6c mRNA and PE6d mRNA showed 2.0-fold and 1.7-fold higher efficiencies than PEmax in the liver. Therefore, we applied PE6c for the next studies.
[0203] We compared different types of pegRNA modifications in cells and in vivo, including classical end-modified pegRNA (EM-pegRNA) , end-modified epegRNA (EM-epegRNA) , and highly modified epegRNA (HM-epegRNA, as the evopreQ1 is heavily modified) . In Hepa1-6 cells, HM-epegRNA achieved the highest editing efficiency (FIG. 20d) . Consistent with the in vitro results, EM-epegRNA outperformed EM-pegRNA, while HM-epegRNA demonstrated the highest efficiency, with an average prime editing efficiency of 51.6%in bulk livers (FIG. 21c) . These findings indicate that heavily modified epegRNA significantly boosts in vivo prime editing efficiencies via non-viral delivery. In vivo delivery of MLH1-SB further improves prime editing
[0204] PE efficiency is limited by mismatch repair (MMR) . Inhibiting key components of the MMR complex through expression of dominant-negative MLH1 (MLH1dn) can enhance PE efficiency in cells. Recently, an MLH1 small binder (MLH1-SB) designed using generative AI was developed to disrupt the dimerization interface between MLH1 and PMS2, leading to higher editing efficiency. We examined the effect of both MLH1dn and MLH1-SB in cells and in vivo. In Hepa1-6 cells, co-delivery of MLH1dn or MLH1-SB encoding mRNA did not improve prime editing efficiency (FIG. 22a) . In contrast, when MLH1dn or MLH1-SB mRNA was co-delivered with PE6c mRNA in vivo, MLH1-SB, but not MLH1dn significantly improved editing efficiencies, reaching an averaging 64.0%in bulk livers (FIG. 21d) .
[0205] Adult mouse liver cells are typically quiescent, dividing at a low frequency, and quiescent cells have low nucleotide levels. It has been shown that supplementing deoxynucleotides and Vpx, a protein encoded by HIV-2 and simian immunodeficiency virus, can enhance prime editing efficiency in hematopoietic stem and progenitor cells (HSPCs) . In addition, Vpx alone improves editing efficiency in quiescent HSPCs. Therefore, we examined the effect of Vpx alone and its combination with MLH1-SB. In Hepa1-6 cells, adding mRNA encoding Vpx alone did not enhance editing efficiency, but the combination of Vpx and MLH1-SB slightly but significantly improved editing efficiency (FIG. 22b) . The combination of Vpx and MLH1-SB encoding mRNA, along with PE6c encoding mRNA and heavily modified epegRNA resulted in the highest editing efficiency, reaching an average of 67.9%in bulk liver (FIG. 21e, FIG. 22c) . Since LNPs primarily target hepatocytes, which make up approximately 70%of liver cells, our data suggest that most hepatocyte genomes have been precisely edited by PE in mouse liver.
[0206] We then measured Pcsk9 and cholesterol levels in mouse serum. Compared to wild-type (WT) mice, mice injected with heavily modified epegRNA and PE6c LNPs showed an average reduction of 88.6%in Pcsk9 levels and a 32.7%reduction in cholesterol (FIG. 21f-g) . Addition of MLH1-SB and Vpx further reduced Pcsk9 levels by 96%and cholesterol by 50.7%in mice (FIG. 21f-g) . These reductions in physiological markers further confirm the editing capability of PE via non-viral delivery.
[0207] In this study, we systematically evaluated the chemical modification of epegRNA subdomains, extending its scope from end modifications to heavily modified epegRNA. This rationally designed modification significantly enhanced prime editing efficiency both in vitro and in vivo (FIG. 13-21) . Notably, 10 pmol of the heavily modified epegRNA achieved similar editing outcomes compared to the 90 pmol dose required for end-modified epegRNA. This represents a nine-fold reduction in dose, thereby expanding the therapeutic window for prime editing-based therapies (FIG. 13f, FIG. 18b) .
[0208] We demonstrate here systemic delivery using non-viral vectors can reach highly efficient prime editing of hepatocytes in vivo, with more than a ten-fold increase in editing efficiencies compared to published datasets and our initial tests using end-modified pegRNA and PEmax mRNA (FIG. 21) . Here, we achieved nearly 70%editing efficiency in bulk liver, establishing a strong foundation for potent in vivo prime editing via non-viral delivery. Notably, the editing efficiency in bulk liver via LNP (67.9%) exceeds the recently reported improvements in prime editing using viral delivery (up to 46%) , highlighting its significant translational potential. This advantage is further supported by the cost-effectiveness of LNP production and the transient expression of gene editors.
[0209] We investigated the chemical modification of individual components of epegRNA, including the spacer, scaffold, RTT, PBS, and evopreQ1, and assessed their effects on prime editing efficiency. Our findings indicate that extensive modifications of each component generally enhance editing efficiency (FIG. 13-17) . Reduced modification of the spacer decreases prime editing efficiency (FIG. 15) , while excessive modifications disrupt the interaction of the spacer with Cas9. To balance these effects, we retained a three-nucleotide modification at the 5′end of the spacer. Although RTT modifications enhanced editing efficiencies, they also significantly increased indel formation (FIG. 19a-c) . We believe that chemical modifications in the RTT region, acting as an RNA template, might introduce errors during reverse transcription, thereby elevating indel rates. In contrast, extensive modifications of the PBS, scaffold, and evopreQ1 regions improved prime editing efficiencies without increasing the proportion of indels (FIG. 19a-c) . Nevertheless, our findings suggest that the intensive chemical modification of the protective RNA motif can play a pivotal role in enhancing in vivo efficiency of a therapeutic RNA. This approach of utilizing chemically modified protective RNA motifs, therefore, can be broadly adapted to improve the performance of other functional RNA-based systems.
[0210] We were surprised to find that, while PE6c and PE6d mRNA exhibited reduced editing efficiencies compared to PEmax in murine hepatoma cell line Hepa1-6, they outperformed PEmax in the mouse liver (FIG. 20c, FIG. 21b) . Furthermore, mRNA encoding MLH1-SB showed no significant effect in Hepa1-6 cells but significantly enhanced editing efficiency in the mouse liver (FIG. 21d, FIG. 22a) . The distinct results for MLH1-SB may be attributed to the rapid dilution of MLH1 mRNA in the fast-dividing Hepa1-6 cells, which does not occur in vivo.
[0211] The performance of extensive modifications to the invariant regions, including those targeting the evopreQ1 and scaffold regions, appears suitable for various loci and forms of prime editing systems (FIG. 13, 14, 18, 21) . Notably, in vitro data align well with in vivo observations, suggesting that these modifications may have universal applicability for enhancing prime editing efficiencies. For the variable regions of PBS and RTT, chemical modifications can indeed improve efficiency (FIG. 15, 16) .
[0212] In conclusion, our study demonstrates that the systematic delivery of prime editors using non-viral vectors can achieve highly efficient in vivo editing. This success is driven by the use of heavily modified epegRNAs and the appropriate prime editor variants. We anticipate that heavily modified epegRNAs could also enhance ex vivo therapies employing prime editing. Furthermore, we showed that MMR inhibitor mRNA can be multiplexed with a proper version of PE mRNA to further boost editing efficiencies in vivo via non-viral delivery. The efficient in vivo prime editing achieved through the transient expression of prime editors lays a solid foundation for its clinical translation. ***
[0213] The present disclosure is not to be limited in scope by the specific embodiments described which are intended as single illustrations of individual aspects of the disclosure, and any compositions or methods which are functionally equivalent are within the scope of this disclosure. It will be apparent to those skilled in the art that various modifications and variations can be made in the methods and compositions of the present disclosure without departing from the spirit or scope of the disclosure. Thus, it is intended that the present disclosure cover the modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.
[0214] All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
Claims
1.A method for ligating two or more single-stranded target nucleic acid fragments, comprising:(a) incubating, in a sample, the two or more single-stranded target nucleic acid fragments with one or more single-stranded splint DNA fragment (s) at a denaturing temperature;(b) reducing the temperature of the sample to allow the target nucleic acid fragments and the splint DNA fragment (s) to form an at least partially duplex molecule by virtue of their sequence complementarity, wherein there is no gap between each adjacent target nucleic acid fragment in the duplex molecule;(c) incubating the sample at a ligation temperature, in the presence of a nucleic acid ligase, to allow the ligase to ligate the target nucleic acid fragments in the duplex molecule to form a ligated nucleic acid fragment; and(d) repeating steps (a) - (c) for at least one time to form more ligated nucleic acid fragments.2.The method of claim 1, wherein the target nucleic acid fragments are RNA fragments, RNA fragments with modified nucleic acids or RNA / DNA chimera fragments.3.The method of claim 1 or 2, wherein the denaturing temperature is 55 ℃ to 98 ℃, preferably 60 ℃ to 80 ℃, and more preferably 65 ℃ to 75 ℃.4.The method of any preceding claim, wherein the ligation temperature is 20 ℃ to 42 ℃, preferably 32 ℃ to 39 ℃, more preferably 35 ℃ to 38 ℃.5.The method of any preceding claim, wherein, in the duplex molecule, one of the strands includes only the target nucleic acid fragments and the other strand includes only the splint DNA fragment (s) .6.The method of claim 5, wherein the duplex molecule includes a single splint DNA fragment.7.The method of claim 6, wherein the duplex molecule includes two or three target nucleic fragments.8.The method of claim 7, wherein the duplex molecule includes two or more splint DNA fragments.9.The method of any preceding claim, further comprising (e) digesting the splint DNA fragment (s) .10.The method of any preceding claim, wherein at least one of the target nucleic fragments is chemically modified or comprises a non-natural nucleotide.11.The method of claim 10, wherein at least one of the target nucleic fragments comprises a chemically modified ribose selected from the group consisting of 2’-O-methyl (2’-O-Me) , 2’-Fluoro (2’-F) , 2’-deoxy-2’-fluoro-beta-D-arabino-nucleic acid (2’F-ANA) , 4’-S, 4’-SFANA, 2’-azido, UNA, 2’-O-methoxy-ethyl (2’-O-ME) , 2’-O-Allyl, 2’-O-Ethylamine, 2’-O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, and N-MeO-aminooxy BNA.12.The method of claim 10, wherein at least one of the target nucleic fragments comprises a phosphodiester linkage selected from the group consisting of Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’, 5’-amide, N3’-phosphoramidate (NP) , Phosphodiester (PO) , or 2’, 5’-phosphodiester (2’, 5’-PO) .13.The method of any preceding claim, wherein the target nucleic acid fragments, following ligation, constitute a pegRNA, which is optionally an engineered pegRNA (epegRNA) .14.The method of any preceding claim, wherein the target nucleic acid fragments, following ligation, have a total length of at least 150 nt, or at least 180 nt, 200 nt, 250 nt, or 300 nt.15.The method of any preceding claim, wherein each target nucleic acid fragment and each splint DNA fragment have a molar ratio, in the sample, that is from 0.5: 1.5 to 1.5: 0.5, or preferably from 0.8: 1.2 to 1.2: 0.8.16.The method of any preceding claim, wherein the nucleic acid ligase is T4 RNA ligase 2 (T4 Rnl2) .17.The method of any preceding claim, wherein at each step (c) when it is repeated, an amount of 0.5 to 5 μg / μl of the nucleic acid ligase is added.18.The method of any preceding claim, wherein the steps (a) - (c) are repeated once or twice in step (d) .19.The method of any preceding claim, wherein, in step (c) , the sample is incubated at the ligation temperature for 15 minutes to 3 hours, preferably 20 minutes to 90 minutes, more preferably 25 to 65 minutes.20.The method of any preceding claim, wherein each target nucleic acid fragment has a length of 10 nt to 200 nt.21.A prime editing guide RNA (pegRNA) , comprising (a) a single guide RNA (sgRNA) , (b) a reverse transcriptase (RT) template, (c) a primer binding site (PBS) , and (d) an RNA aptamer or an RNA pseudoknot motif, wherein the pegRNA is at least 125 nt in length, preferably at last 150nt, 180 nt, 200 nt, 250 nt or 300 nt in length.22.The pegRNA of claim 21, which is chemically modified or comprises a non-natural nucleotide.23.The pegRNA of claim 22, wherein at least 20%of the 50 3’ end nucleotides are modified, optionally at least 30%, 40%, 50%, 60%, 70%, 80%, 90%or 95%of the 50 3’ end nucleotides are modified.24.The pegRNA of claim 22, wherein at least 5 of the 15 3’ end nucleotides are modified, optionally at least 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of the 15 3’ end nucleotides are modified.25.The pegRNA of claim 23 or 24, wherein the modification is a ribose modification selected from 2'-O-methyl (2'-O-Me) , 2’-fluoro (2’-F) , and the combination thereof, preferably wherein the modification is a ribose modification with 2'-O-Me.26.The pegRNA of claim 22, wherein the pegRNA comprises a chemically modified ribose selected from the group consisting of 2’-O-methyl (2’-O-Me) , 2’-Fluoro (2’-F) , 2’-deoxy-2’-fluoro-beta-D-arabino-nucleic acid (2’F-ANA) , 4’-S, 4’-SFANA, 2’-azido, UNA, 2’-O-methoxy-ethyl (2’-O-ME) , 2’-O-Allyl, 2’-O-Ethylamine, 2’-O-Cyanoethyl, Locked nucleic acid (LAN) , Methylene-cLAN, N-MeO-amino BNA, and N-MeO-aminooxy BNA.27.The pegRNA of claim 22, wherein the pegRNA comprises a phosphodiester linkage selected from the group consisting of Phosphorothioate (PS) , Boranophosphate, phosphodithioate (PS2) , 3’, 5’-amide, N3’-phosphoramidate (NP) , Phosphodiester (PO) , or 2’, 5’-phosphodiester (2’, 5’-PO) .28.The pegRNA of any one of claims 21-27, wherein the RNA aptamer is selected from the group consisting of an MS2 aptamer, a tetrahymena thermophila group I intron, a theophylline aptamer, an ATP aptamer, a guanine quadruplex aptamer, a thrombin aptamer (TBA) , a vascular endothelial growth factor (VEGF) aptamer, an aptamer against HIV-1 reverse transcriptase, a hemin binding aptamer (HBA) , a riboswitch, a ribozyme.29.The pegRNA of any one of claims 21-27, wherein the RNA pseudoknot motif is selected from the group consisting of an H-type pseudoknot, a kissing-loop pseudoknot, a triple helix pseudoknot, a complex pseudoknot, a frame-shift pseudoknot, a riboswitch pseudoknot, a G-quadruplex pseudoknot, an intermolecular pseudoknot, wherein the RNA pseudoknot motif is preferably evopreQ1.30.The pegRNA of any one of claims 21-27, wherein the (d) comprises at least a stem-loop, and wherein no more than 50%of the nucleotides in the stem-loop are modified, preferably wherein no more than 40%, 30%, 20%or 10%of the nucleotides in the stem-loop are modified.31.The pegRNA of any one of claims 21-30, wherein the sgRNA comprises a spacer and a scaffold comprising, from 5’ to 3’, a first, a second, a third and a fourth stem-loops, and wherein the first and third stem-loops comprise modifications.32.The pegRNA of claim 31, wherein each of the first and third stem-loops comprises at least 2, 3, 4, 5, 6, 7, or 8 modified ribose’s selected from 2'-O-methyl (2'-O-Me) ribose, 2’-fluoro (2’-F) ribose, and the combination thereof, preferably wherein the modified ribose’s are 2'-O-methyl (2'-O-Me) ribose’s.33.The pegRNA of claim 31 or 32, wherein the fourth stem-loop comprises no more than 5, 4, 3, 2 or 1 modified ribose’s.34.The pegRNA of any one of claims 21-33, wherein the spacer comprises three consecutive phosphorothioate (PS) linkages at the 5’ end.35.The pegRNA of any one of claim 21-34, wherein (i) the spacer comprises three consecutive phosphorothioate (PS) linkages at the 5’ end, (ii) the first and third stem-loops of the scaffold of the sgRNA each comprises at least 8 nucleotides with 2'-O-methyl (2'-O-Me) modifications, and (iii) at least 20%of the 50 3’ end nucleotides of the pegRNA are modified with 2'-O-methyl (2'-O-Me) .36.The pegRNA of claim 35, wherein (i) the spacer comprises only three consecutive PS linkages at the 5’ end, (ii) the fourth stem-loop of the scaffold of the sgRNA comprises no more than 4 modified ribose’s, and (iii) the step-loop of (d) comprises no more than 40%modified ribose’s.
Citation Information
Patent Citations
Pilot editing system and gene editing method
CN116396952A
Simultaneous genome deletion and insertion based on guided editing
CN117321199A
SINGLE pegRNA-MEDIATED LARGE INSERTIONS
WO2023212594A2