Methods and compositions for simultaneous double-strand editing of target double-stranded nucleotide sequences.

JP7914010B2Active Publication Date: 2026-09-01THE BROAD INST INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022567857
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-20
Filing Date
2021-05-07
Publication Date
2026-09-01
Estimated Expiration
2041-05-07

Smart Images

  • Figure 0007914010000404
    Figure 0007914010000404
  • Figure 0007914010000405
    Figure 0007914010000405
  • Figure 0007914010000406
    Figure 0007914010000406
Patent Text Reader

Abstract

The present disclosure provides systems, compositions, and methods for simultaneously editing both strands of a double-stranded DNA sequence at a target site to be edited. In some aspects, the system includes first and second prime editor complexes, each of which includes (1) a prime editor comprising (i) a nucleic acid programmable DNA-binding protein (napDNAbp) and (ii) a polypeptide with RNA-dependent DNA polymerase activity; and (2) a pegRNA comprising a spacer sequence, a gRNA core, a DNA synthesis template, and a primer-binding site, wherein the DNA synthesis template encodes a desired DNA sequence or its complement, and the desired DNA sequence and its complement form a duplex containing an edited portion, which is incorporated into the target site to be edited. In some aspects, the system includes first, second, third, and fourth prime editor complexes, each of which includes a prime editor and a pegRNA. Methods for simultaneously editing both strands of a double-stranded DNA sequence at a target site to be edited are also provided. Further provided herein are pharmaceutical compositions, polynucleotides, vectors, cells, and kits.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Government support This invention was made with government support under grants U01AI142756, RM1HG009490, R01EB022376, and R35GM118062, awarded by the National Institutes of Health. The government has certain rights in this invention.

[0002] Incorporation by related applications and references This application claims priority to U.S. Provisional Application No. 63 / 022,397 filed 8 May 2020 and U.S. Provisional Application No. 63 / 116,785 filed 20 November 2020. The entire contents of each of these applications are incorporated into this application by reference.

[0003] This U.S. provisional application also covers the following applications, namely U.S. provisional application No. 62 / 820,813 filed on March 19, 2019 (agent reference number B1195.70074US00), U.S. provisional application No. 62 / 858,958 filed on June 7, 2019 (agent reference number B1195.70074US01), and U.S. provisional application No. 62 / 889,996 filed on August 21, 2019 ( (Agent reference number B1195.70074US02), U.S. Provisional Application No. 62 / 922,654 filed on August 21, 2019 (Agent reference number B1195.70083US00), U.S. Provisional Application No. 62 / 913,553 filed on October 10, 2019 (Agent reference number B1195.70074US03), U.S. Provisional Application No. 62 / 973 filed on October 10, 2019 U.S. Provisional Application No. 62 / 931,195 (Agent Reference Number B1195.70083US01), filed November 5, 2019 (Agent Reference Number B1195.70074US04), U.S. Provisional Application No. 62 / 944,231 (Agent Reference Number B1195.70074US05), filed December 5, 2019 (Agent Reference Number B1195.70074US05), U.S. Provisional Application No. 62 This also refers to and incorporates by reference U.S. Provisional Application No. 62 / 991,069 filed March 17, 2020 (Agent reference number B1195.70074US06), and U.S. Provisional Application No. 63 / 100,548 filed March 17, 2020 (Agent reference number B1195.70083US03). In addition, this U.S. provisional application incorporates by reference the following international PCT applications filed on 19 May 2020: PCT / US20 / 23721; PCT / US20 / 23730; PCT / US20 / 23713; PCT / US20 / 23712; PCT / US20 / 23727; PCT / US20 / 23724; PCT / US20 / 23725; PCT / US20 / 23728; PCT / US20 / 23732; PCT / US20 / 23723; PCT / US20 / 23553; and PCT / US20 / 23583. [Background technology]

[0004] Background of the present invention Pathogenic single-nucleotide mutations are estimated to contribute to approximately 50% of human diseases with a genetic component. 7 Unfortunately, despite decades of gene therapy exploration, treatment options for patients with these genetic disorders remain extremely limited. 8 Perhaps the simplest solution to this therapeutic challenge is the direct correction of a single nucleotide mutation in the patient's genome, which may address the root cause of the disease and provide lasting benefits. Such a strategy was previously unthinkable, but the CRISRP / Cas system 9 Recent improvements in genome editing capabilities, brought about by the advent of CRISPR, now put this therapeutic approach within reach. By simply designing a guide RNA (gRNA) sequence containing approximately 20 nucleotides complementary to the target DNA sequence, virtually any conceivable genomic region can be specifically accessed by CRISPR-related (Cas) nucleases. 1,2 To date, several monomeric bacterial Cas nuclease systems have been identified and adapted for genome editing applications. 10 This natural diversity of Cas nuclease, along with the ever-growing group of modified variants, 11~14 This will create fertile ground for developing new genome editing technologies.

[0005] Although gene disruption using CRISPR is now a mature technique, high-precision editing of single base pairs in the human genome remains a significant challenge. 3 Homology-directed repair (HDR) has long been used in human cells and other organisms to insert, modify, or exchange DNA sequences at double-strand break (DSB) sites using donor DNA repair templates that encode the desired edits. 15. However, existing HDR has extremely low efficiency in most human cell types, specifically in non-dividing cells, and the competing non-homologous end joining (NHEJ) primarily leads to insertion-deletion (indel) by-products 16 . Another problem relates to the generation of DSBs, which can cause large chromosomal rearrangements and deletions at the target locus, or 17 activate the p53 axis leading to growth arrest and apoptosis 18,19 .

[0006] Several approaches have been explored to address these drawbacks of HDR. For example, repair of single-strand DNA breaks (nicks) using oligonucleotide donors has been shown to reduce indel formation, but the yield of the desired repair product remains low 20 Other strategies attempt to bias repair towards HDR over NHEJ using small molecules and biological reagents 21~23 However, the effectiveness of these methods can be cell-type dependent, and perturbation of normal cellular states can lead to undesirable and unpredictable effects.

[0007] In recent years, the present inventors led by Professor David Liu developed base editing as a technology that edits target nucleotides without creating DSBs or relying on HDR 4~6,24~27Direct modification of DNA bases by Cas fusion deaminase enables highly efficient C·G→T·A or A·T→G·C base pair conversions in short target windows (~5-7 bases). As a result, base editors have been rapidly adopted by the scientific community. However, the following factors limit their generality for high-precision genome editing: (1) "bystander editing" of non-target C or A bases on the target window is observed; (2) a mixture of target nucleotide products is observed; (3) the target base must be located 15±2 nucleotides upstream of the PAM sequence; and (5) repair of small insertion and deletion mutations is not possible.

[0008] Therefore, the development of programmed editing factors that can flexibly introduce any desired single nucleotide change and / or incorporate base pair insertions or deletions (e.g., insertions or deletions of at least 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more base pairs) and / or alter or modify nucleotide sequences at target sites with high specificity and efficiency would substantially expand the scope and therapeutic potential of CRISPR-based genome editing technologies. [Overview of the project]

[0009] This invention describes a novel platform for genome editing called "multiple flap prime editing" (e.g., encompassing "double flap prime editing" and "quadruple flap prime editing"), representing a revolutionary advancement of "prime editing" or "classical prime editing," as described by the inventors in Anzalone, AV et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019), incorporated herein by reference. Classical prime editing, in various embodiments, polymerizes a single 3' flap at a nick site, which is then incorporated into a target nucleic acid on the same strand. The multiple flap prime editing systems described herein involve different constructs, systems, and methodologies, which, in various embodiments, generate pairs or multiple pairs of 3' flaps on different strands, which form a double helix containing the desired edit, and which are then incorporated into a target nucleic acid molecule, for example, at a specific locus or editing site on the genome. In various aspects, a pair or more of 3' flaps form a double helix because, once generated by the prime editing factors described herein, they contain reverse complementary sequences that anneal to each other. The double helix is ​​then incorporated into a target site by a cell-driven mechanism that spontaneously replaces endogenous double helix sequences located between adjacent nick sites. In some embodiments, the novel double helix sequence may be introduced at one or more locations (e.g., at adjacent genomic loci or at locations on two different chromosomes) and may contain one or more target sequences, such as protein-coding sequences, peptide-coding sequences, or RNA-coding sequences.In one embodiment, the novel double-stranded sequence incorporated by the multi-flap prime editing system is a recombinase site, e.g., Bxb1 recombinase attB (38 bp) and / or attP (50 bp) site, or Hin recombinase, Gin recombinase, Tn3 recombinase, β-six recombinase, CinH recombinase, ParA recombinase, γδ recombinase, φC31 recombinase, TP901 recombinase, TG1 It may contain recombinase sites recognized by recombinases, φBT1 recombinase, R4 recombinase, φRV1 recombinase, φFC1 recombinase, MR11 recombinase, A118 recombinase, U153 recombinase, and gp29 recombinase, Cre recombinase, FLP recombinase, R recombinase, lambda recombinase, HK101 recombinase, HK022 recombinase, and pSAM2 recombinase.

[0010] In recent years, the inventors have developed prime editing, which enables insertion, deletion, or replacement of genomic DNA sequences without requiring error-prone double-strand DNA breaks. Prime editing uses an engineered Cas9 nickase-reverse transcriptase fusion protein (PE1 or PE2) paired with an engineered prime editing guide RNA (pegRNA), which both guides Cas9 to the target genomic site and encodes information for incorporating the desired edit. Prime editing proceeds through a multi-step editing process: 1) The Cas9 domain binds to and nicks a target genomic DNA site defined by the spacer sequence of the pegRNA; 2) The reverse transcriptase domain uses the nicked genomic DNA as a primer to initiate the synthesis of an edited DNA strand using the manipulated extension on the pegRNA as a template for reverse transcription, which generates a single-stranded 3' flap containing the edited DNA sequence; 3) Cellular DNA repair eliminates the 3' flap intermediate by displacing the 5' flap species, which occurs through the entry of the edited 3' flap, the excision of the 5' flap containing the original DNA sequence, and the ligation of a new 3' flap to incorporate the edited DNA strand, forming a heteroduplex of one edited and one unedited strand; 4) Cellular DNA repair completes the editing process by replacing the unedited strand of the heteroduplex with the edited strand as a template for repair.

[0011] The efficient integration of a desired edit requires that the newly synthesized 3' flap contains a portion of the sequence homologous to the genomic DNA site. This homology allows the edited 3' flap to compete with the endogenous DNA strand (the corresponding 5' flap) for integration into the DNA double helix. Since the edited 3' flap will contain less sequence homology than the endogenous 5' flap, competition is expected to favor the 5' flap strand. Therefore, a possible limiting factor in the efficiency of prime editing may be the efficiency of the 3' flap insertion of the endogenous DNA and the subsequent displacement and replacement of the 5' flap strand. Moreover, successful 3' flap insertion and 5' flap removal integrate the edit into only one strand of the double-stranded DNA genome. Permanent integration of the edit requires cellular DNA repair to replace the undedited complementary DNA strand using the edited strand as a template. Cells can be made to prefer replacing the unedited strand with the edited strand by introducing a nick on the unedited strand adjacent to the edit using a secondary sgRNA (PE3 system) (step 4 above), although this process still relies on the second stage of DNA repair. For edits that require equilibration of long 5' and 3' flap intermediates or involve long non-homologous regions such as long insertions or long deletions, these DNA repair steps can be particularly inefficient. Further development of prime editing will advance this field.

[0012] In various aspects, this specification describes multiple flap prime editing systems (including, for example, double prime editing systems and quadruple prime editing systems). These systems address the challenges associated with flap equilibration and subsequent incorporation of edits into the unedited complementary genomic DNA strand by simultaneously editing both DNA strands. In a double flap prime editing system, for example, two pegRNAs are used to target opposing strands of a genomic site, leading to the synthesis of two complementary 3' flaps containing the edited DNA sequence (Figure 91). Unlike classical prime editing, there is no requirement that the pair of edited DNA strands (3' flaps) directly compete with the 5' flap on the endogenous genomic DNA, because the complementary edited strands are instead available for hybridization. Since both strands of the double helix are synthesized as edited DNA, the double flap prime editing system eliminates the need for replacement of the unedited complementary DNA strand required by classical prime editing. Instead, cellular DNA repair mechanisms only need to excise the paired 5' flap (original genomic DNA) and ligate the paired 3' flap (edited DNA) to the locus. Therefore, there is no need to include sequences homologous to the genomic DNA on the newly synthesized DNA strand, allow selective hybridization of the new strand, or facilitate editing that contains minimal genomic homology. Nuclease-activated versions of prime editing factors that cut both strands of DNA can also be used to accelerate the removal of the original DNA sequence. A quadruple-flap prime editing system using four pegRNAs offers similar advantages.

[0013] Multiple flap prime editing, such as classical prime editing, is a versatile and precise genome editing method that uses a nucleic acid-programmable DNA-binding protein ("napDNAbp") in cooperation with polymerase (i.e., in the form of a fusion protein, or otherwise provided in trans with napDNAbp) to directly write new genetic information into specific DNA sites. However, the prime editing system is programmed with a prime editing (PE) guide RNA ("PEgRNA") that identifies both the target site and template for the synthesis of the desired edit in the form of a replacement DNA strand via an extension (either DNA or RNA) that is modified onto (e.g., at the 5' or 3' end of the guide RNA, or within the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleic acid base substitution) shares the same sequence as the endogenous strand of the target site to be edited (with the exception that it contains the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand of the target site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a "search-and-replace" genome editing technique, as the prime editing factors described herein not only search for and locate the desired target sites to be edited, but also simultaneously encode a replacement strand containing the desired edits to be incorporated in place of the corresponding target sites on the endogenous DNA strand.

[0014] The multiflap-prime editing factors of this disclosure relate in part to the discovery that the mechanism of target-primed reverse transcription (TPRT) or "prime editing" can be utilized or employed to perform highly efficient and genetically flexible high-precision CRISPR / Cas-based genome editing (as illustrated, for example, in various embodiments in Figures 1A–1F). TPRT is naturally used by mobile DNA elements such as mammalian non-LTR retrotransposons and bacterial group II introns. 28,29The inventors herein use Cas protein-reverse transcriptase fusions or related systems in which a specific DNA sequence is targeted with guide RNA, a single-strand nick is generated at the target site, and the nicked DNA is used as a primer for reverse transcription of a modified reverse transcriptase template incorporated by the guide RNA. However, although this concept began with multiple flap prime editing factors that use reverse transcriptase as a component of DNA polymerase, the prime editing factors described herein are not limited to reverse transcriptase and may effectively encompass the use of DNA polymerase. In fact, this application sometimes refers to multiple flap prime editing factors that have “reverse transcriptase” throughout, but it is explained herein that reverse transcriptase is the only type of DNA polymerase that can work in multiple flap prime editing. Therefore, wherever “reverse transcriptase” is referred to herein, those skilled in the art will understand that any suitable DNA polymerase may be used instead of reverse transcriptase. Therefore, in one aspect, the multiple flap prime editing factor may include Cas9 (or equivalent napDNAbp) programmed to target a DNA sequence by linking the DNA sequence to a specialized guide RNA (i.e., PEgRNA) containing a spacer sequence that anneals with a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an elongation encoding a replacement strand of DNA containing the desired genetic modification, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer the information from the PEgRNA to the target DNA, the mechanism of multiple flap prime editing involves nicking a target site in one strand of DNA, exposing a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime the direct DNA polymerization of the edit-encoding elongation onto the PEgRNA into the target site. In various embodiments, the elongation—which provides a template for polymerization of the replacement strand containing the edit—can be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (such as reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase.

[0015] In classical prime editing, the newly synthesized strand (i.e., the replacement DNA strand containing the desired edit) formed by the prime editing factor disclosed herein will be homologous to the target sequence of the genome (i.e., have the same sequence), except that it contains the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) DNA strand is also sometimes referred to as a single-strand DNA flap, which competes for hybridization with a complementary homologous endogenous DNA strand, thereby replacing the corresponding endogenous strand. In some embodiments, the system may be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with a Cas9 domain, or provided trans with a Cas9 domain). The error-prone reverse transcriptase can introduce the change during single-strand DNA flap synthesis. Thus, in some embodiments, the error-prone reverse transcriptase can be used to introduce a nucleotide change into target DNA. Depending on the error-prone reverse transcriptase used with the system, the changes may be random or non-random.

[0016] In classical prime editing, the degradation of a hybridized intermediate (including a single-strand DNA flap synthesized by a reverse transcriptase hybridized with the endogenous DNA strand) can encompass the removal of the replaced flap (e.g., by the 5'-terminus DNA flap endonuclease, FEN1), the ligation of the synthesized single-strand DNA flap to the target DNA, and the assimilation of the desired nucleotide changes as a result of the cell's DNA repair and / or replication process. Because templated DNA synthesis imparts single-nucleotide precision to any nucleotide modification, including insertions and deletions, the breadth of this approach is extremely broad, and it is foreseeable that it could be used for countless applications in basic science and therapeutics.

[0017] In some aspects, this specification provides pairs of prime editing factors each comprising a nucleic acid programmed DNA-binding protein (napDNAbp) and a DNA polymerase. In some embodiments, each prime editing factor can perform genome editing by reverse transcription, which is primed by a target in the presence of an extended guide RNA.

[0018] In some aspects, this specification provides pairs of prime editing factors each comprising a nucleic acid-programmed DNA-binding protein (napDNAbp) and a DNA polymerase, the DNA polymerase being provided trans with napDNAbp. In various embodiments, each prime editing factor can perform genome editing by reverse transcription, which is primed by a target in the presence of an extended guide RNA.

[0019] In some aspects, this specification provides pairs of prime editing factors each comprising a nucleic acid-programmed DNA-binding protein (napDNAbp) and a reverse transcriptase. In various embodiments, each prime editing factor can perform genome editing by reverse transcription, which is primed by a target in the presence of an extended guide RNA.

[0020] In some aspects, this specification provides a pair of prime editing factors comprising a nucleic acid-programmed DNA-binding protein (napDNAbp) and a reverse transcriptase, the reverse transcriptase being provided in trans with napDNAbp. In various embodiments, each prime editing factor can perform genome editing by reverse transcription, which is primed by a target in the presence of an extended guide RNA.

[0021] In one embodiment, napDNAbp possesses nickase activity. napDNAbp may also be the Cas9 protein or its functional equivalent, such as nuclease-active Cas9, nuclease-inactive Cas9 (dCas9), or Cas9 nickase (nCas9).

[0022] In one embodiment, napDNAbp is selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaut, and optionally possesses nickase activity.

[0023] In another embodiment, each prime editing factor in a dual prime editing factor can bind to a target DNA sequence when complexed with an extended guide RNA.

[0024] In other embodiments, the target DNA sequence includes a target strand and a complementary non-target strand.

[0025] In another embodiment, the binding of the complexed prime editing factor to the extended guide RNA forms an R-loop. The R-loop may include (i) an RNA-DNA hybrid comprising the extended guide RNA and the target strand, and (ii) a complementary non-target strand.

[0026] In another embodiment, the complementary non-target chain is nicked to form a reverse transcriptase prime sequence with a free 3' end.

[0027] In various embodiments, the elongated guide RNA comprises (a) a guide RNA and (b) RNA elongation at the 5' or 3' end of the guide RNA or at an intramolecular position of the guide RNA. The RNA elongation may comprise (i) a reverse transcription template sequence containing the desired nucleotide changes, (ii) a reverse transcription primer binding site, and (iii) optionally a linker sequence. In various embodiments, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to the nick site, the single-stranded DNA flap containing the desired nucleotide changes.

[0028] In various embodiments, the RNA elongation is of a length of at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.

[0029] In another embodiment, the single-stranded DNA flap can hybridize to an endogenous DNA sequence adjacent to the nick site, thereby incorporating a desired nucleotide change. In yet another embodiment, the single-stranded DNA flap replaces an endogenous DNA sequence having a free 5' end and adjacent to the nick site. In one embodiment, the replaced endogenous DNA having a 5' end is excised by the cell.

[0030] In various embodiments, cellular repair of a single-stranded DNA flap results in the incorporation of a desired nucleotide change, thereby forming a desired product.

[0031] In various other embodiments, the desired nucleotide changes are incorporated into an editing window located between approximately -4 and +10 in the PAM sequence.

[0032] In another embodiment, the desired nucleotide change is incorporated into an editing window where the nick site is between approximately -5 and +5, or between approximately -10 and +10, or between approximately -20 and +20, or between approximately -30 and +30, or between approximately -40 and +40, or between approximately -50 and +50, or between approximately -60 and +60, or between approximately -70 and +70, or between approximately -80 and +80, or between approximately -90 and +90, or between approximately -100 and +100, or between approximately -200 and +200.

[0033] In various embodiments, each napDNAbp of the dual prime editing factor contains the amino acid sequence of SEQ ID NO: 18. In various other embodiments, the napDNAbp contains an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs: 26-39, 42-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); SEQ ID NOs: 77-86 (CP-Cas9); SEQ ID NOs: 18-25 and 87-88 (SpCas9); and SEQ ID NOs: 62-72 (Cas12).

[0034] In other embodiments, the disclosed compositions of prime editing factors and / or double prime editing factors may comprise any one of the amino acid sequences of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. In other embodiments, the reverse transcriptase may contain an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. These sequences may be naturally occurring reverse transcriptase sequences from, for example, retroviruses or retrotransposons, or the sequences may be recombinant.

[0035] In various other embodiments, the prime editing factors of the dual prime editing factors disclosed herein may include various structural configurations. For example, in embodiments in which the prime editing factor is provided as a fusion protein, each of the dual prime editing factor fusion proteins may include the structure NH2-[napDNAbp]-[reverse transcriptase]-COOH; or NH2-[reverse transcriptase]-[napDNAbp]-COOH, where each of the ]-[ indicates the presence of any linker sequence.

[0036] In various embodiments, the linker sequence includes an amino acid sequence that is at least 80%, 85%, 90%, 95%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769, or to any one of the linker amino acid sequences of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769.

[0037] In various embodiments, the desired nucleotide change incorporated into the target DNA may be a single nucleotide change (e.g., a transition or transversion), an insertion of one or more nucleotides, or a deletion of one or more nucleotides.

[0038] In certain cases, the insertion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.

[0039] In some other cases, the deletion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.

[0040] In another aspect, the disclosure provides an elongated guide RNA comprising a guide RNA and at least one RNA elongation. The RNA elongation may be located at the 3' end of the guide RNA. In another embodiment, the RNA elongation may be located at the 5' end of the guide RNA. In yet another embodiment, the RNA elongation may be located at an intramolecular position on the guide RNA. However, preferably, the intramolecular positioning of the elongated portion does not obstruct the function of the protospacer.

[0041] In various embodiments, prime editing factor guide RNA (PEgRNA) can bind to napDNAbp and direct the napDNAbp towards a target DNA sequence. The target DNA sequence may include a target strand and a complementary non-target strand, where the guide RNA hybridizes with the target strand to form an RNA-DNA hybrid and an R-loop.

[0042] In various embodiments of the prime editing factor guide RNA, at least one RNA elongation includes a DNA synthesis template. In various other embodiments, the RNA elongation further includes a reverse transcription primer binding site. In yet another embodiment, the RNA elongation includes a linker or spacer that binds the RNA elongation to the guide RNA.

[0043] In various embodiments, the RNA elongation may have a length of at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.

[0044] In another embodiment, the DNA synthesis template (i.e., the editing template for Figure 27) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.

[0045] In yet another embodiment, the reverse transcription primer binding site sequence (i.e., primer binding site for Figure 27) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.

[0046] In another embodiment, any linker or spacer has a length of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.

[0047] In various embodiments of the extended guide RNA, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to a nick site, the single-stranded DNA flap containing the desired nucleotide changes. The single-stranded DNA flap may replace the endogenous single-stranded DNA at the nick site. The replaced endogenous single-stranded DNA at the nick site may have a 5' end, forming an endogenous flap that can be excised by the cell. In various embodiments, excision of the 5' endogenous flap may drive product formation because removing the 5' endogenous flap promotes hybridization of the single-stranded 3' DNA flap to the corresponding complementary DNA strand and the incorporation or assimilation of the desired nucleotide changes carried by the single-stranded 3' DNA flap into the target DNA.

[0048] In various forms of the extended guide RNA, cellular repair of the single-stranded DNA flap results in the incorporation of desired nucleotide changes, thereby forming the desired product.

[0049]

[0043]

[0042] In one embodiment, PEgRNA is sequence numbers 101-104, 181-183, 223-234, 237-244, 277, 324-330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 394, 429-442, 499-505, 641-649, 678-692, Nucleotide sequences of 735-736, 757-761, 776-777, 2997-3103, 3113-3121, 3305-3455, 3479-3493, 3522-3540, 3549-3556, 3628-3698, 3755-3810, 3874, 3890-3901, 3905-3911, 3913-3929, and 3972-3989, or at least 85%, or at least 90%, or The sequence identity of sequence numbers 101-104, 181-183, 223-234, 237-244, 277, 324-330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 394, 429-442, 499-505, and 641 is at least 95%, or at least 98%, or at least 99%. Includes a nucleotide sequence having any one of the following: ~649, 678~692, 735~736, 757~761, 776~777, 2997~3103, 3113~3121, 3305~3455, 3479~3493, 3522~3540, 3549~3556, 3628~3698, 3755~3810, 3874, 3890~3901, 3905~3911, 3913~3929, and 3972~3989.

[0050] In another aspect of the present invention, this specification provides a complex comprising a prime editing factor described herein and any of the extended guide RNAs described above.

[0051] In another further aspect of the present invention, this specification provides a complex comprising napDNAbp and an extended guide RNA. napDNAbp may be Cas9 nickase, or may be an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequences of SEQ ID NOs. 42-57 (Cas9 nickase) and 65 (AsCas12a nickase), or one of the amino acid sequences of SEQ ID NOs. 42-57 (Cas9 nickase) and 65 (AsCas12a nickase).

[0052] In various embodiments involving the complex, the extended guide RNA can direct napDNAbp to the target DNA sequence. In various embodiments, the reverse transcriptase may be supplied trans, i.e., from a source different from the complex itself. For example, the reverse transcriptase may be supplied to the same cell containing the complex by introducing separate vectors that individually encode the reverse transcriptase.

[0053] In another aspect, the disclosure provides a system comprising first and second prime-editing factor complexes, each complex comprising a prime-editing factor and a prime-editing guide RNA (PEgRNA). In some embodiments, each prime-editing factor comprises a nucleic acid-programmed DNA-binding protein (napDNAbp) and a polypeptide having RNA-dependent DNA polymerase activity, and each PEgRNA comprises a spacer sequence, a gRNA core, a DNA synthesis template, and a primer-binding site. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence containing the edited portion. The two encoded single-stranded DNA sequences may be complementary to each other and may form a double helix incorporated into the target site to be edited. In some embodiments, the two encoded single-stranded DNA sequences may contain regions of complementarity to each other. In some aspects, the two single-stranded DNA sequences to be encoded may contain complementary regions to each other, which are at least 2 bp, at least 3 bp, at least 4 bp, at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp in length. In some aspects, the prime editing factor is provided as a fusion protein. In some aspects, the components of the prime editing factor (i.e., napDNAbp and polypeptides having RNA-dependent DNA polymerase activity) are provided in trans.

[0054] In another aspect, the disclosure provides systems comprising first, second, third, and fourth prime-editing factor complexes, each complex comprising a prime-editing factor and a prime-editing guide RNA (PEgRNA). In some embodiments, each prime-editing factor comprises a nucleic acid-programmed DNA-binding protein (napDNAbp) and a polypeptide having RNA-dependent DNA polymerase activity, and each PEgRNA comprises a spacer sequence, a gRNA core, a DNA synthesis template, and a primer-binding site. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence containing the edited portion. The two encoded single-stranded DNA sequences may be complementary to each other and may form a double helix incorporated into the target site to be edited. In some embodiments, the two encoded single-stranded DNA sequences may contain regions of complementarity to each other. In some aspects, the two single-stranded DNA sequences to be encoded may contain complementary regions to each other, which are at least 2 bp, at least 3 bp, at least 4 bp, at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp in length. In some aspects, the prime editing factor is provided as a fusion protein. In some aspects, the components of the prime editing factor (i.e., napDNAbp and polypeptides having RNA-dependent DNA polymerase activity) are provided in trans.

[0055] In some embodiments, each napDNAbp is a Cas9 domain or a variant thereof. In some embodiments, each napDNAbp is a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase domain, or a variant thereof. In some embodiments, each napDNAbp is independently selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaut, and optionally has nickase activity. In various embodiments, each napDNAbp contains one amino acid sequence of sequence numbers 2-65, or an amino acid sequence that is at least 80%, 85%, 90%, 95%, or 99% identical to one of sequence numbers 2-65.

[0056] In some embodiments, the polypeptide having RNA-dependent DNA polymerase activity is a reverse transcriptase. In one embodiment, the polypeptide having RNA-dependent DNA polymerase activity comprises one amino acid sequence of SEQ ID NOs. 37, 68-79, 82-98, 81, 98, and 110, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity with respect to one of SEQ ID NOs. 37, 68-79, 82-98, 81, 98, and 110.

[0057] In some embodiments, each prime editing factor may include a linker that ligates napDNAbp and reverse transcriptase. In some embodiments, the linker includes one amino acid sequence of sequence numbers 119-128, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity with one of sequence numbers 119-128. Each linker may have a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 38, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids.

[0058] In some embodiments, each PEgRNA may independently contain one nucleotide sequence from sequence numbers 192-203, or a nucleotide sequence having 80%, 85%, 90%, 95%, or 99% sequence identity with respect to one of sequence numbers 192-203.

[0059] In various embodiments, each spacer sequence of each PEgRNA may bind to a specific binding site on a double-stranded DNA sequence adjacent to the target site to be edited. In some embodiments, the binding of a spacer sequence of one prime editing factor complex to one strand of a double-stranded DNA sequence and the binding of a spacer sequence of another prime editing factor complex to the opposite strand of the double-stranded DNA sequence result in nicking of both DNA strands at a nick site proximal to the PAM sequence on each strand.

[0060] In another aspect, the present disclosure provides polynucleotides. In some embodiments, a polynucleotide may encode any of the complexes described herein. In some embodiments, a polynucleotide may encode any of the PEgRNAs described herein.

[0061] In yet another aspect, this specification provides polynucleotides. In one embodiment, a polynucleotide may encode any of the prime editing factors disclosed herein. In certain other embodiments, a polynucleotide may encode any of the napDNAbp disclosed herein. In yet another embodiment, a polynucleotide may encode any of the reverse transcriptases disclosed herein. In yet another embodiment, a polynucleotide may encode any of the extended guide RNAs, any of the reverse transcription template sequences, any of the reverse transcription primer sites, or any linker sequence disclosed herein.

[0062] Still in other aspects, this specification provides vectors comprising polynucleotides described herein. Therefore, in some embodiments, the vector comprises a polynucleotide for encoding a prime editing factor comprising napDNAbp and reverse transcriptase (i.e., expressed as a fusion protein or in trans). In some embodiments, the vector comprises a polynucleotide for encoding any of the complexes described herein. In other embodiments, the vector comprises polynucleotides separately encoding napDNAbp and reverse transcriptase. Still in other embodiments, the vector may comprise a polynucleotide for encoding an extended guide RNA. In various embodiments, the vector may comprise one or more polynucleotides encoding napDNAbp, reverse transcriptase, and extended guide RNA on the same or separate vectors. In some embodiments, the vector comprises a polynucleotide for encoding any of the pegRNAs described herein.

[0063] Still in other respects, this specification provides cells comprising the prime editing factor and extended guide RNA described herein. Cells may be transformed by a vector comprising the prime editing factor, napDNAbp, reverse transcriptase, and extended guide RNA. These genetic elements may be contained on the same vector or on different vectors. In some embodiments, cells comprise any of the systems or complexes described herein. Cells may be transformed by a polynucleotide encoding any of the systems, complexes, and / or pegRNAs disclosed herein, or by a vector comprising a polynucleotide encoding any of the systems, complexes, or pegRNAs disclosed herein.

[0064] In other aspects, this specification provides pharmaceutical compositions. In one embodiment, a pharmaceutical composition comprises one or more napDNAbp, a prime editing factor, a reverse transcriptase, and an extended guide RNA. In one embodiment, a pharmaceutical composition comprises any of the systems and / or complexes described herein. In one embodiment, a pharmaceutical composition comprises any of the prime editing factors, systems, or complexes described herein and a pharmaceutically acceptable excipient. In another embodiment, a pharmaceutical composition comprises any of the extended guide RNAs described herein and a pharmaceutically acceptable excipient. Still in yet another embodiment, a pharmaceutical composition comprises any of the extended guide RNAs described herein in combination with any of the prime editing factors described herein and a pharmaceutically acceptable excipient. Still in yet another embodiment, a pharmaceutical composition comprises any of the polynucleotide sequences encoding one or more napDNAbp, a prime editing factor, a reverse transcriptase, and an extended guide RNA, or any of the vectors disclosed herein. Still in yet another embodiment, the various components disclosed herein can be separated into one or more pharmaceutical compositions. For example, the first pharmaceutical composition may contain a prime editing factor or napDNAbp, the second pharmaceutical composition may contain a reverse transcriptase, and the third pharmaceutical composition may contain an extended guide RNA.

[0065] In further aspects, this disclosure provides kits. In one embodiment, a kit comprises one or more polynucleotides encoding one or more components comprising a prime editing factor, napDNAbp, reverse transcriptase, and an extended guide RNA. The kit may also comprise a vector, cells, and isolated preparations of polypeptides comprising any of the prime editing factors, napDNAbp, or reverse transcriptases disclosed herein.

[0066] In another aspect, the Disclosure provides methods using the disclosed compositions and encompasses methods using any of the systems described herein to simultaneously edit both complementary strands of a double-stranded DNA sequence at a target site. In some embodiments, the methods include contacting a double-stranded DNA sequence with any of the systems disclosed herein.

[0067] In one aspect, the disclosure provides a method comprising contacting first and second prime-editing factor complexes with a double-stranded DNA sequence at a target site, each complex comprising a prime-editing factor and a prime-editing guide RNA (PEgRNA). In some embodiments, each prime-editing factor comprises a nucleic acid-programmed DNA-binding protein (napDNAbp) and a polypeptide having RNA-dependent DNA polymerase activity, and each PEgRNA comprises a spacer sequence, a gRNA core, a DNA synthesis template, and a primer-binding site. In some embodiments, each prime-editing factor is provided as a fusion protein. In some embodiments, the components of the prime-editing factor are provided in trans. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence containing the edited portion. The two encoded single-stranded DNA sequences may be complementary to each other and may form a double helix incorporated into the target site to be edited. Various elements of the prime-editing factor complex may include any of the embodiments of the system disclosed herein.

[0068] In another aspect, the disclosure provides a method comprising contacting first, second, third, and fourth prime editing factor complexes with a double-stranded DNA sequence at a target site, each complex comprising a prime editing factor and a prime editing guide RNA (PEgRNA). In some embodiments, each prime editing factor comprises a nucleic acid programmed DNA-binding protein (napDNAbp) and a polypeptide having RNA-dependent DNA polymerase activity, and each PEgRNA comprises a spacer sequence, a gRNA core, a DNA synthesis template, and a primer binding site. In some embodiments, each prime editing factor is provided as a fusion protein. In some embodiments, the components of the prime editing factor are provided in trans. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence. The two encoded single-stranded DNA sequences may be complementary to each other and may form a double helix incorporated into the target site to be edited. Various elements of the prime editing factor complex may include any of the embodiments of the system disclosed herein.

[0069] In some embodiments, the method provided herein allows inversion of the target DNA sequence. In some embodiments, a first single-stranded DNA sequence encoded by a first DNA synthesis template and a second single-stranded DNA sequence encoded by a second DNA synthesis template are at opposite ends of the target DNA sequence, and a third single-stranded DNA sequence encoded by a third DNA synthesis template and a fourth single-stranded DNA sequence encoded by a fourth DNA synthesis template are at opposite ends of the same target DNA sequence.

[0070] In some embodiments, the methods provided herein further include providing a circular DNA donor. In some embodiments, a first single-stranded DNA sequence encoded by a first DNA synthesis template and a third single-stranded DNA sequence encoded by a third DNA synthesis template are located at opposite ends of a target DNA sequence, and a second single-stranded DNA sequence encoded by a second DNA synthesis template and a fourth single-stranded DNA sequence encoded by a fourth DNA synthesis template are located on a circular DNA donor. In some embodiments, the portion of the circular DNA donor between the second and fourth single-stranded DNA sequences replaces the target DNA sequence between the first and third single-stranded DNA sequences.

[0071] In some embodiments, the method provided herein allows translocation of a target DNA sequence from a first nucleic acid molecule (e.g., a first chromosome) to a second nucleic acid molecule (e.g., a second chromosome). In some embodiments, a first single-stranded DNA sequence encoded by a first DNA synthesis template and a third single-stranded DNA sequence encoded by a third single-stranded DNA synthesis template are located on the first nucleic acid molecule, and a second single-stranded DNA sequence encoded by a second DNA synthesis template and a fourth single-stranded DNA sequence encoded by a fourth DNA synthesis template are located on the second nucleic acid molecule. In some embodiments, a portion of the first nucleic acid molecule between the first and third single-stranded DNA sequences is incorporated into the second nucleic acid molecule. In some embodiments, a portion of the second nucleic acid molecule between the second and fourth single-stranded DNA sequences is incorporated into the first nucleic acid molecule.

[0072] In another aspect, this disclosure provides a pair of PEgRNAs for use in multiple flap prime editing. In some embodiments, the pair comprises a first PEgRNA and a second PEgRNA, each PEgRNA independently comprising a spacer sequence, a gRNA core, a DNA synthesis template, and a primer binding site. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence. In various embodiments, the multiple flap prime editing factor is used in conjunction with a pair of PEgRNAs that target separate prime editing factors to both sides of a target site, with each pair of PEgRNAs encoding a 3' nucleic acid flap containing nucleic acid sequences that are each other's reverse complements. In various embodiments, the 3' flaps containing the reverse complementary sequences can anneal to each other to form a double helix containing the desired edit or nucleic acid sequence encoded by the PEgRNAs. The double helix is ​​then incorporated into the target site by replacement of a corresponding endogenous double helix located between adjacent nick sites.

[0073] In another aspect, the disclosure provides a plurality of PEgRNAs for use in multiple flap prime editing. In some embodiments, the plurality comprises first, second, third, and fourth PEgRNAs. In some embodiments, each of the four PEgRNAs independently comprises a spacer sequence, a gRNA core, a DNA synthesis template, and a primer binding site. In some embodiments, each DNA synthesis template encodes a single-stranded DNA sequence. The two encoded single-stranded DNA sequences may be complementary to each other.

[0074] In various aspects, this disclosure provides polynucleotides encoding any pair or more of the PEgRNAs described herein. In one aspect, this disclosure provides vectors encoding polynucleotides encoding any pair or more of the PEgRNAs described herein. In yet another aspect, this disclosure provides cells comprising a vector encoding a polynucleotide encoding any pair or more of the PEgRNAs described herein. In yet another aspect, this disclosure provides a pharmaceutical composition comprising cells comprising any pair or more of the PEgRNAs described herein, a vector encoding any pair or more of the PEgRNAs described herein, or a vector encoding any pair or more of the PEgRNAs described herein. In one embodiment, the pharmaceutical composition comprises a pharmaceutical excipient.

[0075] In one embodiment, the method relates to a method for incorporating a desired nucleotide change onto a double-stranded DNA sequence. The method first involves contacting the double-stranded DNA sequence with a complex comprising a prime editing factor and an extended guide RNA, wherein the prime editing factor comprises napDNAbp and reverse transcriptase, and the extended guide RNA comprises a reverse transcription template sequence containing the desired nucleotide change. In some embodiments, each prime editing factor is provided as a fusion protein. In some embodiments, the components of the prime editing factor are provided in trans. Next, the method involves nicking the double-stranded DNA sequence on a non-target strand, thereby generating a free single-stranded DNA having a 3' end. Then, the method involves hybridizing the 3' end of the free single-stranded DNA with a reverse transcription template sequence, thereby priming the reverse transcriptase domain. Then, the method involves polymerizing the DNA strand from the 3' end, thereby generating a single-stranded DNA flap containing the desired nucleotide change. Furthermore, the method involves replacing the endogenous DNA strand adjacent to the cut site with a single-stranded DNA flap, thereby incorporating the desired nucleotide changes into the double-stranded DNA sequence.

[0076] In another embodiment, the disclosure provides a method for introducing one or more changes into the nucleotide sequence of a DNA molecule at a target locus, comprising contacting the DNA molecule with a nucleic acid-programmed DNA-binding protein (napDNAbp) and a guide RNA that causes the napDNAbp to target the target locus, wherein the guide RNA comprises a reverse transcriptase (RT) template sequence containing at least one desired nucleotide change. The method then involves forming an exposed 3' end in the DNA strand at the target locus, and then priming reverse transcription by hybridizing the exposed 3' end with the RT template sequence. Next, a single-strand DNA flap containing at least one desired nucleotide change based on the RT template sequence is synthesized or polymerized by reverse transcriptase. Finally, at least one desired nucleotide change is incorporated into the corresponding endogenous DNA, thereby introducing one or more changes into the nucleotide sequence of the DNA molecule at the target locus.

[0077] In further other embodiments, the Disclosure provides a method for introducing one or more changes into the nucleotide sequence of a DNA molecule at a target locus by targeted-primed reverse transcription, the method comprising: (a) contacting the DNA molecule at the target locus with a guide RNA comprising i) a prime editing factor including a nucleic acid-programmed DNA-binding protein (napDNAbp) and reverse transcriptase, and (ii) an RT template containing the desired nucleotide changes; (b) causing targeted-primed reverse transcription of the RT template to generate single-stranded DNA containing the desired nucleotide changes; and (c) incorporating the desired nucleotide changes into the DNA molecule at the target locus through DNA repair and / or replication processes.

[0078] In one embodiment, the step of replacing an endogenous DNA strand includes: (i) creating a sequence mismatch by hybridizing a single-stranded DNA flap with an endogenous DNA strand adjacent to the cleavage site; (ii) excising the endogenous DNA strand; and (iii) repairing the mismatch to form a desired product containing the desired nucleotide changes in both strands of DNA.

[0079] In various embodiments, the desired nucleotide change may be a single nucleotide substitution (e.g., a transition or transversion), deletion, or insertion. For example, the desired nucleotide change may be (1) a G→T substitution, (2) a G→A substitution, (3) a G→C substitution, (4) a T→G substitution, (5) a T→A substitution, (6) a T→C substitution, (7) a C→G substitution, (8) a C→T substitution, (9) a C→A substitution, (10) an A→T substitution, (11) an A→G substitution, or (12) an A→C substitution.

[0080] In other embodiments, the desired nucleotide change may be (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair.

[0081] In another embodiment, the method introduces a desired nucleotide change, which is an insertion. In certain cases, the insertion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.

[0082] In another aspect, the method introduces a desired nucleotide change, which is a deletion. In some other cases, the deletion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.

[0083] In various embodiments, the desired nucleotide alteration modifies disease-associated genes. Disease-associated genes may be associated with monogenic disorders selected from the group consisting of: adenosine deaminase (ADA) deficiency; alpha-1 antitrypsin deficiency; cystic fibrosis; Duchenne muscular dystrophy; galactosemia; hemochromatosis; Huntington's disease; maple syrup urine disease; Marfan syndrome; neurofibromatosis type 1; onychoplasia; phenylketonuria; severe combined immunodeficiency; sickle cell anemia; Smith-Lemle-Oppitz syndrome; and Tay-Sachs disease. In other embodiments, disease-associated genes may be associated with polygenic disorders selected from the group consisting of: heart disease; hypertension; Alzheimer's disease; arthritis; diabetes mellitus; cancer; and obesity.

[0084] The methods disclosed herein may involve a fusion protein having napDNAbp, which is a nuclease-inactive (dead) Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In other embodiments, napDNAbp and reverse transcriptase may be provided in separate constructs rather than encoded as a single fusion protein. Thus, in some embodiments, reverse transcriptase may be provided trans to napDNAbp (rather than as a fusion protein).

[0085] In various embodiments of the method, napDNAbp may include the amino acid sequences of SEQ ID NOs. 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467 and 482-487 (Cas9); (SpCas9); SEQ ID NOs. 77-86 (CP-Cas9); SEQ ID NOs. 18-25 and 87-88 (SpCas9); and SEQ ID NOs. 62-72 (Cas12). napDNAbp may also include amino acid sequences that are at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs. 77-86 (CP-Cas9); SEQ ID NOs. 18-25 and 87-88 (SpCas9); and SEQ ID NOs. 62-72 (Cas12).

[0086] In various embodiments of the method, the reverse transcriptase may contain any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. The reverse transcriptase may also contain an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766.

[0087] The methods include sequence numbers 101-104, 181-183, 223-234, 237-244, 277, 324-330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 394, 429-442, 499-505, 641-649, 678-692, 735-736, 757-761, 776-777, 2997-3103, 3113-3121, 3305-34 The method may involve the use of PEgRNA containing nucleotide sequences 55, 3479-3493, 3522-3540, 3549-3556, 3628-3698, 3755-3810, 3874, 3890-3901, 3905-3911, 3913-3929, and 3972-3989, or nucleotide sequences having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% sequence identity to them. The method may involve the use of an elongated guide RNA containing RNA elongation at its 3' end, where the RNA elongation contains a reverse transcription template sequence.

[0088] The method may involve the use of an elongated guide RNA containing RNA elongation at its 5' end, where the RNA elongation contains a reverse transcription template sequence.

[0089] The method may involve the use of an elongated guide RNA, including RNA elongation, to position the guide RNA within the molecule, and the RNA elongation includes a reverse transcription template sequence.

[0090] The method may include the use of an elongated guide RNA having one or more RNA elongations that are at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0091] It should be understood that the above concepts, and the additional concepts discussed below, can be arranged in any preferred combination (however, this disclosure is not limited thereto). Furthermore, other advantages and novel features of this disclosure will become apparent from the detailed description of various non-limiting aspects below, when considered in conjunction with the attached figures. [Brief explanation of the drawing]

[0092] Simple description of the drawing The following drawings form part of this specification and are included to further demonstrate certain aspects of the disclosure that can be better understood by referring to one or more of these drawings in combination with the detailed description of the particular aspects presented herein.

[0093] [Figure 1A]Figure 1A provides a schematic diagram of an exemplary process for introducing single-nucleotide alterations, and / or insertions, and / or deletions, into a DNA molecule (e.g., a genome) using a fusion protein comprising a reverse transcriptase fused with a Cas9 protein conjugated with an extended guide RNA molecule. In this embodiment, the guide RNA contains the reverse transcriptase template sequence by being extended at its 3' end. The schematic diagram shows how the reverse transcriptase (RT), fused with Cas9 nickase and conjugated with the guide RNA (gRNA), binds to the DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which is used as a template for the synthesis of a new DNA strand encoding the desired edit. The editing process shown may be referred to as targeted-primed reverse transcription editing (TRT editing), or "primed editing."

[0094] [Figure 1B] Figure 1B provides the same diagram as Figure 1A, with the exception that the prime editing factor complex is more commonly represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNA. In the formula, "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid programmed DNA-binding protein (e.g., SpCas9), "PEgRNA" refers to the prime editing guide RNA, and ]-[ refers to any linker. Elsewhere, as shown in Figures 3A-3G, for example, PEgRNA includes a 5' elongation arm containing the primer binding site and the DNA synthesis template. Although not shown, it is intended that the elongation arm of PEgRNA (i.e., this contains the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase intended in this configuration will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase may be RNA-dependent DNA polymerase (e.g., reverse transcriptase). When the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase.

[0095] [Figure 1C] Figure 1C provides a schematic diagram of an exemplary process for introducing a single nucleotide change and / or insertion and / or deletion into a DNA molecule (e.g., a genome), using a fusion protein containing reverse transcriptase fused to a Cas9 protein as a complex with an elongated guide RNA molecule. In this embodiment, the guide RNA is elongated at its 5' end to encompass the reverse transcriptase template sequence. The schematic diagram shows how reverse transcriptase (RT) fused to Cas9 nickase, as a complex with guide RNA (gRNA), binds to the DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which is used as a template for the synthesis of a new DNA strand encoding the desired edit. The editing process shown may be called primed reverse transcription editing (TRT editing) or equivalently "primed editing".

[0096] [Figure 1D]Figure 1D provides the same depiction as Figure 1C, except that the prime editing factor complex is more commonly represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNAPEgRNA, where "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid programmed DNA-binding protein (e.g., SpCas9), "PEgRNA" refers to the prime editing guide RNA, and "[]-[" refers to any linker. Elsewhere, as shown in Figures 3A-3G for example, the PEgRNA includes a 3' elongation arm containing the primer binding site and the DNA synthesis template. Although not shown, it is intended that the elongation arm of the PEgRNA (i.e., containing the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase intended in this configuration would depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). If the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase. In various embodiments, PEgRNA can be modified or synthesized to incorporate DNA-based DNA synthesis templates.

[0097] [Figure 1E] Figure 1E is a schematic diagram illustrating an illustrative process of how a synthesized single DNA strand (containing the desired nucleotide changes) is degraded so that the desired nucleotide changes are incorporated into the DNA. As shown, subsequent synthesis of the edited strand (or "mutagenic strand"), equilibrium with the endogenous strand, flap cleavage of the endogenous strand, and ligation lead to the incorporation of the DNA edit after the degradation of the mismatched DNA double helix through the action of the endogenous DNA repair and / or replication processes.

[0098] [Figure 1F]Figure 1F is a schematic diagram showing how incorporating "opposite strand nicking" into the degradation method shown in Figure 1E can help promote the formation of the desired product-pair-restore product. In opposite strand nicking, a second Cas9 / gRNA complex is used to introduce a second nick onto the opposite strand of the initially nicked strand. This induces the endogenetic DNA repair and / or replication process to preferentially replace the unedited strand (i.e., the strand containing the second nick site).

[0099] [Figure 1G]Figure 1G provides another schematic diagram of an exemplary process for introducing single-nucleotide alterations, and / or insertions, and / or deletions into a DNA molecule (e.g., a genome) at a target locus using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an elongated guide RNA. This process is sometimes referred to as prime editing. The elongated guide RNA involves elongation at the 3' or 5' end of the guide RNA, or at some location within the molecule within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts the DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) to one strand of the DNA at the target locus (R-loop strand, or PAM-containing strand, or non-target DNA strand, or protospacer strand), thereby creating a usable 3' end on one strand of the target locus. In one embodiment, the nick is created on the DNA strand corresponding to the R loop strand, i.e., the strand that does not hybridize with the guide RNA sequence. In step (c), the 3'-end DNA strand interacts with the extended portion of the guide RNA to prime the reverse transcription. In some embodiments, the 3'-end DNA strand hybridizes with a specific RT prime sequence on the extended portion of the guide RNA. In step (d), reverse transcriptase is introduced to synthesize a single strand of DNA from the 3' end of the primed site to the 3' end of the guide RNA. This forms a single-strand DNA flap containing the desired nucleotide change (e.g., a single base change, insertion, deletion, or a combination thereof). In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-strand DNA flap so that the desired nucleotide change is incorporated into the target locus. This process can drive the formation of the desired product by removing the corresponding 5' endogenous DNA flap once the 3' single-strand DNA flap has penetrated and hybridized with a complementary sequence on the other strand. The process can also be driven to the formation of a product with a nicked second strand, as illustrated in Figure 1F.This process may introduce at least one or more of the following genetic changes: transversion, transition, deletion, and insertion.

[0100] [Figure 1H] Figure 1H is a schematic diagram illustrating the types of gene changes that can occur in the prime editing process as described herein. The types of nucleotide changes that can be achieved by prime editing include deletions (including short and long deletions), single nucleotide changes (including transitions and transversions), and insertions (including short and long ones).

[0101] [Figure 1I] Figure 1I is a schematic diagram illustrating a temporary nick to the second strand, exemplified by PE3b (PE3b = PE2 prime editing factor fusion protein + PEgRNA + guide RNA for nicking the second strand). Temporary nicking to the second strand is a variant of nicking to the second strand that facilitates the formation of the desired edited product. The term "temporary" refers to the fact that the second strand nick to the unedited strand occurs only after the desired edit has been incorporated into the edited strand. This avoids simultaneous nicking on both strands, which would lead to double-strand DNA breaks.

[0102] [Figure 1J-1K]Figures 1J-1K illustrate variations of prime editing as envisioned herein, in which a napDNAbp (e.g., SpCas9 nickase) is replaced with any programmed nuclease domain, such as a zinc finger nuclease (ZFN) or a transcription activator-like effector nuclease (TALEN). Therefore, the preferred nuclease does not necessarily need to be "programmed" by the nucleic acid target molecule (e.g., guide RNA), but rather may be programmed, particularly by defining the specificity of the DNA-binding domain, such as the nuclease. Just as with prime editing with the napDNAbp moiety, the alternative programmed nuclease is preferably modified to cleave only one strand of the target DNA. In other words, the programmed nuclease should preferably function as a nickase. Once a programmed nuclease is selected (e.g., a ZFN or TALEN), additional functions may be modified to enable it to operate according to a prime editing-like mechanism. For example, a programmed nuclease may be modified by coupling it with an RNA or DNA elongation arm (e.g., via a chemical linker), where the elongation arm includes a primer-binding site (PBS) and a DNA synthesis template. The programmed nuclease may also be coupled with a polymerase (e.g., via a chemical linker or an amino acid linker), the properties of which will depend on whether the elongation arm is DNA or RNA. In the case of an RNA elongation arm, the polymerase may be an RNA-dependent DNA polymerase (e.g., a reverse transcriptase). In the case of a DNA elongation arm, the polymerase may be a DNA-dependent DNA polymerase (e.g., a prokaryotic polymerase encompassing Pol I, Pol II, or Pol III, or a eukaryotic polymerase encompassing Pol a, Pol b, Pol g, Pol d, Pol e, or Pol z).The system may also include other functions that are added as a fusion with a programmed nuclease or added in trans to facilitate the overall reaction (e.g., (a) a helicase that unwinds the DNA at the cleavage site to create a cleavage strand with a 3' end that can be used as a primer, (b) a flap-end nuclease (e.g., FEN1) that removes the endogenous strand on the cleavage strand to help facilitate the reaction toward the replacement of the endogenous strand with the synthesized strand, or (c) an nCas9:gRNA complex that creates a second-site nick on the opposite strand (which may also help facilitate the incorporation of synthetic repair through favorable cellular repair of the unedited strand)). In a similar manner to priming editing with napDNAbp, such complexes with different programmed nucleases may be used to synthesize and then permanently incorporate the replacement strand of the newly synthesized DNA with the edit of interest into the target site of the DNA.

[0103] [Figure 1L]Figure 1L depicts, in one embodiment, the anatomical features of target DNA that may be edited by prime editing. The target DNA includes a “non-target strand” and a “target strand.” The target strand is the strand that will anneal with the PEgRNA spacer of the prime editing factor complex that recognizes the PAM site (in this case, NGG recognized by a standard SpCas9-based prime editing factor). The target strand may also be referred to as the “non-PAM strand” or “unedited strand.” In contrast, the non-target strand (i.e., the strand containing the protospacer and the PAM sequence of NGG) may also be referred to as the “PAM strand” or “edited strand.” In various embodiments, the nick site of the PE complex (e.g., in SpCas9-based PE) will be located in the protospacer on the PAM strand. The location of the nick will be a feature of the specific Cas9 that forms the PE. For example, in a SpCas9-based PE, the nick site is located in the phosphodiester bond between bases 3 (at position -3 relative to position 1 in the PAM sequence) and 4 (at position -4 relative to position 1 in the PAM sequence). The nick site in the protospacer forms a free 3' hydroxyl group that complexes with the primer binding site of the PEgRNA elongation arm, as shown in the figure below, providing a substrate to initiate the polymerization of a single strand of DNA encoding the DNA synthesis template of the PEgRNA elongation arm. This polymerization reaction is catalyzed in the 5'→3' direction by the polymerase of the PE fusion protein (e.g., reverse transcriptase). Polymerization is terminated before reaching the gRNA core (e.g., by a polymerization termination signal or inclusion of a secondary structure that functions to terminate the polymerization activity of the PE), producing a single-strand DNA flap extended from the original 3' hydroxyl group of the nicked PAM strand. The DNA synthesis template encodes a single strand of DNA homologous to the 5' end of the endogenous DNA immediately following the nicking site on the PAM strand, and incorporates the desired nucleotide changes (e.g., single base substitutions, insertions, deletions, inversions).The desired editing location can be any position downstream of the nick site on the PAM chain, such as positions +1, +2, +3, +4 (start of PAM site), +5 (PAM site position 2), +6 (PAM site position 3), +7, +8, +9, +10, +11, +12, +13, +14, +15, +16, +17, +18, +19, + 20, +21, +22, +23, +24, +25, +26, +27, +28, +29, +30, +31, +32, +33, +34, +35, +36, +37, +38, +39, +40, +41, +42, +43, +44, +45, +46, +47, +48, ​​+49, +50, +51, +52, +53, +54, +55 +56, +57, +58, +59, +60, +61, +62, +63, +64, +65, +66, +67, +68, +69, +70, +71, +72, +73, +74, +75, +76, +77, +78, +79, +80, +81, +82, +83, +84, +85, +86, +87, +88, +89, +90, +91, +92, +93, +94, +95, +96, +97, +98, +99, +100, +101, +102, +103, +104, +105, +106, +107, +108, +109, +110, +111, +112, +113, +114, +115, +116, +117, +118, +119, +120, + +121, +122, +123, +124, +125, +126, +127, +128, +129, +130, +131, +132, +133, +134, +135, +136, +137, +138, +139, +140, +141, +142, +143, +144, +145, +146, +147, +148, +149, or +150 or more (relative to the downstream position of the nick site). Once the 3'-end single strand DNA (containing the edit of interest) replaces the endogenous 5'-end single strand DNA, the DNA repair and replication processes will result in the permanent incorporation of the edit site on the PAM strand, followed by the correction of mismatches on the non-PAM strand present at the edit site. Thus, the edit will spread to both strands of DNA on the target DNA site. It should be understood that the references to "edited strands" and "unedited" strands are simply intended to accurately describe (delineate) the DNA strands involved in the PE mechanism.The "edited strand" is the strand that becomes edited only after the 5' end of the single-stranded DNA immediately downstream of the nick site is replaced with a synthesized 3' end of single-stranded DNA containing the desired edit. The "unedited" strand is the paired strand with the edited strand, but it too becomes edited (in particular the edit of interest) through repair and / or replication to become complementary to the edited strand.

[0104] [Figure 1M]Figure 1M illustrates the mechanism of prime editing, showing the anatomical features of the target DNA, the prime editing factor complex, and the interaction between PEgRNA and the target DNA. Firstly, the prime editing factor, which includes a fusion protein containing a polymerase (e.g., reverse transcriptase) and napDNAbp (e.g., SpCas9 nickase, e.g., SpCas9 with an inactivating mutation in the HNH nuclease domain (e.g., H840A) or an activating mutation in the RuvC nuclease domain (D10A)), is complexed with PEgRNA and DNA containing the target DNA to be edited. PEgRNA includes a spacer, a gRNA core (also known as the gRNA backbone or gRNA main chain) (which binds to the napDNAbp), and an elongation arm. The elongation arm can be at the 3' end, the 5' end, or anywhere else within the PEgRNA molecule. As shown, the elongation arm is at the 3' end of the PEgRNA. The elongation arm contains a primer-binding site and a DNA synthesis template (including both the edit of interest and the homologous region (i.e., the homologous arm)) that is homologous to the single-stranded DNA at the 5' end immediately following the nick site on the PAM strand in the 3'→5' direction. As shown, once a nick is introduced, which produces a free 3' hydroxyl group immediately upstream of the nick site, the region immediately upstream of the nick site on the PAM strand anneals with a complementary sequence at the 3' end of the elongation arm, referred to as the "primer-binding site," creating a short double-stranded region with an available 3' hydroxyl end, thereby forming a substrate for the polymerase of the prime editing factor complex. The polymerase (e.g., reverse transcriptase) then polymerizes the DNA strand from the 3' hydroxyl end to the end of the elongation arm. The sequence of the single-stranded DNA is encoded by the DNA synthesis template, which is the portion of the elongation arm (i.e., excluding the primer-binding site) that is "read" by the polymerase that synthesizes the new DNA. This polymerization effectively extends to the original 3' hydroxyl-terminus sequence of the initial nick site. The DNA synthesis template encodes a single strand of DNA that includes not only the desired edit but also a region homologous to the endogenous DNA single strand immediately downstream of the nick site on the PAM strand.Next, the single strand of DNA at the 3' end that is encoded (i.e., the 3' single-strand DNA flap) replaces the corresponding homologous single strand of endogenous 5' end DNA immediately downstream of the nick site on the PAM strand, forming a DNA intermediate with a 5' single-strand DNA flap, which is removed by the cell (e.g., by flap endonuclease). The 3' single-strand DNA flap, which anneals with the complement of the endogenous 5' single-strand DNA flap, is ligated with the endogenous strand after the 5' DNA flap has been removed. The desired edit in the 3' single-strand DNA flap, which has just been annealed and ligated, forms a mismatch with the complementary strand, and after DNA repair and / or a series of replications, the desired edit is permanently incorporated into both strands.

[0105] [Figure 2] Figure 2 shows three Cas complexes (SpCas9, SaCas9, and LbCas12a) that can be used in the prime editing factors described herein, and their PAM, gRNA, and DNA cleavage characteristics. The figure shows the design of the complexes involving SpCas9, SaCas9, and LbCas12a.

[0106] [Figure 3]Figures 3A–3F show designs of the manipulated 5' prime edit factor gRNA (Figure 3A), 3' prime edit factor gRNA (Figure 3B), and intramolecular extension (Figure 3C). The extended guide RNA (or extended gRNA) may also be referred to in this application as PEgRNA or "prime edit guide RNA." Figures 3D and 3E provide additional embodiments of the 3' and 5' prime edit factor gRNAs (PEgRNAs), respectively. Figure 3F illustrates the interaction between the 3'-end prime edit factor guide RNA and the target DNA sequence. The embodiments of Figures 3A–3C illustrate the arrangement of the reverse transcription template sequence (i.e., or more broadly, the DNA synthesis template, as indicated, since RT is the only type of polymerase that can be used in the context of prime edit factors), primer binding sites, and exemplary linker sequences in the extended portions of the 3', 5', and intramolecular versions, as well as the general arrangement of spacers and core regions. The disclosed prime editing process is not limited to these configurations of the extended guide RNA. An aspect of Figure 3D provides the structure of an exemplary PEgRNA intended in this application. The PEgRNA comprises three main component elements ordered in the 5'→3' direction: a spacer, a gRNA core, and an extension arm at the 3' end. The extension arm may be further divided in the 5'→3' direction into the following structural elements: a primer binding site (A), an editing template (B), and a homologous arm (C). In addition, the PEgRNA may contain an optional 3' end modification region (e1) and an optional 5' end modification region (e2). Furthermore, the PEgRNA may contain a transcription termination signal at the 3' end of the PEgRNA (not shown). These structural elements are further defined in this application. The illustration of the PEgRNA structure is not intended to be limiting and encompasses variations in the arrangement of elements. For example, the optional sequence modifications (e1) and (e2) may be located within or between any of the other regions shown. It is not limited to being located at the 3' and 5' ends.In some embodiments, PEgRNA may include secondary RNA structures such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., MS2 aptamers that recruit and bind the MS2cp protein). For example, such secondary structures may be located within spacers, gRNA cores, or elongation arms, particularly within the e1 and / or e2 modification regions. In addition to secondary RNA structures, PEgRNA may include chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modification regions), where "N" can be any nucleic acid base. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linkers may function to prevent reverse transcription of the sgRNA backbone or core. In addition, in certain embodiments (see, for example, Figure 72(c)), the elongation arms (3) may consist of RNA or DNA and / or contain one or more nucleic acid base analogs (e.g., this may add functionality such as temperature resilience). Furthermore, the orientation of the elongation arm (3) may be the natural 5'→3' direction, or it may be synthesized in the opposite direction, 3'→5' (relative to the orientation of the overall PEgRNA molecule). It should also be noted that those skilled in the art will have the ability to select a suitable DNA polymerase for use in prime editing, depending on the properties of the nucleic acid material of the elongation arm (i.e., DNA or RNA). This may be implemented either as a fusion with napDNAbp or as a separate part in trans, to synthesize a 3' single-stranded DNA flap encoded by a desired template encompassing the desired edit. For example, if the elongation arm is RNA, the DNA polymerase may be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the elongation arm is DNA, the DNA polymerase may be a DNA-dependent DNA polymerase. In various embodiments, the provision of DNA polymerase can be trans, for example, an RNA-protein recruitment domain (e.g., an MS2 hairpin incorporated on PEgRNA (e.g., in the e1 or e2 region or elsewhere) and an MS2cp protein fused to the DNA polymerase).This is achieved by using a primer binding site (which co-localizes the DNA polymerase to PEgRNA). It should also be noted that the primer binding site generally does not form part of the template used by the DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation “DNA synthesis template” refers to a region or portion of the elongation arm (3) used as a template by the DNA polymerase to encode the desired 3' single-stranded DNA flap containing a homologous region to the 5' endogenous single-stranded DNA flap replaced by the 3' single-stranded DNA product of the edit and primed DNA synthesis. In some embodiments, the DNA synthesis template includes an “edit template” and a “homologous arm” or one or more homologous arms, for example, before and after the edit template. The edit template can be as small as a single-nucleotide substitution, or it can be a DNA insertion or inversion. In addition, the edit template can also include a deletion, which can be manipulated by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template may also include the e2 region or a portion thereof. For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any portion of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of the e2 is actually used as a template will depend on its composition and whether that composition disrupts DNA polymerase function.

[0107] [Figure 3E]The aspect of Figure 3E provides a structure of another PEgRNA intended in this application. The PEgRNA comprises three main component elements ordered in the 5'→3' direction: a spacer, a gRNA core, and an elongation arm at the 3' end. The elongation arm may be further divided in the 5'→3' direction into the following structural elements: a primer binding site (A), an editing template (B), and a homologous arm (C). In addition, the PEgRNA may include an optional 3' end modification region (e1) and an optional 5' end modification region (e2). Furthermore, the PEgRNA may include a transcription termination signal at the 3' end (not shown). These structural elements are further defined in this application. The illustration of the PEgRNA structure is not intended to be limiting and encompasses variations in the arrangement of elements. For example, any sequence modifications (e1) and (e2) may be located within or between any of the other regions shown, and are not limited to being located at the 3' and 5' ends. In some embodiments, PEgRNA may include secondary RNA structures such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., MS2 aptamers that recruit and bind the MS2cp protein). These secondary structures can be located anywhere on the PEgRNA molecule. For example, such secondary structures may be located within spacers, the gRNA core, or the elongation arms, particularly within the e1 and / or e2 modification regions. In addition to secondary RNA structures, PEgRNA may include chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modification regions), where "N" can be any nucleic acid base. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linkers may function to prevent reverse transcription of the sgRNA backbone or core. In addition, in certain embodiments (see, for example, Figure 72(c)), the elongation arm (3) may consist of RNA or DNA and / or may contain one or more nucleic acid base analogs (for example, this may add functionality such as temperature resilience). Furthermore, the orientation of the elongation arm (3) may be the natural 5'→3' direction, or it may be synthesized in the opposite direction, 3'→5' (relative to the orientation of the overall PEgRNA molecule).It will also be noted that those skilled in the art will have the ability to select a suitable DNA polymerase for use in prime editing, depending on the properties of the nucleic acid material of the elongation arm (i.e., DNA or RNA). This can be implemented either as a fusion with napDNAbp or as a separate part in trans, to synthesize a 3' single-stranded DNA flap encoded by a desired template encompassing the desired edit. For example, if the elongation arm is RNA, the DNA polymerase may be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the elongation arm is DNA, the DNA polymerase may be a DNA-dependent DNA polymerase. In various embodiments, the provision of the DNA polymerase may be in trans, for example, by the use of an RNA-protein recruitment domain (e.g., an MS2 hairpin incorporated on PEgRNA (e.g., in or elsewhere in the e1 or e2 region) and an MS2cp protein fused to the DNA polymerase, thereby colocalizing the DNA polymerase to PEgRNA). It should also be noted that primer binding sites generally do not form part of the template used by a DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation “DNA synthesis template” refers to a region or portion of the elongation arm (3) used as a template by a DNA polymerase to encode the desired 3' single-stranded DNA flap, which contains a homologous region to the 5' endogenous single-stranded DNA flap replaced by the 3' single-stranded DNA product of the edit and prime-edited DNA synthesis. In some embodiments, the DNA synthesis template includes an “edit template” and a “homologous arm” or one or more homologous arms, for example, before and after the edit template. The edit template can be as small as a single-nucleotide substitution, or it may be a DNA insertion or inversion. In addition, the edit template may also include a deletion, which can be manipulated by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template may also include the e2 region or a portion thereof.For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any portion of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of e2 is actually used as a template will depend on its composition and whether that composition disrupts DNA polymerase function.

[0108] [Figure 3F]The schematic diagram in Figure 3F illustrates the typical interaction of PEgRNA with a target site on double-stranded DNA and the associated generation of a 3' single-stranded DNA flap containing the desired gene alteration. The double-stranded DNA is shown by an upper strand oriented 3'→5' (i.e., the target strand) and a lower strand oriented 5'→3' (i.e., the PAM strand or non-target strand). The upper strand, containing the complements of the "protospacer" and the PAM sequence, is called the "target strand" because it is the strand targeted by the PEgRNA spacer and anneals to it. The complementary lower strand is called the "non-target strand," "PAM strand," or "protospacer strand" because it contains the PAM sequence (e.g., NGG) and the protospacer. Although not shown, the illustrated PEgRNA would complex with the Cas9 or equivalent domain of a prime editing factor fusion protein. As shown in the schematic diagram, the PEgRNA spacer anneals to the complementary region of the protospacer on the target strand. This interaction forms a DNA / RNA hybrid between the spacer RNA and the complement of the protospacer DNA, inducing the formation of an R-loop on the protospacer. As taught elsewhere in this application, the Cas9 protein (not shown) then induces a nick on the non-target strand as shown. This then leads to the formation of a 3' ssDNA flap region immediately upstream of the nick site, which interacts with the 3' end of the PEgRNA at the primer binding site according to *z*. The 3' end of the ssDNA flap (i.e., the reverse transcriptase primer sequence) anneals to the primer binding site (A) on the PEgRNA, thereby priming the reverse transcriptase. Next, the reverse transcriptase (for example, provided in trans or cis as a fusion protein attached to a Cas9 construct) polymerizes a single strand of DNA encoded by the DNA synthesis template (which includes the editing template (B) and homologous arm (C)). Polymerization continues toward the 5' end of the elongation arm.The polymerized strands of ssDNA form an ssDNA 3' end flap, which, as described elsewhere (for example, as shown in Figure 1G), invades endogenous DNA, displacing the corresponding endogenous strand (which is removed as a DNA flap at the 5' end of the endogenous DNA), and incorporates the desired nucleotide edits (single nucleotide base pair changes, deletions, and insertions (including entire genes)) by the naturally occurring DNA repair / replication rounds.

[0109] [Figure 3G]Figure 3G illustrates yet another embodiment of prime editing as envisioned in this application. In particular, the upper schematic diagram illustrates one embodiment of a prime editing factor (PE). This comprises a fusion protein of napDNAbp (e.g., SpCas9) and polymerase (e.g., reverse transcriptase), which are linked by a linker. The PE forms a complex with PEgRNA by binding to the gRNA core of PEgRNA. In the embodiment shown, the PEgRNA has a 3' extension arm, which, starting at the 3' end, contains a primer-binding site (PBS), followed by a DNA synthesis template. The lower schematic diagram illustrates a variant of the prime editing factor called a “transprime editing factor (tPE)”. In this embodiment, the DNA synthesis template and PBS are decoupled from the PEgRNA and presented by a separate molecule called a transprime editing factor RNA template ("tPERT"). This contains an RNA-protein recruitment domain (e.g., MS2 hairpin). PE itself is further modified to include a fusion with the rPERT recruiting protein ("RP"), which is a protein that specifically recognizes and binds to the RNA-protein recruiting domain. In the example where the RNA-protein recruiting domain is the MS2 hairpin, the corresponding rPERT recruiting protein could be the MS2cp of the MS2 tagging system. The MS2 tagging system is based on the innate interaction of the MS2 bacteriophage coat protein ("MCP" or "MS2cp") with stem-loop or hairpin structures present on the phage genome, i.e., "MS2 hairpins" or "MS2 aptamers". In the case of transprime editing, the RP-PE:gRNA complex "recruits" a tPERT having a suitable RNA-protein recruiting domain to colocalize with the PE:gRNA complex, thereby providing trans-PBS and DNA synthesis templates for use in prime editing, as shown in the example illustrated in Figure 3H.

[0110] [Figure 3H]Figure 3H illustrates the process of transprime editing. In this embodiment, the transprime editing factor comprises a “PE2” prime editing factor (i.e., a fusion of Cas9(H840A) and variant MMLV RT) fused to an MS2cp protein (i.e., a recruiting protein of the type that recognizes and binds to the MS2 aptamer) and complexed with sgRNA (i.e., a standard guide RNA in contrast to PEgRNA). The transprime editing factor binds to target DNA and introduces a nick into the non-target strand. The MS2cp protein recruits tPERT trans-prime by specific interaction with the RNA-protein recruiting domain on the tPERT molecule. tPERT then co-localizes with the transprime editing factor, thereby trans-prime providing PBS and DNA synthesis template functions for use by reverse transcriptase polymerase to synthesize a single-stranded DNA flap having a 3' end and containing the desired genetic information encoded by the DNA synthesis template.

[0111] [Figure 4A] Figures 4A-4E demonstrate the in vitro TPRT assay (i.e., prime editing assay). Figure 4A shows a schematic diagram of the fluorescently labeled DNA substrate, gRNA template extension by the RT enzyme, and PAGE. [Figure 4B-C] Figure 4B shows TPRT (i.e., prime editing) with pre-nicked substrates, dCas9, and 5' elongated gRNAs of different synthetic template lengths. Figure 4C shows the RT reaction with pre-nicked DNA substrates in the absence of Cas9. [Figure 4D-E] Figure 4D shows TPRT (i.e., prime editing) of Cas9(H840A) and 5' extended gRNA to a full-length dsDNA substrate. Figure 4E shows the 3' extended gRNA template with a pre-nicked full-length dsDNA substrate. M-MLV RT is present in all reactions.

[0112] [Figure 5]Figure 5 shows in vitro validation results using 5' elongated gRNA with varying synthetic template lengths. A fluorescently labeled (Cy5) DNA target was used as the substrate and pre-nicked in this experimental setup. The Cas9 used in these experiments was the catalytically inactive Cas9 (dCas9), and the RT used was Superscript III, a commercial RT derived from Moloney's mouse leukemia virus (M-MLV). dCas9:gRNA complexes were formed from purified components. The fluorescently labeled DNA substrate was then added along with dNTPs and the RT enzyme. After incubation at 37°C for 1 hour, the reaction product was analyzed by denatured urea-polyacrylamide gel electrophoresis (PAGE). The gel image shows elongation to a length consistent with the original DNA strand ~ reverse transcription template length.

[0113] [Figure 6] Figure 6 shows the in vitro validation results using 5' elongated gRNA with varying synthetic template lengths, which are very similar to those shown in Figure 5. However, the DNA substrate in this experiment did not contain pre-mixed nicks. The Cas9 used in these experiments was Cas9 nickase (SpyCas9 H840A mutant), and the RT used was Superscript III, a commercial RT derived from Moloney's mouse leukemia virus (M-MLV). The reaction products were analyzed by denatured urea polyacrylamide gel electrophoresis (PAGE). As shown in the gel, nickase efficiently cleaves DNA strands when standard gRNA is used (gRNA_0, lane 3).

[0114] [Figure 7]Figure 7 demonstrates that 3' extension does not significantly induce Cas9 nickase activity, supporting DNA synthesis. Pre-nicked substrates (black arrows) are converted to RT products almost quantitatively when either dCas9 or Cas9 nickase is used (lanes 4 and 5). More than 50% conversion to RT products (red arrow) is observed for full-length substrates (lane 3). Cas9 nickase (SpyCas9 H840A mutant), catalytically inactive Cas9 (dCas9), and Superscript III (commercial RT derived from Moloney's mouse leukemia virus (M-MLV)) are used.

[0115] [Figure 8] Figure 8 demonstrates a dual-color experiment used to determine whether the RT reaction preferentially results in cis (binding within the same complex) to the gRNA. Two separate experiments were performed for 5'-extended gRNA and 3'-extended gRNA. Products were analyzed by PAGE. Product ratios were calculated as (Cy3cis / Cy3trans) / (Cy5trans / Cy5cis).

[0116] [Figure 9A-B] Figures 9A–9D demonstrate the flap model substrate. Figure 9A shows a dual FP reporter for flap-specific (-directed) mutagenesis. Figure 9B shows arrest codon repair in HEK cells. [Figure 9C-D] Figure 9C shows sequenced yeast clones after flap repair. Figure 9D shows testing of different flap characteristics in human cells.

[0117] [Figure 10]Figure 10 demonstrates prime editing on a plasmid substrate. A dual fluorescent reporter plasmid was constructed for expression in yeast (S. cerevisiae). Expression of this construct in yeast produces only GFP. In vitro prime editing induces point mutations and transforms yeast with the parent plasmid or a plasmid nicked in vitro with Cas9(H840A). Colonies are visualized by fluorescence imaging. Yeast dual FP plasmid transformants are shown. Transdevelopment of the parent plasmid or a plasmid nicked in vitro with Cas9(H840A) results in only green GFP-expressing colonies. Prime editing with 5'-extended gRNA or 3'-extended gRNA produces mixed green and yellow colonies. The latter expresses both GFP and mCherry. More yellow colonies are observed with 3'-extended gRNA. Positive controls without stop codons are also not shown.

[0118] [Figure 11] Figure 11 shows prime editing on a plasmid substrate similar to the experiment in Figure 10, but instead of incorporating a point mutation within a stop codon, prime editing incorporates a single nucleotide insertion (left) or deletion (right) that repairs a frameshift mutation, enabling downstream mCherry synthesis. Both experiments used 3' elongated gRNAs.

[0119] [Figure 12] Figure 12 shows the edit products of prime editing on plasmid substrates, characterized by Sanger sequencing. Colonies from TRT transformation were individually selected and analyzed by Sanger sequencing. Precise editing was observed by sequencing the selected colonies. Green colonies contained plasmids with the original DNA sequence, while yellow colonies contained precise mutations designed by the prime editing gRNA. No other point mutations or indels were observed.

[0120] [Figure 13]Figure 13 illustrates the potential scope of the new prime editing technology and shows a comparison with deaminase-mediated base editing technology.

[0121] [Figure 14] Figure 14 shows a schematic diagram of editing in human cells.

[0122] [Figure 15] Figure 15 demonstrates the elongation of the primer binding site in gRNA.

[0123] [Figure 16] Figure 16 shows gRNA truncated for adjacent targeting.

[0124] [Figure 17A-C] Figures 17A-17C are graphs showing the %T→A conversion at target nucleotides after transfection of components in human embryonic kidney (HEK) cells. Figure 17A shows data presenting the results using N-terminal fusion of wild-type MLV reverse transcriptase to Cas9(H840A) nickase (32-amino acid linker). Figure 17B is similar to Figure 17A, except for the C-terminal fusion of the RT enzyme. Figure 17C is similar to Figure 17A, but the linker between MLV RT and Cas9 is 60 amino acids long instead of 32.

[0125] [Figure 18] Figure 18 shows high-purity T→A editing at the HEK3 site by high-throughput amplicon sequencing. The sequencing analysis output displays the most abundant genotype of the edited cells.

[0126] [Figure 19] Figure 19 shows the indel ratio (blue bars) alongside the editing efficiency at the target nucleotide (orange bars). WT refers to the wild-type MLV RT enzyme. Mutant enzymes (M1-M4) contain the mutations listed on the right. The editing ratio was quantified by high-throughput sequencing of genomic DNA amplicons.

[0127] [Figure 20] Figure 20 shows the editing efficiency of a target nucleotide when a single-strand nick is introduced into a complementary DNA strand adjacent to the target nucleotide. We tested introducing the nick at various intervals from the target nucleotide (triangles). The editing efficiency at the target base pair (blue bars) is shown alongside the indel formation rate (orange bars). The "none" example does not contain guide RNA for introducing the nick into the complementary strand. The editing rate was quantified by high-throughput sequencing of genomic DNA amplicons.

[0128] [Figure 21] Figure 21 demonstrates processed high-throughput sequencing data showing the overall absence of desired T→A transversion mutations and other major genome editing byproducts.

[0129] [Figure 22]Figure 22 provides a schematic diagram of an exemplary process for performing targeted mutagenesis on a target locus (i.e., prime editing with an error-prone reverse transcriptase) using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an extended guide RNA. This process is sometimes referred to as a form of prime editing for targeted mutagenesis. The extended guide RNA involves extension at the 3' or 5' end of the guide RNA, or at some intramolecular location within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts a DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus to be mutageneised. In step (b), a nick is introduced into one strand of the DNA at the target locus (e.g., by a nuclease or chemical agent), thereby creating a usable 3' end on one strand of the DNA at the target locus. In one embodiment, the nick is created in the strand of DNA corresponding to the R loop, i.e., the strand that is not hybridized with the guide RNA sequence. In step (c), the 3'-terminus DNA strand interacts with the elongation region of the guide RNA to prime the reverse transcription. In one embodiment, the 3'-terminus DNA strand hybridizes with a specific RT prime sequence on the elongation region of the guide RNA. In step (d), an error-prone reverse transcriptase is introduced, which synthesizes a single strand of mutagenic DNA from the 3'-terminus of the primed site to the 3'-terminus of the guide RNA. Exemplary mutations are indicated by an asterisk "*". This forms a single-strand DNA flap containing the desired mutagenic region. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-strand DNA flap (containing the mutagenic region) so that the desired mutagenic region is incorporated into the target locus. This process can be driven to the formation of the desired product by removing the corresponding 5' endogenous DNA flap once the 3' single-strand DNA flap has entered and hybridized with a complementary sequence on the other strand. The process can also be driven to the formation of a product in which nicks are introduced into the second chain, as illustrated in Figure 1F.Following the endogenous DNA repair and / or replication process, the mutagenic region becomes integrated into both strands of DNA at the DNA locus.

[0130] [Figure 23] Figure 23 is a schematic diagram of trinucleotide repeat reduction in gRNA design and TPRT genome editing (i.e., prime editing) for reducing trinucleotide repeat sequences. Trinucleotide repeat expansion is associated with numerous human diseases, including Huntington's disease, fragile X syndrome, and Friedreich's ataxia. The most common trinucleotide repeat contains the CAG triplet, but GAA triplets (Friedreich's ataxia) and CGG triplets (Fragile X syndrome) also occur. Inheriting a predisposition to expansion or acquiring an already expanded parental allele increases the likelihood of developing the disease. The pathogenic expansion of trinucleotide repeats can hypothetically be corrected using prime editing. The region upstream of the repeat region can be nicked by an RNA-guided nuclease and then used to prime the synthesis of a new DNA strand containing a healthy number of repeats (which depends on the specific gene and disease). Following the repeat sequence, a series of short homologous sequences (red strands) are added that match the identity of the adjacent sequence at the other end of the repeat. The intrusion of the newly synthesized strand, followed by the replacement of the newly synthesized flap with endogenous DNA, leads to the reduced repeat allele.

[0131] [Figure 24] Figure 24 is a schematic diagram showing the precise 10-nucleotide deletion in prime editing. The guide RNA targeted by the HEK3 locus was designed with a reverse transcription template encoding a 10-nucleotide deletion after the nick site. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing.

[0132] [Figure 25]Figure 25 is a schematic diagram illustrating gRNA design for peptide tagging of genes at endogenous genomic loci and peptide tagging in TPRT genome editing (i.e., prime editing). The FlAsH and ReAsH tagging systems contain a genetically encoded peptide comprising two parts: (1) a fluorophore-biarsenic probe and (2) a tetracysteine ​​motif, exemplified by the sequence FLNCCPGCCMEP (SEQ ID NO: 1). When expressed in cells, the tetracysteine ​​motif-containing protein can be fluorescently labeled with the fluorophore-biarsenic probe (see reference: J.Am.Chem.Soc., 2002, 124(21), pp6063-6076; DOI: 10.1021 / ja017687n). The "sortagging" system employs a bacterial saltase enzyme to covalently conjugate a labeled peptide probe to a protein containing a suitable peptide substrate (see reference: Nat.Chem.Biol.2007 Nov;3(11):707-8. DOI:10.1038 / nchembio.2007.31). FLAG tags (DYKDDDDK (SEQ ID NO: 2)), V5 tags (GKPIPNPLLGLDST (SEQ ID NO: 3)), GCN4 tags (EELLSKNYHLENEVARLKK (SEQ ID NO: 4)), HA tags (YPYDVPDYA (SEQ ID NO: 5)), and Myc tags (EQKLISEEDL (SEQ ID NO: 6)) are commonly used as epitope tags for immunoassays. The pi-clamp encodes a peptide sequence (FCPF (SEQ ID NO: 622)) that can be labeled with a pentafluoro-aromatic substrate (Reference: Nat. Chem. 2016 Feb;8(2):120-8. doi:10.1038 / nchem.2413).

[0133] [Figure 26A]Figure 26A shows the precise insertion of His6 and FLAG tags into genomic DNA. Guide RNAs targeting the HEK3 locus were designed using reverse transcription templates encoding either an 18nt His tag insertion or a 24nt FLAG tag insertion. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing. Note that the full-length 24nt FLAG tag sequence is cut off from the frame (sequencing confirmed the precise insertion of the full length). [Figure 26B] Figure 26B is a schematic diagram summarizing the various applications of protein / peptide tagging, which include (a) solubilizing or insolubilizing proteins, (b) altering or tracking the intracellular localization of proteins, (c) extending the half-life of proteins, (d) facilitating protein purification, and (e) facilitating protein detection.

[0134] [Figure 27] Figure 27 shows an overview of prime editing by incorporating protective mutations into PRNPs that prevent or halt the progression of prion diseases. The PEgRNA sequences on the left correspond to residues 1-20 of SEQ ID NO: 810 (i.e., 5' of the sgRNA backbone), and on the right correspond to residues 21-43 of SEQ ID NO: 810 (i.e., 3' of the sgRNA backbone).

[0135] [Figure 28A] Figure 28A is a schematic diagram of PE-based insertions of sequences encoding RNA motifs. [Figure 28B] Figure 28B is a list (not exhaustive) of some example motifs that could potentially be inserted, and their functions.

[0136] [Figure 29]Figure 29A illustrates the prime editing factor. Figure 29B shows possible modifications to genome, plasmid, or viral DNA guided by PE. Figure 29C shows an example scheme for inserting a peptide loop library onto a defined protein (GFP in this case) using a library of PEgRNAs. Figure 29D shows an example of possible programmable deletion or N- or C-terminal shortening of protein codons using different PEgRNAs. Deletions are expected to occur with minimal generation of frameshift mutations.

[0137] [Figure 30] Figure 30 shows possible schemes for repetitive codon insertion in continuous evolutionary systems such as PACE.

[0138] [Figure 31] Figure 31 shows a description of the manipulated gRNA. It shows the gRNA core, a ~20nt spacer that matches the sequence of the target gene, a reverse transcription template with an immunogenic epitope nucleotide sequence, and a primer binding site that matches the sequence of the target gene.

[0139] [Figure 32] Figure 32 is a schematic diagram illustrating the use of prime editing as a means of inserting known immunogenic epitopes into endogenous or exogenous genomic DNA, resulting in modification of the corresponding proteins.

[0140] [Figure 33]Figure 33 is a schematic diagram illustrating PEgRNA design for primer-binding sequence insertion and primer-binding insertion into genomic DNA using prime editing to determine off-target editing. In this embodiment, prime editing is performed in living cells, tissues, or animal models. As a first step, a suitable PEgRNA is designed. The upper schematic diagram shows an exemplary PEgRNA that may be used in this aspect. A spacer on the PEgRNA (labeled "protospacer") is complementary to one of the strands of the genomic target. The PE:PEgRNA complex (i.e., the PE complex) incorporates a single-stranded 3' end flap into the nick site, which contains the encoded primer-binding sequence and a homologous region (encoded by the homologous arm of the PEgRNA) that is complementary to the region just downstream of the cut site (indicated in red). Through flap insertion and DNA repair / replication processes, the synthesized strand is incorporated into the DNA, thereby incorporating the primer-binding site. This process can occur not only at the desired genomic target but also at other genomic sites that may interact with PEgRNA in an off-target manner (i.e., PEgRNA guides the PE complex to other off-target sites due to the complementarity of the spacer region to other genomic sites that are not the intended genomic site). Therefore, primer-binding sequences can be incorporated not only at the desired genomic target but also at other off-target genomic sites elsewhere on the genome. To detect the insertion of these primer-binding sequences at both the intended genomic target site and the off-target genomic site, genomic DNA (post-PE) can be isolated, fragmented, and ligated into adapter nucleotides (shown in red). Next, PCR can be performed using PCR oligonucleotides that anneal to the adapter and the inserted primer-binding sequence to amplify the on-target and off-target genomic DNA regions into which the primer-binding sequence was inserted by PE. High-throughput sequencing can then be performed to align the sequences and identify the insertion sites of the primer-binding sequences inserted by PE at either the on-target or off-target site.

[0141] [Figure 34] Figure 34 is a schematic diagram illustrating the precise insertion of genes by PE.

[0142] [Figure 35] Figure 35A is a schematic diagram showing the innate insulin signaling pathway. Figure 35B is a schematic diagram showing FKBP12-tagged insulin receptor activation controlled by FK1012.

[0143] [Figure 36] Figure 36 shows low molecular weight monomers. Reference: Bump FK506 Mimic(2)107.

[0144] [Figure 37] Figures 37A-37B show the low molecular weight dimers. References: FK1012 495,96; FK1012 5108; FK1012 6107; AP1903 7107; Cyclosporine A dimer 898; FK506-Cyclosporine A dimer (FkCsA) 9100.

[0145] [Figure 38]Figures 38A–38F provide an overview of prime editing and feasibility studies in vitro and in yeast cells. Figure 38A shows 75,122 known pathogenic human gene variants from ClinVar (accessed July 2019), classified by type. Figure 38B shows that the prime editing complex consists of a prime editing factor (PE) protein complexed with prime editing guide RNA (PEgRNA) and containing a DNA nicking domain, e.g., Cas9 nickase, which is guided by RNA fused to an engineered reverse transcriptase domain. The PE:PEgRNA complex binds to a target DNA site, enabling large and varied precision DNA editing at a wide range of DNA locations before and after the protospacer adjacent motif (PAM) of the target site. Figure 38C shows that DNA target binding causes the PE:PEgRNA complex to nick the PAM-containing DNA strand. The resulting free 3' end hybridizes to the primer-binding site of the PEgRNA. The reverse transcriptase domain catalyzes primer extension using the PEgRNA RT template, resulting in a newly synthesized DNA strand (3' flap) containing the desired edits. Equilibrium between the edited 3' flap and the unedited 5' flap containing the original DNA, followed by cellular 5' flap cleavage and ligation, and DNA repair or replication to degrade the heterodouble-stranded DNA, results in stably edited DNA. Figure 38D shows an in vitro 5'-extended PEgRNA primer extension assay using a pre-nicked dsDNA substrate containing a 5'Cy5-labeled PAM strand, dCas9, and a commercially available M-MLV RT variant (RT, Superscript III). dCas9 was complexed with PEgRNAs containing RT templates of various lengths, and then added to the DNA substrate along with the indicated components. The reaction was incubated at 37°C for 1 hour and then analyzed by denatured urea PAGE, visualized for Cy5 fluorescence.Figure 38E shows primer extension performed as shown in Figure 38D using 3'-extended PEgRNA pre-complexed with dCas9 or Cas9 H840A nickase and a pre-nicked or unnicked 5'Cy5-labeled dsDNA substrate. Figure 38F shows yeast colonies transformed with PEGRNA, Cas9 nickase, and a GFP-mCherry fusion reporter plasmid edited in vitro by RT. Plasmids containing nonsense or frameshift mutations between GFP and mCherry were edited with 5'-extended or 3'-extended PEgRNA that restored mCherry translation by transversion mutation, 1 bp insertion, or 1 bp deletion. Cells double-positive for GFP and mCherry (yellow) reflect successful editing.

[0146] [Figure 39]Figures 39A–39D show prime editing of human genomic DNA by PE1 and PE2. Figure 39A shows that PEgRNA contains a spacer sequence, an sgRNA backbone, and a 3' extension containing a reverse transcription (RT) template (purple) with a primer binding site (green) and the base(s) to be edited (single or plural) (red). The primer binding site hybridizes to the PAM-containing DNA strand immediately upstream of the nicking site. With the exception of the encoded edit, the RT template is homologous to the DNA sequence downstream of the nicking. Figure 39B shows the incorporation of T·A→A·T transversion editing at the HEK3 site in HEK293T cells using Cas9 H840A nickase (PE1) fused to wild-type M-MLV reverse transcriptase and PEgRNAs of various primer binding site lengths. Figure 39C shows that the use of engineered quintuplet mutant M-MLV reverse transcriptases (D200N, L603W, T306K, W313F, T330P) to PE2 substantially improves prime editing transversion efficiency at five genomic sites in HEK293T cells and small insertion and small deletion editing in HEK3. Figure 39D compares PE2 editing efficiency at five genomic sites in HEK293T cells with various RT template lengths. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0147] [Figure 40]Figures 40A–40C demonstrate that the PE3 and PE3b systems increase prime editing efficiency by nicking the unedited strand. Figure 40A provides an overview of prime editing by PE3. After the initial synthesis of the edited strand, DNA repair will remove either the newly synthesized strand containing the edit (3' flap excision) or the original genomic DNA strand (5' flap excision). The 5' flap excision leaves behind a DNA heteroduplex containing one edited strand and one unedited strand. Mismatch repair mechanisms or DNA replication may degrade the heteroduplex, yielding either an edited or unedited product. Nicking the unedited strand favors the repair of that strand, resulting in the preferred generation of a stable double-strand DNA containing the desired edit. Figure 40B shows the effect of complementary strand nicking on PE3-mediated prime editing efficiency and indel formation. "None" refers to the PE2 control, which does not nick the complementary strand. Figure 40C compares editing efficiency with PE2 (no complementary strand nick), PE3 (general complementary strand nick), and PE3b (edit-specific complementary strand nick). All edit yields reflect the percentage of total sequencing reads containing the intended edit and free of indels among all treated cells without sorting. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0148] [Figure 41]Figures 41A–41K show PE3-targeted insertions, deletions, and all 12 types of point mutations at seven endogenous human genome loci in HEK293T cells. Figure 41A is a graph showing all 12 types of 1-nucleotide transitions and transversion edits at HEK3 sites from position +1 to +8 (counting PEgRNA-induced nick positions between position +1 and -1) using a 10nt RT template. Figure 41B is a graph showing long-range PE3 transversion edits at HEK3 sites using a 34nt RT template. Figures 41C–41H are graphs showing all 12 types of transitions and transversion edits at various positions on the prime editing window for (Figure 41C) RNF2, (Figure 41D) FANCF, (Figure 41E) EMX1, (Figure 41F) RUNX1, (Figure 41G) VEGFA, and (Figure 41H) DNMT1. Figure 41I is a graph showing PE3-targeted 1 and 3 bp insertions and 1 and 3 bp deletions at seven endogenous genomic loci. Figure 41J is a graph showing targeted, precise deletions of 5–80 bp at HEK3 target sites. Figure 41K is a graph showing combined edits of insertions and deletions, insertions and point mutations, deletions and point mutations, and double point mutations at three endogenous genomic loci. All edit yields reflect the percentage of total sequencing reads containing the intended edits and no indels among all treated cells without sorting. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0149] [Figure 42]Figures 42A–42H show a comparison of prime editing, base editing, and off-target editing by Cas9 and PE3 at known Cas9 off-target sites. Figure 42A shows the total editing efficiency of C·G→T·A at the same target nucleotide for PE2, PE3, BE2max, and BE4max in endogenous HEK3, FANCF, and EMX1 sites of HEK293 T cells. Figure 42B shows the indel frequency from the treatment in Figure 42A. Figure 42C shows the editing efficiency of precise C·G→T·A editing (bystander editing or no indel) for PE2, PE3, BE2max, and BE4max in HEK3, FANCF, and EMX1. In EMX1, precise PE combination editing of all possible combinations of C·G→T·A conversion at three target nucleotides is also shown. Figure 42D shows the total A·T→G·C editing efficiency for PE2, PE3, ABEdmax, and ABEmax in HEK3 and FANCF. Figure 42E shows the precise A·T→G·C editing efficiency without bystander editing or indels for HEK3 and FANCF. Figure 42F shows the indel frequency from the treatment in Figure 42D. Figure 42G shows the mean triplicate editing efficiency (percentage sequencing reads with indels) in HEK293T cells for Cas9 nuclease at four on-target and 16 known off-target sites. The 16 off-target sites examined were the top four previously reported off-target sites for each of the four on-target sites: 118,159. For each on-target site, Cas9 was paired with either an sgRNA or each of four PEgRNAs that recognize the same protospacer. Figure 42H shows the mean triplicate on-target and off-target editing efficiencies and indel efficiencies (in parentheses below) in HEK293T cells for PE2 or PE3 paired with each PEgRNA (Figure 42G). On-target editing yield reflects the percentage of total sequencing reads that contain the intended edit but do not contain indels, out of all treated cells without sorting. Off-target editing yield reflects off-target locus modifications consistent with prime editing.The values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0150] [Figure 43]Figures 43A–43I show prime editing, pathogenic transversion, insertion or deletion mutation incorporation and modification, as well as a comparison of prime editing and HDR in various human cell lines and primary mouse cortical neurons. Figure 43A is a graph showing incorporation (by T·A→A·T transversion) and modification (by A·T→T·A transversion) of the pathogenic E6V mutation in the HBB of HEK293T cells. Modification to either wild-type HBB or HBB containing a silent mutation that blocks PEgRNA PAM is shown. Figure 43B is a graph showing incorporation (by 4bp insertion) and modification (by 4bp deletion) of the pathogenic HEXA 1278+TATC allele in HEK293T cells. Modification to either wild-type HEXA or HEXA containing a silent mutation that blocks PEgRNA PAM is shown. Figure 43C is a graph showing the incorporation of the protective G127V variant of PRNP in HEK293T cells via G·C→T·A transversion. Figure 43D is a graph showing prime editing in other human cell lines, including K562 (leukemia myeloid cells), U2OS (osteosarcoma cells), and HeLa (cervical cancer cells). Figure 43E is a graph showing the incorporation of the G·C→T·A transversion mutation in DNMT1 of mouse primary cortical neurons, using a binary mitotic intein PE3 lentiviral system. In this system, the N-terminal half is Cas9(1-573) fused to the N-intine and GFP-KASH via a P2A autocleavage peptide, and the C-terminal half is the C-intine fused to the remainder of PE2. The PE2 halves are expressed from a human synapsin promoter that is highly specific to mature neurons. Sorted values ​​reflect edits or indels from GFP-positive nuclei, while unsorted values ​​are from all nuclei. Figure 43F compares the HDR editing efficiencies mediated by PE3 and Cas9 at endogenous genomic loci in HEK293T cells. Figure 43G compares the HDR editing efficiencies mediated by PE3 and Cas9 at endogenous genomic loci in K562, U2OS, and HeLa cells.Figure 43H compares PE3 and Cas9-mediated HDR indel byproduct generation in HEK293T, K562, U2OS, and HeLa cells. Figure 43I shows targeted insertions of His6 tags (18 bp), FLAG epitope tags (24 bp), or extended LoxP sites (44 bp) in HEK293T cells by PE3. All edit yields reflect the percentage of total sequencing reads out of all treated cells that contain the intended edits but do not contain indels. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicates.

[0151] [Figure 44]Figures 44A–44G show in vitro prime editing validation studies using fluorescently labeled DNA substrates. Figure 44A shows electrophoretic mobility shift assays with dCas9, 5'-extended PEgRNA, and 5'-Cy5 labeled DNA substrate. PEgRNAs 1–5 contain a 15nt linker sequence (linker A in PEgRNA 1, linker B in PEgRNAs 2–5), a 5nt PBS sequence, and RT templates of 7nt (PEgRNAs 1 and 2), 8nt (PEgRNA 3), 15nt (PEgRNA 4), and 22nt (PEgRNA 5) between the spacer and PBS. The PEgRNAs used are those shown in Figures 44E and 44F; complete sequences are listed in Tables 2A–2C. Figure 44B shows an in vitro nicking assay of Cas9 H840A using 5'-extended and 3'-extended PEgRNAs. Figure 44C shows Cas9-mediated indel formation in HEK293T cells in HEK3 using 5'-extended and 3'-extended PEgRNA. Figure 44D shows an overview of the prime editing in vitro biochemical assay. 5'-Cy5-labeled pre-nicked and unnicked dsDNA substrates were tested. sgRNA, 5'-extended PEgRNA, or 3'-extended PEgRNA were pre-complexed with dCas9 or Cas9 H840A nickase and then combined with dsDNA substrate, M-MLV RT, and dNTPs. The reaction was allowed to proceed for 1 hour at 37°C prior to separation by denatured urea PAGE and visualization by Cy5 fluorescence. Figure 44E shows that the primer extension reaction using 5'-extended PEgRNA, pre-nicked DNA substrate, and dCas9 leads to significant conversion to the RT product. Figure 44F shows the primer extension reaction using an unnicked DNA substrate and Cas9 H840A nickase, and 5'-extended PEgRNA as shown in Figure 44B. Product yield is greatly reduced compared to a pre-nicked substrate. Figure 44G shows, by denatured urea PAGE, that the in vitro primer extension reaction using 3'-PEgRNA produces a single, clear product.The RT product band was excised, eluted from the gel, and then subjected to homopolymer tailing with terminal transferase (TdT) using either dGTP or dATP. The tailed product was extended with poly-T or poly-C primers, and the resulting DNA was sequenced. Sanger traces indicate that three nucleotides derived from the gRNA backbone were reverse transcribed (added to the DNA product as the last 3' nucleotide). Note that PEgRNA backbone insertions are considerably rarer in mammalian cell prime editing experiments than in vitro (Figures 56A-56D). Possible causes include the inability of the tethered reverse transcriptase to access the Cas9-bound guide RNA backbone, and / or cellular excision of the 3' end of a mismatch in the 3' flap containing the PEgRNA backbone sequence.

[0152] [Figure 45]Figures 45A–45G show cellular repair of 3' DNA flaps in yeast from in vitro prime-edit reactions. Figure 45A shows that a binary fluorescent protein reporter plasmid contains GFP and mCherry open reading frames separated by target sites encoding in-frame stop codons, +1 frameshifts, or -1 frameshifts. Prime-edit reactions were performed in vitro with Cas9 H840A nickase, PEgRNA, dNTPs, and M-MLV reverse transcriptase, and then the cells were transformed into yeast. Colonies containing the unedited plasmid produce GFP but not mCherry. Yeast colonies containing the edited plasmid produce both GFP and mCherry as fusion proteins. Figure 45B shows superposition of GFP and mCherry fluorescence in yeast colonies transformed with reporter plasmids containing a stop codon between GFP and mCherry (unedited negative control; top) or without a stop codon or frameshift between GFP and mCherry (pre-edited positive control; bottom). Figures 45C–45F show visualizations of mCherry and GFP fluorescence from yeast colonies transformed by in vitro prime editing reaction products. Figure 45C shows stop codon correction by T·A→A·T transversion using 3'-extended or 5'-extended PEgRNA, as shown in Figure 45D. Figure 45E shows +1 frameshift correction of 1 bp deletion using 3'-extended PEgRNA. Figure 45F shows -1 frameshift correction by 1 bp insertion using 3'-extended PEgRNA. Figure 45G shows Sanger DNA sequencing traces from plasmids isolated from GFP-only colonies in Figure 45B and GFP and mCherry double-positive colonies in Figure 45C.

[0153] [Figure 46]Figures 46A-46F show correct editing versus indel generation by PE1. Figure 46A shows the efficiency of T·A→A·T transversion editing and indel generation by PE1 at the +1 position of HEK3, using PEgRNA containing a 10nt RT template and a PBS sequence in the range of 8-17nt. Figure 46B shows the efficiency of G·C→T·A transversion editing and indel generation by PE1 at the +5 position of EMX1, using PEgRNA containing a 13nt RT template and a PBS sequence in the range of 9-17nt. Figure 46C shows the efficiency of G·C→T·A transversion editing and indel generation by PE1 at the +5 position of FANCF, using PEgRNA containing a 17nt RT template and a PBS sequence in the range of 8-17nt. Figure 46D shows the efficiency of C·G→A·T transversion editing and indel generation by PE1 at the +1 position of RNF2, using PEgRNA containing an 11nt RT template and a PBS sequence in the range of 9–17nt. Figure 46E shows the efficiency of G·C→T·A transversion editing and indel generation by PE1 at the +2 position of HEK4, using PEgRNA containing a 13nt RT template and a PBS sequence in the range of 7–15nt. Figure 46F shows PE1-mediated +1T deletion, +1A insertion, and +1CTT insertion at the HEK3 site, using 13nt PBS and 10nt RT templates. The PEgRNA sequences used are those used in Figure 39C (see Tables 3A–3R). Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0154] [Figure 47]Figures 47A-47S show the evaluation of M-MLV RT variants for prime editing. Figure 47A shows the abbreviations for the prime editing factor variants used in this figure. Figure 47B shows targeted insertion and deletion editing by PE1 at the HEK3 locus. Figures 47C–47H show a comparison of 18 prime editing factor constructs containing the M-MLV RT variant in their ability to incorporate the following edits: +2G·C→C·G in HEK3 (as shown in Figure 47C), 24bp FLAG insertion in HEK3 (as shown in Figure 47D), +1C·G→A·T in RNF2 (as shown in Figure 47E), +1G·C→C·G in EMX1 (as shown in Figure 47F), +2T·A→A·T in HBB (as shown in Figure 47G), and +1G·C→C·G in FANCF (as shown in Figure 47H). Figures 47I–47N show a comparison of four prime editing factor constructs containing the M-MLV variant in their ability to incorporate the edits shown in Figures 47C–47H in a second round of independent experiments. Figures 47O–47S show the PE2 editing efficiency at five genomic loci with varying PBS lengths. Figure 47O shows the +1T·A→A·T variation in HEK3. Figure 47P shows the +5G·C→T·A variation in EMX1. Figure 47Q shows the +5G·C→T·A variation in FANCF. Figure 47R shows the +1C·G→A·T variation in RNF2. Figure 47S shows the +2G·C→T·A variation in HEK4. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0155] [Figure 48]Figures 48A–48C illustrate the design features of PEgRNA PBS and RT template sequences. Figure 48A shows the efficiency of PE2-mediated +5G·C→T·A transversion editing in VEGFA of HEK293T cells (blue line) as a function of RT template length. Indels (gray line) are plotted for comparison. The sequences below the graph show the last nucleotide for editing that serves as a template for synthesis by PEgRNA. G nucleotides (using C on PEgRNA as a template) are highlighted; to maximize prime editing efficiency, RT templates ending in C should be avoided during PEgRNA design. Figure 48B shows the +5G·C→T·A transversion editing and indels of DNMT1 as in Figure 48A. Figure 48C shows the +5G·C→T·A transversion editing and indels of RUNX1 as in Figure 48A. Values ​​and error bars reflect the mean and sd of three independent biological replicates.

[0156] [Figure 49]Figures 49A–49B show the effects of PE2, PE2 R110S K103L, Cas9 H840A nickase, and dCas9 on cell viability. HEK293T cells were transfected with plasmids encoding PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9, along with a HEK3-targeted PEgRNA plasmid. Cell viability was measured every 24 hours and over 3 days post-transfection using the CellTiter-Glo2.0 assay (Promega). Figure 49A shows viability measured by luminescence at 1, 2, or 3 days post-transfection. Values ​​and error bars reflect the mean and sem of three independent biological replicas performed using technical triplicates, respectively. Figure 49B shows the percentage edits and indels for PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9, along with the HEK3-targeted PEgRNA plasmid encoding +5G→A editing. Editing efficiency was measured from treated cells on day 3 post-transfection, alongside those used to assay viability in Figure 49A. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0157] [Figure 50]Figures 50A–50B show PE3-mediated HBB E6V and HEXA 1278+TATC modification by various PEgRNAs. Figure 50A shows a screen of 14 PEgRNAs for modification of the HBB E6V allele in HEK293T cells by PE3. All PEgRNAs evaluated convert the HBB E6V allele back to wild-type HBB without introducing any silent PAM mutations. Figure 50B shows a screen of 41 PEgRNAs for modification of the HEXA 1278+TATC allele in HEK293T cells by PE3 or PE3b. HEXA-labeled PEgRNAs modify the pathogenic allele by a shifted 4bp deletion that blocks PAM and leaves a silent mutation. HEXA-labeled PEgRNAs modify the pathogenic allele back to wild-type. Entries ending in "b" use an editing-specific nicking sgRNA in combination with PEgRNA (PE3b system). The values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0158] [Figure 51]Figures 51A–51G show PE3 activity and a comparison of PE3 and Cas9-initiated HDRs in human cell lines. Figures 51A (HEK293T cells), 51B (K562 cells), 51C (U2OS cells), and 51D (HeLa cells) show the efficiency of generating correct editing (no indels) and indel frequencies for PE3 and Cas9-initiated HDRs. Each grouped editing comparison incorporates the same editing by PE3 and Cas9-initiated HDRs. The untargeted controls are PE3 and PEgRNAs targeting non-target loci. Figure 51E shows control experiments with untargeted PEgRNA + PE3 and dCas9 + sgRNA compared to wild-type Cas9 HDR experiments. We have confirmed that common contaminant ssDNA donor HDR templates, which artificially increase apparent HDR efficiency, do not contribute to the HDR measurements in Figures 51A–51D. Figures 51F–51G show example HEK3 site allele tables from genomic DNA samples isolated from K562 cells after editing with HDR initiated by PE3 or Cas9. Alleles were sequenced by Illumina MiSeq and analyzed by CRISPResso2.178 The reference HEK3 sequence from this region is shown above. The allele tables are shown for a negative control of untargeted PEgRNA, +1CTT insertions in HEK3 using PE3, and +1CTT insertions in HEK3 using HDR initiated by Cas9. Allele frequencies and corresponding Illumina sequencing read counts are shown for each allele. All alleles observed with frequencies ≥0.20% are shown. The values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0159] [Figure 52]Figures 52A–52D show the distribution of pathogenic insertions, duplications, deletions, and indel lengths in the ClinVar database. The ClinVar variant summary was downloaded from NCBI on July 15, 2019. The lengths of reported insertions, deletions, and duplications were calculated using appropriate identifying information such as reference and alternative alleles, variant start and stop positions, or variant names. Variants that did not report any of the above information were excluded from the analysis. The length of reported indels (single variants containing both insertions and deletions relative to the reference genome) was calculated by determining the number of mismatches or gaps in the best pairwise alignment between the reference and alternative alleles.

[0160] [Figure 53]Figures 53A–53E show examples of FACS gating for GFP-positive cell sorting. Below is an example of the original batch analysis file, outlining the sorting strategies used to generate HEXA 1278+TATC and HBB E6V HEK293T cell lines. Image data were generated by a Sony LE-MA900 cytometer using Cell Sorter software v.3.0.5. Graphic 1 shows a gating plot of non-GFP-expressing cells. Graphic 2 shows an example of sorting of P2A-GFP-expressing cells used to isolate the HBB E6V HEK293T cell line. HEK293T cells were initially gated collectively using FSC-A / BSC-A (gate A), and then sorted by singlet using FSC-A / FSC-H (gate B). Live cells were sorted by gating DAPI-negative cells (gate C). Cells exhibiting GFP fluorescence levels higher than those of negative control cells were sorted using EGFP as the fluorochrome (gate D). Figure 53A shows HEK293T cells (GFP-negative). Figure 53B shows a representative plot of FACS gating for cells expressing PE2-P2A-GFP. Figure 53C shows the genotype of HEXA 1278+TATC homozygous HEK293T cells. Figures 53D-53E show the allele table for the HBB E6V homozygous HEK293T cell line.

[0161] [Figure 54] Figure 54 is a schematic diagram summarizing the PEgRNA cloning procedure.

[0162] [Figure 55]Figures 55A–55G are schematic diagrams of PEgRNA designs. Figure 55A shows a simple diagram of PEgRNA with a labeled domain (left) and bound to nCas9 at a genomic site (right). Figure 55B shows various types of modifications to PEgRNA that are expected to increase activity. Figure 55C shows PEgRNA modifications to increase transcription of longer RNAs through promoter options and 5', 3' processing and termination. Figure 55D shows the elongation of the P1 system, which is an example of skeletal modification. Figure 55E shows that the incorporation of synthetic modifications on or elsewhere on the PEgRNA can increase activity. Figure 55F shows that the designed incorporation of minimal secondary structures on the template can prevent the formation of longer, more inhibitory secondary structures. Figure 55G shows a fission PEgRNA with a second template sequence anchored by an RNA element at the 3' end of the PEgRNA (left). Incorporation of elements at the 5' or 3' end of PEgRNA may enhance RT binding.

[0163] [Figure 56] Figures 56A–56D show the incorporation of PEgRNA backbone sequences into target loci. HTS data were analyzed for PEgRNA backbone insertions as described in Figures 60A–60B. Figure 56A shows the analysis of the EMX1 locus. The percentage of total sequencing reads containing one or more PEgRNA backbone nucleotides on an insertion adjacent to the RT template (left); the percentage of total sequencing reads containing a PEgRNA backbone insertion of a specified length (center); and the cumulative total percentage of PEgRNA insertions containing a specified length at best on the X axis are shown. Figure 56B shows the same for FANCF as in Figure 56A. Figure 56C shows the same for HEK3 as in Figure 56A. Figure 56D shows the same for RNF2 as in Figure 56A. Values ​​and error bars reflect the mean and sd of three independent biological replicas.

[0164] [Figure 57]Figures 57A–57I show the effects of PE2, PE2-dRT, and Cas9 H840A nickase on transcriptome-wide RNA abundance. Analysis of ribosomal RNA-depleted cellular RNA isolated from HEK293T cells expressing PE2, PE2-dRT, or Cas9 H840A nickase and PRNP-targeted or HEXA-targeted PEgRNA. RNA corresponding to 14,410 genes and 14,368 genes were detected in PRNP and HEXA samples, respectively. Figures 57A–57F show volcano plots presenting the ~-fold change in -log10 FDR-adjusted p-value versus log2 transcript abundance for each (Aeach) RNA. (Figure 57A) PE2 vs. PE2-dRT using PRNP-targeted PEgRNA, (Figure 57B) PE2 vs. Cas9 H840A using PRNP-targeted PEgRNA, (Figure 57C) PE2-dRT vs. Cas9 H840A using PRNP-targeted PEgRNA, (Figure 57D) PE2 vs. PE2-dRT using HEXA-targeted PEgRNA, (Figure 57E) PE2 vs. Cas9 H840A using HEXA-targeted PEgRNA, and (Figure 57F) PE2-dRT vs. Cas9 H840A using HEXA-targeted PEgRNA are compared. Red dots indicate genes showing a statistically significant change in relative abundance of ≥2 times (FDR-adjusted p<0.05). Figures 57G-57I are Venn diagrams of upwardly controlled and downwardly controlled transcripts (≥2x change), comparing PRNP and HEXA samples for (Figure 57G) PE2 vs. PE2-dRT, (Figure 57H) PE2 vs. Cas9 H840A, and (Figure 57I) PE2-dRT vs. Cas9 H840A.

[0165] [Figure 58] Figures 58A–58B show typical FACS gating for neuronal nucleus sorting. The nuclei were sequentially gated based on the DyeCycle Ruby signal, FSC / SSC ratio, SSC width / SSC height ratio, and GFP / DyeCycle ratio.

[0166] [Figure 59]Figures 59A to 59G show a protocol for cloning 3'-extended PEgRNA onto a mammalian U6 expression vector by Golden Gate assembly. Figure 59A shows an overview of the cloning. Figure 59B shows "Step 1: Digestion of pU6-PEgRNA-GG-vector plasmid (Component 1)". Figure 59C shows "Steps 2 and 3: Ordering and annealing of oligonucleotide parts (Components 2, 3, and 4)". Figure 59D shows "Step 2.b.ii.: sgRNA backbone phosphorylation (unnecessary if the oligonucleotides were purchased phosphorylated)". Figure 59E shows "Step 4: PEgRNA assembly". Figure 59F shows "Steps 5 and 6: Transformation of the assembled plasmid". Figure 59G shows a diagram summarizing the PEgRNA cloning protocol.

[0167] [Figure 60] Figures 60A to 60B show a Python script for quantifying PEgRNA backbone incorporation. A custom Python script was generated to characterize and quantify PEgRNA insertion at the target genomic locus. The script iteratively matches increasing-length text strings taken from a reference sequence (the guide RNA backbone sequence) to sequencing reads in a fastq file, and counts the number of sequencing reads that match the search query. Each sequential text string corresponds to an additional nucleotide of the guide RNA backbone. Exact length incorporation and cumulative incorporation of up to a defined length were calculated in this manner. To guarantee alignment and accurate counting of short sgRNA slices, 5 to 6 bases at the 3' end of the new DNA strand synthesized by reverse transcriptase are included at the start of the reference sequence.

[0168] [Figure 61] Figure 61 is a graph showing the percentage of total sequencing reads with the intended edit for SaCas9(N580A)-MMLV RT HEK3 +6C>A. Values for correct editing and indels are shown.

[0169] [Figure 62] Figures 62A–62B illustrate the importance of protospacers for the efficient incorporation of desired edits in precise positioning by prime editing. Figure 62A is a graph showing the percentage of total sequencing reads in which the target T·A base pair was converted to A·T for various HEK3 loci. Figure 62B shows the sequence analysis illustrating this.

[0170] [Figure 63] Figure 63 is a graph showing SpCas9 PAM variants with PAM editing (N=3). The percentage of total sequencing reads with targeted PAM editing is shown for SpCas9(H840A)-VRQR-MMLV RTs where NGA>NTA and for SpCas9(H840A)-VRER-MMLV RTs where NGCG>NTCG. The PEgRNA primer binding site (PBS) length, RT template (RT) length, and PE system used are listed.

[0171] [Figure 64] Figures 64A–64F are schematic diagrams illustrating the introduction of various site-specific recombinase (SSR) targets into the genome using PE. Figure 64A provides a general schematic diagram of the insertion of a recombinase target sequence by a prime editing factor. Figure 64B shows how a single SSR target inserted by PE can be used as a site for genomic incorporation of a DNA donor template. Figure 64C shows how tandem insertion of SSR target sites can be used to delete a portion of the genome. Figure 64D shows how tandem insertion of SSR target sites can be used to invert a portion of the genome. Figure 64E shows how insertion of two SSR target sites in two distal chromosomal regions can result in a chromosomal translocation. Figure 64F shows how insertion of two different SSR target sites on the genome can be used to replace a cassette from a DNA donor template.

[0172] [Figure 65] Figure 65 illustrates 1) PE-mediated synthesis of an SSR target site on the human cell genome and 2) the use of that SSR target site to incorporate a DNA donor template containing a GFP expression marker. Once successfully incorporated, GFP causes the cell to fluoresce. See Example 17 for further details.

[0173] [Figure 66] Figure 66 illustrates one embodiment of a prime editing factor provided as two halves of a prime editing factor protein. These regenerate the whole prime editing factor through self-splicing of the halves of the fission intein located at the end or beginning of each half of the prime editing factor protein.

[0174] [Figure 67]Figures 67A–67B illustrate the mechanism of intein removal from the polypeptide sequence and peptide bond reformation between the N-terminal and C-terminal extension sequences. Figure 67A illustrates the general mechanism of two half-proteins, each containing half of the intein sequence. These yield a fully functional intein when they come into contact within the cell. This then undergoes self-splicing and excision. The excision process results in the formation of a peptide bond between the N-terminal protein half (or "N-extine") and the C-terminal protein half (or "C-extine") to form a whole single polypeptide containing the N-extine and C-extine portions. In various embodiments, the N-extine may correspond to the N-terminal half of a split prime-editor fusion protein, and the C-extine may correspond to the C-terminal half of a split prime-editor. (b) shows the chemical mechanism of intein excision and the reformation of the peptide bond linking the N-extine half (red half) and the C-extine half (blue half). Since this involves the splicing action of two distinct components provided in trans, the excision of fission inteins (i.e., N-intine and C-intine in fission intein configurations) can also be called "trans-splicing."

[0175] [Figure 68A] Figure 68A demonstrates that when co-transfected into HEK293T cells, the delivery of both halves of the mitotic inteins of SpPE (SEQ ID NOs. 3875, 3876) in the linker maintains activity at three test loci. [Figure 68B] Figure 68B demonstrates that when co-transfected into HEK293T cells, the delivery of both halves of the mitotic intein of SaPE2 (e.g., SEQ ID NO: 443 and SEQ ID NO: 450) replicates the activity of full-length SaPE2 (SEQ ID NO: 134). Residues indicated in quotation marks are SaCas

[0176] The sequence of amino acids 741-743 of SaCas9 (the first residue of the C-terminal extension) is important for the intein trans-splicing reaction. "SMP" is the native residue. We also mutated these to the "CFN" consensus splicing sequence. As measured by the prime editing percentage, the consensus sequence has been shown to produce the highest rearrangement.

[0177] [Figure 68C] Figure 68C provides data showing that various disclosed PE ribonucleoprotein complexes (high-concentration PE2, high-concentration PE3, and low-concentration PE3) can be delivered in this manner.

[0178] [Figure 69] Figure 69 shows a bacteriophage plaque assay to determine the efficacy of PE in PANCE. Plaques (dark circles) indicate phages capable of successfully infecting E. coli. Increasing the concentration of L-rhamnose resulted in increased PE expression and increased plaque formation. Plaque sequencing revealed the presence of genome editing incorporated by PE.

[0179] [Figure 70]Figures 70A–70I provide examples of edited target sequences as illustrative step-by-step instructions for designing PEgRNA and nicking sgRNA for prime editing. Figure 70A: Step 1. Determine the target sequence and edit. Retrieve the sequence of the target DNA region (~200 bp) centered on the location of the desired edit (point mutation, insertion, deletion, or combination thereof). Figure 70B: Step 2. Position the target PAM. Identify the PAM proximal to the edit location. Take care to look for PAMs on both strands. A PAM close to the edit site is preferred, but it is possible to incorporate the edit using a protospacer and PAM that nick ≥30 nt from the edit site. Figure 70C: Step 3. Position the nick site. For each PAM to be considered, identify the corresponding nick site. In Sp Cas9 H840A nickase, the cleavage occurs between the 3rd and 4th bases of the 5' of the NGG PAM on the PAM-containing strand. All edited nucleotides must be located at 3' of the nick site. Therefore, a suitable PAM must place the nick at 5' of the target edit on the PAM-containing strand. In the example shown below, there are two possible PAMs. For simplicity, the remaining steps demonstrate the design of PEgRNA using only PAM 1. Figure 70D: Step 4. Design the spacer sequence. The protospacer of Sp Cas9 corresponds to the 20 nucleotides at 5' of the NGG PAM on the PAM-containing strand. Efficient Pol III transcription initiation requires that G be the first transcribed nucleotide. If the first nucleotide of the protospacer is G, then the spacer sequence of PEgRNA is simply the protospacer sequence. If the first nucleotide of the protospacer is not G, then the spacer sequence of PEgRNA is G, then the protospacer sequence. Figure 70E: Step 5. Design the primer binding site (PBS). Identify the DNA primers on the PAM-containing strand using the starting allele sequence. The 3' end of the DNA primer is the nucleotide just upstream of the nick site (i.e., the fourth base at the 5' end of the NGG PAM in SpCas9).A general design principle for use with PE2 and PE3 is that a PEgRNA primer-binding site (PBS) containing 12-13 nucleotides complementary to the DNA primer can be used for sequences with a GC content of ~40-60%. For sequences with a low GC content, longer (14-15 nt) PBS should be tested. For sequences with a higher GC content, shorter (8-11 nt) PBS should be tested. Regardless of GC content, the optimal PBS sequence should be determined empirically. To design a PBS sequence of length p, take the reverse complement of the first p nucleotide at 5' of the nick site on the PAM-containing strand using the starting allele sequence. Figure 70F: Step 6. Design the RT template. The RT template encodes homology to the designed edit and the sequences adjacent to the edit. The optimal RT template length varies based on the target site. For short-range editing (positions +1 to +6), it is recommended to test short (9-12 nt), medium (13-16 nt), and long (17-20 nt) RT templates. For long-range editing (positions +7 and beyond), it is recommended to use RT templates that extend at least 5 nt (preferably 10 nt or more) beyond the editing site to allow sufficient 3' DNA flap homology. For long-range editing, several RT templates should be screened to identify a functional design. For larger insertions and deletions (≧5 nt), incorporating greater 3' homology (~20 nt or more) onto the RT template is recommended. Editing efficiency is typically impaired when the RT template encodes the synthesis of G (corresponding to C in PEgRNA RT templates) as the last nucleotide on the DNA product being reverse-transcribed. Since many RT templates support efficient prime editing, it is recommended to avoid G as the last nucleotide synthesized when designing RT templates. To design an RT template sequence of length r, the desired allele sequence is used, and the reverse complement of the first r nucleotide at 3' of the nick site on the strand originally containing PAM is taken. Note that, compared to SNP editing, insertion or deletion editing using the same length RT template will not contain the same homology. Figure 70G: Step 7. Assemble the complete PEgRNA sequence.Concatenate the PEgRNA components in the following order (5'→3'): spacer, backbone, RT template, and PBS. Figure 70H: Step 8. Design the nicking sgRNA for PE3. Identify the PAM on the unedited strand upstream and downstream of the edit. The optimal nicking site is highly locus-dependent and should be determined empirically. Generally, a nick placed 40-90 nucleotides 5' opposite the nick induced by the PEgRNA leads to higher edit yields and fewer indels. The nicking sgRNA has a spacer sequence that matches a 20nt protospacer on the starting allele. It has a 5'G addition if the protospacer does not start with G. Figure 70I: Step 9. Design the PE3b nicking sgRNA. If the PAM is present on the complementary strand and its corresponding protospacer overlaps with the sequence targeted for editing, this edit may be a candidate for the PE3b system. In the PE3b system, the spacer sequence of the nicking sgRNA matches the sequence of the desired allele to be edited, rather than the starting allele. The PE3b system operates efficiently when the nucleotide(s) to be edited are located within the seed region (~10nt adjacent to the PAM) of the nicking sgRNA protospacer. This prevents nicking of the complementary strand until after the incorporation of the edited strand, thus preventing competition between the PEgRNA and sgRNA for binding to the target DNA. PE3b also avoids simultaneous nicking on both strands, and therefore significantly reduces indel formation while maintaining high editing efficiency. The PE3b sgRNA should have a spacer sequence that matches the 20nt protospacer of the desired allele, with an additional 5'G if required.

[0180] [Figure 71A]Figure 71A shows the nucleotide sequence of a SpCas9 PEgRNA molecule (upper strand). It terminates with "UUU" at the 3' end and does not contain a toehold loop element. The lower part of the figure illustrates the same SpCas9 PEgRNA molecule, which is further modified to contain a toehold loop element having the sequence 5'-"GAAANNNNN"-3' inserted immediately upstream of the "UUU" 3' end. "N" can be any nucleobase.

[0181] [Figure 71B] Figure 71B demonstrates that the efficiency of prime editing in HEK cells or EMX cells is increased when using PEgRNA containing a toehold loop element, while the percentage of indel formation does not change significantly.

[0182] [Figure 72]Figures 72A-72C illustrate alternative PEgRNA configurations that can be used for prime editing. Figure 72A illustrates the PE2:PEgRNA configuration for prime editing. This configuration involves PE2 (a fusion protein containing Cas9 and reverse transcriptase) complexed with PEgRNA (as also described in Figures 1A-1I and / or Figures 3A-3E). In this configuration, the reverse transcription template is incorporated onto the 3' elongation arm of the sgRNA to form PEgRNA, and the DNA polymerase enzyme is reverse transcriptase (RT) directly fused to Cas9. Figure 72B illustrates the MS2cp-PE2:sgRNA+tPERT configuration. This configuration includes a PE2 fusion (Cas9+reverse transcriptase), which is further fused to the MS2 bacteriophage coat protein (MS2cp) to form the MS2cp-PE2 fusion protein. To achieve prime editing, the MS2cp-PE2 fusion protein is complexed with sgRNA, which targets the complex to a specific target site on the DNA. The embodiment then involves the introduction of a transprime editing RNA template ("tPERT"), which acts in place of PEgRNA by providing a primer-binding site (PBS) and a DNA synthesis template via a separate molecule, namely tPERT. This also comprises an MS2 aptamer (stem-loop). The MS2cp protein recruits tPERT by binding to the MS2 aptamer of the molecule. Figure 72C illustrates alternative designs of PEgRNA that can be achieved by known methods for the chemosynthesis of nucleic acid molecules. For example, chemosynthesis may be used to synthesize a hybrid RNA / DNA PEgRNA molecule for use in prime editing, where the elongation arm of the hybrid PEgRNA is DNA instead of RNA. In such embodiments, DNA-dependent DNA polymerase may be used instead of reverse transcriptase to synthesize the 3' DNA flap containing the desired genetic alteration formed by prime editing. In another embodiment, the extension arm may be synthesized to encompass a chemical linker, which prevents DNA polymerase (e.g., reverse transcriptase) from using the sgRNA backbone or backbone as a template.In another embodiment, the elongation arm may include a DNA synthesis template oriented in the opposite direction relative to the overall orientation of the PEgRNA molecule. For example, as shown for a PEgRNA having an elongation attached to the 3' end of the sgRNA backbone and oriented 5'→3', the DNA synthesis template oriented in the opposite direction, i.e., 3'→5'. This embodiment may be advantageous for PEgRNA embodiments having an elongation arm positioned at the 3' end of the gRNA. By reversing the orientation of the elongation arm, once it reaches the 5' end of the new orientation of the elongation arm, DNA synthesis by polymerase (e.g., reverse transcriptase) will terminate, and therefore there will be no risk of using the gRNA core as a template.

[0183] [Figure 73] Figure 73 demonstrates prime editing using the tPERT and MS2 recruitment system (also known as MS2 tagging technology). An sgRNA that targets the prime editing factor protein (PE2) to a target gene locus is expressed in combination with tPERT containing a primer binding site (13nt or 17nt PBS), an RT template encoding His6 tag insertion and homologous arms, and an MS2 aptamer (located at the 5' or 3' end of the tPERT molecule). Either the prime editing factor protein (PE2) or a fusion of PE2 and MS2cp at the N-terminus was used. Editing was performed with or without complementary strand nicking sgRNA, as in the previously developed PE3 system (labeled "PE2+nicking" or "PE2" on the x-axis, respectively). This is also called "second strand nicking" as defined herein.

[0184] [Figure 74]Figure 74 demonstrates the MS2 aptamer expression of reverse transcriptase in trans and its recruitment by the MS2 aptamer system. The PEgRNA contains an MS2 RNA aptamer inserted into one of the two sgRNA backbone hairpins. Wild-type M-MLV reverse transcriptase is expressed as an N-terminal or C-terminal fusion to the MS2 coat protein (MCP). Editing occurs at the HEK3 site in HEK293T cells.

[0185] [Figure 75] Figure 75 provides bar graphs comparing the efficiency (i.e., "percentage of total sequencing reads with defined edits or indels") of PE2, PE2-trunc, PE3, and PE3-trunc for different target sites in various cell lines. The data show that prime editing factors containing truncated RT variants were approximately as efficient as prime editing factors containing untruncated RT proteins.

[0186] [Figure 76] Figure 76 demonstrates the editing efficiency of the intein mitotic prime editing factor. HEK239T cells were transfected with plasmids encoding full-length PE2 or intein mitotic PE2, PEgRNA, and nicking guide RNA. The consensus sequence (most of the amino-terminal residues of the C-terminal extension) is indicated. Percentage editing at two sites is shown: HEK3 +1CTT insertion and PRNP +6G→T. Replicate n=3 independent transfections.

[0187] [Figure 77]Figure 77 demonstrates the editing efficiency of the intein mitotic prime editing factor. Editing was assessed by targeted deep sequencing of bulk cortical and GFP+ subpopulations after delivery of nuclear-localized GFP:KASH at 5E10vg and small amounts of 1E10 per half SpPE3 to P0 mice via ICV injection. The editing factor and GFP were packaged in AAV9 with an EFS promoter. Mice were harvested 3 weeks after injection, and GFP+ nuclei were isolated by flow cytometry. Individual data points are shown for 1-2 mice per condition analyzed.

[0188] [Figure 78] Figure 78 demonstrates the editing efficiency of the intein fission prime editing factor. Specifically, the figure illustrates the AV fission SpPE3 construct used in Example 20. Cotransduction with AAV particles expressing SpPE3-N and SpPE3-C separately reproduces PE3 activity. Note that the N-terminal genome contains a U6-sgRNA cassette expressing nickel sgRNA, and the C-terminal genome contains a U6-PEgRNA cassette expressing PEgRNA.

[0189] [Figure 79]Figure 79 shows the editing efficiency of certain optimized linkers. In particular, the data shows the editing efficiency of the PE2 construct with the current linker (labeled PE2; white box) compared to various versions in which the linker is replaced by the indicated sequence, for representative PEgRNAs for transition, transversion, insertion, and deletion editing at the HEK3, EMX1, FANCF, and RNF2 loci. The replacement linkers are labeled "1×SGGS" (SEQ ID NO: 174), "2×SGGS" (SEQ ID NO: 446), "3×SGGS" (SEQ ID NO: 3889), "1×XTEN" (SEQ ID NO: 171), "No Linker", "1×Gly", "1×Pro", "1×EAAAK" (SEQ ID NO: 3968), "2×EAAAK" (SEQ ID NO: 3969), and "3×EAAAK" (SEQ ID NO: 3970). Editing efficiency is measured in bar graph format relative to the "control" editing efficiency of PE2. The linker for PE2 is SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (Sequence ID 127). All editing was performed in the context of the PE3 system, which means the PE2 editing construct plus the addition of an optimal secondary sgRNA nicking guide. See Example 21.

[0190] [Figure 80] Taking an average efficiency of ~ times relative to PE2 produces the graph shown, indicating that using the 1×XTEN (sequence ID 171) linker sequence improves editing efficiency by an average of 1.14 times (n=15). See Example 21.

[0191] [Figure 81] Figure 81 illustrates the transcription levels of PEgRNA from different promoters, as described in Example 22.

[0192] [Figure 82] Figure 82 illustrates the impact of different types of structural modifications on PEgRNA on the relative editing efficiency compared to unmodified PEgRNA.

[0193] [Figure 83] Figure 83 illustrates a PE experiment targeting HEK3 gene editing. PE3 was used to specifically target a 10nt insertion at position +1 relative to the nick site.

[0194] [Figure 84A] Figure 84A illustrates an exemplary PEgRNA having a spacer, gRNA core, and extension arm (RT template + primer binding site). The 3' end of the PEgRNA is modified by a tRNA molecule coupled via a UCU linker. The tRNA can encompass various post-transcriptional modifications; however, these modifications are not required.

[0195] [Figure 84B] Figure 84B illustrates the structure of a tRNA that may be used to modify the PEgRNA structure. See Example 22. P1 may be of variable length. P1 may be elongated to help prevent RNAseP processing of the PEgRNA-tRNA fusion.

[0196] [Figure 85] Figure 85 illustrates a PE experiment targeting editing of the FANCF gene. The G→T conversion at position +5 relative to the nick site was specifically targeted, and the PE3 construct was used.

[0197] [Figure 86] Figure 86 illustrates a PE experiment targeting HEK3 gene editing. The PE3 construct was used to specifically target the insertion of a 71nt FLAG tag at position +1 relative to the nick site.

[0198] [Figure 87] Results from screening N2A cells in which pegRNA incorporates 1412Adel, along with details of primer-binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels).

[0199] [Figure 88] Results from screening N2A cells in which pegRNA incorporates 1412Adel, along with details of primer-binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels).

[0200] [Figure 89] Figure 89 illustrates the results of editing at the proximal (proxy) locus of the β-globin gene and at HEK3 in healthy HSCs. The concentrations of editing factors versus pegRNA and nicking gRNA were varied.

[0201] [Figure 90] Figure 90 provides a schematic diagram of one aspect of double flap prime editing. The DNA target sequence is acted upon by two prime editing complexes (guided by pegRNA-A and pegRNA-B). The two pegRNAs target opposing strands of the double helix. Each prime editing factor (PE2·pegRNA) nicks a single DNA strand and then synthesizes a 3' DNA flap using the pegRNA as a template. The action of the two prime editing factor complexes results in the production of an intermediate containing two 3' flaps on opposing strands of DNA. The two 3' flaps are complementary to each other at their 3' ends. Annealing of the 3' ends of the 3' flaps results in the formation of two double-strand structures: one double helix is ​​made from a paired 3' flap containing a new DNA sequence (red), and the other double helix is ​​made from a paired 5' flap containing the original DNA sequence (black). The excision of the intervening original DNA double helix (black paired 5' flap) creates a double-nicked DNA species containing the desired new DNA sequence (red) that replaces the original DNA sequence. Ligation of both nicks completes the editing process.

[0202] [Figure 91]Figures 91A–91B provide results from Example 7. Figure 91A shows the Crispresso2 output allele table (aligned for the desired product) for the replacement of a 90bp sequence with a novel 22bp sequence at the HEK3 site in HEK293T cells using a double prime editing factor. The desired product accounts for over 80% of the sequenced reads. For comparison, the reference starting allele is shown above the sequenced allele. Figure 91B shows the sequences of the pegRNAs used to achieve the sequence replacement shown in Figure 91A. pegRNA1 and pegRNA2 target different strands of the DNA double helix, generating a 5' displaced nick as depicted in Figure 90.

[0203] [Figure 92] Figure 92 shows a design embodiment of a pegRNA design for double flap prime editing. Two pegRNAs, shown in the figure as pegRNA A and pegRNA B, are used for double flap prime editing. Each pegRNA contains a spacer sequence (dark blue) which guides the prime editing complex to the target DNA site. The two pegRNAs target opposing strands of the DNA double helix. Like other pegRNAs, the double flap prime editing pegRNAs contain a 3' extension with a primer-binding sequence (PBS, green) that anneals to a nicked genomic DNA strand to initiate reverse transcription, and a reverse transcription template (RT template, light blue) that serves as a template for the synthesis of new DNA by reverse transcriptase. Unlike pegRNAs used in classical prime editing, which require the newly synthesized edited 3' flap to compete with the endogenous 5' flap, it is not necessary to encode homology to the target site on the RT template. Instead, it is only required that the 3' ends of the two synthesized 3' flaps contain complementarity to each other (i.e., the 3' ends of the 3' flaps are inversely complementary sequences to each other). This complementarity allows the two 3' flaps to anneal, facilitating the formation of the desired edited DNA sequence.

[0204] [Figure 93]Figures 93A–93B show the results of Example 7 of using double-flap prime editing to incorporate the attB and attP sites of BxB1 by double-flap prime editing. Figure 93A shows the incorporation of a 38 bp BxB1 attB site in HEK3. Six pegRNAs were constructed, three targeting the (+) strand (A1, A2, and A3) and three targeting the (-) strand (B1, B2, and B3). These differed in the amount of attB sequence encoded on the RT template, leading to a different number of complementary nucleotides between the two flaps. A 3x3 matrix of pegRNAs was evaluated for the incorporation of the attB sequence in the genomic positioning of the target in HEK293T cells. Figures 93A and 93B show the incorporation of a 50 bp BxB1 attP site in HEK3. Six pegRNAs were constructed, three targeting the (+) strand (A1, A2, and A3) and three targeting the (-) strand (B1, B2, and B3). These differed in the amount of attP sequence encoded on the RT template, leading to a different number of complementary nucleotides between the two flaps. A 3x3 matrix of pegRNAs was evaluated for the incorporation of the attP sequence in the genomic positioning of the target in HEK293T cells. For both edits in Figure 93A and Figure 93B, incorporation of the attB or attP site occurs with a 90 bp concomitant deletion of the genomic DNA sequence positioned between the two nick sites.

[0205] [Figure 94] Figure 94 shows the results of incorporating the attB and attP sites of BxB1 at the human safeharbour locus using double-flap prime editing. The image shows the incorporation of the attP site of BxB1 at the AAVS1 locus (left) or the attB site of BxB1 at the CCR5 locus (right) in HEK293T cells. Correct editing is shown in blue (AAVS1) and red (CCR5), while indel byproducts are shown in gray.

[0206] [Figure 95]Figure 95 provides a schematic diagram of genomic sequence inversion by quadruple flap prime editing. A region of genomic DNA is targeted for inversion (green and orange segments). Four pegRNAs are delivered to the cell along with the PE2 prime editing factor. One pair of pegRNAs targets a single genomic DNA strand, serving as a template for the synthesis of two complementary DNA flaps (A and A', blue), and the second pair targets another genomic DNA strand, serving as a template for the synthesis of two complementary DNA flaps with orthogonal DNA sequences (B and B', pink). The complementary flaps anneal to form a 3' overhang double helix. The 5' overhang double helix is ​​excised by endocellular repair enzymes. The nick is ligated, producing a product allele containing the inverted DNA sequence (green and orange segments) and the pegRNA-templated sequence at the inversion junction (blue and pink segments).

[0207] [Figure 96] Figure 96 shows the results of amplicon sequencing of the AAVS1 inversion junction. CRISPResso2 analysis output of a 2.7kb inversion at the AAVS1 locus in HEK293T cells using a quadruple flap-primed editing strategy. Expected PCR amplification and sequencing of the inversion junction showed the desired product with the attP or attB sequence of Bxb1 inserted into the inversion junction.

[0208] [Figure 97]Figures 97A-97B show the results of targeted incorporation of a circular DNA plasmid into the genome using quadruple flap prime editing. (Figure 97A) A region of genomic DNA and a region of plasmid DNA are targeted for quadruple flap prime editing incorporation. Four pegRNAs are delivered to the cell along with the PE2 prime editing factor. Two pegRNAs serve as templates for complementary sequences, one targeting a single genomic DNA strand and the other targeting a single plasmid DNA strand (generating a blue flap). The other two pegRNAs target genomic DNA and plasmid DNA strands opposite to those of the first two pegRNAs, and they serve as templates for the synthesis of two complementary DNA flaps that are orthogonal to the first pair (pink). The complementary flaps anneal to form a 3' overhang double helix. The 5' overhang double helix is ​​excised by endocellular repair enzymes. Nicks are ligated, producing product alleles (green and orange segments) containing the incorporated plasmid DNA sequence and pegRNA-templated sequences at the incorporation junction (blue and pink segments). (Figure 97B) CRISPResso2 analysis of expected junction amplicon sequencing, showing the plasmid backbone and genomic DNA sequences bridged by the pegRNA-templated attP sequence.

[0209] [Figure 98] Figures 98A–98B show the results of targeted chromosomal translocation by quadruple flap prime editing. (Figure 98A) pegRNA targets two regions on different chromosomes. Complementary 3' DNA flaps bridge the two chromosomal sequences and guide the orientation of the translocation. (Figure 98B) Targeted translocation between the MYC and TIMM44 loci in HEK239 T cells. CRISPResso2 analysis output from amplicon sequencing of the expected junction from the translocation between the MYC locus on chromosome 8 and the TIMM44 locus on chromosome 19. The majority of the sequencing reads correspond to the desired allele sequence.

[0210] [Figure 99]Figure 99 shows the incorporation of the attB and attP sites of BxB1 by double flap prime editing at the IDS locus. HEK293T cells were transfected with different pairs of PE2 and pegRNA (e.g., in the first column, pegRNA A1_a and pegRNA B2_a, which have templates for incorporation of the attP site in the forward direction). Efficiency was measured by HTS. This data demonstrates that double flap editing can successfully insert the target sequence into the IDS locus with an efficiency of up to ~80%.

[0211] [Figure 100] Figures 100A and 100B illustrate double-flap-mediated duplication. Figure 100A is a schematic diagram showing double-flap-mediated duplication at the AAVS1 locus in 293T cells. Figure 100B shows the results of using double-flap pegRNA in PE2 to induce gene sequence duplication in AAVS1.

[0212] [Figure 101] Figure 101 shows a novel translocation MYC-CCR5 induced by multiple flaps. MYC-CCR5 translocation was induced by quadruple flap pegRNA and PE2. MYC-CCR5 translocation events were induced by quadruple pegRNA. Four different sets of pegRNA were tested in HEK293T cells. The derived translocation junction product between chr8 and chr3 was amplified by junction primers. The graph shows the percentage of reads aligned to the expected junction allele. The results indicate that quadruple flaps can mediate translocations of the MYC and CCR5 genes, with product purity of nearly 100% at junction 1 and ~50% at junction 2. The representative allele plot shows the sequence aligned to the expected allele sequence at junction 1.

[0213] [Figure 102]Figures 102A–102B show double flap and multiple flap editing in other human cell lines. Figure 102A shows double flap editing in four different human cell lines. HEK293T and HeLa cells were transfected with double pegRNA and PE2 to edit three different genomic loci (IDS, MYC, and TIMM44). U2OS and K562 cells were nucleofected with the same components. Double flap editing showed robust editing at the target loci in all four human cell lines, particularly in HEK293T and K562 cells. The cellular mechanisms enabling double flap editing are conserved in many human cell types. Figure 102B shows that multiple flap (quadriple pegRNA) led to a 2.7kb inversion at AAVS1 in HeLa cells.

[0214] [Figure 103] Figures 103A-103B show the results of HTS-mediated inversion efficiency measurements at the CCR5 locus. The percentage of expected inversion-edited alleles was measured by HTS. Four quadruple pegRNA sets, PE2, were transfected into HEK293T cells. Figure 103A shows double-flap-mediated sequence duplication (~100 nt) at the CCR5 locus in HEK293T cells. Editing efficiency achieved ~1.5% according to 300-cycle paired-end sequencing analysis by HTS. Figure 103B shows quadruple-flap-mediated sequence inversion (~95-117 nt) at the CCR5 locus in HeLa cells. Editing efficiency achieved ~1.2% according to 300-cycle paired-end sequencing analysis by HTS. These results indicate that multiple flaps can successfully mediate duplication and inversion accurately at the CCR5 locus. Editing specificity is high when the target sequence is duplicated (percentage of indels <2%).

[0215] [Figure 104]Figures 104A–104E illustrate targeted cellular repair pathways for double flap editing. In Figures 104A–104D, HEK293T cells were transfected with pegRNA and PE2 along with plasmids expressing Exo1, Fen1, red fluorescent protein (control), DNA2, Mlh1 neg, and a P53 inhibitor. Editing efficiency was measured by HTS. Editing efficiency was compared between candidate and RFP controls. In Figure 104E, HEK293T cells were transfected for each target locus with siRNA plasmids as well as pegRNA and PE2. Untargeted siRNA (siNT) was used as a control. Editing efficiency was measured by HTS. Editing efficiency was compared at each target locus between each siRNA knockdown and siNT control. HEK293T cells were transfected with double pegRNA, PE2, and plasmids expressing Exo1, Fen1, red fluorescent protein (control), DNA2, Mlh1 neg, and a P53 inhibitor, respectively. Editing efficiency was measured by HTS. Editing efficiency was compared between candidate and RFP controls. Two-sided paired Student's t-tests were used to measure statistical differences between each treatment and RFP control (P<0.05, *; P<0.01, **; P<0.001, ***). Overexpression of FEN1 improved double flap editing efficiency at all four target loci (MYC, TIMM44, IDS, CCR5).

[0216] [Figure 105]Figures 105A–105B show sequence duplication mediated by a double flap at the AAVS1 locus. Figure 105A shows a schematic diagram of sequence duplication mediated by a double flap at the AAVS1 locus. Figure 105B shows that a ~300 bp sequence duplication was introduced into the AAVS1 locus of 293T cells by using double pegRNAs that generate two unique 3' flap structures. The expected allele is amplified with specific primers and attached to HTS. ~94% of the reads are aligned to the expected allele with the duplication. The duplication product is not observed in untreated samples.

[0217] [Figure 106] Figure 106 shows targeted IDS genome sequence inversion by quadruple flap prime editing. It has been shown that ~13% of Hunter syndrome patients have IDS gene sequence inversion (Bondeson et al., Human Molecular Genetics, 1995). Quadruple flap prime editing was applied to induce this pathogenic inversion of ~40kb of the IDS genome sequence in HEK293T cells. Six sets of quadruple pegRNAs were tested by transfecting HEK293T cells with pegRNA and PE2. Primers were used to specifically amplify the inverted sequences at junctions "ab" and "cd". ~95% of the expected inverted allele sequences were observed at IDS_QF1 at both junctions. Other sets of pegRNAs also produced high percentages of the expected allele sequences at both junctions. Inverted junction products were not observed in untreated samples.

[0218] [Figure 107]Figure 107 shows that 3' motif modification of PegRNA improves double flap editing efficiency at IDS loci. To further improve double flap editing efficiency, the pseudoknot evoPreQ1 motif was introduced to protect the 3' end of the pegRNA. By comparing the editing efficiency produced by unmodified and evoPreQ1-modified double pegRNA, an overall increase in editing efficiency with modified pegRNA at the target IDS locus is observed. The improvement in double flap editing efficiency can reach up to 5.3 times.

[0219] [Figure 108] Figures 108A–108C provide an overview of twin PE and twin PE-mediated sequence replacement. Figure 108A shows that the twin PE system targets a genomic DNA sequence containing two protospacer sequences on opposing strands of DNA. The PE2·pegRNA complex targets each protospacer, generating a single-stranded nick and reverse transcribing a template encoded by the pegRNA containing the desired insertion sequence. After the synthesis and release of the 3' DNA flap, a hypothetical intermediate exists, possessing an annealed 3' flap containing the edited DNA sequence and an annealed 5' flap containing the original DNA sequence. Excision of the original DNA sequence contained on the 5' flap, followed by ligation of the 3' flap to the corresponding excision site, produces the desired edited product. Figure 108B shows an example of twin PE-mediated replacement of a 90bp sequence at HEK site 3 by the 38bp Bxb1 attB sequence. Figure 108C shows the evaluation of twin PE in HEK293T cells for the incorporation of the 38 bp BxB1 attB site or the 50 bp Bxb1 attP site shown in Figure 108B at HEK site 3. PegRNAs of varying insertion lengths are used as templates. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0220] [Figure 109]Figures 109A–109E show targeted sequence insertions, deletions, and recodings using twin PE in human cells. Figure 109A shows the insertion of an FKBP coding sequence fragment by PE3 (12 bp, 36 bp, 108 bp, or 321 bp) or twin PE (108 bp) at HEK site 3 in HEK293T cells. Figure 109B shows sequence recoding on exons 4 and 7 in PAHs of HEK293T cells using twin PE. The 64 bp target sequence in exon 4 was edited using a 24, 36, or 59 bp overlap flap, the 46 bp target sequence in exon 7 was edited using a 22 or 42 bp overlap flap, or the 64 bp sequence in exon 7 was edited using a 24 or 47 bp overlap flap. Editing activity was compared using standard pegRNA or epegRNA containing the 3'evoPreQ1 motif. Figure 109C is a schematic diagram of three different types of double flap deletion strategies considered to perform targeted deletions. The "Basic Anchor (BA)" twin PE strategy allows for flexible deletions starting from an arbitrary position 3' of one nick site and ending at another nick site. The "Hybrid Anchor (HA)" twin PE strategy allows for flexible deletions of sequences at an arbitrarily selected position between two nick sites. The "Prime Del (PD)" strategy tested here allows for deletions of sequences starting from one nick site and ending at another nick site. Figure 109D shows sequence deletions at HEK site 3 in HEK293T cells using the BA-twin PE, HA-twin PE, or PD strategies targeting the same protospacer pair. Editing activity was compared using standard pegRNA or epegRNA containing the 3'evoPreQ1 motif. Figure 109E shows the deletion of exon 51 sequence at the DMD locus in HEK293T cells using BA-twin PE, PD, paired Cas9 nuclease, or attB sequence substitution mediated by twin PE. Values ​​and error bars reflect the mean and sd of three independent biological replicas. At least two independent biological replicas were performed in the DMD exon 51 skipping experiments.

[0221] [Figure 110] Figures 110A–110E show site-specific genomic incorporation of DNA cargo by twin PE and BxB1 recombinases in human cells. Figure 110A shows screening of twin PE·pegRNA pairs for incorporation of the attP sequence of BxB1 at the AAVS1 locus in HEK293T cells. Figure 110B shows screening of twin PE·pegRNA pairs for incorporation of the attB sequence of BxB1 at the CCR5 locus in HEK293T cells. Figure 110C shows single-transfection knock-in of a 5.6kb DNA donor using twin PE·pegRNA pairs targeting CCR5 (four leftmost bars) or AAVS1 (three rightmost bars). The twin PE·pegRNAs incorporate attB into CCR5 or attP into AAVS1. BxB1 then incorporates the donor with the corresponding attachment site into the genomic attachment site. Figure 110D shows the optimization of single-transfection knock-in in CCR5 using a 531 / 584 twin PE·pegRNA pair. The identity of the edit by template (attB vs. attP), the identity of the central dinucleotide (wild-type GT vs. orthogonal mutant GA), and the length of the overlap between flaps were varied to identify the combination that supported the highest knock-in efficiency. Figure 110E shows the insertion of the attB sequence of BxB1 into intron 1 of ALB in HEK293T and Huh7 cell lines. Figure 110F shows a comparison of single-transfection knock-in efficiencies in CCR5 and ALB in HEK293T and Huh7 cell lines.

[0222] [Figure 111]Figures 111A–111E show site-specific genomic sequence inversions in human cells using twin PE and Bxb1 recombinases. Figure 111A is a schematic diagram of the recombinant hotspots in IDS and IDS2 leading to pathogenic 39kb inversions, as well as the combined twin PE-Bxb1 strategy for incorporating or correcting IDS inversion mutations. Figure 111B shows screening of pegRNA pairs in IDS and IDS2 for incorporation of attP or attB recombinant site insertions at IDS and IDS2 loci by specific DNA targets. Figure 111C shows DNA sequencing analysis of attP or attB insertions by sequential DNA transfection. Figure 111D shows the purity of the inversion product at inverted junctions 1 and 2 (sequential transfection), indicating successful inversions at the two junctions. Figure 111E shows the quantification of inversion efficiency at the junction (sequential transfection and "one-pot" RNA nucleofection).

[0223] [Figure 112] Figure 112 shows the recoding of exon 10, 11, and 12 sequences of PAH in HEK293T cells by twin PE. The 64 bp target sequence in exon 10 was edited using a 28 bp overlap flap, the 61 bp and 55 bp target sequences in exon 11 were edited using a 25 bp overlap flap, or the 68 bp and 58 bp sequences in exon 12 were edited using 27 and 24 bp overlap flaps, respectively. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicas.

[0224] [Figure 113] Figure 113 shows transfection of HEK293T clone cell lines containing homozygous attB site insertions with BxBI plasmids and attP-containing donor DNA plasmids. Knock-in efficiencies measured by ddPCR were between 12 and 17% at the target site.

[0225] [Figure 114] Figure 114 shows HTS measurements of expected junctional sequences containing attL and attR recombinant products after one-pot knock-in mediated by twin PE and BxBI. Product purity ranges from 71 to 95%. Values ​​and error bars reflect the mean and standard deviation of three independent biological replicates.

[0226] [Figure 115A] Figure 115A shows the attB insertion efficiency mediated by twin PE due to the reduced flap overlap length of double pegRNA.

[0227] [Figure 115B] Figure 115B shows PCR products amplified by specific primers to capture recombination between the donor DNA and pegRNA plasmid shown on an agarose gel. Recombination between the donor DNA and pegRNA plasmid was reduced by smaller flap overlap.

[0228] [Figure 116] Figure 116 is a schematic diagram of the PCR strategy developed to quantify the IDS inverse efficiency.

[0229] definition Unless otherwise specified, all technical and scientific terms used in this application have the meanings commonly understood by practitioners of the art to which this invention pertains. The following references provide many common definitions of terms used in this invention to those skilled in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Unless otherwise specified, the following terms used in this application have the meanings attributed thereto.

[0230] Antisense chain In genetics, the "antisense" strand of a segment within a double-stranded DNA is considered the template strand, extending in a 3'→5' orientation (runs). Conversely, the "sense" strand is a segment within the double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of the DNA that extends from 3' to 5'. In the case of a protein-coding DNA segment, the sense strand is the DNA strand that has the same sequence as the mRNA that takes the antisense strand as its template during transcription and eventually (typically, but not always) undergoes translation to become a protein. Thus, the antisense strand carries the RNA that is later translated into a protein, while the sense strand has a makeup that is nearly identical to that of the mRNA. Note that for each segment of dsDNA, there are likely to be two sets of sense and antisense strands, depending on the direction from which one is read (since sense and antisense are relative to the viewpoint). The gene product or mRNA ultimately indicates that one of the strands of one of the segments of the dsDNA is referred to as either sense or antisense.

[0231] Dual-specific ligands As used in this application, the terms “bispecific ligand” or “bispecific moiety” refer to a ligand that binds to two different ligand-binding domains. In one embodiment, the ligand is a small molecule compound, peptide, or polypeptide. In another embodiment, the ligand-binding domain is a “dimerizing domain,” which may be incorporated onto a protein as a peptide tag. In various embodiments, two proteins, each containing the same or different dimerizing domains, may be induced to dimerize by the binding of each dimerizing domain to a bispecific ligand. As used in this application, the term “bispecific ligand” may also be referred to as a “chemical inducer of dimerization” or “CID.”

[0232] Cas9 The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease containing the Cas9 domain or a fragment thereof (for example, a protein containing the active or inactive DNA cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). “Cas9 domain” as used herein is a protein fragment containing the active or inactive cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9. “Cas9 protein” is the full-length Cas9 protein. Cas9 nucleases are also sometimes referred to as casn1 nucleases or CRISPR (clustered and regularly arranged short palindromic sequence repeats). C stestered R egularly I nterspaced S hort P alindromic RAlso referred to as epeat))-related nucleases. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the preceding mobile element, and a targeted entry nucleic acid. The CRISPR cluster is transcribed and processed into processed CRISPR RNA (crRNA). In the type II CRISPR system, the modification processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for the ribonuclease 3-aided processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endonuclease-likely cleaves a linear or circular dsDNA target complementary to the spacer. Target strands that are not complementary to crRNA are first cleaved endonucleaseically and then 3'-5' exonucleaseically. Normally, DNA binding and cleavage require both proteins and both RNAs. However, single guide RNAs ("sgRNAs," or simply "gNRAs") can be modified to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (this entire content is incorporated herein by reference). Cas9 helps distinguish self versus non-self by recognizing short motifs within CRISPR repeat sequences (PAMs or protospacer-adjacent motifs).Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III."Deltcheva E., Chylinski K., Sharma See CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (the entire contents of each of these are incorporated herein by reference). Cas9 orthologs have been described in various species, including S. pyogenes and S. thermophilus, but are not limited to these. Additional preferred Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Furthermore, such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems" (2013) RNA Biology 10:5, 726-737 (the entire content of which is incorporated herein by reference). In some embodiments, the Cas9 nuclease contains one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0233] A Cas9 domain with an inactivated nuclease may interchangeably be referred to as the “dCas9” protein (where the nuclease represents the “inactive” form of Cas9). Methods for generating Cas9 domains (or fragments thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5): 1173-83 (the entire contents of each of these are incorporated herein by reference)). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, proteins containing fragments of Cas9 are provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA-cleaving domain of Cas9. In some embodiments, the protein or fragment containing Cas9 is referred to as a "Cas9 variant." The Cas9 variant shares homology with Cas9 or its fragment.For example, a Cas9 variant is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, at least approximately 99.8% identical, or at least approximately 99.9% identical to a wild-type Cas9 (e.g., SpCas9 with sequence number 18). In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9 (for example, SpCas9 of SEQ ID NO: 18). In some embodiments, a Cas9 variant includes a fragment of Sequence ID X Cas9 (e.g., a gRNA-binding domain or a DNA-cleaving domain) such that its fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of Sequence ID X). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18).

[0234] cDNA The term "cDNA" refers to a strand of DNA copied from an RNA template. cDNA is complementary to the RNA template.

[0235] cyclic permutant As used herein, the term “circular permutant” refers to a protein or polypeptide (e.g., Cas9) that contains a circular permutation, which is a change in the structural configuration of a protein involving a change in the order of amino acids in the amino acid sequence of the protein. In other words, a circular permutant is a protein in which the N-terminus and C-terminus are altered compared to the wild-type counterpart, for example, the C-terminal half of the wild-type protein is replaced by a new N-terminal half. A circular permutation (or CP) is essentially a morphological rearrangement of the protein's primary sequence, where its N-terminus and C-terminus are joined, often with a peptide linker, but simultaneously splitting the sequence at different positions to create new adjacent N-terminus and C-terminus. The result is a protein structure that may have a different but often the same or overall similar three-dimensional (3D) shape, possibly including improved or altered features (including reduced sensitivity to proteolysis, improved catalytic activity, altered substrate or ligand binding, and / or improved thermal stability). Circulating substitution proteins can occur naturally (e.g., concanavalin A and lectins). In addition, circulating substitutions can result from post-translational modifications or may be modified using recombinant techniques.

[0236] Circularly replaced Cas9 The term “circularly permuted Cas9” refers to any Cas9 protein or variant that has been generated as a circularly permuted form, thereby undergoing local rearrangement of its N-terminus and C-terminus. Such circularly permuted Cas9 proteins ("CP-Cas9"), or their variants, retain their ability to bind to DNA when complexed with guide RNA (gRNA). See Oakes et al., “Protein Engineering of Cas9 for enhanced function,” Methods Enzymol, 2014, 546:491-511 and Oakes et al., “CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification,” Cell, January 10, 2019, 176:254-267 (each of which is incorporated herein by reference). This disclosure uses any previously known CP-Cas9 or any new CP-Cas9, provided that the resulting circulating replacement protein retains its ability to bind to DNA when complexed with guide RNA (gRNA). Exemplary CP-Cas9 proteins are sequence numbers 77-86.

[0237] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of preceding infections caused by viruses that have invaded prokaryotes. These DNA snippets are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with arrays of CRISPR-related proteins (including Cas9 and its homologs) and CRISPR-related RNAs, they effectively form a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of precrRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for the ribonuclease 3-assisted processing of precrRNA. Next, Cas9 / crRNA / tracrRNA endo-cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endo-cut and then 3'-5' exo-trimmed. In nature, DNA binding and cleavage typically require both proteins and both RNAs. However, single guide RNAs ("sgRNA" or simply "gNRA") can be manipulated to incorporate aspects of both crRNA and tracrRNA into a single RNA species guide RNA. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entire content of this is incorporated here by reference. Cas9 recognizes short motifs (PAM or protospacer adjacency motifs) on CRISPR repeat sequences to help distinguish between self and non-self.CRISPR biology and Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP,Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III."Deltcheva E., Chylinski See K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entire contents of each of these are incorporated herein by reference. Cas9 orthologs have been described in various species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems" (2013) RNA Biology 10:5, 726-737; the entire contents of this publication are incorporated herein by reference.

[0238] In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of precrRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for the ribonuclease 3-assisted processing of the precrRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves linear or circular nucleic acid targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endo-cut and then 3'-5' exo-trimmed. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single guide RNAs ("sgRNA" or simply "gRNA") can be manipulated to incorporate aspects of both crRNA and tracrRNA into a single RNA species guide RNA.

[0239] Generally, the “CRISPR system” collectively refers to the transcripts and other elements involved in the expression of CRISPR-related (“Cas”) genes or that lead to their activity, and includes the Cas gene, tracr (trans-activated CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (in the context of endogenous CRISPR systems, encompassing “direct repeats” and partial direct repeats processed by tracrRNA), guide sequences (also called “spacers” in the context of endogenous CRISPR systems), or sequences encoding other sequences and transcripts from the CRISPR locus. The tracrRNA of the system is (fully or partially) complementary to the tracr mate sequence present on the guide RNA.

[0240] DNA synthesis template As used herein, the term “DNA synthesis template” refers to a region or portion of the elongation arm of a PEgRNA that is utilized by the polymerase of a prime editing factor as a template strand (containing the desired edit and then encoding a 3' single-strand DNA flap that replaces the corresponding endogenous DNA strand at a target site through the mechanism of prime editing). In various embodiments, the DNA synthesis template is shown in Figure 3A (in the context of a PEgRNA including a 5' elongation arm), Figure 3B (in the context of a PEgRNA including a 3' elongation arm), Figure 3C (in the context of an internal elongation arm), Figure 3D (in the context of a 3' elongation arm), and Figure 3E (in the context of a 5' elongation arm). The elongation arm contains the DNA synthesis template and may consist of DNA or RNA. In the case of RNA, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of DNA, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase. In various embodiments (as depicted, for example, in Figures 3D-3E), the DNA synthesis template (4) may include the “editing template” and the “homologous arm,” as well as all or part of any 5'-end modified region e2. That is, depending on the nature of the e2 region (for example, whether it encompasses a hairpin, toe loop, or stem / loop secondary structure), the polymerase may encode nothing, or part or all of the e2 region. In other words, in the case of the 3' extension arm, the DNA synthesis template (3) may include a portion of the extension arm (3) extending from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core, which may act as a template for the synthesis of a single strand of DNA by a polymerase (for example, reverse transcriptase). In the case of a 5' elongated arm, the DNA synthesis template (3) may encompass the portion of the elongated arm (3) extending from the 5' end of the PEgRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of PEgRNAs having either a 3' elongated arm or a 5' elongated arm.The embodiments described herein (for example, Figure 71A) refer to an “RT template” that encompasses the editing template and homologous arms, i.e., the sequences of the PEgRNA elongation arms that are actually used as templates during DNA synthesis. The term “RT template” is equivalent to the term “DNA synthesis template.”

[0241] In the case of transpriming (for example, Figures 3G and 3H), the primer binding site (PBS) and DNA synthesis template can be manipulated into separate molecules called transpriming factor RNA templates (tPERT).

[0242] Dimerization domain The term “dimerization domain” refers to a ligand-binding domain that binds to the binding domain of a bispecific ligand. The “first” dimerization domain binds to the first binding site of the bispecific ligand, and the “second” dimerization domain binds to the second binding site of the same bispecific ligand. When the first dimerization domain is fused to the first protein (e.g., via PE as discussed in this application) and the second dimerization domain is fused to the second protein (e.g., via PE as discussed in this application), the first and second proteins dimerize in the presence of the bispecific ligand, and the bispecific ligand has at least one portion that binds to the first dimerization domain and at least another portion that binds to the second dimerization domain.

[0243] downstream As used herein, the terms “upstream” and “downstream” are relative terms defining the linear positions of at least two elements located within a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5’→3’ direction. In particular, if the first element is located somewhere at 5’ relative to the second element, the first element is upstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 5’ side of a nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, if the first element is located somewhere at 3’ relative to the second element, the first element is downstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 3’ side of a nick site, the SNP is downstream of the Cas9-induced nick site. Nucleic acid molecules can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis is the same for single-stranded and double-stranded nucleic acid molecules, except where it is necessary to choose which strand of a double-stranded molecule to consider, where the terms upstream and downstream refer only to single-stranded nucleic acid molecules. Often, the strand of double-stranded DNA that can be used to determine the relative positions of at least two elements is the “sense” strand or the “coding” strand. In genetics, the “sense” strand is the intra-double-stranded DNA segment extending from 5' to 3' that is complementary to the antisense or template strand of the DNA extending from 3' to 5'. Thus, for example, if an SNP nucleic acid base is on the 3' side of the promoter on the sense or coding strand, the SNP nucleic acid base is “downstream” of the promoter sequence in the genomic DNA (which is double-stranded).

[0244] Editing template The term “edit template” refers to the portion of the elongation arm on a single-stranded 3' DNA flap that encodes the desired edit, synthesized by a polymerase, e.g., DNA-dependent DNA polymerase or RNA-dependent DNA polymerase (e.g., reverse transcriptase). The embodiments described herein (for example, Figure 71A) refer to the “RT template” which refers to both the edit template and the homologous arm together, i.e., the sequence of the PEgRNA elongation arm that is actually used as a template during DNA synthesis. The term “RT edit template” is also equivalent to the term “DNA synthesis template,” where the RT edit template reflects the use of a prime editing factor with a polymerase that is a reverse transcriptase, while the DNA synthesis template more broadly reflects the use of a prime editing factor with either polymerase.

[0245] Effective amount As used in this application, the term “effective amount” means the amount of a bioactive agent sufficient to induce a desired biological response. For example, in some embodiments, the effective amount of a prime editing factor (PE) may mean the amount of editing factor sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, the effective amount of a fusion protein of a prime editing factor (PE) provided herein, e.g., comprising a nickase Cas9 domain and a reverse transcriptase, may mean the amount of fusion protein sufficient to induce editing of a target site specifically bound to and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, e.g., a fusion protein, nuclease, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or polynucleotide may vary depending on various factors, e.g., the desired biological response, e.g., a specific allele, genome, or target site to be edited, the cell or tissue to be targeted, and the agent to be used.

[0246] Reverse transcriptase prone to errors As used herein, the term “error-prone” reverse transcriptase (or more broadly, any polymerase) refers to a reverse transcriptase (or more broadly, any polymerase) that is naturally occurring or derived from another reverse transcriptase (e.g., wild-type M-MLV reverse transcriptase) having an error rate lower than that of wild-type M-MLV reverse transcriptase. The error rate of wild-type M-MLV reverse transcriptase has been reported to be between 15,000 (higher) and 27,000 (lower) for a 1 in 1 error. The error rate of 1 in 15,000 is 6.7 × 10⁻⁶. -5 This corresponds to an error rate of 1 in 27,000. -5 This corresponds to an error rate of 6.7 × 10⁻¹⁰. See Boutabout et al. (2001) "DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Ty1," Nucleic Acids Res 29(11):2217-2222 (which is incorporated herein by reference). Therefore, for the purposes of this application, the term "error-prone" means an error greater than 1 (6.7 × 10⁻¹⁰) in the incorporation of 15,000 nucleic acid bases. -5 Or higher), for example, 1 error in 14,000 nucleic acid bases (7.14 × 10⁻¹⁴). -5 (or higher), 1 error (7.7 × 10) within 13,000 nucleic acid bases or fewer. -5 (or higher), 1 error (7.7 × 10) within 12,000 nucleic acid bases or less. -5 (or higher), 1 error (9.1 × 10⁻¹⁰) within 11,000 nucleic acid bases or fewer. -5 (or higher), 1 error (1 × 10⁻¹⁰) within 10,000 nucleic acid bases or fewer. -4(or 0.0001 or higher), 1 error within 9,000 nucleic acid bases or less (0.00011 or higher), 1 error within 8,000 nucleic acid bases or less (0.00013 or higher), 1 error within 7,000 nucleic acid bases or less (0.00014 or higher), 1 error within 6,000 nucleic acid bases or less (0.00016 or higher), 1 error within 5,000 nucleic acid bases or less (0.0002 or higher), 1 error within 4,000 nucleic acid bases or less This refers to RTs that have an error rate of (0.00025 or higher), 1 error in 3,000 nucleic acid bases or less (0.00033 or higher), 1 error in 2,000 nucleic acid bases or less (0.00050 or higher), 1 error in 1,000 nucleic acid bases or less (0.001 or higher), 1 error in 500 nucleic acid bases or less (0.002 or higher), or 1 error in 250 nucleic acid bases or less (0.004 or higher).

[0247] Extain As used in this application, the term "extine" refers to a polypeptide sequence that is flanked by an intein and ligated to another extine during the protein splicing process to form a mature, spliced ​​protein. Typically, an intein is flanked by two extine sequences, which are ligated together when the intein catalyzes its own excision. An extine is therefore a protein analog to an exon found on mRNA. For example, a polypeptide containing an intein may have the structure extine(N)-intine-extine(C). After excision of the intein and splicing of the two extines, the resulting structure is extine(N)-extine(C) and free intein. In various configurations, the extines may be separate proteins (e.g., half of a Cas9 or PE fusion protein) each fused to a fission intein, and excision of the fission intein triggers splicing of the extine sequences together.

[0248] Extension arm The term “elongation arm” refers to a nucleotide sequence component of PEgRNA that provides several functions, including a primer-binding site and an editing template for reverse transcriptase. In some embodiments, for example, in Figure 3D, the elongation arm is located at the 3' end of the guide RNA. In other embodiments, for example, in Figure 3E, the elongation arm is located at the 5' end of the guide RNA. In some embodiments, the elongation arm also includes a homologous arm. In various embodiments, the elongation arm includes the following components in the 5'→3' direction: a homologous arm, an editing template, and a primer-binding site. Since the polymerization activity of reverse transcriptase is in the 5'→3' direction, the preferred arrangement of the homologous arm, editing template, and primer-binding site is in the 5'→3' direction so that the reverse transcriptase, once primed by the annealed primer sequence, uses the editing template as a complementary template strand to polymerize single-strand DNA. Further details, such as the length of the elongation arm, are described elsewhere in this specification.

[0249] The elongation arm may also be described as comprising two regions in general: a primer-binding site (PBS) and a DNA synthesis template, as illustrated in Figure 3G (top row). The primer-binding site is nicked by the prime-editing factor complex, so that when its 3' end is exposed to the nicked endogenous strand, it binds to the primer sequence formed from the endogenous DNA strand at the target site. As described herein, the binding of the primer sequence to the primer-binding site on the elongation arm of PEgRNA creates a double-stranded region with an exposed 3' end (i.e., 3' of the primer sequence), which then provides a substrate for polymerase to begin polymerizing a single strand of DNA along the length of the DNA synthesis template from the exposed 3' end. The sequence of the single-stranded DNA product is a complement to the DNA synthesis template. Polymerization continues to 5' of the DNA synthesis template (or elongation arm) until polymerization terminates. Therefore, the DNA synthesis template represents a portion of the elongation arm encoded by the polymerase of the prime editing factor complex into a single-strand DNA product (i.e., a 3' single-strand DNA flap containing the desired gene editing information), which replaces the corresponding endogenous DNA strand at the target site immediately downstream of the PE-induced nick site. While not bound by theory, polymerization of the DNA synthesis template continues to the 5' end of the elongation arm until a termination event occurs. Polymerization may terminate in various ways, including, but is not limited to, (a) reaching the 5' end of the PEgRNA (e.g., in the case of the 5' elongation arm, the DNA polymerase simply runs out of the template), (b) reaching an RNA secondary structure that cannot be passed through (e.g., a hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a signal on the nucleic acid morphology such as supercoiled DNA or RNA.

[0250] Flap end nucleases (for example, FEN1) As used herein, the term “flap endonuclease” refers to an enzyme that catalyzes the removal of a 5' single-strand DNA flap. These are naturally occurring enzymes that handle the removal of 5' flaps formed during cellular processes (including DNA replication). The prime editing methods described herein may utilize endogenously supplied flap endonucleases or those supplied trans to remove the 5' flap of endogenous DNA formed at the target site during prime editing. Flap endonucleases are known in the art and can be found in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10):4507-4519 and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2):198-211 and Balakrishnan et al., "Flap Endonuclease 1," Annu Rev Biochem, 2013, Vol 82:119-138 (each of these incorporated herein by reference). An exemplary flap endonuclease is FEN1, which can be represented by the following amino acid sequence: [Table 1]

[0251] functional equivalent The term “functional equivalent” refers to a second biomolecule that is functionally equivalent but structurally not equivalent to a first biomolecule. For example, a “Cas9 equivalent” refers to a protein that has the same or substantially the same function as Cas9 but does not necessarily have the same amino acid sequence. In the context of this disclosure, “protein X, or its functional equivalent” refers throughout this specification. In this regard, a “functional equivalent” of protein X accepts any homolog, paralog, fragment, naturally occurring version, modified version, mutant version, or synthetic version that carries the equivalent function of protein X.

[0252] Fusion protein The term “fusion protein,” as used herein, refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein may be positioned at the amino-terminus (N-terminus) or carboxy-terminus (C-terminus) of the fusion protein, thus forming an “amino-terminus fusion protein” or a “carboxy-terminus fusion protein,” respectively. The protein may contain different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9, which directs the protein toward binding to a target site) and a nucleic acid-cleaving domain or catalytic domain of a nucleic acid-editing protein. Another example includes Cas9 or its equivalent fused with reverse transcriptase. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference)).

[0253] Gene of interest (GOI) The term "GOI" refers to a gene that codes for a biomolecule of interest (e.g., a protein or RNA molecule). The protein of interest may include any intracellular, membrane, or extracellular protein, such as nuclear proteins, transcription factors, nuclear transporters, intracellular organelle-associated proteins, membrane receptors, catalytic proteins, and enzymes, therapeutic proteins, membrane proteins, membrane transport proteins, signaling proteins, or immunological proteins (e.g., IgG or other antibody proteins), etc. The gene of interest may also code for RNA molecules, including, but is not limited to, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), antisense RNA, guide RNA, microRNA (miRNA), small interfering RNA (siRNA), and cell-free RNA (cfRNA).

[0254] Guide RNA ("gRNA") As used herein, the term “guide RNA” mostly refers to a specific type of guide nucleic acid commonly associated with the Cas protein of CRISPR-Cas9, which binds to Cas9 and directs the Cas9 protein to a specific sequence in the DNA molecule (including complementarity with the protospace sequence of the guide RNA). However, the term also accepts equivalent guide nucleic acid molecules that bind to Cas9 equivalents, homologs, orthologues, or paralogs, whether naturally occurring or not (e.g., modified or recombinant), and that are separately programmed to localize the Cas9 equivalent to a specific target nucleotide sequence. The Cas9 equivalent may also include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), encompassing Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299) (this content is incorporated herein by reference). Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing appropriate guide RNA sequences are also provided herein. When used herein, “guide RNA” may also be referred to as “existing guide RNA” to distinguish it from a modified form of guide RNA called “prime editing guide RNA” (or “PEgRNA”) invented for the prime editing methods and compositions disclosed herein.

[0255] Guide RNA or PEgRNA may contain a variety of structural elements, but are not limited to, the following:

[0256] Spacer sequence - A sequence within a guide RNA or PEgRNA (approximately 20 nt in length) that binds to a protospacer in the target DNA.

[0257] The gRNA core (or gRNA backbone or main chain sequence) refers to the sequence within the gRNA responsible for Cas9 binding, and it does not include the spacer / targeting sequence used to guide Cas9 to the target DNA.

[0258] Elongation arm - single-strand elongation of the 3' or 5' end of PEgRNA. This includes a primer binding site and a DNA synthesis template sequence encoding a single-strand DNA flap containing the desired gene alteration by polymerase (e.g., reverse transcriptase). This is then incorporated onto endogenous DNA by replacing the corresponding endogenous strand, thereby incorporating the desired gene alteration.

[0259] Transcriptional terminator-guide RNA or PEgRNA may contain a transcription termination sequence at the 3' end of the molecule.

[0260] Homologous arms The term “homologous arm” refers to the portion of an elongation arm that codes for the resulting single-strand DNA flap (which is incorporated into the target DNA site by replacing the endogenous strand) encoded by the reverse transcriptase. The single-strand DNA flap portion coded by the homologous arm is complementary to the unedited strand of the target DNA sequence, facilitating its replacement for the endogenous strand and annealing with the single-strand DNA flap at that location, thereby incorporating the edit. This component is further defined elsewhere. By definition, since it is coded by the polymerase of the prime editing factor described herein, the homologous arm is part of the DNA synthesis template.

[0261] host cell When used herein, the term “host cell” refers to a cell capable of hosting, replicating, and expressing a vector described herein, for example, a vector comprising a nucleic acid molecule encoding a fusion protein including Cas9 or a Cas9 equivalent and reverse transcriptase.

[0262] Intein As used in this application, the term "intene" refers to a self-processing polypeptide domain found in organisms from all domains of life. Inteins (intermediate proteins) perform an intrinsic self-processing event known as protein splicing. In this process, they excise themselves from a larger precursor polypeptide by cleaving two peptide bonds and ligate extein (external protein) sequences that flank in the process by forming new peptide bonds. Since intein genes are found embedded in frame with other protein-coding genes, this rearrangement occurs post-translation (or co-translationally). Furthermore, intein-mediated protein splicing is spontaneous; it requires only the folding of the intein domain and no external factors or energy sources. This process is also known as cis-protein splicing, in contrast to the innate process of trans-protein splicing by "fission inteins." Inteins are protein equivalents of self-splicing RNA introns (see Perler et al., Nucleic Acids Res. 22:1125-1127 (1994)), and they catalyze their own excision from precursor proteins and have associated fusions with flanking protein sequences known as exteins (reviewed in Perler et al., Curr. Opin. Chem. Biol. 1:292-299 (1997); Perler, FBCell 92(1):1-4 (1998); Xu et al., EMBO J. 15(19):5146-5153 (1996)).

[0263] The term "protein splicing" as used in this application refers to the process in which the inner region (intein) of a precursor protein is cleaved, and the flanking regions (extines) of the protein are ligated together to form a mature protein. This natural process has been observed in numerous proteins from both prokaryotes and eukaryotes (Perler, FB, Xu, MQ, Paulus, H. Current Opinion in Chemical Biology 1997, 1, 292-299; Perler, FB Nucleic Acids Research 1999, 27, 346-347). An intein unit contains the necessary components required to catalyze protein splicing, and in many cases, it contains an endonuclease domain involved in intein mobility (Perler, FB, Davis, EO, Dean, GE, Gimble, FS, Jack, WE, Neff, N., Noren, CJ, Thomas, J., Belfort, M. Nucleic Acids Research 1994, 22, 1127-1127). However, the resulting protein is ligated and not expressed as a separate protein. Protein splicing can also occur in trans by fission inteins expressed on separate polypeptides. These spontaneously combine to form a single intein, which then undergoes the protein splicing process to ligate to a separate protein.

[0264] The elucidation of the mechanism of protein splicing has led to numerous intein-based applications (Comb, et al., U.S. Patent No. 5,496,714; Comb, et al., U.S. Patent No. 5,834,247; Camarero and Muir, J. Am. Chem. Soc., 121:5597-5598 (1999); Chong, et al., Gene, 192:271-281 (1997); Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al. J. Am. Chem. Soc., 121:1100-1101 (1999); Evans, et al. al.,J.Biol.Chem.,274:18359-18363(1999);Evans,et al.,J.Biol.Chem.,274:3923-3926(1999);Evans,et al.,Protein Sci.,7:2256-2264(1998);Evans,et al. al.,J.Biol.Chem.,275:9091-9094(2000);Iwai and Pluckthun,FEBS Lett.459:166-172(1999);Mathys,et al.,Gene,231:1-13(1999);Mills,et al.,Proc.Natl.Acad.Sci.USA 95:3543-3548(1998);Muir,et al.,Proc.Natl.Acad.Sci.USA 95:6705-6710(1998);Otomo,et al.,Biochemistry 38:16040-16044(1999);Otomo,et al.,J.Biolmol.NMR 14:105-114(1999);Scott,et al. al.,Proc.Natl.Acad.Sci.USA 96:13638-13643(1999);Severinov and Muir,J.Biol.Chem.,273:16205-16209(1998);Shingledecker,et al.,Gene,207:187-195(1998);Southworth,et al.,EMBO J.17:918-926(1998);Southworth,et al.,Biotechniques,27:110-120(1999);Wood,et al.,Nat.Biotechnol.,17:889-892(1999);Wu,et al.,Proc.Natl.Acad.Sci.USA 95:9226-9231(1998a);Wu,et al.,Biochim Biophys Acta 1387:422-432(1998b);Xu,et al.,Proc.Natl.Acad.Sci.USA 96:388-393(1999);Yamazaki,et al., J. Am. Chem. Soc., 120:5591-5592 (1998)). Each reference is incorporated herein by reference.

[0265] Ligand-dependent inteins As used in this application, the term "ligand-dependent intein" refers to an intein containing a ligand-binding domain. Typically, the ligand-binding domain is inserted into the amino acid sequence of the intein, resulting in a structural intein (N)-ligand-binding domain-intein (C). Typically, ligand-dependent inteins exhibit little to no protein splicing activity in the absence of a suitable ligand, and a significant increase in protein splicing activity in the presence of a ligand. In some embodiments, ligand-dependent inteins exhibit no observable splicing activity in the absence of a ligand, but exhibit splicing activity in the presence of a ligand. In some embodiments, ligand-dependent inteins exhibit observable protein splicing activity in the absence of a ligand, and in the presence of a suitable ligand, exhibit protein splicing activity that is at least 5 times, at least 10 times, at least 50 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, at least 500 times, at least 1000 times, at least 1500 times, at least 2000 times, at least 2500 times, at least 5000 times, at least 10000 times, at least 20000 times, at least 25000 times, at least 50000 times, at least 100000 times, at least 500000 times, or at least 1000000 times greater than the activity observed in the absence of the ligand. In some embodiments, the increase in activity is dose-dependent by at least one order of magnitude, at least two orders of magnitude, at least three orders of magnitude, at least four orders of magnitude, or at least five orders of magnitude, allowing for fine-tuning of the intein activity by adjusting the ligand concentration.Suitable ligand-dependent inteins are known in the art and are provided below and published in U.S. Patent Application No. US2014 / 0065711A1; Mootz et al., "Protein splicing triggered by a small molecule." J.Am.Chem.Soc. 2002; 124, 9044-9045; Mootz et al., "Conditional protein splicing: a new tool to control protein structure and function in vitro and in vivo." J.Am.Chem.Soc. 2003; 125, 10561-10569; Buskirk et al., Proc.Natl.Acad.Sci.USA. 2004; 101, 10505-10510); Skretas & Wood, "Regulation of protein activity with small-molecule-controlled inteins." Protein This includes what is described in Sci.2005;14,523-532;Schwartz, et al., "Post-translational enzyme activation in an animal via optimized conditional protein splicing."Nat.Chem.Biol.2007;3,50-54;Peck et al., Chem.Biol.2011;18(5),619-630; the entire contents of each are incorporated here by reference. An example sequence is as follows: [Table 2-1] [Table 2-2]

[0266] Linker As used in this application, the term "linker" refers to a molecule that links two other molecules or parts. A linker can be an amino acid sequence in the case of a linker linking two fusion proteins. For example, Cas9 can be fused to reverse transcriptase via an amino acid linker sequence. A linker can also be a nucleotide sequence in the case of linking two nucleotide sequences together. For example, in this case, a conventional guide RNA is linked to the RNA elongation of a prime editing guide RNA, which may include an RT template sequence and an RT primer binding site, via a spacer or linker nucleotide sequence. In other embodiments, a linker can be an organic molecule, a group, a polymer, or a chemical part. In some embodiments, the linker is 5-100 amino acids long, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids long. Longer or shorter linkers are also intended.

[0267] isolated "Isolated" means modified or removed from its natural state. For example, a nucleic acid or peptide that is naturally present in a living animal is not "isolated," but the same nucleic acid or peptide that has been partially or completely separated from its natural state in the coexisting material is "isolated." Isolated nucleic acids or proteins may exist in a substantially purified form or in a non-native environment (e.g., a host cell).

[0268] In some embodiments, the gene of interest is encoded by an isolated nucleic acid. As used herein, the term “isolated” refers to the characteristic of a material as provided herein that it has been removed from its original or native environment (e.g., the natural environment if it is naturally occurring). Thus, a naturally occurring polynucleotide or protein or polypeptide present in a living animal is not isolated, but the same polynucleotide or polypeptide that has been separated from some or all of the coexisting material in a natural system by human intervention is isolated. Artificial or modified materials, such as nucleic acid constructs that do not exist naturally, including the expression constructs and vectors described herein, are also consequently referred to as isolated. The material does not need to be purified to be isolated. Consequently, the material may be part of a vector and / or part of a composition, and such vector or composition is still isolated in that it is not part of the environment in which the material is found in nature.

[0269] MS2 Tagging Technology In various embodiments (for example, as illustrated in Figures 72-73 and the embodiments of Example 19), the term “MS2 tagging technique” refers to a combination of an “RNA-protein interaction domain” (also known as an “RNA-protein mobilization domain or protein”) paired with an RNA-binding protein that specifically recognizes and binds to a particular hairpin structure, e.g., a specific hairpin structure. These types of systems can be utilized to mobilize various functionalities to a prime editing factor complex bound to a target site. MS2 tagging techniques are based on the innate interaction of the MS2 bacteriophage coat protein (“MCP” or “MS2cp”) with stem-loop or hairpin structures present on the phage genome, i.e., “MS2 hairpins.” In the case of prime editing, MS2 tagging techniques involve introducing an MS2 hairpin onto a desired RNA molecule involved in prime editing (e.g., PEgRNA or tPERT). This then constitutes a specific interactable binding target for an RNA-binding protein that recognizes and binds to its structure. In the case of the MS2 hairpin, it is recognized and bound by the MS2 bacteriophage coat protein (MCP). Then, when the MCP is fused to another protein (e.g., reverse transcriptase or other DNA polymerase), the MS2 hairpin can be used to trans-"mobilize" the other protein to the target site occupied by the prime editing complex.

[0270] The prime editing factors described herein may, as an aspect, incorporate any known RNA-protein interaction domain to recruit or “colocalize” a specific function of interest to the prime editing factor complex. Reviews of other modular domains of RNA-protein interactions are described in the art, for example, in Johansson et al., "RNA recognition by the MS2 phage coat protein," Sem Virol., 1997, Vol.8(3):176-185; Delebecque et al., "Organization of intracellular reactions with rationally designed RNA assemblies," Science, 2011, Vol.333:470-474; Mali et al., "Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering," Nat. Biotechnol., 2013, Vol.31:833-838; and Zalatan et al., "Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds," Cell, 2015, Vol.160:339-350 (each of these is incorporated herein by reference in its entirety). Other systems also include PP7 hairpins that specifically recruit PCP proteins and "com" hairpins that specifically recruit Com proteins. See Zalatan et al.

[0271] The nucleotide sequence of the MS2 hairpin (i.e., referred to as the "MS2 aptamer") is as follows: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (Sequence ID 763).

[0272] The amino acid sequence of MCP or MS2cp is as follows: GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (Sequence ID 764).

[0273] The MS2 hairpin (or "MS2 aptamer") may also be referred to as a type of "RNA effector recruitment domain" (or equivalently, an "RNA-binding protein recruitment domain" or simply a "recruitment domain") because it is a physical structure (e.g., a hairpin) incorporated onto PEgRNA or tPERT that effectively recruits other effector functions (e.g., RNA-binding proteins with various functions, such as DNA polymerase or other DNA-modifying enzymes) to the thus modified PEgRNA or rPERT, and thus co-localizes the effector function in trans to the prime editing mechanism. This application is not intended to be limited to any specific RNA effector recruitment domain, but may encompass any available such domain that includes the MS2 hairpin. Example 19 and Figure 72(b) illustrate the use of a prime editing factor containing an MS2 aptamer linked to the DNA synthesis domain (i.e., the tPERT molecule) and an MS2cp protein fused to PE2 to achieve colocalization between the prime editing factor complex (MS2cp-PE2:sgRNA complex) bound to a target DNA site and the DNA synthesis domain of the tPERT molecule.

[0274] napDNAbp As used herein, the term “nucleic acid programmed DNA-binding protein” or “napDNAbp” (with Cas9 being an example) refers to a protein that uses RNA:DNA hybridization to target and bind to a specific sequence in a DNA molecule. Each napDNAbp is bound to at least one guide nucleic acid (e.g., guide RNA) that localizes the napDNAbp to a DNA sequence containing a DNA strand (i.e., the target strand) complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of guide RNA). In other words, the guide nucleic acid “programs” the napDNAbp (e.g., Cas9 or an equivalent) to localize and bind to the complementary sequence.

[0275] While not strictly theoretical, the binding mechanism of the napDNAbp-guide RNA complex generally involves the step where napDNAbp forms an R-loop that induces unwinding of the double-stranded DNA target (thus separating the strands in the region bound by napDNAbp). The guide RNA protospacer then hybridizes with the "target strand," replacing the complementary "non-target strand," thereby forming the single-stranded region of the R-loop. In some embodiments, napDNAbp contains one or more nuclease activities, which then cleave the DNA to eliminate various types of lesions. For example, napDNAbp may contain nuclease activity that cleaves the non-target strand at a first site and / or the target strand at a second site. Depending on the nuclease activity, the target DNA may be cleaved, forming a "double-strand break" where both strands are cut. In other embodiments, the target DNA can only be cleaved at a single site; i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with various nuclease activities include “Cas9 nickase” (“nCas9”) and inactivated Cas9 (“inactive Cas9” or “dCas9”) which have no nuclease activity whatsoever. Exemplary sequences of these and other napDNAbps are provided herein.

[0276] Nikkaze The term "nickase" refers to Cas9 with one of its two inactivated nuclease domains. This enzyme is capable of cleaving only one strand of target DNA.

[0277] Nuclear localization sequence (NLS) The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that facilitates the import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, an NLS sequence is described in Plank et al.’s international PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published May 31, 2001, as WO / 2001 / 038547 (this content is incorporated herein by reference with respect to its disclosure of exemplary nuclear localization sequences). In some embodiments, an NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 16) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 17).

[0278] nucleic acid molecule The term “nucleic acid” as used herein refers to polymers of nucleotides. Polymers include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (for example, 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyluridine, C5 propynylcytidine, C5 methylcytidine, 7 deazaadenosine, 7 deazaguanosine, 8 oxoadenosine, 8 oxoguanosine, O(6) methylguanine, 4-acetylcytidine) These may include 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N phosphoramidite linkages).

[0279] PEgRNA As used herein, the terms “prime editing guide RNA,” “PEgRNA,” or “extended guide RNA” refer to a specialized form of guide RNA modified to include one or more additional sequences for performing the prime editing methods and compositions described herein. As described herein, prime editing guide RNA includes one or more “extended regions” of nucleic acid sequences. The extended regions may include, but are not limited to, single-stranded RNA or DNA. Furthermore, the extended region may occur at the 3' end of the existing guide RNA. In other configurations, the extended region may occur at the 5' end of the existing guide RNA. In other configurations, the extended region may occur in a region within the terminal molecule of the existing guide RNA (e.g., in a gRNA core region that binds to and / or ligates to a napDNAbp). The extended region contains a “DNA synthesis template” that encodes single-stranded DNA (by the polymerase of a prime editing factor), but the molecule is then designed to be (a) homologous to the endogenous target DNA to be edited, and (b) contain at least one desired nucleotide change (e.g., a transition, transversion, deletion, or insertion) to be induced into or incorporated into the endogenous target DNA. The extended region may also contain other functional sequence elements (but not limited to “primer binding sites” and “spacer or linker” sequences) or other structural elements (but not limited to aptamers, stem-loops, hairpins, toe-loops (e.g., a 3' toe-loop), or RNA-protein recruitment domains (e.g., an MS2 hairpin)). When used herein, “primer binding sites” include a sequence that hybridizes with a single-stranded DNA sequence having a 3' end generated from R-loop nicked DNA.

[0280] In one embodiment, PEgRNA is represented by Figure 3A, showing PEgRNA having a 5' extension arm, a spacer, and a gRNA core. The 5' extension further includes a reverse transcriptase template, a primer binding site, and a linker in the 5'→3' direction. As shown, the reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editing factor described herein is a different type of polymerase than RT.

[0281] In one other embodiment, PEgRNA is represented by Figure 3B, showing PEgRNA having a 3' extension arm, a spacer, and a gRNA core. The 3' extension further includes a reverse transcriptase template and a primer binding site in the 5'→3' direction. As shown, the reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editing factor described herein is a different type of polymerase than RT.

[0282] In yet another embodiment, PEgRNA is represented by Figure 3D, showing PEgRNA having a spacer (1), a gRNA core (2), and an elongation arm (3) in the 5'→3' direction. The elongation arm (3) is located at the 3' end of the PEgRNA. The elongation arm (3) further includes a "primer binding site" (A), an "editing template" (B), and a "homologous arm" (C) in the 5'→3' direction. The elongation arm (3) may also include any modification regions at the 3' and 5' ends, which may be the same sequence or different sequences. In addition, the 3' end of the PEgRNA may include a transcription terminator sequence. These sequence elements of PEgRNA are further described and defined herein.

[0283] In another embodiment, PEgRNA is represented by Figure 3E, which shows PEgRNA having an elongation arm (3), a spacer (1), and a gRNA core (2) in the 5'→3' direction. The elongation arm (3) is located at the 5' end of the PEgRNA. The elongation arm (3) further includes a "primer binding site" (A), an "editing template" (B), and a "homologous arm" (C) in the 3'→5' direction. The elongation arm (3) may also include any modification regions at the 3' and 5' ends, which may be the same sequence or different sequences. The 3' end of the PEgRNA may include a transcription terminator sequence. These sequence elements of PEgRNA are further described and defined herein.

[0284] PE1 In this application, "PE1" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and wild-type MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)]+desired PEgRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 123, which is shown below; [ka] [ka]

[0285] PE2 In this application, "PE2" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and the variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+desired PEgRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 134, which is shown as follows: [ka] [ka]

[0286] PE3 In this application, "PE3" refers to PE2, plus a second-strand nicking guide RNA that complexes with PE2, and introduces a nick onto the unedited DNA strand to induce preferential replacement of the strand to be edited.

[0287] PE3b In this application, "PE3b" refers to PE3, but the second strand nicking guide RNA is designed for temporal control so that the second strand nick is not introduced until after the desired edit has been incorporated. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand and not the original allele. Using this strategy, hereafter referred to as PE3b, the mismatch between the protospacer and the unedited allele should not favor nicking by the sgRNA until after the editing event on the PAM strand occurs.

[0288] Short PE (PE-short) The term used in this application Short PE " refers to the PE construct fused to the C-terminal truncated reverse transcriptase, and has the following amino acid sequence: [ka] [ka]

[0289] Peptide tags The term "peptide tag" refers to a peptide amino acid sequence that genetically fuses with a protein sequence to confer one or more functions to the protein, thereby facilitating the manipulation of proteins for various purposes such as visualization, purification, solubilization, and separation. Peptide tags can encompass various types of tags classified by purpose or function, and may include "affinity tags" (to facilitate protein purification), "solubilization tags" (to assist in the correct folding of proteins), "chromatographic tags" (to alter the chromatographic properties of proteins), "epitope tags" (to bind to high-affinity antibodies), and "fluorescence tags" (to facilitate the visualization of proteins in cells or in vitro).

[0290] polymerase As used herein, the term “polymerase” refers to an enzyme that synthesizes nucleotide chains, which may be used in relation to the prime editing factor systems described herein. A polymerase may be a “template-dependent” polymerase (i.e., a polymerase that synthesizes nucleotide chains based on the order of nucleotide bases of a template chain). A polymerase may also be a “template-independent” polymerase (i.e., a polymerase that synthesizes nucleotide chains without the requirement of a template chain). A polymerase may further be classified as a “DNA polymerase” or an “RNA polymerase”. In various embodiments, the prime editing factor system includes a DNA polymerase. In various embodiments, the DNA polymerase may be a “DNA-dependent DNA polymerase” (i.e., the template molecule is a strand of DNA). In such cases, the DNA template molecule may be PEgRNA, and the extension arm includes a strand of DNA. In such cases, PEgRNA may be called a chimeric or hybrid PEgRNA, which comprises an RNA portion (i.e., a guide RNA component encompassing a spacer and a gRNA core) and a DNA portion (i.e., an elongation arm). In various other embodiments, DNA polymerase may be an "RNA-dependent DNA polymerase" (i.e., the template molecule is a strand of RNA). In such cases, PEgRNA is RNA, i.e., encompasses RNA elongation. The term "polymerase" may also refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Generally, the enzyme will initiate synthesis at the 3' end of a primer annealed to a polynucleotide template sequence (e.g., a primer sequence annealed to the primer-binding site of PEgRNA, for example) and proceed toward the 5' end of the template strand. "DNA polymerase" catalyzes the polymerization of deoxynucleotides. The term DNA polymerase as used in this application with reference to DNA polymerase encompasses "its functional fragment". "The functional fragment" refers to a portion of either wild-type or mutant DNA polymerase that encapsulates less than the entire amino acid sequence of the polymerase, while retaining the ability to catalyze polynucleotide polymerization under at least one set of conditions.Such functional fragments may exist as separate entities, or they may be components of larger polypeptides such as fusion proteins.

[0291] Prime editing and multiple flap prime editing factors As used herein, the terms “prime editing” or “classical prime editing” refer to a novel approach to gene editing that utilizes a specialized guide RNA containing napDNAbp, polymerase (e.g., reverse transcriptase), and a DNA synthesis template for encoding (or deleting) desired new genetic information to be incorporated into a target DNA sequence. Certain aspects of prime editing are described in other figures, specifically in Figures 1A–1H and 72(a)–72(c). Classical prime editing is described in our publication Anzalone, A.V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157 (2019) (which is incorporated herein by reference in its entirety).

[0292] Prime editing represents a platform for genome editing, a versatile and precise genome editing method that directly writes new genetic information to a defined DNA site, using a nucleic acid-programmable DNA-binding protein ("napDNAbp") that works in conjunction with polymerase (i.e., provided in the form of a fusion protein or otherwise in trans with napDNAbp). The prime editing system is programmed by prime editing (PE) guide RNA ("PEgRNA"), which defines the target site and serves as a template for the synthesis of the desired edit in the form of a replacement DNA strand by an manipulated extension (either DNA or RNA) on the guide RNA (e.g., at the 5' or 3' end or in the interior of the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleic acid substitution) shares (or is homologous to) the same sequence as the endogenous strand of the target site to be edited (immediately downstream of the nick site) (with the exception that it contains the desired edit). By DNA repair and / or replication mechanisms, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a “search and replace” genome editing technology, because the prime editing factors described herein not only search for and locate the desired target site to be edited, but also simultaneously encode a replacement strand containing the desired edit that is incorporated in place of the endogenous DNA strand at the corresponding target site. The prime editing factors of this disclosure relate, in part, to the discovery that the mechanism of primed reverse transcription (TPRT) or “prime editing” to a prime target can be utilized or adapted to perform precise CRISPR / Cas-based genome editing with high efficiency and gene flexibility (as illustrated, for example, in various aspects of Figure 1A–1F). In nature, TPRT is used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial group II introns. 28,29In this application, the inventors used a Cas protein-reverse transcriptase fusion or related system to target a specific DNA sequence with guide RNA, generate a single-stranded nick at the target site, and use the nicked DNA as a primer for reverse transcription of the manipulated reverse transcriptase template incorporated into the guide RNA. However, while the concept begins with a prime editing factor that uses reverse transcriptase as a DNA polymerase component, the prime editing factors described herein are not limited to reverse transcriptase and may substantially encompass the use of any DNA polymerase. Indeed, although this application may refer to a prime editing factor having “reverse transcriptase” everywhere, it is hereby made clear that reverse transcriptase is only one type of DNA polymerase that can function in prime editing. Therefore, whenever this specification refers to “reverse transcriptase,” those skilled in the art should understand that any suitable DNA polymerase may be used instead of reverse transcriptase. Therefore, in one aspect, the prime editing factor may include Cas9 (or equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., PEgRNA) containing a spacer sequence that anneals to a complementary protospacer on the target DNA. The specialized guide RNA also contains new genetic information in the form of an elongation encoding a replacement strand of DNA containing the desired genetic alteration, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer the information from the PEgRNA to the target DNA, the prime editing mechanism involves nicking a target site on one strand of DNA to expose a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime the DNA polymerization of the edit-encoding elongation on the PEgRNA to the target site. In various embodiments, the elongation providing a template for the polymerization of the edit-containing replacement strand may be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase. The newly synthesized strand formed by the prime editing factors disclosed herein (i.e., the replacement DNA strand containing the desired edit) will be homologous to the genomic target sequence (i.e., have the same sequence) with the exception of the inclusion of the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be called a single-stranded DNA flap. This will compete for hybridization with a complementary homologous endogenous DNA strand, thereby replacing the corresponding endogenous strand. In some embodiments, the system may be combined with the use of an error-prone reverse transcriptase enzyme (e.g., provided as a fusion protein with a Cas9 domain, or provided trans to a Cas9 domain). The error-prone reverse transcriptase enzyme may introduce alterations during the synthesis of the single-stranded DNA flap. Therefore, in some embodiments, the error-prone reverse transcriptase enzyme may be used to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used in the system, the changes can be random or non-random. Degradation of the hybridized intermediate (including a single-stranded DNA flap synthesized by a reverse transcriptase hybridized to an endogenous DNA strand) may encompass removal of the replaced flap resulting from the endogenous DNA (e.g., by the 5' end DNA flap endonuclease FEN1), ligation of the synthesized single-stranded DNA flap onto the target DNA, and assimilation of the desired nucleotide changes as a result of cellular DNA repair and / or replication processes. Since template-based DNA synthesis offers one-nucleotide precision for any nucleotide modification, including insertions and deletions, the scope of this approach is extremely broad, and foreseeably, it could have countless applications in basic research and therapeutics.

[0293] In various embodiments, prime editing operates by bringing a target DNA molecule (to which a nucleotide sequence change is to be introduced) into contact with a nucleic acid programmed DNA-binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). Referring to Figure 1G, the prime editing guide RNA (PEgRNA) includes an extension at the 3' or 5' end of the guide RNA or at an intramolecular position of the guide RNA, encoding the desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex comes into contact with the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced into one of the strands of DNA at the target locus (e.g., by a nuclease or chemical agent), thereby generating an available 3' end on one of the strands at the target locus. In one embodiment, the nick is generated on the DNA strand corresponding to the R-loop strand, i.e., the strand that does not hybridize to the guide RNA sequence, i.e., the “non-target strand”. However, the nick can be introduced on either strand. That is, the nick can be introduced on the R-loop “target strand” (i.e., the strand that hybridizes to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand that forms the single-stranded portion of the R-loop that is complementary to the target strand). In step (c), the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime the reverse transcription (i.e., “RT primed to the prime target”). In one embodiment, the 3' end DNA strand hybridizes to a specific RT prime sequence on the extended portion of the guide RNA, i.e., the “reverse transcriptase prime sequence” or “primer binding site” on the PEgRNA. In step (d), the reverse transcriptase (or other preferred DNA polymerase) is introduced. This synthesizes a single strand of DNA from the 3' end of the primed region toward the 5' end of the prime editing guide RNA. DNA polymerase (e.g., reverse transcriptase) can be fused to napDNAbp, or alternatively, provided trans to napDNAbp.This forms a single-stranded DNA flap containing a desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion, or a combination thereof) that is otherwise homologous to the endogenous DNA at or adjacent to the nicking site. In step (e), napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-stranded DNA flap, resulting in the integration of the desired nucleotide change into the target locus. This process can be driven toward the formation of the desired product by removing the corresponding 5' endogenous DNA flap that forms once the 3' single-stranded DNA flap invades and hybridizes with the endogenous DNA sequence. Without being constrained by theory, the cell's endogenous DNA repair and replication processes degrade mismatched DNA and incorporate nucleotide changes (one or more) to form the desired modified product. The process can also be driven toward product formation by "second-strand nicking," as illustrated in Figure 1F. This process can introduce at least one of the following gene changes: transversion, transition, deletion, and insertion.

[0294] The terms “prime editing factor (PE) system” or “prime editing factor (PE)” or “PE system” or “PE editing system” refer to compositions relating to genome editing methods using primed reverse transcription (TPRT) to a prime target as described herein, and include, but are not limited to, napDNAbp, reverse transcriptase, fusion proteins (e.g., napDNAbp and reverse transcriptase), prime editing guide RNA, and complexes comprising the fusion protein and prime editing guide RNA, as well as auxiliary elements, such as second-strand nicking components (e.g., second-strand sgRNA) and 5' endogenous DNA flap removal endonucleases (e.g., FEN1) to help drive the prime editing process toward the formation of the edited product.

[0295] In the embodiments described so far, PEgRNA constitutes a single molecule comprising a guide RNA (which itself contains a spacer sequence and a gRNA core or backbone) and a 5' or 3' elongation arm containing a primer binding site and a DNA synthesis template (see, for example, Figure 3D). However, PEgRNA can also take the form of two individual molecules, consisting of a guide RNA and a transprime editing factor RNA template (tPERT). This essentially houses the elongation arm (in particular, containing the primer binding site and DNA synthesis domain) and the RNA-protein recruitment domain (e.g., an MS2 aptamer or hairpin) on the same molecule. This then colocalizes or recruits to a modified prime editing factor complex containing a tPERT recruiting protein (e.g., an MS2cp protein that binds to an MS2 aptamer). See Figures 3G and 3H for examples of tPERTs that can be used for prime editing.

[0296] In the "double flap prime editing system," two pegRNAs are used to target opposing strands of a genomic site, leading to the synthesis of two complementary 3' flaps containing the edited DNA sequence (Figure 91). Unlike classical prime editing, there is no requirement that the pair of edited DNA strands (3' flaps) directly compete with the 5' flap on the endogenous genomic DNA, because the complementary edited strands are available for hybridization instead. Since both strands of the double helix are synthesized as edited DNA, the double flap prime editing system eliminates the need to replace the unedited complementary DNA strand required by classical prime editing. Instead, cellular DNA repair mechanisms only need to excise the paired 5' flap (original genomic DNA) and ligate the paired 3' flap (edited DNA) to the locus. Therefore, it is not necessary to include sequences homologous to the genomic DNA on the newly synthesized DNA strand, allowing for selective hybridization of the new strand and facilitating editing that contains minimal genomic homology. Nuclease-activated versions of prime editing factors that cut both strands of DNA can also be used to accelerate the removal of the original DNA sequence.

[0297] Like classical prime editing, multi-flap prime editing (including double-flap and quadruple-flap prime editing) is a versatile and precise genome editing method. It uses a nucleic acid-programmable DNA-binding protein ("napDNAbp") that works in conjunction with polymerase (i.e., provided in the form of a fusion protein with napDNAbp, or otherwise in trans) to directly write new genetic information to a defined DNA site, and the prime editing system is programmed by prime editing (PE) guide RNA ("PEgRNA"). This both defines the target site and serves as a template for the synthesis of the desired edit in the form of a replacement DNA strand, through an manipulated extension (either DNA or RNA) on the guide RNA (e.g., at the 5' or 3' end, or within the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleic acid substitution) shares the same sequence as the endogenous strand of the target site to be edited (with the exception that it contains the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand of the target site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a “search-and-replace” genome editing technology, because the prime editing factors described herein not only search for and locate the desired target site to be edited, but also simultaneously encode a replacement strand containing the desired edit, which is then incorporated in place of the endogenous DNA strand at the corresponding target site.

[0298] Prime editing factor The dual prime editing system described herein comprises a pair of prime editing factors. The term “prime editing factor” refers to the fusion construct described herein, comprising napDNAbp (e.g., Cas9 nickase) and reverse transcriptase, which is capable of performing prime editing on a target nucleotide sequence in the presence of PEgRNA (or “extended guide RNA”). The term “prime editing factor” may also refer to a fusion protein, or a fusion protein complexed with PEgRNA, and / or a fusion protein further complexed with sgRNA to nick the second strand. In some embodiments, the prime editing factor may also refer to a complex comprising a fusion protein (reverse transcriptase fused with napDNAbp), PEgRNA, and regular guide RNA, which is capable of proceeding to the nicking step to the second site of the unedited strand as described herein. In other embodiments, the reverse transcriptase component of the “primer editing factor” may be provided in trans.

[0299] The dual-flap prime editing system described in this application includes a pair of prime editing factors. The quadruple-flap prime editing system described in this application includes four prime editing factors.

[0300] Primer binding site The term "primer binding site" or "PBS" refers to a nucleotide sequence located on PEgRNA as a component of the elongation arm (typically at the 3' end of the elongation arm) that helps bind to a primer sequence formed after Cas9 nicking of the target sequence by a prime editing factor. As detailed elsewhere, when the Cas9 nickase component of the prime editing factor nicks one strand of the target DNA sequence, a 3'-end ssDNA flap is formed, which anneals to the primer binding site on the PEgRNA and acts as a primer sequence that primes reverse transcription.

[0301] promoter The term “promoter” is recognized in the art and refers to a nucleic acid molecule having a sequence that is recognized by a cellular transcription mechanism and has the ability to initiate the transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is active only in the presence of specific conditions. For example, a conditional promoter may be active only in the presence of a specific protein that connects a protein associated with a regulatory element on the promoter to the basic transcription mechanism, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters is inducible promoters, which require the presence of a small molecule “inducer” for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. Various constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to identify various such promoters useful for carrying out the present invention. This is not limited to this.

[0302] Protospacer As used in this application, the term "protospacer" refers to a sequence (~20 bp) on DNA adjacent to a PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand of it, i.e., the "target strand" relative to the "non-target strand" of the target DNA sequence). For Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence called NGG on the non-target strand, found directly downstream of the target sequence on the genomic DNA. Those skilled in the art will understand that the literature of the current art sometimes refers to "protospacer" as a ~20 nt target-specific guide sequence on the guide RNA itself, rather than simply calling it a "spacer." Therefore, in some cases, the term "protospacer" as used in this application may be interchangeable with the term "spacer." The context surrounding the appearance of either "protospacer" or "spacer" will help inform the reader whether the term refers to a gRNA or DNA target.

[0303] Protospacer adjacent motif (PAM) As used in this application, the term “protospacer adjacent sequence” or “PAM” refers to a DNA sequence of approximately 2–6 base pairs that is a key targeting component of the Cas9 nuclease. Typically, the PAM sequence is located on either strand and downstream in the 5'→3' direction of the site cut by Cas9. A standard PAM sequence (i.e., the PAM sequence associated with the Cas9 nuclease or SpCas9 of Streptococcus pyogenes) is 5'-NGG-3', where “N” is any nucleic acid base followed by two guanine (”G") nucleic acid bases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, may be modified to alter its PAM specificity by causing the nuclease to recognize alternative PAM sequences.

[0304] For example, with reference to the standard SpCas9 amino acid sequence of SEQ ID NO: 18, the PAM sequence can be modified by introducing one or more mutations, including (a) D1135V, R1335Q, and T1337R "VQR variants" that alter PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q, and T1337R "EQR variants" that alter PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E, and T1337R "VRER variants" that alter PAM specificity to NGCG. In addition, the D1135E variant of standard SpCas9 still recognizes NGG, but it is more selective compared to the wild-type SpCas9 protein.

[0305] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have varying PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. And yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and not meant to be limiting. Furthermore, it will be understood that non-SpCas9s bind to a variety of PAM sequences. This makes them useful when a suitable SpCas9 PAM sequence is not present at the desired target's cut site. Moreover, non-SpCas9s may possess other characteristics that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is about 1 kilobase smaller than SpCas9. Therefore, it can be packaged in adeno-associated viruses (AAV). Further reference can be made to Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5):891-899 (which is incorporated into this application by reference).

[0306] Recombinase As used in this application, the term "recombinase" refers to a site-specific enzyme that mediates the recombination of DNA between recombinase-recognition sequences, resulting in the excision, incorporation, inversion, or exchange (e.g., translocation) of DNA fragments between recombinase-recognition sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvers and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, without limitation, Hin, Gin, Tn3, β-six, CinH, ParA, γδ, Bxb1, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, without limitation, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. The names serine and tyrosine recombinases are derived from the conserved nucleophilic amino acid residues that recombinases use to attack DNA and which become covalently ligated to the DNA during strand exchange. Recombinases have numerous applications, including gene knockout / knock-in production and gene therapy applications. For example, Brown et al.,"Serine recombinases as tools for genome engineering."Methods.2011;53(4):372-9;Hirano et al.,"Site-specific recombinases as tools for heterologous gene integration."Appl.Microbiol.Biotechnol.2011;92(2):227-39;Chavez and Calos,"Therapeutic applications of the ΦC31 integrase system."Curr.Gene Ther.2011;11(5):375-81;Turan and Bode,"Site-specific recombinases:from tag-and-target- to tag-and-exchange-based genomic modifications."FASEB J.2011;25(12):4088-107;Venken and Bellen,"Genome-wide manipulations of Drosophila melanogaster with transposons, Flp recombinase, and ΦC31 integrase."Methods Mol.Biol.2012;859:203-28;Murphy,"Phage recombinases and their applications."Adv.Virus Res.2012;83:367-414;Zhang et al.,"Conditional gene manipulation: Cre-ating a new biological era."J.Zhejiang Univ.Sci.B.2012;13(7):511-24;Karpenshif and Bernstein,"From yeast to mammals:recent advances in genetic control of homologous recombination."DNA See Repair(Amst).2012;1;11(10):781-8 (the full contents of each are incorporated herein by reference). The recombinases provided herein are not intended to be exclusive examples of recombinases that may be used in embodiments of the present invention. The methods and compositions of the present invention may be extended by mining a database of novel orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity (for example, Groth et al., "Phage integrases: biology and applications." J.Mol.Biol.2004;335,667-678; Gordley et al., "Synthesis of programmable integrases." Proc.Natl.Acad.Sci.US A.See 2009;106,5053-5058 (the full contents of each are incorporated herein by reference). Other examples of recombinases useful for the methods and compositions described herein are known to those skilled in the art, and it is expected that any new recombinases discovered or produced will be capable of being used in different embodiments of the present invention. In some embodiments, the catalytic domain of the recombinase is fused to a nuclease (e.g., dCas9 or a fragment thereof) that is programmable by RNA inactivating the nuclease, and as a result, the recombinase domain does not contain a nucleic acid binding domain or is incapable of binding to a target nucleic acid (e.g., the recombinase domain is manipulated so that it does not have specific DNA binding activity). Recombinases lacking DNA binding activity and methods for modifying them are known, including Klippel et al., "Isolation and characterization of unusual gin mutants." EMBO J. 1988; 7: 3983-3989; Burke et al., "Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation." Mol Microbiol. 2004; 51: 937-948; Olorunniji et al., "Synapsis and catalysis by activated Tn3 resolvase mutants." Nucleic Acids Res. 2008; 36: 7181-7191; Rowland et al., "Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome." Mol Microbiol. 2009; 74: 282-298; Akopian et al., "Chimeric recombinases with designed DNA sequence recognition."Proc Natl Acad Sci USA.2003;100:8688-8691;Gordley et al.,"Evolution of programmable zinc finger-recombinases with activity in human cells.J Mol Biol.2007;367:802-813;Gordley et al.,"Synthesis of programmable integrases."Proc Natl Acad Sci USA.2009;106:5053-5058;Arnold et al.,"Mutants of Tn3 resolvase which do not require accessory binding sites for recombination activity."EMBO J.1999;18:1407-1414;Gaj et al.,"Structure-guided reprogramming of serine recombinase DNA sequence specificity."Proc Natl Acad Sci USA.2011;108(2):498-503; and Proudfoot et al.,"Zinc finger recombinases This includes what is described in "with adaptable DNA sequence specificity." PLoS One. 2011;6(4):e19537 (the full contents of each are incorporated herein by reference). For example, serine recombinases of the resolverase-invertase group, such as Tn3 and γδ resolvers, and Hin and Gin invertases, have molecular structures with autonomous catalytic and DNA-binding domains (for example, Grindley et al., "Mechanism of site-specific recombination." Ann Rev Biochem.See 2006;75:567-605 (its entire contents are incorporated by reference). Therefore, after isolating "activated" recombinase mutants that do not require any auxiliary factors (e.g., DNA binding activity), the catalytic domains of these recombinases are compliant with programmable nucleases (e.g., dCas9 or its fragments) by RNA that has inactivated the nuclease, as described herein (see, for example, Klippel et al., "Isolation and characterisation of unusual gin mutants." EMBO J.1988;7:3983-3989; Burke et al., "Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation. Mol Microbiol.2004;51:937-948; Olorunniji et al., "Synapsis and catalysis by activated Tn3 resolvase mutants." Nucleic Acids Res.2008;36:7181-7191; Rowland et al., "Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome."Mol Microbiol.2009;74:282-298;Akopian et al.,"Chimeric recombinases with designed DNA sequence recognition."Proc Natl Acad Sci USA.(See 2003;100:8688-8691). In addition, many other natural serine recombinases are known that have an N-terminal catalytic domain and a C-terminal DNA-binding domain (e.g., Phi C31 integrase, TnpX transposese, IS607 transposese), and their catalytic domains can be selected to operate programmable site-specific recombinases as described herein (see, for example, Smith et al., "Diversity in the serine recombinases." Mol Microbiol. 2002;44:299-307 (its entire contents are incorporated by reference)). Similarly, core catalytic domains of tyrosine recombinases (e.g., Cre, λ-integrase) are known and can be similarly selected to manipulate programmable site-specific recombinases as described herein (for example, Guo et al., "Structure of Cre recombinase complexed with DNA in a site-specific recombination synapse." Nature. 1997; 389: 40-46; Hartung et al., "Cre mutants with altered DNA binding properties." J Biol Chem 1998; 273: 22884-22891; Shaikh et al., "Chimeras of the Flp and Cre recombinases: Tests of the mode of cleavage by Flp and Cre. J Mol Biol. 2000; 302: 27-48; Rongrong et al., "Effect of deletion mutation on the recombination activity of Cre recombinase." Acta Biochim) Pol.2005;52:541-544;Kilbride et al.,"Determinants of product topology in a hybrid Cre-Tn3 resolvase site-specific recombination system."J Mol Biol.2006;355:185-195;Warren et al.,"A chimeric cre recombinase with regulated directionality."Proc Natl Acad Sci USA.2008 105:18278-18283;Van Duyne,"Teaching Cre to follow directions."Proc Natl Acad Sci USA.2009 Jan 6;106(1):4-5;Numrych et al.,"A comparison of the effects of single-base and triple-base changes in the integrase arm-type binding sites on the site-specific recombination of bacteriophage λ."Nucleic Acids Res.1990;18:3953-3959;Tirumalai et al.,"The recognition of core-type DNA sites by λ integrase."J Mol Biol.1998;279:513-527;Aihara et See "A conformational switch controls the DNA cleavage activity of λ integrase," Mol Cell. 2003;12:187-198; Biswas et al., "A structural basis for allosteric control of DNA recombination by λ integrase," Nature. 2005;435:1059-1066; and Warren et al., "Mutations in the amino-terminal domain of λ-integrase have differential effects on integrative and excisive recombination," Mol Microbiol. 2005;55:1104-1112 (the full contents of each are incorporated by reference)).

[0307] Recombinase recognition sequence In this application, or the terms “recombinase-recognized sequence,” as used interchangeably as “RRS,” “recombinase target sequence,” or “recombinase site,” refer to a nucleotide sequence target that is recognized by a recombinase and undergoes strand exchange with another DNA molecule having an RRS. This results in the excision, incorporation, inversion, or exchange of DNA fragments between recombinase-recognized sequences. In various embodiments, a multi-strand prime editing factor may incorporate one or more recombinase sites on a target sequence or on more than one target sequence. When more than one recombinase site is incorporated by a multi-strand prime editing factor, the recombinase sites may be incorporated into adjacent target sites or non-adjacent target sites (e.g., separate chromosomes). In various embodiments, a single incorporated recombinase site may be used as a “landing site” for a recombinase-mediated reaction between a recombinase site on the genome and an exogenously supplied nucleic acid molecule, such as a second recombinase site on a plasmid. This enables the targeted incorporation of the desired nucleic acid molecule. In another embodiment, where two recombinase sites are inserted into adjacent regions of DNA (separated by, for example, 25-50 bp, 50-100 bp, 100-200 bp, 200-300 bp, 300-400 bp, 400-500 bp, 500-600 bp, 600-700 bp, 700-800 bp, 800-900 bp, 900-1000 bp, 1000-2000 bp, 2000-3000 bp, 3000-4000 bp, 4000-5000 bp, or more), the recombinase sites may be used for recombinase-mediated excision or inversion of the intervening sequence, or for recombinase-mediated cassette exchange with foreign DNA having the same recombinase sites. When two or more recombinase sites are incorporated by multiple flap-prime editing factors on two different chromosomes, a translocation of the intervening sequence from its position on the first chromosome to the second may occur.

[0308] Recombine or rearrange The term “recombination” is used in the context of nucleic acid modification (e.g., genome modification) to refer to the process by which two or more nucleic acid molecules or two or more regions of a single nucleic acid molecule are modified by the action of a recombinase protein (e.g., the recombinase fusion protein of the present invention provided herein). Recombination can result in, for example, insertion, inversion, excision, or translocation of nucleic acids on or between one or more nucleic acid molecules.

[0309] reverse transcriptase The term “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have been primarily used to transcribe mRNA into cDNA, which can then be cloned onto vectors for further manipulation. Myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific to the RNA strand of RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading, so transcriptional errors cannot be corrected by reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). Detailed studies of AMV reverse transcriptase activity and its associated RNase H activity are presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is that derived from Moloney's mouse leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases that substantially lack RNase H activity have also been described. For example, see USPat. No. 5,244,797.The present invention intends to utilize any such reverse transcriptase or its variants or variants.

[0310] In addition, the present invention intends to use a reverse transcriptase that is prone to errors, i.e., a reverse transcriptase that may be called an error-prone reverse transcriptase, or a reverse transcriptase that does not support high-fidelity incorporation of nucleotides during polymerization. During the synthesis of a single-stranded DNA flap based on guide RNA and an incorporated RT template, an error-prone reverse transcriptase may introduce one or more nucleotides that are mismatched with the RT template sequence, thereby introducing changes in the nucleotide sequence through the erroneous polymerization of the single-stranded DNA flap. These errors introduced during the synthesis of the single-stranded DNA flap are then incorporated into the double-stranded molecule by hybridization to the corresponding endogenous target strand, removal of the replaced endogenous strand, ligation, and then another round of the endogenous DNA repair and / or sequencing process.

[0311] Reverse transcription As used in this application, the term "reverse transcription" refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, reverse transcription may be "error-prone reverse transcription." This refers to the characteristics of certain reverse transcriptases in which their DNA polymerization activity is prone to errors.

[0312] PACE The term "phage-assisted sequential evolution (PACE)" as used in this application refers to sequential evolution using a phage as a viral vector. The general concept of PACE technology is illustrated, for example, by the international PCT application PCT / US2009 / 056194 filed on September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; the international PCT application PCT / US2011 / 066747 filed on December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; the US application, US No. 9,023,594 issued on May 5, 2015; and the international PCT application PCT / US2015 / 012022 filed on January 20, 2015, published as WO 2012 / 088381 on September 11, 2015. Published as 2015 / 134121; and described in the international PCT application PCT / US2016 / 027795 filed on 15 April 2016, published as WO 2016 / 168631 on 20 October 2016 (the contents of each of these in whole are incorporated herein by reference).

[0313] Phage In this application, the term "phage," used interchangeably with the term "bacteriophage," refers to a virus that infects bacterial cells. Typically, a phage consists of an outer protein capsid containing genetic material. The genetic material may be linear or circular ssRNA, dsRNA, ssDNA, or dsDNA. Phages and phage vectors are well known to those skilled in the art, and some non-limiting examples of phages useful for performing the PACE method provided herein are λ (lysogen), T2, T4, T7, T12, R17, M13, MS2, G4, P1, P2, P4, phi X174, N4, Φ6, and Φ29. In one embodiment, the phage utilized in the present invention is M13. Additional suitable phages and host cells will be apparent to those skilled in the art. The present invention is not limited in this respect.For further examples of suitable phages and host cells, please refer to: Elizabeth Kutter and Alexander Sulakvelidze: Bacteriophages: Biology and Applications. CRC Press; 1st edition (December 2004), ISBN: 0849313368; Martha RJ Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 1: Isolation, Characterization, and Interactions (Methods in Molecular Biology). Humana Press; 1st edition (December, 2008), ISBN: 1588296822; Martha RJ Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 2: Molecular and Applied Aspects (Methods in Molecular Biology). Humana Press; 1st edition (December See (2008), ISBN:1603275649; all of these, with regard to the disclosure of suitable phages and host cells, as well as methods and protocols for the isolation, culture, and manipulation of such phages, are incorporated herein by reference in their entirety).

[0314] Proteins, peptides, and polypeptides The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acid long. A protein, peptide, or polypeptide may refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified by the addition of a chemical entity, such as a carbohydrate group, hydroxyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference)).

[0315] Protein splicing As used in this application, the term "protein splicing" refers to the process in which an intein (or, as may be, a fission intein) sequence is excised from an amino acid sequence, and the remaining fragment of the amino acid sequence, the extein, is ligated by an amide bond to form a continuous amino acid sequence. The term "trans" protein splicing refers to the specific case in which the intein is a fission intein and they are located on different proteins.

[0316] Second-chain nicking The degradation of heteroduplex DNA (i.e., containing one edited and one unedited strand) formed as a result of prime editing determines the long-term editing outcome. In words, the goal of prime editing is to degrade the heteroduplex DNA (the edited strand paired with the endogenous unedited strand) formed as an intermediate of priming by permanently incorporating the edited strand onto the complementary endogenous strand. To help drive the degradation of heteroduplex DNA in a manner favorable to the permanent incorporation of the edited strand onto the DNA molecule, a “second strand nicking” approach may be used in this application. The concept of “second strand nicking” as used in this application refers to the introduction of a second nick, preferably on the unedited strand, downstream of the first nick (i.e., the initial nick site providing the free 3' end for use of reverse transcriptase on the extended portion of the guide RNA for priming). In one embodiment, the first and second nicks are on opposing strands. In other embodiments, the first and second nicks are on opposing strands. In yet another embodiment, the first nick is on the non-target strand (i.e., the strand forming the single-stranded portion of the R-loop), and the second nick is on the target strand. In yet another embodiment, the first nick is on the strand being edited, and the second nick is on the strand not being edited. The second nick may be located at least 5 nucleotides downstream of the first nick, or at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the first nick, or more.In one embodiment, the second nick may be introduced on the unedited strand at a distance of approximately 5–150 nucleotides, or approximately 5–140, or approximately 5–130, or approximately 5–120, or approximately 5–110, or approximately 5–100, or approximately 5–90, or approximately 5–80, or approximately 5–70, or approximately 5–60, or approximately 5–50, or approximately 5–40, or approximately 5–30, or approximately 5–20, or approximately 5–10 from the site of the PEgRNA-induced nick. In one embodiment, the second nick is introduced at a distance of 14–116 nucleotides from the PEgRNA-induced nick. Without being constrained by theory, the second nick directs the cell's endogenous DNA repair and replication processes toward replacement or editing of the unedited strand, thereby pe...

Claims

1. The system: (a) a prime editing factor or one or more polynucleotides encoding a prime editing factor, wherein the prime editing factor comprises a nucleic acid programmed DNA-binding protein (napDNAbp) and a polypeptide comprising RNA-dependent DNA polymerase activity, wherein the napDNAbp comprises a RuvC nuclease domain and an HNH nuclease domain, and wherein the HNH nuclease domain comprises one or more mutations that reduce or eliminate its nuclease activity; (b) A first prime editing guide RNA (first PEGRNA) or one or more polynucleotides encoding the first PEGRNA, where the first PEGRNA is (i) A first spacer sequence that is complementary to a first binding site on the first strand of a double-stranded DNA sequence upstream of the target site relative to the second strand, (ii) A first gRNA core that can complex with a prime editing factor, (iii) A first DNA synthesis template encoding a first single-stranded DNA sequence, and (iv) A first primer binding site complementary to the region of the second chain upstream of the first cleavage site. including; and (c) A second prime editing guide RNA (second PEGRNA) or one or more polynucleotides encoding the second PEGRNA, where the second PEGRNA is (i) A second spacer sequence complementary to a second binding site on the second strand of a double-stranded DNA sequence downstream of the target site relative to the second strand; (ii) A second gRNA core that can complex with a prime editing factor, (iii) A second DNA synthesis template encoding a second single-stranded DNA sequence, and (iv) A second primer binding site complementary to the region of the first chain upstream of the second cleavage site; Including; Here, the prime editing factor can cleave the first strand at the second cleavage site when complexed with the second PEGRNA, and can cleave the second strand at the first cleavage site when complexed with the first PEGRNA. Here, the first single-stranded DNA sequence and the second single-stranded DNA sequence are reverse complements across the complementary regions of each single-stranded DNA sequence. Here, the complementary region is at least 13 nucleotides long and forms a double helix, and Here, the first single-stranded DNA sequence includes a first edit compared to the second strand of the target site, starting at a position of 3 nucleotides or less from the first cleavage site. A system for simultaneously editing both strands of a double-stranded DNA sequence at a target site to be edited.

2. The system according to claim 1, comprising a second edit compared to the first strand of the target site, wherein the second single-stranded DNA sequence starts at a position of 3 nucleotides or less from the second cleavage site.

3. The system according to claim 1, wherein the complementary region is 22 to 38 nucleotides long.

4. The system according to claim 1, wherein the complementary region encapsulates the 3' ends of the first and second single-stranded DNA sequences.

5. The system according to claim 1, wherein the first single-stranded DNA sequence is the reverse complement of the second single-stranded DNA sequence.

6. The system according to claim 2, wherein a double helix formed by complementary regions of a first single-stranded DNA sequence and a second single-stranded DNA sequence comprises a first edit and a second edit.

7. (i) The first edit involves insertion, deletion, substitution, or a combination thereof, and / or (ii) The second edit includes, and / or involves insertion, deletion, substitution, or a combination thereof. (iii) The first single-stranded DNA sequence and the second single-stranded DNA sequence have the same length. The system according to claim 2.

8. The system according to claim 2, wherein the first or second edit includes a recombinase recognition sequence.

9. The system according to claim 8, further comprising a recombinase capable of recognizing a recombinase recognition sequence, or one or more polynucleotides encoding a recombinase, wherein optionally the recombinase is a serine recombinase.

10. The system according to claim 9, wherein the recombinase is Bxb1.

11. The system according to claim 10, wherein the recombinase recognition sequence includes sequence number 536 or sequence number 537.

12. The system according to claim 9, further comprising a donor template or a polynucleotide encoding a donor template.

13. The system according to claim 12, wherein the recombinase recognition sequence includes sequence number 536 and the donor template includes sequence number 537, or the recombinase recognition sequence includes sequence number 537 and the donor template includes sequence number 536.

14. The system according to claim 1, wherein the first editing starts at a position of 2 nucleotides or less from the first cleavage site.

15. The system according to claim 2, wherein the second editing starts at a position of 2 nucleotides or less from the second cleavage site.

16. The system according to claim 1, wherein the first editing begins at the first cutting site.

17. The system according to claim 2, wherein the second editing begins at the second cutting site.

18. The system according to claim 1, wherein the first single-stranded DNA sequence and / or the second single-stranded DNA sequence do not exhibit sequence homology with respect to the DNA sequence at the target site.

19. The system according to claim 1, wherein napDNAbp includes one amino acid sequence of sequence numbers 18, 21, 23, 25-39, 42-61, 75, and 77-88, or an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity with one of sequence numbers 18, 21, 23, 25-39, 42-61, 75, and 77-88.

20. The system according to claim 1, wherein the polypeptide containing RNA-dependent DNA polymerase activity is a reverse transcriptase, wherein optionally the polypeptide containing RNA-dependent DNA polymerase activity includes one amino acid sequence of SEQ ID NOs: 89-100, 106-122, 128, 132, 139, 143, 149, 154, 159, 700-736, 738-742, and 763-766, or an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity with one of SEQ ID NOs: 89-100, 106-122, 128, 132, 139, 143, 149, 154, 159, 700-736, 738-742, and 763-766.

21. The system according to claim 4, wherein the complementary region is 22 to 38 nucleotides long.

22. The system according to claim 21, wherein the first edit includes a recombinase recognition sequence.

23. The system according to claim 22, wherein the recombinase recognition sequence includes sequence number 536 or sequence number 537.

24. An in vitro or ex vivo method for simultaneously editing both strands of a double-stranded DNA sequence at a target site to be edited, comprising contacting the double-stranded DNA sequence with any one of the systems described in claims 1 to 23.

25. The system according to claims 1 to 23 for use as a pharmaceutical.

Citation Information

Patent Citations

  • High-throughput precision genome editing

    WO2018049168A1

  • Materials and methods for treatment of usher syndrome type 2a

    WO2019123429A1