Edited Methods and compositions for editing nucleotide sequences

Prime editing addresses the limitations of CRISPR by using target-primed reverse transcription to achieve precise and efficient nucleotide modifications, enhancing genome editing capabilities.

JP7786948B2Active Publication Date: 2025-12-16THE BROAD INST INC +1

Patent Information

Application Number
JP2021556862
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-17
Filing Date
2020-03-19
Publication Date
2025-12-16
Estimated Expiration
2040-03-19

AI Technical Summary

Technical Problem

Current genome editing technologies, such as CRISPR-based methods, face challenges in precisely editing single nucleotide mutations with high efficiency and specificity, particularly in non-dividing human cells, and often result in unwanted chromosomal rearrangements or cell apoptosis.

Method used

The development of prime editing, which uses a nucleic acid-programmed DNA-binding protein (napDNAbp) in conjunction with a polymerase to write new genetic information directly into a specific DNA site through target-primed reverse transcription, allowing for precise and efficient incorporation of desired nucleotide changes or modifications.

Benefits of technology

Prime editing enables high-precision genome editing with flexibility to introduce single nucleotide changes, insertions, or deletions, reducing bystander editing and target site disruptions, thus expanding the therapeutic potential of CRISPR-based technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007786948000488
    Figure 0007786948000488
  • Figure 0007786948000489
    Figure 0007786948000489
  • Figure 0007786948000490
    Figure 0007786948000490
Patent Text Reader

Abstract

The present disclosure provides compositions and methods for primed editing of target DNA molecules (e.g., genomes), which enable the incorporation of nucleotide changes and / or targeted mutagenesis. Nucleotide changes can include single nucleotide changes (e.g., any transition or any transversion), the insertion of one or more nucleotides, or the deletion of one or more nucleotides. More specifically, the present disclosure provides a fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a polymerase (e.g., reverse transcriptase), which is guided to a specific DNA sequence by a modified guide RNA termed PEgRNA. The PEgRNA is modified to include an extended portion (relative to the standard guide RNA) that provides a DNA synthesis template sequence. This encodes a single-stranded DNA flap that is homologous to the strand of the targeted endogenous DNA sequence to be edited, but which contains the desired one or more nucleotide changes, and which becomes incorporated into the target DNA molecule after synthesis by the polymerase (e.g., reverse transcriptase). Various methods utilizing prime editing are also disclosed herein, including, among others, treating trinucleotide repeat contraction diseases, incorporating targeted peptide tags, treating prion diseases by incorporating protective mutations, engineering genes encoding RNA for the incorporation of RNA tags to control RNA function and expression, constructing high performance gene libraries using prime editing, inserting immune epitopes onto proteins using prime editing, using prime editing to insert inducible dimerization domains onto protein targets, and delivery methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Government support This invention was made with government support under Grant Nos. U01AI142756, RM1HG009490, R01EB022376, and R35GM118062 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0002] Related Applications and Incorporation by Reference This U.S. provisional application is a continuation of the following applications: U.S. Provisional Application No. 62 / 820,813, filed March 19, 2019 (Attorney Docket No. B1195.70074US00), U.S. Provisional Application No. 62 / 858,958, filed June 7, 2019 (Attorney Docket No. B1195.70074US01), U.S. Provisional Application No. 62 / 889,996, filed August 21, 2019 (Attorney Docket No. U.S. Provisional Application No. 62 / 922,654, filed August 21, 2019 (Attorney Docket No. B1195.70074US02), U.S. Provisional Application No. 62 / 922,654, filed August 21, 2019 (Attorney Docket No. B1195.70083US00), U.S. Provisional Application No. 62 / 913,553, filed October 10, 2019 (Attorney Docket No. B1195.70074US03), U.S. Provisional Application No. 62 / 973,558, filed October 10, 2019 No. (Attorney Docket No. B1195.70083US01), U.S. Provisional Application No. 62 / 931,195, filed November 5, 2019 (Attorney Docket No. B1195.70074US04), U.S. Provisional Application No. 62 / 944,231, filed December 5, 2019 (Attorney Docket No. B1195.70074US05), U.S. Provisional Application No. 62 / 974, filed December 5, 2019 No. 537, filed March 17, 2020 (Attorney Docket No. B1195.70083US02), U.S. Provisional Application No. 62 / 991,069, filed March 17, 2020 (Attorney Docket No. B1195.70074US06), and U.S. Provisional Application No. (serial number not available at the time of filing this) filed March 17, 2020 (Attorney Docket No. B1195.70083US03), which are incorporated by reference. [Background technology]

[0003] Background of the Invention Pathogenic single-nucleotide mutations contribute, by some estimates, to approximately 50% of human diseases with a genetic component 7 Unfortunately, treatment options for patients with these genetic disorders remain extremely limited, despite decades of gene therapy research. 8 Perhaps the simplest solution to this therapeutic challenge is the direct correction of single-nucleotide mutations in the patient's genome, which would address the underlying cause of the disease and may provide lasting benefit. Such a strategy has never been considered before, but the CRISRP / Cas system has proven to be a promising approach. 9 Recent improvements in genome editing capabilities brought about by the advent of CRISPR now put this therapeutic approach within reach. By simple design of guide RNA (gRNA) sequences containing ~20 nucleotides complementary to the target DNA sequence, almost any conceivable genomic site can be specifically accessed by CRISPR-associated (Cas) nucleases. 1,2 To date, several monomeric bacterial Cas nuclease systems have been identified and adapted for genome editing applications. 10 This natural diversity of Cas nucleases, along with a growing collection of engineered variants, 11~14 , providing fertile ground for developing new genome editing technologies.

[0004] Although CRISPR gene disruption is now a mature technique, precisely editing single base pairs in the human genome remains a major challenge. 3 Homology-directed repair (HDR) has long been used in human cells and other organisms to insert, correct, or exchange DNA sequences at sites of double-strand breaks (DSBs) using a donor DNA repair template that encodes the desired edit. 15However, existing HDR has extremely low efficiency in most human cell types, particularly in non-dividing cells, and incompatible non-homologous end joining (NHEJ) leads primarily to insertion-deletion (indel) by-products. 16 Another issue concerns the generation of DSBs, which can result in large chromosomal rearrangements and deletions at the target locus. 17 or activate the p53 axis, which leads to growth arrest and apoptosis 18,19 .

[0005] Several approaches have been explored to address these shortcomings of HDR. For example, repair of single-strand DNA breaks (nicks) with oligonucleotide donors has been shown to reduce indel formation, but the yield of the desired repair product remains low. 20 Other strategies have attempted to bias repair toward HDR over NHEJ using small molecules and biological reagents. 21~23 However, the effectiveness of these methods may be cell type dependent, and perturbation of normal cellular conditions may lead to unwanted and unpredictable effects.

[0006] Recently, the present inventors, led by Professor David Liu and colleagues, developed base editing as a technique to edit targeted nucleotides without creating DSBs or relying on HDR. 4~6,24~27Direct modification of DNA bases by Cas-fused deaminases allows for highly efficient C·G → T·A or A·T → G·C base pair conversions within short target windows (~5–7 bases). As a result, base editors have been rapidly adopted by the scientific community. However, the following factors limit their generality for high-precision genome editing: (1) "bystander editing" of non-target C or A bases within the target window is observed; (2) target nucleotide product mixtures are observed; (3) the target base must be located 15 ± 2 nucleotides upstream of the PAM sequence; and (5) repair of small insertion and deletion mutations is not possible.

[0007] Thus, the development of programmable editors that can flexibly introduce any desired single nucleotide change and / or incorporate base pair insertions or deletions (e.g., insertions or deletions of at least 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more base pair insertions or deletions) and / or alter or modify nucleotide sequences at target sites with high specificity and efficiency would substantially expand the scope and therapeutic potential of CRISPR-based genome editing technologies. Summary of the Invention

[0008] SUMMARY OF THE INVENTION The present invention describes an entirely new platform for genome editing called "prime editing." Prime editing is a versatile and precise genome editing method that uses a nucleic acid-programmed DNA-binding protein ("napDNAbp") in conjunction with a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with the napDNAbp) to write new genetic information directly into a specific DNA site. The prime editing system is programmed with a prime editing (PE) guide RNA ("PEgRNA") that specifies both the target site and a template for synthesis of the desired edit in the form of a replacement DNA strand via an extension (either DNA or RNA) engineered onto the guide RNA (e.g., at the 5' or 3' end of the guide RNA, or internally). The replacement strand, containing the desired edit (e.g., a single nucleobase substitution), shares the same sequence as the endogenous strand at the target site to be edited (except that it encompasses the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand at the target site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing may be considered a "search-and-replace" genome editing technique, in that prime editing elements as described herein not only search for and locate the desired target site to be edited, but also simultaneously encode a replacement strand containing the desired edit that is incorporated in place of the corresponding target site in the endogenous DNA strand.

[0009] The prime editing elements of the present disclosure relate, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) or "prime editing" can be exploited or employed to perform high-precision CRISPR / Cas-based genome editing (e.g., as depicted in various embodiments in Figures 1A-1F) with high efficiency and genetic flexibility. TPRT is naturally used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial group II introns.28,29Herein, we use a Cas protein-reverse transcriptase fusion or related system that targets a specific DNA sequence with a guide RNA, generates a single-stranded nick at the target site, and uses the nicked DNA as a primer for reverse transcription of an engineered reverse transcriptase template incorporated with the guide RNA. However, although the concept originates with prime editors that use reverse transcriptase as a DNA polymerase component, the prime editors described herein are not limited to reverse transcriptase and may in fact encompass the use of any DNA polymerase. Indeed, although the application sometimes refers to prime editors with "reverse transcriptase" throughout, it is explained herein that reverse transcriptase is the only type of DNA polymerase that can function in prime editing. Thus, wherever "reverse transcriptase" is referred to herein, those skilled in the art will understand that any suitable DNA polymerase may be used in place of reverse transcriptase. Thus, in one aspect, a prime editor may comprise a Cas9 (or equivalent napDNAbp) programmed to target a DNA sequence by linking it to a specialized guide RNA (i.e., a PEgRNA) containing a spacer sequence that anneals to a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing the desired genetic alteration that is used to replace the corresponding endogenous DNA strand at the target site. To transfer information from the PEgRNA to the target DNA, the mechanism of prime editing involves nicking the target site in one strand of DNA, exposing a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime DNA polymerization of the edit-encoding extension on the PEgRNA directly into the target site. In various embodiments, the extension—which provides a template for polymerization of the replacement strand containing the edit—can be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editor may be a DNA-dependent DNA polymerase.

[0010] The newly synthesized strand (i.e., the replacement DNA strand containing the desired edit) formed by the primed editors disclosed herein will be homologous to (i.e., have the same sequence as) the target sequence in the genome, except for the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA, sometimes referred to as a single-stranded DNA flap, competes with the complementary homologous endogenous DNA strand for hybridization, thereby displacing the corresponding endogenous strand. In some embodiments, the system can be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with a Cas9 domain or provided in trans with the Cas9 domain). The error-prone reverse transcriptase can introduce modifications during single-stranded DNA flap synthesis. Thus, in some embodiments, an error-prone reverse transcriptase can be utilized to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used with the system, the modifications can be random or non-random.

[0011] Degradation of the hybridized intermediate (including the single-stranded DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can involve the resulting removal of the displaced flap from the endogenous DNA (e.g., with the 5'-terminal DNA flap endonuclease, FEN1), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis confers single-nucleotide precision on any nucleotide modification, including insertions and deletions, the breadth of this approach is extremely broad, and it is foreseeable that it could be used for countless applications in basic science and therapeutics.

[0012] In one aspect, the present disclosure provides a fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a reverse transcriptase. In various embodiments, the fusion protein is capable of performing genome editing by primed reverse transcription at a primed target in the presence of an extended guide RNA.

[0013] In some embodiments, the napDNAbp has nickase activity. The napDNAbp can also be a Cas9 protein or a functional equivalent thereof, such as a nuclease-active Cas9, a nuclease-inactive Cas9 (dCas9), or a Cas9 nickase (nCas9).

[0014] In some embodiments, the napDNAbp is selected from the group consisting of: Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaute, optionally having nickase activity.

[0015] In other embodiments, the fusion protein is capable of binding to a target DNA sequence when complexed with an extended guide RNA.

[0016] In still other embodiments, the target DNA sequence comprises a target strand and a complementary non-target strand.

[0017] In other embodiments, binding of the complexed fusion protein to the extended guide RNA forms an R-loop, which may comprise (i) an RNA-DNA hybrid comprising the extended guide RNA and the target strand, and (ii) a complementary non-target strand.

[0018] In still other embodiments, the complementary non-target strand is nicked to form a reverse transcriptase prime sequence with a free 3' end.

[0019] In various embodiments, the extended guide RNA comprises (a) a guide RNA and (b) an RNA extension at the 5' or 3' end of the guide RNA or at an intramolecular position of the guide RNA. The RNA extension may comprise (i) a reverse transcription template sequence containing the desired nucleotide change, (ii) a reverse transcription primer binding site, and (iii) optionally a linker sequence. In various embodiments, the reverse transcription template sequence may encode a single-stranded DNA flap that is complementary to the endogenous DNA sequence adjacent to the nick site, and the single-stranded DNA flap contains the desired nucleotide change.

[0020] In various embodiments, the RNA extension is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides in length.

[0021] In still other embodiments, the single-stranded DNA flap hybridizes to the endogenous DNA sequence adjacent to the nick site, thereby incorporating the desired nucleotide change. In still other embodiments, the single-stranded DNA flap has a free 5' end and replaces the endogenous DNA sequence adjacent to the nick site. In some embodiments, the replaced endogenous DNA with the 5' end is excised by the cell.

[0022] In various embodiments, cellular repair of the single-stranded DNA flap results in incorporation of the desired nucleotide change, thereby forming the desired product.

[0023] In various other embodiments, the desired nucleotide changes are incorporated over an editing window between about -4 and +10 of the PAM sequence.

[0024] In still other embodiments, the desired nucleotide change is incorporated over an editing window that is between about -5 and +5 of the nick site, or between about -10 and +10 of the nick site, or between about -20 and +20 of the nick site, or between about -30 and +30 of the nick site, or between about -40 and +40 of the nick site, or between about -50 and +50 of the nick site, or between about -60 and +60 of the nick site, or between about -70 and +70 of the nick site, or between about -80 and +80 of the nick site, or between about -90 and +90 of the nick site, or between about -100 and +100 of the nick site, or between about -200 and +200 of the nick site.

[0025] In various embodiments, the napDNAbp comprises the amino acid sequence of SEQ ID NO: 18. In various other embodiments, the napDNAbp comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 26-39, 42-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs: 77-86 (CP-Cas9); SEQ ID NOs: 18-25 and 87-88 (SpCas9); and SEQ ID NOs: 62-72 (Cas12).

[0026] In other embodiments, the reverse transcriptase of the disclosed fusion proteins and / or compositions can comprise any one of the amino acid sequences of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. In still other embodiments, the reverse transcriptase can comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. These sequences can be naturally occurring reverse transcriptase sequences, for example, from a retrovirus or retrotransposon, or the sequences can be recombinant.

[0027] In various other embodiments, the fusion proteins disclosed herein can comprise various structural configurations. For example, a fusion protein can comprise the structure NH-[napDNAbp]-[reverse transcriptase]-COOH; or NH-[reverse transcriptase]-[napDNAbp]-COOH, where each "]-[" indicates the presence of an optional linker sequence.

[0028] In various embodiments, the linker sequence comprises an amino acid sequence of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769, or an amino acid sequence that is at least 80%, 85%, or 90%, or 95%, or 99% identical to any one of the linker amino acid sequences of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769.

[0029] In various embodiments, the desired nucleotide change incorporated onto the target DNA can be a single nucleotide change (eg, a transition or transversion), an insertion of one or more nucleotides, or a deletion of one or more nucleotides.

[0030] In certain cases, the insertion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0031] In certain other cases, the deletion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0032] In another aspect, the present disclosure provides an extended guide RNA comprising a guide RNA and at least one RNA extension. The RNA extension may be located at the 3' end of the guide RNA. In other embodiments, the RNA extension may be located at the 5' end of the guide RNA. In yet other embodiments, the RNA extension may be located at an intramolecular position on the guide RNA. However, preferably, the intramolecular positioning of the extended portion does not block the function of the protospacer.

[0033] In various embodiments, a prime editor guide RNA (PEgRNA) is capable of binding to the napDNAbp and directing the napDNAbp to a target DNA sequence, which may include a target strand and a complementary non-target strand, where the guide RNA hybridizes to the target strand to form an RNA-DNA hybrid and an R-loop.

[0034] In various embodiments of the prime editor guide RNA, at least one RNA extension comprises a DNA synthesis template. In various other embodiments, the RNA extension further comprises a reverse transcription primer binding site. In yet other embodiments, the RNA extension comprises a linker or spacer that joins the RNA extension to the guide RNA.

[0035] In various embodiments, the RNA extensions can be at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0036] In other embodiments, the DNA synthesis template (i.e., with respect to Figure 27, the editing template) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0037] In still other embodiments, the reverse transcription primer binding site sequence (i.e., with respect to Figure 27, the primer binding site) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0038] In other embodiments, any linker or spacer is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0039] In various embodiments of the extended guide RNA, the reverse transcription template sequence can encode a single-stranded DNA flap that is complementary to the endogenous DNA sequence adjacent to the nick site, and the single-stranded DNA flap contains the desired nucleotide change. The single-stranded DNA flap can replace the endogenous single-stranded DNA at the nick site. The replaced endogenous single-stranded DNA at the nick site can have a 5' end, forming an endogenous flap that can be excised by the cell. In various embodiments, excision of the 5'-end endogenous flap can drive product formation because removal of the 5'-end endogenous flap promotes hybridization of the single-stranded 3' DNA flap to the corresponding complementary DNA strand and incorporation or assimilation of the desired nucleotide change carried by the single-stranded 3' DNA flap into the target DNA.

[0040] In various embodiments of the extended guide RNA, cellular repair of the single-stranded DNA flap results in the incorporation of the desired nucleotide change, thereby forming the desired product.

[0041] In one embodiment, the PEG RNA is 101-104, 131, 181-183, 222-234, 237-244, 277, 324-330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 394, 429-442, 499-505, 641-649, 678-6 92, 735–736, 738, 757–761, 776–777, 2997–3103, 3113–3121, 3305–3455, 3479–3493, 3522–3540, 3549–3556, 3628–3698, 3755–3810, 3874, 3890–3901, 3905–3911, 3913–3929, and 3972–3989 or SEQ ID NO: 101-104, 181-183, 223-234, 237-244, 277, 324-330, 332, 334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 394, 429-442, 499-505, 641-649, 678-6 92, 735–736, 757–761, 776–777, 2997–3103, 3113–3121, 3305–3455, 3479–3493, 3522–3540, 3549–3556, 3628–3698, 3755–3810, 3874, 3890–3901, 3905–3911, 3913–3929, and 3972–3989 or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity to any one of the following:

[0042] In another aspect of the present invention, the description provides a complex comprising a fusion protein described herein and any of the extended guide RNAs described above.

[0043] In yet another aspect, the present invention provides a complex comprising a napDNAbp and an extended guide RNA. The napDNAbp can be a Cas9 nickase or can be an amino acid sequence at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 42-57 (Cas9 nickase) and 65 (AsCas12a nickase).

[0044] In various embodiments involving the complex, the extended guide RNA can direct the napDNAbp to the target DNA sequence. In various embodiments, the reverse transcriptase can be provided in trans, i.e., from a source different from the complex itself. For example, the reverse transcriptase can be provided to the same cell containing the complex by introducing a separate vector encoding the reverse transcriptase.

[0045] In yet another aspect, the present disclosure provides polynucleotides. In some embodiments, the polynucleotides may encode any of the fusion proteins disclosed herein. In certain other embodiments, the polynucleotides may encode any of the napDNAbps disclosed herein. In yet further embodiments, the polynucleotides may encode any of the reverse transcriptases disclosed herein. In still other embodiments, the polynucleotides may encode any of the extended guide RNAs, any of the reverse transcription template sequences, any of the reverse transcription primer sites, or any of the optional linker sequences disclosed herein.

[0046] In yet another aspect, the present specification provides a vector comprising a polynucleotide described herein. Thus, in some embodiments, the vector comprises a polynucleotide encoding a fusion protein comprising a napDNAbp and a reverse transcriptase. In other embodiments, the vector comprises polynucleotides separately encoding the napDNAbp and the reverse transcriptase. In still other embodiments, the vector may comprise a polynucleotide encoding an extended guide RNA. In various embodiments, the vector may comprise one or more polynucleotides encoding the napDNAbp, the reverse transcriptase, and the extended guide RNA on the same or separate vectors.

[0047] In yet another aspect, the description provides a cell comprising the fusion protein and extended guide RNA described herein. The cell can be transformed with a vector comprising the fusion protein, napDNAbp, reverse transcriptase, and extended guide RNA. These genetic elements can be contained on the same vector or on different vectors.

[0048] In another aspect, the present disclosure provides pharmaceutical compositions. In some embodiments, the pharmaceutical composition comprises one or more of a napDNAbp, a fusion protein, a reverse transcriptase, and an extended guide RNA. In some embodiments, the pharmaceutical composition comprises a fusion protein described herein and a pharmaceutically acceptable excipient. In other embodiments, the pharmaceutical composition comprises any extended guide RNA described herein and a pharmaceutically acceptable excipient. In yet other embodiments, the pharmaceutical composition comprises any extended guide RNA described herein in combination with any fusion protein described herein and a pharmaceutically acceptable excipient. In yet another embodiment, the pharmaceutical composition comprises any polynucleotide sequence encoding one or more of a napDNAbp, a fusion protein, a reverse transcriptase, and an extended guide RNA. In yet other embodiments, the various components disclosed herein may be separated into one or more pharmaceutical compositions. For example, a first pharmaceutical composition may comprise a fusion protein or napDNAbp, a second pharmaceutical composition may comprise a reverse transcriptase, and a third pharmaceutical composition may comprise an extended guide RNA.

[0049] In yet a further aspect, the present disclosure provides kits. In one embodiment, the kits include one or more polynucleotides encoding one or more components, including a fusion protein, a napDNAbp, a reverse transcriptase, and an extended guide RNA. The kits may also include isolated preparations of vectors, cells, and polypeptides, including any of the fusion proteins, napDNAbp, or reverse transcriptases disclosed herein.

[0050] In another aspect, the present disclosure provides methods of using the disclosed compositions of matter.

[0051] In one embodiment, a method is provided for incorporating a desired nucleotide change into a double-stranded DNA sequence. The method first involves contacting the double-stranded DNA sequence with a complex comprising a fusion protein and an extended guide RNA, where the fusion protein comprises napDNAbp and a reverse transcriptase, and the extended guide RNA comprises a reverse transcription template sequence containing the desired nucleotide change. The method then involves nicking the double-stranded DNA sequence on the non-target strand, thereby generating a free single-stranded DNA having a 3' end. The method then involves hybridizing the 3' end of the free single-stranded DNA to the reverse transcription template sequence, thereby priming the reverse transcriptase domain. The method then involves polymerizing the DNA strand from the 3' end, thereby generating a single-stranded DNA flap containing the desired nucleotide change. The method then involves replacing the endogenous DNA strand adjacent to the cut site with the single-stranded DNA flap, thereby incorporating the desired nucleotide change into the double-stranded DNA sequence.

[0052] In another aspect, the present disclosure provides a method for introducing one or more changes into the nucleotide sequence of a DNA molecule at a target locus, comprising contacting the DNA molecule with a nucleic acid programmable DNA-binding protein (napDNAbp) and a guide RNA that targets the napDNAbp to the target locus, where the guide RNA comprises a reverse transcriptase (RT) template sequence containing at least one desired nucleotide change. The method then involves forming an exposed 3' end in the DNA strand at the target locus and then priming reverse transcription by hybridizing the exposed 3' end with the RT template sequence. Next, a single-stranded DNA flap containing at least one desired nucleotide change based on the RT template sequence is synthesized or polymerized by the reverse transcriptase. Finally, the at least one desired nucleotide change is incorporated into the corresponding endogenous DNA, thereby introducing one or more changes into the nucleotide sequence of the DNA molecule at the target locus.

[0053] In yet another aspect, the present disclosure provides a method for introducing one or more changes in the nucleotide sequence of a DNA molecule at a target locus by target-primed reverse transcription, the method comprising: (a) contacting the DNA molecule at the target locus with i) a fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a reverse transcriptase, and (ii) a guide RNA comprising an RT template comprising the desired nucleotide change; (b) allowing target-primed reverse transcription of the RT template to generate a single-stranded DNA comprising the desired nucleotide change; and (c) incorporating the desired nucleotide change into the DNA molecule at the target locus through DNA repair and / or replication processes.

[0054] In some embodiments, the step of replacing the endogenous DNA strand comprises: (i) hybridizing a single-stranded DNA flap to the endogenous DNA strand adjacent to the cleavage site, thereby creating a sequence mismatch; (ii) excising the endogenous DNA strand; and (iii) repairing the mismatch to form a desired product containing the desired nucleotide changes in both strands of DNA.

[0055] In various embodiments, the desired nucleotide change can be a single nucleotide substitution (e.g., a transition or transversion change), a deletion, or an insertion. For example, the desired nucleotide change can be: (1) a G to T substitution, (2) a G to A substitution, (3) a G to C substitution, (4) a T to G substitution, (5) a T to A substitution, (6) a T to C substitution, (7) a C to G substitution, (8) a C to T substitution, (9) a C to A substitution, (10) an A to T substitution, (11) an A to G substitution, or (12) an A to C substitution.

[0056] In other embodiments, the desired nucleotide change can convert (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair.

[0057] In still other embodiments, the method introduces a desired nucleotide change that is an insertion. In certain cases, the insertion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0058] In other embodiments, the method introduces a desired nucleotide change that is a deletion. In certain other cases, the deletion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0059] In various embodiments, the desired nucleotide change corrects a disease-associated gene. The disease-associated gene may be associated with a monogenic disorder selected from the group consisting of adenosine deaminase (ADA) deficiency, alpha-1 antitrypsin deficiency, cystic fibrosis, Duchenne muscular dystrophy, galactosemia, hemochromatosis, Huntington's disease, maple syrup urine disease, Marfan syndrome, neurofibromatosis type 1, pachyonychia congenita, phenylketonuria, severe combined immunodeficiency, sickle cell disease, Smith-Lemli-Opitz syndrome, and Tay-Sachs disease. In other embodiments, the disease-associated gene may be associated with a polygenic disorder selected from the group consisting of heart disease, high blood pressure, Alzheimer's disease, arthritis, diabetes, cancer, and obesity.

[0060] The methods disclosed herein may involve a fusion protein having a napDNAbp that is a nuclease-dead Cas9 (dCas9), a Cas9 nickase (nCas9), or a nuclease-active Cas9. In other embodiments, the napDNAbp and reverse transcriptase are not encoded as a single fusion protein, but rather can be provided in separate constructs. Thus, in some embodiments, the reverse transcriptase can be provided in trans to the napDNAbp (rather than as a fusion protein).

[0061] In various embodiments involving the methods, the napDNAbp may comprise the amino acid sequences of SEQ ID NOs: 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs: 77-86 (CP-Cas9); SEQ ID NOs: 18-25 and 87-88 (SpCas9); and SEQ ID NOs: 62-72 (Cas12). The napDNAbp may also comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs: 77-86 (CP-Cas9); SEQ ID NOs: 18-25 and 87-88 (SpCas9); and SEQ ID NOs: 62-72 (Cas12).

[0062] In various embodiments involving the methods, the reverse transcriptase may comprise any one of the amino acid sequences set forth in SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. The reverse transcriptase may also comprise an amino acid sequence at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence set forth in any one of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766.

[0063] The method comprises: 101~104、 131, 181~183、 222 ~234、237~244、277、324~330、332、334、336、338、340、342、344、346、348、350、352、354、356、358、360、362、364、366、368、 394, 429 ~ 442 、499~505、 641 ~ 649, 678 ~ 692, 735~736、 738, 757~761、776~777、 2997 ~ 3103, 3113 ~ 3121, 3305~ 3455, 3479 ~ 3493, 3522 ~ 3540, 3549 ~ 3556, 3628 ~ 3698, 3755 ~3810、3874、3890~3901、3905~3911、3913~3929 , and 3972~3989 or a nucleotide sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% sequence identity thereto. The method may involve the use of an extended guide RNA comprising an RNA extension at its 3' end, wherein the RNA extension comprises a reverse transcriptase template sequence.

[0064] The method can include the use of an extended guide RNA that includes an RNA extension at its 5' end, where the RNA extension includes a reverse transcription template sequence.

[0065] The method may include the use of an extended guide RNA comprising an RNA extension positioned intramolecularly of the guide RNA, the RNA extension comprising a reverse transcription template sequence.

[0066] The methods may involve the use of extended guide RNAs having one or more RNA extensions that are at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.

[0067] It should be understood that the above concepts, and additional concepts discussed below, may be arranged in any suitable combination (although the disclosure is not limited in this respect). Furthermore, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying figures. [Brief explanation of the drawings]

[0068] Brief description of the drawings The following drawings form part of the present specification and are included to further demonstrate certain aspects of the disclosure that may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0069] [Figure 1A] Figure 1A provides a schematic of an exemplary process for introducing single-nucleotide changes, insertions, and / or deletions into a DNA molecule (e.g., a genome) using a fusion protein comprising a reverse transcriptase fused to a Cas9 protein complexed with an extended guide RNA molecule. In this embodiment, the guide RNA is extended at its 3' end to include a reverse transcriptase template sequence. The schematic shows how a reverse transcriptase (RT) fused to a Cas9 nickase and complexed with a guide RNA (gRNA) binds to a DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which is used as a template for synthesis of a new DNA strand encoding the desired edit. The editing process shown may also be referred to as target-primed reverse transcription editing (TRT editing), or "primed editing."

[0070] [Figure 1B]Figure 1B provides the same diagram as Figure 1A, except that the prime editor complex is more commonly represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNA. Here, "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid-programmable DNA-binding protein (e.g., SpCas9), "PEgRNA" refers to the prime editing guide RNA, and "]-[" refers to an optional linker. As described elsewhere, e.g., in Figures 3A-3G, the PEgRNA includes a 5' extension arm that includes a primer binding site and a DNA synthesis template. Although not shown, it is contemplated that the extension arm of the PEgRNA (i.e., which includes the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase contemplated in this configuration will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). When the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase.

[0071] [Figure 1C]Figure 1C provides a schematic of an exemplary process for introducing single nucleotide changes and / or insertions and / or deletions into a DNA molecule (e.g., a genome) using a fusion protein comprising a reverse transcriptase fused to a Cas9 protein in complex with an extended guide RNA molecule. In this embodiment, the guide RNA is extended at its 5' end to encompass a reverse transcriptase template sequence. The schematic shows how a reverse transcriptase (RT) fused to a Cas9 nickase, in complex with a guide RNA (gRNA), binds to a DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which serves as a template for synthesis of a new DNA strand encoding the desired edit. The editing process shown can be referred to as target-primed reverse transcription editing (TRT editing) or, equivalently, "primed editing."

[0072] [Figure 1D]Figure 1D provides the same depiction as Figure 1C, except that the prime editor complex is more generally represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNAPEgRNA, where "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid-programmable DNA-binding protein (e.g., SpCas9), and "PEgRNA" refers to the prime editing guide RNA, and "]-[" refers to an optional linker. As depicted elsewhere, e.g., in Figures 3A-3G, the PEgRNA includes a 3' extension arm that includes a primer binding site and a DNA synthesis template. Although not shown, it is contemplated that the extension arm of the PEgRNA (i.e., including the primer binding site and DNA synthesis template) can be DNA or RNA. The specific polymerase contemplated in this configuration will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). If the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase. In various embodiments, the PEG RNA can be modified or synthesized to incorporate a DNA-based DNA synthesis template.

[0073] [Figure 1E] IE is a schematic depicting an exemplary process of how a synthesized single strand of DNA (containing a desired nucleotide change) is degraded so that the desired nucleotide change is incorporated into the DNA. As shown, subsequent synthesis of an edited strand (or "mutagenized strand"), equilibration with the endogenous strand, flap cleavage of the endogenous strand, and ligation leads to the incorporation of the DNA edit after degradation of the mismatched DNA duplex through the action of endogenous DNA repair and / or replication processes.

[0074] [Figure 1F]Figure 1F is a schematic illustrating that "opposite strand nicking" can be incorporated into the degradation method of Figure 1E to help drive the formation of the desired product versus the restored product. In opposite strand nicking, a second Cas9 / gRNA complex is used to introduce a second nick on the strand opposite the originally nicked strand. This directs the endogenous cellular DNA repair and / or replication processes to preferentially replace the unedited strand (i.e., the strand containing the second nick site).

[0075] [Figure 1G]FIG. 1G provides another schematic diagram of an exemplary process for introducing a single nucleotide change, insertion, and / or deletion into a DNA molecule (e.g., a genome) at a target locus using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an extended guide RNA. This process is sometimes referred to as prime editing. The extended guide RNA includes an extension at the 3' or 5' end of the guide RNA, or at an intramolecular location within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts the DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) into one strand of the DNA (the R-loop strand, or the PAM-containing strand, or the non-target DNA strand, or the protospacer strand) at the target locus, thereby creating an available 3' end on one of the strands at the target locus. In some embodiments, a nick is created in the strand of DNA corresponding to the R-loop strand, i.e., the strand that is not hybridized with the guide RNA sequence. In step (c), the 3'-terminal DNA strand interacts with the extended portion of the guide RNA to prime reverse transcription. In some embodiments, the 3'-terminal DNA strand hybridizes with a specific RT prime sequence on the extended portion of the guide RNA. In step (d), a reverse transcriptase is introduced to synthesize a single strand of DNA from the 3' end of the primed site to the 3' end of the guide RNA. This forms a single-stranded DNA flap containing the desired nucleotide change (e.g., a single base change, an insertion, a deletion, or a combination thereof). In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve degradation of the single-stranded DNA flap so that the desired nucleotide change is incorporated into the target locus. This process can drive the formation of the desired product by removing the corresponding 5' endogenous DNA flap once the 3' single-stranded DNA flap invades and hybridizes with the complementary sequence on the other strand. The process can also drive the formation of a nicked product on the second strand, as illustrated in Figure 1F.This process may introduce at least one or more of the following genetic alterations: transversions, transitions, deletions, and insertions.

[0076] [Figure 1H] Figure 1H is a schematic diagram depicting the types of genetic changes that can occur in the prime editing process described herein. The types of nucleotide changes that can be achieved by prime editing include deletions (including short and long deletions), single nucleotide changes (including transitions and transversions), and insertions (including short and long ones).

[0077] [Figure 1I] Figure 1I is a schematic depicting temporal second-strand nicking, exemplified by PE3b (PE3b = PE2 prime editor fusion protein + PEgRNA + second-strand nicking guide RNA). Temporal second-strand nicking is a variant of second-strand nicking to facilitate the formation of the desired edited product. The term "temporary" refers to the fact that second-strand nicking of the unedited strand occurs only after the desired edit has been incorporated into the edited strand. This avoids simultaneous nicking on both strands, which would lead to a double-stranded DNA break.

[0078] [Figure 1JK]Figures 1J-1K depict variations of prime editing contemplated herein in which napDNAbp (e.g., SpCas9 nickase) is replaced with any programmable nuclease domain, such as a zinc finger nuclease (ZFN) or transcription activator-like effector nuclease (TALEN). Therefore, it is contemplated that a suitable nuclease need not necessarily be "programmed" by a nucleic acid target molecule (e.g., a guide RNA), but rather may be programmed by defining the specificity of the DNA-binding domain, such as a nuclease. Just as with prime editing with napDNAbp moieties, alternative programmable nucleases are preferably modified to cleave only one strand of the target DNA. In other words, the programmable nuclease should preferably function as a nickase. Once a programmable nuclease is selected (e.g., a ZFN or TALEN), additional functionality may be engineered to enable it to operate according to a prime editing-like mechanism. For example, a programmable nuclease may be modified by coupling it to an RNA or DNA extension arm (e.g., via a chemical linker), where the extension arm includes a primer binding site (PBS) and a DNA synthesis template. A programmable nuclease may also be coupled to a polymerase (e.g., via a chemical or amino acid linker), although the nature of the polymerase will depend on whether the extension arm is DNA or RNA. In the case of an RNA extension arm, the polymerase may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of a DNA extension arm, the polymerase may be a DNA-dependent DNA polymerase (e.g., a prokaryotic polymerase including Pol I, Pol II, or Pol III, or a eukaryotic polymerase including Pol a, Pol b, Pol g, Pol d, Pol e, or Pol z).The system may also include other functions added as fusions with the programmable nuclease or added in trans to facilitate the overall reaction (e.g., (a) a helicase that unwinds the DNA at the cut site, creating a cut strand with a 3' end that can be used as a primer; (b) a flap endonuclease (e.g., FEN1) that serves to remove the endogenous strand on the cut strand, driving the reaction to replace the endogenous strand with the synthesized strand; or (c) an nCas9:gRNA complex that creates a second-site nick on the opposite strand (which may also serve to drive incorporation of synthetic repair over preferred cellular repair of the non-edited strand). In a manner similar to priming editing with napDNAbp, such complexes with other programmable nucleases can be used to synthesize and then permanently incorporate a newly synthesized replacement strand of DNA bearing the edit of interest into a target site in DNA.

[0079] [Figure 1L]Figure 1L depicts anatomical features of target DNA that may be edited by prime editing in one embodiment. The target DNA includes a "non-target strand" and a "target strand." The target strand is the strand that becomes annealed with the spacer of the PE gRNA of the prime editor complex that recognizes the PAM site (in this case, NGG, recognized by standard SpCas9-based prime editors). The target strand may also be referred to as the "non-PAM strand" or "non-edited strand." In contrast, the non-target strand (i.e., the strand containing the protospacer and NGG PAM sequence) may also be referred to as the "PAM strand" or "edited strand." In various embodiments, the nick site of the PE complex (e.g., in SpCas9-based PE) will be in the protospacer on the PAM strand. The location of the nick will be a characteristic of the specific Cas9 forming the PE. For example, in the case of SpCas9-based PE, the nick site is located in the phosphodiester bond between bases 3 (position "-3" relative to position 1 of the PAM sequence) and 4 (position "-4" relative to position 1 of the PAM sequence). The nick site in the protospacer forms a free 3' hydroxyl group that complexes with the primer-binding site of the extension arm of the PE-gRNA, as shown in the diagram below, providing a substrate for initiating polymerization of a single strand of DNA encoding the DNA synthesis template for the extension arm of the PE-gRNA. This polymerization reaction is catalyzed by the polymerase (e.g., reverse transcriptase) of the PE fusion protein in the 5'→3' direction. Polymerization terminates (e.g., by the inclusion of a polymerization termination signal or secondary structure that functions to terminate PE polymerization activity) before reaching the gRNA core, producing a single-stranded DNA flap that is extended from the original 3' hydroxyl group of the nicked PAM strand. The DNA synthesis template encodes a single-stranded DNA that is homologous to the 5'-terminal single strand of the endogenous DNA immediately following the nick site on the PAM strand, incorporating the desired nucleotide change (e.g., single-base substitution, insertion, deletion, inversion).The desired editing position can be at any position downstream of the nick site on the PAM strand, including positions +1, +2, +3, +4 (start of the PAM site), +5 (position 2 of the PAM site), +6 (position 3 of the PAM site), +7, +8, +9, +10, +11, +12, +13, +14, +15, +16, +17, +18, +19, +20, +21, +22, +23, +24, +25, +26, +27, +28, +29, +30, +31, +32, +33, +34, +35, +36, +37, +38, +39, +40, +41, +42, +43, +44, +45, +46, +47, +48, ​​+49, +50, +51, +52, +53, +54, +55, +56, +57, +58, +59, +60, +61, +62, +63, +64, +65, +66, +67, +68, +69, +70, +71, +72, +73, +74, +75, +76, +77, +78, +79, +80, +81, +82, +83, +84, +85, +86, +87, +88, +89, +90, +91, +92, +93, +94, +95, +96, +97, +98, +99, +100, +101, +102, +103, +104, +105, +106, +107, +108, +109, +110, +111, +112, +113, +114, + 20, +21, +22, +23, +24, +25, +26, +27, +28, +29, +30, +31, +32, +33, +34, +35, +36, +37, +38, +39, +40, +41, +42, +43, +44, +45, +46, +47, +48, ​​+49, +50, +51, +52, +53, +54, +55 , +56, +57, +58, +59, +60, +61, +62, +63, +64, +65, +66, +67, +68, +69, +70, +71, +72, +73, +74, +75, +76, +77, +78, +79, +80, +81, +82, +83, +84, +85, +86, +87, +88, +89, +90, +91, +92, +93, +94, +95, +96, +97, +98, +99, +100, +101, +102, +103, +104, +105, +106, +107, +108, +109, +110, +111, +112, +113, +114, +115, +116, +117, +118, +119, +120, + The nick site may include 121, +122, +123, +124, +125, +126, +127, +128, +129, +130, +131, +132, +133, +134, +135, +136, +137, +138, +139, +140, +141, +142, +143, +144, +145, +146, +147, +148, +149, or +150 or more (relative to the downstream position of the nick site). Once the 3'-terminal single-stranded DNA (containing the target edit) replaces the endogenous 5'-terminal single-stranded DNA, DNA repair and replication processes will result in permanent incorporation of the edited site on the PAM strand and then correction of the mismatch on the non-PAM strand present at the edit site. In this way, the edit will spread to both strands of DNA at the target DNA site. It will be understood that references to "edited strand" and "non-edited" strand are merely intended to delineate the strand of DNA involved in the PE mechanism.The "edited strand" is the strand that first becomes edited by replacing the 5'-single-stranded DNA immediately downstream of the nick site with a synthesized 3'-single-stranded DNA containing the desired edit. The "non-edited" strand, which is the counterpart to the edited strand, also becomes edited (particularly with the edit of interest) through repair and / or replication to become complementary to the edited strand.

[0080] [Figure 1M]Figure 1M depicts the mechanism of prime editing, showing the target DNA, the prime editor complex, and the anatomical features of the interaction between the PEgRNA and the target DNA. First, a prime editor, including a fusion protein having a polymerase (e.g., reverse transcriptase) and a napDNAbp (e.g., SpCas9 nickase, e.g., SpCas9 with an inactivating mutation in the HNH nuclease domain (e.g., H840A) or an activating mutation in the RuvC nuclease domain (D10A)), is complexed with a PEgRNA and a DNA containing the target DNA to be edited. The PEgRNA includes a spacer, a gRNA core (also known as the gRNA backbone or gRNA backbone), which binds to the napDNAbp, and an extension arm. The extension arm can be at the 3' end, 5' end, or anywhere within the PEgRNA molecule. As shown, the extension arm is at the 3' end of the PEgRNA. The extender arm contains a primer binding site and a DNA synthesis template (containing both the edit of interest and a homologous region (i.e., the homologous arm)) that is homologous in the 3' to 5' direction to the single-stranded DNA at the 5' end immediately following the nick site on the PAM strand. As shown, once a nick is introduced, thereby generating a free 3' hydroxyl group immediately upstream of the nick site, the region immediately upstream of the nick site on the PAM strand anneals to a complementary sequence at the 3' end of the extender arm, termed the "primer binding site," creating a short double-stranded region with an available 3' hydroxyl end, which forms a substrate for the polymerase of the primed editor complex. The polymerase (e.g., reverse transcriptase) then polymerizes a strand of DNA from the 3' hydroxyl end to the end of the extender arm. The sequence of the single-stranded DNA is encoded by the DNA synthesis template, and it is the portion of the extender arm (i.e., excluding the primer binding site) that is "read" by the polymerase to synthesize new DNA. This polymerization effectively extends to the sequence originally 3' hydroxyl-terminus of the initial nick site. The DNA synthesis template encodes a single strand of DNA containing not only the desired edit, but also a region of homology to the single strand of endogenous DNA immediately downstream of the nick site on the PAM strand.The encoded 3'-single strand of DNA (i.e., the 3'-single-stranded DNA flap) then displaces the corresponding homologous endogenous 5'-single strand of DNA immediately downstream of the nick site on the PAM strand, forming a DNA intermediate with a 5'-single-stranded DNA flap, which is removed by the cell (e.g., by a flap endonuclease). The 3'-single-stranded DNA flap that anneals to the complement of the endogenous 5'-single-stranded DNA flap is ligated to the endogenous strand after the 5'-single-stranded DNA flap is removed. The desired edit in the 3'-single-stranded DNA flap, which has just been annealed and ligated, forms a mismatch with the complementary strand and undergoes DNA repair and / or rounds of replication, thereby permanently incorporating the desired edit on both strands.

[0081] [Figure 2] Figure 2 shows three Cas complexes (SpCas9, SaCas9, and LbCas12a) that can be used with the prime editors described herein, along with their PAM, gRNA, and DNA cleavage features. The diagram shows the complex designs involving SpCas9, SaCas9, and LbCas12a.

[0082] [Figure 3]Figures 3A-3F show the design of an engineered 5' prime editor gRNA (Figure 3A), 3' prime editor gRNA (Figure 3B), and intramolecular extension (Figure 3C). The extended guide RNA (or extended gRNA) may also be referred to herein as a PEgRNA or "prime editing guide RNA." Figures 3D and 3E provide additional embodiments of 3' and 5' prime editor gRNAs (PEgRNAs), respectively. Figure 3F illustrates the interaction between the 3' prime editor guide RNA with a target DNA sequence. The embodiments of Figures 3A-3C illustrate exemplary placement of a reverse transcription template sequence (i.e., or more broadly referred to as a DNA synthesis template, as indicated, because RT is only one type of polymerase that can be used in the context of a prime editor), primer binding site, and optional linker sequence, as well as the general placement of spacer and core regions, in the extended portions of the 3', 5', and intramolecular versions. The disclosed prime editing process is not limited to these configurations of the extended guide RNA. The embodiment in Figure 3D provides an exemplary PEgRNA structure contemplated herein. The PEgRNA comprises three main components, ordered from 5' to 3': a spacer, a gRNA core, and an extension arm at the 3' end. The extension arm can be further divided into the following structural elements from 5' to 3': a primer binding site (A), an editing template (B), and a homology arm (C). Additionally, the PEgRNA can comprise an optional 3'-end modification region (e1) and an optional 5'-end modification region (e2). Furthermore, the PEgRNA can comprise a transcription termination signal at the 3' end of the PEgRNA (not shown). These structural elements are further defined herein. The illustrated structure of the PEgRNA is not meant to be limiting and encompasses variations in the placement of elements. For example, the optional sequence modifications (e1) and (e2) can be located within or between any of the other regions shown. It is not limited to being located at the 3' and 5' ends.In certain embodiments, PEGRNAs may contain secondary RNA structures, such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., MS2 aptamers that recruit and bind MS2cp proteins). For example, such secondary structures may be located within the spacer, gRNA core, or extension arms, particularly within the e1 and / or e2 modified regions. In addition to secondary RNA structures, PEGRNAs may contain chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modified regions), where "N" can be any nucleobase. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linker may function to prevent reverse transcription of the sgRNA backbone or core. Additionally, in certain embodiments (see, e.g., Figure 72(c)), the extension arm (3) may be composed of RNA or DNA and / or may include one or more nucleobase analogs (e.g., which may add functionality such as temperature resilience). It should be further noted that the orientation of the extension arm (3) can be the natural 5' to 3' direction, or can be synthesized in the opposite direction, 3' to 5' (relative to the orientation of the overall PEG RNA molecule). It is also noted that those skilled in the art will be able to select an appropriate DNA polymerase for use in prime editing depending on the nature of the nucleic acid material of the extension arm (i.e., DNA or RNA). This can be implemented either as a fusion with napDNAbp or provided in trans as a separate moiety to synthesize a 3' single-stranded DNA flap encoded by the desired template encompassing the desired edit. For example, if the extension arm is RNA, the DNA polymerase can be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the extension arm is DNA, the DNA polymerase can be a DNA-dependent DNA polymerase. In various embodiments, provision of the DNA polymerase can be in trans, for example, an MS2 hairpin incorporated onto the RNA-protein recruitment domain (e.g., PE) (e.g., in the e1 or e2 region or elsewhere) and an MS2cp protein fused to the DNA polymerase.(whereby the DNA polymerase is co-localized with the PEG RNA). It is also noted that the primer binding site generally does not form part of the template used by the DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation "DNA synthesis template" refers to the region or portion of the extension arm (3) used as a template by the DNA polymerase to encode the desired 3' single-stranded DNA flap containing the edit and a region of homology to the 5' endogenous single-stranded DNA flap that is replaced by the 3' single-stranded DNA product of the primed editing DNA synthesis. In some embodiments, the DNA synthesis template includes an "editing template" and a "homologous arm," or one or more homologous arms, for example, before and after the editing template. The editing template can be as small as a single nucleotide substitution, or it can be an insertion or inversion of DNA. In addition, the editing template can also include deletions, which can be engineered by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template can also include the e2 region or a portion thereof. For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any portion of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of the e2 region is actually used as a template will depend on its composition and whether that composition disrupts DNA polymerase function.

[0083] [Figure 3E]The embodiment of Figure 3E provides another PEgRNA structure contemplated herein. PEgRNAs comprise three main components, ordered from 5' to 3': a spacer, a gRNA core, and an extension arm at the 3' end. The extension arm can be further divided into the following structural elements, from 5' to 3': a primer binding site (A), an editing template (B), and a homology arm (C). In addition, PEgRNAs can comprise an optional 3'-end modification region (e1) and an optional 5'-end modification region (e2). Furthermore, PEgRNAs can comprise a transcription termination signal at the 3' end of the PEgRNA (not shown). These structural elements are further defined herein. The illustrated structure of PEgRNAs is not meant to be limiting and encompasses variations in the placement of elements. For example, the optional sequence modifications (e1) and (e2) can be located within or between any of the other regions shown. They are not limited to being located at the 3' and 5' ends. In certain embodiments, PEGRNAs may contain secondary RNA structures, such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., the MS2 aptamer, which recruits and binds the MS2cp protein). These secondary structures may be located anywhere on the PEGRNA molecule. For example, such secondary structures may be located within the spacer, gRNA core, or extension arms, particularly within the e1 and / or e2 modified regions. In addition to secondary RNA structures, PEGRNAs may contain chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modified regions), where "N" can be any nucleobase. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linker may function to prevent reverse transcription of the sgRNA backbone or core. Additionally, in certain embodiments (see, e.g., Figure 72(c)), extension arm (3) may be comprised of RNA or DNA and / or may include one or more nucleobase analogs (e.g., which may add functionality such as temperature resilience). Still further, the orientation of extension arm (3) may be in the natural 5' to 3' direction or may be synthesized in the opposite orientation, 3' to 5' (relative to the orientation of the overall PEG RNA molecule).It should also be noted that those skilled in the art will be able to select an appropriate DNA polymerase for use in prime editing depending on the nature of the nucleic acid material of the extension arm (i.e., DNA or RNA). This can be implemented either as a fusion with napDNAbp or provided in trans as a separate moiety to synthesize a 3' single-stranded DNA flap encoded by the desired template encompassing the desired edit. For example, if the extension arm is RNA, the DNA polymerase can be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the extension arm is DNA, the DNA polymerase can be a DNA-dependent DNA polymerase. In various embodiments, the DNA polymerase can be provided in trans, for example, by using an RNA-protein recruitment domain (e.g., an MS2 hairpin (e.g., in the e1 or e2 region or elsewhere) incorporated on the PEgRNA and an MS2cp protein fused to the DNA polymerase, thereby co-localizing the DNA polymerase to the PEgRNA). It is also noted that the primer binding site generally does not form part of the template used by a DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation "DNA synthesis template" refers to the region or portion of the extension arm (3) used as a template by a DNA polymerase to encode the desired 3' single-stranded DNA flap containing the edit and a region of homology to the 5' endogenous single-stranded DNA flap that is replaced by the 3' single-stranded DNA product of primed editing DNA synthesis. In some embodiments, the DNA synthesis template includes an "editing template" and a "homologous arm," or one or more homologous arms, for example, before and after the editing template. The editing template can be as small as a single nucleotide substitution, or it can be an insertion or inversion of DNA. In addition, the editing template can also include a deletion, which can be engineered by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template can also include the e2 region or a portion thereof.For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any part of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of the e2 region is actually used as a template will depend on its organization and whether that organization disrupts DNA polymerase function.

[0084] [Figure 3F]The schematic in Figure 3F illustrates the interaction of a typical PEG-RNA with a target site in double-stranded DNA and the concomitant generation of a 3' single-stranded DNA flap containing the desired genetic alteration. The double-stranded DNA is depicted with an upper strand (i.e., the target strand) in a 3' to 5' orientation and a lower strand (i.e., the PAM strand or non-target strand) in a 5' to 3' orientation. The upper strand contains the complement of the "protospacer" and the complement of the PAM sequence and is referred to as the "target strand" because it is the strand that is targeted by and anneals to the spacer of the PEG-RNA. The complementary lower strand is referred to as the "non-target strand" or "PAM strand" or "protospacer strand" because it contains the PAM sequence (e.g., NGG) and the protospacer. Although not shown, the depicted PEG-RNA would complex with Cas9 or an equivalent domain of a prime editor fusion protein. As shown in the schematic, the spacer of the PEgRNA anneals to the complementary region of the protospacer on the target strand. This interaction forms a DNA / RNA hybrid between the spacer RNA and the complement of the protospacer DNA, inducing the formation of an R-loop on the protospacer. As taught elsewhere herein, the Cas9 protein (not shown) then induces a nick on the non-target strand as shown. This then leads to the formation of a 3' ssDNA flap region immediately upstream of the nick site, which interacts with the 3' end of the PEgRNA at the primer binding site according to *z*. The 3' end of the ssDNA flap (i.e., the reverse transcriptase primer sequence) anneals to the primer binding site (A) on the PEgRNA, thereby priming the reverse transcriptase. Next, a reverse transcriptase (e.g., provided in trans or provided in cis as a fusion protein attached to the Cas9 construct) then polymerizes the single strand of DNA encoded by the DNA synthesis template (including the editing template (B) and the homology arm (C)). Polymerization continues toward the 5' end of the extension arm.The polymerized strand of ssDNA forms a ssDNA 3' end flap, which invades the endogenous DNA as described elsewhere (e.g., as shown in Figure 1G), displacing the corresponding endogenous strand (which is removed as a DNA flap at the 5' end of the endogenous DNA), and incorporating the desired nucleotide edit (single nucleotide base pair change, deletion, or insertion (including entire genes)) through naturally occurring rounds of DNA repair / replication.

[0085] [Figure 3G]Figure 3G illustrates yet another embodiment of prime editing contemplated herein. In particular, the upper schematic diagram illustrates one embodiment of a prime editor (PE), which comprises a fusion protein of a napDNAbp (e.g., SpCas9) and a polymerase (e.g., reverse transcriptase), connected by a linker. PE forms a complex with the PEgRNA by binding to the gRNA core of the PEgRNA. In the embodiment shown, the PEgRNA comprises a 3' extension arm, which, starting at the 3' end, contains a primer binding site (PBS) followed by a DNA synthesis template. The lower schematic diagram illustrates a variant of a prime editor referred to as a "trans-prime editor (tPE)." In this embodiment, the DNA synthesis template and PBS are decoupled from the PEgRNA and presented by a separate molecule referred to as a trans-prime editor RNA template ("tPERT"), which contains an RNA-protein recruitment domain (e.g., an MS2 hairpin). The PE itself is further modified to include a fusion to the rPERT recruitment protein ("RP"). This is a protein that specifically recognizes and binds to the RNA-protein recruitment domain. In instances where the RNA-protein recruitment domain is an MS2 hairpin, the corresponding rPERT recruitment protein can be MS2cp of the MS2 tagging system. The MS2 tagging system is based on the natural interaction of the MS2 bacteriophage coat protein ("MCP" or "MS2cp") with a stem-loop or hairpin structure present on the phage genome, i.e., the "MS2 hairpin" or "MS2 aptamer." In the case of trans-prime editing, the RP-PE:gRNA complex "recruits" tPERT bearing the appropriate RNA-protein recruitment domain to colocalize with the PE:gRNA complex, thereby providing the RP and DNA synthesis template in trans for use in prime editing, as shown in the example illustrated in Figure 3H.

[0086] [Figure 3H]Figure 3H illustrates the process of trans-prime editing. In this embodiment, the trans-prime editor comprises a "PE2" prime editor (i.e., a fusion of Cas9(H840A) and a variant MMLV RT) fused to an MS2cp protein (i.e., a type of recruitment protein that recognizes and binds the MS2 aptamer) and complexed with an sgRNA (i.e., a standard guide RNA, as opposed to a PEgRNA). The trans-prime editor binds to the target DNA and nicks the non-target strand. The MS2cp protein recruits tPERT in trans through a specific interaction with the RNA-protein recruitment domain on the tPERT molecule. tPERT becomes co-localized with the trans-prime editor, thereby providing the PBS and DNA synthesis template functions in trans for use by the reverse transcriptase polymerase to synthesize a single-stranded DNA flap with a 3' end and containing the desired genetic information encoded by the DNA synthesis template.

[0087] [Figure 4A] Figures 4A-4E demonstrate an in vitro TPRT assay (i.e., a prime editing assay). Figure 4A is a schematic representation of the fluorescently labeled DNA substrate, extension of the gRNA template by the RT enzyme, and PAGE. [Figure 4B-C] Figure 4B shows TPRT (i.e., prime editing) with pre-nicked substrates, dCas9, and 5'-extended gRNAs of different synthetic template lengths. Figure 4C shows the RT reaction with pre-nicked DNA substrates in the absence of Cas9. [Figure 4D-E] Figure 4D shows TPRT (i.e., prime editing) on ​​a full-length dsDNA substrate with Cas9(H840A) and a 5'-extended gRNA. Figure 4E shows a 3'-extended gRNA template with a pre-nicked full-length dsDNA substrate. M-MLV RT is present in all reactions.

[0088] [Figure 5]Figure 5 shows in vitro validation results using 5'-extended gRNAs with varying synthetic template lengths. A fluorescently labeled (Cy5) DNA target was used as the substrate and was pre-nicked in this set of experiments. The Cas9 used in these experiments was catalytically inactive Cas9 (dCas9), and the RT used was Superscript III, a commercially available RT derived from Moloney murine leukemia virus (M-MLV). The dCas9:gRNA complex was formed from purified components. The fluorescently labeled DNA substrate was then added along with dNTPs and the RT enzyme. After a 1-hour incubation at 37°C, the reaction products were analyzed by denaturing urea-polyacrylamide gel electrophoresis (PAGE). The gel image shows an extension of the original DNA strand to a length consistent with the reverse transcription template length.

[0089] [Figure 6] Figure 6 shows in vitro validation results using 5'-extended gRNAs of varying synthetic template length, similar to those shown in Figure 5. However, the DNA substrate was not pre-nicked in this set of experiments. The Cas9 used in these experiments was Cas9 nickase (SpyCas9 H840A mutant), and the RT used was Superscript III, a commercial RT derived from Moloney murine leukemia virus (M-MLV). Reaction products were analyzed by denaturing urea-polyacrylamide gel electrophoresis (PAGE). As shown on the gel, the nickase efficiently cleaves the DNA strand when a standard gRNA is used (gRNA_0, lane 3).

[0090] [Figure 7]Figure 7 demonstrates that 3' extension supports DNA synthesis without significant Cas9 nickase activity. Pre-nicked substrates (black arrows) are nearly quantitatively converted to RT products when either dCas9 or Cas9 nickase is used (lanes 4 and 5). Greater than 50% conversion to RT products (red arrows) is observed for the full-length substrate (lane 3). Cas9 nickase (SpyCas9 H840A mutant), catalytically inactive Cas9 (dCas9), and Superscript III (a commercial RT derived from Moloney murine leukemia virus (M-MLV)) are used.

[0091] [Figure 8] Figure 8 demonstrates a dual-color experiment used to determine whether the RT reaction occurs preferentially in cis with the gRNA (bound in the same complex). Two separate experiments were performed for the 5'-extended gRNA and the 3'-extended gRNA. Products were analyzed by PAGE. Product ratios were calculated as (Cy3cis / Cy3trans) / (Cy5trans / Cy5cis).

[0092] [Figure 9A-B] Figures 9A-9D demonstrate the flap model substrate. Figure 9A shows a dual FP reporter for flap-directed mutagenesis. Figure 9B shows stop codon repair in HEK cells. [Figure 9C-D] Figure 9C shows yeast clones sequenced after flap repair, and Figure 9D shows examination of different flap features in human cells.

[0093] [Figure 10]Figure 10 demonstrates prime editing on a plasmid substrate. A dual-fluorescent reporter plasmid was constructed for yeast (S. cerevisiae) expression. Expression of this construct in yeast produces only GFP. An in vitro prime editing reaction induces point mutations, and the parental plasmid or a plasmid in vitro nicked with Cas9(H840A) is transformed into yeast. Colonies are visualized by fluorescent imaging. Yeast dual-FP plasmid transformants are shown. Transformation of the parental plasmid or a plasmid in vitro nicked with Cas9(H840A) results in only green GFP-expressing colonies. Prime editing reactions with 5'- or 3'-extended gRNAs produce mixed green and yellow colonies, the latter expressing both GFP and mCherry. More yellow colonies are observed with the 3'-extended gRNA. A positive control containing no stop codon is also not shown.

[0094] [Figure 11] Figure 11 shows prime editing on a plasmid substrate similar to the experiment in Figure 10, but instead of incorporating a point mutation in the stop codon, the prime edit incorporates a single nucleotide insertion (left) or deletion (right) that repairs the frameshift mutation and allows synthesis of mCherry downstream. Both experiments used 3'-extended gRNAs.

[0095] [Figure 12] Figure 12 shows the editing products of prime editing on the plasmid substrate, characterized by Sanger sequencing. Colonies from the TRT transformation were individually selected and analyzed by Sanger sequencing. Correct editing was observed by sequencing selected colonies. Green colonies contained the plasmid with the original DNA sequence, while yellow colonies contained the correct mutation designed by the prime-editing gRNA. No other point mutations or indels were observed.

[0096] [Figure 13]Figure 13 illustrates the potential scope of the new prime-editing technology and provides a comparison with deaminase-mediated base editing techniques.

[0097] [Figure 14] FIG. 14 shows a schematic diagram of editing in human cells.

[0098] [Figure 15] Figure 15 demonstrates the extension of the primer binding site in the gRNA.

[0099] [Figure 16] Figure 16 shows gRNAs truncated for adjacent targeting.

[0100] [Figures 17A-C] Figures 17A-17C are graphs displaying % T→A conversion at the target nucleotide after transfection of the constructs in human embryonic kidney (HEK) cells. Figure 17A shows data presenting the results using an N-terminal fusion of wild-type MLV reverse transcriptase to the Cas9 (H840A) nickase (32 amino acid linker). Figure 17B is similar to Figure 17A except for the C-terminal fusion of the RT enzyme. Figure 17C is similar to Figure 17A, except that the linker between the MLV RT and Cas9 is 60 amino acids long instead of 32 amino acids.

[0101] [Figure 18] Figure 18 shows high-purity T→A editing at the HEK3 site by high-throughput amplicon sequencing. The sequencing analysis output displays the most abundant genotypes in edited cells.

[0102] [Figure 19] Figure 19 shows the editing efficiency (orange bars) at the target nucleotide alongside the indel rate (blue bars). WT refers to the wild-type MLV RT enzyme. Mutant enzymes (M1-M4) contain the mutations listed to the right. Editing rates were quantified by high-throughput sequencing of genomic DNA amplicons.

[0103] [Figure 20] Figure 20 shows the editing efficiency of a target nucleotide when a single-stranded nick is introduced into the complementary DNA strand adjacent to the target nucleotide. Nicking at various distances from the target nucleotide was tested (triangles). The editing efficiency at the target base pair (blue bar) is shown alongside the indel formation rate (orange bar). The "none" example does not contain a guide RNA that nicks the complementary strand. The editing rate was quantified by high-throughput sequencing of genomic DNA amplicons.

[0104] [Figure 21] Figure 21 demonstrates the processed high-throughput sequencing data showing the desired T→A transversion mutation and the general absence of other major genome editing by-products.

[0105] [Figure 22]Figure 22 provides a schematic diagram of an exemplary process for performing targeted mutagenesis at a target locus using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an error-prone reverse transcriptase (i.e., priming with an error-prone RT). This process is sometimes referred to as a prime-editing embodiment for targeted mutagenesis. The extended guide RNA includes extension at the 3' or 5' end of the guide RNA, or at an intramolecular location within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts a DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus to be mutagenized. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) into one strand of DNA at the target locus, thereby creating an available 3' end on one strand of the target locus. In one embodiment, the nick is created in the strand of DNA corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence. In step (c), the 3'-terminal DNA strand interacts with the extended portion of the guide RNA to prime reverse transcription. In some embodiments, the 3'-terminal DNA strand hybridizes to a specific RT prime sequence on the extended portion of the guide RNA. In step (d), an error-prone reverse transcriptase is introduced, which synthesizes a single strand of mutagenized DNA from the 3' end of the primed site to the 3' end of the guide RNA. Exemplary mutations are indicated with an asterisk (*). This forms a single-stranded DNA flap containing the desired mutagenized region. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve degradation of the single-stranded DNA flap (containing the mutagenized region) so that the desired mutagenized region is integrated into the target locus. This process can be driven to the desired product formation by removing the corresponding 5' endogenous DNA flap formed once the 3'-terminal single-stranded DNA flap invades and hybridizes with a complementary sequence on the other strand. The process can also be driven to form a second-strand nicked product, as illustrated in Figure 1F.Following endogenous DNA repair and / or replication processes, the mutagenic region becomes incorporated into both strands of DNA at the DNA locus.

[0106] [Figure 23] Figure 23 shows a schematic diagram of gRNA design for reducing trinucleotide repeat sequences and trinucleotide repeat reduction with TPRT genome editing (i.e., prime editing). Trinucleotide repeat expansions are associated with a number of human diseases, including Huntington's disease, fragile X syndrome, and Friedreich's ataxia. The most common trinucleotide repeat contains a CAG triplet, but GAA triplets (Friedreich's ataxia) and CGG triplets (Fragile X syndrome) also occur. Inheriting a predisposition to expansion or acquiring an already expanded parental allele increases the likelihood of acquiring disease. Pathogenic trinucleotide repeat expansions could hypothetically be corrected using prime editing. A region upstream of the repeat region could be nicked by an RNA-guided nuclease and then used to prime the synthesis of a new DNA strand containing a healthy number of repeats (which depends on the specific gene and disease). The repeat sequence is followed by a short stretch of homology (red strand) that matches the identity of the sequence adjacent to the other end of the repeat. Invasion of the newly synthesized strand, followed by displacement of the endogenous DNA with the newly synthesized flap, leads to a shortened repeat allele.

[0107] [Figure 24] Figure 24 is a schematic diagram showing the precise 10-nucleotide deletion in prime editing. Guide RNAs targeting the HEK3 locus were designed with a reverse transcription template encoding a 10-nucleotide deletion after the nick site. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing.

[0108] [Figure 25]Figure 25 is a schematic diagram showing gRNA design for peptide tagging of genes at endogenous genomic loci and peptide tagging in TPRT genome editing (i.e., prime editing). The FlAsH and ReAsH tagging systems comprise two parts: (1) a fluorophore-biarsenical probe and (2) a genetically encoded peptide containing a tetracysteine ​​motif, exemplified by the sequence FLNCCPGCCMEP (SEQ ID NO: 1). When expressed in cells, tetracysteine ​​motif-containing proteins can be fluorescently labeled with the fluorophore-arsenical probe (see reference: J. Am. Chem. Soc., 2002, 124(21), pp. 6063-6076. DOI: 10.1021 / ja017687n). The "sortagging" system employs bacterial sortase enzymes to covalently conjugate labeled peptide probes to proteins containing suitable peptide substrates (see reference: Nat. Chem. Biol. 2007 Nov;3(11):707-8. DOI: 10.1038 / nchembio.2007.31). FLAG tags (DYKDDDDK (SEQ ID NO: 2)), V5 tags (GKPIPNPLLGLDST (SEQ ID NO: 3)), GCN4 tags (EELLSKNYHLENEVARLKK (SEQ ID NO: 4)), HA tags (YPYDVPDYA (SEQ ID NO: 5)), and Myc tags (EQKLISEEDL (SEQ ID NO: 6)) are commonly employed as epitope tags for immunoassays. Pi-clamp encodes a peptide sequence (FCPF (SEQ ID NO: 622)) that can be labeled with pentafluoro-aromatic substrates (Reference: Nat. Chem. 2016 Feb;8(2):120-8. doi:10.1038 / nchem.2413).

[0109] [Figure 26A]Figure 26A shows the precise incorporation of His6 and FLAG tags into genomic DNA. Guide RNAs targeting the HEK3 locus were designed with reverse transcription templates encoding either an 18nt His tag insertion or a 24nt FLAG tag insertion. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing. Note that the full 24nt sequence of the FLAG tag is truncated (sequencing confirmed the full length was inserted correctly). [Figure 26B] FIG. 26B shows a schematic summarizing various applications involving protein / peptide tagging, including (a) solubilizing or insolubilizing proteins, (b) altering or tracking protein subcellular localization, (c) extending protein half-life, (d) facilitating protein purification, and (e) facilitating protein detection.

[0110] [Figure 27] Figure 27 shows an overview of prime editing by incorporating protective mutations into PRNP that prevent or halt the progression of prion disease. The PEGRNA sequence corresponds to SEQ ID NO: 4082 on the left (i.e., 5' of the sgRNA backbone) and SEQ ID NO: 4083 on the right (i.e., 3' of the sgRNA backbone).

[0111] [Figure 28A] FIG. 28A is a schematic representation of PE-based insertion of sequences encoding RNA motifs. [Figure 28B] Figure 28B is a non-exhaustive list of some example motifs that can potentially be inserted, and their functions.

[0112] [Figure 29]Figure 29A is a diagram of a prime editor. Figure 29B shows possible PE-directed modifications to genome, plasmid, or viral DNA. Figure 29C shows an example scheme for the insertion of a library of peptide loops onto a defined protein (in this case, GFP) by a library of PEgRNAs. Figure 29D shows examples of possible programmable deletions of codons or N- or C-terminal truncations of a protein using different PEgRNAs. Deletions would be expected to occur with minimal generation of frameshift mutations.

[0113] [Figure 30] Figure 30 shows a possible scheme for the recursive insertion of codons in a continuous evolution system such as PACE.

[0114] [Figure 31] Figure 31 is a representation of an engineered gRNA, showing the gRNA core, a ∼20 nt spacer that matches the sequence of the target gene, a reverse transcription template with an immunogenic epitope nucleotide sequence, and a primer binding site that matches the sequence of the target gene.

[0115] [Figure 32] Figure 32 is a schematic diagram of using prime editing as a means to insert known immunogenic epitopes onto endogenous or exogenous genomic DNA, resulting in the modification of the corresponding protein.

[0116] [Figure 33]Figure 33 is a schematic diagram showing PEgRNA design for primer binding sequence insertion and primer-bound insertion onto genomic DNA using prime editing to determine off-target editing. In this embodiment, prime editing is performed in living cells, tissues, or animal models. As a first step, an appropriate PEgRNA is designed. The upper schematic shows an exemplary PEgRNA that can be used in this aspect. The spacer on the PEgRNA (labeled "protospacer") is complementary to one of the strands of the genomic target. The PE:PEgRNA complex (i.e., the PE complex) incorporates a single-stranded 3'-end flap at the nick site, which contains the encoded primer binding sequence and a homologous region (encoded by the homologous arm of the PEgRNA) that is complementary to the region just downstream of the cut site (shown in red). Through flap invasion and DNA repair / replication processes, the synthesized strand becomes incorporated onto the DNA, thereby incorporating the primer binding site. This process can occur not only at the desired genomic target but also at other genomic sites that may interact with the PEgRNA in an off-target manner (i.e., due to the complementarity of the spacer region to other genomic sites other than the intended genomic site, the PEgRNA guides the PE complex to other off-target sites). Therefore, primer binding sequences can be incorporated not only at the desired genomic target but also at other off-target genomic sites on the genome. To detect the insertion of these primer binding sites at both the intended genomic target site and the off-target genomic site, genomic DNA (post-PE) can be isolated, fragmented, and ligated to adapter nucleotides (shown in red). Next, PCR can be performed using PCR oligonucleotides that anneal to the adapter and the inserted primer binding sequence to amplify the on-target and off-target genomic DNA regions where the primer binding sites were inserted by PE. High-throughput sequencing and sequence alignment can then be performed to identify the insertion point of the primer binding sequence inserted by PE at either the on-target or off-target site.

[0117] [Figure 34] FIG. 34 is a schematic diagram showing precise insertion of genes by PE.

[0118] [Figure 35] Figure 35A is a schematic diagram showing the native insulin signaling pathway, and Figure 35B is a schematic diagram showing FKBP12-tagged insulin receptor activation controlled by FK1012.

[0119] [Figure 36] Figure 36 shows the small molecule monomer. Reference: Bump FK506 Mimic (2) 107.

[0120] [Figure 37] Figures 37A-37B show small molecule dimers. References: FK1012 495,96; FK1012 5108; FK1012 6107; AP1903 7107; Cyclosporine A dimer 898; FK506-Cyclosporine A dimer (FkCsA) 9100.

[0121] [Figure 38]Figures 38A-38F provide an overview of prime editing and feasibility studies in vitro and in yeast cells. Figure 38A shows 75,122 known pathogenic human genetic variants from ClinVar (accessed July 2019), categorized by type. Figure 38B shows that the prime editing complex consists of a prime editor (PE) protein containing an RNA-guided DNA nicking domain, such as Cas9 nickase, fused to an engineered reverse transcriptase domain and complexed with a prime editing guide RNA (PEgRNA). The PE:PEgRNA complex binds to the target DNA site, enabling a wide range of precise DNA editing at a wide range of DNA positions before and after the target site's protospacer adjacent motif (PAM). Figure 38C shows that upon DNA target binding, the PE:PEgRNA complex nicks the PAM-containing DNA strand. The resulting free 3' end hybridizes to the primer-binding site of the PEgRNA. The reverse transcriptase domain catalyzes primer extension using the PEgRNA RT template, resulting in a newly synthesized DNA strand (3' flap) containing the desired edit. Equilibration between the edited 3' flap and the unedited 5' flap containing the original DNA, followed by cellular 5' flap cleavage and ligation, and DNA repair or replication to resolve the heteroduplex DNA, results in a stably edited DNA. Figure 38D shows an in vitro 5'-extended PEgRNA primer extension assay using a pre-nicked dsDNA substrate containing a 5' Cy5-labeled PAM strand, dCas9, and a commercially available M-MLV RT variant (RT, Superscript III). dCas9 was complexed with PEgRNA containing RT templates of various lengths and then added to the DNA substrate along with the indicated components. Reactions were incubated at 37°C for 1 hour and then analyzed by denaturing urea-PAGE and visualized for Cy5 fluorescence.Figure 38E shows primer extension performed as in Figure 38D using 3'-extended PEgRNA pre-complexed with dCas9 or Cas9 H840A nickase and pre-nicked or unnicked 5' Cy5-labeled dsDNA substrates. Figure 38F shows yeast colonies transformed with PEgRNA, Cas9 nickase, and a GFP-mCherry fusion reporter plasmid edited in vitro with RT. Plasmids containing nonsense or frameshift mutations between GFP and mCherry were edited with 5'- or 3'-extended PEgRNA that restored mCherry translation by transversion mutation, 1-bp insertion, or 1-bp deletion. GFP and mCherry double-positive cells (yellow) reflect successful editing.

[0122] [Figure 39]Figures 39A-39D show primed editing of genomic DNA in human cells with PE1 and PE2. Figure 39A shows that the PEgRNA contains a spacer sequence, an sgRNA backbone, and a 3' extension containing a reverse transcription (RT) template (purple) containing a primer binding site (green) and the base(s) to be edited (red). The primer binding site hybridizes to the PAM-containing DNA strand immediately upstream of the site of nicking. With the exception of the encoded edit, the RT template is homologous to the DNA sequence downstream of the nick. Figure 39B shows the incorporation of a T·A to A·T transversion edit at the HEK3 site in HEK293T cells using Cas9 H840A nickase (PE1) fused to wild-type M-MLV reverse transcriptase and PEgRNAs of various primer binding site lengths. Figure 39C shows that the use of an engineered quintuple mutant M-MLV reverse transcriptase (D200N, L603W, T306K, W313F, T330P) in PE2 substantially improves prime-editing transversion efficiency at five genomic sites in HEK293T cells and small insertion and deletion editing in HEK3. Figure 39D compares PE2 editing efficiency with various RT template lengths at five genomic sites in HEK293T cells. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0123] [Figure 40]Figures 40A-40C show that the PE3 and PE3b systems increase prime editing efficiency by nicking the non-edited strand. Figure 40A is an overview of prime editing with PE3. After initial synthesis of the edited strand, DNA repair will remove either the newly synthesized strand containing the edit (3' flap excision) or the original genomic DNA strand (5' flap excision). 5' flap excision leaves behind a DNA heteroduplex containing one edited strand and one non-edited strand. Mismatch repair mechanisms or DNA replication can resolve the heteroduplex, yielding either an edited or non-edited product. Nicking the non-edited strand favors repair of that strand, resulting in the preferential generation of stable duplex DNA containing the desired edit. Figure 40B shows the effect of complementary strand nicking on prime editing efficiency and indel formation mediated by PE3. "None" refers to the PE2 control, which does not nick the complementary strand. Figure 40C compares the editing efficiency of PE2 (no complementary strand nick), PE3 (general complementary strand nick), and PE3b (editing-specific complementary strand nick). All editing yields reflect the percentage of total sequencing reads containing the intended edit and no indels among all treated cells without sorting. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0124] [Figure 41]Figures 41A-41K show PE3-mediated targeted insertions, deletions, and all 12 types of point mutations at seven endogenous human genomic loci in HEK293T cells. Figure 41A is a graph showing all 12 types of single-nucleotide transition and transversion editing at positions +1 to +8 (counting the positioning of the PEgRNA-induced nick as between positions +1 and -1) of the HEK3 site using a 10-nt RT template. Figure 41B is a graph showing long-range PE3 transversion editing at the HEK3 site using a 34-nt RT template. Figures 41C-41H are graphs showing all 12 types of transition and transversion editing at various positions over the prime editing window for (Figure 41C) RNF2, (Figure 41D) FANCF, (Figure 41E) EMX1, (Figure 41F) RUNX1, (Figure 41G) VEGFA, and (Figure 41H) DNMT1. Figure 41I is a graph showing targeted 1- and 3-bp insertions and 1- and 3-bp deletions by PE3 at seven endogenous genomic loci. Figure 41J is a graph showing targeted precision deletions of 5 to 80 bp at HEK3 target sites. Figure 41K is a graph showing combined editing of insertions and deletions, insertions and point mutations, deletions and point mutations, and double point mutations at three endogenous genomic loci. All editing yields reflect the percentage of total sequencing reads containing the intended edit and no indels among all treated cells without sorting. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0125] [Figure 42]Figures 42A-42H show a comparison of primed and base editing by Cas9 and PE3 at known Cas9 off-target sites, as well as off-target editing. Figure 42A shows the total C·G to T·A editing efficiency at the same target nucleotide for PE2, PE3, BE2max, and BE4max at endogenous HEK3, FANCF, and EMX1 sites in HEK293T cells. Figure 42B shows the indel frequency from the treatments in Figure 42A. Figure 42C shows the editing efficiency of precise C·G to T·A editing (without bystander edits or indels) for PE2, PE3, BE2max, and BE4max in HEK3, FANCF, and EMX1. Also shown in EMX1 is precise PE combinatorial editing of all possible combinations of C·G to T·A conversions at the three target nucleotides. Figure 42D shows the total A·T to G·C editing efficiency for PE2, PE3, ABEdmax, and ABEmax in HEK3 and FANCF. Figure 42E shows the precise A·T to G·C editing efficiency without bystander editing or indels for HEK3 and FANCF. Figure 42F shows the indel frequency from the treatments in Figure 42D. Figure 42G shows the average triplicate editing efficiency (percentage sequencing reads with indels) in HEK293T cells for the Cas9 nuclease at four on-target and 16 known off-target sites. The 16 off-target sites observed were the top four previously reported off-target sites (118,159) for each of the four on-target sites. For each on-target site, Cas9 was paired with either an sgRNA or one of four PEgRNAs that recognize the same protospacer. Figure 42H shows the average triplicate on-target and off-target editing efficiencies and indel efficiencies (below in brackets) for PE2 or PE3 paired with each PEgRNA (Figure 42G) in HEK293T cells. On-target editing yield reflects the percentage of total sequencing reads containing the intended edit and no indels among all treated cells without sorting.Off-target editing yields reflect off-target locus modifications consistent with prime editing. Values ​​and error bars reflect the mean and sd of three independent biological replicates.

[0126] [Figure 43]Figures 43A-43I show a comparison of prime editing, incorporation and correction of pathogenic transversion, insertion, or deletion mutations, and prime editing and HDR in various human cell lines and primary mouse cortical neurons. Figure 43A is a graph showing incorporation (by a T·A to A·T transversion) and correction (by an A·T to T·A transversion) of a pathogenic E6V mutation in HBB of HEK293T cells. Correction to either wild-type HBB or HBB containing a silent mutation that blocks the PEGRNA PAM is shown. Figure 43B is a graph showing incorporation (by a 4-bp insertion) and correction (by a 4-bp deletion) of the pathogenic HEXA 1278+TATC allele in HEK293T cells. Correction to either wild-type HEXA or HEXA containing a silent mutation that blocks the PEGRNA PAM is shown. Figure 43C is a graph showing incorporation of the protective G127V variant of PRNP in HEK293T cells by a G·C to T·A transversion. Figure 43D is a graph showing prime editing in other human cell lines, including K562 (leukemia myeloid cells), U2OS (osteosarcoma cells), and HeLa (cervical cancer cells). Figure 43E is a graph showing incorporation of a G·C to T·A transversion mutation in DNMT1 in mouse primary cortical neurons using the dual split-intein PE3 lentiviral system. In this, the N-terminal half is Cas9(1-573) fused to the N intein and to GFP-KASH via the P2A self-cleaving peptide, and the C-terminal half is the C intein fused to the remainder of PE2. Both halves of PE2 are expressed from the human synapsin promoter, which is highly specific to mature neurons. Sorted values ​​reflect editing or indels from GFP-positive nuclei, and unsorted values ​​are from all nuclei. Figure 43F shows a comparison of HDR editing efficiencies mediated by PE3 and Cas9 at endogenous genomic loci in HEK293T cells. Figure 43G shows a comparison of HDR editing efficiencies mediated by PE3 and Cas9 at endogenous genomic loci in K562, U2OS, and HeLa cells.Figure 43H compares HDR indel by-product generation mediated by PE3 and Cas9 in HEK293T, K562, U2OS, and HeLa cells. Figure 43I shows targeted insertion of a His6 tag (18 bp), a FLAG epitope tag (24 bp), or an extended LoxP site (44 bp) in HEK293T cells by PE3. All editing yields reflect the percentage of total sequencing reads containing the intended edit and no indels among all treated cells. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0127] [Figure 44]Figures 44A-44G show in vitro prime-editing validation studies with fluorescently labeled DNA substrates. Figure 44A shows electrophoretic mobility shift assays with dCas9, 5'-extended PEG-RNAs, and 5'-Cy5-labeled DNA substrates. PEG-RNAs 1 to 5 contain a 15-nt linker sequence between the spacer and PBS (linker A for PEG-RNA 1 and linker B for PEG-RNAs 2 to 5), a 5-nt PBS sequence, and 7-nt (PEG-RNAs 1 and 2), 8-nt (PEG-RNA 3), 15-nt (PEG-RNA 4), and 22-nt (PEG-RNA 5) RT templates. PEG-RNAs are those used in Figures 44E and 44F; the complete sequences are listed in Tables 2A-2C. Figure 44B shows in vitro nicking assays of Cas9 H840A using 5'- and 3'-extended PEG-RNAs. Figure 44C shows Cas9-mediated indel formation in HEK293T cells using 5'- and 3'-extended PEG-RNA in HEK3. Figure 44D shows an overview of the prime editing in vitro biochemical assay. 5'-Cy5-labeled pre-nicked and unnicked dsDNA substrates were tested. sgRNA, 5'-extended PEG-RNA, or 3'-extended PEG-RNA was pre-complexed with dCas9 or Cas9 H840A nickase and then combined with dsDNA substrate, M-MLV RT, and dNTPs. Reactions were allowed to proceed for 1 hour at 37°C prior to separation by denaturing urea-PAGE and visualization by Cy5 fluorescence. Figure 44E shows that primer extension reactions using 5'-extended PEG-RNA, pre-nicked DNA substrate, and dCas9 lead to significant conversion to RT product. Figure 44F shows a primer extension reaction using an unnicked DNA substrate and Cas9 H840A nickase with a 5'-extended PEG-RNA as in Figure 44B. Product yield is greatly reduced compared to pre-nicked substrates. Figure 44G shows by denaturing urea PAGE that an in vitro primer extension reaction using a 3'-PEG-RNA produces a single, clear product.The RT product band was excised, eluted from the gel, and then subjected to homopolymeric tailing with terminal deoxyribonuclease (TdT) using either dGTP or dATP. The tailed product was extended with a poly-T or poly-C primer, and the resulting DNA was sequenced. Sanger traces indicate that three nucleotides from the gRNA backbone were reverse transcribed (added to the DNA product as the final 3' nucleotides). Note that in mammalian cell primed editing experiments, PEG-RNA backbone insertions were significantly rarer than in vitro (Figures 56A-56D). This is likely due to the inability of the tethered reverse transcriptase to access the Cas9-bound guide RNA backbone and / or cellular excision of the mismatched 3' end of the 3' flap containing the PEG-RNA backbone sequence.

[0128] [Figure 45]Figures 45A-45G show cellular repair in yeast of a 3' DNA flap from an in vitro prime editing reaction. Figure 45A shows that a dual fluorescent protein reporter plasmid contains GFP and mCherry open reading frames separated by target sites encoding an in-frame stop codon, a +1 frameshift, or a -1 frameshift. Prime editing reactions were performed in vitro with Cas9 H840A nickase, PEG RNA, dNTPs, and M-MLV reverse transcriptase and then transformed into yeast. Colonies containing the unedited plasmid produce GFP but not mCherry. Yeast colonies containing the edited plasmid produce both GFP and mCherry as fusion proteins. Figure 45B shows an overlay of GFP and mCherry fluorescence of yeast colonies transformed with reporter plasmids containing either a stop codon between GFP and mCherry (unedited negative control, top) or no stop codon or frameshift between GFP and mCherry (pre-edited positive control, bottom). Figures 45C-45F show visualization of mCherry and GFP fluorescence from yeast colonies transformed with in vitro primed editing reaction products. Figure 45C shows stop codon correction by a T·A to A·T transversion using a 3'-extended PEgRNA or a 5'-extended PEgRNA, as shown in Figure 45D. Figure 45E shows +1 frameshift correction of a 1 bp deletion using a 3'-extended PEgRNA. Figure 45F shows -1 frameshift correction by a 1 bp insertion using a 3'-extended PEgRNA. Figure 45G shows Sanger DNA sequencing traces from plasmids isolated from the GFP-only colony in Figure 45B and the GFP and mCherry double-positive colony in Figure 45C.

[0129] [Figure 46]Figures 46A-46F show correct editing versus indel generation by PE1. Figure 46A shows the efficiency of transversion editing and indel generation of T·A to A·T by PE1 at the +1 position of HEK3 using a 10-nt RT template and a PBS sequence ranging from 8-17 nt. Figure 46B shows the efficiency of transversion editing and indel generation of G·C to T·A by PE1 at the +5 position of EMX1 using a 13-nt RT template and a PBS sequence ranging from 9-17 nt. Figure 46C shows the efficiency of transversion editing and indel generation of G·C to T·A by PE1 at the +5 position of FANCF using a 17-nt RT template and a PBS sequence ranging from 8-17 nt. Figure 46D shows the efficiency of PE1-mediated C·G to A·T transversion editing and indel generation at the +1 position of RNF2, using a PEgRNA containing an 11-nt RT template and a PBS sequence ranging from 9 to 17 nt. Figure 46E shows the efficiency of PE1-mediated G·C to T·A transversion editing and indel generation at the +2 position of HEK4, using a PEgRNA containing a 13-nt RT template and a PBS sequence ranging from 7 to 15 nt. Figure 46F shows the +1T deletion, +1A insertion, and +1CTT insertion mediated by PE1 at the HEK3 site, using a 13-nt PBS and 10-nt RT template. The sequences of the PEgRNAs are those used in Figure 39C (see Tables 3A-3R). Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0130] [Figure 47]Figures 47A-47S show evaluation of M-MLV RT variants for prime editing. Figure 47A shows abbreviations for prime editor variants used in this figure. Figure 47B shows targeted insertion and deletion editing by PE1 at the HEK3 locus. Figures 47C-47H show a comparison of 18 prime editor constructs containing M-MLV RT variants for their ability to incorporate a +2G·C to C·G transversion edit in HEK3 as shown in Figure 47C, a 24-bp FLAG insertion in HEK3 as shown in Figure 47D, a +1C·G to A·T transversion edit in RNF2 as shown in Figure 47E, a +1G·C to C·G transversion edit in EMX1 as shown in Figure 47F, a +2T·A to A·T transversion edit in HBB as shown in Figure 47G, and a +1G·C to C·G transversion edit in FANCF as shown in Figure 47H. Figures 47I-47N show a comparison of four prime editor constructs containing M-MLV variants for their ability to incorporate the edits shown in Figures 47C-47H in a second round of independent experiments. Figures 47O-47S show PE2 editing efficiency at five genomic loci with various PBS lengths. Figure 47O shows the +1T·A to A·T variation in HEK3. Figure 47P shows the +5G·C to T·A variation in EMX1. Figure 47Q shows the +5G·C to T·A variation in FANCF. Figure 47R shows the +1C·G to A·T variation in RNF2. Figure 47S shows the +2G·C to T·A variation in HEK4. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0131] [Figure 48]Figures 48A-48C show design features of the PEgRNA PBS and RT template sequences. Figure 48A shows the efficiency of PE2-mediated +5G·C to T·A transversion editing in VEGFA of HEK293T cells as a function of RT template length (blue line). Indels (gray line) are plotted for comparison. The sequence below the graph indicates the final nucleotide for editing, which serves as a template for synthesis by the PEgRNA. The G nucleotide (C on the PEgRNA serves as a template) is highlighted; to maximize prime editing efficiency, RT templates ending in C should be avoided during PEgRNA design. Figure 48B shows +5G·C to T·A transversion editing and indels in DNMT1 as shown in Figure 48A. Figure 48C shows +5G·C to T·A transversion editing and indels in RUNX1 as shown in Figure 48A. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0132] [Figure 49]Figures 49A-49B show the effects of PE2, PE2 R110S K103L, Cas9 H840A nickase, and dCas9 on cell viability. HEK293T cells were transfected with plasmids encoding PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9 together with a HEK3-targeting PEgRNA plasmid. Cell viability was measured every 24 hours for 3 days post-transfection using the CellTiter-Glo 2.0 assay (Promega). Figure 49A shows viability measured by luminescence 1, 2, or 3 days post-transfection. Values ​​and error bars reflect the mean and s.e.m. of three independent biological replicates, each performed in technical triplicate. Figure 49B shows percent editing and indels for PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9, along with HEK3-targeting PEgRNA plasmids encoding +5G to A edits. Editing efficiency was measured from treated cells alongside those used to assay viability in Figure 49A, on day 3 post-transfection. Values ​​and error bars reflect the mean and SD of three independent biological replicates.

[0133] [Figure 50]Figures 50A-50B show PE3-mediated HBB E6V correction and HEXA 1278 + TATC correction by various PEgRNAs. Figure 50A shows a screen of 14 PEgRNAs for correction of the HBB E6V allele in HEK293T cells with PE3. All PEgRNAs evaluated convert the HBB E6V allele back to wild-type HBB without introducing any silent PAM mutations. Figure 50B shows a screen of 41 PEgRNAs for correction of the HEXA 1278 + TATC allele in HEK293T cells with PE3 or PE3b. PEgRNAs labeled with HEXA correct the pathogenic allele with a shifted 4-bp deletion that blocks the PAM and leaves a silent mutation. PEgRNAs labeled with HEXA correct the pathogenic allele back to wild-type. Entries ending in "b" use an editing-specific nicking sgRNA in combination with a PEgRNA (PE3b system). Values ​​and error bars reflect the mean and sd of three independent biological replicates.

[0134] [Figure 51]Figures 51A-51G show a comparison of PE3 activity and HDR initiated by PE3 and Cas9 in human cell lines. The efficiency of generating correct edits (no indels) and indel frequency for HDR initiated by PE3 and Cas9 in HEK293T cells (Figure 51A), K562 cells (Figure 51B), U2OS cells (Figure 51C), and HeLa cells (Figure 51D) is shown. Comparisons of each bracketed edit incorporate identical edits by HDR initiated by PE3 and Cas9. The non-targeting control is PE3 and a PEgRNA targeting a non-target locus. Figure 51E shows a control experiment with a non-targeting PEgRNA + PE3 and dCas9 + sgRNA compared to a wild-type Cas9 HDR experiment. This confirms that common contaminant ssDNA donor HDR templates, which falsely increase apparent HDR efficiency, do not contribute to the HDR measurements in Figures 51A-51D. Figures 51F-51G show an example HEK3 site allele table from a genomic DNA sample isolated from K562 cells after editing by PE3- or Cas9-initiated HDR. Alleles were sequenced by Illumina MiSeq and analyzed by CRISPResso2. The reference HEK3 sequence from this region is at the top. Allele tables are shown for the non-targeting PEgRNA negative control, the +1 CTT insertion in HEK3 using PE3, and the +1 CTT insertion in HEK3 using Cas9-initiated HDR. The allele frequency and corresponding Illumina sequencing read counts are shown for each allele. All alleles observed at frequencies ≥ 0.20% are shown. Values ​​and error bars reflect the mean and sd of three independent biological replicates.

[0135] [Figure 52]Figures 52A-52D show the distribution of pathogenic insertion, duplication, deletion, and indel lengths on the ClinVar database. The ClinVar variant summary was downloaded from NCBI on July 15, 2019. The lengths of reported insertions, deletions, and duplications were calculated using appropriate identifying information for the reference and alternate alleles, variant start and stop positions, or variant names. Variants that did not report any of the above information were excluded from the analysis. The lengths of reported indels (single variants encompassing both an insertion and a deletion relative to the reference genome) were calculated by determining the number of mismatches or gaps in the best pairwise alignment between the reference and alternate alleles.

[0136] [Figure 53]Figures 53A-53E show example FACS gating for GFP-positive cell sorting. Below is an example of the original batch analysis file, outlining the sorting strategy used to generate the HEXA 1278+TATC and HBB E6V HEK293T cell lines. Image data was generated on a Sony LE-MA900 cytometer using Cell Sorter software v.3.0.5. Graphic 1 shows a gating plot of cells not expressing GFP. Graphic 2 shows an example sorting of P2A-GFP-expressing cells used to isolate the HBB E6V HEK293T cell line. HEK293T cells were initially gated on populations using FSC-A / BSC-A (gate A) and then sorted for singlets using FSC-A / FSC-H (gate B). Live cells were sorted by gating on DAPI-negative cells (gate C). Cells with GFP fluorescence levels above that of negative control cells were sorted using EGFP as the fluorochrome (Gate D). Figure 53A shows HEK293T cells (GFP negative). Figure 53B shows a representative plot of FACS gating for cells expressing PE2-P2A-GFP. Figure 53C shows the genotype of HEXA 1278 + TATC homozygous HEK293T cells. Figures 53D-5E show the allele table of the HBB E6V homozygous HEK293T cell line.

[0137] [Figure 54] Figure 54 is a schematic diagram summarizing the PEG RNA cloning procedure.

[0138] [Figure 55]Figures 55A-55G are schematic diagrams of PEG RNA design. Figure 55A shows a simple diagram of a PEG RNA with its domains labeled (left) and bound to nCas9 at a genomic site (right). Figure 55B shows various types of modifications of PEG RNA expected to increase activity. Figure 55C shows modifications of PEG RNA to increase transcription of longer RNAs through promoter selection and 5', 3' processing and termination. Figure 55D shows the lengthening of the P1 system, an example of a backbone modification. Figure 55E shows that incorporation of synthetic modifications on the template region or elsewhere on the PEG RNA can increase activity. Figure 55F shows that the designed incorporation of minimal secondary structure on the template can prevent the formation of longer, more inhibitory secondary structures. Figure 55G shows a split PEG RNA with a second template sequence anchored by an RNA element at the 3' end of the PEG RNA (left). Incorporation of elements at the 5' or 3' end of the PEGRNA can enhance RT binding.

[0139] [Figure 56] Figures 56A-56D show integration of PEG RNA backbone sequences into target loci. HTS data were analyzed for PEG RNA backbone sequence insertions as described in Figures 60A-60B. Figure 56A shows an analysis of the EMX1 locus. Shown are the % of total sequencing reads containing one or more PEG RNA backbone sequence nucleotides on the insert adjacent to the RT template (left); the percentage of total sequencing reads containing PEG RNA backbone sequence insertions of a specified length (center); and the cumulative total percentage of PEG RNA insertions up to encompassing the specified length on the x-axis. Figure 56B shows the same as Figure 56A for FANCF. Figure 56C shows the same as Figure 56A for HEK3. Figure 56D shows the same as Figure 56A for RNF2. Values ​​and error bars reflect the mean and sd of three independent biological replicates.

[0140] [Figure 57]Figures 57A-57I show the effects of PE2, PE2-dRT, and Cas9 H840A nickase on transcriptome-wide RNA abundance. Analysis of ribosomal RNA-depleted cellular RNA isolated from HEK293T cells expressing PE2, PE2-dRT, or Cas9 H840A nickase and PRNP- or HEXA-targeted PEgRNA. RNAs corresponding to 14,410 and 14,368 genes were detected in the PRNP and HEXA samples, respectively. Figures 57A-57F show volcano plots displaying the -log10 FDR-adjusted p-value versus the log2-fold change in transcript abundance for each RNA. (Figure 57A) Compares PE2 vs. PE2-dRT with PRNP-targeted PEgRNA, (Figure 57B) PE2 vs. Cas9 H840A with PRNP-targeted PEgRNA, (Figure 57C) PE2-dRT vs. Cas9 H840A with PRNP-targeted PEgRNA, (Figure 57D) PE2 vs. PE2-dRT with HEXA-targeted PEgRNA, (Figure 57E) PE2 vs. Cas9 H840A with HEXA-targeted PEgRNA, and (Figure 57F) PE2-dRT vs. Cas9 H840A with HEXA-targeted PEgRNA. Red dots indicate genes showing a statistically significant ≥ 2-fold change in relative abundance (FDR-adjusted p<0.05). Figures 57G-57I are Venn diagrams of up- and down-regulated transcripts (≧2-fold change) comparing PRNP and HEXA samples for (Figure 57G) PE2 vs. PE2-dRT, (Figure 57H) PE2 vs. Cas9 H840A, and (Figure 57I) PE2-dRT vs. Cas9 H840A.

[0141] [Figure 58] Figures 58A-58B show representative FACS gating for neuronal nuclei sorting. Nuclei were sequentially gated based on DyeCycle Ruby signal, FSC / SSC ratio, SSC width / SSC height ratio, and GFP / DyeCycle ratio.

[0142] [Figure 59]Figures 59A-59G show the protocol for cloning 3'-extended PEG-RNA into a mammalian U6 expression vector by Golden Gate assembly. Figure 59A shows an overview of the cloning. Figure 59B shows "Step 1: Digesting the pU6-PEG-RNA-GG-vector plasmid (Component 1)." Figure 59C shows "Steps 2 and 3: Ordering and annealing oligonucleotide parts (Components 2, 3, and 4)." Figure 59D shows "Step 2.b.ii.: sgRNA backbone phosphorylation (not necessary if oligonucleotides are purchased phosphorylated)." Figure 59E shows "Step 4: PEG-RNA assembly." Figure 59F shows "Steps 5 and 6: Transformation of assembled plasmid." Figure 59G shows a diagram summarizing the PEG-RNA cloning protocol.

[0143] [Figure 60] Figures 60A-60B show a Python script for quantifying PEGRNA backbone incorporation. A custom python script was generated to characterize and quantify PEGRNA insertions at target genomic loci. The script iteratively matches text strings of increasing length taken from the reference sequence (guide RNA backbone sequence) with the sequencing reads in the fastq file and counts the number of sequencing reads that match the search query. Each successive text string corresponds to an additional nucleotide in the guide RNA backbone sequence. The exact length incorporation and cumulative incorporation up to a specified length were calculated in this manner. To ensure alignment and accurate counting of short slices of sgRNA, the start of the reference sequence encompasses 5 to 6 bases at the 3' end of the new DNA strand synthesized by reverse transcriptase.

[0144] [Figure 61] Figure 61 is a graph showing the percent of total sequencing reads with the defined edit for SaCas9(N580A)-MMLV RT HEK3 +6C>A. Values ​​for correct edits and indels are shown.

[0145] [Figure 62] Figures 62A-62B show the importance of the protospacer for efficient incorporation of the desired edit in precise positioning by prime editing. Figure 62A is a graph showing the percentage of total sequencing reads in which the targeted T·A base pair was converted to A·T for various HEK3 loci. Figure 62B is a sequence analysis showing this.

[0146] [Figure 63] Figure 63 is a graph showing SpCas9 PAM variants of PAM editing (N=3). The percentage of total sequencing reads with targeted PAM editing is shown for SpCas9(H840A)-VRQR-MMLV RT, where NGA>NTA, and for SpCas9(H840A)-VRER-MMLV RT, where NGCG>NTCG. The PEG RNA primer binding site (PBS) length, RT template (RT) length, and PE system used are listed.

[0147] [Figure 64] Figures 64A-64F are schematic diagrams showing the introduction of various site-specific recombinase (SSR) targets into a genome using PE. Figure 64A provides a general schematic diagram of the insertion of a recombinase target sequence by a primed editor. Figure 64B shows how a single SSR target inserted by PE can be used as a site for genomic integration of a DNA donor template. Figure 64C shows how tandem insertion of SSR target sites can be used to delete a portion of a genome. Figure 64D shows how tandem insertion of SSR target sites can be used to invert a portion of a genome. Figure 64E shows how insertion of two SSR target sites in two distal chromosomal regions can result in a chromosomal translocation. Figure 64F shows how insertion of two different SSR target sites into a genome can be used to exchange a cassette from a DNA donor template. See Example 17 for further details.

[0148] [Figure 65] Figure 65 shows 1) PE-mediated synthesis of an SSR target site on the genome of a human cell and 2) use of that SSR target site to incorporate a DNA donor template containing a GFP-expressing marker. Once successfully incorporated, the GFP causes the cell to fluoresce. See Example 17 for further details.

[0149] [Figure 66] Figure 66 illustrates one embodiment of a prime editor provided as two PE half proteins that regenerate the entire prime editor by self-splicing of the split-intein halves located at the end or beginning of each of the prime editor half proteins.

[0150] [Figure 67]Figure 67 illustrates the mechanism of intein removal from a polypeptide sequence and peptide bond reformation between the N- and C-terminal extein sequences. (a) illustrates the general mechanism of two protein halves, each containing half of an intein sequence. When these come into contact within a cell, they result in a fully functional intein. This then undergoes self-splicing and excision. The excision process results in the formation of a peptide bond between the N-terminal protein half (or "N-extein") and the C-terminal protein half (or "C-extein"), forming a whole single polypeptide containing the N-extein and C-extein portions. In various embodiments, the N-extein can correspond to the N-terminal half of a split prime editor fusion protein, and the C-extein can correspond to the C-terminal half of a split prime editor. (b) illustrates the chemical mechanism of intein excision and peptide bond reformation linking the N-extein half (red half) and the C-extein half (blue half). Because it involves the splicing action of two separate components provided in trans, excision of a split intein (i.e., the N intein and C intein in a split intein configuration) can also be referred to as "trans-splicing."

[0151] [Figure 68A] Figure 68A demonstrates that delivery of both split-intein halves of SpPE (SEQ ID NO: 762) in a linker maintains activity at the three test loci when co-transfected into HEK293T cells.

[0152] [Figure 68B] Figure 68B demonstrates that delivery of both split-intein halves of SaPE2 (e.g., SEQ ID NO: 443 and SEQ ID NO: 450) recapitulates the activity of full-length SaPE2 (SEQ ID NO: 134) when co-transfected into HEK293T cells.

[0153] PullThe residues indicated by symbols are SaCa s 9 The sequence of amino acids 741-743 (the first residues of the C-terminal extein) of α-glucanase 1 (SMP) is important for the intein trans-splicing reaction. "SMP" is the native residue. We also mutated these to the "CFN" consensus splicing sequence. The consensus sequence has been shown to yield the highest rearrangements, as measured by prime editing percentage.

[0154] [Figure 68C] Figure 68C provides data showing that various disclosed PE ribonucleoprotein complexes (high concentration PE2, high concentration PE3, and low concentration PE3) can be delivered in this manner.

[0155] [Figure 69] Figure 69 shows a bacteriophage plaque assay to determine PE efficacy in PANCE. Plaques (dark circles) indicate phage capable of successfully infecting E. coli. Increasing the concentration of L-rhamnose results in increased expression of PE and increased plaque formation. Plaque sequencing revealed the presence of genome editing incorporated by PE.

[0156] [Figure 70]Figures 70A-70I provide examples of edited target sequences as an illustration of step-by-step instructions for designing PEGRNAs and nicking sgRNAs for prime editing. Figure 70A: Step 1. Define the target sequence and edit. Retrieve the sequence of the target DNA region (~200 bp) centered around the location of the desired edit (point mutation, insertion, deletion, or a combination thereof). Figure 70B: Step 2. Locate the target PAM. Identify a PAM proximal to the edit location. Be careful to look for PAMs on both strands. While a PAM close to the edit location is preferred, it is possible to incorporate an edit using a protospacer and PAM that positions the nick ≥30 nt from the edit location. Figure 70C: Step 3. Locate the nick site. For each PAM to be considered, identify the corresponding nick site. For SpCas9 H840A nickase, cleavage occurs on the PAM-containing strand between the third and fourth bases 5' of the NGG PAM. All edited nucleotides must be 3' to the nick site. Therefore, an appropriate PAM must place a nick 5' to the target edit site on the PAM-containing strand. In the example shown below, there are two possible PAMs. For simplicity, the remaining steps demonstrate the design of a PEG RNA using only PAM 1. Figure 70D: Step 4. Design the spacer sequence. The protospacer of SpCas9 corresponds to the 20 nucleotides 5' to the NGG PAM on the PAM-containing strand. Efficient Pol III transcription initiation requires G to be the first transcribed nucleotide. If the first nucleotide of the protospacer is G, the spacer sequence of the PEG RNA is simply the protospacer sequence. If the first nucleotide of the protospacer is not G, the spacer sequence of the PEG RNA is G followed by the protospacer sequence. Figure 70E: Step 5. Design the primer binding site (PBS). The starting allele sequence is used to identify a DNA primer on the PAM-containing strand. The 3' end of the DNA primer is the nucleotide just upstream of the nick site (i.e., the fourth base 5' of the NGG PAM for Sp Cas9).As a general design principle for use with PE2 and PE3, a PEG RNA primer binding site (PBS) containing 12 to 13 nucleotides of complementarity to the DNA primer can be used for sequences containing ~40-60% GC content. For sequences with low GC content, longer PBSs (14 to 15 nt) should be tested. For sequences with higher GC content, shorter PBSs (8 to 11 nt) should be tested. Regardless of GC content, the optimal PBS sequence must be determined empirically. To design a PBS sequence of length p, use the starting allele sequence to take the reverse complement of the first p nucleotides 5' to the nick site on the PAM-containing strand. Figure 70F: Step 6. Design the RT template. The RT template encodes the designed edit and homology to the sequence flanking the edit. The optimal RT template length varies based on the target site. For short-range editing (positions +1 to +6), we recommend testing short (9 to 12 nt), medium (13 to 16 nt), and long (17 to 20 nt) RT templates. For long-range editing (positions +7 and beyond), we recommend using an RT template that extends at least 5 nt (preferably 10 nt or more) beyond the position of editing to allow for sufficient 3' DNA flap homology. For long-range editing, several RT templates should be screened to identify functional designs. For larger insertions and deletions (≥5 nt), we recommend incorporating greater 3' homology (~20 nt or more) into the RT template. Editing efficiency is typically compromised when the RT template encodes the synthesis of a G (corresponding to a C in the PEG RNA RT template) as the last nucleotide on the reverse-transcribed DNA product. Because many RT templates support efficient prime editing, we recommend avoiding G as the last synthesized nucleotide when designing an RT template. To design an RT template sequence of length r, use the desired allele sequence and take the reverse complement of the first r nucleotides 3' to the nick site on the strand that originally contained the PAM. Note that compared to SNP editing, insertion or deletion editing using an RT template of the same length will not contain identical homology. Figure 70G: Step 7. Assemble the complete PEG RNA sequence.Concatenate the PEgRNA components in the following order (5' to 3'): spacer, backbone, RT template, and PBS. Figure 70H: Step 8. Design a nicking sgRNA for PE3. Identify PAMs on the non-edited strand upstream and downstream of the editing site. The optimal nicking position is highly locus-dependent and must be empirically determined. Generally, a nick placed 40 to 90 nucleotides 5' opposite the PEgRNA-induced nick leads to higher editing yields and fewer indels. The nicking sgRNA has a spacer sequence that matches the 20-nt protospacer on the starting allele. If the protospacer does not begin with a G, add a 5' G. Figure 70I: Step 9. Design a PE3b nicking sgRNA. If a PAM is present on the complementary strand and its corresponding protospacer overlaps with the sequence targeted for editing, this edit may be a candidate for the PE3b system. In the PE3b system, the spacer sequence of the nicking sgRNA matches the sequence of the desired edited allele, not the starting allele. The PE3b system operates efficiently when the nucleotide(s) to be edited fall within the seed region (~10 nt adjacent to the PAM) of the nicking sgRNA protospacer. This prevents nicking of the complementary strand until after incorporation of the edited strand, preventing competition between the PEgRNA and sgRNA for binding to the target DNA. PE3b also avoids simultaneous nicking on both strands, thus significantly reducing indel formation while maintaining high editing efficiency. The PE3b sgRNA should have a spacer sequence that matches the 20 nt protospacer of the desired allele, with an additional 5' G if needed.

[0157] [Figure 71A]Figure 71A shows the nucleotide sequence of the SpCas9 PEGRNA molecule (top), which terminates at the 3' end with "UUU" and does not contain a toe loop element. The lower portion of the figure illustrates the same SpCas9 PEGRNA molecule, but further modified to contain a toe loop element with the sequence 5'-"GAAANNNNN"-3' inserted immediately prior to the "UUU" 3' end. "N" can be any nucleobase.

[0158] [Figure 71B] Figure 71B shows the results of Example 18, demonstrating that the efficiency of prime editing in HEK or EMX cells is increased using a PEGRNA containing a toe loop element, while the percentage of indel formation remains largely unchanged.

[0159] [Figure 72]Figures 72A-72C illustrate alternative PEgRNA configurations that can be used for prime editing. Figure 72A illustrates the PE2:PEgRNA embodiment of prime editing. This embodiment involves PE2 (a fusion protein comprising Cas9 and reverse transcriptase) complexed with PEgRNA (as also depicted in Figures 1A-1I and / or 3A-3E). In this embodiment, a reverse transcription template is incorporated onto the 3' extension arm of an sgRNA to create PEgRNA, and the DNA polymerase enzyme is reverse transcriptase (RT) fused directly to Cas9. Figure 72B illustrates the MS2cp-PE2:sgRNA+tPERT embodiment. This embodiment includes a PE2 fusion (Cas9+reverse transcriptase), which is further fused to the MS2 bacteriophage coat protein (MS2cp) to form an MS2cp-PE2 fusion protein. To achieve prime editing, the MS2cp-PE2 fusion protein is complexed with the sgRNA, which targets the complex to a specific target site on the DNA. Then, embodiments involve the introduction of a transprime editing RNA template ("tPERT"), which acts in place of the PEgRNA by providing a primer binding site (PBS) and a DNA synthesis template via a separate molecule, i.e., tPERT. It also contains an MS2 aptamer (stem-loop). The MS2cp protein recruits tPERT by binding to the molecule's MS2 aptamer. Figure 72C illustrates alternative designs of PEgRNA that can be achieved by known methods for chemically synthesizing nucleic acid molecules. For example, chemical synthesis can be used to synthesize hybrid RNA / DNA PEgRNA molecules for use in prime editing, where the extension arm of the hybrid PEgRNA is DNA instead of RNA. In such embodiments, a DNA-dependent DNA polymerase can be used instead of reverse transcriptase to synthesize a 3' DNA flap containing the desired genetic change formed by prime editing. In another embodiment, the extension arm can be synthesized to include a chemical linker, which prevents a DNA polymerase (e.g., reverse transcriptase) from using the sgRNA scaffold or backbone as a template.In yet another embodiment, the extension arm can comprise a DNA synthesis template that has a reverse orientation relative to the overall orientation of the PEG RNA molecule. For example, as shown for a PEG RNA with an extension attached to the 3' end of the sgRNA backbone and oriented 5' to 3', the DNA synthesis template faces the opposite direction, i.e., 3' to 5'. This embodiment can be advantageous for PEG RNA embodiments with an extension arm positioned at the 3' end of the gRNA. By reversing the orientation of the extension arm, DNA synthesis by a polymerase (e.g., reverse transcriptase) will terminate once it reaches the 5' end of the new orientation of the extension arm, and therefore there will be no risk of using the gRNA core as a template.

[0160] [Figure 73] Figure 73 demonstrates prime editing with tPERT and the MS2 recruitment system (also known as MS2 tagging technology). An sgRNA targeting the prime editor protein (PE2) to the target locus is expressed in combination with tPERT containing a primer binding site (13nt or 17nt PBS), a RT template encoding a His6 tag insert and homology arms, and an MS2 aptamer (located at the 5' or 3' end of the tPERT molecule). Either the prime editor protein (PE2) or a fusion of MS2cp to the N-terminus of PE2 was used. As with the previously developed PE3 system, editing was performed with or without a complementary strand nicking sgRNA (labeled "PE2+nick" or "PE2" on the x-axis, respectively). This is also referred to as "second strand nicking" and is defined herein.

[0161] [Figure 74]Figure 74 demonstrates the expression of the MS2 aptamer of reverse transcriptase in trans and its recruitment by the MS2 aptamer system. The PEGRNA contains the MS2 RNA aptamer inserted into either one of two sgRNA backbone hairpins. Wild-type M-MLV reverse transcriptase is expressed as an N- or C-terminal fusion to the MS2 coat protein (MCP). Editing is at the HEK3 site in HEK293T cells.

[0162] [Figure 75] Figure 75 provides a bar graph comparing the efficiency (i.e., "% of total sequencing reads with the defined edit or indel") of PE2, PE2-trunc, PE3, and PE3-trunc for different target sites in various cell lines. The data show that prime editors containing truncated RT variants were about as efficient as prime editors containing the untruncated RT protein.

[0163] [Figure 76] Figure 76 demonstrates the editing efficiency of the intein-split prime editor of Example 20. HEK239T cells were transfected with plasmids encoding full-length PE2 or intein-split PE2, PEgRNA, and a nicking guide RNA. The consensus sequence (most amino-terminal residue of the C-terminal extein) is indicated. Percent editing at two sites is shown: HEK3 +1 CTT insertion and PRNP +6 G to T. Replicates n = 3 independent transfections. See Example 20.

[0164] [Figure 77]Figure 77 demonstrates the editing efficiency of the intein split-primed editing factor of Example 20. Editing was assessed by delivery of 5E10vg per half of SpPE3 and a small amount of 1E10 nuclear-localized GFP:KASH to P0 mice via ICV injection, followed by targeted deep sequencing of bulk cortical and GFP+ subpopulations. The editing factor and GFP were packaged into AAV9 with an EFS promoter. Mice were harvested 3 weeks after injection, and GFP+ nuclei were isolated by flow cytometry. Individual data points are shown for 1-2 mice per condition analyzed. See Example 20.

[0165] [Figure 78] Figure 78 demonstrates the editing efficiency of the intein-split prime editor of Example 20. Specifically, the figure illustrates the AV-split SpPE3 construct used in Example 20. Co-transduction with AAV particles separately expressing SpPE3-N and SpPE3-C recapitulates PE3 activity. Note that the N-terminal genome contains a U6-sgRNA cassette expressing a nicking sgRNA, and the C-terminal genome contains a U6-PEgRNA cassette expressing a PEgRNA. See Example 20.

[0166] [Figure 79]Figure 79 shows the editing efficiency of certain optimized linkers discussed in Example 21. In particular, the data show the editing efficiency of the PE2 construct with the current linker (labeled PE2; white boxes) compared to various versions in which the linker was replaced by the indicated sequence for representative PEgRNAs for transition, transversion, insertion, and deletion edits at the HEK3, EMX1, FANCF, and RNF2 loci. The replacement linkers are referred to as "1xSGGS" (SEQ ID NO: 174), "2xSGGS" (SEQ ID NO: 446), "3xSGGS" (SEQ ID NO: 3889), "1xXTEN" (SEQ ID NO: 171), "no linker," "1xGly," "1xPro," "1xEAAAK" (SEQ ID NO: 3968), "2xEAAAK" (SEQ ID NO: 3969), and "3xEAAAK" (SEQ ID NO: 3970). Editing efficiency is measured in bar graph format relative to the "control" editing efficiency of PE2. The linker for PE2 is SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 127). All edits were made in the context of the PE3 system. That is, this refers to the PE2 editing construct plus the addition of an optimal secondary sgRNA nicking guide. See Example 21.

[0167] [Figure 80] Taking the average 2-fold potency relative to PE2 yields the graph shown, which indicates that use of the 1×XTEN (SEQ ID NO: 171) linker sequence improves editing efficiency by an average of 1.14-fold (n=15). See Example 21.

[0168] [Figure 81] Figure 81 illustrates the transcription levels of PEG RNA from different promoters, as described in Example 22.

[0169] [Figure 82] As illustrated in Example 22, the impact of different types of modifications on the PEGRNA structure on the editing efficiency relative to unmodified PEGRNA.

[0170] [Figure 83] Figure 83 illustrates a PE experiment that targeted editing of a HEK3 gene. Insertion of a 10 nt insertion at position +1 relative to the nick site was specifically targeted and PE3 was used. See Example 22.

[0171] [Figure 84A] Figure 84A illustrates an exemplary PEG-gRNA having a spacer, gRNA core, and extension arm (RT template + primer binding site). This is modified at the 3' end of the PEG-gRNA with a tRNA molecule coupled via a UCU linker. The tRNA may contain various post-transcriptional modifications; however, these modifications are not required.

[0172] [Figure 84B] Figure 84B illustrates the structure of a tRNA that can be used to modify the PEGRNA structure. See Example 22. P1 can be variable in length. P1 can be elongated to help prevent RNAse P processing of the PEGRNA-tRNA fusion.

[0173] [Figure 85] Figure 85 illustrates a PE experiment targeting editing of the FANCF gene. The G to T conversion at position +5 relative to the nick site was specifically targeted and used the PE3 construct. See Example 22.

[0174] [Figure 86] Figure 86 illustrates a PE experiment that targeted editing of a HEK3 gene. Insertion of a 71 nt FLAG tag insertion at position +1 relative to the nick site was specifically targeted and used the PE3 construct. See Example 22.

[0175] [Figure 87] Results from screening N2A cells for pegRNA incorporating 1412Adel, with details of primer binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels). See Example 23.

[0176] [Figure 88] Results from screening N2A cells for pegRNA incorporating 1412Adel, with details of primer binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels). See Example 23.

[0177] [Figure 89] Figure 89 illustrates the results of editing at the proxy locus of the β-globin gene in healthy HSCs and in HEK3. Varying concentrations of editing factors versus pegRNA and nicking gRNA. See Example 23.

[0178] definition Unless otherwise defined, all technical and scientific terms used in this application have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Unless otherwise defined, the following terms used in this application have the meanings ascribed to them.

[0179] antisense strand In genetics, the "antisense" strand of a segment of double-stranded DNA is considered to be the template strand, running in a 3' to 5' direction. In contrast, the "sense" strand is the 5' to 3' running segment of double-stranded DNA that is complementary to the antisense or template strand of the 3' to 5' running DNA. In the case of a DNA segment encoding a protein, the sense strand is the strand of DNA with the same sequence as the mRNA that takes the antisense strand as its template during transcription and ultimately (typically, but not always) is translated into the protein. Thus, the antisense strand carries the RNA that is subsequently translated into protein, while the sense strand possesses nearly identical makeup to the mRNA. Note that for each segment of dsDNA, there will likely be two pairs: sense and antisense (since sense and antisense are relative perspectives), depending on the direction from which one is read. The gene product or mRNA ultimately refers to either strand of a segment of dsDNA, which may be designated as sense or antisense.

[0180] Dual-specific ligands As used herein, the term "dual-specific ligand" or "dual-specific moiety" refers to a ligand that binds to two different ligand-binding domains. In some embodiments, the ligand is a small molecule compound or a peptide or polypeptide. In other embodiments, the ligand-binding domain is a "dimerization domain," which can be incorporated onto a protein as a peptide tag. In various embodiments, two proteins, each containing the same or different dimerization domain, can be induced to dimerize by binding of each dimerization domain to a dual-specific ligand. As used herein, a "dual-specific ligand" can equivalently be referred to as a "chemical inducer of dimerization" or "CID."

[0181] Cas9 The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease comprising a Cas9 domain or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). A "Cas9 domain," as used herein, is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nuclease is also sometimes referred to as casn1 nuclease or CRISPR (clustered regularly interspaced short palindromic repeats (CRISPR)). C Lustered R regularly I Interspaced S hort P alindromic RCRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to the preceding mobile element, and a targeting invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, corrective processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for ribonuclease 3-aided processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved endonucleolytically, followed by 3'-5' exonucleolytic excision. In nature, DNA binding and cleavage typically require both a protein and an RNA. However, single guide RNAs ("sgRNAs," or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat (the PAM or protospacer adjacent motif) to help distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski K., Sharma See CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences also include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0182] A Cas9 domain with inactivated nuclease may be interchangeably referred to as a "dCas9" protein (representing a nuclease-inactivated Cas9). Methods for generating a Cas9 domain (or a fragment thereof) with an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83 (the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof.For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18). In some embodiments, the Cas9 variant comprises a fragment of SEQ ID NO: 18 Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18).

[0183] cDNA The term "cDNA" refers to a strand of DNA copied from an RNA template. The cDNA is complementary to the RNA template.

[0184] circular permutants As used herein, the term "circular permutant" refers to a protein or polypeptide (e.g., Cas9) containing a circular permutation, which is a change in the structural configuration of a protein accompanied by a change in the order of amino acids appearing in the protein's amino acid sequence. In other words, a circular permutant is a protein with altered N- and C-termini compared to its wild-type counterpart, e.g., the C-terminal half of the wild-type protein becomes a new N-terminal half. A circular permutation (or CP) is essentially a topological rearrangement of a protein's primary sequence, with its N- and C-termini connected, often with a peptide linker, while simultaneously splitting the sequence at a different location to create new adjacent N- and C-termini. The result is a protein structure that may have different connections but often the same or an overall similar three-dimensional (3D) shape, potentially including improved or altered characteristics, including reduced susceptibility to proteolysis, improved catalytic activity, altered substrate or ligand binding, and / or improved thermostability. Circularly permuted proteins can occur naturally (e.g., concanavalin A and lectins). In addition, circular permutations can occur as a result of post-translational modifications or can be engineered using recombinant techniques.

[0185] Circularly permuted Cas9 The term "circularly permuted Cas9" refers to any Cas9 protein or variant thereof that has been generated as a circular permutation, thereby locally rearranging its N- and C-termini. Such circularly permuted Cas9 proteins ("CP-Cas9"), or variants thereof, retain the ability to bind to DNA when complexed with a guide RNA (gRNA). See Oakes et al., "Protein Engineering of Cas9 for Enhanced Function," Methods Enzymol, 2014, 546:491-511 and Oakes et al., "CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification," Cell, January 10, 2019, 176:254-267 (each of which is incorporated herein by reference). The present disclosure contemplates the use of any previously known CP-Cas9 or new CP-Cas9, so long as the resulting circularly permuted protein retains the ability to bind to DNA when complexed with a guide RNA (gRNA). Exemplary CP-Cas9 proteins are set forth in SEQ ID NOS: 77-86.

[0186] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of prior infection by viruses that have invaded prokaryotes. These snippets of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, effectively forming a prokaryotic immune defense system in conjunction with an array of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endolytically cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, the target strand not complementary to the crRNA is first endolytically cut, then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require both proteins and RNAs. However, single-guide RNAs ("sgRNAs" or simply "gNRAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into the guide RNA of a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) on CRISPR repeats to help distinguish self from non-self.CRISPR biology and Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP,Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski See K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0187] In certain types of CRISPR systems (e.g., type II CRISPR systems), proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endolytically cleaves linear or circular nucleic acid targets complementary to the RNA. Specifically, the target strand not complementary to the crRNA is first endolytically cut and then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require a protein and both RNAs. However, single-guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both the crRNA and tracrRNA into the guide RNA of a single RNA species.

[0188] Generally, a "CRISPR system" collectively refers to the transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated ("Cas") genes, and includes sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or an active partial tracrRNA), tracr mate sequences (which, in the context of an endogenous CRISPR system, encompass "direct repeats" and partial direct repeats processed by the tracrRNA), guide sequences (also referred to as "spacers" in the context of an endogenous CRISPR system), or other sequences and transcripts from the CRISPR locus. The tracrRNA of the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA.

[0189] DNA synthesis template As used herein, the term "DNA synthesis template" refers to the region or portion of the extension arm of a PEgRNA that is utilized by the polymerase of a prime editor as the template strand (encoding a 3' single-stranded DNA flap that contains the desired edit and then replaces the corresponding strand of endogenous DNA at the target site through the mechanism of prime editing). In various embodiments, DNA synthesis templates are shown in Figure 3A (in the context of a PEgRNA comprising a 5' extension arm), Figure 3B (in the context of a PEgRNA comprising a 3' extension arm), Figure 3C (in the context of an internal extension arm), Figure 3D (in the context of a 3' extension arm), and Figure 3E (in the context of a 5' extension arm). The extension arm encompasses the DNA synthesis template and may be composed of DNA or RNA. In the case of RNA, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of DNA, the polymerase of the prime editor can be a DNA-dependent DNA polymerase. In various embodiments (e.g., as depicted in Figures 3D-3E), the DNA synthesis template (4) can include all or part of the "editing template" and "homology arm," as well as any 5'-terminal modification region e2. That is, depending on the nature of the e2 region (e.g., whether it includes a hairpin, toe-loop, or stem / loop secondary structure), the polymerase can encode none, part, or all of the e2 region. In other words, in the case of a 3' extension arm, the DNA synthesis template (3) can include a portion of the extension arm (3) extending from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core, which can serve as a template for synthesis of a single strand of DNA by a polymerase (e.g., reverse transcriptase). In the case of a 5' extension arm, the DNA synthesis template (3) can include a portion of the extension arm (3) extending from the 5' end of the PEGRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of PEGRNAs with either a 3' extension arm or a 5' extension arm.Certain embodiments described herein (e.g., Figure 71A) refer to an "RT template" that encompasses the editing template and homology arms, i.e., the sequences of the PEGRNA extension arms that are actually used as templates during DNA synthesis. The term "RT template" is equivalent to the term "DNA synthesis template."

[0190] In the case of trans-prime editing (e.g., Figures 3G and 3H), the primer binding site (PBS) and DNA synthesis template can be engineered into separate molecules termed trans-prime editor RNA template (tPERT).

[0191] Dimerization domain The term "dimerization domain" refers to a ligand-binding domain that binds to a binding domain of a dual-specific ligand. A "first" dimerization domain binds to a first binding moiety of a dual-specific ligand, and a "second" dimerization domain binds to a second binding moiety of the same dual-specific ligand. When a first dimerization domain is fused to a first protein (e.g., via PE as discussed herein) and a second dimerization domain is fused to a second protein (e.g., via PE as discussed herein), the first and second proteins dimerize in the presence of the dual-specific ligand, and the dual-specific ligand has at least one moiety that binds to the first dimerization domain and at least another moiety that binds to the second dimerization domain.

[0192] downstream As used herein, the terms "upstream" and "downstream" are relative terms that define the linear positions of at least two elements located in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in a 5' to 3' direction. In particular, if a first element is located anywhere 5' relative to the second element, the first element is upstream of the second element in the nucleic acid molecule. For example, if a SNP is located 5' to the nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, if a first element is located anywhere 3' relative to the second element, the first element is downstream of the second element in the nucleic acid molecule. For example, if a SNP is located 3' to the nick site, the SNP is downstream of the Cas9-induced nick site. The nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The terms upstream and downstream refer only to a single strand of a nucleic acid molecule, except when it is necessary to select which strand of a double-stranded molecule is considered, and the analysis is the same for single-stranded nucleic acid molecules and double-stranded molecules. Often, the strand of double-stranded DNA that can be used to determine the relative positions of at least two elements is the "sense" strand or "coding" strand. In genetics, the "sense" strand is the segment within double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of DNA that extends from 3' to 5'. Thus, for example, if a SNP nucleobase is located 3' from the promoter on the sense strand or coding strand, the SNP nucleobase is "downstream" of the promoter sequence in genomic DNA (which is double-stranded).

[0193] Editing Template The term "editing template" refers to the portion of the extension arm that encodes the desired edit on the single-stranded 3' DNA flap synthesized by a polymerase, e.g., a DNA-dependent DNA polymerase, an RNA-dependent DNA polymerase (e.g., reverse transcriptase). Certain embodiments described herein (e.g., Figure 71A) refer to an "RT template," which refers to both the editing template and the homologous arm together, i.e., the sequence of the PEG RNA extension arm that is actually used as a template during DNA synthesis. The term "RT editing template" is also equivalent to the term "DNA synthesis template," except that an RT editing template reflects the use of a primed editor with a polymerase that is a reverse transcriptase, while a DNA synthesis template more broadly reflects the use of a primed editor with either polymerase.

[0194] Effective dose As used herein, the term "effective amount" refers to an amount of a bioactive agent sufficient to induce a desired biological response. For example, in some embodiments, an effective amount of a prime editor (PE) can refer to the amount of the editor sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, an effective amount of a prime editor (PE) provided herein, such as a fusion protein comprising a nickase Cas9 domain and a reverse transcriptase, can refer to the amount of the fusion protein sufficient to induce editing of a target site specifically bound and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, such as a fusion protein, nuclease, hybrid protein, protein dimer, protein (or protein dimer) and polynucleotide complex, or polynucleotide, can vary depending on various factors, such as the desired biological response, e.g., the specific allele, genome, or target site to be edited, the cell or tissue to be targeted, and the agent to be used.

[0195] Error-prone reverse transcriptase As used herein, the term "error-prone" reverse transcriptase (or, more broadly, any polymerase) refers to a reverse transcriptase (or, more broadly, any polymerase) that is naturally occurring or derived from another reverse transcriptase (e.g., wild-type M-MLV reverse transcriptase) that has an error rate less than that of wild-type M-MLV reverse transcriptase. The error rate of wild-type M-MLV reverse transcriptase has been reported to range from 15,000 (higher) to 27,000 (lower) errors per 1 error. An error rate of 1 in 15,000 is 6.7 x 10 -5 The error rate of 1 in 27,000 is 3.7 x 10 -5 See Boutabout et al. (2001) "DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Ty1," Nucleic Acids Res 29(11):2217-2222, which is incorporated herein by reference. Thus, for purposes of this application, the term "error-prone" refers to a rate of more than 1 error (6.7 x 10) in the incorporation of 15,000 nucleobases. -5 or higher), e.g., 1 error in 14,000 nucleobases (7.14 × 10 -5 or higher), 1 error (7.7 × 10) in 13,000 nucleobases or less -5 or higher), 1 error (7.7 × 10) in 12,000 nucleobases or less -5 or higher), 1 error in 11,000 nucleobases or less (9.1 × 10 -5 or higher), 1 error (1 × 10) in 10,000 nucleobases or less -4or 0.0001 or higher), 1 error in 9,000 nucleobases or less (0.00011 or higher), 1 error in 8,000 nucleobases or less (0.00013 or higher), 1 error in 7,000 nucleobases or less (0.00014 or higher), 1 error in 6,000 nucleobases or less (0.00016 or higher), 1 error in 5,000 nucleobases or less (0.0002 or higher), 1 error in 4,000 nucleobases or less "RTs" refer to those RTs that have an error rate that is one error in 3,000 or fewer nucleic acid bases (0.00025 or higher), one error in 3,000 or fewer nucleic acid bases (0.00033 or higher), one error in 2,000 or fewer nucleic acid bases (0.00050 or higher), or one error in 1,000 or fewer nucleic acid bases (0.001 or higher), or one error in 500 or fewer nucleic acid bases (0.002 or higher), or one error in 250 or fewer nucleic acid bases (0.004 or higher).

[0196] Extein The term "extein" as used herein refers to a polypeptide sequence that is flanked by an intein and ligated to another extein during the protein splicing process to form a mature spliced ​​protein. Typically, an intein is flanked by two extein sequences, which are ligated together when the intein catalyzes its own excision. An extein is thus a protein analog to an exon found on an mRNA. For example, a polypeptide containing an intein can have the structure extein(N)-intein-extein(C). After excision of the intein and splicing of the two exteins, the resulting structure is extein(N)-extein(C) and a free intein. In various configurations, the exteins can be separate proteins (e.g., half of a Cas9 or PE fusion protein), each fused to a split-intein, and excision of the split-intein causes the extein sequences to be spliced ​​together.

[0197] Extension arm The term "extension arm" refers to a nucleotide sequence component of a PEG RNA that serves several functions, including a primer binding site and an editing template for reverse transcriptase. In some embodiments, e.g., in FIG. 3D, the extension arm is located at the 3' end of the guide RNA. In other embodiments, e.g., in FIG. 3E, the extension arm is located at the 5' end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In various embodiments, the extension arm comprises the following components in a 5' to 3' direction: a homology arm, an editing template, and a primer binding site. Because the polymerization activity of reverse transcriptase is in a 5' to 3' direction, the preferred arrangement of the homology arm, editing template, and primer binding site is in a 5' to 3' direction such that the reverse transcriptase, once primed by the annealed primer sequence, polymerizes single-stranded DNA using the editing template as the complementary template strand. Further details, such as the length of the extension arm, are described elsewhere herein.

[0198] The extension arm may also be described as generally comprising two regions: a primer binding site (PBS) and a DNA synthesis template, as illustrated in Figure 3G (top row). The primer binding site binds to a primer sequence formed from the endogenous DNA strand of the target site when it becomes nicked by a prime editor complex, thereby exposing its 3' end on the nicked endogenous strand. As described herein, binding of the primer sequence to the primer binding site on the extension arm of the PEGRNA creates a double-stranded region with an exposed 3' end (i.e., 3' of the primer sequence), which then provides a substrate for a polymerase that begins polymerizing single strands of DNA from the exposed 3' end along the length of the DNA synthesis template. The sequence of the single-stranded DNA product is the complement of the DNA synthesis template. Polymerization continues 5' of the DNA synthesis template (or extension arm) until polymerization terminates. Thus, the DNA synthesis template represents the portion of the extension arm that is encoded by the polymerase of the primed editor complex into a single-stranded DNA product (i.e., a 3' single-stranded DNA flap containing the desired gene editing information), which replaces the corresponding endogenous DNA strand at the target site immediately downstream of the PE-induced nick site. Without being bound by theory, polymerization of the DNA synthesis template continues toward the 5' end of the extension arm until a termination event occurs. Polymerization can terminate in a variety of ways, including, but not limited to, (a) reaching the 5' end of the PE-gRNA (e.g., in the case of a 5' extension arm, the DNA polymerase simply runs out of template), (b) reaching an impassable RNA secondary structure (e.g., a hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a signal on a nucleic acid form such as supercoiled DNA or RNA.

[0199] Flap endonucleases (e.g., FEN1) As used herein, the term "flap endonuclease" refers to an enzyme that catalyzes the removal of 5' single-stranded DNA flaps. These are naturally occurring enzymes that process the removal of 5' flaps formed during cellular processes, including DNA replication. The prime editing methods described herein may utilize endogenously supplied flap endonucleases or those provided in trans to remove 5' flaps of endogenous DNA formed at the target site during prime editing. Flap endonucleases are known in the art and can be found in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10):4507-4519 and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2):198-211, and Balakrishnan et al., "Flap Endonuclease 1," Annu Rev Biochem, 2013, Vol 82:119-138, each of which is incorporated herein by reference. An exemplary flap endonuclease is FEN1, which can be represented by the following amino acid sequence: [Table 1]

[0200] functional equivalent The term "functional equivalent" refers to a second biomolecule that is functionally equivalent, but not necessarily structurally equivalent, to a first biomolecule. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9, but does not necessarily have the same amino acid sequence. In the context of this disclosure, the specification refers throughout to "protein X, or a functional equivalent thereof." In this context, a "functional equivalent" of protein X embraces any homolog, paralog, fragment, naturally occurring version, modified version, mutated version, or synthetic version that bears the equivalent function of protein X.

[0201] fusion proteins The term "fusion protein," as used herein, refers to a hybrid polypeptide comprising protein domains from at least two different proteins. A protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) portion of the protein, thereby forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. A protein may contain different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9, which directs the protein to bind to a target site) and a nucleic acid-cleavage domain or catalytic domain of a nucleic acid-editing protein. Another example includes Cas9 fused to a reverse transcriptase or its equivalent. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference).

[0202] Gene of interest (GOI) The term "gene of interest" or "GOI" refers to a gene encoding a biomolecule of interest (e.g., a protein or an RNA molecule). The protein of interest may include any intracellular protein, membrane protein, or extracellular protein, such as a nuclear protein, a transcription factor, a nuclear membrane transporter, an intracellular organelle-associated protein, a membrane receptor, a catalytic protein, and an enzyme, a therapeutic protein, a membrane protein, a membrane transport protein, a signal transduction protein, or an immunological protein (e.g., an IgG or other antibody protein), etc. The gene of interest may also encode an RNA molecule, including, but not limited to, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), antisense RNA, guide RNA, microRNA (miRNA), small interfering RNA (siRNA), and cell-free RNA (cfRNA).

[0203] Guide RNA ("gRNA") As used herein, the term "guide RNA" refers to a specific type of guide nucleic acid, most commonly associated with the Cas protein of CRISPR-Cas9, that binds to Cas9 and directs the Cas9 protein to a specific sequence in a DNA molecule (including complementarity with the protospace sequence of the guide RNA). However, the term also embraces equivalent guide nucleic acid molecules, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), that bind to a Cas9 equivalent, homolog, ortholog, or paralog and otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. Cas9 equivalents may also include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are also provided herein. As used herein, "guide RNA" may also be referred to as "existing guide RNA" to contrast with the modified form of guide RNA called "prime editing guide RNA" (or "PEgRNA") invented for the prime editing methods and compositions disclosed herein.

[0204] Guide RNAs or PEGRNAs may contain a variety of structural elements, including but not limited to:

[0205] Spacer sequence - a sequence in the guide RNA or PEgRNA (having a length of approximately 20 nt) that binds to the protospacer in the target DNA.

[0206] gRNA core (or gRNA backbone or backbone sequence) - refers to the sequence within the gRNA responsible for Cas9 binding, this does not include the spacer / targeting sequence used to guide Cas9 to the target DNA.

[0207] Extension Arm—A single-stranded extension at the 3' or 5' end of the PEG RNA, which contains a primer binding site and a DNA synthesis template sequence encoding a single-stranded DNA flap containing the desired genetic alteration by a polymerase (e.g., reverse transcriptase), which is then incorporated onto the endogenous DNA by displacing the corresponding endogenous strand, thereby incorporating the desired genetic alteration.

[0208] Transcription terminator-guide RNAs or PEgRNAs may contain a transcription termination sequence 3' to the molecule.

[0209] Homologous arms The term "homologous arm" refers to the portion or portions of the extension arm that encode the resulting reverse transcriptase-encoded single-stranded DNA flap (which is incorporated into the target DNA site by displacing the endogenous strand). The portion of the single-stranded DNA flap encoded by the homologous arm is complementary to the non-edited strand of the target DNA sequence, facilitating its displacement of the endogenous strand and its annealing with the single-stranded DNA flap in its place, thereby incorporating the edit. This component is further defined elsewhere. By definition, the homologous arm is part of the DNA synthesis template because it is encoded by the polymerase of the prime editor described herein.

[0210] host cell The term "host cell," as used herein, refers to a cell that can host, replicate, and express a vector described herein, e.g., a vector containing a nucleic acid molecule encoding a fusion protein comprising Cas9 or a Cas9 equivalent and a reverse transcriptase.

[0211] Intein As used herein, the term "intein" refers to a self-processing polypeptide domain found in organisms from all domains of life. Inteins (intervening proteins) carry out a unique self-processing event known as protein splicing, in which they excise themselves from a larger precursor polypeptide by cleaving two peptide bonds, ligating flanking extein (external protein) sequences together in the process by forming new peptide bonds. Because intein genes are found embedded in-frame with other protein-coding genes, this rearrangement occurs post-translationally (or co-translationally). Furthermore, intein-mediated protein splicing is spontaneous; it requires only the folding of the intein domain, not external factors or energy sources. This process is also known as cis protein splicing, in contrast to the natural process of trans protein splicing by "split inteins." Inteins are the protein equivalents of self-splicing RNA introns (see Perler et al., Nucleic Acids Res. 22:1125-1127 (1994)), which catalyze their own excision from precursor proteins and have the concomitant fusion of flanking protein sequences known as exteins (reviewed in Perler et al., Curr. Opin. Chem. Biol. 1:292-299 (1997); Perler, FB Cell 92(1):1-4 (1998); Xu et al., EMBO J. 15(19):5146-5153 (1996)).

[0212] As used herein, the term "protein splicing" refers to the process by which internal regions (inteins) of precursor proteins are excised and flanking regions (exteins) of the protein are ligated to form the mature protein. This natural process has been observed in numerous proteins from both prokaryotes and eukaryotes (Perler, F.B., Xu, MQ, Paulus, H. Current Opinion in Chemical Biology 1997, 1, 292-299; Perler, F.B., Nucleic Acids Research 1999, 27, 346-347). Intein units contain the necessary components required to catalyze protein splicing and often contain an endonuclease domain responsible for intein mobility (Perler, FB, Davis, EO, Dean, GE, Gimble, FS, Jack, WE, Neff, N., Noren, CJ, Thomer, J., Belfort, M. Nucleic Acids Research 1994, 22, 1127-1127). However, the resulting proteins are linked and not expressed as separate proteins. Protein splicing can also be performed in trans by split inteins expressed on separate polypeptides. These spontaneously combine to form a single intein, which then undergoes the protein splicing process to link separate proteins.

[0213] Elucidation of the mechanisms of protein splicing has led to numerous intein-based applications (Comb, et al., U.S. Patent No. 5,496,714; Comb, et al., U.S. Patent No. 5,834,247; Camarero and Muir, J. Amer. Chem. Soc., 121:5597-5598 (1999); Chong, et al., Gene, 192:271-281 (1997); Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al., J. Am. Chem. Soc., 121:1100-1101 (1999); Evans, et al. al.,J.Biol.Chem.,274:18359-18363(1999);Evans,et al.,J.Biol.Chem.,274:3923-3926(1999);Evans,et al.,Protein Sci.,7:2256-2264(1998);Evans,et al. al.,J.Biol.Chem.,275:9091-9094(2000);Iwai and Pluckthun,FEBS Lett.459:166-172(1999);Mathys,et al.,Gene,231:1-13(1999);Mills,et al.,Proc.Natl.Acad.Sci.USA 95:3543-3548(1998);Muir,et al.,Proc.Natl.Acad.Sci.USA 95:6705-6710(1998);Otomo,et al.,Biochemistry 38:16040-16044(1999);Otomo,et al.,J.Biolmol.NMR 14:105-114(1999);Scott,et al. al.,Proc.Natl.Acad.Sci.USA 96:13638-13643(1999);Severinov and Muir,J.Biol.Chem.,273:16205-16209(1998);Shingledecker,et al.,Gene,207:187-195(1998);Southworth,et al.,EMBO J.17:918-926(1998);Southworth,et al.,Biotechniques,27:110-120(1999);Wood,et al.,Nat.Biotechnol.,17:889-892(1999);Wu,et al.,Proc.Natl.Acad.Sci.USA 95:9226-9231(1998a);Wu,et al.,Biochim Biophys Acta 1387:422-432(1998b);Xu,et al.,Proc.Natl.Acad.Sci.USA 96:388-393(1999);Yamazaki,et al. al., J. Am. Chem. Soc., 120:5591-5592 (1998)). Each reference is incorporated herein by reference.

[0214] Ligand-dependent inteins As used herein, the term "ligand-dependent intein" refers to an intein that includes a ligand-binding domain. Typically, the ligand-binding domain is inserted into the amino acid sequence of the intein, resulting in the structure intein(N)-ligand-binding domain-intein(C). Typically, a ligand-dependent intein exhibits no or minimal protein splicing activity in the absence of an appropriate ligand, and exhibits significantly increased protein splicing activity in the presence of a ligand. In some embodiments, a ligand-dependent intein exhibits no observable splicing activity in the absence of a ligand, but exhibits splicing activity in the presence of a ligand. In some embodiments, the ligand-dependent intein exhibits observable protein splicing activity in the absence of ligand, and in the presence of an appropriate ligand, protein splicing activity that is at least 5-fold, at least 10-fold, at least 50-fold, at least 100-fold, at least 150-fold, at least 200-fold, at least 250-fold, at least 500-fold, at least 1000-fold, at least 1500-fold, at least 2000-fold, at least 2500-fold, at least 5000-fold, at least 10,000-fold, at least 20,000-fold, at least 25,000-fold, at least 50,000-fold, at least 100,000-fold, at least 500,000-fold, or at least 1,000,000-fold greater than the activity observed in the absence of ligand. In some embodiments, the increase in activity is dose-dependent over at least one order of magnitude, at least two orders of magnitude, at least three orders of magnitude, at least four orders of magnitude, or at least five orders of magnitude, allowing for fine-tuning of intein activity by adjusting the concentration of the ligand.Suitable ligand-dependent inteins are known in the art and include those provided below and those disclosed in published U.S. patent application no. US2014 / 0065711A1; Mootz et al., "Protein splicing triggered by a small molecule." J. Am. Chem. Soc. 2002; 124, 9044-9045; Mootz et al., "Conditional protein splicing: a new tool to control protein structure and function in vitro and in vivo." J. Am. Chem. Soc. 2003; 125, 10561-10569; Buskirk et al., Proc. Natl. Acad. Sci. USA. 2004; 101, 10505-10510); Skretas & Wood, "Regulation of protein activity with small-molecule-controlled inteins." Protein Sci. 2005;14,523-532; Schwartz, et al., "Post-translational enzyme activation in an animal via optimized conditional protein splicing." Nat. Chem. Biol. 2007;3,50-54; Peck et al., Chem. Biol. 2011;18(5),619-630, the entire contents of each of which are incorporated herein by reference. Exemplary sequences are as follows: [Table 2-1] [Table 2-2]

[0215] Linker The term "linker" as used herein refers to a molecule that connects two other molecules or moieties. A linker can be an amino acid sequence, in the case of a linker that connects two fusion proteins. For example, Cas9 can be fused to a reverse transcriptase enzyme via an amino acid linker sequence. A linker can also be a nucleotide sequence, in the case of linking two nucleotide sequences together. For example, in this case, a conventional guide RNA is linked to the RNA extension of a prime-edited guide RNA, which can include an RT template sequence and an RT primer binding site, via a spacer or linker nucleotide sequence. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.

[0216] isolated "Isolated" means altered or removed from its natural state. For example, a nucleic acid or peptide that is naturally present in a living animal is not "isolated," but the same nucleic acid or peptide partially or completely separated from its natural state from coexisting materials is "isolated." An isolated nucleic acid or protein can exist in a substantially purified form or can exist in a non-native environment (e.g., a host cell).

[0217] In some embodiments, the gene of interest is encoded by an isolated nucleic acid. As used herein, the term "isolated" refers to the characteristic of a material as provided herein, which is removed from its original or native environment (e.g., the natural environment if it occurs in nature). Thus, a naturally occurring polynucleotide or protein or polypeptide present in a living animal is not isolated, but the same polynucleotide or polypeptide separated from some or all of the coexisting materials in a natural system by human intervention is isolated. Artificial or modified materials, such as non-naturally occurring nucleic acid constructs, such as the expression constructs and vectors described herein, are also consequently referred to as isolated. A material does not need to be purified to be isolated. Consequently, a material may be part of a vector and / or part of a composition, and such a vector or composition is still isolated in that it is not part of the environment in which the material is found in nature.

[0218] MS2 tagging technology In various embodiments (e.g., as illustrated in Figures 72-73 and the embodiment of Example 19), the term "MS2 tagging technology" refers to the combination of an "RNA-protein interaction domain" (also known as an "RNA-protein recruitment domain or protein") paired with an RNA-protein interaction domain, e.g., an RNA-binding protein that specifically recognizes and binds to a particular hairpin structure. These types of systems can be exploited to recruit various functionalities to a prime editor complex bound to a target site. MS2 tagging technology is based on the natural interaction of the MS2 bacteriophage coat protein ("MCP" or "MS2cp") with a stem-loop or hairpin structure, i.e., an "MS2 hairpin," present on the genome of the phage. In the case of prime editing, MS2 tagging technology involves introducing an MS2 hairpin onto the desired RNA molecule (e.g., PEgRNA or tPERT) involved in prime editing. This then constitutes a specific, interactive binding target for the RNA-binding protein that recognizes and binds to that structure. In the case of the MS2 hairpin, it is recognized and bound by the MS2 bacteriophage coat protein (MCP), and if the MCP is fused to another protein (e.g., reverse transcriptase or other DNA polymerase), the MS2 hairpin can be used to "recruit" other proteins in trans to the target site occupied by the prime editing complex.

[0219] The prime editors described herein may incorporate, as an aspect, any known RNA-protein interaction domain to recruit or "co-localize" a particular function of interest to the prime editor complex. Reviews of other modular domains of RNA-protein interactions are described in the art, for example, in Johansson et al., "RNA recognition by the MS2 phage coat protein," Sem Virol., 1997, Vol. 8(3):176-185; Delebecque et al., "Organization of intracellular reactions with rationally designed RNA assemblies," Science, 2011, Vol. 333:470-474; Mali et al., "Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering," Nat. Biotechnol., 2013, Vol. 31:833-838; and Zalatan et al., "Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds," Cell, 2015, Vol. 160:339-350 (each of which is incorporated herein by reference in its entirety). Other systems include the PP7 hairpin, which specifically recruits the PCP protein, and the "com" hairpin, which specifically recruits the Com protein. See Zalatan et al.

[0220] The nucleotide sequence of the MS2 hairpin (i.e., referred to as the "MS2 aptamer") is: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 763).

[0221] The amino acid sequence of MCP or MS2cp is: GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (SEQ ID NO: 764).

[0222] An MS2 hairpin (or "MS2 aptamer") can also be referred to as a type of "RNA effector recruitment domain" (or equivalently, an "RNA-binding protein recruitment domain" or simply "recruitment domain") because it is a physical structure (e.g., a hairpin) that is incorporated onto a PEgRNA or tPERT, which effectively recruits other effector functions (e.g., RNA-binding proteins with various functions, such as DNA polymerases or other DNA-modifying enzymes) to the so-modified PEgRNA or rPERT, thus co-localizing the effector functions in trans with the primed editing machinery. This application is in no way intended to be limited to any particular RNA effector recruitment domain, but can encompass any available such domain, including an MS2 hairpin. Example 19 and Figure 72(b) illustrate the use of a prime editor comprising an MS2 aptamer linked to a DNA synthesis domain (i.e., a tPERT molecule) and an MS2cp protein fused to PE2 to cause colocalization of the prime editor complex (MS2cp-PE2:sgRNA complex) bound to a target DNA site with the DNA synthesis domain of the tPERT molecule.

[0223] napDNAbp As used herein, the term "nucleic acid-programmed DNA-binding protein" or "napDNAbp" (of which Cas9 is an example) refers to a protein that uses RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is linked to at least one guide nucleic acid (e.g., a guide RNA) that localizes the napDNAbp to a DNA sequence that includes a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of the guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to the complementary sequence.

[0224] Without being bound by theory, the binding mechanism of the napDNAbp-guide RNA complex generally involves the napDNAbp forming an R-loop that induces unwinding of the double-stranded DNA target (thereby separating the strands in the region bound by the napDNAbp). The guide RNA protospacer then hybridizes to the "target strand." This displaces the "non-target strand" that is complementary to the target strand, thereby forming the single-stranded region of the R-loop. In some embodiments, the napDNAbp contains one or more nuclease activities, which then cleave the DNA to eliminate various types of lesions. For example, the napDNAbp may contain nuclease activities that cleave the non-target strand at a first location and / or the target strand at a second location. In response to the nuclease activity, the target DNA can be cleaved to form a "double-stranded break" in which both strands are cleaved. In other embodiments, the target DNA can be cleaved at only a single site, i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with various nuclease activities include "Cas9 nickase" ("nCas9") and inactivated Cas9 that does not have any nuclease activity ("inactivated Cas9" or "dCas9"). Exemplary sequences for these and other napDNAbps are provided herein.

[0225] Nickase The term "nickase" refers to Cas9 with one of its two nuclease domains inactivated, allowing the enzyme to cleave only one strand of target DNA.

[0226] Nuclear localization sequence (NLS) The term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that facilitates the import of a protein into a cell nucleus, e.g., by nuclear transport. Nuclear localization sequences are known in the art and would be apparent to one skilled in the art. For example, NLS sequences are described in International PCT Application PCT / EP2000 / 011690 to Plank et al., filed November 23, 2000, and published May 31, 2001 as WO / 2001 / 038547 (the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences). In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 16) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 17).

[0227] nucleic acid molecule The term "nucleic acid," as used herein, refers to a polymer of nucleotides. The polymer may be composed of natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyluridine, C5 propynylcytidine, C5 methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)methylguanine, 4-acetylcytidine, and the like). , 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudouridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N phosphoramidite linkages).

[0228] PEG-RNA As used herein, the terms "prime editing guide RNA" or "PEgRNA" or "extended guide RNA" refer to specialized forms of guide RNA that have been modified to include one or more additional sequences for practicing the prime editing methods and compositions described herein. As described herein, a prime editing guide RNA includes one or more "extended regions" of nucleic acid sequence. The extended region may include, but is not limited to, a single strand of RNA or DNA. Additionally, the extended region may occur at the 3' end of an existing guide RNA. In other configurations, the extended region may occur at the 5' end of an existing guide RNA. In yet other configurations, the extended region may occur at an intramolecular region at the end of an existing guide RNA (e.g., in the gRNA core region that attaches and / or binds to the napDNAbp). The extended region comprises a "DNA synthesis template" encoding a single-stranded DNA (by the polymerase of the prime editor), which molecule is then (a) designed to be homologous to the endogenous target DNA to be edited, and (b) contains at least one desired nucleotide change (e.g., a transition, transversion, deletion, or insertion) to be introduced or incorporated into the endogenous target DNA. The extended region may also contain other functional sequence elements (such as, but not limited to, a "primer binding site" and a "spacer or linker" sequence) or other structural elements (such as, but not limited to, an aptamer, a stem-loop, a hairpin, a toe-loop (e.g., a 3' toe-loop), or an RNA-protein recruitment domain (e.g., an MS2 hairpin)). As used herein, a "primer binding site" includes a sequence that hybridizes to a single-stranded DNA sequence having a 3' end generated from the R-loop nicked DNA.

[0229] In one embodiment, a PEGRNA is represented by Figure 3A, which shows a PEGRNA with a 5' extension arm, a spacer, and a gRNA core. The 5' extension further comprises, in a 5' to 3' direction, a reverse transcriptase template, a primer binding site, and a linker. As shown, the reverse transcriptase template can also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editor described herein is not RT but another type of polymerase.

[0230] In certain other embodiments, the PEGRNA is represented by Figure 3B, which shows a PEGRNA with a 3' extension arm, a spacer, and a gRNA core. The 3' extension further comprises a reverse transcriptase template in the 5' to 3' direction, a primer binding site. As shown, the reverse transcriptase template can also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editor described herein is not RT but another type of polymerase.

[0231] In yet another embodiment, the PEgRNA is represented by Figure 3D, which shows a PEgRNA having, in a 5' to 3' direction, a spacer (1), a gRNA core (2), and an extension arm (3). The extension arm (3) is at the 3' end of the PEgRNA. The extension arm (3) further comprises, in a 5' to 3' direction, a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may also comprise optional modification regions at the 3' and 5' ends, which may be the same or different sequences. Additionally, the 3' end of the PEgRNA may comprise a transcription terminator sequence. These sequence elements of the PEgRNA are further described and defined herein.

[0232] In yet another embodiment, the PEgRNA is represented by Figure 3E, which shows a PEgRNA having, in a 5' to 3' direction, an extension arm (3), a spacer (1), and a gRNA core (2). The extension arm (3) is at the 5' end of the PEgRNA. The extension arm (3) further comprises, in a 3' to 5' direction, a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may also comprise optional modification regions at the 3' and 5' ends, which may be the same or different sequences. The 3' end of the PEgRNA may also comprise a transcription terminator sequence. These sequence elements of the PEgRNA are further described and defined herein.

[0233] PE1 As used herein, "PE1" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and wild-type MMLV RT, having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)]+desired PE gRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 123, which is shown as follows: [ka] [ka]

[0234] PE2 As used herein, "PE2" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and a variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+desired PE gRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 134, which is shown as follows: [ka] [ka]

[0235] PE3 As used herein, "PE3" refers to PE2 plus a second strand nicking guide RNA that complexes with PE2 to introduce a nick on the non-edited DNA strand to direct preferential displacement of the edited strand. PE3b

[0236] As used herein, "PE3b" refers to PE3, but the second-strand nicking guide RNA is designed for temporal control by not introducing a second-strand nick until after the desired edit has been incorporated. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand, not the original allele. Using this strategy, hereafter referred to as PE3b, the mismatch between the protospacer and the non-edited allele should not favor nicking by the sgRNA until after the editing event on the PAM strand has occurred.

[0237] Short PE (PE-short) As used in this application, Short PE " refers to a PE construct fused to a C-terminally truncated reverse transcriptase and has the following amino acid sequence: [ka] [ka]

[0238] Peptide tags The term "peptide tag" refers to a peptide amino acid sequence that is genetically fused to a protein sequence to confer one or more functions onto the protein, thereby facilitating manipulation of the protein for various purposes such as visualization, purification, solubilization, isolation, etc. Peptide tags can include various types of tags classified by purpose or function, which may include "affinity tags" (to facilitate protein purification), "solubilization tags" (to assist in correct folding of the protein), "chromatography tags" (to alter the chromatographic properties of the protein), "epitope tags" (to bind to high-affinity antibodies), and "fluorescent tags" (to facilitate visualization of the protein in cells or in vitro).

[0239] polymerase As used herein, the term "polymerase" refers to an enzyme that synthesizes a nucleotide strand, which can be used in connection with the prime editor system described herein. The polymerase can be a "template-dependent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand based on the order of nucleotide bases in a template strand). The polymerase can also be a "template-independent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand without the requirement of a template strand). Polymerases can also be categorized as "DNA polymerases" or "RNA polymerases." In various embodiments, the prime editor system comprises a DNA polymerase. In various embodiments, the DNA polymerase can be a "DNA-dependent DNA polymerase" (i.e., whereby the template molecule is a strand of DNA). In such cases, the DNA template molecule can be a PEG-RNA, and the extension arm comprises a strand of DNA. In such cases, the PEGRNA may be referred to as a chimeric or hybrid PEGRNA, which comprises an RNA portion (i.e., a guide RNA component including a spacer and a gRNA core) and a DNA portion (i.e., an extension arm). In various other embodiments, the DNA polymerase may be an "RNA-dependent DNA polymerase" (i.e., according to which the template molecule is a strand of RNA). In such cases, the PEGRNA is RNA, i.e., involves RNA extension. The term "polymerase" may also refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Generally, the enzyme will initiate synthesis at the 3' end of a primer annealed to a polynucleotide template sequence (e.g., a primer sequence annealed to the primer binding site of the PEGRNA) and proceed toward the 5' end of the template strand. A "DNA polymerase" catalyzes the polymerization of deoxynucleotides. The term DNA polymerase, as used herein with reference to a DNA polymerase, includes "functional fragments thereof." A "functional fragment thereof" refers to any portion of a wild-type or mutant DNA polymerase encompassing less than the entire amino acid sequence of the polymerase, which portion retains the ability to catalyze the polymerization of polynucleotides under at least one set of conditions.Such a functional fragment can exist as a separate entity, or it can be a component of a larger polypeptide such as a fusion protein.

[0240] Prime Edit As used herein, the term "prime editing" refers to a novel approach to gene editing that uses napDNAbp, a polymerase (e.g., reverse transcriptase), and a specialized guide RNA that contains a DNA synthesis template to encode the desired new genetic information (or delete genetic information) that is then incorporated into the target DNA sequence. Certain embodiments of prime editing are described in Figures 1A-1H and 72(a)-72(c), among other figures.

[0241] Prime editing represents an entirely new platform for genome editing, a versatile and precise genome editing method that directly writes new genetic information at defined DNA sites. It uses a nucleic acid-programmable DNA-binding protein ("napDNAbp") that works in conjunction with a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with napDNAbp). The prime editing system is programmed by a prime editing (PE) guide RNA ("PEgRNA"), which defines the target site and serves as a template for synthesis of the desired edit in the form of a replaced DNA strand by an extension (either DNA or RNA) engineered onto the guide RNA (e.g., at the 5' or 3' end of the guide RNA or internally). The replacement strand, containing the desired edit (e.g., a single nucleobase substitution), shares the same sequence as (or is homologous to) the endogenous strand at the target site to be edited (immediately downstream of the nick site), except that it encompasses the desired edit. By DNA repair and / or replication mechanisms, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a "search and replace" genome editing technology, because the prime editors described herein not only search for and locate the desired target site to be edited, but also simultaneously encode a replacement strand containing the desired edit that is incorporated in place of the endogenous DNA strand at the corresponding target site. The prime editors of the present disclosure relate, in part, to the discovery that the mechanism of prime target-primed reverse transcription (TPRT) or "prime editing" can be exploited or adapted to perform precise CRISPR / Cas-based genome editing with high efficiency and genetic flexibility (e.g., as illustrated in various embodiments in Figures 1A-1F). In nature, TPRT is used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial group II introns. 28,29In this application, the inventors have used a Cas protein-reverse transcriptase fusion or related system to target a specific DNA sequence with a guide RNA, generate a single-stranded nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered reverse transcriptase template incorporated into the guide RNA. However, although the concept begins with a prime editor using reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptase and can encompass the use of virtually any DNA polymerase. Indeed, although this application may refer to a prime editor with a "reverse transcriptase" throughout, it is expressly stated herein that reverse transcriptase is only one type of DNA polymerase that can function in prime editing. Therefore, whenever this specification refers to a "reverse transcriptase," those skilled in the art should understand that any suitable DNA polymerase can be used in place of reverse transcriptase. Thus, in one aspect, a prime editor can include Cas9 (or the equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., a PEgRNA) containing a spacer sequence that anneals to a complementary protospacer on the target DNA. The specialized guide RNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing the desired genetic alteration, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer information from the PEgRNA to the target DNA, the mechanism of prime editing involves nicking the target site on one strand of DNA to expose a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime DNA polymerization of the extension encoding the edit on the PEgRNA directly to the target site. In various embodiments, the extension that provides the template for polymerization of the replacement strand containing the edit can be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editor can be a DNA-dependent DNA polymerase. The newly synthesized strand formed by the prime editor disclosed herein (i.e., the replacement DNA strand containing the desired edit) will be homologous to (i.e., have the same sequence as) the genomic target sequence, except for the inclusion of the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA can also be referred to as a single-stranded DNA flap. It will compete for hybridization with a complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In some embodiments, the system can be combined with the use of an error-prone reverse transcriptase enzyme (e.g., provided as a fusion protein with a Cas9 domain or provided in trans relative to the Cas9 domain). The error-prone reverse transcriptase enzyme can introduce mutations during the synthesis of the single-stranded DNA flap. Therefore, in some embodiments, an error-prone reverse transcriptase can be utilized to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used in the system, the change can be random or non-random. The degradation of the hybridized intermediate (including the single-stranded DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can involve the removal of the resulting displaced flap of the endogenous DNA (for example, by the 5'-end DNA flap endonuclease FEN1), the ligation of the synthesized single-stranded DNA flap to the target DNA, and the assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because template-directed DNA synthesis offers single-nucleotide precision for any nucleotide modification, including insertion and deletion, the scope of this approach is very broad and can foreseeably be used in countless applications in basic research and therapeutics.

[0242] In various embodiments, prime editing operates by contacting a target DNA molecule (in which a nucleotide sequence change is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). Referring to FIG. 1G, the prime editing guide RNA (PEgRNA) contains an extension at the 3' or 5' end of the guide RNA or at an intramolecular position of the guide RNA that encodes the desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) into one of the strands of the DNA at the target locus, thereby creating an available 3' end on one of the strands at the target locus. In some embodiments, the nick is generated on the strand of DNA corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the "non-target strand." However, the nick can be introduced into either strand. That is, the nick can be introduced into the R-loop "target strand" (i.e., the strand that hybridizes to the protospacer of the extended gRNA) or the "non-target strand" (i.e., the strand that forms the single-stranded portion of the R-loop that is complementary to the target strand). In step (c), the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime reverse transcription (i.e., "RT primed target"). In some embodiments, the 3' end DNA strand hybridizes to a specific RT prime sequence in the extended portion of the guide RNA, i.e., the "reverse transcriptase prime sequence" or "primer binding site" on the PEGRNA. In step (d), a reverse transcriptase (or other suitable DNA polymerase) is introduced. This synthesizes a single strand of DNA from the 3' end of the primed site toward the 5' end of the primed editing guide RNA. A DNA polymerase (e.g., reverse transcriptase) can be fused to the napDNAbp or alternatively provided in trans relative to the napDNAbp.This forms a single-stranded DNA flap containing the desired nucleotide change (e.g., a single-base change, insertion, or deletion, or a combination thereof) that is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve degradation of the single-stranded DNA flap, resulting in the desired nucleotide change being incorporated into the target locus. This process can be driven toward desired product formation by removing the corresponding 5' endogenous DNA flap, which forms once the 3' single-stranded DNA flap invades and hybridizes with the endogenous DNA sequence. Without being bound by theory, the cell's endogenous DNA repair and replication processes degrade the mismatched DNA and incorporate the nucleotide change(s) to form the desired altered product. The process can also be driven toward product formation by "second-strand nicking," as illustrated in Figure 1F. This process can introduce at least one or more of the following genetic changes: transversions, transitions, deletions, and insertions.

[0243] The terms "prime editor (PE) system" or "prime editor (PE)" or "PE system" or "PE editing system" refer to compositions involved in the methods of genome editing using prime-target-primed reverse transcription (TPRT) described herein, and include, but are not limited to, napDNAbp, reverse transcriptase, a fusion protein (e.g., comprising napDNAbp and reverse transcriptase), a prime editing guide RNA, and a complex comprising the fusion protein and the prime editing guide RNA, as well as auxiliary elements, such as a second-strand nicking component (e.g., second-strand sgRNA) and a 5' endogenous DNA flap removal endonuclease (e.g., FEN1) to help drive the prime editing process toward edited product formation.

[0244] In the embodiments described thus far, the PEGRNA constitutes a single molecule comprising a guide RNA (which itself comprises a spacer sequence and a gRNA core or backbone) and a 5' or 3' extension arm comprising a primer binding site and a DNA synthesis template (see, e.g., Figure 3D). However, the PEGRNA can also take the form of two individual molecules: a guide RNA and a trans-prime editor RNA template (tPERT), which essentially houses an extension arm (including, inter alia, a primer binding site and a DNA synthesis domain) and an RNA-protein recruitment domain (e.g., an MS2 aptamer or hairpin) on the same molecule. This becomes co-localized or recruited to an engineered prime editor complex comprising a tPERT recruitment protein (e.g., an MS2cp protein that binds to the MS2 aptamer). See Figures 3G and 3H for examples of tPERT that can be used for prime editing.

[0245] Prime editing factor The term "prime editor" refers to a fusion construct described herein comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase, which is capable of performing prime editing on a target nucleotide sequence in the presence of a PEgRNA (or "extended guide RNA"). The term "prime editor" may also refer to a fusion protein, or a fusion protein complexed with a PEgRNA, and / or a fusion protein further complexed with an sgRNA that nicks the second strand. In some embodiments, a prime editor may also refer to a complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a PEgRNA, and a regular guide RNA, capable of directing the nicking step to the second site of the non-edited strand as described herein. In other embodiments, the reverse transcriptase component of a "primer editor" may be provided in trans.

[0246] Primer binding site The term "primer binding site" or "PBS" refers to a nucleotide sequence located on the PEgRNA as a component of the extension arm (typically at the 3' end of the extension arm) that serves to bind to a primer sequence formed after Cas9 nicking of the target sequence by the prime editor. As detailed elsewhere, when the Cas9 nickase component of the prime editor nicks one of the strands of the target DNA sequence, a 3'-terminal ssDNA flap is formed that anneals to the primer binding site on the PEgRNA and serves as a primer sequence to prime reverse transcription. Figures 27 and 28 show embodiments of primer binding sites located on the 3' and 5' extension arms, respectively.

[0247] promoter The term "promoter" is recognized in the art and refers to a nucleic acid molecule having a sequence capable of being recognized by the cellular transcription machinery and initiating transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is only active under certain conditions. For example, a conditional promoter may be active only in the presence of a specific protein that connects proteins associated with the promoter's regulatory elements to the basal transcription machinery, or in the absence of an inhibitory molecule. A subclass of conditionally active promoters is inducible promoters, which require the presence of a small molecule "inducer" for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to identify a variety of such promoters useful in carrying out the present invention. This is not limiting in this regard.

[0248] Protospacer As used herein, the term "protospacer" refers to a sequence (~20 bp) on DNA adjacent to the PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand of it, i.e., the "target strand" versus the "non-target strand" of the target DNA sequence). For Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease from S. pyogenes recognizes the PAM sequence NGG on the non-target strand, which is found directly downstream of the target sequence on genomic DNA. Those skilled in the art will understand that the state of the art literature sometimes refers to the "protospacer" as the ~20 nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a "spacer." Therefore, in some cases, the term "protospacer" as used herein may be used interchangeably with the term "spacer." The context surrounding the appearance of either "protospacer" or "spacer" will help inform the reader as to whether the term refers to a gRNA or a DNA target.

[0249] Protospacer adjacent motif (PAM) As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is the key targeting component of the Cas9 nuclease. Typically, the PAM sequence is located on either strand, downstream in the 5' to 3' direction from the site cut by Cas9. The standard PAM sequence (i.e., the PAM sequence associated with Streptococcus pyogenes Cas9 nuclease or SpCas9) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, can be modified to alter the nuclease's PAM specificity by allowing the nuclease to recognize alternative PAM sequences.

[0250] For example, with reference to the standard SpCas9 amino acid sequence of SEQ ID NO: 18, the PAM sequence can be modified by introducing one or more mutations, including (a) the D1135V, R1335Q, and T1337R "VQR variant," which alters the PAM specificity to NGAN or NGNG; (b) the D1135E, R1335Q, and T1337R "EQR variant," which alters the PAM specificity to NGAG; and (c) the D1135V, G1218R, R1335E, and T1337R "VRER variant," which alters the PAM specificity to NGCG. In addition, the D1135E variant of standard SpCas9 still recognizes NGG, but more selectively than the wild-type SpCas9 protein.

[0251] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. Additionally, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. In yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and are not meant to be limiting. Furthermore, it will be understood that non-SpCas9s bind to a variety of PAM sequences, making them useful when a suitable SpCas9 PAM sequence is not present at the desired target cut site. Furthermore, non-SpCas9s may have other characteristics that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is approximately 1 kilobase smaller than SpCas9. Therefore, it can be packaged into adeno-associated virus (AAV). Further reference may be made to Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5):891-899, which is incorporated herein by reference. Recombinase

[0252] The term "recombinase" as used herein refers to a site-specific enzyme that mediates the recombination of DNA between recombinase recognition sequences, resulting in the excision, incorporation, inversion, or exchange (e.g., translocation) of DNA fragments between the recombinase recognition sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, without limitation, Hin, Gin, Tn3, β-six, CinH, ParA, γδ, Bxb1, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, without limitation, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. The names of serine and tyrosine recombinases are derived from the conserved nucleophilic amino acid residues that the recombinase uses to attack DNA and become covalently linked to the DNA during strand exchange. Recombinases have numerous applications, including the generation of gene knockouts / knockins and gene therapy applications. For example, Brown et al., “Serine recombinases as tools for genome engineering.”Methods.2011;53(4):372-9;Hirano et al., “Site-specific recombinases as tools for heterologous gene integration.”Appl.Microbiol.Biotechnol.2011;92(2):227-39;Chavez and Calos, “Therapeutic applications of the ΦC31 integrase system.”Curr.Gene Ther.2011;11(5):375-81;Turan and Bode,“Site-specific recombinases:from tag-and-target- to tag-and-exchange-based genomic modifications.”FASEB J.2011;25(12):4088-107;Venken and Bellen,“Genome-wide manipulations of Drosophila melanogaster with transposons, Flp recombinase, and ΦC31 integrase.”Methods Mol.Biol.2012;859:203-28;Murphy,“Phage recombinases and their applications.”Adv.Virus Res.2012;83:367-414;Zhang et al.,“Conditional gene manipulation: Cre-ating a new biological era.”J.Zhejiang Univ.Sci.B.2012;13(7):511-24;Karpenshif and Bernstein,“From yeast to mammals:recent advances in genetic control of homologous recombination.”DNA Repair (Amst). 2012;1;11(10):781-8 (the entire contents of each are incorporated herein by reference in their entirety). The recombinases provided herein are not meant to be the only examples of recombinases that may be used in embodiments of the present invention. The methods and compositions of the present invention can be extended by mining databases for new orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity (see, e.g., Groth et al., "Phage integrases: biology and applications." J. Mol. Biol. 2004;335,667-678; Gordley et al., "Synthesis of programmable integrases." Proc. Natl. Acad. Sci. USA.2009;106,5053-5058 (the entire contents of each are incorporated herein by reference in their entirety). Other examples of recombinases useful in the methods and compositions described herein will be known to those of skill in the art, and it is anticipated that any new recombinases discovered or produced will be capable of being used in different embodiments of the present invention. In some embodiments, the catalytic domain of a recombinase is fused to an RNA-programmable nuclease (e.g., dCas9 or a fragment thereof) that inactivates the nuclease, such that the recombinase domain does not contain a nucleic acid binding domain or is incapable of binding to a target nucleic acid (e.g., the recombinase domain is engineered so that it does not have specific DNA-binding activity). Recombinases lacking DNA-binding activity and methods for modifying them are known, see Klippel et al., "Isolation and characterization of unusual gin mutants," EMBO J. 1988;7:3983-3989; Burke et al., "Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation," Mol Microbiol. 2004;51:937-948; Olorunniji et al., "Synapsis and catalysis by activated Tn3 resolvase mutants," Nucleic Acids Res. 2008;36:7181-7191; Rowland et al., "Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome," Mol Microbiol. 2009;74:282-298; Akopian et al., "Chimeric recombinases with designed DNA sequence recognition.”Proc Natl Acad Sci USA.2003;100:8688-8691;Gordley et al.,“Evolution of programmable zinc finger-recombinases with activity in human cells.J Mol Biol.2007;367:802-813;Gordley et al.,“Synthesis of programmable integrases.”Proc Natl Acad Sci USA.2009;106:5053-5058;Arnold et al.,“Mutants of Tn3 resolvase which do not require accessory binding sites for recombination activity.”EMBO J.1999;18:1407-1414;Gaj et al.,“Structure-guided reprogramming of serine recombinase DNA sequence specificity.”Proc Natl Acad Sci USA.2011;108(2):498-503; and Proudfoot et al., “Zinc finger recombinases with adaptable DNA sequence specificity.” PLoS One. 2011;6(4):e19537 (the entire contents of each are incorporated herein by reference in their entirety). For example, serine recombinases of the resolvase-invertase group, such as Tn3 and γδ resolvases and Hin and Gin invertases, have molecular structures with autonomous catalytic and DNA-binding domains (see, e.g., Grindley et al., “Mechanism of site-specific recombination.” Ann Rev Biochem.2006;75:567-605, the entire contents of which are incorporated by reference. Thus, for example, after isolation of "activated" recombinase mutants that do not require any auxiliary factors (e.g., DNA binding activity), the catalytic domains of these recombinases are amenable to recombination with RNA-programmable nucleases (e.g., dCas9 or fragments thereof) that inactivate the nuclease, as described herein (see, e.g., Klippel et al., "Isolation and characterization of unusual gin mutants," EMBO J. 1988;7:3983-3989; Burke et al., "Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation." Mol Microbiol. 2004;51:937-948; Olorunniji et al., "Synapsis and catalysis by activated Tn3 resolvase mutants," Nucleic Acids Res. 2008;36:7181-7191; Rowland et al. al., “Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome.”Mol Microbiol.2009;74:282-298;Akopian et al., “Chimeric recombinases with designed DNA sequence recognition.”Proc Natl Acad Sci USA.2003;100:8688-8691). In addition, many other naturally occurring serine recombinases with N-terminal catalytic domains and C-terminal DNA-binding domains are known (e.g., phiC31 integrase, TnpX transposase, IS607 transposase), and their catalytic domains can be selected to engineer programmable site-specific recombinases as described herein (see, e.g., Smith et al., "Diversity in the serine recombinases." Mol Microbiol. 2002;44:299-307, the entire contents of which are incorporated by reference). Similarly, the core catalytic domains of tyrosine recombinases (e.g., Cre, λ integrase) are known and can be similarly selected to engineer programmable site-specific recombinases as described herein (see, e.g., Guo et al., "Structure of Cre recombinase complexed with DNA in a site-specific recombination synapse." Nature. 1997;389:40-46; Hartung et al., "Cre mutants with altered DNA binding properties." J Biol Chem 1998;273:22884-22891; Shaikh et al., "Chimeras of the Flp and Cre recombinases: Tests of the mode of cleavage by Flp and Cre." J Mol Biol. 2000;302:27-48; Rongrong et al., "Effect of deletion mutations on the recombination activity of Cre recombinase." Acta Biochim Pol.2005;52:541-544;Kilbride et al., “Determinants of product topology in a hybrid Cre-Tn3 resolvase site-specific recombination system.”J Mol Biol.2006;355:185-195;Warren et al.,“A chimeric cre recombinase with regulated directionality.”Proc Natl Acad Sci USA.2008 105:18278-18283;Van Duyne,“Teaching Cre to follow directions.”Proc Natl Acad Sci USA.2009 Jan 6;106(1):4-5;Numrych et al.,“A comparison of the effects of single-base and triple-base changes in the integrase arm-type binding sites on the site-specific recombination of bacteriophage λ.”Nucleic Acids Res.1990;18:3953-3959;Tirumalai et al.,“The recognition of core-type DNA sites by λ integrase.”J Mol Biol. 1998;279:513-527; Aihara et al., "A conformational switch controls the DNA cleavage activity of λ integrase." Mol Cell. 2003;12:187-198; Biswas et al., "A structural basis for allosteric control of DNA recombination by λ integrase." Nature. 2005;435:1059-1066; and Warren et al., "Mutations in the amino-terminal domain of λ-integrase have differential effects on integrative and excisive recombination." Mol Microbiol. 2005;55:1104-1112 (the entire contents of each are incorporated by reference).

[0253] Recombinase recognition sequence As used herein, the term "recombinase recognition sequence" or, equivalently, "RRS" or "recombinase target sequence" refers to a nucleotide sequence target that is recognized by a recombinase and undergoes strand exchange with another DNA molecule that contains an RRS, resulting in excision, incorporation, inversion, or exchange of a DNA fragment between the recombinase recognition sequences.

[0254] Recombination or recombination The terms "recombining" or "recombination" are used in the context of nucleic acid modification (e.g., genome modification) to refer to a process in which two or more nucleic acid molecules or two or more regions of a single nucleic acid molecule are modified by the action of a recombinase protein (e.g., the recombinase fusion proteins of the invention provided herein). Recombination can result in, among other things, an insertion, inversion, excision, or translocation of nucleic acid on or between one or more nucleic acid molecules. Recombinase Recognition Sequences

[0255] reverse transcriptase The term "reverse transcriptase" describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have primarily been used to transcribe mRNA into cDNA, which can then be cloned onto a vector for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand of an RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading, so transcription errors cannot be corrected by the reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity was presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase extensively used in molecular biology is the reverse transcriptase originating from Moloney murine leukemia virus (M-MLV). See, e.g., Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). An M-MLV reverse transcriptase substantially lacking RNase H activity has also been described. See, for example, US Pat. No. 5,244,797.The present invention contemplates the use of any such reverse transcriptase or variant or mutant thereof.

[0256] In addition, the present invention contemplates the use of reverse transcriptases that are prone to errors, or that can be referred to as error-prone reverse transcriptases, or reverse transcriptases that do not support high-fidelity nucleotide incorporation during polymerization.During the synthesis of a single-stranded DNA flap based on a guide RNA and an incorporated RT template, the error-prone reverse transcriptase may introduce one or more nucleotides that are mismatched with the RT template sequence, thereby introducing changes into the nucleotide sequence through the error-prone polymerization of the single-stranded DNA flap.These errors introduced during the synthesis of the single-stranded DNA flap are then incorporated into the double-stranded molecule by hybridization with the corresponding endogenous target strand, removal of the displaced endogenous strand, ligation, and then another round of endogenous DNA repair and / or sequencing processes.

[0257] Reverse transcription As used herein, the term "reverse transcription" refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription may be "error-prone reverse transcription," which refers to the property of certain reverse transcriptases that their DNA polymerization activity is error-prone.

[0258] PACE The term "phage-assisted continuous evolution (PACE)" as used herein refers to continuous evolution using phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published March 11, 2010 as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, published June 28, 2012 as WO 2012 / 088381; U.S. Patent Application No. 9,023,594, issued May 5, 2015; International PCT Application No. PCT / US2015 / 012022, filed January 20, 2015, published September 11, 2015 as WO 2012 / 088381; and International PCT application PCT / US2016 / 027795, filed April 15, 2016, published as WO 2016 / 168631, the entire contents of each of which are incorporated herein by reference.

[0259] phage The term "phage," used interchangeably with the term "bacteriophage" herein, refers to a virus that infects bacterial cells. Typically, phages consist of an outer protein capsid that encapsulates genetic material. The genetic material can be ssRNA, dsRNA, ssDNA, or dsDNA, in either linear or circular form. Phages and phage vectors are well known to those of skill in the art; non-limiting examples of phages useful for practicing the PACE method provided herein include λ (lysogen), T2, T4, T7, T12, R17, M13, MS2, G4, P1, P2, P4, phiX174, N4, Φ6, and Φ29. In one embodiment, the phage utilized in the present invention is M13. Additional suitable phages and host cells will be apparent to those of skill in the art. The present invention is not limited in this respect.Exemplary descriptions of additional suitable phages and host cells can be found in Elizabeth Kutter and Alexander Sulakvelidze: Bacteriophages: Biology and Applications. CRC Press; 1st edition (December 2004), ISBN: 0849313368; Martha R.J. Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 1: Isolation, Characterization, and Interactions (Methods in Molecular Biology), Humana Press; 1st edition (December, 2008), ISBN: 1588296822; Martha R.J. Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 2: Molecular and Applied Aspects (Methods in Molecular Biology), Humana Press; 1st edition (December, 2008), ISBN: 1588296822. 2008), ISBN: 1603275649; all of which are incorporated herein by reference in their entireties for their disclosure of suitable phages and host cells, as well as methods and protocols for the isolation, culture, and manipulation of such phages.

[0260] Proteins, peptides, and polypeptides The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide may refer to an individual protein or a group of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified by the addition of a chemical entity, such as, for example, a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference).

[0261] Protein splicing The term "protein splicing" as used herein refers to the process in which the sequence of an intein (or split-intein, as the case may be) is excised from within an amino acid sequence and the exteins of the remaining fragment of the amino acid sequence are ligated by amide bonds to form a continuous amino acid sequence. The term "trans" protein splicing refers to the specific case in which the inteins are split-inteins and are located on different proteins.

[0262] Second strand nicking The degradation of heteroduplex DNA (i.e., containing one edited and one non-edited strand) formed as a result of prime editing determines the long-term editing outcome. In other words, the goal of prime editing is to degrade the heteroduplex DNA (the edited strand paired with the endogenous non-edited strand) formed as a PE intermediate by permanently incorporating the edited strand onto the complementary endogenous strand. To help drive the degradation of heteroduplex DNA in favor of the permanent incorporation of the edited strand into the DNA molecule, a "second-strand nicking" approach can be used herein. The term "second-strand nicking" as used herein refers to the introduction of a second nick, preferably on the non-edited strand, downstream of the first nick (i.e., the original nick site that provides a free 3' end for use in priming reverse transcriptase on the extended portion of the guide RNA). In some embodiments, the first nick and the second nick are on opposite strands. In other embodiments, the first nick and the second nick are on opposite strands. In yet another embodiment, the first nick is on the non-target strand (i.e., the strand that forms the single-stranded portion of the R-loop) and the second nick is on the target strand. In still other embodiments, the first nick is on the edited strand and the second nick is on the non-edited strand. The second nick can be located at least 5 nucleotides downstream from the first nick, or at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 nucleotides or more downstream from the first nick.The second nick, in some embodiments, can be introduced on the non-edited strand between about 5-150 nucleotides, or between about 5-140 nucleotides, or between about 5-130 nucleotides, or between about 5-120 nucleotides, or between about 5-110 nucleotides, or between about 5-100 nucleotides, or between about 5-90 nucleotides, or between about 5-80 nucleotides, or between about 5-70 nucleotides, or between about 5-60 nucleotides, or between about 5-50 nucleotides, or between about 5-40 nucleotides, or between about 5-30 nucleotides, or between about 5-20 nucleotides, or between about 5-10 nucleotides, from the site of the non-edited strand. In one embodiment, the second nick is introduced between 14-116 nucleotides from the non-edited strand. Without being bound by theory, the second nick directs the cell's endogenous DNA repair and replication processes toward replacing or editing the non-edited strand, thereby permanently incorporating the edited sequence on both strands and dissolving the heteroduplex formed as a result of PE. In some embodiments, the edited strand is the non-target strand and the non-edited strand is the target strand, while in other embodiments, the edited strand is the target strand and the non-edited strand is the non-target strand.

[0263] Sense strand In genetics, a "sense" strand is a 5' to 3' segment of double-stranded DNA that is complementary to the 3' to 5' antisense or template strand of DNA. In the case of a DNA segment encoding a protein, the sense strand is the strand of DNA with the same sequence as the mRNA that takes the antisense strand as its template during transcription and ultimately (typically, but not always) is translated into the protein. Thus, the antisense strand carries the RNA that is subsequently translated into protein, while the sense strand possesses nearly identical makeup to the mRNA. Note that for each segment of dsDNA, there will likely be two pairs: sense and antisense (because sense and antisense are relative perspectives), depending on which direction one is read from. The gene product or mRNA ultimately represents either strand of a segment of dsDNA, which may be designated as sense or antisense.

[0264] In the context of PEGRNA, the first step is the synthesis of a single-stranded complementary DNA (i.e., the 3' ssDNA flap to be incorporated) oriented in the 5' to 3' direction, but the synthesized strand is off-templated for the PEGRNA extension arm. Whether the 3' ssDNA flap should be considered the sense strand or the antisense strand depends on the direction of transcription, where it is fully accepted that both strands of DNA can serve as templates for transcription (but not simultaneously). Thus, in some embodiments, the 3' ssDNA flap (extending generally in the 5' to 3' direction) will serve as the sense strand, since it is the coding strand. In other embodiments, the 3' ssDNA flap (extending generally in the 5' to 3' direction) will serve as the antisense strand and thus will be a transcription template.

[0265] Spacer sequence As used herein, the term "spacer sequence" in relation to a guide RNA or PEgRNA refers to a portion of the guide RNA or PEgRNA of approximately 20 nucleotides that contains a nucleotide sequence complementary to a protospacer sequence in the target DNA sequence. The spacer sequence anneals to the protospacer sequence, forming an ssRNA / ssDNA hybrid structure at the target site and a corresponding R-loop ssDNA structure in the endogenous DNA strand complementary to the protospacer sequence.

[0266] subject The term "subject," as used herein, refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human animal of the order Primates. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cattle, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically modified, e.g., a genetically modified non-human subject. The subject may be of either sex and at any stage of development.

[0267] Split intein Inteins are most often found as a continuous domain, but some naturally occur in split forms: in this case, the two fragments are expressed as separate polypeptides and must associate before splicing can occur, a process known as protein trans-splicing.

[0268] An exemplary split intein is the Ssp DnaE intein, which contains two subunits, DnaE-N and DnaE-C. The two distinct subunits are encoded by separate genes, dnaE-n and dnaE-c, which encode the DnaE-N and DnaE-C subunits, respectively. DnaE is a naturally occurring split intein in Synechocytis sp. PCC6803 that can direct the trans-splicing of two separate proteins, each containing a fusion with either DnaE-N or DnaE-C.

[0269] Additional naturally occurring or engineered split intein sequences are known or can be created from the whole intein sequences described herein or available in the art. Examples of split intein sequences can be found in Stevens et al., "A promiscuous split intein with expanded protein engineering applications," PNAS, 2017, Vol. 114:8538-8543; Iwai et al., "Highly efficient protein trans-splicing by a naturally split DnaE intein from Nostoc punctiforme," FEBS Lett, 580:1853-1858, each of which is incorporated herein by reference. Additional split intein sequences can be found, for example, in WO2013 / 045632, WO2014 / 055782, WO2016 / 069774, and EP2877490, the contents of each of which are incorporated herein by reference.

[0270] In addition, protein splicing in trans has been described in vivo and in vitro (Shingledecker, et al., Gene 207:187 (1998); Southworth, et al., EMBO J. 17:918 (1998); Mills, et al., Proc. Natl. Acad. Sci. USA 95:3543-3548 (1998); Lew, et al., J. Biol. Chem. 273:15887-15890 (1998); Wu, et al., Biochim. Biophys. Acta 35732:1 (1998b); Yamazaki, et al., J. Am. Chem. Soc. 120:5591 (1998); Evans, et al. al., J. Biol. Chem. 275:9091 (2000); Otomo, et al., Biochemistry 38:16040-16044 (1999); Otomo, et al., J. Biol. Mol. NMR 14:105-114 (1999); Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999)), providing the opportunity to express a protein as two inactive fragments that subsequently undergo ligation to form a functional product, as shown, for example, in Figures 66 and 67 for the formation of an entire PE fusion protein from two separately expressed halves.

[0271] target site The term "target site" refers to a sequence on a nucleic acid molecule that is edited by a prime editor (PE) disclosed herein. Additionally, a target site refers to a sequence on a nucleic acid molecule to which a complex of a prime editor (PE) and a gRNA binds.

[0272] tPERT See definition for "trans prime editing factor 'RNA type (tPERT)."

[0273] Temporary second strand nicking As used herein, the term "transient second-strand nicking" refers to a variant of second-strand nicking in which the incorporation of a second nick in the unedited strand occurs only after the desired edit has been incorporated into the edited strand. This avoids simultaneous nicking on both strands, which can lead to double-stranded DNA breaks. The second-strand nicking guide RNA is designed for temporal control so that the second-strand nick is not introduced until after the desired edit has been incorporated. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand but not the original allele. Using this strategy, the mismatch between the protospacer and the unedited allele should prevent nicking by the sgRNA until after the editing event on the PAM strand has occurred.

[0274] Trans PrimeEdit As used herein, the term "trans prime editing" refers to a modified form of prime editing that utilizes a split PEG RNA, i.e., in which the PEG RNA is separated into two separate molecules: an sgRNA and a trans prime editing RNA template (tPERT). The sgRNA serves to target the prime editor (or more generally, the napDNAbp component of the prime editor) to a desired genomic target site, while once tPERT is recruited in trans to the prime editor through the interaction of binding domains located on the prime editor and tPERT, tPERT is used by a polymerase (e.g., reverse transcriptase) to write new DNA sequences into the target locus. In one embodiment, the binding domain may include an RNA-protein recruitment moiety, such as the MS2 aptamer located on tPERT and the MS2cp protein fused to the prime editor. An advantage of trans prime editing is that by separating the DNA synthesis template from the guide RNA, longer templates can potentially be used.

[0275] An embodiment of trans-prime editing is shown in Figures 3G and 3H. Figure 3G shows, on the left, the composition of a trans-prime editor complex ("RP-PE:gRNA complex"), which includes napDNAbp fused to each of a polymerase (e.g., reverse transcriptase) and an rPERT recruitment protein (e.g., MS2sc), complexed with a guide RNA. Figure 3G further shows a separate tPERT molecule, including a DNA synthesis template and an extended arm feature of PE-gRNA encompassing a primer-binding sequence. The tPERT molecule also includes an RNA-protein recruitment domain (which, in this case, is a stem-loop structure and may be, for example, an MS2 aptamer). As depicted in the process in Figure 3H, the RP-PE:gRNA complex binds to and nicks the target DNA sequence. The recruitment protein (RP) then recruits and colocalizes tPERT to the primed editor complex bound to the DNA target site, allowing the primer-binding site to bind to the primer sequence on the nicked strand, and subsequently allowing a polymerase (e.g., RT) to synthesize a single strand of DNA onto the DNA synthesis template, up to 5' of tPERT.

[0276] Although tPERT is shown in Figures 3G and 3H as containing a PBS and a DNA synthesis template on the 5' end of the RNA-protein recruitment domain, other configurations of tPERT may also be designed with the PBS and DNA synthesis template located on the 3' end of the RNA-protein recruitment domain. However, tPERT with a 5' extension has the advantage that synthesis of a single strand of DNA will naturally terminate at the 5' end of tPERT, thus eliminating the risk of using any part of the RNA-protein recruitment domain as a template during the DNA synthesis step of prime editing.

[0277] Trans-primed editor RNA template (tPERT) As used herein, "trans prime editor RNA template (tPERT)" refers to the components used in trans prime editing. This modified version of prime editing operates by separating the PE-gRNA into two separate molecules: a guide RNA and a tPERT molecule. The tPERT molecule is programmed to colocalize with the prime editor complex at the target DNA site and bring the primer binding site and DNA synthesis template to the prime editor in trans. For example, see Figure 3G for an embodiment of a trans prime editor (tPE) showing a two-component system including (1) an RP-PE:gRNA complex and (2) tPERT, which contains a primer binding site and a DNA synthesis template linked to an RNA-protein recruitment domain. Here, the RP (recruitment protein) component of the RP-PE:gRNA complex recruits tPERT to the target site to be edited, thereby linking the PE and DNA synthesis template to the prime editor in trans. In other words, tPERT is modified to contain (all or part of) the extension arm of a PEG RNA, which encompasses a primer binding site and a DNA synthesis template.

[0278] Transitions As used herein, "transition" refers to the interconversion of purine nucleobases (A⇔G) or pyrimidine nucleobases (C⇔T). This class of interconversion involves nucleobases of similar shape. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both a transition and a transversion in the same target DNA molecule. These changes involve A⇔G, G⇔A, C⇔T, or T⇔C. In the context of double-stranded DNA with Watson-Crick paired nucleobases, a transversion refers to the following base pair exchanges: A:T⇔G:C, G:G⇔A:T, C:G⇔T:A, or T:A⇔C:G. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.

[0279] Transversion As used herein, "transversion" refers to the interconversion of a purine nucleobase to a pyrimidine nucleobase, or vice versa, and thus involves the interconversion of dissimilar nucleobases. These changes involve T⇔A, T⇔G, C⇔G, C⇔A, A⇔T, A⇔C, G⇔C, and G⇔T. In the context of double-stranded DNA with Watson-Crick paired nucleobases, transversion refers to the following base pair exchanges: T:A⇔A:T, T:A⇔G:C, C:G⇔G:C, C:G⇔A:T, A:T⇔T:A, A:T⇔C:G, G:C⇔C:G, and G:C⇔T:A. The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.

[0280] treatment The terms "treatment," "treat," and "treating" refer to a clinical intervention aimed at halting, alleviating, delaying the onset of, or inhibiting the progression of, a disease or disorder as described herein, or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating" refer to a clinical intervention aimed at halting, alleviating, delaying the onset of, or inhibiting the progression of, a disease or disorder as described herein, or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.

[0281] Trinucleotide repeat disorders As used herein, "trinucleotide repeat disorders" (or alternatively, "expansion repeat disorders" or "repeat expansion disorders") refer to a group of genetic disorders caused by a type of mutation (a trinucleotide repeat in a gene or intron) known as a "trinucleotide repeat expansion." Trinucleotide repeats were once thought to be commonplace iterations in the genome, but in the 1990s these disorders were classified. These seemingly "benign" stretches of DNA can sometimes expand and cause disease. Several defining features are common to disorders caused by trinucleotide repeat expansions. First, the mutant repeats exhibit instability in both somatic cells and the germline, and more frequently they expand rather than contract with successive transmission. Second, an earlier age of onset and increased severity (prognosis) of the phenotype in subsequent generations generally correlate with larger repeat length. Finally, the parental origin of a disease allele can often affect the expectation that paternal inheritance carries a greater risk of spread for many of these disorders.

[0282] Triplet expansions are thought to be caused by slippage during DNA replication. Due to the repetitive nature of DNA sequences in these regions, "loop-out" structures can also form during DNA replication, maintaining complementary base pairing between the parent and daughter strands being synthesized. If a loop-out structure is formed from a sequence on the daughter strand, this will result in an increase in the number of repeats. However, if a loop-out structure is formed on the parent strand, a decrease in the number of repeats occurs. These repeat expansions appear to be more common than reductions. In general, the larger the expansions, the more likely they are to cause disease or increase disease severity. This characteristic leads to the predictive features seen in trinucleotide repeat disorders. Prediction explains the tendency for an earlier age of onset and increasing symptom severity through successive generations of affected families due to these repeat expansions.

[0283] Nucleotide repeat disorders may encompass disorders in which the triplet repeat occurs in non-coding regions (i.e., non-coding trinucleotide repeat disorders) or disorders in which the triplet repeat occurs in coding regions.

[0284] The prime editor (PE) system described herein may be used to treat nucleotide repeat disorders, which may include Fragile X Syndrome (FRAXA), Fragile XE MR (FRAXE), Friedreich's Ataxia (FRDA), Myotonic Dystrophy (DM), Spinocerebellar Ataxia Type 8 (SCA8), and Spinocerebellar Ataxia Type 12 (SCA12), among others.

[0285] Upstream As used herein, the terms "upstream" and "downstream" are relative terms that define the linear positions of at least two elements located in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in a 5' to 3' direction. In particular, if a first element is located anywhere 5' relative to the second element, the first element is upstream of the second element in the nucleic acid molecule. For example, if a SNP is located 5' to the nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, if a first element is located anywhere 3' relative to the second element, the first element is downstream of the second element in the nucleic acid molecule. For example, if a SNP is located 3' to the nick site, the SNP is downstream of the Cas9-induced nick site. The nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. Unless it is necessary to select which strand of a double-stranded molecule is considered, the terms upstream and downstream refer only to a single strand of a nucleic acid molecule, and the analysis is the same for single-stranded nucleic acid molecules and double-stranded molecules. Often, the strand of double-stranded DNA that can be used to determine the relative positions of at least two elements is the "sense" strand or "coding" strand. In genetics, a "sense" strand is a segment within double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of DNA that extends from 3' to 5'. Thus, for example, if a SNP nucleobase is 3' to the promoter on the sense or coding strand, the SNP nucleobase is "downstream" of the promoter sequence in genomic DNA (which is double-stranded).

[0286] variant As used herein, the term "variant" should be interpreted as meaning something that exhibits qualities with a pattern that deviates from those found in nature; for example, a variant Cas9 is a Cas9 that contains one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. The term "variant" encompasses homologous proteins that have at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% percent identity with a reference sequence and have the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses mutants, truncations, or domains of the reference sequence that exhibit the same or substantially the same functional activity(ies) as the reference sequence.

[0287] vector The term "vector," as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that can enter a host cell and mutate or replicate within the host cell (and then transfer the replicative form of the vector into another host cell). Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the present disclosure.

[0288] Wild type As used herein, the term "wild-type" is a term of art understood by those skilled in the art and means the typical form of an organism, strain, gene, or characteristic as occurring in nature, as distinguished from mutant or variant forms.

[0289] 5' endogenous DNA flap As used herein, the term "endogenous 5' DNA flap" refers to the strand of DNA located immediately downstream of the nick site induced by PE on the target DNA. Nicking of the target DNA strand by PE exposes a 3' hydroxyl group upstream of the nick site and a 5' hydroxyl group downstream of the nick site. The endogenous strand terminating in the 3' hydroxyl group is used to prime the DNA polymerase of a primed editor (e.g., the DNA polymerase is a reverse transcriptase). The endogenous strand downstream of the nick site and beginning with the exposed 5' hydroxyl group is referred to as the "5' endogenous DNA flap" and is ultimately removed and replaced by a newly synthesized replacement strand (i.e., a "3' replacement DNA flap") encoded by extension of the PE gRNA.

[0290] 5' endogenous DNA flap removal As used herein, the term "5' endogenous DNA flap removal" or "5' flap removal" refers to the removal of a 5' endogenous DNA flap formed when a single-stranded DNA flap synthesized by RT competitively invades and hybridizes with endogenous DNA, displacing the endogenous strand in the process. Removal of this displaced endogenous strand can drive a reaction to form a desired product containing the desired nucleotide change. The cell's own DNA repair enzymes may catalyze the removal or excision of the 5' endogenous flap (e.g., flap endonucleases such as EXO1 or FEN1). Host cells may also be transformed to express one or more enzymes that catalyze the removal of the 5' endogenous flap, thereby driving the process to form the product (e.g., flap endonucleases). Flap endonucleases are known in the art and can be found in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10):4507-4519 and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2):198-211, each of which is incorporated herein by reference.

[0291] 3' replacement DNA flap As used herein, the term "3' replacement DNA flap," or simply "replacement DNA flap," refers to the strand of DNA synthesized by a prime editor and encoded by the extension arm of the prime editor PEgRNA. More specifically, the 3' replacement DNA flap is encoded by the polymerase template of the PEgRNA. The 3' replacement DNA flap contains the same sequence as the 5' endogenous DNA flap, except that it also contains an edited sequence (e.g., a single nucleotide change). The 3' replacement DNA flap anneals to the target DNA, replacing or displacing the 5' endogenous DNA flap (which can be excised, for example, by a 5' flap endonuclease such as FEN1 or EXO1), and is then ligated to join the 3' end of the 3' replacement DNA flap to the exposed 5' hydroxyl end of the endogenous DNA (exposed after excision of the 5' endogenous DNA flap), thereby reforming a phosphodiester bond to incorporate the 3' replacement DNA flap and form a heteroduplex DNA containing one edited strand and one unedited strand. The DNA repair process degrades the heteroduplex and copies the information in the edited strand to the complementary strand, thereby permanently incorporating the edit into the DNA. This degradation process can be further driven to completion by nicking the unedited strand, i.e., via "second-strand nicking" as described herein. DETAILED DESCRIPTION OF THE INVENTION

[0292] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS The adoption of the clustered regularly interspaced short palindromic repeats (CRISPR) system for genome editing has revolutionized the life sciences. 1~3While gene disruption using CRISPR is now routine, the precise integration of single-nucleotide edits remains a major challenge, despite being necessary to study or correct numerous disease-causing mutations. Homologous recombination repair (HDR) can achieve such edits but suffers from low efficiency (often <5%), the requirement for a donor DNA repair template, and the deleterious effects of double-strand DNA break (DSB) formation. Recently, Professor David Liu's laboratory developed base editing, which achieves efficient single-nucleotide edits without DSBs. Base editors (BEs) combine the CRISPR system with base-modifying deaminase enzymes to convert targeted C·G or A·T base pairs to A·T or G·C, respectively. 4~6 Although already widely used by researchers worldwide, current BEs can only correct four of 12 possible base pair changes and are unable to correct small insertions or deletions. Furthermore, the target range of base editing is limited by the editing of non-targeted C or A bases adjacent to the target base ("bystander editing") and by the requirement that a PAM sequence be present 15 ± 2 bp from the target base. Therefore, overcoming these drawbacks would significantly expand the basic research and therapeutic applications of genome editing.

[0293] This disclosure proposes a new high-precision editing approach that offers many of the benefits of base editing—namely, avoidance of double-strand breaks and a donor DNA repair template—while overcoming its major drawbacks. The proposed approach described herein achieves direct integration of an edited DNA strand at a target genomic site using target-primed reverse transcription (TPRT). In the design discussed herein, a CRISPR guide RNA (gRNA) will be modified to carry a reverse transcriptase (RT) template sequence that encodes a single-stranded DNA containing the desired nucleotide change. The target site DNA nicked by the CRISPR nuclease (Cas9) will serve as a primer for reverse transcription of the template sequence on the modified gRNA, allowing for the direct integration of any desired nucleotide edits.

[0294] Consequently, the present invention relates, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) or "prime editing" can be exploited or employed to perform highly efficient and genetically flexible, high-precision CRISPR / Cas-based genome editing (e.g., as depicted in various embodiments in Figures 1A-1F). The inventors herein propose using a Cas9 protein-reverse transcriptase fusion to target a specific DNA sequence with a modified guide RNA (an "extended guide RNA"), generate a single-stranded nick at the target site, and use the nicked DNA as a primer for a modified reverse transcriptase template incorporated into the extended gRNA. The newly synthesized strand will be homologous to the genomic target sequence except for the inclusion of the desired nucleotide change (e.g., a single-nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized strand of DNA, sometimes referred to as a single-stranded DNA flap, competes with the complementary homologous endogenous DNA strand for hybridization, thereby displacing the corresponding endogenous strand. Degradation of this hybridized intermediate can involve the resulting removal of the displaced flap from the endogenous DNA (e.g., by 5'-end DNA flap endonuclease, FEN1), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis offers high precision down to a single nucleotide, the breadth of this approach is extremely broad, and it is foreseeable that it could be used for countless applications in basic science and therapeutics.

[0295] [1] napDNAbp The prime and trans-prime editors described herein can include nucleic acid programmable DNA binding proteins (napDNAbp).

[0296] In one aspect, the napDNAbp can be associated or complexed with at least one guide nucleic acid (e.g., guide RNA or PEGRNA). This allows the napDNAbp to localize to a DNA sequence containing a DNA strand (i.e., target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., the spacer of the guide RNA that anneals to the protospacer of the DNA target). In other words, the guide nucleic acid "programs" the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to the complementary sequence of the protospacer on the DNA.

[0297] Any suitable napDNAbp can be used in the prime editing elements described herein. In various embodiments, the napDNAbp can be any Class 2 CRISPR-Cas system, including any Type II, Type V, or Type VI CRISPR-Cas enzyme. The rapid development of CRISPR-Cas as a tool for genome editing has led to the constant development of nomenclature systems used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologues. This application refers to CRISPR-Cas enzymes by nomenclature systems that may be old and / or new. Those skilled in the art will be able to identify the specific CRISPR-Cas enzymes referenced in this application based on the nomenclature system used, regardless of whether it is an old (i.e., "legacy") or new nomenclature system. The CRISPR-Cas nomenclature system is extensively described in Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?," The CRISPR Journal, Vol. 1, No. 5, 2018, the entire contents of which are incorporated herein by reference. The specific CRISPR-Cas nomenclature system used in any given instance in this application is in no way limiting, and one of skill in the art would be able to identify which CRISPR-Cas enzyme is being referenced.

[0298] For example, the following Type II, Type V, and Type VI Class 2 CRISPR-Cas enzymes have the following art-recognized old (i.e., legacy) and new names: Each of these enzymes and / or their variants can be used in the prime editors described herein: [Table 3]

[0299] Without being bound by theory, the mechanism of action of certain napDNAbps contemplated herein involves the formation of an R-loop, which induces the napDNAbp to unwind the double-stranded DNA target, thereby separating the strands of the region bound by the napDNAbp. The guide RNA spacer then hybridizes to the "target strand" at the protospacer sequence. This displaces the "non-target strand," which is complementary to the target strand, forming the single-stranded region of the R-loop. In some embodiments, the napDNAbp contains one or more nuclease activities, which then cut the DNA, leaving various types of scars. For example, the napDNAbp can contain nuclease activities that cut the non-target strand at a first position and / or cut the target strand at a second position. Depending on the nuclease activity, the target DNA can be cut to form a "double-stranded break," whereby both strands are cut. In other embodiments, the target DNA can be cut only at a single site; i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with different nuclease activities include "Cas9 nickase" ("nCas9") and inactivated Cas9 with no nuclease activity ("inactivated Cas9" or "dCas9").

[0300] The description below of various napDNAbps that can be used in conjunction with the prime editors of the present disclosure is not meant to be limiting in any way. Prime editors can include standard SpCas9 or any orthologous Cas9 protein or any variant Cas9 protein, whether known or created or evolved by directed evolution or other mutational processes, including any naturally occurring variant, mutant, or otherwise engineered version of Cas9. In various embodiments, the Cas9 or Cas9 variant has nickase activity, i.e., cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has inactive nuclease activity, i.e., is an "inactive" Cas9 protein. Other variant Cas9 proteins that can be used have a smaller molecular weight than standard SpCas9 (e.g., for easier delivery) or have an altered or rearranged primary amino acid structure (e.g., a circular permutation format).

[0301] Prime editors described herein can also include Cas9 equivalents, including Cas12a (Cpf1) and Cas12b1 proteins, which are the result of convergent evolution. The napDNAbp (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) used herein can also contain various modifications that alter / enhance their PAM specificity. Finally, the present application contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a reference SpCas9 standard sequence or a reference Cas9 equivalent (e.g., Cas12a (Cpf1)).

[0302] napDNAbp may be a CRISPR (clustered regularly interspaced short palindromic repeats)-associated nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to an ancestral mobile element, and a target invading nucleic acid. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, correct processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endolytically cut, and then 3'-5' exolytically trimmed. In nature, DNA binding and cleavage typically require both proteins and RNAs. However, single-guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. For example, see Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference.

[0303] In some embodiments, the napDNAbp directs cleavage of one or both strands at the location of the target sequence, e.g., on the target sequence and / or on the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a napDNAbp that is mutated relative to the corresponding wild-type enzyme such that the mutated napDNAbp lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. For example, an aspartate-to-alanine substitution (D10A) on the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaving a single strand). Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A, with reference to the equivalent amino acid positions in the standard SpCas9 sequence or other Cas9 variants or Cas9 equivalents.

[0304] The term "Cas protein" as used herein refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of the essential basic functions required for the disclosed methods, namely, (i) possession of nucleic acid-programmed binding of the Cas protein to target DNA and (ii) the ability to nick the target DNA sequence on one strand. Cas proteins contemplated herein include CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)), homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and may include Cas9 equivalents from any Class 2 CRISPR system (e.g., Type II, V, VI), including Cas12a (Cpf1), Cas12e (CasX), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9, and the like. Cas13a (C2c2), Cas13d, Cas13c (C2c7), Cas13b (C2c6), and Cas13b. Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299) and Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?," The CRISPR Journal, Vol. 1, No. 5, 2018, the contents of which are incorporated herein by reference.

[0305] The term "Cas9" or "Cas9 nuclease" or "Cas9 portion" or "Cas9 domain" encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not meant to be particularly limiting and may be referred to as "Cas9 or equivalent." Exemplary Cas9 proteins are further described herein and / or in the art and are incorporated herein by reference. This disclosure is not limited with respect to the particular Cas9 used in the prime editor (PE) of the present invention.

[0306] The Cas9 nuclease sequences and structures described in this application are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP,Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012) (the entire contents of each of which are incorporated herein by reference).

[0307] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not meant to be limiting. Primer editors of the present disclosure may use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent.

[0308] A. Standard wild-type SpCas9 In one embodiment, the primer editor constructs described herein can include the "standard SpCas9" nuclease from S. pyogenes. This is widely used as a genome engineering tool and is classified as a type II subgroup of enzymes in Class 2 CRISPR-Cas systems. The Cas9 protein is a large, multidomain protein containing two distinct nuclease domains. Point mutations can be introduced into Cas9 to abolish one or both nuclease activities, resulting in nickase Cas9 (nCas9) or inactive Cas9 (dCas9), respectively, which still retain their ability to bind DNA in a manner programmed by an sgRNA. In principle, when fused to another protein or domain, Cas9 or its variants (e.g., nCas9) can target the protein to virtually any DNA sequence simply by coexpression with the appropriate sgRNA. As used herein, the standard SpCas9 protein refers to the wild-type protein from Streptococcus pyogenes, which has the following amino acid sequence: [Table 4-1] [Table 4-2] [Table 4-3]

[0309] Prime editors described herein can include standard SpCas9 or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with the wild-type Cas9 sequence provided above. These variants can include SpCas9 variants containing one or more mutations, including any of the known mutations reported in SwissProt Accession No. Q99ZW2 (SEQ ID NO: 18) entry. These include: [Table 5] Includes.

[0310] Other wild-type SpCas9 sequences that can be used in the present disclosure include: [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7] [Table 6-8] Includes.

[0311] The prime editors described herein can include any of the above SpCas9 sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0312] B. Wild-type Cas9 ortholog In other embodiments, the Cas9 protein can be a wild-type Cas9 ortholog from another bacterial species that differs from the standard Cas9 from S. pyogenes. For example, the following Cas9 orthologs can be used in conjunction with the prime editor constructs described herein: In addition, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the orthologs below can also be used in the present prime editors. [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5] [Table 7-6] [Table 7-7]

[0313] The prime editors described herein can include any of the above Cas9 orthologous sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0314] The napDNAbp can include any suitable homolog and / or ortholog of a naturally occurring enzyme, such as Cas9. Cas9 homologs and / or orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Preferably, the Cas moiety is configured as a nickase (e.g., mutated, engineered, or otherwise naturally occurring), i.e., capable of cleaving only one strand of the target. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure. Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA-cleavage domain; i.e., Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 3. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any one of the Cas9 orthologs in the table above.

[0315] C. Inactive Cas9 variants In certain embodiments, the prime editors described herein can include an inactive Cas9, e.g., an inactive SpCas9, which lacks nuclease activity due to one or more mutations that inactivate both nuclease domains of Cas9, namely, the RuvC domain (which cleaves the non-protospacer DNA strand) and the HNH domain (which cleaves the protospacer DNA strand). Nuclease inactivation can be due to one or more mutations that result in one or more substitutions and / or deletions in the amino acid sequence of the encoded protein or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0316] As used herein, the term "dCas9" refers to nuclease-inactive or nuclease-dead Cas9 or a functional fragment thereof, and encompasses any naturally occurring dCas9 from any organism, any naturally occurring dCas9 equivalent or a functional fragment thereof, any dCas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered dCas9. The term dCas9 is not meant to be particularly limiting and may be referred to as "dCas9 or equivalent." Exemplary dCas9 proteins and methods for making dCas9 proteins are further described herein and / or in the art and are incorporated herein by reference.

[0317] In other embodiments, dCas9 corresponds to or includes a Cas9 amino acid sequence with one or more mutations that partially or completely inactivate Cas9 nuclease activity. In other embodiments, Cas9 variants with mutations other than D10A and H840A are provided, which can result in complete or partial inactivation of endogenous Cas9 nuclease activity (e.g., nCas9 or dCas9, respectively). Such mutations include, by way of example, other amino acid substitutions at D10 and H820, or other substitutions on the nuclease domains of Cas9 (e.g., substitutions on the HNH nuclease subdomain and / or RuvC1 subdomain), with reference to a wild-type sequence, such as Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1). In some embodiments, variants or homologs of Cas9 (e.g., variants of Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 20)) are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to NCBI Reference Sequence: NC_017053.1. In some embodiments, variants of dCas9 (e.g., variants of NCBI reference sequence: NC_017053.1 (SEQ ID NO: 20)) are provided having an amino acid sequence that is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more, shorter or longer than NC_017053.1 (SEQ ID NO: 20).

[0318] In one embodiment, the inactive Cas9 can be based on the standard SpCas9 sequence of Q99ZW2 and can have the following sequence, including D10X and H810X substitutions (underlined and bold), where X can be any amino acid, or the variant can be a variant of SEQ ID NO: 40 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0319] In one embodiment, the inactive Cas9 can be based on the standard SpCas9 sequence of Q99ZW2 and can have the following sequence, including the D10A and H810A substitutions (underlined and bold), or can be a variant of SEQ ID NO: 41 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. [Table 8-1] [Table 8-2]

[0320] D. Cas9 nickase variants In one embodiment, the prime editor described herein comprises a Cas9 nickase. The term "Cas9 nickase" or "nCas9" refers to a variant of Cas9 that can introduce single-strand breaks into double-stranded DNA molecular targets. In some embodiments, the Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., standard SpCas9) comprises two distinct nuclease domains: the RuvC domain (which cleaves the non-protospacer DNA strand) and the HNH domain (which cleaves the protospacer DNA strand). In one embodiment, the Cas9 nickase comprises a mutation in the RuvC domain that inactivates RuvC nuclease activity. For example, mutation of aspartic acid (D)10, histidine (H)983, aspartic acid (D)986, or glutamic acid (E)762 has been reported as a loss-of-function mutation in the RuvC nuclease domain and results in a functional Cas9 nickase (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, which is incorporated herein by reference). Thus, nickase mutations in the RuvC domain can include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In some embodiments, the nickase can be D10A, or (of) H983A, or D986A, or E762A, or a combination thereof.

[0321] In various embodiments, the Cas9 nickase can have a mutation in the RuvC nuclease domain and can have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5] [Table 9-6]

[0322] In another embodiment, the Cas9 nickase comprises a mutation in the HNH domain that inactivates HNH nuclease activity. For example, mutation of histidine (H)840 or asparagine (R)863 has been reported as a loss-of-function mutation in the HNH nuclease domain and produces a functional Cas9 nickase (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, which is incorporated herein by reference). Thus, nickase mutations in the HNH domain can include H840X and R863X, where X is any amino acid other than the wild-type amino acid. In some embodiments, the nickase can be H840A or R863A, or a combination thereof.

[0323] In various embodiments, the Cas9 nickase may have a mutation in the HNH nuclease domain and may have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table 10-1] [Table 10-2] [Table 10-3]

[0324] In some embodiments, the N-terminal methionine is removed from the Cas9 nickase or from any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, methionine-minus Cas9 nickases include the following sequences, or variants thereof having amino acid sequences having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table 11-1] [Table 11-2] [Table 11-3]

[0325] E. Other Cas9 variants In addition to inactive Cas9 and Cas9 nickase variants, the Cas9 proteins used herein can also include other "Cas9 variants" that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least ab...

Claims

1. 1. A system for prime editing, comprising: a prime editing factor comprising Cas9 nickase and reverse transcriptase, and Prime-edited guide RNA (PEgRNA) Including, Here, the primed editing factor complexed with PEGRNA is binding to a target DNA sequence, including a target strand and a complementary non-target strand; (i) forming an RNA-DNA hybrid comprising the PEG RNA and the target strand, and (ii) an R-loop comprising the complementary non-target strand; and Nicking the complementary non-target strand to form a reverse transcriptase-primed sequence with a free 3' end can be done, wherein the PEGRNA comprises (a) a guide RNA and (b) an RNA extension arm that comprises, in the 5' to 3' direction, (i) a reverse transcription template sequence containing one or more desired nucleotide changes, and (ii) a reverse transcription primer binding site that is at least 7 nucleotides in length; wherein the guide RNA and the RNA extension arm are in a single molecule; wherein (1) the Cas9 nickase and reverse transcriptase are connected to form a fusion protein, or (2) the reverse transcriptase is recruited to the Cas9 nickase via an RNA-protein or protein-protein interaction domain; and wherein the reverse transcription primer binding site is capable of hybridizing to the free 3' end of the non-target strand of the nicked target DNA sequence; The system.

2. The system described in claim 1, wherein the Cas9 nickase has an amino acid substitution of H840A, N854A, or N863A with reference to the amino acid sequence of SEQ ID NO:

18.

3. The reverse transcriptase (i) a reverse transcriptase from a retrovirus or retrotransposon, or (ii) Moloney murine leukemia virus reverse transcriptase (M-MLV RT) or variant M-MLV RT The system according to claim 1 or 2, wherein:

4. (i) The variant M-MLV RT has the following amino acid substitution: P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, and D653X relative to SEQ ID NO: 89, where X is any amino acid substitution; including one or more of: (ii) the variant M-MLV RT has the following amino acid mutation: For SEQ ID NO: 89, P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, and D653N including one or more of: (iii) the variant M-MLV RT comprises the following amino acid substitutions relative to SEQ ID NO: 89: D200N, T330P, and / or L603W; (iv) the variant M-MLV RT comprises the following amino acid substitutions relative to SEQ ID NO: 89: D200N, T306K, W313F, T330P, and / or L603W; (v) the reverse transcriptase comprises an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 89; or the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 89; (vi) the reverse transcriptase comprises an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 122; or the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 122; or (vii) the variant M-MLV RT is C-terminally truncated and contains a D200N, T306K, W313F, and / or T330P mutation relative to the amino acid sequence of SEQ ID NO: 89; The system of claim 3 .

5. the truncated variant M-MLV RT comprises an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 766; or the truncated variant M-MLV RT comprises the sequence of SEQ ID NO: 766; The system of claim 4.

6. The fusion protein is 2 -[Cas9 nickase]-[reverse transcriptase]-COOH, or NH 2 The system of any one of claims 1 to 5, comprising the structure -[reverse transcriptase]-[Cas9 nickase]-COOH, where each "]-[" indicates the presence of an optional peptide linker.

7. 7. The system of any one of claims 1 to 6, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO: 134 or an amino acid sequence having at least 90%, at least 95%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO:

134.

8. 8. The system of any one of claims 1 to 7, wherein the RNA extension arm is at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides in length.

9. (i) the one or more desired nucleotide changes are within an editing window between −4 and +10 of the PAM sequence; and / or (ii) the one or more desired nucleotide changes are one or more single base nucleotide changes, one or more nucleotide deletions, one or more nucleotide insertions, or a combination thereof; and / or (iii) one or more desired nucleotide changes are present at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 9, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides downstream, A system according to any one of claims 1 to 8.

10. 10. The system of claim 9, wherein the insertion or deletion of one or more nucleotides is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50 nucleotides in length.

11. The reverse transcription template sequence is (i) comprises a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to a target DNA sequence; and / or (ii) encoding a replacement strand that is homologous to the target DNA sequence downstream of the nick, except for one or more desired nucleotide changes; A system according to any one of claims 1 to 10.

12. The system of any one of claims 1 to 11, wherein the primer binding site is 3 to 20 nucleotides in length.

13. The system of any one of claims 1 to 12, wherein the primer binding site is 7 to 17 nucleotides in length.

14. The system of claim 1, wherein the Cas9 nickase and the reverse transcriptase are connected to form a fusion protein.

15. 14. The system of any one of claims 1 to 13, wherein the reverse transcriptase is recruited to the Cas9 nickase via an RNA-protein or protein-protein interaction domain.

16. 16. The system of any one of claims 1 to 13 and 15, wherein the reverse transcriptase is recruited to the Cas9 nickase via an MS2 tagging system.

17. One or more polynucleotides encoding the system of any one of claims 1 to 16.

18. 18. One or more vectors comprising one or more polynucleotides of claim 17.

19. A cell comprising the system of any one of claims 1 to 16, one or more polynucleotides of claim 17, or one or more vectors of claim 18.

20. (i) a system according to any one of claims 1 to 16, one or more polynucleotides according to claim 17, or one or more vectors according to claim 18, and (ii) a pharmaceutically acceptable excipient A pharmaceutical composition comprising:

21. 21. A system according to any one of claims 1 to 16, one or more polynucleotides according to claim 17, one or more vectors according to claim 18, a cell according to claim 19, or a pharmaceutical composition according to claim 20, for use in medicine.

Citation Information

Patent Citations

  • High-throughput precision genome editing

    WO2018049168A1

  • Cytosine to guanine base editor

    WO2018165629A1

  • RNA-guided endonuclease fusion polypeptides and methods of use thereof

    WO2019051097A1

Cited By

  • Methods and compositions for editing edited nucleotide sequences

    JP2026062659A