Improved Prime Editor and Usage
Patent Information
- Application Number
- JP2024507108
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-13
- Filing Date
- 2022-08-05
- Publication Date
- 2025-08-20
AI Technical Summary
Existing prime editing technologies face challenges in achieving efficient integration of edited DNA strands into target genomic sites and reducing the frequency of indel byproducts.
Development of improved prime editors with engineered Cas9 and reverse transcriptase domains, including fusion proteins and uncoupled components, optimized for enhanced specificity and efficiency in genomic editing processes.
The modified prime editors demonstrate significantly increased editing efficiency, up to 10-fold improvement, and reduced indel formation compared to canonical systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] government support This invention was made with Government support under Grant Nos. R01EB031172, R01EB022376, U01Al142756, RM1HG009490 and R35GM118062 awarded by the National Institutes of Health. The Government has certain rights in the invention.
[0002] Related Applications This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. USSN63 / 388,888, filed July 13, 2022, and U.S. Provisional Application No. USSN63 / 230,688, filed August 6, 2021, which are incorporated herein by reference.
[0003] Incorporation by Reference This application references and incorporates by reference the entire contents of each of the following patent applications directed to prime editing that were previously filed by one or more of the present inventors: U.S. Provisional Application No. USSN 62 / 820,813, filed March 19, 2019; U.S. Provisional Application No. USSN 62 / 858,958, filed June 7, 2019; U.S. Provisional Application No. USSN 62 / 889,996, filed August 21, 2019; U.S. Provisional Application No. USSN 62 / 897,996, filed August 21, 2019; U.S. Provisional Application No. USSN 62 / 899,996, filed August 21, 2019; No. 2 / 922,654; U.S. Provisional Application No. USSN62 / 913,553, filed October 10, 2019; U.S. Provisional Application No. USSN62 / 973,558, filed October 10, 2019; U.S. Provisional Application No. USSN62 / 931,195, filed November 5, 2019; U.S. Provisional Application No. USSN62 / 944,231, filed December 5, 2019; U.S. Provisional Application No. USSN62 / 974,537, filed March 1, 2020 No. 62 / 991,069, filed on March 7; No. 63 / 100,548, filed on March 17, 2020; No. 17 / 300,668, filed on September 17, 2021; International PCT Application No. PCT / US2020 / 023721, filed on March 19, 2020; International PCT Application No. PCT / US2020 / 023553, filed on March 19, 2020; International PCT Application No. PCT / US2020 / 023553, filed on March 19, 2020 CT Application No. PCT / US2020 / 023583; U.S. Patent Application No. USSN17 / 219,635, filed March 31; International PCT Application No. PCT / US2020 / 023730, filed March 19, 2020; International PCT Application No. PCT / US2020 / 023713, filed March 19, 2020; U.S. Patent Application No. USSN17 / 219,672, filed March 31, 2021; U.S. Patent Application No. USSNNo. 17 / 751,599; International PCT Application No. PCT / US2020 / 023712, filed March 19, 2020; International PCT Application No. PCT / US2020 / 023727, filed March 19, 2020; International PCT Application No. PCT / US2020 / 023724, filed March 19, 2020; U.S. Patent Application No. USSN17 / 440,68, filed September 17, 2021 No. 2; International PCT Application No. PCT / US2020 / 023725, filed March 19, 2020; International PCT Application No. PCT / US2020 / 023728, filed March 19, 2020; International PCT Application No. PCT / US2020 / 023732, filed March 19, 2020; and International PCT Application No. PCT / US2020 / 023723, filed March 19, 2020.
[0004] This application also references, and incorporates by reference, the entire contents of each of the following patent applications directed to prime editing that were previously filed by one or more of the present inventors: International PCT Application No. PCT / US2022 / 012054, filed January 11, 2022; U.S. Provisional Application No. USSN63 / 255,897, filed October 14, 2021; and U.S. Provisional Application No. USSN63 / 231,230, filed August 9, 2021. , U.S. Provisional Application No. USSN63 / 194,913 filed May 28, 2021, U.S. Provisional Application No. USSN63 / 194,865 filed May 28, 2021, U.S. Provisional Application No. USSN63 / 176,202 filed April 16, 2021, U.S. Provisional Application No. USSN63 / 176,180 filed April 16, 2021, and U.S. Provisional Application No. USSN63 / 136,194 filed January 11, 2021.
[0005] This application additionally references and incorporates by reference the entire contents of each of the following patent applications directed to prime editing that were previously filed by one or more of the present inventors: International PCT Application No. PCT / US2021 / 052097, filed September 24, 2021; U.S. Provisional Application No. USSN63 / 231,231, filed August 9, 2021; U.S. Provisional Application No. USSN63 / 091,272, filed October 13, 2020; U.S. Provisional Application No. USSN63 / 083,067, filed September 24, 2020; and U.S. Provisional Application No. USSN63 / 182,633, filed April 30, 2021.
[0006] This application additionally references and incorporates by reference the entire contents of each of the following patent applications directed to prime editing previously filed by one or more of the present inventors: International PCT Application No. PCT / US2021 / 031439, filed May 7, 2021, U.S. Provisional Application No. 63 / 022,397, filed May 8, 2020, and U.S. Provisional Application No. 63 / 116,785, filed November 20, 2020. [Background technology]
[0007] 2. Background of the Invention Recent developments in prime editing allow for the insertion, deletion, or replacement of genomic DNA sequences without the need for error-prone double-strand DNA breaks. See Anzalone et al., “Search-and-replace genome editing without double-strand breaks or donor DNA,” Nature, 2019, Vol. 576, pp. 149-157, the contents of which are incorporated herein by reference. Prime editing may use an engineered Cas9 nickase-reverse transcriptase fusion protein (e.g., PE1 or PE2) paired with an engineered prime editing guide RNA (pegRNA) that not only directs Cas9 to the target genomic site but also encodes information for installing the desired edit. Prime editing proceeds via a multi-step editing process: 1) the Cas9 domain binds to and nicks the target genomic DNA site specified by the spacer sequence of the pegRNA; 2) the reverse transcriptase domain uses the nicked genomic DNA as a primer to initiate synthesis of an edited DNA strand using an engineered extension on the pegRNA as a template for reverse transcription - this generates a single-stranded 3' flap containing the edited DNA sequence; 3) cellular DNA repair resolves the 3' flap intermediate by transfer of the 5' flap seed that occurs via invasion by the edited 3' flap, excision of the 5' flap containing the original DNA sequence, and ligation of a new 3' flap to incorporate the edited DNA strand, forming a heteroduplex of one edited and one unedited strand; and 4) cellular DNA repair replaces the unedited strand in the heteroduplex using the edited strand as a template for repair, completing the editing process.
[0008] Prime editing represents a powerful tool for genome editing, but modifications that enhance the specificity and efficiency of the prime editing process would help advance the technology.In particular, modifications that facilitate more efficient integration of the edited DNA strand synthesized by the prime editor into the target genome site are desirable.It is also desirable to reduce the frequency of indel by-products that may form as a result of prime editing.Such further modifications to prime editing would advance the technology. Summary of the Invention
[0009] Summary of the Invention The present disclosure describes an improved prime editor system that includes a prime editor fusion protein that includes an engineered Cas9 domain, an engineered reverse transcriptase domain, or a combination of an engineered Cas9 domain and an engineered reverse transcriptase domain. In the case of a prime editor system, the components of the prime editor (i.e., the Cas9 domain and the RT domain) can be provided as individual elements (i.e., uncoupled or unfused). In the case of a prime editor fusion protein, the components of the prime editor (i.e., the Cas9 domain and the RT domain) are provided as a fusion protein.
[0010] In various embodiments, the engineered Cas9 domain of the prime editor system or fusion protein disclosed herein may comprise a variant Cas9 sequence of SEQ ID NO:178, SEQ ID NO:179, or SEQ ID NO:180, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO:178, SEQ ID NO:179, or SEQ ID NO:180.
[0011] In various embodiments, the prime editor system or fusion protein provided herein is a nucleic acid-programmed DNA binding protein (napDNAbp) and a nucleic acid sequence encoding ... The reverse transcriptase may include Mason-Pfizer Monkey Virus (MPMV) reverse transcriptase or a variant thereof, POK11ERV reverse transcriptase or a variant thereof, Simian Retrovirus Type 2 (SRV2) reverse transcriptase or a variant thereof, Woolly Monkey Sarcoma Virus (WMSV) reverse transcriptase or a variant thereof, Vp96 reverse transcriptase or a variant thereof, Vc95 reverse transcriptase or a variant thereof, Ec48 reverse transcriptase or a variant thereof, Gs reverse transcriptase or a variant thereof, Er reverse transcriptase or a variant thereof, Ne144 reverse transcriptase or a variant thereof, Tf1 reverse transcriptase or a variant thereof, or Rs09415 reverse transcriptase ("CRISPR-RT") or a variant thereof.
[0012] In various other embodiments, the engineered RT domain of a prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on the MMLV RT wild type of SEQ ID NO: 33, and may include a variant of SEQ ID NOs: 172-177 or 183-184, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 172-177 or 183-184.
[0013] In still various other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on Ec48 RT and may include variants of SEQ ID NOs: 188-195, 256, and 257, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 188-195, 256, and 257.
[0014] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on Tf1 RT and may include a variant of SEQ ID NOs: 196-213, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 196-213.
[0015] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on PERV RT and may include a variant of SEQ ID NO:214-215 or 234-238, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO:214-215 or 234-238.
[0016] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on AVIRE RT wild type (SEQ ID NO: 216), and may include a variant of SEQ ID NOs: 217-221, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 217-221.
[0017] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on KORV RT wild type (SEQ ID NO: 222), and may include a variant of SEQ ID NOs: 223-227, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 223-227.
[0018] In yet other embodiments, the engineered RT domain of a prime editor system or fusion protein disclosed herein can include a variant RT sequence based on WMSV RT wild type (SEQ ID NO: 228), and can include a variant of SEQ ID NOs: 229-233, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 229-233.
[0019] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on Ne144 RT wild type (SEQ ID NO: 239) and may include a variant of SEQ ID NO: 240, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO: 240.
[0020] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a variant RT sequence based on Vc95 RT wild type (SEQ ID NO: 241), and may include a variant of SEQ ID NO: 242, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO: 242.
[0021] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on Gs RT wild type (SEQ ID NO: 60), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 159-171.
[0022] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein may comprise a five mutant variant RT sequence based on AVIRE RT, KORV RT, and WMSV RT, and may include a variant of SEQ ID NOs: 243-245, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 243-245.
[0023] In yet other embodiments, the engineered RT domain of a prime editor system or fusion protein disclosed herein can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to a variant RT sequence of Tfl-rat4 (SEQ ID NO:251), Tflevo3.1 (SEQ ID NO:252), Tflevo+rat-1 (SEQ ID NO:254), Tf1evo+rat2 (SEQ ID NO:255), Ec48-v2 (SEQ ID NO:256), Ec48-evo3 (SEQ ID NO:257), or any of SEQ ID NOs:251-257.
[0024] In other embodiments, the disclosure describes improved prime editors and prime editor systems, including prime editor fusion proteins including PEmax of SEQ ID NO: 2, which may be encoded by the nucleic acid sequence of SEQ ID NO: 1 and may be modified with any one of the variant Cas9 domains or variant RT domains disclosed herein. The disclosure also provides other improved prime editor variants, including fusion proteins of SEQ ID NOs: 2-8, as well as fusion proteins including evolved nucleic acid programmable DNA binding proteins of SEQ ID NOs: 9-32 and reverse transcriptases of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241. The present disclosure also contemplates fusion proteins having an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to any one of SEQ ID NOs: 2 and 3-8. The present disclosure also contemplates evolved nucleic acid programmed DNA binding proteins having an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to any one of SEQ ID NOs: 9-32. Additionally, the present disclosure contemplates reverse transcriptases having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241.
[0025] In addition, the disclosure provides nucleic acid molecules encoding and / or expressing the evolved and / or modified prime editors described herein, as well as expression vectors or constructs for expressing the evolved and / or modified prime editors described herein, host cells comprising the nucleic acid molecules and expression vectors, and compositions for delivering and / or administering the nucleic acid-based embodiments described herein. In addition, the disclosure provides isolated evolved and / or modified prime editors as described herein, as well as compositions comprising the isolated evolved and / or modified prime editors. Still further, the disclosure provides methods for making evolved and / or modified prime editors, as well as methods for using evolved and / or modified prime editors or nucleic acid molecules encoding evolved and / or modified prime editors in applications including editing nucleic acid molecules, e.g. genomes, preferably in a sequence context-independent manner (i.e., the desired editing site does not require a specific sequence context) with improved efficiency compared to prime editors that form the state of the art. In embodiments, the method of making provided herein is an improved phage-assisted continuous evolution (PACE) system that may be utilized to evolve one or more components of a prime editor (e.g., a Cas9 domain or a reverse transcriptase domain). The present specification also provides a method for efficiently editing a single nucleic acid base of a target nucleic acid molecule, e.g., a genome, using the prime editing system described herein (e.g., in the form of an isolated evolved and / or engineered prime editor described herein or a vector or construct encoding the same), preferably in a sequence context-independent manner, to perform prime editing.Further still, the present specification provides therapeutic methods for treating a genetic disease and / or altering or changing a genetic trait or condition by contacting a target nucleic acid molecule, e.g., a genome, with a prime editing system (e.g., in the form of an isolated evolved and / or engineered prime editor protein or a vector encoding same) and performing prime editing to treat the genetic disease and / or change a genetic trait (e.g., eye color).
[0026] The inventors have surprisingly found that the editing efficiency of prime editing can be significantly increased (e.g., 2-fold increase, 3-fold increase, 4-fold increase, 5-fold increase, 6-fold increase, 7-fold increase, 8-fold increase, 9-fold increase, or 10-fold or more increase) when one or more components of the canonical prime editor (i.e., PE2) are modified. Modifications may include modified amino acid sequences of one or more components (e.g., the Cas9 component, the reverse transcriptase component, or the linker).
[0027] The inventors have developed prime editing, which allows for the insertion, deletion or replacement of genomic DNA sequences without the need for error-prone double-stranded DNA breaks. Prime editing may use an engineered Cas9 nickase-reverse transcriptase fusion protein (e.g., PE1 or PE2) paired with an engineered prime editing guide RNA (pegRNA) that encodes information to direct Cas9 to a target genomic site and install the desired edit. Prime editing proceeds via a multi-step editing process: 1) the Cas9 domain binds to and nicks the target genomic DNA site specified by the spacer sequence of the pegRNA; 2) the reverse transcriptase domain uses the nicked genomic DNA as a primer to initiate synthesis of an edited DNA strand using an engineered extension on the pegRNA as a template for reverse transcription - this generates a single-stranded 3' flap containing the edited DNA sequence; 3) cellular DNA repair resolves the 3' flap intermediate by transfer of the 5' flap seed that occurs via invasion by the edited 3' flap, excision of the 5' flap containing the original DNA sequence, and ligation of a new 3' flap to incorporate the edited DNA strand, forming a heteroduplex of one edited and one unedited strand; and 4) cellular DNA repair replaces the unedited strand in the heteroduplex using the edited strand as a template for repair, completing the editing process.
[0028] Efficient incorporation of the desired edit requires that the newly synthesized 3' flap contains a portion of sequence that is homologous to the genomic DNA site. This homology allows the edited 3' flap to compete with the endogenous DNA strand (the corresponding 5' flap) for incorporation into the DNA duplex. Because the edited 3' flap contains less sequence homology than the endogenous 5' flap, competition is expected to favor the 5' flap strand. Thus, a potential limiting factor in the efficiency of prime editing may be the inability of the edit-containing 3' flap to effectively invade and transfer to the 5' flap strand. Moreover, successful 3' flap invasion and 5' flap removal only incorporates the edit into one strand of the double-stranded DNA genome. Permanent installation of the edit requires cellular DNA repair to use the edited strand as a template to replace the unedited complementary DNA strand. Although cells can be engineered to favor replacement of the unedited strand over the edited strand by using a secondary sgRNA (i.e., the PE3 system) to introduce nicks in the unedited strand adjacent to the edit (step 4 above), this process still relies on a second step of DNA repair.
[0029] The napDNAbp and the polymerase of the prime editor may be joined together to form a fusion protein. In some embodiments, the napDNAbp and the polymerase of the prime editor are joined by a linker to form a fusion protein. In certain embodiments, the linker comprises an amino acid sequence of any one of SEQ ID NOs: 79-93, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 79-93. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 38, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids in length.
[0030] In other embodiments, the linker comprises a SGGSx2-NLS linker, which in certain embodiments corresponds to the amino acid sequence SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS (SEQ ID NO: 79). SV40 -SGGSx2 may be included.
[0031] The components used in the method (for example, prime editor, pegRNA) may be encoded on a DNA vector. In some embodiments, the prime editor, pegRNA is encoded on one or more DNA vectors. In certain embodiments, the one or more DNA vectors include AAV or lentivirus DNA vectors. In some embodiments, the AAV vector is serotype 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
[0032] The prime editor utilized in the method of the present disclosure may be further conjugated to an additional component. In certain embodiments, the second linker is an autohydrolyzable linker. In certain embodiments, the second linker comprises an amino acid sequence of any one of SEQ ID NOs: 79-93, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity with any one of SEQ ID NOs: 79-93. In some embodiments, the second linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acids in length.
[0033] In some embodiments, the one or more modifications to the nucleic acid molecule installed at the target site include one or more transitions, one or more transversions, one or more insertions, one or more deletions, or one or more inversions.In certain embodiments, the one or more transitions are selected from the group consisting of (a) T to C; (b) A to G; (c) C to T; and (d) G to A.In certain embodiments, the one or more transversions are selected from the group consisting of (a) T to A; (b) T to G; (c) C to G; (d) C to A; (e) A to T; (f) A to C; (g) G to C; and (h) G to T. In certain embodiments, the one or more modifications include changing (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair. In some embodiments, the one or more modifications comprise an insertion or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides.
[0034] The disclosed method may be used to make corrections to one or more disease-associated genes. In some embodiments, the one or more modifications include corrections to disease-associated genes. In certain embodiments, the disease-associated genes are associated with polygenic disorders selected from the group consisting of heart disease; hypertension; Alzheimer's disease; arthritis; diabetes; cancer; and obesity. In certain embodiments, the disease-associated genes are associated with monogenic disorders selected from the group consisting of adenosine deaminase (ADA) deficiency; alpha-1 antitrypsin deficiency; cystic fibrosis; Duchenne muscular dystrophy; galactosemia; hemochromatosis; Huntington's disease; maple syrup urine disease; Marfan syndrome; neurofibromatosis type 1; congenital pachyonychia; phenylketonuria; severe combined immunodeficiency; sickle cell disease; Smith-Lemli-Opitz syndrome; trinucleotide repeat disorder; prion disease; and Tay-Sachs disease.
[0035] In another aspect, the present disclosure provides a composition for editing a nucleic acid molecule by prime editing. In some embodiments, the composition comprises a prime editor, a pegRNA, and the composition is capable of installing one or more modifications to the nucleic acid molecule at a target site.
[0036] The composition may increase the efficiency of prime editing and / or reduce the frequency of indel formation. In some embodiments, the efficiency of prime editing is increased by at least 1.5-fold, at least 2.0-fold, at least 2.5-fold, at least 3.0-fold, at least 3.5-fold, at least 4.0-fold, at least 4.5-fold, at least 5.0-fold, at least 5.5-fold, at least 6.0-fold, at least 6.5-fold, at least 7.0-fold, at least 7.5-fold, at least 8.0-fold, at least 8.5-fold, at least 9.0-fold, at least 9.5-fold, or at least 10.0-fold compared to prime editing with PE2. In some embodiments, the frequency of indel formation is reduced by at least 1.5-fold, at least 2.0-fold, at least 2.5-fold, at least 3.0-fold, at least 3.5-fold, at least 4.0-fold, at least 4.5-fold, at least 5.0-fold, at least 5.5-fold, at least 6.0-fold, at least 6.5-fold, at least 7.0-fold, at least 7.5-fold, at least 8.0-fold, at least 8.5-fold, at least 9.0-fold, at least 9.5-fold, or at least 10.0-fold compared to editing with PE2.
[0037] The prime editor utilized in the composition of the present disclosure comprises multiple components. In some embodiments, the prime editor comprises a napDNAbp and a polymerase. In some embodiments, the napDNAbp is a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase domain or a variant thereof. In certain embodiments, the napDNAbp is selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaute, optionally having nickase activity. In certain embodiments, the napDNAbp comprises an amino acid sequence of any one of SEQ ID NOs: 9-32, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 9-32. In certain embodiments, the napDNAbp comprises the amino acid sequence of SEQ ID NO: 10 (i.e., napDNAbp of PE1 and PE2), or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to SEQ ID NO: 10. In some embodiments, the polymerase is a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the polymerase is a reverse transcriptase. In certain embodiments, the reverse transcriptase comprises an amino acid sequence of any one of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241.
[0038] The napDNAbp and the polymerase of the prime editor may be joined together to form a fusion protein. In some embodiments, the napDNAbp and the polymerase of the prime editor are joined by a linker to form a fusion protein. In certain embodiments, the linker comprises an amino acid sequence of any one of SEQ ID NOs: 79-93, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 79-93. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 38, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids in length.
[0039] The components used in the compositions disclosed herein may be encoded on DNA vector. In some embodiments, the prime editor, pegRNA, is encoded on one or more DNA vectors. In certain embodiments, the one or more DNA vectors comprise AAV or lentivirus DNA vectors. In some embodiments, the AAV vector is serotype 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
[0040] The prime editor utilized in the composition of the present disclosure may also be further linked to additional components. In some embodiments, the prime editor as a fusion protein is further joined by a second linker. In certain embodiments, the second linker is an autohydrolyzable linker. In certain embodiments, the second linker comprises an amino acid sequence of any one of SEQ ID NOs: 79-93, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 79-93. In some embodiments, the second linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acids in length.
[0041] In some embodiments, the one or more modifications to the nucleic acid molecule installed at the target site include one or more transitions, one or more transversions, one or more insertions, one or more deletions, or one or more inversions.In certain embodiments, the one or more transitions are selected from the group consisting of (a) T to C; (b) A to G; (c) C to T; and (d) G to A.In certain embodiments, the one or more transversions are selected from the group consisting of (a) T to A; (b) T to G; (c) C to G; (d) C to A; (e) A to T; (f) A to C; (g) G to C; and (h) G to T. In certain embodiments, the one or more modifications include changing (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair. In some embodiments, the one or more modifications comprise an insertion or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides.
[0042] The compositions of the present disclosure may be used to make corrections to one or more disease-related genes. In some embodiments, the one or more modifications include corrections to disease-related genes. In certain embodiments, the disease-related genes are associated with polygenic disorders selected from the group consisting of heart disease; hypertension; Alzheimer's disease; arthritis; diabetes; cancer; and obesity. In certain embodiments, the disease-related genes are associated with monogenic disorders selected from the group consisting of adenosine deaminase (ADA) deficiency; alpha-1 antitrypsin deficiency; cystic fibrosis; Duchenne muscular dystrophy; galactosemia; hemochromatosis; Huntington's disease; maple syrup urine disease; Marfan syndrome; neurofibromatosis type 1; congenital pachyonychia; phenylketonuria; severe combined immunodeficiency; sickle cell disease; Smith-Lemli-Opitz syndrome; trinucleotide repeat disorder; prion disease; and Tay-Sachs disease.
[0043] In another aspect, the present disclosure provides a polynucleotide for editing a DNA target site by prime editing.In some embodiments, the polynucleotide comprises a nucleic acid sequence encoding napDNAbp, a polymerase, and the napDNAbp and the polymerase are capable of installing one or more modifications in the DNA target site in the presence of pegRNA.
[0044] The prime editor utilized in the polynucleotides of the present disclosure comprises multiple components (e.g., napDNAbp and a polymerase). In some embodiments, the napDNAbp is a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase domain or variants thereof. In certain embodiments, the napDNAbp is selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaute, optionally having nickase activity. In certain embodiments, the napDNAbp comprises an amino acid sequence of any one of SEQ ID NOs: 2-8, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 2-8. In certain embodiments, the napDNAbp comprises the amino acid sequence of SEQ ID NO: 10 (i.e., napDNAbp of PE1 and PE2), or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to SEQ ID NO: 10. In some embodiments, the polymerase is a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the polymerase is a reverse transcriptase. In certain embodiments, the reverse transcriptase comprises an amino acid sequence of any one of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241.
[0045] The napDNAbp and the polymerase of the prime editor may be joined together to form a fusion protein. In some embodiments, the napDNAbp and the polymerase of the prime editor are joined by a linker to form a fusion protein. In certain embodiments, the linker comprises an amino acid sequence of any one of SEQ ID NOs: 9-32, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 9-32. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 38, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids in length.
[0046] The polynucleotide disclosed herein may comprise vector.In some embodiments, the polynucleotide is a DNA vector.In certain embodiments, the DNA vector is an AAV or lentivirus DNA vector.In some embodiments, the AAV vector is serotype 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
[0047] The prime editor encoded by the polynucleotide of the present disclosure may also be further conjugated to additional components. In certain embodiments, the second linker comprises an autohydrolyzable linker. In certain embodiments, the second linker comprises an amino acid sequence of any one of SEQ ID NOs: 79-93, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 99% sequence identity to any one of SEQ ID NOs: 79-93. In some embodiments, the second linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 amino acids in length.
[0048] In some embodiments, the one or more modifications to the nucleic acid molecule installed at the target site include one or more transitions, one or more transversions, one or more insertions, one or more deletions, or one or more inversions.In certain embodiments, the one or more transitions are selected from the group consisting of (a) T to C; (b) A to G; (c) C to T; and (d) G to A.In certain embodiments, the one or more transversions are selected from the group consisting of (a) T to A; (b) T to G; (c) C to G; (d) C to A; (e) A to T; (f) A to C; (g) G to C; and (h) G to T. In certain embodiments, the one or more modifications include changing (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair. In some embodiments, the one or more modifications comprise an insertion or deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides.
[0049] The polynucleotides of the present disclosure may be used to make corrections to one or more disease-associated genes. In some embodiments, the one or more modifications include corrections to disease-associated genes. In certain embodiments, the disease-associated genes are associated with polygenic disorders selected from the group consisting of heart disease; hypertension; Alzheimer's disease; arthritis; diabetes; cancer; and obesity. In certain embodiments, the disease-associated genes are associated with monogenic disorders selected from the group consisting of adenosine deaminase (ADA) deficiency; alpha-1 antitrypsin deficiency; cystic fibrosis; Duchenne muscular dystrophy; galactosemia; hemochromatosis; Huntington's disease; maple syrup urine disease; Marfan syndrome; neurofibromatosis type 1; congenital pachyonychia; phenylketonuria; severe combined immunodeficiency; sickle cell disease; Smith-Lemli-Opitz syndrome; trinucleotide repeat disorder; prion disease; and Tay-Sachs disease.
[0050] In another aspect, the disclosure provides a cell. In some embodiments, the cell comprises any of the polynucleotides described herein.
[0051] In another aspect, the present disclosure provides a pharmaceutical composition.In some embodiments, the pharmaceutical composition comprises any of the compositions disclosed herein.In some embodiments, the pharmaceutical composition comprises any of the compositions disclosed herein and a pharmaceutically acceptable excipient.In some embodiments, the pharmaceutical composition comprises any of the polynucleotides disclosed herein.In some embodiments, the pharmaceutical composition comprises any of the polynucleotides disclosed herein and a pharmaceutically acceptable excipient.
[0052] In another aspect, the present disclosure provides a kit.In some embodiments, the kit comprises any of the compositions disclosed herein, pharmaceutical excipients, and instructions for editing DNA target sites by prime editing.In some embodiments, the kit comprises any of the polynucleotides disclosed herein, pharmaceutical excipients, and instructions for editing DNA target sites by prime editing.
[0053] The following drawings form part of this specification and are included to further demonstrate certain aspects of the present disclosure that may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. [Brief description of the drawings]
[0054] [Figure 1] Figure 1 provides a schematic showing the optimization of the PE2 protein. SEQ ID NO:80 is shown.
[0055] [Diagram 2] Figure 2 shows the fold change in frequency of intended editing using PE2 and various other PE constructs in HEK293T cells (low plasmid dose) at a range of gene targets (HEK3, EMX1, RNF2, FANCF, FUNX1, DNMT1, VEGFA, HEK4, PRNP, APOE, CXCR4, HEK3).
[0056] [Diagram 3] FIG. 3 shows the fold change in frequency of intended editing using PE3 and various prime editor constructs in HeLa cells at a range of gene targets (HEK3, FANCF, RUNX1, VEGFA).
[0057] [Figure 4] FIG. 4 shows a comparison of prime editing in HEK293T versus HeLa editing using various PE constructs.
[0058] [Diagram 5] FIG. 5 shows optimization of the NLS construct of PE3 in HeLa cells.
[0059] [Figure 6] FIG. 6 provides a schematic showing the final PEmax construct corresponding to SEQ ID NO:2.
[0060] [Figure 7] FIG. 7 shows that PEmax increases indels in addition to intended edits.
[0061] [Figure 8A] 8A-8C show the expression of PEmax. [Figure 8B] 8A-8B show screening of prime editor variants to maximize editing efficiency in HeLa cells. All PE constructs carry the Cas9 H840A mutation. NLSSV40 indicates the bipartite SV40 NLS. *NLSSV40 contains a 1-aa deletion outside of the PKKKRKV (SEQ ID NO: 94) NLSSV40 consensus sequence. All individual values of n=3 independent biological replicates are shown. [Figure 8C] FIG. 8C shows a comparison of PE3max (a PE3 editing system containing PEmax protein) and PE3 (a PE3 editing system containing PE2 protein) in HeLa cells (average of n=3 independent biological replicates).
[0062] [Figure 9] Figure 9 shows that PEmax constructs enhance editing in disease-relevant gene targets and cell types. Figure 9 provides a schematic of the PE2 and PEmax editor constructs. bpNLSSV40, bipartite SV40 NLS. MMLV RT, Moloney murine leukemia virus reverse transcriptase pentamutant. GS codon, human codon optimized in GenScript.
[0063] [Figure 10] Figure 10 provides a schematic of the phage-assisted continuous evolution (PACE) circuit of prime editors. The PACE circuit is useful for disease-specific evolution, evolution of different prime editor domains, and evolution of all editors.
[0064] [Figure 11] FIG. 11 shows the editing efficiency of evolved Gs mutants in HEK293T cells.
[0065] [Figure 12] Figure 12 shows the editing efficiency of evolved PE2 reverse transcriptase (RT) mutants in HEK293T cells at low doses (75 ng editor). The evolved mutants provide a large benefit at low doses.
[0066] [Figure 13] FIG. 13 provides a schematic of the PACE circuit for Cas9 and reverse transcriptase evolution.
[0067] [Figure 14] FIG. 14 shows the editing efficiency of the Cas9 mutant prime editor in HEK293T cells.
[0068] [Figure 15] FIG. 15 shows the editing efficiency of evolved prime editor mutants in N2A cells.
[0069] [Figure 16] Figure 16 shows that native reverse transcriptase exhibits detectable primed editing activity at the RNF2 and HEK3 sites in HEK293T cells. M-MLV* is an engineered pentamutant variant of M-MLV RT.
[0070] [Figure 17] Figure 17 shows that retroviral reverse transcriptases exhibit prime editing activity. Unique retroviral reverse transcriptase (RT) enzymes exhibit prime editing activity in HEK293T cells at the FANCF and HEK3 loci. MMTV, PERV, AVIRE, KORV and WMSV perform better than wild-type (WT) M-MLV enzyme.
[0071] [Figure 18]Figure 18 shows a comparison of PERV pentamer with PE2. The pentamutant, an engineered version of PERV retroviral RT (21.6), shows improved performance over the WT enzyme. 21.6 has comparable editing to the pentamutant, an engineered version of M-MLV RT (PE2), for FANCF+5 G to T, HEK3+1 His ins, and HEK3+1 FLAG ins editing, but less editing for VEGFA+2 G to A, RNF2+1 C to A, EMX1+5 G to T, and DNMT1-15 deletion editing.
[0072] [Figure 19] Figure 19 shows that the yeast retrotransposon RT enzyme Tf1 RT shows prime editing activity in HEK293T cells. The yeast retrotransposon RT enzyme Tf1 shows prime editing activity in HEK293T cells. Tf1 has higher editing than WT M-MLV reverse transcriptase, but lower activity than the pentamutant engineered enzyme (PE2).
[0073] [Figure 20] Figure 20 shows that mutants S297Q and K118R improve editing activity. A structurally derived rationally designed variant of Tf1 (with S297Q and K118R mutations) shows improved editing over the WT enzyme. The double mutant outperforms the WT enzyme by 1.3-4.2 fold at the four sites tested. PE2 outperforms the rationally designed mutants. Increasing contact between RT and the RNA-DNA substrate improves the outcome of PE.
[0074] [Figure 21] Figure 21 shows the editing efficiency of Tf1 20bp PANCE mutants in HEK293T cells. Tf1 variants (evolved using PANCE) 5.27, 5.59 and 5.60 show improved editing compared to the WT enzyme Tf1 variants in HEK293T cells. Variants 5.59 and 5.60 have comparable editing to PE2 at the sites tested.
[0075] [Figure 22] Figure 22 shows the editing efficiency of evolved Tf1 mutants in N2a cells. Shown is editing using Tf1 variants (evolved using PACE or PANCE) 5.27, 5.47, 5.59 and 5.60 in mouse Neuro2a cells. WT and evolved Tf1 variants (5.47 and 5.60) show higher editing than PE2 at the Dnmt1 locus.
[0076] [Diagram 23] FIG. 23 shows that unique small bacterial reverse transcriptases exhibit prime editing activity in HEK293T cells.
[0077] [Figure 24] Figure 24 shows the editing efficiency of Ec48 20bp PANCE mutants in HEK293T cells. Ec48 variants (evolved using PANCE) 3.8, 3.35, 3.36 and 3.38 show improved editing compared to the WT Ec48 enzyme in HEK293T cells.
[0078] [Diagram 25] Figure 25 shows the editing efficiency of evolved Ec48 mutants in N2a cells. Ec48 variants (evolved using PACE or PANCE) 3.8, 3.23, 3.35, 3.36, 3.37 and 3.38 were used in mouse Neuro2a cells. The evolved Ec48 variants show equivalent editing at the Dnmt1 locus as PE2.
[0079] [Figure 26] FIG. 26 provides the structural components of PEmax from the N-terminus to the C-terminus.
[0080] [Figure 27A]FIG. 27A illustrates a strategy for improving a prime editor, e.g., PE2, that involves (a) PACE evolution of the Cas9 domain, (b) PACE evolution of the RT domain, and (c) replacing the RT domain with an alternative RT domain.
[0081] [Figure 27B] Figure 27B provides a list of embodiments of the prime editors disclosed herein that include a PACE-evolved Cas9 domain and an MMLV domain or variant thereof. The amino acid substitution (e.g., "T128N") refers to the amino acid position of the wild-type MMLV protein of SEQ ID NO: 33.
[0082] [Figure 28] FIG. 28 provides a list of alternative reverse transcriptase domains described in Example 2 herein that can be used in place of the MMLV domain of PE2 or in another prime editor.
[0083] [Figure 29] FIG. 29 shows that incorporation of PE2 mutations into the retroviruses RT AVIRE, KORV, WMSV and PERV improves average prime editing activity compared to the WT enzyme at four different loci in HEK293T cells.
[0084] [Diagram 30] Figure 30 shows that incorporation of all five mutations into PERV-RT improves activity by 6.6-fold compared to the WT enzyme across nine different edits in HEK293T cells. (The 21.6 mutations are D199N, T305K, W312F, E329P, L602W).
[0085] [Figure 31A] Figures 31A-31D show the construction and validation of the PE-PACE circuit of Figure 10. Figure 31A shows the initial overnight growth of PE2 RT phage within the circuit. [Figure 31B]FIG. 31B shows an overnight growth screen of pegRNA. [Figure 31C] FIG. 31C shows overnight growth of PE1 and PE2 in circuits with optimized pegRNA. [Figure 31D] Figure 31D shows PANCE selection of PE1 RT phage. Green shaded circles are drift, no selection pressure was applied.
[0086] [Diagram 32] FIG. 32 provides a summary of the mutations in M-MLV RT introduced by PANCE of PE1.
[0087] [Diagram 33] Figures 33A-33B are modified PE-PACE circuits. Figure 33A shows that phage growth decreases as expression of T7 RNAP is reduced either through the RBS or the promoter. This improves stringency. Figure 33B shows pegRNA optimization for the 20 bp insertion PE-PACE circuit. Numbers on the x-axis indicate different pegRNAs.
[0088] [Diagram 34] Figure 34 is a bar graph showing that evolved variants of Tf1 (evolved using PANCE), 5.27, 5.59 and 5.60, show improved editing in HEK293T cells compared to the WT enzyme Tf1 variants. Variants 5.59 and 5.60 have equivalent editing to PE2 at the sites tested above.
[0089] [Diagram 35] FIG. 35 shows the editing activity of seven unique small bacterial RT enzymes active in HEK293T cells.
[0090] [Diagram 36] FIG. 36 shows that the evolved variant 38.14 is on average 23-fold superior to the WT enzyme across the four loci in HEK293T cells.
[0091] [Figure 37] FIG. 37 shows that the Vc95 variant (L11M+S75A+V97M+N146D+N245T) is on average 7-fold superior to the WT enzyme across the four loci.
[0092] [Figure 38] Figures 38A-38B show the evolution of Gs RT. Mammalian prime editing in HEK293T cells for Gs RT mutants derived from (A) PANCE or (B) PACE.
[0093] [Figure 39] Figure 39 shows the PE-PACE evolution of Cas9. The bar graph compares the editing efficiency of PE2 in HEK293T cells with three evolved prime editors using the PE-PACE system of Figure 13. The evolved editors contain modifications to the Cas9(H840A) component of PE2.
[0094] [Diagram 40] FIG. 40 shows structure-guided engineering of Tf1 reverse transcriptase, where variants I260L, E274R, R288Q and Q293K showed improved editing over WT in HEK293T cells.
[0095] [Diagram 41] FIG. 41 shows structure-guided engineering of 28 Tf1 reverse transcriptase mutants, where variants K118R, S188K, I64L, I64W, N316Q, K321R, L133N showed improved editing over WT in HEK293T cells.
[0096] [Diagram 42]FIG. 42 shows the editing capacity of rationally designed Tf1 variants containing mutation combinations (5.19=wild type Tf1+K118R+S297Q; 5.618=K118R+S297Q+S188K+I64L+I260L+R288Q; 5.59=E22K+P70T+G72V+M102I+K106R+A139T+L158Q+F269L+A363V+K413E+S492N), with variant 5.618 showing comparable editing to the best evolved variant 5.59 in HEK293T cells.
[0097] [Diagram 43] FIG. 43 shows the editing capabilities of Tf1 variants containing combinations of mutations derived from rational design and evolution approaches (5.59=E22K+P70T+G72V+M102I+K106R+A139T+L158Q+F269L+A363V+K413E+S492N; 5.618=K118R+S297Q+S188K+I64L+I260L+R288Q; 5.612=5.59+K118R+S297Q+S188K+I64L+I260L), with variant 5.59 showing further improved activity in HEK293T cells and Tf1 variant 5.612 showing improved activity over PE2.
[0098] [Figure 44A] Figures 44A-44B show an exemplary evolution approach that yielded Ec48 reverse transcriptase variants. Figure 44A shows the genotype of Ec48 after selection using PANCE against a higher stringency strain. [Figure 44B] FIG. 44B shows the use of a more stringent promoter called ProB, which contains sd8 regulatory sequences and the Syn 4.0 regulatory sequences combined with a 20 bp deletion, which was used in place of ProD, which contains a 20 bp deletion.
[0099] [Diagram 45]Figure 45 shows the editing capabilities of Ec48 mutants in HEK293T cells, with variants 3.500 (E60K+K87E+E165D+D243N+R267I+E279K+K318E+K343N) and 3.501 (E60K+K87E+S151T+E165D+D243N+R267I+E279K+V303M+K318E+K343N) outperforming the best previously characterized evolved variant 3.35 (E54K+K87E+D243N+R267I+E279K+K318E).
[0100] [Figure 46] Figure 46 shows improved editing efficiency of a Tf1-based prime editor using five mutations (K118R, S188K, I260L, S297Q and R288Q) predicted by structure-guided engineering.
[0101] [Figure 47] Figure 47 shows improved editing of the Tf1-based prime editor when combining mutations to generate rat1 (K118R+S188K), rat2 (K118R+S188K+I260L), rat3 (K118R+S188K+I260L+S297Q) and rat4 (K118R+S188K+I260L+S297Q+R288Q) variants.
[0102] [Figure 48] Figure 48 shows improved editing of the Tf1-based prime editor using the Tf1evo3.1 and Tf1evo3.2 variants.
[0103] [Figure 49] Figure 49 shows that combining rational mutations with the best evolved variants results in slightly improved editing on average at specific sites.
[0104] [Figure 50A]Figures 50A-50B show the improvement of the editing efficiency of an Ec48-based prime editor using five mutations predicted by structure-guided engineering. Figure 50A shows the editing efficiency of the T189N EC48 mutant. [Figure 50B] Figure 50B shows the editing efficiencies of R378K, K307R, T385R, L182N and R315K mutants.
[0105] [Figure 51] Figure 51 shows the improved editing efficiency of the Ec48-based prime editor when mutations are combined to generate the Ec48-v2 (R315K+L182N+T189N) variant.
[0106] [Figure 52] Figure 52 shows that the Ec48-evo3 variant shows further improvement in editing efficiency.
[0107] [Figure 53] Figure 53 shows the editing efficiency, expressed as editing percentage, at the indicated target genes of Tf1 and Ec48 variants in the PEmax configuration.
[0108] [Figure 54] FIG. 54 shows an overview of the improvement of short RTT editing performed in N2A cells by the indicated M-MLV mutants.
[0109] [Figure 55A] Figures 55A-55B show an overview of the improvement of long RTT editing by the indicated M-MLV mutants. Figure 55A shows the improvement over full-length PE2max in HEK293T cells. [Figure 55B] Figure 55B shows improvement compared to truncated PE2max in HEK293T cells.
[0110] [Figure 56]Figure 56 shows additional PACE and PANCE evolved and engineered Cas9 mutants that improve mammalian prime editing in N2A cells.
[0111] [Figure 57A] Figures 57A-57C show the Tay-Sachs disease circuit. Figure 57A shows the circuit setup demonstrating where the pathogenic fragment is inserted in T7 RNAP. [Figure 57B] Figure 57B shows the sequence of the mutation-containing T7 region before prime editing. [Figure 57C] Figure 57C shows the resulting sequencing after prime editing, in which the correct frame is restored.
[0112] [Figure 58A] Figures 58A-58B show the editing efficiency, expressed as percent editing, of Ec48 and Gs variants. Figure 58A shows the editing efficiency of Ec48-3.35, Ec48-3.500 and Ec48-TSD1 variants. [Figure 58B] Figure 58B shows the editing efficiencies of Gs811, Gs813, Gs814, Gs815, Gs816, Gs-TSD1, Gs-TSD2 and Gs-TSD3 variants.
[0113] [Figure 59] Figure 59 shows the improved editing capabilities of the pentamutant versions of each retroviral RT enzyme over the individual mutants. For AVIRE RT, KORV RT and WMSV RT, five mutations that improved editing were combined, resulting in an additive effect on editing efficiency. The final variants PERV_penta, AVIRE_penta, KORV_penta and WMSV_penta demonstrated approximately 4-fold to 7-fold improvement in editing efficiency on average across five edits. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0114] definition Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide those of ordinary skill in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2 nd ed.1994);The Cambridge Dictionary of Science and Technology(Walker ed.,1988);The Glossary of Genetics,5 th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.
[0115] Cas9 The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease that includes a Cas9 domain or a fragment thereof (for example, a protein that includes an active or inactive DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). A "Cas9 domain" as used herein is a protein fragment that includes an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nuclease is also sometimes referred to as casn1 nuclease or CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to the preceding mobile element, and a targeting invading nucleic acid. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, corrective processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the spacer. Target strands that are not complementary to the crRNA are first cut endonucleolytically and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require a protein and both RNAs. However, single guide RNAs ("sgRNAs" or simply "gNRAs") can be engineered to incorporate both crRNA and tracrRNA aspects into a single RNA species.See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes,” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor Rnase III.” Deltcheva et al., “Complete genome sequence of an M1 strain of Streptococcus pyogenes,” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor Rnase III,” Deltcheva Jinek M., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, including Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5,726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0116] The nuclease-inactivated Cas9 domain may be interchangeably referred to as a "dCas9" protein (representing a nuclease-inactivated Cas9). Methods for generating a Cas9 domain (or a fragment thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5):1173-83, the contents of each of which are incorporated herein by reference in their entirety). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, whereas the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S.pyogenes Cas9 (Jinek et al., Science.337:816-821(2012); Qi et al., Cell.28;152(5):1173-83(2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, the protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant shares homology with Cas9 or a fragment thereof.For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:9). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 9). In some embodiments, the Cas9 variant comprises a fragment of Cas9 of SEQ ID NO:9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:9). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO:9).
[0117] The wild-type canonical Streptococcus pyogenes Cas9 (SpCas9) sequence reference herein has the following amino acid sequence:
Table 1
[0118] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of preceding infection by viruses that have invaded prokaryotes. The snippets of DNA are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with an array of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, effectively comprise a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the RNA. Specifically, the target strand that is not complementary to the crRNA is first cut by an endonuclease, and then 3'-5' is trimmed by an exonuclease. In nature, DNA binding and cleavage typically requires a protein and both RNA. However, single guide RNAs ("sgRNAs" or simply "gNRAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into the guide RNA of a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 helps distinguish self from non-self by recognizing a short motif in the CRISPR repeats (the PAM or protospacer adjacent motif).CRISPR biology, as well as Cas9 nuclease sequences and structure, are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes,” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor Rnase III,” Deltcheva, J. M. et al ... Jinek M., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to the skilled artisan based on this disclosure, including Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.
[0119] In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular nucleic acid targets that are complementary to the RNA. Specifically, the target strand that is not complementary to the crRNA is first cut by an endonuclease and then trimmed by a 3'-5 exonuclease. In nature, DNA binding and cleavage typically requires a protein and both RNAs. However, single guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into the guide RNA of a single RNA species.
[0120] In general, a "CRISPR system" collectively refers to the transcripts and other elements that accompany or direct the expression of CRISPR-associated ("Cas") genes, and includes sequences encoding Cas genes, tracr (transactivating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (which, in the context of an endogenous CRISPR system, encompass "direct repeats" and partial direct repeats processed by tracrRNA), guide sequences (also referred to as "spacers" in the context of an endogenous CRISPR system), or other sequences and transcripts from the CRISPR locus. The tracrRNA of the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA.
[0121] DNA synthesis template As used herein, the term "DNA synthesis template" refers to the region or portion of the extension arm of the PEgRNA that is utilized by the polymerase of the prime editor as the template strand (containing the desired edit and encoding the 3' single-stranded DNA flap that then replaces the corresponding endogenous DNA strand at the target site by the mechanism of prime editing). The extension arm encompasses the DNA synthesis template and may be composed of DNA or RNA. In the case of RNA, the polymerase of the prime editor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of DNA, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. In various embodiments, the DNA synthesis template may include the "editing template" and the "homology arm", as well as all or a portion of the optional 5'-terminal modification region e2. That is, depending on the nature of the e2 region (e.g., whether it includes a hairpin, a two-loop, or a stem / loop secondary structure), the polymerase may code for none, or a portion or all of the e2 region. In other words, in the case of a 3' extension arm, the DNA synthesis template may include a portion of the extension arm that extends from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core, which may act as a template for synthesis of a single strand of DNA by a polymerase (e.g., reverse transcriptase). In the case of a 5' extension arm, the DNA synthesis template may include a portion of the extension arm that extends from the 5' end of the PEgRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of the PEgRNA having either a 3' extension arm or a 5' extension arm. Certain embodiments described herein refer to an "RT template" that includes the editing template and the homology arm, i.e., the sequence of the PEgRNA extension arm that is actually used as a template during DNA synthesis. The term "RT template" is equivalent to the term "DNA synthesis template".
[0122] Editing template The term "editing template" refers to the portion of the extension arm that encodes the desired edit in the single-stranded 3' DNA flap synthesized by a polymerase, e.g., a DNA-dependent DNA polymerase, an RNA-dependent DNA polymerase (e.g., a reverse transcriptase). Certain embodiments described herein refer to the "RT template" which refers to both the editing template and the homologous arm together, i.e., the sequence of the PEgRNA extension arm that is actually used as a template during DNA synthesis. The term "RT editing template" is also equivalent to the term "DNA synthesis template", where RT editing template reflects the use of a prime editor with a polymerase that is a reverse transcriptase, and DNA synthesis template reflects the use of a prime editor with any polymerase more broadly.
[0123] Extension arm The term "extension arm" refers to a nucleotide sequence component of a PEgRNA that serves several functions, including a primer binding site and an editing template for reverse transcriptase. In some embodiments, the extension arm is located at the 3' end of the guide RNA. In other embodiments, the extension arm is located at the 5' end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In various embodiments, the extension arm includes the following components in the 5' to 3' direction: a homology arm, an editing template, and a primer binding site. Since the polymerization activity of reverse transcriptase is in the 5' to 3' direction, the preferred arrangement of the homology arm, the editing template, and the primer binding site is in the 5' to 3' direction, such that the reverse transcriptase, once primed by the annealed primer sequence, polymerizes a single strand of DNA using the editing template as a complementary template strand. Further details, such as the length of the extension arm, are described elsewhere herein.
[0124] The extension arm may also be generally described as comprising two regions: for example, a primer binding site (PBS) and a DNA synthesis template. The primer binding site binds to a primer sequence formed from the endogenous DNA strand of the target site when it is nicked by the prime editor complex, thereby exposing its 3' end on the nicked endogenous strand. As described herein, binding of the primer sequence to the primer binding site on the extension arm of the PEgRNA creates a double-stranded region with an exposed 3' end (i.e., 3' of the primer sequence), which then provides a substrate for a polymerase to initiate general polymerization of DNA from the exposed 3' end along the length of the DNA synthesis template. The sequence of the single-stranded DNA product is the complement of the DNA synthesis template. Polymerization continues 5' of the DNA synthesis template (or extension arm) until polymerization terminates. Thus, the DNA synthesis template represents the portion of the extension arm that is encoded by the polymerase of the prime editor complex into the single-stranded DNA product (i.e., the 3' single-stranded DNA flap containing the desired gene editing information) and ultimately replaces the corresponding endogenous DNA strand at the target site immediately downstream of the PE-induced nick site. Without being bound by theory, polymerization of the DNA synthesis template continues to the 5' end of the extension arm until an event known as termination. Polymerization may terminate in a variety of ways, including, but not limited to, (a) reaching the 5' end of the PEgRNA (e.g., in the case of a 5' extension arm, the DNA polymerase simply runs out of the template), (b) reaching an RNA secondary structure that cannot be threaded (e.g., a hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a nucleic acid topology signal such as supercoiled DNA or RNA.
[0125] Fusion proteins The term "fusion protein" as used herein refers to a hybrid polypeptide that contains protein domains from at least two different proteins. A protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) portion of the protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein", respectively. The protein may contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which directs the protein to bind to a target site) and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. Another example includes Cas9 or its equivalent for reverse transcriptase. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins that include peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4), the entire contents of which are incorporated herein by reference. th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0126] Guide RNA ("gRNA") As used herein, the term "guide RNA" refers to a specific type of guide nucleic acid that is generally associated with the Cas protein of CRISPR-Cas9 and associates with Cas9 to direct Cas9 protein to a specific sequence in a DNA molecule (including the complementarity of the protospace sequence of the guide RNA).However, this term also includes equivalent guide nucleic acid molecules that associate with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or not naturally occurring (e.g., engineered or recombinant), and separately program Cas9 equivalents to localize to a specific target nucleotide sequence.Cas9 equivalents may also include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system) and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are also provided herein. As used herein, "guide RNA" may also be referred to as "conventional guide RNA," in contrast to a modified form of guide RNA termed "prime editing guide RNA" (or "PEgRNA").
[0127] A guide RNA or PEgRNA may comprise a variety of structural elements, including but not limited to:
[0128] Spacer sequence - a sequence in the guide RNA or PEgRNA (having a length of about 20 nts) that has the same sequence as the protospacer in the target DNA.
[0129] gRNA core (or gRNA scaffold or backbone sequence) - refers to the sequence within the gRNA responsible for Cas9 binding, this does not include the 20 bp spacer / targeting sequence used to guide Cas9 to the target DNA.
[0130] Extension Arm - A single-stranded extension at the 3' or 5' end of the PEGRNA that contains a primer binding site and a DNA synthesis template sequence that encodes a single-stranded DNA flap containing the genetic alteration of interest by a polymerase (e.g., reverse transcriptase), which then incorporates into the endogenous DNA by displacing the corresponding endogenous strand, thereby installing the desired genetic alteration.
[0131] Transcription terminator-guide RNA or PEgRNA may include a transcription termination sequence 3' to the molecule.
[0132] host cell The term "host cell," as used herein, refers to a cell that can host, replicate, and express the vectors described herein, e.g., vectors comprising a nucleic acid molecule encoding an MLH1 variant and a fusion protein comprising Cas9 or a Cas9 equivalent and reverse transcriptase.
[0133] Linker The term "linker" as used herein refers to a molecule that connects two other molecules or moieties. A linker can be an amino acid sequence in the case of a linker that joins two fusion proteins. For example, Cas9 can be fused to reverse transcriptase by an amino acid linker sequence. A linker can also be a nucleotide sequence in the case of joining two nucleotide sequences together. For example, in the case of the present invention, a conventional guide RNA is linked to the RNA extension of a primed editing guide RNA, which may include a RT template sequence and a RT primer binding site, via a spacer or linker nucleotide sequence. In other embodiments, the linker is an organic molecule, a group, a polymer, or a chemical moiety. In some embodiments, the linker is 5-100 amino acids long, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids long. Longer or shorter linkers are also contemplated. In certain embodiments, the linker is an autohydrolyzable linker (e.g., the 2A autocleaving peptide, further described herein). Autohydrolyzable linkers, such as the 2A self-cleaving peptide, can induce ribosome skipping during protein translation, such that the ribosome is unable to make a peptide bond between two genes or gene fragments.
[0134] napDNAbp As used herein, the term "nucleic acid programmed DNA binding protein" or "napDNAbp" (of which Cas9 is an example) refers to a protein that uses RNA:DNA hybridization to target and bind to a specific sequence in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that includes a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., the protospacer of the guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to the complementary sequence.
[0135] Without being bound by theory, the binding mechanism of the napDNAbp-guide RNA complex generally involves the formation of an R-loop in which the napDNAbp induces the unwinding of the double-stranded DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the "target strand". It displaces the "non-target strand" that is complementary to the target strand, forming a single-stranded region of the R-loop. In some embodiments, the napDNAbp contains one or more nuclease activities and then cuts the DNA to eliminate various types of lesions. For example, the napDNAbp may contain a nuclease activity that cuts the non-target strand at a first location and / or cuts the target strand at a second location. In response to the nuclease activity, the target DNA may be cut to form a "double-stranded break" in which both strands are cut. In other embodiments, the target DNA may be cut at only a single site, i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with different nuclease activities include "Cas9 nickase" ("nCas9"), and inactivated Cas9 that has no nuclease activity ("inactivated Cas9" or "dCas9"). Exemplary sequences for these and other napDNAbps are provided herein.
[0136] Nickase The term "nickase" refers to Cas9 with one of the two nuclease domains inactivated, allowing the enzyme to cleave only one strand of the target DNA.
[0137] nucleic acid molecule The term "nucleic acid" as used herein refers to a polymer of nucleotides. The polymer may be any combination of naturally occurring nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyluridine, C5 propynylcytidine, C5 methylcytidine, 7 deazaadenosine, 7 deazaguanosine, 8 oxoadenosine, 8 oxoguanosine, O(6) methylguanine, 4-acetylcytidine, 5-acetylcytidine, 6-acetylcytidine, 7-acetylcytidine, 8 ... , 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudouridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N phosphoramidite linkages).
[0138] PACE The term "phage-assisted continuous evolution (PACE)" as used herein refers to continuous evolution employing phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, published March 11, 2010 as WO 2010 / 028347; International Application No. PCT / US2011 / 066747, filed December 22, 2011, published June 28, 2012 as WO 2012 / 088381; U.S. Patent Application No. 9,023,599, published May 5, 2015; and U.S. Patent Application No. 10,233,661, filed December 22, 2011, published June 28, 2012 as WO 2012 / 088381. No. 4, International PCT application PCT / US2015 / 012022, filed January 20, 2015, which published on September 11, 2015 as WO 2015 / 134121, and International PCT application PCT / US2016 / 027795, filed April 15, 2016, which published on October 20, 2016 as WO 2016 / 168631, the entire contents of each of which are incorporated herein by reference.
[0139] PEgRNA As used herein, the term "prime edited guide RNA" or "PEgRNA" or "extended guide RNA" refers to a specialized form of guide RNA that has been modified to include one or more additional sequences to practice the prime editing methods and compositions described herein. As described herein, a prime edited guide RNA includes one or more "extended regions" of a nucleic acid sequence. The extended region may include, but is not limited to, a single strand of RNA or DNA. Additionally, the extended region may occur at the 3' end of a conventional guide RNA. In other configurations, the extended region may occur at the 5' end of a conventional guide RNA. In yet other configurations, the extended region may occur at the terminal region of a conventional guide RNA, for example, at the gRNA core region that associates with and / or binds to the napDNAbp. The extended region contains a "DNA synthesis template" that encodes a single-stranded DNA (by the polymerase of the prime editor) and then (a) is designed to be homologous to the endogenous target DNA to be edited, and (b) contains at least one desired nucleotide change (e.g., transition, transversion, deletion, or insertion) to be introduced or incorporated into the endogenous target DNA. The extended region may also contain other functional sequence elements (such as, but not limited to, a "primer binding site" and a "spacer or linker" sequence), or other structural elements (such as, but not limited to, an aptamer, a stem loop, a hairpin, a two-loop (e.g., a 3' two-loop), or an RNA-protein recruitment domain (e.g., an MS2 hairpin). As used herein, a "primer binding site" includes a sequence that hybridizes with a single-stranded DNA sequence having a 3' end generated from a DNA having an R-loop nick.
[0140] In certain embodiments, the PEgRNA has a 5' extension arm, a spacer, and a gRNA core. The 5' extension further comprises, from 5' to 3', a reverse transcriptase template, a primer binding site, and a linker. The reverse transcriptase template may be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editor described herein is not RT but another type of polymerase.
[0141] In certain other embodiments, the PEgRNA has a 5' extension arm, a spacer, and a gRNA core. The 5' extension further comprises, in the 5' to 3' direction, a reverse transcriptase template, a primer binding site, and a linker. The reverse transcriptase template may be more broadly referred to as a "DNA synthesis template," in which the polymerase of the prime editor described herein is not RT but another type of polymerase.
[0142] In yet another embodiment, the PEgRNA has, from 5' to 3', a spacer (1), a gRNA core (2), and an extension arm (3). The extension arm (3) is at the 3' end of the PEgRNA. The extension arm (3) further comprises, from 3' to 5', a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may comprise any modification region at the 3' and 5' ends, which may be the same or different sequences. In addition, the 3' end of the PEgRNA may comprise a transcription terminator sequence. These sequence elements of the PEgRNA are further described and defined herein.
[0143] In yet another embodiment, the PEgRNA has, from 5' to 3', an extension arm (3), a spacer (1), and a gRNA core (2). The extension arm (3) is at the 5' end of the PEgRNA. The extension arm (3) further comprises, from 3' to 5', a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may comprise any modification region at the 3' and 5' ends, which may be the same or different sequences. The PEgRNA may comprise a transcription terminator sequence at the 3' end. These sequence elements of the PEgRNA are further described and defined herein.
[0144] PE1 As used herein, "PE1" refers to a PE complex comprising the following structure: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(wt)]+ a fusion protein comprising Cas9(H840A) with a desired PE gRNA and wild-type MMLV RT, where the PE fusion has the amino acid sequence of SEQ ID NO:3, as shown below; [ka]
[0145] PE2 As used herein, "PE2" refers to a PE complex comprising a fusion protein containing Cas9(H840A) and a variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+a desired PEgRNA, wherein the PE fusion has the amino acid sequence of SEQ ID NO:4, as shown below. [ka]
[0146] PE3 As used herein, "PE3" refers to PE2 and a second strand nicking guide RNA that complexes with PE2 and introduces a nick on the non-edited DNA strand to induce preferential displacement of the edited strand.
[0147] PE3b As used herein, "PE3b" refers to PE3, but the second strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit. This is accomplished by designing a gRNA with a spacer sequence that matches only the edited strand and not the original allele. Using this strategy, hereafter referred to as PE3b, mismatches between the protospacer and the unedited allele should not favor nicking by the sgRNA until after the editing event on the PAM strand has occurred.
[0148] PEmax As used herein, "PEmax" refers to a PE complex comprising a fusion protein comprising Cas9 (R221K N39K H840A) and a variant MMLV RT pentamutant (D200N T306K W313F T330P L603W) having the following structure: [Bipartite NLS]-[Cas9(R221K)(N394K)(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)]-[Bipartite NLS]-[NLS]+desired PE gRNA, where the PE fusion has the amino acid sequence of SEQ ID NO:2, and the nucleic acid sequence of SEQ ID NO:1, as shown below: [ka] [ka] [ka] [ka]
change
change
[0149] Polymerase As used herein, the term "polymerase" refers to an enzyme that synthesizes a nucleotide strand that may be used in connection with the prime editor system described herein. The polymerase may be a "template-dependent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand based on the order of nucleotide bases of a template strand). The polymerase may also be a "template-independent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand without the requirement of a template strand). The polymerase may be further categorized as a "DNA polymerase" or an "RNA polymerase". In various embodiments, the prime editor system comprises a DNA polymerase. In various embodiments, the DNA polymerase may be a "DNA-dependent DNA polymerase" (i.e., whereby the template molecule is a strand of DNA). In such cases, the DNA template molecule may be a PEgRNA, and the extension arm comprises a strand of DNA. In such cases, the PEgRNA may be referred to as a chimeric or hybrid PEgRNA, which comprises an RNA portion (i.e., a guide RNA component that includes a spacer and a gRNA core) and a DNA portion (i.e., an extension arm). In various other embodiments, the DNA polymerase may be an "RNA-dependent DNA polymerase" (i.e., whereby the template molecule is a strand of RNA). In such cases, the PEgRNA is RNA, i.e., RNA elongation is involved. The term "polymerase" may refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Generally, the enzyme will initiate synthesis at the 3' end of a primer annealed to a polynucleotide template sequence (such as, for example, a primer sequence annealed to the primer binding site of the PEgRNA) and proceed toward the 5' end of the template strand. A "DNA polymerase" catalyzes the polymerization of deoxynucleotides. As used herein with reference to a DNA polymerase, the term DNA polymerase includes "functional fragments thereof."A "functional fragment thereof" refers to any portion of a wild-type or mutant DNA polymerase that encompasses less than the entire amino acid sequence of the polymerase and that retains the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. Such functional fragments may exist as separate entities or may be components of a larger polypeptide, such as a fusion protein.
[0150] PrimeEdit As used herein, the term "prime editing" refers to an approach for gene editing that uses napDNAbp, a polymerase (e.g., reverse transcriptase), and a specialized guide RNA that contains a DNA synthesis template to code for the desired new genetic information (or delete genetic information) that is then incorporated into the target DNA sequence. Classical prime editing is described in the inventors' publication Anzalone, AV et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019), which is incorporated herein by reference in its entirety.
[0151] Prime editing represents a platform for genome editing, a versatile and precise genome editing method that writes new genetic information directly at specific DNA sites, using a nucleic acid programmable DNA binding protein ("napDNAbp") that acts in association with a polymerase (i.e., provided in the form of a fusion protein or otherwise in trans with napDNAbp), and the prime editing system is programmed by a prime editing (PE) guide RNA ("PEgRNA") that identifies the target site and serves as a template for the synthesis of the desired edit in the form of a replaced DNA strand by an extension (either DNA or RNA) engineered onto the guide RNA (e.g., at the 5' or 3' end of the guide RNA, or at an internal portion). The replaced strand containing the desired edit (e.g., a single nucleobase substitution) shares the same (or is homologous to) sequence as the endogenous strand (immediately downstream of the nick site) of the target site to be edited (except that it encompasses the desired edit). By DNA repair and / or replication mechanisms, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing may be considered a "search and replace" genome editing technology, because the prime editor described herein not only searches for and locates the desired target site to be edited, but also simultaneously encodes a replacement strand containing the desired edit that is installed in place of the endogenous DNA strand at the corresponding target site. The prime editor of the present disclosure relates, in part, to the discovery that the mechanism of targeted prime reverse transcription (TPRT) or "prime editing" can be exploited or adapted to perform precise CRISPR / Cas-based genome editing with high efficiency and genetic flexibility. In nature, TPRT is used by mobile DNA elements such as mammalian non-LTR retrotransposons and bacterial group II introns.The inventors herein use a Cas protein-reverse transcriptase fusion or related system in trans to target a specific DNA sequence with a guide RNA, generate a single-stranded nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered reverse transcriptase template incorporated into the guide RNA. However, although the concept begins with a prime editor using reverse transcriptase as a DNA polymerase component, the prime editors described herein are not limited to reverse transcriptase and may encompass the use of virtually any DNA polymerase. Indeed, although the present application may refer throughout to a prime editor having a "reverse transcriptase," it is here shown that reverse transcriptase is the only type of DNA polymerase that may act in prime editing. Thus, the present specification refers to "reverse transcriptase," and one of skill in the art should recognize that any suitable DNA polymerase may be used in place of reverse transcriptase. Thus, in one aspect, the prime editor may comprise Cas9 (or equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specific guide RNA (i.e., PEgRNA) that contains a spacer sequence that anneals to a complementary protospacer of the target DNA. The specific guide RNA also contains new genetic information in the form of an extension that codes for a replacement strand of DNA containing the desired genetic change, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer information from the PEgRNA to the target DNA, the mechanism of prime editing involves nicking one strand of DNA to the target site to expose a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to directly prime DNA polymerization of the editing code extension on the PEgRNA to the target site. In various embodiments, the extension that provides a template for polymerization of the replacement strand containing the edit can be formed from RNA or DNA. In the case of RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as reverse transcriptase). In the case of DNA extension, the polymerase of the prime editor may be a DNA-dependent DNA polymerase.The newly synthesized strand (i.e., the replacement DNA strand containing the desired edit) formed by the prime editor disclosed herein will be homologous to (i.e., have the same sequence as) the genomic target sequence, except for the inclusion of the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be referred to as a single-stranded DNA flap, which competes for hybridization with the complementary, homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In certain embodiments, the system can be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with a Cas9 domain or provided in trans to a Cas9 domain). The error-prone reverse transcriptase can introduce changes during the synthesis of the single-stranded DNA flap. Thus, in certain embodiments, an error-prone reverse transcriptase can be utilized to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used with the system, the changes can be random or non-random. Degradation of the hybridized intermediate (including the single-stranded DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can include removal of the resulting displaced flap of the endogenous DNA (e.g., by 5'-end DNA flap endonuclease FEN1), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis provides single-base accuracy for any nucleotide modification, including insertions and deletions, the scope of this approach is very broad and can foreseeably be used for countless applications in basic research and therapeutics.
[0152] In various embodiments, prime editing is actuated by contacting a target DNA molecule (in which a nucleotide sequence change is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). In various embodiments, the prime editing guide RNA (PEgRNA) includes an extension at the 3' or 5' end of the guide RNA, or at an intramolecular location in the guide RNA, encoding the desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex is contacted with the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) in one of the strands of the DNA at the target locus, thereby generating an available 3' end in one of the strands of the target locus. In certain embodiments, the nick is generated in the strand of DNA corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the "non-target strand." However, the nick can be introduced into either strand. That is, the nick could be introduced into the "target strand" of the R-loop (i.e., the strand hybridized to the protospacer of the extended gRNA) or into the "non-target strand" (i.e., the strand that forms the single-stranded portion of the R-loop and is complementary to the target strand). In step (c), the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime reverse transcription (i.e., "target primed RT"). In certain embodiments, the 3' end DNA strand hybridizes to a specific RT priming sequence of the extended portion of the guide RNA, i.e., the "reverse transcriptase priming sequence" or "primer binding site" on the PEgRNA. In step (d), a reverse transcriptase (or other suitable DNA polymerase) is introduced, which synthesizes a single strand of DNA from the 3' end of the primed site toward the 5' end of the primed edited guide RNA. A DNA polymerase (eg, reverse transcriptase) may be fused to the napDNAbp or alternatively provided in trans relative to the napDNAbp.This forms a single-stranded DNA flap containing the desired nucleotide change (e.g., a single base change, insertion, or deletion, or a combination thereof) that is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve degradation of the single-stranded DNA flap, so that the desired nucleotide change becomes incorporated into the target locus. This process can be driven toward the desired product formation by removing the corresponding 5' endogenous DNA flap that forms once the 3' single-stranded DNA flap invades and hybridizes to the endogenous DNA sequence. Without being bound by theory, the cell's endogenous DNA repair and replication processes degrade the mismatched DNA and incorporate the nucleotide change to form the desired altered product. The process can also be driven toward product formation by "second strand nicking". This process may introduce at least one or more of the following genetic changes: transversion, transition, deletion, and insertion.
[0153] The term "prime editor (PE) system" or "prime editor (PE)" or "PE system" or "PE editing system" refers to a composition involving the method of genome editing using targeted primed reverse transcription (TPRT) described herein, including, but not limited to, napDNAbp, reverse transcriptase, a fusion protein (e.g., comprising napDNAbp and reverse transcriptase), a prime editing guide RNA, and a complex comprising the fusion protein and the prime editing guide RNA, as well as auxiliary elements, such as a second strand nicking component (e.g., second strand sgRNA) and a 5' endogenous DNA flap removal endonuclease (e.g., FEN1) to help drive the prime editing process towards edited product formation.
[0154] In the embodiments described thus far, the PEgRNA constitutes a single molecule comprising the guide RNA (which itself comprises a spacer sequence and a gRNA core or scaffold) and a 5' or 3' extension arm comprising a primer binding site and a DNA synthesis template, however, the PEgRNA may also take the form of two individual molecules, composed of a guide RNA and a trans-prime editor RNA template (tPERT), which essentially houses an extension arm (comprising, inter alia, the primer binding site and the DNA synthesis domain) and an RNA-protein recruitment domain (e.g., an MS2 aptamer or hairpin) on the same molecule, which becomes co-localized or recruited to an engineered prime editor complex comprising a tPERT recruitment protein (e.g., the MS2cp protein which binds to the MS2 aptamer).
[0155] Prime Editor The term "prime editor" refers to a construct comprising napDNAbp (e.g., Cas9 nickase) and reverse transcriptase capable of performing prime editing on a target nucleotide sequence in the presence of PEgRNA (or "extended guide RNA"). The term "prime editor" may refer to a fusion protein, or a fusion protein complexed with PEgRNA, and / or a fusion protein further complexed with a second strand nicking sgRNA. In some embodiments, a prime editor may also refer to a complex comprising a fusion protein (reverse transcriptase fused to napDNAbp), PEgRNA, and a normal guide RNA capable of directing the second site nicking step of the non-edited strand as described herein. In some embodiments, the term prime editor refers to a napDNAbp and reverse transcriptase that are provided in trans or are not otherwise fused to each other.
[0156] Primer binding site The term "primer binding site" or "PBS" refers to a nucleotide sequence located on the PEgRNA (typically at the 3' end of the extension arm) as a component of the extension arm, which serves to bind to a primer sequence formed after Cas9 nicking of the target sequence by the prime editor. As detailed elsewhere, when the Cas9 nickase component of the prime editor nicks one strand of the target DNA sequence, a 3' terminal ssDNA flap is formed, which serves as a primer sequence that anneals to the primer binding site on the PEgRNA to prime reverse transcription.
[0157] Protospacer As used herein, the term "protospacer" refers to a sequence (about 20 bp) on DNA adjacent to a PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, to one strand of it, i.e., the "target strand" to the "non-target strand" of the target DNA sequence). For Cas9 to function, a specific protospacer adjacent motif (PAM) is also required, which differs depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes the PAM sequence NGG on the non-target strand found directly downstream of the target sequence on genomic DNA. Those skilled in the art will recognize that the state of the art literature sometimes refers to the "protospacer" as the approximately 20 nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a "spacer". Thus, in some cases, as used herein, the term "protospacer" may be used interchangeably with the term "spacer." The context of the description surrounding the appearance of either "protospacer" or "spacer" will help inform the reader as to whether the term refers to a gRNA or DNA target.
[0158] Protospacer adjacent motif (PAM) As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is the key targeting component of the Cas9 nuclease. Typically, the PAM sequence is on either strand and is downstream in the 5' to 3' direction of the site cut by Cas9. The canonical PAM sequence (i.e., the PAM sequence associated with Streptococcus pyogenes Cas9 nuclease or SpCas9) is 5'-NGG-3', where "N" is any nucleobase followed by two guanine ("G") nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, can be modified to alter the PAM specificity of the nuclease such that the nuclease recognizes alternative PAM sequences.
[0159] For example, with reference to the canonical SpCas9 amino acid sequence of SEQ ID NO: 9, the PAM sequence can be modified by introducing one or more mutations, including (a) D1135V, R1335Q, and T1337R "VQR variants" that change the PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q, and T1337R "EQR variants" that change the PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E, and T1337R "VRER variants" that change the PAM specificity to NGCG. In addition, the D1135E variant of canonical SpCas9 still recognizes NGG, but more selectively than the wild-type SpCas9 protein.
[0160] It will also be recognized that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) may have various PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. In yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and are not meant to be limiting. In addition, it will be further recognized that non-SpCas9s bind to various PAM sequences, making them useful when a suitable SpCas9 PAM sequence is not present at the site of the desired target to be cut. Moreover, non-SpCas9s may have other features that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is approximately 1 kilobase smaller than SpCas9 and can therefore be packaged into an adeno-associated virus (AAV). Further reference may be made to Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5):891-899, which is incorporated herein by reference.
[0161] Reverse transcriptase The term "reverse transcriptase" describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptases have been used primarily to transcribe mRNA into cDNA, which can then be cloned into vectors for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme has 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand of an RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading, so transcription errors cannot be corrected by the reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity is presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase that is extensively used in molecular biology is the reverse transcriptase originating from Moloney Murine Leukemia Virus (M-MLV or "MMLV"). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases that are substantially devoid of RNase H activity have also been described, see, e.g., U.S. Patent No. 5,244,797.The present invention contemplates the use of any such reverse transcriptase or variants or mutants thereof.
[0162] In addition, the present invention contemplates the use of reverse transcriptase that may be called error-prone or error-prone reverse transcriptase, or reverse transcriptase that does not support high fidelity incorporation of nucleotides during polymerization.During the synthesis of single-stranded DNA flap based on the RT template incorporated with guide RNA, error-prone reverse transcriptase can introduce one or more nucleotides that are mismatched with the RT template sequence, thereby introducing changes in nucleotide sequence by error-prone polymerization of single-stranded DNA flap.These errors that are introduced during the synthesis of single-stranded DNA flap can then be incorporated on double-stranded molecule by hybridization with corresponding endogenous target strand, removal of displaced endogenous strand, ligation, and then another round of endogenous DNA repair and / or sequencing process.
[0163] The present disclosure provides, in some embodiments, a prime editor comprising MMLV RT.
[0164] Reverse transcription As used herein, the term "reverse transcription" refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription can be "error-prone reverse transcription." This refers to the property of certain reverse transcriptase enzymes that their DNA polymerization activity is error-prone.
[0165] Proteins, peptides, and polypeptides The terms "protein", "peptide" and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a group of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified by the addition of a chemical entity, such as, for example, a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may also be a single molecule or a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (2002), the entire contents of which are incorporated herein by reference. 4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0166] Spacer sequence As used herein, the term "spacer sequence" in reference to a guide RNA or PEgRNA refers to a portion of the guide RNA or PEgRNA of about 20 nucleotides that contains a nucleotide sequence that shares the same sequence as a protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R-loop ssDNA structure on the endogenous DNA strand.
[0167] target site The term "target site" refers to a sequence within a nucleic acid molecule that is edited by a prime editor (PE) disclosed herein. Additionally, a target site refers to a sequence within a nucleic acid molecule to which a complex of a prime editor (PE) and a gRNA binds.
[0168] variant As used herein, the term "variant" should be interpreted as meaning to exhibit qualities having a pattern that deviates from those that occur in nature, and as an example, a variant Cas9 is a Cas9 that contains one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. The term "variant" encompasses homologous proteins that have at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% percent identity with the reference sequence and have the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses mutants, truncations, or domains of the reference sequence that exhibit the same or substantially the same functional activity(ies) as the reference sequence.
[0169] vector The term "vector" as used herein refers to a nucleic acid that can be modified to encode a gene of interest, enter a host cell, mutate or replicate within the host cell, and then transfer the replicative form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the present disclosure.
[0170] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS The present disclosure provides compositions and methods for prime editing with improved editing efficiency and / or reduced indel formation. In particular, the present disclosure provides improved prime editor proteins in which one or more components, including the napDNAbp domain and / or reverse transcriptase domain, are modified (e.g., amino acid sequence changes relative to a starting point prime editor such as PE1 or PE2). As illustrated in the examples and described herein, various strategies can be used to obtain variant or engineered protein components, such as variant napDNAbp domains and variant RT domains, e.g., PACE and PANCE evolution methods, and domain replacement with replacement homologous domains (e.g., see diagram in Figure 27A).
[0171] The present disclosure describes an improved prime editor system that includes a prime editor protein that includes an engineered Cas9 domain, an engineered reverse transcriptase domain, or a combination of an engineered Cas9 domain and an engineered reverse transcriptase domain. In the case of a prime editor system, the components of the prime editor (i.e., the Cas9 domain and the RT domain) can be provided as individual elements (i.e., uncoupled or unfused). In the case of a prime editor fusion protein, the components of the prime editor (i.e., the Cas9 domain and the RT domain) are provided as a fusion protein.
[0172] In various embodiments, the engineered Cas9 domain of the prime editor system or fusion protein disclosed herein comprises a variant Cas9 sequence of SEQ ID NO: 178, SEQ ID NO: 179, or SEQ ID NO: 180, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% or up to 100% sequence identity to any of SEQ ID NO: 178, SEQ ID NO: 179, or SEQ ID NO: 180. provided that the amino acid sequence contains at least one substitution relative to wild-type Cas9 selected from the group consisting of D23G, H99Q, H99R, E102K, E102S, E102R, N175K, D177G, K218R, N309D, I312V, E471K, G485S, K562N, D608N, I632V, D645N, D645E, R654C, G687D, G715E, H721Y, R753K, R753G, H754R, K775R, E790K, T804A, K918A, K1003R, M1021Y, E1071K, and E1260D.
[0173] In various embodiments, the prime editor system or fusion protein provided herein is a nucleic acid-programmed DNA binding protein (napDNAbp) and a nucleic acid sequence encoding ... The reverse transcriptase may include Mason-Pfizer Monkey Virus (MPMV) reverse transcriptase or a variant thereof, POK11ERV reverse transcriptase or a variant thereof, Simian Retrovirus Type 2 (SRV2) reverse transcriptase or a variant thereof, Woolly Monkey Sarcoma Virus (WMSV) reverse transcriptase or a variant thereof, Vp96 reverse transcriptase or a variant thereof, Vc95 reverse transcriptase or a variant thereof, Ec48 reverse transcriptase or a variant thereof, Gs reverse transcriptase or a variant thereof, Er reverse transcriptase or a variant thereof, Ne144 reverse transcriptase or a variant thereof, Tf1 reverse transcriptase or a variant thereof, or Rs09415 reverse transcriptase ("CRISPR-RT") or a variant thereof.
[0174] In various other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can include a variant RT sequence based on the MMLV RT wild type of SEQ ID NO: 33, and can include a variant of SEQ ID NOs: 172-177 or 183-184, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 172-177 or 183-184, and can include ... comprises at least one of residues 13I, 19I, 32T, 38V, 60Y, 111L, 120R, 126Y, 128N, 128F, 128H, 129S, 132S, 138R, 157F, 175Q, 175S, 200S, 200Y, 200N, 200C, 222F, 223A, 223M, 223T, 223W, 223Y, 234I, 246I, 249S, 287A, 292T, 302A, 302K, 306K, 316R, 346K, 373N, 388C, 402A, 445N, 457I, and 462S.
[0175] In still various other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on Ec48 RT, and can have a similar identity to variants of SEQ ID NOs: 188-195, 256, and 257, or any of SEQ ID NOs: 188-195, 256, and 257 at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. %, or up to 100% sequence identity, and the amino acid sequence includes at least one of residues 36V, 54K, 60K, 87E, 151T, 165D, 182N, 189N, 205K, 214L, 243N, 267I, 277F, 279K, 303M, 307R, 315K, 317S, 318E, 324Q, 326E, 328K, 343N, 372K, 378K, and 385.
[0176] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on Tf1 RT, and has a similar sequence identity to a variant of SEQ ID NOs: 196-213 and 251-255, or to any of SEQ ID NOs: 196-213 and 251-255 at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. %, or up to 100% sequence identity, and the amino acid sequence includes at least one of residues 14A, 22K, 64L, 64W, 70T, 72V, 102I, 106R, 118R, 133N, 139T, 158Q, 188K, 260L, 269L, 274R, 288Q, 293K, 297Q, 316Q, 321R, 356E, 363V, 413E, 423V, and 492N.
[0177] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on PERV RT, and can include a variant of SEQ ID NOs: 214-215 or 234-238, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 214-215 or 234-238, wherein the amino acid sequence comprises at least one of residues 199N, 305K, 312F, 329P, and 602W.
[0178] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on AVIRE RT wild type (SEQ ID NO: 216), and can include a variant of SEQ ID NOs: 217-221, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 217-221, wherein the amino acid sequence comprises at least one of residues 199N, 305K, 312F, 329P, and 604W.
[0179] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on KORV RT wild type (SEQ ID NO: 222), and can include a variant of SEQ ID NOs: 223-227, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 223-227, wherein the amino acid sequence comprises at least one of residues 197N, 303K, 310F, 327P, and 599W.
[0180] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can include a variant RT sequence based on WMSV RT wild type (SEQ ID NO: 228), and can include a variant of SEQ ID NOs: 229-233, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs: 229-233, wherein the amino acid sequence includes at least one of residues 197N, 303K, 311F, 327P, and 599W.
[0181] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on Ne144 RT wild type (SEQ ID NO: 239) and can include a variant of SEQ ID NO: 240, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO: 240, wherein the amino acid sequence comprises at least one of residues 157T, 165T, and 288V.
[0182] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise a variant RT sequence based on Vc95 RT wild type (SEQ ID NO: 241), and can include a variant of SEQ ID NO: 242, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NO: 242, wherein the amino acid sequence comprises at least one of residues 11M, 75A, 97M, 146D, and 245T.
[0183] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to a variant RT sequence based on Gs RT wild type (SEQ ID NO: 60), or to any of SEQ ID NOs: 159-171, wherein the amino acid sequence is and at least one of residues 12D, 16E, 16V, 17P, 20G, 37R, 37P, 38H, 40C, 41N, 41S, 45R, 67T, 67R, 72E, 73V, 78V, 93R, 123V, 126F, 129G, 162N, 190L, 206V, 233K, 234V, 263G, 264S, 267M, 279E, 287I, 291K, 309T, 344S, 358S, 360S, 363G, 374A, and 412H.
[0184] In yet other embodiments, the engineered RT domain of the prime editor system or fusion protein disclosed herein can include five mutant variant RT sequences based on AVIRE RT, KORV RT, and WMSV RT, and can include variants of SEQ ID NOs:243-245, or amino acid sequences having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to any of SEQ ID NOs:243-245, where AVIRE RT includes residues 199N, 305K, 312F, 329P, and 604W, KORV RT includes residues 197N, 303K, 310F, 327P, and 599W, and WMSV RT includes residues 197N, 303K, 311F, 327P, and 599W.
[0185] In yet other embodiments, the engineered RT domain of a prime editor system or fusion protein disclosed herein can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%, or up to 100% sequence identity to a variant RT sequence of Tfl-rat4 (SEQ ID NO:251), Tflevo3.1 (SEQ ID NO:252), Tflevo+rat-1 (SEQ ID NO:254), Tf1evo+rat2 (SEQ ID NO:255), Ec48-v2 (SEQ ID NO:256), Ec48-evo3 (SEQ ID NO:257), or any of SEQ ID NOs:251-257, provided that the sequence comprises at least one of the amino acid substitutions provided in this disclosure.
[0186] In other embodiments, the disclosure describes improved prime editors and prime editor systems, including prime editor fusion proteins including PEmax of SEQ ID NO: 2, which may be encoded by the nucleic acid sequence of SEQ ID NO: 1 and may be modified with any one of the variant Cas9 domains or variant RT domains disclosed herein. The disclosure also provides other improved prime editor variants, including fusion proteins of SEQ ID NOs: 2-8, as well as fusion proteins including evolved nucleic acid programmable DNA binding proteins of SEQ ID NOs: 9-32 and reverse transcriptases of SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241. The present disclosure also contemplates fusion proteins having an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to any one of SEQ ID NOs: 2 and 3-8. The present disclosure also contemplates evolved nucleic acid programmed DNA binding proteins having an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to any one of SEQ ID NOs: 9-32. Additionally, the present disclosure contemplates reverse transcriptases having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to SEQ ID NOs: 33-46, 48, 49, 51-53, 55-57, 59, 60, 63-78, 185, 216, 222, 228, 239, and 241.
[0187] In addition, the disclosure provides nucleic acid molecules encoding and / or expressing the evolved and / or modified prime editors described herein, as well as expression vectors or constructs for expressing the evolved and / or modified prime editors described herein, host cells comprising the nucleic acid molecules and expression vectors, and compositions for delivering and / or administering the nucleic acid-based embodiments described herein. In addition, the disclosure provides isolated evolved and / or modified prime editors as described herein, as well as compositions comprising the isolated evolved and / or modified prime editors. Still further, the disclosure provides methods for making evolved and / or modified prime editors, as well as methods for using evolved and / or modified prime editors or nucleic acid molecules encoding evolved and / or modified prime editors in applications including editing nucleic acid molecules, e.g. genomes, preferably in a sequence context-independent manner (i.e., the desired editing site does not require a specific sequence context) with improved efficiency compared to prime editors that form the state of the art. In embodiments, the method of making provided herein is an improved phage-assisted continuous evolution (PACE) system that may be utilized to evolve one or more components of a prime editor (e.g., a Cas9 domain or a reverse transcriptase domain). The present specification also provides a method for efficiently editing a single nucleic acid base of a target nucleic acid molecule, e.g., a genome, using the prime editing system described herein (e.g., in the form of an isolated evolved and / or engineered prime editor described herein or a vector or construct encoding the same), preferably in a sequence context-independent manner, to perform prime editing.Further still, the present specification provides therapeutic methods for treating a genetic disease and / or altering or changing a genetic trait or condition by contacting a target nucleic acid molecule, e.g., a genome, with a prime editing system (e.g., in the form of an isolated evolved and / or engineered prime editor protein or a vector encoding same) and performing prime editing to treat the genetic disease and / or change a genetic trait (e.g., eye color).
[0188] Thus, the present disclosure provides a method for editing a nucleic acid molecule by prime editing, which involves contacting the nucleic acid molecule with a modified prime editor and a pegRNA, thereby increasing editing efficiency and / or reducing indel formation to install one or more modifications to the nucleic acid molecule at the target site. The present disclosure further provides a polynucleotide for editing a DNA target site by prime editing, which comprises a nucleic acid sequence encoding a modified prime editor protein comprising a modified napDNAbp and / or polymerase domain, wherein the napDNAbp and polymerase domain are capable of installing one or more modifications at the DNA target site with increased editing efficiency and / or reduced indel formation in the presence of a pegRNA. The present disclosure further provides vectors, cells and kits comprising the compositions and polynucleotides of the present disclosure, as well as methods of making such vectors, cells and kits, and methods of delivering such compositions, polynucleotides, vectors, cells and kits to cells in vitro, ex vivo (e.g., during cell-based therapy to modify cells outside the body) and in vivo.
[0189] Modified Prime Editor The present disclosure provides modified prime editors and prime editor fusion proteins, such as, but not limited to, PEmax, and can further include variants of PEmax in which one or both of the napDNAbp domain and the RT domain are replaced with one of the engineered Cas9 or RT variants disclosed herein.
[0190] In one embodiment, the modified prime editor fusion protein is PEmax (SEQ ID NO:2) or an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least up to 100% sequence identity to SEQ ID NO:2. PEmax has the amino acid sequence of SEQ ID NO:2 and the nucleic acid sequence of SEQ ID NO:1.
[0191] PEmax (of SEQ ID NO:2) comprises, from N-terminus to C-terminus, (a) a bipartite SV40 NLS domain (SEQ ID NO:95), (b) a SpCas9 based on wild-type SpCas9 of SEQ ID NO:10 with amino acid substitutions R221K, N394K, and H840A relative to said sequence, (c) a linker sequence, (d) a GenScript codon-optimized MMLV RT pentamutant based on wild-type MMLV RT of SEQ ID NO:33 with amino acid substitutions D200N T306K W313F T330P L603W relative to said sequence, (e) a linker, (f) a bipartite SV40 NLS domain, (g) a linker, and (h) a c-Myc NLS domain. These amino acid sequences are provided as follows: Sequence of the PEmax component of SEQ ID NO:2: ⇒Double-jointed SV40 NLS MKRTADGSEFESPKKKRKV (SEQ ID NO: 95) ⇒SpCas9 R221K N394K H840A ⇒Linker=(SGGSx2-Bipartite SV40 NLS-SGGSx2): SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS (SEQ ID NO:79) ⇒GenScript codon-optimized MMLV RT five mutants (D200N T306K W313F T330P L603W) (SEQ ID NO:34) ⇒Other linker sequences SGGS (SEQ ID NO:81) ⇒Double-jointed SV40 NLS KRTADGSEFESPKKKRKV (SEQ ID NO:97) ⇒Other linker sequences GSG (SEQ ID NO:82) ⇒c-Myc NLS PAAKRVKLD (SEQ ID NO:98)
[0192] Prime editors contemplated herein include, in some embodiments, systems in which a nucleic acid programmable DNA binding protein (napDNAbp) and a reverse transcriptase domain (RT) are provided in trans such that they can be separately localized and / or targeted to the desired DNA editing site to perform their prime editing functions. In other embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) and the reverse transcriptase domain (RT) are provided as a fusion protein.
[0193] In those embodiments in which the nucleic acid programmable DNA binding protein (napDNAbp) and the reverse transcriptase domain (RT) are provided in the form of a fusion protein, the modified prime editor disclosed herein may comprise any suitable structural configuration. For example, the fusion protein may comprise a napDNAbp fused to a polymerase (e.g., a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase, e.g., a reverse transcriptase) from the N-terminus to the C-terminus. In other embodiments, the fusion protein may comprise a polymerase (e.g., a reverse transcriptase) fused to the napDNAbp from the N-terminus to the C-terminus. The fused domains may optionally be joined by a linker, e.g., an amino acid sequence. In other embodiments, the fusion protein has the structure NH 2 -[napDNAbp]-[polymerase]-COOH; or NH 2 -[polymerase]-[napDNAbp]-COOH, with each instance of "]-[" indicating the presence of an optional linker sequence. In embodiments where the polymerase is a reverse transcriptase, the fusion protein has the structure NH 2 -[napDNAbp]-[RT]-COOH; or NH 2It may include -[RT]-[napDNAbp]-COOH, where each instance of "]-[" indicates the presence of an optional linker sequence.
[0194] In various embodiments, the modified prime editor may be based on PE1, and one or more components of PE1 are replaced with variant domains. For example, the PE1 SpCas9 domain may be replaced with a modified SpCas9 domain. Or, the RT domain may be replaced with a modified RT domain (e.g., a codon-optimized variant).
[0195] PE1 encompasses the H840A mutation (i.e., Cas9 nickase) and M-MLV RT wild type, as well as Cas9 variants that contain an N-terminal NLS sequence (19 amino acids) and an amino acid linker (32 amino acids) that connects the C-terminus of the Cas9 nickase domain to the N-terminus of the RT domain. The PE1 fusion protein has the following structure: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(wt)]. The amino acid sequences of PE1 and its individual components are as follows: [Table 2-1] [Table 2-2] [Table 2-3]
[0196] In various other embodiments, the modified prime editor protein can be based on PE2, and one or more components of PE2 can be replaced with a variant domain. For example, the PE2 SpCas9 domain can be replaced with a modified SpCas9 domain. Or, the RT domain of PE2 can be replaced with a modified RT domain (e.g., a codon-optimized variant).
[0197] PE2 encompasses a Cas9 variant containing the H840A mutation (i.e., Cas9 nickase) and M-MLV RT containing the mutations D200N, T330P, L603W, T306K, and W313F, as well as an N-terminal NLS sequence (19 amino acids) and an amino acid linker (33 amino acids) that joins the C-terminus of the Cas9 nickase domain to the N-terminus of the RT domain. The PE2 fusion protein has the following structure: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]. The amino acid sequence of PE2 is as follows: [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4]
[0198] In yet other embodiments, modified prime editor proteins disclosed herein may be based on other prime editor protein sequences, where one or more components of such fusions are replaced with variant domains. Such starting point prime editor proteins may include: [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4]
[0199] In yet another embodiment, the prime editor used in the present disclosure may comprise PEmax, which is a complex comprising Cas9(R221K N39K H840A) and variant MMLV RT pentamutant (D200N T306K W313F T330P L603W) having the following structure: [Bipartite NLS]-[Cas9(R221K)(N394K)(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)]-[Bipartite NLS]-[NLS]+desired PEgRNA, where the PE fusion has the amino acid sequence of SEQ ID NO:2, as shown below. [Table 5-1] [Table 5-2]
[0200] In various embodiments, a prime editor protein utilized in the methods and compositions contemplated herein may include variants of any of the sequences disclosed above having an amino acid sequence that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a prime editor sequence disclosed herein.
[0201] napDNAbp domain and its engineered variants In various embodiments, the modified prime editor proteins disclosed herein, including PEmax, comprise a nucleic acid programmable DNA binding protein (napDNAbp).
[0202] In various embodiments, the modified prime editor protein may comprise a napDNAbp domain having a wild-type Cas9 sequence, including, for example, the canonical Streptococcus pyogenes Cas9 sequence of SEQ ID NO:9.
[0203] In other embodiments, the modified prime editor protein may comprise a napDNAbp domain with a modified Cas9 sequence that includes a nickase variant of Streptococcus pyogenes Cas9 (of SEQ ID NO: 9) of SEQ ID NO: 12, which has an H840A substitution relative to wild-type SpCas9, for example, as shown below: [Table 6]
[0204] In one embodiment, the napDNAbp component or domain of the prime editor designated "PEmax" comprises the following amino acid sequence based on the canonical SpCas9 amino acid sequence of SEQ ID NO: 9 with the following substitutions: R221K, N394K, and H840A. SpCas9 R221K N394K H840A:
[0205] The modified prime editor protein may further comprise one or more mutations in the napDNAbp (e.g., Cas9) domain that result in improved editing efficiency. For example, the present disclosure describes the development of improved prime editor proteins using PACE. In some embodiments, the prime editor (e.g., a fusion protein, or a prime editor in which napDNAbp and reverse transcriptase are provided in trans) may comprise one or more of the following mutations: D23G, H99Q, H99R, E102K, E102S, E102R, N175K, D177G, K218R, N309D, I312V, E471K, G485S, K562N, D608N, D706N, D705N, D707N, D708N, D709N, D710N, D711N, D712N, D71 ... , I632V, D645N, D645E, R654C, G687D, G715E, H721Y, R753K, R753G, H754R, K775R, E790K, T804A, K918A, K1003R, M1021Y, E1071K, and E1260D. In some embodiments, such Cas9 variants comprise a single mutation, the single mutation being selected from D23G, H99Q, H99R, E102K, E102S, E102R, N175K, D177G, K218R, N309D, I312V, E471K, G485S, K562N, D608N, I632V, D645N, D645E, R654C, G687D, G715E, H721Y, R753K, R753G, H754R, K775R, E790K, T804A, K918A, K1003R, M1021Y, E1071K, and E1260D. In some embodiments, the Cas9 variant comprises a R753G mutation. In certain embodiments, the Cas9 variant comprises an H721Y mutation and an R753G mutation; an E102K mutation and an R753G mutation; or an E102K mutation, an H721Y mutation, and an R753G mutation. In certain embodiments, the Cas9 variant comprises the amino acid sequence of any one of SEQ ID NOs: 178-180.
[0206] In some embodiments, an improved prime editor protein used in the compositions and methods described herein comprises a mutation at position R753X, where X is any amino acid, relative to the amino acid sequence of a wild-type Cas9 from Streptococcus pyogenes: [Table 7]
[0207] In some embodiments, the R753X mutation is an R753G mutation. [Table 8]
[0208] Improved prime editor proteins utilized in the methods and compositions described herein may comprise any of the modified Cas9 sequences above, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95% or at least 99% sequence identity thereto, provided that the variant comprises one of the amino acid substitutions provided herein. The proteins described herein may also include any Cas9 protein (including, by way of example, those described below) that comprises a mutation corresponding to R753X or R753G at the relevant position of the amino acid sequence.
[0209] The present disclosure contemplates modification of any Cas9 protein known in the art with one or more of the mutations described herein (i.e., R221K, N394K, R753G and / or H840A), and combinations of any modified Cas9 protein with one or more of the PEmax construct features described herein (e.g., optimized MMLV RT pentamutant, NLS, linkers, etc.).
[0210] In some embodiments, the improved prime editor proteins described herein include any of the following other wild-type SpCas9 sequences, which may be modified by one or more mutations described herein at the corresponding amino acid positions: [Table 9-1] [Table 9-2] [Table 9-3] [Table 9-4] [Table 9-5] [Table 9-6] [Table 9-7] [Table 9-8]
[0211] The improved prime editor proteins utilized in the methods and compositions described herein may include any of the above SpCas9 sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0212] In other embodiments, the Cas9 protein can be a wild-type Cas9 ortholog from another bacterial species that differs from the canonical Cas9 from S. pyogenes. For example, modified versions of the following Cas9 orthologs can be used in conjunction with the PEmax constructs utilized in the methods and compositions described herein by making mutations at positions corresponding to R 221 K, N 394 K, R 753 G, and / or H 840 A of wild-type SpCas9. In addition, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the orthologs below may also be used with a prime editor. [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4] [Table 10-5] [Table 10-6] [Table 10-7] [Table 10-8] [Table 10-9] [Table 10-10]
[0213] Prime editors utilized in the methods and compositions described herein may include any of the above Cas9 orthologous sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0214] The napDNAbp used in the PEmax constructs described herein may include any suitable homologue and / or orthologue or naturally occurring enzyme, such as Cas9. Cas9 homologues and / or orthologues have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. The Cas moiety may be configured as a nickase (e.g., mutagenized, recombinantly engineered, or otherwise derived from nature), i.e., capable of cleaving only one strand of the target double-stranded DNA. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5,726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain; i.e., Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any one of the Cas9 orthologs in the table above.
[0215] The present disclosure also contemplates the inclusion of the following additional napDNAbp in the prime editors provided herein. Any suitable napDNAbp may be used in the prime editors utilized in the methods and compositions described herein. In various embodiments, the napDNAbp may be any class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a tool for genome editing, the nomenclature used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologues, has been constantly developed. The present application refers to CRISPR-Cas enzymes with nomenclature that may be old and / or new. Those skilled in the art can identify the specific CRISPR-Cas enzymes referred to in the present application based on the nomenclature used, whether it is old (i.e., "legacy") nomenclature or new nomenclature. CRISPR-Cas nomenclature is discussed extensively in Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?," The CRISPR Journal, Vol. 1. No. 5, 2018, the entire contents of which are incorporated herein by reference. The particular CRISPR-Cas nomenclature used in any given instance of this application is in no way limiting, and one of skill in the art will be able to identify which CRISPR-Cas enzyme is being referenced.
[0216] For example, the following Class 2 CRISPR-Cas enzymes, Types II, V, and VI, have the following old (i.e., legacy) and new art-recognized names: Each of these enzymes and / or variants thereof may be used with the prime editors utilized in the methods and compositions described herein. [Table 11]
[0217] The following description of various napDNAbps that may be used in conjunction with the prime editors utilized in the methods and compositions of the present disclosure is not meant to be limiting in any way. The prime editor may include canonical SpCas9 or any orthologous Cas9 protein or any variant Cas9 protein that is known or can be created or evolved by directed evolution or otherwise mutagenesis processes, including any naturally occurring variant, mutant, or otherwise engineered version of Cas9. In various embodiments, the Cas9 or Cas9 variant has nickase activity, i.e., cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has an inactive nuclease, i.e., is an "inactive" Cas9 protein. Other variant Cas9 proteins that may be used are those that have a smaller molecular weight than canonical SpCas9 (e.g., for easier delivery) or have a modified or rearranged primary amino acid structure (e.g., a circular permutation format).
[0218] The prime editors utilized in the methods and compositions described herein may also include Cas9 equivalents, including Cas12a (Cpf1) and Cas12b1 proteins that are the result of convergent evolution. napDNAbps (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) used herein may contain various modifications that alter / enhance their PAM specificity. Finally, the present application contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a reference SpCas9 canonical sequence or a reference Cas9 equivalent (e.g., Cas12a (Cpf1)).
[0219] In some embodiments, the napDNAbp directs cleavage of one or both strands at the location of the target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a napDNAbp that is mutated with respect to the corresponding wild-type enzyme, such that the mutated napDNAbp lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. For example, an aspartic acid to alanine substitution (D10A) on the RuvC I catalytic domain of Cas9 from S.pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (that cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, but are not limited to, H840A, N854A, and N863A, with reference to the canonical SpCas9 sequence or the equivalent amino acid location in other Cas9 variants or Cas9 equivalents.
[0220] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of the essential basic functions required for the methods of the disclosure, namely, (i) possession of nucleic acid-programmed binding of the Cas protein to target DNA and (ii) the ability to nick the target DNA sequence on one strand. Cas proteins contemplated herein include CRISPR Cas9 proteins, and Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)), homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and may include Cas9 equivalents from any Class 2 CRISPR system (e.g., types II, V, VI), including Cas12a (Cpf1), Cas12e (CasX), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9, C2c10, C2c11, C2c12, C2c13, C2c14, C2c15, C2c16, C2c17, C2c18, C2c19, C2c20, C2c210, C2c22, C2c23, C2c24, C2c25, C2c26, C2c27, C2c28, C2c29, C2c30, C2c31, C2c32, C2c33, C2c34, C2c35, C2c36, C2c37, C2c38, C2c39, C2c40, C2c41, C2c42, C2c43, C2c44, C2c45, C2c46, C2c47, C2c48, C2c49 ... Cas13a (C2c2), Cas13d, Cas13c (C2c7), Cas13b (C2c6), and Cas13b. Additional Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016;353(6299) and Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?," The CRISPR Journal, Vol. 1. No. 5, 2018, the contents of which are incorporated herein by reference.
[0221] The term "Cas9" or "Cas9 nuclease" or "Cas9 portion" or "Cas9 domain" encompasses any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered Cas9. The term Cas9 is not meant to be particularly limiting and may be referred to as "Cas9 or equivalent." Exemplary Cas9 proteins are further described herein and / or in the art and are incorporated herein by reference. The present disclosure is not limited with respect to the particular Cas9 employed in the prime editor utilized in the methods and compositions described herein.
[0222] For the latest updates, please contact Cas9 Thanksgiving is a great place to stay ( These and all of the rest of the books are listed as “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc.Natl.Acad.Sci.USA98:4658-4663(2001); CM、Gonzales K.、Chao Y.、Pirzada ZA、Eckert MR、Vogel J.、Charpentier E.,Nature 471:602-607(2011); K.,Fonfara I.,Hauer M.,Doudna JA,Charpentier E.Science 337:816-821(2012).
[0223] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not meant to be limiting. Prime editors utilized in the methods and compositions of the disclosure may use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent.
[0224] A. Wild-type canonical SpCas9 In one embodiment, the prime editor construct utilized in the methods and compositions described herein may comprise the "canonical SpCas9" nuclease from S.pyogenes, which is widely used as a tool for genome engineering and is classified as a type II subgroup of enzymes of class 2 CRISPR-Cas systems. The Cas9 protein is a large multi-domain protein that contains two separate nuclease domains. Point mutations to abolish one or both nuclease activities can be introduced on Cas9, resulting in a nickase Cas9 (nCas9) or an inactive Cas9 (dCas9), respectively, which still retains its ability to bind to DNA in a manner programmed by sgRNA. In principle, when fused to another protein or domain, Cas9 or its variants (e.g., nCas9) can target the protein to virtually any DNA sequence simply by co-expression with the appropriate sgRNA. As used herein, the canonical SpCas9 protein refers to the wild-type protein from Streptococcus pyogenes having the following amino acid sequence: [Table 12-1] [Table 12-2]
[0225] Prime editors utilized in the methods and compositions described herein may include canonical SpCas9 or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the wild-type Cas9 sequence provided above. These variants may include SpCas9 variants containing one or more mutations, and may include any of the known mutations reported in SwissProt Accession No. Q99ZW2 (SEQ ID NO: 9) entry, including the following: [Table 13]
[0226] B. Wild-type Cas9 orthologs In other embodiments, the Cas9 protein may be a wild-type Cas9 ortholog from another bacterial species that differs from the canonical Cas9 from S. pyogenes. For example, the following Cas9 orthologs may be used in conjunction with the prime editor constructs utilized in the methods and compositions described herein. In addition, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the orthologs below may also be used with the prime editor. [Table 14-1] [Table 14-2] [Table 14-3] [Table 14-4] [Table 14-5] [Table 14-6]
[0227] Prime editors utilized in the methods and compositions described herein may include any of the above Cas9 orthologous sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0228] The napDNAbp may include any suitable homologue and / or orthologue or naturally occurring enzyme, such as Cas9. Cas9 homologues and / or orthologues have been described in various species, including but not limited to S. pyogenes and S. thermophilus. Preferably, the Cas moiety is configured as a nickase (e.g., mutagenized, recombinantly engineered, or otherwise derived from nature), i.e., capable of cleaving only one strand of the target double-stranded DNA. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5,726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 3. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any one of the Cas9 orthologs in the table above.
[0229] C. Inactive Cas9 variants In certain embodiments, the prime editor utilized in the methods and compositions described herein may include an inactive Cas9, e.g., an inactive SpCas9, which does not have nuclease activity due to one or more mutations that inactivate both nuclease domains of Cas9, i.e., the RuvC domain (which cleaves the non-protospacer DNA strand) and the HNH domain (which cleaves the protospacer DNA strand). The nuclease inactivation may be due to one or more mutations that result in one or more substitutions and / or deletions in the amino acid sequence of the encoded protein or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0230] As used herein, the term "dCas9" refers to nuclease-inactive Cas9 or nuclease-inactive Cas9, or a functional fragment thereof, including any naturally occurring dCas9 from any organism, any equivalent or functional fragment of naturally occurring dCas9, any engineered dCas9 variant or functional fragment thereof, any dCas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of naturally occurring or engineered dCas9. The term dCas9 is not meant to be particularly limiting and may be referred to as "dCas9 or equivalent." Exemplary dCas9 proteins and methods for making dCas9 proteins are further described herein and / or in the art and are incorporated herein by reference.
[0231] In other embodiments, dCas9 corresponds to or includes a Cas9 amino acid sequence with one or more mutations that partially or entirely inactivate Cas9 nuclease activity. In other embodiments, Cas9 variants with mutations other than D10A and H840A are provided, which may result in full or partial inactivation of endogenous Cas9 nuclease activity (e.g., nCas9 or dCas9, respectively). Such mutations include, for example, other amino acid substitutions in D10 and H840, or other substitutions in the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain), with reference to the wild-type sequence, such as Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1). In some embodiments, variants or homologs of Cas9 (e.g., variants of Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 16)) are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to NCBI Reference Sequence: NC_017053.1. In some embodiments, variants of dCas9 (e.g., variants of NCBI reference sequence: NC_017053.1 (SEQ ID NO: 16)) are provided that have an amino acid sequence that is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more, shorter or longer than NC_017053.1 (SEQ ID NO: 16).
[0232] In one embodiment, the inactive Cas9 may be based on the canonical SpCas9 sequence of Q99ZW2 and may have the following sequence including D10X and H810X substitutions (underlined and bold), where X may be any amino acid, or the variant may be a variant of SEQ ID NO: 260 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0233] In one embodiment, the inactive Cas9 may be based on the canonical SpCas9 sequence of Q99ZW2 and may have the following sequence including the D10A and H810A substitutions (underlined and bold), or may be a variant of SEQ ID NO: 261 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. [Table 15]
[0234] D. Cas9 Nickase Variants In one embodiment, the prime editor utilized in the methods and compositions described herein comprises Cas9 nickase. The term "Cas9 nickase" or "nCas9" refers to a variant of Cas9 that can introduce single-stranded breaks on double-stranded DNA molecular targets. In some embodiments, Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., canonical SpCas9) comprises two separate nuclease domains, the RuvC domain (which cleaves non-protospacer DNA strands) and the HNH domain (which cleaves protospacer DNA strands). In one embodiment, Cas9 nickase comprises a mutation in the RuvC domain that inactivates RuvC nuclease activity. For example, mutations of aspartic acid (D)10, histidine (H)983, aspartic acid (D)986, or glutamic acid (E)762 have been reported as loss-of-function mutations in the RuvC nuclease domain and creating functional Cas9 nickases (see, e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, incorporated herein by reference). Thus, nickase mutations in the RuvC domain can include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In certain embodiments, the nickase can be D10A, H983A, D986A, or E762A, or a combination thereof.
[0235] In various embodiments, the Cas9 nickase may have a mutation in the RuvC nuclease domain and may have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. [Table 16-1] [Table 16-2] [Table 16-3] [Table 16-4] [Table 16-5]
[0236] In another embodiment, Cas9 nickase comprises a mutation in HNH domain that inactivates HNH nuclease activity. For example, mutation of histidine (H)840 or asparagine (R)863 has been reported as a loss-of-function mutation in HNH nuclease domain and creates functional Cas9 nickase (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, which is incorporated herein by reference). Thus, the nickase mutation in HNH domain can include H840X and R863X, where X is any amino acid other than wild-type amino acid. In certain embodiments, nickase can be H840A or R863A or a combination thereof.
[0237] In various embodiments, the Cas9 nickase may have a mutation in the HNH nuclease domain and may have one of the following amino acid sequences, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. [Table 17-1] [Table 17-2] [Table 17-3]
[0238] In some embodiments, the N-terminal methionine is removed from the Cas9 nickase or from any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, methionine-minus Cas9 nickases include the following sequences, or variants thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table 18-1] [Table 18-2] [Table 18-3]
[0239] E. Other Cas9 variants In addition to inactive Cas9 and Cas9 nickase variants, the Cas9 proteins used herein may also include other "Cas9 variants" that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 protein, including any wild-type or mutant Cas9 (e.g., an inactive Cas9 or Cas9 nickase) disclosed herein or known in the art, or a fragment Cas9, or a circularly permuted Cas9, or other variant of Cas9. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a reference Cas9. In some embodiments, the Cas9 variants comprise a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 (e.g., SEQ ID NO: 9).
[0240] In some embodiments, the present disclosure may utilize a Cas9 fragment that is a fragment of any Cas9 protein disclosed herein and retains their functionality. In some embodiments, the Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0241] In various embodiments, a prime editor utilized in the methods and compositions disclosed herein may comprise one of the Cas9 variants described as follows, or a Cas9 variant thereof that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 variant.
[0242] F. Small Cas9 variants In some embodiments, the prime editor utilized in the methods and compositions contemplated herein may comprise a Cas9 protein that is smaller in molecular weight than the canonical SpCas9 sequence. In some embodiments, the smaller size of the Cas9 variant may facilitate delivery to cells, for example, by expression vectors, nanoparticles, or other means of delivery. In certain embodiments, the smaller size of the Cas9 variant may comprise an enzyme that is classified as the type II enzyme of class 2 CRISPR-Cas system. In some embodiments, the smaller size of the Cas9 variant may comprise an enzyme that is classified as the type V enzyme of class 2 CRISPR-Cas system. In other embodiments, the smaller size of the Cas9 variant may comprise an enzyme that is classified as the type VI enzyme of class 2 CRISPR-Cas system.
[0243] The canonical SpCas9 protein is 1368 amino acids in length with a predicted molecular weight of 158 kilodaltons. The term "small Cas9 variant" as used herein refers to a Cas9 variant that is at least 1300 amino acids long, or at least 1290 amino acids long, or at least 1280 amino acids long, or at least 1270 amino acids long, or at least 1260 amino acids long, or at least 1250 amino acids long, or at least 1240 amino acids long, or at least 1230 amino acids long, or at least 1220 amino acids long, or at least 1210 amino acids long, or at least 1200 amino acids long, or at least 1190 amino acids long, or at least 1180 amino acids long, or at least 1170 amino acids long, or at least 1160 amino acids long, or at least 1150 amino acids long, or at least 1140 amino acids long, or at least 1130 amino acids long. It refers to Cas9 variants, either naturally occurring, engineered or otherwise, that are less than 1120 amino acids, or less than 1110 amino acids, or less than 1100 amino acids, or less than 1050 amino acids, or less than 1000 amino acids, or less than 950 amino acids, or less than 900 amino acids, or less than 850 amino acids, or less than 800 amino acids, or less than 750 amino acids, or less than 700 amino acids, or less than 650 amino acids, or less than 600 amino acids, or less than 550 amino acids, or less than 500 amino acids, but at least more than 400 amino acids, and that retain the necessary functions of the Cas9 protein. Cas9 variants can include those categorized as Type II, Type V, or Type VI enzymes of Class 2 CRISPR-Cas systems.
[0244] In various embodiments, the prime editor utilized in the methods and compositions disclosed herein may comprise one of the small Cas9 variants described below, or Cas9 variants thereof that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference small Cas9 protein. [Table 19-1] [Table 19-2] [Table 19-3]
[0245] G. Cas9 equivalents In some embodiments, the prime editor utilized in the methods and compositions described herein may include any Cas9 equivalent. As used herein, the term "Cas9 equivalent" is a broad term that encompasses any napDNAbp protein that provides the same function as Cas9 in the prime editor, even though its amino acid primary sequence and / or its three-dimensional structure may differ and / or may be unrelated from an evolutionary standpoint. Thus, Cas9 equivalents encompass any Cas9 ortholog, homolog, mutant, or variant described or encompassed herein that is evolutionarily related, but Cas9 equivalents also encompass proteins that may have evolved by convergent evolution processes to have the same or similar function as Cas9, but do not necessarily have any similarity in amino acid sequence and / or three-dimensional structure. Although Cas9 equivalents may be based on proteins that have arisen by convergent evolution, the prime editor utilized in the methods and compositions described herein encompasses any Cas9 equivalent that will provide the same or similar function as Cas9. For example, where Cas9 refers to the Type II enzyme of a CRISPR-Cas system, a Cas9 equivalent could refer to a Type V or Type VI enzyme of a CRISPR-Cas system.
[0246] For example, Cas12e (CasX) is a Cas9 equivalent that reportedly has the same function as Cas9 but evolved by convergent evolution. Thus, the Cas12e (CasX) protein described in Liu et al., "CasX enzymes comprises a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566:218-223, is contemplated for use with the prime editors utilized in the methods and compositions described herein. Additionally, any variants or modifications of Cas12e (CasX) are envisioned and are within the scope of the present disclosure.
[0247] Cas9 is a bacterial enzyme that has evolved in a wide variety of species. However, the Cas9 equivalents contemplated herein can also be obtained from Archaea, which constitute a domain and kingdom of unicellular prokaryotic microorganisms distinct from bacteria. In some embodiments, Cas9 equivalents can refer to Cas12e (CasX) or Cas12d (CasY), as described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Using genome-resolution metagenomics, several CRISPR-Cas systems have been identified, including the first reported Cas9 in the Archaea domain of life. This diverse Cas9 protein was found as part of an active CRISPR-Cas system from the little-studied nanoarchaea. In bacteria, two previously unknown systems, CRISPR-Cas12e and CRISPR-Cas12d, have been discovered, which are the most compact systems discovered to date. In some embodiments, Cas9 refers to Cas12e or a variant of Cas12e. In some embodiments, Cas9 refers to Cas12d or a variant of Cas12d. It should be recognized that other RNA-guided DNA binding proteins may be used as nucleic acid programmed DNA binding proteins (napDNAbp) and are within the scope of this disclosure. See also Liu et al., "CasX enzymes comprises a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566: 218-223. Any of these Cas9 equivalents are contemplated.
[0248] In some embodiments, the Cas9 equivalent comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp is a naturally occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a wild-type Cas portion or any Cas portion provided herein.
[0249] In various embodiments, nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12e (CasX), Cas12d (CasY), Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), Argonaute, and Cas12b1. One example of a nucleic acid programmable DNA binding protein with PAM specificity different from Cas9 is clustered regularly interspaced short palindromic repeats 1 from Prevotella and Francisella (i.e., Cas12a (Cpf1)). Like Cas9, Cas12a (Cpf1) is also a class 2 CRISPR effector, but it is a type V subgroup of enzymes rather than a type II subgroup. Cas12a (Cpf1) has been shown to mediate robust DNA interference with distinct characteristics from Cas9. Cas12a (Cpf1) is a single RNA-guided endonuclease lacking tracrRNA, utilizing T-rich protospacer adjacent motifs (TTN, TTTN, or YTN). Moreover, Cpf1 cleaves DNA by staggered DNA double-strand breaks. Of the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae are shown to have efficient genome editing activity in human cells. Cpf1 protein is known in the art and has been previously described, for example, in Yamano et al., "Crystal structure of Cpf1 in complex with guide RNA and target DNA." Cell (165) 2016, p. 949-962, the entire contents of which are incorporated herein by reference.
[0250] In yet other embodiments, the Cas protein may include any CRISPR associated protein, including, but not limited to, Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm 6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, preferably comprising a nickase mutation (e.g., a mutation corresponding to the D10A mutation in the wild-type Cas9 polypeptide of SEQ ID NO:9).
[0251] In various other embodiments, the napDNAbp can be any of the following proteins: Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, circularly permuted Cas9, or an Argonaute (Ago) domain, or a variant thereof.
[0252] Exemplary Cas9 equivalent protein sequences can include: [Table 20-1] [Table 20-2] [Table 20-3] [Table 20-4]
Table 20-5
[0253]
[0254] The prime editors utilized in the methods and compositions described herein may also include Cas12a(Cpf1)(dCpf1) variants that may be used as programmed DNA binding protein domains by guide nucleotide sequences. Cas12a(Cpf1) protein has a RuvC-like endonuclease domain that is similar to the RuvC domain of Cas9, but does not have the HNH endonuclease domain, and the N-terminus of Cas12a(Cpf1) does not have the recognition lobe of the alpha-helix system of Cas9. In Zetsche et al., Cell, 163, 759-771,2015 (which is incorporated herein by reference), it was shown that the RuvC-like domain of Cas12a(Cpf1) is responsible for cleaving both DNA strands, and inactivation of the RuvC-like domain inactivates Cas12a(Cpf1) nuclease activity. In some embodiments, napDNAbp is the single effector of microbial CRISPR-Cas system. The single effector of microbial CRISPR-Cas system includes, but is not limited to, Cas9, Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2) and Cas12c (C2c3). Typically, microbial CRISPR-Cas system is divided into class 1 and class 2 system. Class 1 system has multi-subunit effector complex, while class 2 system has single protein effector. For example, Cas9 and Cas12a (Cpf1) are class 2 effectors. In addition to Cas9 and Cas12a (Cpf1), three distinct Class 2 CRISPR-Cas systems (Cas12b1, Cas13a, and Cas12c) have been described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems", Mol. Cell, 2015 Nov 5;60(3):385-397, the entire contents of which are incorporated herein by reference.
[0255] The effectors of two of the systems, Cas12b1 and Cas12c, contain a RuvC-like endonuclease domain related to Cas12a. The third system, Cas13a, contains an effector with two predicted HEPN RNase domains. Unlike the production of CRISPR RNA by Cas12b1, the production of mature CRISPR RNA is tracrRNA-independent. Cas12b1 depends on both CRISPR RNA and tracrRNA for DNA cleavage. Bacterial Cas13a has been shown to possess an intrinsic RNase activity for CRISPR RNA maturation that is separate from its RNA-activated single-stranded RNA degradation activity. These RNase functions are distinct from each other and from the CRISPR RNA processing behavior of Cas12a. See, for example, East-Seletsky, et al., "Two distinct RNase activities of CRISPR-Cas13a enable guide-RNA processing and RNA detection", Nature, 2016 Oct 13;538(7624):270-273, the entire contents of which are incorporated herein by reference. In vitro biochemical analysis of Leptotrichia shahii Cas13a shows that Cas13a can be programmed to be guided by a single CRISPR RNA and cleave ssRNA targets with complementary protospacers. Two conserved catalytic residues on the HEPN domain mediate cleavage. Mutation of the catalytic residues generates a catalytically inactive RNA-binding protein. See, e.g., Abudayyeh et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector”, Science, 2016 Aug 5;353(6299), the entire contents of which are incorporated herein by reference.
[0256] The crystal structure of Alicyclobacillus acidoterrastris Cas12b1 (AacC2c1) has been reported in complex with a chimeric single-molecule guide RNA (sgRNA). See, e.g., Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism", Mol. Cell, 2017 Jan 19; 65(2): 310-322, the entire contents of which are incorporated herein by reference. A crystal structure has also been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease", Cell, 2016 Dec 15; 167(7): 1814-1828, the entire contents of which are incorporated herein by reference. The catalytically competent conformation of AacC2c1 is captured by both target and non-target DNA strands, which are independently arranged in a single RuvC catalytic pocket, and the cleavage mediated by C2c1 results in staggered 7-nucleotide cleavage of target DNA. The structural comparison between the C2c1 ternary complex and the previously identified Cas9 and Cpf1 counterparts demonstrates the diversity of the mechanism used by the CRISPR-Cas9 system. In some embodiments, napDNAbp can be C2c1, C2c2, or C2c3 protein. In some embodiments, napDNAbp is C2c1 protein. In some embodiments, napDNAbp is Cas13a protein. In some embodiments, napDNAbp is Cas12c protein.In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12b1 (C2c1), Cas13a (C2c2), or Cas12c (C2c3) protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b1 (C2c1), Cas13a (C2c2), or Cas12c (C2c3) protein.
[0257] H. Cas9 circular permutation In various embodiments, the prime editor utilized in the methods and compositions disclosed herein may comprise a circular permutation of Cas9.
[0258] The term "circularly permuted Cas9" or "circular permutants" of Cas9 or "CP-Cas9" refers to any Cas9 protein or variant thereof that is generated or modified or engineered as a circular permutation variant, meaning that the N-terminus and C-terminus of the Cas9 protein (e.g., wild-type Cas9 protein) are locally rearranged. Such circularly permuted Cas9 proteins or variants thereof retain the ability to bind to DNA when complexed with a guide RNA (gRNA). See Oakes et al., "Protein Engineering of Cas9 for enhanced function," Methods Enzymol, 2014, 546:491-511 and Oakes et al., "CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification," Cell, January 10, 2019, 176:254-267, each of which is incorporated herein by reference. The present disclosure contemplates any previously known CP-Cas9 or employs a new CP-Cas9, so long as the resulting circularly permuted protein retains the ability to bind DNA when complexed with a guide RNA (gRNA). Any of the Cas9 proteins described herein, including any variants, orthologs, or any engineered or naturally occurring Cas9, or equivalents thereof, may be reconstituted as circularly permuted variants. In various embodiments, the circular permutation of Cas9 may have the following structure: N-terminus-[original C-terminus]-[optional linker]-[original N-terminus]-C-terminus.
[0259] As an example, the present disclosure contemplates the following circular permutations of canonical S. pyogenes Cas9 (1368 amino acids of UniProtKB-Q99ZW2 (CAS9_STRP1) (numbering is based on the location of the amino acids in SEQ ID NO: 9)): N-terminus-[1268-1368]-[optional linker]-[1-1267]-C-terminus; N-terminus-[1168-1368]-[any]-[1-1167]-C-terminus; N-terminus-[1068-1368]-[optional linker]-[1-1067]-C-terminus; N-terminus-[968-1368]-[optional linker]-[1-967]-C-terminus; N-terminus-[868-1368]-[optional linker]-[1-867]-C-terminus; N-terminus-[768-1368]-[optional linker]-[1-767]-C-terminus; N-terminus-[668-1368]-[optional linker]-[1-667]-C-terminus; N-terminus-[568-1368]-[optional linker]-[1-567]-C-terminus; N-terminus-[468-1368]-[optional linker]-[1-467]-C-terminus; N-terminus-[368-1368]-[optional linker]-[1-367]-C-terminus; N-terminus-[268-1368]-[optional linker]-[1-267]-C-terminus; N-terminus-[168-1368]-[optional linker]-[1-167]-C-terminus; N-terminus-[68-1368]-[optional linker]-[1-67]-C-terminus; or N-terminus-[10-1368]-[optional linker]-[1-9]-C-terminus, or the corresponding circular permutations of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0260] In certain embodiments, the circularly permuted Cas9 has the following structure (based on the 1368 amino acids of S. pyogenes Cas9 (UniProtKB-Q99ZW2 (CAS9_STRP1)) (numbering is based on the location of the amino acids in SEQ ID NO: 9): N-terminus-[102-1368]-[optional linker]-[1-101]-C-terminus; N-terminus-[1028-1368]-[optional linker]-[1-1027]-C-terminus; N-terminus-[1041-1368]-[optional linker]-[1-1043]-C-terminus; N-terminus-[1249-1368]-[optional linker]-[1-1248]-C-terminus; or N-terminus-[1300-1368]-[optional linker]-[1-1299]-C-terminus, or the corresponding circular permutation of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0261] In yet other embodiments, the circularly permuted Cas9 has the following structure (based on the 1368 amino acids of S. pyogenes Cas9 (UniProtKB-Q99ZW2 (CAS9_STRP1)) (numbering is based on the location of the amino acids in SEQ ID NO: 9): N-terminus-[103-1368]-[optional linker]-[1-102]-C-terminus; N-terminus-[1029-1368]-[optional linker]-[1-1028]-C-terminus; N-terminus-[1042-1368]-[optional linker]-[1-1041]-C-terminus; N-terminus-[1250-1368]-[optional linker]-[1-1249]-C-terminus; or N-terminus-[1301-1368]-[optional linker]-[1-1300]-C-terminus, or the corresponding circular permutations of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0262] In some embodiments, a circular permutation can be formed by linking a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment can correspond to the C-terminal 95% or more of the amino acids of Cas9 (e.g., about amino acids 1300-1368), or the C-terminal 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of Cas9 (e.g., any one of SEQ ID NOs:54-63). The N-terminal portion may correspond to the N-terminal 95% or more of the amino acids of Cas9 (e.g., about amino acids 1-1300), or the N-terminal 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of Cas9 (e.g., of SEQ ID NO:9).
[0263] In some embodiments, a circular permutation may be formed by linking a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment relocated to the N-terminus encompasses or corresponds to the C-terminal 30% or less of the amino acids of Cas9 (e.g., amino acids 1012-1368 of SEQ ID NO:9). In some embodiments, the C-terminal fragment relocated to the N-terminus encompasses or corresponds to the C-terminal 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the amino acids of Cas9 (e.g., Cas9 of SEQ ID NO:9). In some embodiments, the C-terminal fragment relocated to the N-terminus includes or corresponds to the C-terminal 410 or less residues of Cas9 (e.g., Cas9 of SEQ ID NO: 9). In some embodiments, the C-terminal portion relocated to the N-terminus includes or corresponds to the C-terminal 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues of Cas9 (e.g., Cas9 of SEQ ID NO: 9). In some embodiments, the C-terminal portion relocated to the N-terminus includes or corresponds to the C-terminal 357, 341, 328, 120, or 69 residues of Cas9 (e.g., Cas9 of SEQ ID NO:9).
[0264] In other embodiments, a circular permutation Cas9 variant may be defined as a topological rearrangement of the Cas9 primary structure based on S. pyogenes Cas9 of SEQ ID NO: 9: (a) selecting a circular permutation (CP) site corresponding to an amino acid residue within the Cas9 primary structure, which splits the original protein into two halves: an N-terminal region and a C-terminal region; (b) modifying the Cas9 protein sequence (e.g., by genetic engineering techniques) by moving the original C-terminal region (containing the CP site amino acid) to precede the original N-terminal region, thereby forming a new N-terminus of the Cas9 protein that now begins with the CP site amino acid residue. The CP site may be located on any domain of the Cas9 protein, including, for example, the helical II domain, the RuvCIII domain, or the CTD domain. For example, the CP site may be located at the original amino acid residues 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 (relative to the S. pyogenes Cas9 of SEQ ID NO: 9). Thus, once relocated to the N-terminus, the original amino acids 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 would become the new N-terminal amino acids. The nomenclature for these CP-Cas9 proteins is Cas9-CP, respectively. 181 , Cas9-CP 199 , Cas9-CP 230 , Cas9-CP 270 , Cas9-CP 310 , Cas9-CP 1010 , Cas9-CP 1016 , Cas9-CP 1023 , Cas9-CP 1029 , Cas9-CP 1041 , Cas9-CP 1247 , Cas9-CP 1249 , and Cas9-CP 1282This description is not meant to be limited to making CP variants from SEQ ID NO:9, but may be implemented to make CP variants in any Cas9 sequence, either at CP sites corresponding to these positions or at other CP sites altogether. This description is in no way meant to be limited to a particular CP site. Virtually any CP site may be used to form a CP-Cas9 variant.
[0265] An exemplary CP-Cas9 amino acid sequence based on the Cas9 of SEQ ID NO:9 is provided below. In this, the linker sequence is indicated by underlining and the optional methionine (M) residue is indicated in bold. It should be recognized that the present disclosure provides CP-Cas9 sequences that do not include a linker sequence or that include a different linker sequence. It should be recognized that the CP-Cas9 sequence may be based on a Cas9 sequence other than SEQ ID NO:9, and any examples provided herein are not meant to be limiting. An exemplary CP-Cas9 sequence is as follows: [Table 21-1] [Table 21-2] [Table 21-3]
[0266] Cas9 circular permutations may be useful in the prime editing constructs utilized in the methods and compositions described herein. Provided below are exemplary C-terminal fragments of Cas9 based on the Cas9 of SEQ ID NO:2 that may be rearranged to the N-terminus of Cas9. It should be recognized that such C-terminal fragments of Cas9 are exemplary and are not meant to be limiting. These exemplary CP-Cas9 fragments have the following sequences: [Table 22]
[0267] I. Cas9 variants with altered PAM specificity Prime editors utilized in the disclosed methods and compositions may also include Cas9 variants with altered PAM specificity. Some aspects of the disclosure provide Cas9 proteins that are active against target sequences that do not include a canonical PAM (5'-NGG-3', where N is A, C, G, or T) at their 3' end. In some embodiments, the Cas9 protein is active against a target sequence that includes a 5'-NGG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that includes a 5'-NNG-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that includes a 5'-NNA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that includes a 5'-NNC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that includes a 5'-NNT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NGT-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NGA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NGC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NAA-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NAC-3'PAM sequence at its 3' end. In some embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NAT-3'PAM sequence at its 3' end. In still other embodiments, the Cas9 protein is active against a target sequence that comprises a 5'-NAG-3'PAM sequence at its 3' end.
[0268] It should be understood that any of the amino acid mutations described herein from a first amino acid residue (e.g., A) to a second amino acid residue (e.g., T) (e.g., A262T) may also include a mutation from the first amino acid residue to an amino acid residue that is similar to the second amino acid residue (e.g., conserved). For example, a mutation of an amino acid having a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan) may be a mutation to a second amino acid having a different hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan). For example, a mutation from alanine to threonine (e.g., A262T mutation) may also be a mutation from alanine to an amino acid that is similar in size and chemical properties to threonine, such as serine. As another example, a mutation of an amino acid with a positively charged side chain (e.g., arginine, histidine, or lysine) may be a mutation to a second amino acid with a different positively charged side chain (e.g., arginine, histidine, or lysine). As another example, a mutation of an amino acid with a polar side chain (e.g., serine, threonine, asparagine, or glutamine) may be a mutation to a second amino acid with a different polar side chain (e.g., serine, threonine, asparagine, or glutamine). Additional similar amino acid pairs include, but are not limited to: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. Those skilled in the art will recognize that such conservative amino acid substitutions will likely have a mild effect on protein structure and will likely be well tolerated without impairing function. In some embodiments, any amino acid of the amino acid mutations provided herein from one amino acid to threonine may be an amino acid mutation to serine.In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to arginine may be an amino acid mutation to lysine. In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to isoleucine may be an amino acid mutation to alanine, valine, methionine, or leucine. In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to lysine may be an amino acid mutation to arginine. In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to aspartic acid may be an amino acid mutation to glutamic acid or asparagine. In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to valine may be an amino acid mutation to alanine, isoleucine, methionine, or leucine. In some embodiments, any amino acid of the amino acid mutation provided herein from one amino acid to glycine may be an amino acid mutation to alanine. However, it should be recognized that additional conserved amino acid residues will be recognized by those skilled in the art, and any amino acid mutation to other conserved amino acid residues is also within the scope of this disclosure.
[0269] In some embodiments, the Cas9 protein comprises a combination of mutations that exert activity against a target sequence that contains a 5'-NAA-3'PAM sequence at its 3' end. In some embodiments, the combination of mutations is present in any one of the clones listed in Table 1. In some embodiments, the combination of mutations is a conservative mutation of the clones listed in Table 1. In some embodiments, the Cas9 protein comprises a combination of mutations of any one of the Cas9 clones listed in Table 1.
[0270] [Table 23]
[0271] In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 1. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 1.
[0272] In some embodiments, the Cas9 protein exhibits increased activity against a target sequence that does not contain a canonical PAM (5'-NGG-3') at its 3' end compared to the Streptococcus pyogenes Cas9 provided by SEQ ID NO: 9. In some embodiments, the Cas9 protein exhibits at least a 5-fold increased activity against a target sequence having a 3' end that is not directly adjacent to a canonical PAM sequence (5'-NGG-3') compared to the activity of the Streptococcus pyogenes Cas9 provided by SEQ ID NO: 9 against the same target sequence. In some embodiments, the Cas9 protein exerts at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, at least 1,000-fold, at least 5,000-fold, at least 10,000-fold, at least 50,000-fold, at least 100,000-fold, at least 500,000-fold, or at least 1,000,000-fold increased activity against a target sequence that is not immediately adjacent to a canonical PAM sequence (5'-NGG-3') compared to the activity of Streptococcus pyogenes provided by SEQ ID NO: 9 against the same target sequence. In some embodiments, the 3' end of the target sequence is immediately adjacent to an AAA, GAA, CAA, or TAA sequence. In some embodiments, the Cas9 protein comprises a combination of mutations that exert activity against a target sequence that includes a 5'-NAC-3'PAM sequence at its 3' end. In some embodiments, the combination of mutations is present in any one of the clones listed in Table 2. In some embodiments, the combination of mutations is a conservative mutation of the clones listed in Table 2. In some embodiments, the Cas9 protein comprises a combination of mutations of any one of the Cas9 clones listed in Table 2.
[0273] [Table 24-1] [Table 24-2]
[0274] In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 2. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any one of the variants in Table 2.
[0275] In some embodiments, the Cas9 protein exhibits increased activity against a target sequence that does not contain a canonical PAM (5'-NGG-3') at its 3' end compared to the Streptococcus pyogenes Cas9 provided by SEQ ID NO: 9. In some embodiments, the Cas9 protein exhibits at least a 5-fold increased activity against a target sequence having a 3' end that is not directly adjacent to a canonical PAM sequence (5'-NGG-3') compared to the activity of the Streptococcus pyogenes Cas9 provided by SEQ ID NO: 9 against the same target sequence. In some embodiments, the Cas9 protein exerts at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, at least 1,000-fold, at least 5,000-fold, at least 10,000-fold, at least 50,000-fold, at least 100,000-fold, at least 500,000-fold, or at least 1,000,000-fold increased activity against a target sequence that is not immediately adjacent to a canonical PAM sequence (5'-NGG-3') compared to the activity of Streptococcus pyogenes against the same target sequence as provided by SEQ ID NO: 9. In some embodiments, the 3' end of the target sequence is immediately adjacent to an AAC, GAC, CAC, or TAC sequence.
[0276] In some embodiments, the Cas9 protein comprises a combination of mutations that exert activity against a target sequence that contains a 5'-NAT-3'PAM sequence at its 3' end. In some embodiments, the combination of mutations is present in any one of the clones listed in Table 3. In some embodiments, the combination of mutations is a conservative mutation of the clones listed in Table 3. In some embodiments, the Cas9 protein comprises a combination of mutations of any one of the Cas9 clones listed in Table 3.
[0277] [Table 25]
[0278] The above description of various napDNAbps that may be used in relation to the prime editor is not meant to be limiting in any way. The prime editor may include canonical SpCas9 or any orthologous Cas9 protein or any variant Cas9 protein, which may be known or may be created or evolved by directed evolution or other mutagenesis processes, including any naturally occurring variant, mutant, or otherwise engineered version of Cas9. In various embodiments, the Cas9 or Cas9 variant has nickase activity, i.e., cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has an inactive nuclease, i.e., is an "inactive" Cas9 protein. Other variant Cas9 proteins that may be used are those that have a smaller molecular weight than the canonical SpCas9 (e.g., for easier delivery) or have a modified or rearranged primary amino acid structure (e.g., a circular permutation format). The prime editors utilized in the methods and compositions described herein may also include Cas9 equivalents, including Cas12a / Cpf1 and Cas12b proteins that are the result of convergent evolution. napDNAbps (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) used herein may contain various modifications that alter / enhance their PAM specificity. Finally, the present application contemplates any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a reference SpCas9 canonical sequence or a reference Cas9 equivalent (e.g., Cas12a / Cpf1). In certain embodiments, the Cas9 variant with expanded PAM capabilities is SpCas9(H840A)VRQR (SEQ ID NO:294), which has the following amino acid sequence (V, R, Q, R substitutions relative to SpCas9(H840A) are shown in bold and underlined. In addition, the methionine residue of SpCas9(H840) has been removed in SpCas9(H840A)VRQR): [Table 26]
[0279] In another specific embodiment, the Cas9 variant with expanded PAM capabilities is SpCas9(H840A)VRER, which has the following amino acid sequence (V, R, E, R substitutions relative to SpCas9(H840A) of SEQ ID NO: 12 are shown in bold and underlined. In addition, the methionine residue of SpCas9(H840) has been removed in SpCas9(H840A)VRER): [Table 27]
[0280] In some embodiments, the napDNAbp that functions for non-canonical PAM sequences is an Argonaute protein. One example of such a nucleic acid programmable DNA binding protein is the Argonaute protein (NgAgo) from Natronobacterium gregoryi. NgAgo is an endonuclease guided by ssDNA. NgAgo will bind and guide about 24 nucleotides of 5' phosphorylated ssDNA (gDNA) to its target site and create a DNA double-strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Using nuclease-inactive NgAgo (dNgAgo) can greatly expand the bases that can be targeted. The characterization and use of NgAgo is described in Gao et al., Nat Biotechnol., 2016 Jul;34(7):768-73. PubMed PMID:27136078; Swarts et al., Nature.507(7491)(2014):258-61; and Swarts et al., Nucleic Acids Res.43(10)(2015):5120-9, each of which is incorporated herein by reference.
[0281] In some embodiments, napDNAbp is a prokaryotic homolog of Argonaute protein. Prokaryotic homologs of Argonaute proteins are known, for example, Makarova K., et al., "Prokaryotic homologs of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile genetic elements", Biol Direct. 2009 Aug 25; 4: 29. doi: 10.1186 / 1745-6150-4-29, the entire contents of which are incorporated herein by reference. In some embodiments, napDNAbp is a Marinintoga piezophila Argonaute (MpAgo) protein. CRISPR-associated Marinintoga piezophila Argonaute (MpAgo) protein cleaves single-stranded target sequences using 5'-phosphorylated guides. 5' guides are used by all known Argonautes. The crystal structure of the MpAgo-RNA complex shows a guide strand binding site that includes residues that block 5' phosphate interactions. This data suggests the evolution of an Argonaute subclass with noncanonical specificity for 5' hydroxylated guides. See, e.g., Kaya et al., "A bacterial Argonaute with noncanonical guide RNA specificity", Proc Natl Acad Sci US A. 2016 Apr 12;113(15):4057-62, the entire contents of which are incorporated herein by reference. It should be recognized that other Argonaute proteins may be used and are within the scope of the present disclosure.
[0282] Some aspects of the present disclosure provide Cas9 domains with different PAM specificities. Typically, Cas9 proteins, such as Cas9 from S.pyogenes (spCas9), require a canonical NGG PAM sequence to bind to a particular nucleic acid region. This may limit the ability to edit a desired base in a genome. In some embodiments, the base editing fusion proteins provided herein may need to be precisely positioned, for example, where the target base is placed within a 4-base region (e.g., an "editing window") that is approximately 15 bases upstream of the PAM. See Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Thus, in some embodiments, any of the fusion proteins provided herein may contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-canonical PAM sequences have been described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015); and Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each of which are incorporated herein by reference.
[0283] For example, a napDNAbp domain with altered PAM specificity, such as a domain having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to wild-type Francisella novicida Cpf1 (D917, E1006, and D1255) (SEQ ID NO: 296) having the following amino acid sequence: [Table 28]
[0284] Additional napDNAbp domains with altered PAM specificity, such as domains having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to wild-type Geobacillus thermodenitrificans Cas9 (SEQ ID NO:31) having the following amino acid sequence: [Table 29]
[0285] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a nucleic acid programmable DNA binding protein that does not require a canonical (NGG) PAM sequence. In some embodiments, the napDNAbp is an Argonaute protein. One example of such a nucleic acid programmable DNA binding protein is the Argonaute protein (NgAgo) from Natronobacterium gregoryi. NgAgo is an endonuclease guided by ssDNA. NgAgo will bind and guide approximately 24 nucleotides of 5' phosphorylated ssDNA (gDNA) to its target site and create a DNA double strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Using a nuclease-inactive NgAgo (dNgAgo) can greatly expand the bases that can be targeted. The characterization and use of NgAgo is described in Gao et al., Nat Biotechnol., 34(7):768-73(2016), PubMed PMID:27136078; Swarts et al., Nature, 507(7491):258-61(2014); and Swarts et al., Nucleic Acids Res. 43(10)(2015):5120-9, each of which is incorporated herein by reference. The sequence of Natronobacterium gregoryi Argonaute is provided by SEQ ID NO:297.
[0286] The disclosed fusion proteins may include a napDNAbp domain having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to a wild-type Natronobacterium gregoryi Argonaute (SEQ ID NO: 297) having the following amino acid sequence: [Table 30]
[0287] In addition, any available method may be utilized to obtain or construct variant or mutant Cas9 proteins. The term "mutation" as used herein refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the original residue, then the position of the residue in the sequence, and the identity of the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual ( 4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). Mutations can encompass a variety of categories, such as single nucleotide polymorphisms, microduplication regions, indels, and inversions, and are in no way meant to be limiting. Mutations can encompass "loss-of-function" mutations, which are the normal consequence of a mutation reducing or abolishing protein activity. Most loss-of-function mutations are recessive because, in heterozygotes, the second chromosomal copy carries an unmutated version of the gene that encodes a fully functional protein, the presence of which compensates for the effect of the mutation. Mutations also encompass "gain-of-function" mutations, which confer abnormal activity to a protein or cell that would not otherwise be present under normal conditions. Many gain-of-function mutations are in regulatory sequences rather than coding regions, and thus can have a number of consequences. For example, a mutation can lead to one or more genes being expressed in the wrong tissues, and these tissues gain a function that they normally lack. By their nature, gain-of-function mutations are usually dominant.
[0288] Mutations can be introduced into the reference Cas9 protein using site-directed mutagenesis. Older methods of site-directed mutagenesis known in the art rely on subcloning the sequence to be mutated into a vector, such as the M13 bacteriophage vector, that allows the isolation of a single-stranded DNA template. In these methods, a mutagenic primer (i.e., a primer that can anneal to the site to be mutated but has one or more mismatched nucleotides at the site to be mutated) is annealed to the single-stranded template, and then the complement of the template is polymerized starting from the 3' end of the mutagenic primer. The resulting duplex is then transverted into a host bacterium, and plaques are screened for the desired mutation. More recently, site-directed mutagenesis has used PCR methodologies, which have the advantage of not requiring a single-stranded template. In addition, methods that do not require subcloning have been developed. When PCR-based site-directed mutagenesis is performed, several issues must be considered. First, in these methods, it is desirable to reduce the number of PCR cycles to prevent the spread of undesired mutations introduced by the polymerase. Second, selection must be used to reduce the number of unmutated parental molecules remaining in the reaction. Third, extended-length PCR methods are preferred to allow the use of a single PCR primer set. Fourth, due to the non-template-dependent end-extension activity of some thermostable polymerases, it is often necessary to incorporate an end-polishing step into the procedure prior to blunt-end ligation of the PCR-generated mutant products.
[0289] Mutations may also be introduced by directed evolution processes, such as phage-assisted continuous evolution (PACE) or phage-assisted discontinuous evolution (PANCE). The term "phage-assisted continuous evolution (PACE)" as used herein refers to continuous evolution employing phages as viral vectors. The general concept of PACE technology is described, for example, in International PCT Application PCT / US2009 / 056194, filed September 8, 2009, published as International Publication No. WO 2010 / 028347 on March 11, 2010; International Application PCT / US2011 / 066747, filed December 22, 2011, published as International Publication No. WO 2012 / 088381 on June 28, 2012; U.S. Patent Application No. 9,023,599, published May 5, 2015; and U.S. Patent Application No. 10,233,661, filed December 22, 2011; and U.S. Patent Application No. 10,233,661, published May 5, 2015. 4, International PCT application PCT / US2015 / 012022 filed on January 20, 2015, published as WO 2015 / 134121 on September 11, 2015, and International PCT application PCT / US2016 / 027795 filed on April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference. Variant Cas9 may also be obtained by phage-assisted non-linear evolution (PANCE). As used herein, refers to non-linear evolution using phage as a viral vector. PANCE is a simplified technique for rapid in vivo directed evolution that uses serial flask transfer of evolving "selection phage" (SP) containing the gene of interest to be evolved with new E. coli host cells, whereby the gene contained in the SP is continuously evolved while allowing the gene in the host E. coli to be held constant. Serial flask transfer has long served as a widely accessible approach for the laboratory evolution of microorganisms, and more recently, a related approach has been developed for bacteriophage evolution. The PANCE system features lower stringency than the PACE system.
[0290] Any of the above references relating to Cas9 or Cas9 equivalents are hereby incorporated by reference in their entirety unless already so claimed.
[0291] Reverse transcriptase domain and its modified variants In various embodiments, the improved prime editors disclosed herein include a polymerase (e.g., a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase, such as a reverse transcriptase) or variants thereof, which may be provided as a fusion protein or in trans with napDNAbp or other programmable nucleases. In various embodiments, the improved prime editors disclosed herein include an optimized evolved reverse transcriptase, as further described below.
[0292] In some embodiments, the improved prime editor protein comprises an MMLV reverse transcriptase that comprises one or more amino acid substitutions. A wild-type MMLV reverse transcriptase is provided by the following sequence: [Table 31]
[0293] The reverse transcriptase used in the improved prime editor described herein may contain one or more mutations compared to the wild-type amino acid sequence. In some embodiments, the reverse transcriptase is the MMLV pentamutant described above (i.e., containing the amino acid substitutions D200N, T306K, W313F, T330P, and L603W).
[0294] In some embodiments, the disclosure provides MMLV reverse transcriptase variants, and prime editors comprising MMLV reverse transcriptase variants (e.g., fusion proteins and prime editors in which napDNAbp and reverse transcriptase are provided in trans), the variants being T13I, V19I, A32T, G38V, S60Y, P111L, K120R, H126Y, T128N, T128F, T128H, V129S, P132S, G138R, C157 The MMLV reverse transcriptase variant comprises one or more mutations to SEQ ID NO: 33 selected from the group consisting of: F, P175Q, P175S, D200S, D200Y, D200N, D200C, Y222F, V223A, V223M, V223T, V223W, V223Y, L234I, T246I, N249S, T287A, P292T, E302A, E302K, T306K, G316R, E346K, K373N, W388C, V402A, K445N, M457I, and A462S. In some embodiments, the MMLV reverse transcriptase variant comprises two or more of these mutations, three or more of these mutations, four or more of these mutations, or five or more of these mutations.
[0295] In some embodiments, the MMLV reverse transcriptase variant used in the prime editor provided herein comprises a single mutation compared to SEQ ID NO: 33. In some embodiments, the single mutation is selected from the group consisting of T13I, G38V, K120R, H126Y, T128N, T128F, T128H, V129S, P132S, P175Q, P175S, D200C, D200Y, V223M, V223T, V223W, V223Y, L234I, P292T, G316R, K373N, M457I, and V402A.
[0296] In certain embodiments, the MMLV reverse transcriptase variants used in the prime editors provided herein comprise the following groups of mutations relative to the amino acid sequence of SEQ ID NO: 33: D200Y and E302A; D200Y, V223A, and M457I; V223M, T306K, and A462S; D200N and E302K; D200Y and E302K; T128N and V223 A; V19I, A32T, and D200Y; D200S, V223A, E346K, and W388C; S60Y, V223A, and N249S; P111L, V223A, T287A, and G316R; S60Y, G138R, and V223A; S60Y, Y222F, V223A, and K445N; or S60Y, C157F, V223A, and T246I. In certain embodiments, the MMLV reverse transcriptase variant used in the prime editor provided herein has an amino acid sequence similar to any one of SEQ ID NOs: 35-42, 172-177, 183, and 184, or an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 35-42, 172-177, 183, and 184. and the amino acid sequence comprises at least one of residues 13I, 19I, 32T, 38V, 60Y, 111L, 120R, 126Y, 128N, 128F, 128H, 129S, 132S, 138R, 157F, 175Q, 175S, 200S, 200Y, 200N, 200C, 222F, 223A, 223M, 223T, 223W, 223Y, 234I, 246I, 249S, 287A, 292T, 302A, 302K, 306K, 316R, 346K, 373N, 388C, 402A, 445N, 457I, and 462S.
[0297] In other examples, the proteins described herein may comprise an MMLV reverse transcriptase that comprises one or more substitutions at amino acid positions V19, A32, S60, P111, T128, G138R, C157F, D200, Y222, V223, T246, N249, T287, G316, E346, W388, and / or K445. In some embodiments, the proteins described herein comprise an MMLV reverse transcriptase that comprises one or more substitutions selected from the group consisting of V19I, A32T, S60Y, P111L, T128N, G138R, C157F, D200S, D200Y, Y222F, V223A, T246I, N249S, T287A, G316R, E346K, W388C, and K445N. In certain embodiments, the proteins described herein include an MMLV reverse transcriptase that includes any one of the following:
[0298] The following groups of amino acid substitutions:
[0299] T128N and V223A; V19I, A32T, and D200Y; D200S, V223A, E346K, and W388C; S60Y, V223A, and N249S; P111L, V223A, T287A, and G316R; S60Y, G138R, and V223A; S60Y, Y222F, V223A, and K445N; or S60Y, C157F, V223A, and T246I.
[0300] Exemplary evolved reverse transcriptases are as follows: [Table 32-1] [Table 32-2] [Table 32-3] [Table 32-4] [Table 32-5] [Table 32-6] [Table 32-7]
[0301] In the improved prime editors disclosed herein, the use of a reverse transcriptase comprising an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the evolved variants described herein is also contemplated by the present disclosure, provided that the RT sequence comprises one of the amino acid substitutions disclosed herein.
[0302] The present disclosure also contemplates the use of any wild-type reverse transcriptase in the improved prime editors described herein. Exemplary wild-type reverse transcriptases that may be used include, but are not limited to, the following sequences, or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto: [Table 33-1] [Table 33-2] [Table 33-3] [Table 33-4] [Table 33-5] [Table 33-6]
[0303] The use of a reverse transcriptase comprising an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the above enzymes in the improved prime editor proteins disclosed herein is also contemplated by the present disclosure.
[0304] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor in which each component is provided in trans), wherein the reverse transcriptase is an AVIRE reverse transcriptase of SEQ ID NO: 216, or an AVIRE reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 216, wherein the AVIRE reverse transcriptase variant comprises one or more mutations selected from the group consisting of D199N, T305K, W312F, G329P, and L604W. In some embodiments, the AVIRE reverse transcriptase variant comprises two or more of these mutations, three or more of these mutations, four or more of these mutations, or five or more of these mutations. In some embodiments, the AVIRE reverse transcriptase variant comprises the mutation D199N. In some embodiments, the AVIRE reverse transcriptase variant comprises a mutation T305K. In some embodiments, the AVIRE reverse transcriptase variant comprises a mutation W312F. In some embodiments, the AVIRE reverse transcriptase variant comprises a mutation G329P. In some embodiments, the AVIRE reverse transcriptase variant comprises a mutation L604W.
[0305] In certain embodiments, the AVIRE reverse transcriptase variant comprises an amino acid sequence of any one of SEQ ID NOs:217-221, or an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs:217-221, wherein the amino acid sequence comprises at least one of residues 199N, 305K, 312F, 329P, and 604W. AVIRE-RT(D199N): (SEQ ID NO:217) AVIRE-RT(T305K): APLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPVRVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLREVNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIFAFEWADAEEGESGQLTWTRLPQGFKNSPTLFDEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQAACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQIPVPKTKRQVREFLGKIGYCRLWIPGFAELAQPLYAATRGGNDPLVWGEKEEEAFQSLKLALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAGWPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSLTSTRPDLTDQPLAQAEATLFTDGSSYIRDGKRYAGAAVVTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNIYTDSRYAFATLHVHGMIYRERGLLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKGHQKDDAPTSTGNRRADEVAREVAIRPLSTQATISDAPDMPDTETPQYSNVEEALG (SEQ ID NO: 218) AVIRE-RT(W312F): (SEQ ID NO:219) AVIRE-RT(G329P): (SEQ ID NO: 220) AVIRE-RT(L604W): (SEQ ID NO:221)
[0306] In certain embodiments, the AVIRE reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO:243, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:243, wherein the amino acid sequence comprises residues 199N, 305K, 312F, 329P, and 604W: AVIRE_penta: (SEQ ID NO:243)
[0307] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor, where each component is provided in trans), where the reverse transcriptase is the KORV reverse transcriptase of SEQ ID NO: 222, or a KORV reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 222, where the KORV reverse transcriptase variant comprises one or more mutations selected from the group consisting of D197N, T303K, W310F, E327P, and L599W. In some embodiments, the KORV reverse transcriptase variant comprises two or more of these mutations, three or more of these mutations, four or more of these mutations, or five or more of these mutations. In some embodiments, the KORV reverse transcriptase variant comprises the mutation D197N. In some embodiments, the KORV reverse transcriptase variant comprises a mutation T303K. In some embodiments, the KORV reverse transcriptase variant comprises a mutation W310F. In some embodiments, the KORV reverse transcriptase variant comprises a mutation E327P. In some embodiments, the KORV reverse transcriptase variant comprises a mutation L599W.
[0308] In certain embodiments, the KORV reverse transcriptase variant comprises an amino acid sequence of any one of SEQ ID NOs:223-227, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs:223-227, wherein the amino acid sequence comprises at least one of residues 197N, 303K, 310F, 327P, and 599W: KORV-RT D197N: MNLEEEYRLHEKPVPPSIDPSWLQLFPMVWAEKAGMGLANQVPPVVVELKSDASPVAVRQYPMSKEAREGIRPHIQRFLDLGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSSLPPSHTWYSVLDLKDAFFCLKLHPNSQPLFAFEWRDPEKGNTGQLTWTRLPQGFKNSPTLFNEALHRDLASFRALNPQVVMLQYVDDLLVAAPTYRDCKEGTRRLLQELSKLGYRVSAKKAQLCREEVTYLGYLLKGGKRWLTPARKATVMKIPTPTTPRQVREFLGTAGFCRLWIPGFASLAAPLYPLTREKVPFTWTEAHQEAFGRIKEALLSAPALALPDLTKPFALYVDEKEGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPTCLKAIAAVALLLKDADKLTLGQNVLVIAPHNLESIVRQPPDRWMTNARMTHYQSLLLNERVSFAPPAILNPATLLPVESDDTPIHICSEILAEETGTRPDLRDQPLPGVPAWYTDGSSFIMDGRRQAGAAIVDNKRTVWASNLPEGTSAQKAELIALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGKDIKNKEEILALLEAIHLPKRVAIIHCPGHQRGTDPVATGNRKADEAAKQAAQSTRILTETTKNQEHFEPTRGK(SEQ ID NO: 223) KORV-RT T303K: MNLEEEYRLHEKPVPPSIDPSWLQLFPMVWAEKAGMGLANQVPPVVVELKSDASPVAVRQYPMSKEAREGIRPHIQRFLDLGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSSLPPSHTWYSVLDLKDAFFCLKLHPNSQPLFAFEWRDPEKGNTGQLTWTRLPQGFKNSPTLFDEALHRDLASFRALNPQVVMLQYVDDLLVAAPTYRDCKEGTRRLLQELSKLGYRVSAKKAQLCREEVTYLGYLLKGGKRWLTPARKATVMKIPTPTTPRQVREFLGKAGFCRLWIPGFASLAAPLYPLTREKVPFTWTEAHQEAFGRIKEALLSAPALALPDLTKPFALYVDEKEGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPTCLKAIAAVALLLKDADKLTLGQNVLVIAPHNLESIVRQPPDRWMTNARMTHYQSLLLNERVSFAPPAILNPATLLPVESDDTPIHICSEILAEETGTRPDLRDQPLPGVPAWYTDGSSFIMDGRRQAGAAIVDNKRTVWASNLPEGTSAQKAELIALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGKDIKNKEEILALLEAIHLPKRVAIIHCPGHQRGTDPVATGNRKADEAAKQAAQSTRILTETTKNQEHFEPTRGK(SEQ ID NO: 224) KORV-RT W310F: MNLEEEYRLHEKPVPPSIDPSWLQLFPMVWAEKAGMGLANQVPPVVVELKSDASPVAVRQYPMSKEAREGIRPHIQRFLDLGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSSLPPSHTWYSVLDLKDAFFCLKLHPNSQPLFAFEWRDPEKGNTGQLTWTRLPQGFKNSPTLFDEALHRDLASFRALNPQVVMLQYVDDLLVAAPTYRDCKEGTRRLLQELSKLGYRVSAKKAQLCREEVTYLGYLLKGGKRWLTPARKATVMKIPTPTTPRQVREFLGTAGFCRLFIPGFASLAAPLYPLTREKVPFTWTEAHQEAFGRIKEALLSAPALALPDLTKPFALYVDEKEGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPTCLKAIAAVALLLKDADKLTLGQNVLVIAPHNLESIVRQPPDRWMTNARMTHYQSLLLNERVSFAPPAILNPATLLPVESDDTPIHICSEILAEETGTRPDLRDQPLPGVPAWYTDGSSFIMDGRRQAGAAIVDNKRTVWASNLPEGTSAQKAELIALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGKDIKNKEEILALLEAIHLPKRVAIIHCPGHQRGTDPVATGNRKADEAAKQAAQSTRILTETTKNQEHFEPTRGK(SEQ ID NO: 225) KORV-RT E327P: MNLEEEYRLHEKPVPPSIDPSWLQLFPMVWAEKAGMGLANQVPPVVVELKSDASPVAVRQYPMSKEAREGIRPHIQRFLDLGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSSLPPSHTWYSVLDLKDAFFCLKLHPNSQPLFAFEWRDPEKGNTGQLTWTRLPQGFKNSPTLFDEALHRDLASFRALNPQVVMLQYVDDLLVAAPTYRDCKEGTRRLLQELSKLGYRVSAKKAQLCREEVTYLGYLLKGGKRWLTPARKATVMKIPTPTTPRQVREFLGTAGFCRLWIPGFASLAAPLYPLTRPKVPFTWTEAHQEAFGRIKEALLSAPALALPDLTKPFALYVDEKEGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPTCLKAIAAVALLLKDADKLTLGQNVLVIAPHNLESIVRQPPDRWMTNARMTHYQSLLLNERVSFAPPAILNPATLLPVESDDTPIHICSEILAEETGTRPDLRDQPLPGVPAWYTDGSSFIMDGRRQAGAAIVDNKRTVWASNLPEGTSAQKAELIALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGKDIKNKEEILALLEAIHLPKRVAIIHCPGHQRGTDPVATGNRKADEAAKQAAQSTRILTETTKNQEHFEPTRGK(SEQ ID NO: 226) KORV-RT L599W: (SEQ ID NO:227)
[0309] In certain embodiments, the KORV reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO:244, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:244, wherein the amino acid sequence comprises residues 197N, 303K, 310F, 327P, and 599W: KORV_penta: (query number 244)
[0310] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor, where each component is provided in trans), where the reverse transcriptase is a WMSV reverse transcriptase of SEQ ID NO: 228, or a WMSV reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 228, where the WMSV reverse transcriptase variant comprises one or more mutations selected from the group consisting of D197N, T303K, W311F, E327P, and L599W. In some embodiments, the WMSV reverse transcriptase variant comprises two or more of these mutations, three or more of these mutations, four or more of these mutations, or five or more of these mutations. In some embodiments, the WMSV reverse transcriptase variant comprises the mutation D197N. In some embodiments, the WMSV reverse transcriptase variant comprises a mutation T303K. In some embodiments, the WMSV reverse transcriptase variant comprises a mutation W311F. In some embodiments, the WMSV reverse transcriptase variant comprises a mutation E327P. In some embodiments, the WMSV reverse transcriptase variant comprises a mutation L599W.
[0311] In certain embodiments, the WMSV reverse transcriptase variant comprises an amino acid sequence of any one of SEQ ID NOs:229-233, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs:229-233, wherein the amino acid sequence comprises at least one of residues 197N, 303K, 311F, 327P, and 599W: WMSV-RT-D197N: (SEQ ID NO:229) WMSV-RT T303K: (SEQ ID NO: 230) WMSV-RT W311F: (SEQ ID NO:231) WMSV-RT E327P: (SEQ ID NO:232) WMSV-RT L599W: (SEQ ID NO:233)
[0312] In certain embodiments, the WMSV reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO:245, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:245, wherein the amino acid sequence comprises residues 197N, 303K, 311F, 327P, and 599W: WMSV_penta: (SEQ ID NO:245)
[0313] In some embodiments, the domain comprising RNA-dependent DNA polymerase activity comprises a PERV reverse transcriptase. For example, the improved prime editor protein described herein may comprise a PERV reverse transcriptase comprising one or more mutations relative to the amino acid sequence of SEQ ID NO: 45. In some embodiments, the PERV reverse transcriptase comprises one or more mutations selected from the group consisting of D199N, T305K, W312F, E329P and L602W relative to the amino acid sequence of SEQ ID NO: 45. In some embodiments, the PERV reverse transcriptase comprises mutations D199N, T305K, W312F, E329P and L602W relative to the amino acid sequence of SEQ ID NO: 45. In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor in which each component is provided in trans), wherein the reverse transcriptase is a PERV reverse transcriptase of SEQ ID NO: 45, or a PERV reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 45, wherein the PERV reverse transcriptase variant comprises one or more mutations selected from the group consisting of D199N, T305K, W312F, E329P, and L602W. In some embodiments, the PERV reverse transcriptase variant comprises two or more of these mutations, three or more of these mutations, four or more of these mutations, or five or more of these mutations. In some embodiments, the PERV reverse transcriptase variant comprises the mutation D199N. In some embodiments, the PERV reverse transcriptase variant comprises a mutation T305K. In some embodiments, the PERV reverse transcriptase variant comprises a mutation W312F. In some embodiments, the PERV reverse transcriptase variant comprises a mutation E329P. In some embodiments, the PERV reverse transcriptase variant comprises a mutation L602W.
[0314] In certain embodiments, the reverse transcriptase variant comprises an amino acid sequence of any one of SEQ ID NOs:214 and 234-238, or an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs:214 and 234-238, wherein the amino acid sequence comprises at least one of residues 199N, 305K, 312F, 329P, and 602W. PERV variant 21: (SEQ ID NO:214) PERV-RT D199N: TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPI(SEQ ID NO: 234) PERV-RT T305K: TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPI(SEQ ID NO: 235) PERV-RT W313F: TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPI(SEQ ID NO: 236) PERV-RT E329P: TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKPKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPI(SEQ ID NO: 237) PERV-RT L602W: (SEQ ID NO:238)
[0315] In certain embodiments, the PERV reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO:215, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:215, wherein the amino acid sequence comprises residues 199N, 305K, 312F, 329P, and 602W: PERV variant 21.6 (pentamutant containing D199N, T305K, W312F, E329P, and L602W substitutions): (SEQ ID NO:215)
[0316] In some embodiments, the domain comprising RNA-dependent DNA polymerase activity comprises Tf1 reverse transcriptase. For example, the improved prime editor protein described herein may comprise a Tf1 reverse transcriptase comprising one or more mutations relative to the amino acid sequence of SEQ ID NO: 55. In some embodiments, the Tf1 reverse transcriptase comprises one or more mutations selected from the group consisting of V14A, E22K, P70T, G72V, M102I, K106R, K118R, A139T, L158Q, F269L, S297Q, K356E, A363V, K413E, I423V and S492N relative to the amino acid sequence of SEQ ID NO: 55. In certain embodiments, the Tf1 reverse transcriptase comprises any one of the following amino acid substitutions relative to the amino acid sequence of SEQ ID NO: 55: K118R and S297Q; V14A, L158Q, F269L, and K356E; K106R, L158Q, F269L, A363V, and I423V; E22K, P70T, G72V, M102I, K106R, A139T, L158Q, F269L, A363V, K413E, and S492N; or P70T, G72V, M102I, K106R, L158Q, F269L, A363V, K413E, and S492N.
[0317] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor in which each component is provided in trans), wherein the reverse transcriptase is a Tf1 reverse transcriptase of SEQ ID NO: 171, or a Tf1 reverse transcriptase having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 171. and a Tf1 reverse transcriptase variant comprising, relative to SEQ ID NO: 171, one or more mutations selected from the group consisting of V14A, E22K, I64L, I64W, P70T, G72V, M102I, K106R, K118R, L133N, A139T, L158Q, S188K, I260L, F269L, E274R, R288Q, Q293K, S297Q, N316Q, K321R, K356E, A363V, K413E, I423V, and S492N. In some embodiments, the Tf1 reverse transcriptase variant comprises a single mutation, the single mutation comprising an I64L mutation, an I64W mutation, a K118R mutation, an L133N mutation, an S188K mutation, an I260L mutation, an E274R mutation, an R288Q mutation, a Q293K mutation, an S297Q mutation, an N316Q mutation, or a K321R mutation.
[0318] In some embodiments, the Tf1 reverse transcriptase variant has the following mutation group relative to the amino acid sequence of SEQ ID NO: 171: K118R and S297Q; V14A, L158Q, F269L, and K356E; E22K, P70T, G72V, M102I, K106R, A139T, L158Q, F269L, A363V, K413E, and S492N; P70T, G72V, M102I, K106R, L158Q, F269L, A363V, K413E, and S492N; K106R, L15 8Q, F269L, A363V, and I423V;K118R, S297Q, S188K, I64L, I260L, and R288Q;E22K, P70T, G72V, M102I, K106R, A139T, L158Q, F269L, A363V, K413E, S492N, K118R, S297Q, S188K, I64L, and I260L;K118R and S188K; K118R, S188K, and I260L; K118R, S188K, I260L, and S297Q; or Contains one of K118R, S188K, I260L, R288K, and S297Q.
[0319] In certain embodiments, the Tf1 reverse transcriptase variant has an amino acid sequence similar to any one of SEQ ID NOs: 196-213 and 251-255, or a sequence similar to any one of SEQ ID NOs: 196-213 and 251-255, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or an amino acid sequence that is at least 99% identical, wherein the amino acid sequence includes at least one of residues 14A, 22K, 64L, 64W, 70T, 72V, 102I, 106R, 118R, 133N, 139T, 158Q, 188K, 260L, 269L, 274R, 288Q, 293K, 297Q, 316Q, 321R, 356E, 363V, 413E, 423V, and 492N. Tf1 variant 5.131: ISSSKHTLSQMNKVSNIVKEPKLPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKSGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNI YPLPLIEQLLTKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 196) Tf1 variant 5.27: ISSSKHTLSQMNKASNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKEILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 197) Tf1 variant 5.47: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPRKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTVEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 198) Tf1 variant 5.59: ISSSKHTLSQMNKVSNIVKEPKLPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKSGIIRESKAINACPVIFVPRKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLTKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 199) Tf1 variant 5.60: (Sequence number 200) Tf1 variant 5.612: ISSSKHTLSQMNKVSNIVKEPKLPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPLRNYPLTPVKMQAMNDEINQGLKSGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNI YPLPLIEQLLTKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFLGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 201) Tf1 variant 5.618: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYRPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 202) Tf1 variant S188K: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 203) Tf1 variant I260L: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFLGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 204) Tf1 variant R288Q: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNQKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 205) Tf1 variant Q293K: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRKFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 206) Tf1 variant I64L: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPLRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 207) Tf1 variant I64W: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPWRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 208) Tf1 variant N316Q: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLQKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 209) Tf1 variant K321R: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKRDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 210) Tf1 variant L133N: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPNIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 211) Tf1 variant K118R: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYRPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 212) Tf1 variant K118R:Tf1 variant S297Q: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGFTPCQENIDKVLQWKQPKNRKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 213) Tf1-rat4: MISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVDYRPLNKYVKPN IYPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPRGVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQ VKFLGYHISEKGFTPCQENIDKVLQWKQPKNQKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (sequence number 251) Tf1evo3.1: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKSGIIRESKAINACPVIFVPRKEGTLRMVVDYKPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYCINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 252) Tf1evo3.2: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNV YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 253) Tf1evo+rat-1: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNV YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFLGYHISEKGLTPCQENIDKVLQWKQPKNQKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 254) Tf1evo+rat2: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVKMQAMNDEINQGLKSGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNI YPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGIKTAPAHFQYCINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQAKCEFHQSQV KFLGYHISEKGLTPCQENIDKVLQWKQPKNQKELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (sequence number 255)
[0320] In some embodiments, the domain comprising RNA-dependent DNA polymerase activity comprises Ec48 reverse transcriptase. For example, the improved prime editor proteins described herein may comprise an Ec48 reverse transcriptase comprising one or more mutations relative to the amino acid sequence of SEQ ID NO:59. In some embodiments, the Ec48 reverse transcriptase comprises one or more mutations selected from the group consisting of A36V, E54K, K87E, R205K, V214L, D243N, R267I, S277F, E279K, N317S, K318E, H324Q, K326E, E328K and R372K relative to the amino acid sequence of SEQ ID NO:59. In certain embodiments, the Ec48 reverse transcriptase comprises any one of the following amino acid substitutions relative to the amino acid sequence of SEQ ID NO:59: R267I, K318E, K326E, E328K, and R372K; K87E, R205K, V214L, D243N, R267I, N317S, K318E, H324Q, and K326E; E54K, K87E, D243N, R267I, E279K, and K318E; A36V, K87E, R205K, D243N, R267I, E279K, and K318E; E54K, K87E, D243N, R267I, E279K, and K318E; or E54K, K87E, D243N, R267I, S277F, E279K, and K318E.
[0321] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor, in which each component is provided in trans), wherein the reverse transcriptase is an Ec48 reverse transcriptase of SEQ ID NO:59, or has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO:59. An Ec48 reverse transcriptase variant, the Ec48 reverse transcriptase variant comprising one or more mutations selected from the group consisting of A36V, E54K, E60K, K87E, S151T, E165D, L182N, T189N, R205K, V214L, D243N, R267I, S277F, E279K, V303M, K307R, R315K, N317S, K318E, H324Q, K326E, E328K, K343N, R372K, R378K, and T385R relative to SEQ ID NO:59. In some embodiments, the Ec48 reverse transcriptase variant comprises a single mutation, wherein the single mutation is a L182N mutation, a T189N mutation, a K307R mutation, a R315K mutation, a R378K mutation, or a T385R mutation.
[0322] In some embodiments, the Ec48 reverse transcriptase variant has the following mutations relative to the amino acid sequence of SEQ ID NO: R267I, K318E, K326E, E328K, and R372K; K87E, R205K, V214L, D243N, R267I, N317S, K318E, H324Q, and K326E; E54K, K87E, D243N, R267I, E279K, and K318E; A36V, K87E, R205K, D243N, R267I, E279K, and K318E; E5 4K, K87E, D243N, R267I, E279K, and K318E; E54K, K87E, D243N, R267I, S277F, E279K, and K318E; E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N; E60K, K87E, S151T, E165D, D243N, R267I, E279K, V303M, K318E, and K343N; or R315K, L182N, and T189N.
[0323] In certain embodiments, the Ec48 reverse transcriptase variant has an amino acid sequence similar to any one of SEQ ID NOs: 188-195, 256 and 257, or a sequence similar to any one of SEQ ID NOs: 188-195, 256 and 257, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or an amino acid sequence that is at least 99% identical, wherein the amino acid sequence includes at least one of residues 36V, 54K, 60K, 87E, 151T, 165D, 182N, 189N, 205K, 214L, 243N, 267I, 277F, 279K, 303M, 307R, 315K, 317S, 318E, 324Q, 326E, 328K, 343N, 372K, 378K, and 385R. Ec48 variant 3.23: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQKKGL VYTRLLDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDEVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVSELGRVGQEEYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 188) Ec48 variant 3.35 (or Ec48-evo2): GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDKKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 189) Ec48 variant 3.36: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKVLSISVEELKAIAELSLDEKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQKKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 190) Ec48 variant 3.37: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDKKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 191) Ec48 variant 3.38: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDKKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPFDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 192) Ec48 variant 3.500: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALDYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 193) Ec48 variant 3.501: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRTVFEEILHIKDEALDYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSMAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 194) Ec48 variant 3.8 (or Ec48-evo1): GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINKRIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHDLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDEVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEEYKSFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKKKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 195) Ec48-v2: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKEIPKIDGSKRIVYSLHPKMRLLQSRINKRIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDDFVVQGANTSSYIANLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHDLPINKHKTKIFHCSSEPIKVHGLRVDYDSPRLPSDEVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGKVNKLGRVGHEKYESFKKQLQAIKPMPSKRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 256) Ec48-evo3: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTLKKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKRDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALDYLVDICTKDDFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGL VYTRLVDDITVSSKISNYDFSQMQSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDVAVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLKPL (SEQ ID NO: 257)
[0324] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor, where each component is provided in trans), wherein the reverse transcriptase is a Ne144 reverse transcriptase of SEQ ID NO: 239, or a Ne144 reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 239, wherein the Ne144 reverse transcriptase variant comprises one or more mutations selected from the group consisting of A157T, A165T, and G288V relative to SEQ ID NO: 239. In some embodiments, the Ne144 reverse transcriptase variant comprises mutations A157T, A165T, and G288V.
[0325] In certain embodiments, the Ne144 reverse transcriptase variant comprises an amino acid sequence of SEQ ID NO:240, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:240, wherein the amino acid sequence comprises at least one of residues 157T, 165T, and 288V: Ne144 RT 38.14: AGQPTSREALYERIRSTSKEEVILEEMIRLGFWPAQGAVPHDPAEEIRRRGELERQLSELREKSRKLYNEKALIAEQRKQRLAESRRKQKETKARRERERQERAQKWAQRKAGEILFLGEDVSGGMSHKT CDAELIKREGVPAIASAEELARAMGITLKELRFLTYNRKVSRVTHYRRFLLPKKTGGLRLISAPMPRLKRAQAWALEHIFNKLSFEPAAHGFVAGRSIVSNARPHVGADVVVNLDLKDFFPTVSFPRVKGA LRHLGYSESVATALALVCTEPEVDEVVLDGTTWYVARGERFLPQGSPCSPAITNLLCRRLDRRLHGLAQALGFVYTRYADDLTFSGRGEAAESKRVGKLLRGAADIVAHEGFVVHPDKTRVMRRGRRQEVTGVVVNDKTSVPRDELRKFRATLYQIEKDGPADKRWGNGGDVLAAVHGYACFVAMVDPSRGQPLLARARALLAKHGGPSKPPGGSGPRAPTPVQPTANAPEAPKPVAPATPAAPAKKGWKLF (SEQ ID NO: 240)
[0326] In some embodiments, the disclosure provides a reverse transcriptase, and a prime editor comprising the reverse transcriptase (e.g., a fusion protein or a prime editor in which each component is provided in trans), wherein the reverse transcriptase is a Vc95 reverse transcriptase of SEQ ID NO: 241, or a Vc95 reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 241, wherein the Vc95 reverse transcriptase variant comprises one or more mutations selected from the group consisting of L11M, S75A, V97M, N146D, and N245T relative to SEQ ID NO: 241. In some embodiments, the Vc95 reverse transcriptase variant comprises mutations L11M, S75A, V97M, N146D, and N245T.
[0327] In certain embodiments, the Vc95 reverse transcriptase variant comprises an amino acid sequence of SEQ ID NO:242, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:242, wherein the amino acid sequence comprises at least one of residues 11M, 75A, 97M, 146D, and 245T: Vc95 RT variant-25.8: NILTTLREQLMTNNVIMPQEFERLEVRGSHAYKVYSIPKRKAGRRTIAHPSSKLKICQRHLNAILNPLLKVHDASYAYVKGRSIKDNALVHSHSAYMLKMDFQNFFNSITPTILRQCLIQNDILLSVNELEKLEQLIFWNPSKKRDGKLILSVGSPISPLISNAIMYPFDKIINDICTKHGINYTRYADDITFSTNIKNTLNKLPEIVEQLIIQTYAGRIIINKRKTVFSSKKHNRHVTGITLTTDSKISIGIRSRKRYISSLVFKYINKNLDIDEINHMKGMLAFAYNIEPIYIHRLSHKYKVNIVEKILRGSN (SEQ ID NO: 242)
[0328] In some embodiments, the disclosure provides a reverse transcriptase and a prime editor (e.g., a fusion protein or a prime editor in which each component is provided in trans) comprising a reverse transcriptase, wherein the reverse transcriptase is a Gs reverse transcriptase of SEQ ID NO: 60, or a Gs reverse transcriptase variant having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 60, wherein the Gs reverse transcriptase variant has the sequence Relative to sequence number 60, the sequence includes one or more mutations selected from the group consisting of N12D, A16E, A16V, L17P, V20G, L37R, L37P, R38H, Y40C, I41N, I41S, W45R, I67T, I67R, G72E, G73V, G78V, Q93R, A123V, Y126F, E129G, K162N, P190L, D206V, R233K, A234V, R263G, P264S, R267M, K279E, R287I, R291K, P309T, R344S, R358S, R360S, E363G, V374A, and Q412H. In some embodiments, the Gs reverse transcriptase variant comprises two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more of these mutations.
[0329] In some embodiments, the Gs reverse transcriptase variant has the following mutations relative to the amino acid sequence of SEQ ID NO: 60: L17P and D206V; N12D, L37R, and G78V; A16E, L37P, and A123V; A16V, R38H, W45R, Y126F, and Q412H; A16V, R38H, W45R, and R291K; N12D, L37R, G72E, E129G, P264S, R344S, and R360S; N12D, Y40C, I67T, G73V, Q93R, R287I, and R358S; N12D, Y40C, I67T, G73V, Q93R, and R358S; N12D, I41N, P190L, A234V, and K279E; N12D, L37R, R267M, P309T, R358S, and E363G; A16V, V20G, I41S, R233K, and P264S; L17P, V20G, I41S, I67R, R263G, P264S, and V374A; or L17P, V20G, I41S, I67R, K162N, R263G, and P264S.
[0330] In certain embodiments, the Gs reverse transcriptase variant comprises an amino acid sequence of any one of SEQ ID NOs: 159-171, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 159-171, wherein the amino acid sequence is Includes at least one of: 12D, 16E, 16V, 17P, 20G, 37R, 37P, 38H, 40C, 41N, 41S, 45R, 67T, 67R, 72E, 73V, 78V, 93R, 123V, 126F, 129G, 162N, 190L, 206V, 233K, 234V, 263G, 264S, 267M, 279E, 287I, 291K, 309T, 344S, 358S, 360S, 363G, 374A, and 412H. Gs variant containing L17P+D206V EANQGAPGIDGVSTDQLRDYIRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSFGFRPGRNAHDAVRQ AQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDVLDKELEKRGLKFCRYADD CNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 159) Gs variant N12D+L37R+G78V ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQRRDYIRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLVIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 160) Gs A16E+L37P+A123V ALLERILARDNLITELKRVEAANQGAPGIDGVSTDQPRDYIRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQVQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 161) Gs variant A16V+R38H+W45R+Y126F+Q412H ALLERILARDNLITVLKRVEAANQGAPGIDGVSTDQLHDYIRAHRSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGFIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTHRYFELRQG (SEQ ID NO: 162) Gs A16V+R38H+W45R+R291K ALLERILARDNLITVLKRVEAANQGAPGIDGVSTDQLHDYIRAHRSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQKLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 163) Gs variant 814 N12D+L37R+G72E+E129G+P264S+R344S+R360S ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQRRDYIRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGEGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQGGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRSWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRSRLRLCQWLQWKRVRTSIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 164) Gs variant 815 N12D+Y40C+I67T+G73V+Q93R+R287I+R358S ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQLRDCIRAHWSTIHAQLLAGTYRPAPVRRVETPKPGGVTRQLGIPTVVDRLIQQAILRELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPISIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVSTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 165) Gs variant 816 N12D+Y40C+I67T+G73V+Q93R+R358S ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQLRDCIRAHWSTIHAQLLAGTYRPAPVRRVETPKPGGVTRQLGIPTVVDRLIQQAILRELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVSTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 166) Gs variant 817 N12D+I41N+P190L+A234V+K279E ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQLRDYNRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTLQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRVGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKRAFLGFSFTPEREARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 167) Gs variant 818 N12D+L37R+R267M+P309T+R358S+E363G ALLERILARDDLITALKRVEAANQGAPGIDGVSTDQRRDYIRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRPWKMAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMTERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVSTRIRGLRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 168) Gs variant 819 A16V+V20G+I41S+R233K+P264S ALLERILARDNLITVLKRGEANQGAPGIDGVSTDQLRDYSRAHWSTIHAQLLAGTYRPAPVRRVEIPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLKAGQRVKQSIQRFLEKTLKLKVNEEKSAVDRSWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 169) Gs variant 820 L17P+V20G+I41S+I67R+R263G+P264S+V374A ALLERILARDNLITAPKRGEANQGAPGIDGVSTDQLRDYSRAHWSTIHAQLLAGTYRPAPVRRVERPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKRVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDGSWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAAMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 170) Gs variant 821 L17P+V20G+I41S+I67R+K162N+R263G+P264S ALLERILARDNLITAPKRGEANQGAPGIDGVSTDQLRDYSRAHWSTIHAQLLAGTYRPAPVRRVERPKPGGGTRQLGIPTVVDRLIQQAILQELTPIFDPDFSSSSF GFRPGRNAHDAVRQAQGYIQEGYRYVVDMDLEKFFDRVNHDILMSRVARKVKDKNVLKLIRAYLQAGVMIEGVKVQTEEGTPQGGPLSPLLANILLDDLDKELEKRG LKFCRYADDCNIYVKSLRAGQRVKQSIQRFLEKTLKLKVNEEKSAVDGSWKRAFLGFSFTPERKARIRLAPRSIQRLKQRIRQLTNPNWSISMPERIHRVNQYVMGWIGYFRLVETPSVLQTIEGWIRRRLRLCQWLQWKRVRTRIRELRALGLKETAVMEIANTRKGAWRTTKTPQLHQALGKTYWTAQGLKSLTQRYFELRQG (SEQ ID NO: 171)
[0331] As illustrated in FIG. 27A, the present disclosure provides, in part, an engineered and PACE for prime editing. 2 Provided are evolved RT variants. To date, the only RT enzyme that has been utilized for prime editing in mammalian cells is M-MLV RT. M-MLV RT is a large enzyme (2.2 kB), which poses a barrier to many in vivo delivery methods, such as adeno-associated virus (AAV). Because RT enzymes vary greatly in their size and enzymatic activity, the alternative enzymes disclosed herein provide unique advantages for prime editing (e.g., smaller size or improved editing). These improvements result in prime editors that are more efficient and more easily delivered for therapeutic applications.
[0332] In various other embodiments, the modified prime editor protein, including PEmax, comprises a reverse transcriptase domain. In some embodiments, the reverse transcriptase domain is a variant of the wild-type MMLV reverse transcriptase having the amino acid sequence of SEQ ID NO: 34.
[0333] For example, PEmax of SEQ ID NO:2 is based on the wild-type MMLV reverse transcriptase domain of SEQ ID NO:33 (specifically, the GenScript codon-optimized MMLV reverse transcriptase having the nucleotide sequence of SEQ ID NO:33) and contains a variant reverse transcriptase domain of SEQ ID NO:34, which contains the amino acid substitutions D200N T306K W313F T330P L603W compared to the wild-type MMLV RT of SEQ ID NO:34. The amino acid sequence of the variant RT of PEmax is SEQ ID NO:34.
[0334] Modified prime editors may also include other variant RTs. In various embodiments, modified prime editors described herein (wherein the RT is provided either as a fusion partner or in trans) may include variant RTs that include one or more of the following mutations: P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, or D653N at the corresponding amino acid positions in the wild-type M-MLV RT of SEQ ID NO: 33 or in another wild-type RT polypeptide sequence.
[0335] Provided below are several exemplary reverse transcriptases that may be provided as fusions to napDNAbp proteins or as individual proteins in accordance with various embodiments of the present disclosure. Exemplary reverse transcriptases include variants having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the following wild-type or partial enzymes: [Table 34-1] [Table 34-2] [Table 34-3] [Table 34-4] [Table 34-5] [Table 34-6] [Table 34-7]
[0336] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that contains one or more of the following mutations: P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X at the corresponding amino acid positions in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid.
[0337] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a P51X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is L.
[0338] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an S67X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0339] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an E69X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0340] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an L139X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is P.
[0341] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a T197X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is A.
[0342] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a D200X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is N.
[0343] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an H204X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is R.
[0344] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an F209X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is N.
[0345] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an E302X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0346] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an E302X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is R.
[0347] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a T306X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0348] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a F309X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is N.
[0349] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a W313X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is F.
[0350] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a T330X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is P.
[0351] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an L345X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is G.
[0352] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an L435X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is G.
[0353] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an N454X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0354] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a D524X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is G.
[0355] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an E562X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is Q.
[0356] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a D583X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is N.
[0357] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an H594X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is Q.
[0358] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an L603X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is W.
[0359] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes an E607X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is K.
[0360] In various other embodiments, a prime editor described herein (wherein the RT is provided either as a fusion partner or in trans) can include a variant RT that includes a D653X mutation at the corresponding amino acid position in the wild-type M-MLV RT of SEQ ID NO: 33 or on another wild-type RT polypeptide sequence, where "X" can be any amino acid. In certain embodiments, X is N.
[0361] Several exemplary reverse transcriptases that may be provided as fusions to napDNAbp proteins or as individual proteins according to various embodiments of the present disclosure are provided below. Exemplary reverse transcriptases include variants having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the wild-type or partial enzymes represented by SEQ ID NOs: 33-34 and 63-78.
[0362] The Prime Editor (PE) system described herein contemplates any publicly available reverse transcriptases described or disclosed in any of the following US patents (each of which is incorporated by reference in its entirety): US Patent Nos. 10,202,658; 10,189,831; 10,150,955; 9,932,567; 9,783,791; 9,580,698; 9,534,201; and 9,458,484, and any variants thereof that may be made using known methods for installing mutations or evolving proteins. The following references describe reverse transcriptases in the art. Each of these disclosures is incorporated by reference herein in its entirety.
[0363] Herzig, E., Voronin, N., Kucherenko, N. & Hizi, AA Novel Leu92 Mutant of HIV-1 Reverse Transcriptase with a Selective Deficiency in Strand Transfer Causes a Loss of Viral Replication. J. Virol. 89, 8119-8129 (2015).
[0364] Mohr, G. et al.A Reverse Transcriptase-Cas1 Fusion Protein Contains a Cas6 Domain Required for Both CRISPR RNA Biogenesis and RNA Spacer Acquisition.Mol.Cell 72,700-714.e8(2018).
[0365] Zhao, C., Liu, F. & Pyle, AMan ultraprocessive, accurate reverse transcriptase encoded by a metazoan group II intron.RNA 24,183-195(2018).
[0366] Zimmerly、S.&Wu、L.An Unexplored Diversity of Reverse Transcriptases in Bacteria.Microbiol Spectr 3,MDNA3-0058-2014(2015).
[0367] Ostertag、E.M.&Kazazian Jr、H.H.Biology of Mammalian L1 Retrotransposons.Annual Review of Genetics 35,501-538(2001).
[0368] Perach、M.&Hizi、A.Catalytic Features of the Recombinant Reverse Transcriptase of Bovine Leukemia Virus Expressed in Bacteria.Virology 259,176-189(1999).
[0369] Lim,D.et al.Crystal structure of the moloney murine leukemia virus RNase H domain.J.Virol.80,8379-8389(2006).
[0370] Zhao、C.&Pyle、A.M.Crystal structures of a group II intron maturase reveal a missing link in spliceosome evolution.Nature Structural&Molecular Biology 23,558-565(2016).
[0371] Griffiths,D.J.Endogenous retroviruses in the human genome sequence.Genome Biol.2,REVIEWS1017(2001).
[0372] Baranauskas、A.et al.Generation and characterization of new highly thermostable and processive M-MuLV reverse transcriptase variants.Protein Eng Des Sel 25,657-668(2012).
[0373] Zimmerly、S.、Guo、H.、Perlman、P.S.&Lambowltz、A.M.Group II intron mobility occurs by target DNA-primed reverse transcription.Cell 82,545-554(1995).
[0374] Feng,Q.,Moran,J.V.,Kazazian,H.H.&Boeke,J.D.Human L1 retrotransposon encodes a conserved endonuclease required for retrotransposition.Cell 87,905-916(1996).
[0375] Berkhout,B.,Jebbink,M.&Zsiros,J.Identification of an Active Reverse Transcriptase Enzyme Encoded by a Human Endogenous HERV-K Retrovirus.Journal of Virology 73,2365-2375(1999).
[0376] Kotewicz、M.L.、Sampson、C.M.、D'Alessio、J.M.&Gerard、G.F.Isolation of cloned Moloney murine leukemia virus reverse transcriptase lacking ribonuclease H activity.Nucleic Acids Res 16,265-277(1988).
[0377] Arezi,B.&Hogrefe,H.Novel mutations in Moloney Murine Leukemia Virus reverse transcriptase increase thermostability through tighter binding to template-primer.Nucleic Acids Res 37,473-481(2009).
[0378] Blain,S.W.&Goff,S.P.Nuclease activities of Moloney murine leukemia virus reverse transcriptase.Mutants with altered substrate specificities.J.Biol.Chem.268,23585-23592(1993).
[0379] Xiong,Y.&Eickbush,T.H.Origin and evolution of retroelements based upon their reverse transcriptase sequences.EMBO J 9,3353-3362(1990).
[0380] Herschhorn,A.&Hizi,A.Retroviral reverse transcriptases.Cell.Mol.Life Sci.67,2717-2747(2010).
[0381] Taube、R.、Loya、S.、Avidan、O.、Perach、M.&Hizi、A.Reverse transcriptase of mouse mammary tumour virus:expression in bacteria、purification and biochemical characterization.Biochem.J.329(Pt 3),579-587(1998).
[0382] Liu、M.et al.Reverse Transcriptase-Mediated Tropism Switching in Bordetella Bacteriophage.Science 295,2091-2094(2002).
[0383] Luan、D.D.、Korman、M.H.、Jakubczak、J.L.&Eickbush、T.H.Reverse transcription of R2Bm RNA is primed by a nick at the chromosomal target site:a mechanism for non-LTR retrotransposition.Cell 72,595-605(1993).
[0384] Nottingham、R.M.et al.RNA-seq of human reference RNA samples using a thermostable group II intron reverse transcriptase.RNA 22,597-613(2016).
[0385] Telesnitsky、A.&Goff、S.P.RNase H domain mutations affect the interaction between Moloney murine leukemia virus reverse transcriptase and its primer-template.Proc.Natl.Acad.Sci.U.S.A.90,1276-1280(1993).
[0386] Halvas,E.K.,Svarovskaia,E.S.&Pathak,V.K.Role of Murine Leukemia Virus Reverse Transcriptase Deoxyribonucleoside Triphosphate-Binding Site in Retroviral Replication and In Vivo Fidelity.Journal of Virology 74,10349-10358(2000).
[0387] Nowak、E.et al.Structural analysis of monomeric retroviral reverse transcriptase in complex with an RNA / DNA hybrid.Nucleic Acids Res 41,3874-3887(2013).
[0388] Stamos、J.L.、Lentzsch、A.M.&Lambowitz、A.M.Structure of a Thermostable Group II Intron Reverse Transcriptase with Template-Primer and Its Functional and Evolutionary Implications.Molecular Cell 68,926-939.e4(2017).
[0389] Das、D.&Georgiadis、M.M.The Crystal Structure of the Monomeric Reverse Transcriptase from Moloney Murine Leukemia Virus.Structure 12,819-829(2004).
[0390] Avidan、O.、Meer、M.E.、Oz、I.&Hizi、A.The processivity and fidelity of DNA synthesis exhibited by the reverse transcriptase of bovine leukemia virus.European Journal of Biochemistry 269、859-867(2002).
[0391] Gerard,G.F.et al.The role of template-primer in protection of reverse transcriptase from thermal inactivation.Nucleic Acids Res 30,3118-3129(2002).
[0392] Monot、C.et al.The Specificity and Flexibility of L1 Reverse Transcription Priming at Imperfect T-Tracts.PLOS Genetics 9、e1003499(2013).
[0393] Mohr、S.et al.Thermostable group II intron reverse transcriptase fusion proteins and their use in cDNA synthesis and next-generation RNA sequencing.RNA 19、958-970(2013).
[0394] Any of the above references relating to reverse transcriptase are hereby incorporated by reference in their entirety unless already so stated.
[0395] Additional Domains A. Linker The modified PE fusion proteins described herein may include one or more linkers.
[0396] As defined above, the term "linker" as used herein refers to a chemical group or molecule that connects two molecules or moieties, for example, the binding domain and cleavage domain of a nuclease. In some embodiments, the linker connects the gRNA binding domain of a programmable nuclease and the catalytic domain of a polymerase (for example, reverse transcriptase) by RNA. In some embodiments, the linker connects dCas9 and reverse transcriptase. Typically, the linker is placed between or flanked by two groups, molecules, or other moieties, and is connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (for example, a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0397] The linker may be as simple as a covalent bond, or it may be a polymeric linker that is many atoms in length. In certain embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.
[0398] In some other embodiments, the linker comprises the amino acid sequence (GGGGS) n(SEQ ID NO: 84), (G) n (SEQ ID NO: 85), (EAAAK) n (SEQ ID NO: 86), (GGS) n (SEQ ID NO: 87), (SGGS) n (SEQ ID NO: 81), (XP) n (SEQ ID NO: 88), or any combination thereof, where n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n (SEQ ID NO: 87), where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ I...
Claims
1. (a) a nucleic acid programmable DNA binding protein (napDNAbp); and (b) an amino acid sequence having at least 80%, at least 90%, or at least 95% sequence identity to SEQ ID NO: 33, or a truncation of SEQ ID NO: 33 lacking the RNase H domain, and further comprising the following amino acids: T13I, V19I, A32T, G38V, S60Y, P111L, K120R, H126Y, T128N, T128F, T128H, V129S, P132S, G138R, C157F, P 175Q, P175S, D200S, D200Y, D200C, Y222F, V223A, V223M, V223T, V223W, V223Y, L234I, T246I, N249S, T287A, P292T, E302A, E302K, G316R, E346K, K373N, W388C, V402A, K445N, M457I, and A462S.
2. The MMLV reverse transcriptase variant has the following mutations relative to the amino acid sequence of SEQ ID NO: 33: D200Y and E302A; D200Y, V223A, and M457I; V223M, T306K, and A462S; D200N and E302K; D200Y and E302K; T128N and V223A; V19I, A32T, and D200Y; D200S, V22 3A, E346K, and W388C; S60Y, V223A, and N249S; P111L, V223A, T287A, and G316R; S60Y, G138R, and V223A; S60Y, Y222F, V223A, and K445N; or S60Y, C157F, V223A, and T246I.
3. 2. The prime editor of Claim 1, wherein the one or more mutations comprise: (i) T13I, G38V, K120R, H126Y, P132S, P175Q, P175S, L234I, P292T, G316R, K373N, V402A, or M457I; or (ii) T128F, T128H, T128N, V129S, D200C, V223M, V223T, V223W, or V223Y.
4. The prime editor of claim 1, wherein the MMLV reverse transcriptase variant comprises the amino acid sequence of any one of SEQ ID NOs: 35-42, 172-177, 183, and 184.
5. 10. The prime editor of any preceding claim, wherein the napDNAbp: (i) is a Cas protein; (ii) is a Cas9 nickase (nCas9); (iii) comprises the amino acid sequence of SEQ ID NO:9, SEQ ID NO:10, or SEQ ID NO:11, or an amino acid sequence at least 80%, at least 90%, or at least 95% identical to the sequence of SEQ ID NO:9, SEQ ID NO:10, or SEQ ID NO:11; and / or (iv) D23G, H99Q, H99R, E102K, E102S, E102R, N175K, D177G, K218R, N309D, I312V, 2. The prime editor of claim 1, comprising a Cas9 variant comprising one or more mutations with respect to SEQ ID NO:9 or SEQ ID NO:11 selected from the group consisting of E471K, G485S, K562N, D608N, I632V, D645N, D645E, R654C, G687D, G715E, H721Y, R753K, R753G, H754R, K775R, E790K, T804A, K918A, K1003R, M1021Y, E1071K, and E1260D; further optionally, wherein the Cas9 variant comprises the amino acid sequence of any one of SEQ ID NOs: 178-180.
6. an amino acid sequence having at least 80%, at least 90%, or at least 95% sequence identity to SEQ ID NO: 33, or a truncation of SEQ ID NO: 33 lacking the RNase H domain, and further comprising: T13I, V19I, A32T, G38V, S60Y, P111L, K120R, H126Y, T128N, T128F, T128H, V129S, P132S, G138R, C157F, P175Q, P175S, 1. A MMLV reverse transcriptase variant comprising one or more mutations relative to SEQ ID NO: 33 selected from the group consisting of: D200S, D200Y, D200C, Y222F, V223A, V223M, V223T, V223W, V223Y, L234I, T246I, N249S, T287A, P292T, E302A, E302K, G316R, E346K, K373N, W388C, V402A, K445N, M457I, and A462S.
7. The following mutations were made to the amino acid sequence of SEQ ID NO: 33: D200Y and E302A; D200Y, V223A, and M457I; V223M, T306K, and A462S; D200N and E302K; D200Y and E302K; T128N and V223A; V19I, A32T, and D200Y; D200S, V223A, E346K, and and W388C; S60Y, V223A, and N249S; P111L, V223A, T287A, and G316R; S60Y, G138R, and V223A; S60Y, Y222F, V223A, and K445N; or S60Y, C157F, V223A, and T246I.
8. 7. The MMLV reverse transcriptase variant of claim 6, wherein the one or more mutations include: (i) T13I, G38V, K120R, H126Y, P132S, P175Q, P175S, L234I, P292T, G316R, K373N, V402A, or M457I; or (ii) T128F, T128H, T128N, V129S, D200C, V223M, V223T, V223W, or V223Y.
9. The MMLV reverse transcriptase variant of claim 6, comprising the amino acid sequence of any one of SEQ ID NOs: 35-42, 172-177, 183, and 184.
10. A complex comprising the prime editor of any one of claims 1 to 5 and a prime editing guide RNA (PEGRNA).
11. One or more polynucleotides encoding the prime editor of any one of claims 1 to 5.
12. A vector comprising one or more polynucleotides according to claim 11.
13. A pharmaceutical composition comprising the prime editor of any one of claims 1 to 5.
14. A pharmaceutical composition comprising the conjugate of claim 10.
15. A pharmaceutical composition comprising one or more polynucleotides according to claim 11.
16. A pharmaceutical composition comprising the vector of claim 12.
17. 6. An in vitro or ex vivo method comprising contacting a nucleic acid molecule with the prime editor of any one of claims 1 to 5, and a prime editing guide RNA (PEGRNA).
18. The prime editor of any one of claims 1 to 5 for use as a drug.
19. The conjugate of claim 10 for use as a drug.
20. One or more polynucleotides according to claim 11 for use as a medicament.
21. The vector according to claim 12 for use as a drug.
22. 14. A pharmaceutical composition according to claim 13 for use as a medicament.