Methods and compositions for editing nucleotide sequences
By using guided editing technology, PEgRNA is combined with Cas protein fusion enzyme to achieve precise nucleotide replacement or insertion at the target site, solving the problems of low genome editing efficiency and numerous byproducts in existing technologies, and realizing efficient and precise genome editing.
Patent Information
- Application Number
- CN202080037186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-17
- Filing Date
- 2020-03-19
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2040-03-19
AI Technical Summary
Existing technologies are insufficient for efficiently and accurately editing single nucleotide mutations in the human genome. Traditional methods are inefficient and produce byproducts that affect cell growth and survival.
Using guided editing technology, PEgRNA is combined with Cas protein fusion enzyme, and reverse transcriptase or DNA polymerase performs precise nucleotide replacement or insertion at the target site. Error-prone reverse transcriptase is used to introduce the desired changes, forming a single-stranded DNA flap and achieving efficient editing through cellular repair mechanisms.
It enables efficient and precise single or multiple nucleotide changes, expands the scope and therapeutic potential of genome editing, is applicable to multiple cell types, and reduces adverse effects.
Smart Images

Figure CN114729365B_ABST
Abstract
Description
[0001] Government Support
[0002] This application was made with government support under Grant Numbers U01 AI 142756, RM1 HG009490, R01 EB022376, and R35 GM118062 awarded by the National Institutes of Health. The government has certain rights in the application.
[0003] Related Applications and References Incorporated by Reference
[0004] This U.S. provisional application is related to and incorporates by reference the following applications, U.S. Provisional Application No. 62 / 820,813 (Attorney Docket No. B1195.70074US00), filed March 19, 2019, U.S. Provisional Application No. 62 / 858,958 (Attorney Docket No. B1195.70074US01), filed June 7, 2019, U.S. Provisional Application No. 62 / 889,996 (Attorney Docket No. B1195.70074US02), filed August 21, 2019, U.S. Provisional Application No. 62 / 922,654 (Attorney Docket No. B1195.70083US00), filed August 21, 2019, U.S. Provisional Application No. 62 / 913,553 (Attorney Docket No. B1195.70074US03), filed October 10, 2019, U.S. Provisional Application No. 62 / 973,558 (Attorney Docket No. B1195.70083US01), filed October 10, 2019, U.S. Provisional Application No. 62 / 931,195 (Attorney Docket No. B1195.70074US04), filed November 5, 2019, U.S. Provisional Application No. 62 / 944,231 (Attorney Docket No. B1195.70074US05), filed December 5, 2019, U.S. Provisional Application No. 62 / 974,537 (Attorney Docket No. B1195.70083US02), filed December 5, 2019, U.S. Provisional Application No. 62 / 991,069 (Attorney Docket No. B1195.70074US06), filed March 17, 2020, and U.S. Provisional Application No. (number not available as of the filing of this document) (Attorney Docket No. B1195.70083US03), filed March 17, 2020.
[0005] Incorporation of Sequence Listing
[0006] This specification includes the sequence listing (2 parts) submitted concurrently on compact disc pursuant to 37 CFR § 1.52(e). Pursuant to the requirements of 37 CFR § 1.52(e)(5), the applicant expressly incorporates by reference all information and material contained on the compact disc into this specification, designated as “B119570083WO00-SEQ.txt,” created on March 19, 2020, and having a size of 371.109 MB. By this reference, the sequence listing constitutes a part of the specification. The compact disc does not contain other files. SUMMARY
[0008] Pathogenic single nucleotide mutations cause approximately 67% of human diseases with a genetic component 7 Unfortunately, despite decades of gene therapy exploration, treatment options for these patients with genetic diseases remain extremely limited 8 Perhaps one of the most straightforward solutions to this therapeutic challenge is to directly correct the single nucleotide mutation in the patient’s genome, which would address the root cause of the disease and potentially provide long-lasting benefit. Although this strategy was previously unimaginable, the recent advent of the CRISRP / Cas system 9 has brought improvements in genome editing capabilities that are now within reach. By directly designing a guide RNA (gRNA) sequence comprising approximately 20 nucleotides complementary to the target DNA sequence, CRISPR-associated (Cas) nucleases can be specifically accessed to almost any possible genomic site 1,2 To date, several monomeric bacterial Cas nuclease systems have been identified and adapted for genome editing applications 10 This natural diversity of Cas nucleases, coupled with an increasing number of engineered variants 11-14 provides fertile ground for the development of new genome editing technologies.
[0009] While gene disruption with CRISPR is now a mature technology, the precise editing of a single base pair in the human genome remains a major challenge 3 Homology directed repair (HDR) has long been used in human cells and other organisms to insert, correct, or exchange DNA sequences at double-strand break (DSB) sites using a repair template of donor DNA encoding the desired edit 15 However, traditional HDR is very inefficient in most human cell types, especially in non-dividing cells, and competitive non-homologous end joining (NHEJ) predominantly leads to indel byproducts 16 Other issues associated with the generation of DSBs, which can lead to large chromosomal rearrangements and deletions at the target locus 17 , or activation of the p53 axis leading to growth arrest and apoptosis18,19 .
[0010] Several approaches have been explored to address the shortcomings of HDR. For example, repair of single-stranded DNA breaks (nicks) with oligonucleotide donors has been shown to reduce indel formation, but the yield of the desired repair product remains low 20 . Other strategies attempt to bias repair towards HDR rather than NHEJ using small molecules and biological reagents 21-23 . However, the effectiveness of these approaches is dependent on cell type, and perturbation of normal cellular states can lead to adverse and unpredictable effects.
[0011] Recently, Liu et al. developed base editing as a technology to edit target nucleotides without creating DSBs or relying on HDR 4 -6,24-27 Direct modification of DNA bases by Cas fusion deaminases allows C·G to T·A, or A·T to G·C, base pairs to be efficiently converted within a short target window (~5-7 bases). As a result, base editors have rapidly been adopted by the scientific community. However, several factors can limit their universality for precise genome editing.
[0012] Accordingly, the development of programmable editors capable of introducing any desired single or multiple nucleotide changes, which can install nucleotide insertions or deletions (e.g., at least 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more base pairs insertions or deletions), and / or can alter or modify the nucleotide sequence of a target site with high specificity and high efficiency, would greatly expand the scope and therapeutic potential of CRISPR-based genome editing technologies. SUMMARY
[0014] The present disclosure discloses new compositions (e.g., new PEgRNAs and PE complexes comprising them) and methods of using prime editing (PE) to repair therapeutic targets, such as those identified in the ClinVar database, using PEgRNAs designed using specialized algorithms described herein. Accordingly, in one aspect, the present application discloses algorithms for large-scale prediction of PEgRNA sequences that can be used to repair therapeutic targets (e.g., those included in the ClinVar database). Further, the present application discloses predicted sequences of therapeutic PEgRNAs designed using the disclosed algorithms and can be used with prime editing to repair therapeutic targets.
[0015] The algorithms and predicted PEgRNA sequences disclosed herein generally relate to prime editing. Accordingly, the present disclosure also provides a description of various components and aspects of prime editing, including suitable napDNAbps (e.g., Cas9 nickases) and polymerases (e.g., reverse transcriptases), as well as other suitable components (e.g., linkers, NLSs) and PE fusion proteins, which can be used with the therapeutic PEgRNAs disclosed herein.
[0016] As disclosed herein, prime editing is a general and precise method of genome editing that uses a nucleic acid programmable DNA binding protein (“napDNAbp”) bound to a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans to the napDNAbp) to directly write new genetic information into a specific prime editing DNA prime editing site, where the prime editing system is programmed with a prime editing (PE) guide RNA (“PEgRNA”), both of which specify a target site and a template to synthesize a desired edit in the form of a replacement strand, engineered onto the guide RNA by extension (DNA or RNA) (e.g., at the 5’ or 3’ end, or at an internal portion of the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleobase replacement) shares the same sequence as the endogenous strand of the target site to be edited (except that it includes the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand of the target site is replaced with the newly synthesized replacement strand containing the desired edit. In certain instances, prime editing can be considered a “search and replace” genome editing technique, as the prime editors described herein not only search for and locate the desired target site to be edited, but also, at the same time, encode a replacement strand containing the desired edit, which is installed in place of the corresponding target site endogenous DNA strand. The prime editors of the present disclosure are in part related to the discovery that the mechanism of target prime reverse transcription (TPRT) or “prime editing” can be harnessed or adapted to efficiently perform CRISPR / Cas-based precise genome editing and genetic flexibility (e.g., as described in the Examples below). FIG. 1A-1TPRT is naturally used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial group II introns.28 29The inventors herein use a Cas protein- reverse transcriptase fusion or related system to target a specific DNA sequence with a guide RNA, make a single-strand nick at the target site, and use the nicked DNA as a primer for engineered reverse transcriptase template reverse transcription that integrates with the guide RNA. However, while this concept started with a prime editor using a reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptases, but can include virtually any DNA polymerase. In fact, while the application refers to prime editors with "reverse transcriptases" throughout, the reverse transcriptase is presented herein only as one example of a DNA polymerase that can work with a prime editor. Thus, wherever the specification refers to a "reverse transcriptase," one of ordinary skill in the art will understand that any suitable DNA polymerase can be used in place of the reverse transcriptase. Thus, in one aspect, a prime editor can comprise a Cas9 (or equivalent napDNAbp) that is programmed to target a DNA sequence by associating it with a specified guide RNA (i.e., PEgRNA) that contains a spacer sequence that anneals to a complementary pre-spacer in the target DNA. The specified guide RNA also contains new genetic information in the form of an extension that encodes a DNA replacement strand containing a desired genetic alteration that is used to replace the corresponding endogenous DNA strand at the target site. To transfer the information from the PEgRNA to the target DNA, the mechanism of prime editing includes cleaving the target site on one strand of the DNA to expose a 3'-hydroxyl. The exposed 3'-hydroxyl can then be used to direct DNA polymerization of the edit-encoding extension on the PEgRNA directly to the target site. In various embodiments, the extension, which provides a template for the polymerization of the replacement strand containing the edit, can be formed from RNA or DNA. In the case of RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (e.g., a reverse transcriptase). In the case of DNA extension, the polymerase of the prime editor can be a DNA-dependent DNA polymerase.
[0017] The newly synthesized strand formed by the guide editor disclosed herein (i.e., the replacement DNA strand containing the desired edit) will be homologous to the genomic target sequence (i.e., have the same sequence) except for containing the desired nucleotide change (e.g., a single nucleotide change, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replacement) DNA strand can also be referred to as a single-stranded DNA flap, which will compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In certain embodiments, the system can be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with the Cas9 domain, or provided in trans to the Cas9 domain). The error-prone reverse transcriptase can introduce changes during the synthesis of the single-stranded DNA flap. Thus, in certain embodiments, the error-prone reverse transcriptase can be utilized to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used with the system, the changes can be random or non-random.
[0018] The resolution of the hybridization intermediate (comprising the single-stranded DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can include removal of the replacement flap of the resulting endogenous DNA (e.g., using the 5’ end DNA flap endonuclease FENl), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change due to cellular DNA repair and / or replication processes. Because the templated DNA synthesis provides single-nucleotide precision for any nucleotide modification (including insertions and deletions), the scope of this approach is very broad, and it is envisioned that it can be used for countless applications in basic science and therapeutics.
[0019] Algorithms and methods for designing therapeutic PEgRNAs
[0020] In one aspect, the disclosure relates to a new algorithm for designing therapeutic PEgRNAs, in particular, large-scale design as opposed to one-off PEgRNA design exercises.
[0021] Accordingly, some aspects relate to a computerized method for determining a sequence of a guide editor guide RNA (PEgRNA). The method includes using at least one computer hardware processor to access data indicative of an input allele, an output allele, and a fusion protein comprising a nucleic acid programmable DNA binding protein and a polymerase (e.g., a reverse transcriptase). The method includes determining a PEgRNA sequence based on the input allele, the output allele, and the fusion protein, wherein the PEgRNA sequence is designed to associate with the fusion protein to change the input allele to the output allele, including determining one or more of the following features of the PEgRNA sequence: a spacer that is complementary to a target nucleotide sequence in the input allele (i.e., a spacer, as defined in FIG. 27 a gRNA backbone for interacting with the fusion protein (i.e., a gRNA backbone, as defined in FIG. 27a gRNA core defined in the middle); and an extension (i.e., an extension arm as shown in FIG. 27 The PEgRNA can also include a DNA synthesis template (as shown in FIG. 27 The PEgRNA can also include a DNA synthesis template (as shown in FIG. 27 The PEgRNA can also include a DNA synthesis template (as shown in FIG. 27 The PEgRNA can also include a DNA synthesis template (as shown in
[0022] In some examples, the method includes determining the spacer and the extension, and determining that the spacer is at the 5' end of the PEgRNA and the extension is at the 3' end of the PEgRNA structure.
[0023] In some examples, the method includes determining the spacer and the extension, and determining that the spacer is at the 5' end of the PEgRNA and the extension is at the 3' end of the PEgRNA structure.
[0024] In some examples, accessing data indicative of the input allele and the output allele includes accessing a database containing a set of input alleles and associated output alleles. Accessing the database can include accessing the ClinVar database of the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov / clinvar / ), which contains a plurality of entries, each entry containing an input allele from a set of input alleles and an output allele from a set of output alleles (e.g., a wild-type or an allele with a desired activity). Determining the PEgRNA sequence can include determining one or more PEgRNA sequences for each input allele and associated output allele in the set.
[0025] In some examples, accessing data indicative of the fusion protein includes determining the fusion protein from a plurality of fusion proteins.
[0026] In some examples, the fusion protein includes a Cas9 protein. The fusion protein can include a Cas9-NG protein, a Cas9-NGG, a saCas9-KKH, or a SpCas9 protein.
[0027] In some examples, changing the input allele to the output allele includes a single nucleotide change, an insertion of one or more nucleotides, a deletion of one or more nucleotides, or a combination thereof.
[0028] In some embodiments, the method comprises determining a spacer, wherein the spacer comprises a nucleotide sequence of between 1 and 40 nucleotides. In some embodiments, the method comprises determining a spacer, wherein the spacer comprises a nucleotide sequence of between 5 and 35 nucleotides. In some embodiments, the method comprises determining a spacer, wherein the spacer comprises a nucleotide sequence of between 10 and 30 nucleotides. In some embodiments, the method comprises determining a spacer, wherein the spacer comprises a nucleotide sequence of between 15 and 25 nucleotides. In some examples, the method comprises determining a spacer, wherein the spacer comprises a nucleotide sequence of about 20 nucleotides. The method can comprise determining the spacer based on the location of the variation in the corresponding protospacer nucleotide sequence. The variation can be installed in an edit window of between about protospacer position -15 to protospacer position +39. The variation can be installed in an edit window of between about protospacer position -10 to protospacer position +34. The variation can be installed in an edit window of between about protospacer position -5 to protospacer position +29. The variation can be installed in an edit window of between about protospacer position -1 to protospacer position +27.
[0029] In some examples, the method can comprise: determining a set of initial candidate protospacers based on the input allele and the fusion protein, wherein each initial candidate protospacer comprises a PAM of the fusion protein in the input allele; determining one or more initial candidate protospacers from the set of initial candidate protospacers, wherein each comprises an incompatible nicking position; removing the determined one or more initial candidate protospacers from the set to generate a set of remaining candidate protospacers; and wherein determining the PEgRNA structure comprises determining a plurality of PEgRNA structures, wherein each of the PEgRNA structures comprises a different spacer determined based on a corresponding protospacer from the set of remaining candidate protospacers.
[0030] In some examples, the method comprises determining the extension and the DNA synthesis template (e.g., RT template sequence), wherein the DNA synthesis template (e.g., RT template sequence) comprises about 1 nucleotide to about 40 nucleotides. In some examples, the method comprises determining the extension and the DNA synthesis template (e.g., RT template sequence), wherein the DNA synthesis template (e.g., RT template sequence) comprises about 3 nucleotides to about 38 nucleotides. In some examples, the method comprises determining the extension and the DNA synthesis template (e.g., RT template sequence), wherein the DNA synthesis template (e.g., RT template sequence) comprises about 5 nucleotides to about 36 nucleotides. In some examples, the method comprises determining the extension and the DNA synthesis template (e.g., RT template sequence), wherein the DNA synthesis template (e.g., RT template sequence) comprises about 7 nucleotides to about 34 nucleotides.
[0031] In some examples, determining the PEgRNA comprises determining the spacer based on the input allele and / or the fusion protein; and determining the DNA synthesis template (e.g., RT template sequence) based on the spacer.
[0032] In some examples, the DNA synthesis template (e.g., RT template sequence) encodes a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to a nick site, wherein the single-stranded DNA flap comprises a desired nucleotide change. The single-stranded DNA flap is capable of hybridizing to the endogenous DNA sequence adjacent to the nick site, thereby resulting in installation of the desired nucleotide change. The single-stranded DNA flap is capable of displacing the endogenous DNA sequence adjacent to the nick site. Cellular repair of the single-stranded DNA flap results in installation of the desired nucleotide change, thereby forming the desired product.
[0033] In some examples, the fusion protein, when complexed with the PEgRNA, is capable of binding to a target DNA sequence. The target DNA sequence comprises a target strand in which a change occurs and a complementary non-target strand.
[0034] In some examples, the input allele comprises a pathogenic DNA mutation and the output allele comprises a corrected DNA sequence.
[0035] Some embodiments relate to a system comprising at least one processor; and at least one computer-readable storage medium having instructions encoded thereon that, when executed, cause the at least one processor to perform a computerized method for determining a PEgRNA structure.
[0036] Some embodiments relate to at least one computer-readable storage medium having instructions encoded thereon that, when executed, cause at least one processor to perform a computerized method for determining a PEgRNA sequence. Some embodiments relate to a base editing method using a PEgRNA determined according to the computerized method for determining a PEgRNA.
[0037] Therapeutic PEgRNAs
[0038] In another aspect, the disclosure provides therapeutic PEgRNAs that have been designed using the algorithms disclosed herein, such as FIG. 27 and FIG. 28 as shown.
[0039] For example, PEgRNAs useful in the methods disclosed herein can be designed using the algorithms disclosed herein, such as FIG. 27An example illustration is provided. The figure provides the structure of one embodiment of PEgRNAs encompassed herein, which can be designed according to the methods defined in Example 2. The PEgRNA comprises three main constituent elements arranged in 5’ to 3’ direction, namely: a spacer, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in 5’ to 3’ direction, namely: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, the PEgRNA can comprise an optional 3’ end modifier region (el) and an optional 5’ end modifier region (e2). Further, the PEgRNA can comprise a transcription termination signal (not depicted) at the 3’ end of the PEgRNA. The description of the PEgRNA structure is not meant to be limiting, but to encompass variations in the arrangement of elements. For example, the optional sequence modifications (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends. FIG. 27 The PEgRNAs shown can be designed by the algorithms disclosed herein.
[0040] In another example illustration, FIG. 28 An example illustration is provided. The figure provides the structure of one embodiment of PEgRNAs encompassed herein, which can be designed according to the methods defined in Example 2. The PEgRNA comprises three main constituent elements arranged in 5’ to 3’ direction, namely: a spacer, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in 5’ to 3’ direction, namely: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, the PEgRNA can comprise an optional 3’ end modifier region (el) and an optional 5’ end modifier region (e2). Further, the PEgRNA can comprise a transcription termination signal (not depicted) at the 3’ end of the PEgRNA. The description of the PEgRNA structure is not meant to be limiting, but to encompass variations in the arrangement of elements. For example, the optional sequence modifications (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends. FIG. 27 The PEgRNAs shown can be designed by the algorithms disclosed herein.
[0041] In various embodiments, the present disclosure provides therapeutic PEgRNAs of SEQ ID NOS: 1-135514 and 813085-880462 designed against ClinVar database entries using the algorithms disclosed herein.
[0042] In various other embodiments, exemplary PEgRNAs designed against the ClinVar database using the algorithms disclosed herein are included in the sequence listing, which forms a part of the present specification. The sequence listing includes the complete PEgRNA sequences of SEQ ID NOS: 1-135514 and 813085-880462. Each of these complete PEgRNAs consists of a spacer (SEQ ID NOS: 135515-271028 and 880463-947840) and an extension arm (SEQ ID NOS: 271029-406542 and 947841-1015218). In addition, each PEgRNA comprises a gRNA core, e.g., as defined by SEQ ID NOS: 1361579-1361580. The extension arms of SEQ ID NOS: 271029-406542 and 947841-1015218 further each comprise a primer binding site (SEQ ID NOS: 406543-542056 and 1015219-1082596), an editing template (SEQ ID NOS: 542057-677570 and 1082597-1149974), and a homology arm (SEQ ID NOS: 677571-813084 and 1149975-1217352). The PEgRNAs optionally can comprise a 5’-terminal modifier region and / or a 3’-terminal modifier region. The PEgRNAs can further comprise a reverse transcription termination signal at the 3’ of the PEgRNA (e.g., SEQ ID NOS: 1361560-1361566). This application includes the design and use of all of these sequences.
[0043] In various embodiments, the prime editor guide RNA comprises (a) a guide RNA and (b) an RNA extension at the 5’ or 3’ end of the guide RNA, or at an intramolecular location in the guide RNA, as shown in FIG. 3A The RNA extension can comprise (i) a DNA synthesis template comprising a desired nucleotide change, (ii) a reverse transcription primer binding site, and (iii) an optional linker sequence. In various embodiments, the DNA synthesis template encodes a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to the nick site, wherein the single-stranded DNA flap comprises the desired nucleotide change.
[0044] In various embodiments, the RNA extension arm is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides in length.
[0045] In certain embodiments, the prime editor guide RNA comprises the nucleotide sequence of SEQ ID NO: 1361548-1361581, or a nucleotide sequence having at least 85%, or at least 90%, or at least 95%, or at least 98%, or at least 99% sequence identity to any one of SEQ ID NO: 1361548-1361581.
[0046] In some embodiments, the prime editor guide RNA (PEgRNA) comprises a variant of the nucleotide sequence of SEQ ID NO: 1361548-1361581, comprising at least one mutation as compared to the nucleotide sequence of SEQ ID NO: 1361548-1361581. In some embodiments, the variant comprises more than 1 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more) mutations as compared to the nucleotide sequence of SEQ ID NO: 1361548-1361581.
[0047] In another aspect, the disclosure provides a prime editor guide RNA comprising a guide RNA and at least one RNA extension (i.e., extension arm, according to FIG. 27 ) to the 3’ end of the guide RNA. In other embodiments, the RNA extension is to the 5’ end of the guide RNA. In other embodiments, the RNA extension is to an intra-molecular position within the guide RNA, preferably the intra-molecular positioning of the extension portion does not disrupt the function of the protospacer.
[0048] In various embodiments, the prime editor guide RNA (PEgRNA) is capable of binding a napDNAbp and directing the napDNAbp to a target DNA sequence. The target DNA sequence can comprise a target strand and a complementary non-target strand, wherein the guide RNA hybridizes to the target strand to form an RNA-DNA hybrid and an R-loop.
[0049] In various embodiments of the guide editor guide RNA, the at least one RNA extension comprises a DNA synthesis template. In various other embodiments, the RNA extension further comprises a reverse transcription primer binding site. In other embodiments, the RNA extension comprises a linker or spacer that connects the RNA extension to the guide RNA.
[0050] In various embodiments, the RNA extension can be at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.
[0051] In other embodiments, the DNA synthesis template (i.e., the editing template, according to FIG. 27 ) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.
[0052] In other embodiments, the reverse transcription primer binding site sequence (i.e., the primer binding site, according to FIG. 27the length of the at least one nucleic acid sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0053] In various embodiments, the length of the optional linker or spacer is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0054] The designed PEgRNAs disclosed herein can complex with a prime editor fusion protein.
[0055] In one aspect, the present specification provides a prime editor fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a reverse transcriptase. In various embodiments, the fusion protein is capable of genome editing by target- guided reverse transcription in the presence of a prime editor guide RNA (PEgRNA).
[0056] In some embodiments, the napDNAbp is selected from the group consisting of Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2C3, and Argonaute, and optionally has nickase activity.
[0057] In other embodiments, the fusion protein is capable of binding a target DNA sequence (e.g., genomic DNA) when complexed with a prime editor guide RNA as described herein.
[0058] In other embodiments, the target DNA sequence comprises a target strand and a complementary non-target strand.
[0059] In other embodiments, the fusion protein binds to a prime editor guide RNA complex to form an R-loop. The R-loop can comprise (i) an RNA-DNA hybrid comprising the prime editor guide RNA and the target strand, and (ii) the complementary non-target strand.
[0060] In other embodiments, the complementary non-target strand is cleaved to form a reverse transcriptase primer sequence having a free 3’ end.
[0061] In other embodiments, the single-stranded DNA flap hybridizes to an endogenous DNA sequence adjacent to the nick site, thereby installing the desired nucleotide change. In other embodiments, the single-stranded DNA flap displaces an endogenous DNA sequence adjacent to the nick site and having a free 5’ end. In some embodiments, the displaced endogenous DNA having a 5’ end is excised by the cell.
[0062] In various embodiments, cellular repair of the single-stranded DNA flap results in installation of the desired nucleotide change, thereby forming the desired product.
[0063] In various other embodiments, the desired nucleotide change is installed in an editing window between about -4 to +10 of the PAM sequence.
[0064] In other embodiments, the desired nucleotide change is installed in an editing window between about -5 to +5 nucleotides of the nick site, or between about -10 to +10 of the nick site, or between about -20 to +20 of the nick site, or between about -30 to +30 of the nick site, or between about -40 to +40 of the nick site, or between about -50 to +50 of the nick site, or between about -60 to +60 of the nick site, or between about -70 to +70 of the nick site, or between about -80 to +80 of the nick site, or between about -90 to +90 of the nick site, or between about -100 to +100 of the nick site, or between about -200 to +200 of the nick site.
[0065] In various embodiments, the napDNAbp comprises the amino acid sequence of SEQ ID NO: 1361421. In various other embodiments, the napDNAbp comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 1361421-1361484 and 1361593-1361596.
[0066] In other embodiments, the reverse transcriptase of the disclosed fusion proteins and / or compositions can comprise the amino acid sequence of any one of SEQ ID NOS: 1361485-1361514 and 1361597-1361598. In other embodiments, the reverse transcriptase can comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOS: 1361485-1361514 and 1361597-1361598. These sequences can be naturally occurring reverse transcriptase sequences, for example from a retrovirus or a retrotransposon, or these sequences can be non-naturally occurring or engineered.
[0067] In various other embodiments, the fusion proteins disclosed herein can comprise various structural configurations. For example, the fusion protein can comprise the structure NH2- [napDNAbp]-[reverse transcriptase]-COOH; or NH2-[reverse transcriptase]-[napDNAbp]-COOH, wherein each instance of "]-[" indicates the presence of an optional linker sequence.
[0068] In various embodiments, the linker sequence comprises the amino acid sequence of SEQ ID NOS: 1361520-1361530, 1361585, and 1361603, or an amino acid sequence that is at least 80%, 85%, or 90%, or 95%, or 99% identical to any one of the linker amino acid sequences of SEQ ID NOS: 1361520-1361530, 1361585, and 1361603.
[0069] In various embodiments, the desired nucleotide change(s) that is / are incorporated into the target DNA can be a single nucleotide change (e.g., a transition or a transversion), an insertion of one or more nucleotides, a deletion of one or more nucleotides, or a combination thereof.
[0070] In certain instances, the insertion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.
[0071] In certain other cases, the deletion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.
[0072] In various embodiments of the prime editor guide RNA, the DNA synthesis template (i.e., the edit template, according to FIG. 27 ) can encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to the nick site, wherein the single-stranded DNA flap comprises the desired nucleotide change. The single-stranded DNA flap can replace the endogenous single-stranded DNA at the nick site. The endogenous single-stranded DNA that is replaced at the nick site can have a 5’ end and form an endogenous flap that can be excised by the cell. In various embodiments, excision of the 5’ endogenous flap can help drive product formation, as removal of the 5’ endogenous flap facilitates hybridization of the single-stranded 3’ DNA flap to the corresponding complementary DNA strand, as well as incorporation or assimilation of the desired nucleotide change carried by the single-stranded 3’ DNA flap into the target DNA.
[0073] In various embodiments of the prime editor guide RNA, cellular repair of the single-stranded DNA flap results in installation of the desired nucleotide change, thereby forming the desired product.
[0074] In yet another aspect of the application, the specification provides a complex comprising a fusion protein described herein and any of the above-described prime editor guide RNAs (PEgRNA).
[0075] In other aspects of the application, the specification provides a complex comprising a napDNAbp (e.g., Cas9) and a prime editor guide RNA. The napDNAbp can be a Cas9 nickase (e.g., spCas9), or can be an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of any one of SEQ ID NOS: 1361421-1361484 and 1361593-1361596.
[0076] In various embodiments involving complexes, the prime editor guide RNA is capable of directing the napDNAbp to a target DNA sequence. In various embodiments, the reverse transcriptase can be provided in trans, i.e., from a source different from the complex itself. For example, the reverse transcriptase can be provided to the same cell with the complex by introducing a separate vector encoding the reverse transcriptase alone.
[0077] In another aspect, the present description provides pharmaceutical compositions (e.g., the fusion proteins described herein, PEgRNAs of SEQ ID NOs: 1-135, 514). In some embodiments, the pharmaceutical compositions comprise one or more of a napDNAbp, a fusion protein, a reverse transcriptase, and a prime editor guide RNA. In some embodiments, the fusion proteins described herein and a pharmaceutically acceptable excipient. In other embodiments, the pharmaceutical compositions comprise any of the extension guide RNAs described herein and a pharmaceutically acceptable excipient. In other embodiments, the pharmaceutical compositions comprise any of the extension guide RNAs described herein in combination with any of the fusion proteins described herein and a pharmaceutically acceptable excipient. In other embodiments, the pharmaceutical compositions comprise any polynucleotide sequence encoding one or more of a napDNAbp, a fusion protein, a reverse transcriptase, and a prime editor guide RNA. In other embodiments, the various components disclosed herein can be isolated into one or more pharmaceutical compositions. For example, a first pharmaceutical composition can comprise a fusion protein or a napDNAbp, a second pharmaceutical composition can comprise a reverse transcriptase, and a third pharmaceutical composition can comprise a prime editor guide RNA.
[0078] In yet another aspect, the present disclosure provides kits. In one embodiment, the kit comprises one or more polynucleotides encoding one or more components, including a fusion protein, a napDNAbp, a reverse transcriptase, and a prime editor guide RNA (e.g., any of SEQ ID NOs: 1-135 514 or 813085-880462). The kits can also comprise isolated preparations of vectors, cells, and polypeptides, including any of the fusion proteins, napDNAbps, or reverse transcriptases disclosed herein.
[0079] In yet another aspect, the present disclosure provides methods of using the disclosed PEgRNA repertoire.
[0080] In one embodiment, the method involves a method of installing a desired nucleotide change in a double-stranded DNA using a PEgRNA disclosed herein. The method first comprises contacting a double-stranded DNA sequence with a complex comprising a fusion protein as described herein and a prime editor guide RNA, wherein the fusion protein comprises a napDNAbp and a reverse transcriptase, and wherein the prime editor guide RNA comprises a DNA synthesis template comprising the desired nucleotide change. The napDNAbp cleaves the double-stranded DNA sequence on the non-target strand, thereby generating a free single-stranded DNA with a 3’ end. After cleavage, the 3’ end of the free single-stranded DNA hybridizes to the DNA synthesis template, thereby priming the reverse transcriptase domain. The reverse transcriptase then facilitates DNA polymerization from the 3’ end, thereby generating a single-stranded DNA flap comprising the desired nucleotide change. The single-stranded DNA flap then replaces the endogenous DNA strand near the cleavage site, thereby installing the desired nucleotide change in the double-stranded DNA sequence.
[0081] In other embodiments, the present disclosure provides a method for introducing one or more changes in the nucleotide sequence of a DNA molecule at a target locus, comprising contacting the DNA molecule with a nucleic acid programmable DNA binding protein (napDNAbp) and a guide RNA that targets the napDNAbp to the target locus, wherein the guide RNA comprises a reverse transcriptase (RT) template sequence comprising at least one desired nucleotide change. The napDNAbp exposes a 3’ end in a DNA strand at the target locus, which hybridizes to a DNA synthesis template (e.g., the RT template sequence) to prime reverse transcription. Next, a single-stranded DNA flap comprising at least one desired nucleotide change based on the DNA synthesis template (e.g., the RT template sequence) is synthesized or polymerized by a reverse transcriptase. Finally, the at least one desired nucleotide change is incorporated into the corresponding endogenous DNA, thereby introducing one or more changes in the nucleotide sequence of the DNA molecule at the target locus.
[0082] In other embodiments, the present disclosure provides a method of introducing one or more changes in the nucleotide sequence of a DNA molecule at a target locus by target-directed reverse transcription, the method comprising: contacting the DNA molecule at the target locus with (i) a fusion protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a reverse transcriptase, and (ii) a guide RNA comprising an RT template comprising a desired nucleotide change (e.g., any one of SEQ ID NOs: 1-135514 or 813085-880462); such contacting facilitates target primed reverse transcription of the RT template to generate a single-stranded DNA comprising the desired nucleotide change, and incorporation of the desired nucleotide change into the DNA molecule at the target locus by DNA repair and / or replication processes.
[0083] In some embodiments, the step of replacing the endogenous DNA strand comprises: (i) hybridizing a single-stranded DNA flap to the endogenous DNA strand proximal to the cleavage site to create a sequence mismatch; (ii) excising the endogenous DNA strand; (iii) repairing the mismatch to form a desired product comprising a desired nucleotide change in both strands of the DNA.
[0084] The methods disclosed herein can involve a fusion protein having a napDNAbp that is a nuclease dead Cas9 (dCas9), a Cas9 nickase (nCas9), or a nuclease active Cas9. In other embodiments, the napDNAbp and reverse transcriptase are not encoded as a single fusion protein, but can be provided in separate constructs. Thus, in some embodiments, the reverse transcriptase can be provided in trans relative to the napDNAbp (rather than by way of a fusion protein).
[0085] In various embodiments involving methods, the napDNAbp can comprise the amino acid sequence of any one of SEQ ID NOS: 1361421 (Cas9). The napDNAbp can also comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOS: 1361421.
[0086] In various embodiments involving methods, the reverse transcriptase can comprise any one of the amino acid sequences of SEQ ID NOS: 1361485-1361514 and 1361597-1361598. The reverse transcriptase can also comprise an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOS: 1361485-1361514 and 1361597-1361598.
[0087] The method can involve using an extension RNA having a nucleotide sequence of SEQ ID NOS: 271029-406542 and 947841-1015218, or a nucleotide sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% sequence identity thereto.
[0088] The method can include using a prime editor guide RNA comprising an RNA extension at the 3’ end, wherein the RNA extension comprises a DNA synthesis template, e.g., FIG. 3B The PEgRNA (having the following components described from 5’ to 3’: spacer; gRNA core; reverse transcription template; primer binding site) shown in The method can involve using a prime editor guide RNA comprising an RNA extension at the 3’ end, wherein the RNA extension comprises a DNA synthesis template, e.g.,
[0089] The method can comprise using a prime editor guide RNA comprising an RNA extension at the 5’ end, wherein the RNA extension comprises a DNA synthesis template, e.g. FIG. 3A The PEgRNA (with the following components described from 5’ to 3’: reverse transcription template; primer binding site; linker; spacer; gRNA core) shown in
[0090] The method can comprise using a prime editor guide RNA comprising an RNA extension at an intramolecular position of the guide RNA, wherein the RNA extension comprises a DNA synthesis template.
[0091] The method can comprise using a prime editor guide RNA having one or more RNA extensions, the RNA extensions being at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.
[0092] It should be appreciated that the foregoing concepts and additional concepts discussed below can be arranged in any suitable combination, as the disclosure is not limited in this respect. Moreover, upon considering the following detailed description of various non-limiting embodiments in conjunction with the accompanying drawings, other advantages and novel features of the disclosure will become apparent to those skilled in the art. BRIEF DESCRIPTION OF DRAWINGS
[0094] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The application can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0095] FIG. 1A.1A schematic diagram is provided illustrating an exemplary process for introducing single nucleotide changes, insertions, and / or deletions into a DNA molecule (e.g., a genome) using a fusion protein comprising a reverse transcriptase fused to a napDNAbp (e.g., Cas9) protein and a guide RNA. In this embodiment, the guide RNA extends at the 3' end to include a DNA synthesis template. The schematic diagram illustrates how a reverse transcriptase (RT) fused to a Cas9 nickase and complexed with the guide RNA (gRNA) binds to a DNA target site and creates a nick on a PAM-containing DNA strand adjacent to the target nucleotide. The RT template uses the nicked DNA as a primer for synthesizing DNA from the gRNA, which serves as a template for synthesizing a new DNA strand encoding the desired edit. The editing process illustrated can be referred to as target-guided reverse transcription editing (guided editing). FIG. 1A.2 Provided with FIG. 1A.1 The same notation applies, except that the guide editor complex is more generally represented as [napDNAbp]-[P]:PEgRNAPEgRNA or [P]-[napDNAbp]:PEgRNAPEgRNA, where: "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid-programmable DNA-binding protein (e.g., SpCas9), and "PEgRNAPEgRNA" refers to the guide editing RNA, and "]-[" refers to an optional adapter. As described elsewhere, for example, FIG. 3A -3G, PEgRNA. PEgRNA includes a 5' extension arm containing a primer binding site and a DNA synthesis template. Although not shown, it is anticipated that the extension arm of PEgRNA (i.e., containing the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase covered in this conformation will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase example is an RNA-dependent DNA polymerase (e.g., reverse transcriptase). If the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase. In various embodiments, PEgRNA can be engineered or synthesized to incorporate a DNA-based DNA synthesis template.
[0096] FIG. 1B.1A schematic diagram is provided illustrating an exemplary process for introducing single nucleotide changes, insertions, and / or deletions into a DNA molecule (e.g., a genome) using a fusion protein containing a reverse transcriptase fused to nap DNAbp (e.g., Cas9) and a guide editor (gRNA). In this embodiment, the guide RNA is extended at the 5' end to include a DNA synthesis template. The schematic diagram shows how a reverse transcriptase (RT) fused to a Cas9 nicking enzyme and complexed with guide RNA (gRNA) binds to a DNA target site and creates a nick on a PAM-containing DNA strand adjacent to the target nucleotide. The canonical PAM sequence is 5'-NGG-3', but different PAM sequences can be associated with different Cas9 proteins or equivalent proteins from different organisms. Furthermore, any given Cas9 nuclease, such as SpCas9, can be modified to alter the protein's PAM specificity to recognize alternative PAM sequences. The RT enzyme uses the nicked DNA as a primer for synthesizing DNA from the gRNA, which serves as a template for synthesizing a new DNA strand encoding the desired edit. The editing process shown can be referred to as target-guided reverse transcription editing (TPRT editor or guide editor). FIG. 1B.2 Provided with FIG. 1B.1 The same notation applies, except that the guide editor complex is more generally represented as [napDNAbp]-[P]:PEgRNAPEgRNA or [P]-[napDNAbp]:PEgRNAPEgRNA, where: "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid-programmable DNA-binding protein (e.g., SpCas9), and "PEgRNAPEgRNA" refers to the guide editing RNA, and "]-[" refers to optional adapters. As described elsewhere, for example, FIG. 3A -3G, PEgRNA. PEgRNA includes a 3' extension arm containing a primer binding site and a DNA synthesis template. Although not shown, it is expected that the extension arm of PEgRNA (i.e., containing the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase covered in this conformation will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase example is an RNA-dependent DNA polymerase (e.g., reverse transcriptase). If the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase.
[0097] FIG. 1CThis is a schematic diagram illustrating an exemplary process of how a synthesized single-stranded DNA (containing the desired nucleotide changes) is broken down to incorporate the desired nucleotide changes, insertions, and / or deletions into the DNA. As shown, after the synthesis of the edited strand (or "mutated strand"), balancing with the endogenous strand, flap cleavage of the endogenous strand, and ligation leading to the resolution of mismatched DNA duplexes through the action of endogenous DNA repair and / or replication processes, DNA editing is incorporated.
[0098] FIG. 1D It shows that "anti-chain nick formation" can be incorporated. FIG. 1C This is a schematic diagram illustrating the degradation methods used to help drive the formation of the desired product against the formation of the reverse product. In reverse strand nick formation, a second napDNAbp / gRNA complex (e.g., a Cas9 / gRNA complex) is used to introduce a second nick from the initial nick strand onto the reverse strand. This induces endogenous cellular DNA repair and / or replication processes to preferentially replace the unedited strand (i.e., the strand containing the second nick site).
[0099] FIG. 1EAnother schematic of an exemplary process for using a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editor guide RNA (e.g., prime editing) for introducing at least one nucleotide change (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more), insertion, and / or deletion into a target locus of a DNA molecule (e.g., a genome) is provided. The prime editor guide RNA comprises an extension at the 3’ or 5’ end of the guide RNA or at an intramolecular location in the guide RNA. In step (a), the napDNAbp / gRNA complex contacts the DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) in one of the DNA strands of the target locus (R-loop strand, or PAM-containing strand, or non-target DNA strand, or pre-interval region strand), thereby creating an available 3’ end in one of the strands of the target locus. In certain embodiments, the nick is created in the DNA strand corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence. In step (c), the 3’ end DNA strand interacts with the extension portion of the guide RNA to direct reverse transcription. In some embodiments, the 3’ end DNA strand hybridizes to a specific primer binding site on the extension portion of the guide RNA. In step (d), a reverse transcriptase is introduced, which synthesizes single-stranded DNA from the 3’ end of the prime site towards the 3’ end of the guide RNA. This forms a single-stranded DNA flap comprising the desired nucleotide change (e.g., single or multiple base changes, insertions, deletions, or combinations thereof). In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve resolution of the single-stranded DNA flap, such that the desired nucleotide change is incorporated into the target locus. The process can be driven towards the desired product formation by removal of the corresponding 5’ endogenous DNA flap, which forms once the 3’ single-stranded DNA flap invades and hybridizes to the complementary sequence on the other strand. The process can be driven towards product formation with a second strand nick, as shown. The process can introduce at least one or more of the following genetic changes: transversions, transitions, deletions, and insertions. FIG. 1D
[0100] FIG. 1F is a schematic depicting the types of genetic changes that can be utilized in a prime editing (prime editing) process described herein. The types of nucleotide changes that can be achieved by prime editing include deletions (including short and long deletions), single and / or multiple nucleotide changes, and insertions (including short and long insertions).
[0101] FIG. 1G Figure 2 is a schematic depicting an example of a temporal second strand nicking by the guide editor complex example. The temporal second strand nicking formation is a variant of the second strand nicking formation to facilitate the formation of the desired editing product. The term "temporal" refers to the fact that the second strand nick of the unedited strand only occurs after the desired edit has been installed in the edited strand. This avoids concurrent nicks on both strands that can lead to double stranded DNA breaks.
[0102] FIG. 1HVariants of prime editing considered herein are depicted that replace the napDNAbp (e.g., SpCas9 nickase) with any programmable nuclease domain such as a zinc finger nuclease (ZFN) or a transcription activator-like effector nuclease (TALEN). Thus, it is contemplated that a suitable nuclease does not necessarily need to be “programmed” by a nucleic acid targeting molecule (e.g., a guide RNA), but rather can be programmed by defining the specificity of the DNA binding domain, such as especially the nuclease. Just as with prime editing with a napDNAbp moiety, it is preferred that such alternative programmable nucleases are modified so that only one strand of the target DNA is cleaved. In other words, the programmable nuclease should preferably function as a nickase. Once a programmable nuclease (e.g., ZFN or TALEN) is selected, then additional functionality can be engineered into the system to allow it to operate in a prime editing-like mechanism. For example, the programmable nuclease can be modified by coupling (e.g., via a chemical linker) an RNA or DNA extension arm to it, where the extension arm comprises a primer binding site (PBS) and a DNA synthesis template. The programmable nuclease can also be coupled (e.g., via a chemical or amino acid linker) to a polymerase, the nature of which will depend on whether the extension arm is DNA or RNA. In the case of an RNA extension arm, the polymerase can be an RNA-dependent DNA polymerase (e.g., a reverse transcriptase). In the case of a DNA extension arm, the polymerase can be a DNA-dependent DNA polymerase (e.g., a prokaryotic polymerase, including Pol I, Pol II, or Pol III, or a eukaryotic polymerase, including Pol a, Pol b, Pol g, Pol d, Pol e, or Pol z). The system can also include additional functionality added as a fusion to the programmable nuclease, or added in trans to facilitate the overall reaction (e.g., (a) a helicase to unwind the DNA at the cleavage site to make the cleaved strand with a 3’ end available as a primer, (b) a FEN1 to help remove the endogenous strand on the cleaved strand to drive the reaction towards replacement of the endogenous strand with the synthesized strand, or (c) a nCas9:gRNA complex to create a second site nick on the opposite strand, which can help drive integration of the synthesized repair by favored cellular repair of the non-edited strand). In a manner analogous to prime editing with a napDNAbp, such complexes with other programmable nucleases can be used for synthesis, and then the newly synthesized DNA carrying the edit of interest is permanently installed into the target site of the DNA.
[0103] FIG. 1IStructural features of a target DNA that can be edited by prime editing in one embodiment are depicted. The target DNA comprises a "non-target strand" and a "target strand." The target strand is the strand that anneals to the spacer of the PEgRNA of the prime editor complex that recognizes the PAM site (in this case, NGG, which is recognized by a canonical SpCas9-based prime editor). The target strand can also be referred to as the "non-PAM strand" or "non-editing strand." In contrast, the non-target strand (i.e., the strand comprising the protospacer and NGG PAM sequence) can be referred to as the "PAM strand" or "editing strand." In various embodiments, the nicking site of the PE complex will be in the protospacer on the PAM strand (e.g., for SpCas9-based PE). The location of the nick will be characteristic of the particular Cas9 forming the PE. For example, for SpCas9-based PE, the nicking site is in the phosphodiester bond between bases three ("-3" position relative to position 1 of the PAM sequence) and four ("-4" position relative to position 1 of the PAM sequence). The nicking site in the protospacer forms a free 3' hydroxyl, as shown in the figure below, which complexes with the primer binding site of the extension arm of the PEgRNA and provides a substrate to begin polymerizing single-stranded DNA encoding by the DNA synthesis template of the PEgRNA extension arm. This polymerization reaction is catalyzed by the polymerase (e.g., reverse transcriptase) of the PE fusion protein in the 5' to 3' direction. The polymerization is terminated before reaching the gRNA core (e.g., by inclusion of a polymerization termination signal or secondary structure, the function of which is to terminate the polymerization activity of the PE), resulting in a single-stranded DNA flap extending from the original 3' hydroxyl of the nicked PAM strand. The DNA synthesis template encodes single-stranded DNA that is homologous to the endogenous 5' end single-stranded DNA immediately following the nicking site on the PAM strand and incorporates the desired nucleotide change (e.g., single base substitution, insertion, deletion, inversion).The position of the desired edit can be anywhere downstream of the nick site on the PAM strand, which can include positions +1, +2, +3, +4 (start of the PAM site), +5 (position 2 of the PAM site), +6 (position 3 of the PAM site), +7, +8, +9, +10, +11, +12, +13, +14, +15, +16, +17, +18, +19, +20, +21, +22, +23, +24, +25, +26, +27, +28, +29, +30, +31, +32, +33, +34, +35, +36, +37, +38, +39, +40, +41, +42, +43, +44, +45, +46, +47, +48, +49, +50, +51, +52, +53, +54, +55, +56, +57, +58, +59, +60, +61, +62, +63, +64, +65, +66, +67, +68, +69, +70, +71, +72, +73, +74, +75, +76, +77, +78, +79, +80, +81, +82, +83, +84, +85, +86, +87, +88, +89, +90, +91, +92, +93, +94, +95, +96, +97, +98, +99, +100, +101, +102, +103, +104, +105, +106, +107, +108, +109, +110, +111, +112, +113, +114, +115, +116, +117, +118, +119, +120, +121, +122, +123, +124, +125, +126, +127, +128, +129, +130, +131, +132, +133, +134, +135, +136, +137, +138, +139, +140, +141, +142, +143, +144, +145, +146, +147, +148, +149, or +150 or more (relative to the downstream position of the nick site). Once the 3' terminal single-stranded DNA (containing the edit of interest) replaces the endogenous 5' terminal single-stranded DNA, the DNA repair and replication processes will result in the permanent installation of the edit on the PAM strand, and then the correction will exist on the non-PAM strand at the edit site. In this way, the edit will extend to both DNA strands on the target DNA site. It should be understood that the reference to "edit strand" and "non-edit" strand is only intended to depict the DNA strands involved in the PE mechanism. The "edit strand" is the strand that is first edited by replacing the 5' terminal single-stranded DNA immediately downstream of the nick site with a synthetic 3' terminal single-stranded DNA containing the desired edit. The "non-edit" strand is the strand that pairs with the edit strand, but it itself is also edited through repair and / or replication to be complementary to the edited strand, particularly the edit of interest.
[0104] FIG. 1J A mechanism of prime editing is depicted showing the target DNA, the prime editor complex, and structural features of the interactions between the PEgRNA and the target DNA. First, a prime editor comprising a fusion protein of a polymerase (e.g., a reverse transcriptase) and a napDNAbp (e.g., a SpCas9 nickase, e.g., SpCas9 with an inactivating mutation in the HNH nuclease domain (e.g., H840A) or an inactivating mutation in the RuvC nuclease domain (D10A)) complexes with a PEgRNA and a DNA with a target DNA to be edited. The PEgRNA comprises a spacer, a gRNA core (a.k.a. gRNA scaffold or gRNA backbone) (which binds to the napDNAbp), and an extension arm. The extension arm can be at the 3’ end, the 5’ end, or somewhere within the PEgRNA molecule. As shown, the extension arm is at the 3’ end of the PEgRNA. The extension arm comprises, in the 3’ to 5’ direction, a primer binding site and a DNA synthesis template (including the edit of interest) and a homology region (i.e., a homology arm) that is directly homologous to the 5’ end single-stranded DNA immediately following the nick site on the PAM strand. As shown, once the nick is introduced, creating a free 3’ hydroxyl immediately upstream of the nick site, the region immediately upstream of the nick site on the PAM strand anneals to the complementary sequence at the 3’ end of the extension arm, called the “primer binding site,” creating a short double-stranded region with a free 3’ hydroxyl end, which forms a substrate for the polymerase of the prime editing complex. The polymerase (e.g., reverse transcriptase) then polymerizes a DNA strand from the 3’ hydroxyl end to the end of the extension arm. The sequence of the single-stranded DNA is encoded by the DNA synthesis template, which is “read” by the polymerase to synthesize the extension arm portion of the new DNA (i.e., not including the primer binding site). This polymerization effectively extends the sequence of the original 3’ hydroxyl end of the initial nick site. The DNA synthesis template encodes a DNA single strand that not only contains the desired edit, but also a region that is homologous to the endogenous single-stranded DNA immediately downstream of the nick site on the PAM strand. Next, the encoded 3’ end DNA single strand (i.e., 3’ single-stranded DNA flap) displaces the corresponding homologous endogenous 5’ end DNA single strand immediately downstream of the nick site on the PAM strand, forming a DNA intermediate with a 5’ end single-stranded DNA flap, which is removed by the cell (e.g., by a flap endonuclease). The 3’ end single-stranded DNA flap, which anneals to the complement of the endogenous 5’ end single-stranded DNA flap, is ligated to the endogenous strand upon removal of the 5’ DNA flap. The desired edit in the now annealed and ligated 3’ end single-stranded DNA flap forms a mismatch with the complementary strand, which undergoes DNA repair and / or a round of replication, permanently installing the desired edit on both strands.
[0105] FIG. 2Three Cas complexes to be tested and their PAM, gRNA, and DNA cleavage characteristics are shown. The figure shows the design of complexes involving SpCas9, SaCas9, and LbCas12a.
[0106] FIG. 3A-3C Designs of engineered 5' extended gRNAs FIG. 3A , 3' extended gRNAs FIG. 3B , and intramolecular extensions FIG. 3C are shown, each of which can be used to guide editing. The embodiments depict an exemplary arrangement of the DNA synthesis template, primer binding site, and optional linker sequences in the extended portion of the 3', 5', and intramolecular extension gRNAs, as well as the arrangement of the protospacer and core regions. The disclosed TPRT process is not limited to these configurations of the prime editor guide RNAs.
[0107] FIG. 4A-4E An in vitro TPRT assay is demonstrated. FIG. 4A is a schematic of fluorescently labeled DNA substrate gRNA templated extension by RT enzyme, polyacrylamide gel electrophoresis (PAGE) analysis of reverse transcriptase product. FIG. 4B TPRT with pre-formed nicked substrates, dCas9, and 5'-extended gRNAs with different editing template lengths is shown. FIG. 4C RT reactions with pre-formed nicked DNA substrates in the absence of Cas9 are shown. FIG. 4D TPRT on full dsDNA substrates with Cas9 (H840A) and 5' extended gRNAs is shown. FIG. 4E 3' extended gRNA templates with pre-formed nicked and intact dsDNA substrates are shown. All reactions employ M-MLV RT.
[0108] FIG. 5 In vitro validation using 5' extended gRNAs with different lengths of editing templates is shown. Fluorescently labeled (Cy5) DNA targets were used as substrates and were pre-nicked in this set of experiments. The Cas9 used in these experiments was a catalytically dead Cas9 (dCas9) and the RT used was Superscript III, a commercially available RT derived from Moloney-murine leukemia virus (M-MLV). The dCas9:gRNA complex was formed from purified components. The fluorescently labeled DNA substrate was then added with dNTPs and the RT enzyme. After incubation at 37°C for 1 hour, the reaction products were analyzed by denaturing urea-polyacrylamide gel electrophoresis (PAGE). The gel image shows that the length of the extended DNA strand is consistent with the length of the reverse transcription template.
[0109] FIG. 6In vitro validation using 5' extended gRNAs with different length editing templates is shown, which are very similar to those shown in FIG. 5 However, in this set of experiments, the DNA substrate was not pre- nicked. The Cas9 used in these experiments was a Cas9 nickase (SpyCas9 H840A mutant), and the RT used was Superscript III, a commercially available RT derived from Moloney murine leukemia virus (M-MLV). Reaction products were analyzed by denaturing urea-polyacrylamide gel electrophoresis (PAGE). As shown in the gel, the nickase can efficiently cleave the DNA strand when using gRNAs (gRNA_0, lane 3).
[0110] FIG. 7 3' extension is shown to support DNA synthesis and does not significantly affect Cas9 nickase activity. When using dCas9 or Cas9 nickase, pre-nicked substrates (black arrow) are almost quantitatively converted to RT products (lanes 4 and 5). More than 50% conversion of RT products (red arrow) is observed with intact substrates (lane 3). Cas9 nickase (SpyCas9 H840A mutant), catalytically dead Cas9 (dCas9), Superscript III, a commercially available RT derived from Moloney murine leukemia virus (M-MLV) were used.
[0111] FIG. 8 Dual color experiments are shown to determine whether RT reactions preferentially occur with gRNAs in cis (bound in the same complex). Both 5'-extended and 3'-extended gRNAs were performed in two separate experiments. Products were analyzed by PAGE. Product ratios were calculated as (Cy3 cis / Cy3 trans) / (Cy5 trans / Cy5 cis).
[0112] FIG. 9A-9D Flap model substrates are shown. FIG. 9A Dual FP reporter for flap directed mutagenesis is shown. FIG. 9B Termination codon repair in HEK cells is shown. FIG. 9C Yeast clones sequenced after flap repair are shown. FIG. 9D Testing of different flap features in human cells is shown.
[0113] FIG. 10Directed editing on a plasmid substrate is shown. A dual fluorescent reporter plasmid was constructed for yeast (S. cerevisiae) expression. Expression of this construct in yeast only produces GFP. In vitro TRT reactions introduce point mutations and either the parental plasmid or in vitro Cas9 (H840A) nicked plasmid is transformed into yeast. Colonies are visualized by fluorescent imaging. Yeast dual FP plasmid transformants are shown. Transformation of the parental plasmid or in vitro Cas9 (H840A) nicked plasmid only produces green GFP expressing colonies. TRT reactions with 5'-extended or 3'-extended gRNAs produce a mixture of green and yellow colonies. The latter express both GFP and mCherry. More yellow colonies are observed with 3'-extended gRNAs. A positive control without a stop codon is also shown.
[0114] FIG. 11 Directed editing on a plasmid substrate is shown, similar to the experiment in FIG. 10 , but instead of installing a point mutation in the stop codon, directed editing installs a single nucleotide insertion (left) or deletion (right) that repairs the frameshift mutation and allows synthesis of downstream mCherry. Both experiments use 3' extended gRNAs.
[0115] FIG. 12 The editing products of directed editing on a plasmid substrate are shown, characterized by Sanger sequencing. Individual colonies from TRT transformations are selected and analyzed by Sanger sequencing. The precise edits are observed by sequencing the selected colonies. Green colonies contain plasmids with the original DNA sequence, while yellow colonies contain the precise mutations designed by the directed editing gRNAs. No other point mutations or indels are observed.
[0116] FIG. 13 The potential scope of the new directed editing technology is shown and compared to deaminase-mediated base editor technology.
[0117] FIG. 14 A schematic of editing in human cells is shown.
[0118] FIG. 15 The extension of primer binding sites in gRNAs is shown.
[0119] FIG. 16 Truncated gRNAs for adjacent targeting are shown.
[0120] FIG. 17A-17C A graph showing the % T to A conversion at the target nucleotide after transfection of components in human embryonic kidney (HEK) cells is shown. FIG. 17A Data is shown presenting results using wild type MLV reverse transcriptase fused to the N-terminus of Cas9 (H840A) nickase (32 amino acid linker).FIG. 17B Similar to FIG. 17A But C-terminal fusion to RT enzyme. FIG. 17C Similar to FIG. 17A But linker between MLV RT and Cas9 is 60 amino acids long instead of 32 amino acids.
[0121] FIG. 18 Shows high purity T to A editing at HEK3 locus by high throughput amplicon sequencing. Output of sequencing analysis shows the most abundant genotype of edited cells.
[0122] FIG. 19 Shows editing efficiency at target nucleotides (left bar of each pair) as well as indel rate (right bar of each pair). WT refers to wild type MLV RT enzyme. Mutant enzymes (M1 to M4) contain mutations listed on the right. Editing rate was quantified by high throughput sequencing of genomic DNA amplicons.
[0123] FIG. 20 Shows editing efficiency of target nucleotides when introducing a single strand nick in the complementary DNA strand adjacent to the target nucleotide. Nicks at different distances from the target nucleotide were tested (orange triangles). Editing efficiency of the target base pair (blue bars) is shown together with indel formation rate (orange bars). The "none" example does not contain a guide RNA with a nicked complementary strand. Editing rate was quantified by high throughput sequencing of genomic DNA amplicons.
[0124] FIG. 21 Demonstrates processed high throughput sequencing data showing the expected T to A transversion mutation and the general absence of other major genomic editing byproducts.
[0125] FIG. 22A schematic of an exemplary process for targeted mutagenesis at a target locus using a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editor guide RNA with error-prone reverse transcriptase is provided. This process can be referred to as an embodiment of prime editing for targeted mutagenesis. The prime editor guide RNA comprises an extension at the 3’ or 5’ end of the guide RNA or at an intramolecular location in the guide RNA. In step (a), the napDNAbp / gRNA complex is contacted with a DNA molecule, and the gRNA directs the napDNAbp to bind to the target locus to be mutagenized. In step (b), a nick is introduced in one of the DNA strands at the target locus (e.g., by a nuclease or chemical agent), thereby creating an available 3’ end in one of the strands at the target locus. In certain embodiments, the nick is created in the DNA strand corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence. In step (c), the 3’ end DNA strand interacts with the extended portion of the guide RNA to prime reverse transcription. In some embodiments, the 3’ end DNA strand hybridizes to a specific primer binding site on the extended portion of the guide RNA. In step (d), an error-prone reverse transcriptase is introduced, which synthesizes a mutagenized single strand of DNA from the 3’ end of the priming site to the 3’ end of the guide RNA. Exemplary mutations are indicated with asterisks “*”. This forms a single-stranded DNA flap comprising the desired mutagenized region. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve resolution of the single-stranded DNA live flap (comprising the mutagenized region), such that the desired mutagenized region is incorporated into the target locus. This process can be driven toward the desired product formation by removal of the corresponding 5’ endogenous DNA flap, which forms once the 3’ single-stranded DNA flap invades and hybridizes to the complementary sequence on the other strand. This process can also be driven toward product formation by a second strand nick formation, as shown in FIG. 1D After endogenous DNA repair and / or replication processes, the mutagenized region is incorporated into both DNA strands of the DNA locus.
[0126] FIG. 23is a schematic of gRNA design for tri-nucleotide repeat sequence reduction and tri- nucleotide repeat reduction using TPRT genome editing. Tri-nucleotide repeat expansions are associated with many human diseases, including Huntington's disease, fragile X syndrome, and Friedreich's ataxia. The most common tri-nucleotide repeats contain CAG triplets, but GAA triplets (Friedreich's ataxia) and CGG triplets (fragile X syndrome) also occur. Inheritance of an expanded predisposition, or acquisition of a parent allele that has already expanded, increases the likelihood of acquiring the disease. Pathogenic expansion of tri-nucleotide repeats can hypothetically be corrected using prime editing. The region upstream of the repeat region can form a nick using an RNA-guided nuclease, which is then used to prime synthesis of a new DNA strand containing a healthy number of repeats (depending on the specific gene and disease). Following the repeat sequence, a small stretch of homology is added that matches the identity of the sequence adjacent to the other end of the repeat (red strand). Invasion of the newly synthesized strand, and subsequent replacement of the endogenous DNA with the newly synthesized flap, results in a reduced repeat allele.
[0127] FIG. 24 is a schematic showing a precise 10 nucleotide deletion using prime editing. The guide RNA targeting the HEK3 locus was designed with a reverse transcription template that encodes a 10 nucleotide deletion following the nicking site. Amplicon sequencing was used to assess editing efficiency in HEK cells transfected.
[0128] FIG. 25is a schematic showing gRNA design for peptide-tagging genes at endogenous genomic loci and peptide-tagging using TPRT genome editing. The FlAsH and ReAsH tagging systems comprise two parts: (1) fluorophore-bisarsenic probes, and (2) gene-encoded peptides containing a four-cysteine motif, exemplified by the sequence FLNCCPGCCMEP (SEQ ID NO: 1361586). When expressed in cells, proteins containing a four-cysteine motif can be fluorescently labeled with fluorophore-arsenic probes (see reference: J. Am. Chem. Soc, 2002, 124(21), pp 6063-6076. DOI: 10.1021 / ja017687n). The "sortagging" system employs bacterial sortases to covalently bind labeled peptide probes to proteins containing suitable peptide substrates (see reference: Nat. Chem. Biol. 2007 Nov;3(11):707-8. DOI: 10.1038 / nchembio.2007.31). FLAG tag (DYKDDDDK (SEQ ID NO: 1361587)), V5 tag (GKPIPNPLLGLDST (SEQ ID NO: 1361588)), GCN4 tag (EELLSKNYHLENEVARLKK (SEQ ID NO: 1361589)), HA tag (YPYDVPDYA (SEQ ID NO: 1361590)), and Myc tag (EQKLISEEDL (SEQ ID NO: 1361591)) are commonly employed as epitope tags for immunoassays. The pi-clamp encoding peptide sequence (FCPF (SEQ ID NO: 1361592)) which can be labeled with a pentafluoroaromatic substrate (reference: Nat. Chem. 2016 Feb;8(2):120-8. doi: 10.1038 / nchem.2413).
[0129] FIG. 26 The precise installation of His6 tag and FLAG tag in genomic DNA is shown. Guide RNAs targeting the HEK3 locus were designed with reverse transcription templates that encode either an 18-nt His tag insertion or a 24-nt FLAG tag insertion. Amplicon sequencing was used to assess editing efficiency of transfected HEK cells. Note that the full 24-nt sequence of the FLAG tag is viewed outside the frame (sequencing confirms full and precise insertion).
[0130] FIG. 27The structure of embodiments of PEgRNAs contemplated herein is provided and can be designed according to the methods defined in Example 2. A PEgRNA comprises three main component elements arranged in a 5’ to 3’ direction, namely: a spacer, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in a 5’ to 3’ direction, namely: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, a PEgRNA can comprise an optional 3’ end modifier region (el) and an optional 5’ end modifier region (e2). Still further, a PEgRNA can comprise a transcription termination signal (not depicted) at the 3’ end of the PEgRNA. These structural elements are further defined herein. The depiction of the PEgRNA structure is not meant to be limiting, but rather to encompass variations in the arrangement of elements. For example, the optional sequence modifiers (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends. In certain embodiments, a PEgRNA PEgRNA can comprise secondary RNA structures such as, but not limited to, hairpins, stems / loops, toe loops, RNA binding protein recruiting domains (e.g., MS2 aptamer that recruits and binds the MS2 cp protein). For example, such secondary structures can be located within the spacer, gRNA core, or extension arm, particularly within the el and / or e2 modifier regions. In addition to secondary RNA structures, a PEgRNA PEgRNA can comprise (e.g., within the el and / or e2 modifier regions) a chemical linker or a poly(N) linker or tail, where “N” can be any nucleobase. In some embodiments (e.g., as shown in FIG. 72(c)), the chemical linker can function to prevent reverse transcription of the sgRNA scaffold or core. Furthermore, in certain embodiments (e.g., see FIG. 72(c)), the extension arm (3) can comprise RNA or DNA, and / or can include one or more nucleobase analogs (e.g., which can add functionality such as temperature elasticity). Still further, the orientation of the extension arm (3) can be the natural 5’ to 3’ direction, or synthesized in the opposite direction of 3’ to 5’ (relative to the orientation of the entire PEgRNA PEgRNA molecule). It is also noted that one of ordinary skill in the art will be able to select an appropriate DNA polymerase for the guided editing depending on the nature of the nucleic acid material (i.e., DNA or RNA) of the extension arm, which can be implemented as a fusion with the napDNAbp, or provided in trans as a separate moiety to synthesize the desired template-encoded 3’ single-stranded DNA flap including the desired edit. For example, if the extension arm is RNA, then the DNA polymerase can be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the extension arm is DNA, then the DNA polymerase can be a DNA-dependent DNA polymerase.In various embodiments, the provision of the DNA polymerase can be in trans, for example by using an RNA-protein recruitment domain (e.g., installed on the PEgRNA (e.g., in the el or e2 region, or elsewhere, and the MS2cp protein is fused to the DNA polymerase, thereby co-localizing the DNA polymerase to the PEgRNA). It is also noted that the primer binding site generally does not form part of the template used by the DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3’ single-stranded DNA flap including the desired edit. Thus, the designation “DNA synthesis template” refers to the region or portion of the PEgRNA that is used by the DNA polymerase as a template to encode the extension arm (3) of the desired 3’ single-stranded DNA flap containing the edit. In some embodiments, the DNA synthesis template includes the “editing template” and the “homology arm.” In other embodiments, the DNA synthesis template can also include the e2 region or portions thereof. For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, the function of the DNA polymerase can terminate before any portion of the e2 region is actually encoded into the DNA. Some or even all of the e2 region will also have the potential to be encoded into the DNA. How much of e2 actually serves as a template will depend on its constitution and whether that constitution interrupts DNA polymerase function.
[0131] FIG. 28Another embodiment of the PEgRNA contemplated herein is provided, and it can be designed according to the methods defined in Example 2. The PEgRNA comprises three main component elements arranged in a 5’ to 3’ direction, namely: a spacer region, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in a 5’ to 3’ direction, namely: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, the PEgRNA can comprise an optional 3’ end modifier region (el) and an optional 5’ end modifier region (e2). Still further, the PEgRNA can comprise a transcription termination signal (not depicted) on the 3’ end of the PEgRNA. These structural elements are further defined herein. The depiction of the PEgRNA structure is not meant to be limiting, but rather to encompass variations in the arrangement of the elements. For example, the optional sequence modifiers (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends. In certain embodiments, the PEgRNA PEgRNA can comprise secondary RNA structures such as, but not limited to, hairpins, stems / loops, toe loops, RNA binding protein recruiting domains (e.g., MS2 aptamer that recruits and binds the MS2 cp protein). These secondary structures can be located anywhere in the PEgRNA PEgRNA molecule. For example, such secondary structures can be located within the spacer region, the gRNA core, or the extension arm, particularly within the el and / or e2 modifier regions. In addition to secondary RNA structures, the PEgRNA PEgRNA can comprise (e.g., within the el and / or e2 modifier regions) a chemical linker or a poly(N) linker or tail, where “N” can be any nucleobase. In some embodiments (e.g., as shown in FIG. 27 Example 2), the chemical linker can function to prevent reverse transcription of the sgRNA scaffold or core. In addition, in certain embodiments (e.g., see FIG. 28), the extension arm (3) can comprise RNA or DNA, and / or can include one or more nucleobase analogs (e.g., which can add functionality, such as temperature elasticity). Still further, the orientation of the extension arm (3) can be the natural 5’ to 3’ direction, or synthesized in the opposite direction of 3’ to 5’ (relative to the orientation of the entire PEgRNA PEgRNA molecule). It should also be noted that one of ordinary skill in the art will be able to select an appropriate DNA polymerase for directed editing, which can be implemented as a fusion with the napDNAbp, or provided in trans as a separate moiety, depending on the nature of the nucleic acid material (i.e., DNA or RNA) of the extension arm, to synthesize the desired template-encoded 3’ single-stranded DNA flap including the desired edit. For example, if the extension arm is RNA, the DNA polymerase can be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the extension arm is DNA, the DNA polymerase can be a DNA-dependent DNA polymerase. In various embodiments, the provision of the DNA polymerase can be in trans, for example by using an RNA-protein recruitment domain (e.g., installed on the PEgRNA PEgRNA (e.g., in the el or e2 region, or elsewhere, and the MS2cp protein is fused to the DNA polymerase, thereby co-localizing the DNA polymerase to the PEgRNA PEgRNA). It should also be noted that the primer binding site generally does not form part of the template used by the DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3’ single-stranded DNA flap including the desired edit. Thus, the designation “DNA synthesis template” refers to the region or portion of the extension arm (3) that is used as a template by the DNA polymerase to encode the desired 3’ single-stranded DNA flap containing the edit. In some embodiments, the DNA synthesis template includes the “edit template” and the “homology arm.” In other embodiments, the DNA synthesis template can also include the e2 region or a portion thereof. For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, the function of the DNA polymerase can terminate before any portion of the e2 region is actually encoded into the DNA. Some or even all of the e2 region will also have the potential to be encoded into the DNA. How much of the e2 is actually used as a template will depend on its composition and whether that composition interrupts DNA polymerase function.
[0132] FIG. 29is a schematic depicting the interaction of a typical PEgRNA with a double stranded DNA target site and the concomitant production of a 3' single stranded DNA flap containing the genetic change of interest. The double stranded DNA is shown as the top strand (i.e., the target strand) oriented 3' to 5' and the lower strand (i.e., the PAM strand or non-target strand) oriented 5' to 3'. The top strand contains the complement of the "protospacer" and the complement of the PAM sequence, which is referred to as the "target strand". Because it is the strand to which the PEgRNA's spacer targets, and which it anneals to the PEgRNA's spacer. The complementary lower strand is referred to as the "non-target strand" or "PAM strand" or "protospacer strand" because it contains the PAM sequence (e.g., NGG) and the protospacer. Although not shown, the depicted PEgRNA will be complexed with a Cas9 or equivalent. The domains of the guide editor fusion protein. As shown in the schematic, the PEgRNA's spacer anneals to the complementary region of the protospacer on the target strand, which is referred to as the protospacer, which is located just downstream of the PAM sequence, which is approximately 20 nucleotides in length. This interaction forms a DNA / RNA hybrid between the complementary sequences of the spacer RNA and the protospacer DNA, and induces the formation of an R-loop at the region opposite the protospacer. As taught elsewhere herein, the Cas9 protein (not shown) then induces a nick in the non-target strand, as shown. This then results in the formation of a 3' ssDNA flap region immediately upstream of the nick site, which interacts with the 3' end of the PEgRNA at the primer binding site according to *z*. The 3' end of the ssDNA flap (i.e., the reverse transcriptase primer sequence) anneals to the primer binding site on the PEgRNA (A), thereby priming the reverse transcriptase. Next, the reverse transcriptase (e.g., provided in trans or provided in cis as a fusion protein, attached to the Cas9 construct) then polymerizes a DNA strand encoded by the DNA synthesis template (comprising the edit template (B) and the homology arm (C)). The polymerization continues towards the 5' end of the extension arm. The polymerized strand of ssDNA forms a ssDNA 3' end flap, as described elsewhere (e.g., as shown in FIG. 1E the endogenous DNA, displaces the corresponding endogenous strand (which is removed as a 5' DNA flap of the endogenous DNA), and installs the desired nucleotide edit (single nucleotide base pair change, deletion, insertion (including entire genes)) through the naturally occurring DNA repair / replcation wheel.
[0133] FIG. 30The disclosure of PEgRNAs is aided by understanding the sequence table. The figure shows two exemplary PEgRNA sequences (SEQ ID NO: 135529 (top) and SEQ ID NO: 135880 (bottom)) and how various disclosed subsets of sequences are positioned thereon. For SEQ ID NO: 135529, the corresponding sequences are the spacer (SEQ ID NO: 271043), the extension arm (SEQ ID NO: 406557), the primer binding site (SEQ ID NO: 542071), the editing template (SEQ ID NO: 677585), and the homology arm (SEQ ID NO: 813099). For SEQ ID NO: 135880, the corresponding sequences are the spacer (SEQ ID NO: 880463), the extension arm (SEQ ID NO: 947841), the primer binding site (SEQ ID NO: 1015219), the editing template (SEQ ID NO: 1082597), and the homology arm (SEQ ID NO: 1149975).
[0134] FIG. 31 is a flowchart showing an exemplary high-level computerized method 3100 for determining extended gRNA structures according to some embodiments of the present disclosure. At step 3102, a computing device (e.g., in conjunction with the computing device 3400 described FIG. 34 The computing device 3400 described accesses data indicative of an input allele, an output allele, and a fusion protein comprising a nucleic acid programmable DNA binding protein and a reverse transcriptase. While step 3102 describes accessing all three of the input allele, the output allele, and the fusion protein in one step, this is for illustrative purposes, and it should be understood that such data can be accessed using one or more steps without departing from the spirit of the technology described herein. Accessing data can include receiving data, storing data, accessing a database, etc.
[0135] FIG. 32 is a flowchart showing an exemplary computerized method 3200 for determining components of an extended gRNA structure (including extended components) according to some embodiments. It should be understood that, FIG. 32 is intended to be illustrative, thus, the techniques for determining an extended gRNA can include more or fewer steps than those shown in FIG. 32
[0136] FIG. 33 is a flowchart showing an exemplary computerized method 3300 for determining a set of extended gRNA structures for each mutation entry in a database, according to some embodiments. At step 3302, a computing device accesses a database comprising a set of mutation entries (e.g., the ClinVar database, which is accessible at www.ncbi.nlm.nih.gov / clinvar / ), the mutation entries each comprising an input allele representing a mutation and an output allele representing a wildtype sequence.
[0137] FIG. 34 is an illustrative implementation of a computer system 3400 that can be used to perform any aspect of the technology and embodiments disclosed herein. The computer system 3400 can include one or more processors 3410 and one or more non-transitory computer-readable storage media (e.g., a memory 3420 and one or more non-volatile storage media 3430) and a display 3440. The processor(s) 3410 can control writing data to and reading data from the memory 3420 and the non-volatile storage device 3430 in any suitable manner, as the aspects of the application described herein are not limited in this respect.
[0138] FIG. 35A is a schematic of PE-based insertion of sequences encoding RNA motifs related to Example 3.
[0139] FIG. 35B is a list (not exhaustive) of some example motifs that can potentially be inserted and their functions related to Example 3.
[0140] FIG. 36 Bar graphs comparing the efficiency (i.e., “percent of total sequencing reads with the indicated edit or indel”) of PE2, PE2-trunc, PE3, and PE3-trunc on different target sites in various cell lines are provided. The data show that guide editors comprising truncated RT variants are roughly as efficient as guide editors comprising non-truncated RT proteins.
[0141] FIG. 37A Nucleotide sequences of SpCas9 PE gRNA molecules are shown (top), which terminate at the 3’ end in “UGU” and do not contain a toe-loop element. The bottom half of the figure depicts the same SpCas9 PE gRNA molecule, but further modified to contain a toe-loop element with the sequence 5’-“GAAANNNNN”-3’ inserted before the “UUU” 3’ end. “N” can be any nucleobase.
[0142] FIG. 37BResults from Example 4 are shown, which demonstrate that the use of PE gRNAs containing toe-loop elements improves the efficiency of prime editing in HEK cells or EMX cells, while the percentage of indel formation is essentially unchanged.
[0143] FIG. 38 One embodiment of a prime editor provided as two PE half-proteins is depicted, which regenerate into a full prime editor through the action of split-intein halves at the end or beginning of each prime editor half-protein that self-splice to form a full functional intein, which then undergoes self-splicing and excision.
[0144] FIG. 39 depicts the mechanism of removing an intein from a polypeptide sequence and reforming a peptide bond between N-terminal and C-terminal extein sequences. (a) depicts the general mechanism of two half-proteins, each containing half of an intein sequence, which when brought into contact within a cell, produce a full functional intein, which then undergoes self-splicing and excision. The excision process results in the formation of a peptide bond between the N-terminal protein half (or “N-extein”) and the C-terminal protein half (or “C-extein”) to form a complete single polypeptide comprising N-extein and C-extein portions. In various embodiments, the N-extein can correspond to the N-terminal half of a split prime editor fusion protein, and the C-extein can correspond to the C-terminal half of a split prime editor. (b) shows the chemical mechanics of intein excision and reformation of a peptide bond joining the N-extein half (red half) and the C-extein half (blue half). Excision of a split intein (i.e., N-intein and C-intein in split intein configuration) can also be referred to as “trans-splicing” because it involves the splicing action of two separate components provided in trans.
[0145] DEFINITIONS
[0146] Antisense strand
[0147] In genetics, the “anti-sense” strand of a segment within double-stranded DNA is the template strand, and is considered to run in the 3’ to 5’ direction. In contrast, the “sense” strand is the segment of double-stranded DNA that runs from 5’ to 3’, which is complementary to the anti-sense or template strand of DNA (from 3’ to 5’). In the case of a DNA segment that encodes a protein, the sense strand is the DNA strand that has the same sequence as the mRNA, which is templated by the anti-sense strand during transcription, and ultimately translated (usually, not always) into a protein. Thus, the anti-sense strand is responsible for generating the RNA that is later translated into a protein, while the sense strand has nearly the same composition as the mRNA. Note that for each segment of dsDNA, there can be two sets of sense and anti-sense, depending on the direction of reading (since sense and anti-sense are relative to the perspective). It is ultimately the gene product or mRNA that determines which strand of a dsDNA segment is called sense or anti-sense.
[0148] Cas9
[0149] The term“Cas9” or“Cas9 nuclease” refers to a protein comprising a Cas9 domain or fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or a gRNA binding domain of Cas9). As used herein, a“Cas9 domain” is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A“Cas9 protein” is a full-length Cas9 protein. Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. CRISPR is a system of adaptive immunity that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacer regions, sequences complementary to antecedent mobile elements, and targets invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer regions. The target strand non-complementary to the crRNA is first endonucleolytically cleaved, then 3’-5’ exonucleolytically trimmed. In nature, DNA binding and cleavage typically require a protein and two RNAs. However, a single guide RNA (“sgRNA”, or simply“gNRA”) can be engineered so as to integrate aspects of both the crRNA and the tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012)), which is incorporated by reference herein in its entirety. Cas9 recognizes a short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar EZ., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471 :602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus.Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0150] A nuclease-inactivated Cas9 domain is interchangeably referred to as a "dCas9" protein (for nuclease- "dead" Cas9). Methods for generating Cas9 domains (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 includes two subdomains, an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant." A Cas9 variant has homology to Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421). In some embodiments, a Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421).In some embodiments, a Cas9 variant comprises a fragment of SEQ ID NO: X Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% in amino acid length to the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421).
[0151] cDNA
[0152] The term "cDNA" refers to a DNA strand copied from an RNA template. The cDNA is complementary to the RNA template.
[0153] Circular permutant
[0154] As used herein, the term "cyclopermutant" refers to a protein or polypeptide (e.g., Cas9) comprising a cyclopermutant, which is an alteration in the protein structure configuration involving a change in the order of amino acids that occur in the protein sequence. In other words, a cyclopermutant is a protein with altered N- and C-termini compared to the wild-type counterpart, e.g., the wild-type C-terminal half of the protein becomes the new N-terminal half. A cyclopermutant (or CP) is essentially a topological rearrangement of the primary sequence of a protein, typically using a peptide linker to connect its N and C termini, while splitting its sequence at different positions to create new adjacent N and C-termini. The result is a protein structure with different connectivity, but can often have the same overall similar three-dimensional (3D) shape, and can include improved or altered characteristics, including reduced proteolytic susceptibility, increased catalytic activity, altered substrate or ligand binding, and / or improved thermal stability. Cyclopermutant proteins can occur in nature (e.g., concanavalin A and lectins). In addition, cyclopermutations can occur as a result of post-translational modifications, or can be engineered using recombinant techniques.
[0155] Circular permutant Cas9
[0156] The term "circularly permuted Cas9" refers to any Cas9 protein or variant thereof that occurs in a circularly permuted form whereby by rearrangement of the primary sequence of the protein, its N- and C-termini have been reconfigured. Such circularly permuted Cas9 proteins ("CP-Cas9") or variants thereof retain the ability to bind DNA when complexed with a guide RNA (gRNA). See Oakes et al., "Protein Engineering of Cas9 for enhanced function," Methods Enzymol, 2014, 546:491-511 and Oakes et al., "CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification," Cell, January 10, 2019, 176:254-267, each incorporated herein by reference. The present disclosure contemplates any previously known CP-Cas9 or use of a new CP-Cas9 so long as the resulting circularly permuted protein retains the ability to bind DNA when complexed with a guide RNA (gRNA). Exemplary CP-Cas9 proteins are SEQ ID NOS: 1361475-1361484.
[0157] DNA synthesis template
[0158] As used herein, the term "DNA synthesis template" refers to the region or portion of the extension arm of a PEgRNA that is used as a template strand by the polymerase of the prime editor to encode the 3' single-stranded DNA flap that contains the desired edit, which is then substituted for the corresponding endogenous DNA strand at the target site by the prime editing mechanism. In various embodiments, the DNA synthesis template is shown in FIG. 3A (in the case of a PEgRNA comprising a 5' extension arm), FIG. 3B (in the case of a PEgRNA comprising a 3' extension arm), FIG. 3C(3) can include all or part of the 3' extension arm (3) that extends from the 5' end of the primer binding site (PBS) to the 3' end of the primer binding site (PBS). The gRNA core can serve as a template for the polymerase (e.g., reverse transcriptase) to synthesize single-stranded DNA. In the case of a 5' extension arm, the DNA synthesis template (3) can include a portion of the extension arm (3) that spans from the 5' end of the PEgRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template does not include the primer binding site (PBS) of the PEgRNA with either a 3' extension arm or a 5' extension arm. Certain embodiments described herein (e.g., FIG. 71A) refer to an "RT template" which includes the editing template and the homology arm, i.e., the sequence synthesis of the PEgRNA extension arm that is actually used as a template in the DNA process. The term "RT template" is equivalent to the term "DNA synthesis template."
[0159] Downstream
[0160] As used herein, the terms "upstream" and "downstream" are relative terms that define position in a 5 '-to-3' direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is located somewhere 5' of the second element. For example, if a SNP is located 5' of a nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is located somewhere 3' of the second element. For example, if a SNP is located 3' of a nick site, the SNP is downstream of the Cas9-induced nick site. The nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. Analysis of single-stranded nucleic acid molecules and double-stranded molecules is the same because the upstream and downstream terms refer only to the single strand of the nucleic acid molecule, it is just a matter of choosing which strand of a double-stranded molecule is being considered. Generally, the strand of double-stranded DNA that can be used to determine the positional relationship of at least two elements is the "sense" or "coding" strand. In genetics, the "sense" strand is the segment of double-stranded DNA that runs from 5' to 3' and is complementary to the DNA's anti-sense or template strand (from 3' to 5'). Thus, for example, if a SNP nucleobase is 3' of a promoter on the sense or coding strand, the SNP nucleobase is "downstream" of the promoter sequence in genomic DNA (double-stranded).
[0161] CRISPR
[0162] CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent fragments of previous infections by viruses that have invaded prokaryotes. Prokaryotic cells use the DNA fragments to detect and destroy DNA from similar viruses to protect against subsequent attacks by similar viruses and effectively synthesize a range of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNA, a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), an endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA acts as a guide for ribonuclease 3 to assist in processing the pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the RNA. Specifically, the target strand that is not complementary to the crRNA is first endonucleolytically cleaved, and then 3’-5’ exonucleolytic trimming. In nature, DNA binding and cleavage typically requires a protein and two RNAs. However, a single guide RNA (“sgRNA,” or simply “gNRA”) can be engineered so as to integrate aspects of both the crRNA and the tracrRNA into one single RNA species - the guide RNA. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), which is incorporated by reference herein in its entirety. Cas9 recognizes a short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.CRISPR biology as well as Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471 :602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to S. pyogenes and S. thermophilus.Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0163] The term "Cas9" or "Cas9 nuclease" refers to a protein comprising a Cas9 domain or fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or a gRNA binding domain of Cas9). As used herein, a "Cas9 domain" is a fragment of a protein comprising an active or inactive cleavage domain of Cas9 and / or a gRNA binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. CRISPR is a system of adaptive immunity that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacer regions, sequences complementary to preceding mobile elements, and targets invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer regions. The target strand non-complementary to the crRNA is first endonucleolytically cleaved, then 3'-5' exonucleolytically trimmed. In nature, DNA binding and cleavage typically require a protein and two RNAs. However, a single guide RNA ("sgRNA", or simply "gRNA") can be engineered to integrate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471 :602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including, but not limited to, S. pyogenes and S. thermophilus.Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0164] A nuclease-inactivated Cas9 domain is interchangeably referred to as a "dCas9" protein (for nuclease- "dead" Cas9). Methods for generating Cas9 domains (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 includes two subdomains, an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant." A Cas9 variant has homology to Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421). In some embodiments, a Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421).In some embodiments, the Cas9 variant comprises a fragment of SEQ ID NO: 1361421 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% in amino acid length to the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 1361421).
[0165] Edit template
[0166] The term "editing template" refers to the portion of the extended arm that encodes the desired edit in the single-stranded 3' DNA flap synthesized by the polymerase, e.g., DNA-dependent DNA polymerase, RNA-dependent DNA polymerase (e.g., reverse transcriptase). Certain embodiments described herein (e.g., FIG. 71A) refer to "RT template," which refers to both the editing template and the homology arm, i.e., the sequence synthesis of the PEgRNA extended arm that is actually used as a template in the DNA process. The term "RT editing template" is also equivalent to the term "DNA synthesis template," but where the RT editing template reflects the use of a primary editor with a reverse transcriptase polymerase, where the DNA synthesis template more broadly reflects the use of a primary editor with any polymerase.
[0167] Error prone
[0168] As used herein, the term "error-prone" reverse transcriptase (or more broadly, any polymerase) refers to a reverse transcriptase (or more broadly, any polymerase) that is naturally occurring or derived from another reverse transcriptase (e.g., wild-type M-MLV reverse transcriptase) that has an error rate less than that of the wild-type M-MLV reverse transcriptase. The error rate of wild-type M-MLV reverse transcriptase has been reported to range from 15,000 (higher) to 27,000 (lower). An error rate of 1 in 15,000 corresponds to an error rate of 6.7 x 10 - An error rate of 1 in 27,000 corresponds to an error rate of 3.7 x 10 -5error rate. See Boutabout et al. (2001) "DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Ty1," Nucleic Acids Res 29(11):2217-2222, incorporated herein by reference. Thus, for the purposes of this application, the term "error-prone" refers to those RTs with an error rate greater than 1 error in 15,000 nucleobases incorporated (6.7 x 10 -5 or higher), 1 error in 14,000 nucleobases incorporated (7.14 x 10 -5 or higher), 1 error in 13,000 or fewer nucleobases incorporated (7.7 x 10 -5 or higher), 1 error in 12,000 or fewer nucleobases incorporated (7.7 x 10 -5 or higher), 1 error in 11,000 or more nucleobases incorporated (9.1 x 10 -5 or higher), 1 error in 10,000 or fewer nucleobases incorporated (1 x 10 -4 or 0.0001 or higher), 1 error in 9,000 or fewer nucleobases incorporated (0.00011 or higher), 1 error in 8,000 or fewer nucleobases incorporated (0.00013 or higher), 1 error in 7,000 or fewer nucleobases incorporated (0.00014 or higher), 1 error in 6,000 or fewer nucleobases incorporated (0.00016 or higher), 1 error in 5,000 or fewer nucleobases incorporated (0.0002 or higher), 1 error in 4,000 or fewer nucleobases incorporated (0.00025 or higher), 1 error in 3,000 nucleobases or fewer (0.00033 or higher), 1 error in 2,000 nucleobases or fewer (0.00050 or higher), or 1 error in 1,000 or fewer nucleobases (0.001 or higher), or 1 error in 500 or fewer nucleobases (0.002 or higher), or 1 error in 250 or fewer nucleobases (0.004 or higher).
[0169] Extension arm
[0170] The term "extension arm" refers to a nucleotide sequence component of a PEgRNA that provides multiple functions, including a primer binding site and an editing template for a reverse transcriptase. In some embodiments, for example, in FIG. 3D, the extension arm is located at the 3' end of the guide RNA. In other embodiments, for example, in FIG. 3E, the extension arm is located at the 5' end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In various embodiments, the extension arm comprises, in the 5' to 3' direction, the following components: a homology arm, an editing template, and a primer binding site. Because the polymerization activity of a reverse transcriptase is in the 5' to 3' direction, the preferred arrangement of the homology arm, editing template, and primer binding site is in the 5' to 3' direction, such that the reverse transcriptase, once primed by the annealed primer sequence, uses the editing template as a complementary template strand to polymerize single-stranded DNA. Further details, such as the length of the extension arm, are described elsewhere herein.
[0171] The extension arm can also be described as generally comprising two regions: a primer binding site (PBS) and a DNA synthesis template, as shown in FIG. 3G (top). When the primer binding site is cleaved by the prime editor complex, the primer binding site binds to a primer sequence formed by the endogenous DNA strand of the target site, thereby exposing the 3' end on the endogenous nicked strand. As described herein, the binding of the primer sequence to the primer binding site on the PEgRNA extension arm creates a duplex region with the exposed 3' end (i.e., of the primer sequence) that then provides a substrate for a polymerase to polymerize single-stranded DNA along the length of the DNA synthesis template, starting from the exposed 3' end. The sequence of the single-stranded DNA product is the complement of the DNA synthesis template. Polymerization continues in the 5' direction of the DNA synthesis template (or extension arm) until polymerization terminates. Thus, the DNA synthesis template represents a portion of the extension arm that is encoded by the polymerase of the prime editor complex into a single-stranded DNA product (i.e., a 3' single-stranded DNA flap containing the desired genetic edit information) and ultimately replaces the corresponding endogenous DNA strand of the target site located downstream of the PE-induced nick site. Without being bound by theory, polymerization of the DNA synthesis template continues toward the 5' end of the extension arm until a termination event. Polymerization can terminate in a variety of ways, including but not limited to (a) reaching the 5' end of the PEgRNA (e.g., in the case of a 5' extension arm, where the DNA polymerase simply runs out of template), (b) reaching an insurmountable RNA secondary structure (e.g., a hairpin or stem loop), or (c) reaching a replication termination signal, such as a specific nucleotide sequence that blocks or inhibits the polymerase, or a nucleic acid topological signal, such as supercoiled DNA or RNA.
[0172] Effective amount
[0173] As used herein, the term "effective amount" refers to the amount of a biologically active agent that is sufficient to elicit a desired biological response. For example, in some embodiments, an effective amount of an initiating editor can refer to an amount of an editor that is sufficient to edit a nucleotide sequence of a target site, e.g., a genome. In some embodiments, an effective amount of an initiating editor provided herein, e.g., an effective amount of a fusion protein comprising a nickase Cas9 domain and a reverse transcriptase, can refer to an amount of the fusion protein that is sufficient to induce editing of a target site to which the fusion protein specifically binds. And edited by the fusion protein. Those skilled in the art will appreciate that an effective amount of an agent, e.g., a fusion protein, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide, can vary depending on various factors, e.g., depending on the desired biological response, e.g., a particular allele, genome, or target site to be edited, the cell or tissue targeted, and the agent used.
[0174] Functional equivalent
[0175] The term "functional equivalent" refers to a second biological molecule that is functionally equivalent to a first biological molecule but not necessarily structurally equivalent. For example, a "Cas9 equivalent" refers to a protein that has the same or substantially the same function as Cas9 but not necessarily the same amino acid sequence. In the context of the present disclosure, the specification refers to "protein X or a functional equivalent thereof" throughout. In this context, a "functional equivalent" of protein X includes any homolog, paralog, fragment, naturally occurring, engineered, mutated, or synthetic form of protein X that has equivalent function.
[0176] Fusion protein
[0177] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein can be located in the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) protein of the fusion protein, forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. A protein can comprise different domains, for example, a nucleic acid binding domain (e.g., a gRNA binding domain of Cas9 that directs the protein to bind to a target site) and a nucleic acid cleavage domain or a nucleic acid catalytic domain. Acid editing protein. Another example includes Cas9 or its equivalent of reverse transcriptase. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising a peptide linker. Methods of recombinant protein expression and purification are well known, including those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0178] Gene product
[0179] As used herein, the term "gene product" refers to any product encoded by a nucleic acid sequence. Thus, a gene product can be, for example, a primary transcript, a mature transcript, a processed transcript, or a protein or peptide encoded by a transcript. Examples of gene products thus include mRNAs, rRNAs, tRNAs, hairpin RNAs, microRNAs (miRNAs), shRNAs, siRNAs, as well as peptides and proteins, e.g., reporter proteins or therapeutic proteins.
[0180] Gene of interest (GOI)
[0181] The term "gene of interest" or "GOI" refers to a gene that encodes a biomolecule of interest (e.g., a protein or an RNA molecule). Proteins of interest can include any intracellular protein, membrane protein, or extracellular protein, such as a nuclear protein, transcription factor, nuclear membrane transport protein, intracellular organelle-associated protein, membrane receptor, catalytic protein and enzyme, therapeutic protein, membrane protein, membrane transport protein, signal transduction protein, or immune protein (e.g., IgG or other antibody protein), and the like. Genes of interest can also encode RNA molecules, including but not limited to messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), antisense RNA, guide RNA, microRNA (miRNA), small interfering RNA (siRNA), and cell-free RNA (cfRNA).
[0182] Guide RNA ("gRNA")
[0183] As used herein, the term "guide RNA" is a specific type of guide nucleic acid that is typically associated with the Cas protein of CRISPR-Cas9 and is associated with Cas9 to direct the Cas9 protein to a specific sequence comprising a DNA molecule that is complementary to the original spatial sequence of the guide RNA. As described elsewhere, a PEgRNA is a subcategory of guide RNA that further comprises an extension arm at the 3' or 5' end of the guide, enabling the molecule to be used with the primary editors disclosed herein. The term "guide RNA" also includes equivalent guide nucleic acid molecules that are associated with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and otherwise program Cas9 equivalents to locate to a particular target nucleotide sequence. Cas9 equivalents can include other napDNAbps from any type of CRISPR system (e.g., Type II, V, VI), including Cpf1 (Type V CRISPR-Cas system), C2cl (Type V CRISPR-Cas system), C2c2 (Type VI CRISPR-Cas system), and C2c3 (Type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated by reference herein. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are provided herein. As used herein, "guide RNA" can also be referred to as "traditional guide RNA" to contrast it with the modified form of guide RNA referred to as "primary editor guide RNA" (or "PEgRNA"), which has been invented for use with the primary editing methods and compositions disclosed herein.
[0184] A guide RNA or PEgRNA can comprise various structural elements, including but not limited to:
[0185] Spacer sequence - a sequence (having a length of about 10 to about 40 (e.g., about 10, about 15, about 20, about 25, about 30) nucleotides) in a guide RNA or PEgRNA that binds to a protospacer in a target DNA (as defined below).
[0186] gRNA core (or gRNA backbone or backbone sequence) - refers to the sequence within a gRNA that is responsible for napDNAbp (e.g., Cas9) binding that does not include the spacer / targeting sequence used to direct the napDNAbp (e.g., Cas9) to target DNA.
[0187] Extension arm - the portion of the extension arm that encodes a portion of the resulting reverse transcriptase encoded single-stranded DNA flap that will integrate into the target DNA site by replacement of the endogenous strand. The portion of the single-stranded DNA flap encoded by the extension arm is complementary to the non-edited strand of the target DNA sequence, which facilitates replacement of the endogenous strand and annealing of the single-stranded DNA flap in its place, thereby installing the edit. This component is further defined elsewhere.
[0188] Homology arm - the portion of the extension arm that encodes a portion of the resulting reverse transcriptase encoded single-stranded DNA flap that will integrate into the target DNA site by replacement of the endogenous strand. The portion of the single-stranded DNA flap encoded by the homology arm is complementary to the non-edited strand of the target DNA sequence, which facilitates replacement of the endogenous strand and annealing of the single-stranded DNA flap in its place, thereby installing the edit. This component is further defined elsewhere.
[0189] Edit template - the portion of the extension arm that encodes the desired edit in the reverse transcriptase synthesized single-stranded DNA flap. This component is further defined elsewhere.
[0190] Primer binding site - the portion of the extension arm that anneals to a primer sequence, which is formed from the target DNA strand that is nicked by Cas9 mediated nuclease action. This component is further defined elsewhere.
[0191] Transcriptional terminator - the guide RNA or PEgRNA can contain a transcriptional termination sequence at the 3' of the molecule. Typically the transcriptional terminator sequence (e.g., SEQ ID NOs: 1361560-1361565) is about 70 to about 125 nucleotides in length, but shorter and longer transcriptional terminator sequences are also contemplated and any sequence known in the art can be used.
[0192] Flap endonuclease (e.g., FEN1)
[0193] As used herein, the term "flap endonuclease" refers to enzymes that catalyze the removal of 5' single-stranded DNA flaps. These are naturally occurring enzymes that carry out the removal of 5' flaps formed during cellular processes, including DNA replication. The prime editing methods described herein can utilize flap endonucleases provided endogenously or in trans to remove 5' flaps of endogenous DNA formed at the target site during prime editing. Flap endonucleases are known in the art and can be found described in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10): 4507-4519 and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2): 198-211, and Balakrishnan et al., "Flap Endonuclease 1," Annu Rev Biochem, 2013, Vol 82: 119-138 (each of which is incorporated herein by reference). An exemplary flap endonuclease is FEN1, which can be represented by the following amino acid sequence:
[0194]
[0195] Fusion protein
[0196] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) protein of the fusion protein, forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. A protein can comprise different domains, such as a nucleic acid binding domain (e.g., a gRNA binding domain of Cas9 that directs the protein to bind to a target site) and a nucleic acid editing protein (e.g., an RT domain) and a nucleic acid cleavage domain (e.g., Cas9 nickase, napDNAbp) or catalytic domain. Another example includes a napDNAbp (e.g., RNA Cas9) or an equivalent thereof fused to a reverse transcriptase. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0197] Homology arm
[0198] The term "homology arm" refers to the portion of the extension arm that includes the sequence of the single-stranded DNA flap encoded by the reverse transcriptase that is produced thereby that will integrate into the target DNA site by replacement of the endogenous strand. The portion of the single-stranded DNA flap encoded by the homology arm is complementary to the unedited strand of the target DNA sequence, which facilitates replacement of the endogenous strand and annealing of the single-stranded DNA flap in its place, thereby installing the edit. This component is further defined elsewhere.
[0199] Host cell
[0200] As used herein, the term "host cell" refers to a cell that can serve as a host, replicate, and express a vector described herein, such as a vector comprising a nucleic acid molecule encoding a fusion protein comprising a napDNAbp or napDNAbp equivalent (e.g., Cas9 or equivalent) and a reverse transcriptase.
[0201] Isolated
[0202] "Isolated" means removed from the natural state or environment. For example, a nucleic acid or peptide naturally present in a living animal is not "isolated," but the same nucleic acid or peptide partially or completely removed from the natural co-matcrials of its environment is "isolated." An isolated nucleic acid or protein can exist in substantially purified form, or can exist in a non-native environment such as a host cell.
[0203] In some embodiments, a gene of interest is encoded by an isolated nucleic acid. As used herein, the term "isolated" refers to the property of a material as provided herein removed from its original or natural environment (e.g., the natural environment if it is naturally occurring). Thus, a naturally occurring polynucleotide or protein or polypeptide that is present in a living animal is not isolated, but the same polynucleotide or polypeptide that is separated from some or all of the coexisting materials of the natural system by the hand of man is isolated. Thus, an artificial or engineered material, such as a non-naturally occurring nucleic acid construct, such as an expression construct and vector described herein, is also therefore referred to as isolated. The material need not be purified to be isolated. Thus, the material can be part of a vector and / or part of a composition, and still be isolated in that such vector or composition is not part of the environment in which the material is found in nature.
[0204] nanDNAbp
[0205] As used herein, the term "nucleic acid programmable DNA binding protein" or "napDNAbp", with Cas9 as an example, refers to a protein that uses RNA:DNA hybridization to target and bind to specific sequences in DNA molecules. Each napDNAbp is associated with at least one guide nucleic acid (e.g., a guide RNA) that localizes the naDNAbp to a DNA sequence comprising a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid or portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp (e.g., Cas9 or equivalent) to locate and bind to the complementary sequence.
[0206] Without being bound by theory, the binding mechanism of napDNAbp-guide RNA complexes generally includes a step of forming an R-loop, whereby the napDNAbp induces the unwinding of the double-stranded DNA target, separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the "target strand." This displaces the "non-target strand," which is complementary to the target strand, forming a single-stranded region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cleave the DNA leaving various types of lesions. For example, the napDNAbp can comprise a nuclease activity that cleaves the non-target strand at a first location, and / or cleaves the target strand at a second location. Depending on the nuclease activity, the target DNA can be cleaved to form a "double-strand break," cleaving both strands. In other embodiments, the target DNA can be cleaved at only a single site, i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with different nuclease activities include "Cas9 nickases" ("nCas9") and inactive Cas9s without nuclease activity ("dead Cas9" or "dCas9"). Exemplary sequences of these and other napDNAbps are provided herein.
[0207] Linker
[0208] As used herein, the term "linker" refers to a molecule that connects two other molecules or moieties. Linkers are well known in the art and can comprise any suitable nucleic acid or amino acid combination to facilitate proper function of the structure they connect. A linker can be a series of amino acids. In the case of a linker connecting two fusion proteins, the linker can be an amino acid sequence. For example, a napDNAbp (e.g., Cas9) can be fused to a reverse transcriptase through an amino acid linker sequence. In the case of linking two nucleotide sequences together, a linker can also be a nucleotide sequence. For example, in the current case, a traditional guide RNA is linked to the RNA extension of a prime editor guide RNA through a spacer or linker nucleotide sequence, which can comprise a DNA synthesis template (e.g., an RT template sequence) and a primer binding site. In other embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, a linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, a linker is 5-100 nucleotides in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, 150-200, 200-300, 300-500, 500-1000, 1000-2000, or 2000-5000 nucleotides in length. Longer or shorter linkers can also be contemplated.
[0209] Nicking enzyme
[0210] The term "nickase" refers to a napDNAbp (e.g., Cas9) in which one of the two nuclease domains is inactivated. Such an enzyme is capable of cutting only one strand of the target DNA.
[0211] Nuclear localization sequence (NLS)
[0212] The term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that facilitates import of a protein into the nucleus, for example, by nuclear transport factors. Nuclear localization sequences are known in the art and will be apparent to the skilled person. For example, nuclear localization sequences are described in Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, published as WO 2001 / 038547 on May 31, 2001, the disclosure of exemplary nuclear localization sequences thereof incorporated herein by reference. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 1361531) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 1361533).
[0213] Nucleic acid molecule
[0214] As used herein, the term "nucleic acid" refers to a polymer of nucleotides (i.e., more than one (e.g., 2, 3, 4, etc.) nucleotides. The polymer can include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2- aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 5-methyladenosine, 5- methyluridine, 5-methylcytidine, 5-hydroxyluridine, dimethyluridine, methyltrityluridine, 1- methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabino-nose, and hexose), or modified phosphate groups (e.g., phosphorothioates and 5'-N phosphoramidite linkages).
[0215] Nucleobase
[0216] As used herein, the term "nucleoside base", also known as "nitrogenous base" or often simply "base", is a nitrogenous biological compound that forms a nucleoside, which in turn is a component of a nucleotide, with all these monomers constituting the basic building blocks of nucleic acids. The ability of nucleoside bases to form base pairs and stack with one another directly leads to long chain helical structures, such as ribonucleic acids (RNA) and deoxyribonucleic acids (DNA).
[0217] Five nucleobases, which are adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), can be referred to as primary or canonical. They function as the basic units of the genetic code, with bases A, G, C, and T present in DNA, and A, G, C, and U present in RNA. Thymine and uracil are identical except that T includes a missing methyl group of U. DNA and RNA can also contain modified nucleobases. For example, for the adenosine and guanosine nucleobases, alternative nucleobases can include hypoxanthine, xanthine, or 7-methylguanine, which correspond to alternative nucleobases for inosine, xanthosine, and 7-methylguanosine, respectively. Further, for example, for the cytosine, thymine, or uridine nucleobases, alternative nucleobases can include 5,6 dihydro uracil, 5-methylcytosine, or 5-hydroxymethylcytosine, which correspond to alternative nucleobases for dihydrouridine, 5-methylcytidine, and 5-hydroxymethylcytidine, respectively. Nucleobases can also include nucleobase analogs, of which a large number are known in the art. In general, analogous nucleobases confer, among other things, different base pairing and base stacking properties. Examples include universal bases, which can pair with all four canonical bases, and phosphate-sugar backbone analogs, such as PNA, which affect the properties of the chain (PNA can even form triple helices). Nucleic acid analogs are also referred to as "xenonucleic acids" and represent one of the major pillars of exobiology, the design of life forms based on alternative biochemistry to that of nature. Artificial nucleic acids include peptide nucleic acid (PNA), morpholino, and locked nucleic acid (LNA), as well as glycol nucleic acid (GNA) and threose nucleic acid (TNA). Each of these is distinguished from naturally occurring DNA or RNA by a change in the molecular backbone. Example analogs are (e.g., 2- aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5- methylcytidine, C5 bromouridine, C5 fluoro uridine, C5 iodo uridine, C5 propynyl uridine, C5 propynyl methyl cytidine, C57 deazoadenosine, 7 deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 0(6) methyl guanine, 4-acetylcytidine, 5-(carboxyhydroxylmethyl) uridine, dihydrouridine, methyl pseudouridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, thymidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-0-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N phosphoramidate linkages).
[0218] PEgRNA
[0219] As used herein, the term "prime editor guide RNA" or "PEgRNA" or "extended guide RNA" refers to a specialized form of a guide RNA that has been modified to include one or more additional sequences for use in the prime editing methods, compositions, and systems described herein. As described herein, a prime editor guide RNA comprises one or more "extension regions" of nucleic acid sequence. The extension region(s) can include, but are not limited to, single-stranded RNA. In addition, the extension region(s) can occur at the 3' end of a traditional guide RNA. In other arrangements, the extended regions can occur at the 5' end of a traditional guide RNA. In other arrangements, the extension region(s) can occur at an internal region of the traditional guide RNA rather than one of the ends, e.g., in the gRNA core region that associates and / or binds with the napDNAbp. The extension region(s) comprise a "reverse transcriptase template sequence" which is a single-stranded RNA molecule that encodes a single-stranded complementary DNA (cDNA) that in turn is designed to (a) be homologous to the endogenous target DNA to be edited and (b) comprise at least one desired nucleotide change (e.g., transition, transversion, deletion, insertion, or combination thereof) to be introduced or integrated into the endogenous target DNA. The extension region(s) can also comprise other functional sequence elements such as, but not limited to, a "primer binding site" and / or a "spacer or linker" sequence. As used herein, a "primer binding site" comprises a sequence that hybridizes to a single-stranded DNA sequence that has a 3' end generated from the nicked DNA of the R-loop and comprises a primer for a reverse transcriptase.
[0220] In some embodiments, the PEgRNA is represented by FIG. 3A which shows a PEgRNA having a 5' extension arm, a spacer region, and a gRNA core. The 5' extension further comprises, in the 5' to 3' direction, a reverse transcriptase template, a primer binding site, and a linker.
[0221] In some embodiments, the PEgRNA is represented by FIG. 3B which shows a PEgRNA having a 3' extension arm, a spacer region, and a gRNA core. The 3' extension further comprises, in the 5' to 3' direction, a reverse transcriptase template, and a primer binding site.
[0222] In other embodiments, the PEgRNA is represented by FIG. 27denotes a PEgRNA having a spacer (1), a gRNA core (2), and an extension arm (3) in the 5' to 3' direction. The extension arm (3) is located at the 3' end of the PEgRNA. The extension arm (3) further comprises, in the 5' to 3' direction, a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) can also comprise optional modifier regions at the 3' and 5' ends, which can be the same sequence or different sequences. In addition, the 3' end of the PEgRNA can comprise a transcriptional terminator sequence. These sequence elements of the PEgRNA are further described and defined herein. In addition, the specification discloses exemplary PEgRNAs in the accompanying sequence listing, which are designed according to the methods disclosed herein.
[0223] In still other embodiments, the PEgRNA is of the form FIG. 28 denotes a PEgRNA having an extension arm (3), a spacer (1), and a gRNA core (2) in the 5' to 3' direction. The extension arm (3) is located at the 5' end of the PEgRNA. The extension arm (3) further comprises, in the 3' to 5' direction, a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) can also comprise optional modifier regions at the 3' and 5' ends, which can be the same sequence or different sequences. The PEgRNA can also comprise a transcriptional terminator sequence at the 3' end. These sequence elements of the PEgRNA are further described and defined herein.
[0224] Peptide tag
[0225] The term "peptide tag" refers to a peptide amino acid sequence that is genetically fused to a protein sequence to confer one or more functions to the protein that facilitate manipulation of the protein for various purposes, such as visualization, identification, localization, purification, solubilization, separation, etc. Peptide tags can include various types of tags that are classified by purpose or function, which can include "affinity tags" (to facilitate protein purification), "solubilization tags" (to assist in proper folding of the protein), "chromatography tags" (to alter the chromatographic properties of the protein), "epitope tags" (to bind to high affinity antibodies), "fluorescent tags" (to facilitate visualization of the protein in cells or in vitro).
[0226] PE1
[0227] As used herein, “PE1” refers to a PE complex comprising a fusion protein comprising Cas9(H840A) and a wild-type MMLV RT having the structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)]+ the required PEgRNA, where the PE fusion has the amino acid sequence of SEQ ID NO: 1361515, as shown below;
[0228]
[0229] Key:
[0230] Nuclear localization sequence (NLS) top : (SEQ ID NO: 1361532), Bottom: (SEQ ID NO: 1361541)
[0231] Cas9(H840A) (SEQ ID NO: 1361454)
[0232] 33-amino acid linker (SEQ ID NO: 1361528)
[0233] M-MLV Reverse Transcriptase (SEQ ID NO: 1361485).
[0234] PE2
[0235] As used herein, “PE2” refers to a PE complex comprising a fusion protein comprising Cas9(H840A) and a MMLV RT variant having the structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+ the required PEgRNA, where the PE fusion has the amino acid sequence of SEQ ID NO: 1361516, as shown below:
[0236]
[0237] Key:
[0238] Nuclear localization sequence (NLS) top : (SEQ ID NO: 1361532), Bottom: (SEQ ID NO: 1361541)
[0239] Cas9(H840A) (SEQ ID NO: 1361454)
[0240] 33-amino acid linker (SEQ ID NO: 1361528)
[0241] M-MLV reverse transcriptase (SEQ ID NO: 1361514).
[0242] PE3
[0243] As used herein, "PE3" refers to PE2 plus a second strand nick guide RNA that complexes with PE2 and introduces a nick in the non-edited DNA strand to induce preferential replacement of the edited strand.
[0244] PE3b
[0245] As used herein, "PE3b" refers to PE3, but where the second strand nick guide RNA is designed for temporal control such that the second strand nick is not introduced until after the desired edit is installed. This is achieved by designing the gRNA with one spacer sequence that only matches the edited strand and not the original allele. Using this strategy (hereafter PE3b), the mismatch between the protospacer and the unedited allele should disfavor nicking by the sgRNA until after the editing event on the PAM strand has occurred.
[0246] PE-short
[0247] As used herein, "PE-short" refers to a PE construct fused to a C-terminally truncated reverse transcriptase with the following amino acid sequence:
[0248]
[0249] Key:
[0250] Nuclear localization sequence (NLS) top: (SEQ ID NO: 1361532), bottom: (SEQ ID NO: 1361541)
[0251] Cas9 (H840A) (SEQ ID NO: 1361454)
[0252] 33-amino acid linker 1 (SEQ ID NO: 1361528)
[0253] M-MLV truncated reverse transcriptase
[0254] (SEQ ID NO: 1361597)
[0255] Percent identity
[0256] "Percent identity," "sequence identity," "% identity," or "% sequence identity" of a sequence (e.g., nucleic acid or amino acid), as they can be used interchangeably herein, is a measure of the similarity between two sequences (e.g., nucleic acid or amino acid). The percent identity of genomic DNA sequences, intron and exon sequences, and amino acid sequences between humans and other species varies by species type, with chimpanzee and human having the highest percent identity in each category. Percent identity can be determined using the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as in Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul et al., J. Mol. Biol. 215:403-10, 1990. BLAST protein searches can be performed with the XBLAST program, score = 50, wordlength = 3, to obtain amino acid sequences homologous to protein molecules of interest. Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997, where appropriate. When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used. When stating or referencing a percent identity or a range thereof (e.g., at least, greater than, between, etc.), unless otherwise specified, the endpoints are to be included and ranges (e.g., at least 70% identity) are to include all ranges within the referenced range (e.g., at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95.5%, at least 96%, at least 96.5%, at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% identity) and all increments thereof (e.g., tenths of a percent (i.e., 0.1%), hundredths of a percent (i.e., 0.01%), etc.).
[0257] Prime editor
[0258] The term "prime editor" refers to a fusion construct described herein comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase and capable of prime editing a target nucleotide sequence in the presence of a PEgRNA (or "prime editing guide RNA"). The term "prime editor" can refer to the fusion protein or the fusion protein in complex with a PEgRNA. In some embodiments, a prime editor can also refer to a complex comprising the fusion protein (reverse transcriptase fused to a napDNAbp), a PEgRNA, and a conventional guide RNA capable of directing a second site nicking step of the non-edited strand, as described herein. In certain embodiments, the reverse transcriptase component of the "prime editor" is provided in trans.
[0259] Primer binding site
[0260] The term "primer binding site" or "PBS" refers to a nucleotide sequence on a PEgRNA (typically at the 3' end of the extension arm) that serves as an extension arm component and binds to a primer sequence formed after a target sequence is cleaved by a napDNAbp (e.g., Cas9) of a prime editor. As detailed elsewhere, when the Cas9 nickase component of a prime editor cleaves one strand of a target DNA sequence, a 3' ended ssDNA flap is formed that anneals to the primer binding site on the PEgRNA as a primer sequence to prime reverse transcription. FIG. 27 and 28 show embodiments of primer binding sites on the 3' and 5' extension arms, respectively.
[0261] Protein, peptide, and polypeptide
[0262] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of whatever size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by the addition of a chemical entity, such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, and the like. A protein, peptide, or polypeptide can also be a single molecule or can be a multimeric complex. A protein, peptide, or polypeptide can be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. Any protein provided herein can be produced by any method known in the art. For example, a protein provided herein can be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising peptide linkers. Methods of recombinant protein expression and purification are well known, including those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), which is incorporated by reference herein in its entirety.
[0263] Operably linked
[0264] As used herein, the term "operably linked" refers to a functional linkage between a regulatory sequence and a heterologous nucleic acid sequence (e.g., a transgene), resulting in expression of the heterologous nucleic acid sequence (e.g., a transgene). For example, a first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is in a functional relationship with the second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Generally, operably linked nucleic acid sequences are contiguous and, where necessary to join two protein coding regions, in the same reading frame.
[0265] Promoter
[0266] The term "promoter" is art-recognized and refers to a nucleic acid molecule having a sequence recognized by the transcription machinery of a cell and capable of controlling the transcription of a downstream gene. Promoters can be constitutively active, meaning that the promoter is always active in a given cellular environment, or conditionally active, meaning that the promoter has activity only in the presence of a particular condition. For example, a conditional promoter can have activity only in the presence of a particular protein that links the protein associated with a regulatory element in the promoter to the basic transcription machinery, or only in the absence of an inhibitory molecule. One subset of conditionally active promoters are inducible promoters, which require the presence of a small molecule "inducer" for activity. Examples of inducible promoters include, but are not limited to, arabinose inducible promoters, Tet-on promoters, and tamoxifen inducible promoters. A skilled artisan is well-versed in a variety of constitutive, conditional, and inducible promoters, and will be able to determine a variety of such promoters for use in practicing the present application, which is not limited in this respect.
[0267] Protospacer adjacent motif (PAM)
[0268] As used herein, the term "protospacer adjacent motif" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is an important targeting component of Cas9 nucleases. Typically, the PAM sequence is on either strand and is downstream of the 5' to 3' direction of the Cas9 cleavage site. A standard PAM sequence (i.e., the PAM sequence associated with the Cas9 nuclease of Streptococcus pyogenes or SpCas9) is 5'-NGG-3', where "N" is any nucleobase, followed by two guanine ("G") nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms, for example, 5'-NG-3', where "N" is any nucleobase, followed by one guanine ("G") nucleobase, or 5'-KKH-3', where two lysine ("K") are followed by one histidine ("H"). Furthermore, any given Cas9 nuclease, such as SpCas9, can be modified to change the PAM specificity of the nuclease, such that the nuclease recognizes an alternative PAM sequence.
[0269] For example, with reference to the reference standard SpCas9 amino acid sequence SEQ ID NO: 1361421 (SpCas9 M1 QQ99ZW2 wild type), the PAM sequence can be modified by introducing one or more mutations including (a) D1135V, R1335Q, and T1337R “VQR variant” which changes PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q, and T1337R “EQR variant” which changes PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E, and T1337R “VRER variant” which changes PAM specificity to NGCG. In addition, the D1135E variant of the standard SpCas9 can still recognize NGG, but is more selective compared to the wild type SpCas9 protein.
[0270] It is also understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilus (StCas9) recognizes NNAGAAW. In another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These examples are not meant to be limiting. It is further understood that non-SpCas9 bind to a variety of PAM sequences, which makes them useful when a suitable SpCas9 PAM sequence is not present at the desired target cleavage site. In addition, non-SpCas9 can have other features that make them more useful than SpCas9. For example, Cas9 from Staphylococcus aureus (SaCas9) is about 1 kb smaller than SpCas9 and thus can be packaged into an adeno-associated virus (AAV). Further reference can be made to Shah et al., “Protospacer recognition motifs: mixed identities and functional diversity,” RNA Biology, 10(5):891-899 (incorporated herein by reference).
[0271] Protospacer
[0272] As used herein, the term "protospacer" refers to the sequence (~20 bp) in DNA adjacent to the PAM (protospacer adjacent motif) sequence that has the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the original spacer sequence on the target DNA (specifically, one of its strands, the "target strand" to the "non-target strand" of the target sequence). To function, Cas9 also requires a specific protospacer adjacent motif (PAM), which varies by bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease is derived from S. pyogenes, which recognizes a PAM sequence of NGG on the non-target strand downstream of the genomic DNA target sequence. The skilled artisan will appreciate that the literature in the art sometimes refers to the "protospacer" as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than as the "spacer." Thus, in some instances, the term "protospacer" as used herein can be used interchangeably with the term "spacer." The context of the description in which "protospacer" or "spacer" appears will help inform the reader as to whether the term is refined to the gRNA or the DNA target. Both uses of these terms are acceptable, as the art uses both terms in each of these ways.
[0273] Reverse transcriptase
[0274] The term "reverse transcriptase" describes a class of polymerases characterized by RNA-dependent DNA polymerase activity. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptases have been used primarily to transcribe mRNA into cDNA, which can then be cloned into vectors for further manipulation. Avian myeloblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473: 1 (1977)). The enzyme has 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand of an RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Reverse transcriptases cannot correct transcriptional errors because known viral reverse transcriptases lack the 3'-5' exonuclease activity required for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of AMV reverse transcriptase activity and its associated RNase H activity is presented by Berger et al., Biochemistry 22: 2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is reverse transcriptase derived from Moloney murine leukemia virus (M-MLV). See, e.g., Gerard, G.R., DNA 5: 271-279 (1986) and Kotewicz, M.L., et al., Gene 35: 249-258 (1985). M-MLV reverse transcriptases that are essentially lacking RNase H activity have also been described. See, e.g., U.S. Patent No. 5,244,797. The present invention contemplates the use of any such reverse transcriptases, or variants or mutants thereof.
[0275] Further, the present disclosure encompasses the use of error-prone reverse transcriptases, i.e., that can be referred to as error-prone reverse transcriptases or reverse transcriptases that do not support high-fidelity nucleotide incorporation during polymerization. During the process of single-stranded DNA flap synthesis based on a RT template integrated with a guide RNA, an error-prone reverse transcriptase can introduce one or more nucleotides that are mismatched to the DNA synthesis template (e.g., the RT template sequence), thereby introducing an error polymerization through the single-stranded DNA flap to alter the nucleotide sequence. These errors introduced during single-stranded DNA flap synthesis are then propagated through hybridization to the corresponding endogenous target strand, removal of the endogenous displacement strand, ligation, and then through multiple rounds of endogenous DNA repair and / or replication.
[0276] Reverse transcription
[0277] As used herein, the term “reverse transcription” refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, reverse transcription can be “error-prone reverse transcription,” which refers to the property of certain reverse transcriptases that are error-prone in their DNA polymerization activity.
[0278] Sense strand
[0279] In genetics, the “sense” strand is the piece of double-stranded DNA from 5’ to 3’ that is complementary to the antisense or template strand of DNA, from 3’ to 5’. In the case of a DNA segment that encodes a protein, the sense strand is the DNA strand that has the same sequence as the mRNA, which is templated by the antisense strand during transcription and eventually undergoes (usually, not always) translation into a protein. Thus, the antisense strand is responsible for generating the RNA that is later translated into a protein, while the sense strand has nearly the same composition as the mRNA. Note that for each piece of dsDNA, there can be two sets of sense and antisense, depending on the direction of reading (since sense and antisense are relative to the perspective). It is the end product or mRNA that determines which strand of a piece of dsDNA is called sense or antisense.
[0280] In the context of a PEgRNA, the first step is to synthesize a single-stranded complementary DNA (i.e., a 3’ ssDNA flap, which is incorporated) oriented in the 5’ to 3’ direction, which is templated off of the PEgRNA extension arm. Whether the 3’ ssDNA flap should be considered the sense strand or the antisense strand depends on the direction of transcription, as it is generally accepted that both DNA strands can serve as templates for transcription (but not simultaneously). Thus, in some embodiments, the 3’ ssDNA flap (which extends generally in the 5’ to 3’ direction) will be the sense strand, as it is the coding strand. In other embodiments, the 3’ ssDNA flap (which extends generally in the 5’ to 3’ direction) will be the antisense strand and thus serve as the transcription template.
[0281] Second strand nick
[0282] As used herein, the concept refers to the introduction of a second nick at a position downstream of the first nick (i.e., the initial nick site that provides a free 3’ end for the reverse transcriptase to initiate extension). A portion of the guide RNA). In some embodiments, the first and second nicks are on opposite strands. In other embodiments, the first and second nicks are on opposite strands. In yet another embodiment, the first nick is on the non-target strand (i.e., the strand that forms the single-stranded portion of the R-loop) and the second nick is on the target strand. The second nick is at least 5 nucleotides downstream of the first nick, or at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more nucleotides downstream of the first nick. Without being bound by theory, the second nick induces the cell’s endogenous DNA repair and replication processes to replace the unedited strand. In some embodiments, the edited strand is the non-target strand and the unedited strand is the target strand. In other embodiments, the edited strand is the target strand and the unedited strand is the non-target strand.
[0283] Spacer sequence
[0284] As used herein, the term “spacer sequence” in relation to a guide RNA or PEgRNA refers to a portion of about 10 to about 40 (e.g., about 10, about 15, about 20, about 25, about 30) nucleotides of the guide RNA or PEgRNA that comprises a nucleotide sequence that is complementary to a protospacer sequence in the target DNA sequence. The spacer sequence anneals to the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R-loop ssDNA structure of the endogenous DNA strand that is complementary to the protospacer sequence.
[0285] Subject
[0286] As used herein, the term “subject” refers to an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, such as a genetically engineered non-human subject. The subject can be of any gender and at any stage of development.
[0287] Target site
[0288] The term "target site" refers to a sequence within a nucleic acid molecule that is edited by a prime editor disclosed herein. The target site also refers to a sequence within a nucleic acid molecule to which a complex of a base editor and a gRNA binds.
[0289] Temporal second strand nick
[0290] As used herein, the term "time second strand nick" refers to a variant of the second strand nick whereby the installation of the second nick in the unedited strand only occurs after the desired edit is installed in the edited strand. This avoids concurrent nicks on both strands that can lead to a double stranded DNA break. The second strand nick guide RNA is designed for temporal control so that the second strand nick is introduced only after the desired edit is installed. This is achieved by designing a gRNA with a spacer sequence that only matches the edited strand and not the original allele. Using this strategy, the mismatch between the original spacer and the unedited allele should disfavor nicking by the sgRNA until after the editing event on the PAM strand has occurred.
[0291] tPERT
[0292] See definition of "trans guide editor RNA template (tPERT)".
[0293] Temporal second strand nick
[0294] As used herein, the term "time second strand nick" refers to a variant of the second strand nick whereby the installation of the second nick in the unedited strand only occurs after the desired edit is installed in the edited strand. This avoids concurrent nicks on both strands that can lead to a double stranded DNA break. The second strand nick guide RNA is designed for temporal control so that the second strand nick is introduced only after the desired edit is installed. This is achieved by designing a gRNA with a spacer sequence that only matches the edited strand and not the original allele. Using this strategy, the mismatch between the original spacer and the unedited allele should disfavor nicking by the sgRNA until after the editing event on the PAM strand has occurred.
[0295] Trans prime editing
[0296] As used herein, the term“trans-guided editing” refers to a modified version of prime editing that utilizes a split PEgRNA, i.e., where the PEgRNA is split into two separate molecules: a sgRNA and a trans-guided editing RNA template (tPERT). The sgRNA is used to target the prime editor (or, more generally, the napDNAbp component of the prime editor) to the desired genomic target site, while the tPERT is used by a polymerase (e.g., a reverse transcriptase) to write the new DNA sequence. The tPERT enters the target locus upon being recruited to the prime editor in trans, through the interaction of binding domains located on the prime editor and the tPERT. In one embodiment, the binding domains can include an RNA-protein recruiting moiety, such as an MS2 aptamer located on the tPERT and an MS2 cp protein fused to the prime editor. One advantage of trans-guided editing is that by separating the DNA synthesis template from the guide RNA, a longer template can be used.
[0297] One embodiment of trans-guided editing is shown in FIGS. 3G and 3H. FIG. 3G shows the composition of the trans-guided editor complex on the left (“RP-PE:gRNA complex”), which comprises a napDNAbp fused to each of a polymerase (e.g., reverse transcriptase) and a rPERT recruiting protein (e.g., MS2sc), and complexed with a guide RNA. FIG. 3G further shows a separate tPERT molecule, which comprises the extended arm features of a PEgRNA, including a DNA synthesis template and a primer binding sequence. The tPERT molecule also includes an RNA protein recruiting domain (in this case, it is a stem loop structure, which can be, for example, an MS2 aptamer). The process is shown as described in the figure. FIG. 3H, the RP-PE:gRNA complex binds to and cleaves the target DNA sequence. The recruiting protein (RP) then recruits the tPERT to co-localize to the prime editing complex bound to the DNA target site, such that the primer binding site binds to the primer sequence on the nicked strand, and subsequently, the polymerase (e.g., RT) is allowed to synthesize single-stranded DNA from the DNA synthesis template 5’ of the tPERT.
[0298] While the tPERT is shown in FIG. 3G and FIG. 3H with the PBS and DNA synthesis template included 5’ of the RNA-protein recruiting domain, tPERTs in other configurations can be designed with the PBS and DNA synthesis template located 3’ of the RNA-protein recruiting domain. However, the advantage of a tPERT with a 5’ extension is that the synthesis of single-stranded DNA will naturally terminate at the 5’ end of the tPERT, thus not risking the use of any portion of the RNA protein recruiting domain as a template during the DNA synthesis phase of prime editing.
[0299] Trans prime editor RNA template (tPERT)
[0300] As used herein, “trans guide editor RNA template (tPERT)” refers to a component used for trans guide editing, a modified version of prime editing that operates by separating the PEgRNA into two distinct molecules: a guide RNA and a tPERT molecule. The tPERT molecule is programmed to co-localize with the prime editor complex at the target DNA site, introducing a primer binding site and a DNA synthesis template in trans to the prime editor. See, e.g., FIG. 3G for one embodiment of a trans guide editor (tPE), which shows a two-component system RNA-protein recruitment domain comprising (1) a RP-PE:gRNA complex and (2) a tPERT comprising a primer binding site and a DNA synthesis template, where the RP (recruiting protein) component of the RP-PE:gRNA complex recruits the tPERT to the target site to be edited, thereby associating the PBS and DNA synthesis template with the prime editor in trans. In other words, the tPERT is designed to comprise (all or part of) the extension arm of the PEgRNA, which includes a primer binding site and a DNA synthesis template.
[0301] Transition
[0302] As used herein, “transversion” refers to the interchange of a purine nucleobase or a pyrimidine nucleobase . Such exchanges involve nucleobases of similar shape. The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing transversions and transitions in the same target DNA molecule. These changes involve or In the case of double-stranded DNA with Watson-Crick pairing nucleobases, a transition refers to the following base pair exchange: or The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing transversions and transitions in the same target DNA molecule, as well as other nucleotide changes, including deletions and insertions.
[0303] Transversion
[0304] As used herein, “transversion” refers to the interchange of a purine nucleobase with a pyrimidine nucleobase, or vice versa, thus involving the interchange of nucleobases of different shape. These changes involve and In the case of double-stranded DNA with Watson-Crick pairing nucleobases, a transition refers to the following base pair exchange: and The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing transitions and transversions in the same target DNA molecule, as well as other nucleotide changes, including deletions and insertions.
[0305] Treatment
[0306] The terms“treatment,”“treat,” and“treating” refer to clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit a disease or disorder, or one or more symptoms thereof, as described herein. As used herein, the terms“treatment,”“treat,” and“treating” refer to clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment can be performed after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment can be administered in the absence of symptoms, e.g., to prevent or delay onset of symptoms or inhibit onset or progression of a disease. For example, treatment can be performed on an individual who is predisposed to developing the disease (e.g., based on a history of symptoms and / or based on genetic or other predisposing factors) prior to the development of symptoms. Treatment can also continue after symptoms have resolved, e.g., to prevent or delay their recurrence.
[0307] Trinucleotide repeat disorder
[0308] As used herein, a“trinucleotide repeat disorder” (or alternatively, an“expanding repeat disorder” or“repeat expansion disorder”) refers to a group of genetic disorders caused by“trinucleotide repeat expansion,” which is a mutation of a specific trinucleotide repeat in a particular gene or intron. Trinucleotide repeats were once thought to be common repeats in the genome, but these diseases were clarified in the 1990s. These apparently“benign” stretches of DNA sometimes expand and cause disease. Diseases caused by trinucleotide repeat expansion share several common features. First, the mutant repeats show instability in both somatic and germline cells, and more often, they expand rather than contract in successive transmission. Second, earlier age of onset and severity of the expected phenotype in offspring are often associated with larger repeat lengths. Finally, parental origin of the disease allele often influences the expectation, with paternal transmission of many of the diseases having a greater risk of expansion.
[0309] Trinucleotide repeat amplification is thought to be caused by slippage during DNA replication. Due to the repetitive nature of DNA sequences in these regions, "loopout" structures can form during DNA replication while maintaining complementary base pairing between the synthesizing parent and daughter strands. If the loopout structure forms on the daughter strand, this results in an increase in the number of repeats. However, if the loopout structure forms on the parent strand, the number of repeats decreases. These expansions of repeats appear to be more common than reductions. Generally, the larger the expansion, the more likely they are to cause disease or increase its severity. This property leads to the expected characteristics seen in trinucleotide repeat disorders. Expectations describe a trend of decreasing age of onset and increasing symptom severity across successive generations in affected families due to the expansion of these repeats.
[0310] Nucleotide duplication syndromes can include those in which triplet duplications occur in non-coding regions (i.e., non-coding trinucleotide duplication syndromes) or coding regions.
[0311] The guide editor (PE) system described in this article can be used to treat nucleotide duplication disorders, including Fragile X syndrome (FRAXA), Fragile XE MR (FRAXE), Freidreich ataxia (FRDA), myotonic dystrophy (DM), spinocerebellar ataxia type 8 (SCA8), and spinocerebellar ataxia type 12 (SCA12).
[0312] Prime editing or "prime editing (PE)"
[0313] As used herein, the term "guided editing" or "guided editing (PE)" refers to the use of techniques as described in this application and in [the context of the application]. FIG. 1A-1 The J implementation illustrates a novel method for gene editing using napDNAbps and specialized guide RNA. TPRT stands for "target-initiated reverse transcription" because, in one implementation, a target DNA molecule is used to initiate DNA strand synthesis via a reverse transcriptase (or another polymerase). In various implementations, guided editing is performed by contacting the target DNA molecule (to which a change in its nucleotide sequence needs to be introduced) with a nucleic acid-programmable DNA-binding protein (napDNAbp) complexed with a guide RNA. Reference FIG. 1Eguide RNA at the 3’ or 5’ end of the guide RNA or an intramolecular position of the guide RNA, and encodes the desired nucleotide change (e.g., single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) in one of the DNA strands of the target locus, thereby creating an available 3’ end in one of the strands of the target locus. In some embodiments, the nick is created in the DNA strand corresponding to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the “non-target strand.” However, the nick can be introduced in either strand. That is, the nick can be introduced into the “target strand” (i.e., the strand that is hybridized to the spacer of the extended gRNA) or the “non-target strand” (i.e., the strand that forms the single-stranded portion of the gene). The R-loop and is complementary to the target strand). In step (c), the 3’ end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime reverse transcription (i.e., “target primed RT”). In some embodiments, the 3’ end DNA strand hybridizes to a specific primer binding site on the extended portion of the guide RNA, i.e., the “reverse transcriptase priming sequence.” In step (d), a reverse transcriptase is introduced, synthesizing a single-stranded DNA from the 3’ end of the priming site to the 3’ end of the priming editor guide RNA. This forms a single-stranded DNA flap comprising the desired nucleotide change (e.g., single base change, insertion, or deletion, or a combination thereof), and is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve resolution of the single-stranded DNA flap, such that the desired nucleotide change is incorporated into the target locus. The process can push product formation by removing the corresponding 5’ endogenous DNA flap, once the 3’ single-stranded DNA flap invades and hybridizes to the endogenous DNA sequence, the desired product is formed. Without being bound by theory, cellular endogenous DNA repair and replication processes resolve the mismatched DNA to incorporate the nucleotide change to form the desired altered product. The process can also push product formation by a “second strand nick,” as shown in FIG. 1D
[0314] The term "prime editor (PE) system" or "prime editor" or "PE system" or "PE editing system" refers to compositions involved in the genome editing methods described herein using target primed reverse transcription (TPRT), including but not limited to napDNAbps, reverse transcriptases, fusion proteins (e.g., comprising napDNAbps and reverse transcriptases), prime editor guide RNAs, and complexes comprising fusion proteins and prime editor guide RNAs, as well as ancillary elements, such as a second strand nicking component and a 5' endogenous DNA flap removal endonuclease that help drive the primary editing process toward the formation of an edited product.
[0315] Upstream
[0316] As used herein, the terms "upstream" and "downstream" are relative terms that define the linear position of at least two elements in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in a 5' to 3' direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is located somewhere 5' of the second element. For example, if a SNP is located 5' of a nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is located somewhere 3' of the second element. For example, if a SNP is located 3' of a nick site, the SNP is downstream of the Cas9-induced nick site. A nucleic acid molecule can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. Analysis of single-stranded nucleic acid molecules and double-stranded molecules is the same because the terms upstream and downstream refer only to the single strand of the nucleic acid molecule, except where it is necessary to choose which strand of a double-stranded molecule is being considered. Generally, the strand of double-stranded DNA that can be used to determine the position correlation of at least two elements is the "sense" or "coding" strand. In genetics, the "sense" strand is the segment of double-stranded DNA from 5' to 3' that is complementary to the antisense or template strand (from 3' to 5') of DNA. Thus, for example, if a SNP nucleobase is 3' of a promoter on the sense or coding strand, the SNP nucleobase is "downstream" of the promoter sequence in genomic DNA (double-stranded).
[0317] Variant
[0318] As used herein, the term "variant" shall be taken to mean exhibiting a characteristic that deviates from the pattern found in nature, e.g., a variant Cas9 is a Cas9 that comprises one or more amino acid residue changes compared to the wild-type Cas9 amino acid sequence. The term "variant" encompasses a homologous protein having at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% identity to a reference sequence, and having the same or substantially the same functional activity as the reference sequence. The term also encompasses mutants, truncations, or domains of a reference sequence, and which exhibit the same or substantially the same functional activity(s) as the reference sequence.
[0319] Vector
[0320] As used herein, the term "vector" refers to a nucleic acid that can be modified to encode a gene of interest and is capable of entering a host cell, where it can undergo mutation and replication, and then transferring the replicated form of the vector to another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, as well as conjugative plasmids. Other suitable vectors will be apparent to one of skill in the art based on the present disclosure.
[0321] Wild type
[0322] As used herein, the term "wild-type" is art-recognized by the skilled artisan and refers to the typical form of an organism, strain, gene, or characteristic, as it occurs in nature, as distinguished from a mutant or variant form.
[0323] 5' endogenous DNA flap removal
[0324] As used herein, the term "5' endogenous DNA flap removal" or "5' flap removal" refers to
[0325] When the RT-synthesized single-stranded DNA flap competitively invades and hybridizes with the endogenous DNA, the resulting 5' endogenous DNA flap is removed, in the process displacing the endogenous strand. Removal of this endogenous displaced strand can drive the reaction toward formation of the desired product containing the desired nucleotide change. Cellular DNA repair enzymes can catalyze the removal or excision of the 5' endogenous flap (e.g., flap endonucleases such as EXOl or FENl). In addition, the host cell can be transformed to express one or more enzymes that catalyze the removal of the 5' endogenous flap, driving the process toward product formation (e.g., flap endonucleases). Flap endonucleases are known in the art and can be found in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10): 4507-4519 and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2): 198-211 (each incorporated herein by reference).
[0326] 5' endogenous DNA flap
[0327] As used herein, the term "5' endogenous DNA flap" refers to the strand of DNA located immediately downstream of the PE-induced nick site in the target DNA. The nick of the target DNA strand by PE exposes a 3' hydroxyl on the upstream side of the nick site and a 5' hydroxyl on the downstream side of the nick site. The endogenous strand ending in a 3' hydroxyl is used to prime the DNA polymerase of the prime editor (e.g., where the DNA polymerase is a reverse transcriptase). The endogenous strand on the downstream side of the nick site begins with an exposed 5' hydroxyl, referred to as a "5'p endogenous DNA flap," which is ultimately removed and replaced by a newly synthesized replacement strand encoded by the extension of the PEgRNA (i.e., a "3'f replacement DNA flap").
[0328] 3' substitution DNA flap
[0329] As used herein, the term "3' substitution DNA flap" or simply "substitution DNA flap" refers to a strand of DNA synthesized by a prime editor and encoded by the extension arm of a prime editor PEgRNA. More specifically, the 3' substitution DNA flap is encoded by the polymerase template of the PEgRNA. The 3' substitution DNA flap contains the same sequence as the 5' endogenous DNA flap, except that it also contains an edited sequence (e.g., a single nucleotide change). The 3' substitution DNA flap recombines with the target DNA, displacing or substituting the 5' endogenous DNA flap (e.g., can be excised by a 5' flap endonuclease, such as FEN1 or EXOl), and then ligates to add the 3' end of the 3' substitution DNA flap to the exposed 5' hydroxyl end of the endogenous DNA exposed upon excision of the 5' endogenous DNA flap, thereby reforming a phosphodiester bond and installing the 3' substitution DNA flap to form a heteroduplex DNA containing one edited strand and one unedited strand. The DNA repair process resolves the heteroduplex nucleic acid molecule by copying the information in the edited strand to the complementary strand, thereby permanently installing the edit into the DNA. This resolution process can be further completed by cleaving the unedited strand, i.e., by a "second strand nick", as described herein.
[0330] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0331] The present disclosure discloses new compositions (e.g., new PEgRNAs and PE complexes comprising them) and methods of using prime editing (PE) to repair therapeutic targets, such as those identified in the ClinVar database, using PEgRNAs designed using specialized algorithms described herein. Accordingly, the present application discloses algorithms for large-scale prediction of PEgRNA sequences that can be used to repair therapeutic targets (e.g., those included in the ClinVar database). In addition, the present application discloses predicted sequences of therapeutic PEgRNAs designed using the disclosed algorithms and can be used with prime editing to repair therapeutic targets.
[0332] The algorithms and predicted PEgRNA sequences disclosed herein generally relate to prime editing. Accordingly, the present disclosure also provides a description of various components and aspects of prime editing, including suitable napDNAbps (e.g., Cas9 nickases) and reverse transcriptases, as well as other suitable components (e.g., linkers, NLSs) and PE fusion proteins, which can be used with the therapeutic PEgRNAs disclosed herein.
[0333] Genome editing with the clustered regularly interspaced short palindromic repeats (CRISPR) system has revolutionized the life sciences 1-3Despite the now routine use of CRISPR for gene disruption, the precise installation of single nucleotide edits remains a significant challenge, despite being necessary for research or correction of a large number of disease-causing mutations. Homology directed repair (HDR) enables such edits, but is inefficient (typically <5%), requires a donor DNA repair template, and has deleterious effects of double-strand DNA break (DSB) formation. Recently, the laboratory of Professor David Liu developed base editing, enabling highly efficient single nucleotide edits without the need for DSBs. Base editors (BEs) combine the CIRSPR system with base-modifying deaminases to convert target C·G or A·T base pairs to A·T or G·C 4-6 Despite being widely used by researchers around the world, current BEs support only four of the twelve possible base pair conversions, and are unable to correct small insertions or deletions. Furthermore, the targeting range of base editing is limited by editing of non-target C or A bases adjacent to the target base ("bystander editing") and the requirement that the PAM sequence be present 15±2 bp from the target base. Thus, overcoming these limitations would greatly expand the basic research and therapeutic applications of genome editing.
[0334] The present disclosure proposes a new method of precise editing that provides many of the benefits of base editing - i.e. avoidance of double-strand breaks and donor DNA repair templates - while overcoming its major limitations. The proposed method described herein uses target-primed reverse transcription (TPRT) to enable direct installation of edited DNA strands at targeted genomic sites. In the design discussed herein, a CRISPR guide RNA (gRNA) would be designed to carry a reverse transcriptase (RT) template sequence that encodes a single-stranded DNA containing the desired nucleotide change. The target site DNA cleaved by the CRISPR nuclease (Cas9) would serve as a primer for reverse transcription of the template sequence on the modified gRNA, allowing direct incorporation of any desired nucleotide edit.
[0335] Accordingly, the present invention is directed, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) can be harnessed or adapted for CRISPR / Cas-based precise genome editing with high efficiency and genetic flexibility (e.g., as FIG. 1A-1Various embodiments of G). The inventors have proposed herein the use of a napDNAbp-polymerase fusion (e.g., a Cas9 nickase fused to a reverse transcriptase) to target a specific DNA sequence with a modified guide RNA (“extended guide RNA” or PEgRNA), to make a single-stranded nick at the target site, and to use the cleaved DNA as a primer to synthesize DNA by a polymerase (e.g., reverse transcriptase) based on the DNA that is part of the PEgRNA. The newly synthesized strand will be homologous to the genomic target sequence, except for containing the desired nucleotide change (e.g., a single nucleotide change, a deletion, or an insertion, or a combination thereof). The newly synthesized DNA strand can be referred to as a single-stranded DNA flap, which will compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. Resolution of this hybridization intermediate can include removal of the displacement flap of endogenous DNA that results therefrom (e.g., with a 5’ end DNA flap endonuclease, FEN1), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis provides single-nucleotide precision, the scope of this approach is very broad, foreseeably useful for countless applications in basic science and therapeutics.
[0336] I. Therapeutic PEgRNAs
[0337] The prime editor (PE) systems described herein encompass the use of any suitable prime editor guide RNA or PEgRNA. The inventors have discovered that by using a guide RNA comprising a special configuration of a DNA synthesis template, which encodes a desired nucleotide change through a polymerase (e.g., reverse transcriptase), the mechanism of target primed reverse transcription (TPRT) can be harnessed or adapted to perform precise and versatile CRISPR / Cas-based genome editing. This application refers to such a specially configured guide RNA as a “prime editor guide RNA” (or PEgRNA), as the DNA synthesis template can be provided as an extension of a standard or conventional guide RNA molecule. This application encompasses any suitable configuration or arrangement of a prime editor guide RNA.
[0338] In various embodiments, the disclosure provides therapeutic PEgRNAs of SEQ ID NOs: 1-135514 and 813085-880462 designed using the algorithms disclosed herein against ClinVar database entries.
[0339] In various other embodiments, exemplary PEgRNAs designed against the ClinVar database using the algorithms disclosed herein are included in the sequence listing, which forms a part of the present specification. The sequence listing includes the complete PEgRNA sequences of SEQ ID NOS: 1-135514 and 813085-880462. Each of these complete PEgRNAs consists of a spacer (SEQ ID NOS: 135515-271028 and 880463-947840) and an extension arm (SEQ ID NOS: 271029-406542 and 947841-1015218). In addition, each PEgRNA comprises a gRNA core, e.g., as defined by SEQ ID NOS: 1361579-1361580. The extension arms of SEQ ID NOS: 271029-406542 and 947841-1015218 each further comprise a primer binding site (SEQ ID NOS: 406543-542056 and 1015219-1082596), an editing template (SEQ ID NOS: 542057-677570 and 1082597-1149974), and a homology arm (SEQ ID NOS: 677571-813084 and 1149975-1217352). The PEgRNAs optionally can comprise a 5’-terminal modifier region and / or a 3’-terminal modifier region. The PEgRNAs can also comprise a reverse transcription termination signal at the 3’ of the PEgRNA (e.g., SEQ ID NOS: 1361560-1361566). This application encompasses the design and use of all of these sequences.
[0340] FIG. 3A An embodiment of a prime editor guide RNA (referred to as a “PEgRNA” or “extended gRNA”) that can be used in the prime editor (PE) systems disclosed herein is shown, where the conventional guide RNA (green portion) includes a spacer and a gRNA core region, which binds to the napDNAbp. In this embodiment, the guide RNA includes an extended RNA segment at the 5’ end, i.e., a 5’ extension. In this embodiment, the 5’ extension includes a DNA synthesis template, a primer binding site, and optionally a 5-20 nucleotide linker sequence. As shown in FIG. 1A, the primer binding site hybridizes to the free 3’ end that is formed after a nick is made in the non-target strand of the R-loop, thereby directing a polymerase (e.g., a reverse transcriptase) to perform DNA polymerization in the 5’ to 3’ direction.
[0341] FIG. 3BAnother embodiment of a prime editor guide RNA useful in the prime editor (PE) systems disclosed herein is shown, where the traditional guide RNA (green portion) includes a ~20 nt spacer and gRNA core, which binds to the napDNAbp. In this embodiment, the guide RNA includes an extended RNA segment at an intramolecular location within the gRNA core, i.e., an intramolecular extension. In this embodiment, the intramolecular extension includes a DNA synthesis template and a primer binding site. The primer binding site hybridizes to the free 3’ end formed after a nick is made in the non-target strand of the R-loop, thereby directing the polymerase to perform DNA polymerization in the 5’ to 3’ direction.
[0342] FIG. 3C Another embodiment of a prime editor guide RNA useful in the prime editor (PE) systems disclosed herein is shown, where the traditional guide RNA (green portion) includes a ~20 nt spacer and gRNA core, which binds to the napDNAbp. In this embodiment, the guide RNA includes an extended RNA segment at an intramolecular location within the gRNA core, i.e., an intramolecular extension. In this embodiment, the intramolecular extension includes a DNA synthesis template and a primer binding site. The primer binding site hybridizes to the free 3’ end formed after a nick is made in the non-target strand of the R-loop, thereby directing the polymerase to perform DNA polymerization in the 5’ to 3’ direction.
[0343] In one embodiment, the location of the intramolecular RNA extension is in the spacer of the guide RNA. In another embodiment, the location of the intramolecular RNA extension is in the gRNA core. In yet another embodiment, the location of the intramolecular RNA extension is anywhere within the guide RNA molecule other than the spacer, or at a position that disrupts the spacer.
[0344] In one embodiment, the intramolecular RNA extension is inserted downstream of the 3’ end of the spacer. In another embodiment, the intramolecular RNA extension is inserted at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides downstream of the 3’ end of the spacer.
[0345] In other embodiments, the intramolecular RNA extension is inserted into a gRNA, which refers to a guide RNA corresponding to or comprising a portion of a tracrRNA that binds and / or interacts with a Cas9 protein or its equivalent (i.e., a different napDNAbp). Preferably, the insertion of the intramolecular RNA extension does not disrupt or minimally disrupts the interaction between the tracrRNA portion and the napDNAbp.
[0346] The length of the RNA extension can be any useful length. In various embodiments, the length of the RNA extension is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0347] The DNA synthesis template (e.g., RT template sequence) can also be any useful length. For example, the length of the DNA synthesis template (e.g., RT template sequence) can be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0348] In other embodiments, the length of the reverse transcription primer binding site sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, or at least 100 nucleotides
[0349] In other embodiments, the length of the optional linker or spacer region is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, or at least 400 nucleotides.
[0350] In some embodiments, the DNA synthesis template (e.g., the RT template sequence) encodes a single-stranded DNA molecule that is homologous to the non-target strand (and thus complementary to the corresponding site of the target strand) but includes one or more nucleotide changes. The nucleotide changes can include one or more single base nucleotide changes, one or more deletions, one or more insertions, and combinations thereof.
[0351] As shown in FIG. 1E the single-stranded DNA product of the DNA synthesis template (e.g., the RT template sequence) is homologous to the non-target strand and contains one or more nucleotide changes. The single-stranded DNA product of the DNA synthesis template (e.g., the RT template sequence) equilibrates hybridization with the complementary target strand sequence, thereby displacing the homologous endogenous target strand sequence. In some embodiments, the displaced endogenous strand can be referred to as a 5’ endogenous DNA flap species (e.g., see FIG. 1C). This 5' endogenous DNA flap species can be removed by a 5' flap endonuclease (e.g., FEN1), and the single-stranded DNA product now hybridized to the endogenous target strand can be ligated, establishing a mismatch between the endogenous sequence and the newly synthesized strand. This mismatch can be resolved by the cell's innate DNA repair and / or replication processes.
[0352] In various embodiments, the nucleotide sequence of the DNA synthesis template (e.g., the RT template sequence) corresponds to the nucleotide sequence of the non-target strand, which is replaced as a 5' flap species and overlaps with the site to be edited.
[0353] In various embodiments of the prime editor guide RNA, the DNA synthesis template can encode a single-stranded DNA flap complementary to the endogenous DNA sequence adjacent to the nick site, wherein the single-stranded DNA flap comprises the desired nucleotide alteration. The single-stranded DNA flap can replace the endogenous single-stranded DNA at the nick site. The endogenous single-stranded DNA replaced at the nick site can have a 5' end and form an endogenous flap, which can be excised by the cell. In various embodiments, excision of the 5' endogenous flap can help drive product formation, as removal of the 5' endogenous flap facilitates hybridization of the single-stranded 3' DNA flap to the corresponding complementary DNA strand, and incorporation or assimilation of the desired nucleotide alteration carried into the target DNA by the single-stranded 3' DNA flap.
[0354] In various embodiments of the prime editor guide RNA, cellular repair of the single-stranded DNA flap results in installation of the desired nucleotide alteration, forming the desired product.
[0355] In other embodiments, the desired nucleotide alteration is installed in an editing window of between about -5 to +5 of the nick site, or between about -10 to +10 of the nick site, or between about -20 to +20 of the nick site, or between about -30 to +30 of the nick site, or between about -40 to +40 of the nick site, or between about -50 to +50 of the nick site, or between about -60 to +60 of the nick site, or between about -70 to +70 of the nick site, or between about -80 to +80 of the nick site, or between about -90 to +90 of the nick site, or between about -100 to +100 of the nick site, or between about -200 to +200 of the nick site.
[0356] In various aspects, the prime editor guide RNA is a modified version of a guide RNA. The guide RNA can be naturally occurring, expressed from a coding nucleic acid, or chemically synthesized. Methods for obtaining or otherwise synthesizing guide RNAs and for determining appropriate sequences of guide RNAs are well known in the art, including a spacer that interacts and hybridizes to the target strand of the genomic target site of interest.
[0357] In various embodiments, the specific design aspects of the guide RNA sequence will depend on the nucleotide sequence of the genomic target site of interest (i.e., the desired site to be edited) and the type of napDNAbp (e.g., Cas9 protein) present in the prime editor (PE) system described herein, as well as other factors such as PAM sequence position, percent G / C content in the target sequence, extent of microhomology region, secondary structure, etc.
[0358] Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity to a target polynucleotide sequence to hybridize to the target sequence and direct sequence-specific binding of a napDNAbp (e.g., Cas9, Cas9 homolog, or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned, is about or greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch Algorithm, algorithms based on the Burrows- Wheeler transform (e.g., the Burrows Wheeler Aligner), ClustalW, ClustalX, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a guide sequence is about or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length.
[0359] In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of a guide sequence to direct sequence-specific binding of a base editor to a target sequence can be assessed by any suitable assay. For example, components of a base editor, including a guide sequence to be tested, can be provided to a host cell having a corresponding target sequence, such as by transfection with a vector encoding the components of a base editor disclosed herein, followed by assessment of preferential cleavage within the target sequence, such as by a Surveyor assay as described herein.
[0360] Similarly, cleavage of a target polynucleotide sequence can be assessed in vitro by providing the target sequence, components of a base editor, including a guide sequence to be assayed and a control guide sequence that is different from the test guide sequence, and comparing the rate of cleavage of the target sequence between the binding or assay and control guide sequence reactions. Other assays are possible and will occur to those of skill in the art.
[0361] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within a cell genome. Exemplary target sequences include those that are unique in the target genome. For example, for S. pyogenes Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG (SEQ ID NO: 1361548), where NNNNNNNNNNNNXGG (SEQ ID NO: 1361551) (N is A, G, T or C; and X can be anything) occurs only once in the genome. A unique target sequence in the genome can include a S. pyogenes Cas9 target site of the form MMMMMMMMNNNNNNNNNNNXGG (SEQ ID NO: 1361550), where NNNNNNNNNNNNXGG (SEQ ID NO: 1361559) (N is A, G, T or C; and X can be anything) occurs only once in the genome. For S. thermophilus CRISPR1 Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW (SEQ ID NO: 1361552), where NNNNNNNNNNNNNXXAGAAW (SEQ ID NO: 1361553) (N is A, G, T or C; X can be anything; W is A or T) occurs only once in the genome. A unique target sequence in the genome can include a S. thermophilus CRISPR1 Cas9 of the form MMMMMMMMNNNNNNNNNNNXXAGAAW (SEQ ID NO: 1361554), where NNNNNNNNNNNNXXAGAAW (SEQ ID NO: 1361555) (N is A, G, T or C; X can be anything; W is A or T) occurs only once in the genome. For S. pyogenes Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG (SEQ ID NO: 1361556), where NNNNNNNNNNNNXGGXG (SEQ ID NO: 1361557) (N is A, G, T or C; and X can be anything) occurs only once in the genome. A unique target sequence in the genome can include a S. pyogenes Cas9 target site of the form MMMMMMMMNNNNNNNNNNNXGGXG (SEQ ID NO: 1361558), where NNNNNNNNNNNNXGGXG (SEQ ID NO: 1361559) (N is A, G, T or C; and X can be anything) occurs only once in the genome.In each of these sequences, "M" can be A, G, T, or C, and need not be considered when the sequence is identified as unique.
[0362] In some embodiments, the guide sequence is selected to reduce the extent of secondary structure within the guide sequence. Secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs base calculations on the minimization of Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151-62). Further algorithms can be found in U.S. Application Serial No. 61 / 836,080; Broad Reference BI-2013 / 004A, incorporated herein by reference.
[0363] Generally, a tracr mate sequence comprises any sequence that has sufficient complementarity to a tracr sequence to promote one or more of: (1) excision of a guide sequence flanked by the tracr mate sequence in a cell containing the corresponding tracr sequence; (2) formation of a complex at a target sequence, where the complex comprises the tracr mate sequence hybridized to the tracr sequence. Generally, the degree of complementarity is with reference to optimal alignment of the tracr mate sequence and the tracr sequence along the length of the shorter of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm, and can further account for secondary structure, such as self-complementarity within the tracr sequence or the tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr mate sequence, along the length of the shorter of the two, when optimally aligned, is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more. In some embodiments, the tracr sequence is about or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the tracr mate sequence are contained within a single transcript, such that hybridization between the two results in a transcript with secondary structure, such as a hairpin. A preferred loop-forming sequence for a hairpin structure is four nucleotides in length, most preferably having the sequence GAAA. However, longer or shorter loop sequences can be used, as can alternative sequences. The sequence preferably includes a nucleotide triad (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In one embodiment of the application, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In preferred embodiments, the transcript has two, three, four, or five hairpins. In another embodiment of the application, the transcript has at most five hairpins. In some embodiments, the single transcript also includes a transcriptional termination sequence; preferably, this is a polyT sequence, such as six T nucleotides.Additional non-limiting examples of single polynucleotides comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed 5' to 3'), where "N" represents a base of the guide sequence, the first block of lower case letters represents a tracr mate sequence, the second block of lower case letters represents a tracr sequence, and the final poly-T sequence represents a transcriptional terminator: (1) NNNNNNNNgtttttgtactctcaagatttaGAAAtaaatcttgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 1361560); (2) NNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 1361561); (3) NNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgtTTTTT (SEQ ID NO: 1361562); (4) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcTTTTTT (SEQ ID NO: 1361563); (5) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAATAGcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcTTTTTT (SEQ ID NO: 1361564); and (6) NNNNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtccgttatcaTTTTTTTT (SEQ ID NO: 1361565). In some embodiments, sequences (1) to (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from the transcript comprising the tracr mate sequence.
[0364] It will be apparent to those skilled in the art that, in order to target a target site, e.g., a site comprising a point mutation to be edited, with any fusion protein comprising a Cas9 domain and a single-stranded DNA binding protein, as disclosed herein, the fusion protein is typically co-expressed with a guide RNA (e.g., sgRNA). As explained in greater detail elsewhere herein, the guide RNA typically comprises a tracrRNA framework that allows for Cas9 binding and a guide sequence that confers sequence specificity to the Cas9:nucleic acid editing enzyme / domain fusion protein.
[0365] In some embodiments, the guide RNA comprises the structure 5'-[guide sequence]- guuuuagagcuagaaauagcaaguuaaaauaaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3' (SEQ ID NO: 1361566), wherein the guide sequence comprises a sequence complementary to a target sequence. The guide sequence is typically 20 nucleotides in length. Based on the present disclosure, sequences of suitable guide RNAs for targeting a Cas9:nucleic acid editing enzyme / domain fusion protein to a particular genomic target site will be apparent to those skilled in the art. Such suitable guide RNA sequences typically comprise a guide sequence that is complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of a target nucleotide to be edited. Some exemplary guide RNA sequences suitable for targeting any of the provided fusion proteins to a particular target sequence are provided herein. Additional guide sequences are well known in the art and can be used with the base editors described herein.
[0366] In other embodiments, the PEgRNA can include those described by the structure depicted in FIG. 27 comprising a guide RNA and a 3' extension arm.
[0367] FIG. 27A structure of one embodiment of PEgRNA encompassed herein is provided, which can be designed according to the methods defined in Example 2. The PEgRNA comprises three main constituent elements arranged in 5’ to 3’ direction, namely: a spacer, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in 5’ to 3’ direction, namely: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, the PEgRNA can comprise an optional 3’ terminal modifier region (el) and an optional 5’ terminal modifier region (e2). Further, the PEgRNA can comprise a transcription termination signal (not depicted) at the 3’ end of the PEgRNA. The description of the PEgRNA structure is not meant to be limiting, but to encompass variations in the arrangement of elements. For example, the optional sequence modifications (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends.
[0368] In other embodiments, the PEgRNA can include those depicted by the structure shown in FIG. 28 comprise a guide RNA and a 5’ extension arm.
[0369] FIG. 28 A structure of another embodiment of PEgRNA encompassed herein is provided, which can be designed according to the methods defined in Example 2. The PEgRNA comprises three main constituent elements arranged in 5’ to 3’ direction, namely: a spacer, a gRNA core, and an extension arm at the 3’ end. The extension arm can be further divided into the following structural elements in 5’ to 3’ direction, namely: a primer binding site (A) (SEQ ID NOs: 406543-542056 and 1015219-1082596), an editing template (B) (SEQ ID NOs: 542057-677570 and 1082597-1149974), and a homology arm (C) (SEQ ID NOs: 677571-813084 and 1149975-1217352). In addition, the PEgRNA can comprise an optional 3’ terminal modifier region (el) and an optional 5’ terminal modifier region (e2). Further, the PEgRNA can comprise a transcription termination signal (not depicted) at the 3’ end of the PEgRNA. The description of the PEgRNA structure is not meant to be limiting, but to encompass variations in the arrangement of elements. For example, the optional sequence modifications (el) and (e2) can be located within or between any of the other regions shown, and are not limited to being located at the 3’ and 5’ ends.
[0370] PEgRNAs can also include additional design improvements that can modify the properties and / or characteristics of PEgRNAs to improve the efficiency of prime editing. In various embodiments, these improvements can fall into one or more of a number of different categories, including but not limited to: (1) design to enable efficient expression of functional PEgRNAs from non-polymerase III (pol III) promoters, which would enable expression of longer PEgRNAs without the onerous sequence requirements; (2) improvements to the core, Cas9-binding PEgRNA scaffold, which can improve efficacy; (3) modification of PEgRNAs to improve RT processivity, enabling insertion of longer sequences at the target genomic locus; and (4) addition of RNA motifs at the 5’ or 3’ end of PEgRNAs to improve PEgRNA stability, enhance RT processivity, prevent PEgRNA misfolding, or recruit other factors important for genome editing.
[0371] In one embodiment, PEgRNAs can be designed with a pol III promoter to improve expression of longer PEgRNAs with larger extension arms. sgRNAs are typically expressed from a U6 snRNA promoter. This promoter recruits pol III to express relevant RNAs, and is used to express short RNAs that remain within the nucleus. However, the processing power of pol III is not strong enough to express RNAs longer than a few hundred nucleotides at the levels required for efficient genome editing. Furthermore, pol III can stop or terminate at the extension of the U, which can limit the sequence diversity that can be inserted using PEgRNAs. Other promoters that recruit polymerase II (such as pCMV) or polymerase I (such as the U1 snRNA promoter) have been examined for their ability to express longer sgRNAs. However, these promoters are often partially transcribed, which would result in additional sequence 5’ of the spacer in the expressed PEgRNA, which has been shown to cause a dramatic reduction in Cas9:sgRNA activity in a site-dependent manner. Furthermore, while pol III-transcribed PEgRNAs can simply terminate in a run of 6-7 Us, PEgRNAs transcribed from pol II or pol I require a different termination signal. Often, such signals also result in polyadenylation, leading to the inadvertent transport of PEgRNAs from the nucleus. Similarly, RNAs expressed from pol II promoters such as pCMV are typically 5’ capped, which also leads to their nuclear export.
[0372] Previously, Rinn and colleagues screened a variety of expression platforms for the production of long-chain non-coding RNA- (lncRNA)-tagged sgRNAs 183 . These platforms included expression from pCMV and termination at the ENE element from the human MALAT1 ncRNA 184 , the PAN ENE element from KSHV185 Or from the 3' frame of U1snRNA 186 The RNA. Notably, MALAT1 ncRNA and PANENE form a triple helix to protect the polyA tail. 184、187 These constructs can also enhance RNA stability. These expression systems are also expected to be able to express longer PEgRNAs.
[0373] In addition, a series of methods were designed to cleave the pol II promoter portion that will be transcribed as part of PEgRNA, by adding self-cleaving ribozymes (such as hammerheads). 188 (hammerhead 188 ),pistol 189 (pistol 189 ),ax 189 (hatchet 189 hair clip 190 VS 191 twister 192 or twin sister 192 Ribozymes or other self-cleaving elements are used to process transcription guides, or by Csy4 193 The hairpins are identified and lead to guide processing. Furthermore, it is hypothesized that incorporating multiple ENE motifs could result in improved PEgRNA expression and stability, as previously demonstrated with KSHV PAN RNA and elements. 185 It is also anticipated that circularization of PEgRNA in the form of circular intron RNA (ciRNA) may lead to enhanced RNA expression and stability, as well as nuclear localization. 194 .
[0374] In various implementations, PEgRNA may include various of the above-described elements, as exemplified by the following sequences.
[0375] Non-restrictive example 1 - PEgRNA expression platform consisting of pCMV, Csy4 hairpin, PEgRNA, and MALAT1 ENE
[0376]
[0377] Non-restrictive example 2 - PEgRNA expression platform consisting of pCMV, Csy4 hairing, PEgRNA, and PAN ENE
[0378]
[0379] Non-restrictive example 3 - PEgRNA expression platform consisting of pCMV, Csy4 hairing, PEgRNA, and 3xPAN ENE
[0380]
[0381] Non-limiting Example 4 - PEgRNA expression platform consisting of pCMV, Csy4 hairpin, PEgRNA, and 3’ box
[0382]
[0383] Non-limiting Example 5 - PEgRNA expression platform consisting of pU1, Csy4 hairpin, PEgRNA, and 3’ box
[0384]
[0385]
[0386] In various other embodiments, PEgRNAs can be improved by introducing improvements to the scaffold or core sequence. This can be done by introducing known.
[0387] The core, Cas9-binding PEgRNA scaffold can be improved to enhance PE activity. Several such approaches have been demonstrated. For example, the first pairing element of the scaffold (P1) contains a GTTTT-AAAAC pairing element. Such Ts runs have been shown to cause pol III stalling and premature termination of RNA transcription. Reasonable mutation of one of the T-A pairs to a G-C pair in this portion of P1 has been shown to enhance sgRNA activity, suggesting that this approach is also feasible for PEgRNAs. In addition, increasing the length of P1 has also been shown to enhance sgRNA folding and lead to improved activity195, suggesting that it is another avenue to improve PEgRNA activity. Examples of improvements to the core include:
[0388] PEgRNA containing a 6nt extension to P1
[0389]
[0390] PEgRNA containing a T-A to G-C mutation within P1
[0391]
[0392] In various other embodiments, PEgRNAs can be improved by introducing modifications to the editing template region. As the size of the insert templated by the PEgRNA increases, it is more likely to be degraded by endonucleases, undergo spontaneous hydrolysis, or fold into secondary structures that cannot be reverse transcribed by the RT or disrupt the PEgRNA scaffold folding and subsequent Cas9-RT binding. Thus, modifications to the PEgRNA template can be required to affect large insertions, such as insertions of entire genes. Some strategies to do so include incorporating modified nucleotides in the synthetic or semi-synthetic PEgRNA that make the RNA more resistant to degradation or hydrolysis, or less likely to adopt inhibitory secondary structures 196 Such modifications can include 8-aza-7-deazaguanosine, which reduces RNA secondary structure in G-rich sequences; locked nucleic acids (LNAs), which can reduce degradation and enhance certain kinds of RNA secondary structure; 2'-O-methyl, 2'-fluoro, or 2'-O-methoxyethoxy modifications, which can enhance RNA stability. Such modifications can also be included elsewhere in the PEgRNA to enhance stability and activity. Alternatively or additionally, the template of the PEgRNA can be designed to both encode the desired protein product and be more likely to adopt simple secondary structures that can be unwound by the RT. Such simple structures will act as a thermodynamic sink, making it less likely that more complex structures that prevent reverse transcription will arise. In such a design, the PE would be used to initiate transcription, and a separate template RNA would be recruited to the target site by a RNA-binding protein fused to Cas9 or a RNA recognition element (such as an MS2 aptamer) on the PEgRNA itself. The RT can bind directly to this separate template RNA, or initiate reverse transcription on the original PEgRNA before switching to a second template. Such an approach can enable long insertions by preventing misfolding of the PEgRNA when adding long templates, as well as not requiring the Cas9 to be dissociated from the genome for long insertions to occur, which can inhibit PE-based long insertions.
[0393] (iv) Installing extra RNA motifs at the 5' or 3' end
[0394] In other embodiments, PEgRNAs can be improved by introducing extra RNA motifs at the 5' and 3' ends of the PEgRNA. Several such motifs - such as the PAN ENE from KSHV and the ENE from MALAT1 discussed above - serve as a possible means of terminating expression of longer PEgRNAs from non-pol III promoters. These elements form RNA triplexes that engulf polyA tails, causing them to remain within the nucleus 184,187 However, by forming complex structures of closed-ended nucleotides at the 3' end of the PEgRNA, these structures can also help prevent exonuclease-mediated degradation of the PEgRNA.
[0395] Other structural elements inserted at the 3' end can also enhance RNA stability, although not terminating transcription from non-pol III promoters. Such motifs can include hairpins or RNA quadruplexes, which would seal the 3' end 197 , or self-cleaving ribozymes (such as HDV), which would result in the formation of a 2'-3'-cyclic phosphate at the 3' end and potentially result in a PEgRNA less likely to be degraded by exonucleases 198 . Inducing PEgRNA circularization by incomplete splicing - forming ciRNAs - can also increase PEgRNA stability and result in PEgRNAs remaining in the nucleus 194 .
[0396] Additional RNA motifs can also improve RT processivity or enhance PEgRNA activity by enhancing RT binding to DNA-RNA duplexes. Adding natural sequences bound by RT in its cognate retroviral genome can enhance RT activity 199 . This can include natural primer binding sites (PBS), polypurine tracts (PPT), or kissing loops involved in retroviral genome dimerization and transcription initiation 199 .
[0397] Adding dimerization motifs - such as kissing loops or GNRA tetraloops / tetraloop receptor pairs 200 - at the 5' and 3' ends of PEgRNAs can also result in efficient circularization of PEgRNAs, improving stability. Furthermore, it is expected that adding these motifs can enable physical separation of the PEgRNA spacer and primer, preventing spacer occlusion that can hinder PE activity. Short 5' extensions of PEgRNAs that form small toehold hairpins can also advantageously compete with the annealing region of PEgRNAs that bind the spacer. Finally, kissing loops can also be used to recruit other template RNAs to the genomic site and enable RT activity to switch from one RNA to another. Example modifications include, but are not limited to:
[0398] PEgRNA-HDV fusion
[0399]
[0400] PEgRNA-MMLV kissing loop
[0401]
[0402] PEgRNA-VS ribozyme kissing loop,
[0403]
[0404] PEgRNA-GNRA tetraloop / tetraloop receptor
[0405]
[0406] PEgRNA template switch secondary RNA-HDV fusion
[0407]
[0408] The PEgRNA scaffold can be further improved through directed evolution in a similar manner to the already improved SpCas9 and base editors. Directed evolution can enhance the recognition of PEgRNAs by Cas9 or evolved Cas9 variants. Furthermore, different PEgRNA scaffold sequences can be optimal at different genomic loci, either enhancing PE activity at the relevant site, reducing off-target activity, or both. Finally, evolution of PEgRNA scaffolds with additional RNA motifs almost certainly will improve the activity of the fusion PEgRNAs relative to the unevolved fusion RNAs. For example, evolution of an allosteric ribozyme consisting of a c-di-GMP-I aptamer and a hammerhead ribozyme resulted in a dramatic improvement in activity 202 , suggesting that evolution will also improve the activity of hammerhead-PEgRNA fusions. Furthermore, while Cas9 is currently generally intolerant of 5' extensions of sgRNAs, directed evolution can yield mutations that mitigate this intolerance, allowing the utilization of additional RNA motifs.
[0409] The present disclosure encompasses any such ways to further improve the efficacy of the prime editing systems disclosed herein.
[0410] II. Algorithms and methods for designing therapeutic PEgRNAs
[0411] As described herein, the inventors have discovered and appreciated that prime editing using PEgRNAs can be used to install a variety of nucleotide changes, including insertions (of any length, including entire genes or protein-coding regions), deletions (of any length), and correction of pathogenic mutations. However, there is no existing technology to determine and / or predict PEgRNA structures, including specifying the various components of a PEgRNA, such as the spacer, gRNA core, and extension arm (as well as the extension components described herein). The inventors have developed computerized techniques for determining PEgRNA structures, including determining extended gRNA structures. Each extended gRNA structure can be determined based on an input allele (e.g., representing a pathogenic mutation), an output allele (e.g., representing a corrected wild-type sequence), and a fusion protein (e.g., a CRISPR system for prime editing, including the relative positions of the PAM motif and the prime editor nick). The difference between the input allele and the output allele represents the desired edit (e.g., a single nucleotide change, an insertion, a deletion, etc.). The determined structure can be created and used to perform base editing to change the input allele to the output allele, as further described herein.
[0412] FIG. 31 is a flowchart showing an exemplary high-level computerized method 3100 for determining an extended gRNA structure, according to some embodiments. At step 3102, a computing device (e.g., in conjunction with the computing device 3400 described below) accesses data indicative of an input allele, an output allele, and a fusion protein, the fusion protein including a nucleic acid programmable DNA binding protein and a pseudo-transcriptase. While step 3102 describes accessing all three of the input allele, the output allele, and the fusion protein in one step, this is for illustrative purposes, and it should be appreciated that such data can be accessed using one or more steps without departing from the spirit of the technology described herein. Accessing the data can include receiving the data, storing the data, accessing a database, etc. FIG. 34
[0413] At step 3104, the computing device determines an extended gRNA structure based on the input allele, the output allele, and the fusion protein accessed in step 3102. The extended gRNA structure is designed to associate with the fusion protein to change the input allele to the output allele. When the fusion protein is complexed with the extended gRNA, it is able to bind to a target DNA sequence, including a target strand where the change occurs and a complementary non-target strand. As described herein, the input allele can represent a pathogenic DNA mutation, and the output allele can represent a corrected DNA sequence.
[0414] Changing the input allele to the output allele can include a single nucleotide change, insertion of one or more nucleotides, deletion of one or more nucleotides, and / or any other change designed to achieve the output allele. In particular, exemplary classes of edits that can be induced by a single PEgRNA include single nucleotide substitutions, insertions from 1 nt up to about 40 nt, deletions from 1 nt up to about 30 nt, and combinations thereof. For example, prime editing can support these types of changes from spacer position -3 (e.g., immediately 3’ of the nick) to spacer position +27 (e.g., 30 nt 3’ of the nick in the input allele). For example, editing can be performed at spacer position -4 using a SpCas9 system with prime editing (e.g., this can result from an incidental RuvC cleavage between spacer positions -5 and -4). The type of change, the number of nt changes, and / or the location of the change can be configurable parameters that can be used by computerized techniques to determine the extended gRNA structure.
[0415] As discussed in connection with FIG. 3A-3B and FIG. 27-28 An extended gRNA can include various components, such as a spacer of the extended gRNA that is complementary to a target nucleotide sequence in the input allele, a gRNA backbone for interacting with a fusion protein, and an extension. With further reference to step 3104, the computing device determines one or more of the spacer, the gRNA backbone, and the extension. In some embodiments, while this technique can include determining any combination of the spacer, the gRNA backbone, and / or the extension, in some embodiments one or more of such components and / or aspects of such components are known (e.g., predetermined, pre-specified, fixed, etc.) and thus can not be determined as part of step 3104.
[0416] As described herein, the gRNA extension can include various components. For example, as shown in FIG. 3A-3B and 27-28, the extension can include one or more of an RT template (which includes an RT edit template and a homology arm), a primer binding site, an RT termination signal, an optional 5’ end modifier region, and an optional 3’ end modifier region. FIG. 32 is a flowchart showing an exemplary computerized method 3200 for determining components of an extended gRNA structure, including components of the extension, according to some embodiments. It will be appreciated that, FIG. 32 is intended to be illustrative, and thus the techniques for determining an extended gRNA can include more or fewer steps than those shown in FIG. 32 .
[0417] At step 3202, the computing device determines a set of protospacers in the input alleles on both strands that are compatible with the PAM motif of the selected CRISPR system. In some embodiments, the computing device determines an initial set of protospacers and filters out protospacers whose associated nicking positions are incompatible with guide editing of the output alleles to generate a set of remaining candidate protospacers. For example, the computing device can determine that a protospacer is incompatible because the nick is located 3’ of the desired edit on the strand. As another example, the computing device can determine that the distance between the nick and the desired edit is too great (e.g., greater than a user-defined threshold, e.g., 30 nt, 35 nt, etc.).
[0418] At step 3204, the computing device selects a protospacer from the determined set of protospacers. At step 3206, the computing device determines a spacer and an edit template sequence using the protospacer sequence of the input alleles, the location of the nick, and the sequence of the desired edit. The spacer can comprise a nucleotide sequence of about 20 nucleotides.
[0419] At step 3208, the computing device selects one or more sets of parameters, where each set of parameters includes values for primer binding site length (e.g., can vary in number of nts, such as from about 8 nts to 17 nts), homology arm length (e.g., can vary in number of nts, such as from about 2 nts to 33 nts), and gRNA backbone sequence. For example, the gRNA backbone sequence can be GTTT AAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 1361579), GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 1361580), and / or other gRNA backbone sequences, such as gRNA backbone sequences that preserve wild-type RNA secondary structure.
[0420] At step 3210, the computing device determines a set of parameters selected at step 3208. At step 3212, the computing device determines a homology arm, a primer binding site sequence, and a gRNA backbone using the selected set of parameters. At step 3214, the computing device then forms a resulting PEgRNA sequence by joining the spacer, the gRNA backbone, the PEgRNA extension arm, which includes the homology arm and the editing template. In addition, the extension arm can include a terminator signal, which is a sequence that triggers termination of reverse transcription. Such terminator sequences can include, for example, TTITTTGTTTT (SEQ ID NO: 1361581). In some embodiments, the PEgRNA extension arm can be considered to include the termination signal. In other embodiments, the PEgRNA extension arm can be considered to exclude the termination signal, but where the extension arm is attached to the termination signal as an element located outside of the extension arm.
[0421] The method 3200 proceeds to step 3216, and the computing device determines whether there are more sets of parameters. If yes, the method proceeds to step 3210 and the computing device selects another set of parameters. If no, the method proceeds to step 3218 and the computing device determines whether there are more protospacers. If yes, the method returns to step 3204 and the computing device selects another protospacer from the set of protospacers. If no, the method proceeds to step 3220 and ends.
[0422] As described herein, the extended DNA synthesis template (e.g., the RT template sequence) includes the desired nucleotide change that changes the input allele to the output allele, and includes the RT editing template (e.g., determined in step 3206) and the homology arm (e.g., determined in step 3212). As also described herein, the DNA synthesis template (e.g., the RT template sequence) encodes a single-stranded DNA flap that is complementary to the endogenous DNA sequence adjacent to the nick site. The single-stranded DNA flap contains the desired nucleotide change (e.g., a single nucleotide change, one or more nucleotide insertions, one or more nucleotide deletions, etc.). In some base editing deployments, the single-stranded DNA flap can hybridize to the endogenous DNA sequence adjacent to the nick site to install the desired nucleotide change. In some base editing deployments, the single-stranded DNA flap replaces the endogenous DNA sequence adjacent to the nick site. Cellular repair of the single-stranded DNA flap can result in installation of the desired nucleotide change to form the desired output allele product. The DNA synthesis template (e.g., the RT template sequence) can have a variable number of nucleotides, and can range from about 7 nucleotides to 34 nucleotides.
[0423] While FIG. 32As shown in FIG. 31, the computing device can be configured to determine other components of the extended gRNA. For example, in some embodiments, the computing device is configured to determine an RT termination signal adjacent to the RT template. In some embodiments, the computing device can be configured to determine a first modification adjacent to the RT termination signal. In some embodiments, the computing device is configured to determine a second modification adjacent to the primer binding site.
[0424] The extended gRNA components can be arranged in different structural arrangements, such as those shown in FIGS. 32A-32D. FIGS. 3A-3B and FIGS. 27-28 For example, with reference to FIG. 32A, the extension is located at the 5’ end of the extended gRNA structure, the spacer is located 3’ of the extension and 5’ of the gRNA core. As another example, with reference to FIG. 32B, the spacer is located at the 5’ end of the extended gRNA structure (and 5’ of the gRNA core), and the extension is located at the 3’ end of the extended gRNA structure (and 3’ of the gRNA core). FIG. 3A FIG. 3B For example, with reference to FIG. 32A, the extension is located at the 5’ end of the extended gRNA structure, the spacer is located 3’ of the extension and 5’ of the gRNA core. As another example, with reference to FIG. 32B, the spacer is located at the 5’ end of the extended gRNA structure (and 5’ of the gRNA core), and the extension is located at the 3’ end of the extended gRNA structure (and 3’ of the gRNA core).
[0425] In some embodiments, the computing device accesses a database comprising a set of input alleles and associated output alleles. For example, the computing device can access a database provided by ClinVar that includes hundreds of thousands of mutations, each comprising an allele representing a pathogenic mutation and an allele representing a corrected wild-type sequence. This technique can be used to determine one or more extended gRNA structures for each database entry. FIG. 33 FIG. 33 is a flowchart showing an exemplary computerized method 3300 for determining a set of extended gRNA structures for each mutation entry in a database, according to some embodiments. At step 3302, the computing device accesses a database comprising a set of mutation entries (e.g., a ClinVar database), each mutation entry comprising an input allele representing a mutation and an output allele representing a corrected wild-type sequence.
[0426] At step 3304, the computing device accesses a set of one or more fusion proteins. In some embodiments, this technique can include generating a set of extended gRNA structures for a single fusion protein and / or combinations of different fusion proteins (e.g., for different Cas9 proteins). The computing device can be configured to access data indicating a plurality of fusion proteins and can create a set of extended gRNA structures for each fusion protein as described herein (e.g., a Cas9-NG protein and a SpCas9 protein).
[0427] At step 3306, the computing device selects a fusion protein from the set of fusion proteins. At step 3308, the computing device selects a mutation entry from the set of entries in the database. The computing device can be configured to, for example, loop through each entry in the database and create a set of extended gRNA structures for that entry (e.g., one set for the particular fusion protein, and / or multiple sets for each of multiple fusion proteins). In some embodiments, the computing device can be configured to generate extended gRNA structures for a subset of entries in the database, such as a preconfigured set, a set of mutations with the highest significance (e.g., those with known therapeutic benefits), etc. In some embodiments, if the database includes entries that are not compatible with some fusion proteins for guided editing, the computing device can be configured to determine which entries in the database are compatible for guided editing using the fusion protein selected in step 3304, and select entries that are compatible with the fusion protein selected in step 3308.
[0428] At step 3310, the computing device determines a set of one or more extended gRNA structures using the techniques described herein. The method proceeds from step 3310 to step 3312, and the computing device determines whether there are additional entries in the database. If so, the computing device returns to step 3308 and selects another entry. If not, the computing device proceeds to step 3314 and determines whether there are more fusion proteins. If so, the computing device returns to step 3306 and selects another fusion protein. If not, the computing device proceeds to step 3316 and ends the method 3300.
[0429] In some embodiments, the techniques can design PEgRNAs with gRNA extensions that contain non-complementary sequences, such as non-complementary sequences that are 5' of the homology arm, 3' of the primer binding site, or both. For example, the non-complementary sequences can be designed to form a kissing loop interaction, act as a protection hairpin for RNA stability, etc.
[0430] In some embodiments, the PEgRNAs can be designed using a strategy that prioritizes among multiple design candidates. For example, the techniques can be designed to avoid PEgRNA extensions where the 5'-most nucleotide is a cytosine (e.g., because it disrupts the natural nucleotide-protein interaction in the sgRNA:Cas9 complex). As another example, the techniques can use RNA secondary structure prediction tools to select a preferred PBS length, flap length, etc. based on other parameters of the extended gRNA, such as the protospacer, the desired edit, etc.
[0431] Example implementations of the computerized techniques described herein for determining extended gRNA structures are as follows:
[0432]
[0433]
[0434]
[0435]
[0436]
[0437]
[0438] Using the techniques described herein, an exemplary sequence list was generated with the ClinVar database of input alleles and corresponding output alleles submitted herewith. Entries in the ClinVar database were first filtered to be annotated as pathogenic or likely pathogenic germline mutations. For these examples, Cas9-NG and SpCas9 were used to identify compatible mutations. Of the filtered mutations, approximately 72,020 unique ClinVar mutations were determined to be compatible with prime editing with Cas9-NG, and approximately 63,496 unique ClinVar mutations were identified as compatible with prime editing with SpCas9 of NGG PAM. It will be appreciated that other and / or additional mutations are correctable if using a prime editor containing different Cas9 variants with different PAM compatibilities.
[0439] In various embodiments, the algorithm was used to design therapeutic PEgRNAs of SEQ ID NOs: 1-135514 and 813085-880462, which were designed against ClinVar database entries using the algorithm disclosed herein.
[0440] In various other embodiments, the algorithm is used to design PEgRNAs for the ClinVar database using the algorithm disclosed herein, which is included in the sequence listing, which forms a part of the specification. The sequence listing includes the complete PEgRNA sequences of SEQ ID NOs: 1-135514 and 813085-880462. Each of these complete PEgRNAs individually comprises a spacer region (SEQ ID NOs: 135515-271028 and 880463-947840) and an extension arm (SEQ ID NOs: 271029-406542 and 947841-1015218). In addition, each PEgRNA comprises a gRNA core, e.g., as defined by SEQ ID NOs: 1361579-1361580. The extension arms of SEQ ID NOs: 271029-406542 and 947841-1015218 further each comprise a primer binding site (SEQ ID NOs: 406543-542056 and 1015219-1082596), an editing template (SEQ ID NOs: 542057-677570 and 1082597-1149974), and a homology arm (SEQ ID NOs: 677571-813084 and 1149975-1217352). The PEgRNAs optionally can comprise a 5’-terminal modifier region and / or a 3’-terminal modifier region. The PEgRNAs can also comprise a reverse transcription termination signal at the 3’ of the PEgRNA (e.g., SEQ ID NOs: 1361560-1361566). The application includes the design and use of all of these sequences.
[0441] Mutations were classified into four categories of clinical significance using allele frequency, number of submitters, whether the submitters’ interpretations conflicted, and whether the mutation was reviewed by an expert panel. Of the 63,496 SpCas9-compatible mutations: 4,627 mutations were identified as the most significant level (4); 13,943 mutations were identified as significant level 3 or 4; and 44,385 mutations were identified as significant level 2, 3, or 4.
[0442] The provided sequence listing lists the individual PEgRNAs for each unique mutation, selected as the PEgRNA with the shortest distance between the cut and the edit site. The PEgRNAs involved are those with a homologous arm length of 13 nt, a primer binding site length of 13 nt, a gRNA cut site of 17 nt, and a gRNA length of 20 nt. Pre-spacer regions with nick sites 20 nt further than the edit site were ignored. The main strand sequence of the gRNA used is GTTTAAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 1361579). The terminator sequence used is TTTTTTGTTTT (SEQ ID NO: 1361581).
[0443] As described herein, the exemplary sequence listings provided are not intended to be limiting. It should be understood that variants of the PEgRNA designs provided may include those described herein, including modifications to the gRNA backbone sequence, primer binding site length, flap length, etc.
[0444] An illustrative implementation of computer system 3400 that can be used to perform any aspect of the techniques and embodiments disclosed herein. FIG. 34 As shown in the diagram. Computer system 3400 may include one or more processors 3410 and one or more non-transitory computer-readable storage media (e.g., memory 3420 and one or more non-volatile storage media 3430) and a display 3440. Processor 3410 can control the writing of data to and reading data from memory 3420 and non-volatile storage devices 3430 in any suitable manner, as the aspects of the invention described herein are not limited thereto. To perform the functions and / or techniques described herein, processor 3410 can execute one or more instructions stored in one or more computer-readable storage media (e.g., memory 3420, storage media, etc.), which can be used as a non-transitory computer-readable storage medium from which instructions are executed by processor 3410.
[0445] In conjunction with the techniques described herein, code for, for example, determining the structure of an extended gRNA can be stored on one or more computer-readable storage media of computer system 3400. Processor 3410 can execute any such code to provide any technique for scheduling the operations described herein. Any other software, program, or instructions described herein can also be stored and executed by computer system 3400. It should be understood that computer code can be applied to any aspect of the methods and techniques described herein. For example, computer code can be applied to interact with an operating system to determine the structure of an extended gRNA through regular operating system processes.
[0446] The various methods or processes outlined herein can be programmed into software that is executable on one or more processors employing any one of a variety of operating systems or platforms. In addition, such software can be written using any of a number of suitable programming languages and / or programming or scripting tools, and also can be compiled as executable machine language code or intermediate code that is executed on a virtual machine or a suitable framework.
[0447] In this respect, various inventive concepts can be embodied as one or more non-transitory computer readable storage media (e.g., computer memory, one or more disk drives, compact disks, optical discs, tape, flash memory, Field Programmable Gate Arrays (FPGAs), or other semiconductor devices in circuit arrangements, etc.) encoded with one or more programs that, when executed on one or more computers or other processors, implement the various embodiments of the present application. The non-transitory computer readable medium or media can be transportable, such that the one or more programs stored thereon can be loaded onto any computer resource to implement various aspects of the present application as described above.
[0448] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of embodiments as described above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform some or all of the methods of the present application need not reside on a single computer or processor, but can be distributed in a modular fashion amongst different computers or processors to implement various aspects of the present application.
[0449] Computer-executable instructions can be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically the functionality of the program modules can be combined or distributed as desired in various embodiments.
[0450] Also, data structures can be stored in any suitable place within a non-transitory computer-readable medium, which can be located in a computer-readable storage medium in a storage device. Such data structures can have fields that are related through location, as exemplified by a record in a table with fields arranged in a particular database design. Such relationships can likewise be achieved by assigning storage for the fields with locations that convey the relationship among the fields. However, any suitable mechanism can be used to establish a relationship among information in fields of a data structure, including the use of pointers, tags or other mechanisms that establish relationship among data elements.
[0451] Various inventive concepts can be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments can be constructed that perform acts in an order different than illustrated, which can include performing some acts simultaneously, even though shown as being performed sequentially in illustrative embodiments.
[0452] III. Utilization of therapeutic PEgRNAs for guide editing of editors
[0453] When complexed with a prime editor, a therapeutic PEgRNA designed according to the algorithms disclosed herein can be used to perform prime editing. A prime editor comprises a napDNAbp fused to a polymerase (e.g., a reverse transcriptase) (or provided in trans), optionally wherein the two domains are connected by a linker and further can comprise one or more NLS. These aspects are further described below.
[0454] A.napDNAbp
[0455] The prime editors described herein can comprise a nucleic acid programmable DNA binding protein (napDNAbp).
[0456] In an aspect, the napDNAbp can be associated or complexed with at least one guide nucleic acid (e.g., a guide RNA or a PEgRNA) that localizes the naDNAbp to a DNA sequence comprising a strand of DNA (i.e., a target strand) that is complementary to the guide nucleic acid or a portion thereof (e.g., a spacer of a guide RNA that recombines with a protospacer of a DNA target). In other words, the guide nucleic acid “programs” the napDNAbp (e.g., Cas9 or an equivalent) to locate and bind to the complement of a protospacer in DNA.
[0457] Any suitable napDNAbp can be used in the prime editors described herein. In various embodiments, the napDNAbp can be any Class 2 CRISPR-Cas system, including any Type II, Type V, or Type VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a genome editing tool, the nomenclature used to describe and / or identify CRISPR-Cas enzymes is continually evolving, e.g., Cas9 and Cas9 orthologs. The present application references CRISPR-Cas enzymes whose nomenclature can be old and / or new. The skilled artisan will be able to identify the particular CRISPR-Cas enzyme referenced in the present application based on the nomenclature used, whether it is old (i.e., “legacy”) or new nomenclature. CRISPR-Cas nomenclature is discussed extensively in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol. 1. No. 5, 2018, the entire contents of which are incorporated herein by reference. The particular CRISPR-Cas nomenclature used in any given example of the present application is not limiting in any way, and the skilled artisan will be able to determine which CRISPR-Cas enzyme is referenced.
[0458] For example, the following Type II, Type V, and Type VI Class 2 CRISPR-Cas enzymes have the following art-recognized old (i.e., legacy) and new names. Each of these enzymes, and / or variants thereof, can be used with the prime editors described herein:
[0459]
[0460] See Makarova et al., The CRISPR Journal, Vol. 1, No. 5, 2018.
[0461] Without being bound by theory, the mechanism of action of certain naDNAbps contemplated herein includes a step of R-loop formation, whereby the napDNAbp induces unwinding of the double-stranded DNA target, separating the strands in the region bound by the napDNAbp. The guide RNA spacer then hybridizes to the "target strand" at the protospacer sequence. This displaces the "non-target strand," which is complementary to the target strand, forming a single-stranded region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cleave the DNA leaving various types of lesions. For example, the napDNAbp can comprise a nuclease activity that cleaves the non-target strand at a first location and / or cleaves the target strand at a second location. Depending on the nuclease activity, the target DNA can be cleaved to form a "double-stranded break," cleaving both strands. In other embodiments, the target DNA can be cleaved at only a single site, i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with different nuclease activities include "Cas9 nickases" ("nCas9") and inactive Cas9s without nuclease activity ("dead Cas9" or "dCas9").
[0462] The following description of various napDNAbps that can be used in conjunction with the presently disclosed prime editors is not meant to be limiting in any way. The prime editors can include canonical SpCas9, or any orthologous Cas9 protein, or any variant Cas9 protein - including any naturally occurring Cas9 variant, mutant, or other engineered version - known or that can be made or evolved through directed evolution or other mutagenesis processes. In various embodiments, the Cas9 or Cas9 variant has nickase activity, i.e., cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has no nuclease activity, i.e., a "dead" Cas9 protein. Other variant Cas9 proteins that can be used are those that have a smaller molecular weight than canonical SpCas9 (e.g., for easier delivery) or have a modified or rearranged primary amino acid structure (e.g., in a circular permutation format).
[0463] The guide editors described herein can also comprise Cas9 equivalents, including Casl2a (Cpf1) and Casl2b1 proteins, which are the result of convergent evolution. The napDNAbps used herein (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) can also contain various modifications that alter / enhance its PAM specificity. Finally, the present application encompasses any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a reference SpCas9 canonical sequence or a reference Cas9 equivalent (e.g., Casl2a (Cpf1)).
[0464] The napDNAbp can be a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. As noted above, CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to previous mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand non-complementary to the crRNA is first cleaved by endonucleolytic means, and then trimmed 3’-5’ by exonucleolytic means. In nature, DNA binding and cleavage typically require a protein and two RNAs. However, a single guide RNA (“sgRNA,” or simply “gRNA”) can be engineered to integrate aspects of both the crRNA and the tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), which is incorporated by reference herein in its entirety.
[0465] In some embodiments, the napDNAbp directs cleavage of one or both strands at the location of the target sequence, e.g., within the target sequence and / or within the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a napDNAbp that is mutated relative to the corresponding wild-type enzyme such that the mutated napDNAbp lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. For example, an aspartate to alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, but are not limited to, H840A, N854A, and N863A, referring to the equivalent amino acid positions in the canonical SpCas9 sequence or other Cas9 variants or Cas9 equivalents.
[0466] As used herein, the term "Cas protein" refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that retains all or a substantial portion of the essential basic functions required for the disclosed methods, i.e., (i) the ability to possess nucleic acid programmable binding of a Cas protein to a target DNA, and (ii) the ability to cleave a target DNA sequence on one strand. Cas proteins encompassed herein include CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickases (nCas9) or nuclease-inactivated Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and can include Cas9 equivalents from any Class 2 CRISPR system (e.g., Type II, V, VI), including Casl2a (Cpfl), Casl2e (CasX), Casl2bl (C2cl), Casl2b2, Casl2c (C2c3), C2c4, C2c8, C2c5, C2clO, C2c9 Casl3a (C2c2), Casl3d, Casl3c (C2c7), Casl3b (C2c6), and Casl3b. Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299) and Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?," The CRISPR Journal, Vol. 1. No. 5, 2018, the entire contents of which are incorporated herein by reference.
[0467] The term "Cas9" or "Cas9 nuclease" or "Cas9 moiety" or "Cas9 domain" includes any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of a naturally occurring or engineered Cas9. The term Cas9 is not intended to be particularly limiting and can be referred to as "Cas9 or equivalent." Exemplary Cas9 proteins are further described herein and / or described in the art and incorporated herein by reference. The present disclosure is not limited as to the particular Cas9 employed in the guide editors (PEs) of the present invention.
[0468] As described herein, Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471 :602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference).
[0469] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not intended to be limiting. The guide editors of the present disclosure can use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent.
[0470] (i) wild-type canonical SpCas9
[0471] In one embodiment, the guide editors described herein can comprise the "canonical SpCas9" nuclease from Streptococcus pyogenes, which has been widely used as a tool for genome engineering and classified as a type II sub-group enzyme of the class 2 CRISPR-Cas system. This Cas9 protein is a large multidomain protein that contains two distinct nuclease domains. Point mutations can be introduced into Cas9 to abolish one or both nuclease activities, resulting in either a nickase Cas9 (nCas9) or a dead Cas9 (dCas9), respectively, but they still retain their ability to bind DNA programmably with sgRNA. In principle, Cas9 or its variants (e.g., nCas9) can target this protein to almost any DNA sequence when fused to another protein or domain, by co-expression with the appropriate sgRNA. As used herein, canonical SpCas9 protein refers to the wild-type protein from Streptococcus pyogenes having the following amino acid sequence:
[0472]
[0473]
[0474]
[0475] The guide editors described herein can include canonical SpCas9 or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the wild-type Cas9 provided above. These variants can include SpCas9 variants containing one or more mutations, including any known mutations reported with SwissProt accession number Q99ZW2 entry, which include:
[0476]
[0477]
[0478] Other wild-type SpCas9 sequences that can be used in the present disclosure include:
[0479]
[0480]
[0481]
[0482] The guide editors described herein can include any of the above SpCas9 sequences, or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0483] (ii) Wild-type Cas9 Orthologs
[0484] In other embodiments, the Cas9 protein can be a wild-type Cas9 ortholog from another bacterial species that is different from the canonical Cas9 from S. pyogenes. For example, the following Cas9 orthologs can be used in conjunction with the guide editors constructs described in this specification. In addition, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the following orthologs can also be used with the guide editors described herein.
[0485]
[0486]
[0487]
[0488]
[0489] The guide editors described herein can include any of the above Cas9 ortholog sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0490] The napDNAbp can include any suitable homolog and / or ortholog or naturally occurring enzyme, such as Cas9. Cas9 homologs and / or orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Preferably, the Cas moiety is configured (e.g., mutagenized, recombineering engineered, or otherwise obtained from nature) to be a nickase, i.e., a dual-suitable Cas9 nuclease capable of cutting only the target single strand, and sequences will be apparent to one of skill in the art based on the present disclosure, such Cas9 nuclease and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., the Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein provided by any of the variants of Table 3. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein provided by any of the Cas9 orthologs in the above table.
[0491] (iii) Dead Cas9 variants
[0492] In some embodiments, the prime editors described herein can include a dead Cas9, e.g., a dead SpCas9, which is devoid of nuclease activity due to one or more mutations that inactivate both nuclease domains of Cas9, i.e., the RuvC domain (which cleaves the non-pre spacer DNA strand) and the HNH domain (which cleaves the pre spacer DNA strand). The nuclease inactivation can be due to one or more substitutions and / or deletions in the encoded protein or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity in the amino acid sequence.
[0493] As used herein, the term "dCas9" refers to a nuclease-inactivated Cas9 or nuclease-dead Cas9, or a functional fragment thereof, including any naturally occurring dCas9, any naturally occurring dCas9 equivalent, or a functional fragment thereof, from any organism, any dCas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of a naturally occurring or engineered dCas9. The term dCas9 does not imply a particular limitation and can be referred to as "dCas9 or equivalent." Exemplary dCas9 proteins and methods for making dCas9 proteins are further described herein and / or are described in the art and incorporated herein by reference.
[0494] In other embodiments, dCas9 corresponds to or comprises, in part or in whole, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. In other embodiments, Cas9 variants having mutations other than D10A and H840A are provided that can result in complete or partial inactivation of endogenous Cas9 nuclease activity (e.g., nCas9 or dCas9, respectively). For example, with reference to a wild-type sequence, such as Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 1361424)), such mutations include other amino acid substitutions at D10 and H820, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or RuvC1 subdomain). In some embodiments, variants or homologs of Cas9 are provided (e.g., from a Streptococcus pyogenes Cas9 variant (NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 1361424)) that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 1361424). In some embodiments, variants of dCas9 are provided (e.g., variants of NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 1361424)) that have a shorter or longer amino acid sequence of about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more than NC_017053.1 (SEQ ID NO: 1361424).
[0495] In an embodiment, the dead Cas9 can be based on the canonical SpCas9 sequence of Q99ZW2 and can have the following sequence, which includes D10A and H810A substitutions (underlined and bolded), or a variant of SEQ ID NO: 1361444 in which there is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity:
[0496]
[0497] (iv) Cas9 nickase variants
[0498] In an embodiment, the prime editors described herein comprise a Cas9 nickase. The terms "Cas9 nickase," "nCas9" refer to a Cas9 variant that is capable of introducing a single-strand break in a target double-stranded DNA molecule. In some embodiments, a Cas9 nickase comprises only a single functional nuclease domain. Wild-type Cas9 (e.g., canonical SpCas9) comprises two independent nuclease domains, a RuvC domain (which cleaves the non-protospacer sequence DNA strand) and an HNH domain (which cleaves the protospacer sequence DNA strand). In an embodiment, a Cas9 nickase comprises a mutation in the RuvC domain that inactivates the RuvC nuclease activity. For example, mutations in aspartate (D) 10, histidine (H) 983, aspartate (D) 986, or glutamate (E) 762 have been reported as loss-of-function mutations for the production of a RuvC nuclease domain and a functional Cas9 nickase (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, incorporated herein by reference). Thus, a nickase mutation in the RuvC domain can include D10X, H983X, D986X, or E762X, where X is any amino acid other than the wild-type amino acid. In some embodiments, the nickase can be D10A, H983A, or D986A, or E762A, or a combination thereof.
[0499] In various embodiments, a Cas9 nickase can have a mutation in the RuvC nuclease domain and have one of the following amino acid sequences, or a variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the amino acid sequence.
[0500]
[0501]
[0502]
[0503] In another embodiment, the Cas9 nickase comprises a mutation in the HNH domain that inactivates HNH nuclease activity. For example, mutations in histidine (H) 840 or asparagine (R) 863 have been reported as loss-of-function mutations for the production of the HNH nuclease domain and functional Cas9 nickase (e.g., Nishimasu et al., "Crystal structure of Cas9 in complex with guide RNA and target DNA," Cell 156(5), 935-949, incorporated herein by reference). Thus, the nickase mutation in the HNH domain can include H840X and R863X, where X is any amino acid other than the wild-type amino acid. In some embodiments, the nickase can be H840A or R863A, or a combination thereof.
[0504] In various embodiments, the Cas9 nickase can have a mutation in the HNH nuclease domain and have one of the following amino acid sequences, or a variant thereof having an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity.
[0505]
[0506]
[0507] (v) Other Cas9 variants
[0508] In addition to dead Cas9 and Cas9 nickase variants, Cas9 proteins used herein can also include other "Cas9 variants" that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 protein, including any wild-type Cas9, or a mutant Cas9 (e.g., a dead Cas9 or Cas9 nickase), or a Cas9 fragment, or a circularly permuted Cas9, or other variants of Cas9 disclosed herein or known in the art. In some embodiments, a Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a reference Cas9. In some embodiments, a Cas9 variant comprises a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of a wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9 (e.g., SEQ ID NO: 1361421).
[0509] In some embodiments, the disclosure can also utilize Cas9 fragments that retain their functionality and are fragments of any Cas9 protein disclosed herein. In some embodiments, a Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0510] In various embodiments, the prime editors disclosed herein can comprise one of the Cas9 variants described below, or a Cas9 variant thereof that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any of the reference Cas9 variants.
[0511] (vi) Small Cas9 variants
[0512] In some embodiments, the prime editors encompassed herein can include Cas9 proteins having a molecular weight that is less than the canonical SpCas9 sequence. In some embodiments, the smaller size of the Cas9 variants can facilitate delivery to a cell, for example, by an expression vector, nanoparticle, or other delivery means. In certain embodiments, the smaller size of the Cas9 variants can include enzymes classified as Type II enzymes of the Class 2 CRISPR-Cas system. In some embodiments, the smaller size of the Cas9 variants can include enzymes classified as Type V enzymes of the Class 2 CRISPR-Cas system. In other embodiments, the smaller size of the Cas9 variants can include enzymes classified as Type VI enzymes of the Class 2 CRISPR-Cas system.
[0513] The canonical SpCas9 protein is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. As used herein, the term "small size Cas9 variant" refers to any Cas9 variant - naturally occurring, engineered, or otherwise - that is less than at least 1300 amino acids, or at least less than 1290 amino acids, or less than 1280 amino acids, or less than 1270 amino acids, or less than 1260 amino acids, or less than 1250 amino acids, or less than 1240 amino acids, or less than 1230 amino acids, or less than 1220 amino acids, or less than 1210 amino acids, or less than 1200 amino acids, or less than 1190 amino acids, or less than 1180 amino acids, or less than 1170 amino acids, or less than 1160 amino acids, or less than 1150 amino acids, or less than 1140 amino acids, or less than 1130 amino acids, or less than 1120 amino acids, or less than 1110 amino acids, or less than 1100 amino acids, or less than 1050 amino acids, or less than 1000 amino acids, or less than 950 amino acids, or less than 900 amino acids, or less than 850 amino acids, or less than 800 amino acids, or less than 750 amino acids, or less than 700 amino acids, or less than 650 amino acids, or less than 600 amino acids, or less than 550 amino acids, or less than 500 amino acids, but at least greater than about 400 amino acids and retains the functions required of a Cas9 protein. Cas9 variants can include those classified as Type II, Type V, or Type VI enzymes of the Class 2 CRISPR-Cas system.
[0514] In various embodiments, the guided editors disclosed herein can comprise one of the small size Cas9 variants described below, or a Cas9 variant thereof that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference small size Cas9 protein.
[0515]
[0516]
[0517] (vii) Cas9 equivalents
[0518] In some embodiments, the prime editors described herein can include any Cas9 equivalent. As used herein, the term "Cas9 equivalent" is a broad term and includes any napDNAbp protein that functions the same as Cas9 in the present prime editors, although its amino acid primary sequence and / or its three-dimensional structure can be different and / or unrelated from an evolutionary perspective. Thus, while a Cas9 equivalent includes any Cas9 ortholog, homolog, mutant, or variant described or encompassed herein that is evolutionarily related, a Cas9 equivalent also includes proteins that can have evolved through a process of convergent evolution to have the same or similar function as Cas9, but they do not necessarily have any similarity in amino acid sequence and / or three-dimensional structure. The prime editors described herein encompass any Cas9 equivalent that will provide the same or similar function as Cas9, although the Cas9 equivalent can be based on a protein that has arisen through convergent evolution. For example, if Cas9 refers to a type II enzyme of the CRISPR-Cas system, then a Cas9 equivalent can refer to a type V or VI enzyme of the CRISPR-Cas system.
[0519] For example, Casl2e (CasX) is a Cas9 equivalent that is reported to have the same function as Cas9 but evolved through convergent evolution. Thus, the Casl2e (CasX) protein described by Liu et al. in "CasX enzymes comprises a distinct family of RNA-guided genome editors," Nature, 2019, Vol. 566: 218-223 is contemplated for use with the prime editors described herein. Further, any variant or modification of Casl2e (CasX) is contemplated and within the scope of the present disclosure.
[0520] Cas9 is a bacterial enzyme that has evolved in a variety of species. However, the Cas9 equivalents encompassed herein can also be obtained from archaea, which constitute a domain and kingdom of single-celled, prokaryotic microorganisms that are distinct from bacteria.
[0521] In some embodiments, Cas9 equivalents can refer to Casl2e (CasX) or Casl2d (CasY), which have been described in, e.g., Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entirety of which is incorporated herein by reference. Using genome-resolved metagenomics, a number of CRISPR-Cas systems were identified, including a Cas9 that was first reported in the archaeal domain of life. This distinct Cas9 protein was found in the poorly studied nanarchaea as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-Casl2e and CRISPR-Casl2d, were discovered, which are among the most compact systems ever discovered. In some embodiments, Cas9 refers to Casl2e, or a variant of Casl2e. In some embodiments, Cas9 refers to Casl2d, or a variant of Casl2d. It will be appreciated that other RNA-guided DNA binding proteins can be used as a nucleic acid programmable DNA bin...
Claims
1. A guide RNA comprising a spacer sequence, a gRNA core, and an extension arm, wherein the extension arm comprises (i) a primer binding site, (ii) an editing template, and (iii) a homology arm; wherein the spacer sequence is 10 to 40 nucleotides in length and is capable of annealing to a target strand of a target DNA sequence; wherein the gRNA core is capable of binding to a Cas protein, wherein the Cas protein has nickase activity; wherein the extension arm is capable of serving as a template for reverse transcription of a DNA single strand and comprises, in the 5’ to 3’ direction, the homology arm, the editing template, and the primer binding site; wherein the primer binding site is 8 to 20 nucleotides in length and is capable of annealing to a 3’-end ssDNA flap that is formed after the Cas protein cleaves a non-target strand of the target DNA sequence and serves as a primer sequence to prime reverse transcription of the DNA single strand; wherein the editing template encodes a desired edit in the DNA single strand; and wherein the homology arm encodes a portion of the DNA single strand that is complementary to the target strand of the target DNA sequence.
2. The guide RNA of claim 1, wherein the spacer sequence is about 20 nucleotides in length.
3. The guide RNA of claim 1, wherein the extension arm is a 5’ extension arm.
4. The guide RNA of claim 1, wherein the extension arm is a 3’ extension arm.
5. The guide RNA of claim 1, wherein the desired edit comprises one or more nucleotide substitutions, one or more deletions, one or more insertions, or a combination thereof.
6. The guide RNA of any one of claims 1-5, further comprising a 5’-terminal modifier region comprising a hairpin sequence, a stem-loop sequence, or a toe-loop sequence.
7. The guide RNA of any one of claims 1-5, further comprising a 3’-terminal modifier region comprising a hairpin sequence, a stem-loop sequence, or a toe-loop sequence.
8. A polynucleotide encoding the guide RNA of any one of claims 1-7.
9. A vector comprising the polynucleotide of claim 8.
10. A pharmaceutical composition comprising: (a) the guide RNA of any one of claims 1-7, the polynucleotide of claim 8, or the vector of claim 9 and (b) a pharmaceutically acceptable excipient.
11. A prime editing complex comprising a Cas protein, a reverse transcriptase, and the guide RNA of any one of claims 1-7, wherein the Cas protein has nickase activity.
12. The prime editing complex of claim 11, wherein the Cas protein and the reverse transcriptase form a fusion protein.
13. The prime editing complex of claim 11, wherein the Cas protein is a Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, or Cas12c protein.
14. The prime editing complex of any one of claims 11-13, wherein the Cas protein is a Cas9 nickase.
15. One or more polynucleotides encoding the prime editing complex of any one of claims 11-14.
16. One or more vectors comprising the one or more polynucleotides of claim 15.
17. A pharmaceutical composition comprising: (a) the prime editing complex of any one of claims 11-14, the one or more polynucleotides of claim 15, or the one or more vectors of claim 16, and (b) a pharmaceutically acceptable excipient.
18. A composition comprising the guide RNA of any one of claims 1-7 and one or more polynucleotides encoding a prime editor comprising a Cas protein and a reverse transcriptase, wherein the Cas protein has nickase activity.
19. A composition comprising the polynucleotide of claim 8 or the vector of claim 9 and one or more polynucleotides encoding a prime editor comprising a Cas protein and a reverse transcriptase, wherein the Cas protein has nickase activity.
20. The composition of claim 18 or 19, wherein the prime editor is a fusion protein of the Cas protein and the reverse transcriptase.
21. The composition of claim 18 or 19, wherein the Cas protein is a Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, or Cas12c protein.
22. The composition of claim 18 or 19, wherein the Cas protein is a Cas9 nickase.
23. Use of the guide RNA of any one of claims 1-7, the polynucleotide of claim 8, the vector of claim 9, the pharmaceutical composition of claim 10 or 17, the prime editing complex of any one of claims 11-14, the one or more polynucleotides of claim 15, the one or more vectors of claim 16, or the composition of any one of claims 18-22 in the manufacture of a medicament.
Citation Information
Patent Citations
Split inteins, conjugates and uses thereof
EP2877490A2
Stabilized reverse transcriptase fusion proteins
US10150955B2
Non-nucleoside reverse transcriptase inhibitors
US10189831B2
Methods for determining hypersusceptibility of HIV-1 to non-nucleoside reverse transcriptase inhibitors
US10202658B2
Regulation of endogenous gene expression in cells using zinc finger proteins
US20030087817A1