Compositions and methods for production of double-stranded DNA in cells for programmable gene integration
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2026-03-12
AI Technical Summary
Existing genome engineering methods for large gene integration, such as PASSIGE, CAST, and quadruple flap prime editing, face inefficiencies in delivering exogenous DNA cargo, leading to cytotoxicity and limited applicability in primary cells due to innate immune system detection.
A system for producing DNA from RNA within cells using a reverse transcriptase and cargo RNA with specific sequences, including a primer binding site and polypurine tract, enabling synthesis of donor DNA molecules up to 10,000 base pairs in length, reducing cytotoxicity and enhancing delivery efficiency.
Facilitates large gene integration with reduced cytotoxicity, allowing for efficient delivery and integration of therapeutic genes in various cell types, including primary cells, by synthesizing donor DNA molecules intracellularly.
Smart Images

Figure US2025038886_12032026_PF_FP_ABST
Abstract
Description
COMPOSITIONS AND METHODS FOR PRODUCTION OF DOUBLE-STRANDED DNA IN CELLS FOR PROGRAMMABLE GENE INTEGRATIONRELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application, U.S.S.N. 63 / 674,589, filed July 23, 2024, and U.S. Provisional Application, U.S.S.N. 63 / 680,392, filed August 7, 2024, each of which is incorporated herein by reference.GOVERNMENT SUPPORT
[0002] This invention was made with government support under grant numbers U19 NS132304, R01 EB031172, RM1 HG009490, and R35 GM118062, awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND OF THE INVENTION
[0003] Genome engineering technologies such as base editing and prime editing have significant therapeutic potential. In principle, base editing and prime editing together can correct almost every known genetic variant associated with human disease. Developing a base or prime editing strategy for each disease variant, however, requires multiple rounds of optimization, often taking months or longer for each target. Improved methods for programmable large gene integration could enable the development of editing strategies that target many disease variants at once, facilitating the development of personalized medicines to benefit larger populations.
[0004] Large gene integration technologies theoretically enable application of a single strategy for one composition of matter to target many disease variants, or enable installation of loss of function genes for cell therapies. Such technologies that have been developed to date include: 1) PASSIGE, which couples with dual flap prime editing to install the landing pad for the serine recombinase Bxbl for site-specific genomic integration of DNA cargo; 2) CAST, which is also an approach for programmable integration of large DNA sequences in human cells that avoids the generation of double strand breaks by using type-I CRISPR- associated transposase; and 3) quad flap prime editing, which utilizes four pegRNAs to simultaneously integrate a plasmid donor into a targeted genomic site through exon swapping (see FIG. 1). All three of these methods require the use of exogenous DNA cargo for precise insertion into the genome. However, the delivery of DNA produced in vitro to target cells can be inefficient and problematic.
[0005] While methods that use plasmid DNA or linearized double-stranded DNA (for example, PCR products) could theoretically allow for the delivery of large pay loads, they also typically cause substantial cytotoxicity, especially in primary cells like T cells. Naked DNA cytotoxicity is usually due to the innate immune system (cGas-STING pathway) that functions to detect cytosolic DNA. This limits the use of targeted insertion technologies for in vivo integration. Alternative means to provide cargo DNA for insertion into genomes are therefore needed.SUMMARY OF THE INVENTION
[0006] The present disclosure describes the development of systems and methods for producing DNA from RNA in cells that are compatible with methods for large gene integration (including prime editing-based methods of large gene integration, such as PASSIGE, CAST, and quadruple flap prime editing as described herein). In some situations, typical methods of prime editing can only be used to insert sequences of up to approximately 100 base pairs in length into a genome, but large gene integration methods such as those described herein can be used to insert sequences thousands of base pairs in length. The systems and methods for DNA synthesis from RNA described herein facilitate these methods by providing a means to synthesize the donor DNA molecule to be inserted into a genome within cells. This is advantageous because delivery of donor DNA molecules to cells directly often results in unacceptable levels of cytotoxicity.
[0007] Thus, in one aspect, the present disclosure provides systems comprising a first polynucleotide encoding a reverse transcriptase, and a second polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoterenhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA.
[0008] In some embodiments, the cargo RNA does not comprise a portion of a viral genome. In some embodiments, the reverse transcriptase is not a wild-type reverse transcriptase.
[0009] In some embodiments, the cargo RNA comprises one or more genes of interest. In some embodiments, the one or more genes of interest comprise one or more therapeutic genes. In some embodiments, the cargo RNA is greater than 100 base pairs in length. In someembodiments, the cargo RNA is up to 3000 base pairs in length. In some embodiments, the cargo RNA is up to 10,000 base pairs in length.
[0010] In some embodiments, the system further comprises one or more prime editors, prime editing guide RNAs (pegRNAs), recombinases, and / or transposases. In some embodiments, the system further comprises one or more polynucleotides encoding one or more prime editors, prime editing guide RNAs (pegRNAs), recombinases, and / or transposases.
[0011] In another aspect, the present disclosure provides polynucleotides comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and further comprises a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA.
[0012] In another aspect, the present disclosure provides vectors comprising any of the polynucleotides provided herein.
[0013] In another aspect, the present disclosure provides compositions comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the composition comprises any components of any of the systems described herein.
[0014] In another aspect, the present disclosure provides cells comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the cell comprises any components of any of the systems described herein.
[0015] In another aspect, the present disclosure provides kits comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the kits comprises any components of any of the systems described herein.
[0016] In another aspect, the present disclosure provides methods comprising expressing in the cell a first polynucleotide encoding a reverse transcriptase, and contacting a second polynucleotide with the reverse transcriptase, wherein the second polynucleotide comprises a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoterenhancer region, and a second region comprising the same or substantially the same sequence as the first region of the sequence 5' of the cargo RNA. In some embodiments, the methodsfurther comprise using a prime editing-based strategy (e.g., PASSIGE, CAST, or quadruple flap prime editing) to insert the double- stranded DNA, or a portion thereof, into a genomic sequence.
[0017] The foregoing concepts, and additional concepts discussed below, may be arranged in any suitable combination, as the present disclosure is not limited in this respect. Further, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying figures.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0019] FIG. 1 provides schematics of prime editing-based methods for therapeutic large gene integration. Such methods are useful, for example, for applying a single strategy to target many disease variants and / or for installing loss-of-function genes for cell therapies. Such technologies include: 1) PASSIGE, which couples with dual flap prime editing to install the landing pad for the serine recombinase Bxbl for site- specific genomic integration of DNA cargo; 2) CAST, which is also an approach for programmable integration of large DNA sequences in human cells that avoids the generation of double strand breaks by using type-I CRISPR-associated transposase; and 3) quad flap prime editing, which utilizes four pegRNAs to simultaneously integrate a plasmid donor into a targeted genomic site through exon swapping. All three of these methods require the use of exogenous DNA cargo for precise insertion into the genome.
[0020] FIG. 2 provides a schematic showing how an RNA-based molecule can be used for delivery of donor DNA to a cell. RNA can be used as a template to generate a doublestranded DNA donor for targeted integration by reverse transcription. This provides advantages of reducing cytotoxicity compared to direct delivery of a DNA donor, synergizing with mRNA delivery technology e.g., lipid nanoparticles (LNPs)), and coupling with prime editing by using the same reverse transcriptase as the prime editor.
[0021] FIG. 3 shows a strategy for RNA-to-DNA conversion using a retron non-coding RNA molecule and a retron-type reverse transcriptase.
[0022] FIG. 4 shows RNA-to-DNA conversion using non-LTR retrotransposons.
[0023] FIG. 5 shows RNA-to-DNA conversion using LTR retrotransposons / retrovirus. U3: enhancer and promoter region for transcription. R: start and termination sites for transcription. U5: polyadenylation signal.
[0024] FIGs. 6A-6B provides a schematic showing more details of the strategy shown in FIG. 5. FIG. 6A shows synthesis of the first DNA strand. FIG. 6B shows synthesis of the second DNA strand. This strategy is advantageous because minimal components must be delivered into cells, since only an RNA template and reverse transcriptase are required.
[0025] FIG. 7 provides a plasmid-based assay that allows detection of RNA to DNA.
[0026] FIGs. 8A-8B show that evolved reverse transcriptases show activity generating DNA from RNA templates.
[0027] FIGs. 9A-9B show that nuclease Pl digestion indicates the presence of doublestranded DNA generated from RNA.
[0028] FIG. 10 shows that divergent PCR confirms the presence of circular DNA products.
[0029] FIG. 11 provides a schematic demonstrating the use of RNA to double- stranded DNA donor conversion for use in combination with evoBxbl recombination.
[0030] FIG. 12 shows that flow cytometry confirms double-stranded production.
[0031] FIG. 13 shows that ddPCR indicates RNA to double-stranded DNA donor conversion and evoBxbl recombination.
[0032] FIG. 14 shows screening of an array of reverse transcriptase systems. Reverse transcriptases marked with an asterisk use LTR reverse transcription mechanism.
[0033] FIG. 15 provides a schematic of the workflow for the screening shown in FIG. 14.
[0034] FIG. 16 shows that wild type reverse transcriptase does not generate double-stranded DNA at the same levels as reverse transcriptase variants.
[0035] FIG. 17 shows that the evolved reverse transcriptases function on varied LTR templates. PE6d in combination with its cognate MMLV LTR architecture displayed the highest editing efficiency.
[0036] FIG. 18 shows that reverse transcriptase-MS2 coat protein (MCP) fusion increases DNA production. Co-transfection of PE6d reverse transcriptase fused with MCP and the cognate MMLV LTR RNA template increase DNA production. MCP likely plays a role in stabilizing the reverse transcriptase and / or binding to other small stem loops from the MMLV LTR RNA template.
[0037] FIGs. 19A-19D show that programmable gene integration technologies require exogenous DNA donors. FIG. 19A shows Cas9-mediated HDR (homology-dependent DNA repair). FIG. 19B shows eePASSIGE (Twin PE x Serine Recombinase BxBl). FIG. 19C shows CASTs (Type I-F CRISPR-associated transposases). FIG. 19D shows Bridge RNAs (RNA-guided recombinase).
[0038] FIG. 20 shows an RNA-based donor system coupled with reverse transcriptases for intracellular conversion of RNA to dsDNA.
[0039] FIG. 21 shows a schematic of bacteria retrons.
[0040] FIG. 22 provides a schematic showing how non-LTR retrotransposons couple reverse transcription and integration.
[0041] FIG. 23 shows LTR-based retroelements: second strand synthesis.
[0042] FIG. 24 shows a method for DNA product detection. A self-splicing intron enabled the plasmid-based method for reverse transcription product detection.
[0043] FIG. 25 shows that no false positive was detected by plasmid-based assay. The structure of the T. thermophila rRNA group I intron (Tet Intron having the nucleotide sequence “uccuguAAUUAGCAAUAUAAUGAAUUGCAGGGGaguucauu” (SEQ ID NO: 31)) is shown in the right panel.
[0044] FIG. 26 shows neo junction amplitude changes in relation to ACTB amplitude.
[0045] FIGs. 27A-27B show use of ssDNA exonucleases to confirm dsDNA product.
[0046] FIG. 28 shows that the processivity of reverse transcriptases affects dsDNA synthesis. Low processivity promotes recombination between viral RNA templates. Reverse transcription is viral capsid dependent. Capsids cage the reverse transcriptases and prevent dissociation of RT from RNA.
[0047] FIGs. 29A-29B show that evolved M-MLV RT with high processivity produced dsDNA from RNA. The evolved M-MLV RT enzyme was from phage-assisted continuous evolution in the context of prime editing.
[0048] FIG. 30 shows a VI RNA-based system with intracellular reverse transcription for DNA delivery. The calculation for copy number of donor molecules: cone. (dsDNA) / cone. (beta- actin) x 3 (Hela cells are triploid).
[0049] FIG. 31 shows V2: Harnessing MS2-MCP interaction for recruiting RTs to RNA templates.
[0050] FIG. 32 shows a V2 RNA-based system with intracellular reverse transcription for DNA delivery.
[0051] FIG. 33 shows V3.1: Inactivation of RNaseH domain improved dsDNA production.
[0052] FIG. 34 shows V3.2: Protein MPNN designed PE6d variants improved dsDNA production.
[0053] FIG. 35 shows V4: Overexpression of DNA repair proteins improved dsDNA production.
[0054] FIGs. 36A-36B show V5: RNA template engineering. Murine leukemia virus orthologs were investigated for RNA template engineering.
[0055] FIG. 37 shows RNA-based transduction.
[0056] FIGs. 38A-38B show that RNA quality and integrity matter for donor molecule production. FIG. 38A shows a PCR-based transcription template. FIG. 39B shows a restriction digest-based transcription template.
[0057] FIGs. 39A-39B show that an all RNA-based system enabled intracellular production of DNA donor molecules. FIG. 39A shows delivery of RNA components by MessengerMax. FIG. 39B shows all RNA delivery of RNA to DNA system by MessengerMax HEK293T cells. Plasmid-based engineering can transit to all RNA delivery system. More than 600 copies of donor molecules were made per cell.
[0058] FIG. 40 shows that IDLV experiments were used for benchmarks of targeted gene integration. Around 200 copies of donor molecules per cell were required to achieve efficient (10%) integration. The calculation for copy number of donor molecules: cone. (dsDNA) / cone. (beta- actin) x 2 (Fibroblasts are diploid).
[0059] FIGs. 41A-41B show that an all RNA-based system enabled targeted gene integration. FIG. 41A provides a schematic demonstrating the use of RNA to doublestranded DNA donor conversion for use in combination with evoBxbl recombination and pre-installed recombinase attachment sites (attP). The recombinase and the pre-install attachment site cell line were incorporated to assess the functionality of produced donors. FIG. 41B shows an all RNA delivery of RNA to DNA system and eeBxbl by MessengerMax. All RNA gene integration achieved approximately 20% integration.
[0060] FIGs. 42A-42B show an all RNA delivery for programmable large gene integration. Neither HEK293T nor N2a cells had sufficient cGAS STING sensing. All RNA delivery of Twin PE RNA to DNA system and eeBxbl by MessengerMax HEK293T cells (FIG. 42A) and N2a cells (FIG. 42B).
[0061] FIG. 43 shows all RNA delivery of eePASSIGE by electroporation and MessengerMax WI-38 primary human fibroblasts. RNA donors showed activity for one-poteePASSIGE in primary fibroblasts WI-38. Electroporation: Twin PE, eeBxbl, and RT. MessengerMax 2 hours afterwards: RNA template. Optimization on the way reached 10%.
[0062] FIG. 44 provides a schematic of the initial investigation of the cGAS STING pathway by measuring the cGAMP level. Upon binding to DNA, cGAS produces the cyclic dinucleotide second messenger cGAMP.
[0063] FIGs. 45A-45D show that RNA donors showed less cGAMP level compared with DNA donors. Upon binding to DNA, cGAS produced the cyclic dinucleotide second messenger cGAMP. Hela cells with sufficient cGAS activity also showed similar trend. cGAMP measurement was normalized to cell number in this assay. FIG. 45A shows percent integration of RNA delivery of eePASSIGE by MessengerMax H1299 cells. FIG. 45B shows cGAMP level of RNA delivery of eePASSIGE by MessengerMax H1299 cells. FIG. 45C shows percent integration of RNA delivery of eePASSIGE by MessengerMax Hela cells. FIG. 45D shows cGAMP level of RNA delivery of eePASSIGE by MessengerMax Hela cells.DEFINITIONS
[0064] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and. Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.Complementary
[0065] The term “complementary” is used herein to refer to two oligonucleotide sequences (e.g., DNA or RNA) comprising bases that hydrogen bond to one another. The degree of complementarity between two oligonucleotide sequences can vary, from complete complementarity to no complementarity. Two sequences may be complementary or partially complementary to one another (e.g.. two sequences may have 100% complementarity, 99% complementarity, 98% complementarity, 97% complementarity, 96% complementarity, 95% complementarity, 90% complementarity, 85% complementarity, 80% complementarity, or less than 80% complementarity). In some embodiments, a sequence is complementary orpartially complementary to only a portion of another sequence. In some embodiments, a sequence is complementary or partially complementary to another sequence under certain conditions (e.g., certain salt concentrations, pHs, etc.).Long Terminal Repeat (LTR) and LTR Elements
[0066] A “long terminal repeat (LTR)” is a pair of identical sequences of DNA, generally several hundred base pairs long, which occur in eukaryotic genomes on either end of a series of genes or pseudogenes that form a retrotransposon, an endogenous retrovirus, or a retroviral provirus. LTR retrotransposons have direct long terminal repeats that range from about 100 bp to over 5 kb in size. In nature, an element flanked by a pair of LTRs will typically encode a reverse transcriptase and an integrase, allowing the element to be copied and inserted at a different location of the genome. In the systems, polynucleotides, and methods of the present disclosure, the element typically encoding a reverse transcriptase and integrase is replaced with a cargo RNA comprising any sequence of interest, as discussed further herein. The multi-step process of reverse transcription results in the placement of two identical LTRs, each consisting of a “U3,” “R,” and “U5” region, at either end of the proviral DNA. The ends of the LTRs subsequently participate in integration of the provirus into the host genome.
[0067] LTRs are typically segmented into the U3 region (“unique 3 ' region”), R region (“regulatory region”), and U5 region (“unique 5 ' region”). The U3 and U5 regions generally contain transcription factor sites. Specifically, the U3 region contains enhancer and promoter elements to facilitate reverse transcription, and the U5 region frequently contains a polyadenylation signal, which in nature plays a role in genome packaging. The R region defines the transcription initiation site and the transcription termination site.Primer Binding Site
[0068] A “primer binding site” or “PBS” refers to a sequence to which a nucleic acid primer molecule (e.g., a tRNA as described herein) binds and acts as a primer for initiation of reverse transcription. In some embodiments, a PBS is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In certain embodiments, a PBS is 18 nucleotides in length. Specific cell tRNAs bind to the PBS and act as primers for reverse-transcription, which occurs in a complex and multi-step process, ultimately producing a double- stranded cDNA molecule. In some embodiments, a PBS has a specific sequence that binds to a tRNA. In certain embodiments, a PBS has a specific sequence that binds to tRNALys.Polypurine Tract (PPT)
[0069] A “polypurine tract” or “PPT” refers to a sequence of repeating purines (z.e., adenines or guanines). In some embodiments, a PPT comprises a sequence of at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 repeating purines in length. In some embodiments, the PPT comprises or consists of RNA. The PPT is resistant to digestion by RNase and acts as a primer for synthesis of the second strand of a double-stranded DNA molecule in the systems and methods described herein.PASSIGE
[0070] ‘ ‘Prime editing-assisted site-specific integrase gene editing” or “PASSIGE” is a platform that uses prime editing and site- specific recombinases to integrate multi-kilobase DNA cargo into targeted sites, for example, in a genome. In PASSIGE, single-flap or dualflap prime editing installs a site-specific recombinase landing site into a target genomic location. The corresponding recombinase then catalyzes the insertion of the cargo DNA into the landing site, resulting in targeted integration. PASSIGE may be performed with a singletransfection by simultaneously delivering the prime editor, pegRNA(s), nicking sgRNA (if desired), donor DNA plasmid, and recombinase, or by using two successive transfections to perform the prime editing step and recombination step at different times. PASSIGE is described further in International PCT Application No. PCT / US2024 / 014998, filed February 8, 2024, and Pandey, S. el al., Nat. Biomed. Eng. (2024), each of which is incorporated herein by reference.Quadruple Flap Prime Editing
[0071] “Quadruple-flap prime editing” refers to a prime editing approach that can directly carry out integration of large sequences (e.g., greater than 100 base pairs) without the need for a recombinase enzyme (as in single-flap and dual-flap prime editing) but with similar versatility. In quadruple-flap prime editing, four different pegRNA sequences are delivered to cells along with the prime editor. One pair of pegRNAs template the synthesis of complementary DNA flaps, as with dual-flap prime editing, while the other pair of pegRNAs template the synthesis of an orthogonal pair of complementary DNA flaps on a second DNA molecule (e.g.. a donor DNA molecule). The location of these flaps, their relative orientation, and the pairing of complementarity dictates which class of rearrangement will occur. The junctions of the rearranged DNA contain the sequence encoded by the pegRNA-templated flaps, which direct the orientation of the rearrangement. Quadruple flap prime editing isdescribed further, for example, in International PCT Application Publication No. WO 2021 / 226558, published November 11, 2021, which is incorporated herein by reference.CAST
[0072] ‘ ‘CAST,” or “CRTS PR-associated transposon,” refers a system for large gene integration that relies on transposases (e.g., TnsA, TnsB, TnsC, and TniQ) and Cas proteins (e.g., Cas5, Cas6, Cas7, and Cas8) for targeted integration of large sequences into genomes. CAST is further described, for example, in International PCT Application Publication No. PCT / US2024 / 015825, filed February 14, 2024, which is incorporated herein by reference. pegRNA
[0073] As used herein, the terms “prime editing guide RNA,” “PEgRNA,” “pegRNA,” or “extended guide RNA” refer to a specialized form of a guide RNA that has been modified to include one or more additional sequences for implementing the prime editing as described herein. As described herein, the prime editing guide RNAs comprise one or more “extended regions,” also referred to herein as “extension arms,” of nucleic acid sequence. The extended regions may comprise, but are not limited to, single-stranded RNA or DNA. Further, the extended regions may occur at the 3' end of a traditional guide RNA. In other arrangements, the extended regions may occur at the 5' end of a traditional guide RNA. In still other arrangements, the extended region may occur at an intramolecular region of the traditional guide RNA, for example, in the gRNA core region which associates and / or binds to the napDNAbp. The extended region comprises a “DNA synthesis template” or “reverse transcriptase template” that encodes (by the polymerase / reverse transcriptase of the prime editor) a single- stranded DNA which, in turn, has been designed to be (a) homologous with the endogenous target DNA to be edited, and (b) which comprises at least one desired nucleotide change (e.g., a transition, a transversion, a deletion, or an insertion) to be introduced or integrated into the endogenous target DNA. The extended region may also comprise other functional sequence elements, such as, but not limited to, a “primer binding site” and a “linker” sequence, or other structural elements, such as, but not limited to, aptamers, stem loops, hairpins, toe-loops (e.g., a 3' toeloop), or an RNA-protein recruitment domain (e.g., MS2 hairpin). As used herein, the “primer binding site” comprises a sequence that hybridizes to a single- strand DNA sequence having a 3' end generated from the nicked DNA of the R-loop.PEI
[0074] As used herein, “PEI” refers to a PE complex comprising a fusion protein comprisingCas9(H840A) and a wild type MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)] + a desired pegRNA, wherein the PE fusion has the following amino acid sequence:MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAK VDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPIN ASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEI TKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHA ILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPW NFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYV TEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVE DRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNF MQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNK VLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLV SDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYD VRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIV WDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDW DPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNF LYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDAT LIHQSITGLYETRIDLSOLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQY PMSQEAREGIKPHIQREEDQGIEVPCQSPWNTPEEPVKKPGTNDYRPVQDEREVNKRVED 1HPTVPNPYNEESGEPPSHQWYTVEDEKDAFFCEREHPTSQPEFAFEWRDPEMG1SGQET WTREPQGFKNSPTEFDEAEHRDEADFRIQHPDEIEEQYVDDEEEAATSEEDCQQGTRAEE QTEGNEGYRASAKKAQICQKQVKYEGYEEKEGQRWETEARKETVMGQPTPKTPRQEREFLGTAGFCREWIPGFAEMAAPLYPLTKTGTLFNX GPDQQKAYQEIKQALLTAPALGLPDLTK PFEEFVDEKQGYAKGVETQKEGPWRRPVAYESKKEDPVAAGWPPCERMVAAIAVETKDAG KETMGQPEVIEAPHAVEAEVKQPPDRWESNARMTHYQAEEEDTDRVQFGPVVAENPATEE PEPEEGEQHNCEDIEAEAHGTRPDETDQPEPDADHTWYTDGSSEEQEGQRKAGAAVTTET EV1WAKAEPAGTSAQRAEE1AETQAEKMAEGKKENVYTDSRYAFATAH1HGE1YRRRGEETSEGKEIKNKDEIEAEEKAEFEPKRESIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTS TLLZEASSPSGGSKRTADGSEFEPKKKRKV (SEQ ID NO: 1)KEY:NUCLEAR LOCALIZATION SEQUENCE (NLS)CAS9(H840A)33-AMINO ACID LINKERM-MLV reverse transcriptasePE2
[0075] As used herein, “PE2” refers to a PE complex comprising a fusion protein comprisingCas9(H840A) and a variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)] + a desired pegRNA, wherein the PE fusion has the amino acid sequence of:MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAK VDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTD KADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPIN ASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFD LAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEI TKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGG ASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHA ILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPW NFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYV TEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVE DRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNF MQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNK VLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGL SELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLV SDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYD VRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIV WDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDW DPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDF LEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNF LYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDAT LIHQSITGLYETRIDLSOLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQY PMSQEAREGIKPHIQREEDQGIEVPCQSPWNTPEEPVKKPGTNDYRPVQDEREVNKRVED 1HPTVPNPYNEESGEPPSHQWYTVEDEKDAFFCEREHPTSQPEFAFEWRDPEMG1SGQET WTREPQGFKNSPTEFNEAEHRDEADFRIQHPDEIEEQYVDDEEEAATSEEDCQQGTRAEE QTEGNEGYRASAKKAQICQKQVKYEGYEEKEGQRWETEARKETVMGQPTPKTPRQEREF EGKAGFCREF1PGFAEMAAPEYPETKPGTEFNWGPDQQKAYQE1KQAEETAPAEGEPDETK PFEEFVDEKQGYAKGVETQKEGPWRRPVAYESKKEDPVAAGWPPCERMVAAIAVETKDAG KETMGQPEV1EAPHAVEAEVKQPPDRWESNARMTHYQAEEEDTDRVQFGPWAENPATEE PEPEEGEQHNCEDIEAEAHGTRPDETDQPEPDADHTWYTDGSSEEQEGQRKAGAAVTTET EV1WAKAEPAGTSAQRAEE1AETQAEKMAEGKKENVYTDSRYAFATAH1HGE1YRRRGWETS EGKEIKNKDEIEAEEKAEFEPKRESIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDT STLLZEASSPSGGSKRTADGSEFEPKKKRKV (SEQ ID NO: 2)KEY:NUCLEAR LOCALIZATION SEQUENCE (NLS)CAS9(H840A)33-AMINO ACID LINKER
[0076] M-MLV reverse transcriptasePE3
[0077] As used herein, “PE3” refers to a prime editing composition comprising a PE2 prime editor and further comprising a second-strand nicking guide RNA that complexes with PE2 and introduces a nick in the non-edit DNA strand in order to induce preferential replacement of the edit strand.PE3b
[0078] As used herein, “PE3b” refers to a prime editing composition comprising PE2 and further comprising a second- strand nicking guide RNA that complexes with PE2 and introduces a nick in the non-edit DNA strand, wherein the second-strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit. This is achieved by designing the second strand nicking guide RNA with a spacer sequence that comprises complementarity to, and only hybridizes with, the edited strand after installation of the desired nucleotide edit(s), but not the endogenous target DNA sequence. Using this strategy, mismatches between the nicking guide RNA spacer and the unedited target DNA should disfavor nicking by the sgRNA until after the editing event on the PAM strand takes place.PE4
[0079] As used herein, “PE4” refers to a prime editing composition comprising a PE2 and further comprising an MLH1 dominant negative protein variant (z.e., wild- type MLH1 with amino acids 754-756 truncated, which may be referred to herein as “MLH1 A754-756” or “MLHldn”). The MLH1 dominant negative protein variant may be expressed in trans in some embodiments. In some embodiments, a PE4 system comprises a fusion protein comprising a PE2 protein and an MLH1 dominant negative protein joined via an optional linker.PE5 and PE5b
[0080] As used herein, “PE5” refers to a prime editing composition comprising a PE3 prime editor and further comprising an MLH1 dominant negative protein variant (z.e., wild-type MLH1 with amino acids 754-756 truncated, which may be referred to as “MLH1 A754-756”or “MLHldn”). The MLH1 dominant negative variant may be expressed in trans in some embodiments. In some embodiments, a PE5 system comprises a fusion protein comprising a PE2 protein and an MLH1 dominant negative protein joined via an optional linker. “PE5b” refers to a prime editing composition comprising a PE3 and an MLH1 dominant negative protein, wherein the second- strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit.This is achieved by designing the second strand nicking guide RNA with a spacer sequence that comprise complementarity to, and hybridize with, only the edited strand after installation of the desired nucleotide edit(s), but not the endogenous target DNA sequence.PE6
[0081] The term “PE6” refers to a suite of prime editors (PE6a, PE6b, PE6c, PE6d, PE6e, PE6f, and PE6g) comprising improved reverse transcriptase and / or Cas9 variants, as described, for example, in International PCT Application No. PCT / US2024 / 030786, filed May 23, 2024, which is incorporated herein by reference. The improved reverse transcriptase and Cas9 domains of the PE6 variants can also be combined with each other to offer cumulative benefits. For example, a PE6 prime editor comprising an improved reverse transcriptase variant of PE6a and an improved Cas9 variant of PE6e is referred to herein as the prime editor “PE6a-e” (or “PE6e-a”). Any possible combination of PE6 prime editors is contemplated by the present disclosure including, for example, PE6a-e, PE6a-f, PE6a-g, PE6b-e, PE6b-f, PE6b-g, PE6c-e, PE6c-f, PE6c-g, PE6d-e, PE6d-f, and PE6d-g.
[0082] Any of the PE6 prime editors may also comprise the architecture of the PEmax protein as provided herein. In some embodiments, any of the PE6 prime editors provided herein may further comprise additional amino acid mutations, e.g., any of those included in PEmax as provided herein.
[0083] In some embodiments, a PE6 protein comprises a reverse transcriptase of the following amino acid sequence (the RT domain of “PE6a”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTL KKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKR DYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALDYLVDICTKDD FVVQGALTSSYIATLCLFAVEGDVVRRAQRKGLVYTRLVDDITVSSKISNYDFSQMQ SHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKLLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDVA VIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLK PL (SEQ ID NO: 3)
[0084] In some embodiments, a PE6 protein comprises a reverse transcriptase of the following amino acid sequence (the RT domain of “PE6b”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGLTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVL RHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDK EMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNF EINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 4)
[0085] In some embodiments, a PE6 protein comprises a reverse transcriptase of the following amino acid sequence (the RT domain of “PE6c”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFLGYHISEKGLTPCQENIDKVLQWKQPKNQKE LRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPV LRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSD KEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFN FEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 5)
[0086] In some embodiments, a PE6 protein comprises a reverse transcriptase comprising the following amino acid sequence (the RT domain of “PE6d”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKA YQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKL DPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD (SEQ ID NO: 6)
[0087] In some embodiments, a PE6 protein comprises a Cas9 protein of the following amino acid sequence (the Cas9 domain of “PE6e”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQR TFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFA WMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTV YNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFD SVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVK VMGRHKPENIVIEMARENQTTQKGQRNSRERMKRIEEGIKELGSQILKEHPVENTQL QNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDK NRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIA RQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKV REINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGK ATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSM PQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQ HKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGA PAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 7)
[0088] In some embodiments, a PE6 protein comprises a Cas9 protein of the following amino acid sequence (the Cas9 domain of “PE6f”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFRRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQR TFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFA WMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTV YNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFD SVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMVEER LKTYAHLFDNKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANR NFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQ LQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSD KNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFI ARQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIG KATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLS MPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLV VAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLF ELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO:8)
[0089] In some embodiments, a PE6 protein comprises a Cas9 protein of the following amino acid sequence (the Cas9 domain of “PE6g”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the following amino acid sequence:
[0090] MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFRRLEESFLVEEDK KHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFL IEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQ LPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQ YADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQL PEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEY FTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIE CFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMV EERLKTYAHLFDNKVMKQLKRCRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGF ANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLAGSPAIKKGILQTVKVVD ELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVE NTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLT RSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDK AGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQ FYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQ EIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRK VLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYS VLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQ LFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLT NLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 9)PE7
[0091] The term “PE7” refers to the PE6 prime editors plus a second strand nicking guideRNA. For example, “PE7a” refers to the PE6a prime editor as provided herein, plus a second strand nicking guide RNA.PEmax
[0092] As used herein, “PEmax” refers to a prime editing composition comprising 1) a fusion protein comprising a Cas9 protein variant Cas9(R221K N39K H840A) and a variant MMLVRT having the following structure: NH2- [bipartite NLS]-[Cas9(R221K)(N394K)(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)]-[bipartite NLS]-[NLS]-COOH and 2) a desired PEgRNA, wherein the fusion protein (referred to as the PEmax protein) has the following amino acid sequence:MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSOLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPOAWAETGGMGLAVRQA PLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGT NDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPT SQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQ YVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEG QRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTL FNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALV KQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILA EAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGT SAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKN KDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLL IENSSPSGGSKRTADGSEFESPKKKRKVGSGPAAKRV LD (SEQ ID NO: 10)KEY:BIPARTITE SV40 NUCLEAR LOCALIZATION SEQUENCE (NLS).CAS9(R221K N39K H840A)SGGSx2-BIPARTITE SV40NLS-SGGSx2 LINKERM-MLV reverse transcriptase(D200N T306K W313F T330P L603W) Other linker sequence BIPARTITE SV40NLS Other linker sequence c-Myc NLS Polymerase
[0093] As used herein, the term “polymerase” refers to an enzyme that synthesizes a nucleotide strand and that may be used in connection with the prime editor systems and the double-stranded DNA synthesis systems and methods described herein. The polymerase can be a “template-dependent” polymerase (i.e., a polymerase that synthesizes a nucleotide strand based on the order of nucleotide bases of a template strand). The polymerase can also be a “template-independent” polymerase (i.e., a polymerase that synthesizes a nucleotide strand without the requirement of a template strand). A polymerase may also be further categorized as a “DNA polymerase” or an “RNA polymerase.” In various other embodiments, the DNA polymerase can be an “RNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of RNA. The term “polymerase” may also refer to an enzyme that catalyzes the polymerization of nucleotide (i.e., the polymerase activity). Generally, the enzyme will initiate synthesis at the 3 ' end of a primer annealed to a polynucleotide template sequence and will proceed toward the 5 ' end of the template strand.Polynucleotide
[0094] The term “polynucleotide,” as used herein, (also referred to as a “nucleic acid”) refers to a polymer of nucleotides. The polymer may include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxy cytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5-methylcytidine, C5 bromouridine, C5fluorouridine, C5 iodouridine, C5 propynyl uridine, C5 propynyl cytidine, C5 methylcytidine, 7 deazaadenosine, 7 deazaguanosine, 8 oxoadenosine, 8 oxoguanosine, 0(6) methylguanine, 4-acetylcytidine, 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudouridine, 1- methyl adenosine, 1-methyl guanosine, N6-methyl adenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-0-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioates and 5' N phosphoramidite linkages). In some embodiments, a polynucleotide is RNA. In some embodiments, a polynucleotide is mRNA.Prime editing
[0095] As used herein, the term “prime editing” refers to an approach for gene editing using napDNAbps, a polymerase (e.g., a reverse transcriptase), and specialized guide RNAs that include a primer binding site and a DNA synthesis template for encoding desired new genetic information (or deleting genetic information) that is then incorporated into a target DNA sequence. Prime editing is described in Anzalone, A. V. et al., Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019), which is incorporated herein by reference. See also International PCT Application, PCT / US2020 / 023721, filed March 19, 2020, and published as WO 2020 / 191239, and International PCT Application, PCT / US2021 / 031439, filed May 7, 2021, and published as WO 2021 / 226558, each of which is incorporated herein by reference.
[0096] Prime editing represents a platform for genome editing that is a versatile and precise method to directly write new genetic information into a specified DNA site using a nucleic acid programmable DNA binding protein (“napDNAbp”) working in association with a polymerase (z.e., in the form of a fusion protein or otherwise provided in trans with the napDNAbp), wherein the prime editing system is programmed with a prime editing (PE) guide RNA (“pegRNA”) that both specifies the target site and templates the synthesis of the desired edit in the form of a replacement DNA strand by way of an extension (either DNA or RNA) engineered onto a guide RNA (e.g., at the 5' or 3' end, or at an internal portion of a guide RNA). The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same sequence as the endogenous strand (or is homologous to it) immediately downstream of the nick site of the target site to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the endogenous strand downstream of the nick site is replaced by the newly synthesizedreplacement strand containing the desired edit. In some cases, prime editing may be thought of as a “search-and-replace” genome editing technology since the prime editors, as described herein, not only search and locate the desired target site to be edited, but at the same time, encode a replacement strand containing a desired edit that is installed in place of the corresponding target site endogenous DNA strand. The prime editors of the present disclosure relate, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) or “prime editing” can be leveraged or adapted for conducting precision CRISPR / Cas-based genome editing with high efficiency and genetic flexibility. TPRT is naturally used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial Group II introns. Cas protein-reverse transcriptase fusions or related systems are used to target a specific DNA sequence with a guide RNA, generate a single strand nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered DNA synthesis template that is integrated with the guide RNA. However, while the concept begins with prime editors that use reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptases but may include the use of virtually any DNA polymerase. Indeed, while the application throughout may refer to prime editors with “reverse transcriptases,” it is set forth here that reverse transcriptases are only one type of DNA polymerase that may work with prime editing. Thus, wherever the specification mentions a “reverse transcriptase,” the person having ordinary skill in the art should appreciate that any suitable DNA polymerase may be used in place of the reverse transcriptase. Thus, in one aspect, the prime editors may comprise Cas9 (or an equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specialized guide RNA (z.e., pegRNA) containing a spacer sequence that anneals to a complementary sequence (the complementary sequence to an endogenous protospacer sequence) in the target DNA. The pegRNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing a desired nucleotide change, which is used to replace a corresponding endogenous DNA strand at the target site. To transfer information from the pegRNA to the target DNA, the mechanism of prime editing involves nicking the target site in one strand of the DNA to expose a 3'-hydroxyl group. The exposed 3 '-hydroxyl group can then be used to prime the DNA polymerization of the editencoding extension on pegRNA directly into the target site. In various embodiments, the extension — which provides the template for polymerization of the replacement strand containing the edit — can be formed from RNA or DNA. In the case of an RNA extension, thepolymerase of the prime editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of a DNA extension, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. The newly synthesized strand (z.e., the replacement DNA strand containing the desired nucleotide edit) that is formed by the prime editor would be homologous to the genomic target sequence (z.e., have the same sequence as), except for the inclusion of one or more desired nucleotide changes (e.g., a single nucleotide substitution, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be referred to as a single strand DNA flap, which would compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. Resolution of the hybridized intermediate (also referred to as a heteroduplex, comprising the single strand DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand with the exception of mismatches at positions where desired nucleotide edits are installed in the edit strand) can include removal of the resulting displaced flap of endogenous DNA (e.g., with a 5' end DNA flap endonuclease, FEN1), ligation of the synthesized single strand DNA flap to the target DNA, and assimilation of the desired nucleotide changes as a result of cellular DNA repair and / or replication processes.
[0097] In various embodiments, prime editing operates by contacting a target DNA molecule (for which a change in the nucleotide sequence is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (pegRNA). In various embodiments, the prime editing guide RNA (pegRNA) comprises an extension at the 3' or 5' end of the guide RNA, or at an intramolecular location in the guide RNA, and encodes the desired nucleotide change (e.g., single nucleotide substitution, insertion, or deletion). First, the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to a target locus. Next, a nick in one of the strands of DNA of the target locus is introduced (e.g., by a nuclease or chemical agent), thereby creating an available 3' end in one of the strands of the target locus. In certain embodiments, the nick is created in the strand of DNA that corresponds to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the “non-target strand.” The nick, however, could be introduced in either of the strands. That is, the nick could be introduced into the R-loop “target strand” (i.e., the strand hybridized to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand forming the single- stranded portion of the R-loop, and which is complementary to the target strand). Inthe next step, the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA in order to prime reverse transcription (z.e., “target-primed RT”). In certain embodiments, the 3' end DNA strand hybridizes to a specific RT priming sequence on the extended portion of the guide RNA, i.e., the “reverse transcriptase priming sequence” or “primer binding site” on the pegRNA. In the next step, a reverse transcriptase (or other suitable DNA polymerase) is introduced that synthesizes a single strand of DNA from the 3' end of the primed site towards the 5' end of the prime editing guide RNA. The DNA polymerase (e.g., reverse transcriptase) can be fused to the napDNAbp or alternatively can be provided in trans to the napDNAbp. This forms a single- strand DNA flap comprising the desired nucleotide change (e.g., the single base change, insertion, or deletion, or a combination thereof) and that is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In the next step, the napDNAbp and guide RNA are released. The final two steps relate to the resolution of the single strand DNA flap such that the desired nucleotide change becomes incorporated into the target locus. This process can be driven towards the desired product formation by removing the corresponding 5' endogenous DNA flap that forms once the 3' single strand DNA flap invades and hybridizes to the endogenous DNA sequence. Without being bound by theory, the cell’s endogenous DNA repair and replication processes resolve the mismatched DNA to incorporate the nucleotide change(s) to form the desired altered product. The process can also be driven towards product formation with “second strand nicking.” This process may introduce at least one or more of the following genetic changes: transversions, transitions, deletions, and insertions.Prime editor
[0098] The term “prime editor” refers to the polypeptide or polypeptide components involved in prime editing as described herein. In some embodiments, a prime editor comprises a fusion construct comprising a napDNAbp (e.g., Cas9 nickase, and / or any of the Cas9 variants provided herein) and a reverse transcriptase (e.g., any of the reverse transcriptase variants provided herein). In some embodiments, a prime editor is capable of carrying out prime editing on a target nucleotide sequence in the presence of a pegRNA (or “extended guide RNA”). In some embodiments, a prime editor comprises a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase provided in trans, i.e., the napDNAbp and the reverse transcriptase are not fused. The in trans napDNAbp and the reverse transcriptase may be tethered via a non-peptide linkage, e.g., an MS2 RNA-protein binding RNA sequence and a MS2 coat protein fused to either the napDNAbp or the reverse transcriptase, or may beunlinked to each other and simply recruited by the pegRNA. In some embodiments, a prime editor composition, system, or complex provided herein comprises a fusion protein or a fusion protein complexed with a pegRNA, and / or further complexed with a second-strand nicking sgRNA. In some embodiments, the prime editor system may also refer to the complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a pegRNA, and a regular guide RNA capable of directing the second-site nicking step of the non-edited strand as described herein.Reverse transcriptase
[0099] The term “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptase has been used primarily to transcribe mRNA into cDNA, which can then be cloned into a vector for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme has 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand for RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Errors in transcription cannot be corrected by reverse transcriptase because known viral reverse transcriptases lack the 3 '-5' exonuclease activity necessary for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNaseH activity has been presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase that is used extensively in molecular biology is reverse transcriptase originating from Moloney murine leukemia virus (M-MLV or “MMLV”). See, e.g., Gerard, G. R., DNA 5:271-279 (1986) and Kotewicz, M. L., et al., Gene 35:249-258 (1985). M-MLV reverse transcriptase substantially lacking in RNase H activity has also been described. See, e.g., U.S. Pat. No. 5,244,797. The invention contemplates the use of any such reverse transcriptases, or variants or mutants thereof.
[0100] In some embodiments, a reverse transcriptase is part of a prime editor (e.g. , is fused to a napDNAbp). In some embodiments, a reverse transcriptase is not a wild-type reverse transcriptase. In some embodiments, a reverse transcriptase is a variant of MMLV reverse transcriptase. In some embodiments, a reverse transcriptase comprises the amino acidsubstitutions T128N, D200C, V223Y, T306K, W313F, and T330P relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11. In some embodiments, a reverse transcriptase comprises a C-terminal truncation relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11 (e.g., as shown in SEQ ID NO: 12). In some embodiments, a reverse transcriptase comprises the amino acid sequence of the reverse transcriptase of PE6d (SEQ ID NO: 6), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of the reverse transcriptase of PE6d (SEQ ID NO: 6). In some embodiments, a reverse transcriptase is not a wild-type Tfl reverse transcriptase. In some embodiments, a reverse transcriptase is a variant of Tfl reverse transcriptase. In some embodiments, a reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase of SEQ ID NO: 13. In some embodiments, a reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase of SEQ ID NO: 13. In some embodiments, a reverse transcriptase comprises the amino acid sequence of the reverse transcriptase of PE6b (SEQ ID NO: 4), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of the reverse transcriptase of PE6b (SEQ ID NO: 4). In some embodiments, a reverse transcriptase comprises the amino acid sequence of the reverse transcriptase of PE6c (SEQ ID NO: 5), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of the reverse transcriptase of PE6c (SEQ ID NO: 5).Variant
[0101] As used herein, the term “variant” should be taken to mean the exhibition of qualities that have a pattern that deviates from what occurs in nature, e.g., a variant reverse transcriptase is a reverse transcriptase comprising one or more changes in amino acid residues (z.e., “substitutions”) as compared to a wild type reverse transcriptase amino acid sequence. The term “variant” encompasses homologous proteins having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a reference sequence and having the same or substantially the same functional activity or activities as the reference sequence. The term also encompasses mutants,truncations, or domains of a reference sequence that display the same or substantially the same functional activity or activities as the reference sequence.Vector
[0102] The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that is able to enter a host cell, mutate, and replicate within the host cell, and then transfer a replicated form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phage, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the instant disclosure.Wild type
[0103] As used herein, the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene, protein, or characteristic as it occurs in nature (z.e., as distinguished from mutant or variant forms).DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0104] The present disclosure describes the development of systems, compositions, polynucleotides, vectors, and methods for producing DNA from RNA in cells. The systems, compositions, polynucleotides, vectors, and methods described herein are compatible with methods for integration of large sequences (e.g., whole genes) into genomes, such as prime editing-based methods of large gene integration (e.g., PASSIGE, CAST, and quadruple flap prime editing). The systems, compositions, polynucleotides, vectors, and methods described herein provide advantages over direct delivery of donor DNA molecules to cells, which can frequently result in cytotoxicity. Additionally, the systems, compositions, polynucleotides, vectors, and methods are advantageous because they are compatible for use with RNA delivery methods, including lipid nanoparticles (LNPs).
[0105] Thus, the present disclosure provides systems for producing double-stranded DNA in a cell. The present disclosure also provides methods for producing double- stranded DNA in a cell. Further provided herein are polynucleotides, vectors, compositions, cells, and kits for use in the systems and methods described herein.Systems
[0106] In one aspect, the present disclosure provides systems for synthesizing doublestranded DNA from an RNA template. In some embodiments, the present disclosure providessystems comprising a first polynucleotide encoding a reverse transcriptase, and a second polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA. In some embodiments, the present disclosure provides systems comprising a reverse transcriptase and a polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA. In some embodiments, the polynucleotide or second polynucleotide consists of RNA (which may include modifications, for example, modified nucleotides). In some embodiments, the cargo RNA does not comprise a portion of a viral genome. In some embodiments, second polynucleotide does not comprise a portion of a viral genome.
[0107] In certain embodiments, the first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA comprise substantially the same sequence as one another. The first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA may comprise substantially the same sequence as one another if they have, for example, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to one another. The first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA may also comprise substantially the same sequence as one another if their sequences differ by fewer than 20 nucleotides, fewer than 19 nucleotides, fewer than 18 nucleotides, fewer than 17 nucleotides, fewer than 16 nucleotides, fewer than 15 nucleotides, fewer than 14 nucleotides, fewer than 13 nucleotides, fewer than 12 nucleotides, fewer than 11 nucleotides, fewer than 10 nucleotides, fewer than 9 nucleotides, fewer than 8 nucleotides, fewer than 7 nucleotides, fewer than 6 nucleotides, fewer than 5 nucleotides, fewer than 4 nucleotides, fewer than 3 nucleotides, fewer than 2 nucleotides, or fewer than 1 nucleotide compared to one another. In certain embodiments, thefirst region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA comprise the same sequence as one another (z.e., the two sequences have 100% sequence identity to one another).
[0108] In some embodiments, the reverse complement of the first region of the sequence 5 ' of the cargo RNA is complementary to the second region of the sequence 3 ' of the cargo RNA. In some embodiments, the reverse complement of the first region of the sequence 5' of the cargo RNA is partially complementary to the second region of the sequence 3 ' of the cargo RNA. The sequences of each of the two regions may be, for example, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementary to one another. The sequences of each of the two regions may also, for example, comprise one or more, two or more, three or more, four or more, or five or more nucleotide mismatches relative to one another. In some embodiments, the reverse complement of the first region of the sequence 5 ' of the cargo RNA is capable of hybridizing to the second region of the sequence 3 ' of the cargo RNA. The sequences of each of the two regions may be capable of hybridizing to each other if they are complementary to one another, or if they are partially complementary to one another. In certain embodiments, the reverse complement of the first region of the sequence 5 ' of the cargo RNA hybridizes to the second region of the sequence 3 ' of the cargo RNA. In certain embodiments, the reverse complement of the first region of the sequence 5 ' of the cargo RNA binds to the second region of the sequence 3 ' of the cargo RNA.
[0109] In some embodiments, the sequence 5 ' of the cargo RNA (including the first region and the poly adenylation signal) comprises a long terminal repeat (LTR). In some embodiments, the sequence 5 ' of the cargo RNA (including the first region and the polyadenylation signal) comprises one or more portions of an LTR. In some embodiments, the sequence 3' of the cargo RNA comprises an LTR. In some embodiments, the sequence 3' of the cargo RNA comprises one or more portions of an LTR. In certain embodiments, the sequence 5 ' of the cargo RNA and the sequence 3 ' of the cargo RNA both comprise an LTR. In certain embodiments, the sequence 5 ' of the cargo RNA and the sequence 3 ' of the cargo RNA both comprise one or more portions of an LTR.
[0110] Each of the first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA may independently comprise any portion of any LTR known in the art (as described, for example, in Thompson et al., Mol. Cell 2016, 62(5), 766- 776, which is incorporated herein by reference). For example, each of the first region of thesequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA may independently comprise an R region of an LTR or any portion thereof, a U5 region of an LTR or any portion thereof, and / or a U3 region of an LTR of any portion thereof. In some embodiments, the first region of the sequence 5 ' of the cargo RNA comprises an R region of an LTR. In some embodiments, the first region of the sequence 5' of the cargo RNA comprises a portion of an R region of an LTR. In some embodiments, the second region of the sequence 3' of the cargo RNA comprises an R region of an LTR. In some embodiments, the second region of the sequence 3 ' of the cargo RNA comprises a portion of an R region of an LTR. In some embodiments, the polyadenylation signal of the sequence 5 ' of the cargo RNA is part of a U5 region of an LTR. In some embodiments, the promoter-enhancer region of the sequence 3 ' of the cargo RNA is part of a U3 region of an LTR.
[0111] Each of the first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA may independently comprise transcription initiation and / or termination sites. In some embodiments, the first region of the sequence 5 ' of the cargo RNA comprises a transcription initiation site. In some embodiments, the first region of the sequence 5 ' of the cargo RNA comprises a transcription termination site. In some embodiments, the second region of the sequence 3 ' of the cargo RNA comprises a transcription initiation site. In some embodiments, the second region of the sequence 3 ' of the cargo RNA comprises a transcription termination site. In certain embodiments, the first region of the sequence 5 ' of the cargo RNA and the second region of the sequence 3 ' of the cargo RNA each comprise transcription initiation and / or termination sites.
[0112] In some embodiments, the second polynucleotide comprises the structure:5 '-[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region]-3 '. In some embodiments, the sequence 5 ' of the cargo RNA further comprises a promoter-enhancer region. In certain embodiments, the promoter-enhancer region of the sequence 5 ' of the cargo RNA is part of a U3 region of an LTR. In some embodiments, the sequence 3 ' of the cargo RNA further comprises a polyadenylation signal. In certain embodiments, the polyadenylation signal of the sequence 3' of the cargo RNA is part of a U5 region of an LTR. In some embodiments, the second polynucleotide comprises the structure:5 '-[promoter-enhancer region] -[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region] -[poly adenylation signal] -3 '. In some embodiments, the sequence 5 ' of the cargo RNA does not comprise apromoter-enhancer region. In some embodiments, the sequence 3' of the cargo RNA does not comprise a polyadenylation signal.
[0113] The cargo RNA of the second polynucleotide may comprise any sequence, which can be used as a template for synthesis of DNA by a reverse transcriptase. In some embodiments, the cargo RNA comprises one or more genes of interest, or portions thereof. In some embodiments, the one or more genes of interest comprise one or more therapeutic genes. In some embodiments, the one or more genes of interest comprise wild-type versions of the one or more genes. In some embodiments, the one or more genes of interest comprise versions of the one or more genes that are associated with a healthy phenotype and / or not associated with disease (e.g., do not comprise a mutation or multiple mutations associated with a disease phenotype or the risk of developing a disease phenotype). In some embodiments, the cargo RNA comprises one gene, two genes, three genes, four genes, or five genes of interest. In some embodiments, the cargo RNA is greater in length than the maximum length of a sequence that can be inserted into a genome using prime editing. In some embodiments, the cargo RNA is greater than 50, greater than 60, greater than 70, greater than 80, greater than 90, greater than 100, greater than 150, greater than 200, greater than 250, greater than 300, greater than 350, greater than 400, greater than 450, or greater than 500 base pairs in length. In certain embodiments, the cargo RNA is greater than 100 base pairs in length. In some embodiments, the cargo RNA is up to 1000, up to 2000, up to 3000, up to 4000, up to 5000, up to 6000, up to 7000, up to 8000, up to 9000, or up to 10,000 base pairs in length. In certain embodiments, the cargo RNA is up to 3000 base pairs in length. In certain embodiments, the cargo RNA is up to 10,000 base pairs in length.
[0114] Any reverse transcriptase known in the art, as well as variants thereof, may be used in conjunction with the systems, compositions, and methods of the present disclosure. In some embodiments, the reverse transcriptase is not a wild-type reverse transcriptase. In some embodiments, the reverse transcriptase is a variant of a wild-type reverse transcriptase. For example, a reverse transcriptase used in the methods and systems described herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more amino acid mutations or substitutions relative to a wild-type reverse transcriptase, or may comprise at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with a wild-type reverse transcriptase (as long asthe reverse transcriptase has less than 100% sequence identity with a wild-type reverse transcriptase).
[0115] In some embodiments, the reverse transcriptase is not a wild-type MMLV reverse transcriptase. In some embodiments, the reverse transcriptase is a variant of MMLV reverse transcriptase. In some embodiments, the reverse transcriptase comprises less than 100%, less than 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 90%, less than 85%, less 80%, less than 75%, or less than 70% sequence identity with, or comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid mutations or substitutions relative to, a wild-type MMLV reverse transcriptase of the sequence:TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEMAAPLYPLTKTGTLFNWGPDQQK AYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKK LDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS NARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLT DQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIAL TQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGLLTSEGKEIKNKDEILALLKA LFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP (SEQ ID NO: 11), or a truncated version of wild-type MMLV reverse transcriptase of the sequence: TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEMAAPLYPLTKTGTLFNWGPDQQK AYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKK LDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS NARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD (SEQ ID NO: 12).
[0116] In some embodiments, the reverse transcriptase comprises an amino acid substitution at T 128 in a wild- type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution T128N in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises an amino acid substitution at D200 in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution D200C in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises an amino acid substitution at V223 in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution V223Y in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises an amino acid substitution at T306 in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution T306K in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises an amino acid substitution at W313 in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution W313F in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises an amino acid substitution at T330 in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitution T330P in a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises amino acid substitutions at T128, D200, V223, T306, W313, and T330 relative to a wild-type MMLV reverse transcriptase, or a truncated version thereof. In certain embodiments, the reverse transcriptase comprises the amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P relative to a wild-type MMLV reverse transcriptase, or a truncated version thereof. In some embodiments, the reverse transcriptase comprises a C-terminal truncation relative to a wild-type MMLV reverse transcriptase. In certain embodiments, the reverse transcriptase comprises a C-terminal truncation at amino acid D497 of wild-type MMLV reverse transcriptase. In some embodiments, the reverse transcriptase comprises the amino acid sequence:TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKA YQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKL DPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD (SEQ ID NO: 6), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to thereto.
[0117] In some embodiments, the reverse transcriptase is not a wild-type Schizosaccharomyces pombe Tfl reverse transcriptase. In some embodiments, the reverse transcriptase is a variant of Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises less than 100%, less than 99%, less than 98%, less than 97%, less than 96%, less than 95%, less than 90%, less than 85%, less 80%, less than 75%, or less than 70% sequence identity with, or comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid mutations or substitutions relative to, a wild-type Tfl reverse transcriptase of the sequence: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVD YKPLNKYVKPNIYPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGFTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVL RHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDK EMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNF EINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 13).
[0118] In some embodiments, the reverse transcriptase comprises an amino acid substitution at P70 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution P70T in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at G72 in a wild-type Tfl reverse transcriptase. In certain embodiments, thereverse transcriptase comprises the amino acid substitution G72V in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at S87 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution S87G in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at M102 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution M102I in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at K106 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution K106R in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at KI 18 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution K118R in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at 1128 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution I128V in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at L158 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution L158Q in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at SI 88 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution S188K in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at 1260 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution I260L in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at F269 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution F269L in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at R288 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution R288Q in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at S297 in a wild-type Tfl reverse transcriptase. In certain embodiments, thereverse transcriptase comprises the amino acid substitution S297Q in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at A363 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution A363V in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at K413 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution K413E in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises an amino acid substitution at S492 in a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitution S492N in a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises amino acid substitutions at P70, G72, S87, M102, K106, KI 18, 1128, L158, F269, A363, K413, and S492 relative to a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises amino acid substitutions at P70, G72, S87, M102, K106, KI 18, 1128, L158, S188, 1260, F269, R288, S297, A363, K413, and S492 relative to a wild-type Tfl reverse transcriptase. In certain embodiments, the reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase. In some embodiments, the reverse transcriptase comprises the amino acid sequence:ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGLTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVL RHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDK EMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNF EINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 4),or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In some embodiments, the reverse transcriptase comprises the amino acid sequence:ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFLGYHISEKGLTPCQENIDKVLQWKQPKNQKE LRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPV LRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSD KEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFN FEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 5), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto.
[0119] In some embodiments, the reverse transcriptase is fused to a stabilizing protein. In certain embodiments, the stabilizing protein is an MS2 coat protein.
[0120] The systems described herein may comprise additional polynucleotides, vectors, and proteins (including fusion proteins). In some embodiments, the system comprises one or more prime editors, either in protein form or as one or more polynucleotides encoding the prime editor. In some embodiments, the system comprises one or more polynucleotides encoding one or more prime editors. In some embodiments, the reverse transcriptase of the system is part of the prime editor (z.e., is fused to the napDNAbp of the prime editor). In some embodiments, the system further comprises one or more prime editing guide RNAs (pegRNAs). In some embodiments, the system further comprises one or more polynucleotides encoding one or more pegRNAs.
[0121] In some embodiments, the system further comprises a recombinase (e.g., for use in PASSIGE). In some embodiments, the system further comprises a polynucleotide encoding a recombinase. In some embodiments, the recombinase comprises a serine recombine. In some embodiments, the recombinase comprises a Bxbl recombinase. In certain embodiments, the Bxbl recombinase comprises the sequence:MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDR KRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAH FDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGRE PQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRA ELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCG NGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLI GSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQ DTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS, or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In some embodiments, the recombinase is a variant of a wild-type Bxbl recombinase. In certain embodiments, the recombinase comprises the sequence (“evoBxbl” comprising a V74A mutation relative to wild type Bxbl recombinase):MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDR KRRPNLARWLAFEEQPFDAIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAH FDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRV DGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGRE PQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRA ELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCG NGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLI GSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQ DTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS, or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the recombinase comprises the sequence (“eeBxbl” comprising V74X, E229X, and V375X mutations relative to wild type Bxbl recombinase):MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDR KRRPNLARWLAFEEQPFDAIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAH FDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRV DGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGRE PQGRKWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRA ELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCG NGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAIELAEVNAELVDLTSLIG SPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQD TAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS(SEQ ID NO: 16), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto.
[0122] In some embodiments, the system further comprises one or more transposases (e.g., for use in CAST). In some embodiments, the system further comprises one or more polynucleotides encoding one or more transposases. In certain embodiments, the transposase comprises the sequence (“Tn6677-TnsA”):MATSLPTPSAITTSALEYAFHTPARNLTKSRGKNIHRYVSVKMSKRITVESTLECDAC YHFDFEPSIVRFCAQPIRFLYYLNGQSHSYVPDFLVQFDTNEFVLYEVKSAYAKNKPD FDVEWEAKVKAATEEGEEEEEVEESDIRDTVVENNEKRMHRYASKDEENNVHNSEE KIIKYNGAQSARCLGEQLGLKGRTVLPILCDLLSRCLLDTRLDKPLSLESRFELASYG (SEQ ID NO: 17), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677-TnsB”): MAKKGFSSFHRKAVSSQDTLESIELVSSANCLESVTYQDISAFPETIAVEINFRLSILRF LARKCETIVAKSIEPHRVELQQNYSRKIPSAITIYRWWLAFRKSDYNPISLAPNIKDRG NRETKVSTVVDSIMEQAVERVISGRKVNVSSAYKRVRRKVRQYNLTHGTKYTYPKY ESVRKRVKKKTPFELLAAGKGERVAKREFRRMGKKILTSSVLERVEIDHTVVDLFAV HEEYRIPLGRPWLTQLVDCYSKAVIGFYLGFEPPSYVSVSLALKNAIQRKDDLISSYES IENEWLCYGIPDLLVTDNGKEFLSKAFDQACESLLINVHQNKVETPDNKPHVERNYG TINTSLLDDLPGKSFSQYLQREGYDSVGEATLTLNEIREIYLIWLVDIYHKKPNQRGT NCPNVAWKKGCQEWEPEEFSGSKDELDFKFAIVDYKQLTKVGITVYKELSYSNDRL AEYRGKKGNHKVQFKYNPECMAVIWVLDEDMNEYFTVNAIDYEYASRVSLWQHK YNMKYQAELNSAEYDEDKEIDAEIKIEEIADRSIVKTNKIRARRRGARHQENSARAKS ISNANPASIQKHEDEIVSADNDDWDIDYV (SEQ ID NO: 18), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677-TnsC”):MSETREARISRAKRAFVSTPSVRKILSYMDRCRDLSDLESEPTCMMVYGASGVGKTT VIKKYLNQNRRESEAGGDIIPVLHIELPDNAKPVDAARELLVEMGDPLALYETDLAR LTKRLTELIPAVGVKLIIIDEFQHLVEERSNRVLTQVGNWLKMILNKTKCPIVIFGMPY SKVVLQANSQLHGRFSIQVELRPFSYQGGRGVFKTFLEYLDKALPFEKQAGLANESL QKKLYAFSQGNMRSLRNLIYQASIEAIDNQHETITEEDFVFASKLTSGDKPNSWKNPFEEGVEVTEDMLRPPPKDIGWEDYLRHSTPRVSKPGRNKNFFE (SEQ ID NO: 19), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016-TnsA”):MYIRNLRKPSPNKNVFKFASTKVSSVVMCESSLEFDACFHHEYNDLIESFGSQPEGFK YEFMGKSLPYTPDALISYTDKTQKYHEYKPYSKIASPLFRAEFAAKRAASLKLGIDLV LVTDRQIRVNPILNNLKLLHRYSGVYGISGIQKELLSFIHKSGVIKLNDISSQVGIPIGET RSFLFGLMHKGLVKADLGCDDLTNNPTLWATP (SEQ ID NO: 20), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (Tn7016-TnsB):MTDFFNEFDESLVPLKPQTPTQYVKLDDANLIQRDLDTFSDTFKNQALQRYKLISTID KKLSRGWTQRNLDPILDELFKGGDVVRPNWRTVARWRKKYIESNGDIASLADKNHK MGNRTNRIKGDDKFFDKALERFLDAKRPTIATAYQYYKDLIVIENESIVEGKIPIISYN AFNKRIKAIPPYAVAVARHGKFKADQWFAYCAAHVPPTRILERVEIDHTPLDLILLDD ELLIPIGRPYLTLLIDVFSGCVLGFHLSYKSPSYVSAAKAITHAIKPKSLDALNIELQND WPCFGKFENLVVDNGAEFWSKNLEHACQSAGINIQYNPVRKPWLKPFIERFFGVMN EYFLPELPGKTFSNILEKEEYKPEKDAIMRFSTFVEEFHRWIADVYHQDSNSRETRIPI KRWQQGFDAYPPLTMNEEEETRFSMLMRISDSRTLTRNGFKYQELMYDSTALADYR KHYPQTKETVKKLIKVDPDDISKIYVYLEELESYLEVPCTDPTGYTDGLSIYEHKTIKK INREVIRESKDSLGLAKARMAIHERVKQEQEVFIESKTKAKITAVKKQAQIADVSNTG TSTIKVSEESAAPVQKHISNDNSDDWDDDLEAFE (SEQ ID NO: 21), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016-TnsC”):MNALTEIQIEKLRNFSDCIVMHPQIKTIFNDFDELRLNRKFQSDQQCMLLIGDTGVGK SHTINHYKKRVLATQNYSRNTMPVLVSRISRGKGLDATLVQMLADLELFGSSQIKKR GYKTDLTKKLVESLIKAQVELLIINEFQELIEFKSVQERQQIANGLKFISEEAKVPIVLV GMPWAAKIAEEPQWASRLVRKRKLEYFSLKNDSKYFRQYLMGLAKKMPFDVPPKL ESKNTTIALFAACRGENRALKHLLLEALKLALSCNEYLENKHFITAYDKFDFFNDKE KLKSKNPFKQDIKDIEIYEVIKNSSYNPNALDPEDMLTDRVFAIVK (SEQ ID NO: 22), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677-TniQ”):MFLQRPKPYSDESLESFFIRVANKNGYGDVHRFLEATKRFLQDIDHNGYQTFPTDITR INPYSAKNSSSARTASFLKLAQLTFNEPPELLGLAINRTNMKYSPSTSAVVRGAEVFP RSLLRTHSIPCCPLCLRENGYASYLWHFQGYEYCHSHNVPLITTCSCGKEFDYRVSGL KGICCKCKEPITLTSRENGHEAACTVSNWLAGHESKPLPNLPKSYRWGLVHWWMGI KDSEFDHFSFVQFFSNWPRSFHSIIEDEVEFNLEHAVVSTSELRLKDLLGRLFFGSIRLP ERNLQHNIILGELLCYLENRLWQDKGLIANLKMNALEATVMLNCSLDQIASMVEQRI LKPNRKSKPNSPLDVTDYLFHFGDIFCLWLAEFQSDEFNRSFYVSRW (SEQ ID NO: 23), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677-Cas8-Cas5”):MQTLKELIASNPDDLTTELKRAFRPLTPHIAIDGNELDALTILVNLTDKTDDQKDLLD RAKCKQKLRDEKWWASCINCVNYRQSHNPKFPDIRSEGVIRTQALGELPSFLLSSSKI PPYHWSYSHDSKYVNKSAFLTNEFCWDGEISCLGELLKDADHPLWNTLKKLGCSQK TCKAMAKQLADITLTTINVTLAPNYLTQISLPDSDTSYISLSPVASLSMQSHFHQRLQ DENRHSAITRFSRTTNMGVTAMTCGGAFRMLKSGAKFSSPPHHRLNSKRSWLTSEH VQSLKQYQRLNKSLIPENSRIALRRKYKIELQNMVRSWFAMQDHTLDSNILIQHLNH DLSYLGATKRFAYDPAMTKLFTELLKRELSNSINNGEQHTNGSFLVLPNIRVCGATA LSSPVTVGIPSLTAFFGFVHAFERNINRTTSSFRVESFAICVHQLHVEKRGLTAEFVEK GDGTISAPATRDDWQCDVVFSLILNTNFAQHIDQDTLVTSLPKRLARGSAKIAIDDFK HINSFSTLETAIESLPIEAGRWLSLYAQSNNNLSDLLAAMTEDHQLMASCVGYHLLEE PKDKPNSLRGYKHAIAECIIGLINSITFSSETDPNTIFWSLKNYQNYLVVQPRSINDETT DKSSL (SEQ ID NO: 24), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677- Cas7”):MKLPTNLAYERSIDPSDVCFFVVWPDDRKTPLTYNSRTLLGQMEAASLAYDVSGQPI KSATAEALAQGNPHQVDFCHVPYGASHIECSFSVSFSSELRQPYKCNSSKVKQTLVQ LVELYETKIGWTELATRYLMNICNGKWLWKNTRKAYCWNIVLTPWPWNGEKVGFE DIRTNYTSRQDFKNNKNWSAIVEMIKTAFSSTDGLAIFEVRATLHLPTNAMVRPSQV FTEKESGSKSKSKTQNSRVFQSTTIDGERSPILGAFKTGAAIATIDDWYPEATEPLRVG RFGVHREDVTCYRHPSTGKDFFSILQQAEHYIEVLSANKTPAQETINDMHFLMANLIKGGMFQHKGD (SEQ ID NO: 25), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn6677- Cas6”): VKWYYKTITFLPELCNNESLAAKCLRVLHGFNYQYETRNIGVSFPLWCDATVGKKIS FVSKNKIELDLLLKQHYFVQMEQLQYFHISNTVLVPEDCTYVSFRRCQSIDKLTAAGL ARKIRRLEKRALSRGEQFDPSSFAQKEHTAIAHYHSLGESSKQTNRNFRLNIRMLSEQ PREGNSIFSSYGLSNSENSFQPVPLI (SEQ ID NO: 26), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016-TniQ”):MAFLFSPKARAFSDESLESYLLRVVSENFFDSYEGLSLAIREELHELDFEAHGAFPVD LKRLNVYHAKHNSHFRMRALGLLETLLDLPRYELQKLALLKSDIKFNSSVALYNNG VDIPLRFIRHHAEEAVDSIPVCSQCLAEEAYIKQSWHIKWVNACTKHQCALLHNCPE CYAPINYIENESITHCSCGFELSCASTSPVNTLSIEHLNKLLDKGERNDSNPLFNNMTL TERFAALLWYQERYSQTDNFCLNDAVNYFSKWPAVFNTELDELSKNAEMKLIDLFN KTEFKFIFGDAILACPSTQKQSESHFIYRALLDYLVTLVESNPKTKKPNAADLLVSVLE AATLLGTSVEQVYRLYQNGILQTAFRHKMNQRINPYKGAFFLRHVIEYKTSFGNDKA RMYLSAW (SEQ ID NO: 27), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016- Cas8-Cas5”):MHLKELLEITDTTERDRSLRRAFSPYTAMIDITGSEAVALIILLNLTYRKNQVDDLLD KKLAKQALKSEDHINKCIKEIAWFHTHNLKYPDIRVSKQNLAVEPPTLHSYVLSSAN YPKAYGWSHNSAKVNFAKLFVSYFKWQNQVSWLAQVLATNSDNWKSAFTSLGLS VKAFKSLCVTVKNSLPEEAIPDSVDRYSRQIRMPYHDGYLAVTPVISHVVQSKIQQA AIDKRARFSNVEFTRPAAVSMLAASLGGVINVLNYPPYIRSKYHGLSNSRAFKLNNG QTVFNVEALLKPELIKALEGIIFSNNALALKQRRQQKVKNIKELRNTLLEWFSPVFEW RLDAIENGYDLEQLESASERLEYKILSLPDNELPSLTIPLFRLLNEMLGGVSMTQRYAF HPKLMSPLKAALQWLLVNLTDQKHVLIEEDDEHYRYLHLSGIRVFDAQALSNPYCS GIPSLTAVWGMIHSYQRKLNEALGTNVRFTSFSWFIRNYSAVAGKKLPELSLQGAQQ SRLKRPGIIDGKYCDLVFDLIIHIDGYEDDLQAVDSKPDILKAHFPSNFAGGVMHQPE LNSNINWCCLYSNENQLFEKLRRLPLSGCWVMPTEHKIQDLDELLLLLNSDSKLSPSMMGYMEETEPMARVGSEEREHCYAEPAIGVVKYEAATSVREKGIGNYFNSAFWME DAQEKFMLMKKV (SEQ ID NO: 28), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016-Cas7”):MELCNILKYDRSLYPGKAVFFYKTADSDFVPLEADINKIRGPKSGFTEAFTPQFSPKNI SPQDETHNNIETEEECYVPPNVEHIFCRFSERVQANSEVPSGCSDPEVFSEEKEEAETF KECGGYKEEAVRYCRNIEIGTWEWRNQNTGNTQIEIKTSKGSCYEIDNTRKEAWESK WASDDEKVEEEESNEIESAETDPNVFWSADITAKIEASFCQEIYPSQIENDKVKQGEA SKQFVKAKCADGRYAVSFNSVKIGAAEQSIDDWWDEDASKRERVHEFGADKEIGVA RRPPDSEQNFYSIFKNTEWYESAEKNCITNKNEKIDPAIYYEFSVEIKGGMFQKKAEA KKA (SEQ ID NO: 29), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto. In certain embodiments, the transposase comprises the sequence (“Tn7016-Cas6”): MQRYYFTVHFEPKQANEAEETGRCISIMHGFIEKHNIEGMGVTFPAWSDSSIGNEIAF VYTDKEIENTEKDQAYFVDMQDCGFFKVSQVEAVPDSCEEVRFIRNQAVAKIFTGES RRREKREQKRAEARGEDFNPKKIEAPREIDIFHRVAMTSKSSQEDYIEHIQKQDVDCQ AEPYFSNYGEASNEKFKGTVPDESPSIDRN (SEQ ID NO: 30), or a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto.Polynucleotides, Vectors, Compositions, Cells, and Kits
[0123] In other aspects, the present disclosure provides polynucleotides and vectors for use in the systems and methods described herein. In some embodiments, the present disclosure provides polynucleotides comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and further comprises a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA. In some embodiments, the polynucleotide consists of RNA (which may include modifications, for example, modified nucleotides). In some embodiments, the cargo RNA does not comprise a portion of a viral genome. In some embodiments, the polynucleotidedoes not comprise a portion of a viral genome. In some embodiments, the polynucleotides and vectors provided herein comprise DNA. In some embodiments, the polynucleotides and vectors provided herein comprise RNA. In some aspects, any of the polynucleotides described herein may be provided in a vector.
[0124] Other aspects of the present disclosure relate to compositions comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the composition is a pharmaceutical composition. The term “pharmaceutical composition,” as used herein, refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises additional agents (e.g., for specific delivery, increasing half-life, or other therapeutic compounds).
[0125] As used herein, the term “pharmaceutically-acceptable carrier” means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the protein, fusion protein, polynucleotide, or vector from one site (e.g., the delivery site) of the body, to another site (e.g., an organ, tissue, or other part of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.). Some examples of materials that can serve as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, glucose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose, and its derivatives, such as sodium carboxymethyl cellulose, methylcellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricating agents, such as magnesium stearate, sodium lauryl sulfate and talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer’s solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates and / or poly anhydrides; (22) bulking agents, such as polypeptides and amino acids; (23) serum component, such as serum albumin, HDL, and LDL; (22) C2-C12 alcohols,such as ethanol; and (23) other non-toxic compatible substances employed in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, perfuming agents, preservatives, and antioxidants can also be present in the formulation. The terms such as “excipient,” “carrier,” “pharmaceutically acceptable carrier,” or the like are used interchangeably herein.
[0126] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, e.g., for gene editing. Suitable routes of administering the pharmaceutical composition described herein include, without limitation: topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, transtympanic, intraorgan, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseus, periocular, intratumoral, intracerebral, and intracerebroventricular administration.
[0127] In some embodiments, the pharmaceutical composition described herein is administered locally to a diseased site e.g., tumor site). In some embodiments, the pharmaceutical composition described herein is administered to a subject by injection, by means of a catheter, by means of a suppository, or by means of an implant, the implant being of a porous, non-porous, or gelatinous material, including a membrane, such as a sialastic membrane, or a fiber.
[0128] In some embodiments, the pharmaceutical composition is formulated in accordance with routine procedures as a composition adapted for intravenous or subcutaneous administration to a subject, e.g., a human. In some embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. Where necessary, the pharmaceutical composition can also include a solubilizing agent and a local anesthetic such as lignocaine to ease pain at the site of the injection. Generally, the ingredients are supplied either separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or water free concentrate in a hermetically sealed container such as an ampoule or sachette indicating the quantity of active agent. Where the pharmaceutical composition is to be administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. Where the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.
[0129] A pharmaceutical composition for systemic administration may be a liquid, e.g., sterile saline, lactated Ringer’s or Hank’s solution. In addition, the pharmaceuticalcomposition can be in solid forms and re-dissolved or suspended immediately prior to use. Lyophilized forms are also contemplated.
[0130] The pharmaceutical composition can be contained within a lipid particle or vesicle, such as a liposome or microcrystal, which is also suitable for parenteral administration. The particles can be of any suitable structure, such as unilamellar or plurilamellar, so long as compositions are contained therein. Proteins, fusion proteins, polynucleotides, or vectors can be entrapped in “stabilized plasmid-lipid particles” (SPLP) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipid and stabilized by a polyethyleneglycol (PEG) coating (Zhang Y. P. et al., Gene Ther. 1999, 6:1438-47). Positively charged lipids, such as N-[l-(2,3-dioleoyloxi)propyl]-N,N,N- trimethyl-amoniummethylsulfate, or “DOTAP,” are particularly preferred for such particles and vesicles. The preparation of such lipid particles is well known. See, e.g., U.S. Patent Nos. 4,880,635; 4,906,477; 4,911,928; 4,917,951; 4,920,016; and 4,921,757; each of which is incorporated herein by reference. Virus-like particles (VLPs) may also be used for delivery of the polynucleotides, proteins, and / or pharmaceutical compositions described herein. See, for example, PCT Publication No. WO2023 / 102537, published June 8, 2023, PCT Publication No. WO2023 / 102538, published June 8, 2023, and PCT Publication No. W02023 / 102550, published June 8, 2023, each of which is incorporated herein by reference.
[0131] The pharmaceutical compositions described herein may be administered or packaged as a unit dose, for example. The term “unit dose” when used in reference to a pharmaceutical composition of the present disclosure refers to physically discrete units suitable as unitary dosage for the subject, each unit containing a predetermined quantity of active material calculated to produce the desired therapeutic effect in association with the required diluent, i.e., carrier or vehicle.
[0132] Further, the pharmaceutical composition can be provided as a pharmaceutical kit comprising (a) a container containing a protein, fusion protein, complex (e.g., ribonucleoprotein complex), polynucleotide, or vector of the invention in lyophilized form, and (b) a second container containing a pharmaceutically acceptable diluent (e.g., sterile water) for injection. The pharmaceutically acceptable diluent can be used for reconstitution or dilution of the lyophilized protein, fusion protein, complex (e.g., ribonucleoprotein complex), polynucleotide, or vector of the invention. Optionally associated with such container(s) can be a notice in the form prescribed by a governmental agency regulating the manufacture, use,or sale of pharmaceuticals or biological products, which notice reflects approval by the agency of manufacture, use, or sale for human administration.
[0133] In another aspect, an article of manufacture containing materials useful for the treatment of the diseases described above is included. In some embodiments, the article of manufacture comprises a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers may be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition that is effective for treating a disease and may have a sterile access port. For example, the container may be an intravenous solution bag or a vial having a stopper pierce-able by a hypodermic injection needle. The active agent in the composition is a protein, fusion protein, polynucleotide, or vector of the invention. In some embodiments, the label on or associated with the container indicates that the composition is used for treating the disease of choice. The article of manufacture may further comprise a second container comprising a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer’s solution, or dextrose solution. It may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts with instructions for use.
[0134] In another aspect, the present disclosure provides cells comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell, such as a human cell). In some embodiments, the cell is in vitro (e.g., a cultured cell). In some embodiments, the cell is in vivo (e.g., in a subject, such as a human subject). In some embodiments, the cell is ex vivo (e.g., isolated from a subject and may be administered back to the same or a different subject). In some embodiments, a host cell is transiently or non-transiently transfected or electroporated with one or more polynucleotides and / or vectors described herein. In some embodiments, a cell is transfected or electroporated as it naturally occurs in a subject. In some embodiments, a cell that is transfected or electroporated is taken from a subject. In some embodiments, the cell is derived from cells taken from a subject, such as a cell line.
[0135] In another aspect, the present disclosure provides kits comprising any of the polynucleotides and / or vectors provided herein. In some embodiments, the kits further comprise one or more prime editors. In some embodiments, the kits further comprise one or more polynucleotides encoding one or more prime editors. In some embodiments, the kits further comprise one or more pegRNAs. In some embodiments, the kits further comprise oneor more polynucleotides encoding one or more pegRNAs. In some embodiments, the kits further comprise one or more recombinases. In some embodiments, the kits further comprise one or more polynucleotides encoding one or more recombinases. In some embodiments, the kits further comprise one or more transposases. In some embodiments, the kits further comprise one or more polynucleotides encoding one or more transposases.
[0136] In some embodiments, the kits may optionally include instructions and / or promotion for use of the components provided. As used herein, “instructions” can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of the disclosure. Instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communications, etc. The written instructions may be in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which can also reflect approval by the agency of manufacture, use, or sale for animal administration. As used herein, “promoted” includes all methods of doing business including methods of education, hospital and other clinical instruction, scientific inquiry, drug discovery or development, academic research, pharmaceutical industry activity including pharmaceutical sales, and any advertising or other promotional activity including written, oral, and electronic communication of any form, associated with the disclosure.Additionally, the kits may include other components depending on the specific application, as described herein.
[0137] The kits may contain any one or more of the components described herein in one or more containers. The components may be prepared sterilely, packaged in a syringe, and shipped refrigerated. Alternatively, they may be housed in a vial or other container for storage. A second container may have other components prepared sterilely. Alternatively, the kits may include the active agents premixed and shipped in a vial, tube, or other container.
[0138] The kits may have a variety of forms, such as a blister pouch, a shrink-wrapped pouch, a vacuum sealable pouch, a sealable thermoformed tray, or a similar pouch or tray form, with the accessories loosely packed within the pouch, one or more tubes, containers, a box, or a bag. The kits may be sterilized after the accessories are added, thereby allowing the individual accessories in the container to be otherwise unwrapped. The kits can be sterilized using any appropriate sterilization techniques, such as radiation sterilization, heat sterilization, or other sterilization methods known in the art. The kits may also include othercomponents, depending on the specific application, for example, containers, cell media, salts, buffers, reagents, syringes, needles, a fabric, such as gauze, for applying or removing a disinfecting agent, disposable gloves, a support for the agents prior to administration, etc. Some aspects of this disclosure provide kits comprising a nucleic acid construct comprising a nucleotide sequence encoding the prime editor systems described herein, or various components thereof (e.g., polynucleotides and / or vectors provided herein). In some embodiments, the nucleotide sequence(s) comprises a heterologous promoter (or more than a single promoter) that drives expression of one or more components.Methods of Use
[0139] In another aspect, the present disclosure provides methods for producing doublestranded DNA in a cell. In some embodiments, the present disclosure provides methods comprising expressing in the cell a first polynucleotide encoding a reverse transcriptase, and contacting a second polynucleotide with the reverse transcriptase, wherein the second polynucleotide comprises a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT), wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal, and wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA. In some embodiments, the cargo RNA does not comprise a portion of a viral genome. In some embodiments, the second polynucleotide does not comprise a portion of a viral genome.
[0140] In some embodiments, the reverse transcriptase initiates polymerization of a first strand of DNA at the PBS. In some embodiments, the reverse transcriptase uses a tRNA in the cell as a primer to initiate reverse transcription of the first strand of DNA. In some embodiments, the second polynucleotide is reverse transcribed to its 5 ' end, and the first strand of DNA is thereby complementary to the sequence 5 ' of the cargo RNA in the second polynucleotide. In some embodiments, the sequence 5' of the cargo RNA in the second polynucleotide is degraded by an RNase while hybridized to the first strand of DNA. In certain embodiments, the reverse transcriptase comprises the RNase. In certain embodiments, the reverse transcriptase comprises RNase activity.
[0141] In some embodiments, the 3 ' end of the first strand of DNA is complementary to the 3 ' end of the second polynucleotide. In some embodiments, the 3 ' end of the first strand ofDNA is partially complementary to the 3 ' end of the second polynucleotide. In some embodiments, the 3 ' end of the first strand of DNA is capable of hybridizing to the 3 ' end of the second polynucleotide. In some embodiments, the 3 ' end of the first strand of DNA binds to the 3 ' end of the second polynucleotide following degradation of the 5 ' end of the second polynucleotide. In some embodiments, the reverse transcriptase extends the first strand of DNA by using it as a primer for reverse transcription of the cargo RNA. In some embodiments, the remaining portions of the second polynucleotide hybridized to the first strand of DNA other than the PPT are degraded by an RNase. In some embodiments, the reverse transcriptase comprises the RNase. In certain embodiments, the reverse transcriptase comprises RNase activity.
[0142] In some embodiments, the PPT acts as a primer for synthesis of a second strand of DNA comprising a portion that is complementary to the 5 ' end of the first strand of DNA and a portion that is complementary to the 3 ' end of the first strand of DNA, thereby promoting circularization of the first strand of DNA. In some embodiments, the second strand of DNA is extended to complete production of the double-stranded DNA. In some embodiments, the 5 ' end and the 3 ' end of each of the first strand of DNA and the second strand of DNA are ligated together (e.g., by a ligase or other DNA repair protein in a cell) to produce a circular double-stranded DNA.
[0143] In some embodiments, the reverse transcriptase is part of a prime editor. In some embodiments, the methods provided herein further comprises using a prime editing-based strategy to insert the double- stranded DNA, or a portion thereof, into a genomic sequence (e.g., in a chromosome). In certain embodiments, the prime editing-based strategy is PASSIGE, CAST, or quadruple flap prime editing.
[0144] The present disclosure also provides for use of the systems, compositions, polynucleotides, and vectors described herein as a medicament.EXAMPLESExample 1. Converting RNA to DNA to Enable Programmable Gene Integration in Mammalian Cells without Exogenous DNA
[0145] Since RNA plays an important role in transferring genetic information, it was hypothesized that instead of directly delivering a naked DNA donor for programmable gene integration, the immune sensing pathway could be evaded by producing the DNA donor inside mammalian cells, particularly in the nucleus, through RNA reverse transcription(FIG. 2). This can not only reduce the cytotoxicity from directly delivering DNA donor but can also synergize with current mRNA delivery technologies (e.g., lipid nanoparticles (LNPs, virus-like particles (VLPs), etc.
[0146] Three major strategies to prime reverse transcription were considered: the 2 ' hydroxyl priming mechanism of the bacterial retron (FIG. 3), target priming for non-LTR retrotransposons and group-II intron (FIG. 4), and the tRNA priming mechanism of retroviruses and LTR retrotransposons (FIG. 5). The third strategy utilizes LTR retrotransposons and retroviruses (FIG. 5). The general structure of the LTR retrotransposons and retrovirus is characterized by long terminal repeats flanking at both ends of a template gene. There are three components in the LTR: 1) at the far left is the U3, which is the enhancer and promoter region for transcription; 2) in the middle is the R region, which defines the start and termination site for transcription; and 3) at the far right is the U5 region, which functions as a poly adenylation signal. LTR based reverse transcription using tRNA priming only requires minimal components to generate dsDNA and creates circular products ready for targeted integration (FIGs. 6A-6B). This third strategy was ultimately used to further develop the methods and systems described herein.
[0147] A previously reported retrotransposition assay (see D. Ribet, et al., Genome Research 2004 and M. Dewannieux, et al., Nature Genetics 2004, each of which is incorporated herein by reference) was repurposed as a plasmid-based assay for detection of RNA to DNA conversion (FIG. 7). Incorporating the neo indicated gene within the RNA template, the tet intron will be spliced out after RNA transcription. The mature RNA will then be reverse transcribed to an intronless copy that can be detected by ddPCR targeting at the exon-exon junction. A co-transfection experiment was designed with the MusD LTR RNA templates and the evolved reverse transcriptase (PE6d variant) to see whether only RNA template and reverse transcriptase are needed to produce dsDNA (FIGs. 8A-8B). The MusD RNA expression was under the promoter of CMV right before the R region of the LTR. R defines the transcription start site, while U3 is normally a transcription promoter that can be replaced by the CMV in order to have a unified expression. Reverse transcriptase was expressed under the EFla promoter. By co-transfecting these two plasmids, DNA production can be detected using the ddPCR method. Using the reverse transcriptases extracted from PEI, PE2, PE6c, and PE6d, there is a correlation between DNA production and reverse transcriptase processivity. RT from PE6d is the most processive, and it was found to produce around 2-fold more DNA compared to beta actin concentration, which is around 4 copies per cells since each cell has two copies of beta actin.
[0148] To confirm that dsDNA was produced, an extra step was added to the protocol: the gDNA lysate was incubated with nuclease Pl before the ddPCR experiment was performed (FIGs. 9A-9B). Nuclease Pl digestion indicates the presence of dsDNA production. To validate the system using the evoRT from PE6d, a divergent PCR experiment was used (FIG. 10). After gDNA extraction, a specialized exonuclease was applied to digest the linear DNA, and then divergent PCR was performed using two primers going in opposite directions. Divergent PCR only produces a product if the circular DNA substrate is present. The assay showed that circular DNA product was produced from the system.
[0149] Overall, expressing RNA with LTR architecture and evolved reverse transcriptase (extracted from PE6d) allows production of DNA inside mammalian cells. Minimal components are required, but the evolved processive reverse transcriptase is key. The reverse transcription product is double- stranded, which is optimal compared to the other strategies described above, which only synthesize single- stranded DNA. The double- stranded DNA product is also presumably in circular form.
[0150] To test conversion of RNA to double- stranded DNA donor for use in combination with evoBxBl recombination, RNA template using the MusDl LTR architecture and carrying a BFP cargo in the reverse direction was produced by in vitro transcription (IVT) and nucleofected into the clonal cell line (FIG. 11). BFP signal corresponded to doublestranded DNA production from RNA reverse transcription (FIG. 12). Another validation to show that the system can produce dsDNA inside mammalian cells from an RNA template coupled with evolved reverse transcriptase (PE6d variant) was performed. Detection of integration by ddPCR showed 0.1% integration efficiency using the endogenous generated double-stranded DNA donor (FIG. 13).
[0151] The developed strategy was validated by testing integration of the double- stranded DNA product. Coupling with evolved processive reverse transcriptase, RNA delivered by IVT to an in vitro cell line can be used to generate double-stranded DNA donor. The doublestranded DNA produced can function as the donor molecule for recombinase integration.
[0152] Next, to optimize the reverse transcriptase, array screening was set up for different reverse transcriptase systems to determine which is optimal (FIG. 14). The reverse transcriptases tested are all compatible for use with prime editing. Arrayed style screening was set up by using the wild type reverse transcriptases with their cognate LTR RNAtemplates, and then double-stranded DNA production was compared again by ddPCR targeting the exon-exon junction (FIG. 15). It was observed that wild type reverse transcriptase is not sufficient to generate double-stranded DNA with cognate LTR RNA templates (FIG. 16). Evolved reverse transcriptases, specifically the one extracted from PE6d, can function on different LTR templates, and the highest efficiency was observed when using cognate MMLV LTR architecture (FIG. 17).
[0153] Reverse transcriptase was also fused with MS2 coating protein (MCP), which was found to facilitate reverse transcription (FIG. 18). Co-transfection of RT(PE6d) fused with MCP and the cognate MMLV LTR RNA template increased DNA production to ~20 copies per cell on average. MCP likely plays a role in stabilizing the reverse transcriptase and / or binding to other small stem loops from the MMLV-LTR RNA template.Example 2. Development of an RNA-based Donor System for Generating Double- Stranded DNA for Programmable Gene Integration in Cells
[0154] This Example is in reference to the data of FIGs. 19A-45D.
[0155] To develop the RNA-based donor system based on the LTR retroelements, a plasmidbased assay was first established for efficient engineering. This assay was inspired by a previously optimized retrotransposition assay and coupled it with droplet digital PCR (ddPCR). The plasmid-based assay incorporates an engineered neomycin indicator cargo with a self- splicing intron in the RNA templates. After transcription, the intron is spliced from the RNA intermediate, which is subsequently reverse transcribed into an intronless dsDNA that can then be detected by ddPCR probes targeting the exon-exon junction. This plasmid-based approach avoids in vitro transcription of RNA templates, substantially simplifying the assay process.
[0156] To identify a reverse transcription system for converting RNA into dsDNA, cells were co-transfected with various wild-type reverse transcriptases and their cognate RNA templates which encoded the indicator cargo. However, reverse-transcribed DNA products were not detected. Although in vitro reverse transcription assays with LTR-retroelements have shown that full-length DNA can be synthesized using tRNA primers, RNA templates, and reverse transcriptases (RTs), efficient endogenous reverse transcription within cells is known to be capsid-dependent, involving Gag proteins. The capsid interior surface can “cage” the low- processivity RTs, preventing dissociation and facilitating key steps like priming and template switching. Given the general low processivity of wild-type RTs, a highly processive,engineered Moloney Murine Leukemia Virus (M-MLV) RT that had been previously evolved using phage-assisted evolution was evaluated next. Upon using this RT variant, dsDNA product generation of approximately 15 copies per cell was observed.
[0157] While this initial system demonstrated that the engineered M-MLV reverse transcriptases could generate dsDNA from an RNA template, the efficiency of the process could be further improved. To address this, rational engineering was used to enhance RT activity. Promising RT mutations were identified, including in the catalytic domain (V223F) and the RNaseH domain (D524A). Fusing the RT with MS2 coat protein (MCP) also enhanced DNA production two-fold, presumably due to increased stability of the RT structure.
[0158] To investigate and engineer the RNA-based donor system to minimize cellular toxicity, a workflow was established to translate the plasmid-based engineering efforts into an all-RNA delivery assay. In brief, both the reverse transcriptase (RT) and the RNA template are produced in RNA form via in vitro transcription (IVT). These two components are then transfected into HEK293T cells using the commercially available MessengerMax liposome reagent. Notably, dsDNA production in human cells was observed by providing only the RT and the RNA template encoding the desired cargo in RNA format, indicating that the plasmid-based assay is an effective strategy for engineering an RNA-based donor system. It was further validated that the improvements from the plasmid-based assay synergistically enabled efficient intracellular donor DNA production, resulting in more than 600 copies of donor molecules produced per triploid genome.
[0159] To validate the functionality of the RNA-based donor system, it was coupled with the site-specific eeBxbl recombinase used in PASSIGE. Since PASSIGE requires the installation of recombinase attachment sites prior to the recombination step, the RNA-based donor system was assessed in cells containing pre-installed recombinase attachment sites (attP). When co-delivered with eeBxbl recombinase into human cells with pre-installed attachment sites, the RNA-based donor system achieved a 17% gene integration efficiency.
[0160] The presence of cytoplasmic DNA in mammalian cells serves as a danger signal that triggers host immune responses, such as type I interferon production. This Example demonstrates successful conversion of RNA into double-stranded DNA (dsDNA) in mammalian cells (HEK293T) within the cytoplasm by providing only an RNA template and the reverse transcriptase enzyme. Using the proposed delivery method, intracellular DNA production alone was determined to be sufficient to reduce the major cytotoxicity associatedwith conventional synthetic DNA delivery by measuring the activation of cGAS pathway through Elisa.EQUIVALENTS AND SCOPE
[0161] In the articles such as “a,” “an,” and “the” may mean one or more than one unless indicated to the contrary or otherwise evident from the context. Embodiments or descriptions that include “or” between one or more members of a group are considered satisfied if one, more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process unless indicated to the contrary or otherwise evident from the context. The invention includes embodiments in which exactly one member of the group is present in, employed in, or otherwise relevant to a given product or process. The invention includes embodiments in which more than one, or all of the group members are present in, employed in, or otherwise relevant to a given product or process.
[0162] Furthermore, the disclosure encompasses all variations, combinations, and permutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the listed claims is introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claims that is dependent on the same base claim. Where elements are presented as lists, e.g., in Markush group format, each subgroup of the elements is also disclosed, and any element(s) can be removed from the group. It should it be understood that, in general, where the invention, or aspects of the invention, is / are referred to as comprising particular elements and / or features, certain embodiments of the disclosure or aspects of the disclosure consist, or consist essentially of, such elements and / or features. For purposes of simplicity, those embodiments have not been specifically set forth in haec verba herein. It is also noted that the terms “comprising” and “containing” are intended to be open and permits the inclusion of additional elements or steps. Where ranges are given, endpoints are included. Furthermore, unless otherwise indicated or otherwise evident from the context and understanding of one of ordinary skill in the art, values that are expressed as ranges can assume any specific value or sub-range within the stated ranges in different embodiments of the invention, to the tenth of the unit of the lower limit of the range, unless the context clearly dictates otherwise.
[0163] This application refers to various issued patents, published patent applications, journal articles, and other publications, all of which are incorporated herein by reference. If there is a conflict between any of the incorporated references and the instant specification, thespecification shall control. In addition, any particular embodiment of the present invention that falls within the prior art may be explicitly excluded from any one or more of the embodiments. Because such embodiments are deemed to be known to one of ordinary skill in the art, they may be excluded even if the exclusion is not set forth explicitly herein. Any particular embodiment of the invention can be excluded from any embodiment, for any reason, whether or not related to the existence of prior art.
[0164] Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. The scope of the present embodiments described herein is not intended to be limited to the above Description, but rather is as set forth in the appended embodiments. Those of ordinary skill in the art will appreciate that various changes and modifications to this description may be made without departing from the spirit or scope of the present invention, as defined in the following embodiments.
Claims
CLAIMSWhat is claimed is:
1. A system comprising a first polynucleotide encoding a reverse transcriptase and a second polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT); wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal; wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA; and wherein the cargo RNA does not comprise a portion of a viral genome.
2. The system of claim 1, wherein the sequence 5 ' of the cargo RNA and / or the sequence 3 ' of the cargo RNA each comprise a long terminal repeat (LTR), or one or more portions thereof.
3. The system of claim 1 or 2, wherein the first region of the sequence 5 ' of the cargo RNA and / or the second region of the sequence 3 ' of the cargo RNA each comprise transcription initiation and / or termination sites.
4. The system of any one of claims 1-3, wherein the first region of the sequence 5 ' of the cargo RNA and / or the second region of the sequence 3 ' of the cargo RNA each comprise an R region of an LTR.
5. The system of any one of claims 1-4, wherein the polyadenylation signal of the sequence 5 ' of the cargo RNA is part of a U5 region of an LTR.
6. The system of any one of claims 1-5, wherein the promoter-enhancer region of the sequence 3 ' of the cargo RNA is part of a U3 region of an LTR.
7. The system of any one of claims 1-6, wherein the reverse complement of the first region of the sequence 5 ' of the cargo RNA is complementary or partially complementary to the second region of the sequence 3 ' of the cargo RNA.
8. The system of any one of claims 1-7, wherein the reverse complement of the first region of the sequence 5 ' of the cargo RNA is capable of hybridizing to the second region of the sequence 3 ' of the cargo RNA.
9. The system of any one of claims 1-8, wherein the second polynucleotide comprises the structure:5 '-[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region]-3 '.
10. The system of any one of claims 1-9, wherein the sequence 5 ' of the cargo RNA further comprises a promoter-enhancer region.
11. The system of claim 10, wherein the promoter-enhancer region of the sequence 5 ' of the cargo RNA is part of a U3 region of an LTR.
12. The system of any one of claims 1-11, wherein the sequence 3 ' of the cargo RNA further comprises a polyadenylation signal.
13. The system of claim 12, wherein the polyadenylation signal of the sequence 3 ' of the cargo RNA is part of a U5 region of an LTR.
14. The system of any one of claims 1-13, wherein the second polynucleotide comprises the structure:5 '-[promoter-enhancer region] -[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region] -[poly adenylation signal] -3 '.
15. The system of any one of claims 1-9, wherein the sequence 5 ' of the cargo RNA does not comprise a promoter-enhancer region, and / or the sequence 3 ' of the cargo RNA does not comprise a polyadenylation signal.
16. A system comprising a reverse transcriptase and a polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA, wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT); wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal; wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA; and wherein the cargo RNA does not comprise a portion of a viral genome.
17. The system of any one of claims 1-16, wherein the reverse transcriptase is not a wildtype reverse transcriptase.
18. The system of any one of claims 1-17, wherein the reverse transcriptase is a variant of MMLV reverse transcriptase.
19. The system of any one of claims 1-18, wherein the reverse transcriptase comprises the amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11.
20. The system of any one of claims 1-19, wherein the reverse transcriptase comprises a C-terminal truncation relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11.
21. The system of any one of claims 1-20, wherein the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 6, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 6.
22. The system of any one of claims 1-17, wherein the reverse transcriptase is not a wildtype Tfl reverse transcriptase.
23. The system of any one of claims 1-17 or 22, wherein the reverse transcriptase is a variant of Tfl reverse transcriptase.
24. The system of any one of claims 1-17, 22, or 23, wherein the reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N; or P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase of SEQ ID NO: 13.
25. The system of any one of claims 1-17 or 22-24, wherein the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 5 or 6, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 5 or 6.
26. The system of any one of claims 1-25, wherein the reverse transcriptase is fused to a stabilizing protein.
27. The system of claim 26, wherein the stabilizing protein is an MS2 coat protein.
28. The system of any one of claims 1-27, wherein the cargo RNA comprises one or more genes of interest.
29. The system of claim 28, wherein the one or more genes of interest comprise one or more therapeutic genes.
30. The system of any one of claims 1-29, wherein the cargo RNA is greater than 100 base pairs in length.
31. The system of any one of claims 1-30, wherein the cargo RNA is up to 3000 base pairs in length.
32. The system of any one of claims 1-30, wherein the cargo RNA is up to 10,000 base pairs in length.
33. The system of any one of claims 1-32, wherein the system comprises a prime editor.
34. The system of any one of claims 1-33, wherein the system comprises one or more polynucleotides encoding a prime editor.
35. The system of claim 33 or 34, wherein the reverse transcriptase in the system is part of the prime editor.
36. The system of any one of claims 1-35 further comprising one or more prime editing guide RNAs (pegRNAs).
37. system of any one of claims 1-36 further comprising one or more polynucleotides encoding one or more pegRNAs.
38. The system of any one of claims 1-37 further comprising a recombinase.
39. The system of any one of claims 1-38 further comprising a polynucleotide encoding a recombinase.
40. The system of any one of claims 1-39 further comprising one or more transposases.
41. The system of any one of claims 1-40 further comprising one or more polynucleotides encoding one or more transposases.
42. A polynucleotide comprising a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA; wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and further comprises a polypurine tract (PPT);wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal; wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA; and wherein the cargo RNA does not comprise a portion of a viral genome.
43. A vector comprising the polynucleotide of claim 42.
44. A composition comprising the polynucleotide of claim 42 or the vector of claim 43.
45. A cell comprising the system of any one of claims 1-41, the polynucleotide of claim 42, or the vector of claim 43.
46. A kit comprising the system of any one of claims 1-41, the polynucleotide of claim 42, or the vector of claim 43.
47. A method for producing double- stranded DNA in a cell comprising: expressing in the cell a first polynucleotide encoding a reverse transcriptase, and contacting a second polynucleotide with the reverse transcriptase, wherein the second polynucleotide comprises a cargo RNA, a sequence 5 ' of the cargo RNA, and a sequence 3 ' of the cargo RNA; wherein the cargo RNA comprises a primer binding site (PBS) at its 5 ' end and a polypurine tract (PPT); wherein the sequence 5 ' of the cargo RNA comprises a first region and a polyadenylation signal; wherein the sequence 3 ' of the cargo RNA comprises a promoter-enhancer region and a second region comprising the same or substantially the same sequence as the first region of the sequence 5 ' of the cargo RNA; and wherein the cargo RNA does not comprise a portion of a viral genome.
48. The method of claim 47, whereby the reverse transcriptase initiates polymerization of a first strand of DNA at the PBS.
49. The method of claim 48, wherein the reverse transcriptase uses a tRNA in the cell as a primer to initiate reverse transcription of the first strand of DNA.
50. The method of claim 49, wherein the second polynucleotide is reverse transcribed to its 5 ' end, and the first strand of DNA is thereby complementary to the sequence 5 ' of the cargo RNA in the second polynucleotide.
51. The method of claim 50, whereby the sequence 5 ' of the cargo RNA in the second polynucleotide is degraded by an RNase while hybridized to the first strand of DNA.
52. The method of claim 51, wherein the 3 ' end of the first strand of DNA is complementary or partially complementary, and / or capable of hybridizing to, the 3 ' end of the second polynucleotide, and wherein the 3 ' end of the first strand of DNA binds to the 3 ' end of the second polynucleotide following degradation of the 5 ' end of the second polynucleotide.
53. The method of claim 52, wherein the reverse transcriptase extends the first strand of DNA by using it as a primer for reverse transcription of the cargo RNA.
54. The method of claim 53, wherein the remaining portions of the second polynucleotide hybridized to the first strand of DNA other than the PPT are degraded by an RNase.
55. The method of any one of claims 51-54, wherein the reverse transcriptase has RNase activity.
56. The method of claim 55, wherein the PPT acts as a primer for synthesis of a second strand of DNA comprising a portion that is complementary to the 5 ' end of the first strand of DNA and a portion that is complementary to the 3 ' end of the first strand of DNA, thereby promoting circularization of the first strand of DNA.
57. The method of claim 56, wherein the second strand of DNA is extended to complete production of the double- stranded DNA.
58. The method of claim 57, wherein the 5 ' end and the 3 ' end of each of the first strand of DNA and the second strand of DNA are ligated together to produce a circular doublestranded DNA.
59. The method of any one of claims 47-58, wherein the reverse transcriptase is part of a prime editor.
60. The method of any one of claims 47-59 further comprising using a prime editingbased strategy to insert the double- stranded DNA, or a portion thereof, into a genomic sequence.
61. The method of claim 60, wherein the prime editing-based strategy is PASSIGE, CAST, or quadruple flap prime editing.
62. The method of any one of claims 47-61, wherein the sequence 5 ' of the cargo RNA and / or the sequence 3 ' of the cargo RNA each comprise a long terminal repeat (LTR), or one or more portions thereof.
63. The method of any one of claims 47-62, wherein the first region of the sequence 5 ' of the cargo RNA and / or the second region of the sequence 3 ' of the cargo RNA each comprise transcription initiation and / or termination sites.
64. The method of any one of claims 47-63, wherein the first region of the sequence 5 ' of the cargo RNA and / or the second region of the sequence 3 ' of the cargo RNA each comprise an R region of an LTR.
65. The method of any one of claims 47-64, wherein the polyadenylation signal of the sequence 5 ' of the cargo RNA is part of a U5 region of an LTR.
66. The method of any one of claims 47-65, wherein the promoter-enhancer region of the sequence 3 ' of the cargo RNA is part of a U3 region of an LTR.
67. The method of any one of claims 47-66, wherein the reverse complement of the first region of the sequence 5 ' of the cargo RNA is complementary or partially complementary to the second region of the sequence 3 ' of the cargo RNA.
68. The method of any one of claims 47-67, wherein the reverse complement of the first region of the sequence 5 ' of the cargo RNA is capable of hybridizing to the second region of the sequence 3 ' of the cargo RNA.
69. The method of any one of claims 47-68, wherein the second polynucleotide comprises the structure:5 '-[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region]-3 '.
70. The method of any one of claims 47-69, wherein the sequence 5 ' of the cargo RNA further comprises a promoter-enhancer region.
71. The method of claim 70, wherein the promoter-enhancer region of the sequence 5 ' of the cargo RNA is part of a U3 region of an LTR.
72. The method of any one of claims 47-71, wherein the sequence 3' of the cargo RNA further comprises a polyadenylation signal.
73. The method of claim 72, wherein the poly adenylation signal of the sequence 3 ' of the cargo RNA is part of a U5 region of an LTR.
74. The method of any one of claims 47-73, wherein the second polynucleotide comprises the structure:5 '-[promoter-enhancer region] -[first region] -[poly adenylation signal] -[cargo RNA comprising 5 ' PBS and PPT]- [promoter-enhancer region] -[second region] -[poly adenylation signal] -3 '.
75. The method of any one of claims 47-69, wherein the sequence 5 ' of the cargo RNA does not comprise a promoter-enhancer region and / or the sequence 3 ' of the cargo RNA does not comprise a polyadenylation signal.
76. The method of any one of claims 47-75, wherein the reverse transcriptase is not a wild-type reverse transcriptase.
77. The method of any one of claims 47-76, wherein the reverse transcriptase is a variant of MMLV reverse transcriptase.
78. The method of any one of claims 47-77, wherein the reverse transcriptase comprises the amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11.
79. The method of any one of claims 47-78, wherein the reverse transcriptase comprises a C-terminal truncation relative to a wild-type MMLV reverse transcriptase of SEQ ID NO: 11.
80. The method of any one of claims 47-79, wherein the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 6, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 6.
81. The method of any one of claims 47-75, wherein the reverse transcriptase is not a wild-type Tfl reverse transcriptase.
82. The method of any one of claims 47-75 or 81, wherein the reverse transcriptase is a variant of Tfl reverse transcriptase.
83. The method of any one of claims 47-75, 81, or 82, wherein the reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N; or P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a wild-type Tfl reverse transcriptase of SEQ ID NO: 13.
84. The method of any one of claims 47-75 or 81-83, wherein the reverse transcriptase comprises the amino acid sequence of SEQ ID NO: 5 or 6, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 5 or 6.
85. The method of any one of claims 47-84, wherein the reverse transcriptase is fused to a stabilizing protein.
86. The method of claim 85, wherein the stabilizing protein is an MS2 coat protein.
87. The method of any one of claims 47-86, wherein the cargo RNA comprises one or more genes of interest.
88. The method of claim 87, wherein the one or more genes of interest comprise one or more therapeutic genes.
89. The method of any one of claims 47-88, wherein the cargo RNA is greater than 100 base pairs in length.
90. The method of any one of claims 47-89, wherein the cargo RNA is up to 3000 base pairs in length.
91. The method of any one of claims 47-89, wherein the cargo RNA is up to 10,000 base pairs in length.
92. Use of the system of any one of claims 1-41, the polynucleotide of claim 42, the vector of claim 43, or the composition of claim 44 as a medicament.
93. The system of any one of claims 18-20 or the method of any one of claims 77-79, wherein the MMLV reverse transcriptase comprises a V223F mutation relative to SEQ ID NO: 11.
94. The system of any one of claims 18-20 or the method of any one of claims 77-79, wherein the MMLV reverse transcriptase comprises a D524A mutation relative to SEQ ID NO: 11.
95. The system of any one of claims 18-20 or the method of any one of claims 77-79, wherein the MMLV reverse transcriptase comprises a V223F mutation and a D524A mutation relative to SEQ ID NO: 11.
96. An MMLV reverse transcriptase variant comprising a D524A mutation relative to SEQ ID NO: 11.
97. An MMLV reverse transcriptase variant comprising a D524A mutation and a V223F mutation relative to SEQ ID NO: 11.
Citation Information
Patent Citations
Protein, preparation method and application thereof, biological material and application thereof
CN118272339A