Precision editing system based on nickase-anti- template and methods of use
Patent Information
- Application Number
- CN202480076961.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-06
- Filing Date
- 2024-10-11
- Publication Date
- 2026-08-07
Smart Images

Figure CN122535696A_ABST
Abstract
Description
[0001] sequence list This application contains a sequence listing submitted electronically in eXentsible Markup Language (XML) format, named 60676WO_CRF_sequencelisting, created on September 30, 2024, and measuring 36,013,000 bytes. The contents of this sequence listing are incorporated herein by reference in their entirety. This application also incorporates, by reference, the sequence listings filed in U.S. Application No. 18 / 087,673, International Application No. PCT / US2023 / 061038 (filed January 20, 2023), and International Application No. PCT / US2023 / 072872 (filed August 24, 2023), including, but not limited to, each of the disclosed reverse transcriptase amino acid and nucleotide sequences of the reverse transcriptases and each of the disclosed reverse transcriptase ncRNA nucleotide sequences described therein. The contents of these sequence listings are incorporated herein by reference in their entirety. Technical Field
[0002] This disclosure generally relates to systems, methods, and pharmaceutical compositions for precise genome editing, including targeted and precise nucleic acid editing, insertion, substitution, and deletion at genomic sites, wherein the systems, methods, and compositions are based on chimeric editing systems comprising one or more components of a lead editor and one or more components of a reverse transcriptome editor. Background Technology
[0003] Gene editing tools encompass a wide range of technologies capable of performing various types of genetic alterations in different contexts. These technologies have evolved over the past few decades, providing a suite of user-programmable editing tools, including ZFN (zinc finger) nuclease editing systems, large-scale nuclease editing systems, and transcription activator-like effector nucleases (TALENS). The past decade has seen an explosive growth in next-generation gene editing systems based on components derived from bacterial immune pathways, including clustered regularly interspaced short palindromic repeats (CRISPR) and associated CRISPR-related proteins (e.g., CRISPR-Cas9) (e.g., Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity”). Science(As described in Volume 337 (6096), pp. 816-821), a wide range of nuclease editors (e.g., Boissel et al., “megaTAL: a rare-cleaving nuclease architecture for therapeutic genome engineering”), Nucleic Acids Research 42: As described on pages 2591-2601) and bacterial reverse transcriptase systems (e.g., as Schubert et al., “High-throughput functional variant screens via invivo production of single-stranded DNA”, PNAS (As described in Volume 118(18), April 27, 2021, pp. 1-10).
[0004] CRISPR-Cas9 has been derivatized in various ways to extend its programmable double-stranded cleavage activity based on guide RNA, forming a range of systems from searching for alternative CRISPR Cas nucleases with different PAM requirements and cleavage properties (e.g., Cas12a, Cas12f, Cas13a, and Cas13b) to base editing (e.g., as Komor et al., “Programmableediting of a target base in genomic DNA without double-stranded DNA cleavage”). Nature May 19, 2016, 533 (7603); pp. 420-424 [Cytosine base editor or CBE] and Gaudelli et al., “Programmable base editing of AT to GC in genomic DNA without DNA cleavage”, Nature (as described in Volume 551, pp. 464-471 [Adenine Base Editor or ABE]) to leader editing (e.g., as Anzalone et al., “Search-and-replace genome editing without double-strand breaks or donor DNA”). Nature(As described in , December 2019, 576 (7789): pp. 149-157) to double primeediting (e.g., as Anzalone et al., “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin primeediting”). Nature Biotechnology (As described in Volume 40, pp. 731-740, December 9, 2021) to epigenetic editing (e.g., as Kunglovski and Jeltsch, “Epigenome Editing: State of the Art, Concepts, and Perspective”). Trends in Genetics (as described in Volume 32, 206, pp. 101-113) and CRISPR-guided integrase editing (e.g., as Yarnell et al., “Drag-and-drop genome insertion of large sequences without double-stranded DNA cleavage using CRISPR-directed integrases”). Nature Biotechnology November 24, 2022 doi.org / 10.1038 / s41587-022-01527-4 (As described in the text).
[0005] Despite the tremendous progress made in the field of gene editing over the past decade through the development of various CRISPR enzyme-based gene editing tools (e.g., base editors and leader editors) and other editing tools (e.g., reverse transcriptases), there remains a pressing need in the field for novel and improved editing systems that surpass the capabilities of existing known editing systems. Summary of the Invention
[0006] In one aspect, this disclosure describes a chimeric gene editing system comprising one or more components of a site-specific gene editing system (e.g., a leader editing system or a leader editor-like system) and one or more components of a reverse transcriptase editing system, thereby producing a novel gene editing system with unique properties, capabilities, and advantages. Furthermore, the provided system includes improved reverse transcriptase and reverse transcriptase editing systems to enhance gene editing specificity and efficiency.
[0007] In some embodiments, the chimeric gene editing system comprises: (a) a nicking enzyme component, (b) a guide RNA compounded with the nicking enzyme component and directed to a target DNA sequence, (c) a polymerase component, and (d) a polymerase template sequence, wherein the polymerase template sequence is provided by a modified reverse transcriptase ncRNA containing the polymerase template sequence (referred herein to as "template-based ncRNA" or "tncRNA"). In several embodiments, the tncRNA comprises a structural configuration altered relative to the wild-type reverse transcriptase ncRNA, wherein the alteration includes: (i) a linker that joins the 5' end of the ncRNA a1 region to the 3' end of the ncRNA a2 region, (ii) a deletion of the wild-type msd region or a portion thereof, or (iii) the insertion of a single-stranded RNA sequence containing the polymerase template in place of the deleted msd sequence. The chimeric gene editing system may also optionally comprise a reverse transcriptase (RT) to convert the template-based ncRNA into homologous msDNA, referred herein to as "template-based msDNA" or "tmsDNA".
[0008] In some other embodiments, the chimeric gene editing system comprises: (a) a nicking enzyme component or a nucleic acid sequence (e.g., mRNA) encoding therein; (b) a guide RNA that is compounded with the nicking enzyme component and guides it to a target DNA sequence; (c) a polymerase component or a nucleic acid sequence (e.g., mRNA) encoding therein; and (d) a polymerase template sequence, wherein the polymerase template sequence is provided by a modified reverse transcriptase ncRNA containing the polymerase template sequence (referred herein to as “template-based ncRNA” or “tncRNA”). In several embodiments, the tncRNA comprises a structural configuration altered relative to the wild-type reverse transcriptase ncRNA, wherein the alteration includes: (i) a connector that joins the 5' end of the ncRNA a1 region to the 3' end of the ncRNA a2 region; (ii) a deletion of the wild-type MSD region or a portion thereof; or (iii) the replacement of the deleted MSD sequence with a single-stranded RNA sequence containing the polymerase template. The chimeric gene editing system may also optionally include a reverse transcriptase RT (or a nucleic acid sequence encoding the reverse transcriptase RT) to convert templated ncRNA into homologous msDNA, referred to herein as “templated msDNA” or “tmsDNA”.
[0009] See the attached diagram. Figure 1A schematic diagram of a leader editor is provided, comprising (a) a reverse transcriptase component fused with (b) a CRISPR Cas9 nicking enzyme component, and (c) a leader editor-guided RNA (pegRNA) complex (i.e., the PE:pegRNA complex). The pegRNA contains (i) a primer binding site (PBS) and (ii) a template sequence. The mechanism of leader editing is described in the art. The PE:pegRNA complex binds to the target DNA at a site specified by a spacer sequence of the pegRNA. The Cas9 nicking enzyme then nicks one strand, generating a 3' flap (“3' flap”) of endogenous DNA. PBS located on the pegRNA binds to the 3' flap, and the edited RNA sequence is reverse transcribed from the 3' flap by the reverse transcriptase of the PE:pegRNA complex targeting the template portion of the pegRNA. The reverse transcription product is efficiently extended from the endogenous 3' flap, forming a newly synthesized single strand of DNA (i.e., the RT product or the edited strand). The edited strand replaces the endogenous strand immediately downstream of the nick on the nicked strand, forming a 5' lobe of the replaced endogenous DNA, which is removed by cellular enzymes. Due to endogenous DNA repair enzymes and a round of DNA replication, the edited DNA is permanently installed on both strands. This process can be enhanced by using an additional guide RNA of normal length (programmed to install a second nick in the unedited strand by recombining with PE). This specific DNA damage is recognized by the cell, prompting the cell to repair the unedited strand to match as the complementary strand to the edited strand. This process places both strands in the edited conformation, thus permanently installing the edit.
[0010] See Figure 2 This describes a reverse transcriptase editing system. The reverse transcriptase encodes and transcribes a single RNA containing a non-coding RNA (ncRNA) portion and a portion encoding a specific reverse transcriptase (RT) (see [link to documentation]). Figure 1 ). Reverse transcript ncRNA ( MSR and msd ( ) is the precursor to the final hybrid molecule, and it initially folds into a typical RNA secondary structure, which is recognized by the accompanying RT. For example Figure 2 As shown, translated RTs typically recognize specific secondary structures in ncRNAs and bind to them. msd The RNA template is located downstream of the region. RT begins at the 2' end of the double-stranded RNA structure (a1 / a2 region) within the ncRNA, immediately adjacent to the conserved guanosine (G) residue. The initiator RNA reverses transcription towards its 5' end. A portion of the ncRNA serves as the template for reverse transcription, and the reverse transcription proceeds to the 5' end. MSRThe transcription process terminates before the locus. During reverse transcription, cellular ribonuclease H degrades the ncRNA segment used as a template, but not other parts of the ncRNA. (Example: ...) Figure 3 As shown, the reverse transcription results in msDNA being covalently attached to the RNA template via 2'-5' phosphodiester bonds, and the 3' end of the msDNA is used for base pairing with the RNA template.
[0011] like Figure 3 As shown, this disclosure generally describes a chimeric gene editing system comprising one or more components of a lead editor system and one or more components of a reverse transcriptase editor. The resulting chimeric editing system is referred to herein as a "nickase / reverse transcriptase template precise editing system".
[0012] This disclosure considers a variety of configurations and interchangeable components that can constitute different implementations of the precise editing system for nickase / reverse transcriptase templates described herein.
[0013] In several implementations, the precise editing system for nickase / reverse transcript templates described herein may include: (1) A nicking enzyme or a nucleic acid sequence encoding a nicking enzyme (e.g., mRNA encoding a nicking enzyme); (2) A guide RNA that combines with the nicking enzyme and guides the nicking enzyme to the target DNA sequence; (3) The polymerase or the nucleic acid sequence encoding the polymerase (e.g., mRNA encoding the polymerase); and (4) A polymerase template sequence for synthesizing the edit strand to be incorporated into the target site, wherein the polymerase template sequence is provided by templated ncRNA.
[0014] In several embodiments, the tncRNA comprises a structural configuration altered relative to the wild-type reverse transcriptase ncRNA, wherein the alteration includes, but is not limited to: (i) a adapter that joins the 5' end of the ncRNA a1 region to the 3' end of the ncRNA a2 region; (ii) deletion of the wild-type MSD region or a portion thereof; and / or (iii) installation of a single-stranded RNA sequence containing a polymerase template in place of the deleted MSD sequence. The polymerase template may also contain a primer binding site that adheres to an endogenous 3' lobe resulting from the introduction of a nick at the target site by a nicking enzyme. The adhesion of the primer binding site to the endogenous 3' lobe creates a polymerization initiation site for synthesizing a new strand of single-stranded DNA (“edited strand” or “edited lobe”) starting from the 3' end of the endogenous lobe and using the polymerase template as an indicator. The edited strand replaces the endogenous strand immediately downstream of the nick on the nicked strand, forming a 5' lobe of the replaced endogenous DNA, which is removed by cellular enzymes. Due to endogenous DNA repair enzymes and DNA replication, the edited lobe is permanently attached to both strands. This process can be enhanced by using an additional guide RNA of normal length (programmed to install a second nick in the unedited strand by recombining with a nicking enzyme). This specific DNA damage is recognized by the cell, prompting it to repair the unedited strand to match the complementary strand of the edited strand. This process puts both strands in the edited conformation, thus permanently installing the edit.
[0015] Figure 3 The chimeric gene editing system shown may also optionally include a reverse transcriptase RT (or a nucleic acid sequence encoding the reverse transcriptase RT) to convert templated ncRNA into homologous msDNA, referred to herein as “templated msDNA” or “tmsDNA”.
[0016] In several implementations of the precise editing system for nickase / reverse transcriptase templates described herein, the polymerase template can be provided directly by templated ncRNA (e.g., ...). Figure 6 (As shown). In other embodiments, the polymerase template of the nicking enzyme / reverse transcriptome template-precise editing system described herein can be provided by templated ncRNA and a reverse transcriptome RT (which converts the ncRNA into templated msDNA), which is then efficiently synthesized by the polymerase and the edited strand is used to provide template functionality for incorporation. This embodiment is depicted in Figure 5 middle.
[0017] In several implementations, the nickase / reverse transcript template precise editing system described herein comprises templated ncRNA and / or templated msDNA, such as Figure 4As shown in Figure A, an unfolded wild-type ncRNA is depicted. The molecule contains an a1 region, an msr region, an msd region, and an a2 region in the 5' to 3' direction. The ncRNA folds as shown in Figure B, where the a1 / a2 regions contain anticomplementary regions, thus forming a double strand. Furthermore, the msr region forms one or more stem-loops, as does the msd region. See also... Figure 2 Figure 1 provides a detailed schematic diagram of how ncRNA is routinely converted into msRNA via reverse transcriptase (RT). In contrast, Figures C and D are components of the precise editing system for nickase / reverse transcriptase templates described herein. Figure C illustrates one embodiment of templated ncRNA. As depicted, the tncRNA contains a structural configuration altered relative to the wild-type reverse transcriptase ncRNA (e.g., as shown in Figure B), said alterations including, but not limited to: (i) a linker that joins the 5' end of the ncRNA a1 region to the 3' end of the ncRNA a2 region; (ii) the deletion of the wild-type msd region or a portion thereof; and (iii) the replacement of the deleted msd sequence with a single-stranded RNA sequence containing a polymerase template ((i) or (ii)). The polymerase template may also contain a primer binding site ((i) or (ii)) that adheres to an endogenous 3' valve generated by introducing a nick at the target site via nickase. Figure D depicts the structure of templated msDNA, generated after reverse transcription of a single-stranded RNA sequence containing a polymerase template.
[0018] Figure 5 An embodiment of the precise editing system for the nicking enzyme / reverse transcriptome template described herein is depicted. In this embodiment, templated ncRNA is converted into templated msDNA via reverse transcriptase (RT). Simultaneously, a complex containing a fusion protein with a nicking enzyme and a polymerase, along with a guide RNA, binds to and cleaves the unedited target DNA, thereby introducing a nick (“nick DNA”) into the DNA. The templated msDNA (containing a polymerase template with PBS) then associates with the nick DNA by adhering the PBS to an endogenous 3' lobe immediately upstream of the nick on the nick strand. The adhered PBS provides a starting point for the polymerase of the fusion protein (e.g., a DNA-dependent DNA polymerase) to synthesize a new DNA strand from the available 3' end of the endogenous 3' lobe, which serves as a template for the polymerase template. This produces an edited 3' lobe that replaces the 5' endogenous lobe formed downstream of the nick site and is incorporated into the DNA following cellular DNA repair and replication processes. In some embodiments, a second nick guide RNA may be provided to introduce a nick into the unedited strand downstream of the original nick site. The second nick site induces the cell to replace the unedited strand, creating reverse complementarity against the 3' edit lobe, thereby introducing the edit into both strands. Replication ensures the edit is permanently incorporated into both strands.
[0019] Figure 6 Another embodiment of the precise editing system for the nicking enzyme / reverse transcriptome template described herein is depicted. In this embodiment, the templated ncRNA directly provides polymerase template function (without needing to convert it to msDNA via reverse transcriptome RT). Simultaneously, a complex containing a fusion protein with both a nicking enzyme and a polymerase, along with a guide RNA, binds to and cleaves the unedited target DNA to introduce a nick (“nick DNA”) into the DNA. The templated ncRNA (containing a polymerase template with PBS) then associates with the nick DNA by adhering the PBS to an endogenous 3' lobe immediately upstream of the nick on the nick strand. The adhered PBS provides a starting point for the polymerase of the fusion protein (e.g., an RNA-dependent DNA polymerase) to synthesize a new DNA strand from the available 3' end of the endogenous 3' lobe, which serves as a template for the polymerase template. This produces an edited 3' lobe that replaces the 5' endogenous lobe formed downstream of the nick site and is incorporated into the DNA following cellular DNA repair and replication processes. In some implementations, a second nick site guide RNA may be provided to introduce a nick into the unedited strand downstream of the original nick site. The second nick site induces the cell to displace the unedited strand, making it anticomplementary to the 3' edit lobe, thereby introducing the edit into both strands. Replication ensures the edit is permanently incorporated into both strands.
[0020] Figures 7A to 7D The mechanism of the precise editing system for nickase / reverse transcriptase templates described in this paper is presented in more detail. This system relies on the use of msDNA (i.e., corresponding to...) Figure 5 (General configuration). Figure 7A The association of a fusion protein containing nickase and polymerase with a guide RNA complex is depicted. Black triangles represent the catalytic activity of a single nuclease. Figure 7B This demonstrates how the guide RNA binds to the target site by adhering to a specific complementary sequence, which in turn triggers the action of a nuclease, thereby introducing a single nick in the top strand (which will then become the "edited strand"). Figure 7C In this embodiment, a 3' valve region is formed immediately upstream of the cut, and this region adheres to the templated msDNA introduced into the PBS containing the msDNA. In this embodiment, the PBS is located at the 3' end of the msDNA, and the template region is located upstream of the PBS. Furthermore, since this embodiment involves msDNA, the polymerase template (containing both PBS and the template) is a single-stranded DNA that binds to the msDNA via 2' to 5' linkages with guanosine. Given that the template is DNA, the polymerase can be a DNA-dependent DNA polymerase. Next, as... Figure 7DAs shown in Figure (1), the polymerase, guided by the template of msDNA, synthesizes a new DNA strand starting from the 3' end of the endogenous valve. Once the complex dissociates from the msDNA, the newly synthesized DNA strand replaces the endogenous DNA, forming a 5' endogenous valve, which is removed by the host enzyme (Figure (2)). Finally, as shown in Figure (3), the editing is incorporated into both strands because the cell repair and replication processes preferentially edit the endogenous strand to match the reverse complement of the first edited strand.
[0021] In several embodiments, the nickase / reverse transcript template-precise editing system described herein includes both a guide RNA (for guiding the nickase to the target site) and a reverse transcript templated ncRNA. In some embodiments, the guide RNA and ncRNA may be fused together, such as, for example... Figure 8A As shown (PBS = primer binding site and RTT = reverse transcriptase template). In other embodiments, the guide RNA and ncRNA can be delivered separately and without conjugation, such as... Figure 8B As shown. In other embodiments, such as Figure 8C As shown, the structure of the ncRNA can be further minimized by removing all or part of the a1 / a2 stem / loop. In the described embodiment, the guide RNA may or may not be fused with the ncRNA. Other configurations for delivering the ncRNA and the guide RNA are described in Figure 8E middle.
[0022] In several other embodiments, the ncRNA used herein may include one or more stabilization modifications, including stable nucleotide analogs, such as 3' hairpin structures, or circularization. Such embodiments are shown in Figure 8D middle.
[0023] In another implementation scheme, such as Figure 9 As shown, the nicking enzyme / reverse transcriptome template precise editing system can be configured as a dual-editing system to produce a pair of complementary edited strands (as shown in Figure (2)), which are then recombined into the target DNA (as shown in Figure (3)). In this embodiment, as shown in Figure (1), the system comprises a first template ncRNA and a first nicking enzyme-polymerase-guide RNA complex (gray dashed circle), which creates a nick at a first site and forms a first single-stranded DNA editing flap. In this case, since ncRNA is used directly (i.e., msDNA is not present), a reverse transcriptome RT is provided to convert the ncRNA into msDNA. Furthermore, since the ncRNA is whole RNA, the polymerase is an RNA-dependent DNA polymerase, enabling it to synthesize the edited DNA strand against the RNA template.
[0024] Figure 10A and Figure 10BAnother implementation scheme is described, which is related to Figure 9 The difference in the system is that each ncRNA has a first PBS and a second PBS. As shown in Figure (1), the first ncRNA contains a polymerase template with PBS side-attached sites (PBS1 and PBS2). Figure (2) depicts that, as a first step, the RT (e.g., reverse transcriptase RT) of the first fusion protein (which contains a nicking enzyme (black dashed circle) and RT) starts from the 3' of the endogenous 3' lobe, synthesizes DNA for the first RTT, extending to cover the entire RTT and PBS2 downstream of the RTT. PBS2 is encoded into the resulting first edited DNA lobe (corresponding to the PBS2' DNA sequence), which adheres to the second 3' endogenous lobe formed at the second nicking site, as shown in Figure (1). Figure 3 As shown. The DNA is then synthesized again by a second editing fusion protein (e.g., by a DNA-dependent DNA polymerase, which could theoretically be an engineered reverse transcriptase RT or another DNA-dependent DNA polymerase). The second synthesis uses the product of the first synthesis as a template. Similarly, in Figure (4), the double-stranded DNA duplex is then incorporated into the DNA, thereby assembling the edited strand.
[0025] exist Figure 11 In the various embodiments described, each component of the nickase / reverse transcript template precision editing system may contain various modifications resulting from optimization, engineering, mutagenesis, orthologous substitution, and other engineering processes to alter, modify, or improve editing functionality, stability and persistence, and / or other characteristics.
[0026] In such Figure 12 In other aspects described, this disclosure provides pharmaceutical compositions, such as lipid nanoparticles (LNPs), comprising suitable cargo elements constituting the components of a nickase / reverse transcript template precise editing system. The nickase / reverse transcript template precise editing system can be delivered as a protein, nucleic acid, or protein-nucleic acid complex using any delivery system. For example, the nickase / reverse transcript template precise editing system can be delivered in a whole-RNA configuration comprising one or more mRNAs encoding a nickase, a polymerase, and optionally a reverse transcriptase RT (i.e., where the polymerase template is msDNA, such that it can be converted from ncRNA), and functional RNA components such as targeting guide RNA, a second nick guide RNA, and ncRNA.
[0027] In another aspect, this disclosure further provides nucleic acid molecules encoding components of a gene editing system (e.g., recombinant ncRNA and / or recombinant reverse transcriptase RT). In another aspect, this disclosure provides a genome editing system comprising recombinant reverse transcriptase components (e.g., recombinant ncRNA and / or recombinant RT), a programmable nickase, and a guide RNA. In another aspect, this disclosure provides nucleic acid molecules encoding said genome editing system and its components, as well as polypeptides constituting components of said genome editing system. In another aspect, this disclosure provides vectors for transferring and / or expressing said genome editing system, for example, under in vitro, ex vivo, and in vivo conditions. In another aspect, this disclosure provides cell delivery compositions and methods, including compositions (e.g., plasmids) for passive and / or active transport to cells, delivery via virus-based recombinant vectors (e.g., AAV and / or lentiviral vectors), delivery via non-virus-based systems (e.g., liposomes and LNPs), and delivery via virus-like particles. Depending on the delivery system employed, the genome editing systems described herein can be delivered in the following forms: DNA (e.g., plasmids or DNA-based viral vectors), RNA (e.g., ncRNA and mRNA delivered via LNPs), mixtures of DNA and RNA, proteins (e.g., virus-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combination of methods can be used to deliver the components of the retrotranscribed genome editing systems disclosed herein. In one embodiment, each component of the retrotranscribed genome editing system is delivered by a whole RNA system, for example, by delivering one or more RNA molecules (e.g., mRNA and / or ncRNA) via one or more LNPs, wherein the one or more RNA molecules constitute ncRNA and guide RNA (as needed) and / or are translated into polypeptide components (e.g., RT and programmable nucleases). In another aspect, this disclosure provides genome editing methods that generate edits at the target editing site by introducing the retrotranscribed genome editing systems described herein into cells containing target editing sites (e.g., under in vitro, in vivo, or ex vivo conditions). In other respects, this disclosure provides formulations comprising any of the foregoing components for delivery to cells and / or tissues, including in vitro, in vivo, and ex vivo delivery; recombinant cells and / or tissues modified by the recombinant reverse transcriptase-based genome modification systems and methods described herein; and methods for modifying cells using the reverse transcriptase-based genome modification systems disclosed herein through genome editing and related DNA donor-dependent methods (e.g., recombinant engineering or cell recording). This disclosure also provides methods for preparing the recombinant reverse transcriptases, reverse transcriptase-based genome modification systems, vectors, compositions, and formulations described herein, as well as pharmaceutical compositions and kits for modifying cells in vitro, in vivo, and under in vitro conditions, comprising the genome editing and / or modification systems disclosed herein.
[0028] In one embodiment, this disclosure or the present invention provides a gene editing system comprising one or more delivery media, wherein: the delivery media comprises RNA cargo; the RNA cargo comprises (a) at least one mRNA molecule encoding (i) a nucleic acid programmable nuclease (e.g., a nickase) and (ii) a polymerase (e.g., a DNA-dependent DNA polymerase, an RNA-dependent DNA polymerase, such as a reverse transcriptase), (b) an engineered templated reverse transcriptase ncRNA, and (c) a guide RNA for the programmable nuclease; and each delivery media contains (a)(i) and / or (a)(ii) and / or (b) and / or (c); thereby one or more delivery media deliver (a)(i), (a)(ii), (b), and (c).
[0029] In one implementation, in the gene editing system, (a)(i) and (a)(ii) comprise a single mRNA molecule that encodes a nucleic acid programmable nuclease and polymerase (e.g., a reverse transcriptase).
[0030] In one implementation, in the gene editing system, (a)(i) and (a)(ii) are encoded and expressed as fusion proteins.
[0031] In one embodiment, in the gene editing system, (a)(i) and (a)(ii) are encoded and expressed as a fusion protein, and the fusion protein comprises the C-terminus of a nucleic acid programmable nuclease fused to the N-terminus of a polymerase (e.g., a reverse transcriptase) (nuclease:polymerase); or, the fusion protein comprises the N-terminus of a nucleic acid programmable nuclease fused to the C-terminus of a reverse transcriptase (polymerase:nuclease fusion).
[0032] In one implementation, in the gene editing system, (a)(i) and (a)(ii) comprise a first mRNA molecule encoding a programmable nuclease for nucleic acids and a second mRNA molecule encoding a polymerase (e.g., reverse transcriptase).
[0033] In one implementation, in the gene editing system, (c) is separate from (a)(i), (a)(ii) and (b), or is provided in trans form.
[0034] In one implementation, in the gene editing system, (b) engineered reverse transcriptase ncRNA and (c) guide RNA fusion, or provided in cis.
[0035] In one implementation, in the gene editing system, (b) engineered reverse transcriptase ncRNA and (c) guide RNA fusion or provided in cis, and the guide RNA is fused to the 5' end of the reverse transcriptase ncRNA.
[0036] In one implementation, in the gene editing system, (b) engineered reverse transcriptase ncRNA and (c) guide RNA fusion or provided in cis, and the guide RNA is fused to the 3' end of the reverse transcriptase ncRNA.
[0037] In one embodiment, in the gene editing system, (b) an engineered reverse transcriptase ncRNA and (c) a guide RNA are fused or provided in cis, and said engineered ncRNA comprises a first guide RNA fused to the 5' end of the reverse transcriptase ncRNA and a second guide RNA fused to the 3' end of the reverse transcriptase ncRNA, wherein the first and second guide RNAs target different sequences. Thus, more broadly, in one embodiment, in the gene editing system, (c) the guide RNA for a programmable nuclease may comprise one or more guides targeting the same or different target sequences. In one embodiment, such a guide RNA may be a single guide RNA or sgRNA; for example, when the nucleic acid programmable nuclease comprises Cas9.
[0038] In one implementation, in the gene editing system, one or more delivery media comprise liposomes or lipid nanoparticles (LNPs).
[0039] In one implementation, in the gene editing system, (a) at least one mRNA molecule encoding (i) a nucleic acid programmable nuclease and (ii) a polymerase (e.g., a reverse transcriptase), and (b) an engineered reverse transcriptase ncRNA, are located in the same delivery medium.
[0040] In one implementation, in the gene editing system, (a) at least one mRNA molecule encoding (i) a nucleic acid programmable nuclease and (ii) a polymerase (e.g., a reverse transcriptase), and (b) an engineered reverse transcriptase ncRNA, are located in a separate delivery medium.
[0041] In one implementation, in the gene editing system, nucleic acid programmable nucleases and polymerases (e.g., reverse transcriptases) are encoded on separate mRNA molecules, and these separate mRNA molecules (a)(i) and (a)(ii) are contained in the same delivery medium.
[0042] In one implementation, in the gene editing system, nucleic acid programmable nucleases and polymerases (e.g., reverse transcriptases) are encoded on separate mRNA molecules, and these separate mRNA molecules (a)(i) and (a)(ii) are contained in different delivery media.
[0043] In one embodiment, in a gene editing system, the engineered reverse transcriptase ncRNA includes a polymerase template comprising PBS and a template sequence containing the intended edit to be integrated into a target sequence in the cell, wherein the polymerase template contains one or more regions homologous to an endogenous sequence downstream of the nick site to facilitate integration of the edited strand. The editing system can be used in animal cells, or mammalian cells (e.g., primates, non-human primates, or domesticated mammals such as cats, dogs, or horses) or human cells; for example, to correct, resolve, treat, or alleviate genetic diseases in animals, mammals, domesticated mammals, cats, dogs, horses, or humans. This operation can be performed in plant cells to introduce mutations that produce advantageous phenotypic traits (e.g., disease resistance or other advantageous plant traits).
[0044] In one embodiment, in the gene editing system, the programmable nucleic acid nuclease includes Cas9 nuclease, TnpB nuclease, or Cas12a nuclease. Preferably, the nuclease is a nicking enzyme, such that only one of the two strands in the target region is cleaved.
[0045] In several embodiments, the templated ncRNA may be derived from baseline or wild-type ncRNA, such as any of the following: those ncRNAs or nucleotide sequences of Table B disclosed in U.S. Application No. 18 / 087,673 or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 (each incorporated herein by reference in its entirety, including its sequence listing), or nucleotide sequences having at least 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with the sequence of any of the above applications in Table B. The polymerase template may be cellularly heterologous. Alternatively, the polymerase template may be cellularly endogenous. For example, the cell may contain sequences typical of people with disease states, and the donor polynucleotide may be sequences typical of people without disease states (e.g., the donor may be used for gene correction or repair of cells to modify cells with mutations or modifications that cause disease states to cells with sequences typical of disease-free states).
[0046] In one embodiment of the gene editing system, the gene editing system may comprise any combination of the above-described embodiments of the gene editing system.
[0047] In one embodiment, this disclosure or invention provides a cell herein, such as a single cell comprising the gene-editing system disclosed herein, as in any of the preceding paragraphs. In one embodiment, the cell (e.g., a single cell) may be a eukaryotic cell. In one embodiment, the eukaryotic cell may be a plant cell, an animal cell, or a mammalian cell, for example, a single plant cell, a single animal cell, or a single mammalian cell. In one embodiment, the mammalian cell (e.g., a single mammalian cell) may be a human cell. In one embodiment, the cell may be a prokaryotic cell, such as a bacterial cell. In this embodiment where the cell is a bacterial cell, the donor polynucleotide may encode antibiotic sensitivity; and therefore, this disclosure and any invention disclosed herein may relate to a method of combating antibiotic-resistant bacteria by sensitizing such bacteria to antibiotics (and a subject given the gene-editing system may then also receive antibiotics to which the bacteria have been sensitized by the gene-editing system).
[0048] In one embodiment, this disclosure or the present invention provides a composition comprising: a) a gene editing system disclosed herein, such as those described in any of the foregoing paragraphs; and b) a pharmaceutically or veterinary-acceptable carrier. In one embodiment, in the composition, the delivery medium may comprise lipid nanoparticles comprising: a) one or more ionizable lipids; b) one or more structural lipids; c) one or more PEGylated lipids; and d) one or more phospholipids. In one embodiment, in the composition, the one or more ionizable lipids comprise ionizable lipids set forth in Table 2 of U.S. Application No. 18 / 087,673 or International Application No. PCT / US2023 / 061038.
[0049] In one embodiment, this disclosure or invention provides herein the use of the gene editing system embodiments and / or compositions disclosed herein, for example, in any of the foregoing paragraphs; for example, for modifying or genetically modifying cells, such as eukaryotic or prokaryotic cells and / or animal cells and / or mammalian cells and / or human cells and / or bacterial cells and / or plant cells (e.g., any cells discussed herein, wherein said cells comprise single isolated cells). In one embodiment, this disclosure or invention provides herein the use of the gene editing system embodiments and / or compositions disclosed herein, for example, in any of the foregoing paragraphs; for example, for treating or addressing a subject's hereditary disease.
[0050] In one embodiment, this disclosure or the present invention provides a method for genetically modifying cells, the method comprising: contacting a gene editing system as discussed herein (e.g., in any of the preceding paragraphs) or a composition as discussed herein (e.g., in any of the preceding paragraphs) (which contains a gene editing system as discussed herein (e.g., in any of the preceding paragraphs), advantageously including a gene editing system encoding a target sequence of a donor polynucleotide, the donor polynucleotide containing a desired edit to be integrated into a target sequence in the cell, the method comprising contacting the composition or the gene editing system with the cell to deliver RNA cargo to the cell, wherein: a nucleic acid programmable nuclease forms a complex with a guide RNA, wherein the guide RNA guides the complex to the target sequence; the nucleic acid programmable nuclease produces a double-strand break in the target sequence; a reverse transcriptase and an engineered reverse transcriptase ncRNA produce msDNA containing the donor polynucleotide; and the donor polynucleotide is integrated into the target sequence; thereby editing the cell for genetic modification. In one embodiment, the cell may be a eukaryotic cell or a prokaryotic cell or an animal cell or a mammalian cell or a human cell or a bacterial cell or a plant cell.
[0051] In another aspect, this disclosure provides a non-coding RNA (ncRNA) variant comprising a reference reverse transcriptase ncRNA having one or more modifications, wherein the reference reverse transcriptase ncRNA comprises, from 5' to 3': an a1 region, a first branch guanine, msr, msd, and an a2 region, and the one or more modifications comprise: (i) a linking of the a1 region to the a2 region; (ii) deletion of at least a portion of the msr; (iii) deletion of at least a portion of the msd; (iv) addition of a single-stranded RNA comprising a polymerase template; or (v) addition of an RNA motif.
[0052] In one embodiment, the one or more modifications include a1 region and a2 region bonding.
[0053] In one embodiment of any of the foregoing embodiments, the one or more modifications include a bonding between regions a1 and a2, wherein the bonding between regions a1 and a2 is formed by joining the 5' end of region a1 and the 3' end of region a2 via a connector. In another embodiment of any of the foregoing embodiments, the one or more modifications include a bonding between regions a1 and a2, wherein the bonding between regions a1 and a2 is formed by directly joining the 5' end of region a1 to the 3' end of region a2 without a connector.
[0054] In any of the foregoing embodiments, the a1 region and the a2 region are partially or completely complementary to each other, and the a1 region and the a2 region form a stem-loop structure after bonding.
[0055] In any of the foregoing embodiments, the ncRNA variant further comprises a second-branched guanosine.
[0056] In any of the foregoing embodiments, the bonding between the a1 region and the a2 region causes ncRNA variant circularization.
[0057] In any of the foregoing embodiments, the one or more modifications further include a break between the msd and a2 regions to generate an ncRNA variant comprising, from the 5' to 3' direction: the a2 region, the bond between the a1 and a2 regions, the a1 region, the first branch guanosine, the msr, and the msd.
[0058] In any of the foregoing embodiments, the one or more modifications comprise the deletion of at least a portion of the MSR. In one embodiment, the deleted portion of the MSR comprises a spacer between two stem-loops in the MSR or a portion of said spacer. In one embodiment, the deleted portion of the MSR comprises a portion of a spacer, thereby preserving 1-15 base pairs remaining between the two stem-loops in the MSR compared to the reference reverse transcriptase ncRNA. In one embodiment, the deletion preserves 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs remaining compared to the reference reverse transcriptase ncRNA. In one embodiment, the deleted portion of the MSR comprises the entire spacer between the two stem-loops in the MSR compared to the reference reverse transcriptase ncRNA. In one implementation, the missing portion of the MSR includes a stem loop or a portion thereof within the MSR.
[0059] In any of the foregoing embodiments, the one or more modifications comprise the deletion of at least a portion of the MSD. In one embodiment, the deleted portion of the MSD comprises a spacer or a portion of the spacer located between the stem loop and the A2 region in the MSD. In one embodiment, the deleted portion of the MSD comprises a portion of the spacer located between the stem loop and the A2 region in the MSD, thereby preserving 1-15 base pairs remaining between the stem loop and the A2 region in the MSD compared to the reference reverse transcriptase ncRNA. In one embodiment, the deletion preserves 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs remaining between the stem loop and the A2 region in the MSD compared to the reference reverse transcriptase ncRNA. In one embodiment, the missing portion in the MSD includes the entire spacer between the stem loop in the MSD and the A2 region. In another embodiment, the missing portion in the MSD also includes the stem loop in the MSD, whereby the missing portion includes the entire spacer between the stem loop in the MSD and the A2 region, as well as the missing stem loop in the MSD.
[0060] In any of the foregoing embodiments, the one or more modifications comprise the addition of a single-stranded RNA containing a polymerase template. In one embodiment, the one or more modifications comprise the addition of a single-stranded RNA containing a polymerase template and a deletion of at least a portion of the MSD, wherein the single-stranded RNA containing the polymerase template is added to the deleted portion of the MSD. In one embodiment, the single-stranded RNA containing the polymerase template comprises a pair of homologous arms specifically targeting the target locus. In one embodiment, each of the homologous arms comprises 5-200 nucleotides. In one embodiment, each of the homologous arms comprises 5 to 100 nucleotides, 5 to 50 nucleotides, 10 to 50 nucleotides, 25 to 50 nucleotides, 5 to 25 nucleotides, 10 to 25 nucleotides, 10 to 20 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, or 40 to 50 nucleotides. In one embodiment, the single-stranded RNA containing the polymerase template further comprises a donor sequence for integration into the target locus. In one embodiment, the donor sequence comprises 2-100 nucleotides, 2-50 nucleotides, 2-25 nucleotides, 5-25 nucleotides, 5-15 nucleotides, 5-10 nucleotides, 10-30 nucleotides, 10-20 nucleotides, or 1-10 nucleotides.
[0061] In any of the foregoing embodiments, the one or more modifications comprise the addition of an RNA motif. In one embodiment, the RNA motif is an MS2 stem-loop or 3' tail of a U7 small nuclear RNA (snRNA).
[0062] In any of the foregoing embodiments, the one or more modifications comprise: (a) a combination of (i) and (ii): the a1 region is linked to the a2 region and at least a portion of the MSR is deleted; (b) a combination of (i) and (iii): the a1 region is linked to the a2 region and at least a portion of the MSD is deleted; (c) a combination of (i) and (iv): the a1 region is linked to the a2 region and a single-stranded RNA containing a polymerase template is added; (d) a combination of (i) and (v): the a1 region is linked to the a2 region and an RNA motif is added; (e) a combination of (ii) and (iii): at least a portion of the MSR is deleted and at least a portion of the MSD is deleted; (f) a combination of (ii) and (iv): at least a portion of the MSR is deleted and a single-stranded RNA containing a polymerase template is added; (g) a combination of (ii) and (v): at least a portion of the MSR is deleted and an RNA motif is added at the 5' or 3' end; or (h) (iv) and (v) combination: add a single-stranded RNA sequence containing the polymerase template and add an RNA motif at the 5' or 3' end.
[0063] In any of the foregoing embodiments, the one or more modifications comprise: (a) a combination of (i), (ii), and (iii): a binding of the a1 region to the a2 region, deletion of at least a portion of the msr, and deletion of at least a portion of the msd; (b) a combination of (i), (ii), and (iv): a binding of the a1 region to the a2 region, deletion of at least a portion of the msr, and addition of a single-stranded RNA containing a polymerase template; (c) a combination of (i), (iii), and (iv): a binding of the a1 region to the a2 region, deletion of at least a portion of the msd, and addition of a single-stranded RNA sequence containing a polymerase template; or (d) a combination of (ii), (iii), and (iv): deletion of at least a portion of the msr, deletion of at least a portion of the msd, and addition of an RNA motif at the 5' or 3' end.
[0064] In any of the foregoing embodiments, the one or more modifications comprise: (i) a1 region to a2 region, (ii) deletion of at least a portion of msr, (iii) deletion of at least a portion of msd, (iv) addition of a single-stranded RNA sequence containing a polymerase template, and (v) addition of an RNA motif at the 5' or 3' end region.
[0065] In any of the foregoing embodiments, the one or more modifications further comprise substitution, insertion, or deletion of one or more nucleotides, or a combination thereof. In one embodiment, the substitution, insertion, or deletion is capable of modulating the activity of a first or second polymerase, generating single-stranded DNA from an ncRNA variant, or the immunogenicity of the ncRNA variant.
[0066] In any of the foregoing embodiments, the reference reverse transcriptant ncRNA comprises a sequence of a naturally occurring reverse transcriptant or a portion thereof.
[0067] In any of the foregoing embodiments, the reference reverse transcriptase ncRNA comprises a sequence selected from the following: SEQ ID NO: 3980-4178, 11231-11429, 4671-4825, 11922-12075, 4980-5143, 12229-12392, 367-368, 427-441, 494-521, 526-527, 536, 626, 649, 660-668, 675, 679, 687-692, 695, 697, 703, 716, 721-722, 751-763, 767, 770-1411, 1456-1462, 7624-7625, 7684-7698, 7751-7778, 7783-7784, 7793, 78 83, 7906, 7917-7925, 7932, 7936, 7944-7949, 7952, 7954, 7960, 7973, 7978-7979, 8008-8020, 8024, 8027-8667, 8712-8718, 1529-1569, 8784-8823, 6697-6701, 13943-13947, 4179-4670, 11430-11921, 4884-4909, 12134-12159, 6919-6972, 14163-14215, 2786-2866, 2887-2938, 10039 -10119, 10140-10191, 4826-4863, 12076-12113, 4864-4875, 12114-12125, 6974-7002, 14217-14244, 2598-2600, 2759-2785, 9851-9853, 10012-10038, 2445-2582, 9699-9836, 1983-2158, 9237-9412, 1612-1982, 8866-9236, 2601-2678, 9854-9931, 2679-2758, 9932-10011, 3442-360 3. 10694-10855, 3604-3708, 10856-10959, 2939-3441, 3709-3979, 5177-5192, 10192-10693, 10960-111230, 12426-12441, 7003-7033, 14245-14275, 7054-7133, 14296-14374, 7034-7049, 14276-14291, 6835-6918, 14079-14162, 6823-6834, 14068-14078, 298-366, 369-373, 442-493522-525、528-535、537、551-554、557、560-625、672-674、680-681、684-686、696、698、702、723-742、764-766、1412-1453、1463-1466、1571-1577、7555-7623、2626-7630、7699-7750、7785-7792、7794、7808-7811、7814、7817-7882、7929-7931、7937-7938、7941-7943、7953、7955、7959、7980-7999、8021-8023、8668-8706、8708-8709、8719-8722、8825-8831、374-426、539-550、555-556、558-559、671、682-683、743、745-750、7631-7683、7796-7807、7812-7813、7815-7816、7928、7939-7940、800、8002-8007、5942-6665、13189-13911、1-297、715、1580-1603、7258-7554、7972、8834-8857、705-714、7962-7971、6681-6694、13927-13940、6788-6803、14033-14048、1469-1526、5147-5151、8725-8781、12396-12400、2159-2428、9413-9682、646-648、7903-7905、2592-2595、9846-9849、676-678、717-720、7933-7935、7974-7977、538、669、704、7795、7926、7961、8710、670、699-701、7927、7956-7958、4917-4979、12167-12228、4910-4916、12160-12166、5195-5941、12444-13188、627-645、650-659、693-694、744、768-769、1451、1455、1467-1468、1527-1528、1570、1578、1579、1604-1611、2429-2444、2583-2591、2596-2597、2867-2886、4876-4883、5144-5146、5152-5176、5193-5194、6666-6680、6695-6696, 6702-6787, 6804-6822, 6973, 7050-7053, 7134-7257, 7884-7902, 7907-7916, 7950-7951, 8001, 8025-8026, 8707, 8711, 8723-8724, 8782-8783, 8824, 8832-8833, 8858-8865, 9683-9698, 9837-9845, 9850, 10120-10139, 12126 -12133, 12393-12395, 12401-12425, 12442-12443, 13912-13926, 13941-13942, 13948-14032, 14049-14067, 14216, 14292-14295, 14375-14498, 16886-17078, 17478-17622, 17677-17756, 14831-14833, 14838, 14847, 14850-15460, 17079 -17477, 17660-17676, 19031-19080, 16414-16516, 17623-17659, 19081-19108, 16397-16413, 16195-16320, 15779-15925, 15476-15778, 16321-16366, 16367-16396, 16705-16814, 16815-16885, 16517-16704, 18949-19030, 14657-14716 14778-14824, 14834, 14835-14836, 14839, 15461-15475, 14717-14777, 14841-14846, 18413-18936, 14499-14656, 18939, 15926-16178, 14837, 17757-18412, 14825-14830, 14840, 14848-14849, 16179-16194, 18937-18938 and 18940-18948 (Form A or Form B of U.S. Application No. 18 / 087,673 or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or Form 31A of PCT Publication WO2024044723A1; SEQ ID NO: 19543-19733 of the PCT).
[0068] In any of the foregoing embodiments, the ncRNA variant comprises a table A or B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or Table 31A of PCT Publication WO2024044723A1 (SEQ ID NO: of the PCT). A sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity.
[0069] In any of the foregoing embodiments, the ncRNA variant comprises a sequence selected from the following: RTX3_6342_msr_stem_var1 (SEQ ID NO: 19644), RTX3_6342_msr_stem_var20 (SEQ ID NO: 19663), RTX3_6342_msr_stem_var5 (SEQ ID NO: 19648), RTX3_6342_a1a2_var6 (SEQ ID NO: 19548), RTX3_6342_a1a2_var10 (SEQ ID NO: 19552), RTX3_6342_a1a2_var15 (SEQ ID NO: 19557), RTX3_6342_a1a2_var16 (SEQ ID NO: 19558), RTX3_6342_a1a2_var19 ... 19561), RTX3_6342_a1a2_var20 (SEQ ID NO: 19562), RTX3_6342_a1a2_var21 (SEQ ID NO: 19563), RTX3_6342_a1a2_var26 (SEQ ID NO: 19568), RTX3_6342_a1a2_var27 (SEQ ID NO: 19569) and RTX3_6342_a1a2_var32 (SEQ ID NO: 19574).In any of the foregoing embodiments, the ncRNA variant comprises, compared to RTX3_6342_WT (SEQ ID NO: 19734), RTX3_6342_msr_stem_var1 (SEQ ID NO: 19644), RTX3_6342_msr_stem_var20 (SEQ ID NO: 19663), RTX3_6342_msr_stem_var5 (SEQ ID NO: 19648), RTX3_6342_a1a2_var6 (SEQ ID NO: 19548), RTX3_6342_a1a2_var10 (SEQ ID NO: 19552), RTX3_6342_a1a2_var15 (SEQ ID NO: 19557), and RTX3_6342_a1a2_var16 (SEQ ID NO: 19557). One or more of the following modifications: RTX3_6342_a1a2_var19 (SEQ ID NO: 19558), RTX3_6342_a1a2_var20 (SEQ ID NO: 19562), RTX3_6342_a1a2_var21 (SEQ ID NO: 19563), RTX3_6342_a1a2_var26 (SEQ ID NO: 19568), RTX3_6342_a1a2_var27 (SEQ ID NO: 19569), or RTX3_6342_a1a2_var32 (SEQ ID NO: 19574).
[0070] In any of the foregoing embodiments, the ncRNA variant comprises a sequence selected from the following: msR spacer Del-1 (SEQ ID NO: 19939), msD spacer Del-1 (SEQ ID NO: 19940), msD spacer Del-2 (SEQ ID NO: 19941), and msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19942). In any of the foregoing embodiments, the ncRNA variant comprises one or more modifications of msR spacer Del-1 (SEQ ID NO: 19939), msD spacer Del-1 (SEQ ID NO: 19940), msD spacer Del-2 (SEQ ID NO: 19941), or msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19942) compared to WT R6342ncRNA (SEQ ID NO: 19734).
[0071] In any of the foregoing embodiments, the ncRNA variant comprises a sequence selected from the following: Alt1 msR spacer Del- (SEQ ID NO: 19947)1, Alt1 msD spacer Del-1 (SEQ ID NO: 19948), Alt1 msD spacer Del-2 (SEQ ID NO: 19949), Alt1 msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19950), Alt1 msR spacer Del-2 / msD Del-3 (SEQ ID NO: 19951), Alt1 msR spacer Del-2 / msD Del-3 MS2 (SEQ ID NO: 19952), and Alt1 msR spacer Del-2 / msDDel-3 U7 (SEQ ID NO: 19953). In any of the foregoing embodiments, the ncRNA variant comprises one or more modifications of Alt1 msR spacer Del- (SEQ ID NO: 19947)1, Alt1 msD spacer Del-1 (SEQ ID NO: 19948), Alt1 msD spacer Del-2 (SEQ ID NO: 19949), Alt1 msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19950), Alt1 msR spacer Del-2 / msD Del-3 (SEQ ID NO: 19951), Alt1 msR spacer Del-2 / msD Del-3 MS2 (SEQ ID NO: 19952), or Alt1 msR spacer Del-2 / msD Del-3 U7 (SEQ ID NO: 19953) compared to WTR6342 ncRNA (SEQ ID NO: 19734).
[0072] In any of the foregoing embodiments, the ncRNA variant comprises a sequence selected from the following: msR spacer del_30HA (SEQ ID NO: 19955), msR spacer del_45HA (SEQ ID NO: 19956), WT PAM mt, (SEQ ID NO: 19957) Alt del1+del2 PAM mt (SEQ ID NO: 19958) or Alt del1+del2+U7 PAM mt (SEQ ID NO: 19959), and Alt del1+del2 25bp ins (SEQ ID NO: 19960). In any of the foregoing embodiments, the ncRNA variant comprises one or more modifications of the msR spacer del_30HA (SEQ ID NO: 19955), msR spacer del_45HA (SEQ ID NO: 19956), WT PAM mt, Alt del1+del2PAM mt (SEQ ID NO: 19957), Alt del1+del2+U7 PAM mt (SEQ ID NO: 19959), or Alt del1+del2 25bp ins (SEQ ID NO: 19960) compared to WTR6342 ncRNA (SEQ ID NO: 19734).
[0073] In one embodiment of any of the foregoing embodiments, the one or more modifications further comprise the addition of a 2'O methyl or thiophosphate bond at the 5' or 3' end of the ncRNA variant. In one embodiment, the one or more modifications further comprise the addition of a 2'O methyl or thiophosphate bond at both the 5' and 3' ends of the ncRNA variant.
[0074] In another aspect, this disclosure provides a reverse transcriptase variant comprising an ncRNA variant of any of the foregoing embodiments, and an RNA encoding a first polymerase, optionally downstream of the msd of the ncRNA variant.
[0075] In one embodiment, the first polymerase is a reverse transcriptase, optionally wherein the reverse transcriptase is derived from a naturally occurring reverse transcripton or reverse transcripton-like sequence.
[0076] In any of the foregoing embodiments, the first polymerase is a reverse transcriptase selected from EcoI-RT, Efe1-RT, Mva1-RT, Cex1-RT, Eco8-RT, Vap1-RT, and Vro1-RT. In one embodiment, the reverse transcriptase is derived from the same naturally occurring reverse transcriptase as the reference reverse transcriptase ncRNA.
[0077] In another aspect, this disclosure provides a chimeric gene editing composition comprising: (a) an ncRNA variant of any of the foregoing embodiments, or a reverse transcriptase variant of any of the foregoing embodiments, (b) a nuclease or a first mRNA encoding the nuclease, (c) a guide RNA (gRNA) associated with the nuclease, and (d) optionally, a second polymerase or a second mRNA encoding the second polymerase.
[0078] In one embodiment, the chimeric gene editing composition comprises a second polymerase or a second mRNA encoding the second polymerase, wherein the second polymerase is a reverse transcriptase. In another embodiment, the chimeric gene editing composition comprises a second polymerase or a second mRNA encoding the second polymerase, wherein the second polymerase is a DNA polymerase.
[0079] In any of the foregoing embodiments, the nuclease is a nicking enzyme. In one embodiment, the nicking enzyme comprises Cas9 nuclease, Cas9(D10A) nuclease, TnpB nuclease, or Cas12a nuclease.
[0080] In any of the foregoing embodiments, the ncRNA variant or the reverse transcriptase variant is directly or indirectly linked to the gRNA. In one embodiment, the ncRNA variant or the reverse transcriptase variant, the gRNA, and the first mRNA encoding the nuclease are directly or indirectly linked.
[0081] In any of the foregoing embodiments, the chimeric gene editing composition further comprises a delivery medium. In one embodiment, the delivery medium encapsulates one or more components selected from: (a) an ncRNA variant or a reverse transcriptase variant, (b) a nuclease or a first mRNA encoding the nuclease, (c) a gRNA associated with the nuclease, and (d) optionally, a second polymerase or a second mRNA encoding the second polymerase. In one embodiment, the delivery medium is a lipid nanoparticle.
[0082] In another aspect, this disclosure provides one or more polynucleotides comprising the coding sequence of an ncRNA variant or a reverse transcriptase variant of any of the foregoing embodiments.
[0083] In one embodiment, the one or more polynucleotides further comprise a coding sequence for a nuclease or a coding sequence for gRNA.
[0084] In one embodiment, the one or more polynucleotides further comprise a nuclease coding sequence and a gRNA coding sequence.
[0085] In any of the foregoing embodiments, the one or more polynucleotides further comprise the coding sequence of a first polymerase.
[0086] In any of the foregoing embodiments, the one or more polynucleotides further comprise the coding sequence for a second polymerase.
[0087] In any of the foregoing embodiments, the one or more polynucleotides further comprise one or more promoters, each of which is operatively linked to: (i) a coding sequence of an ncRNA variant or a reverse transcriptase variant, (ii) a coding sequence of a nuclease, (iii) a coding sequence of a gRNA, (iv) a coding sequence of a first polymerase, or (v) a coding sequence of a second polymerase.
[0088] In another aspect, this disclosure provides a vector comprising one or more polynucleotides from any of the foregoing embodiments.
[0089] In one embodiment, the vector is a plasmid or a viral vector, optionally wherein the viral vector is an AAV or a lentiviral vector.
[0090] In another aspect, this disclosure provides a method for editing target DNA, the method comprising interacting the target DNA with a chimeric gene editing composition of any of the foregoing embodiments, one or more polynucleotides of any of the foregoing embodiments, or a vector of any of the foregoing embodiments.
[0091] In one embodiment, the chimeric gene editing composition, the one or more polynucleotides, or the ncRNA variant in the vector contains a pair of homologous arms that are specifically targeted at the target locus and specifically targeted at the target DNA.
[0092] In any of the foregoing embodiments, the interaction step is performed in vivo or in vitro.
[0093] In another aspect, this disclosure provides a chimeric gene editing composition comprising: (a) a non-coding RNA (ncRNA) containing a single-stranded RNA with a polymerase template, (b) a nicking enzyme or a first mRNA encoding the nicking enzyme, (c) a guide RNA (gRNA) associated with the nicking enzyme, and (d) optionally, a second polymerase or a second mRNA encoding the second polymerase.
[0094] In one embodiment, the chimeric gene editing composition comprises a second polymerase or a second mRNA encoding the second polymerase, wherein the second polymerase is a reverse transcriptase.
[0095] In one embodiment, the chimeric gene editing composition comprises a second polymerase or a second mRNA encoding the second polymerase, wherein the second polymerase is a DNA polymerase.
[0096] In any of the foregoing embodiments, the nicking enzyme comprises Cas9 nuclease, Cas9(D10A) nuclease, TnpB nuclease, or Cas12a nuclease.
[0097] In any of the foregoing embodiments, the ncRNA is directly or indirectly linked to the gRNA. In one embodiment, the ncRNA, the gRNA, and the first mRNA encoding the nickase are directly or indirectly linked. Attached Figure Description
[0098] The following figures form part of this specification and are included to further illustrate certain aspects of this disclosure. A better understanding of this disclosure can be achieved by referring to one or more of these figures and the detailed description of the specific embodiments presented herein.
[0099] Figure 1 Describe the pilot editing system.
[0100] Figure 2 Describe the reverse transcriptase editing system.
[0101] Figure 3 This paper describes the chimeric editing system disclosed herein, namely a precise editing system for nickase reverse transcriptase templates that includes one or more lead editing components and one or more reverse transcriptase editor components.
[0102] Figure 4 A to Figure 4 D compares the structures of wild-type ncRNAs (A and B) with those of templated ncRNAs (C) and homologous templated msDNAs (D) used in the gene editing system described herein.
[0103] Figure 5The procedure for editing using the gene editing system described herein is illustrated, which utilizes ncRNA and reverse transcriptase RT to generate msDNA for use as a polymerase template to edit target DNA.
[0104] Figure 6 The process of editing using the gene editing system described herein is illustrated, which utilizes ncRNA and the reverse transcriptase RT to use the ncRNA as a polymerase template to edit target DNA.
[0105] Figures 7A to 7D The editing mechanism of the gene editing system described in this paper is illustrated.
[0106] Figures 8A to 8C Various hypothetical configurations of the guide RNA and ncRNA components of the gene editing system described in this paper are shown.
[0107] Figure 8D This illustrates the various modifications that can be introduced into the templated ncRNAs described in this paper.
[0108] Figure 8E Various hypothetical configurations of the guide RNA and ncRNA components of the gene editing system described in this paper are shown.
[0109] Figure 9 A dual-editor implementation scheme for the gene editing system described herein is shown.
[0110] Figures 10A to 10B A dual-editor implementation scheme for the gene editing system described herein is shown.
[0111] Figure 11 This is a schematic diagram illustrating various modifications to the gene editing system described in this article.
[0112] Figure 12 This is a schematic diagram of an LNP formulation containing RNA cargo for delivering the gene editing system of this disclosure to cells, tissues, or patients under in vitro, in vivo, or ex vivo conditions.
[0113] Figure 13 This is a scatter plot from pooling screening assays, in which the a1:a2 variant was identified, showing superior performance compared to the WT reverse transcriptase R6342.
[0114] Figure 14 It is a scatter plot, which displays... Figure 13 The various a1a2 variants identified in the pooling screening assay were plotted as the relationship between ssDNA production and melting temperature.
[0115] Figure 15This is a graph summarizing the edit percentages determined by exact inserts, inserts with inserts / deletions, and inserts with SNPs for WT R6342, variants lacking both msr spacers and stems (Del1), variants lacking msd spacers (Del2), and variants lacking both (Del1+Del2).
[0116] Figure 16 This is a set of charts and graphs showing the impact of MSR stem deletion on editing efficiency. Overall, MSR stem deletion leads to almost complete loss of precise gene editing.
[0117] Figure 17 This is a set of charts and graphs showing the effect of altering the a1:a2 stem on editing efficiency. These results indicate that, for ncRNAs, the a1:a2 stem structure, rather than the sequence, is a key feature of the edited output.
[0118] Figure 18 This is a set of charts and graphs showing the impact of deleting msD and msR spacers on editing efficiency. Deleting only the msR spacer does not affect gene editing results, while the msD deletion variant exhibits lower editing efficiency, partly due to reduced insertion / deletion rates in these cell lines. When combined, simultaneous deletion of both msR and msD spacers results in robust editing results comparable to WT, indicating that ncRNAs can be shortened in the msD and msR spacer regions without impairing their function.
[0119] Figure 19 This is a set of charts and graphs showing the impact of fusion-type a1a2 ncRNA configurations on editing efficiency. Fusion-type a1a2 designs (ALT) with both long and short R6342 retrotranscript ncRNAs (R6342L and R6342S, respectively) were generated. Alternative designs with fusion-type a1a2 loops did not affect the editing efficiency of R6342L ncRNA but slightly reduced the editing efficiency of R6342S ncRNA.
[0120] Figure 20 This is a set of graphs and figures showing the effect of deleting the msD and msR spacers on editing efficiency with fusion-type a1a2 ncRNA configurations. Compared to WT, deleting spacer sequences generally reduces editing efficiency, but overall the modified constructs still exhibit considerable activity.
[0121] Figure 21This is a set of charts and graphs showing the effects of msD stem-loop deletion combined with msD and msR spacers. ncRNAs with msD stem-loop deletion, further shortening of the msD spacer, and complete removal of the msR spacer with a fusion a1a2 stem (Alt1 msR spacer Del-2 / msD Del-3) have only a small negative impact on gene editing. Incorporation of an MS2 stem-loop at the 3' end of the ncRNA significantly reduces the editing efficiency of this ncRNA (Alt1 msR spacer Del-2 / msD Del-3 MS2), while a 3' tail from snRNA U7 maintains similar editing efficiency (Alt1 msR spacer Del-2 / msD Del-3 MS2). In summary, this indicates a high level of editing with the deletion of both the msD spacer and stem-loop elements, except for the msR spacer element.
[0122] Figure 22 This is a set of charts and graphs showing that the U7 snRNA sequence strongly promotes ssDNA production in short ncRNAs. ssDNA production of short ncRNAs with missing msD and msR stem-loops was measured. It was found that such ncRNA variants resulted in a 6-fold reduction in ssDNA production. Incorporation of the MS2 stem-loop (Alt1 msR spacer Del-2 / msD Del-3 MS2) promoted ssDNA production, while incorporation of the U7 sequence (Alt1 msR spacer Del-2 / msD Del-3 U7) showed a greater enhancement of ssDNA production.
[0123] Figure 23 This is a set of charts showing that the minimized variant ncRNA with msr spacer deletion allows up to 20% editing for the M(ATG)>T(ACC) mutation installation in the UBA1 gene in human hematopoietic stem cells (HSCs), which is 2.5 times more efficient than the WT version.
[0124] Figure 24 It is a set of charts that show sub-missing MSR / MSD intervals and alternative configurations (such as...). Figure 19 The ncRNA presented allows for up to 40% editing of the PAM mutation installation or 25 bp insertion in the AAVS1 gene in human hematopoietic stem cells (HSCs), which is 2.5 times more efficient than the WT version.
[0125] definition All technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this disclosure pertains. The following references provide general definitions for many of the terms used in this disclosure for those skilled in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale and Margham, The Harper Collins Dictionary of Biology (1991).
[0126] In some instances, general methods in molecular and cellular biochemistry can be found in, for example, the following standard textbooks: *Molecular Cloning: A Laboratory Manual*, 3rd edition (Sambrook et al., HaRBor Laboratory Press 2001); *Short Protocols in Molecular Biology*, 4th edition (Ausubel et al., John Wiley & Sons 1999); *Protein Methods* (Bollag et al., John Wiley & Sons 1996); *Nonviral Vectors for Gene Therapy* (Wagner et al., Academic Press 1999); *Viral Vectors* (Kaplift and Loewy, Academic Press 1995); *Immunology Methods Manual* (I. Lefkovits, Academic Press 1997); and *Cell and Tissue Culture: Laboratory Procedures in Biotechnology* (Doyle and Griffiths, John Wiley & Sons 1998), the contents of which are incorporated herein by reference.
[0127] When numerical ranges are provided, it should be understood that, unless the context explicitly specifies otherwise, every intermediate value (accurate to one-tenth of the lower limit unit) between the upper and lower limits of the range and between any other stated or intermediate values within the stated range is included in this disclosure. The upper and lower limits of these smaller ranges may be independently included within the smaller range and are also included in this disclosure, but are subject to any explicitly excluded limits in the stated range. When a stated range includes one or two limits, ranges excluding those including either or both limits are also included in this disclosure.
[0128] It must be noted that, unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “described” as used herein and in the appended claims include a plural of the referred to. Thus, for example, a reference to “ncRNA” includes multiple ncRNAs, and a reference to “reverse transcriptase” includes a reference to one or more RTs and their equivalents known to those skilled in the art, and so on. Furthermore, it should be noted that the claims may be designed to exclude any optional elements. Therefore, this statement is intended as a priori basis for using exclusive terms such as “solely,” “only,” or using negative limitations in the recitation of claim elements. For example, the claims may be designed to exclude certain RT sequences.
[0129] It should be understood that certain features of this disclosure, for clarity, are described in the context of individual embodiments, or may be provided in combination in a single embodiment. Conversely, various features of this disclosure, for brevity, are described in the context of a single embodiment, or may be provided individually or in any suitable sub-combination. This disclosure expressly covers all combinations of embodiments relating to this disclosure, and is disclosed herein as if each combination were disclosed separately and expressly. Furthermore, all sub-combinations of multiple embodiments and their elements are also expressly included within this disclosure, and are disclosed herein as if each such sub-combination were disclosed separately and expressly.
[0130] Bioactivity As used herein, the term "bioactivity" refers to the characteristic of an agent (e.g., DNA, RNA, or protein) that is active in a biological system (including in vitro and in vivo biological systems), particularly in living organisms (e.g., in mammals, including human and non-human mammals). For example, an agent that has a biological effect on an organism when applied to it is considered to be bioactive.
[0131] protrusion As used herein, the term "bump" refers to a small region of unpaired bases in the "stem" of a nucleotide that interrupts base pairing. A bump may consist of one or two single-stranded or unpaired nucleotides joined at both ends by base-paired nucleotides of the stem. A bump may be symmetrical (i.e., the two unpaired single-stranded regions have the same number of nucleotides), asymmetrical (i.e., the unpaired single-stranded regions have different or unequal numbers of nucleotides), or may contain only one unpaired nucleotide on one strand. A bump may be described as A / B (e.g., "2 / 2 bump" or "1 / 0 bump"), where A represents the number of unpaired nucleotides on the upstream strand of the stem, and B represents the number of unpaired nucleotides on the downstream strand of the stem. In the primary nucleotide sequence, the upstream strand of the bump is closer to 5′ than the downstream strand.
[0132] cDNA As used herein, the term “cDNA” refers to a strand of DNA copied from an RNA template, for example, by reverse transcriptase.
[0133] Homogeneous The term "homologous" refers to two biomolecules that typically interact or coexist in nature.
[0134] Complementary As used herein, the term “complementary” or “substantially complementary” means a nucleic acid (e.g., RNA, DNA) that contains a nucleotide sequence that enables it to “gather” or “hybridize” with another nucleic acid in a sequence-specific, antiparallel (i.e., specific binding of the nucleic acid to the complementary nucleic acid) manner (i.e., forming Watson-Crick base pairs and / or G / U base pairs) under appropriate in vitro and / or in vivo temperature and solution ionic strength conditions. Standard Watson-Crick base pairings include: adenine (A) paired with thymidine (T), adenine (A) paired with uracil (U), and guanine (G) paired with cytosine (C) [DNA, RNA]. Additionally, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization between DNA and RNA molecules (e.g., when the target DNA nucleic acid pairs with a guide RNA base, etc.): guanine (G) may also pair with uracil (U) bases. For example, in the case of tRNA anticodon base pairing with codon bases in mRNA, G / U base pairing is at least partially responsible for the degeneracy (i.e., redundancy) of the genetic code. Therefore, in the context of this disclosure, guanine (G) is considered complementary to uracil (U) and adenine (A). For example, when a G / U base pair can be generated at a given nucleotide position in the dsRNA duplex that guides the RNA molecule, the position is not considered non-complementary but complementary.
[0135] It should be understood that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid for specific hybridization or hybridization. Furthermore, the polynucleotide may hybridize on one or more segments such that intercalated or adjacent segments do not participate in the hybridization event (e.g., protrusions, loops, or hairpins). The polynucleotide may contain 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity with the target region within the target nucleic acid sequence to which it will hybridize. For example, an antisense nucleic acid in which 18 out of 20 nucleotides of an antisense compound are complementary to the target region and will therefore specifically hybridize would represent 90% complementarity. In this example, the remaining non-complementary nucleotides may cluster or be scattered with complementary nucleotides and do not need to be adjacent to each other or adjacent to complementary nucleotides. The percentage of complementarity between specific nucleic acid sequence fragments within a nucleic acid can be determined using any suitable method. Examples of methods include basic local alignment search tools (BLAST) and PowerBLAST programs (e.g., those described in Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), such as using the default Gap program (Wisconsin Sequence Analysis Package, Unix version 8, Genetics Computer Group, University Research Park, Madison Wis.), which uses the Smith and Waterman algorithm as described in Adv. Appl. Math., 1981, 2, 482-489, etc.
[0136] DNA-directed nucleases As used herein, “DNA-guided nuclease” refers to a type of “programmable nuclease” and a specific type of “nucleic acid-guided nuclease.” Examples of DNA-guided nucleases are reported in Varshney et al., DNA-guided genome editing using structure-guided endonucleases, Genome Biology, 2016, 17(1), 187, which is used in the context of this disclosure and is incorporated herein by reference. As used herein, the term “DNA-guided nuclease” or “DNA-guided endonuclease” refers to a nuclease that covalently or non-covalently associates with a guide RNA to form a complex between the guide RNA and the DNA-guided nuclease. The guide RNA contains a spacer sequence that contains a nucleotide sequence complementary to the target DNA strand. Thus, the DNA-guided nuclease indirectly guides or programs itself to locate a specific site in a DNA molecule through its association with the guide RNA, which binds or adheres directly to the target DNA strand via Watson-Crick base pairing through its complementary region.
[0137] DNA regulatory sequences As used herein, the terms “DNA regulatory sequence,” “control element,” and “regulatory element” are used interchangeably and refer to transcriptional and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, protein degradation signals, etc., which provide and / or regulate the transcription of non-coding sequences (e.g., guide RNA) or coding sequences and / or regulate the translation of mRNA into encoded polypeptides.
[0138] donor nucleic acid "Donor nucleic acid" or "donor polynucleotide" or "donor DNA" or "HDR donor DNA" refers to single-stranded DNA to be inserted at a site to be cleaved by a programmable nuclease (e.g., CRISPR / Cas effector protein; TALEN; ZFN; wide range of nucleases) (e.g., after dsDNA cleavage, after nicking the target DNA, after double nicking the target DNA, etc.). The donor polynucleotide must have sufficient homology with the genomic sequence at the target site, for example, with the target site within approximately 200 bases or less, approximately 190 bases or less, approximately 180 bases or less, approximately 170 bases or less, approximately 160 bases or less, approximately 150 bases or less, approximately 140 bases or less, approximately 130 bases or less, approximately 120 bases or less, or approximately 110 bases. The nucleotide sequence (within 70%, 80%, 85%, 90%, 95%, or 100% homology, such as within about 100 or fewer bases of the target site, or within about 90 or fewer bases of the target site, or within about 80 or fewer bases of the target site, or within about 70 or fewer bases of the target site, or within about 60 or fewer bases of the target site, or within about 50 or fewer bases of the target site, or within about 30 bases, about 15 bases, about 10 bases, about 5 bases, or closely adjacent to the target site) supports homology-directed repair between the donor polynucleotide and its homologous genomic sequence.
[0139] coding As used in this article, the DNA sequence “encoding” a specific RNA is the DNA nucleotide sequence transcribed into RNA. DNA polynucleotides can encode RNA (mRNA) that is translated into protein (and therefore both DNA and mRNA encode proteins), or DNA polynucleotides can encode RNA that is not translated into protein (e.g., tRNA, rRNA, microRNA (miRNA), “non-coding” RNA (ncRNA), guide RNA, etc.). In the case of reverse transcripts, the reverse transcript DNA can encode the ncRNA locus (which includes the msr and msd regions) and the reverse transcript RT.
[0140] engineered reverse transcriptase As used herein, the term “engineered retrotran” or equivalently “recombinant retrotran” or “retrotran variant” refers to a retrotran that does not exist in nature. In one embodiment, the engineered reverse transcript may include a wild-type or naturally occurring reverse transcript modified to contain at least one modification, including a single nucleotide substitution, insertion, or deletion, or more than one nucleotide substitution, insertion, or deletion, i.e., up to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or up to 100, or up to 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or up to 2000 nucleotides are replaced, inserted, or deleted from the starting reverse transcript (e.g., wild-type reverse transcript). When more than one nucleotide of the start-point reverse transcript (e.g., wild-type reverse transcript) is substituted, inserted, or deleted, the nucleotides can be continuous or discontinuous. While the engineered reverse transcript as a whole is not naturally occurring, it can include components that are indeed present in nature, such as nucleotide sequences. For example, an engineered reverse transcript may have nucleotide sequences derived from different organisms (e.g., different bacterial species) or from fully synthetic / artificial / recombinant nucleic acid sequences. Thus, an engineered reverse transcript may have bacterial nucleotide sequences, human nucleotide sequences, viral nucleotide sequences, and / or synthetic / artificial / recombinant nucleotide sequences, and / or combinations of such sequences. Examples of modifications to the recombinant reverse transcripts disclosed herein include the insertion of heterologous nucleic acid sequences into the reverse transcript, for example, insertion into ncRNA loci, such as in the MSR or MSD loci. Linking a guide RNA molecule to the 5′ end and / or the 3′ end (i.e., one linked to the 5′ end of the ncRNA and / or one linked to the 3′ end of the ncRNA) also represents another modification contemplated for the recombinant reverse transcripts disclosed herein. In such implementations, the guide RNA molecule can also be classified as, or more broadly referred to as, a type of heterologous nucleic acid sequence used to modify the start site reverse transcriptase.
[0141] exosomes As used herein, the term "exosome" refers to a small membrane-bound vesicle with endocytic origin. Without being bound by theory, exosomes are typically released from the host / progenitor cell into the extracellular environment after fusion of a multivesicular body with the cytoplasmic membrane. Therefore, in addition to engineered components (e.g., engineered reverse transcriptases), exosomes may also include components of the progenitor cell membrane. Exosome membranes are typically layered, consisting of a lipid bilayer with spaces between aqueous nanoparticles.
[0142] expression carrier As used herein, the term "expression vector" or "expression construct" refers to a vector that includes one or more expression control sequences, and an "expression control sequence" is a DNA sequence that controls and regulates the transcription and / or translation of another DNA sequence. Suitable expression vectors include, but are not limited to, plasmids and viral vectors derived from, for example, bacteriophages, baculoviruses, tobacco mosaic viruses, herpesviruses, cytomegaloviruses, retroviruses, vaccinia viruses, adenoviruses, and adeno-associated viruses. Many vectors and expression systems are available from, for example, Novagen (Madison, WI), Clontech (Palo Alto, CA), Stratagene (LaJolla, CA), and Invitrogen / Life Technologies (Carlsbad, CA). This disclosure covers recombinant vectors, which may include viral vectors, bacterial vectors, protozoan vectors, DNA vectors, or recombinants thereof.
[0143] Guide RNA An RNA molecule that binds to a programmable nuclease in a reverse transcriptase-based gene editing system and targets the nuclease to a specific location within a targeted polynucleotide sequence is referred to herein as a "guide RNA" or "guide RNA polynucleotide" (also referred to herein as "guide RNA," "gRNA," or "crRNA"). In some embodiments (depending on the specific nuclease it interacts with), the guide RNA comprises two segments: a "DNA-targeting segment" and a "protein-binding segment." A "segment" refers to a segment / part / region of a molecule, such as an adjacent fragment of nucleotides in RNA. As an illustrative and non-limiting example, the protein-binding segment of the guide RNA may comprise base pairs 5-20 of an RNA molecule of 40 base pairs in length; and the DNA-targeting segment may comprise base pairs 21-40 of an RNA molecule of 40 base pairs in length. Unless otherwise specifically defined in a particular context, the definition of a “segment” is not limited to a specific number of total base pairs, not limited to any specific number of base pairs from a given RNA molecule, not limited to a specific number of independent molecules within a complex, and may include RNA molecule regions of any total length and may include or exclude regions complementary to other molecules.
[0144] The DNA targeting region (or “DNA targeting sequence”) contains a nucleotide sequence complementary to a specific sequence within the targeted polynucleotide sequence (the complementary strand of the targeted polynucleotide sequence), referred to herein as a “protospacer-like” sequence. The protein-binding region (or “protein-binding sequence”) interacts with the site-specific modifying polypeptide. When the site-specific modifying polypeptide is a CRISPR nuclease, site-specific cleavage of the targeted polynucleotide sequence can occur at a location determined by two factors: (i) base pairing complementarity between the guide RNA and the targeted polynucleotide sequence; and (ii) a short motif in the targeted polynucleotide sequence (called a protospacer adjacent motif (PAM)).
[0145] Heterologous nucleic acid sequences As used herein, the term "heteronucleotide" refers to an entity with a genotype that is different from that of the entity to which it is compared or which is introduced or incorporated. For example, a polynucleotide introduced into a different cell type via genetic engineering is a heteropolynucleotide (e.g., DNA or RNA) and, if expressed, may encode a heteropolypeptide. Similarly, a cellular sequence (e.g., a gene or a portion thereof) incorporated into a viral vector is a heteronucleotide sequence relative to the vector. In some embodiments, the heterologous sequence inserted into a wild-type retrotran region does not naturally insert into such a region (e.g., an engineered retrotran with the inserted heterologous sequence does not exist in nature). For example, the heterologous sequence may be from the same bacterial species in which wild-type retrotran are normally present, provided that the heterologous sequence does not naturally insert into the location in the wild-type retrotran where the heterologous sequence was inserted. In some embodiments, the heterologous sequence is a mammalian sequence (e.g., a human sequence) or its reverse complementary sequence. The heterologous nucleic acid sequence introduced into the retrotran may include, but is not limited to, a guide RNA sequence, an HDR donor template, a gene encoding a protein, or a non-coding functional RNA element (e.g., a stem-loop, hairpin, or protrusion).
[0146] Homologous targeted repair As used in this article, "homology-directed repair (HDR)" refers to a specific form of DNA repair that occurs, for example, during double-strand break repair in cells. This process requires nucleotide sequence homology, uses a "donor" molecule to perform template repair on a "target" molecule (i.e., the molecule undergoing a double-strand break), and results in the transfer of genetic information from the donor to the target. If the donor polynucleotide differs from the target molecule and part or all of the donor polynucleotide's sequence is incorporated into the target polynucleotide sequence, homology-directed repair can lead to changes in the target molecule's sequence (e.g., insertions, deletions, mutations).
[0147] same As used herein, the term "identical" refers to two or more identical sequences or subsequences. Additionally, the term "substantially identical" as used herein refers to two or more sequences having the same percentage of consecutive units when compared and aligned over a comparison window or specified region using a comparison algorithm or by manual alignment and visual inspection to obtain maximum correspondence. For example only, two or more sequences may be considered "substantially identical" if consecutive units are approximately 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% identical within a specified region. Such percentages describe the "percentage of identity" of two or more sequences. Sequence identity may exist within a region of at least approximately 75-100 consecutive units, within a region of approximately 50 consecutive units, or, if not specified, throughout the entire sequence. This definition also refers to complementary sequences of the test sequences.
[0148] Alternatively, when a nucleic acid or a fragment thereof hybridizes with another nucleic acid, another nucleic acid strand, or its complementary strand under stringent hybridization conditions, there is a substantial identity or similarity. In the context of nucleic acid hybridization experiments, "stringent hybridization conditions" and "stringent washing conditions" depend on many different physical parameters. Nucleic acid hybridization will be affected by conditions such as salt concentration, temperature, solvent, base composition of the hybrid species, length of the complementary region, and the number of nucleotide base mismatches between the hybrid nucleic acids, as readily understood by those skilled in the art. Those skilled in the art know how to modify these parameters to achieve a specific degree of hybridization stringency.
[0149] Lipid nanoparticles (LNP) As used herein, the term "lipid nanoparticle" or LNP refers to a class of lipid particle delivery systems formed from small solid or semi-solid particles, having an outer lipid layer with a hydrophilic outer surface exposed to a non-LNP environment, an internal space that may be aqueous (vesicle-like) or non-aqueous (microcellular-like), and at least one hydrophobic intermembranous space. The LNP membrane may be layered or non-layered and may contain 1, 2, 3, 4, 5, or more layers. In some embodiments, the LNP may contain nucleic acids (e.g., engineered reverse transcriptases) entering its internal space, entering its intermembranous space, entering its outer surface, or any combination thereof. In some embodiments, the LNP of this disclosure comprises ionizable lipids, structural lipids, polyethylene glycol-modified lipids (also known as PEG lipids), and phospholipids. In alternative embodiments, the LNP comprises ionizable lipids, structural lipids, polyethylene glycol-modified lipids (also known as PEG lipids), and zwitterionic amino acid lipids.
[0150] Further discussion on liposomes can be found, for example, in Tenchov et al., “Lipid Nanoparticles – From Liposomes to mRNA Vaccine Delivery, a Landscape of Diversity and Advancement”. ACS Nano , 2021, 15, pp. 16982-17015 (the contents of which are incorporated by reference).
[0151] connector As used herein, the term "connector" refers to a molecule that links or joins two other molecules or parts. In cases where a connector joins two fusion proteins, the connector can be an amino acid sequence. For example, an RNA-directed nuclease (e.g., Cas12a) can be fused to a reverse transcriptase via an amino acid connector sequence. In cases where two nucleotide sequences are joined together, the connector can also be a nucleotide sequence. For example, in this case, ncRNA can be linked to one or more guide RNAs via a nucleotide sequence connector at its 5′ and / or 3′ ends. In other embodiments, the connector is an organic molecule, group, polymer, or chemical moiety. In some implementations, the linker length is 5-100 amino acids, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter linkers have also been considered.
[0152] Liposomes As used herein, the term "liposome" refers to a class of lipid particle delivery systems comprising small vesicles containing at least one lipid membrane surrounding the interior space of an aqueous nanoparticle that is not typically derived from progenitor / host cells. Liposomes are versatile carrier platforms capable of transporting hydrophobic or hydrophilic molecules, including small molecules, proteins, and nucleic acids, into cells. They are among the first generation of nanoscale drug delivery platforms developed. Numerous liposomal drug formulations have been approved for use in human medicine, such as Doxil, a lipid nanoparticle formulation of the antitumor drug doxorubicin. In some instances, liposomes may include those described in the following literature: Tenchov et al., “LipidNanoparticles – From Liposomes to mRNA Vaccine Delivery, a Landscape of Diversity and Advancement”, ACS Nano, 2021, 15, pp. 16982-17015 (the contents of which are incorporated herein by reference).
[0153] ring As used herein, the term "loop" in polynucleotide refers to a single-stranded segment of one or more nucleotides, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides, wherein the 5' and 3' nucleotides of the loop are each linked to a base-paired nucleotide in the stem.
[0154] microcells As used in this article, the term "microcell" refers to a small particle that does not have an internal space within an aqueous particle.
[0155] Nanoparticles As used in this article, the term "nanoparticle" refers to any nanoscale particle whose size typically ranges from about 1 nm to 1000 nm.
[0156] Nuclear localization sequence (NLS) As used herein, the term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that facilitates the importation of proteins (e.g., RNA-directed nucleases) into the cell nucleus, for example, via nuclear transport. For instance, an NLS sequence is described in Plank et al., International PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001, the contents of which, with respect to exemplary nuclear localization sequences, are incorporated herein by reference.
[0157] Nucleic acid As used herein, the terms “nucleic acid” or “nucleic acid molecule” or “nucleic acid sequence” or “polynucleotide” generally refer to deoxyribonucleic acid or ribonucleic acid oligonucleotides in single-stranded or double-stranded form. The term may also cover oligonucleotides containing known analogs of natural nucleotides. The term may also cover nucleic acid-like structures having a synthetic backbone, such as Eckstein, 1991; Baserga et al., 1992; Milligan, 1993; WO 97 / 03211; WO 96 / 39154; Mata, 1997; Strauss-Soukup, 1997; and Samstag, 1996. The term covers ribonucleic acid (RNA) and DNA, including cDNA (including RT DNA), genomic DNA, synthetic, synthetic (e.g., chemically synthesized) DNA, and / or DNA (or RNA) containing nucleic acid analogs. Nucleotides adenine (A), thymine (T), guanine (G), and cytosine (C) may or may not cover nucleotide modifications, such as methylated and / or hydroxylated nucleotides. For example, cytosine (C) covers 5-methylcytosine and 5-hydroxymethylcytosine.
[0158] Nucleic acid-guided nucleases As used herein, the term "nucleic acid-directed nuclease" or "nucleic acid-directed endonuclease" refers to a nuclease that covalently or non-covalently associates with a guide nucleic acid (e.g., guide RNA or guide DNA) to form a complex between the guide nucleic acid and the nucleic acid-directed nuclease. The guide nucleic acid contains a spacer sequence comprising a nucleotide sequence complementary to the target DNA sequence strand. Thus, the nucleic acid-directed nuclease indirectly directs or programs itself to a specific site in a DNA molecule through its association with the guide nucleic acid, which directly binds to or adheres to the target DNA strand via its complementary region through Watson-Crick base pairing. In some embodiments, the nucleic acid-directed nuclease will include DNA-binding activity (e.g., as in the case of CRISPR Cas9). Most commonly, the nucleic acid-directed nuclease is programmed by association with a guide RNA molecule, and in such cases, the nuclease may be referred to as an "RNA-directed nuclease." When programmed by guide DNA, the nuclease may be referred to as a "DNA-directed nuclease." Nucleic acid-directed, RNA-directed, or DNA-directed nucleases can also be called “programmable nucleases,” which also include other classes of programmable nucleases that associate with specific DNA sequences via amino acid / nucleotide sequence recognition (e.g., zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs)) rather than via guide RNA. Additionally, any nuclease considered herein can be engineered to remove, inactivate, or otherwise eliminate one or more nuclease activities (e.g., by introducing a nuclease-inactivating mutation into the active site of the nuclease). Nucleases modified to remove, inactivate, or otherwise eliminate all nuclease activities can be called “dead” nucleases. Dead nucleases cannot cleave either strand of a double-stranded DNA molecule. Nucleases modified to remove, inactivate, or otherwise eliminate at least one nuclease activity but still retain at least one nuclease activity can be called “nicking” nucleases. Nicking nucleases cleave one strand of a double-stranded DNA molecule but not both strands. For example, CRISPR Cas9 naturally contains two distinct nuclease active domains: the HNH domain and the RuvC domain. The HNH domain cleaves the DNA strand bound to the guide RNA, and the RuvC domain cleaves the protospacer strand. The nicking enzyme Cas9 can be obtained by inactivating either the HNH or RuvC domain. Dead Cas9 can be obtained by inactivating both the HNH and RuvC domains. Other RNA-directed nucleases can be similarly converted into nicking enzymes and / or dead nucleases by inactivating one or more of the existing nuclease domains.
[0159] Operable connection As used herein, the terms “operably linked” or “under transcriptional control” when used in conjunction with a description of a promoter refer to the correct position and orientation of a polynucleotide (e.g., a coding sequence) to control the initiation of RNA polymerase transcription and the coding sequence (e.g., ...) MSR Gene, msd Genes and / or ret The expression of a gene (the coding sequence of the gene). If other transcriptional regulatory elements (e.g., enhancer sequences, transcription factor binding sites) control or regulate gene expression relative to the gene's position, they can also be operatively linked to the gene.
[0160] Programmable nuclease As used herein, the term "programmable nuclease" refers to a polypeptide that has the property of selectively targeting specific desired nucleotide sequences within a nucleic acid molecule (e.g., targeting a specific gene target) due to one or more targeting functions. Such targeting functions may include one or more DNA-binding domains, such as zinc finger domains characteristic of many different types of DNA-binding proteins or TALE domains characteristic of TALEN proteins. This targeting function may also include the ability to associate with and / or form a complex with guide RNA, and then target the guide RNA to a specific site on DNA carrying a sequence complementary to a portion of the guide RNA (i.e., the spacer of the guide RNA). In some embodiments, a programmable nuclease may be a single protein comprising a domain that binds directly (e.g., a ZF protein) or indirectly (e.g., an RNA-directed protein) to a target DNA site, as well as a nuclease domain. In other embodiments, a programmable nuclease may be a synthesis of two or more separate proteins or domains (from different proteins) that together provide the necessary functions of selective DNA binding and nuclease activity. For example, a programmable nuclease may include a (a) nuclease inactive RNA-guided nuclease fused to (b) a nuclease protein or domain (which is still able to bind to the guide RNA, locate to the target DNA, and bind to the target DNA, but cannot cleave or nick the strand), such as the FokI nuclease.
[0161] promoter As used herein, the term "promoter" is a nucleic acid molecule recognized in the art and refers to a sequence that is recognized by the cellular transcription machinery and is capable of initiating transcription of a downstream gene. Promoters can be persistently activated, meaning that a promoter is always active in a given cellular environment, or conditionally active, meaning that a promoter is active only in the presence of specific conditions. For example, a conditional promoter may be active only in the presence of a specific protein that links a protein associated with a regulatory element in the promoter to the basic transcription machinery, or only in the absence of a repressor molecule. Within the promoter sequence are transcription start sites and protein-binding domains responsible for RNA polymerase binding. Eukaryotic promoters typically, but not always, contain a "TATA" box and a "CAT" box. Various promoters, including inducible promoters, can be used to drive the expression of various vectors disclosed herein.
[0162] Recombinant nucleic acid "Recombinant nucleic acid" or "recombinant nucleotide" refers to a molecule constructed by conjugating nucleic acid molecules, which optionally can self-replicate in living cells.
[0163] retrotran As used herein, the term "reverse transcript" refers to a specific type of naturally occurring and unique DNA sequence found in the genomes of many bacteria that typically encodes three distinct components: (a) non-coding RNA ("ncRNA") containing consecutive reverse sequences (msr and msd), (b) a gene encoding reverse transcriptase (RT) (ret), and (c) in many cases, a reverse transcriptase-associated gene whose function is unknown. The unique feature of reverse transcripts is their ability to produce a satellite DNA called msDNA (multiple-copy single-stranded DNA). ncRNA (containing...) MSR and msd Components) and ret Genes are transcribed into a single polycistronic RNA transcript, which is then processed into ncRNA transcripts and encoding... ret The transcript of a gene. The ncRNA then folds into a specific secondary structure. Once translated, the RT then binds to the folded ncRNA and... msd The region undergoes reverse transcription to form a single-stranded cDNA (msDNA), which is covalently linked to the RNA template via a 2'-5' phosphodiester bond and base pairing between the msDNA and the 3' end of the RNA template. See also Figure 2 The figure provided is a schematic diagram of the preparation of msDNA from naturally occurring reverse transcriptases.
[0164] Reverse transcriptome components As used herein, the term “reverse transcript component” refers to the unique elements or features of a reverse transcript, namely (a) a non-coding RNA (“ncRNA”) containing consecutive reverse sequences (msr and msd)), (b) a gene encoding reverse transcriptase (RT) (ret), and (c) in many cases, a reverse transcript-associated gene whose function is unknown.
[0165] RNA-directed nucleases As used herein, “RNA-directed nuclease” refers to a type of “programmable nuclease” and a specific type of “nucleic acid-directed nuclease.” As used herein, the term “RNA-directed nuclease” or “RNA-directed endonuclease” refers to a nuclease that covalently or nonvalently associates with a guide RNA to form a complex between the guide RNA and the RNA-directed nuclease. The guide RNA contains a spacer sequence comprising a nucleotide sequence complementary to the target DNA strand. Thus, the RNA-directed nuclease indirectly directs or programs itself to locate a specific site in the DNA molecule through its association with the guide RNA, which directly binds to or adheres to the target DNA strand via Watson-Crick base pairing through its complementary region.
[0166] Sequence identity As used herein, the term "sequence identity" refers to the overall relevance between aggregate molecules, such as between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or between polypeptide molecules. For example, the percentage of identity between two polynucleotide sequences can be calculated by aligning two sequences for optimal comparison purposes (e.g., vacancies can be introduced in one or both of the first and second nucleic acid sequences to achieve optimal alignment, and non-identical sequences can be ignored for comparison purposes). For example, the length of the sequences aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the length of the reference sequence. Nucleotides at corresponding nucleotide positions are then compared. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percentage of identity between two sequences is a function of the number of shared positions in the sequences, taking into account the number of vacancies and the length of each vacancy that is necessary to achieve optimal alignment of the two sequences. The comparison of sequences and the determination of the percentage of identity between two sequences can be accomplished using mathematical algorithms. For example, the percentage of identity between two nucleotide sequences can be determined using methods described in, for example, those in the following literature: Computational Molecular Biology, Lesk, AM ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW ed., Academic Press, New York, 1993; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Computer Analysis of Sequence Data, Part I, Griffin, AM and Griffin, HG ed., Humana Press, New Jersey, 1994; and Sequence Analysis Primer, Gribskov, M. and Devereux, J. ed., M Stockton Press, New York, 1991; each of these literatures is incorporated herein by reference.For example, the percentage of identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller as described in CABIOS (1989, 4:11-17), which has been incorporated into the ALIGN program (version 2.0) using a PAM120 weighted residue table with a vacancy length penalty of 12 and a vacancy penalty of 4. Alternatively, the percentage of identity between two nucleotide sequences can be determined using the GAP program in the GCG software package utilizing the NWSgapdna.CMP matrix. Methods commonly used to determine the percentage of identity between sequences include, but are not limited to, those disclosed in Carillo, H. and Lipman, D., SIAM J Applied Math., 48:1073 (1988); which is incorporated herein by reference. Techniques for determining identity are incorporated into publicly available computer programs. Exemplary computer software used to determine homology between two sequences may include, but is not limited to, the GCG package (as described in Devereux, J. et al., Nucleic Acids Research, 12(1), 387 (1984)), BLASTP, BLASTN, and FASTA (as described in Altschul, SF et al., J. Molec. Biol., 215, 403 (1990)).
[0167] It should be noted that when this disclosure relates to polypeptides having a percentage of identity with another amino acid sequence (reference amino acid sequence) (including any part of this specification, including Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, and examples), for example having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, or at least 80% with another amino acid sequence (reference amino acid sequence), For polypeptides with at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity, it is advantageous that, among polypeptides having a percentage of identity with a reference amino acid sequence, the conserved regions of the reference amino acid sequence (e.g., those with other reverse transcriptases RT) (For example, those identified herein) are conserved compared to those that are more so, and / or the polypeptide has at least one activity selected from reverse transcriptase; endonuclease activity; ribonuclease activity, or RNA-directed DNase activity, and / or the polypeptide comprises: a. one or more α-helix recognition leaves (REC) and a nuclease leaf (NUC); b. a wedge (WED), α-helix recognition leaf (REC), PAM interaction (PI), RuvC nuclease, helical bridge (BH), and NUC domain; or c. one or more domains selected from RuvC, REC, WED, BH, PI, and NUC domains, and / or the polypeptide recognizes or binds to guide RNA or ncRNA, as applicable.Similarly, when this disclosure relates to a nucleic acid sequence or molecule having a percentage of identity with a nucleic acid sequence having a percentage of identity with another nucleic acid sequence or molecule (reference nucleic acid sequence), for example, having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity with another nucleic acid sequence, it is advantageous that, in the nucleic acid sequence having a percentage of identity with the reference nucleic acid sequence, the conserved regions of the reference nucleic acid sequence (e.g., with other retrotranscriptor ncRNAs) are... Conserved regions (e.g., those identified herein) are retained compared to other reverse transcriptase sequences (e.g., those identified herein), and / or in polypeptides expressed from nucleic acid sequences having a percentage of identity with a reference nucleic acid sequence, the polypeptide contains conserved regions (e.g., conserved compared to other reverse transcriptase sequences (e.g., those identified herein), and / or the polypeptide has at least one activity selected from reverse transcriptase; endonuclease activity; ribonuclease activity, or RNA-directed DNase activity, and / or the polypeptide comprises: a. one or more α-helix recognition leaves (REC) and nuclease leaves (NUC); b. a wedge (WED), α-helix recognition leaf (REC), PAM interaction (PI), RuvC nuclease, helical bridge (BH), and NUC domain; or c. one or more domains selected from RuvC, REC, WED, BH, PI, and NUC domains, and / or the polypeptide recognizes or binds to guide RNA.
[0168] Subjects As used herein, the term "subject" refers to an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, for example, a genetically engineered non-human subject. The subject can be of any sex and at any developmental stage. The terms "individual," "subject," "host," and "patient" are used interchangeably herein.
[0169] stem As used herein, the term "stem" refers to two or more base pairs, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more base pairs, formed by an inverted repeat sequence attached to a "tip," in which the 5′ or "upstream" strand of the stem bends to allow the 3′ or "downstream" strand to pair with the upstream strand. The number of base pairs in the stem is the "length" of the stem. The tip of the stem is typically at least 3 nucleotides, but can also be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleotides. Larger tips with more than 5 nucleotides are also called "loops." A stem that is originally continuous may be interrupted by one or more protrusions as defined herein. The number of unpaired nucleotides in a protrusion is not included in the length of the stem. The location of the protrusion closest to the tip can be described by the number of base pairs between the protrusion and the tip (e.g., the protrusion is 4 bp from the tip). The location of other protrusions farther from the tip (if present) can be described by the number of base pairs in the stem between the protrusion in question and the tip, excluding any unpaired bases from other protrusions in between.
[0170] Synthetic or artificial nucleic acids "Synthetic or artificial nucleic acids" refers to nucleic acids that are sequences that do not exist naturally. Such sequences are not derived from any living organism, or their existence in any living organism is unknown (e.g., based on sequence searches in existing sequence databases). Recombinant and synthetic nucleic acids also include those molecules produced by replication of any of the foregoing. The engineered nucleic acid constructs disclosed herein, such as the engineered reverse transcripts described herein, may be encoded by a single molecule (e.g., encoded by or present on the same plasmid or other suitable vector) or by multiple different molecules (e.g., multiple independently replicated vectors).
[0171] target site As used herein, a “target site” is a polynucleotide (e.g., DNA, such as genomic DNA) that includes a site or specific locus (“target site” or “target sequence”) targeted by the recombinant reverse transcriptome genome modification system disclosed herein. In the context of the reverse transcriptome genome modification system disclosed herein, which includes an RNA-guided nuclease, a target sequence is a sequence that guides a nucleic acid (e.g., guide RNA) to hybridize with. For example, the target site (or target sequence) 5′-GTCAATGGACC-3′ (SEQ ID NO:19933) within a target nucleic acid is targeted (or bound to, hybridized with, or complemented by) the sequence 5′-GGTCCATTGAC-3′ (SEQ ID NO:19934). Suitable hybridization conditions include physiological conditions commonly present in cells. For double-stranded target nucleic acids, the strand of the target nucleic acid that is complementary to and hybridizes with the guide RNA is referred to as the “complementary strand” or “target strand”; while the strand of the target nucleic acid that is complementary to said “target strand” (and therefore not complementary to the guide RNA) is referred to as the “non-target strand” or “non-complementary strand”.
[0172] treat As used herein, the term "treatment / treat / treating" refers to a clinical intervention aimed at reversing or alleviating a disease or condition or one or more symptoms thereof as described herein, delaying its onset, or inhibiting its progression. In some embodiments, treatment may be administered after the onset of one or more symptoms and / or after the disease is diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms or to inhibit the onset or progression of the disease. For example, treatment may be administered to a susceptible individual before the onset of symptoms (e.g., based on a history of symptoms and / or based on genetic or other susceptibility factors). Treatment may also continue after symptom relief, for example, to prevent or delay its recurrence.
[0173] Upstream and downstream As used herein, the terms “upstream” and “downstream” are relative terms defining the linear positions of at least two elements within a nucleic acid molecule (whether single-stranded or double-stranded) oriented in a 5′ to 3′ direction. A first element is allegedly located upstream of a second element within the nucleic acid molecule, where the first element is located somewhere at the 5′ position of the second element. Conversely, a first element is allegedly located downstream of a second element within the nucleic acid molecule, where the first element is located somewhere at the 3′ position of the second element.
[0174] variants As used herein, the term "variant" should be understood to mean exhibiting properties different from those occurring in nature, such as a variant reverse transcriptant RT containing one or more altered amino acid residues compared to the amino acid sequence of the wild-type reverse transcriptant RT. The term "variant" encompasses homologous proteins that share at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% identity with a reference sequence and have the same or substantially the same functional activity as the reference sequence. The term also encompasses mutants, truncated versions, or domains of a reference sequence that exhibit the same or substantially the same functional activity as the reference sequence.
[0175] carrier As used herein, the term "vector" allows or facilitates the transfer of polynucleotides from one environment to another. Vectors are replicons such as plasmids, bacteriophages, or granules that insert another DNA segment into which the inserted segment is replicated (e.g., a subject-engineered reverse transcriptase). Typically, a vector is capable of replication when associated with appropriate control elements. The term "vector" can include cloning and expression vectors, as well as viral vectors and integration vectors.
[0176] wild type As used herein, the term "wild type" is a term understood by those skilled in the art and refers to the typical form of an organism, strain, gene, protein, or trait found in nature, in order to distinguish it from mutant or variant forms. Detailed Implementation
[0177] This disclosure provides systems, methods, and compositions for precise genome editing, including the insertion, replacement, and deletion of nucleic acids at targeted and precise genomic sites, wherein said systems, methods, and compositions are based on novel and / or modified reverse transcripts or components thereof, such as the reverse transcripts RT of Table X in U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. Modified versions of ncRNA in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, and modified versions of RT in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0178] In one aspect, this disclosure provides a non-coding RNA variant (ncRNA variant) comprising a reference reverse transcriptant non-coding RNA (ncRNA) having one or more modifications. The reference reverse transcriptant ncRNA may comprise, from the 5' to 3' direction, an a1 region, a first-branch guanine, an MSR, an MSD, and an a2 region. In some embodiments, the reference reverse transcriptant ncRNA further comprises a second-branch guanine. In some embodiments, one or more modifications comprise: (i) a linking of the a1 region to the a2 region; (ii) deletion of at least a portion of the MSR; (iii) deletion of at least a portion of the MSD; (iv) addition of a single-stranded RNA comprising a polymerase template; or (v) addition of an RNA motif. The ncRNA variant may be incorporated into a recombinant or engineered reverse transcriptant and provides improved functionality and / or properties compared to the reference reverse transcriptant.
[0179] Therefore, in one aspect, this disclosure provides recombinant reverse transcripts comprising one or more genetic modifications that improve the functionality and / or properties of the reverse transcript. Such genetic modifications may include mutations, insertions, deletions, inversions, substitutions, replacements, or translocations of one or more adjacent or non-adjacent nucleotides in a nucleic acid molecule encoding a reverse transcript or a component of a reverse transcript (e.g., ncRNA or reverse transcriptase). In several aspects, the reverse transcript modified with one or more genetic modifications (i.e., a “pre-modified” or “unmodified” reverse transcript or a component of a reverse transcript) is a naturally occurring reverse transcript or a component of a reverse transcript (e.g., naturally occurring ncRNA, or RT, as per Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872) that is capable of promoting homology-dependent recombination (or HDR) within the cell, resulting in a relative increase in the concentration or amount of msDNA containing a DNA donor template. In certain embodiments, the recombinant reverse transcript is based on and / or derived from naturally occurring reverse transcripts, such as any reverse transcript-related sequence provided in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 (introducing one or more genetic modifications into a set of 7257 previously unknown reverse transcripts discovered by the computational methods described herein (e.g., see examples)). In other instances, the recombinant reverse transcript may be based on introducing one or more genetic modifications into a previously available reverse transcript sequence, such as “Mestre et al., Systematic Prediction of Genes Functionally Associated with Bacterial Retrons and Classification of The Encoded Tripartite SystemsAs described in "Nucleic Acids Research, Vol. 48, No. 22, December 16, 2020, pp. 12632-12647" (incorporated hereby by reference), to obtain recombinant reverse transcripts with enhanced capabilities, thereby producing msDNA containing a DNA donor template at increased concentrations or amounts.
[0180] In another aspect, this disclosure further provides nucleic acid molecules encoding recombinant reverse transcripts and / or recombinant reverse transcript components (e.g., recombinant ncRNAs and / or recombinant reverse transcript RTs). In another aspect, this disclosure provides a genome editing system comprising recombinant reverse transcript components (e.g., recombinant ncRNAs and / or recombinant RTs), a programmable nuclease (e.g., RNA-directed nucleases such as CRISPR-Cas proteins, ZFPs, and TALENS), and a guide RNA (in the case of an RNA-directed nuclease used in the genome editing system). In another aspect, this disclosure provides nucleic acid molecules encoding the genome editing system and its components, as well as polypeptides constituting components of the genome editing system. In another aspect, this disclosure provides vectors for transferring and / or expressing the genome editing system, e.g., under in vitro, ex vivo, and in vivo conditions. In another aspect, this disclosure provides cell delivery compositions and methods, including compositions for passive and / or active transport to cells (e.g., plasmids), delivery via virus-based recombinant vectors (e.g., AAVs and / or lentiviral vectors), delivery via non-virus-based systems (e.g., liposomes and LNPs), and delivery via virus-like particles. Depending on the delivery system employed, the reverse transcriptase-based genome editing system described herein can be delivered in the form of DNA (e.g., plasmids or DNA-based viral vectors), RNA (e.g., ncRNA and mRNA delivered via LNPs), mixtures of DNA and RNA, proteins (e.g., virus-like particles), and ribonucleoprotein (RNP) complexes. Any suitable combination of methods for delivering the components of the reverse transcriptase-based genome editing system disclosed herein can be employed. In one embodiment, each of the components of the reverse transcriptase-based genome editing system is delivered via a whole RNA system, such as by delivering one or more RNA molecules (e.g., mRNA and / or ncRNA) via one or more LNPs, wherein said one or more RNA molecules form ncRNA and guide RNA (as needed) and / or are translated into polypeptide components (e.g., RT and programmable nucleases). In another aspect, this disclosure provides methods for generating edits at the target editing site by introducing the reverse transcriptase-based genome editing system described herein into cells containing target editing sites (e.g., under in vitro, in vivo, or ex vivo conditions). In other respects, this disclosure provides formulations comprising any of the foregoing components for delivery to cells and / or tissues (including in vitro, in vivo, and ex vivo delivery), i.e., recombinant cells and / or tissues modified by the recombinant reverse transcriptase-based genome modification systems and methods described herein, as well as methods for modifying cells by genome editing using the reverse transcriptase-based genome modification systems disclosed herein and related DNA donor-dependent methods, such as recombinant engineering or cell recording.This disclosure also provides methods for manufacturing the recombinant reverse transcripts described herein, reverse transcript-based genome modification systems, vectors, compositions, and formulations, as well as pharmaceutical compositions and kits for modifying cells under in vitro, in vivo, and ex vivo conditions, comprising the genome editing and / or modification systems disclosed herein.
[0181] This article describes engineered reverse transcriptases comprising one or more heterologous nucleic acids. These heterologous nucleic acids may, for example, be inserted at or within locations selected from: msd loci MSR upstream of the locus msd upstream of the locus and msd Downstream of the locus. In some embodiments, the engineered reverse transcriptome has structural improvements, at least in terms of the encoded ncRNA and / or reverse transcriptase (RT), compared to the naturally occurring counterpart or wild-type reverse transcriptome, such that when the engineered reverse transcriptome or its encoded ncRNA is delivered to a host cell (e.g., a mammalian host cell), it exhibits various functional improvements compared to its naturally occurring / wild-type reverse transcriptome element.
[0182] Exemplary but non-limiting functional improvements may include any one or more features described herein. For example, in some embodiments, engineered reverse transcriptases may be available in... MSR loci and / or msd The locus contains sequence modifications (e.g., insertion, deletion, and / or substitution of one or more nucleotides) that: i) regulate (e.g., enhance) reverse transcription, continuity, accuracy / fidelity, and / or msDNA production (e.g., in mammalian cells); ii) regulate (e.g., reduce) the production of msDNA encoded by engineered reverse transcriptases in the host (e.g., in hosts containing mammalian cells). MSR loci and / or msd (iii) Immunogenicity of ncRNAs encoding loci; (iv) Nucleotide sequences containing functions that regulate (e.g., inhibit or antagonize) msDNA; and / or (v) regulate (e.g., enhance) the efficiency of targeted genome engineering.
[0183] Therefore, generally speaking, engineered reverse transcriptases are engineered nucleic acid constructs that contain: a) a first polynucleotide encoding non-coding RNA (ncRNA), wherein the first polynucleotide contains: i) encoding multiple copies of single-stranded DNA (msDNA). MSR RNA portion MSR loci; and ii) encoding msDNA msd RNA portion msd Locus; and b) One or more heterologous nucleic acids inserted at or within a location selected from the following: msdloci MSR upstream of the locus msd upstream of the locus and msd Downstream of the locus.
[0184] Engineered nucleic acid constructs (e.g., engineered reverse transcriptases or reverse transcriptase variants) may also contain a second polynucleotide encoding a reverse transcriptase (RT) or a portion thereof, wherein the encoded RT is capable of synthesizing msDNA. msd At least a portion of the DNA copy of the locus.
[0185] In some embodiments, the engineered reverse transcriptase of this disclosure encodes a reverse transcriptase (RT) or a functional domain thereof, comprising: i) a polypeptide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, or at least 8% with respect to the polypeptide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. The following are polypeptides with sequence identity of at least 0%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%; and / or ii) polypeptides listed in any of the following Table C of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. In some embodiments, the RT does not contain any polypeptides listed in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0186] In some embodiments, the engineered reverse transcriptase of this disclosure encodes a reverse transcriptase (RT) or a functional domain thereof, comprising: i) a polynucleotide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, or at least 8% with the polynucleotide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. 0%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity of polynucleotides; and / or ii) common polynucleotide sequences listed in Table C of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. In some implementations, the polynucleotide encoding RT does not contain the polynucleotide listed in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0187] In some embodiments, the engineered reverse transcriptase of this disclosure encodes ncRNA, comprising: (I) an ncRNA listed in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or an ncRNA similar to that listed in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. ncRNAs with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity.
[0188] In some embodiments, the engineered reverse transcriptase of this disclosure encodes ncRNA and reverse transcriptase (RT) or their functional domains, wherein the ncRNA and the RT or their functional domains are as described above.
[0189] Specifically, in such embodiments, the ncRNA may comprise: (I) an ncRNA listed in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a molecular weight of at least 5 ppm with the ncRNA listed in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. ncRNAs with sequence identity of at least 0%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%.
[0190] Furthermore, in such embodiments, the reverse transcriptase (RT) or its functional domain comprises: (A) i) a polypeptide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, or at least 98%. The RT does not contain polypeptides with at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity; and / or ii) polypeptides listed in Table C of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872; optionally, the RT does not contain polypeptides listed in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872; or (B) i) The polynucleotides listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, or at least 95% of the ...90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at The polynucleotide encoding RT does not contain the polynucleotides in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0191] In some implementations, the engineered nucleic acid construct includes: 1) MSR 1) Locus (the msrRNA portion encoding msDNA); 2) Encoding msDNA msd RNA portion msd 3) The sequence encoding the reverse transcriptase (RT), wherein... msd RNA can be reverse transcribed by reverse transcriptase (RT) to form msDNA; and 4) heterologous nucleic acids inserted at or within the following locations: msd locus, upstream of msr locus, upstream of or downstream of msd locus; wherein the engineered nucleic acid construct is designed based on and / or similar to the following structures: secondary structures of wild-type or co-reverse transcripts encoding wild-type or co-reverse transcript ncRNAs, which are covered by: a) Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 and / or Figure 2-2 The sequence and / or structure depicted by any of SEQ ID NO of 7; or a variant of b) a) having: i) at most 1, 2, or 3 (e.g., at most 1) nucleotide variations per 10 red letter nucleotides; ii) at most 4, 5, or 6 (e.g., at most 1 or 2) nucleotide variations per 10 black letter nucleotides; and / or iii) at most 7, 8, or 9 (e.g., at most 3 or 4) nucleotide variations per 10 gray letter nucleotides; and / or optionally further comprising: i) 7, 8, 9, or 10 (e.g., 9 or 10) nucleotides per 10 red circle-marked nucleotides; ii) 6, 7, 8, 9, or 10 (e.g., 8, 9, or 10) nucleotides per 10 black circle-marked nucleotides. iii) 4, 5, 6, 7, 8, 9, or 10 (e.g., 6, 7, 8, 9, or 10) nucleotides per 10 gray-circled nucleotides; and / or iv) 2, 3, 4, 5, 6, 7, 8, 9, or 10 (e.g., 4, 5, 6, 7, 8, 9, or 10) nucleotides per 10 white-circled nucleotides; wherein the ncRNA does not contain an ncRNA associated with the sequence of Table X in U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0192] Engineered nucleic acid constructs (e.g., engineered reverse transcriptases) can be used in MSR loci and / or msdThe locus contains one or more sequence modifications (e.g., insertion, deletion, and / or substitution of one or more nucleotides) that: a) regulate (e.g., enhance) reverse transcription, continuity, accuracy / fidelity, and / or msDNA production (e.g., in mammalian cells); b) regulate (e.g., reduce) the production of reverse transcriptase by engineered reverse transcriptase (e.g. MSR loci and / or msd (c) Immunogenicity of ncRNAs encoded by loci in the host (e.g., hosts containing mammalian cells); (d) Regulation (e.g., permanent or temporary inhibition) of msDNA function; and / or (e.g., regulation) of the efficiency of targeted genome editing / engineering.
[0193] In some implementations, the engineered nucleic acid construct (e.g., engineered reverse transcript) is designed based on and / or similar to the following structures: a secondary structure of a wild-type or co-reverse transcript encoding a wild-type or co-reverse transcript ncRNA, which is covered by: a) the sequence of any of the ncRNA sequences in Table B of U.S. Patent Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 and / or Figure 2-2 The structure described in any of 7; or a variation of b) a) having: i) at most 1, 2, or 3 (e.g., at most 1) nucleotide variation per 10 red-letter nucleotides; ii) at most 4, 5, or 6 (e.g., at most 1 or 2) nucleotide variation per 10 black-letter nucleotides; and / or iii) at most 7, 8, or 9 (e.g., at most 3 or 4) nucleotide variation per 10 gray-letter nucleotides; and / or optionally also including: i) the presence of 7, 8, 9, or 10 (e.g., 9 or 1) nucleotides marked with red circles. ii) 6, 7, 8, 9 or 10 (e.g. 8, 9 or 10) nucleotides per 10 black circles; iii) 4, 5, 6, 7, 8, 9 or 10 (e.g. 6, 7, 8, 9 or 10) nucleotides per 10 gray circles; and / or iv) 2, 3, 4, 5, 6, 7, 8, 9 or 10 (e.g. 4, 5, 6, 7, 8, 9 or 10) nucleotides per 10 white circles.
[0194] Another aspect of this disclosure provides a vector system comprising a vector containing an engineered reverse transcriptase as described herein.
[0195] Another aspect of this disclosure provides a single host cell containing either the engineered reverse transcriptase or the vector system described herein.
[0196] Another aspect of this disclosure provides a pharmaceutical composition comprising the engineered reverse transcriptase or the vector system described herein.
[0197] Another aspect of this disclosure provides a delivery medium comprising an engineered reverse transcriptant or an ncRNA encoded by an engineered reverse transcriptant as described herein, a vector or vector system as described herein, a host cell as described herein, or a pharmaceutical composition as described herein.
[0198] Another aspect of this disclosure provides a kit comprising an engineered reverse transcriptant or an ncRNA encoded by an engineered reverse transcriptant as described herein, and instructions for genetically modifying cells optionally using the engineered reverse transcriptant or an ncRNA encoded by an engineered reverse transcriptant as described herein.
[0199] Another aspect of this disclosure provides a method for modifying a target DNA sequence in a host cell (e.g., a mammalian cell), the method comprising introducing an engineered reverse transcriptant of this disclosure, an ncRNA encoded by the engineered reverse transcriptant of this disclosure, or a vector / vector system described herein into a host cell (e.g., a mammalian cell) to generate msDNA in the host cell (e.g., the mammalian cell), wherein at least a portion of the heterologous nucleic acid in the msDNA is integrated into a target DNA sequence in the genome of the host (e.g., mammalian) cell. Optionally, the target sequence is recognized by a suitable nuclease such as a CRISPR / Cas effector enzyme, ZFN, TALEN, a broad-spectrum nuclease, TnpB, IscB, or a restriction endonuclease (RE), and the nuclease generates a double-strand break (DSB) to facilitate / enhance the insertion of the heterologous nucleic acid portion into the target sequence. Furthermore, optionally, the target sequence modified / inserted with the heterologous nucleic acid portion is no longer recognized by the nuclease to generate a DSB again.
[0200] Another aspect of this disclosure provides the use of engineered reverse transcripts in the various methods described herein.
[0201] Another aspect of this disclosure provides a genome editing system comprising: a) a nuclease capable of acting at a target site on a genome (e.g., the human genome), such as a CRISPR / Cas effector enzyme, ZFN, TALEN, a wide-ranging nuclease, TnpB, IscB, or a restriction endonuclease (RE); and b) an engineered reverse transcriptase as described herein, or an ncRNA encoded therefrom, or a vector or vector system containing or encoding therefrom. Optionally, the nuclease may be linked to one or more elements of the engineered reverse transcriptase or the encoded ncRNA. For example, in one embodiment, the nuclease may be linked (e.g., fused or conjugated) to the reverse transcriptase of the engineered reverse transcriptase as described herein. In another embodiment, the nuclease may be conjugated / binded to a nucleic acid guide sequence (e.g., a single guide RNA of a Cas enzyme) to form a complex, wherein the guide sequence is linked to the ncRNA and / or msDNA of the engineered reverse transcriptase as described herein.
[0202] Another aspect of this disclosure provides an enhanced genome editing system comprising the genome editing system of this disclosure, said system being linked to a biomolecule capable of modulating host DNA repair for, for example, to modulate (e.g., enhance) the incorporation of heterologous nucleic acid sequences into a genome (e.g., the human genome).
[0203] Based on the general aspects of this disclosure described herein, specific aspects and embodiments of this disclosure will be further described in the following sections. It should be understood that any embodiment of this disclosure, including those described only in the embodiments or claims, or only in some section below, may be combined with any one or more other embodiments of this disclosure, unless such combination is expressly excluded or inappropriate.
[0204] A. Chimeric gene editing compositions This disclosure describes a chimeric gene editing system that combines one or more site-specific gene editing components (e.g., leader editing or CRISPR / Cas9) with components of one or more reverse transcriptase editing systems (including reverse transcriptase RT and modified reverse transcriptase ncRNA). The chimeric gene editing composition comprises: (a) ncRNA, optionally containing a polymerase template sequence and a first polymerase; (b) a nuclease or a first mRNA encoding said nuclease; (c) a guide RNA (gRNA) associated with said nuclease; and (d) optionally, a second polymerase or a second mRNA encoding the second polymerase.
[0205] In some embodiments, the chimeric gene editing composition comprises: (a) a nicking enzyme component, (b) a guide RNA compounded with said nicking enzyme component and directed to a target DNA sequence, (c) a polymerase component, and (d) a polymerase template sequence, wherein said polymerase template sequence is provided by a modified reverse transcriptase ncRNA (referred to herein as “template-based ncRNA” or “tncRNA”) containing the polymerase template sequence. In several embodiments, the chimeric gene editing composition comprises an ncRNA variant having an altered structural configuration relative to a reference reverse transcriptase ncRNA (e.g., wild-type reverse transcriptase ncRNA), said alteration including: (i) a connector that joins the 5' end of the ncRNA a1 region to the 3' end of the ncRNA a2 region, (ii) a deletion of the wild-type MSD region or a portion thereof, (iii) a deletion of the wild-type MSR region or a portion thereof, (iv) mounting a single-stranded RNA sequence containing the polymerase template, for example, in place of the deleted MSD sequence, and / or (iv) the addition of an RNA motif. The chimeric gene editing system may also optionally include a reverse transcriptase RT to convert templated ncRNA into homologous msDNA, referred to herein as “templated msDNA” or “tmsDNA”.
[0206] The gene editing system described herein may also include any templated ncRNA derived from any ncRNA known in the art, including any of those described in Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 (each incorporated herein by reference in its entirety, including its sequence listing), or a sequence having at least 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a sequence from Table B of any of the aforementioned applications.
[0207] In 1984, in *Myxococcus faecalis* (… Myxococcus xanthus The first discovery of retrotranscriptors was in bacteria when a short, multi-copy single-stranded DNA (msDNA) abundant in bacterial cells was identified. Since then, numerous naturally occurring retrotranscriptors have been found in many prokaryotes, such as bacteria.
[0208] Reverse transcripts are encoded and transcribed as a single RNA molecule, which contains a non-coding RNA (ncRNA) portion and a portion encoding a specific reverse transcriptase (RT). Reverse transcript ncRNA ( MSR and msd The ncRNA is the precursor to the final hybrid molecule, and it initially folds into a typical RNA secondary structure, which is recognized by the accompanying RT. The translated RT typically recognizes specific secondary structures within the ncRNA and binds... msdThe RNA template is located downstream of the region. RT begins at the 2' end of the double-stranded RNA structure (a1 / a2 region) within the ncRNA, immediately adjacent to the conserved guanosine (G) residue. The initiator RNA reverses transcription towards its 5' end. A portion of the ncRNA serves as the template for reverse transcription, and the reverse transcription proceeds to the 5' end. MSR The transcription terminates before the locus. During reverse transcription, cellular ribonuclease H degrades the ncRNA segment used as a template, but not the rest of the ncRNA. The resulting msDNA is covalently linked to the RNA template via a 2'-5' phosphodiester bond, and base pairing is performed using the 3' end of the msDNA with the RNA template. For a general or typical structure of the reverse transcriptase coding sequence, see [link to relevant documentation]. Figure 2 Including RT-coded sequences and MSR and msd The locus, and the reverse transcription of the initial ncRNA transcript to synthesize msDNA.
[0209] Templated retrotranscript ncRNAs can be constructed by modifying the start-site retrotranscript DNA sequence (msr / msd region) encoding the ncRNA (such as any of those in Table B above). The start-site retrotranscript DNA sequence encoding the ncRNA (refer to retrotranscript ncRNA) can be modified in various ways and may include one or more modifications.
[0210] The reverse transcriptase ncRNA sequence may be further modified to include at least one nucleotide modification, including a single nucleotide substitution, insertion, or deletion, or more than one nucleotide substitution, insertion, or deletion, i.e., at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62. 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or up to 100, or up to 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or up to 2000 nucleotides may be substituted, inserted, or deleted from the starting site retrotran (e.g., wild-type retrotran). When more than one nucleotide in a start-point reverse transcript (e.g., a wild-type reverse transcript) is substituted, deleted, or inserted, the nucleotides can be continuous or discontinuous. While an engineered reverse transcript as a whole is not naturally occurring, it can include components that are indeed present in nature, such as nucleotide sequences. For example, an engineered reverse transcript may have nucleotide sequences derived from different organisms (e.g., different bacterial species) or from fully synthetic / artificial / recombinant nucleic acid sequences. Thus, an engineered reverse transcript may have bacterial nucleotide sequences, human nucleotide sequences, viral nucleotide sequences, and / or synthetic / artificial / recombinant nucleotide sequences, and / or combinations of such sequences. Examples of modifications to the recombinant reverse transcripts disclosed herein include the insertion of heterologous nucleic acid sequences into the reverse transcript, for example, insertion into ncRNA loci, such as in the MSR or MSD loci. Linking a guide RNA molecule to the 5′ end and / or the 3′ end (i.e., one linked to the 5′ end of the ncRNA and / or one linked to the 3′ end of the ncRNA) also represents another modification contemplated for the recombinant reverse transcripts disclosed herein. In such implementations, the guide RNA molecule can also be categorized as, or more broadly, as a type of heterologous nucleic acid sequence used to modify the initiation site reverse transcriptase. Examples of such modifications are depicted in... Figures 8A to 8E middle.
[0211] In addition to DNA encoding ncRNA, DNA encoding RT can also be modified to obtain recombinant RT. For example, DNA encoding RT can be modified to include at least one nucleotide modification, including a single nucleotide substitution, insertion, or deletion, or more than one nucleotide substitution, insertion, or deletion, i.e., up to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or up to 100, or up to 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or up to 2000 nucleotides are replaced, inserted, or deleted from the start-point retrotran (e.g., wild-type retrotran) within the RT gene.
[0212] Such modifications to the DNA encoding ncRNA and / or RT can modulate the function of ncRNA and / or RT in a variety of ways, including: i) regulating (e.g., enhancing) reverse transcription, continuity, accuracy / fidelity, and / or msDNA production (e.g., in mammalian cells); ii) regulating (e.g., reducing) the ncRNA encoded by engineered reverse transcriptases in the host (e.g., a host containing mammalian cells). MSR loci and msd (iii) Immunogenicity of loci; (iv) Regulation (e.g., permanent or temporary inhibition) of msDNA function; and / or (v) Regulation (e.g., enhancement) of the efficiency of targeted genome editing / engineering.
[0213] In one embodiment, this disclosure provides a recombinant reverse transcript having the following general structure: a) MSR a) locus; b) encoding msDNA msd RNA portion msd locus; c) sequence encoding reverse transcriptase (RT) (optionally trans-formed with ncRNA), wherein msdRNA can be reverse transcribed by a reverse transcriptase (RT) (e.g., in a host cell such as a mammalian cell) to form msDNA; and / or d) a polymerase template nucleic acid that can be transcribed. In some embodiments, the reverse transcriptase has a general structure comprising the following from the 5' to 3' direction: an a1 region, a first branch guanine, msr, msd, and an a2 region.
[0214] The engineered reverse transcripts disclosed herein may optionally be further structurally modified to include one or more heterologous nucleic acids. The engineered reverse transcripts may be further modified to provide various functional improvements, such as (but not limited to) enhancing the production of msDNA in cells (e.g., mammalian cells, including human cells).
[0215] In some embodiments, the engineered reverse transcriptase of this disclosure encodes a reverse transcriptase (RT) or a domain thereof, comprising: i) a polypeptide listed in Table A above, or ii) a polypeptide having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with a polypeptide listed in Table A above.
[0216] In some embodiments, the engineered reverse transcriptase of this disclosure encodes a reverse transcriptase (RT) or a functional domain thereof, comprising: i) a polynucleotide listed in Table A above, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, or at least 80% of the polynucleotides listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. 1) Polynucleotides with sequence identity of at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%, and / or ii) common polynucleotide sequences listed in Table A above.
[0217] In some embodiments, the templated ncRNA is derived from ncRNAs comprising: (I) ncRNAs listed in Table B above, or (II) ncRNAs having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with the ncRNAs in Table B above.
[0218] The engineered ncRNAs disclosed herein may differ in size from the ncRNAs on which they are based, and the proportion of ncRNAs retained in the engineered ncRNAs may vary. In some embodiments, the amount of ncRNA retained is about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 88%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, or 50% to 80%, or 60% to 85%, or 70% to 90%, or 80% to 95%, or 85% to 98%, or 90% to 99%, or all of the ncRNAs.
[0219] In some embodiments, the engineered reverse transcriptase of this disclosure encodes an ncRNA and a reverse transcriptase (RT) or a domain thereof, wherein the ncRNA and the RT or the domain thereof are as described above.
[0220] In some instances, engineered reverse transcripts are designed based on an evolutionary clade defined as reverse transcripts / reverse transcript RTs, where the reverse transcript is associated with a ternary system consisting of ncRNA, RT, and an additional protein or RT fusion domain with multiple enzymatic functions, for example, as described by Mestre et al. Systematic Prediction of Genes Functionally Associated with Bacterial Retrons and Classification of The Encoded Tripartite SystemsThe clades are described in Nucleic Acids Research, Vol. 48, No. 22, December 16, 2020, pp. 12632-12647 (incorporated hereby by reference). While the clades are primarily based on naturally occurring ncRNAs and retrotrans / retrotran RTs, as well as additional protein or RT fusion domains, they are not limited to naturally occurring sequences for the purpose of serving as templates for engineered retrotrans. Conversely, the clades may also encompass non-naturally occurring ncRNAs and RTs, including but not limited to recombinant, modified or altered, chimeric, hybrid, synthetic, artificial, etc.
[0221] Therefore, according to this disclosure, retrotranscriptors can be considered phylogenetically related based on the alignment of their RTs, with a Neighbor-Joining algorithm success rate of at least 75% (at least 1000 repetitions) and a Poisson-corrected distance measurement not exceeding 0.05. Furthermore, or additionally, retrotranscriptors can be considered phylogenetically related when / if the same RT or closely related RTs can recognize the secondary structure of the retrotranscriptor's ncRNA and reverse transcribe the retrotranscriptor to produce msDNA.
[0222] In some implementations, sequence alignment or secondary structure generation between different reverse transcript sequences (e.g., ncRNA and / or RT (protein and / or nucleic acid) sequences) is based on software known to those skilled in the art.
[0223] Retrotran ncRNA sequences within the same clade (including MSR and msd (Sequences) may be highly conserved at some positions and less conserved at others.
[0224] Unless otherwise explicitly stated, both a1 and a2 regions are single-stranded and substantially anticomplementary to each other, forming a stem optionally interrupted by symmetrical or asymmetrical protrusions, with one or more optional 5′ and / or 3′ protrusions / unpaired nucleotides, wherein the a1 region typically terminates before the conserved branched guanosine (G) that provides the 2′-OH of the reverse transcription primer (e.g., immediately adjacent to the 5′ end).
[0225] In some implementations, sequence variations include mutated, reduced, or eliminated protrusions in the a1 / a2 stem region, including sequence variations in one (i.e., a1 or a2) strand or both a1 and a2 strands.
[0226] For example, in some embodiments, sequence changes include deleting nucleotides from a1, a2, or both a1 and a2, thereby reducing the size of the bump, or making a symmetrical bump asymmetrical or vice versa, or eliminating the bump.
[0227] In some implementations, the sequence change includes replacing / substituting nucleotides in a1, a2, or both a1 and a2, such that previously unpaired bases in the protrusion form base pairs.
[0228] In some implementations, the sequence variation includes replacing an unpaired purine base with one or more unpaired pyrimidine bases.
[0229] In some implementations, the sequence change includes replacing an unpaired pyrimidine base with one or more unpaired purine bases.
[0230] In some implementations, the sequence change includes replacing an unpaired purine base (e.g., A or G) with another unpaired purine base (e.g., G or A).
[0231] In some embodiments, the sequence change includes replacing an unpaired pyrimidine base (e.g., T / U or C) with another unpaired pyrimidine base (e.g., C or T / U, respectively).
[0232] In some implementations, sequence variations include lengthening or shortening a1, a2, or both a1 and a2.
[0233] For example, the length of a1 can be shortened by deleting a 5′ overhang, deleting any upstream protruding nucleotide, or deleting a base involved in base pairing. Similarly, the length of a1 can be lengthened by adding a 5′ overhang, adding any upstream protruding nucleotide, or adding a base involved in base pairing.
[0234] In some implementations, the length of a2 can be shortened by deleting a 5′ overhang, deleting any downstream protruding nucleotide, or deleting a base involved in base pairing. Similarly, the length of a2 can be lengthened by adding a 5′ overhang, adding any downstream protruding nucleotide, or adding a base involved in base pairing.
[0235] Retrotran and ncRNA variants The ncRNA variants provided herein contain one or more modifications compared to the reference reverse transcriptase ncRNA. These modifications may include: (i) a1-a2 region binding; (ii) deletion of at least a portion of the msr; (iii) deletion of at least a portion of the msd; (iv) addition of a single-stranded RNA containing a polymerase template; and / or (v) addition of an RNA motif.
[0236] In some embodiments, the ncRNAs disclosed herein can be modified by introducing additional RNA motifs into the ncRNA, for example, at the 5′ and 3′ ends of the ncRNA, or even at a position between the two (e.g., at...). MSRor msd ncRNAs can be modified to enhance transcription production and / or stability and / or function (e.g., RT-DNA production). Such structures may include, but are not limited to, RNA hairpins, RNA stem-loops, RNA tetrads, cap structures, and poly(A) tails or ribozyme functions. Furthermore, ncRNAs may be modified to include one or more nuclear localization sequences.
[0237] Additional RNA motifs can also improve the RT continuity of ncRNAs or enhance ncRNA activity by strengthening RT binding. Adding dimerizing motifs (e.g., kissing loops or GNRA tetracyclic / tetracyclic receptor pairs) to the 5′ and 3′ ends of ncRNAs can also lead to efficient circularization of ncRNAs, thereby improving stability. Furthermore, it is envisioned that the addition of these motifs could enable the physical separation of ncRNA components, for example… MSR and msd Separation of regions. Short 5' or 3' extensions of ncRNAs that form small foothold hairpins at either end or both ends can also advantageously compete for adhesion along the length of the ncRNA to complementary regions. Finally, kissing loops can also be used to recruit other RNAs or proteins to genomic sites and to facilitate the exchange of RT activity from one RNA to another.
[0238] ncRNAs can be further improved through directed evolution in a manner similar to that used to improve protein function. Directed evolution can enhance ncRNA recognition by RT and / or reduce ectopic targeting and / or insertion / deletion and / or improve the efficiency of precise editing.
[0239] This disclosure considers any such means to further improve the stability and / or functionality of the ncRNAs disclosed herein.
[0240] In some embodiments, the RNA (including guide RNA and ncRNA) used in the compositions of this disclosure has undergone chemical or biological modifications to make it more stable. Exemplary modifications to RNA include base deletions (e.g., by deletion or by substituting one nucleotide for another) or base modifications, such as chemical modifications of bases. The phrase “chemical modification” as used herein includes modifications that introduce chemical properties different from those seen in naturally occurring RNA, such as covalent modifications, such as the introduction of modified nucleotides (e.g., nucleotide analogs, or side groups not naturally present in such mRNA molecules).
[0241] Other suitable polynucleotide modifications to the RNA that may be incorporated into the compositions of this disclosure include, but are not limited to, 4'-thio-modified bases: 4'-thio-adenosine, 4'-thio-guanosine, 4'-thio-cytidine, 4'-thio-uridine, 4'-thio-5-methyl-cytidine, 4'-thio-pseudouridine and 4'-thio-2-thiouridine, pyridine-4-ketoribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-pseudouridine, 5-propynyluridine, 1-propynyl-pseudouridine, 5-nitrogenyluridine, 5-carboxymethyluridine, 1-propynyl-pseudouridine, 5-nitrogenyluridine, 1-propynyl-pseudouridine, 5-nitrogenyluridine, 1-carboxymethyl ... Sulfonate methyluridine, 1-Taurate methyl-pseudouridine, 5-Taurate methyl-2-thio-uridine, 1-Taurate methyl-4-thio-uridine, 5-Methyl-uridine, 1-Methyl-pseudouridine, 4-Thio-1-methyl-pseudouridine, 2-Thio-1-methyl-pseudouridine, 1-Methyl-1-deazon-pseudouridine, 2-Thio-1-methyl-1-deazon-pseudouridine, dihydrouridine, dihydropseudouridine, 2-Thio-dihydrouridine, 2-Thio-dihydropseudouridine, 2-Methoxyuridine, 2-Methoxy-4-thio-uridine, 4-Methoxy-pseudouridine, 4-Methoxy-2-thio-pseudouridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4- Acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudocytidine, pyrrole-cytidine, pyrrole-pseudocytidine, 2-thio-cytidine, 2-thio-5-methylcytidine, 4-thio-pseudocytidine, 4-thio-1-methyl-pseudocytidine, 4-thio-1-methyl-1-deazo-pseudocytidine, 1-methyl-1-deazo-pseudocytidine, zebularine, 5-aza-zebularine, 5-methylzebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methylcytidine, 4-methoxy-pseudocytidine, 4 -Methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deazo-adenine, 7-deazo-8-aza-adenine, 7-deazo-2-aminopurine, 7-deazo-8-aza-2-aminopurine, 7-deazo-2,6-diaminopurine, 7-deazo-8-aza-2,6-diaminopurine, 1-methyladenine, N6-methyladenine, N6-isopentenyladenine, N6-(cis-hydroxyisopentenyl)adenine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenine, N6-glycinylcarbamoyladenine, N6-threonylcarbamoyladenine, 2-methylthio-N6-threonylcarbamoyladenine, N6,N6-Dimethyladenosine, 7-methyladenosine, 2-methylthioadenosine and 2-methoxyadenosine, inosine, 1-methylinosine, wyosine, wybutosine, 7-deazoguanosine, 7-deazo-8-azaguanosine, 6-thioguanosine, 6-thio-7-deazoguanosine, 6-thio-7-deazo-8-azaguanosine, 7-methylguanosine, 6-thio-7-methylguanosine, 7-methylinosine, 6-methoxyguanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxoguanosine, 7-methyl-8-oxoguanosine, 1-methyl-6-thioguanosine, N2-methyl-6-thioguanosine and N2,N2-dimethyl-6-thioguanosine and combinations thereof. The term "modification" also includes, for example, incorporating non-nucleotide-linked or modified nucleotides into the mRNA sequence of this disclosure (e.g., modifying one or both of the 3' and 5' ends of an mRNA molecule encoding a functional protein or enzyme). Such modifications include adding bases to the mRNA sequence (e.g., including poly-A tails or longer poly-A tails), altering the 3' UTR or 5' UTR, causing the mRNA to complex with an agent (e.g., a protein or complementary nucleic acid molecule), and including elements that alter the structure of the RNA molecule (e.g., its formation of secondary structures).
[0242] In some implementations, RNA (e.g., ncRNA) includes a 5' cap structure. The 5' cap is typically added as follows: first, an RNA terminal phosphatase removes one terminal phosphate group from the 5' nucleotide, leaving two terminal phosphate groups; then, guanosine triphosphate (GTP) is added to the terminal phosphate ester via guanylate transferase, creating a 5'5'5 triphosphate bond; and then the 7-nitrogen of guanine is methylated by a methyltransferase. Examples of cap structures include, but are not limited to, m7G(5')ppp(5'(A,G(5')ppp(5')A and G(5')ppp(5')G. Naturally occurring cap structures contain a 7-methylguanosine bridging the 5' end of the first transcribed nucleotide via a triphosphate bridge, forming a dinucleotide cap of m7G(5')ppp(5')N, where N is any nucleoside. In vivo, the cap is added enzymatically; the cap is added to the cell nucleus and catalyzed by guanylate transferase. The cap is added to the 5' end of the RNA immediately after transcription initiation. The terminal nucleoside is typically guanosine and is oriented in the opposite direction to all other nucleotides, i.e., G(5')ppp(5')GpNpNp.
[0243] Additional cap analogs include, but are not limited to, chemical structures selected from the group consisting of: m7GpppG, m7GpppA, m7GpppC; unmethylated cap analogs (e.g., GpppG); dimethylated cap analogs (e.g., m2,7GpppG), trimethylated cap analogs (e.g., m2,2,7GpppG), dimethylated symmetrical cap analogs (e.g., m7Gpppm7G), or anti-reverse cap analogs (e.g., ARCA; m7,2'OmeGpppG, m72'dGpppG, m7,3'OmeGpppG, m7,3'dGpppG and their tetraphosphate derivatives), for example, as described in Jemielity, J. et al., “Novel 'anti-reverse' cap analogs with superior translational properties”, RNA, 9: 1108-1122 (2003).
[0244] Typically, the presence of a "tail" serves to protect RNA (e.g., ncRNA) from exonuclease degradation. Poly-A or poly-U tails are considered to stabilize natural messengers and synthetic sense RNA. Therefore, in some embodiments, long poly-A or poly-U tails can be added to RNA molecules to make the RNA more stable. A variety of techniques recognized in the art can be used to add poly-A or poly-U tails. For example, a long poly-A tail can be added to synthetic or in vitro transcribed RNA using a poly-A polymerase, as described in Yokoe et al., Nature Biotechnology. 1996; 14: 1252-1256. Transcription vectors can also encode long poly-A tails. Alternatively, a poly-A tail can be added by transcription directly from PCR products. In some instances, an RNA ligase can also be used to ligate a poly-A tail to the 3' end of sense RNA, as described in, for example, Molecular Cloning A Laboratory Manual, 2nd edition, edited by Sambrook, Fritsch, and Maniatis (Cold Spring Harbor Laboratory Press: 1991).
[0245] Typically, the length of the poly-A or poly-U tail can be at least about 10, 50, 100, 200, 300, 400, or at least 500 nucleotides. In some embodiments, the poly-A tail at the 3' end of the mRNA typically comprises about 10 to 300 adenosine nucleotides (e.g., about 10 to 200 adenosine nucleotides, about 10 to 150 adenosine nucleotides, about 10 to 100 adenosine nucleotides, about 20 to 70 adenosine nucleotides, or about 20 to 60 adenosine nucleotides). In some embodiments, the mRNA includes a 3' poly(C) tail structure. A suitable poly(C) tail at the 3' end of the mRNA typically comprises about 10 to 200 cytosine nucleotides (e.g., about 10 to 150 cytosine nucleotides, about 10 to 100 cytosine nucleotides, about 20 to 70 cytosine nucleotides, about 20 to 60 cytosine nucleotides, or about 10 to 40 cytosine nucleotides). PolyC tails can be added to or replace polyA or polyU tails.
[0246] The RNA (e.g., ncRNA) according to this disclosure can be synthesized using any of a variety of known methods. For example, the RNA according to this disclosure can be synthesized via in vitro transcription (IVT). Briefly, IVT typically uses a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system that may include DTT and magnesium ions, and a suitable RNA polymerase (e.g., T3, T7, or SP6 RNA polymerase), deoxyribonuclease I, pyrophosphatase, and / or RNase inhibitors. The exact conditions will vary depending on the specific application. An improved method for IVT of ncRNA is disclosed in Example 5 herein.
[0247] In one specific implementation (as illustrated in Example 6 herein), the ncRNA may include an MS2 modification, a specific RNA hairpin structure recognized in nature by an MS2-binding protein. This domain can help stabilize the ncRNA and improve editing efficiency. Other similar modifications are contemplated in this disclosure. In some instances, other such MS2-like domains may include, for example, those described in the following: Johansson et al., “RNA recognition by the MS2 phage coat protein”, SemVirol., 1997, Vol. 8(3): 176-185; Delebecque et al., “Organization of intracellular reactions with rationally designed RNA assemblies”, Science, 2011, Vol. 333: 470-474; Mali et al., “Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering”, Nat. Biotechnol., 2013, Vol. 31: 833-838; and Zalatan et al., “Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds”, Cell, 2015, Vol. 160: 339-350, each of which is incorporated herein by reference in its entirety. In some instances, other systems include the PP7 hairpin, which specifically recruits the PCP protein, and the “com” hairpin, which specifically recruits the Com protein, as described, for example, by Zalatan et al. The nucleotide sequence of the MS2 hairpin (or equivalently referred to as the “MS2 aptamer”) is: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO:19935).
[0248] polymerase In several embodiments, the chimeric gene editing composition comprises one or more polymerases, which may be reverse transcriptases (RTs). In such embodiments, the reverse transcriptase RT is responsible for converting template ncRNA into homologous templated msDNA, and for acting as an RNA-dependent DNA polymerase that functions at a nick site on the target DNA (see attached figure for details of the mechanism of action).
[0249] The reverse transcriptant RT that can be used with the gene editing system described herein can be any reverse transcriptant RT described in the art, including any of those described in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872 (each incorporated herein by reference in its entirety, including its sequence listing), or a sequence having at least 75%, 80%, 85%, 90%, 95%, 99%, or 100% sequence identity with a sequence from Table A of any of the aforementioned applications.
[0250] Reverse transcriptase (RT, also known as RNA-guided DNA polymerase) is an enzyme present in all three domains of life; it is a DNA polymerase that uses RNA as a template. The reverse transcriptase disclosed herein is used to transfer the template... msd RNA is reverse transcribed into single-stranded msDNA.
[0251] The reverse transcriptases or their functional domains that may be used in this disclosure include prokaryotic and eukaryotic RTs, provided that the RTs can function within the host to generate donor polynucleotide sequences from an RNA template (e.g., an RNA template from a reverse transcriptase ncRNA).
[0252] In some embodiments, suitable RT sequences (including amino acid sequences and encoded polynucleotide sequences) are provided in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0253] In some implementations, the nucleotide sequence of the natural or wild-type RT is modified, for example, using known codon optimization techniques, to optimize its expression in the desired host.
[0254] In some embodiments, the present disclosure uses the RT domain of reverse transcriptase, provided it is compatible with the engineered reverse transcriptase of the present disclosure. The domain may only include RNA-dependent DNA polymerase activity. In some embodiments, the RT domain is non-mutagenic, i.e., it does not cause mutations in donor polynucleotides (e.g., during the reverse transcriptase process). In some embodiments, the source of the RT domain may be a non-reverse transcriptase RT, such as a viral RT or a human endogenous RT. In some embodiments, the RT domain is a reverse transcriptase RT or a diversity-generating retroelement (DGR) RT. In some embodiments, the mutagenicity of the RT may be lower than that of the corresponding wild-type RT. In some embodiments, the RT is not mutagenic.
[0255] In some implementations, reverse transcriptase is produced by reverse transcriptase. ret Genes encode, and these genes may be accompanied by homologous genes. MSR and msd It is a locus that specifically recognizes the secondary structure of homologous ncRNA transcripts.
[0256] In some implementations, RTs can be derived from prokaryotic or eukaryotic cells. Most reverse transcriptases (80%) phylogenetically belong to three major lineages: group II introns, diversity generation reversal elements (DGRs), and reverse transcriptases. Other clades of RTs include abortion-infected (Abi) RTs, CRISPR-Cas-associated RTs, group II-like (G2L), group unknown (UG), and rvt elements.
[0257] In some implementations, the RT gene is a homologous RT, a retrotranscript RT from a species belonging to the same species or clade as the homologous RT, or a retrotranscript RT not belonging to the same clade as the homologous RT, such as an irrelevant RT or engineered RT. In some instances, non-retrotran-related RTs are RTs from group II introns, diversity generation reversal elements (DGRs), abortion infection (Abi) RTs, CRISPR-Cas-related RTs, group II-like (G2L) RTs, group unknown (UG) RTs, and rvt elements, for example, as described in: Mestere et al., Nucleic Acids Research, Vol. 48, No. 22, December 16, 2020, pp. 12632-12647; and Mestere et al., UG / Abi: "A Highly Diverse Family of Prokaryotic Reverse Transcriptases Associated With Defense Functions,” doi.org / 10.1101 / 2021.12.02.470933 (incorporated into this paper by reference).
[0258] In some embodiments, the RT is derived from a clade associated with a retrotran / retrotran-like sequence. In some embodiments, the RT is selected from the RTs provided in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. In some embodiments, the RT is not associated with the sequence identified in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0259] In the prokaryotic reverse transcriptase system, RT genes are typically located at ncRNAs ( MSR and msdDownstream of the RT locus. In engineered reverse transcriptases, the RT location may differ from that in natural or wild-type reverse transcriptases. In some implementations, the RT gene can be provided cis-, for example in... MSR locus or msd Upstream or downstream of the locus. In some embodiments, the RT gene is provided in trans form, for example, separately in a vector of the vector system described herein, wherein the gene encodes... MSR and msd The sequence of ncRNA is provided in different vectors of the vector system described herein.
[0260] In some implementations, RT is modified (e.g., inserted, deleted, and / or substituted one or more nucleotides) or codon optimized to enhance activity or continuity.
[0261] In some implementations, the recessive termination signal is removed from the RT, thereby allowing the generation of longer ssDNA.
[0262] In some implementations, the RT is derived from a reverse transcriptase encoding msDNA, for example, the RT described in US 6,017,737; US 5,849,563; US 5,780,269; US 5,436,141; US 5,405,775; US 5,320,958; CA 2,075,515; all of which are incorporated herein by reference in their entirety.
[0263] In some embodiments, the engineered reverse transcriptome also contains a polynucleotide (e.g., a DNA molecule) encoding a reverse transcriptase (RT) or a portion thereof. In some embodiments, the encoded RT or a portion thereof is capable of synthesizing msDNA. msd At least a portion of the DNA copy of the locus.
[0264] In some embodiments, the polynucleotide encoding RT (e.g., a DNA molecule) comprises, or has with respect to, the polynucleotides listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or has at least Polynucleotides with sequence identity of 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%.
[0265] In some embodiments, the polynucleotide encoding RT encodes a polypeptide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or having a content of at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, or at least 80% of the polypeptide listed in Table A of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872. A polypeptide with at least 5%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity; and / or a polypeptide of Table C in U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0266] In some implementations, the polynucleotide encoding RT does not contain the polynucleotides listed in Table X of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872.
[0267] Once translated, RT combines msd The ncRNA template downstream of the locus forms an RT-RNA complex and initiates reverse transcription of the RNA toward its 5′ end. Therefore, in some aspects, this disclosure relates to an engineered nuclease construct comprising: a. a non-coding RNA (ncRNA) comprising: i) encoding a multi-copy single-stranded DNA (msDNA) MSR RNA portion MSR loci; and ii) encoding msDNA msd RNA portion msd a. A locus; b. A heterologous nucleic acid inserted at or within a location selected from the following: msd loci MSR upstream of the locus msd upstream of the locus and msd Downstream of the locus; and c. Reverse transcriptase (RT) or its domain comprising: i) a polypeptide listed in Table A above, or a polypeptide having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with a polypeptide listed in Table C above. In some embodiments, RT does not contain a polypeptide listed in Table X above.
[0268] In some respects, this disclosure relates to an engineered nucleic acid-enzyme construct comprising: a) a non-coding RNA (ncRNA) comprising: i) encoding a multi-copy single-stranded DNA (msDNA). MSR RNA portion MSR loci; and ii) encoding msDNA msd RNA portion msd locus, b) a heterologous nucleic acid inserted at or within the following locations: msd Locus; MSR Upstream of the gene locus; msd Upstream of the locus; and msd Downstream of the locus; and c) reverse transcriptase (RT) or a portion thereof, wherein the RT is capable of synthesizing msDNA. msd A copy of at least a portion of the DNA of the locus, wherein the ncRNA and / or the RT is any of the contents of this disclosure as described herein.
[0269] In some respects, this disclosure relates to an engineered nucleic acid-enzyme construct comprising: a) a non-coding RNA (ncRNA) comprising: i) encoding a multi-copy single-stranded DNA (msDNA). MSR RNA portion MSR loci; and ii) encoding msDNA msd RNA portion msd a) A locus; b) A heterologous nucleic acid inserted at or within a location selected from the following: msd loci MSR Upstream of the locus msd upstream of the locus and msd Downstream of the locus; and c) reverse transcriptase (RT) or its domain: wherein the RT comprises: i) the RT listed in Table A above, or the RT having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with the RT listed in Table C above; and / or ii) the common sequences listed in Table C above; and wherein the RT does not contain sequences from Table X above.
[0270] In some embodiments of the nuclease constructs described herein, the ncRNA comprises: i) an ncRNA listed in Table B above, or ii) an ncRNA having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with an ncRNA listed in Table B above; and / or optionally, wherein the ncRNA is not an ncRNA derived from a reverse transcriptase of Table X above.
[0271] In some embodiments, RT is linked to components such as RNA-directed nucleases and non-RNA-directed nucleases. This linking can be achieved via peptide bonds or short linker peptides in the fusion protein. Suitable linker peptides include flexible linkers, such as those containing G or S repeat sequences, such as G4S (SEQ ID NO: 19511) repeat units or GS repeat units, having 1-20 repeats (SEQ ID NO: 19936), for example, 1, 2, 3, 4, 5, 6, 7, or 8 repeats.
[0272] In some implementations, RT is chemically linked or conjugated to RNA-guided and non-RNA-guided nucleases via non-peptide bonds. Such protein conjugates can be delivered directly to host cells, either together with the nucleic acid components of the engineered reverse transcriptase described herein, or separately.
[0273] In some implementations, RT is linked to DNA repair regulatory biomolecules (such as NHEJ peptide inhibitors).
[0274] nuclease The chimeric gene editing compositions disclosed herein contain a nuclease. In some embodiments, the nuclease is a double-stranded endonuclease or a nicking enzyme.
[0275] In some embodiments, the chimeric gene editing system comprises: (a) a nicking enzyme component, (b) a guide RNA compounded with said nicking enzyme component and directed to a target DNA sequence, (c) a polymerase component, and / or (d) a polymerase template sequence, wherein said polymerase template sequence is provided by a modified reverse transcriptase ncRNA containing the polymerase template sequence (referred herein to as "template-based ncRNA" or "tncRNA"). In several embodiments, said tncRNA comprises a structural configuration altered relative to wild-type reverse transcriptase ncRNA, said alteration including: (i) a linker that joins the 5' end of ncRNA a1 region to the 3' end of ncRNA a2 region; (ii) deletion of the wild-type msd region or a portion thereof; and / or (iii) replacement of the deleted msd sequence with a single-stranded RNA sequence containing the polymerase template. The chimeric gene editing system may also optionally include a reverse transcriptase RT to convert the template-based ncRNA into homologous msDNA, referred herein to as "template-based msDNA" or "tmsDNA".
[0276] The cutter enzyme component of the chimeric editing system described herein can be any programmable CRISPR cutter enzyme, including those based on: CRISPR type I, II, or III Cas nucleases; type II nucleases (e.g., Cas9); type V nucleases (e.g., Cpfl); or type VI nucleases (e.g., C2c2). Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8al, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csxl2), Cas1O, Cas1Od, CasF, CasG, CasH, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), and Cse4. (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 and Cul966, and their homologs or modified versions.
[0277] In some implementations, a type II CRISPR system Cas9 endonuclease is used. Cas9 nucleases from any species, or biologically active fragments, variants, analogs, or derivatives thereof that retain Cas9 endonuclease activity (i.e., catalyzing site-directed cleavage of DNA to produce double-strand breaks), can be used to perform genome modifications as described herein. Cas9 need not be physically derived from an organism, but can be prepared synthetically or recombinantly. In some instances, Cas9 sequences from various bacterial species are well known in the art and are listed in the National Center for Biotechnology Information (NCBI) database, for example, NCBI entries for Cas9 from *Streptococcus pyogenes* (…). Streptococcus pyogenes (WP002989955, WP_038434062, WP_011528583); Campylobacter jejuni ( Campylobacter jejuni (WP_022552435, YP 002344900), Campylobacter coli ( Campylobacter coli (WP 060786116); Campylobacter fetus ( Campylobacter fetus (WP 059434633); Corynebacterium ulcerans ( Corynebacterium ulcerans(NC_015683, NC_017317); Corynebacterium diphtheriae ( Corynebacterium diphtheria (NC_016782, NC_016786); Enterococcus faecalis ( Enterococcus faecalis (WP 033919308); Aphididae ( Spiroplasmasyrphidicola (NC 021284); Prevotella intermedia ( Prevotellaintermedia (NC017861); Taiwan Spiroplasma (NC 021846); Dolphin Streptococcus ( Streptococcusiniae (NC 021314); Baltic Bellerella ( Belliellabaltica (NC 018010); Curvularia truncatulata ( Psychroflexustorquisl (NC O 18721); Streptococcus thermophilus ( Streptococcus themophilus (YP 820832), Streptococcus mutans ( Streptococcus mutans (WP 061046374, WP 024786433); harmless Listeria ( Listeriainnocua (NP 472073); Listeria monocytogenes ( Listeriamonocytogenes (WP 061665472); Legionella pneumophila ( Legionella pneumophila (WP062726656); Staphylococcus aureus ( Staphylococcusaureus (WP_001573634); Tulazella Francisella ( Francisellatularensis ) (WP_032729892, WP_014548420), Enterococcus faecalis (WP 033919308); Lactobacillus rhamnosus ( Lactobacillusrhamnosus (WP 048482595, WP_032965177); and Neisseria meningitidis ( Neisseriameningitidis(WP_061704949, YP_002342100); all of the sequences (as entered on the date of submission of this application) are incorporated herein by reference in their entirety. These sequences or their variants (containing sequences with at least approximately 70-100% sequence identity, including any percentage of identity within this range, such as 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity) can be used for genome editing, as described herein and in some examples, such as Fonfara et al., (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al., (2015) J. Bacterid. 198(5): 797-807; Shmakov et al., (2015) Mol. Cell. 60(3):385- 397; and Chylinski et al., (2014) Nucleic Acids Res. 42(10):6091-6105) as described.
[0278] In another implementation, bacteria from the genus *Prevotella* are used. Prevotella ) and Francisella spp. ( FrancisellaCRISPR nuclease 1 (Cpfl or Cas12a). Cpfl is another class II CRISPR / Cas system RNA-guided nuclease, similar to Cas9 and usable in a similar manner. Unlike Cas9, Cpfl does not require tracrRNA and its guide RNA depends solely on crRNA, which provides the advantage of using shorter guide RNAs than Cas9 when targeting. Cpfl is capable of cleaving DNA or RNA. Compared to the G-rich PAM sites recognized by Cas9, the PAM sites recognized by Cpfl have the sequence 5′-YTN-3′ (where “Y” is pyrimidine and “N” is any nucleobase) or 5′-TTN-3′. Cpfl cleavage of DNA produces double-strand breaks, with sticky ends having 4 or 5 nucleotide overhangs. In some instances, additional information about Cpfl instances is disclosed in Ledford et al., (2015) Nature. 526 (7571):17-17; Zetsche et al., (2015) Cell.163 (3):759-771; Murovec et al., (2017) Plant Biotechnol. J. 15(8):917-926; Zhang et al., (2017) Front. Plant Sci. 8:177; Fernandes et al., (2016) Postepy Biochem.62(3):315-326 (incorporated hereby by reference).
[0279] C2c1 (Cas12b) is another usable class II CRISPR / Cas system RNA-guided nuclease. Similar to Cas9, C2c1 relies on both crRNA and tracrRNA to guide to the target site, for example, those disclosed in Shmakov et al., (2015) MolCell. 60(3):385-397; Zhang et al., (2017) Front Plant Sci. 8:177 (incorporated hereby by reference).
[0280] In one aspect, a programmable DNA-binding domain of a nucleic acid sequence can associate or complex with at least one guide nucleic acid (e.g., guide RNA or pegRNA) that localizes the DNA-binding domain to a DNA sequence containing a DNA strand complementary to the guide nucleic acid (i.e., the target strand), or a portion thereof (e.g., a spacer of guide RNA attached to a protospacer of a DNA target). In other words, the guide nucleic acid “programs” the DNA-binding domain (e.g., Cas9 or its equivalent) to localize and bind to a complementary sequence of a protospacer in DNA.
[0281] The lead editor described herein can use any suitable nucleic acid sequence with a programmable DNA-binding domain. In several embodiments, the nucleic acid sequence-programmable DNA-binding domain can be any Class 2 CRISPR-Cas system, including any type II, V, or VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a genome editing tool, the nomenclature used to describe and / or identify CRISPR-Cas enzymes (e.g., Cas9 and Cas9 orthologs) is also constantly evolving. In some instances, CRISPR-Cas nomenclature is discussed in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol. 1, No. 5, 2018, the entire contents of which are incorporated herein by reference.
[0282] Not bound by theory, the mechanisms of action of some CRISPR Cas enzymes covered herein include the step of forming an R-loop, through which the Cas protein induces the unwinding of a double-stranded DNA target, thereby separating the strands in the Cas protein-binding region. A guide RNA spacer then hybridizes with the “target strand” in a region complementary to the original spacer sequence of the DNA. In some embodiments, the Cas protein may include one or more nuclease activities that then cleave the DNA, leaving various types of damage. For example, the Cas protein may contain nuclease activity that can cleave a non-target strand at a first location and / or cleave the target strand at a second location. Depending on the nuclease activity, the target DNA may be cleaved to form a “double-strand break,” thereby cutting both strands. In other embodiments, the target DNA may be cleaved at only a single site, i.e., one strand of the DNA is “cut.” Exemplary Cas proteins with different nuclease activities include “Cas9 cleavage enzyme” (“nCas9”) and inactivated Cas9 (“dead Cas9” or “dCas9”) that do not have nuclease activity.
[0283] The following description of various Cas proteins that can be used in conjunction with the LNP delivery gene editing systems disclosed herein is in no way intended to limit the scope. The gene editing systems may comprise typical SpCas9, or any orthologous Cas9 protein, or any variant Cas9 protein, including any naturally occurring variants, mutants, or other engineered versions of Cas9 that are known or that can be prepared or evolved through directed evolution or other mutagenesis processes. In several embodiments, the Cas9 or Cas9 variant has cleavage enzyme activity, i.e., it cleaves only one strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variant has an inactive nuclease, i.e., a “dead” Cas9 protein. Other variant Cas9 proteins that may be used are those with a smaller molecular weight than typical SpCas9 (e.g., for ease of delivery) or those with modified or rearranged primary amino acid structures.
[0284] The gene editing system described herein may also include Cas9 equivalents, including Cas12a (Cpf1) and Cas12b1 proteins. Cas proteins available herein (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) may also contain various modifications that alter / enhance their PAM specificity. This disclosure covers any Cas9, Cas9 variant, or Cas9 equivalent that has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity with a reference Cas9 sequence, such as the reference SpCas9 canonical sequence of Streptococcus pyogenes M1 (accession number Q99ZW2) (SEQ ID NO: 2027). The Cas proteins covered in this document include CRISPR Cas9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or non-natural (e.g., engineered or recombinant), and may include Cas9 equivalents from any Class 2 CRISPR system (e.g., type II, type V, type VI), including Cas12a (Cpf1), Cas12e (CasX), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9, Cas13a (C2c2), Cas13d, Cas13c (C2c7), Cas13b (C2c6), and Cas13b. In some instances, other Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299) and Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol. 1, No. 5, 2018, the contents of which are incorporated herein by reference.
[0285] The terms “Cas9” or “Cas9 nuclease” or “Cas9 moiety” or “Cas9 domain” encompass any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of Cas9 (natural or engineered). The term Cas9 is not intended to be particularly limiting and may be referred to as “Cas9 or equivalent”. Exemplary Cas9 proteins are further described in the art and are incorporated herein by reference. As described herein, Cas9 nuclease sequences and structures are well known to those skilled in the art.In some instances, the Cas9 nuclease sequence and structure can be those described below: "Complete genome sequence of an M1 strain of Streptococcuspyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III." Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of these references are incorporated herein by reference.
[0286] In some embodiments, the polynucleotide-programmable nucleotide-binding domain of the nucleobase editor itself comprises one or more domains. In one embodiment, the polynucleotide-programmable nucleotide-binding domain comprises one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide-binding domain comprises an endonuclease or an exonuclease. In some embodiments, the endonuclease cleaves one strand of a double-stranded nucleobase molecule. In some embodiments, the endonuclease cleaves both strands of a double-stranded nucleobase molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a ribonuclease.
[0287] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain is capable of cleaving zero, one, or both strands of the target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide binding domain includes a nicking enzyme domain. The term "nicking enzyme" as used herein refers to a polynucleotide programmable nucleotide binding domain containing a nuclease domain capable of cleaving only one strand of a double-stranded nucleobase molecule (e.g., DNA). In some embodiments, the nicking enzyme is derived from the fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by introducing one or more mutations into an active polynucleotide programmable nucleotide binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain includes a nicking enzyme domain derived from Cas9.
[0288] In some embodiments, the Cas9-derived nickase has one or more mutations in the RuvC-1 domain. In one embodiment, the Cas9-derived nickase has a D10A mutation in the RuvC-1 domain. In some embodiments, the Cas9-derived nickase has one or more mutations in the REC Lobe domain. In one embodiment, the Cas9-derived nickase has N497A, R661A, and / or Q695A mutations in the REC Lobe domain. In some embodiments, the Cas9-derived nickase has one or more mutations in the HNH domain. In one embodiment, the Cas9-derived nickase has H840A, N863A, and / or D839A mutations in the HNH domain.
[0289] In some embodiments, in the spCas9-derived nickase, residue H840 retains catalytic activity and is thus capable of cleaving one strand of the nucleobase duplex. In some embodiments, the Cas9-derived nickase domain may contain the H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the Cas9-derived nickase domain may contain the N863A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nickase is derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide binding domain by removing all or part of the nuclease domains that are not desired for nickase activity. In some embodiments, when the polynucleotide programmable nucleotide binding domain contains a Cas9-derived nickase domain, the Cas9-derived nickase domain contains all or part of the deletion of the RuvC domain or the HNH domain.
[0290] In some embodiments, the nucleobase editing system is or includes a CRISPR-Cas editor or Cas9 disclosed and described in one or more of the following documents: U.S. Application Publications US2015 / 0045546A1, US2019 / 0264232A1, US2018 / 0258417A1, and PCT Publications WO2013141680A1 and WO2021173359A1, each of which is incorporated herein by reference in its entirety.
[0291] This document considers any of the above-described CRISPR-Cas editor implementations, or any variations, modifications, or derivatives thereof, delivered by the LNP system disclosed herein, for gene editing in cells, tissues, and / or organs under in vitro, ex vivo, or in vivo conditions. The various components described herein can be configured and delivered in any suitable manner. Any descriptions presented in this section are not intended to be strictly limiting.
[0292] Preview Editor The chimeric gene editing system described in this article may include a lead editor or one or more of its components, as well as a reverse transcriptase editing component.
[0293] Leader editing is a gene editing technique that allows for targeted insertions, deletions, and all transversions and transitions at point mutations in the target genome. Not bound by any particular theory, leader editing searches for and replaces endogenous sequences in target polynucleotides. A spacer sequence of the leader editing guide RNA (“PEgRNA” or “pegRNA”) recognizes and binds to the searched target sequence in the target strand of a double-stranded target polynucleotide (e.g., double-stranded target DNA). The leader editing complex creates a nick in the target DNA on the edit strand, which is the complementary strand to the target strand. The leader editing complex then initiates DNA synthesis using the free 3' end formed at the nick site on the edit strand, where the “primer binding site sequence” (PBS) of the PEgRNA complexes with the free 3' end, and a template for editing the PEgRNA is used as a template to synthesize single-stranded DNA (by reverse transcriptase). As used herein, a “primer binding site” is a single-stranded portion of the PEgRNA containing a region complementary to the PAM strand (i.e., the non-target strand or the edit strand). PBS is complementary or substantially complementary to the sequence on the PAM strand of the double-stranded target DNA immediately upstream of the nick site.
[0294] The term "lead editor (PE)" refers to a polypeptide or polypeptide component involved in lead editing, or any polynucleotide encoding said polypeptide or polypeptide component. In several embodiments, the lead editor includes a polypeptide domain having DNA-binding activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the lead editor also includes a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA-binding activity includes a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity includes a nicking enzyme or a fully active nuclease. As used herein, the term "nicking enzyme" refers to a nuclease capable of cleaving only one strand of a double-stranded DNA target. In some embodiments, the lead editor includes a polypeptide domain that is an inactive nuclease. In some embodiments, the polypeptide domain having programmable DNA-binding activity includes a nucleic acid-directed DNA-binding domain, such as a CRISPR-Cas protein, such as the Cas9 nicking enzyme, the Cpf1 nicking enzyme, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity includes a template-dependent DNA polymerase, such as an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the leader editor contains additional polypeptides involved in leader editing, such as a polypeptide domain with 5' endonuclease activity, such as a 5' endogenous DNA flap endonuclease (e.g., FEN1), to help drive the leader editing process toward the formation of the edited product. In some embodiments, the leader editor also contains RNA-protein recruitment polypeptides, such as the MS2 coat protein.
[0295] The primer editor may be engineered. In some embodiments, the polypeptide components of the primer editor are not naturally present in the same organism or cellular environment. In some embodiments, the polypeptide components of the primer editor may have different sources or originate from different organisms. In some embodiments, the primer editor comprises a DNA-binding domain and a DNA polymerase domain derived from different species. In some embodiments, the primer editor comprises a Cas polypeptide (DNA-binding domain) and a reverse transcriptase polypeptide (DNA polymerase) derived from different species. For example, the primer editor may comprise a Streptococcus pyogenes Cas9 polypeptide and a Moronis murine leukemia virus (M-MLV) reverse transcriptase polypeptide. In some embodiments, the primer editor comprises engineered five-mutant M-MLV RT. In some embodiments, a primer editor engineered to reduce size and improve efficiency is used.
[0296] In some embodiments, the polypeptide domain of the lead editor can be fused or linked via peptide linkers to form a fusion protein. In other embodiments, the lead editor comprises one or more polypeptide domains provided in trans form as separate proteins, which are capable of associating with each other via non-peptide linkages or via aptamers or recruitment sequences. For example, the lead editor may comprise a DNA-binding domain and a reverse transcriptase domain that associate with each other via an RNA-protein recruitment aptamer (e.g., the MS2 aptamer), which can be linked to PEgRNA. The lead editor polypeptide component may be wholly or partially encoded by one or more polynucleotides. In some embodiments, a single polynucleotide, construct, or vector encodes a lead editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or a portion of a domain of the lead editor, or a portion of the lead editor fusion protein. For example, the lead editor fusion protein may comprise an N-terminal portion fused to an inteptide-N and a C-terminal portion fused to an inteptide-C, each of which is individually encoded by an AAV vector.
[0297] Compared to the endogenous double-stranded target DNA sequence, the editing template may contain one or more desired nucleotide edits. Therefore, the newly synthesized single-stranded DNA also contains nucleotide edits encoded by the editing template. Through the removal of the editing target sequence from the editing strand of the double-stranded target DNA and DNA repair mechanisms, the newly synthesized single-stranded DNA replaces the editing target sequence and incorporates the desired nucleotide edits into the double-stranded target DNA.
[0298] In some implementations, the lead editing was first described in Anzalone et al., “Search-and-replace genome editing without double-strand breaks or donor DNA”, Nature, December 2019, 576 (7789): pp. 149-157, which is incorporated herein by reference in its entirety. In some instances, subsequent prime editing has been described and detailed in numerous subsequent publications, including but not limited to: (i) Liu et al., “Prime editing: a search and replace tool with versatile base changes”, YiChuan, 20 Nov 2022, 44(11): 993-1008; (ii) Lu C et al., “Prime Editing: AnAll-Rounder for Genome Editing”. Int J Mol Sci. 30 Aug 2022;23(17):9862; (iii) Velimirovic M, Zanetti LC, Shen MW, Fife JD, Lin L, Cha M, Akinci E, Barnum D, Yu T, Sherwood RI. Peptide fusion improves prime editing efficiency. Nat Commun. 18 Jun 2022;13(1):3512. doi: Peptide fusion improves prime editing efficiency. Nat Commun. 2022 Jun 18;13(1):3512. doi: 10.1038 / s41467-022-31270-y. PMID: 35717416; PMCID: PMC9206660; (v) Habib O, Habib G, Hwang GH, Bae S.Comprehensive analysis of prime editing outcomes in human embryonic stem cells. Nucleic Acids Res. January 25, 2022;50(2):1187-1197. doi: 10.1093 / nar / gkab1295. PMID: 35018468; PMCID: PMC8789035;(vi) Marzec M,. , Hensel G. Prime Editing: A New Way for Genome Editing. Trends Cell Biol. April 2020;30(4):257-259. doi:10.1016 / j.tcb.2020.01.004. Electronic publication on January 27, 2020. PMID: 32001098;(vii) TaoR, Wang Y, Jiao Y, Hu Y, Li L, Jiang L, Zhou L, Qu J, Chen Q, Yao S. Bi-PE:bi-directional priming improves CRISPR / Cas9 prime editing in mammalian cells. Nucleic Acids Res. June 24, 2022;50(11):6423-6434. doi: 10.1093 / nar / gkac506.PMID: 35687127; PMCID: PMC9226529;(viii) Nelson JW, Randolph PB, Shen SP, Everette KA, Chen PJ, Anzalone AV, An M, Newby GA, Chen JC, Hsu A, Liu DR. Engineered pegRNA improve prime editing efficiency. Nat Biotechnol. Mar 2022;40(3):402-410. doi: 10.1038 / s41587-021-01039-7. E-published on 4 October 2021. Erratum in: Nat Biotechnol. 8 December 2021;: PMID: 34608327; PMCID:PMC8930418;(ix) Doman JL, Sousa AA, Randolph PB, Chen PJ, Liu DR. Designing and executing prime editing experiments in mammalian cells. Nat Protoc. Nov 2022;17(11):2431-2468. doi: 10.1038 / s41596-022-00724-4. E-published on August 8, 2022.PMID: 35941224; PMCID: PMC9799714; (x) Jiao Y, Zhou L, Tao R, Wang Y, HuY, Jiang L, Li L, Yao S. Random-PE: an efficient integration of random sequences into mammalian genome by prime editing. Mol Biomed. 2021 Nov 18;2(1):36. doi: 2022 Apr;40(4):374-376. doi: 10.1016 / j.tibtech.2022.01.013. Electronic publication on February 10, 2022. PMID: 35153078; and (xii) Doman JL, Pandey S, Neugebauer ME, An M, Davis JR, Randolph PB, McElroy A, Gao XD, Raguram A, Richter MF, Everette KA, Banskota S, Tian K, Tao YA, Tolar J, Osborn MJ, Liu DR. Phage-assisted evolution and protein engineering yield compact, efficient prime editors. Cell. August 31, 2023; 186(18):3983-4002.e26. doi: 10.1016 / j.cell.2023.07.039. PMID: 37657419; PMCID: PMC10482982, all cited references are incorporated herein by reference.
[0299] In some instances, the lead editor has described and disclosed them in numerous published patent applications, the entire contents, amino acid sequences, nucleotide sequences, and all disclosures of each of which are incorporated herein by reference in their entirety: International Application No. PCT / US2022 / 074628, filed August 5, 2022; International Application No. PCT / US2022 / 074088, filed July 23, 2022; International Application No. PCT / US2022 / 073819, filed July 16, 2022; International Application No. PCT / US2022 / 035613, filed June 29, 2022; International Application No. PCT / US202... Application No. 2 / 036230; International Application No. PCT / US2022 / 032267 filed on June 3, 2022; European Application No. EP21707651.2 filed on February 19, 2021; International Application No. PCT / EP2022 / 062223 filed on May 5, 2022; U.S. Application No. 17 / 219,635 filed on March 31, 2021; International Application No. PCT / CN2022 / 080595 filed on March 14, 2022; International Application No. PCT / US2022 / 023175 filed on April 1, 2022; International Application No. PCT / US2022 / 021 filed on March 25, 2022. International Application No. 879; International Application No. PCT / US2022 / 020392, filed March 15, 2022; U.S. Application No. 17 / 219,672, filed March 31, 2021, now U.S. Patent No. 11,447,770, granted September 20, 2022; International Application No. PCT / CN2022 / 077097, filed February 21, 2022; International Application No. PCT / US2022 / 015260, filed February 4, 2022; International Application No. PCT / KR2022 / 001611, filed January 28, 2022; International Application No. PCT / IN2022 / 050017, filed January 7, 2022. International Application No. PCT / US2022 / 012054, filed January 11, 2022; U.S. Application No. 17 / 427,040, filed July 29, 2021, now U.S. Patent No. 11,384,353, granted July 12, 2022; International Application No. PCT / US2021 / 052097, filed September 24, 2021; International Application No. PCT / KR2021 / 017534, filed November 25, 2021; International Application No. PCT / CN2021 / 130059, filed November 11, 2021; International Application No. PCT / US2021 / 057908, filed November 3, 2021.International Application No. PCT / US2021 / 058079, filed November 4, 2021; International Application No. PCT / KR2021 / 013326, filed September 29, 2021; International Application No. PCT / US2021 / 052097, filed September 24, 2021; International Application No. PCT / KR2021 / 010740, filed August 12, 2021; U.S. Application No. 17 / 427,040, filed July 29, 2021; International Application No. PCT / US2021 / 044924, filed August 6, 2021; International Application No. PCT / KR2021 / 00979, filed July 28, 2021. Application No. 4; International Application No. PCT / US2021 / 031439 filed on May 7, 2021; International Application No. PCT / US2021 / 034996 filed on May 28, 2021; International Application No. PCT / KR2021 / 005244 filed on April 26, 2021; International Application No. PCT / KR2021 / 005031 filed on April 21, 2021; International Application No. PCT / US2020 / 023730 filed on March 19, 2020; International Application No. PCT / US2020 / 023713 filed on March 19, 2020; International Application No. PCT / EP filed on February 19, 2021 International Application No. 2021 / 054228; International Application No. PCT / US2020 / 067535 filed on December 30, 2020; International Application No. PCT / US2020 / 059149 filed on November 5, 2020; International Application No. PCT / US2020 / 055959 filed on October 16, 2020; International Application No. PCT / US2020 / 055156 filed on October 9, 2020; International Application No. PCT / US2020 / 023553 filed on March 19, 2020; International Application No. PCT / US2020 / 023583 filed on March 19, 2020; International Application No. PCT / US2020 / 023730, filed on March 19, 2020; International Application No. PCT / US2020 / 023721, filed on March 19, 2020; International Application No. PCT / US2020 / 023728, filed on March 19, 2020; International Application No. PCT / US2020 / 023732, filed on March 19, 2020; International Application No. PCT / US2020 / 023712, filed on March 19, 2020; International Application No. PCT / US2020 / 023725, filed on March 19, 2020; International Application No. PCT / US2020 / 023713, filed on March 19, 2020;International Application No. PCT / US2020 / 023727, filed March 19, 2020; International Application No. PCT / US2020 / 023724, filed March 19, 2020; International Application No. PCT / US2020 / 023583, filed March 19, 2020; International Application No. PCT / US2020 / 023723, filed March 19, 2020; International Application No. PCT / CN2020 / 074218, filed February 3, 2020; U.S. Application No. 15 / 616,756, filed June 7, 2017, now U.S. Patent No. 10,189,831, granted January 29, 2019; July 1, 2018 International application No. PCT / US2018 / 042040, filed on March 3; U.S. application No. 15 / 164,208, filed May 25, 2016, now U.S. Patent No. 10,150,955, issued December 11, 2018; International application No. PCT / US2017 / 050690, filed September 8, 2017; U.S. application No. 11 / 502,819, filed August 10, 2006, now U.S. Patent No. 9,783,791, issued October 10, 2017; and U.S. application No. 13 / 277,763, filed October 20, 2011, now U.S. Patent No. 9,458,484, issued October 4, 2016.
[0300] In some embodiments, the chimeric gene editing composition comprises a lead editing system or a polynucleotide encoding a lead editing system. In some embodiments, the product comprises a lead editing system component or a polynucleotide encoding a lead editing system component.
[0301] Leader editing is a versatile and precise genome editing method that uses a catalytically impaired Cas fused to an engineered reverse transcriptase (also known as a leader editor) to directly write new genetic information into a designated DNA site. This engineered reverse transcriptase can be programmed with a designated target site and a leader editing guide RNA (“pegRNA”) encoding the desired edit, for example, as described by Anzalone et al., Nature 2019. Leader editing bypasses the need for a DNA donor template by using a leader editor with nicking enzyme or catalytically impaired enzyme activity.
[0302] The lead editing system includes a lead editor. The lead editor (“PE”) contains a catalytically impaired Cas protein fused to an engineered reverse transcriptase that can precisely and permanently edit one or more target nucleobases in a target polynucleotide.
[0303] In some embodiments, the lead editor comprises an engineered Moroni rodent leukemia virus (“M-MLV”) reverse transcriptase (“RT”) fused to a Cas-H840A nickase (referred to as “PE2”). In some embodiments, the lead editor comprises an engineered M-MLV RT fused to a Cas9-H840A nickase. In some embodiments, the lead editor comprises an engineered M-MLV RT fused to a Streptococcus pyogenes Cas9 (spCas9)-H840A nickase. PE modifications include increased PAM flexibility to improve the usability of PE2 editing, expanding the coverage of targetable pathogens in the ClinVar database to 94.4% that can now be lead-edited.
[0304] In some embodiments, the lead editing system also includes a lead editing guide RNA (“pegRNA”). In some embodiments, the cargo contains pegRNA or a polynucleotide encoding pegRNA.
[0305] In some implementations, the lead editing system also includes a second guide RNA that targets the complementary strand, allowing the Cas9 nickase to also nick the unedited strand (referred to as "PE3"), which biases mismatch DNA repair toward the edited sequence. In some implementations, the second guide RNA is designed to recognize the complementary strand of the DNA only after PE3 editing has occurred (referred to as "PE3b"), which reduces insertion / deletion formation.
[0306] In some embodiments, the lead editing system comprises a uracil glycosidase inhibitor. In some embodiments, the lead editing system comprises a Cas9 protein fused to a uracil glycosidase inhibitor. In some embodiments, the cargo comprises a uracil glycosidase inhibitor or a polynucleotide encoding a uracil glycosidase inhibitor. In some embodiments, the cargo comprises a Cas9 protein fused to a uracil glycosidase inhibitor or a polynucleotide encoding a Cas9 protein fused to a uracil glycosidase inhibitor.
[0307] This document considers any of the above-described pilot editor embodiments or variations, modifications, or derivatives thereof, delivered by the LNP system disclosed herein, for gene editing in cells, tissues, and / or organs under in vitro, ex vivo, or in vivo conditions. The various components described herein can be configured and delivered in any suitable manner. Any descriptions presented in this section are not intended to be strictly limiting.
[0308] Guide RNA This disclosure further provides guide RNAs for use in editing methods based on the disclosed nucleic acid-programmable DNA-binding proteins (e.g., Cas9). This disclosure provides guide RNAs designed to recognize target sequences. Such gRNAs can be designed to have guide sequences (or “spacer”) that are complementary to the target sequence. Such gRNAs can be designed to have not only guide sequences that are complementary to the target sequence to be edited, but also a backbone sequence that specifically interacts with the nucleic acid-programmable DNA-binding protein.
[0309] In some embodiments, the guide RNA may be 15-100 nucleotides in length and contain at least 10, at least 15, or at least 20 adjacent nucleotide sequences complementary to the target nucleotide sequence. The guide RNA may contain a spacer sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 adjacent nucleotides complementary to the target nucleotide sequence. In some cases, the guide sequence is in the range of 17-30 nucleotides (nt) (e.g., 17-25, 17-22, 17-20, 19-30, 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some cases, the guide sequence is 17-25 nucleotides (nt) in length (e.g., 17-22, 17-20, 19-25, 19-22, 19-20, 20-25, or 20-22 nt). In some cases, the guide sequence is 17 or more nts in length (e.g., 18 or more, 19 or more, 20 or more, 21 or more, or 22 or more nts; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence is 19 or more nts in length (e.g., 20 or more, 21 or more, or 22 or more nts; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence is 17 nt in length. In some cases, the guide sequence is 18 nt long. In some cases, the guide sequence is 19 nt long. In some cases, the guide sequence is 20 nt long. In some cases, the guide sequence is 21 nt long. In some cases, the guide sequence is 22 nt long. In some cases, the guide sequence is 23 nt long.
[0310] In some cases, the length of the spacer sequence is 15 to 50 nucleotides (e.g., 15 nucleotides (nt) to 20 nt, 20 nt to 25 nt, 25 nt to 30 nt, 30 nt to 35 nt, 35 nt to 40 nt, 40 nt to 45 nt, or 45 nt to 50 nt).
[0311] Subject-guided RNAs can interact with target nucleic acids (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), single-stranded RNA (ssRNA), or double-stranded RNA (dsRNA)) in a sequence-specific manner via hybridization (i.e., base pairing).
[0312] The guide RNA can be modified to hybridize with any desired target sequence within the target nucleic acid (e.g., eukaryotic target nucleic acid, such as genomic DNA) (e.g., while also considering PAM, e.g., when targeting dsDNA). In some cases, the complementarity percentage between the spacer sequence of the guide and the target site of the target nucleic acid is 60% or higher (e.g., 65% or higher, 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the complementarity percentage between the spacer and the target site of the target nucleic acid is 80% or higher (e.g., 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the complementarity percentage between the spacer and the target site of the target nucleic acid is 90% or higher (e.g., 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the complementarity percentage between the spacer and the target site of the target nucleic acid is 100%.
[0313] In some cases, the percentage of complementarity between the spacer sequence and the target site of the target nucleic acid is 100% within a contiguous region of at least 5 nucleotides of the spacer. In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 6 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 7 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 8 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 9 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 10 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 11 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 12 nucleotides of the spacer.In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 13 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 14 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 15 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 16 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 17 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 18 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 19 nucleotides of the spacer.In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 20 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 21 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 22 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%).
[0314] In some cases, the percentage of complementarity between the spacer sequence and the target site of the target nucleic acid is 100% within a contiguous region of at least 5-10 nucleotides of the spacer. In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 7-12 nucleotides of the spacer. In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 8-13 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 9-14 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 10-15 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 11-16 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 12-17 nucleotides of the spacer.In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 13-18 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 14-19 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 15-20 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 16-21 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 17-22 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 18-23 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 19-24 nucleotides of the spacer.In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 20-25 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In other cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher within a contiguous region of at least 21-26 nucleotides of the spacer (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%). In some cases, the percentage of complementarity between the guide sequence and the target site of the target nucleic acid is 60% or higher (e.g., 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 97% or higher, 98% or higher, 99% or higher, or 100%) within a contiguous region of at least 22-27 nucleotides of the spacer.
[0315] In several embodiments, the guide RNA may have a scaffold or core region that complexes with a homologous nucleic acid programmable DNA-binding protein (e.g., CRISPR-Cas9 or Cas12a). In some cases, the guide scaffold may have two complementary nucleotide fragments that hybridize to form a double-stranded RNA duplex (dsRNA duplex). Thus, in some cases, the protein-binding segment of the guide RNA includes a dsRNA duplex. In some implementations, the dsRNA double-stranded region comprises 5–25 base pairs (bp) (e.g., 5–22, 5–20, 5–18, 5–15, 5–12, 5–10, 5–8, 8–25, 8–22, 8–18, 8–15, 8–12, 12–25, 12–22, 12–18, 12–15, 13–25, 13–22, 13–18, 13–15, 14–25, 14–22, 14–18, 14–15, 15–25, 15–22, 15–18, 17–25, 17–22, or 17–18 bp, such as 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the dsRNA double-stranded region comprises 6–15 base pairs (bp) (e.g., 6–12, 6–10, or 6–8 bp, such as 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the double-stranded region comprises 5 or more bp (e.g., 6 or more, 7 or more, or 8 or more bp). In some cases, the double-stranded region comprises 6 or more bp (e.g., 7 or more, or 8 or more bp). In some cases, not all nucleotides in the double-stranded region are paired, and therefore the double-stranded region may include a bump. The term "bump" as used herein refers to a nucleotide fragment (which may be a single nucleotide) that does not form a double helix, but whose 5' and 3' ends are surrounded by nucleotides that do form a double helix, and therefore the bump is considered part of the double-stranded region. In some cases, the dsRNA comprises one or more bumps (e.g., two or more, three or more, four or more bumps). In some cases, the dsRNA duplex includes two or more protrusions (e.g., three or more, four or more protrusions). In other cases, the dsRNA duplex includes one to five protrusions (e.g., one to four, one to three, two to five, two to four, or two to three protrusions).
[0316] Therefore, in some cases, the nucleotide fragments that hybridize with each other in the guide scaffold region to form the dsRNA double helix have 70%-100% complementarity with each other (e.g., 75%-100%, 80%-100%, 85%-100%, 90%-100%, 95%-100% complementarity). In some cases, the nucleotide fragments that hybridize with each other to form the dsRNA double helix have 70%-100% complementarity with each other (e.g., 75%-100%, 80%-100%, 85%-100%, 90%-100%, 95%-100% complementarity). In some cases, the nucleotide fragments that hybridize with each other to form the dsRNA double helix have 85%-100% complementarity with each other (e.g., 90%-100%, 95%-100% complementarity). In some cases, the nucleotide segments that hybridize to form a dsRNA double helix have 70%-95% complementarity with each other (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity). In other words, in some cases, the dsRNA double helix comprises two nucleotide segments that have 70%-100% complementarity with each other (e.g., 75%-100%, 80%-100%, 85%-100%, 90%-100%, 95%-100% complementarity). In some cases, the dsRNA double helix comprises two nucleotide segments that have 85%-100% complementarity with each other (e.g., 90%-100%, 95%-100% complementarity). In some cases, the dsRNA double helix comprises two nucleotide segments that have 70%-95% complementarity with each other (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity).
[0317] In several implementations, the scaffold region of the guide RNA may also include one or more (1, 2, 3, 4, 5, etc.) mutations relative to the naturally occurring scaffold region. For example, in some cases, base pairs may be maintained, while the nucleotides that produce base pairs from each segment may differ. In some cases, the double-stranded region of the subject guide RNA may include more paired bases, fewer paired bases, smaller bumps, larger bumps, fewer bumps, more bumps, or any suitable combination thereof, compared to the naturally occurring double-stranded region of the (naturally occurring guide RNA).
[0318] Examples of various guide RNAs can be found in the art, and in some cases, similar variations to those introduced into the Cas9 guide RNA can also be introduced into the guide RNA of this disclosure (e.g., mutations in the dsRNA double-stranded region, extensions at the 5' or 3' ends to increase stability for interaction with another protein, etc.). In some instances, the guide RNA may be the guide RNA described in the following literature: Jinek et al., Science. Aug 17, 2012; 337(6096):816-21; Chylinski et al., RNA Biol. May 2013; 10(5):726-37; Ma et al., Biomed Res Int. 2013; 2013:270805; Hou et al., Proc Natl Acad Sci US A. Sep 24, 2013; 110(39):15644-9; Jinek et al., Elife. 2013; 2:e00471; Pattanayak et al., Nat Biotechnol. Sep 2013; 31(9):839-43; Qi et al., Cell. Feb 28, 2013; 152(5): 1173-83; Wang et al., Cell. May 9, 2013; 153(4):910-8; Auer et al., Genome Res. October 31, 2013; Chen et al., Nucleic Acids Res. November 1, 2013; 41(20):119; Cheng et al., Cell Res. October 2013; 23(10):1163-71; Cho et al., Genetics. November 2013; 195(3):1177-80; DiCarlo et al., Nucleic Acids Res. April 2013; 41(7):4336-43; Dickinson et al., Nat Methods. October 2013; 10(10):1028-34; Ebina et al., Sci Rep. 2013;3:2510;Fujii et al., Nucleic Acids Res. 1 Nov 2013;41(20):1187;Hu et al., Cell Res. 1 Nov 2013;23(11):1322-5;Jiang et al., Nucleic Acids Res. 1 Nov 2013;41(20):1188;Larson et al., NatProtoc. 1 Nov 2013;8(11):2180-96;Mali et al., Nat Methods.October 2013; 10(10):957-63; Nakayama et al., Genesis. December 2013; 51(12):835-43; Ran et al., Nat Protoc. November 2013; 8(ll):2281-308; Ran et al., Cell. September 12, 2013; 154(6):1380-9; Upadhyay et al., G3 (Bethesda). December 9, 2013; 3(12):2233-8; Walsh et al., Proc NatlAcad Sci US A. September 24, 2013; 110(39):15514-5; Xie et al., Mol Plant. October 9, 2013; Yang et al., Cell. September 12, 2013; 154(6):1370-9; Briner et al., Mol Cell.October 23, 2014; 56(2):333-9; and U.S. patents and patent applications: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; 8,697,359; 20140068797; 20140170753; 20140179006; 20140179770; 201401 86843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046; 20140273037; 20140273226; 20140273230; 201402 73231; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 201403 References 35620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of the above-mentioned references are hereby incorporated in their entirety by citation.
[0319] In one embodiment, the guide RNA (including pegRNA) considered herein comprises non-naturally occurring nucleic acids and / or non-naturally occurring nucleotides and / or nucleotide analogs, and / or chemically modified. Non-naturally occurring nucleic acids may include, for example, a mixture of natural and non-naturally occurring nucleotides. Non-naturally occurring nucleotides and / or nucleotide analogs may be modified at the ribose, phosphate, and / or base moieties. In one embodiment of this disclosure, the guide RNA (including pegRNA) component nucleic acid comprises ribonucleotides and non-ribonucleotides. In one such embodiment, the guide RNA (including pegRNA) component comprises one or more ribonucleotides and one or more deoxyribonucleotides. In one embodiment of this disclosure, the guide RNA (including pegRNA) component comprises one or more non-naturally occurring nucleotides or nucleotide analogs, such as nucleotides having phosphate-thioester bonds, locked nucleic acid (LNA) nucleotides comprising a methylene bridge between the 2' and 4' carbons of the ribose ring, or bridged nucleic acid (BNA).
[0320] Other examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluorine analogs. Other examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromouridine, pseudouridine, inosine, and 7-methylguanosine. Examples of chemical modifications to coRNA include, but are not limited to, the incorporation of 2'-O-methyl (M), 2'-O-methyl 3'-thiophosphate (MS), S-bound ethyl (cEt), or 2'-O-methyl 3'-thiopACE (MSP) at one or more terminal nucleotides. Compared to unmodified oRNA components, such chemically modified oRNA components may exhibit increased stability and activity, although on-target and off-target specificity is unpredictable. In some implementations, the modified nucleotides may be those described in the following literature: Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi: 10.1038 / nbt.3290, published online June 29, 2015; Ragdarm et al., 0215, PNAS, E7110-E7111; Allenson et al., J. Med. Chem.2005, 48:901-904; Bramsen et al., Front. Genet., 2012, 3:154; Deng et al., PNAS, 2015, 112: 11870-11875; Sharma et al., MedChemComm., 2014, 5: 1454-1471; Li et al., Nature Biomedical Engineering, 2017, 1, 0066 D01: 10.1038 / s41551-017-0066). In one embodiment, the 5' and / or 3' ends of the guide RNA (including pegRNA) component are modified with various functional portions, said functional portions including fluorescent dyes, polyethylene glycol, cholesterol, proteins, or detection tags, for example, as described in Kelly et al., 2016, J. Biotech. 233:74-83. In one embodiment, the guide RNA (including pegRNA) component comprises a ribonucleotide in a region that binds to a target sequence and one or more deoxyribonucleotides and / or nucleotide analogs in a region that binds to a nucleic acid programmable DNA-binding protein (e.g., Cas9 nickase).
[0321] In one embodiment, deoxyribonucleotides and / or nucleotide analogs are incorporated into the engineered guide RNA (including pegRNA) component structure. In one embodiment, 3-5 nucleotides at the 3' or 5' end of the guide RNA (including pegRNA) component are chemically modified. In one embodiment, minor modifications, such as 2'-F modifications, are introduced only in the seed region. In one embodiment, a 2'-F modification is introduced at the 3' end of the guide RNA (including pegRNA) component. In one embodiment, three to five nucleotides at the 5' and / or 3' end of the reRNA component are chemically modified with 2'-O-methyl (M), 2'-O-methyl 3'-thiophosphate (MS), S-restricted ethyl (cEt), or 2'-O-methyl 3'-thiophosphate (MSP). Such modifications can improve genome editing efficiency, for example, those described in Hendel et al., Nat. Biotechnol. (2015) 33(9): 985-989. In one embodiment, all phosphodiester bonds of the guide RNA (including pegRNA) component are replaced with phosphate thioesters (PS) to enhance the level of gene disruption. In one embodiment, more than five nucleotides at the 5' and / or 3' ends of the guide RNA (including pegRNA) component are chemically modified with 2'-O-Me, 2'-F, or S-restricted ethyl (cEt). Such chemically modified guide RNA (including pegRNA) components can mediate the enhancement of gene disruption levels, for example, those described in Ragdarm et al., 0215, PNAS, E7110-E7111. In one embodiment of this disclosure, the guide RNA (including pegRNA) component is modified to include a chemical moiety at its 3' and / or 5' ends. Such moiety includes, but is not limited to, amines, azides, alkynes, thiols, dibenzocyclooctyne (DBCO), or rhodamine. In some embodiments, the chemical moiety is conjugated to the guide RNA (including pegRNA) component via a linker, such as an alkyl chain. In one embodiment, the chemical portion of the modified nucleic acid component can be used to attach a guide RNA (including pegRNA) component to another molecule, such as DNA, RNA, protein, or nanoparticles. Such chemically modified guide RNA (including pegRNA) components can be used to identify or enrich cells typically edited by the gene-editing systems described herein.
[0322] In some instances, other guide RNA (including pegRNA) modifications have been described in Kim, DY, Lee, JM, Moon, SB, et al., Efficient CRISPR editing with a hypercompact Cas12f1 and engineered guide RNA delivered by adeno-associated virus. Nat Biotechnol 40,94–102 (2022) in.
[0323] Therefore, in several aspects of this disclosure, the guide RNA (including pegRNA) is modified at one or more locations within the molecule: MS1, the inner penta-(uridine) (UUUUU) sequence in the tracrRNA; MS2, the 3' end of the crRNA; MS3, the 'stem 1' region of the tracrRNA; MS4, the tracrRNA–crRNA complementary region; and MS5, the 'stem 2' region of the tracrRNA.
[0324] Several aspects of this disclosure provide methods and compositions for improving the stability of guide RNA (including pegRNA) via chemical modification, for example, the methods and compositions described in the following literature: Braasch, DA, Jensen, S., Liu, Y., Kaur, K., Arar, K., White, MA et al., (2003). RNA interference in mammalian cells by chemically-modified RNA. Biochemistry 42, 7967–7975. doi:10.1021 / bi0343774. Chiu, YL and Rana, TM (2003). siRNA function in RNAi: achemical modification analysis. RNA 9, 1034–1048. doi: 10.1261 / rna.5103703. Behlke, MA (2008). Chemical modification of siRNA for in vivo use. Oligonucleotides18, 305–319. doi: 10.1089 / oli.2008.0164. Bennett, CF and Swayze, EE (2010). RNA targeting therapeutics: molecular mechanisms ofantisense oligonucleotides as a therapeutic platform. Annu. Rev. Pharmacol. Toxicol. 50, 259–293. doi: 10.1146 / annurev.pharmtox.010909.105654. Deleavey, GF and Damha, MJ (2012). Designing chemically modified oligonucleotides for targeted gene silencing. Chem. Biol. 19, 937–954. doi: 10.1016 / j.chembiol.2012.07.011. Lennox, KA and Behlke, MA (2020). Chemicalmodifications in RNA interference and CRISPR / Cas genome editing agents. Methods Mol. Biol. 2115, 23–55. doi: 10.1007 / 978-1-0716-0290-4_2.
[0325] In some instances, Hendel et al. improved guide RNA stability by chemically modifying the ends of gRNA to reduce degradation by exonucleases and RNases. See Hendel, A., Bak, RO, Clark, JT, Kennedy, AB, Ryan, DE, Roy, S. et al., (2015a). Chemically modified guideRNA enhance CRISPR-Cas genome editing in human primary cells. Nat. Biotechnol. 33, 985–989. doi: 10.1038 / nbt.3290. Chemical modification of gRNA enables more efficient and safer gene editing in primary cells suitable for clinical applications.
[0326] In some instances, reviews of chemical modification types are provided in Allen, Daniel, et al., “Using Synthetically Engineered Guide RNA to Enhance CRISPR Genome Editing Systems in Mammalian Cells.” Frontiers in genome editing Volume 2, 617910. January 28, 2021, doi:10.3389 / fgeed.2020.617910.
[0327] Therefore, in several embodiments of this disclosure, the genome editing system includes a guide RNA (including pegRNA) and also includes one or more chemical modifications selected from (but not limited to) those described in Allen et al.
[0328] In exemplary embodiments, chemical modifications to the guide RNA (including pegRNA) include modifications to the ribose ring and phosphate backbone of the guide RNA (including pegRNA), and modifications at the 2'OH position include 2'-O-Me, 2'-F, and 2'F-ANA. More extensive ribose modifications include combined modifications of 2'F-4'-Cα-OMe and 2',4'-di-Cα-OMe at the 2' and 4' carbons. Phosphodiester modifications include sulfide-based thiophosphate (PS) or acetate-based phosphonoacetate alterations. The combination of ribose and phosphodiester modifications has been replaced by, for example, formulations such as 2'-O-methyl-3'-thiophosphate (MS) or 2'-O-methyl-3'-thioPACE (MSP) and 2'-O-methyl-3'-phosphonoacetate (MP) RNA. Locked and non-locked nucleotides, such as locked nucleic acids (LNA), bridged nucleic acids (BNA), S-restricted ethyl (cEt), and non-locked nucleic acids (UNA), are examples of sterically hindered nucleotide modifications. Modifications that form phosphodiester bonds between the 2' and 5' carbons (2',5'-RNA) of adjacent RNAs and that form butane 4-carbon chains between adjacent RNAs have been described.
[0329] In some embodiments, the ncRNA and guide RNA can be delivered as a single molecule, i.e., the guide RNA is fused to the 5' and / or 3' end of the ncRNA. In some embodiments, the ncRNA may have guide RNA at both ends.
[0330] In other implementations, the guide RNA and ncRNA may be provided and / or delivered as separate components. Separating the guide RNA from the ncRNA can improve editing efficiency.
[0331] In other implementations, the ncRNA-gRNA fusion can be co-delivered with a separate guide RNA.
[0332] B. Delivery systems and delivery methods Overview In another aspect, this disclosure provides compositions for transferring and / or expressing a reverse transcriptase-based gene editing system, for example, under in vitro, ex vivo, and in vivo conditions. In yet another aspect, this disclosure provides cell delivery compositions and methods, including compositions (e.g., plasmids) for passive and / or active transport to cells, delivery via virus-based recombinant vectors (e.g., AAV and / or lentiviral vectors), delivery via non-viral systems (e.g., liposomes and LNPs), and delivery via virus-like particles of the reverse transcriptase-based gene editing system described herein. Depending on the delivery system employed, the reverse transcriptase-based gene editing system described herein can be delivered in the form of DNA (e.g., plasmids or DNA-based viral vectors), RNA (e.g., guide RNA and mRNA delivered via LNPs), mixtures of DNA and RNA, proteins (e.g., virus-like particles), mixtures of DNA or RNA and proteins, and ribonucleoprotein (RNP) complexes. Any suitable combination of methods for delivering the components of the reverse transcriptase-based gene editing system disclosed herein can be employed.
[0333] Reverse transcriptome-based gene editing systems and / or their components can be delivered via any known delivery system, such as those described above, including (a) vector-free (e.g., electroporation), (b) viral delivery systems, and (c) non-viral delivery systems. Viral delivery systems include expression vectors, adeno-associated virus (AAV) vectors, retroviral vectors, lentiviral vectors, etc. The expression constructs can be replicated in living cells or can be synthetically prepared. Non-viral delivery systems include, but are not limited to, lipid particles (e.g., lipid nanoparticles (LNPs)), non-lipid nanoparticles, exosomes, liposomes, microcells, viral particles, stable nucleic acid-lipid particles (SNALP), lipid complexes / polymer complexes, DNA nanowires, gold nanoparticles, iTOP, streptococcal hemolysin O (SLO), multifunctional enveloped nanodevices (MEND), lipid-coated mesoporous silica particles, inorganic nanoparticles, and polymer delivery technologies (e.g., polymer-based particles).
[0334] In some instances, nucleic acid modalities, including the delivery of RNA therapeutics, can be further described in the following: Paunovska K, Loughrey D, Dahlman JE. Drug delivery systems for RNAtherapeutics. Nat Rev Genet. May 2022;23(5):265-280. doi: 10.1038 / s41576-021-00439-4. E-published on January 4, 2022. PMID: 34983972; PMCID: PMC8724758; Hong CA, Nam YS. Functional nanostructures for effective delivery of smallinterfering RNA therapeutics. Theranostics. September 19, 2014;4(12):1211-32. doi:10.7150 / thno.8491. PMID: 25285170; PMCID: PMC4183999; Liu F, Wang C, Gao Y, LiX, Tian F, Zhang Y, Fu M, Li P, Wang Y, Wang F. Current Transport Systems andClinical Applications for Small Interfering RNA (siRNA) Drugs. Mol DiagnTher. 2018 Oct; 22(5):551-569. doi: Int J MolSci. 2022 Feb 22;23(5):2408. doi: 10.3390 / ijms23052408.PMID: 35269550; PMCID: PMC8909959; Zhang M, Hu S, Liu L, Dang P, Liu Y, Sun Z, Qiao B, Wang C. Engineered exosomes from different sources for cancer-targeted therapy. Signal Transduct Target Ther. 2023 Mar 15;8(1):124. doi: 10.1038 / s41392-023-01382-y. PMID: 36922504; PMCID: PMC10017761; Hastings ML, Krainer AR. RNAtherapeutics. RNA. 2023 Apr;29(4):393-395. doi: 10.1261 / rna.079626.123.PMID: 36928165; PMCID: PMC10019368; Miele E, Spinelli GP, Miele E, Di Fabrizio E, Ferretti E, Tomao S, Gulino A. Nanoparticle-based delivery of smallinterfering RNA: challenges for cancer therapy. Int J Nanomedicine. 2012;7:3637-57. doi: 10.2147 / IJN.S23696. Electronic publication on July 20, 2012. PMID: 22915840; PMCID: PMC3418108. Each of these references is incorporated herein by reference in its entirety.
[0335] Gene editing systems based on engineered reverse transcriptases (or vectors containing them) can be introduced into any type of cell, including any cell from prokaryotic, eukaryotic, or archaea organisms, including bacteria, archaea, fungi, protists, plants (e.g., monocots and dicots); and animals (e.g., vertebrates and invertebrates). Examples of animals that can be transfected with engineered reverse transcriptase-based editing systems include, but are not limited to, vertebrates such as fish, birds, mammals (e.g., human and non-human primates, farm animals, pets, and laboratory animals), reptiles, and amphibians.
[0336] Gene editing systems based on engineered reverse transcriptases (or components thereof) can be introduced into single cells or populations of cells. Cells derived from tissues, organs, and biopsies, as well as recombinant cells, genetically modified cells, cells from in vitro cultured cell lines, and artificial cells (e.g., nanoparticles, liposomes, polymer vesicles, or microcapsules encapsulating nucleic acids) can all be transfected using gene editing systems based on engineered reverse transcriptases.
[0337] Gene editing systems based on engineered reverse transcriptases (or components thereof) can be introduced into cell fragments, cell components, or organelles (e.g., mitochondria in animal and plant cells, and chloroplasts (e.g., chloroplasts) in plant cells and algae).
[0338] Cells can be cultured or expanded after transfection using a gene editing system based on engineered reverse transcriptase.
[0339] Methods for introducing nucleic acids into host cells are well known in the art. Common methods include chemically induced transformation, typically using divalent cations (e.g., CaCl2), dextran-mediated transfection, polybrene-mediated transfection, lipofectamine and LT-1-mediated transfection, electroporation, protoplast fusion, encapsulation of nucleic acids in liposomes, and direct microinjection of nucleic acids containing the Cas12a editing system into the cell nucleus, for example, methods described in the following literature: Sambrook et al., (2001) Molecular Cloning, a laboratory manual, 3rd edition, ColdSpring Harbor Laboratories, New York; Davis et al., (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw-Hill; and Chu et al., (1981) Gene 13:197; all of which are incorporated herein by reference in their entirety.
[0340] Plant cells can also be targeted by the reverse transcriptase-based gene editing systems (or components thereof) disclosed herein. In some instances, the methods used for genetic transformation of plant cells may be those described in the following literature: US2022 / 0145296 and US Patent Nos. 8,575,425, 7,692,068, 8,802,934, and 7,541,517; Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12; Bates, GW (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; Christou, P. (1995) Euphytica 85:13-27; Tzfira et al., (2004) TRENDS in Genetics 20:375-383; Yao et al., (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; and Jones et al., (2005) PlantMethods 1:5, all of which are incorporated herein by reference in their entirety.
[0341] According to conventional methods, transformed plant cells can be grown into transgenic organisms, such as plants, for example, those described in McCormick et al., (1986) Plant Cell Reports 5:81-84.
[0342] Plant materials that can be transformed using the reverse transcriptase-based gene editing system (or components thereof) described herein include plant cells, plant protoplasts, plant cell tissue cultures of regenerable plants, plant callus tissue, plant masses, and intact plant cells in plants or plant parts, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, kernels, spikes, rachis, husks, stems, roots, root tips, anthers, etc. Progeny, variants, and mutants of regenerated plants are also included within the scope of this disclosure, provided that these parts contain genetic modifications introduced by the reverse transcriptase-based gene editing system. Further, processed plant products or byproducts retaining the genetic modifications introduced by the reverse transcriptase-based gene editing system are provided.
[0343] The reverse transcriptase-based gene editing system described herein can be used to produce transgenic plants with desired phenotypes, including but not limited to increased disease resistance (e.g., increased resistance to viruses, bacteria, or fungi), increased insect resistance, increased drought resistance, increased yield, and altered fruit ripening characteristics, sugar and oil composition, and color.
[0344] In some implementations involving reverse transcriptase-based gene editing systems, reverse transcriptase... MSR Gene, msd Genes and / or ret Genes can be expressed from vectors in vitro, for example, in an in vitro transcription system. The resulting ncRNA or msDNA can be isolated before packaging and / or formulation for direct delivery to host cells. For example, isolated ncRNA or msDNA can be packaged / formulated in a delivery medium (e.g., lipid nanoparticles as described in other sections).
[0345] In some implementations involving reverse transcriptase-based gene editing systems, reverse transcriptase... MSR Gene, msd Genes and / or ret Genes are expressed in vivo via vectors within cells. Reverse transcripts can be expressed using a single vector or multiple separate vectors. MSR Gene, msd Genes and / or ret Genes are introduced into cells to generate msDNA in host subjects.
[0346] In other implementations, the reverse transcriptase MSR Gene, msd Genes and / or ret Genes, and any other components of the reverse transcriptase-based genome editing system described herein (e.g., trans-director RNA, programmable nucleases (e.g., trans-director RNA)), can be expressed in vivo from RNA delivered to cells. Reverse transcriptases can be delivered using a single vector or multiple separate vectors. MSR Gene, msd Genes and / or ret Genes are introduced into cells to generate msDNA in host subjects.
[0347] Vectors and / or nucleic acid molecules encoding recombinant reverse transcriptase-based genome editing systems or components thereof may include control elements operatively linked to a reverse transcriptase sequence, allowing for the generation of msDNA in vitro or in vivo in a subject species. For example, in embodiments involving reverse transcriptase-based gene editing systems, the reverse transcriptase... MSR Gene, msd Genes and / or ret The gene can be operatively linked to a promoter to allow expression of the reverse transcriptase and / or msDNA product. In some embodiments, a heterologous sequence encoding the desired target product (e.g., a polynucleotide encoding a polypeptide or regulatory RNA, a donor polynucleotide for gene editing, or a protospacer DNA for molecular recording) can be inserted. MSR Genes and / or msd In genes.
[0348] In some implementations, the reverse transcriptase-based gene editing system is generated by a vector system containing one or more vectors.
[0349] Many vectors can be used in vectors or vector systems, including but not limited to linear polynucleotides, polynucleotides associated with ionic or amphoteric compounds, plasmids, and viruses.
[0350] Whole RNA format In several embodiments, the reverse transcriptase-based gene editing system (or components thereof) disclosed herein can be delivered in a “whole RNA” format. As used herein, the term “whole RNA” means that all components of the reverse transcriptase editing system (e.g., reverse transcriptase RT, programmable nuclease, sgRNA, and ncRNA) are delivered and / or administered in RNA form (e.g., coding RNA or non-coding RNA). In some embodiments, the RNA components can be delivered to cells and / or tissues by direct means (e.g., electroporation or transfection). In other embodiments, the RNA components can be delivered to cells and / or tissues by delivery media (e.g., LNP or liposomes).
[0351] In various embodiments, the reverse transcriptase editing system described herein may comprise a coding RNA (e.g., linear or circular mRNA) encoding a reverse transcriptase (e.g., any RT from Table X or Table A above), a coding RNA (e.g., linear or circular mRNA) encoding a programmable nuclease (e.g., Cas9, Cas12a, or TnpB nuclease), a reverse transcriptase ncRNA (e.g., ncRNA from Table B above), and a guide RNA for targeting the programmable nuclease to a specific desired target sequence.
[0352] In some embodiments, the RT and nuclease components may encode the same coding RNA molecule. The protein may also be expressed by a separate coding RNA molecule. In other embodiments, the RT and nuclease components may be fused together into a single fusion polypeptide having an RT domain and a nuclease domain optionally linked by a linker.
[0353] Additionally, in some embodiments, the ncRNA and guide RNA may be fused together into a single RNA molecule. For example, the guide RNA may be located at the 5' end of the ncRNA. In other embodiments, the guide RNA may be located at the 3' end of the ncRNA. In some embodiments, the ncRNA may contain guide RNA at both the 3' and 5' ends.
[0354] In other implementations, the ncRNA and the guide RNA can be separate molecules, i.e., delivered separately.
[0355] In other implementations, the reverse transcriptase editing system may include both an ncRNA-guide RNA fusion and additional guide RNA provided as a separate molecule.
[0356] In several embodiments, different RNA components of the whole RNA reverse transcriptase system can be combined and administered (e.g., directly or within a delivery medium) in different ratios. In some embodiments, the ratio of such RNA components or species can be expressed as a molar ratio.
[0357] For example, the molar ratio of RT-encoded RNA to nuclease-encoded RNA can be approximately 1:1, 1:1.5, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:12, 1:15, or 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0358] In another example, the molar ratio of nuclease-encoded RNA to RT-encoded RNA can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0359] In another example, the molar ratio of ncRNA or ncRNA-guide RNA fusion to guide RNA alone can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0360] In another example, the molar ratio of the guide RNA alone to ncRNA or ncRNA-guide RNA fusion can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0361] In another example, the molar ratio of ncRNA to guide RNA alone can be approximately 1:1, 1:1.5, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:12, 1:15, or 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0362] In another example, the molar ratio of guide RNA to ncRNA can be approximately 1:1, 1:1.5, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:12, 1:15, or 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0363] In another instance, the molar ratio (as the case may be) of the encoding RNA (e.g., encoding RT and / or nuclease) to ncRNA or ncRNA-guide RNA fusion can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0364] In another example, the molar ratio of the coding RNA encoding the retrotran RT to the ncRNA or ncRNA-guide RNA fusion (depending on the case) can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0365] In another example, the molar ratio (depending on the case) of the coding RNA encoding the programmable nuclease to the ncRNA or ncRNA-guide RNA fusion can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0366] In another example, the molar ratio of the coding RNA for the reverse transcriptase RT or nuclease to the guide RNA alone can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0367] In another example, the molar ratio of the guide RNA alone to the RNA encoding the reverse transcriptase RT or nuclease can be approximately 1:1, approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40.
[0368] In some embodiments, the amount of ncRNA-sgRNA relative to RT mRNA is increased. In some embodiments, the RT mRNA:ncRNA-sgRNA ratio is approximately 1:1.5, approximately 1:2, approximately 1:2.5, approximately 1:3, approximately 1:4, approximately 1:5, approximately 1:6, approximately 1:7, approximately 1:8, approximately 1:9, approximately 1:10, approximately 1:12, approximately 1:15, or approximately 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40. In some embodiments, the RT-Cas9 (or Cas9-RT) fusion is encoded by mRNA. In some implementations, the RT-Cas9 mRNA:ncRNA-sgRNA ratio is approximately 1:1.5, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:12, 1:15, or 1:20. Commonly used ranges include 1:1 to 1:2, 1:1.5 to 1:4, 1:2 to 1:4, 1:2 to 1:8, 1:2 to 1:10, 1:3 to 1:9, 1:3 to 1:12, 1:3 to 1:15, 1:4 to 1:8, 1:4 to 1:12, 1:4 to 1:20, 1:5 to 1:10, 1:5 to 1:15, 1:5 to 1:20, 1:10 to 1:20, or 1:10 to 1:40. In some implementations, multiple gene loci are targeted, so ncRNA-sgRNA comprises a mixture of ncRNA-sgRNA species, and the same ratio and range apply.
[0369] RNA-protein or DNA-protein format In several embodiments, the reverse transcriptase-based gene editing system (or components thereof) disclosed herein may comprise a combination of proteins, DNA, and / or RNA for delivery. For example, certain components of the reverse transcriptase editing system (e.g., sgRNA and ncRNA) are delivered in RNA form, and other components (e.g., reverse transcriptase RT, programmable nuclease) are delivered in protein form.
[0370] In some embodiments, the RNA or DNA components are delivered to cells or tissues via the same method as the protein components (e.g., electroporation or transfection, LNP, liposomes, exosomes, etc.). In other embodiments, the RNA or DNA components are delivered to cells or tissues via different types of delivery media.
[0371] In some embodiments, the RNA or DNA component is delivered simultaneously with the protein component. In some embodiments, the RNA or DNA component is delivered separately from the protein component, for example, sequentially.
[0372] Viral vector delivery In several implementations, the reverse transcriptase-based gene editing system described herein can be delivered in a viral vector.
[0373] Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus (AAV) vectors, retroviral vectors, and lentiviral vectors. Expression constructs can replicate in living cells or can be synthesized synthetically.
[0374] In some embodiments, nucleic acids comprising a reverse transcriptase-based gene editing system (or components thereof) are under the transcriptional control of a promoter. In some embodiments, the promoter is capable of initiating transcription of an operatively linked coding sequence via RNA polymerase I, II, or III.
[0375] Exemplary promoters for mammalian cell expression include the SV40 early promoter, CMV promoters (e.g., the CMV immediate early promoter, as described in, for example, U.S. Patent Nos. 5,168,062 and 5,385,839, which are incorporated herein by reference in their entirety), the mouse mammary tumor virus LTR promoter, the adenovirus major late promoter (AdMLP), and the herpes simplex virus promoter. Other non-viral promoters, such as promoters derived from mouse metallothionein genes, may also be used for mammalian expression.
[0376] In some instances, exemplary promoters for plant cell expression include the CaMV 35S promoter (as described, for example, in Odell et al., 1985, Nature 313:810-812); the rice actin promoter (as described, for example, in McElroy et al., 1990, Plant Cell 2:163-171); the ubiquitin promoter (as described, for example, in Christensen et al., 1989, Plant Mol. Biol. 12:619-632; and Christensen et al., 1992, Plant Mol. Biol. 18:675-689); the pEMU promoter (as described, for example, in Last et al., 1991, Theor. Appl. Genet. 81:581-588); and the MAS promoter (as described, for example, in Velten et al., 1984, EMBO J. 3:2723-2730).
[0377] In additional embodiments, the reverse transcriptase-based vector may also contain a tissue-specific promoter so that expression only begins after it has been delivered to a specific tissue. Non-limiting exemplary tissue-specific promoters include the B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase-1 promoter, endothelial glycoprotein promoter, fibronectin promoter, Flt-1 promoter, GFAP promoter, GPIIb promoter, ICAM-2 promoter, INF-b promoter, Mb promoter, Nphsl promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter.
[0378] In some instances, these and other promoters can be obtained from or incorporated into commercially available plasmids using techniques described above, as used by Sambrook et al.
[0379] In some implementations, one or more enhancer elements are used in conjunction with a promoter to increase the expression level of the construct. Examples include early SV40 gene enhancers (as described in Dijkema et al., EMBOJ (1985) 4:761), enhancers / promoters derived from Rous sarcoma virus long terminal repeats (LTRs) (as described in Gorman et al., Proc. Natl. Acad. Sci. USA (1982b) 79:6777), and elements derived from human CMV (as described in Boshart et al., Cell (1985) 41:521), such as elements included in CMV intron A sequences. All such sequences are incorporated herein by reference.
[0380] In one embodiment, an expression vector for expressing a reverse transcriptase-based gene editing system (or a component thereof) comprises a promoter operatively linked to a polynucleotide encoding said component. Components of the reverse transcriptase-based gene editing system can be configured as individual gene transcripts or fusion constructs. For example, a nuclease component can be fused with a reverse transcriptase component. In another example, an ncRNA component can be fused with a guide RNA component. In yet another example, a nuclease component can be fused with a reverse transcriptase component, but the guide RNA and ncRNA are independent. In other embodiments, the guide RNA and ncRNA components can be fused, but the reverse transcriptase and nuclease components are provided separately. Any functional combination of fusion and non-fusion components is covered.
[0381] In some implementations, the vector or vector system also includes a transcription terminator / polyadenylation signal. Examples of such sequences include, but are not limited to, those derived from SV40 (such as those described above by Sambrook et al.), and bovine growth hormone terminator sequences (such as those described in U.S. Patent No. 5,122,458).
[0382] Additionally, 5′-UTR sequences can be placed near the coding sequence to further enhance expression. Such sequences may include UTRs containing internal ribosome entry sites (IRES). Including IRES allows translation of one or more open reading frames from the vector. In some instances, IRES elements attract eukaryotic ribosome translation initiation complexes and promote translation initiation, as described, for example, in Kaufman et al., Nuc. Acids Res. (1991) 19:4485-4490; Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229:295-298; Rees et al., BioTechniques (1996) 20:102-110; Kobayashi et al., BioTechniques (1996) 21:399-402; and Mosser et al., BioTechniques (199722 ISO-161)c. In some instances, multiple IRES sequences are known and include sequences derived from various viruses, such as leader sequences derived from piconemaviruses, such as encephalomyocarditis virus (EMCV) UTRs (e.g., described in Jang et al., Virol. (1989) 63:1651-1660), polio leader sequences, hepatitis A virus leaders, hepatitis C virus IRES, human rhinovirus type 2 IRES (e.g., described in Dobrikova et al., Proc. Natl. Acad. Sci. (2003) 100(251:15125-151301)), IRES elements from foot-and-mouth disease virus (e.g., described in Ramesh et al., Nucl. Acid Res. (1996) 24:2697-2700), and giardiavirus IRES (e.g., Garlapati et al., J Biol. Chem.). (As described in (2004) 279(51):3389-33971), etc.In some instances, various non-viral IRES sequences were also used in this paper, including but not limited to IRES sequences from yeast, as well as type 1 human angiotensin II receptor IRES (e.g., as described in Martin et al., Mol. Cell Endocrinol. (2003) 212:51-61), fibroblast growth factor IRES (FGF-1 IRES and FGF-2 IRES, e.g., as described in Martineau et al., (2004) Mol. Cell. Biol. 24(17): 7622-7635), vascular endothelial growth factor IRES (e.g., as described in Baranick et al., (2008) Proc. Natl. AcadSci. USA 105(12):4733-4738, Stein et al., (1998) Mol. Cell. Biol. 18(6):3112-3119, Bert et al., (2006) RNA 12(6): (described in 1074-1083) and insulin-like growth factor 2 (IRES) (e.g., as described in Pedersen et al., (2002) Biochem. J. 363(Pt l):37-44).
[0383] These elements are commercially available via plasmids sold by companies such as Clontech (Mountain View, CA), Invivogen (San Diego, CA), Addgene (Cambridge, MA), and GeneCopoeia (Rockville, MD). See also IRESite: The database of experimentally verified IRES structures (iresite.org). IRES sequences can be included in vectors, for example, to express various phage recombinant proteins for recombination engineering or RNA-directed nucleases (e.g., Cas9) for HDR, in combination with reverse transcriptase from an expression cassette.
[0384] In some implementations, a polynucleotide encoding a viral self-cleaving 2A peptide (e.g., T2A peptide) can be used to allow the production of multiple protein products (e.g., Cas9, phage recombinant proteins, reverse transcriptases) from a single vector or transcription unit under a single promoter. One or more 2A linker peptides can be inserted between the coding sequences in a polycistronic construct. The self-cleaving 2A peptide allows for the production of co-expressed proteins at equimolar levels from the polycistronic construct. In some instances, 2A peptides derived from various viruses may be used, including but not limited to those derived from foot-and-mouth disease virus, equine rhinitis virus type A, Jhosea asigna virus, and porcine teschovirus, such as those described in the following literature: Kim et al., (2011) PLoS One 6(4): el8556; Trichas et al., (2008) BMC Biol. 6:40; Provost et al., (2007) Genesis 45(10): 625-629; Furler et al., (2001) Gene Ther. 8(11):864-873; all of which are incorporated herein by reference.
[0385] In some embodiments, the expression construct contains a plasmid suitable for transforming a bacterial host. Many bacterial expression vectors are known to those skilled in the art, and the selection of an appropriate vector is a matter of choice. Bacterial expression vectors include, but are not limited to, pACYC177, pASK75, pBAD, pBADM, pBAT, pCal, pET, pETM, pGAT, pGEX, pHAT, pKK223, pMal, pProEx, pQE, and pZA31. In some instances, the bacterial plasmid may contain antibiotic selection markers (e.g., resistance to ampicillin, kanamycin, erythromycin, carbenicillin, streptomycin, or tetracycline), the lacZ gene (β-galactosidase producing blue pigment from an x-gal substrate), fluorescent markers (e.g., GFP. mCherry), or other markers for selecting transforming bacteria, such as those described above by Sambrook et al.
[0386] In other embodiments, the expression construct comprises a plasmid suitable for transforming yeast cells. Yeast expression plasmids typically contain a yeast-specific origin of replication (ORI) and nutrient selection markers (e.g., HIS3, URA3, LYS2, LEU2, TRP1, METIS, ura4+, leul+, ade6+), antibiotic selection markers (e.g., kanamycin resistance), fluorescent markers (e.g., mCherry), or other markers for selecting yeast cells for transformation. Yeast plasmids may further contain components that allow shuttle movement between the bacterial host (e.g., E. coli) and the yeast cell. Many different types of yeast plasmids are available, including yeast integration plasmids (Yip), which lack ORI and integrate into the host chromosome via homologous recombination; yeast replication plasmids (YRp), which contain autonomous replication sequences (ARS) and can replicate independently; yeast centromere plasmids (YCp), which are low-copy vectors containing a portion of the ARS and a portion of the centromere sequence (CEN); and yeast augmentation plasmids (YEp), which are high-copy-number plasmids containing a 2-micrometer circle (natural yeast plasmid) fragment that allows for stable reproduction of 50 or more copies per cell.
[0387] In other implementations, the expression construct does not contain plasmids suitable for transforming yeast cells.
[0388] In other implementations, the expression construct comprises a virus or engineered construct derived from a viral genome. Numerous virus-based systems have been developed for the transfer of genes into mammalian cells. Among some examples, virus-based systems include adenoviruses, retroviruses (g-retroviruses and lentiviruses), poxviruses, adeno-associated viruses, baculoviruses, and herpes simplex viruses, for example those described in the following literature: Wamock et al., (2011) Methods Mol. Biol. 737:1-25; Walther et al., (2000) Drugs 60(2):249-271; and Lundstrom (2003) Trends Biotechnol. 21(3): 117-122 (incorporated herein by reference in its entirety). Certain viruses are capable of entering cells via receptor-mediated endocytosis to integrate into the host cell genome and stably and efficiently express viral genes, making them attractive candidates for the transfer of foreign genes into mammalian cells.
[0389] For example, retroviruses provide a convenient platform for gene delivery systems. Selected sequences can be inserted into vectors and packaged into retroviral particles using techniques known in the art. The recombinant virus can then be isolated and delivered in vivo or in vitro to the cells of a recipient. In some instances, retroviral systems may be described in the following literature: U.S. Patent No. 5,219,740; Miller and Rosman (1989) BioTechniques 7:980-990; Miller, AD (1990) Human Gene Therapy 1:5-14; Scarpa et al. (1991) Virology 180:849-852; Bums et al. (1993) Proc. Natl. Acad. Sci. USA 90:8033-8037; Boris-Lawrie and Temin (1993) Cur. Opin. Genet. Develop. 3:102-109; and Ferry et al. (2011) Curr. Pharm. Des. 17(24): 2516-2527. Lentivirals are a class of retroviruses that are particularly useful for delivering polynucleotides to mammalian cells because they can infect both dividing and non-dividing cells, as described in the following literature: Lois et al., (2002) Science 295:868-872; Durand et al., (2011) Viruses 3(2): 132-159; (incorporated hereby by reference).
[0390] Many adenoviral vectors have also been described. Unlike retroviruses that integrate into the host genome, adenoviruses persist outside the chromosome, thereby minimizing the risk associated with insertional mutagenesis.
[0391] In addition, various adeno-associated virus (AAV) vector systems have been developed for gene delivery. In some instances, AAV vectors can be readily constructed using techniques described in the following literature: U.S. Patents 5,173,414 and 5,139,941; International Publications WO 92 / 01070 (published January 23, 1992) and WO 93 / 03769 (published March 4, 1993); Lebkowski et al., Molec. Cell. Biol. (1988) 8:3988-3996; Vincent et al., Vaccines 90 (1990) (Cold Spring Harbor Laboratory Press); Carter, BJ. Current Opinion in Biotechnology (1992) 3:533-539; Muzyczka, N. Current Topics in Microbiol and Immunol. (1992) 158:97-129; Kotin, RM. Human Gene Therapy (1994). 5:793-801; Shelling and Smith, Gene Therapy (1994) 1:165-169; and Zhou et al., J. Exp. Med. (1994) 179:1867-1875.
[0392] In some instances, another vector system that can be used to deliver nucleic acids encoding components of the Cas12a editing system is the enteric-administered recombinant poxvirus vaccine described by Small, Jr., PA et al. (U.S. Patent No. 5,676,950, issued October 14, 1997, which is incorporated herein by reference).
[0393] Other viral vectors include those derived from the poxvirus family, including vaccinia virus and fowlpox virus. For example, a vaccinia virus recombinant expressing a target nucleic acid molecule (e.g., the Cas12a editing system) can be constructed as follows: First, DNA encoding a specific nucleic acid sequence is inserted into an appropriate vector, adjacent to the vaccinia promoter and flanked by a vaccinia DNA sequence, such as the sequence encoding thymidine kinase (TK). This vector is then used to transfect cells co-infected with vaccinia. Homologous recombination is used to insert a gene encoding the target sequence into the viral genome along with the vaccinia virus promoter. The resulting TK recombinants can be selected by culturing cells in the presence of 5-bromodeoxyuridine and selecting viral plaques resistant to it.
[0394] In some implementations, fowlpox viruses, such as chickenpox virus and canarypox virus, can also be used to deliver target nucleic acid molecules. The use of fowlpox vectors is particularly desirable in humans and other mammalian species because members of the genus *Fungi* can only replicate efficiently in susceptible avian species and are therefore not infectious in mammalian cells. Methods for generating recombinant fowlpox viruses are known in the art and employ genetic recombination, as described above with respect to the generation of vaccinia virus, for example, the methods described in WO 91 / 12882; WO 89 / 03429; and WO 92 / 03545.
[0395] In some instances, molecular conjugation vectors, such as the adenovirus chimeric vectors described in Michael et al., J. Biol. Chem. (1993) 268:6866-6869 and Wagner et al., Proc. Natl. Acad. Sci. USA (1992) 89:6099-6103, can also be used for gene delivery.
[0396] Members of the genus Alphavirus, such as, but not limited to, vectors derived from Sindbis virus (SIN), Semliki forest virus (SFV), and Venezuelan Equine Encephalitis virus (VEE), may also be used as viral vectors to deliver the polynucleotides disclosed herein. In some instances, Sindbis virus-derived vectors may be those described in the following publications: Dubensky et al., (1996) J.Virol. 70:508-519; and International Publications WO 95 / 07995 and WO 96 / 17072; and U.S. Patent No. 5,843,723, issued December 1, 1998, to Dubensky, Jr., TW et al., and U.S. Patent No. 5,789,245, issued August 4, 1998, to Dubensky, Jr., T. W., each of which is incorporated herein by reference. In some instances, chimeric alphavirus vectors containing sequences derived from Sindbis virus and Venezuelan equine encephalitis virus are particularly preferred, for example, those described in the following literature: Perri et al., (2003) J. Virol. 77: 10394-10403 and International Publications WO 02 / 099035, WO 02 / 080982, WO 01 / 81609 and WO 00 / 61772, each of which is incorporated herein by reference in its entirety.
[0397] Vaccinia-based infection / transfection systems can be conveniently used to provide inducible transient expression of target nucleic acids (e.g., Cas12a editing systems) in host cells. In this system, cells are first infected in vitro with a vaccinia virus recombinant encoding a phage T7 RNA polymerase. This polymerase exhibits fine specificity because it transcribes only templates carrying a T7 promoter. Following infection, cells are transfected with the target nucleic acid, driven by the T7 promoter. The polymerase expressed in the cytoplasm by the vaccinia virus recombinant transcribes the transfected DNA into RNA. This method provides high-level, transient, cytoplasmic production of large quantities of RNA, as described, for example, in Elroy-Stein and Moss, Proc. Natl. Acad. Sci. USA (1990) 87:6743-6747; Fuerst et al., Proc. Natl. Acad. Sci. USA (1986) 83:8122-8126.
[0398] In other methods of nucleic acid delivery using vaccinia or fowlpox virus recombinants or other viral vectors, an amplification system can be used that leads to high levels of expression upon introduction into the host cell. Specifically, the T7 RNA polymerase promoter preceding the coding region of t...
Claims
1. A non-coding RNA (ncRNA) variant, said ncRNA variant comprising a reference reverse transcriptase ncRNA having one or more modifications, The reference reverse transcriptase ncRNA, from 5' to 3', comprises: an a1 region, a first branch of guanosine, msr, msd, and an a2 region; and The one or more modifications include: (i) the bonding of the a1 region to the a2 region; (ii) the deletion of at least a portion of the msr; (iii) the deletion of at least a portion of the msd; (iv) the addition of a single-stranded RNA containing a polymerase template; or (v) the addition of an RNA motif.
2. The ncRNA variant of claim 1, wherein one or more modifications comprise a link between the a1 region and the a2 region.
3. The ncRNA variant of claim 1 or 2, wherein the one or more modifications comprise a link between the a1 region and the a2 region, wherein the link between the a1 region and the a2 region is formed via a linker connecting the 5' end of the a1 region and the 3' end of the a2 region.
4. The ncRNA variant of claim 1 or 2, wherein the one or more modifications comprise a link between the a1 region and the a2 region, wherein the link between the a1 region and the a2 region is formed by directly joining the 5' end of the a1 region to the 3' end of the a2 region without a linker.
5. The ncRNA variant of any one of claims 1-4, wherein the a1 region and the a2 region are partially or completely complementary to each other, and the a1 region and the a2 region form a stem-loop structure after the bonding.
6. The ncRNA variant of any one of claims 1-5, wherein the ncRNA variant further comprises a second-branched guanosine.
7. The ncRNA variant according to any one of claims 2-6, wherein the bonding between the a1 region and the a2 region causes the ncRNA variant to become circular.
8. The ncRNA variant of any one of claims 2-7, wherein the one or more modifications further comprise a break between the msd and the a2 region, thereby generating the ncRNA variant comprising, from the 5' to 3' direction: the a2 region, the link between the a1 region and the a2 region, the a1 region, the first branched guanosine, the msr, and the msd.
9. The ncRNA variant of any one of claims 1-8, wherein the one or more modifications comprise at least a deletion of the msr.
10. The ncRNA variant of claim 9, wherein the portion of the msr missing includes a spacer between two stem-loops in the msr or a portion of the spacer.
11. The ncRNA variant of claim 10, wherein the deleted portion of the msr includes a portion of the spacer, thereby preserving 1-15 base pairs remaining between the two stem-loops in the msr compared to the reference reverse transcriptase ncRNA.
12. The ncRNA variant of claim 11, wherein the deletion, compared to the reference reverse transcriptase ncRNA, preserves 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs remaining between the two stem-loops in the msr.
13. The ncRNA variant of claim 10, wherein the portion of the msr missing from the reference reverse transcriptase ncRNA includes the entire spacer between the two stem-loops in the msr.
14. The ncRNA variant of any one of claims 9-13, wherein the portion of the msr missing includes a stem-loop or a portion thereof in the msr.
15. The ncRNA variant of any one of claims 1-14, wherein the one or more modifications comprise at least a deletion of the msd.
16. The ncRNA variant of claim 15, wherein the deleted portion of the msd comprises a spacer or a portion thereof located between the stem loop in the msd and the a2 region.
17. The ncRNA variant of claim 16, wherein the deleted portion in the msd comprises a portion of a spacer located between the stem loop in the msd and the a2 region, thereby preserving 1-15 base pairs remaining between the stem loop in the msd and the a2 region compared to the reference reverse transcript ncRNA.
18. The ncRNA variant of claim 17, wherein the deletion retains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs compared to the reference reverse transcriptase ncRNA.
19. The ncRNA variant of claim 16, wherein the deleted portion in the msd comprises the entire spacer between the stem loop in the msd and the a2 region.
20. The ncRNA variant of any one of claims 16-19, wherein the deleted portion in the MSD further comprises a stem-loop in the MSD, whereby the deletion comprises the entire spacer between the stem-loop in the MSD and the a2 region, as well as the deletion of the stem-loop in the MSD.
21. The ncRNA variant of any one of claims 1-20, wherein the one or more modifications comprise the addition of the single-stranded RNA comprising the polymerase template.
22. The ncRNA variant of claim 21, wherein the one or more modifications comprise adding the single-stranded RNA containing the polymerase template and deleting at least a portion of the msd, and the single-stranded RNA containing the polymerase template is added to the deleted portion of the msd.
23. The ncRNA variant of claim 21 or 22, wherein the single-stranded RNA comprising the polymerase template contains a pair of homologous arms specifically targeting the target locus.
24. The method of claim 23, wherein each of the homologous arms comprises 5-200 nucleotides.
25. The method of claim 24, wherein each of the homologous arms comprises 5 to 100 nucleotides, 5 to 50 nucleotides, 10 to 50 nucleotides, 25 to 50 nucleotides, 5 to 25 nucleotides, 10 to 25 nucleotides, 10 to 20 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, or 40 to 50 nucleotides.
26. The ncRNA variant of any one of claims 21-25, wherein the single-stranded RNA comprising the polymerase template further comprises a donor sequence for integration into a target locus.
27. The ncRNA of claim 26, wherein the donor sequence comprises 2-100 nucleotides, 2-50 nucleotides, 2-25 nucleotides, 5-25 nucleotides, 5-15 nucleotides, 5-10 nucleotides, 10-30 nucleotides, 10-20 nucleotides, or 1-10 nucleotides.
28. The ncRNA variant of any one of claims 1-27, wherein the one or more modifications comprise the addition of the RNA motif.
29. The ncRNA variant of claim 28, wherein the RNA motif is an MS2 stem-loop or 3' tail of a U7 small nuclear RNA (snRNA).
30. The ncRNA variant of any one of claims 1-29, wherein the one or more modifications comprise: a. A combination of (i) and (ii): the bonding between region a1 and region a2, and the absence of at least a portion of the MSR; b. A combination of (i) and (iii): the bonding between region a1 and region a2, and the absence of at least a portion of the MSD; c. Combination of (i) and (iv): the bonding of the a1 region to the a2 region, and the addition of the single-stranded RNA containing the polymerase template; d. Combination of (i) and (v): the bonding of the a1 region to the a2 region, and the addition of the RNA motif; e. A combination of (ii) and (iii): at least a portion of the msr is missing, and at least a portion of the msd is missing; f. Combination of (ii) and (iv): Deleting at least a portion of the msr and adding the single-stranded RNA containing the polymerase template; g. A combination of (ii) and (v): deleting at least a portion of the msr and adding the RNA motif at the 5' or 3' end; or h. Combination of (iv) and (v): Add the single-stranded RNA sequence containing the polymerase template, and add the RNA motif at the 5' or 3' end.
31. The ncRNA variant of any one of claims 1-30, wherein the one or more modifications comprise: a. A combination of (i), (ii) and (iii): the bonding between region a1 and region a2, the absence of at least a portion of the MSR, and the absence of at least a portion of the MSD; b. A combination of (i), (ii) and (iv): the bonding of the a1 region to the a2 region, the deletion of at least a portion of the msr, and the addition of the single-stranded RNA containing the polymerase template; c. A combination of (i), (iii), and (iv): the linking of the a1 region to the a2 region, deletion of at least a portion of the MSD, and addition of the single-stranded RNA sequence containing the polymerase template; or d. A combination of (ii), (iii) and (iv): deleting at least a portion of the msr, deleting at least a portion of the msd, and adding the RNA motif at the 5' end or 3' end.
32. The ncRNA variant of any one of claims 1-31, wherein the one or more modifications comprise: (i) a linking of the a1 region to the a2 region, (ii) deletion of at least a portion of the msr, (iii) deletion of at least a portion of the msd, (iv) addition of the single-stranded RNA sequence comprising the polymerase template, and (v) addition of the RNA motif at the 5' or 3' end region.
33. The ncRNA variant of any one of claims 1-32, wherein the one or more modifications further comprise substitution, insertion, or deletion of one or more nucleotides, or a combination thereof.
34. The ncRNA variant of claim 33, wherein the substitution, insertion, or deletion is capable of regulating the activity of a first polymerase or a second polymerase, generating single-stranded DNA from the ncRNA variant, or the immunogenicity of the ncRNA variant.
35. The ncRNA variant of any one of claims 1-34, wherein the reference reverse transcript ncRNA comprises a sequence of a naturally occurring reverse transcript or a portion thereof.
36. The ncRNA variant of any one of claims 1-35, wherein the reference reverse transcriptase ncRNA comprises a sequence selected from: SEQ ID NO: 3980-4178, 11231-11429, 4671-4825, 11922-12075, 4980-5143, 12229-12392, 367-368, 427-441, 494-521, 526-527, 536, 626, 649, 660-668, 675, 679, 687-692, 695, 697, 703, 716, 721-722, 751-763, 767, 770-1411, 1456-1462, 7624-7625, 7684-7698, 7751-7778, 7783-7784, 7793, 7883, 7906, 7917-7925, 7932, 7936, 7944-7949, 7952, 7954, 7960, 7973, 7978-7979, 8008-8020, 8024, 8027-8667, 8712-8718, 1529-1569, 8784-8823, 6697-6701, 13943-13947, 4179-4670, 11430-11921, 4884-4909, 12134-12159, 6919-6972, 14163-14215, 2786-2866, 2887-2938, 1 0039-10119, 10140-10191, 4826-4863, 12076-12113, 4864-4875, 12114-12125, 6974-7002, 14217-14244, 2598-2600, 2759-2785, 9851-9853, 10012-10038, 2445-2582, 9699-9836, 1983-2158, 9237-9412, 1612-1982, 8866-9236, 2601-2678, 9854-9931, 2679-2758, 9932-10011, 34 42-3603, 10694-10855, 3604-3708, 10856-10959, 2939-3441, 3709-3979, 5177-5192, 10192-10693, 10960-111230, 12426-12441, 7003-7033, 14245-14275, 7054-7133, 14296-14374, 7034-7049, 14276-14291, 6835-6918, 14079-14162, 6823-6834, 14068-14078, 298-366, 369-373,442-493、522-525、528-535、537、551-554、557、560-625、672-674、680-681、684-686、696、698、702、723-742、764-766、1412-1453、1463-1466、1571-1577、7555-7623、2626-7630、7699-7750、7785-7792、7794、7808-7811、7814、7817-7882、7929-7931、7937-7938、7941-7943、7953、7955、7959、7980-7999、8021-8023、8668-8706、8708-8709、8719-8722、8825-8831、374-426、539-550、555-556、558-559、671、682-683、743、745-750、7631-7683、7796-7807、7812-7813、7815-7816、7928、7939-7940、800、8002-8007、5942-6665、13189-13911、1-297、715、1580-1603、7258-7554、7972、8834-8857、705-714、7962-7971、6681-6694、13927-13940、6788-6803、14033-14048、1469-1526、5147-5151、8725-8781、12396-12400、2159-2428、9413-9682、646-648、7903-7905、2592-2595、9846-9849、676-678、717-720、7933-7935、7974-7977、538、669、704、7795、7926、7961、8710、670、699-701、7927、7956-7958、4917-4979、12167-12228、4910-4916、12160-12166、5195-5941、12444-13188、627-645、650-659、693-694、744、768-769、1451、1455、1467-1468、1527-1528、1570、1578、1579、1604-1611、2429-2444、2583-2591、2596-2597、2867-2886、4876-4883、5144-5146、5152-5176、5193-5194、6666-6680、6695-6696, 6702-6787, 6804-6822, 6973, 7050-7053, 7134-7257, 7884-7902, 7907-7916, 7950-7951, 8001, 8025-8026, 8707, 8711, 8723-8724, 8782-8783, 8824, 8832-8833, 8858-8865, 9683-9698, 9837-9845, 9850, 10120-10139, 12126 -12133, 12393-12395, 12401-12425, 12442-12443, 13912-13926, 13941-13942, 13948-14032, 14049-14067, 14216, 14292-14295, 14375-14498, 16886-17078, 17478-17622, 17677-17756, 14831-14833, 14838, 14847, 14850-15460, 17079 -17477, 17660-17676, 19031-19080, 16414-16516, 17623-17659, 19081-19108, 16397-16413, 16195-16320, 15779-15925, 15476-15778, 16321-16366, 16367-16396, 16705-16814, 16815-16885, 16517-16704, 18949-19030, 14657-14716 14778-14824, 14834, 14835-14836, 14839, 15461-15475, 14717-14777, 14841-14846, 18413-18936, 14499-14656, 18939, 15926-16178, 14837, 17757-18412, 14825-14830, 14840, 14848-14849, 16179-16194, 18937-18938 and 18940-18948 (Table A or Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872) or Table 31A of PCT Publication WO2024044723A1 (SEQ ID NO: 19543-19733 of the PCT).
37. The ncRNA variant of any one of claims 1-36, wherein the ncRNA variant comprises Table A or Table B of U.S. Application No. 18 / 087,673, or International Application No. PCT / US2023 / 061038 or International Application No. PCT / US2023 / 072872, or Table 31A of PCT Publication WO2024044723A1 (SEQ ID NO: of the PCT). The sequence in 19543-19733) has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity.
38. The ncRNA variant according to any one of claims 1-37, wherein the ncRNA variant comprises a sequence selected from: RTX3_6342_msr_stem_var1 (SEQ ID NO: 19644), RTX3_6342_msr_stem_var20 (SEQ ID NO: 19663), RTX3_6342_msr_stem_var5 (SEQ ID NO: 19648), RTX3_6342_a1a2_var6 (SEQ ID NO: 19548), RTX3_6342_a1a2_var10 (SEQ ID NO: 19552), RTX3_6342_a1a2_var15 (SEQ ID NO: 19557), RTX3_6342_a1a2_var16 (SEQ ID NO: 19558), RTX3_6342_a1a2_var19 ... 19561), RTX3_6342_a1a2_var20 (SEQ ID NO: 19562), RTX3_6342_a1a2_var21 (SEQ ID NO: 19563), RTX3_6342_a1a2_var26 (SEQ ID NO: 19568), RTX3_6342_a1a2_var27 (SEQ ID NO: 19569) and RTX3_6342_a1a2_var32 (SEQ ID NO: 19574).
39. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises, compared to RTX3_6342_WT (SEQ ID NO: 19734), RTX3_6342_msr_stem_var1 (SEQ ID NO: 19644), RTX3_6342_msr_stem_var20 (SEQ ID NO: 19663), RTX3_6342_msr_stem_var5 (SEQ ID NO: 19648), RTX3_6342_a1a2_var6 (SEQ ID NO: 19548), RTX3_6342_a1a2_var10 (SEQ ID NO: 19552), RTX3_6342_a1a2_var15 (SEQ ID NO: 19557), and RTX3_6342_a1a2_var16 (SEQ ID NO: 19557). One or more of the following modifications: RTX3_6342_a1a2_var19 (SEQ ID NO: 19561), RTX3_6342_a1a2_var20 (SEQ ID NO: 19562), RTX3_6342_a1a2_var21 (SEQ ID NO: 19563), RTX3_6342_a1a2_var26 (SEQ ID NO: 19568), RTX3_6342_a1a2_var27 (SEQ ID NO: 19569), or RTX3_6342_a1a2_var32 (SEQ ID NO: 19574).
40. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises a sequence selected from msR spacer Del-1 (SEQ ID NO: 19939), msD spacer Del-1 (SEQ ID NO: 19940), msD spacer Del-2 (SEQ ID NO: 19941), and msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19942).
41. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises, compared to WT R6342 ncRNA (SEQ ID NO: 19734), one or more modifications of msR spacer Del-1 (SEQ ID NO: 19939), msD spacer Del-1 (SEQ ID NO: 19940), msD spacer Del-2 (SEQ ID NO: 19941), or msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19942).
42. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises a sequence selected from the following: Alt1 msR spacer Del- (SEQ ID NO: 19947)1, Alt1 msD spacer Del-1 (SEQ ID NO: 19948), Alt1 msD spacer Del-2 (SEQ ID NO: 19949), Alt1 msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19950), Alt1 msR spacer Del-2 / msD Del-3 (SEQ ID NO: 19951), Alt1 msR spacer Del-2 / msD Del-3 MS2 (SEQ ID NO: 19952), and Alt1 msR spacer Del-2 / msDDel-3 U7 (SEQ ID NO: 19953).
43. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises, compared to WTR6342 ncRNA (SEQ ID NO: 19734), one or more of the following modifications: Alt1 msR spacer Del- (SEQ ID NO: 19947)1, Alt1 msD spacer Del-1 (SEQ ID NO: 19948), Alt1 msD spacer Del-2 (SEQ ID NO: 19949), Alt1 msR spacer Del-1 / msD spacer Del-2 (SEQ ID NO: 19950), Alt1 msR spacer Del-2 / msD Del-3 (SEQ ID NO: 19951), Alt1 msR spacer Del-2 / msD Del-3 MS2 (SEQ ID NO: 19952), or Alt1 msR spacer Del-2 / msD Del-3 U7 (SEQ ID NO: 19953).
44. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises a sequence selected from: msR spacer del_30HA (SEQ ID NO: 19955), msR spacer del_45HA (SEQ ID NO: 19956), WTPAM mt, (SEQ ID NO: 19957) Alt del1+del2 PAM mt (SEQ ID NO: 19958) or Alt del1+del2+U7 PAM mt (SEQ ID NO: 19959) and Alt del1+del2 25bp ins (SEQ ID NO: 19960).
45. The ncRNA variant according to any one of claims 1-38, wherein the ncRNA variant comprises, compared to WT R6342 ncRNA (SEQ ID NO: 19734), one or more of the following modifications: msR spacer del_30HA (SEQ ID NO: 19955), msR spacer del_45HA (SEQ ID NO: 19956), WT PAM mt, (SEQ ID NO: 19957) Alt del1+del2 PAM mt (SEQ ID NO: 19958), Alt del1+del2+U7 PAM mt (SEQ ID NO: 19959), or Alt del1+del225bp ins (SEQ ID NO: 19960).
46. The ncRNA variant of any one of claims 1-45, wherein the one or more modifications further comprise the addition of a 2'O methyl or thiophosphate bond to the 5' or 3' end of the ncRNA variant.
47. The ncRNA variant of claim 46, wherein the one or more modifications further comprise the addition of the 2'O methyl or thiophosphate bond at both the 5' and 3' ends of the ncRNA variant.
48. A reverse transcriptase variant comprising an ncRNA variant according to any one of claims 1-47, and an RNA encoding a first polymerase, said RNA optionally being downstream of the msd of said ncRNA variant.
49. The reverse transcriptase variant of claim 48, wherein the first polymerase is a reverse transcriptase, optionally wherein the reverse transcriptase is derived from a naturally occurring reverse transcriptase or a reverse transcriptase-like sequence.
50. The reverse transcriptase variant of claim 48 or 49, wherein the first polymerase is a reverse transcriptase selected from the group consisting of EcoI-RT, Efe1-RT, Mva1-RT, Cex1-RT, Eco8-RT, Vap1-RT, and Vro1-RT.
51. The reverse transcriptase variant of claim 50, wherein the reverse transcriptase is derived from the same naturally occurring reverse transcriptase as the reference reverse transcriptase ncRNA.
52. A chimeric gene editing composition, said composition comprising: a. The ncRNA variant of any one of claims 1-47, or the reverse transcriptase variant of any one of claims 48-51. b. A nuclease or the first mRNA encoding said nuclease, c. The guide RNA (gRNA) associated with the nuclease, and d. Optionally, a second polymerase or a second mRNA encoding the second polymerase.
53. The chimeric gene editing composition of claim 52, wherein the composition comprises the second polymerase or the second mRNA encoding the second polymerase, wherein the second polymerase is a reverse transcriptase.
54. The chimeric gene editing composition of claim 52, wherein the composition comprises the second polymerase or the second mRNA encoding the second polymerase, wherein the second polymerase is a DNA polymerase.
55. The chimeric gene editing composition according to any one of claims 52-54, wherein the nuclease is a nicking enzyme.
56. The chimeric gene editing composition of claim 55, wherein the nicking enzyme comprises Cas9 nuclease, Cas9(D10A) nuclease, TnpB nuclease, or Cas12a nuclease.
57. The chimeric gene editing composition of any one of claims 52-56, wherein the ncRNA variant or the reverse transcriptase variant is directly or indirectly linked to the gRNA.
58. The chimeric gene editing composition of claim 57, wherein the ncRNA variant or the reverse transcriptase variant, the gRNA, and the first mRNA encoding the nuclease are directly or indirectly linked.
59. The chimeric gene editing composition according to any one of claims 52-58, wherein the composition further comprises a delivery medium.
60. The chimeric gene editing composition of claim 59, wherein the delivery medium encapsulation is selected from one or more of the following components: a. The ncRNA variant or the reverse transcriptase variant. b. The nuclease or the first mRNA encoding the nuclease. c. gRNA associated with the nuclease, and d. Optionally, the second polymerase or the second mRNA encoding the second polymerase.
61. The chimeric gene editing composition of claim 59 or 60, wherein the delivery medium is a lipid nanoparticle.
62. One or more polynucleotides, said polynucleotide comprising the coding sequence of an ncRNA variant of any one of claims 1-47 or a reverse transcriptase variant of any one of claims 48-51.
63. The one or more polynucleotides of claim 62, wherein the polynucleotide further comprises a coding sequence for a nuclease or a coding sequence for gRNA.
64. The one or more polynucleotides of claim 62, wherein the polynucleotide further comprises a coding sequence for a nuclease and a coding sequence for gRNA.
65. One or more polynucleotides according to any one of claims 62-64, wherein the polynucleotide further comprises the coding sequence of the first polymerase.
66. One or more polynucleotides according to any one of claims 62-65, wherein the polynucleotide further comprises a coding sequence for a second polymerase.
67. One or more polynucleotides according to any one of claims 62-66, the polynucleotide further comprising one or more promoters, wherein each of the one or more promoters is operatively linked to: (i) a coding sequence of the ncRNA variant or the reverse transcriptase variant, (ii) a coding sequence of a nuclease, (iii) a coding sequence of a gRNA, (iv) a coding sequence of the first polymerase, or (v) a coding sequence of the second polymerase.
68. A vector comprising one or more polynucleotides according to any one of claims 62-67.
69. The vector of claim 68, wherein the vector is a plasmid or a viral vector, optionally wherein the viral vector is an AAV or a lentiviral vector.
70. A method for editing target DNA, the method comprising interacting the target DNA with a chimeric gene editing composition of any one of claims 52-61, one or more polynucleotides of any one of claims 62-67, or a vector of any one of claims 68 or 69.
71. The method of claim 70, wherein the chimeric gene editing composition, the one or more polynucleotides, or the ncRNA variant in the vector comprises a pair of homologous arms that are specifically targeted at the target locus and specifically targeted at the target DNA.
72. The method of claim 70 or 71, wherein the interaction step is performed in vivo or in vitro.
73. A chimeric gene editing composition, said composition comprising: a. Non-coding RNA (ncRNA), which consists of single-stranded RNA containing a polymerase template. b. The nicking enzyme or the first mRNA encoding the nicking enzyme. c. The guide RNA (gRNA) associated with the nickase, and d. Optionally, a second polymerase or a second mRNA encoding the second polymerase.
74. The chimeric gene editing composition of claim 73, wherein the composition comprises the second polymerase or the second mRNA encoding the second polymerase, wherein the second polymerase is a reverse transcriptase.
75. The chimeric gene editing composition of claim 73, wherein the composition comprises the second polymerase or the second mRNA encoding the second polymerase, wherein the second polymerase is a DNA polymerase.
76. The chimeric gene editing composition according to any one of claims 73-75, wherein the nicking enzyme comprises Cas9 nuclease, Cas9(D10A) nuclease, TnpB nuclease or Cas12a nuclease.
77. The chimeric gene editing composition according to any one of claims 73-76, wherein the ncRNA is directly or indirectly linked to the gRNA.
78. The chimeric gene editing composition of claim 77, wherein the ncRNA, the gRNA, and the first mRNA encoding the nickase are directly or indirectly linked.
Citation Information
Patent Citations
Method for synthesizing stable single-stranded cdna in eucaryotes by means of a bacterial retron, products and uses therefor
CA2075515A1
Stabilized reverse transcriptase fusion proteins
US10150955B2
Non-nucleoside reverse transcriptase inhibitors
US10189831B2
Methods for generating circular DNA from circular RNA
US10683498B2
Nucleic acid vaccines
US10709779B2