ARRDC1-mediated microendoplasmic reticulum-based delivery of RNA-guided proteins
ARRDC1-mediated microvesicles (ARMMs) provide an effective delivery mechanism for RNA-guided proteins, addressing the challenge of in vivo delivery and enhancing genome editing efficacy for therapeutic applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2026-03-25
AI Technical Summary
Current RNA-guided genome editing technologies face challenges in developing effective delivery mechanisms, particularly for in vivo applications.
Utilizing ARRDC1-mediated microvesicles (ARMMs) to deliver RNA-guided proteins, such as base editors, by loading them into microvesicles for targeted cell delivery.
Enhances the efficiency and effectiveness of RNA-guided genome editing by specifically delivering RNA-guided proteins to cells, offering potential therapeutic applications for a range of diseases.
Smart Images

Figure 2026509876000063 
Figure 2026509876000064 
Figure 2026509876000065
Abstract
Description
Technical Field
[0001] Claiming Priority This application claims the benefit of U.S. Provisional Patent Application No. 63 / 451,898, filed Mar. 13, 2023. The entire content of this application is incorporated herein by reference.
[0002] The present invention provides methods, systems, and compositions for ARMM-mediated delivery of RNA-guided proteins (e.g., base editors) to cells.
Background Art
[0003] RNA-guided genome editing technologies have the potential to treat a wide range of currently incurable diseases. In general, improved delivery mechanisms, particularly for in vivo delivery, are still needed.
Summary of the Invention
[0004] The present disclosure relates to the discovery that RNA-guided gene editors, such as base editors, can be loaded into microvesicles, specifically ARRDC1-mediated microvesicles (ARMM), for delivery to cells.
[0005] Provided herein are, inter alia, arrestin domain-containing protein 1 (ARRDC1)-mediated microvesicles (ARMM) comprising (i) a lipid bilayer and an ARRDC1 protein, fragment thereof, or variant thereof, (ii) an RNA-guided protein, and (iii) a gRNA.
[0006] In some embodiments, the RNA guide protein is a Cas fusion protein. In some embodiments, the RNA guide protein is a base editor. In some embodiments, the base editor is an adenine base editor. In some embodiments, the base editor is a cytosine base editor. In some embodiments, the base editor contains a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity to any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6. In some embodiments, the base editor contains or consists of any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6. In some embodiments, the gRNA is an sgRNA. In some embodiments, the gRNA is a crRNA + tracrRNA. In some embodiments, the RNA guide protein and gRNA form a complex as a ribonucleoprotein (RNP). In some embodiments, the ARRDC1 protein, a fragment thereof, or a variant thereof is fused to an RNA-guiding protein, possibly via an intervening linker and / or cleavage domain. In some embodiments, the RNA-guiding protein is fused to a WW domain, possibly via an intervening linker and / or cleavage domain.
[0007] Furthermore, this specification provides microendoplasmic reticulum-producing cells comprising (i) the ARRDC1 protein, a fragment thereof, or a variant thereof, (ii) an RNA guide protein, and (iii) a nucleic acid construct encoding a gRNA. In some embodiments, the RNA guide protein is a Cas fusion protein. In some embodiments, the RNA guide protein is a base editor. In some embodiments, the base editor is an adenine base editor. In some embodiments, the base editor is a cytosine base editor. In some embodiments, the base editor contains a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity to any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6. In some embodiments, the base editor contains or comprises any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6. In some embodiments, the gRNA is an sgRNA. In some embodiments, the gRNA is a crRNA + tracrRNA. In some embodiments, the ARRDC1 protein, a fragment thereof, or a variant thereof is fused to an RNA-guiding protein, possibly via an intervening linker and / or cleavage domain. In some embodiments, the RNA-guiding protein is fused to a WW domain, possibly via an intervening linker and / or cleavage domain.
[0008] Furthermore, this specification provides a method for delivering a molecule to a target cell, comprising contacting the target cell with any of the microvesicles described herein.
[0009] Furthermore, this specification provides a method for treating a disorder in a patient, comprising administering to the patient one of the microvesicles described herein. The method for treating a disorder in a patient comprises administering to the patient one of the microvesicle-producing cells described herein. In some embodiments, the patient is a mammal. In some embodiments, the mammal is a primate. In some embodiments, the primate is a human.
[0010] Furthermore, this specification provides any of the ARMMs described herein for using or manufacturing a drug. This specification also provides compositions substantially as shown and described. Furthermore, this specification provides methods substantially as shown and described.
[0011] ARMMs are microvessels distinct from exosomes, which are produced by direct plasma membrane budding ("DPMB"), such as in budding viruses. DPMB proceeds through a specific interaction between TSG101 and the tetrapeptide PSAP (SEQ ID NO: 1) motif of the arrestin domain-containing protein ARRDC1 accessory protein, and the arrestin domain-containing protein ARRDC1 accessory protein is localized to the plasma membrane via its arrestin domain. ARMMs are described in detail, for example, in PCT application number PCT / US2013 / 024839, and U.S. specifications 9,737,480, 9,816,080, and 1,026,0055, with the title of the invention, “Arrdc1-Mediated Microvesicles (ARMMs) and Uses Thereof,” filed by Lu, Q. et al. on February 6, 2013 (published on August 15, 2013 as International Publication No. 2013 / 119602), and PCT publication International Publication No. 2018 / 067546, the entire contents of these documents are incorporated by reference thereto. The interaction of ARRDC1 / TSG101 results in the transfer of TSG101 from endosomes to the plasma membrane, mediating the release of microvesicles containing TSG101, ARRDC1, and other cellular components, as well as the molecule of interest.
[0012] Naturally occurring or non-natural target molecules include, but are not limited to, proteins, nucleic acids, and small molecules that can preferably bind to one or more ARMM-related proteins (e.g., ARRDC1), or more specifically, proteins, nucleic acids, and small molecules that can be modified to bind to TSG101 or ARRDC1 or specific motifs in these. These bindings facilitate the integration of molecules into the ARMM, which can then be used to deliver the desired payload (target molecule) into a target cell. For example, but are not limited to, a payload RNA can be fused to a transactivation response (TAR) element, thereby binding it to an ARRDC1 protein fused to an RNA-binding protein such as a Tat protein (e.g., bovine TAT protein). Alternatively, a payload protein can be fused to one or more WW domains that bind to the PPXY (SEQ ID NO: 2) motif of ARRDC1. Molecular binding to ARMM-related proteins (e.g., ARRDC1) facilitates the loading of molecules into the ARMM containing ARRDC1. Alternatively, the molecule can be fused to an ARMM protein (e.g., TSG101 or ARRDC1) to load the payload into the ARMM. The molecule can be fused to the ARMM protein (e.g., TSG101 or ARRDC1) via a linker, which can be cleaved upon delivery to the target cell.
[0013] In some further aspects of the present invention, a method is provided for delivering a molecule (e.g., one or more therapeutic agents) to a target cell type, tissue, or structure by contacting the target with a microendoplasmic reticulum, as described herein.
[0014] In some aspects of the present invention, a method is provided for treating a disorder in a patient by administering to the patient a microvesicle or microvesicle-producing cells as described herein. In other aspects of the present invention, the disorder is any of gain-of-function disorder, loss-of-function disorder, or repeat elongation disorder. In yet another aspect of the present invention, the disorder results from one or more substitutions (e.g., missense or nonsense), insertions, deletions, deletion-insertions, duplications, inversions, frameshifts, or repeat elongations.
[0015] Other advantages, features, and uses of the present invention will become apparent from the detailed description of certain typical non-limiting embodiments, the drawings, non-limiting, practically supported examples, and the claims. [Brief explanation of the drawing]
[0016] [Figure 1] This figure shows the base editing efficiency at two editing target sites in an in vivo study of ARMM containing ARRDC1-ABE and gRNA1. [Figure 2] This figure shows the base editing efficiency at two editing target sites in an in vivo study of ARMM containing ARRDC1-ABE and gRNA2. [Figure 3] Figure 3A shows the base editing efficiency of ARRDC1-CEB and HEK3 gRNA in human HEK293 cells. Figure 3B shows the base editing efficiency of ARRDC1-CEB and HEK2 gRNA in human HEK293 cells. [Figure 4] This figure shows the editing of CBE for GFP. [Figure 5-1] Figures 5A-5C show that base editing using CBE with GFP gRNA can introduce immature stop codons that terminate the translation of nascent GFP transcripts in human HEK293 cells that stably express GFP. Figure 5A: control, Figure 5B: CBE-GFP-gRNA1, Figure 5C: CBE-GFP-doublestop gRNA. [Figure 5-2] (As stated above.) [Figure 5-3] (As stated above.) [Modes for carrying out the invention]
[0017] definition The terms “ARRDC1-mediated microvesicle endoplasmic reticulum” or “ARMM” as used herein refer to microvesicle endoplasmic reticulum containing the ARRDC1 protein or its variants, and / or the TSG101 protein or its variants. ARMMs are described in detail, for example, in PCT application PCT / US2013 / 024839 by Lu et al., filed on 6 February 2013 (published on 15 August 2013 as International Publication No. 2013 / 119602), with the title of the invention, "Arrdc1-Mediated Microvesicles (ARMMs) and Uses Thereof," as well as in U.S. patents 9,737,480, 9,816,080, 1,026,0055, 1,094,5954, 1,100,1817, and in PCT publications International Publications 2018 / 067546, 2021 / 062196, and 2021 / 252924, the entire contents of these documents are incorporated by reference thereto. In some embodiments, the ARMM includes a drug (payload), e.g., nucleic acid, protein, or small molecule, released from a cell (e.g., a producer cell) and present in the cytoplasm or bound to the cell membrane. Typical payloads include, but are not limited to, nucleic acids, proteins, or small molecules present in the cytoplasm or bound to the cell membrane. In some embodiments, the ARMM includes a drug, e.g., nucleic acid, protein, or small molecule, released from a cell (e.g., a transgenic cell) and present in the cytoplasm or bound to the cell membrane. In some embodiments, the ARMM is released from a transgenic cell having a recombinant expression construct containing an introduced gene, and the ARMM includes the gene product encoded by the expression construct, e.g., an RNA transcript and / or protein (e.g., ARRDC1-Tat fusion protein and TAR-payload RNA). In some embodiments, the ARMM is produced synthetically, for example, by contacting a lipid bilayer with the ARRDC1 protein or a variant in a cell-free system in the presence of TSG101 or a variant thereof.In other embodiments, ARMMs are produced synthetically by contacting a lipid bilayer with a HECT domain ligase and VPS4a. In some embodiments, ARMMs lack late endosome markers. Some ARMMs provided herein either do not contain one or more exosome biomarkers or are negative for such markers. Exosome biomarkers are known to those skilled in the art and include, but are not limited to, CD63, Lamp-1, Lamp-2, CD9, HSPA8, GAPDH, CD81, SDCBP, PDCD6IP, ENO1, ANXA2, ACTB, YWHAZ, HSP90AA1, ANXA5, EEF1A1, YWHAE, PPIA, MSN, CFL1, ALDOA, PGK1, EEF2, ANXA1, PKM2, HLA-DRA, and YWHAB. Certain ARMMs provided herein may contain exosome biomarkers. Therefore, some ARMMs may be negative for one or more other exosome biomarkers but positive for one or more different exosome biomarkers. For example, such an ARMM may be negative for CD63 and Lamp-1 but contain PGK1 or GAPDH, or it may be negative for CD63, Lamp-1, CD9, and CD81 but positive for HLA-DRA. In some embodiments, the ARMM contains exosome biomarkers, but at levels lower than those found in exosomes. For example, some ARMMs contain one or more exosome biomarkers at levels less than 1%, less than 5%, less than 10%, less than 20%, less than 30%, less than 40%, or less than 50% of the levels of those biomarkers found in exosomes. As a non-limiting example, in some embodiments, the ARMM may be negative for CD63 and Lamp-1, contain CD9 at a level less than 5% of the level of CD9 typically found in exosomes, and be positive for ACTB. Exosome biomarkers other than those listed above are known in the art, and the present invention is not limited in this respect.
[0018] The term “cargo protein,” as used herein, refers to a protein that can be incorporated into an ARMM, for example, into the liquid phase of the ARMM or into the lipid bilayer of the ARMM. The term “cargo protein to be delivered” refers to any protein that can be delivered to a target, organ, tissue, or cell via its binding to or inclusion in the ARMM. In some embodiments, the cargo protein is delivered to target cells in vitro, in vivo, or ex vivo. In some embodiments, the cargo protein to be delivered is a biologically active agent, i.e., the cargo protein to be delivered has activity in cells, organs, tissues, and / or targets. For example, a protein that, when administered to a target, produces a biological effect on that target is considered biologically active. In certain embodiments, the cargo protein is a nuclease or a variant thereof (e.g., Cas9 protein or a variant thereof). In certain embodiments, the nuclease may be a Cas9 nuclease, a TALE nuclease, a zinc finger nuclease, or any variant thereof. Nucleases containing the Cas9 protein and its variants are described in further detail elsewhere in this specification. In some embodiments, the Cas9 protein or its variants are bound to nucleic acids. For example, a cargo protein may be a Cas9 protein bound to gRNA. In some embodiments, the cargo protein to be delivered is a therapeutic agent.
[0019] As used herein, the term "therapeutic agent" refers to any agent that, when administered to a subject, has an advantageous effect. In some embodiments, the therapeutic agent includes a small molecule, a protein (or peptide), one or more nucleic acids, or an agent conjugated to a small molecule. In some embodiments, the payload to be delivered is a diagnostic agent. In some embodiments, the agent to be delivered is a prophylactic agent. In some embodiments, the agent to be delivered is useful as a contrast agent. In some of these embodiments, the diagnostic agent or contrast agent is biologically active, and in other embodiments, it is not. In some embodiments, the therapeutic agent includes an agent that reduces (knocks down) the expression of one or more genes in an organism (e.g., a subject). In other embodiments, the therapeutic agent includes an agent that inactivates or removes (knocks out) one or more specific genes in an organism (e.g., a subject). In some embodiments, the therapeutic agent delivered to a cell is a transcription factor, a tumor suppressor, a developmental regulator, a growth factor, a metastasis suppressor, an apoptosis-promoting protein, a nuclease, or a recombinase.
[0020] As used herein, the term "therapeutic effect" refers to the outcome of a treatment that is judged to have a desirable and advantageous result. Therapeutic effects include, directly or indirectly, the arrest, reduction, or elimination of the symptoms of a disease. Therapeutic effects also include, directly or indirectly, the arrest, reduction, or elimination of the progression of the symptoms of a disease.
[0021] As used herein, the term “transcription factor” refers to DNA-binding proteins that regulate the transcription of DNA into RNA, for example, by activating or repressing transcription. Some transcription factors influence transcription regulation on their own, while others act in conjunction with other proteins. Some transcription factors can both activate and repress transcription under certain conditions. Typically, transcription factors bind to one or more specific target sequences that are highly similar to a particular consensus sequence within the regulatory region of a target gene. Transcription factors can regulate the transcription of a target gene either alone or in complex with other molecules. Examples of transcription factors include, but are not limited to, Sp1, NF1, CCAAT, GATA, HNF, PIT-1, MyoD, Myf5, Hox, Winged Helix, SREBP, p53, CREB, AP-1, Mef2, STAT, R-SMAD, NF-Kβ, Notch, TUBBY, and NFAT.
[0022] As used herein, the term "binding RNA" refers to ribonucleic acid (RNA) that binds to an RNA-binding protein, such as any of the RNA-binding proteins known in the art and / or described herein. In some embodiments, the binding RNA is an RNA that specifically binds to an RNA-binding protein. A binding RNA that "specifically binds" to an RNA-binding protein has a greater affinity, avidity, binds more readily, and / or for a longer period of time to the RNA-binding protein than its binding to another protein, such as a protein that does not bind RNA or a protein with weak binding to the binding RNA. In some embodiments, the binding RNA is a naturally occurring RNA or a non-natural variant thereof that binds to a specific RNA-binding protein. For example, the binding RNA can be a TAR element, a Rev response element (RRE), MS2 RNA, or any variant thereof that specifically binds an RNA-binding protein. In some embodiments, the binding RNA can be a trans-activation response element (TAR element) or a variant thereof, which is an RNA stem-loop structure found at the 5' end of nascent HIV-1 transcripts and specifically binds to the transcriptional trans-activator (Tat) protein. In some embodiments, the binding RNA is a Rev response element (RRE) or a variant thereof that specifically binds to an accessory protein Rev (e.g., Rev from HIV-1). In some embodiments, the binding RNA is MS2 RNA that specifically binds to the MS2 phage coat protein. The binding RNAs of the disclosure can be designed to specifically bind a protein (e.g., an RNA-binding protein fused to ARRDC1) to facilitate loading of the binding RNA to ARMM (e.g., a binding RNA fused to a payload RNA).
[0023] As used herein, the term "aptamer" refers to a nucleic acid (e.g., RNA, DNA) that binds to a specific target molecule, such as an RNA-binding protein. In some embodiments, nucleic acid (e.g., DNA or RNA) aptamers are engineered to bind to a variety of molecular targets, such as proteins, small molecules, macromolecules, metabolites, carbohydrates, metals, nucleic acids, cells, tissues, and organisms, either through repeated in vitro selection or, instead, via the SELEX (systematic evolution of ligands by exponential enrichment) method.Methods for manipulating aptamers to bind to various molecular targets such as proteins are known in the art, including U.S. Patent No. 637619 and U.S. Patent No. 9061043, Shui, B., et al., “RNA aptamers that functionally interact with green fluorescent protein and its derivatives,” Nucleic Acids Res., Mar; 40(5): e39 (2012), Trujillo, UH, et al., “DNA and RNA aptamers: from tools for basic research towards therapeutic applications,” Comb. Chem. High Throughput Screen, 9(8):619-32 (2006), Srisawat, C., et al., “Streptavidin aptamers: Affinity tags for the study of RNAs and ribonucleoproteins,” RNA, 7:632-641 (2001), and Turk and Gold, “Systematic evolution of ligands by exponential enrichment: RNA ligands to This includes the information found in "bacteriophage T4 DNA polymerase," Science, (1990), and the entire contents of each of these publications are thus incorporated herein by reference.
[0024] The term “RNA-binding protein,” as used herein, refers to a polypeptide molecule that binds to binding RNA, e.g., any binding RNA known in the art and / or described herein. In some embodiments, the RNA-binding protein is a protein that specifically binds to binding RNA. An RNA-binding protein that “specifically binds” to binding RNA binds to binding RNA more readily and / or for a longer period of time with greater affinity, affinity, and for a greater duration than its binding to another RNA, e.g., a control RNA (e.g., RNA with a random nucleic acid sequence) or RNA that binds weakly to the RNA-binding protein. In some embodiments, the RNA-binding protein is a naturally occurring protein or a non-natural variant thereof that binds to a specific RNA. For example, in some embodiments, the RNA-binding protein may be a transcriptional transactivator (Tat) protein that specifically binds to a transactivation response element (TAR element). In some embodiments, the Tat protein is bovine-derived. In some embodiments, the RNA-binding protein is a regulator or variant of a virion expression (Rev) protein (e.g., HIV-1-derived Rev) that specifically binds to a Rev response element (RRE). In some embodiments, the RNA-binding protein is a coat protein of the MS2 bacteriophage that specifically binds to MS2 RNA. RNA-binding proteins useful in this disclosure (e.g., binding proteins fused to ARRDC1) can be designed to specifically bind binding RNA (e.g., binding RNA fused to payload RNA) to facilitate the loading of binding RNA into ARMM.
[0025] The terms “payload,” “payload protein,” “payload nucleic acid,” “payload DNA,” “payload RNA,” or “payload small molecule,” as used herein, refer to proteins, nucleic acids including DNA or RNA, or small molecules that can be incorporated into an ARMM, for example, in the liquid phase of the ARMM or the lipid bilayer of the ARMM. The types of payload proteins, payload nucleic acids, payload DNA, payload RNA, and payload small molecules are known in the art and include those described in U.S. Patent No. 9,737,480, U.S. Patent No. 9,816,080, U.S. Patent No. 1,026,0055, and International Publication No. 2018 / 067546, published by the PCT, the entire contents of each of these documents are thus incorporated in their entirety by reference.
[0026] The payload may be delivered to a target, organ, tissue, or cell via its association with or encapsulation within the ARMM. In some embodiments, the payload is delivered to target cells in vitro, in vivo, or ex vivo. In some embodiments, the payload to be delivered is a biologically active drug, i.e., the payload to be delivered has activity in cells, organs, tissues, and / or targets. For example, a protein, nucleic acid (e.g., DNA or RNA), or small molecule that, when administered to a target, produces a biological effect on that target is biologically active. In some embodiments, the payload to be delivered is a therapeutic drug.
[0027] The term "linker," as used herein, refers to a chemical moiety that links two molecules or moieties, for example, the ARRDC1 protein and the Tat protein, the WW domain and the Tat protein, or the ARRDC1 protein and the Cas9 nuclease. Typically, the linker is located between two groups, molecules, or other moieties, or these two groups, molecules, or other moieties are adjacent to the linker, and the linker links them by covalent bonds. In some embodiments, the linker comprises one or more amino acids (e.g., a peptide or protein). In some embodiments, the linker comprises one nucleotide (e.g., DNA or RNA) or more nucleotides (e.g., a nucleic acid). In some embodiments, the linker is an organic molecule, a functional group, a polymer, or one or more other chemical moieties. In some embodiments, the linker is a cleavable linker, for example, the linker comprises a bond that can be cleaved when exposed to, for example, UV light or a hydrolytic enzyme, for example, a protease or esterase. In some embodiments, the linker is an amino acid of any length having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or more amino acids (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids). In other embodiments, the linker is a chemical bond (e.g., a covalent bond, an amide bond, a disulfide bond, an ester bond, a carbon-carbon bond, a carbon-heteroatom bond, and similar).
[0028] As used herein, the term “animal” refers to any member of the animal kingdom. In some embodiments, the term “animal” refers to a human of any sex at any stage of development. In some embodiments, the term “animal” refers to a non-human animal at any stage of development. In certain embodiments, a non-human animal is a mammal (e.g., rodents, mice, rats, rabbits, monkeys, dogs, cats, sheep, cattle, primates, or pigs). Animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, and helminths. In some embodiments, an animal is a transgenic animal, a genetically modified animal, or a clone. In some embodiments, an animal is a transgenic non-human animal, a genetically modified non-human animal, or a non-human clone.
[0029] As used herein, the terms “conjugated,” “linked,” “attached,” and similar terms mean that, when used with respect to two or more entities, for example chemical parts, molecules, and / or ARMMs, the entities are physically bound or linked to each other, either directly or through one or more additional parts acting as linkers, to form a structure, and that the structure is stable enough to maintain the physically bound state of the entities under the conditions under which the structure is used, for example physiological conditions. ARMM microvesicles are typically bound to drugs, such as nucleic acids, proteins, or small molecules, by mechanisms involving covalent bonds (e.g., via amide bonds) or non-covalent bonds (e.g., between ARRDC1 and the WW domain, or between the Tat protein and the TAR element). In certain embodiments, a drug (e.g., a therapeutic drug, payload protein, payload nucleic acid, or payload small molecule) is covalently bonded to a molecule non-covalently linked to a portion of the ARMM, such as the ARRCD1 protein, TSG101 protein or its variants, or a protein or its variant linked to the lipid bilayer by a covalent bond (e.g., an amide bond). In some embodiments, the linkage is via a linker, such as a cleavable linker. In some embodiments, an entity (e.g., a payload protein, payload nucleic acid, or payload small molecule) is linked to the ARMM by encapsulation within the ARMM, for example. For example, in some embodiments, a molecule present in the cytoplasm of an ARMM-producing cell (e.g., a therapeutic drug, payload protein, payload nucleic acid, or payload small molecule) is linked to the ARMM by encapsulating the drug-containing cytoplasm within the ARMM during ARMM budding. Similarly, membrane proteins or other molecules bound to the cell membrane of ARMM-producing cells can bind to ARMM produced by the cell by encapsulating it within the ARMM membrane during budding.
[0030] As used herein, the expression “biologically active” refers to any characteristic of a substance that is active in cells, organs, tissues, and / or subjects. For example, a substance that, when administered to an organism, produces or causes a biological effect on that organism is biologically active. As an example, a payload RNA can be considered biologically active if, when administered to a subject or cell, it increases or decreases the expression of a gene product. As another example, a nuclease payload protein can be considered biologically active if, when administered to a subject, it increases or decreases the expression of a gene product.
[0031] As used herein, the term “conserved” refers to a nucleotide or amino acid residue in each polynucleotide or amino acid sequence that exists unchanged at the same position in two or more related sequences being compared. A relatively conserved nucleotide or amino acid is a nucleotide or amino acid that is more conserved among related sequences than it is elsewhere in the sequence. In some embodiments, two or more sequences are described as “fully conserved” if they are 100% identical to each other. In some embodiments, two or more sequences are described as “highly conserved” if they are at least 70%, at least 80%, at least 90%, or at least 95% identical to each other. In some embodiments, two or more sequences are described as “highly conserved” if they are about 70%, about 80%, about 90%, about 95%, about 98%, or about 99% identical to each other. In some embodiments, two or more sequences are described as “conserved” if they are at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, or at least 95% identical to one another. In some embodiments, two or more sequences are described as “conserved” if they are approximately 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 98% identical, or at least 99% identical to one another.
[0032] The term "engineered," as used herein, refers to a protein, nucleic acid, complex, substance, or entity designed, produced, prepared, synthesized, and / or manufactured by a human. Thus, an engineered product is a product that does not occur naturally. In some embodiments, an engineered protein or nucleic acid is a protein or nucleic acid designed to meet requirements or to have desired characteristics. For example, a payload RNA can be engineered to bind to ARRDC1 by fusing one or more WW domains to a Tat protein and fusing the payload RNA to a TAR element to facilitate the loading of the payload RNA into the ARMM. In another embodiment, a payload RNA can be engineered to bind to ARRDC1 by fusing a Tat protein to ARRDC1 and fusing the payload RNA to a TAR element to facilitate the loading of the payload RNA into the ARMM. In yet another embodiment, a payload protein can be engineered to bind to ARRDC1 by fusing one or more WW domains to the payload protein to facilitate the loading of the payload protein into the ARMM.
[0033] As used herein, the term “expression” of a nucleic acid sequence means one or more of the following events: (1) production of an RNA transcript from a DNA sequence (e.g., by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, 5' cap formation, and / or 3' end processing); (3) translation of an RNA transcript into a polypeptide or protein; and (4) post-translational modification of a polypeptide or protein.
[0034] The term "operably linked," as used herein, refers to an arrangement of sequences or regions where components are configured to perform their normal or intended functions. Thus, regulatory or control sequences operably linked to a coding sequence can influence the expression of the coding sequence. Regulatory or control sequences do not need to be contiguous with the coding sequence, insofar as they function to direct proper expression or polypeptide production. For example, an intervening sequence that is transcribed but not translated may exist between a promoter sequence and a coding sequence, and the promoter sequence may still be considered operably linked to the coding sequence. Promoter sequences as described herein are DNA regulatory regions located a short distance from the 5' end of a gene that act as an RNA polymerase binding site. Promoter sequences can bind intracellular RNA polymerase and / or initiate transcription of downstream (3' direction) coding sequences. Promoter sequences can be promoters that can initiate transcription in prokaryotes or eukaryotes. Some non-exclusive examples of eukaryotic promoters include the cytomegalovirus (CMV) promoter, the chicken β-actin (CBA) promoter, and the hybrid form of the CBA promoter (CBh).
[0035] As used herein, “fusion protein” includes a second protein moiety, for example, a protein delivered to a target cell, and a first protein moiety, for example, the ARRCD1 protein or a variant thereof, or the TSG101 protein or a variant thereof, linked via a peptide bond. In certain embodiments, the fusion protein is encoded by a single fusion gene.
[0036] As used herein, the term “gene” has its meaning as understood in the art. Those skilled in the art will understand that the term “gene” may include gene regulatory sequences (e.g., promoters, enhancers, etc.) and / or intron sequences. It will also be understood that the definition of a gene includes references to nucleic acids that do not code for proteins but rather code for functional RNA molecules such as gRNA, RNAi drugs, ribozymes, and tRNA. As used in this application, it should be noted that the term “gene” generally refers to a portion of nucleic acids that code for proteins. This term may, as will be apparent to those skilled in the art from the context, include regulatory sequences. This definition is not intended to exclude the application of the term “gene” to expression units that do not code for proteins, but rather to clarify that in most cases the term as used herein refers to nucleic acids that code for proteins.
[0037] As used herein, the terms “gene product” or “expression product” generally refer to RNA transcribed from a gene (before and / or after processing), or polypeptides encoded by RNA transcribed from a gene (before and / or after modification).
[0038] As used herein, the term “green fluorescent protein” (GFP) refers to the protein first isolated from the jellyfish Aequorea victoria, which fluoresces green when exposed to blue light, or to derivatives of such proteins (e.g., high-sensitivity proteins or wavelength-shifted proteins). The amino acid sequence of wild-type GFP is as follows:
[0039] [ka]
[0040] Proteins that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% homologous to SEQ ID NO: 3 are also considered to be green fluorescent proteins.
[0041] As used herein, the term “homologousness” refers to the overall relationship between nucleic acids (e.g., DNA molecules and / or RNA molecules) or polypeptides. In some embodiments, nucleic acids or proteins are considered “homologous” to each other if their sequences are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical. In some embodiments, nucleic acids or proteins are considered “homologous” to each other if their sequences are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical. The term “homologous” necessarily refers to a comparison between at least two sequences (nucleotide sequences or amino acid sequences). According to the present invention, two nucleotide sequences are considered homologous if the polypeptides they encode are at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% identical over at least one sequence of at least about 20 amino acids. In some embodiments, homologous nucleotide sequences are characterized by their ability to encode at least 4-5 uniquely identified amino acids in a sequence. Both the identity and approximate spacing of these amino acids compared to each other must be considered in order to consider sequences homologous. For nucleotide sequences less than 60 nucleotides in length, homology is determined by the ability to encode at least 4-5 uniquely identified amino acids in a sequence. According to the present invention, two protein sequences are considered homologous if these proteins are at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% identical over at least one sequence of at least about 20 amino acids.
[0042] As used herein, the term “identity” refers to the overall relationship between nucleic acids or proteins (e.g., DNA molecules, RNA molecules, and / or polypeptides). The percentage of identity between two nucleic acid sequences can be calculated, for example, by aligning the two sequences for the purpose of optimal comparison (for example, gaps can be introduced in one or both of the first and second nucleic acid sequences for optimal alignment, and non-identical sequences can be ignored for the purpose of comparison). In certain embodiments, the length of the sequences aligned for comparison is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the length of the reference sequence. Nucleotides at corresponding nucleotide positions are thus compared. If there is a nucleotide in the first sequence that is identical to the nucleotide at the corresponding position in the second sequence, then the molecules are identical at that position. The percentage of identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps that need to be introduced for optimal alignment of the two sequences and the length of each gap. The comparison of arrays and the determination of the percentage of identity between two arrays can be performed using mathematical algorithms.For example, the percentage of identity between two nucleotide sequences can be determined using methods such as those described in Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smth, DW, ed., Academic Press, New York, 1993; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991, each of which is incorporated herein by reference. For example, the percentage of identity between two nucleotide sequences can be determined using the Meyers and Miller algorithm (CABIOS, 1989, 4:11-17), which is incorporated into the ALIGN program (version 2.0) using the PAM120 weight residue table, gap length penalty 12, and gap penalty 4. Alternatively, the percentage of identity between two nucleotide sequences can be determined using the GAP program of the GCG software package, which uses the NWSgapdna.CMP matrix. Commonly used methods for determining the percentage of identity between sequences include, but are not limited to, the method disclosed in Carillo, H., and Lipman, D., SIAM J Applied Math., 48:1073 (1988), which is incorporated herein by reference.Techniques for determining identity are systematized in publicly available computer programs. Typical computer software for determining homology between two sequences includes, but is not limited to, the GCG program package (Devereux, J., et al., Nucleic Acids Research, 12(1), 387 (1984)), BLASTP, BLASTN, and FASTA (Atschul, SF, et al., J. Molec. Biol., 215, 403 (1990)).
[0043] As used herein, the term “in vitro” refers to events occurring in an artificial environment rather than within living organisms (e.g., animals, plants, or microorganisms), such as in a test tube or reactor, in a cell culture, or in a petri dish.
[0044] As used herein, the term "in vivo" refers to an event occurring within a living organism (e.g., an animal, plant, or microorganism).
[0045] As used herein, the term “ex vivo” is understood to mean an event outside of a living organism, and therefore a medical procedure in which an organ, cell, or tissue is taken from a living organism for treatment or procedure and then returned to the same or another living organism. In certain embodiments, ex vivo treatment involves inducing one or more gene modifications in the patient’s cells outside the patient’s body to obtain a therapeutic effect therein, and then transferring (e.g., transplanting) those cells to the patient.
[0046] As used herein, the term “isolated” refers to a substance or entity that (1) is separated from at least some of the components to which it was originally bound when it was produced (naturally or in an experimental setting), and / or (2) is produced, prepared, and / or manufactured by human hands. An isolated substance and / or entity may be separated from at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or more than 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or more than 10% of the other components to which it was originally bound. In some embodiments, an isolated substance is of a purity greater than about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more than 99%. As used herein, a substance is “pure” if it substantially has no other components.
[0047] As used herein, the term “nucleic acid” in its broadest sense refers to compounds and / or substances that are or can be incorporated into an oligonucleotide chain via phosphodiester bonds. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to an oligonucleotide chain containing individual nucleotides. As used herein, the terms “oligonucleotide” and “polynucleotide” can be used interchangeably to refer to polymers of nucleotides (e.g., a sequence of at least two nucleotides). In some embodiments, “nucleic acid” encompasses RNA, as well as single-stranded and / or double-stranded DNA and / or complementary DNA (cDNA). Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, i.e., analogs having something other than a phosphodiester backbone. For example, so-called “peptide nucleic acids,” which are known in the art and have peptide bonds instead of phosphodiester bonds in their backbone, are considered to be within the scope of the present invention. The term “nucleotide sequence encoding an amino acid sequence” includes all nucleotide sequences that encode degenerate and / or identical amino acid sequences. Nucleic acid sequences encoding proteins and / or RNA may contain introns. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, or chemically synthesized. Where appropriate, for example in the case of chemically synthesized molecules, nucleic acids may contain nucleoside analogs, such as analogs having chemically modified bases or sugars, modifications to the backbone, etc. Nucleic acid sequences are represented in the 5' to 3' direction unless otherwise indicated. The term “nucleic acid segment” is used herein to refer to a nucleic acid sequence that is part of a longer nucleic acid sequence. In many embodiments, a nucleic acid segment contains at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more residues.In some embodiments, nucleic acids are natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-amino The bases are or include adenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoamidite bonds). In some embodiments, the present invention specifically addresses “unmodified nucleic acids” meaning nucleic acids (e.g., polynucleotides and residues, including nucleotides and / or nucleosides) that have not been chemically modified to facilitate or achieve delivery.
[0048] As used herein, the term “protein” means a sequence of at least two amino acids linked to one or more peptide bonds. Proteins may contain non-amino acid portions (e.g., glycoproteins) and / or may be otherwise processed or modified. Those skilled in the art will understand that a “protein” may be a complete protein chain produced by a cell (with or without a signal sequence) or a functional portion thereof. Those skilled in the art will further understand that a protein may sometimes contain two or more protein chains linked, for example, by one or more disulfide bonds or by other means. Proteins may contain L-amino acids, D-amino acids, or both, and may contain any variety of amino acid modifications or analogs known in the art. Useful modifications include, for example, the addition of carbohydrate groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, amide groups, terminal acetyl groups, linkers, etc., for chemical entities, e.g., conjugation, functionalization, or other modifications (e.g., alpha-amidation). In certain embodiments, protein modification has resulted in a more stable protein (e.g., an extended half-life in vivo). These modifications include protein cyclization and D-amino acid incorporation. None of the modifications substantially interfere with the desired biological activity of the protein. In certain embodiments, protein modification has resulted in a more biologically active protein. In some embodiments, the protein may include native amino acids, non-native amino acids, synthetic amino acids, amino acid analogs, and combinations thereof.
[0049] As used herein, the terms “subject” or “patient” refer to any organism to which a composition according to the present invention may be administered, for example, for experimental, diagnostic, preventive, and / or therapeutic purposes. Typical subjects include animals (e.g., mammals, such as mice, rats, rabbits, farm animals and agricultural animals, pets, non-human primates, and humans). In some embodiments, the subject is a patient who has or is suspected of having a disease or disorder. In other embodiments, the subject is a healthy volunteer.
[0050] As used herein, the terms “disease” and “disorder” mean any condition, pathological state, or disorder that impairs or interferes with the normal functioning of a cell, tissue, or organ.
[0051] As used herein, the terms “treating” or “treatment” mean partially or completely preventing, altering, and / or reducing the occurrence of one or more symptoms or characteristics of a particular disease or adverse condition. In one sense of the present invention, treatment can be performed to prevent or improve a pathological condition in a subject. The therapeutic effects of treatment include, but are not limited to, prevention of the onset or recurrence of a disease, reduction of symptoms, reduction of any direct or indirect pathological consequences of a disease, prevention of metastasis, slowing of the rate of disease progression, improvement or mitigation of symptoms, and remission or improvement of prognosis. Treatment may also be administered to a subject that does not show signs or symptoms of a disease, disorder, or condition, or a subject that shows only early signs or symptoms of a disease or condition, for the purpose of reducing the risk of showing or progressing to more severe effects associated with the disease, disorder, or condition. Treatment can prevent the onset of a disorder or symptoms of a disorder in a subject. Treatment can prevent the physical impairments caused by the condition (e.g., decreased vision, reduced visual acuity, low vision, blindness, reduced gait, and more generally, any treatable pathological condition) by preventing or reversing their progression.
[0052] As used herein, the term “therapeutic dose” means the amount of the drug or payload being delivered (e.g., nucleic acids, proteins, drugs, therapeutic agents, diagnostic agents, prophylactic agents, ARMMs, or ARMMs containing payload proteins or payload RNAs) that, when administered to a subject suffering from or susceptible to a disease, disorder, and / or condition, is sufficient to treat, improve, diagnose, prevent, and / or delay the onset of the disease, disorder, and / or condition. The therapeutic dose can be initially determined from preliminary in vitro studies and / or animal models. The therapeutic dose can also be determined from human data. The dose applied can be adjusted based on the relative bioavailability and potency of the compound being administered. Adjusting the dose to achieve maximum efficacy based on the methods described above and other well-known methods is within the scope of the skills of those skilled in the art. General principles for determining therapeutic efficacy can be found in Chapter 1 of Goodman and Gilman's *The Pharmacological Basis of Therapeutics*, 10th Edition, McGraw-Hill (New York) (2001), which are incorporated in their entirety herein by reference.
[0053] As used herein, the term “targeting ligand” refers to a ligand or a functional portion thereof that binds to a “target receptor” that distinguishes a target cell from other cells. Ligands can be bound by the expression or selective expression of ligand receptors accessible for ligand binding on target cells.Examples of such ligands include GE11 peptide, anti-EGFR nanobody, cRGD (cyclo(RGDfC)), KE108 peptide, octreotide, prostate-specific membrane antigen (PSMA) aptamer, TRC105, chimeric monoclonal antibodies, tumor-specific monoclonal or polyclonal antibodies (e.g., rituximab, trastuzumab, bevacizumab, alemtuzumab, panitumumab, etc., as well as their bioequivalents and parts), arginylglycylaspartate ("RGD"), DARPin, R NA aptamers, DNA aptamers, inteins, exteins, viral and nonviral intercellular fusion proteins ("fusogens") (e.g., VSV-G, syncytin-1, syncytin-2, HAP2, SNARE (e.g., VAMP1, 2, 3, 4, 7, 8)), membrane proteins, e.g., tetraspanins ("TM4SF proteins") (e.g., TSPAN1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 2) This includes peptide ligands identified from library screening (6, 27, 28, 29, 30, 31, 32, and 33, and similar), tumor-specific peptides, tumor-specific aptamers, Fab or scFv (i.e., single-stranded variable region) fragments of antibodies, e.g., Fab fragments of antibodies against EphA2 or other proteins that are specifically expressed or intrinsically accessible on metastatic cancer cells, growth factors, e.g., EGF, FGF, insulin and insulin-like growth factors, as well as homologous polypeptides, somatostatin and its analogues, transferrin, lipoprotein complexes, Arg-Gly-Asp-containing peptides, microtubule-binding sequences (MTAS), various galectins, δ-opioid receptor ligands, cholecystokinin A receptor ligands, ligands specific to angiotensin AT1 or AT2 receptors, peroxisome proliferator-activated receptor γ ligands, and other molecules that specifically bind to receptors selectively expressed on the surface of target cells or infectious organisms, and fragments of any of these molecules.
[0054] The term "target receptor," as used herein, refers to a receptor expressed by a cell capable of binding to a cell-targeting ligand. This receptor may be expressed on the cell surface. This receptor may be a transmembrane receptor. Examples of such target receptors include, but are not limited to, EGFR, α v This includes β3 integrins, somatostatin receptors, folate receptors, prostate-specific membrane antigens, CD105, mannose receptors, estrogen receptors, GM1 gangliosides, and similar substances.
[0055] In some embodiments, the cell membrane-permeable peptide may also be attached to one or more PEG-terminal groups in place of or in addition to the targeted ligand. As used herein, the terms “cell membrane-permeable peptide” (“CPP”), “protein transduction domain” (“PTD”), or “membrane translocation sequence” refer to a short-chain peptide (e.g., 4 to about 40 amino acids) that has the ability to translocate across the cell membrane to access the interior of the cell and deliver various cargo, including covalently and non-covalently conjugated proteins and oligonucleotides, into the cell. In preferred embodiments, the CPP includes an amino acid sequence comprising 1) a relatively abundant positively charged amino acid (e.g., lysine or arginine), 2) an amino acid sequence comprising an alternating pattern of polarly charged amino acids and nonpolar hydrophobic amino acids, or 3) an amino acid sequence comprising a hydrophobic peptide (e.g., a nearly nonpolar residue with a low net charge, or a hydrophobic amino acid group).(For example, U.S. Patent Application Publication No. 2022 / 0177494, Oliveira, EC, et al., “Predicting cell-penetrating peptides using machine learning algorithms and navigating in their chemical space,” Scientific Reports., 11(1):7628 (2021), Derakhshankhah, H., and Jafari, S., “Cell penetrating peptides: A concise review with emphasis on biomedical applications,” Biomedicine & Pharmacotherapy., 108:1090-1096 (2018), Milletti, F., “Cell-penetrating peptides: classes, origin, and current landscape,” Drug Discovery Today, 17(15-16):850-860 (2012), Stalmans, S., et al., “Chemical-functional diversity in cell-penetrating peptides,” PLOS ONE, See 8(8):e71752 (2013), Wagstaff, KM, and Jans, DA, “Protein transduction: cell penetrating peptides and their therapeutic applications,” Current Medicinal Chemistry, 13(12):1371-1387 (2006), the full content of each of these references is incorporated by reference thereto.Examples of CPP peptides include, but are not limited to, TAT cell-penetrating peptide, MAP, penetratin or antennapedia PTD, penetratin-Arg, antitrypsin (358-374), temporin L, maurocalcin, pVEC (cadherin-5), calcitonin, neuromedulin, penetratin, TAT-HA2 fusion peptide, TAT (47-57), SynB1, SynB3, PTD-4, PTD-5, FHV Coat-(35-49), BMV Gag-(7-25), HTLV-II Rex-(4-16), HIV-1 Tat(48-60) or D-Tat, R9-Tat, transportan, SBP or human P1, FBP, MPG(δNLS), Pep-1 or Pep-1-cysteamine, Pep-2, cyclic sequences, polyarginine (R×N(4 < N < 17) chimeras), polylysine (K×N(4 < N < 17) chimeras), (RAca)6R, (RAbu)6R, (RG)6R, (RM)6R, (RT)6R, (RS)6R, R10, (RA)6R, and R7.
[0056] As used herein, “vector” means any nucleic acid, or a particle, cell, or organism containing nucleic acid, that can be used to transfer nucleic acids into a host cell. The term “vector” includes both viral and nonviral products and means for introducing nucleic acids into cells. A “vector” can be used in vitro, ex vivo, or in vivo. A vector capable of directing the expression of an operablely linked gene is referred to herein as an “expression vector.” Nonviral vectors include, for example, plasmids, cosmids, artificial chromosomes (e.g., bacterial or yeast artificial chromosomes), liposomes, electrically charged lipids (cytofectins), DNA-protein complexes, and biopolymers. Viral vectors include, for example, retrovirus, lentivirus, adeno-associated virus, poxvirus, baculovirus, reovirus, vaccinia virus, herpes simplex virus, Epstein-Barr virus, and adenovirus vectors, but are not limited to these. A vector may also include the entire genome sequence or recombinant genome sequence of a virus. The vector may also contain a portion of the genome containing a functional sequence for generating a virus that can infect, enter, or be introduced into a cell in order to deliver nucleic acids into the cell.
[0057] As used herein, the term "WW domain" refers to a protein domain having two basic residues at its C-terminus that mediate protein-protein interactions with short-chain proline-rich or proline-containing motifs. It should be understood that the two basic residues of the WW domain (e.g., any two of H, R, and K) do not necessarily have to be at the absolute C-terminus of the WW protein domain. Rather, these two basic residues may be in the C-terminal portion of the WW protein domain (e.g., the C-terminal half of the WW protein domain). In some embodiments, the WW domain contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 tryptophan (W) residues. In some embodiments, the WW domain contains at least two W residues. In some embodiments, the at least two W residues are 15–25 amino acids apart. In some embodiments, the at least two W residues are 19–23 amino acids apart. In some embodiments, the at least two W residues are 20–22 amino acids apart. WW domains, possessing two basic C-terminal amino acid residues, may have the ability to bind to short-chain proline-rich or proline-containing motifs (e.g., the PPXY(SEQ ID NO: 2) motif). WW domains bind to a variety of different peptide ligands, including those containing core proline-rich sequences, such as the PPXY(SEQ ID NO: 2) motif found in ARRDC1. WW domains can be 30-40 amino acid protein interaction domains with two signature tryptophan residues separated by 20-22 amino acids. The three-dimensional structure of WW domains indicates that they generally fold into triple-stranded antiparallel β-sheets with two ligand-binding grooves.
[0058] The WW domain is found in many eukaryotes and is present in approximately 50 human proteins (Bork, P. & Sudol, M., “The WW domain: a signaling site in dystrophin?” Trends Biochem Sci., 19, 531-533 (1994)). The WW domain can coexist with several other interaction domains, including membrane targeting domains, such as C2 in NEDD4 family proteins, the phosphotyrosine-binding (PTB) domain in FE65 protein, the FF domains in CA150 and FBPIl, and the plextrin homology (PH) domain in PLEKHA5. The WW domain is also linked to various catalytic domains, including the HECT E3 protein-ubiquitin ligase domain in NEDD4 family proteins, the rotomerase domain or peptidyl prolyl isomerase domain in Pinl, and the Rho GAP domains in ArhGAP9 and ArhGAP12.
[0059] The WW domain may be a WW domain that naturally has two basic amino acids at its C-terminus. In some embodiments, the WW domain or WW domain variant may be derived from the human ubiquitin ligases WWP1, WWP2, Nedd4-1, Nedd4-2, Smurf1, Smurf2, ITCH, NEDL1, or NEDL2. Typical amino acid sequences of protein-containing WW domains (underlined WW domains) are listed below. It should be understood that not all representative protein WW domains or WW domain variants may be used in the present invention, described herein, or are limited to them.
[0060] The amino acid sequence of human WWP1 (uniprot.org / uniprot / Q9H0M0). The four underlined WW domains correspond to amino acids 349-382 (WW1), 381-414 (WW2), 456-489 (WW3), and 496-529 (WW4).
[0061] [ka]
[0062] WWI (349-382): ETLPSGWEQRKDPHGRTYYVDHNTRTTTWERPQP (Sequence ID 5).
[0063] WW2 (381-414): QPLPPGWERRVDDRRRVYYVDHNTRTTTWQRPTM (Sequence ID 6).
[0064] WW3 (456-489): ENDPYGPLPPGWEKRVDSTDRVYFVNHNTKTTQWEDPRT(Sequence ID 7).
[0065] WW4 (496-529): EPLPEGWEIRYTREGVRYFVDHNTRTTTFKDPRN (Sequence ID 8).
[0066] Human WWP2 amino acid sequence (uniprot.org / uniprot / O00308). The four underlined WW domains correspond to amino acids 300-333 (WW1), 330-363 (WW2), 405-437 (WW3), and 444-547 (WW4).
[0067] [ka]
[0068] WWI (300-333): DALPAGWEQRELPNGRVYYVDHNTKTTTWERPLP(Sequence ID 10).
[0069] WW2 (330-363): PLPPGWEKRTDPRGRFYYVDHNTRTTTWQRPTA (Sequence ID 11).
[0070] WW3 (405-437): HDPLGPLPPGWEKRQDNGRVYYVNHNTRTTQWEDPRT(Sequence ID 12).
[0071] WW4 (444-477): PALPPGWEMKYTSEGVRYFVDHNTRTTTFKDPRP (Sequence ID 13).
[0072] Human Nedd4-1 amino acid sequence (uniprot.org / uniprot / P46934). The four underlined WW domains correspond to amino acids 610-643 (WW1), 767-800 (WW2), 840-873 (WW3), and 892-925 (WW4).
[0073] [ka]
[0074] WWI (610-643): SPLPPGWEERQDILGRTYYVNHESRRTQWKRPTP (Sequence ID 15).
[0075] WW2 (767-800): SGLPPGWEEKQDERGRSYYVDHNSRTTTWTKPTV (Sequence ID 16).
[0076] WW3 (840-873): GFLPKGWEVRHAPNGRPFFIDHNTKTTTWEDPRL (Sequence ID 17).
[0077] WW4 (892-925): GPLPPGWEERTHTDGRIFYINHNIKRTQWEDPRL (Sequence ID 18).
[0078] Human Nedd4-2 amino acid sequence (>gi|21361472|ref|NP_056092.2|E3 ubiquitin protein ligase NEDD4-like isoform 3 [Homo sapiens]). The four underlined WW domains correspond to amino acids 198-224 (WW1), 368-396 (WW2), 480-510 (WW3), and 531-561 (WW4).
[0079] [ka]
[0080] WWI (198-224): GWEEKVDNLGRTYYVNHNNRTTQWHRP (Sequence ID 20).
[0081] WW2 (368-396): PSGWEERKDAKGRTYYVNHNNRTTTWTRP (Sequence ID 21).
[0082] WW3 (480-510): PPGWEMRIAPNGRPFFIDHNTKTTTWEDPRL(Sequence ID 22).
[0083] WW4 (531-561): PPGWEERIHLDGRTFYIDHNSKITQWEDPRL (Sequence ID 23).
[0084] Human Smurf1 amino acid sequence (uniprot.org / uniprot / Q9HCE7). The two underlined WW domains correspond to amino acids 234-267 (WW1) and 306-339 (WW2).
[0085] [ka]
[0086] WWI (234-267): PELPEGYEQRTTVQGQVYFLHTQTGVSTWHDPRI (Sequence ID 25).
[0087] WW2 (306-339): GPLPPGWEVRSTVSGRIYFVDHNNRTTQFTDPRL (Sequence ID 26).
[0088] Human Smurf2 amino acid sequence (uniprot.org / uniprot / Q9HAU4). The three underlined WW domains correspond to amino acids 157-190 (WW1), 251-284 (WW2), and 297-330 (WW3).
[0089] [ka]
[0090] WWI (157-190): NDLPDGWEERRTASGRIQYLNHITRTTQWERPTR (Sequence ID 28).
[0091] WW2 (251-284): PDLPEGYEQRTTQQGQVYFLHTQTGVSTWHDPRV (Sequence ID 29).
[0092] WW3 (297-330): GPLPPGWEIRNTATGRVYFVDHNNRTTQFTDPRL (Sequence ID 30).
[0093] Human ITCH amino acid sequence (uniprot.org / uniprot / Q96J02). The four underlined WW domains correspond to amino acids 326-359 (WW1), 358-391 (WW2), 438-471 (WW3), and 478-511 (WW4).
[0094] [ka]
[0095] ITCH WW1 (326-359): APLPPGWEQRVDQHGRVYYVDHVEKRTTWDRPEP (Sequence ID 32).
[0096] ITCH WW2 (358-391): EPLPPGWERRVDNMGRIYYVDHFTRTTTWQRPTL (Sequence ID 33).
[0097] ITCH WW3 (438~471): GPLPPGWEKRTDSNGRVYFVNHNTRITQWEDPRS (Sequence ID 34).
[0098] ITCH WW4 (478~511): KPLPEGWEMRFTVDGIPYFVDHNRRTTTYIDPRT (Sequence ID 35).
[0099] Human NEDL1 amino acid sequence (uniprot.org / uniprot / Q76N89). The two underlined WW domains correspond to amino acids 829-862 (WW1) and 1018-1051 (WW2).
[0100] [ka]
[0101] WWI (829-862): PLPPNWEARIDSHGRVFYVDHVNRTTTWQRPTA (Sequence ID 37).
[0102] WW2 (1018-1051): LELPRGWEIKTDQQGKSFFVDHNSRATTFIDPRI (Sequence ID 38).
[0103] Human NEDL2 amino acid sequence (uniprot.org / uniprot / Q9P2P5). The two underlined WW domains correspond to amino acids 807-840 (WW1) and 985-1018 (WW2).
[0104] [ka]
[0105] WWI (807-840): EALPPNWEARIDSHGRIFYVDHVNRTTTWQRPTA (Sequence ID 40).
[0106] WW2 (985-1018): LELPRGWEMKHDHQGKAFFVDHNSRTTTFIDPRL (Sequence ID 41).
[0107] In some embodiments, the WW domain is essentially derived from a WW domain or a WW domain variant. "Essentially derived" means that the domain, peptide, or polypeptide is essentially derived from an amino acid sequence, in which case such an amino acid sequence is present within the domain, peptide, or polypeptide with a small number of additional amino acid residues, e.g., about 1 to about 10 or so additional residues, typically 1 to about 5 additional residues.
[0108] Alternatively, the WW domain may be a modified WW domain containing two basic amino acids at its C-terminus. This technique is publicly known in the art and is described, for example, in Sambrook et al., Molecular Cloning: a Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press (2001). Therefore, those skilled in the art can easily modify an existing WW domain that does not normally have two C-terminal basic residues to contain two basic residues at its C-terminus.
[0109] Basic amino acids are amino acids having a side-chain functional group with a pKa greater than 7, and include lysine, arginine, and histidine, as well as basic amino acids not included in the 20 α-amino acids commonly found in proteins. The two basic amino acids at the C-terminus of the WW domain may be the same basic amino acid or different basic amino acids. In one embodiment, the two basic amino acids are two arginines.
[0110] The term WW domain also includes any variant of a WW domain, provided that it has two basic amino acids at its C-terminus and retains the ability of the WW domain to bind to the PPXY(SEQ ID NO: 2) motif. Such a variant of a WW domain refers to a WW domain that retains the variant's ability to bind to the PPXY(SEQ ID NO: 2) motif (i.e., the PPXY(SEQ ID NO: 2) motif of ARRDC1) and has mutations including point mutations, insertion mutations, and / or deletion mutations in one or more amino acids, but still retains the ability to bind to the PPXY(SEQ ID NO: 2) motif. Variants or derivatives therefore include deletions including truncations and fragments; insertions and additions, e.g., conservative substitutions, site-directed mutations, and allele variants; as well as modifications including one or more non-aminoacyl groups (e.g., sugars, lipids, etc.) covalently bonded to the peptide and post-translational modifications. In causing such changes, substitutions of similar amino acid residues can be made based on the relative similarity of the side chain substituents, such as their size, charge, hydrophobicity, hydrophilicity, and similar characteristics. Such substitutions can be assayed for their effects on peptide function by standard tests.
[0111] The WW domain can be part of a longer protein. Therefore, in various different embodiments, a protein may contain, consist of, or be essentially composed of a WW domain, as defined herein. A polypeptide may be a protein that contains a WW domain as a functional domain within its protein sequence.
[0112] The terms “Cas9” or “Cas9 protein” or “Cas9 polypeptide” refer to a Cas9 protein, as well as an RNA-guided nuclease comprising a fusion protein containing such a Cas9 protein and its variants (e.g., a protein comprising an active, inactive, or modified DNA cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). In some embodiments, the fusion protein includes a fusion protein that modifies the epigenome or regulates transcriptional activity. Variants include deletions or additions, e.g., the addition of one, two, or more nuclear localization sequences (e.g., those derived from SV40 and others known in the art), e.g., the addition of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 such sequences, or a range of such sequences including any two of the aforementioned values.
[0113] In some embodiments, the Cas9 polypeptide is the Cas9 protein found in type II CRISPR-associated systems. Suitable Cas9 polypeptides that may be used in certain embodiments of the present invention include, but are not limited to, Cas9 proteins derived from Streptococcus pyogenes (Sp. Cas9), Francisella novicida, Staphylococcus aureus, Streptococcus thermophilus, Neisseria meningitidis, and variants thereof.
[0114] The Cas9 nuclease is also sometimes described as the cason1 nuclease or CRISPR (clustered, regularly spaced, short palindromic repeats) related nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (e.g., viruses, translocation elements, and conjugative plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the preceding mobile element, and a target entry nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, precise processing of precrRNA requires a transcoded small RNA (tracrRNA), endogenous ribonuclease 3 (mc), and the Cas9 protein. The tracrRNA acts as a guide for the processing of precrRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA cleaves a linear or circular dsDNA target complementary to the spacer by nucleolysis from the inside. The target strand, which is not complementary to the crRNA, is first cleaved by nucleolysis from the inside and then trimmed at the 3'-5' end by nucleolysis from the outside. In nature, DNA binding and cleavage typically require a protein and both RNAs. However, by manipulating a single guide RNA ("sgRNA" or simply "gRNA"), the features of both crRNA and tracrRNA can be incorporated into a single RNA species. (See, for example, M., et al., Science, 337:816-821 (2012), whose entire content is incorporated by reference in this way). Cas9 recognizes short motifs (PAMs or protospacer-adjacent motifs) within CRISPR repeat sequences to help distinguish between self and non-self. The sequence and structure of the Cas9 nuclease are well known to those skilled in the art.(For example, see Ferretti et al., “Complete genome sequence of an M1 strain of Streptococcus pyogenes,” Proc. Natl. Acad. Sci. USA, 98:4658-4663 (2001), “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Deltcheva E., et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature, 471:602-607 (2011), and Jinek, M., et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science, 337:816-821 (2012), the entire contents of each of these publications are incorporated herein by reference.) Additional suitable Cas9 nucleases and sequences will become apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences derived from organisms and loci disclosed in the art.(e.g., Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems,” RNA Biology, 10:5, 726-737 (2013), Karvelis, G., et al., “Harnessing the natural diversity and in vitro evolution of Cas9 to expand the genome editing toolbox,” Current Opinion in Microbiology, 37:88-94 (2017), Komor, AC, et al., “CRISPR-Based Technologies for the Manipulation of Eukaryotic Genomes,” Cell, 168:20-36 (2017), and Murovec, J., et al., “New variants of CRISPR RNA-guided genome editing enzymes,” Plant Biotechnol. J., 15:917-26 See (2017), the entire contents of each of these documents are incorporated herein by reference.
[0115] In some embodiments, the Cas9 polypeptide includes wild-type Cas9, nickasase, or nuclease-inactivated (dCas9, short for nuclease-"dead"Cas9) protein.
[0116] Methods for constructing Cas9 proteins (or variants thereof) with inactive DNA cleavage domains are known (see, for example, Jinek et al., Science. 337:816-821 (2012), Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression,” Cell, 28;152(5):1173-83 (2013), the entire contents of each of these publications are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H841A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science, 337:816-821 (2012); Qi et al., Cell, 28;152(5):1173-83 (2013)). In some embodiments, proteins containing variants of Cas9 are provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9, or (2) the DNA-cleaving domain of Cas9. In some embodiments, proteins containing Cas9 or a variant of it are referred to as "Cas9 variants". Cas9 variants have homology to Cas9 or a variant of it. For example, a Cas9 variant is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or at least approximately 99.9% identical to wild-type Cas9.In some embodiments, the Cas9 variant includes a Cas9 variant (e.g., a gRNA-binding domain or a DNA-cleaving domain) such that the variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding variant of wild-type Cas9. In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, SEQ ID NO: 61 (nucleotide), SEQ ID NO: 62 (amino acid)).
[0117] [ka]
[0118] [ka]
[0119] [ka]
[0120] [ka] (Single underline: HNH domain, Double underline: RuvC domain)
[0121] In some embodiments, wild-type Cas9 corresponds to or includes SEQ ID NO: 63 (nucleotide) and / or SEQ ID NO: 64 (amino acid).
[0122] [ka]
[0123] [ka]
[0124] [ka]
[0125] [ka] (Single underline: HNH domain, Double underline: RuvC domain)
[0126] In some embodiments, dCas9 corresponds to, or partially or completely comprises, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and / or H820A mutations. dCas9 (D10A and H840A) are as follows:
[0127] [ka] (Single underline: HNH domain, Double underline: RuvC domain)
[0128] In other embodiments, dCas9 variants having mutations other than D10A and H820A are provided, which result in, for example, nuclease-inactivating Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H820, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 (e.g., variants of SEQ ID NO: 65) are provided, which are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to SEQ ID NO: 65. In some embodiments, variants of Cas9 (e.g., variants of SEQ ID NO: 65) are provided that have amino acid sequences that are about 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, or more amino acids shorter or longer than SEQ ID NO: 65.
[0129] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acids of the Cas9 protein, for example, one of the sequences provided above. In other embodiments, however, the fusion proteins provided herein do not include the full-length Cas9 sequence, but only fragments thereof. For example, in some embodiments, the Cas9 fusion proteins provided herein include a Cas9 fragment, where the fragment binds crRNA and tracrRNA or sgRNA, but does not include a functional nuclease domain, for example, by including only a truncated version of the nuclease domain or not including a nuclease domain at all. Typical amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and further suitable sequences of Cas9 domains and fragments will be obvious to those skilled in the art. In some of these embodiments, the fusion protein includes a transcriptional activator (e.g., VP64), a transcriptional repressor (e.g., KRAB, SID), a nuclease domain (e.g., FokI), a base editor, a prime editor, a recombinase domain (e.g., Hin, Gin, or Tn3), a deaminase (e.g., cytidine deaminase or adenosine deaminase), or an epigenetic modifier domain (e.g., TET1, p300).
[0130] In some embodiments, Cas9 is Corynebacterium ulcerans (NCBI reference numbers: NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI reference numbers: NC_016782.1, NC_016786.1), Spiroplasma syrphidicola (NCBI reference number: NC_021284.1), Prevotella intermedia (NCBI reference number: NC_017861.1), Spiroplasma taiwanense (NCBI reference number: NC_021846.1), and Streptococcus inie. Neisseria meningitidis (NCBI reference number: NC_021314.1), Bellilla baltica (NCBI reference number: NC_018010.1), Psychroflexus torquisl (NCBI reference number: NC_018721.1), Streptococcus thermophilus (NCBI reference number: YP_820832.1), Listeria innocua (NCBI reference number: NP472073.1), Campylobacter jejuni (NCBI reference number: YP_02344900.1), or Neisseria meningitidis. This refers to Cas9 derived from meningitidis (NCBI reference number: YP_02342100.1).
[0131] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively.
[0132] The terms “RNA programmable nuclease” and “RNA guide nuclease” are used herein without distinction and refer to nucleases that form a complex with (e.g., are bound to or attached to) one or more RNA molecules that are not the target of cleavage. In some embodiments, an RNA programmable nuclease may be described as nuclease:RNA complex if it is a complex with RNA. Examples of RNA programmable nucleases include Cas9 nucleases. Typically, the bound RNA is described as guide RNA (gRNA). A gRNA may exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA existing as a single RNA molecule may be described as a single guide RNA (sgRNA), but “gRNA” is used without distinction to refer to guide RNA existing as a single molecule or as two or more molecules. Typically, a gRNA existing as a single RNA species includes two domains: (1) a domain having homology to the target nucleic acid (e.g., directing the binding of the Cas9 complex to the target), and (2) a domain for binding the Cas9 protein. The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site, resulting in sequence specificity of the nuclease:RNA complex.
[0133] As used herein, the term “recombinase” refers to a site-specific enzyme that mediates DNA recombination between recombinase-recognized sequences, resulting in the excision, integration, inversion, or exchange (e.g., transposition) of DNA fragments between recombinase-recognized sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvers and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, but are not limited to, Hin, Gin, Tn3, β-six, CinH, ParA, γδ, Bxb1, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, but are not limited to, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. The names serine recombinases and tyrosine recombinases derive from the conserved nucleophilic amino acid residues that recombinases use to attack DNA and covalently bind to it during strand exchange. Recombinases have many applications, including gene knockout / knock-in generation and gene therapy. (e.g., Brown et al., “Serine recombinases as tools for genome engineering,” Methods, 53(4):372-379 (2011), Hirano et al., “Site-specific recombinases as tools for heterologous gene integration,” Appl. Microbiol. Biotechnol., 92(2):227-239 (2011), Chavez and Calos, “Therapeutic applications of the ΦC31 integrase system,” Curr. Gene Ther.), 11(5):375-381 (2011), Turan and Bode, “Site-specific recombinases: from tag-and-target- to tag-and-exchange-based genomic modifications,” FASEB J., 25(12):4088-4107 (2011), Venken and Bellen, “Genome-wide manipulations of Drosophila melanogaster with transposons, Flp recombinase, and ΦC31 integrase,” Methods Mol. Biol., 859:203-228 (2012), Murphy, “Phage recombinases and their applications,” Adv. Virus Res., 83:367-414 (2012), Zhang et al., “Conditional gene manipulation: Cre-ating a new biological era,” J. Zhejiang Univ. Sci. B., See 13(7):511-524 (2012), Karpenshif and Bernstein, “From yeast to mammals: recent advances in genetic control of homologous recombination,” DNA Repair (Amst), 1;11(10):781-788 (2012), the full contents of each are incorporated by reference thereto. The recombinases provided herein are not exclusive examples of recombinases that may be used in embodiments of the present invention. The methods and compositions of the present invention can be extended by searching databases for novel orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity. (e.g., Groth et al., “Phage integrases: biology and applications,” J. Mol. Biol.)See, 335, 667-678 (2004), Gordley et al., “Synthesis of programmable integrases,” Proc. Natl. Acad. Sci. USA., 106, 5053-5058 (2009), the full contents of each, (the whole is incorporated by reference thereto). Other examples of recombinases useful in the methods and compositions described herein are known to those skilled in the art, and any new recombinases discovered or created are expected to be usable in different embodiments of the present invention. In some embodiments, the recombinase (or its catalytic domain) is fused to a Cas9 protein (e.g., dCas9).
[0134] In the context of nucleic acid modification (e.g., genome modification), the terms “recombination” and “recombination” are used to refer to the process by which two or more nucleic acid molecules, or two or more regions of one nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can, among other things, result in insertions, inversions, excisions, or transpositions of nucleic acid sequences, either within or between one or more nucleic acid molecules.
[0135] As used herein, the terms “approximately” or “about” refer to a value similar to the reference value mentioned when applied to one or more values of interest. In certain embodiments, unless otherwise stated or the context makes it clear (for example, that such a number is greater than 100% of the possible value), the terms “approximately” or “about” refer to a range of values that fall within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less than 10% of the reference value mentioned.
[0136] General description of the present invention Compositions and methods related to ARMM for producing and using the aforementioned particles are known in the Art, but are not limited to, U.S. Patent No. 9,737,480, U.S. Patent No. 9,816,080, U.S. Patent No. 1,026,0055, International Publication No. 2018 / 067546 of the PCT Application Publication, and U.S. Patent Application Publication No. 2022 / 0119785, U.S. Patent Application Publication No. 2022 / 0170013, U.S. Patent Application Publication No. 2022 / 0220462, and U.S. Patent Application Publication No. 2022 / 0282275, the entire contents of each of these documents are thus incorporated in their entirety by reference.
[0137] General pharmaceutical preparations and compositions Further aspects of this disclosure relate to pharmaceutical compositions comprising either ARMM or microendoplasmic reticulum (e.g., ARMM)-producing cells provided herein. The term “pharmaceutical composition” as used herein means a composition formulated for pharmaceutically acceptable use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises further agents (e.g., for increasing specific delivery, half-life, potency, efficacy, biological effects, and the like, as well as for other therapeutic agents and compounds).
[0138] As used herein, the term “pharmaceutically acceptable carrier” means a pharmaceutically acceptable material, composition, or medium, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, magnesium talc, calcium stearate or zinc stearate, or stearic acid), or solvent encapsulation material, that is involved in the transport or delivery of a compound from one site of the body (e.g., a site of delivery) to another site (e.g., an organ, tissue, system, or part of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense that it is compatible with the other components of the formulation and is not harmful to the cells and tissues of interest (e.g., physiologically compatible, sterile, physiologically pH, etc.). Some examples of materials that can function as pharmaceutically acceptable carriers include, but are not limited to, (1) sugars, e.g., lactose, glucose, and sucrose; (2) starches, e.g., corn starch and potato starch; (3) cellulose and its derivatives, e.g., sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants, e.g., magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients, e.g., cocoa butter and suppository waxes; (9) oils, e.g., peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and malt. (10) Soybean oil, (11) Glycols, e.g., propylene glycol, (12) Polyols, e.g., glycerin, sorbitol, mannitol, and polyethylene glycol (PEG), (13) Esters, e.g., ethyl oleate and ethyl laurate, (14) Agar, (15) Buffers, e.g., magnesium hydroxide and aluminum hydroxide, (16) Alginic acid, (17) Pyrogen-free water, (18) Isotonic saline, (19) Ringer's solution, (20) Ethyl alcohol, (21) pH buffer solution, (22) Polyesters, polycarbonates, and / or polyanhydrides, (23) Fillers, e.g., polypeptides and amino acids, (24) Serum components, e.g., serum albumin, HDL, and LDL, (25) C2-C 12This includes alcohols, such as ethanol, and (24) other non-toxic, suitable substances used in pharmaceutical preparations. Wetting agents, colorants, release agents, coatings, sweeteners, flavoring agents, fragrances, preservatives, antioxidants, and similar substances may also be present in the preparations as appropriate. The terms “excipient,” “carrier,” “pharmaceutically acceptable carrier,” and similar terms are used herein without distinction.
[0139] In some embodiments, the pharmaceutical composition is formulated for delivery to a target, for example, for delivery of a therapeutic agent, payload protein, or payload nucleic acid to cells. Appropriate routes of administration for the pharmaceutical compositions described herein include, but are not limited to, subretinal, choroidal, intravitreal, topical, subcutaneous, transdermal, intradermal, intralesional, intra-articular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, transtympanic, intra-organ, epidural, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intrathecal, intracerebral, intracisional, and intraventricular (see, for example, Hartman, RR, and Kompella, UB, “Intravitreal, Subretinal, and Suprachoroidal Injections: Evolution of Microneedles for Drug Delivery,” J. Ocul. Pharmacol. Ther., 34(1-2):141-153 (2018)).
[0140] In some embodiments, the pharmaceutical compositions described herein are administered topically to the site of the disease. In some embodiments, the pharmaceutical compositions described herein are administered to the subject by injection, by catheter, by suppository, or by implant, the implant being a porous, non-porous, or gelatinous material, including a membrane such as a silastic membrane, or fibers.
[0141] In some embodiments, the composition is formulated according to standard procedures and is suitable for injection, intravenous administration, or subcutaneous administration to a subject, such as a human. In some embodiments, the composition for administration by injection comprises a solution in a sterile isotonic aqueous buffer. If necessary, the formulation may also include a solubilizer and a local anesthetic such as lidocaine to relieve pain at the injection site. When the formulation is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the components can be mixed before administration. When the formulation is administered by infusion, the formulation can be dispensed using an infusion bottle containing sterile pharmaceutical-grade water or saline.
[0142] Preparations for systemic administration (pharmaceutical compositions) may be liquids, such as sterile saline, Ringer's lactate solution, or Hanks' solution. Furthermore, pharmaceutical compositions may be in solid form and can be redissolved or suspended immediately before use. Certain formulations in lyophilized form are also being considered.
[0143] The compositions described herein may be administered or packaged as unit doses. When used in reference to the pharmaceutical compositions of this disclosure, the term “unit dose” refers to a physically distinct unit appropriate as a unit dosage for a subject, each unit containing a predetermined amount of active material calculated to produce a desired therapeutic effect, together with the required diluent, i.e., carrier or medium.
[0144] Furthermore, the composition may be provided as a kit comprising 1) a container containing the ARMM or microendoplasmic reticulum-producing cells of the present invention, and 2) a second container containing a pharmaceutically acceptable diluent for injection (e.g., sterile water). The pharmaceutically acceptable diluent may be used, for example, to restore or dilute the ARMM or microendoplasmic reticulum-producing cells of the present invention. In some cases, such containers may be accompanied by a warning in the form prescribed by the government agency that regulates the manufacture, use, or sale of pharmaceuticals or biological products, which reflects the government agency's approval for manufacture, use, or sale for human administration. In this regard, national and regional regulatory agencies are understood to include, but are not limited to, the U.S. Food and Drug Administration, the U.S. Department of Agriculture, the European Medicines Agency, the UK Medicines and Healthcare Products Regulatory Agency, the National Medicines Administration, and similar agencies.
[0145] In another embodiment, the invention includes a product containing materials useful for the treatment of the diseases described herein. In some embodiments, the product includes a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. Containers may be formed from a variety of materials, such as glass or plastic. It is further understood that a suitable container includes a material that is sufficiently non-reactive and protects the contents within the container. In some embodiments, the container holds a composition effective for the treatment of the diseases described herein and may have a sterile access port. For example, the container may be a bag or vial for intravenous solution having a stopper through which a needle for subdermal injection can pass. The active agent in the composition is a compound of the invention. In some embodiments, a label on or accompanying the container indicates that the composition is used for the treatment of one or more selected diseases. The invention may further include a second container containing a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer's solution, or dextrose solution. The manufactured product may further include other buffers, diluents, filters, needles, syringes, and accompanying documentation with instructions and labeling for acceptable, approved, or permitted use, as well as other materials desirable from a commercial and user perspective.
[0146] In some embodiments, a procedure is considered in which a single therapeutic agent is administered using ARMM-mediated delivery of the agent to target cells, tissues, systems, or mammalian subjects.
[0147] In some other embodiments, a procedure is considered in which two therapeutic agents are administered using ARMM-mediated delivery of each agent to a target cell, tissue, system, or mammalian subject.
[0148] Furthermore, in some other embodiments, procedures are considered in which three or more therapeutic agents are administered using ARMM-mediated delivery of each agent to target cells, tissues, systems, or mammalian subjects.
[0149] In compositions and methods described herein, when two or more therapeutic agents are administered (e.g., co-administered) to a target, cell, tissue, system, or subject, it is considered that each (i.e., two or more) agent is provided for administration within a single ARRDC1-mediated microvesicle endoplasmic reticulum (ARMM) particle.
[0150] Similarly, in compositions and methods described herein, if two or more therapeutic agents are administered (e.g., concurrently) to target cells, tissues, systems, or subjects, it is also to be further considered that each (i.e., two or more) agent is provided for administration within different ARRDC1-mediated microvesicle endoplasmic reticulum (ARMM) particles.
[0151] Kits, vectors, and cells Some aspects of this disclosure provide kits comprising nucleic acid constructs, each comprising a nucleotide sequence encoding one or more of the proteins (e.g., ARRDC1 and TSG101), fusion proteins, and / or nucleic acids provided herein. In some embodiments, the nucleotide sequence encodes one of the proteins, fusion proteins, and / or RNAs provided herein. In some embodiments, the nucleotide sequence comprises a heterologous promoter that drives the expression of one of the proteins, fusion proteins, and / or RNAs provided herein.
[0152] Some aspects of this disclosure provide microendoplasmic reticulum (e.g., ARMM) producing cells comprising any of the proteins, fusion proteins, and nucleic acids (e.g., RNA) provided herein. In some embodiments, the cells specifically comprise nucleotides encoding any of the proteins, fusion proteins, and / or RNA provided herein. In some embodiments, the cells comprise any of the nucleotides or vectors provided herein. In some embodiments, the vector comprises one or more cell-targeting proteins or cell-entering proteins (e.g., viruses and / or human fusogens).
[0153] However, those skilled in the art will understand that, based on this disclosure and knowledge in the art, further proteins, fusion proteins, and RNAs will be revealed.
[0154] The functions and advantages of these and other embodiments of the present invention will be understood more fully from the following examples. The following examples are intended to illustrate the benefits of the present invention and to describe specific embodiments, but are not intended to illustrate the entire scope of the invention.
[0155] Therefore, it will be understood that the examples do not limit the scope of the present invention.
[0156] Detailed description of a specific embodiment of the present invention The present invention provides methods, systems, and compositions for the delivery of target molecules (e.g., therapeutic agents) to cells and tissues via ARRDC1-mediated microvesicles ("ARMMs"). The present invention further relates to compositions and methods for producing, testing, and administering ARMMs. More particularly, the present invention provides compositions and methods for producing, testing, and administering ARMMs comprising one or more therapeutic agents (e.g., biomolecules including, but not limited to, CRISPR / Cas9 and other similar endonucleases, base editors, small molecules, proteins, and nucleic acids (e.g., DNA, RNA, siRNA, mRNA, miRNA, and similar)). The present invention also provides methods for administering therapeutic agents associated with ARMM, including, but not limited to, one or more typical drug regimens (e.g., 1) in vivo administration of ARMM to a patient, 2) ex vivo administration of ARMM to target cells and transplantation of ARMM-treated cells to a patient, and 3) in vivo and ex vivo regimens), which include methods for treating or contacting cells and tissues. Furthermore, the present invention relates to methods for producing (e.g., culturing, clarifying, separating, and concentrating) compositions of the present invention obtained from stable producer cell lines and / or transient cell cultures.
[0157] ARMM Arrestin domain-containing protein 1-mediated microvesicles ("ARMMs") are extracellular vesicles ("EVs") distinct from exosomes. ARMM budding requires ARRDC1, which is localized on the cytoplasmic side of the plasma membrane and, via a tetrapeptide motif, recruits the ESCRT-I complex protein TSG101 to the cell surface, initiating outward membrane budding. Thus, in contrast to exosomes, ARMM biosynthesis occurs in the plasma membrane. ARMMs exhibit several further characteristics that make them a potentially ideal medium for therapeutic delivery. ARRDC1 is not only required for ARMM budding, but also sufficient. Overexpression of the ARRDC1 protein increases intracellular ARMM production, enabling controlled production of ARMMs using state-of-the-art biological manufacturing methods. Furthermore, endogenous proteins such as cell surface receptors can be actively recruited into the ARMM and delivered to recipient cells to initiate intercellular communication, suggesting that exogenous payload molecules can similarly be packaged and delivered via the ARMM.
[0158] In some cases, the ARMM further includes a targeting portion, for example, a targeting portion that targets a specific cell type, such as a targeting portion described herein.
[0159] ARRDC1 In a preferred embodiment, ARRDC1 is a protein containing the PSAP (SEQ ID NO: 1) and PPXY (SEQ ID NO: 2) motifs at its C-terminus and interacting with TSG101 as described herein. It should be understood that the PSAP (SEQ ID NO: 1) and PPXY (SEQ ID NO: 2) motifs do not necessarily have to be at the absolute C-terminus of ARRDC1. Rather, these motifs may be located in the C-terminal portion of the ARRDC1 protein (e.g., the C-terminal half of ARRDC1). The disclosure also includes variants of ARRDC1, e.g., fragments of the ARRDC1 protein, and / or ARRDC1 proteins having some degree of identity to the ARRDC1 protein (e.g., 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity) and capable of interacting with TSG101. Thus, the ARRDC1 protein may be a protein containing the PSAP (SEQ ID NO: 1) and PPXY (SEQ ID NO: 2) motifs and interacting with TSG101. In some embodiments, the ARRDC1 protein is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identical to any one amino acid sequence among SEQ ID NOs. 42-44, contains the PSAP (SEQ ID NO: 1) motif and the PPXY (SEQ ID NO: 2) motif, and interacts with TSG101.In some embodiments, the ARRDC1 protein is at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, and at least It also has 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, at least 390, at least 400, at least 410, at least 420, or at least 430 identical consecutive amino acids, contains the PSAP (SEQ ID NO: 1) motif and the PPXY (SEQ ID NO: 2) motif, and interacts with TSG101. In some embodiments, the ARRDC1 protein has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences shown in SEQ ID NOs: 42-44, and contains the PSAP (SEQ ID NO: 1) motif and the PPXY (SEQ ID NO: 2) motif, and interacts with TSG101. In some embodiments, the ARRDC1 protein contains any one of the amino acid sequences shown in SEQ ID NOs: 42-44. Typical, non-limiting ARRDC1 protein sequences are provided herein, and further, suitable ARRDC1 protein variants according to aspects of the present invention are known in the art. Those skilled in the art will understand that the present invention is not limited in this respect. Typical ARRDC1 sequences include the following (marked with the PSAP (SEQ ID NO: 1) and PPXY (SEQ ID NO: 2) motifs):
[0160] >gi|22748653|ref|NP_689498.1|Arrestin Domain-Containing Protein 1 [Homo sapiens]
[0161] [ka]
[0162] >gi|244798004|ref|NP_001155957.1|Arrestin Domain-Containing Protein 1 Isoform a [Mus musculus]
[0163] [ka]
[0164] >gi|244798112|ref|NP_848495.2|Arrestin domain-containing protein 1 isoform b [Mus musculus]
[0165] [ka]
[0166] TSG101 In certain embodiments, the microendoplasmic reticulum of the present invention further comprises TSG101 (tumor susceptibility gene 101), which belongs to the group of putative inactive homologs of ubiquitin-conjugate enzymes. This protein has a coiled-coil domain that interacts with stasmin, a cytoplasmic phosphoprotein involved in tumorigenesis. TSG101 is a protein that contains a UEV domain and interacts with ARRDC1. As described herein, UEV refers to a ubiquitin E2 variant domain of approximately 145 amino acids. The structure of this domain contains an α / β fold similar to that of a standard E2 enzyme, but has an additional N-terminal helix and further lacks two C-terminal helices. As is often seen in the TSG101 / Vps23 protein, the UEV interacts with ubiquitin molecules and is essential for the transport of many ubiquitinated payloads into the multiendoplasmic reticulum (MVB). Furthermore, the UEV domain can bind to Pro-Thr / Ser-Ala-Pro peptide ligands, which are utilized by viruses such as HIV. Therefore, the TSG101 UEV domain binds to the PTAP tetrapeptide motif within the viral Gag protein involved in viral budding. This disclosure also considers variants of TSG101, e.g., fragments of the TSG101 protein, and / or TSG101 proteins that have some degree of identity to the TSG101 protein (e.g., 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity) and can interact with ARRDC1. Thus, the TSG101 protein may be a protein containing a UEV domain that interacts with ARRDC1. In some embodiments, the TSG101 protein is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 45-47, contains a UEV domain, and interacts with ARRDC1.In some embodiments, the TSG101 protein has an identical sequence of amino acids, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, or at least 390, and any integer between these numbers, and contains a UEV domain that interacts with ARRDC1. In some embodiments, the TSG101 protein has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences shown in SEQ ID NOs.45-47, and contains a UEV domain. In some embodiments, the ARRDC1 protein contains any one of the amino acid sequences shown in SEQ ID NOs.45-47. Typical, non-limiting TSG101 protein sequences are provided herein, and further suitable TSG101 protein sequences, isoforms, and variants are known in the art. Those skilled in the art will understand that the present invention is not limited in this respect. Typical TSG101 sequences include the following (the UEV domains in these sequences contain amino acids 1-145 and are underlined in the following sequences):
[0167] >gi|5454140|ref|NP_006283.1|Tumor susceptibility gene 101 protein [Homo sapiens]
[0168] [ka]
[0169] >gi|11230780|ref|NP_068684.1|Tumor susceptibility gene 101 protein [Mus musculus]
[0170] [ka]
[0171] >gi|48374087|ref|NP_853659.2|Tumor susceptibility gene 101 protein [Norwegian rat (Rattus norvegicus)]
[0172] [ka]
[0173] The structure of the UEV domain is known to those skilled in the art. (For example, see Owen Pornillos et al., Structure and functional interactions of the Tsg101 UEV domain, EMBO J., 21(10): 2397-2406 (2002), whose entire contents are incorporated herein by reference.)
[0174] RNA-guided proteins In some cases, the ARMMs described herein include RNA-guided proteins. RNA-guided proteins such as Cas (e.g., Cas9) proteins, base editors, prime editors, and similar proteins are described in the Art.
[0175] In some cases, the RNA guide protein is a base editor (e.g., a protein comprising a fusion protein containing programmable DNA-binding domains such as a Cas domain and a deaminase domain). In some cases, the deaminase is cytosine deaminase and / or adenosine deaminase. In some embodiments, the adenosine deaminase is TadA or a TadA variant. In some embodiments, TadA is TadA * 8 or TadA * The answer is 9. In some cases, cytosine deaminase is APOBEC or an APOBEC variant. In some cases, the base editor is a cytosine base editor. In some cases, the base editor is an adenine base editor.
[0176] In some cases, the base editor is for International Publication Nos. 2023 / 023515, 2023 / 015223, 2023 / 009505, 2023 / 004409, 2022 / 256578, 2022 / 251687, 2022 / 246266, 2022 / 241270, 2022 / 221699, 2022 / 204574, International Publication Publication No. 2022 / 204268, International Publication No. 2022 / 178307, International Publication No. 2022 / 159472, International Publication No. 2022 / 159463, International Publication No. 2022 / 159475, International Publication No. 2022 / 159421, International Publication No. 2022 / 140252, International Publication No. 2022 / 140239, International Publication No. 2022 / 140238, International Publication No. 2022 / 081890, International Publication No. 2022 / 067089, International Publication No. International Publication No. 2022 / 060871, International Publication No. 2021 / 207712, International Publication No. 2021 / 178725, International Publication No. 2021 / 163587, International Publication No. 2021 / 113494, International Publication No. 2021 / 062227, International Publication No. 2021 / 050512, International Publication No. 2021 / 050571, International Publication No. 2021 / 041885, International Publication No. 2021 / 041945, International Publication No. 2020 / 231863, International Publication No. 20 This is a base editor described in any one of the following publications: 20 / 150534, International Publication 2020 / 051561, International Publication 2020 / 051562, International Publication 2020 / 028823, International Publication 2019 / 217942, International Publication 2019 / 217943, International Publication 2019 / 217941, or International Publication 2019 / 079347, each of which is incorporated in its entirety by reference.
[0177] In some cases, the ARMM described herein comprises a nucleic acid that codes for an RNA-guided protein. In some cases, the nucleic acid is DNA, which codes for mRNA, and the mRNA then codes for an RNA-guided protein. In some cases, the nucleic acid is mRNA that codes for an RNA-guided protein.
[0178] SHRNA In some cases, ARMM contains one or more gRNAs, which are described in the art. In some cases, the gRNA is a single guide RNA (sgRNA). In some cases, the gRNA is crisprRNA + tracrRNA. In some cases, the gRNA is a combination of sgRNA and crisprRNA + tracrRNA.
[0179] Expression construct Some embodiments of the present invention provide expression constructs for encoding one or more gene products that induce or promote ARMM production in cells having such constructs. In some embodiments, the expression constructs described herein encode fusion proteins described herein, e.g., ARRDC1 fusion protein and TSG101 fusion protein. In some embodiments, the expression construct encodes the ARRDC1 protein or a variant thereof, and / or the TSG101 protein or a variant thereof. In some embodiments, overexpression of either or both gene products in cells increases ARMM production in the cell, thus transforming the cell into a microendoplasmic reticulum-producing cell. In some embodiments, such an expression construct includes at least one restriction or recombination site at either the C-terminus or N-terminus of the encoded ARRDC1 or a variant thereof that enables in-frame cloning of the protein sequence to be fused. As another example, the expression construct includes at least one restriction or recombination site at either the C-terminus or N-terminus of one or more encoded WW domains that enables in-frame cloning of the protein sequence to be fused.
[0180] In some embodiments, the expression construct includes (a) a nucleotide sequence encoding the ARRDC1 protein or a variant thereof, operably linked to a heterologous promoter, and (b) a restriction or recombination site located adjacent to the nucleotide sequence encoding ARRDC1, which enables in-frame insertion of a nucleotide sequence encoding a payload protein or an RNA-binding protein sequence or an RNA-binding protein variant sequence with the nucleotide sequence encoding ARRDC1. In some embodiments, the heterologous promoter may be a constitutive promoter, and in some embodiments, the heterologous promoter may be an inductive promoter. Some aspects of the present invention provide an expression construct including (a) a nucleotide sequence encoding the TSG101 protein or a variant thereof, operably linked to a heterologous promoter, and (b) a restriction or recombination site located adjacent to the nucleotide sequence encoding TSG101, which enables in-frame insertion of a nucleotide sequence encoding a payload protein or an RNA-binding protein, a DNA-binding protein, or a variant sequence thereof with the nucleotide sequence encoding TSG101. In some embodiments, the heterogeneous promoter may be a constitutive promoter, and in some embodiments, the heterogeneous promoter may be an inductive promoter.
[0181] Some aspects of the present invention provide an expression construct comprising (a) a nucleotide sequence encoding a WW domain or a variant thereof, operably linked to a heterologous promoter, and (b) a restriction or recombination site located adjacent to the nucleotide sequence encoding the WW domain, which enables in-frame insertion of a payload protein or RNA-binding protein or a variant sequence thereof with respect to the WW domain-encoding nucleotide sequence. In some embodiments, the heterologous promoter may be a constitutive promoter, and in some embodiments, the heterologous promoter may be an inductive promoter. The expression construct may encode a payload protein or an RNA-binding protein fused to at least one WW domain. In some embodiments, the expression construct encodes a payload protein or an RNA-binding protein or a variant thereof fused to at least one WW domain or a variant thereof. Any expression construct described herein may encode any WW domain or a variant thereof. In some embodiments, the heterologous promoter may be a constitutive promoter, and in some embodiments, the heterologous promoter may be an inductive promoter.
[0182] The expression constructs described herein may include any nucleic acid sequence capable of encoding a WW domain or a variant thereof. For example, nucleic acid sequences encoding a WW domain or WW domain variant may be derived from human ubiquitin ligases WWP1, WWP2, Nedd4-1, Nedd4-2, Smurf1, Smurf2, ITCH, NEDL1, or NEDL2. Typical nucleic acid sequences of proteins having a WW domain are listed below. It should be understood that any nucleic acid encoding a WW domain or WW domain variant of a typical protein may be used in the present invention, is described herein, or is not limited thereto.
[0183] Human WWP1 nucleic acid sequence (uniprot.org / uniprot / Q9H0M0).
[0184] [ka]
[0185] [ka]
[0186] Human WWP2 nucleic acid sequence (uniprot.org / uniprot / O00308).
[0187] [ka]
[0188] [ka]
[0189] Human Nedd4-1 nucleic acid sequence (uniprot.org / uniprot / P46934).
[0190] [ka]
[0191] Human Nedd4-2 nucleic acid sequence (>gi|345478679|ref|NM_015277.5|NEDD4L, a developmentally downregulated E3 ubiquitin protein ligase expressed by Homo sapiens neural progenitor cells, transcript variant d, mRNA).
[0192] [ka]
[0193] Human Smurf1 nucleic acid sequence (uniprot.org / uniprot / Q9HCE7).
[0194] [ka]
[0195] Human Smurf2 nucleic acid sequence (uniprot.org / uniprot / Q9HAU4).
[0196] [ka]
[0197] Human ITCH nucleic acid sequence (uniprot.org / uniprot / Q96J02).
[0198] [ka]
[0199] Human NEDL1 nucleic acid sequence (uniprot.org / uniprot / Q76N89).
[0200] [ka]
[0201] [ka]
[0202] [ka]
[0203] Human NEDL2 nucleic acid sequence (uniprot.org / uniprot / Q9P2P5).
[0204] [ka]
[0205] Some aspects of the present invention provide expression constructs that encode any of the proteins, nucleic acids such as RNA, or fusions thereof as described herein.
[0206] Nucleic acids encoding any of the proteins and / or nucleic acids (including RNA) described herein may be contained within any number of nucleic acid vectors known in the art. Suitable vectors for use in the compositions and methods of the present invention include both viral and non-viral products, as well as additional means for introducing nucleic acids into cells.
[0207] The expression of any of the proteins and / or nucleic acids (including RNA) described herein can be controlled by any regulatory sequence (e.g., promoter sequence) known in the art. Regulatory sequences described herein are nucleic acid sequences that regulate the expression of a nucleic acid sequence. Regulatory or regulatory sequences may include sequences involved in the expression of a particular nucleic acid, or other sequences, e.g., heterologous sequences, synthetic sequences, or partially synthetic sequences. Sequences may be of eukaryotic, prokaryotic, or viral origin, stimulating or repressing gene transcription in specific or nonspecific ways and in an inducible or non-inducible way. Regulatory or regulatory regions may include origins of replication, RNA splice sites, introns, chimeric introns or hybrid introns, promoters, enhancers, transcription termination sequences, poly(A) sites, locus regulatory regions, and signaling sequences that direct polypeptides to the secretory pathways of target cells. Heterologous regulatory regions are regulatory regions that are not naturally linked to the expressed nucleic acid to which they ligate. Heterogeneous regulatory regions include regulatory regions derived from different species, regulatory regions derived from different genes, hybrid regulatory sequences, and regulatory sequences that do not occur naturally, all of which are designed by those skilled in the art.
[0208] Typical cells that produce ARMM containing a payload The microendoplasmic reticulum-producing cells of the present invention may be cells containing any of the expression constructs, any of the fusion proteins, or a payload of molecules (e.g., biomolecules, small molecules, proteins, and nucleic acids as described herein). For example, the microendoplasmic reticulum-producing cells of the present invention may contain one or more recombinant expression constructs encoding (1) the ARRDC1 protein, or a variant thereof having the PSAP (SEQ ID NO: 1) motif, and (2) an RNA-binding protein (e.g., Tat protein) bound to the ARRDC1 protein or a variant thereof having the PSAP (SEQ ID NO: 1) motif. In some embodiments, the microendoplasmic reticulum-producing cells may, under the control of a heterologous promoter, contain one or more recombinant expression constructs encoding (1) the ARRDC1 protein, or a variant thereof having the PSAP (SEQ ID NO: 1) motif, and (2) a payload protein, e.g., an RNA-binding protein fused to at least one WW domain, or a variant thereof. In certain embodiments, the expression construct in the microendoplasmic reticulum-producing cells encodes a payload protein having one or more WW domains or a variant thereof. In a particular embodiment, the expression construct in microendoplasmic reticulum-producing cells encodes an RNA-binding protein, such as a therapeutic RNA, which specifically binds to (e.g., a therapeutic RNA).
[0209] Any of the expression constructs described herein can be stably inserted into the cellular genome. In some embodiments, the expression construct is maintained within the cell but not inserted into the cellular genome. In some embodiments, the expression construct is contained within a vector, e.g., a plasmid vector, cosmid vector, viral vector, or artificial chromosome. In some embodiments, the expression construct further comprises additional sequences or elements that promote the maintenance and / or replication of the expression construct in microendoplasmic reticulum-producing cells or enhance the expression of the fusion protein within the cell. Such additional sequences or elements include, for example, origins of replication, antibiotic resistance cassettes, poly(A) sequences, and / or transcription isolators. Some expression constructs suitable for the production of microendoplasmic reticulum-producing cells according to embodiments of the present invention are described elsewhere in this specification. Methods and reagents for producing further expression constructs suitable for the production of microendoplasmic reticulum-producing cells according to embodiments of the present invention will become apparent to those skilled in the art based on this disclosure. In some embodiments, the microendoplasmic reticulum-producing cells are mammalian cells, e.g., mouse cells, rat cells, hamster cells, rodent cells, or non-human primate cells. In some embodiments, the microendoplasmic reticulum-producing cells are human cells.
[0210] Those skilled in the art can utilize prior art, such as molecular biology or cell biology, virology, microbiology, and recombinant DNA technology. Typical techniques are fully described in the literature. For example, the present invention can be carried out and used in accordance with the following general texts: Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, and Sambrook et al., Third Edition (2001), DNA Cloning: A Practical Approach, Volumes I and II (DN Glover ed. 1985), Oligonucleotide Synthesis (MJ Gaited. 1984), Nucleic Acid Hybridization (BD Hames & SJ Higgins eds. (1985)), Transcription and Translation Hames & Higgins, eds. (1984); Animal Cell Culture (RI. Freshney, ed. (1986)), Immobilized Cells And Enzymes (IRL Press, (1986)), Gennaro et al., (eds.) Remington's Pharmaceutical Sciences, 18th edition; B. Perbal, A Practical Guide To Molecular Cloning (1984), FM Ausubel et al., (eds.), Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (updates through 2001), Colligan et al., (eds.), Current Protocols in Immunology, John Wiley & Sons, Inc.(updates through 2001), W. Paul et al., (eds.) Fundamental Immunology, Raven Press; EJ Murray et al., (ed.) Methods in Molecular Biology: Gene Transfer and Expression Protocols, The Humana Press Inc. (1991) (especially vol.7), and JE Celis et al., Cell Biology: A Laboratory Handbook, Academic Press (1994). .
[0211] Delivery of ARMM containing payload molecules The microvesicle endoplasmic reticulum of the present invention (e.g., an ARMM containing any of the expression constructs and / or payloads of molecules (e.g., therapeutic agents, biomolecules, small molecules, proteins, and nucleic acids (e.g., DNA, RNA), DNA plasmids, siRNA, shRNA, mRNA)) may optionally further include a targeting moiety. The targeting moiety can be used to target the delivery of the ARMM to a specific cell type, resulting in the release of the contents of the ARMM into the cytoplasm of the specific target cell type. The targeting moiety may be a viral envelope protein or a portion thereof, which typically functions to assist the attachment and entry of the virus into cells. The viral envelope protein may enable targeting of cells in the CNS. Examples of viral envelope proteins include, but are not limited to, vesicular stomatitis virus G protein (VSV-G, Genbank commissioned and version number: AJ318514.1) or rabies virus glycoprotein (RVG, Genbank commissioned and version number: M38452.1). The VSV-G protein facilitates viral entry by mediating viral attachment to LDL receptors (LDLRs) or LDLR family members present on target cells. Following binding, the VSV-G-LDLR complex is immediately endocytotic, subsequently mediating the fusion of the viral envelope with the endosomal membrane. VSV-G enters cells via partially clathrin-coated vesicles, which contain more clathrin and clathrin adapters than conventional vesicles. VSV-G is a common coat protein in vector expression systems used to introduce genetic material into in vitro systems or animal models, primarily due to its extremely broad targeting. RVG is a trimer, surface-exposed viral coat protein known to utilize nicotinic acetylcholine receptors and low-affinity nerve growth factor receptors for viral entry. In some embodiments, viral envelope proteins (e.g., VSV-G, RVG) facilitate the binding (e.g., targeting) of ARMM to target cells.
[0212] In some embodiments, producer cells (e.g., Expi293 cells) are optionally transfected with approximately 0.5–25% of the total plasmid VSV-G expression construct per transfection. In some further embodiments, producer cells are optionally transfected with approximately 10–25% of the total plasmid VSV-G expression construct per transfection.
[0213] The targeting moiety can selectively bind to the surface antigen of the target cell. For example, the targeting moiety may be a membrane-bound immunoglobulin, integrin, receptor, receptor ligand, aptamer, small molecule, or variant thereof. Any number of cell surface proteins may also be included in ARMM to facilitate the binding of ARMM to the target cell and / or the uptake of ARMM into the target cell. Integrins, receptor tyrosine kinases, G protein-coupled receptors, and membrane-bound immunoglobulins suitable for use with embodiments of the present invention are obvious to those skilled in the art, and the present invention is not limited in this respect. For example, in some embodiments, the integrin is α1β1, α2β1, α4β1, α5β1, α6β1, αLβ2, αMβ2, αIIbβ3, αVβ3, αVβ5, αVβ6, or α6β4 integrin. In some embodiments, the receptor tyrosine kinase is an EGF receptor (ErbB family), insulin receptor, PDGF receptor, FGF receptor, VEGF receptor, HGF receptor, Trk receptor, Eph receptor, AXL receptor, LTK receptor, TIE receptor, ROR receptor, DDR receptor, RET receptor, KLG receptor, RYK receptor, or MuSK receptor. In some embodiments, the G protein-coupled receptor is a rhodopsin-like receptor, secretin receptor, metabotropic glutamate / pheromone receptor, cyclic AMP receptor, frizzled / smoothed receptor, CXCR4, CCR5, or beta-adrenergic receptor.
[0214] Further molecules, such as synthetic small molecules or naturally occurring products, can be modified to bind to ARMM proteins (e.g., TSG101 or ARRDC1) for targeting purposes. This binding may facilitate their incorporation into ARMM, which can then be used to increase the delivery of ARMM to target cells. By incorporating a cleavable linker, the small molecule can be released into or upon delivery to target cells. As a non-limiting example, the small molecule can be linked to biotin, allowing it to bind to the ARRDC1 protein fused to streptavidin. As another non-limiting example, small molecules can be linked to synthetic high-affinity ligands that specifically bind to mutant forms of FKBP12, such as FKBP12(F36V) (Yang, W., et al., “Investigating protein-ligand interactions with a mutant FKBP possessing a designed specificity pocket” J. Med. Chem., 43(6):1135-1142 (2000)), which then bind to the ARRDC1 protein fused to FKBP12(F36V). The binding of small molecules to ARMM proteins (e.g., TSG101 or ARRDC1) facilitates the loading of small molecules into ARMMs containing ARRDC1.
[0215] Some aspects of the present invention relate to the recognition that an ARMM is taken up by a target cell, and as a result of the uptake of the ARMM, the contents of the ARMM are released into the cytoplasm of the target cell. In some embodiments, the payload is a drug that causes a desired change in the target cell, such as a change in cell survival, proliferation rate, differentiation stage, cell identity, chromatin state, transcription rate of one or more genes, transcription profile, or post-transcriptional changes in gene compression of the target cell, and similar changes. Those skilled in the art will understand that the drug to be delivered is selected according to the desired effect on the target cell (e.g., a base editor).
[0216] In some embodiments, target cells are obtained, and the payload is delivered to the cells ex vivo by a system or method provided herein. In some embodiments, the treated cells are selected for those in which the desired gene is expressed or suppressed. In some embodiments, the treated cells having the desired payload are returned to the target from which the cells were obtained.
[0217] In some embodiments, the ARMM further includes a detectable label. In certain embodiments, the detectable label ARMM makes it possible to label target cells without genetic manipulation. Detectable labels suitable for direct delivery to target cells are known in the art and include, but are not limited to, fluorescent proteins, fluorescent dyes, membrane-bound dyes, and enzymes, such as membrane-bound enzymes or cytoplasmic enzymes, which catalyze a reaction that produces a detectable reaction product. Suitable detectable labels according to some aspects of the present invention further include membrane-bound antigens, such as membrane-bound ligands, which can be detected by commonly available antibodies or antigen conjugates. Detectable label ARMMs are used in a variety of diagnostic and analytical methods and applications.
[0218] In some embodiments, an ARMM is provided that comprises a payload RNA encoding a transcription factor, transcription repressor, fluorescent protein, kinase, phosphatase, protease, ligase, chromatin regulator, recombinase, and similar. In some embodiments, an ARMM is provided that comprises a payload RNA that inhibits the expression of a transcription factor, transcription repressor, fluorescent protein, kinase, phosphatase, protease, ligase, chromatin regulator, or recombinase. In some embodiments, the payload RNA is therapeutic RNA. In some embodiments, the payload RNA is RNA that brings about a change in the state or identity of a target cell. For example, in some embodiments, the payload RNA encodes a reprogramming factor. Appropriate transcription factors, transcription repressors, fluorescent proteins, kinases, phosphatases, proteases, ligases, chromatin regulators, recombinases, and reprogramming factors may be encoded by a payload RNA that binds to a binding RNA to facilitate their incorporation into the ARMM, and their functions can be tested by any method known to those skilled in the art, and the present invention is not limited in this respect.
[0219] Methods for isolating ARMMs described herein are also provided. One typical method involves utilizing conventional techniques to recover culture medium or supernatant from a cell culture containing microendoplasmic reticulum-producing cells. In some embodiments, the cell culture contains cells obtained from a subject, e.g., cells suspected of exhibiting a pathological phenotype, e.g., a hyperproliferative phenotype. In some embodiments, the cell culture contains genetically engineered cells that produce ARMMs, e.g., cells expressing a recombinant protein, e.g., a recombinant ARRDC1 or TSG101 protein, e.g., ARRDC1 or TSG101 protein optionally fused to an RNA-binding protein (e.g., Tat protein) or a variant thereof. In some embodiments, the supernatant is pre-cleaned of cell debris by centrifugation, e.g., two consecutive centrifugations with increasing G values (e.g., 500 G and 2000 G). In some embodiments, the method includes passing the supernatant through a 0.2 μm filter to remove all large cell debris and whole cells. In some embodiments, the supernatant is ultracentrifuged for 2 hours at, for example, 120,000 G, depending on the volume of the centrifugated material. The resulting pellet contains microendoplasmic reticulum. In some embodiments, exosomes are depleted from the microendoplasmic reticulum pellet by staining and / or sorting (e.g., by FACS or MACS) using exosome markers described herein. The isolated or concentrated ARMM can be suspended in culture medium or a suitable buffer as described herein.
[0220] ARMM-mediated delivery method of payload to cells Some aspects of the present invention provide a method for delivering a drug (e.g., one or more therapeutic drugs) to target cells. Target cells can be brought into contact with ARMM in various ways. For example, target cells can be brought into direct contact with ARMM as described herein, or with ARMM isolated from microendoplasmic reticulum-producing cells. Contact can be made in vitro by administering ARMM to target cells in a culture dish, or in vivo by administering ARMM to a subject. In some embodiments, ARMM is produced from cells obtained from a subject. In some embodiments, ARMM produced from cells obtained from a particular subject is administered to the same subject. Conversely, in some other embodiments, ARMM produced from cells obtained from a subject is administered to different subjects. For example, cells can be obtained from a subject and manipulated to express one or more of the constructs provided herein (e.g., payload RNA bound to binding RNA, ARRDC1 protein, ARRDC1 protein fused to an RNA-binding protein, RNA-binding protein fused to a WW domain, Cas9 protein, base editors, guide and regulatory sequences, and similar).
[0221] Alternatively, target cells may be brought into contact with microendoplasmic reticulum-producing cells as described herein, for example, in vitro by co-culturing the target cells and microendoplasmic reticulum-producing cells, or in vivo by administering microendoplasmic reticulum-producing cells to a subject having target cells. Thus, the method may include bringing target cells into contact with microendoplasmic reticulum and an ARMM containing, for example, one of the payloads of the target to be delivered as described herein. Target cells can be brought into contact with microendoplasmic reticulum-producing cells as described herein, or with isolated microendoplasmic reticulum having a lipid bilayer, ARRDC1 protein or a variant thereof, a payload (e.g., a therapeutic agent), and optionally a viral envelope protein.
[0222] It should be understood that target cells can originate from anything, for example, from living organisms. In some embodiments, target cells are mammalian cells. Some non-limiting examples of mammalian cells include, but are not limited to, mouse cells, rat cells, hamster cells, rodent cells, and non-human primate cells. In some embodiments, target cells are human cells. It should also be understood that target cells can be of any cell type. In other cases, target cells can be any differentiated cell type found in the subject. In some embodiments, target cells are in vitro cells, and the method involves administering microvesicles to cells in vitro, or co-culturing target cells with microvesicle-producing cells in vitro. In some embodiments, target cells are cells within the subject, and the method involves administering microvesicles or microvesicle-producing cells to the subject. In some embodiments, the subject is a mammalian subject, for example, a rodent, mouse, rat, hamster, or non-human primate. In some embodiments, the subject is a human subject.
[0223] In some embodiments, the target cell is a pathological cell. In some embodiments, the target cell is a cancer cell. In some embodiments, the microendoplasmic reticulum is associated with a binder that selectively binds an antigen on the surface of the target cell. In some embodiments, the compositions and methods of the present invention include one or more targeting ligands that bind (e.g., bind) to one or more targeting receptors. In some embodiments, the antigen of the target cell is a cell surface antigen. In some embodiments, the binder is a membrane-bound immunoglobulin, integrin, receptor, receptor ligand, or lectin, among other suitable candidate molecules and candidate moieties. Suitable surface antigens of the target cell are known to those skilled in the art to be suitable binders that specifically bind such antigens. Methods for producing membrane-bound binders, such as membrane-bound immunoglobulins, membrane-bound antibodies, or antibody fragments that specifically bind surface antigens expressed on the surface of cells, are also known to those skilled in the art. The selection of the binder will, of course, depend on the identity or type of the target cell. The various types of optical tissues and cell surface antigens that are specifically expressed on cells that can be targeted by ARMMs containing membrane-bound binders will be apparent to those skilled in the art. It will be understood that the present invention is not limited in this regard.
Example
[0224] [Example 1] ARRDC1-TEV-G4S-ABE8v.1 ARRDC1 was cloned in-frame with the ABE8 base editor into the plasmid backbone to obtain an expression construct having the following DNA sequence (the base editor portion of the construct is shown in bold):
[0225]
Chemical formula
[0228] The expressed polyprotein has the following structure: [ARRDC]-[linker]-[TEV protease cleavage sequence]-[linker]-[base editor (ABE8)], and the following amino acid sequence (the base editor portion of the polyprotein is shown in bold):
[0229] [ka]
[0230] [Example 2] ARRDC1-TEV-G4S-ABE8v.2 ARRDC1 was cloned into the plasmid backbone in frame with the ABE8 base editor to obtain an expression construct with the following DNA sequence:
[0231] [ka]
[0232] [ka]
[0233] The expressed polyprotein has the following structure: [ARRDC]-[linker]-[TEV protease cleavage sequence]-[linker]-[base editor (ABE8)], and the following amino acid sequence:
[0234] [ka]
[0235] [Example 3] ARRDC1-mini-3NES-HIV-G4S-ABE8 DNA sequence
[0236] [ka]
[0237]
Chem.
[0238]
Chem.
[0239] Amino acid sequence
[0240]
Chem.
[0241] [Example 4] ARRDC1-TEV-G4S-CBE The DNA sequence in which the CBE portion of the construct is shown in boldface.
[0242]
Chem.
[0243]
Chem.
[0244]
Chem.
[0245] The amino acid sequence in which the CBE portion of the polyprotein is shown in boldface.
[0246]
Chem.
[0247] [Example 5] ARRDC1-TEV-G4S-TadCBEd-V106W The DNA sequence where the CBE portion of the construct is shown in bold.
[0248] [ka]
[0249] [ka]
[0250] [ka]
[0251] The amino acid sequence of the CBE portion of the polyprotein is shown in bold.
[0252] [ka]
[0253] [Example 6] ARRDC1-Mini-3NES-HIV-G4S-CBE DNA sequence
[0254] [ka]
[0255] [ka]
[0256] [ka]
[0257] amino acid sequence
[0258] [ka]
[0259] [Example 7] ARRDC1-Cas9 ARRDC1-Cas9 expression constructs were created by cloning a full-length DNA fragment of wild-type ARRDC1 and Cas9 into a pcDNA3.1(+) vector using the G4S linker located at the C-terminus of ARRDC1.
[0260] [Example 8] Production, purification, and characterization of ARMM ARMMs were generated by transiently transfecting wild-type Expi293 cells with expression constructs containing nucleic acid sequences encoding the ARRDC1 polynucleotide (see Examples 1-7) and expression constructs containing nucleic acid sequences encoding the gRNA using Expifectamine.
[0261] To purify ARMM, CM was treated by tangential flow filtration (TFF), followed by ultracentrifugation and concentration. ARMM was stored at 4°C.
[0262] To quantify concentration and particle size, NTA was performed using a Particle Metrix Zetaview with the following settings: sensitivity 82.4, shutter speed 100, minimum area 10, maximum area 3000, and minimum brightness 20. Particles were diluted in distilled water to 50-200 particles on the screen, and this value falls within the instrument's linear detection range.
[0263] Western blot analysis was used to quantify ARRDC1 and the payload.
[0264] To quantify gRNAs packaged in ARMMs, a custom Taqman probe was designed for the scaffold sequence within single gRNAs (sgRNAs). The gRNA quantification process involved three steps: RNA sample preparation, cDNA synthesis, and droplet digital PCR (ddPCR). The relative amount of gRNA in the ARMM was reported as a normalized enrichment ratio relative to the control ARMM.
[0265] The results of the quantification of A1-ABE8, various gRNAs, and A1-Cas9 are shown in the table below.
[0266] [Table 1]
[0267] [Example 9] In vivo study to test ARMMs containing ARRDC1-ABE8 and gRNA1 Two C57BL / 6 mice were intranasally administered two 50 μL doses of a solution containing 2E11 ARMM particles (containing ARRDC1-ABE8 and gRNA1) at 30-minute intervals. After 96 hours, the samples were collected and subjected to genomic DNA extraction. Targeted PCR and targeted sequencing were performed to amplify the genomic region containing the target genome editing site. NGS data showing the base editing efficiency at the two editing target sites are shown in Figure 1.
[0268] [Example 10] In vivo study to test ARMM having ARRDC1-ABE8 and gRNA2. Two C57BL / 6 mice were intranasally administered one dose of 50 μL of a solution containing 4E11 ARMM particles (containing ARRDC1-ABE8 and gRNA2).
[0269] After 96 hours, the samples were collected and subjected to genomic DNA extraction. Targeted PCR and targeted sequencing were performed to amplify the genomic region containing the target genome editing site. Figure 2 shows NGS data indicating the base editing efficiency at the two editing target sites.
[0270] [Example 11] In vitro testing of the activity of ARRDC1-CBE constructs using HEK2 gRNA and HEK3 gRNA. On day 1, HEK293 cells were placed in a 6-well plate containing 10% DMEM, with 2 × 10⁶ cells per well. 5 They were sown individually.
[0271] On day 2, 1 ug of DNA (0.5 ug of ARRDC1-CBE + 0.5 ug of HEK2 or HEK3) was mixed with 1.5 uL of transfection reagent (expifectamine) and applied to cells according to the manufacturer's recommendations.
[0272] On the fourth day, cells were collected and subjected to genomic DNA extraction. Genomic DNA was extracted using the DNeasy Blood & Tissue Kit (Qiagen) as recommended by the manufacturer.
[0273] On day 5, targeted PCR was performed to amplify the genomic region containing the genome editing site. One microliter of gDNA solution was used as a template in the PCR reaction using Q5 High-Fidelity 2X Master Mix (NEB, M0492) according to the manufacturer's recommendations. The PCR product was confirmed by electrophoresis and purified using the Qiagen PCR Purification Kit (Qiagen, 28181). To evaluate the editing, the PCR product was subjected to Sanger sequencing or NGS analysis.
[0274] Sanger sequencing data showing the base editing efficiency by CBE using HEK2 gRNA and HEK3 gRNA in human HEK293 cells are shown in Figures 3A and 3B.
[0275] [Example 12] In vitro evaluation of the introduction of a stop codon to terminate translation using an ARRDC1-CBE construct and a GFP-targeting gRNA. On day 1, HEK293 cells expressing GFP were seeded at a rate of 2e5 cells per well in a 6-well plate containing 10% DMEM.
[0276] On day 2, 1 µg of DNA (0.5 µg of ARRDC1-CBE + 0.5 µg of GFP-gRNA) was mixed with 1.5 µL of transfection reagent (expifectamine) and applied to cells according to the manufacturer's recommendations.
[0277] On days 4 and 6, 20% of the cells in each well were reseed.
[0278] On day 8, cells were collected, and 50% of the cells were analyzed by flow cytometry to quantify GFP-positive cells. The remaining 50% of the cells were used for genomic DNA extraction.
[0279] On day 9, genomic DNA was extracted using the DNeasy Blood & Tissue Kit according to the manufacturer's recommendations. Targeted PCR was performed to amplify the genomic region containing the genome editing site. One microliter of gDNA solution was used as a template in the PCR reaction using Q5 High-Fidelity 2X Master Mix (NEB, M0492) according to the manufacturer's recommendations. The PCR product was confirmed by electrophoresis on an E gel and purified using the Qiagen PCR Purification Kit (Qiagen, 28181). To evaluate the editing, the PCR product was subjected to Sanger sequencing or NGS analysis.
[0280] Figure 4 shows Sanger sequencing data indicating the base editing efficiency of CBE using GFP gRNA in human HEK293 cells that stably express GFP.
[0281] Flow cytometry data showing GFP-negative cells in HEK293 cells that stably express GFP, treated with ARMM loaded with CBE and GFP targeting gRNA. Figures 5A, 5B, and 5C show "C to T" base editing by CBE using GFP gRNA, which can introduce an immature stop codon that terminates the translation of the nascent GFP transcript.
[0282] References All publications, patents, and sequence database entries referenced herein, including those listed above, are incorporated by reference in whole, as if specifically and individually indicated therein that each individual publication or patent is incorporated by reference. In case of any conflict, this application shall prevail, including all definitions herein.
[0283] Equivalents and range Those skilled in the art will recognize or confirm many equivalents of the specific embodiments of the invention described herein by means of ordinary testing. The scope of the invention is not limited to the foregoing description, but rather as set forth in the appended claims.
[0284] In the claims, articles such as “a,” “an,” and “the” may mean one or more unless otherwise indicated or the context makes it clear. Unless otherwise indicated or the context makes it clear, a claim or statement containing “or” between one or more components of the group is considered to satisfy the condition if one, two or more, or all of the components of the group are present in, used in, or otherwise related to a given product or process. The present invention includes embodiments in which exactly one component of the group is present in, used in, or otherwise related to a given product or process. The present invention also includes embodiments in which two or more, or all of the components of the group are present in, used in, or otherwise related to a given product or process.
[0285] Furthermore, terms such as “first,” “second,” etc., may be used herein to describe various elements, but these elements are not limited by these terms. These terms are used solely to distinguish one element from another. Thus, the “first” element discussed below may also be referred to as the “second” element without departing from the teachings of this disclosure. The order of operations (or actions / steps) is not limited to the order shown in the claims or drawings unless otherwise specifically indicated.
[0286] Furthermore, it is understood that the present invention encompasses all variations, combinations, and rearrangements in which one or more limitations, elements, clauses, descriptive terms, etc., derived from one or more claims or relevant parts thereof are introduced into another claim. For example, any claim dependent on another claim may be amended to include one or more limitations described in any other claim dependent on the same basic claim. Furthermore, where a claim references a composition, it is understood that, unless otherwise indicated or unless it would be obvious to a person skilled in the art that this would result in a contradiction or inconsistency, it includes a method of using the composition for any of the purposes disclosed herein and a method of preparing the composition according to the manufacturing method disclosed herein or other methods known in the art.
[0287] Where elements are presented enumerated, for example in Markush group format, each subgroup of the elements is also disclosed, and any element may be excluded from the group. Note that the term “includes” is intended to be open and allow for the inclusion of additional elements or steps. In general, where the present invention or an aspect of the present invention is described as including certain elements, features, steps, etc., it should be understood that certain embodiments or aspects of the present invention consist of or are essentially derived from such elements, features, steps, etc. For the purpose of simplification, these embodiments are not specifically shown in this specification in these exact words. Thus, in each embodiment of the present invention that includes one or more elements, features, steps, etc., the present invention also provides embodiments consisting of or being essentially derived from these elements, features, steps, etc.
[0288] Where a range is indicated, it includes endpoints. Furthermore, unless otherwise indicated, or unless otherwise evident from the context and / or the understanding of those skilled in the art, it is understood that values expressed as a range may represent any specific value within the range described in different embodiments of the invention, up to one-tenth of the lower limit of that range, unless otherwise evident from the context. Unless otherwise indicated, or unless otherwise evident from the context and / or the understanding of those skilled in the art, values expressed as a range may represent any subrange within a given range, where the endpoint of the subrange is expressed with a precision equivalent to one-tenth of the lower limit of that range.
[0289] Furthermore, it is understood that any embodiment of the present invention may be expressly excluded from any one or more claims. Where a range is indicated, any value within that range may be expressly excluded from any one or more claims. Any embodiment, element, feature, use, or aspect of the composition and / or method of the present invention may be excluded from any one or more claims. For the sake of brevity, not all embodiments in which one or more elements, features, uses, or aspects are excluded are expressly shown herein.
Claims
1. Arrestin domain-containing protein 1 (ARRDC1)-mediated microendoplasmic reticulum (ARMM), (i) Lipid bilayer and ARRDC1 protein, fragment thereof, or variant thereof, (ii) RNA-guided proteins, and (iii) gRNA The microvesicle endoplasmic reticulum (ARMM) includes the aforementioned microvesicle endoplasmic reticulum.
2. The ARMM according to claim 1, wherein the RNA guide type protein is a Cas fusion protein.
3. The ARMM according to claim 1 or claim 2, wherein the RNA guide protein is a base editor.
4. The ARMM according to any one of claims 1 to 3, wherein the base editor is an adenine base editor.
5. The ARMM according to any one of claims 1 to 3, wherein the base editor is a cytosine base editor.
6. The ARMM according to any one of claims 1 to 5, wherein the base editor includes a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6.
7. The ARMM according to claim 6, wherein the base editor comprises or consists of one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6.
8. The ARMM according to any one of claims 1 to 7, wherein the gRNA is sgRNA.
9. The ARMM according to any one of claims 1 to 7, wherein the gRNA is crRNA + tracrRNA.
10. The ARMM according to any one of claims 1 to 9, wherein the RNA guide protein and gRNA together form a complex as a ribonucleoprotein (RNP).
11. The ARMM according to any one of claims 1 to 10, wherein the ARRDC1 protein, a fragment thereof, or a variant thereof is fused to an RNA guide protein, optionally via an intervening linker and / or cleavage domain.
12. The ARMM according to any one of claims 1 to 10, wherein the RNA guide protein is optionally fused to the WW domain via an intervening linker and / or cleavage domain.
13. (i) ARRDC1 protein, its fragments, or its variants, (ii) RNA-guided proteins, and (iii) gRNA A microendoplasmic reticulum-producing cell containing a nucleic acid construct that codes for [the specified character].
14. The cell according to claim 13, wherein the RNA guide type protein is a Cas fusion protein.
15. The cell according to claim 13 or claim 14, wherein the RNA guide type protein is a base editor.
16. The cell according to any one of claims 13 to 15, wherein the base editor is an adenine base editor.
17. The cell according to any one of claims 13 to 15, wherein the base editor is a cytosine base editor.
18. The cell according to any one of claims 13 to 17, wherein the base editor comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with any one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6.
19. The cell according to claim 18, wherein the base editor comprises or consists of one of the base editors of Example 1, Example 2, Example 3, Example 4, Example 5, or Example 6.
20. The cell according to any one of claims 13 to 19, wherein the gRNA is sgRNA.
21. The cell according to any one of claims 13 to 19, wherein the gRNA is crRNA + tracrRNA.
22. The cell according to any one of claims 13 to 21, wherein the ARRDC1 protein, a fragment thereof, or a variant thereof is fused to an RNA guide protein, optionally via an intervening linker and / or cleavage domain.
23. The cell according to any one of claims 13 to 21, wherein an RNA guide protein is fused to the WW domain, optionally via an intervening linker and / or cleavage domain.
24. A method for delivering a molecule to a target cell, comprising contacting the target cell with a microvesicle according to any one of claims 1 to 12.
25. A method for treating a disorder in a patient, comprising administering a microvesicle according to any one of claims 1 to 12 to the patient.
26. A method for treating a disorder in a patient, comprising administering to the patient a microendoplasmic reticulum-producing cell according to any one of claims 13 to 23.
27. The method according to claim 25 or 26, wherein the patient is a mammal.
28. The method according to claim 27, wherein the mammal is a primate.
29. The method according to claim 15, wherein the primate is a human.
30. A cell according to any one of claims 1 to 12, further comprising a fusogen, or a cell according to any one of claims 13 to 23, further comprising a nucleic acid encoding a fusogen.
31. The ARMM or cell according to claim 30, wherein the fusogen is VSV-G.
32. ARMM according to any one of claims 1 to 12, 30, and 31, or cells according to any one of claims 13 to 23, 30, and 31, for use with or manufacture of a drug.
33. A composition that is substantially as shown and described.
34. In essence, the method is as shown and described.